User Tools

Site Tools


mistudio:logic_editor:database:feature_extractor_request_creator

This is an old revision of the document!


Feature Extractor Request

The Feature Extractor Request component, implemented as FeatureExtractorRequestCreator, builds a Feature Extractor request without requiring the request JSON to be written manually.

It is available from the Database category in the MIStudio Logic Editor.

The generated request can be connected to a Feature Extractor (Wide Schema) component, which performs the database query and feature calculations.

Overview

Feature Extractor Request separates the configuration of an analysis from the values that may change at runtime.

Properties define items such as:

  • Database table
  • Timestamp column
  • Grouping columns
  • Tag
  • Feature group
  • Metrics
  • Contiguity settings
  • Filter definitions

Input ports provide runtime values such as:

  • Start time
  • End time
  • Trigger
  • Clear
  • Row filter values

The component generates a JSON request that can be sent to a Feature Extractor.

A typical workflow is:

Runtime values
      |
      v
Feature Extractor Request
      |
      | Request JSON
      v
Feature Extractor (Wide Schema)
      |
      v
Feature Result JSON

Adding the Component

In the Logic Editor:

  1. Open the Database category.
  2. Add Feature Extractor Request to the Logic Editor canvas.
  3. Select the component.
  4. Configure its properties in the Basic property panel.

Configuring the Data Source

Database Connection

Set Database Connection to the DatabaseConnectionManager that represents the database being used.

This connection is used by the Feature Extractor Request property editors to discover available tables and columns.

The Feature Extractor Request component does not execute the database query itself.

For information about database connections, see:

Table

Select the database table containing the source data.

The table list is populated using metadata from the selected Database Connection.

For example:

process_log

Timestamp Column

Select the column containing the timestamp for each sample.

After a table is selected, the Timestamp Column list is populated with columns from that table.

For example:

TS

The generated request contains a data source similar to:

"dataSource": {
  "dataShape": "WIDE",
  "table": "process_log",
  "timestampColumn": "TS"
}

The Feature Extractor Request component currently generates requests using the WIDE data shape.

Selecting the Measurement

Tag

Tag identifies the measurement column that will be analyzed.

For example, a process table containing temperature measurements may use:

TEMP

Feature Group Name

Feature Group Name assigns a name to the group of calculated features.

For example:

hpo_temperature

The same feature group name can later be used by components such as Feature Comparator.

Feature Group Description

An optional description can be added to explain the purpose of the feature group.

Selecting Metrics

Click Metrics to open the metric selection editor.

Metrics are organized into several categories.

Statistical

Examples include:

  • count
  • mean
  • median
  • min
  • max
  • range
  • stddev
  • variance
  • percentile
  • iqr
  • mad

Time Domain

Examples include:

  • first
  • last
  • delta
  • percent_change
  • slope
  • rate_of_change

Amplitude

Additional metrics are available for amplitude-based analysis, including RMS, peak values, energy, crossing rates, and other calculations.

Frequency Domain

Frequency-domain metrics are also available for applications such as vibration and oscillation analysis.

Examples include:

  • Dominant frequency
  • Dominant period
  • Spectral power
  • Spectral centroid
  • Spectral entropy
  • Band power

Autocorrelation and Threshold Metrics

Additional metric groups include autocorrelation and threshold-based calculations.

Threshold metrics require threshold configuration that is not currently generated by Feature Extractor Request.

Group By Columns

Group By Columns defines the columns that identify groups of samples.

For example:

JOB_NAME
JOB_REV
MODULE
STEP_NO
WF_NO

Each selected column becomes a grouping dimension in the generated request.

Conceptually:

Timestamp Column
    = When did this sample occur?

Tag
    = What value is being analyzed?

Group By Columns
    = Which process run does this sample belong to?

If no Group By Columns are configured, the request is still valid, but all matching rows are treated as one group for the selected time range.

Contiguity

Grouping columns identify matching process context, but matching rows may still represent multiple executions separated by time.

The contiguity settings can split a group when the time gap between samples becomes too large.

Contiguity Max Gap

Defines the largest allowed time gap between samples before a new partition is created.

Example:

60s

Contiguity Min Samples

Defines the minimum number of samples required in a resulting partition.

Example:

4

With these settings, the generated request includes:

"groupOptions": {
  "contiguity": {
    "mode": "SPLIT_ON_GAP",
    "maxGap": "60s",
    "minSamples": 4
  }
}

Row Filters

Row Filters limit which source rows are included before grouping and metric calculation.

The Row Filters editor defines:

  • Column
  • Operator

The comparison value is not stored directly in the Row Filters property.

Instead, MIStudio creates an input port for each configured filter.

For example:

Filter Generated Input Port
JOB_NAME = ? filter_JOB_NAME_eq
JOB_REV = ? filter_JOB_REV_eq
TEMP != ? filter_TEMP_ne

A Constant, graphic value, data source, or another logic component can provide the comparison value.

For example:

"PMMA 2 R1"
      |
      v
filter_JOB_NAME_eq

The generated request then contains:

{
  "column": "JOB_NAME",
  "operator": "=",
  "value": "PMMA 2 R1"
}

Supported Filter Operators

Operator Generated Port Token
= eq
!= ne
> gt
>= ge
< lt
⇐ le
LIKE like
IN in
NOT IN notin

For example:

TEMP != ?

creates:

filter_TEMP_ne

If a configured filter input has not received a good value, that filter is omitted from the generated request and a warning is reported.

Time Range

A Feature Extractor request requires a time range.

There are two primary ways to provide one.

Absolute Time Range

Connect values to both:

  • Start Time
  • End Time

When both inputs contain good values, the component generates:

"timeRange": {
  "mode": "ABSOLUTE",
  "start": "...",
  "end": "..."
}

Date/time values containing an offset may be normalized to UTC in the generated request.

For example:

2025-03-07T15:35:20-07:00

may be generated as:

2025-03-07T22:35:20Z

These represent the same instant.

Latest Window Duration

Instead of supplying Start Time and End Time, Latest Window Duration can define a recent time window.

Examples include:

30m
24h
7d

This is used when both Start Time and End Time do not contain good values.

If only one of Start Time or End Time has a value, the absolute range is incomplete and a warning is generated.

Window Mode

Window Mode controls whether the extractor operates over the full partition or a relative portion of it.

When Whole partition / FULL is selected, the entire partition is used.

Relative window modes can also use Window Start Offset and Window End Offset.

Input Ports

The component includes these fixed input ports:

Input Purpose
Trigger Builds and publishes the request.
Clear Clears stored runtime time/filter values so they are not reused.
Start Time Supplies the start of an absolute time range.
End Time Supplies the end of an absolute time range.

Additional filter input ports are created automatically when Row Filters are configured.

For example:

filter_JOB_NAME_eq
filter_JOB_REV_eq
filter_MODULE_eq
filter_STEP_NO_eq
filter_WF_NO_eq
filter_TEMP_ne

Outputs

Output

The primary output contains the generated Feature Extractor request JSON.

This output can be connected to the request input of Feature Extractor (Wide Schema).

Example:

Feature Extractor Request.Output
             |
             v
Feature Extractor.requestJson

MessageStrings

MessageStrings contains validation errors and warnings generated while building the request.

This output is useful during configuration and troubleshooting.

Request Preview

Request Preview displays the request implied by the current properties and runtime values.

It can be used to inspect the generated request before sending it to a Feature Extractor.

Validation messages may also be included with the preview.

Validation and Warnings

Errors prevent a valid request from being emitted.

Examples include:

  • Missing table
  • Missing timestamp column
  • No time range and no Latest Window Duration
  • Invalid Latest Window Duration
  • Start Time occurring after End Time
  • Missing Tag
  • No valid Metrics

Warnings may still allow a request to be generated.

Examples include:

  • A filter input has not received a value, so that filter is omitted.
  • Only one of Start Time or End Time contains a good value.
  • No Group By Columns are configured.
  • A duplicate Group By Column was ignored.
  • A selected metric is not recognized.
  • A relative window is configured without an offset.

Example: Golden Batch Temperature Request

The following configuration can be used to analyze temperature samples from a process log.

Property Example Value
Database Connection Databricks DatabaseConnectionManager
Table process_log
Timestamp Column TS
Tag TEMP
Feature Group Name hpo_temperature
Group By Columns JOB_NAME, JOB_REV, MODULE, STEP_NO, WF_NO
Contiguity Max Gap 60s
Contiguity Min Samples 4

Selected metrics:

count
mean
median
min
max
range
stddev
variance
first
last
delta
slope
rate_of_change

Example row filters:

Column Operator Runtime Value
JOB_NAME = PMMA 2 R1
JOB_REV = v7.0
MODULE = HPO 2
STEP_NO = 1
WF_NO = 1
TEMP != 0

Example time range:

Start Time: 2025-03-07T15:35:20-07:00
End Time:   2025-03-07T15:37:00-07:00

The resulting request contains the selected table, time range, grouping definitions, filters, feature group, and metrics.

Notes

  • Feature Extractor Request generates requests for wide-schema data.
  • The component currently generates one feature group containing one tag.
  • Row filter values are supplied through generated input ports rather than being stored directly in the Row Filters property.
  • Filter rows without a good runtime value are omitted from the generated request.
  • The component itself builds the request. The Feature Extractor performs the database query and feature calculation.
mistudio/logic_editor/database/feature_extractor_request_creator.1789772549.txt.gz · Last modified: by tputman

Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki