Table of Contents
Feature Extractor Request Creator
The Feature Extractor Request Creator component builds a Feature Extractor request without requiring the request JSON to be written manually.
In the Logic Editor palette, it appears as Feature Extractor Request.
It is available from the Database category in the MIStudio Logic Editor.
The generated request can be connected to a Feature Extractor (Wide Schema) component, which performs the database query and feature calculations.
Overview
Feature Extractor Request separates the configuration of an analysis from the values that may change at runtime.
Properties define items such as:
- Database table
- Timestamp column
- Grouping columns
- Tag
- Feature group
- Metrics
- Contiguity settings
- Filter definitions
Input ports provide runtime values such as:
- Start time
- End time
- Trigger
- Clear
- Row filter values
The component generates a JSON request that can be sent to a Feature Extractor.
A typical workflow is:
Runtime values
|
v
Feature Extractor Request
|
| Request JSON
v
Feature Extractor (Wide Schema)
|
v
Feature Result JSON
Adding the Component
In the Logic Editor:
- Open the Database category.
- Add Feature Extractor Request to the Logic Editor canvas.
- Select the component.
- Configure its properties in the Basic property panel.
Configuring the Data Source
Database Connection
Set Database Connection to the DatabaseConnectionManager that represents the database being used.
This connection is used by the Feature Extractor Request property editors to discover available tables and columns.
The Feature Extractor Request component does not execute the database query itself.
For information about database connections, see:
Table
Select the database table containing the source data.
The table list is populated using metadata from the selected Database Connection.
For example:
process_log
Timestamp Column
Select the column containing the timestamp for each sample.
After a table is selected, the Timestamp Column list is populated with columns from that table.
For example:
TS
The generated request contains a data source similar to:
"dataSource": {
"dataShape": "WIDE",
"table": "process_log",
"timestampColumn": "TS"
}
The Feature Extractor Request component currently generates requests using the WIDE data shape.
Selecting the Measurement
Tag
Tag identifies the measurement column that will be analyzed.
For example, a process table containing temperature measurements may use:
TEMP
Feature Group Name
Feature Group Name assigns a name to the group of calculated features.
For example:
hpo_temperature
The same feature group name can later be used by components such as Feature Comparator.
Feature Group Description
An optional description can be added to explain the purpose of the feature group.
Selecting Metrics
Click Metrics to open the metric selection editor.
Metrics are organized into several categories.
Statistical
Examples include:
countmeanmedianminmaxrangestddevvariancepercentileiqrmad
Time Domain
Examples include:
firstlastdeltapercent_changesloperate_of_change
Amplitude
Additional metrics are available for amplitude-based analysis, including RMS, peak values, energy, crossing rates, and other calculations.
Frequency Domain
Frequency-domain metrics are also available for applications such as vibration and oscillation analysis.
Examples include:
- Dominant frequency
- Dominant period
- Spectral power
- Spectral centroid
- Spectral entropy
- Band power
Autocorrelation and Threshold Metrics
Additional metric groups include autocorrelation and threshold-based calculations.
Threshold metrics require threshold configuration that is not currently generated by Feature Extractor Request.
Group By Columns
Group By Columns defines the columns that identify groups of samples.
For example:
JOB_NAME JOB_REV MODULE STEP_NO WF_NO
Each selected column becomes a grouping dimension in the generated request.
Conceptually:
Timestamp Column
= When did this sample occur?
Tag
= What value is being analyzed?
Group By Columns
= Which process run does this sample belong to?
If no Group By Columns are configured, the request is still valid, but all matching rows are treated as one group for the selected time range.
Contiguity
Grouping columns identify matching process context, but matching rows may still represent multiple executions separated by time.
The contiguity settings can split a group when the time gap between samples becomes too large.
Contiguity Max Gap
Defines the largest allowed time gap between samples before a new partition is created.
Example:
60s
Contiguity Min Samples
Defines the minimum number of samples required in a resulting partition.
Example:
4
With these settings, the generated request includes:
"groupOptions": {
"contiguity": {
"mode": "SPLIT_ON_GAP",
"maxGap": "60s",
"minSamples": 4
}
}
Row Filters
Row Filters limit which source rows are included before grouping and metric calculation.
The Row Filters editor defines:
- Column
- Operator
The comparison value is not stored directly in the Row Filters property.
Instead, MIStudio creates an input port for each configured filter.
For example:
| Filter | Generated Input Port |
|---|---|
JOB_NAME = ? | filter_JOB_NAME_eq |
JOB_REV = ? | filter_JOB_REV_eq |
TEMP != ? | filter_TEMP_ne |
A Constant, graphic value, data source, or another logic component can provide the comparison value.
For example:
"PMMA 2 R1"
|
v
filter_JOB_NAME_eq
The generated request then contains:
{
"column": "JOB_NAME",
"operator": "=",
"value": "PMMA 2 R1"
}
Supported Filter Operators
| Operator | Generated Port Token |
|---|---|
= | eq |
!= | ne |
> | gt |
>= | ge |
< | lt |
⇐ | le |
LIKE | like |
IN | in |
NOT IN | notin |
For example:
TEMP != ?
creates:
filter_TEMP_ne
If a configured filter input has not received a good value, that filter is omitted from the generated request and a warning is reported.
Time Range
A Feature Extractor request requires a time range.
There are two primary ways to provide one.
Absolute Time Range
Connect values to both:
- Start Time
- End Time
When both inputs contain good values, the component generates:
"timeRange": {
"mode": "ABSOLUTE",
"start": "...",
"end": "..."
}
Date/time values containing an offset may be normalized to UTC in the generated request.
For example:
2025-03-07T15:35:20-07:00
may be generated as:
2025-03-07T22:35:20Z
These represent the same instant.
Latest Window Duration
Instead of supplying Start Time and End Time, Latest Window Duration can define a recent time window.
Examples include:
30m 24h 7d
This is used when both Start Time and End Time do not contain good values.
If only one of Start Time or End Time has a value, the absolute range is incomplete and a warning is generated.
Window Mode
Window Mode controls whether the extractor operates over the full partition or a relative portion of it.
When Whole partition / FULL is selected, the entire partition is used.
Relative window modes can also use Window Start Offset and Window End Offset.
Input Ports
The component includes these fixed input ports:
| Input | Purpose |
|---|---|
| Trigger | Builds and publishes the request. |
| Clear | Clears stored runtime time/filter values so they are not reused. |
| Start Time | Supplies the start of an absolute time range. |
| End Time | Supplies the end of an absolute time range. |
Additional filter input ports are created automatically when Row Filters are configured.
For example:
filter_JOB_NAME_eq filter_JOB_REV_eq filter_MODULE_eq filter_STEP_NO_eq filter_WF_NO_eq filter_TEMP_ne
Outputs
Output
The primary output contains the generated Feature Extractor request JSON.
This output can be connected to the request input of Feature Extractor (Wide Schema).
Example:
Feature Extractor Request.Output
|
v
Feature Extractor.requestJson
MessageStrings
MessageStrings contains validation errors and warnings generated while building the request.
This output is useful during configuration and troubleshooting.
Request Preview
Request Preview displays the request implied by the current properties and runtime values.
It can be used to inspect the generated request before sending it to a Feature Extractor.
Validation messages may also be included with the preview.
Validation and Warnings
Errors prevent a valid request from being emitted.
Examples include:
- Missing table
- Missing timestamp column
- No time range and no Latest Window Duration
- Invalid Latest Window Duration
- Start Time occurring after End Time
- Missing Tag
- No valid Metrics
Warnings may still allow a request to be generated.
Examples include:
- A filter input has not received a value, so that filter is omitted.
- Only one of Start Time or End Time contains a good value.
- No Group By Columns are configured.
- A duplicate Group By Column was ignored.
- A selected metric is not recognized.
- A relative window is configured without an offset.
Example: Golden Batch Temperature Request
The following configuration can be used to analyze temperature samples from a process log.
| Property | Example Value |
|---|---|
| Database Connection | Databricks DatabaseConnectionManager |
| Table | process_log |
| Timestamp Column | TS |
| Tag | TEMP |
| Feature Group Name | hpo_temperature |
| Group By Columns | JOB_NAME, JOB_REV, MODULE, STEP_NO, WF_NO |
| Contiguity Max Gap | 60s |
| Contiguity Min Samples | 4 |
Selected metrics:
count mean median min max range stddev variance first last delta slope rate_of_change
Example row filters:
| Column | Operator | Runtime Value |
|---|---|---|
JOB_NAME | = | PMMA 2 R1 |
JOB_REV | = | v7.0 |
MODULE | = | HPO 2 |
STEP_NO | = | 1 |
WF_NO | = | 1 |
TEMP | != | 0 |
Example time range:
Start Time: 2025-03-07T15:35:20-07:00 End Time: 2025-03-07T15:37:00-07:00
The resulting request contains the selected table, time range, grouping definitions, filters, feature group, and metrics.
Notes
- Feature Extractor Request generates requests for wide-schema data.
- The component currently generates one feature group containing one tag.
- Row filter values are supplied through generated input ports rather than being stored directly in the Row Filters property.
- Filter rows without a good runtime value are omitted from the generated request.
- The component itself builds the request. The Feature Extractor performs the database query and feature calculation.
