Learning Data Processing Device for Time-Series Outlier Removal

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The accuracy of determination models is hindered by the quality of learning data, which often includes abnormal measured values and data from periods when the object being measured is not operational, requiring improved methods to preprocess and refine this data.

Innovation Solution

A learning data processing device and method that employs statistical outlier removal processes to identify and eliminate abnormal values and data from non-operational periods, using calculated upper and lower limit values based on statistical measures, thereby enhancing the quality of learning data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If learning data includes all measured values from time-series data, then the quantity of learning data is maximized, but the quality of learning data deteriorates due to inclusion of abnormal values and non-operational period data

Engineering Contradiction:
Improvequantity of learning dataVSAvoidquality of learning data
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent extracts and removes abnormal measured values and measured values from non-operational periods from the time-series data. Specifically, it calculates statistical values (mean and standard deviation) of measured values during operational periods, determines upper and lower limit values based on these statistics, and removes values falling outside these limits along with values from non-operational periods, thereby purifying the learning data while maintaining sufficient quantity

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the time-series data into different periods: operational periods and non-operational periods. It separately processes each segment by calculating statistical values only for operational periods and applying removal criteria selectively, ensuring that only relevant data is used for learning while preserving the overall data structure and quantity

Inventive Principle:
Principle #1Segmentation

2Reliability

If statistical outlier removal process is applied to time-series data, then the quality of learning data is improved by removing abnormal values, but the complexity of data processing increases

Engineering Contradiction:
Improvequality of learning dataVSAvoidcomplexity of data processing
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a self-service approach where the data processing system automatically calculates statistical values (mean and standard deviation) from the time-series data itself, determines appropriate upper and lower limit values without external intervention, and autonomously removes outliers and non-operational period data. This automated self-processing reduces the need for complex manual intervention while maintaining high data quality

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the parameters for data selection by dynamically calculating statistical values (mean and standard deviation) from the actual data and using these to determine removal criteria. Instead of using fixed thresholds, it adapts the removal parameters to the specific characteristics of the time-series data, simplifying the processing logic while improving data quality

Inventive Principle:
Principle #35Parameter changes

3Duration of action of moving object

If measured values from non-operational periods are included in learning data, then the duration of available data is maximized, but the accuracy of learning model deteriorates

Engineering Contradiction:
Improveduration of available dataVSAvoidaccuracy of learning model
Core Design Contradiction:
Duration of action of moving objectVSMeasurement precision

Solution Approach 1:

The patent performs preliminary action by identifying and removing measured values from non-operational periods before the learning model is trained. It calculates statistical values and determines removal criteria in advance, then applies these criteria to filter out inappropriate data points, ensuring that only data from operational periods is used for learning and thereby maintaining high model accuracy while maximizing the use of available operational data

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3889850A1Learning data processing device, learning data processing method and non-transitory computer-readable medium
Publication Date: 2021.10.06 YOKOGAWA ELECTRIC CORP
  • EP3889850A1 patent drawingFigure 1~2
  • EP3889850A1 patent drawingFigure 3
  • EP3889850A1 patent drawingFigure 4~5

AI summary

The learning data processing device 10 includes the data processing unit 12 configured to generate learning data used in the learning device 30 that generates a learning model on the basis of time-series data including at least one kind of measured value. The data processing unit 12 executes at least one of a first removal process in which a statistical value of measured values included in one or multiple predetermined periods of the time-series data and at least one of an outlier determination upper limit value or an outlier determination lower limit value based on the statistical value are calculated, and, of measured values included in one or multiple predetermined periods, measured values that are at least one of those greater than or equal to the outlier determination upper limit value or those less than or equal to the outlier determination lower limit value are removed from the time-series data, or a second removal process in which, of measured values included in the time-series data, measured values satisfying a predetermined condition are removed from the time-series data.