Learning Data Processing Device for Time-Series Outlier Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The accuracy of determination models is hindered by the quality of learning data, which often includes abnormal measured values and data from periods when the object being measured is not operational, requiring improved methods to preprocess and refine this data.
Innovation Solution
A learning data processing device and method that employs statistical outlier removal processes to identify and eliminate abnormal values and data from non-operational periods, using calculated upper and lower limit values based on statistical measures, thereby enhancing the quality of learning data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If learning data includes all measured values from time-series data, then the quantity of learning data is maximized, but the quality of learning data deteriorates due to inclusion of abnormal values and non-operational period data
Solution Approach 1:
The patent extracts and removes abnormal measured values and measured values from non-operational periods from the time-series data. Specifically, it calculates statistical values (mean and standard deviation) of measured values during operational periods, determines upper and lower limit values based on these statistics, and removes values falling outside these limits along with values from non-operational periods, thereby purifying the learning data while maintaining sufficient quantity
Solution Approach 2:
The patent segments the time-series data into different periods: operational periods and non-operational periods. It separately processes each segment by calculating statistical values only for operational periods and applying removal criteria selectively, ensuring that only relevant data is used for learning while preserving the overall data structure and quantity
2Reliability
If statistical outlier removal process is applied to time-series data, then the quality of learning data is improved by removing abnormal values, but the complexity of data processing increases
Solution Approach 1:
The patent implements a self-service approach where the data processing system automatically calculates statistical values (mean and standard deviation) from the time-series data itself, determines appropriate upper and lower limit values without external intervention, and autonomously removes outliers and non-operational period data. This automated self-processing reduces the need for complex manual intervention while maintaining high data quality
Solution Approach 2:
The patent changes the parameters for data selection by dynamically calculating statistical values (mean and standard deviation) from the actual data and using these to determine removal criteria. Instead of using fixed thresholds, it adapts the removal parameters to the specific characteristics of the time-series data, simplifying the processing logic while improving data quality
3Duration of action of moving object
If measured values from non-operational periods are included in learning data, then the duration of available data is maximized, but the accuracy of learning model deteriorates
Solution Approach 1:
The patent performs preliminary action by identifying and removing measured values from non-operational periods before the learning model is trained. It calculates statistical values and determines removal criteria in advance, then applies these criteria to filter out inappropriate data points, ensuring that only data from operational periods is used for learning and thereby maintaining high model accuracy while maximizing the use of available operational data
Data Source
Figure 1~2
Figure 3
Figure 4~5
AI summary
The learning data processing device 10 includes the data processing unit 12 configured to generate learning data used in the learning device 30 that generates a learning model on the basis of time-series data including at least one kind of measured value. The data processing unit 12 executes at least one of a first removal process in which a statistical value of measured values included in one or multiple predetermined periods of the time-series data and at least one of an outlier determination upper limit value or an outlier determination lower limit value based on the statistical value are calculated, and, of measured values included in one or multiple predetermined periods, measured values that are at least one of those greater than or equal to the outlier determination upper limit value or those less than or equal to the outlier determination lower limit value are removed from the time-series data, or a second removal process in which, of measured values included in the time-series data, measured values satisfying a predetermined condition are removed from the time-series data.