Preprocessing Method for Homogeneous Machine Learning Data Blocks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Inhomogeneous machine sensor data sets, due to variations in machine operation and sensor quality, lead to poor quality and non-comparable machine learning models for monitoring, resulting in erroneous and inconsistent results.
Innovation Solution
A preprocessing method that groups data sets into homogeneous blocks by adjusting record counts and normalizing operating parameter values within predetermined ranges, ensuring consistent training and usage of machine learning algorithms for monitoring.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If machine sensor data is collected from various operating conditions and applications, then the data coverage and applicability are improved, but the data homogeneity and model comparability deteriorate
Solution Approach 1:
The patent segments the collected sensor data into discrete data blocks, where each block contains a fixed number of complete data sets. This segmentation approach allows the system to handle diverse operating conditions while maintaining consistent block structures for training different machine learning models, thus resolving the contradiction between data coverage and data homogeneity.
Solution Approach 2:
The patent applies normalization to transform sensor values into a standard range, changing the parameter scale of different sensors to be comparable. This parameter transformation enables data from various operating conditions and sensor types to be processed uniformly, maintaining data homogeneity while preserving the diversity of operating scenarios.
2Adaptability or versatility
If data sets with varying amounts of records are used for training, then the adaptability to different operating scenarios is improved, but the model consistency and comparability worsen
Solution Approach 1:
The patent divides the data into blocks with a predetermined fixed number of complete data sets. This segmentation ensures that each training block has consistent size and structure, allowing for reproducible and comparable model training across different operating scenarios while still covering diverse conditions through the selection of appropriate data blocks.
3Reliability
If sensor data with varying value ranges and qualities is processed directly, then the data authenticity is preserved, but the machine learning model performance and reliability deteriorate
Solution Approach 1:
The patent applies normalization to transform sensor values into a standard range while preserving the underlying patterns and relationships in the data. This parameter transformation improves model performance by ensuring consistent input scales for machine learning algorithms, while the normalization process is designed to maintain the authenticity of the underlying operational patterns.
Solution Approach 2:
The normalization process acts as an intermediary between the raw sensor data and the machine learning models. It transforms the data into a suitable format for processing without losing the essential information, thus improving model performance while preserving data authenticity through a controlled transformation step.
4Adaptability or versatility
If data blocks with inconsistent lengths are used for training, then the flexibility in handling different operating conditions is improved, but the training efficiency and model comparability worsen
Solution Approach 1:
The patent segments data into blocks with a predetermined fixed number of complete data sets, creating uniform training units. This segmentation improves training efficiency by enabling consistent batch processing and model training, while the selection of diverse data blocks maintains the ability to handle different operating conditions.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a preprocessing method for providing homogeneous data blocks from temporally ordered, inhomogeneous data sets containing values of recorded operating parameters of a machine, in order to obtain data blocks suitable for monitoring the machine by machine learning-based algorithms. The method comprises forming data blocks from the data sets such that each data block includes the data sets that lie within a respective time period, adjusting the number of data sets in the data blocks so that the data blocks contain a predetermined number of complete data sets, and normalizing the data sets in the data blocks so that the data sets have values for predetermined operating parameters that lie within predetermined value ranges.