Machine Learning Training With Selected Time-Series Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for training machine learning models for data-driven decision-making are time-consuming and prone to errors, especially when dealing with limited data volumes, which limits their applicability in real-time applications.
Innovation Solution
A method involving the capture of measurement data using sensors, which are then processed to create time series data. This data is used to train multiple instances of the same machine learning model, each with different classification data units and selected portions of the data, allowing for improved decision-making and automation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple machine learning models are trained on complete datasets to improve decision-making quality, then the reliability of decision-making improves, but the training time and computational resources increase significantly
Solution Approach 1:
The patent extracts and utilizes only the selected portion of data that is most relevant for training, rather than processing complete datasets. This extraction approach maintains training effectiveness while significantly reducing the time and computational resources required, directly resolving the contradiction between decision-making quality and training time.
Solution Approach 2:
The patent applies partial action by training models on selected portions of data rather than complete datasets. The system determines that processing a carefully selected subset of data is sufficient to achieve the required decision-making quality, avoiding the excessive time and resource consumption that would result from processing all available data.
2Reliability
If thousands of datasets are used to train AI models to reduce errors, then the reliability of the model improves, but the complexity of the system and data processing requirements increase
Solution Approach 1:
The system extracts only the essential selected portion of data needed for effective training, eliminating the need to manage and process thousands of complete datasets. This extraction strategy maintains model reliability while significantly reducing system complexity and data processing requirements.
Solution Approach 2:
Instead of using complete datasets and hoping to achieve good results, the patent inverts the approach by first selecting a specific portion of data and then training models on this curated subset. This inversion leads to more efficient training with fewer resources while maintaining or improving model reliability.
3Ease of operation
If traditional analysis tools are developed for monitored systems to enable data-driven decision-making, then the decision-making capability improves, but the adaptability to system changes deteriorates due to complex reconfiguration requirements
Solution Approach 1:
The patent creates a universal training framework that can handle different monitored systems and data types through the same selected portion extraction and model training process. This universal approach enables the system to adapt to changes in monitored systems without requiring complex reconfiguration, as the core training methodology remains consistent across different applications.
Data Source
AI summary
The invention relates to a method for training machine learning models, having the steps of: detecting data in the form of time series data using one or more computers, said data being obtained by means of one or more measuring devices (60-62), in each case in the form of a sensor for measuring a physical variable; receiving multiple classification data units relating to the data using the one or more computers; receiving a selected part of the data using the one or more computers for each of the classification data units; and training multiple machine learning models using the one or more computers, in each case on the basis of at least one of the classification data units and the at least one corresponding selected part of the data, wherein the multiple machine learning models represent multiple instances of the same machine learning model.


