Data-Based System Model Training State Check Using k-NN Distance Distributions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data-based system models trained with data from test bench measurements or simulations often fail to accurately represent real-world operating conditions, leading to reduced reliability in model outputs due to deviations from actual field operations.
Innovation Solution
A method involving the creation of a k-Nearest Neighbor tree to compare training data from deviating scenarios with operational data from real-world operations, determining distance distributions, and iteratively adding training data to improve representation, using techniques like Earth Mover's Distance or Kullback-Leibler divergence to assess and refine the training data's accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If training data is obtained from test bench measurements or simulations, then the effort and cost for data collection is reduced, but the accuracy and reliability of the system model in representing real-world operations deteriorates
Solution Approach 1:
The patent implements a feedback mechanism by comparing the distribution of training data with operational data from real-world scenarios. Distance metrics (e.g., Earth Mover's Distance, Kullback-Leibler divergence) quantify the discrepancy between training data distributions and actual operational data distributions. This feedback loop enables iterative refinement of training data by identifying and adding samples that reduce the distributional gap, thereby improving model reliability while maintaining efficient data collection processes.
Solution Approach 2:
The patent performs preliminary analysis of operational data distributions before finalizing the training dataset. By pre-comparing test bench/simulation data with anticipated real-world operational characteristics, the method identifies gaps in advance and supplements training data accordingly. This preliminary action ensures that training data is optimized for real-world performance before model training begins, preventing reliability issues rather than addressing them after the fact.
2Reliability
If training data is obtained from field operations, then the accuracy and reliability of the system model is improved, but the effort, time, and cost for data collection increases significantly
Solution Approach 1:
The patent creates a virtual representation of real-world operational data distributions by comparing statistical characteristics and distance metrics between available training data and target operational data. Instead of collecting extensive field data, the method copies the essential distributional properties of real-world operations into the training dataset through selective augmentation and synthesis, achieving reliable model training with minimal field data collection.
Solution Approach 2:
The patent transforms the approach from direct field data collection to parameter-based data characterization. By focusing on distributional parameters (means, variances, distance metrics) rather than raw field data, the method enables efficient training data generation that captures the essential characteristics of real-world operations without requiring prolonged field measurements.
3Adaptability or versatility
If training data from deviating scenarios is used, then the system model may cover more operating ranges, but the representation accuracy of real-world operating points deteriorates
Solution Approach 1:
The patent applies local quality by treating different regions of the data space differently. Instead of uniformly sampling all operating ranges, the method identifies regions where training data adequately represents real-world operations and regions where it does not. By focusing refinement efforts on specific under-represented local areas through distance-based analysis, the method improves representation accuracy in critical regions while maintaining broad adaptability across the full operating range.
Data Source
AI summary
A method is for providing training data for training a data-based system model for operating a technical system by defining a data point determined from input variables for determining at least one output variable depending on which the technical system is operating. The method includes providing training data that are determined with a scenario other than a real operation of the technical system, the training data are defined for data points determined from the input variables, capturing operational data points determined from the input variables in real-world operation of the technical system, and splitting the training data into training data points and validation data points. The method further includes determining a k-Nearest Neighbor tree from the training data points, and determining a first distribution of distance values of distances between each of the validation data points and a predetermined number of next training data points of the training data points.

