Hierarchical Training Data Structuring for Multi-Device ML Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine-learning systems face challenges in constructing hierarchical training data sets that are applicable to specific controlled devices, leading to inefficiencies due to data anomalies from differing free inputs, design variables, and condition variables.

Innovation Solution

A system comprising a central computer system and database constructs hierarchical training data sets by aggregating data from multiple controlled devices, prioritizing real data over simulated values, and using machine-learning to determine optimal control inputs, thereby minimizing data anomalies and improving training accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If training data is aggregated from multiple controlled devices with differing free inputs, design variables, and condition variables, then the quantity and diversity of training data increases, but data anomalies increase and training accuracy decreases

Engineering Contradiction:
Improvequantity of training dataVSAvoidtraining accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments the training data aggregation process into hierarchical levels (individual device level, group level, and fleet level). Data is processed and normalized at each level before aggregation, allowing the system to maintain data diversity while reducing anomalies through structured segmentation of the data collection and processing workflow.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by allowing different data processing and normalization techniques to be applied to different segments of data based on their specific characteristics. Each controlled device or group can have customized preprocessing steps applied to their data before aggregation, ensuring that local data quality issues are addressed while maintaining overall data diversity.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If real data from multiple controlled devices is used for training, then the relevance to specific conditions improves, but data heterogeneity and processing complexity increase

Engineering Contradiction:
Improverelevance to specific conditionsVSAvoiddata processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by pre-processing and normalizing data at the source (individual devices and groups) before aggregation. Data cleaning, transformation, and standardization operations are performed in advance during data collection phases, reducing the complexity of processing heterogeneous real data when it reaches the central training system.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediary processing layers (data normalization modules, feature extraction components) that act as mediators between diverse real data sources and the machine learning training process. These intermediaries standardize data formats and extract relevant features, making heterogeneous real data more manageable while preserving condition-specific relevance.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If simulated data is used to supplement training data, then data completeness improves, but data anomalies and reduced effectiveness occur

Engineering Contradiction:
Improvecompleteness of training dataVSAvoideffectiveness of training
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent merges real data and simulated data through a hierarchical aggregation process where both data types are collected, normalized, and combined at multiple levels. The system integrates simulated data to fill gaps in real data coverage while applying consistency checks and validation rules to maintain reliability, ensuring that simulated data supplements rather than contaminates the training set.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent uses simulated data as copies or representations of real operational scenarios where real data is unavailable. Rather than using simulated data directly, the system creates realistic data copies through simulation models that replicate actual device behavior patterns, then validates these copies against known real data characteristics before incorporating them into the training set.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11900226B2Systems for constructing hierarchical training data sets for use with machine-learning and related methods therefor
Publication Date: 2024.02.13 SOURCE GLOBAL PBC
  • US11900226B2 patent drawing
  • US11900226B2 patent drawing
  • US11900226B2 patent drawing

AI summary

Some embodiments include a system operable to construct hierarchical training data sets for use with machine-learning for multiple controlled devices. Other embodiments of related systems and methods are also provided.