Modeling Data Coverage via Hyper-Quadrant Density Inspection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing modeling systems face challenges in evaluating and improving data coverage, leading to insufficiently broad models that fail to generalize well, particularly in underrepresented regions of the solution space, resulting in inaccurate predictions and prolonged convergence times.

Innovation Solution

A method and system for modifying data coverage in modeling systems, which involves obtaining and evaluating data records, selecting input parameters, detecting data coverage conditions, and modifying data distribution using techniques like hyper-quadrant density inspection and symmetric random scatter to ensure uniform data coverage, allowing for the generation of computational models indicative of interrelationships between input and output parameters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If training data are collected or simulated without coverage evaluation, then data collection is simple and fast, but data coverage is insufficient and models fail to generalize well to underrepresented regions

Engineering Contradiction:
Improvemodel generalization capabilityVSAvoiddata evaluation and modification system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary evaluation of data coverage before model training by dividing the modeling space into hyper-quadrants and assessing data density in each region. This allows identification of underrepresented regions before committing to model training, enabling proactive data collection or generation in sparse areas to improve model generalization capability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The modeling space is segmented into multiple hyper-quadrants based on independent parameters, allowing systematic evaluation of data distribution across different regions. This segmentation enables targeted identification of underrepresented regions and facilitates focused data collection or generation efforts in specific areas rather than treating the entire space uniformly.

Inventive Principle:
Principle #1Segmentation

2Reliability

If manual evaluation and selection of training data is performed, then data coverage can be improved, but the process is time-consuming and computationally expensive

Engineering Contradiction:
Improvedata coverage qualityVSAvoidmodeling process convergence time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system automatically evaluates data coverage by computing density metrics in hyper-quadrants and identifies underrepresented regions without requiring manual intervention. The system then autonomously generates suggestions for additional training data cases or modifies existing data weights, eliminating the need for time-consuming manual evaluation and selection processes while maintaining high data coverage quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces manual mechanical evaluation processes with automated computational methods. Instead of human experts manually reviewing and selecting training data, the system uses algorithmic density evaluation, hyper-quadrant analysis, and automated case generation suggestions to assess and improve data coverage, significantly reducing time and computational resources required.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If data are uniformly distributed in modeling space, then modeling efficiency and accuracy are improved, but collected data are naturally denser in certain regions leading to non-uniform distribution

Engineering Contradiction:
Improvemodeling efficiencyVSAvoiddata distribution uniformity
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system applies different data handling strategies to different regions of the modeling space based on local data density characteristics. In underrepresented hyper-quadrants, the system generates suggestions for additional training cases or increases data weights to achieve uniform coverage. In well-represented regions, no modification is needed. This localised approach achieves uniform data distribution across the entire modeling space while maintaining efficiency.

Inventive Principle:
Principle #3Local quality

4Measurement precision

If existing methods like scatter plot selection are used, then misclassified data can be removed, but overall data coverage in the modeling space is not improved and additional data generation suggestions are not provided

Engineering Contradiction:
Improvedata selection accuracyVSAvoidcoverage information in underrepresented regions
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system moves beyond traditional scatter plot visualization in two dimensions to a multi-dimensional hyper-quadrant framework that encompasses the entire modeling space. By evaluating data density across all hyper-quadrants simultaneously, the system identifies underrepresented regions in higher dimensions that scatter plots cannot detect, providing comprehensive coverage analysis and generating targeted suggestions for additional training cases in sparse regions.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS8086640B2System and method for improving data coverage in modeling systems
Publication Date: 2011.12.27 CATERPILLAR INC
  • US8086640B2 patent drawing
  • US8086640B2 patent drawing
  • US8086640B2 patent drawing

AI summary

A method for modifying data coverage in a modeling system is disclosed. The method may include obtaining data records relating to a plurality of input variables and one or more output parameters and selecting a plurality of input parameters from the plurality of input variables. The method may further include evaluating a coverage of the data records in a modeling space and modifying the coverage of the data records, if a data coverage condition is detected. The method may also include generating a computational model indicative of interrelationships between the plurality of input parameters and the one or more output parameters based on the data records.