Data Model Discovery via Complexity Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Advanced data modeling techniques require significant technical skill and resources, often producing overly complex models that are difficult to interpret and compare, with users seeking simpler, more interpretable models that balance accuracy and complexity.

Innovation Solution

Implementing a method that allows users to specify structural preferences and complexity values, guiding the data model discovery process through genetic programming, and allocating computer resources based on model interpretability to generate more interpretable and efficient data models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If advanced data modeling techniques are used to improve prediction accuracy, then model accuracy is improved, but model complexity increases making them difficult to interpret

Engineering Contradiction:
Improveprediction accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system changes parameters by allowing users to specify complexity values for different structural patterns and adjusting selection probabilities based on complexity scores, enabling dynamic control over the balance between accuracy and interpretability in model generation

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system dynamically adjusts the model selection process by calculating complexity scores and probabilistically selecting models based on both accuracy and complexity metrics, allowing the modeling process to adapt between competing objectives during execution

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If complex models are generated to capture detailed patterns, then model accuracy is improved, but resource consumption increases

Engineering Contradiction:
Improvemodel accuracyVSAvoidprocessing resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system changes parameters by incorporating complexity values that reflect resource consumption characteristics, allowing users to control the trade-off between accuracy and resource usage through parameter specification in the model generation process

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system applies partial action by probabilistically selecting models rather than exhaustively evaluating all possible models, using complexity scores to guide selective evaluation that achieves satisfactory accuracy without consuming excessive resources

Inventive Principle:
Principle #16Partial or excessive action

3Ease of operation

If users specify structural preferences to improve interpretability, then model interpretability is improved, but model discovery time increases

Engineering Contradiction:
Improvemodel interpretabilityVSAvoidmodel discovery time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-calculating complexity scores for candidate models based on their structural patterns, allowing the selection process to efficiently evaluate interpretability without time-consuming detailed analysis during model discovery

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes parameters by allowing users to specify complexity values for different structural patterns, enabling the model discovery process to converge faster on interpretable models by guiding the search toward preferred structures

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20220237516A1Data modeling systems and methods
Publication Date: 2022.07.28 DATAROBOT INC
  • US20220237516A1 patent drawing
  • US20220237516A1 patent drawing
  • US20220237516A1 patent drawing

AI summary

Data modeling systems and methods are described. A data modeling method may include receiving user input specifying a structure of at least a portion of a data model and a complexity value associated with the structure; (a) generating one or more data models; (b) determining complexity scores for the respective data models; (c) for each of the data models: determining whether to select the respective data model for evaluation based, at least in part, on the complexity score of the respective data model, and if the respective data model is selected for evaluation, evaluating an accuracy of the respective data model for one or more data sets; and repeating steps (a)-(c) until one or more specified termination criteria are satisfied, wherein a first of the generated data models includes the specified structure, and wherein the complexity score for the first data model is determined based, at least in part, on the complexity value associated with the structure.