Automatic Modeling Farmer for Big Data Variable Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for developing large-scale models using big data are inefficient and labor-intensive, as they require manual processing and struggle to handle the complexity and volume of high-volume, high-velocity, and high-variety data sets, limiting the speed and accuracy of decision-making and model optimization.

Innovation Solution

An automatic modeling farmer system that accesses disparate data sources, builds multiple test models, selects the most predictive variables, and generates a master model using data processors, enabling streamlined model development, evaluation, and reporting, thereby reducing manual work and increasing efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual processing methods are used for developing large-scale models, then model development can be completed with existing tools, but the process becomes labor-intensive and inefficient

Engineering Contradiction:
Improvemodel development speedVSAvoidmanual processing requirement
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The system enables self-service automated model development through the modeling farmer that automatically accesses data sources, builds test models, selects variables, and generates master models without requiring manual intervention at each step, thereby increasing productivity while reducing manual labor requirements

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processing with automated computational systems. The modeling farmer substitutes human operators with an automated engine that systematically performs data access, model building, variable selection, and master model generation, transforming labor-intensive manual processes into efficient automated operations

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If traditional data processing applications are used, then existing tools can handle current data volumes, but they struggle with high-volume, high-velocity, and high-variety big data sets

Engineering Contradiction:
Improvedata volume capacityVSAvoidhandling capability for diverse data types
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The modeling farmer is designed as a universal system that can access and process multiple types of disparate data sources simultaneously. It performs multiple functions including data access, model building, variable selection, and master model generation, enabling the system to handle high-volume, high-velocity, and high-variety big data sets with adaptability to different data formats and sources

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If multiple test models are automatically built from disparate data sources, then model accuracy can be improved through comprehensive variable selection, but the complexity of the modeling process increases

Engineering Contradiction:
Improvepredictive variable accuracyVSAvoidmodeling process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The modeling process is segmented into distinct automated stages: data access from disparate sources, test model building with predetermined variables, variable selection through comparison of predictive power, master dataset generation, and master model building. This segmentation manages complexity by breaking down the complex process into manageable automated steps while maintaining high predictive accuracy

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10083263B2Automatic modeling farmer
Publication Date: 2018.09.25 FAIR ISAAC & CO INC
  • US10083263B2 patent drawing
  • US10083263B2 patent drawing
  • US10083263B2 patent drawing

AI summary

Data can be accessed from a plurality of disparate data sources from at least one database. A plurality of test models can be automatically built by a model building engine. Each test model can have predetermined predictive variables. A final set of predictive variables can be determined by a variable selector from the predetermined predictive variables in the plurality of test models by comparing the predictive power of the predictive variables across the plurality of test models. A master dataset can be generated from the disparate data sources. A master model can be built from the master dataset. The master model can combine the final set of predictive variables from the plurality of disparate data sources. The master model can characterize a quantitative estimate of the probability that an entity will display a defined behavior. Related apparatus, systems, techniques, and articles are also described.