Automatic Modeling Farmer for Big Data Variable Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for developing large-scale models using big data are inefficient and labor-intensive, as they require manual processing and struggle to handle the complexity and volume of high-volume, high-velocity, and high-variety data sets, limiting the speed and accuracy of decision-making and model optimization.
Innovation Solution
An automatic modeling farmer system that accesses disparate data sources, builds multiple test models, selects the most predictive variables, and generates a master model using data processors, enabling streamlined model development, evaluation, and reporting, thereby reducing manual work and increasing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual processing methods are used for developing large-scale models, then model development can be completed with existing tools, but the process becomes labor-intensive and inefficient
Solution Approach 1:
The system enables self-service automated model development through the modeling farmer that automatically accesses data sources, builds test models, selects variables, and generates master models without requiring manual intervention at each step, thereby increasing productivity while reducing manual labor requirements
Solution Approach 2:
The patent replaces manual mechanical processing with automated computational systems. The modeling farmer substitutes human operators with an automated engine that systematically performs data access, model building, variable selection, and master model generation, transforming labor-intensive manual processes into efficient automated operations
2Quantity of substance
If traditional data processing applications are used, then existing tools can handle current data volumes, but they struggle with high-volume, high-velocity, and high-variety big data sets
Solution Approach 1:
The modeling farmer is designed as a universal system that can access and process multiple types of disparate data sources simultaneously. It performs multiple functions including data access, model building, variable selection, and master model generation, enabling the system to handle high-volume, high-velocity, and high-variety big data sets with adaptability to different data formats and sources
3Measurement precision
If multiple test models are automatically built from disparate data sources, then model accuracy can be improved through comprehensive variable selection, but the complexity of the modeling process increases
Solution Approach 1:
The modeling process is segmented into distinct automated stages: data access from disparate sources, test model building with predetermined variables, variable selection through comparison of predictive power, master dataset generation, and master model building. This segmentation manages complexity by breaking down the complex process into manageable automated steps while maintaining high predictive accuracy
Data Source
AI summary
Data can be accessed from a plurality of disparate data sources from at least one database. A plurality of test models can be automatically built by a model building engine. Each test model can have predetermined predictive variables. A final set of predictive variables can be determined by a variable selector from the predetermined predictive variables in the plurality of test models by comparing the predictive power of the predictive variables across the plurality of test models. A master dataset can be generated from the disparate data sources. A master model can be built from the master dataset. The master model can combine the final set of predictive variables from the plurality of disparate data sources. The master model can characterize a quantitative estimate of the probability that an entity will display a defined behavior. Related apparatus, systems, techniques, and articles are also described.


