Progressive Sampling Attribute Selection for Predictive Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional predictive modeling systems face challenges in efficiently generating accurate models due to the high computational burden and user frustration caused by the need for large data samples and numerous attributes, often limiting the number of attributes which reduces model effectiveness and omits relevant data.
Innovation Solution
The progressive sampling attribute selection system iteratively samples data to identify focused, relevant attributes, reducing the number of discrete data points and computational resources required, by conducting coarse and refined sampling to generate accurate predictive models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional predictive modeling systems analyze a large number of attributes and data samples to generate accurate models, then model accuracy is improved, but computing resources and time required increase significantly
Solution Approach 1:
The patent segments the attribute selection process into multiple iterative stages. In each iteration, the system processes a subset of attributes and data samples, progressively refining the model. This segmentation allows the system to handle large numbers of attributes (e.g., 50+ attributes) without requiring all data to be processed simultaneously, thereby reducing the computational burden while maintaining model accuracy through progressive refinement.
Solution Approach 2:
The system performs preliminary attribute sampling and evaluation before committing to full model training. By pre-identifying promising attributes through coarse sampling and statistical analysis, the system prepares a refined subset of attributes that are most likely to contribute to model accuracy. This preliminary action filters out irrelevant attributes early, reducing the data volume required for subsequent modeling stages.
2Reliability
If the number of data samples is increased to improve model accuracy, then predictive reliability is improved, but the burden on remote servers and communication bandwidth increases
Solution Approach 1:
The patent applies partial action by processing a carefully selected subset of data samples rather than requiring all available data. The system uses statistical sampling methods to identify a sufficient number of representative samples that capture the essential patterns needed for reliable predictions. This approach achieves acceptable predictive reliability with significantly fewer data samples, reducing server processing power and communication bandwidth requirements.
3Productivity
If conventional systems limit the number of attributes to reduce computational burden, then computing resources are conserved, but relevant data is omitted and model effectiveness decreases
Solution Approach 1:
The patent implements dynamic attribute selection where the set of attributes considered evolves iteratively. The system begins with a broader attribute set, evaluates their contribution to predictive power, and dynamically refines the attribute subset across iterations. This dynamic approach allows the model to adaptively include relevant attributes while excluding irrelevant ones, maintaining model effectiveness without requiring all possible attributes to be processed simultaneously.
Solution Approach 2:
The system incorporates feedback mechanisms where each iteration's results inform subsequent attribute selection. By evaluating model performance metrics after processing each attribute subset, the system receives feedback on which attributes contribute most to predictive accuracy. This feedback drives the progressive refinement of the attribute set, ensuring that relevant data is retained while computational resources are used efficiently.
Data Source
AI summary
The present disclosure includes methods and systems for generating digital predictive models by progressively sampling a repository of data samples. In particular, one or more embodiments of the disclosed systems and methods identify initial attributes for predicting a target attribute and utilize the initial attributes to identify a coarse sample set. Moreover, the disclosed systems and methods can utilize the coarse sample set to identify focused attributes pertinent to predicting the target attribute. Utilizing the focused attributes, the disclosed systems and methods can identify refined data samples and utilize the refined data samples to identify final attributes and generate a digital predictive model.


