Predictive Segments from Sampled Data for Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing predictive modeling systems require well-defined input-output pairs, which are not readily available in cases where the data consists of samples with and without the event of interest, leading to inefficiencies and inaccuracies in modeling and recommendation processes.
Innovation Solution
A method and system that creates predictive segments by comparing distributions of sample data with and without the event of interest, allowing for the synthesis of functional input-output pairs and enabling recommendations based on demographic, geographic, and behavioral characteristics, even when traditional input-output pairs are not defined.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional predictive modeling techniques (linear regression, logistic regression, neural networks, CART) are used, then modeling can be performed when functional input-output pairs are available, but the system cannot handle cases where such pairs are not readily available and must be synthesized from sample distributions
Solution Approach 1:
The patent segments the modeling process into two distinct phases: (1) a pre-processing phase that synthesizes functional input-output pairs from sample distributions using clustering and density estimation, and (2) a modeling phase that applies traditional predictive techniques to the synthesized pairs. This segmentation allows the system to handle diverse data formats while maintaining the simplicity of established modeling methods.
Solution Approach 2:
The patent performs preliminary action by pre-processing the sample data to synthesize functional input-output pairs before applying predictive modeling. This pre-processing step includes clustering samples, estimating density functions, and generating representative input-output pairs, which prepares the data in a format suitable for traditional modeling techniques and eliminates the need to modify the modeling algorithms themselves.
2Measurement precision
If clustering techniques (K-means, vector quantization) are used to generate input-output pairs, then predictive segments can be created, but the methods are computationally expensive requiring large numbers of iterative calculations to adjust clusters to convergence
Solution Approach 1:
The patent applies partial action by performing clustering only to the extent necessary to generate sufficient input-output pairs for modeling, rather than requiring complete convergence of clusters. The system can stop the iterative clustering process once enough representative samples are obtained, reducing computational time while maintaining adequate predictive accuracy.
Solution Approach 2:
The patent performs preliminary sampling and clustering to generate input-output pairs before the main modeling process. By pre-generating these pairs with controlled iteration limits, the system prepares sufficient data in advance without requiring exhaustive computational convergence, thus reducing overall computational time while maintaining model quality.
3Adaptability or versatility
If clustering techniques are used, then groupings can be defined, but determination of the number of clusters is difficult and may require trial and error, particularly given the non-guarantee of the predictability of the clusters
Solution Approach 1:
The patent incorporates feedback mechanisms where the quality of synthesized input-output pairs is evaluated based on model performance metrics. The cluster number and density estimation parameters are adjusted based on feedback from cross-validation results, allowing automatic optimization without manual trial and error. This feedback loop guides the selection of clustering parameters that maximize predictive accuracy.
Solution Approach 2:
The system performs self-service by automatically determining optimal cluster numbers through internal validation procedures. The algorithm evaluates multiple cluster configurations and selects the one that produces the best predictive performance on validation data, eliminating the need for manual parameter tuning and making the system easier to operate across different scenarios.
4Quantity of substance
If more clusters are created to improve predictive coverage, then more input-output pairs can be generated, but the computational expense and complexity of determining optimal cluster numbers increases
Solution Approach 1:
The patent applies partial action by generating a sufficient but not excessive number of input-output pairs through clustering. Rather than creating all possible clusters to ensure complete coverage, the system generates enough pairs to achieve adequate model performance and stops when this threshold is met, thus balancing quantity of data with computational efficiency.
Data Source
AI summary
A system and method is disclosed which predicts the relative occurrence or presence of an event or item based on sample data consisting of samples which contain and samples which do not contain the event or item. The samples also consist of any number of descriptive attributes, which may be continuous variables, binary variables, or categorical variables. Given the sampled data, the system automatically creates statistically optimal segments from which a functional input/output relationship can be derived. These segments can either be used directly in the form of a lookup table or in some cases as input data to a secondary modeling system such as a linear regression module, a neural network, or other predictive system.


