Predictive Segments from Sampled Data for Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing predictive modeling systems require well-defined input-output pairs, which are not readily available in cases where the data consists of samples with and without the event of interest, leading to inefficiencies and inaccuracies in modeling and recommendation processes.

Innovation Solution

A method and system that creates predictive segments by comparing distributions of sample data with and without the event of interest, allowing for the synthesis of functional input-output pairs and enabling recommendations based on demographic, geographic, and behavioral characteristics, even when traditional input-output pairs are not defined.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional predictive modeling techniques (linear regression, logistic regression, neural networks, CART) are used, then modeling can be performed when functional input-output pairs are available, but the system cannot handle cases where such pairs are not readily available and must be synthesized from sample distributions

Engineering Contradiction:
Improveability to handle different data formatsVSAvoidmodeling system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the modeling process into two distinct phases: (1) a pre-processing phase that synthesizes functional input-output pairs from sample distributions using clustering and density estimation, and (2) a modeling phase that applies traditional predictive techniques to the synthesized pairs. This segmentation allows the system to handle diverse data formats while maintaining the simplicity of established modeling methods.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary action by pre-processing the sample data to synthesize functional input-output pairs before applying predictive modeling. This pre-processing step includes clustering samples, estimating density functions, and generating representative input-output pairs, which prepares the data in a format suitable for traditional modeling techniques and eliminates the need to modify the modeling algorithms themselves.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If clustering techniques (K-means, vector quantization) are used to generate input-output pairs, then predictive segments can be created, but the methods are computationally expensive requiring large numbers of iterative calculations to adjust clusters to convergence

Engineering Contradiction:
Improvepredictive segment accuracyVSAvoidcomputational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by performing clustering only to the extent necessary to generate sufficient input-output pairs for modeling, rather than requiring complete convergence of clusters. The system can stop the iterative clustering process once enough representative samples are obtained, reducing computational time while maintaining adequate predictive accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent performs preliminary sampling and clustering to generate input-output pairs before the main modeling process. By pre-generating these pairs with controlled iteration limits, the system prepares sufficient data in advance without requiring exhaustive computational convergence, thus reducing overall computational time while maintaining model quality.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If clustering techniques are used, then groupings can be defined, but determination of the number of clusters is difficult and may require trial and error, particularly given the non-guarantee of the predictability of the clusters

Engineering Contradiction:
Improvemodel applicability to different scenariosVSAvoidease of determining cluster parameters
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent incorporates feedback mechanisms where the quality of synthesized input-output pairs is evaluated based on model performance metrics. The cluster number and density estimation parameters are adjusted based on feedback from cross-validation results, allowing automatic optimization without manual trial and error. This feedback loop guides the selection of clustering parameters that maximize predictive accuracy.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs self-service by automatically determining optimal cluster numbers through internal validation procedures. The algorithm evaluates multiple cluster configurations and selects the one that produces the best predictive performance on validation data, eliminating the need for manual parameter tuning and making the system easier to operate across different scenarios.

Inventive Principle:
Principle #25Self-service

4Quantity of substance

If more clusters are created to improve predictive coverage, then more input-output pairs can be generated, but the computational expense and complexity of determining optimal cluster numbers increases

Engineering Contradiction:
Improvenumber of input-output pairsVSAvoidcomputational processing time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent applies partial action by generating a sufficient but not excessive number of input-output pairs through clustering. Rather than creating all possible clusters to ensure complete coverage, the system generates enough pairs to achieve adequate model performance and stops when this threshold is met, thus balancing quantity of data with computational efficiency.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10460347B2Extracting predictive segments from sampled data
Publication Date: 2019.10.29 CERTONA CORP
  • US10460347B2 patent drawing
  • US10460347B2 patent drawing
  • US10460347B2 patent drawing

AI summary

A system and method is disclosed which predicts the relative occurrence or presence of an event or item based on sample data consisting of samples which contain and samples which do not contain the event or item. The samples also consist of any number of descriptive attributes, which may be continuous variables, binary variables, or categorical variables. Given the sampled data, the system automatically creates statistically optimal segments from which a functional input/output relationship can be derived. These segments can either be used directly in the form of a lookup table or in some cases as input data to a secondary modeling system such as a linear regression module, a neural network, or other predictive system.