Reject Inference for Bias-Reduced Lead Scoring Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional lead scoring models are inaccurate and inefficient due to their reliance on biased training data, which only includes information from accepted prospects, neglecting rejected prospects and thus limiting the dataset and introducing bias, leading to suboptimal resource allocation in engaging potential customers.

Innovation Solution

The implementation of a reject inference model to generate synthetic data for rejected prospects, expanding the dataset and updating the scoring model by modifying its parameters with this synthetic data, thereby incorporating missing information and improving accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional lead scoring models use only data from accepted prospects for training, then the model training process is simpler, but the accuracy and reliability of the scoring model deteriorates due to bias and limited data

Engineering Contradiction:
Improvescoring accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by generating synthetic outcome data for rejected prospects before updating the scoring model. This advance preparation of data allows the model to be trained on a more complete dataset without increasing the complexity of the training process itself, thereby improving scoring accuracy while maintaining manageable complexity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary component - the synthetic data generation system - that bridges the gap between rejected prospects (who lack outcome data) and the scoring model training process. This intermediary generates plausible outcome data for rejected prospects, enabling their inclusion in training without directly complicating the core model architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If the system utilizes only data from accepted prospects, then the data processing requirements are reduced, but the productivity and efficiency of the lead scoring system deteriorates due to larger required datasets

Engineering Contradiction:
Improvelead identification efficiencyVSAvoiddata volume required
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The system creates copies of rejected prospect data by generating synthetic outcome data that mirrors the structure and characteristics of actual outcome data from accepted prospects. These synthetic copies enable the model to learn from rejected prospects without requiring additional real-world data collection, thereby improving productivity while maintaining efficient data volume

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the parameter state of rejected prospect data by assigning synthetic outcome values that transform them from incomplete records into usable training samples. This parameter transformation allows the system to utilize a larger effective dataset for the same volume of collected data, improving lead identification efficiency without increasing data storage requirements

Inventive Principle:
Principle #35Parameter changes

3Reliability

If conventional scoring models incorporate data from rejected prospects, then the dataset size increases improving accuracy, but the complexity of handling and processing this data increases

Engineering Contradiction:
Improvemodel reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the complexity of handling rejected prospect data by separating it into a distinct synthetic data generation module. This extraction allows the main scoring model to receive clean, standardized training data while the complexity of inferring outcomes for rejected prospects is contained in a separate, manageable component

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system segments the data processing workflow into distinct stages: collecting accepted prospect outcomes, generating synthetic outcomes for rejected prospects, and updating the scoring model. This segmentation of the complex task into manageable segments reduces overall system complexity while enabling comprehensive use of both accepted and rejected prospect data, thereby improving model reliability

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11514515B2Generating synthetic data using reject inference processes for modifying lead scoring models
Publication Date: 2022.11.29 ADOBE INC
  • US11514515B2 patent drawing
  • US11514515B2 patent drawing
  • US11514515B2 patent drawing

AI summary

Methods, systems, and non-transitory computer readable storage media are disclosed for using reject inference to generate synthetic data for modifying lead scoring models. For example, the disclosed system identifies an original dataset corresponding to an output of a lead scoring model that generates scores for a plurality of prospects to indicate a likelihood of success of prospects of the plurality of prospects. In one or more embodiments, the disclosed system selects a reject inference model by performing simulations on historical prospect data associated with the original dataset. Additionally, the disclosed system uses the selected reject inference model to generate an imputed dataset by generating synthetic outcome data representing simulated outcomes of rejected prospects in the original dataset. The disclosed system then uses the imputed dataset to modify the lead scoring model by modifying at least one parameter of the lead scoring model using the synthetic outcome data.