Reject Inference for Bias-Reduced Lead Scoring Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional lead scoring models are inaccurate and inefficient due to their reliance on biased training data, which only includes information from accepted prospects, neglecting rejected prospects and thus limiting the dataset and introducing bias, leading to suboptimal resource allocation in engaging potential customers.
Innovation Solution
The implementation of a reject inference model to generate synthetic data for rejected prospects, expanding the dataset and updating the scoring model by modifying its parameters with this synthetic data, thereby incorporating missing information and improving accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional lead scoring models use only data from accepted prospects for training, then the model training process is simpler, but the accuracy and reliability of the scoring model deteriorates due to bias and limited data
Solution Approach 1:
The system performs preliminary actions by generating synthetic outcome data for rejected prospects before updating the scoring model. This advance preparation of data allows the model to be trained on a more complete dataset without increasing the complexity of the training process itself, thereby improving scoring accuracy while maintaining manageable complexity
Solution Approach 2:
The patent introduces an intermediary component - the synthetic data generation system - that bridges the gap between rejected prospects (who lack outcome data) and the scoring model training process. This intermediary generates plausible outcome data for rejected prospects, enabling their inclusion in training without directly complicating the core model architecture
2Productivity
If the system utilizes only data from accepted prospects, then the data processing requirements are reduced, but the productivity and efficiency of the lead scoring system deteriorates due to larger required datasets
Solution Approach 1:
The system creates copies of rejected prospect data by generating synthetic outcome data that mirrors the structure and characteristics of actual outcome data from accepted prospects. These synthetic copies enable the model to learn from rejected prospects without requiring additional real-world data collection, thereby improving productivity while maintaining efficient data volume
Solution Approach 2:
The patent changes the parameter state of rejected prospect data by assigning synthetic outcome values that transform them from incomplete records into usable training samples. This parameter transformation allows the system to utilize a larger effective dataset for the same volume of collected data, improving lead identification efficiency without increasing data storage requirements
3Reliability
If conventional scoring models incorporate data from rejected prospects, then the dataset size increases improving accuracy, but the complexity of handling and processing this data increases
Solution Approach 1:
The patent extracts the complexity of handling rejected prospect data by separating it into a distinct synthetic data generation module. This extraction allows the main scoring model to receive clean, standardized training data while the complexity of inferring outcomes for rejected prospects is contained in a separate, manageable component
Solution Approach 2:
The system segments the data processing workflow into distinct stages: collecting accepted prospect outcomes, generating synthetic outcomes for rejected prospects, and updating the scoring model. This segmentation of the complex task into manageable segments reduces overall system complexity while enabling comprehensive use of both accepted and rejected prospect data, thereby improving model reliability
Data Source
AI summary
Methods, systems, and non-transitory computer readable storage media are disclosed for using reject inference to generate synthetic data for modifying lead scoring models. For example, the disclosed system identifies an original dataset corresponding to an output of a lead scoring model that generates scores for a plurality of prospects to indicate a likelihood of success of prospects of the plurality of prospects. In one or more embodiments, the disclosed system selects a reject inference model by performing simulations on historical prospect data associated with the original dataset. Additionally, the disclosed system uses the selected reject inference model to generate an imputed dataset by generating synthetic outcome data representing simulated outcomes of rejected prospects in the original dataset. The disclosed system then uses the imputed dataset to modify the lead scoring model by modifying at least one parameter of the lead scoring model using the synthetic outcome data.


