Dual-Prediction Model for Probabilistic Data Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data labeling systems face challenges in accurately labeling candidates for complex tasks, particularly in medical contexts where definitive tests are lacking, leading to incorrect candidate selection for clinical trials and interventions, which can be costly and harmful.

Innovation Solution

A dual-prediction model system is employed to generate candidate positive-label probabilities based on historical data, using a candidate label probabilistic model and a historical record prediction model, allowing for the inclusion of previously discarded candidates and improving prediction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional data labeling systems use a defined source of truth with definitive tests, then labeling accuracy is improved, but the system cannot handle complex tasks where definitive tests are not readily available

Engineering Contradiction:
Improvelabeling accuracyVSAvoidability to handle complex tasks
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces a probabilistic label assignment mechanism that acts as an intermediary between the defined source of truth and complex tasks without definitive tests. Instead of requiring direct definitive labeling, the system uses probability thresholds and multiple prediction models to assign labels, enabling handling of complex tasks while maintaining labeling accuracy through the intermediary probabilistic layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the parameter from binary definitive labels to probabilistic labels with confidence scores. By transforming the labeling output from certain/uncertain to probability-based, the system can handle complex tasks where definitive tests are unavailable while maintaining accuracy through statistical confidence thresholds and model predictions.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If candidates who do not definitively fit into one camp or another are discarded, then labeling purity is improved, but the candidate pool size is reduced

Engineering Contradiction:
Improvelabeling purityVSAvoidcandidate pool size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies partial action by using probability thresholds rather than requiring definitive classification. Candidates are included in the pool if they exceed a probability threshold (e.g., >0.5), rather than requiring definitive proof. This partial inclusion criterion maintains labeling purity by filtering out clearly incorrect candidates while increasing the candidate pool size by including those with sufficient probabilistic evidence.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system incorporates feedback through iterative model training where predicted probabilities are used to refine future predictions. The dual-prediction model system uses feedback from historical data and previous predictions to continuously improve accuracy, allowing the system to confidently include candidates with probabilistic labels while maintaining high purity through feedback-driven model refinement.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If a dual-prediction model system is used to generate candidate positive-label probabilities, then prediction accuracy is improved, but system complexity increases

Engineering Contradiction:
Improveprediction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the prediction system into two distinct prediction models: a first prediction model for initial labeling and a second prediction model for refinement. This segmentation allows each model to specialize in specific aspects of prediction, improving overall accuracy while managing complexity through modular architecture where each model can be independently trained and optimized.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system adds a probabilistic dimension to the prediction output, transforming binary classifications into probability distributions. By operating in the probability space dimension rather than just the binary classification dimension, the system achieves higher prediction accuracy while the added complexity is managed through standardized probability thresholding and calibration techniques.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12002585B2Apparatus, computer program product, and method for predictive data labelling using a dual-prediction model system
Publication Date: 2024.06.04 OPTUM INC
  • US12002585B2 patent drawing
  • US12002585B2 patent drawing
  • US12002585B2 patent drawing

AI summary

Various embodiments of the disclosure provide apparatuses, systems, and computer program products for predictive data labelling using a dual-model system. Embodiments provide various advantages in accuracy of predicted labels, for example in various contexts such as medical data analysis for difficult to diagnose diseases. An example provided apparatus is configured to generate a positive, neutral, and negative candidate identifier sets and corresponding positive, neutral, and negative candidate index sets based in part on applying a candidate selection rule set to a candidate data set; train a candidate label probabilistic model based at least in part on a candidate label training subset associated with the candidate data set associated with the positive and negative candidate identifiers; generate a candidate positive-label probability set using at least the candidate label probabilistic model; train a historical record prediction model to predict the candidate positive-label probability set; and utilize the historical record prediction model.