Guideline-Based Cross-Reference Embeddings for Unstructured Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models struggle to seamlessly integrate structured and unstructured data while ensuring data relevance and safeguarding against biases, particularly in handling protected health information (PHI) or personally identifiable information (PII), leading to inefficiencies and potential prejudices in decision-making.

Innovation Solution

The use of guideline-specific cross-reference data objects generated by extracting questions from guidelines, assigning scores to passages, identifying top-ranking passages, and generating embeddings to train predictive models, thereby reducing data requirements and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional machine learning models analyze extensive unstructured data directly, then comprehensive data coverage is achieved, but training time and computing resources are significantly increased

Engineering Contradiction:
Improvepredictive accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary extraction of relevant information from unstructured data using NLP techniques before the main training process. By pre-processing the unstructured data to identify and extract pertinent features, the system reduces the computational burden during training while maintaining predictive accuracy. This preliminary action allows the model to focus on processed, relevant data rather than analyzing all raw unstructured data during training.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary processing layer that transforms unstructured data into structured representations using natural language processing and information extraction techniques. This intermediary layer acts as a bridge between raw unstructured data and the machine learning model, converting complex text data into manageable features that can be efficiently processed during training, thereby reducing training time without sacrificing accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If traditional machine learning models analyze extensive unstructured data directly, then comprehensive data coverage is achieved, but computing resources are significantly increased

Engineering Contradiction:
Improvepredictive accuracyVSAvoidcomputing resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system extracts only the relevant portions of unstructured data that are pertinent to the prediction task using NLP and information extraction techniques. By taking out and isolating the essential information from large volumes of unstructured data, the system reduces the computational resources required for training while maintaining the accuracy benefits of comprehensive data analysis. This selective extraction avoids processing redundant or irrelevant information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary processing to identify and extract relevant features from unstructured data before the main training computation. This pre-extraction step reduces the volume of data that needs to be processed during computationally intensive training phases, thereby reducing overall computing resource consumption while preserving the accuracy benefits of using unstructured data.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If all unstructured data is incorporated into machine learning models, then data completeness is achieved, but data bias and prejudices are increased

Engineering Contradiction:
Improvedecision-making reliabilityVSAvoiddata bias
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system extracts relevant information from unstructured data while deliberately excluding potentially biased or harmful content. By selectively taking out only the pertinent information needed for the prediction task and filtering out biased portions, the system maintains decision-making reliability based on comprehensive data analysis while reducing the impact of data biases and prejudices.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent addresses data bias by using NLP techniques to identify and neutralize biased content in unstructured data. The system converts the potential harm of biased data into a benefit by detecting biased portions and either excluding them or adjusting their influence in the training process, thereby maintaining the reliability benefits of comprehensive data analysis while mitigating harmful biases.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

4Measurement precision

If meticulous design and experimentation are performed to combine structured and unstructured data, then integration quality is improved, but development complexity is increased

Engineering Contradiction:
Improvedata integration qualityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the data integration process into distinct modules: structured data processing, unstructured data processing with NLP, feature extraction, and model training. By dividing the complex integration task into manageable segments with clearly defined interfaces, the system achieves high data integration quality while reducing development complexity. Each segment can be independently developed, tested, and maintained.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs a universal NLP-based information extraction framework that can handle multiple types of unstructured data (text, documents, records) through a common processing pipeline. This multi-functional approach allows the system to integrate both structured and unstructured data using consistent methods, improving integration quality while reducing the complexity that would arise from implementing separate processing logic for each data type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250355923A1Machine learning techniques for guideline-based extraction of relevant information from unstructured data
Publication Date: 2025.11.20 OPTUM INC
  • US20250355923A1 patent drawing
  • US20250355923A1 patent drawing
  • US20250355923A1 patent drawing

AI summary

Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for processing a constrained query by (i) generating a cross-reference document data object by (a) extracting a plurality of questions from a guideline document, (b) assigning a rank to each of a plurality of passages from an unstructured data object for each of the plurality of questions, (c) generating a plurality of answers for the plurality of questions based on top ranking passages of the plurality of passages for each question to retrieval machine learning model, and (d) combining the plurality of answers, (ii) generating one or more cross-reference embeddings based on the cross-reference document data object, and (iii) training a predictive machine learning model based on the one or more cross-reference embeddings.