Unstructured Text Problem List Identification via OCR and NLP

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional automated systems face challenges in accurately identifying problem lists in medical records due to varying terminology, organization, and formatting, leading to inaccurate and time-consuming processing, often resulting in over-inclusive or irrelevant data extraction.

Innovation Solution

A computer-implemented method using optical character recognition and machine learning techniques to identify problem list sections in unstructured medical records by recognizing problem list words and headings, associating relevant text, and outputting a list of potential problems, while removing irrelevant data to enhance processing efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional automated systems are used to identify problem lists in medical records, then processing can be automated, but accuracy deteriorates due to varying terminology, organization, and formatting

Engineering Contradiction:
Improveautomation of problem list identificationVSAvoidaccuracy of problem list identification
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent segments the medical record into distinct sections (e.g., problem list, history of present illness, past medical history) and processes each section separately using specialized NLP models. This segmentation allows the system to focus computational resources on relevant sections, improving accuracy while maintaining automation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts processing parameters based on document characteristics. Different NLP models are selected and applied based on the specific section being processed, the complexity of the document, and identified patterns. This parameter adaptation enables high accuracy across varying terminologies and formats while preserving automated operation.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If conventional automated systems process all sections of medical records, then comprehensive data extraction is achieved, but processing time increases due to irrelevant data

Engineering Contradiction:
Improvecompleteness of data extractionVSAvoidprocessing time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent extracts and isolates only the relevant problem list section from the medical record, separating it from irrelevant content such as history of present illness and past medical history. This extraction approach maintains completeness of the required data while significantly reducing processing time by eliminating unnecessary data from the automated processing pipeline.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary identification and classification of document sections before applying NLP processing. By pre-segmenting the medical record and identifying problem list sections in advance, the system prepares only the relevant portions for detailed analysis, reducing overall processing time while maintaining data completeness.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If conventional automated systems are used, then processing can be performed, but computational resources are consumed excessively due to lack of efficient processing techniques

Engineering Contradiction:
Improveprocessing capabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent applies different computational intensities to different sections of the medical record. High-complexity NLP models are applied only to the problem list section where detailed analysis is needed, while simpler processing is applied to other sections. This local quality approach maintains processing capability for critical tasks while reducing overall computational resource consumption.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system applies NLP processing partially only to the extent necessary for identifying problem lists, rather than processing the entire medical record with full computational intensity. This partial action approach achieves the required productivity for problem list identification while significantly reducing unnecessary computational resource consumption on irrelevant sections.

Inventive Principle:
Principle #16Partial or excessive action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach improves the accuracy and efficiency of identifying problem lists, reducing computational load and facilitating the generation of Hierarchical Condition Categories (HCCs) by focusing on relevant sections within medical records.

Implementation Method 1

generating, by the one or more processors and based on applying an optical character recognition algorithm to the electronic document, unstructured text

Methodology Applied
Scientific EffectOptical character recognition:

Data Source

PatentUS20240331434A1Systems and methods for section identification in unstructured data
Publication Date: 2024.10.03 OPTUM INC
  • US20240331434A1 patent drawing
  • US20240331434A1 patent drawing
  • US20240331434A1 patent drawing

AI summary

A computer-implemented method for identifying a problem list section from an electronic document includes receiving, by one or more processors, the electronic document, generating, by the one or more processors and based on applying an optical character recognition algorithm to the electronic document, unstructured text, and identifying, by the one or more processors, one or more problem list words in the unstructured text, the one or more problem list words belonging in a dataset for identifying a presence of a problem list section. The method also includes associating, by the one or more processors, a portion of the unstructured text that corresponds to the one or more problem list words in the unstructured text with the problem list section and outputting, by the one or more processors, at least a portion of the problem list section.