Unstructured Data Search Algorithm for Medical Information Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing medical data storage systems face challenges in automatically extracting and classifying key information from unstructured data, such as medical reports and images, which are essential for patient care but difficult to process due to their non-standardized format.

Innovation Solution

A method and system that utilize an unstructured data search algorithm to identify and classify information within medical data sources by matching user-defined data patterns, such as keywords or image filters, and outputting the results in a structured format with classification scores, enabling automatic extraction and categorization of relevant information without manual search.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual search methods are used to extract information from unstructured medical data, then information extraction accuracy can be maintained through human judgment, but time consumption and labor costs increase significantly

Engineering Contradiction:
Improveinformation extraction accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces an intermediary system comprising natural language processing algorithms, machine learning models, and structured extraction frameworks that act as a mediator between unstructured medical data and structured information requirements. This intermediary automatically parses free-text clinical notes, identifies relevant entities and relationships, and structures the extracted information according to predefined schemas, thereby eliminating manual search while maintaining extraction accuracy through multiple validation layers.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If unstructured data is stored in free-form format, then data entry flexibility and ease of use are improved, but automatic extraction and classification of useful information becomes difficult

Engineering Contradiction:
Improvedata entry flexibilityVSAvoidautomatic extraction capability
Core Design Contradiction:
Ease of operationVSExtent of automation

Solution Approach 1:

The patent segments unstructured medical text into distinct informational components using natural language processing techniques. The system divides free-text clinical notes into sentences, then further segments them into phrases and individual entities (such as medications, diagnoses, procedures, and patient attributes). Each segment is tagged with semantic labels and structured according to predefined data models, enabling automatic extraction while preserving the flexibility of free-form data entry.

Inventive Principle:
Principle #1Segmentation

3Extent of automation

If structured data formats are used for all medical records, then automatic data extraction becomes straightforward using standard query languages, but data entry flexibility and the ability to capture complex clinical narratives are reduced

Engineering Contradiction:
Improveautomatic data extractionVSAvoiddata entry flexibility
Core Design Contradiction:
Extent of automationVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic data architecture that adapts between unstructured and structured formats based on the operational context. During data entry, the system accepts flexible free-form text input. During data retrieval and analysis, the same data is automatically transformed into structured formats through NLP processing and machine learning classification, enabling both data entry flexibility and automatic extraction capability through dynamic format conversion.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8751495B2Automated patient/document identification and categorization for medical data
Publication Date: 2014.06.10 CERNER INNOVATION INC
  • US8751495B2 patent drawing
  • US8751495B2 patent drawing
  • US8751495B2 patent drawing

AI summary

A method, including receiving a data source selection from a user or software application, the data source including medical information of a plurality of patients, receiving, from the user or software application, a data pattern that is related to a concept to be explored in the data source, querying the data source to find information that approximately matches the data pattern; and receiving the information from the data source, wherein the information includes unstructured data, assigning a classification to individual parts of the information based on the part's relationship to the data pattern, and outputting the classified information to the user or software application.