Automated Clinical Data Standardization via Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Clinical medical data for retrospective research is highly variable in format and content due to differences in data storage formats and human error, requiring significant manual manipulation to generate accurate reports, which can introduce additional errors.

Innovation Solution

Systems and methods for extracting and standardizing clinical data from multiple health data sources using machine-learning algorithms to identify and correct inconsistencies, and generate relevant reports based on user profiles and data analysis needs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual data manipulation is used to standardize clinical data, then data accuracy can be improved, but time consumption and error introduction increase

Engineering Contradiction:
Improvedata accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual data manipulation with automated machine learning algorithms that perform data standardization and quality control. The system uses natural language processing and classification models to automatically extract, standardize, and validate clinical data from multiple sources, eliminating the need for manual intervention while maintaining high accuracy through algorithmic processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-validation through machine learning quality control checks that automatically identify and correct data inconsistencies. The machine learning models continuously learn from data patterns to improve their own accuracy in standardizing clinical data without requiring external manual correction, enabling the system to serve itself in the data standardization process

Inventive Principle:
Principle #25Self-service

2Loss of time

If automated machine learning is used to standardize clinical data, then time consumption is reduced, but data complexity and processing requirements increase

Engineering Contradiction:
Improvetime consumptionVSAvoidsystem complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent divides the complex data standardization process into separate modular components: data extraction modules for each data source, machine learning classification models for data standardization, quality control checks for validation, and report generation modules for output. This segmentation allows each component to handle specific tasks independently, reducing overall system complexity while maintaining automated processing capabilities

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces standardized data formats and classification models as intermediaries between diverse clinical data sources and the final report generation. These intermediaries translate varied data structures from different sources into a unified standardized format, simplifying the processing requirements and enabling automated handling without direct complex interactions between all system components

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If clinical data is collected from multiple data sources, then data comprehensiveness is improved, but data inconsistency and variability increase

Engineering Contradiction:
Improvedata comprehensivenessVSAvoiddata consistency
Core Design Contradiction:
Quantity of substanceVSStability of the object's composition

Solution Approach 1:

The patent implements universal machine learning classification models that can process and standardize data from multiple different clinical data sources simultaneously. The models are trained to recognize and standardize various data formats and structures from different sources into a consistent unified format, enabling the system to handle diverse data comprehensiveness while maintaining consistency through the universal standardization process

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11416472B2Automated computing platform for aggregating data from a plurality of inconsistently configured data sources to enable generation of reporting recommendations
Publication Date: 2022.08.16 UNITEDHEALTH GROUP INC
  • US11416472B2 patent drawing
  • US11416472B2 patent drawing
  • US11416472B2 patent drawing

AI summary

Methods, apparatus, systems, computing devices, computing entities, and/or the like for generating medical research reports automatically collect data from a plurality of separate health data storage systems, standardize the received data to support at least a requested report type, apply one or more machine-learning quality control check to identify potentially inaccurate data included within the received data, and to generate the requested report based at least in part on the standardized, refined data. Moreover, one or more recommended additional reports supported by the refined data set is identified and recommended to a user based at least in part on user attributes and reports initially requested.