LLM Data Normalization for Medical Device Quality Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The integration and analysis of diverse quality management data in medical devices and other industries face challenges due to differences in data formats, languages, and scalability, leading to misclassification of customer feedback and potential regulatory noncompliance.

Innovation Solution

A system utilizing large language models (LLMs) to preprocess and normalize data, automatically classify customer feedback, and flag potential safety issues, ensuring compliance with regulatory standards by transforming raw data into a canonical format and providing quality predictions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If data from diverse sources (clinical trials, electronic health records, manufacturing records) are integrated, then comprehensive quality analysis is achieved, but data integration complexity and format standardization challenges increase

Engineering Contradiction:
Improvequality information completenessVSAvoiddata integration complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments the complex data integration process into distinct modules: data ingestion module that accepts multiple formats, preprocessing module that cleans and standardizes data, normalization module that converts to canonical format, and analysis module that generates insights. This segmentation allows each module to handle specific tasks independently, reducing overall integration complexity while maintaining comprehensive information capture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary normalization layer that acts as a mediator between diverse data sources and the analysis engine. This intermediary component converts various data formats (PDFs, Excel files, database records) into a unified canonical format, enabling seamless integration without requiring changes to source systems or complex custom integration logic for each data type.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If large volumes of medical device data (70,000 100-page pdfs) are processed, then complete quality documentation is analyzed, but processing time and computational resources increase

Engineering Contradiction:
Improvequality documentation completenessVSAvoiddata processing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary actions by implementing automated preprocessing steps that occur before main analysis: PDF parsing and text extraction are performed upfront to convert scanned documents into searchable text; data validation and cleaning are executed in advance to identify and correct errors; and initial normalization is completed beforehand to structure data for faster querying. These preliminary actions reduce the computational burden during actual quality analysis.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual mechanical processing of physical documents with automated computational systems. Optical character recognition (OCR) technology substitutes for manual transcription of handwritten notes in PDFs; automated parsing algorithms replace manual data extraction from various formats; and machine learning models substitute for human analysts in identifying quality issues. This substitution dramatically reduces processing time from weeks to minutes while maintaining thoroughness.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If manual classification of customer feedback is performed, then accurate categorization is achieved, but misclassification errors and regulatory noncompliance risks occur

Engineering Contradiction:
Improvefeedback classification accuracyVSAvoidregulatory compliance reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system implements feedback mechanisms where classification results are continuously validated against regulatory criteria and historical data. The normalization module incorporates feedback loops that check classified feedback against compliance rules, and the analysis engine uses feedback from identified patterns to refine classification accuracy over time. This ensures both high precision in categorization and reliable regulatory compliance through automated verification.

Inventive Principle:
Principle #23Feedback

4Ease of manufacture

If traditional data analysis methods are used, then existing processes are maintained, but scalability and handling of expanding data volumes become critical concerns

Engineering Contradiction:
Improveprocess simplicityVSAvoiddata scalability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal data processing framework that can handle multiple data types and formats through a single standardized interface. The normalization layer provides multi-functionality by accepting PDFs, Excel files, database records, and other formats while applying the same processing logic. This universal approach maintains process simplicity through standardized workflows while achieving excellent scalability as the system can accommodate growing data volumes without requiring fundamentally different processing methods.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250291841A1Quality Management Data Analysis with Machine Learning Models
Publication Date: 2025.09.18 RAREBIT INC
  • US20250291841A1 patent drawing
  • US20250291841A1 patent drawing
  • US20250291841A1 patent drawing

AI summary

A system may receive, from a client device, a query requesting quality information of a target device. The system may access a set of data records associated with the target device, pre-process the set of data records for extracting raw data associated with the target device from the set of data records, and convert the pre-processed data to normalized data using a first large language model (LLM). The system may apply a second LLM to the normalized data for generating an output result, which includes the requested quality information of the target device. Applying the second LLM may include: retrieving contextual information related to the target device; generating a prompt to the second LLM, the prompt comprising at least the normalized data, the retrieved contextual information, and the query requesting quality information of the target device; and providing the generated prompt to the second LLM to receive the output result.