Unstructured Data Normalization via Ensemble Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data standardization solutions face efficiency and reliability challenges in processing unstructured data, particularly in achieving accurate and efficient predictive analysis due to the complexity and variability of unstructured data formats.

Innovation Solution

A method and system utilizing a combination of natural language processing (NLP) and structured data classification machine learning models, along with an ensemble model, to generate and merge classification labels for unstructured data elements, determining distance measure differences and assigning labels based on thresholds to enhance predictive accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional data standardization methods are used to process unstructured data, then the processing can be performed with simpler computational resources, but the predictive accuracy and reliability of the analysis deteriorates due to the complexity and variability of unstructured data formats

Engineering Contradiction:
Improvepredictive accuracyVSAvoidcomputational resource requirements
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the classification process into multiple specialized machine learning models: an NLP model for text-based unstructured data, a structured data classification model for tabular data, and an ensemble model that integrates their outputs. This segmentation allows each model to specialize in specific data types, improving predictive accuracy while managing computational complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs a composite modeling approach by combining predictions from heterogeneous machine learning models (NLP model, structured data model, and ensemble model) to classify unstructured data elements. This composite system leverages the strengths of different model types to achieve higher reliability than any single model could provide alone.

Inventive Principle:
Principle #40Composite materials

2Reliability

If multiple machine learning models are combined to improve classification accuracy of unstructured data, then the predictive accuracy improves, but the training time and computational resources required increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-processing unstructured data elements to extract relevant features and characteristics before feeding them to the classification models. The NLP model performs preliminary text analysis and the structured data model prepares tabular data in advance, reducing the computational burden during the final ensemble classification stage and optimizing training efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies partial action by selectively applying different classification pathways based on the data type. The ensemble model only processes data elements that require integration of multiple model outputs, while simpler data types can be classified by individual models alone, reducing overall training time while maintaining accuracy for complex cases.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If unstructured data is standardized using conventional methods, then the storage requirements remain lower, but the efficiency and reliability of predictive analysis deteriorates

Engineering Contradiction:
Improvepredictive analysis efficiencyVSAvoidcomputational resources
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent creates a universal classification framework that handles multiple data types (text, tabular, and other unstructured formats) through a common ensemble model architecture. This multi-functional system processes diverse data formats through standardized pipelines, improving predictive analysis efficiency across different data types while managing computational resources through unified processing logic.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically adjusts processing parameters based on the characteristics of unstructured data elements. The ensemble model modifies classification thresholds, feature weights, and model combination strategies according to the specific data type and complexity, optimizing predictive efficiency while adapting computational resource usage to match the actual analysis requirements.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12032590B1Machine learning techniques for normalization of unstructured data into structured data
Publication Date: 2024.07.09 OPTUM INC
  • US12032590B1 patent drawing
  • US12032590B1 patent drawing
  • US12032590B1 patent drawing

AI summary

Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for classifying unstructured data by: (i) generating probability scores of natural language classification labels for classifying unstructured data elements using an NLP-based model, (ii) generating probability scores of structured data classification labels for classifying the unstructured data elements using a classification-based model, and (iii) assigning classifications labels based on: a) the probability scores of the natural language classification labels if a distance measure difference associated with the natural language classification labels is greater than a predetermined distance, or b) a determination using an ensemble model if the distance measure difference is less than a predetermined distance.