Unstructured Data Normalization via Ensemble Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data standardization solutions face efficiency and reliability challenges in processing unstructured data, particularly in achieving accurate and efficient predictive analysis due to the complexity and variability of unstructured data formats.
Innovation Solution
A method and system utilizing a combination of natural language processing (NLP) and structured data classification machine learning models, along with an ensemble model, to generate and merge classification labels for unstructured data elements, determining distance measure differences and assigning labels based on thresholds to enhance predictive accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data standardization methods are used to process unstructured data, then the processing can be performed with simpler computational resources, but the predictive accuracy and reliability of the analysis deteriorates due to the complexity and variability of unstructured data formats
Solution Approach 1:
The patent segments the classification process into multiple specialized machine learning models: an NLP model for text-based unstructured data, a structured data classification model for tabular data, and an ensemble model that integrates their outputs. This segmentation allows each model to specialize in specific data types, improving predictive accuracy while managing computational complexity through modular architecture.
Solution Approach 2:
The patent employs a composite modeling approach by combining predictions from heterogeneous machine learning models (NLP model, structured data model, and ensemble model) to classify unstructured data elements. This composite system leverages the strengths of different model types to achieve higher reliability than any single model could provide alone.
2Reliability
If multiple machine learning models are combined to improve classification accuracy of unstructured data, then the predictive accuracy improves, but the training time and computational resources required increases
Solution Approach 1:
The patent implements preliminary action by pre-processing unstructured data elements to extract relevant features and characteristics before feeding them to the classification models. The NLP model performs preliminary text analysis and the structured data model prepares tabular data in advance, reducing the computational burden during the final ensemble classification stage and optimizing training efficiency.
Solution Approach 2:
The system applies partial action by selectively applying different classification pathways based on the data type. The ensemble model only processes data elements that require integration of multiple model outputs, while simpler data types can be classified by individual models alone, reducing overall training time while maintaining accuracy for complex cases.
3Productivity
If unstructured data is standardized using conventional methods, then the storage requirements remain lower, but the efficiency and reliability of predictive analysis deteriorates
Solution Approach 1:
The patent creates a universal classification framework that handles multiple data types (text, tabular, and other unstructured formats) through a common ensemble model architecture. This multi-functional system processes diverse data formats through standardized pipelines, improving predictive analysis efficiency across different data types while managing computational resources through unified processing logic.
Solution Approach 2:
The system dynamically adjusts processing parameters based on the characteristics of unstructured data elements. The ensemble model modifies classification thresholds, feature weights, and model combination strategies according to the specific data type and complexity, optimizing predictive efficiency while adapting computational resource usage to match the actual analysis requirements.
Data Source
AI summary
Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for classifying unstructured data by: (i) generating probability scores of natural language classification labels for classifying unstructured data elements using an NLP-based model, (ii) generating probability scores of structured data classification labels for classifying the unstructured data elements using a classification-based model, and (iii) assigning classifications labels based on: a) the probability scores of the natural language classification labels if a distance measure difference associated with the natural language classification labels is greater than a predetermined distance, or b) a determination using an ensemble model if the distance measure difference is less than a predetermined distance.


