Model Stack for Classifying Imperfect Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern database systems face challenges in managing 'bad' or 'imperfect' data, especially in Big Data contexts, where data is often unstructured and difficult to process, leading to issues with data analysis, storage, and consistency.

Innovation Solution

A data classification system that uses self-learning models to classify and enrich data records, fitting them to known taxonomies, thereby promoting data integrity and consistency, and enabling querying for various usages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional data processing applications are used to manage Big Data, then data storage capacity is sufficient, but data processing capability becomes inadequate

Engineering Contradiction:
Improvedata processing capabilityVSAvoiddata processing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the data processing task into multiple specialized models stacked in sequence. Each model performs a specific function (classification, confidence assessment, data type determination), dividing the complex problem of managing imperfect Big Data into manageable, specialized components that collectively achieve high processing capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the parameter of data quality by applying multiple transformation models in sequence. The classification model transforms unstructured data into categorized data, the confidence model transforms uncertainty into probability scores, and the data type model transforms ambiguous data into typed data, progressively improving data quality parameters.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If manual data classification methods are used, then data accuracy can be maintained, but processing time and resource consumption increase significantly

Engineering Contradiction:
Improvedata classification speedVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The confidence model provides feedback on the reliability of classifications made by the classification model. This feedback mechanism allows the system to identify and reprocess low-confidence classifications, maintaining high accuracy while automating the majority of high-volume, high-confidence classifications for improved productivity.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs self-service through automated model stacking where the classification model, confidence model, and data type model work autonomously in sequence. The system self-corrects and self-improves by using the confidence scores to determine which data requires additional processing, eliminating the need for manual intervention in most cases.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If multiple data sources are integrated, then data comprehensiveness improves, but data consistency and integrity deteriorate

Engineering Contradiction:
Improvedata source compatibilityVSAvoiddata consistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent creates a universal data processing framework that handles multiple data sources and formats through the same model stack. The classification model and data type model are designed to work with diverse input types (unstructured, structured, semi-structured data), providing universal adaptability while maintaining consistent output standards that ensure data integrity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system applies local quality control by processing different data sources through specialized models that adapt to their specific characteristics. The confidence model assesses the quality of each classification locally, and low-confidence results are flagged for additional processing, ensuring that each data source maintains its specific quality requirements while contributing to the overall consistent dataset.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12020172B2System and/or method for generating clean records from imperfect data using model stack(s) including classification model(s) and confidence model(s)
Publication Date: 2024.06.25 XEEVA INC
  • US12020172B2 patent drawing
  • US12020172B2 patent drawing
  • US12020172B2 patent drawing

AI summary

Techniques relating to managing “bad” or “imperfect” data being imported into a database system are described herein. A lifecycle technology solution helps receive data from a variety of different data sources of a variety of known and/or unknown formats, standardize it, fit it to a known taxonomy through model-assisted classification, store it to a database in a manner that is consistent with the taxonomy, and allow it to be queried for a variety of different usages. Auto-classification, enrichment, clustering model and model stacks, and/or other disclosed techniques, may be used in these and/or other regards.