Associative Memory Data Quality Checking System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data-driven classification and management systems face challenges in ensuring data quality, accuracy, and efficiency, particularly due to high complexity, human error, and the need for manual rule-based systems, which are costly and time-consuming. Additionally, existing systems struggle with handling free text data and developing diverse control sets for associative memory systems, leading to inaccuracies and inefficiencies.

Innovation Solution

A computer-implemented data-driven classification and quality checking system utilizing an associative memory model with machine learning algorithms to categorize data, calculate quality rating metrics, and normalize domain-specific terms, while also providing a control set with diverse data for accurate scoring and classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If rule-based data standardization systems are used to improve data quality, then data accuracy can be maintained through coded rules, but the system becomes expensive to maintain and difficult to derive rules manually

Engineering Contradiction:
Improvedata accuracyVSAvoidsystem maintenance complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces manual rule-based data standardization systems with a machine learning-based automated system. The system uses trained models to automatically classify and standardize data without requiring manual creation and maintenance of coded rules, thereby maintaining data accuracy while reducing system complexity and maintenance burden.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The machine learning system enables the data classification process to serve itself by automatically learning patterns from training data and applying them to new data without continuous human intervention. The system self-updates through retraining on new data, eliminating the need for manual rule derivation and maintenance.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If pattern-based systems are used to handle data classification, then data categorization can be achieved, but the process becomes laborious and time consuming particularly for highly specialized data sets

Engineering Contradiction:
Improvedata categorization capabilityVSAvoidclassification time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-training machine learning models on extensive training datasets before actual data classification is needed. The trained models capture patterns and relationships in advance, enabling rapid classification of new data without time-consuming manual pattern matching during operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces laborious manual pattern-based classification with automated machine learning-based classification. The system uses algorithms to automatically identify and apply patterns from training data, dramatically reducing the time required for data categorization compared to manual pattern-based approaches.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If overall accuracy evaluation is used for predictive models, then model performance can be assessed, but high accuracy requirements become expensive, difficult, and sometimes unattainable when all decisions are assumed to be equally difficult

Engineering Contradiction:
Improvemodel accuracy assessmentVSAvoidmodel deployment efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies local quality by evaluating model accuracy at the individual decision level rather than using a single overall accuracy metric. The system assigns confidence scores to each prediction and allows different accuracy thresholds for different types of decisions, enabling deployment of models that may not achieve high overall accuracy but perform adequately for each specific decision context.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes the evaluation parameter from a single overall accuracy metric to multiple decision-specific confidence scores and accuracy measures. This allows flexible threshold setting for different applications and reduces the burden of achieving uniformly high accuracy across all decisions, improving model deployment efficiency.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If analysts review entries record-by-record to classify and score data, then each record can be individually assessed, but the process is tedious and analysts cannot spend time performing in-depth analysis to generate deep understanding

Engineering Contradiction:
Improveindividual record assessment accuracyVSAvoidanalyst time for scoring tasks
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual record-by-record analysis with automated machine learning-based classification and scoring systems. The system individually assesses each record using trained models, maintaining assessment accuracy while eliminating the tedious manual work, thereby freeing analysts to perform in-depth analysis and generate deep understanding of underlying issues.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10089581B2Data driven classification and data quality checking system
Publication Date: 2018.10.02 THE BOEING CO
  • US10089581B2 patent drawing
  • US10089581B2 patent drawing
  • US10089581B2 patent drawing

AI summary

A computer implemented data driven classification and data quality checking system is provided. The system has an interface application enabled to receive data and has an associative memory software. The system has a data driven associative memory model configured to categorize one or more fields of received data and to analyze the received data. The system has a data quality rating metric associated with the received data. The system has a machine learning data quality checker for the received data, and is configured to add the received data to a pool of neighboring data, if the data quality rating metric is greater than or equal to a data quality rating metric threshold. The machine learning data quality checker is configured to generate and communicate an alert of a potential error in the received data, if the data quality rating metric is less than the data quality rating metric threshold.