Missing Data Identification Using Co-occurrence Matrices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current test data management systems are inefficient in identifying and addressing gaps in test data sets, relying on manual reviews and subjective decisions, which are labor-intensive and prone to errors, especially in large enterprise environments.

Innovation Solution

The system automates the identification of missing data by transforming data values into categorical data, generating co-occurrence matrices, and determining expected frequencies to accurately pinpoint missing data locations, and then generates new data to fill gaps while maintaining data consistency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual review of test data sets is used to identify missing data, then human judgment can be applied to determine data gaps, but the process becomes labor-intensive and time-consuming

Engineering Contradiction:
Improveaccuracy of missing data identificationVSAvoidtime required for manual review
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical review processes with automated computational analysis. The system uses processors to automatically analyze test data sets, identify missing data locations, and generate reports, eliminating the need for human reviewers to manually examine large volumes of test data while maintaining or improving identification accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service automated analysis where the test data management system independently identifies missing data without requiring human intervention. The automated process performs data analysis, gap identification, and report generation autonomously, freeing human operators from time-consuming manual review tasks.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If manual review of test data sets is used to identify missing data, then subjective decision making can occur, but consistency and reliability are compromised

Engineering Contradiction:
Improveflexibility in data gap determinationVSAvoidconsistency of missing data identification
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent establishes objective parameters and criteria for identifying missing data, such as data density thresholds, completeness requirements, and statistical measures. These quantifiable parameters replace subjective human judgment, ensuring consistent and reliable identification of data gaps across different reviewers and situations while maintaining the flexibility to adjust parameters based on specific testing requirements.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If test data management professionals manually identify and fill missing data, then data gaps can be addressed, but labor costs and time costs increase significantly

Engineering Contradiction:
Improvecompleteness of test data setVSAvoidefficiency of data completion process
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system performs preliminary automated identification and analysis of missing data before the data completion process begins. By pre-identifying all data gaps, analyzing their characteristics, and preparing completion strategies in advance, the system enables more efficient data completion operations and reduces the overall time and labor required for the entire process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system provides automated feedback about missing data locations, quantities, and characteristics to guide the data completion process. This feedback mechanism enables targeted and efficient data generation or acquisition efforts, reducing unnecessary manual work and improving overall productivity in completing test data sets.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11341187B2Method, apparatus, and computer-readable medium for missing data identification
Publication Date: 2022.05.24 INFORMATICA CORP
  • US11341187B2 patent drawing
  • US11341187B2 patent drawing
  • US11341187B2 patent drawing

AI summary

A system, method and computer-readable medium for missing data identification, including identifying columns in tables of a database, generating categorical columns of categorical data by transforming data values in the columns into categorical data values, generating a co-occurrence matrix corresponding to a pair of categorical columns in the categorical columns, determining an expected frequency of co-occurrence corresponding to each unique pair of categorical data values based at least in part on a marginal totals corresponding to categorical data values in the co-occurrence matrix, and identifying one or more locations of missing data based at least in part on the count of co-occurrence of each unique pair of categorical data values and the expected frequency of co-occurrence corresponding to each unique pair of categorical data values.