Missing Data Generation in Database Tables
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current test data management systems are inefficient in identifying and addressing gaps in test data sets, relying on manual reviews and subjective decisions, which are labor-intensive and prone to errors, especially in large enterprise environments.
Innovation Solution
The system identifies missing data by analyzing data sets to quantify gaps and generates new data that is consistent with the existing data set, using techniques such as categorical transformation, co-occurrence matrix analysis, and visualization tools to pinpoint missing data locations and fill gaps effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual review of test data sets is used to identify missing data, then data quality can be assessed, but labor costs and time consumption increase significantly
Solution Approach 1:
The patent replaces the manual mechanical review process with an automated computational system that uses data mining algorithms and machine learning models to identify missing data patterns, eliminating the need for human analysts to manually examine test data sets while maintaining or improving detection accuracy
Solution Approach 2:
The patent introduces an intermediary automated analysis system that acts as a bridge between raw test data and quality assessment results, using intermediate computational layers including data preprocessing, pattern recognition algorithms, and automated reporting mechanisms to facilitate efficient quality evaluation
2Reliability
If manual creation of data entries is used to fill gaps, then data completeness can be improved, but labor costs increase
Solution Approach 1:
The patent enables the test data set to serve itself by implementing automated data generation capabilities that use machine learning models to synthesize missing data entries based on learned patterns from existing data, eliminating the need for manual data creation while ensuring consistency with the overall data distribution
Solution Approach 2:
The patent transforms the approach to data completion by changing from manual parameter entry to automated parameter synthesis, using computational algorithms that generate data values based on statistical distributions and contextual relationships learned from the data set
3Reliability
If comprehensive test data coverage is pursued, then testing reliability improves, but data management complexity increases
Solution Approach 1:
The patent segments the complex task of achieving comprehensive test data coverage into manageable components including automated identification of missing data patterns, targeted generation of specific data types, and incremental integration of generated data, allowing systematic progression toward complete coverage without overwhelming complexity
Solution Approach 2:
The patent creates a universal automated system that can handle multiple types of data generation tasks across different test scenarios using the same underlying machine learning framework, reducing management complexity by providing a single platform for diverse data completion needs
Data Source
AI summary
A system, method and computer-readable medium for generation of missing data including transmitting indicators corresponding to locations of missing data in columns in tables of a database, each location of missing data corresponding to categorical values of categorical columns and each location of missing data being identified based on an expected count of data values at the corresponding location, receiving a selection of at least one indicator corresponding to at least one location of missing data, the at least one location of missing data corresponding to two or more categorical values of two or more categorical columns in the categorical columns, and generating sets of data records in at least one table in the tables of the database, each set of data records having two or more column values in two or more columns that correspond to the two or more categorical values of the two or more categorical columns.


