Sample Imprint for Optimized Dataset Comparison Testing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current comparison testing methods are inefficient in terms of time and resource utilization, as they often rely on entire datasets, which can be comprehensive but not practical for timely and efficient analysis, especially when updating applications from legacy to new versions.
Innovation Solution
The use of a sample imprint comprising combinatorial strings to select a representative dataset for parallel testing, optimizing the dataset to include specific attributes and reducing duplicate data, thereby improving processing efficiency and accuracy without overwhelming the system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If entire datasets are used for comparison testing, then measurement precision is improved, but productivity deteriorates
Solution Approach 1:
The patent segments the entire dataset into multiple strata based on common attributes (e.g., data types, value ranges, frequency distributions). Instead of processing one monolithic dataset, the system divides it into manageable segments that can be sampled independently, allowing for both comprehensive coverage and efficient processing.
Solution Approach 2:
The patent extracts a representative sample from each stratum of the dataset using optimized sampling techniques. This extraction process selects only the necessary portions of data needed for accurate comparison testing, eliminating the need to process the entire dataset while maintaining testing precision.
2Reliability
If entire datasets are used for comparison testing, then reliability is improved, but loss of time worsens
Solution Approach 1:
The patent performs preliminary stratification of the dataset before sampling, organizing data into meaningful groups based on attributes relevant to the testing objectives. This preliminary action ensures that the subsequent sampling process is more efficient and that the selected samples are representative, reducing overall testing time while maintaining reliability.
Solution Approach 2:
The patent changes the parameter of data representation by transforming the entire dataset into a condensed sample format that preserves key characteristics. By changing from processing all data points to processing optimized samples, the system maintains testing reliability while significantly reducing the time required.
3Productivity
If optimized sample datasets are used, then productivity is improved, but measurement precision deteriorates
Solution Approach 1:
The patent applies local quality by ensuring that each stratum or segment of the sampled data maintains the specific characteristics and quality attributes of the original population. Instead of uniform sampling, the system tailors the sampling approach to preserve local data qualities, ensuring that regional variations and attribute-specific patterns are accurately represented in the sample.
4Productivity
If optimized sample datasets are used, then productivity is improved, but reliability worsens
Solution Approach 1:
The patent creates a universal sampling framework that can be applied across different datasets and testing scenarios. The optimized sample dataset structure is designed to maintain reliability across multiple applications, allowing the same sampling methodology to ensure consistent testing reliability whether used for regression testing, performance testing, or validation across different system configurations.
Data Source
AI summary
An optimized test data selection strategy references a sampling file that identifies data attributes that serve as the basis of the test data selection strategy. By analyzing fields and the corresponding field values of the sample imprint, a total number of test data selected for inclusion into a sample dataset is reduced. The test data selection strategy provides an efficient methodology for implementing a data comparison testing process.


