Sampling Historical Users for Document Testing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Electronic document preparation systems face inefficiencies in generating sample data sets that cover all use cases, leading to resource-intensive and often inadequate testing processes, which can result in delays and inaccuracies.
Innovation Solution
The system generates training sets by executing previous software code for historical users, grouping them based on executed code sections, and sampling a small number of users from each group to create a representative set that covers virtually all user attributes, ensuring accurate testing with reduced resource usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a very large number of historical users are used for testing, then the coverage of user attributes is improved, but the computing and human resources required increase significantly
Solution Approach 1:
The patent segments the historical user population into distinct groups based on their attributes and characteristics. By dividing the large dataset into meaningful segments, the system can select representative samples from each segment rather than processing all users, thus maintaining attribute coverage while reducing computational resources required for testing
Solution Approach 2:
The patent creates synthetic test data that copies and replicates the characteristics and patterns found in historical user data. This allows the system to generate sufficient test cases with rare attributes without needing to process actual large volumes of historical user records, thereby reducing computing resources while maintaining testing reliability
2Productivity
If a smaller random sample of historical users is used for testing, then the computing and human resources are reduced, but the coverage of rare user attributes becomes inadequate
Solution Approach 1:
The patent performs preliminary analysis of historical user data to identify distinct user groups, rare attributes, and important patterns before selecting the test sample. This preliminary action ensures that the subsequent smaller sample is strategically chosen to include rare attributes, eliminating the need for large-scale testing while maintaining comprehensive coverage
Solution Approach 2:
The patent changes the selection parameters from random sampling to stratified or targeted sampling based on identified user groups and attribute frequencies. By adjusting the sampling parameters to prioritize rare attributes, the system achieves adequate coverage with a smaller sample size, improving testing efficiency without sacrificing reliability
3Reliability
If traditional testing processes are used to ensure accurate handling of all user attributes, then the reliability of the system is improved, but the time and expense increase significantly
Solution Approach 1:
The patent creates synthetic test cases that replicate rare and edge-case user attributes without requiring actual processing of large volumes of historical user data. This copying approach maintains testing accuracy for rare scenarios while dramatically reducing the time and expense associated with traditional comprehensive testing processes
Data Source
AI summary
A method and system generate sample data set for efficiently and accurately testing a new calculation for preparing a portion of an electronic document for users of an electronic document preparation system. The method and system receive the new calculation and gather historical use data related to previously prepared electronic documents for a large number of historical users. The method and system group the historical users into groups based on which sections of a previous version of electronic document preparation software were executed for each historical user in preparing electronic documents for the historical users. The groups are then sampled by selecting a small number of historical users from each group.


