Cloud Dataset Performance Estimation via Sandbox Replication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Benchmarking operational performance of large datasets in a cloud computing environment is time-consuming, requiring significant time to build and populate datasets and execute performance tests, which hinders accurate performance estimation for both sandbox and production organizations.
Innovation Solution
A system estimates the time to perform operations on prospective datasets by examining execution times on similar datasets of varying sizes, extrapolating to estimate performance on larger datasets, using batch processes and calculating elapsed times for specific actions to provide accurate estimates without the need for extensive data setup and testing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If performance benchmarking is conducted on large datasets through actual execution of operations, then measurement precision is improved, but loss of time increases significantly
Solution Approach 1:
The patent creates a sandbox organization that replicates the structure and configuration of a production organization, allowing performance characteristics to be copied from the sandbox to the production environment. This copying approach eliminates the need to build and populate actual large datasets in production, thereby reducing time loss while maintaining measurement precision through accurate replication of performance characteristics.
Solution Approach 2:
The patent performs performance benchmarking in advance within a sandbox organization before deploying to production. By conducting preliminary performance tests on representative datasets in the sandbox environment, the system obtains performance measurements that can be applied to production without requiring actual execution on production datasets, thus resolving the time loss issue while preserving measurement precision.
2Reliability
If actual performance tests are executed on very large datasets, then reliability of performance estimates is improved, but loss of time increases due to extensive CPU cycles required
Solution Approach 1:
The system copies performance characteristics from a sandbox organization with representative data to a production organization. By measuring performance on copied or representative data structures in the sandbox environment and applying these measurements to production, the patent achieves reliable performance estimates without requiring extensive CPU cycles on actual production datasets.
Solution Approach 2:
Performance measurements are obtained in advance through preliminary benchmarking in the sandbox organization. These preliminary measurements provide reliable estimates for production performance without requiring the actual execution of CPU-intensive operations on production data, thereby reducing time loss while maintaining reliability.
3Measurement precision
If benchmarking is performed on increasingly larger datasets to ensure accuracy, then measurement precision is improved, but productivity decreases due to extended testing duration
Solution Approach 1:
The patent uses a sandbox organization that copies the essential characteristics and configuration of the production organization. This copying approach allows accurate performance measurement on representative data without requiring progressively larger actual datasets, thereby maintaining measurement precision while significantly improving the productivity of the benchmarking process by eliminating the need for extensive data growth and testing duration.
Data Source
AI summary
Estimating a time to perform an operation on a prospective data set of a selected size that includes a plurality of data entities and relationships between the data entities. A number of data sets of different size each comprising a number of like data entities and like relationships between the like data entities are received as input. A number of actions performed on a subset of the number of like data entities and like relationships between the like data entities that substantially comprise the operation are provided as output. For each of the number of data sets of different size, an elapsed time to perform a batch process for each of the number of actions on the subset of the number of like data entities and like relationships between the like data entities that comprise the operation is calculated. Finally, an elapsed time to perform the operation on the prospective data set based on its selected size and the elapsed times to perform, for each of the number of data sets of different size, the batch process for each of the number of actions on the subset of the number of like data entities and like relationships between the like data entities that comprise the operation is estimated, and provided as output.


