ML-Based Data Storage Performance Test Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional performance testing of data storage systems is bottlenecked by the need for manual analysis, which is time-consuming and prone to human error, especially as systems become complex and generate large volumes of test result data, making it difficult to differentiate valid from invalid test results.
Innovation Solution
A method employing machine learning technology to automate performance testing by using selected features to train a model that assesses the validity of test results, reducing the need for manual review and improving assessment quality by distinguishing valid from invalid performance test results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual analysis is used to review performance test results, then assessment quality can be maintained through human judgment, but testing time increases significantly and human error occurs
Solution Approach 1:
The patent replaces the mechanical system of manual human review with an automated machine learning-based analysis system. The ML model processes performance test results, operating parameters, and performance data to automatically determine test validity, eliminating the need for manual human judgment while maintaining consistent assessment quality across large volumes of test data.
Solution Approach 2:
The system enables self-service automation where the performance testing system automatically validates its own test results through the ML model. The automated analysis system processes test data, compares it against learned patterns from training data, and independently determines validity without requiring external human intervention, thereby reducing both time and human error.
2Reliability
If manual review processes are used to validate test results, then human judgment can identify valid tests, but the process becomes bottlenecked and difficult to scale with complex systems
Solution Approach 1:
The patent substitutes manual human review processes with an automated machine learning system that processes performance test results. The ML model is trained on historical performance data and operating parameters to learn patterns of valid versus invalid tests, enabling reliable automatic differentiation of test validity even as system complexity increases and test volumes grow.
Solution Approach 2:
The system transforms the validation process by changing from subjective human judgment to objective algorithmic parameter-based analysis. The ML model analyzes multiple parameters including operating parameters, performance data, and test configuration data to objectively determine validity, providing consistent and scalable reliability assessment independent of human factors.
3Quantity of substance
If more comprehensive performance testing is conducted on complex systems, then more valid performance data can be obtained, but the volume of test result data increases making manual analysis difficult
Solution Approach 1:
The patent replaces manual analysis operations with automated machine learning-based processing. The ML system efficiently handles large volumes of performance test result data by automatically processing, analyzing, and validating tests based on learned patterns from training data, making the analysis of comprehensive performance testing data tractable and scalable.
Solution Approach 2:
The system uses training data that copies historical performance patterns and characteristics to create a reference model. This trained ML model then serves as a reusable template for automatically evaluating new test results, enabling the system to handle increasing volumes of performance data without proportionally increasing analysis difficulty or time requirements.
Data Source
AI summary
In performance testing a data storage system, operating parameters and performance data are recorded as the system executes performance tests over a test period, where the performance data includes measures of a performance characteristic (e.g., latency) across a range of I/O operation rates or I/O data rates for each performance test. Subsets of recorded operating parameters and performance data are selected and applied to a machine learning model to train and use the model, and the model provides a model output indicative for each performance test of a level of validity of the corresponding performance data. Based on the model output indicating at least a predetermined level of validity for a given performance test, the performance data for the performance test are incorporated into a record of validated performance data for the data storage system, usable for benchmarking, regression analysis, hardware qualification, etc.


