Storage system performance determination method and device, electronic equipment and medium

The random stress test model simulates the sudden load changes in the storage system, solving the problem that the storage system performance cannot be accurately evaluated in the prior art, and achieving higher stability and reliability.

CN120540955APending Publication Date: 2025-08-26INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510682398.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The use of fixed or periodic load models in the prior art cannot accurately simulate the sudden and random input and output pressure of the storage system in the real business, resulting in performance bottlenecks or system crashes.

Method used

The random stress testing model is used to simulate randomly arriving request flow and burst load changes through the Poisson process, superimposed Hearst model and wavelet transform model, and combined with the historical pressure values ​​of the storage system, it is converted into input and output operations to test performance.

Benefits of technology

It can more accurately simulate pressure changes in real business scenarios, discover storage system problems in advance, avoid performance bottlenecks and system crashes, and improve the stability and reliability of the system under random pressure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540955A_ABST
    Figure CN120540955A_ABST
Patent Text Reader

Abstract

The invention discloses a storage system performance determination method and device, electronic equipment and a medium, and relates to the technical field of storage systems. In the method, the pressure in a real business scene is simulated based on a random pressure test model, and the problems existing in the storage system can be found in advance, so that a user can process the problems existing in the storage system in advance, and the system is prevented from reaching a performance bottleneck or collapsing; secondly, according to the historical pressure value and a random pressure test model or according to the historical pressure value, determining simulated pressure values of the storage system at different moments; the historical pressure value of the storage system reflects the actual pressure condition of the storage system, and the random pressure test model can simulate the actual pressure change condition, so that the determined pressure value can better conform to a real business scene; after the simulated pressure value is converted into the corresponding input and output operation, the determined performance of the storage system can accurately represent the performance of the storage system facing the real pressure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of storage systems, and in particular to a method, device, electronic equipment, and medium for determining storage system performance. Background Art

[0002] With the continuous development of information technology, data centers in industries such as finance have increasingly higher requirements for storage systems' ability to cope with random burst loads.

[0003] To ensure the stability of storage systems in actual applications, related technologies use fixed or periodic load models (such as a constant number of Input / Output Operations Per Second (IOPS)), such as step or sine wave stress tests. However, because these fixed or periodic load models are used, real-world businesses often experience sudden and irregular changes in Input / Output (IO) pressure. Therefore, these methods cannot simulate the sudden and random nature of pressure. As a result, when faced with sudden and random pressure in actual operation, the system may reach performance bottlenecks or even crash.

[0004] Therefore, how to simulate the changes in real-world pressure and test the performance of the storage system under pressure is a technical problem that people in this field urgently need to solve. Summary of the Invention

[0005] The present invention provides a method, device, electronic device and medium for determining storage system performance, so as to at least solve the problem in related technologies that it is impossible to simulate changes in pressure in reality and thus it is impossible to accurately test the performance of the storage system when facing real pressure.

[0006] The present invention provides a method for determining storage system performance, comprising:

[0007] Obtaining a pre-established random stress test model; wherein the random stress test model is determined by at least one or more of a model for characterizing and simulating a randomly arriving request flow, a model for measuring long-range dependencies of a time series, and a model for simulating sudden load changes;

[0008] Get the historical pressure value of the storage system;

[0009] Determine simulated pressure values ​​of the storage system at different times based on historical pressure values ​​and a random pressure test model, or determine simulated pressure values ​​of the storage system at different times based on historical pressure values;

[0010] Convert the simulated pressure values ​​of the storage system at different times into corresponding input and output operations;

[0011] The values ​​of the performance parameters of the storage system after the input and output operations are performed are obtained, and the performance of the storage system is determined according to the values ​​of the performance parameters of the storage system.

[0012] The present invention also provides a device for determining storage system performance, comprising:

[0013] A first acquisition module is configured to acquire a pre-established random stress test model; wherein the random stress test model is determined by at least one or more of a model for characterizing and simulating a randomly arriving request flow, a model for measuring long-range dependencies of a time series, and a model for simulating sudden load changes;

[0014] A second acquisition module is used to obtain historical pressure values ​​of the storage system;

[0015] A determination module, configured to determine simulated pressure values ​​of the storage system at different times based on historical pressure values ​​and a random pressure test model, or to determine simulated pressure values ​​of the storage system at different times based on historical pressure values;

[0016] A conversion module, used to convert the analog pressure values ​​of the storage system at different moments into corresponding input and output operations;

[0017] The acquisition and determination module is used to acquire the value of the performance parameter of the storage system after the input and output operations are performed, and determine the performance of the storage system according to the value of the performance parameter of the storage system.

[0018] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned methods for determining storage system performance when executing the computer program.

[0019] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods for determining storage system performance are implemented.

[0020] The beneficial effects of the present invention are that, first, compared with the method of using a fixed or periodic load model to test the performance of a storage system, the method provided by the present invention uses a random stress test model, and random stress is more consistent with the stress in real business scenarios. After simulating the stress in real business scenarios based on the random stress test model, problems with the storage system can be discovered in advance, so that users can deal with the problems in the storage system in advance, avoiding the system from reaching performance bottlenecks or even system crashes when facing sudden and random stress in actual operation; secondly, in order to simulate the pressure changes in reality, in this method, a pre-established random stress test model and the historical stress values ​​of the storage system are obtained; then, the simulated stress values ​​of the storage system at different times are determined based on the historical pressure values ​​and the random stress test model, or, the simulated stress values ​​of the storage system at different times are determined based on the historical pressure values. Since the historical stress values ​​of the storage system reflect the actual stress conditions of the storage system, the random stress testing model can simulate the actual changes in pressure. Therefore, the stress values ​​determined in this method can better conform to real business scenarios. Secondly, since the obtained simulated stress values ​​conform to real business scenarios, the simulated stress values ​​are converted into corresponding input and output operations. The performance of the storage system determined after executing the input and output operations can more accurately represent the performance of the storage system when facing real pressure.

[0021] The storage system performance determination device, electronic device, and computer-readable storage medium provided by the present invention have the same or corresponding technical features as the above-mentioned storage system performance determination method, and have the same effects as above. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] Figure 1 A flowchart of a method for determining storage system performance provided by an embodiment of the present invention;

[0024] Figure 2 A flowchart of a storage system performance testing method based on random pressure mutation is provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0026] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0027] To test storage system performance, related technologies use fixed or periodic load models, such as step-by-step or sinusoidal stress testing. In real-world environments, especially high-concurrency scenarios like e-commerce flash sales, I / O load can experience random, sudden peaks and valleys. Therefore, using fixed or periodic load models has limitations in simulating real-world scenarios and cannot replicate the sudden, irregular I / O load fluctuations found in real-world businesses. This makes it difficult to accurately assess storage system performance under extreme or sudden stress, leading to discrepancies between test results and actual production environments. For example, in e-commerce flash sales, user requests surge instantly (e.g., a 100-fold increase within one second of the sale), followed by a rapid decline due to stock sell-outs. Traditional step-by-step stress testing methods cannot replicate this "spike-and-dip" waveform. Another example is the surge in traffic caused by hot social media events, or the high-frequency trading requests in financial trading systems. These scenarios share the randomness and unpredictability of stress, making them difficult to capture with traditional testing methods. These tests may fail to simulate sudden and random fluctuations, leading to performance bottlenecks or even system crashes in real-world operations.

[0028] Therefore, the present invention provides a testing method that can generate a stress model containing random peaks and troughs to more realistically simulate extreme situations, help discover potential problems in the storage system in advance, and thus optimize the system architecture.

[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods. Figure 1 A flow chart of a method for determining storage system performance provided by an embodiment of the present invention is shown as follows: Figure 1 As shown, the method includes:

[0030] S10: Obtain a pre-established random stress test model;

[0031] S11: Obtain historical pressure values ​​of the storage system;

[0032] S12: Determine simulated pressure values ​​of the storage system at different times based on historical pressure values ​​and a random pressure test model, or determine simulated pressure values ​​of the storage system at different times based on historical pressure values;

[0033] S13: converting the simulated pressure values ​​of the storage system at different times into corresponding input and output operations;

[0034] S14: Obtain values ​​of performance parameters of the storage system after the input / output operation is performed, and determine the performance of the storage system according to the values ​​of the performance parameters of the storage system.

[0035] As described above, the use of fixed or periodic load models has limitations when simulating real scenarios and cannot simulate sudden and irregular IO pressure changes in real business. That is, it is difficult to accurately evaluate the performance of the storage system under extreme or sudden pressure. Therefore, a random stress testing model is established in an embodiment of the present invention. The random stress testing model is determined by at least one or more of a model for characterizing a simulated randomly arriving request stream, a model for measuring the long-range dependence of a time series, and a model for simulating sudden load changes. In order to make the established random stress testing model more consistent with the pressure changes in real business scenarios, in implementation, the random stress testing model is determined by a model for characterizing a simulated randomly arriving request stream, a model for measuring the long-range dependence of a time series, and a model for simulating sudden load changes.

[0036] There are no restrictions on the models used to characterize and simulate randomly arriving request flows; examples include the Poisson process and the Markovian arrival process (MAP). There are no restrictions on the models used to measure long-range dependencies in time series; examples include the Hurst exponent model and fractal analysis. There are no restrictions on the models used to simulate sudden load changes; examples include the wavelet transform model and the Markov modulated model.

[0037] The Poisson process model is suitable for quickly simulating performance tests under regular loads, offering simplicity and efficiency. The additive Hurst model is suitable for simulating complex load patterns with long-range dependencies, providing more realistic test results. The wavelet transform model is suitable for simulating complex scenarios involving both sudden and gradual load changes, offering strong adaptability and multi-resolution analysis. Therefore, when establishing a random stress testing model, the time-varying Poisson process model was selected to represent the simulated randomly arriving request stream; the additive Hurst model was used to measure the long-range dependencies of time series; and the wavelet transform model was used to simulate sudden load changes.

[0038] Specifically, in the Poisson process, at different moments in the Poisson process model Pressure data The expression is:

[0039] ;

[0040] in, represents a non-homogeneous Poisson distribution;

[0041] On the basis of the Poisson process, a self-similar pressure fluctuation model is generated by combining the long-range correlation fractal noise (such as fractional Brownian motion) controlled by the superposition Hurst exponent. The pressure data at different moments in the self-similar pressure fluctuation model The expression is as follows:

[0042] ;

[0043] in, Represents the long-range correlation noise controlled by the Hurst exponent. The Hurst exponent is a parameter used to measure the long-range dependence of a time series, and its value range is usually between 0 and 1:

[0044] Indicates that the sequence has no long-range dependencies (similar to standard Brownian motion).

[0045] Indicates that the sequence is persistent (i.e., past growth trends tend to continue, and past growth indicates future growth, such as traffic from hot events).

[0046] Indicates that the sequence is anti-persistent (trends are prone to reversal, such as a sudden drop in traffic after a short promotion).

[0047] The value of Hurst index H determines the long-range dependency of the time series. For example, in the e-commerce flash sale scenario, user requests may be persistent ( ), meaning that requests during peak hours last longer, while the opposite is true during off-peak hours. By adjusting the H value, you can simulate different business scenarios.

[0048] Based on the Poisson process model and the superimposed Hurst model, a wavelet transform model is added. The sudden peaks and troughs are injected through the wavelet transform.

[0049] Add wavelet transform to inject peaks and troughs, Morlet wavelet The expression is as follows:

[0050] ;

[0051] Pressure data at different moments in the random pressure model determined by the Poisson process model, superposition Hurst model and wavelet transform model The expression is:

[0052] ;

[0053] in, represents the amplitude of the wave, Represents the Morlet wavelet basis function, which is used to control the shape of peaks / troughs; Indicates the time position, used to control the time when the peak / trough occurs. Represents the scale parameter, which is used to control the width of the waveform. represents the sequence number of the wavelet basis function, Indicates the number of wavelet basis functions considered in the wavelet transform.

[0054] After determining the random stress test model, to ensure that the determined stress values ​​are more consistent with real business scenarios, the first step in implementation is to obtain historical stress values ​​for the storage system. These historical stress values ​​can be considered as empirical load values ​​in customer scenarios.

[0055] In implementation, obtaining historical pressure values ​​of the storage system includes:

[0056] Get the maximum and minimum pressure values ​​of all historical pressure values ​​in the storage system.

[0057] The maximum pressure value is the maximum pressure value in the flash sale scenario of the e-commerce platform. , after a period of time after the flash sale ends, it will drop to a lower value, that is, the minimum pressure value The control pressure intensity range is .

[0058] After obtaining the random stress test model and the historical stress values ​​of the storage system, the simulated stress values ​​of the storage system at different times can be determined based on the historical stress values ​​and the random stress test model, or the simulated stress values ​​of the storage system at different times can be determined based on the historical stress values.

[0059] Specifically, determining the simulated pressure values ​​of the storage system at different times based on the historical pressure values ​​includes:

[0060] When it is detected that the current time is less than the first preset time, the simulated pressure values ​​of the storage system at different times are determined according to the maximum pressure value; the first preset time is the time when the historical pressure value of the storage system changes from less than the preset pressure value to greater than the preset pressure value;

[0061] The simulated stress values ​​of the storage system at different times determined based on historical stress values ​​and the random stress test model include:

[0062] When it is detected that the current time is greater than or equal to the first preset time, the product of the maximum pressure value and the value of the random pressure test model at the current time is obtained, and the product is used as the simulated pressure value of the storage system at the current time.

[0063] The customer scenario events are as follows:

[0064] Phase 1 (t0-t1): The pressure value is less than the preset pressure value (i.e. steady-state IO load , for example, 5000 TPS);

[0065] Phase 2 (t1-t2): Market volatility events IO load surges and drops, such as 15000TPS, 2000 TPS.

[0066] Phase 3 (after t2): Pressure is released.

[0067] The first preset time is time t1. There is no limitation on the preset pressure value, which is determined according to actual conditions.

[0068] The composite pressure model that determines the simulated pressure values ​​of the storage system at different times based on historical pressure values ​​and the simulated pressure values ​​of the storage system at different times based on historical pressure values ​​and random pressure test models The expression is:

[0069] ;

[0070] It should be noted that This is the first preset time, which also represents the preheating time, which is generally 300s.

[0071] After calculating the pressure value at each moment (ie, the pressure value used in simulation, referred to as the simulated pressure value) according to the composite pressure model, the simulated pressure values ​​of the storage system at different moments can be converted into corresponding input and output operations.

[0072] Related testing methods lack the ability to test complex mutations (such as network latency jitter and concurrent disk queue congestion), resulting in significant deviations between the system's performance in real-world failures and laboratory data. For example, during a promotional period, the storage service crashed due to a metadata service avalanche. Subsequent analysis revealed that the stress test failed to simulate the combined effects of cache failure and log write blockage. Therefore, to test stress under complex mutations and better simulate real-world business scenarios, the simulated stress values ​​of the storage system at different times were converted into corresponding input and output operations, including the following:

[0073] Obtain a pre-established input / output model library, wherein different input / output operations, read / write ratios, block sizes, and queue depths are randomly set in the input / output model library;

[0074] Obtain the pressure value of the storage system under each input and output operation in the input and output model library;

[0075] The target input-output model is obtained from the input-output model library according to the simulated pressure value of the storage system at the current moment, so as to convert the simulated pressure values ​​of the storage system at different moments into corresponding input-output operations.

[0076] The input and output model library randomly sets different input and output operations, read-write ratios, block sizes, and queue depths. This is a mixed IO model. Based on the stages of customer scenario events, different IO operations are used at different stages to achieve the performance pressure value at time t. .

[0077] Phase 1 (t0-t1): Steady-state IO load Convert to a hybrid IO model and simulate it using IO reading and writing tools such as vdbench or FIO;

[0078] Phase 2 (t1-t2): Simulates a sudden increase in storage system performance to The spike simulates the sudden spike caused by the surge of customers. The data drops sharply to the trough, and the OLTP model is simulated by IO read and write tools such as vdbench or FIO.

[0079] In practice, the storage system may experience failures. To further simulate the situation where the storage system is under random stress and failures simultaneously, the performance of the storage system is tested in this scenario. In addition, in practice, there may be situations where the failure injection is not successful. Therefore, in order to inject the failure into the storage system and ensure that the failure is successfully injected into the storage system, in implementation, after converting the simulated stress values ​​of the storage system at different times into corresponding input and output operations, before obtaining the values ​​of the performance parameters of the storage system after executing the input and output operations and determining the performance of the storage system based on the values ​​of the performance parameters of the storage system, the following steps are also included:

[0080] Obtain a pre-established storage system failure mode library; wherein the failure modes in the storage system failure mode library are generated based on at least a destructive testing strategy, a continuous oscillation strategy, and / or a chaos injection strategy; the destructive testing strategy at least includes forcibly triggering the storage system's degradation / recovery mechanism; the continuous oscillation strategy at least includes simulating the impact of pressure oscillation within a preset duration on hard disk wear leveling; and the chaos injection strategy at least includes superimposing network delay jitter and hardware failure simulation;

[0081] Obtain different types of target failure modes from a storage system failure mode library;

[0082] Control storage system failures sequentially according to the order of target failure modes;

[0083] Obtain target features to be presented by the storage system under the current target failure mode;

[0084] After controlling a storage system failure based on a current target failure mode, obtaining current characteristics of the storage system;

[0085] When it is detected that the current feature is identical to the target feature, simulating an input / output operation using an input / output read / write tool, obtaining a value of a performance parameter of the storage system after the input / output operation is performed, and determining the performance of the storage system based on the value of the performance parameter of the storage system;

[0086] If it is detected that the current signature is different from the target signature, the process returns to the step of controlling the storage system failure based on the target failure mode.

[0087] This method injects multi-dimensional faults into the test scenario. Using a library of storage failure patterns, the following faults are randomly injected: a destructive test mode that forcibly triggers the storage system's degradation / recovery mechanisms; a sustained oscillation mode that simulates the effects of prolonged stress on wear leveling of solid-state drives (SSDs); and a chaos injection mode that superimposes interference factors such as network latency jitter and hardware failure simulations. During the test phase, storage failure patterns are injected to verify reliability and stability under random stress.

[0088] It should be noted that faults can be injected into the storage system at any stage of the customer scenario event. The storage system performs input and output operations under random stress and fault conditions to test the storage system performance.

[0089] There are no restrictions on the performance parameters of the storage system. Table 1 lists the key performance indicators. Related tests typically focus only on steady-state performance (such as average throughput) and ignore the system's self-recovery speed after a sudden pressure drop (such as the recovery time objective (RTO)). This can cause the system to experience secondary failures due to delayed resource release after a sudden pressure drop (such as memory leaks caused by untimely connection pool recovery). Therefore, the key performance indicators in this invention include failure recovery time.

[0090] Table 1

[0091]

[0092] To make the composite pressure model more accurate, in implementation, the method for determining storage system performance also includes:

[0093] Obtain the total number of test cases and the number of defects that detected storage system anomalies in all test cases;

[0094] Obtaining the ratio of the number of defects to the total number and determining the current defect density based on the ratio;

[0095] When it is detected that the current defect density is less than the defect density threshold, the parameters in the Poisson process and the superimposed Hurst model are increased to update the composite pressure model;

[0096] A new composite stress model is used to test storage system performance.

[0097] There is no limit on the defect density threshold, which can be customized by each manufacturer, for example, 0.1. Increase the parameters in the Poisson process and superimposed Hurst models, for example, by 30% (empirical value).

[0098] In this method, the composite stress model is updated according to the test results, and new defects are discovered through the new composite stress model. The stress test model is updated repeatedly, thereby improving the accuracy of the updated stress test model.

[0099] After the pressure is relieved (that is, after the system returns to a steady state), in order to obtain the performance of the storage system, after determining the performance of the storage system based on the values ​​of the storage system performance parameters, the following steps are also required:

[0100] Obtaining the current time and determining whether the current time is greater than a second preset time; wherein the second preset time is the time when the historical pressure value of the storage system changes from greater than the preset pressure value to less than the preset pressure value; and the second preset time is after the first preset time;

[0101] If so, stop simulating the input and output operations through the input and output read and write tools, and simulate the log monitoring process through the input and output tools to monitor the values ​​of the performance parameters of the storage system;

[0102] If not, return to the step of determining the simulated pressure values ​​of the storage system at different times according to the historical pressure values ​​and the random pressure test model, or determining the simulated pressure values ​​of the storage system at different times according to the historical pressure values.

[0103] The second preset time is t2 as described above. That is, in stage 3, the current IO is stopped and the log IO model is started using the IO read and write tool to monitor the performance of the storage system.

[0104] After determining the performance of the storage system based on the values ​​of its performance parameters, potential problems may be discovered. To prevent potential problems from affecting real services and improve the reliability of the storage system, after obtaining the values ​​of the storage system's performance parameters after performing input and output operations and determining the performance of the storage system based on the values ​​of the storage system's performance parameters, the following steps are also performed:

[0105] When a performance anomaly of the storage system is detected, obtain current anomaly information of the storage system;

[0106] Determine the exception handling strategy corresponding to the current exception information based on the pre-established correspondence between the exception information and the exception handling strategy;

[0107] Handle the storage system according to the exception handling policy.

[0108] In this method, after a potential problem in the storage system is detected, the abnormality of the storage system is processed, thereby avoiding the impact of the potential problem on real business and improving the reliability of the storage system.

[0109] In order to enable those skilled in the art to better understand the storage system performance testing method based on random pressure mutation of the present invention, the solution of the present invention is described again below with reference to the accompanying drawings and specific embodiments. Figure 2 A flow chart of a storage system performance testing method based on random pressure mutation provided by an embodiment of the present invention is shown in FIG. Figure 2 As shown, the method includes:

[0110] S15: Obtaining the minimum pressure value, the maximum pressure value and the composite pressure model;

[0111] S16: Using a composite pressure model to calculate the pressure value of the storage system at each moment in real time;

[0112] S17: Converting the pressure value at each moment into a specific input-output model through a load converter;

[0113] S18: Start random stress testing through the input and output operation tool to simulate the input and output stress model at different times;

[0114] S19: Randomly inject storage faults during this period to check the impact of the faults;

[0115] S20: Test result inspection and feedback;

[0116] S21: Fix defects and update stored procedures.

[0117] The following uses a mixed business stress test on an all-flash array as an example to verify the stability of the all-flash array under mixed loads (OLTP transactions + data analysis), simulating the sudden stress changes caused by market fluctuations in financial trading scenarios.

[0118] 1. Establishing customer pressure model:

[0119] Divide the customer scenario into multiple stages and substitute them into the subsequent pressure function to calculate the pressure values ​​at different time points. For example, the customer scenario stages can be divided into:

[0120] Phase 1 (t0-t1): Daily operation, performance is stable;

[0121] Phase 2 (t1-t2): Fluctuations caused by a surge in market customers, sudden spikes or dips; injection of unexpected storage failures;

[0122] Phase 3 (after t2): Pressure relief, return to steady state, and monitoring of storage system processes.

[0123] 2. Establishing a composite pressure model: The embodiment of establishing a composite pressure model has been described in detail above and will not be repeated here.

[0124] 3. Convert the load converter to a testable IO model:

[0125] Convert the stress value obtained in step 2 into a testable IO model (random read / write ratio, block size, queue depth), and use different IO operations at different stages to achieve the performance stress value at time t .

[0126] Phase 1 (t0-t1): Steady-state IO load Convert to a mixed I / O model (70% of the OLTP model load and 30% of the OLAP model load). Use I / O read / write tools such as vdbench or FIO to simulate the OLTP model and storage system performance.

[0127] Phase 2 (t1-t2): Simulates a sudden increase in storage system performance to The spike simulates the sudden spike caused by the surge of customers. The data drops sharply to the trough, and the OLTP model is simulated by IO read and write tools such as vdbench or FIO.

[0128] Phase 3 (after t2): Stop I / O operations and use I / O reading and writing tools such as vdbench or FIO to simulate the log monitoring process.

[0129] 4. Inject storage failure modes during the testing phase to verify reliability and stability under random stress.

[0130] By storing the fault mode library, relevant faults that affect performance are injected. Table 2 is a table of system reliability failure modes and their corresponding test methods.

[0131] Table 2

[0132]

[0133] 5. Stress test model injection process:

[0134] Phase 1 (t0-t1): Use I / O read / write tools to start OLTP+OLAP testing to simulate steady-state performance.

[0135] Phase 2 (t1-t2): Use I / O read / write tools to increase OLTP performance to 15,000 TPS and simulate market volatility events. Use a storage failure pattern library to inject failures and superimpose storage failures to test reliability and stability.

[0136] Phase 3 (after t2): Stop the current IO and use the IO read and write tools to start the log IO model.

[0137] 6. Test results and optimization:

[0138] The problems discovered during the testing process were categorized. Random pressure mutation scenarios often resulted in problems such as a sharp increase in write amplification, the superposition of garbage collection and IO performance spikes leading to IO processing timeouts, and small-scale IO timeouts caused by network jitter.

[0139] For the case of a sudden increase in write amplification: If the write amplification problem is caused by insufficient storage device capacity, you can consider increasing the capacity of the storage device to provide more space for the garbage collection mechanism and reduce the frequency of garbage collection.

[0140] To address the issue of IO processing timeouts caused by the superposition of garbage collection and IO performance spikes: Upgrade the hardware performance of the storage device, such as using higher-performance flash memory chips and increasing the cache capacity of the storage device, to improve the storage device's processing capabilities when garbage collection and IO operations are performed simultaneously.

[0141] Regarding the small-scale IO timeout problem caused by network jitter: In the distributed storage system, a redundant network design is adopted. When a network link fails or jitters, it can automatically switch to the backup link to ensure that communication between storage nodes is not affected.

[0142] The present invention provides a performance testing method for random pressure mutations based on a composite pressure model (i.e., based on composite mathematical transformations). By introducing a random pressure model based on mathematical transformations (such as Poisson process, fractal noise, and wavelet transform), the method realistically simulates random pressure mutations, verifies the storage system's ability to respond to random burst loads (such as flash sale scenarios and batch log writing), discovers potential problems in advance, and evaluates the stability and reliability of the storage system under the composite pressure model. After the test is completed, the test model is fed back to the system to update the influencing factors in the model, making the test model more accurate. The fixed load (such as constant IOPS) or periodic model (such as sine wave and step-by-step increase) commonly used in existing storage tests is no longer used. Instead, the composite pressure model is used to simulate sudden and irregular IO pressure changes in real business.

[0143] Compared with related testing methods, this method uses a composite model of non-stationary Poisson process + fractal noise + wavelet transform to accurately reproduce the random mutation characteristics of business scenarios such as flash sales and hot events. The waveform similarity between the pressure curve and the actual production data reaches 92%.

[0144] Through multi-dimensional interference simulation, it supports the simultaneous injection of hardware-level (power supply fluctuations, temperature changes), software-level (network packet loss) and service-level (mixed load) interference, covering 6 major categories and 27 subcategories of fault scenarios.

[0145] Compared with related performance stress test scenarios, after using this test model for stress testing, the storage system's IOPS increased by 27% under sudden stress scenarios.

[0146] Furthermore, this approach unifies test benchmarks and provides a quantifiable set of stress model parameters (such as the range of the fractal noise Hurst exponent and a library of wavelet basis functions), promoting the development of industry testing standards. Mathematical modeling creates a highly realistic stress testing environment suitable for scenarios such as storage device R&D verification and cloud computing service level agreement (SLA) evaluation.

[0147] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0148] An embodiment of the present invention further provides a device for determining storage system performance, comprising:

[0149] A first acquisition module is configured to acquire a pre-established random stress test model; wherein the random stress test model is determined by at least one or more of a model for characterizing and simulating a randomly arriving request flow, a model for measuring long-range dependencies of a time series, and a model for simulating sudden load changes;

[0150] A second acquisition module is used to obtain historical pressure values ​​of the storage system;

[0151] A determination module, configured to determine simulated pressure values ​​of the storage system at different times based on historical pressure values ​​and a random pressure test model, or to determine simulated pressure values ​​of the storage system at different times based on historical pressure values;

[0152] A conversion module, used to convert the analog pressure values ​​of the storage system at different moments into corresponding input and output operations;

[0153] The acquisition and determination module is used to acquire the value of the performance parameter of the storage system after the input and output operations are performed, and determine the performance of the storage system according to the value of the performance parameter of the storage system.

[0154] In some embodiments, the second acquisition module specifically includes:

[0155] A third acquisition module is used to obtain the maximum pressure value and the minimum pressure value of all historical pressure values ​​of the storage system;

[0156] Determine the module specifically for:

[0157] a determination submodule, configured to determine, when detecting that the current moment is less than a first preset moment, a simulated pressure value of the storage system at a different moment based on the maximum pressure value; wherein the first preset moment is a moment when the historical pressure value of the storage system changes from less than a preset pressure value to greater than a preset pressure value;

[0158] Determine the module specifically for:

[0159] The fourth acquisition module is used to obtain the product of the maximum pressure value and the value of the random pressure test model at the current moment when it is detected that the current moment is greater than or equal to the first preset moment, and use the product as the simulated pressure value of the storage system at the current moment.

[0160] In some embodiments, the conversion module specifically includes:

[0161] A fifth acquisition module is used to acquire a pre-established input-output model library, wherein different input-output operations, read-write ratios, block sizes, and queue depths are randomly set in the input-output model library;

[0162] A sixth acquisition module is used to obtain the pressure value of the storage system under each input and output operation in the input and output model library;

[0163] The acquisition and conversion module is used to obtain the target input-output model from the input-output model library according to the simulated pressure value of the storage system at the current moment, so as to convert the simulated pressure values ​​of the storage system at different moments into corresponding input-output operations.

[0164] In some embodiments, the means for storing system performance further comprises:

[0165] A seventh acquisition module is configured to acquire a pre-established storage system failure mode library; wherein the failure modes in the storage system failure mode library are generated based on at least a destructive testing strategy, a sustained oscillation strategy, and / or a chaos injection strategy; the destructive testing strategy includes at least forcibly triggering a degradation / recovery mechanism for the storage system; the sustained oscillation strategy includes at least simulating the impact of pressure oscillation within a preset duration on hard disk wear leveling; and the chaos injection strategy includes at least superimposing network delay jitter and hardware failure simulation.

[0166] an eighth acquisition module, configured to acquire different types of target failure modes from a storage system failure mode library;

[0167] A control module, configured to sequentially control storage system failures according to the order of target failure modes;

[0168] A ninth acquisition module, configured to acquire target features to be presented by the storage system under the current target failure mode;

[0169] a tenth acquisition module, configured to acquire current characteristics of the storage system after controlling a storage system failure based on a current target failure mode;

[0170] The detection module is used to simulate input and output operations through input and output reading and writing tools and trigger the acquisition and determination module when it detects that the current feature is the same as the target feature; when it detects that the current feature is different from the target feature, it returns to the trigger control module.

[0171] In some embodiments, the means for storing system performance further comprises:

[0172] a determination module, configured to obtain a current time and determine whether the current time is greater than a second preset time; wherein the second preset time is the time when the historical pressure value of the storage system changes from greater than a preset pressure value to less than a preset pressure value; and the second preset time is after the first preset time; if so, triggering the stop and monitoring module; if not, returning to the trigger determination module;

[0173] The stop and monitoring module is used to stop simulating input and output operations through the input and output reading and writing tools, and simulate the log monitoring process through the input and output tools to monitor the values ​​of performance parameters of the storage system.

[0174] In some embodiments, the means for storing system performance further comprises:

[0175] an eleventh obtaining module, configured to obtain current abnormality information of the storage system when a performance abnormality of the storage system is detected;

[0176] A first determining module is used to determine the exception handling strategy corresponding to the current exception information based on a pre-established correspondence between the exception information and the exception handling strategy;

[0177] The processing module is used to process the storage system according to the exception handling strategy.

[0178] For descriptions of features in the embodiments corresponding to the apparatus for storing system performance, reference may be made to the relevant descriptions of the embodiments corresponding to the method for storing system performance, which will not be detailed here.

[0179] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the method for storing system performance.

[0180] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned embodiments of the method for storing system performance when running.

[0181] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0182] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned storage system performance method embodiments are implemented.

[0183] An embodiment of the present invention also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned storage system performance method embodiments.

[0184] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0185] The above is a detailed introduction to the method, device, electronic device, and medium for determining storage system performance provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core concept of the present invention. It should be pointed out that, for those skilled in the art, several improvements and modifications can be made to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the scope of protection of the present invention.

Claims

1. A method for determining storage system performance, characterized in that: include: Obtaining a pre-established random stress test model; wherein the random stress test model is determined by at least one or more of a model for characterizing a simulated randomly arriving request flow, a model for measuring long-range dependencies of a time series, and a model for simulating sudden load changes; Get the historical pressure value of the storage system; Determining simulated pressure values ​​of the storage system at different times based on historical pressure values ​​and the random pressure test model, or determining simulated pressure values ​​of the storage system at different times based on the historical pressure values; Convert the simulated pressure values ​​of the storage system at different times into corresponding input and output operations; The values ​​of the performance parameters of the storage system after the input and output operations are performed are obtained, and the performance of the storage system is determined according to the values ​​of the performance parameters of the storage system.

2. The method for determining storage system performance according to claim 1, wherein: The random stress testing model is determined by a model for characterizing and simulating a randomly arriving request flow, a model for measuring the long-range dependency of a time series, and a model for simulating sudden load changes; wherein, the model for characterizing and simulating a randomly arriving request flow is a model of a Poisson process that varies with time; the model for measuring the long-range dependency of a time series is a superimposed Hurst model; and the model for simulating sudden load changes is a wavelet transform model.

3. The method for determining storage system performance according to claim 2, wherein: Obtaining the historical pressure value of the storage system includes: Get the maximum and minimum pressure values ​​of all historical pressure values ​​in the storage system; Determining the simulated pressure values ​​of the storage system at different times according to the historical pressure values ​​includes: When it is detected that the current time is less than a first preset time, the simulated pressure values ​​of the storage system at different times are determined according to the maximum pressure value; wherein the first preset time is the time when the historical pressure value of the storage system changes from less than the preset pressure value to greater than the preset pressure value; Determining the simulated pressure values ​​of the storage system at different times according to the historical pressure values ​​and the random pressure test model includes: When it is detected that the current moment is greater than or equal to the first preset moment, the product of the maximum pressure value and the value of the random pressure test model at the current moment is obtained, and the product is used as the simulated pressure value of the storage system at the current moment.

4. The method for determining storage system performance according to any one of claims 1 to 3, wherein: The converting of the simulated pressure values ​​of the storage system at different times into corresponding input and output operations includes: Obtaining a pre-established input-output model library; wherein the input-output model library randomly sets different input-output operations, read-write ratios, block sizes, and queue depths; Obtain the pressure value of the storage system under each input and output operation in the input and output model library; The target input-output model is obtained from the input-output model library according to the simulated pressure value of the storage system at the current moment, so as to convert the simulated pressure values ​​of the storage system at different moments into corresponding input-output operations.

5. The method for determining storage system performance according to claim 3, wherein: After converting the simulated pressure values ​​of the storage system at different moments into corresponding input / output operations, and before obtaining the values ​​of the performance parameters of the storage system after the input / output operations are performed and determining the performance of the storage system according to the values ​​of the performance parameters of the storage system, the method further includes: Obtain a pre-established storage system failure mode library; wherein the failure modes in the storage system failure mode library are generated based on at least a destructive testing strategy, a continuous oscillation strategy, and / or a chaos injection strategy; the destructive testing strategy at least includes forcibly triggering a degradation / recovery mechanism for the storage system; the continuous oscillation strategy at least includes simulating the impact of pressure oscillation within a preset duration on hard disk wear leveling; and the chaos injection strategy at least includes superimposing network delay jitter and hardware failure simulation; Obtain different types of target failure modes from a storage system failure mode library; Control storage system failures sequentially according to the order of target failure modes; Obtain target features to be presented by the storage system under the current target failure mode; After controlling a storage system failure based on a current target failure mode, obtaining current characteristics of the storage system; When it is detected that the current feature is identical to the target feature, simulating an input / output operation using an input / output read / write tool, obtaining a value of a performance parameter of the storage system after the input / output operation is performed, and determining the performance of the storage system according to the value of the performance parameter of the storage system; In the case where it is detected that the current feature is different from the target feature, the process returns to the step of controlling the storage system failure based on the target failure mode.

6. The method for determining storage system performance according to claim 5, wherein: After determining the performance of the storage system according to the value of the performance parameter of the storage system, the method further includes: Obtaining the current time and determining whether the current time is greater than a second preset time; wherein the second preset time is the time when the historical pressure value of the storage system changes from greater than the preset pressure value to less than the preset pressure value; and the second preset time is after the first preset time; If so, stop simulating the input and output operations through the input and output read and write tools, and simulate the log monitoring process through the input and output tools to monitor the values ​​of the performance parameters of the storage system; If not, return to the step of determining the simulated pressure values ​​of the storage system at different times according to the historical pressure values ​​and the random pressure test model, or determining the simulated pressure values ​​of the storage system at different times according to the historical pressure values.

7. The method for determining storage system performance according to claim 1, wherein: After obtaining the value of the performance parameter of the storage system after the input / output operation is performed and determining the performance of the storage system according to the value of the performance parameter of the storage system, the method further includes: When a performance anomaly of the storage system is detected, obtain current anomaly information of the storage system; Determine the exception handling strategy corresponding to the current exception information based on the pre-established correspondence between the exception information and the exception handling strategy; Handle the storage system according to the exception handling policy.

8. A device for determining storage system performance, characterized in that: include: A first acquisition module is configured to acquire a pre-established random stress test model; wherein the random stress test model is determined by at least one or more of a model for characterizing a simulated randomly arriving request flow, a model for measuring long-range dependencies of a time series, and a model for simulating sudden load changes; A second acquisition module is used to obtain historical pressure values ​​of the storage system; a determination module, configured to determine simulated pressure values ​​of the storage system at different times based on historical pressure values ​​and the random pressure test model, or to determine simulated pressure values ​​of the storage system at different times based on the historical pressure values; A conversion module, used to convert the analog pressure values ​​of the storage system at different moments into corresponding input and output operations; The acquisition and determination module is used to acquire the value of the performance parameter of the storage system after the input and output operations are performed, and determine the performance of the storage system according to the value of the performance parameter of the storage system.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for determining storage system performance according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method for determining storage system performance according to any one of claims 1 to 7 are implemented.