Training method for performance prediction model of storage system, and computing device
By training a performance prediction model based on the impact of back-end task configuration, the problem of inaccurate prediction of storage system performance in the prior art is solved, and more accurate performance prediction and service quality reflection are achieved.
Patent Information
- Application Number
- PCT/CN2024/133952
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-07
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-12
AI Technical Summary
The prior art cannot accurately predict the performance of storage systems when running workloads, especially the impact of background tasks on performance, resulting in the performance prediction value that cannot accurately reflect the service quality of the storage system.
By training a performance prediction model that predicts the performance of the storage system when running a workload based on the degree of impact of background task configuration on system performance. The method includes determining the load configuration and background task configuration of the workload, inputting it as input data into the performance prediction model to improve the accuracy of performance prediction.
By considering the impact of background tasks on performance, the accuracy of the performance prediction of storage system is improved, so that the performance prediction value more accurately reflects the service quality of the storage system, and solves the problem of inaccurate prediction values in the prior art.
Smart Images

Figure CN2024133952_12062025_PF_FP_ABST
Abstract
Description
A training method for a performance prediction model of a storage system and a computing device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 7, 2023, with application number 202311683871.3 and application name “A training method and computing device for a performance prediction model of a storage system”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a training method and computing device for a performance prediction model of a storage system. Background Art
[0003] Storage systems and their user services are becoming increasingly complex. To ensure that the quality of service (QoS) of storage systems running user service workloads matches user service requirements, it is often necessary to predict the performance of storage systems running user service workloads.
[0004] In related technologies, a performance fitting formula is pre-determined based on storage system expert experience. This formula is then used to process measured data from the storage system to generate performance predictions for the storage system under load. However, the performance predictions generated using this performance fitting formula cannot accurately reflect the service quality of the storage system. Summary of the Invention
[0005] The present application provides a training method and computing device for a performance prediction model of a storage system. The performance prediction model predicts the performance of the storage system when running workloads based on the degree of impact of background task configuration on system performance, which can improve the accuracy of performance prediction.
[0006] In a first aspect, the present application provides a method for predicting the performance of a storage system. The method comprises: determining a workload configuration of the storage system, and determining a background task configuration of the storage system when the storage system executes the workload; inputting input data into a performance prediction model, and determining a performance prediction value output by the performance prediction model for the storage system when executing the workload. The input data includes the workload configuration and the degree of impact of the background task configuration on the performance of the storage system.
[0007] The workload may include a read / write (input / output, IO) task of the storage system, and the number of workloads may include one or more.
[0008] The background task configuration includes configuration parameters for one or more background tasks of the storage system. Configuration parameters may include, for example, status and / or concurrency, with the status including enabled or disabled. The impact level of the background task configuration represents the extent to which the performance of the storage system is reduced due to the background task configuration.
[0009] In specific application scenarios, the background tasks of a storage system will occupy the system resources of the storage system (such as computing resources and storage resources). The system resources of the storage system cannot be fully used to run workloads, resulting in a performance degradation of the storage system when running workloads. When the above solution uses a pre-trained performance prediction model to predict the performance of the storage system when running workloads, it considers the impact of the storage system's background task configuration on the performance of the storage system when running workloads. That is, the background task configuration is used as an input to the performance prediction model. This can improve the accuracy of the storage system performance prediction and enable the performance prediction value to more accurately reflect the service quality of the storage system.
[0010] In a possible implementation of the first aspect, the degree of influence of the background task configuration can be determined based on a first measured performance value of the storage system running a historical workload under the background task configuration, and a second measured performance value of the storage system running a historical workload alone.
[0011] The degree of impact of the background task configuration can be specifically determined using a pre-established background task impact model based on regression analysis. The background task impact model based on regression analysis can be trained using samples corresponding to different background task configurations. The samples corresponding to each background task configuration include the actual performance values of the storage system running each historical workload under the background task configuration, and the actual performance values of the storage system running only each historical workload.
[0012] The impact of the background task configuration can also be obtained by inputting the load configuration and the background task configuration into a pre-trained background task impact model, and the background task impact model outputs the impact. The background task impact model represents the mapping relationship between the background task configuration and the impact, and can be obtained by training based on the load configuration and background task configuration of the historical workload of the storage system, as well as the impact of the background task configuration.
[0013] In a possible implementation of the first aspect, when the storage system includes multiple workloads, the method further includes: determining performance prediction values of the multiple storage systems when running the workloads based on load configurations of the multiple workloads and performance prediction values of each storage system when running the workloads. The performance prediction values may include: the number of read / write operations per second (IOPS), latency, and / or throughput.
[0014] The above solution can be applied to scenarios where multiple workloads are running together, thereby predicting the performance of a storage system running these mixed workloads. In scenarios where multiple workloads are running together, each workload corresponds to the same background task configuration. The performance predictions for each workload running alone can be used to derive the performance value for the mixed operation, improving the accuracy of the performance predictions for the mixed operation of multiple workloads.
[0015] In one possible implementation of the first aspect, determining a workload configuration of the storage system includes receiving a workload configuration from a user interface of the storage system. A user of the storage system may set the workload configuration through the user interface. Of course, the user may set the load configuration of one or more workloads through the user interface. The load configuration may include I / O size, read-write ratio, randomness, concurrency, mix ratio, and / or runtime.
[0016] The above solution can support users to set workload configuration, and predict the performance of the storage system when running the workload based on the load configuration set by the user, so that the storage system can better meet the user's QoS requirements for the storage system.
[0017] In a possible implementation of the first aspect, the background task configuration includes: a first background task configuration and / or a second background task configuration. The first background task configuration includes a configuration that cannot be modified by the user, such as but not limited to a garbage collection task (GC) configuration, an overload control task configuration, and a data reconstruction task configuration. The GC configuration may include the status and / or concurrency of the GC, and the overload control task configuration and the data reconstruction task configuration may include a status, wherein the aforementioned status includes on or off. The second background task configuration includes a configuration that can be modified by the user, such as but not limited to a concurrency of a remote replication task, and / or a concurrency of a deduplication task.
[0018] The storage system can run multiple types of background tasks, for example, a first background task whose configuration cannot be modified by the user and a second background task whose configuration can be modified by the user. The above solution can further improve the accuracy of the performance prediction value by predicting performance based on multiple background task configurations.
[0019] In a possible implementation of the first aspect, the background task configuration of the storage system when the storage system runs the workload is determined, including: determining the first background task configuration according to the load configuration of the workload. Taking the first background task configuration including the status of the GC and the status of the overload control task as an example, the storage system can determine the status of the GC and the status of the overload control task according to the running time in the load configuration of the workload. Specifically, when the running time is greater than a preset threshold corresponding to the GC, the GC status is determined to be on, otherwise it is off; similarly, when the running time is greater than a preset threshold corresponding to the overload control task, the overload control task status is determined to be on, otherwise it is off.
[0020] In a possible implementation of the first aspect, when the GC state is determined to be enabled, the capacity of the storage system occupied by IO can be determined based on the workload load configuration, and then the GC concurrency can be determined based on a predetermined correspondence between the capacity of the storage system occupied by IO and the GC concurrency. The process of determining the capacity of the storage system occupied by IO can include: determining the IO rate corresponding to the concurrency in the workload load configuration based on the correspondence between the IO concurrency and the IO rate, then determining the IO write capacity per unit time based on the IO rate and the product of the IO size in the workload load configuration and the read-write ratio, and finally obtaining the capacity of the storage system occupied by IO based on the IO write capacity per unit time and the running time.
[0021] In the above solution, the first background task configuration is determined according to the workload configuration, so that the first background task can better serve the workload, thereby improving the QoS of the storage system.
[0022] In a possible implementation of the first aspect, the method of determining the background task configuration of the storage system when the storage system runs a workload further includes: receiving a requirement for a second background task from a user interface of the storage system; and determining the second background task configuration based on the requirement for the second background task. The user of the storage system can set the requirement for the second background task through the user interface. Taking the concurrency of the remote replication task as an example, the user can set the requirement for the remote replication task through the user interface, and then the storage system determines the concurrency of the remote replication task based on the requirement for the remote replication task. The requirement for the remote replication task can include the rate requirement for the remote replication task. Taking the concurrency of the deduplication task as an example, the user can set the requirement for the deduplication task through the user interface, and then the storage system determines the concurrency of the deduplication task based on the requirement for the deduplication task. The requirement for the remote replication task can include the rate requirement for the deduplication task.
[0023] In the above solution, the second background task configuration is determined according to the requirements of the second background task set by the user, so that the second background task can better serve the workload, thereby improving the QoS of the storage system when running the workload.
[0024] In a possible implementation of the first aspect, the input data further includes operating indicators and / or hardware configuration of the storage system, wherein the operating indicators may include central processing unit (CPU) utilization, hard disk utilization, and / or network utilization.
[0025] In the above scheme, the impact of the storage system's operating indicators, hardware configuration, and background task configuration on the storage system's performance is used as the model input. This can improve the model's prediction accuracy and enable the performance prediction value output by the model to more accurately reflect the QoS of the storage system when running workloads.
[0026] In a second aspect, the present application provides a method for predicting storage system performance. The method includes: displaying a first user interface for setting a load configuration for a storage system workload; and displaying a second user interface for displaying a performance prediction value for the storage system when running the workload. The performance prediction value is determined using a performance prediction model based on input data, including the workload configuration and the degree of impact of background task configuration on the storage system's performance when the storage system runs the workload.
[0027] In the above solution, the user can set the workload configuration in the first display interface, and the user can intuitively see the performance prediction value of the storage system when running the workload under the background task configuration through the second display interface, thereby improving the user experience.
[0028] In a possible implementation of the second aspect, the background task configuration includes: a first background task configuration and / or a second background task configuration. The first background task configuration includes a configuration that cannot be modified by the user, and the second background task configuration includes a configuration that can be modified by the user.
[0029] In a possible implementation of the second aspect, the first user interface is further used to obtain requirements of the second background task, and the requirements of the second background task are used to determine the second background task configuration.
[0030] In a third aspect, the present application provides a method for training a performance prediction model. The method includes: obtaining a first performance measurement value of a storage system running a historical workload under a background task configuration, and a second performance measurement value of the storage system running the historical workload alone; determining the degree of influence of the background task configuration on the performance of the storage system based on the first performance measurement value and the second performance measurement value, wherein the degree of influence represents the extent of the performance degradation of the storage system when running the workload caused by the background task configuration; inputting input data into a performance prediction model to determine the performance prediction value of the storage system when running the historical workload output by the performance prediction model, wherein the input data includes the degree of influence of the load configuration and the background task configuration on the performance of the storage system; and updating the model parameters of the performance prediction model based on the performance prediction value of the storage system when running the historical workload and the first performance measurement value.
[0031] In a possible implementation manner of the third aspect, the input data further includes operating indicators and / or hardware configuration of the storage system when running the historical workload under the background task configuration.
[0032] In a possible implementation manner of the third aspect, the background task configuration includes: a first background task configuration that cannot be modified by a user of the storage system, and / or a second background task configuration that can be modified by the user.
[0033] In a possible implementation of the third aspect, when the storage system runs multiple historical workloads, the method further includes: determining the performance prediction value of the storage system running multiple historical workloads based on the load configurations of the multiple historical workloads and the performance prediction value of the storage system when running each historical workload; and updating the model parameters of the performance prediction model based on the performance prediction value of the storage system running multiple historical workloads and the first performance measured value.
[0034] In a fourth aspect, the present application provides a storage system performance prediction device, which includes a determination module and a prediction module.
[0035] The determination module is used to determine the load configuration of the workload of the storage system, and to determine the background task configuration when the storage system runs the workload.
[0036] Among them, the prediction module is used to input input data into the performance prediction model to determine the performance prediction value of the storage system when running the workload output by the performance prediction model. The input data includes the load configuration of the workload and the degree of influence of the background task configuration of the storage system when running the workload on the performance of the storage system.
[0037] In a possible implementation of the fourth aspect, when the storage system includes multiple workloads, the prediction module is also used to determine the performance prediction value of the storage system when running multiple workloads based on the load configuration of the multiple workloads and the performance prediction value of the storage system when running each workload.
[0038] In a possible implementation of the fourth aspect, the determining module is specifically configured to: receive a workload configuration from a user interface of the storage system, wherein a user of the storage system can set the workload configuration through the user interface.
[0039] In a possible implementation of the fourth aspect, the background task configuration includes: a first background task configuration and / or a second background task configuration, wherein the first background task configuration includes a configuration that cannot be modified by the user, and the second background task configuration includes a configuration that can be modified by the user.
[0040] In a possible implementation of the fourth aspect, the determination module is specifically configured to determine the first background task configuration according to a load configuration of the workload.
[0041] In one possible implementation of the fourth aspect, the determining module is specifically configured to: receive a requirement for the second background task from a user interface of the storage system; and determine a configuration of the second background task based on the requirement of the second background task. A user of the storage system may set the requirement for the second background task through the user interface.
[0042] In a possible implementation of the fourth aspect, the input data further includes operating indicators and / or hardware configuration of the storage system.
[0043] In a fifth aspect, the present application provides a storage system performance prediction device, which includes a configuration module and a display module.
[0044] The configuration module is used to display a first user interface, and the first user interface is used to set a load configuration of a workload of the storage system;
[0045] Among them, the display module is used to display a second user interface, which is used to display the performance prediction value of the storage system when running the workload. The performance prediction value is determined based on input data using a performance prediction model. The input data includes the load configuration of the workload and the degree of influence of the background task configuration when the storage system runs the workload on the performance of the storage system. The degree of influence represents the extent of the performance degradation of the storage system when running the workload due to the background task configuration.
[0046] In a possible implementation of the fifth aspect, the second user interface is further used to display the performance upper limit and / or performance lower limit of the storage system.
[0047] In a possible implementation of the fifth aspect, the background task configuration includes: a first background task configuration and / or a second background task configuration. The first background task configuration includes a configuration that cannot be modified by the user, and the second background task configuration includes a configuration that can be modified by the user.
[0048] In a possible implementation of the fifth aspect, the first user interface is further used to obtain requirements of the second background task, and the requirements of the second background task are used to determine the second background task configuration.
[0049] In a sixth aspect, the present application provides a device for training a performance prediction model of a storage system, comprising: a determination module and an update module.
[0050] The determination module is configured to obtain a first measured performance value of the storage system running a historical workload under a background task configuration, and a second measured performance value of the storage system running the historical workload alone, and determine the degree of impact of the background task configuration on the performance of the storage system based on the first measured performance value and the second measured performance value;
[0051] Among them, the update module is used to input input data into the performance prediction model, determine the performance prediction value output by the performance prediction model when the storage system runs the historical workload under the background task configuration, and update the model parameters of the performance prediction model according to the first performance prediction value and the actual performance value. The input data includes the degree of influence of the load configuration and the background task configuration.
[0052] In a possible implementation of the sixth aspect, when the storage system runs multiple historical workloads, the update module is also used to: determine the performance prediction value of the storage system running multiple historical workloads based on the load configuration of the multiple historical workloads and the performance prediction value of the storage system when running each historical workload; and update the model parameters of the performance prediction model based on the performance prediction value of the storage system running multiple historical workloads and the first performance measured value.
[0053] In a possible implementation of the sixth aspect, the input data further includes operating indicators and / or hardware configuration of the storage system when running the historical workload under the background task configuration.
[0054] In a possible implementation of the sixth aspect, the background task configuration includes: a first background task configuration that cannot be modified by a user of the storage system, and / or a second background task configuration that can be modified by the user.
[0055] In a seventh aspect, the present application provides a storage device. The storage device includes: a processor and a memory, the processor being configured to execute a computer program stored in the memory to implement the performance prediction method provided in the first aspect or any possible implementation of the first aspect, or the performance prediction method provided in the second aspect or any possible implementation of the second aspect, or the performance prediction method provided in the third aspect or any possible implementation of the third aspect.
[0056] In an eighth aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium includes instructions that, when executed on a computer, cause the computer to execute the performance prediction method provided in the first aspect or any possible implementation of the first aspect, or the performance prediction method provided in the second aspect or any possible implementation of the second aspect, or the performance prediction method provided in the third aspect or any possible implementation of the third aspect.
[0057] In a ninth aspect, the present application provides a computer program product. The computer program product includes program code, and when a computer executes the computer program product, the computer executes the performance prediction method provided in the first aspect or any possible implementation of the first aspect, or the performance prediction method provided in the second aspect or any possible implementation of the second aspect, or the performance prediction method provided in the third aspect or any possible implementation of the third aspect.
[0058] Any of the storage systems, storage devices, computer storage media, or computer program products provided above is used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding schemes in the corresponding methods provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] FIG1 is a schematic diagram of a service level agreement (SLA) signed between an enterprise and a user according to an embodiment of the present application;
[0060] FIG2 is a graph showing changes in the number of read and write operations per second (IOPS) of a storage system provided by an embodiment of the present application;
[0061] 3 is a diagram showing changes in IOPS of services and background tasks before and after background tasks are enabled in a storage system according to an embodiment of the present application;
[0062] FIG4 is a schematic diagram of the structure of a performance prediction system provided in an embodiment of the present application;
[0063] FIG5 is a flow chart of a method for training a performance prediction model provided in an embodiment of the present application;
[0064] FIG6 a is a schematic diagram of updating model parameters of a performance prediction model in a single load operation scenario provided by an embodiment of the present application;
[0065] FIG6 b is a schematic diagram of updating model parameters of a performance prediction model in a multi-load mixed operation scenario provided by an embodiment of the present application;
[0066] FIG7 is a flowchart of a performance prediction method applied to a storage system provided in an embodiment of the present application;
[0067] FIG8 is a schematic diagram of a load configuration and background task configuration for determining a workload provided by an embodiment of the present application;
[0068] FIG9 a is a schematic diagram of a performance prediction process in a single load operation scenario provided by an embodiment of the present application;
[0069] FIG9 b is a schematic diagram of a performance prediction process in a multi-load mixed operation scenario provided by an embodiment of the present application;
[0070] FIG10 is a flowchart of a performance prediction method applied to a terminal device provided in an embodiment of the present application;
[0071] FIG11a is a schematic diagram of a first user interface for setting business requirements provided by an embodiment of the present application;
[0072] FIG11b is a schematic diagram of a second user interface for displaying performance prediction values provided in an embodiment of the present application;
[0073] FIG11c is a schematic diagram of a third user interface for setting QoS of a storage system provided in an embodiment of the present application;
[0074] FIG12 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0075] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0076] A service-level agreement (SLA) is a mutually agreed-upon agreement or contract between a service provider and its users. This SLA includes the agreed-upon quality of service (QoS). QoS can represent the performance of a company's service system when providing services to users. This embodiment of the application uses a storage system as the service system.
[0077] Figure 1 is a schematic diagram of an SLA signed between an enterprise and a user, as provided in an embodiment of the present application. As shown in Figure 1, after obtaining the user's business requirements, the enterprise determines a performance forecast for the storage system when running the service based on the business requirements. Based on this performance forecast, the enterprise then determines an SLA and sends it to the user. Thus, after signing the SLA with the enterprise, the user can deploy their services on the storage system based on the SLA.
[0078] Specifically, the performance prediction value can represent the maximum performance value of the storage system when running a service. This maximum performance value can be included in the SLA. Thus, after signing an SLA with the enterprise, users can set upper and lower limits for services deployed on the storage system based on this maximum performance value.
[0079] Business performance can be expressed through one or more metrics, including read and write operations per second (IOPS), latency, and throughput. Taking IOPS as an example, as shown in Figure 2, the SLA can stipulate the maximum IOPS limit for a user's business when running on the storage system. When deploying a business, users can set upper and lower IOPS limits based on the maximum IOPS limit.
[0080] In related technologies, a performance fitting formula is used to process measured data from a storage system to obtain performance predictions for services running on the storage system. However, this solution requires the performance fitting formula to be determined based on storage system expert experience. This expert experience may not be compatible with other storage systems, resulting in a performance fitting formula that can only be applied to the current storage system and not to other storage systems, resulting in poor versatility.
[0081] In addition, the performance of the storage system when running services will also be affected by the background tasks of the storage system. Figure 3 shows the changes in the IOPS of the services and background tasks before and after the background tasks of the storage system are enabled. In Figure 3, when the background tasks are not enabled, the IOPS of the services is higher. When the background tasks are enabled and in the regulation stage, the IOPS of the services and background tasks are in a fluctuating state. After the background task regulation is completed, the IOPS of the services and background tasks are in a stable state. As can be seen from Figure 3, when the background tasks change from being enabled to being stable, the IOPS of the services has dropped significantly. It can be seen that background tasks have an impact on the performance of the system when running services. However, the solutions of the above-mentioned related technologies do not take into account the impact of background tasks on service performance, and the performance prediction values obtained cannot accurately reflect the service quality QoS of the storage system running the service.
[0082] To this end, an embodiment of the present application provides a performance prediction method for a storage system, which can solve the above problems.
[0083] In the performance prediction method provided in an embodiment of the present application, the load configuration of the workload of the storage system is first determined, as well as the background task configuration of the storage system when the storage system runs the workload. Then, a performance prediction model is used to determine the performance prediction value of the storage system when running the workload based on the load configuration and the background task configuration. The performance prediction model is pre-trained based on training samples of one or more storage systems. A background task configuration may include configuration parameters of one or more background tasks, and the configuration parameters of the background task may include the status and / or concurrency of the background task.
[0084] This solution uses a pre-trained performance prediction model to predict the performance of a storage system under workload conditions. This solution considers the impact of the storage system's background task configuration on the storage system's performance under workload conditions. This background task configuration is used as an input to the performance prediction model. This improves the accuracy of performance predictions, ensuring that the predicted performance value more accurately reflects the service quality of the storage system under workload conditions. Furthermore, this performance prediction model is pre-trained based on training samples from one or more storage systems, making it more universal. This addresses the limited universality of the aforementioned related technologies, which rely on performance fitting formulas determined by expert experience.
[0085] Figure 4 is a schematic diagram of the structure of a performance prediction system 400 provided in an embodiment of the present application. As shown in Figure 4, the performance prediction system 400 includes a terminal device 410 and a storage system 420. Specifically, the terminal device 410 and the storage system 420 cooperate with each other to implement the performance prediction method of the embodiment of the present application.
[0086] In FIG. 4 , a terminal device 410 may include a configuration module 411 and a presentation module 412 .
[0087] The configuration module 411 may be used to send the service requirements of the user for setting the user service to the storage system 420 .
[0088] The display module 412 may be configured to receive a performance prediction value of the storage system 420 when the storage system runs a workload of a user's business based on business demand prediction, and display the performance prediction value to the user.
[0089] In Figure 4, the storage system 420 may include: a determination module 421, a prediction module 422, and an update module 423. The storage system 420 may be specifically used to execute the performance prediction method and the model training method.
[0090] When executing the performance prediction method, the determination module 421 can be used to receive the business requirements sent by the configuration module 411 in the terminal device 410, and determine the load configuration of the user business workload and the background task configuration when the storage system runs the workload based on the business requirements; the prediction module 422 can be used to determine the performance prediction value of the storage system when running the workload based on the load configuration and background task configuration of the workload, using a pre-trained performance prediction model, and send the performance prediction value to the terminal device.
[0091] When executing the model training method, the determination module 421 can be used to determine the training sample based on the business requirements sent by the configuration module 411 in the terminal device 410, and / or the access records recorded by the storage system 420; the update module 423 can be used to update the model parameters of the performance prediction model based on the training sample, and send the updated performance prediction model to the prediction module 422.
[0092] It should be noted that the storage system 420 may include one or more storage devices, and the determination module 421, prediction module 422 and update module 423 can be deployed in one storage device or in a distributed manner in multiple storage devices. The embodiments of the present application do not impose specific restrictions on this.
[0093] The following first introduces the training method of the above-mentioned performance prediction model.
[0094] Figure 5 is a flow chart of a method for training a performance prediction model provided by an embodiment of the present application. The training method can be executed by the determination module 421 and the update module 423 in the storage system 420 shown in Figure 4. As shown in Figure 5, the method may include S501-S503.
[0095] In S501 , the determination module 421 determines the degree of impact of the background task configuration on the performance of the storage system.
[0096] In this embodiment, the determination module 421 may first collect training samples for the performance prediction model. The training samples may include the load configuration of the historical workload, the background task configuration when the storage system 420 runs the historical workload, a first measured performance value when the historical workload is run under the background task configuration, and a second measured performance value when the storage system 420 runs the historical workload alone. Specifically, the storage system 420 running the historical workload alone means that the storage system 420 runs the historical workload without enabling any background tasks.
[0097] Specifically, the storage system 420 may include an access record module for recording access records of various historical workloads run by the storage system 420 , as well as background task configurations and first performance measured values when running each historical workload.
[0098] Determination module 421 can analyze and compile historical workload access records to obtain a load profile for the historical workload. Workload access records can include information such as the workload's I / O size, I / O type, offset address, and timestamp. The I / O type can be read or write. The following describes the analysis and compiling process of workload access records by determination module 421, using a workload configuration including I / O size, read-write ratio, randomness, concurrency, run time, and mixed ratio as an example.
[0099] Regarding the IO size of the workload, the determination module 421 may determine the IO size by calculating an average value of the IO sizes in multiple access records of the workload.
[0100] Regarding the read-write ratio of the workload, the determination module 421 may determine the number of reads and the number of writes of the workload according to the IO types in multiple access records of the workload, and then determine the ratio by calculating the number of reads to the number of writes.
[0101] Regarding the randomness of a workload, determination module 421 can determine the number of random accesses for the workload based on the offset addresses and timestamps in multiple access records of the workload, and then calculate the ratio of the random access number to the total access number of the workload to obtain the randomness. The total access number of the workload is the number of access records of the workload. Specifically, each access record corresponds to an IO of the workload, and random access means that the offset addresses of two adjacent IO accesses are also adjacent. Two adjacent IOs refer to two IOs accessed consecutively by the workload, i.e., the timestamps of the two IOs are temporally continuous, i.e., there are no other IOs of the workload between the timestamps of the two IO accesses; and adjacent offset addresses of two IO accesses refer to the offset addresses of the two IO accesses being continuous. Determining the number of random accesses for the workload based on the offset addresses and timestamps in multiple access records of the workload can include: first initializing the number of random accesses for the workload to 0; then determining a target IO adjacent to the current IO; if the offset address in the access record corresponding to the current IO is also adjacent to the offset address in the access record corresponding to the target IO, the number of random accesses for the workload is increased by 1, and this process is repeated until all access records of the workload are analyzed.
[0102] Regarding the running time, the determination module 421 may determine the running time of the workload according to the timestamps in each access record of the workload. The running time may be, for example, the difference between the maximum timestamp and the minimum timestamp.
[0103] Regarding the concurrency of the workload, the determination module 421 may count the number of access records within the target period according to the timestamps in the multiple access records of the workload, and use the number as the concurrency of the workload.
[0104] In this embodiment, after collecting training samples, determination module 421 can use the background task impact model to analyze the performance impact of the background task configurations in the training samples to determine the degree of impact of the background task configurations on the performance of the storage system. Using the background task impact model, the performance of the storage system running the workload can be more accurately determined, improving the accuracy of workload performance prediction.
[0105] Specifically, during an offline phase, determination module 421 can independently run the historical workload on storage system 420 based on the load configuration of the historical workload obtained above, and obtain a second measured performance value. Then, based on the obtained first and second measured performance values, a pre-established background task impact model based on regression analysis is used to determine the degree of impact of the corresponding background task configuration on the performance of the storage system. The background task impact model based on regression analysis is used to model the performance impact of background tasks on the workload running on the storage system, and can specifically employ a regression model such as a random forest model.
[0106] The above-mentioned background task impact model based on regression analysis can be pre-established according to the following formula (1). Since the impact of background task configuration on the performance of storage system running workload is variable, when determining the training sample of the background task impact model, N experiments can be conducted on different historical workloads based on the same background task configuration to obtain the first performance value P corresponding to the background task configuration of the N experiments. i,o And close the second performance value P of the background task configuration i,c , and then calculate the average value to get the impact degree corresponding to the background task configuration.
[0107] In formula (1), N represents the number of experiments, and i represents the i-th experiment. There is no theoretical limit to the value of N, and a higher value is better, because the more experiments, the more closely the results match the mean of the actual performance distribution. The impact of background task configuration on the IOPS, latency, and throughput of the storage system's workload can be calculated using formula (1).
[0108] In a specific implementation, a storage system may start multiple background tasks when running a workload. The determination module can train a background task impact model for each background task, thereby quantifying the impact of the background task on performance in a fine-grained manner. In addition, the configuration parameters of a background task may be different when running different workloads. When training the background task impact model, the determination module can quantify the performance impact of the background task under each configuration parameter according to formula (1) to obtain the degree of impact of the background task under different configuration parameters. Taking GC as an example, GC may include different concurrency numbers when it is turned on. Therefore, the determination module can collect the first performance measured value corresponding to the different concurrency numbers of GC, and then train the background task impact model corresponding to GC according to formula 1 based on the first performance measured value corresponding to the different concurrency numbers of GC and the second performance measured value corresponding to GC.
[0109] In other embodiments, the background task impact model can also be a machine learning model that characterizes the mapping relationship between the load configuration and the background task configuration, and the degree of influence of the background task configuration. That is, the input data of the background task impact model includes the load configuration and the background task configuration of the workload, and the output data includes the degree of influence of the background task configuration. Specifically, the training samples of the background task impact model can be obtained by running experiments on the storage system. Specifically, the training samples of the background task impact model can be obtained by running different workloads under different background task configurations on the storage system, and collecting the first performance measured values of the storage system when running the workload under different background task configurations.
[0110] In this embodiment, the training sample may further include the hardware configuration and operating indicators of the storage system 420 when running historical workloads. The hardware configuration and operating indicators may be recorded by the access recording module of the storage system 420 when the workload is running.
[0111] The access records, background task configurations, hardware configurations, operating indicators, and performance measurements recorded by the access record module may be stored in a memory in any one or more storage devices in the storage system 420. The memory may include a device with data storage function such as a hard disk and a memory.
[0112] In S502 , the updating module 423 determines input data, inputs the input data into the performance prediction model, and outputs a performance prediction value when the storage system runs a historical workload.
[0113] In this step, the input data may include the load configuration of the historical workload and the impact degree corresponding to the background task configuration. As shown in FIG6a , the update module 423 may input the input data corresponding to the historical workload into the performance prediction model to obtain the performance prediction value output by the performance prediction model.
[0114] When running multiple workloads, update module 423 can input the input data corresponding to each workload into the performance prediction model, obtain the performance prediction value of storage system 420 when running each workload, and then determine the performance prediction value of storage system 420 when running multiple workloads based on the performance prediction value corresponding to each workload. As shown in Figure 6b, when the multiple workloads are load 1, load 2, and load 3, the performance prediction value corresponding to load 1 is output data 1, the performance prediction value corresponding to load 2 is output data 2, and the performance prediction value corresponding to load 3 is output data 3.
[0115] In other embodiments, the input data of the performance prediction model may also include load configuration of historical workload, background task configuration, degree of influence of background task configuration, operating indicators, hardware configuration, etc., which can further improve the prediction accuracy of the model.
[0116] In other embodiments, the input data of the performance prediction model may also only include the load configuration of historical workloads.
[0117] In S503 , the updating module 423 updates the model parameters of the performance prediction model according to the performance prediction value and the first performance measured value when the storage system runs the historical workload under the background task configuration.
[0118] In this step, when the storage system 420 runs a historical workload, as shown in Figure 6a, the update module 423 can calculate the loss value between the performance prediction value and the first performance measured value based on the loss function of the performance prediction model, and then update the model parameters of the performance prediction model according to the gradient descent method based on the loss value. In this embodiment, the performance prediction model can adopt one of the regression models such as the random forest model. The model parameters of the random forest model may include the number of sub-models n_estimators, the name of the loss function criterion, the maximum depth of the tree max_depth, the minimum number of samples for leaf node splitting min_samples_split, the minimum number of samples for leaf nodes min_samples_leaf, the minimum weight for leaf node splitting min_weight_fraction_leaf, the maximum number of features during splitting max_features, the maximum number of leaf nodes of the tree max_leaf_nodes, the maximum number of samples of the tree max_samples, etc.
[0119] When the storage system 420 runs multiple historical workloads, the update module 423 can determine the performance prediction value of the storage system when running multiple historical workloads based on the performance prediction value of each historical workload, then calculate the loss value between the performance prediction value of the storage system when running multiple historical workloads and the first performance measured value of the storage system when running multiple historical workloads, and update the model parameters using the gradient descent method based on the loss value. The specific process of determining the performance prediction value of the storage system when running multiple historical workloads can be referred to the description of the performance prediction method embodiment shown in Figure 7, and will not be repeated here.
[0120] In other embodiments, the input data of the performance prediction model may also include workload configuration and storage system operation status.
[0121] The following introduces a performance prediction method based on the performance prediction model determined by the method shown in FIG5 .
[0122] Figure 7 is a flow chart of a performance prediction method provided by an embodiment of the present application. The performance prediction method can be applied to the storage system 420 and executed by the determination module 421 and the prediction module 422 in the storage system 420. As shown in Figure 7, the method may include the following steps S701 and S702.
[0123] In S701 , the determination module 421 determines the load configuration of the workload of the storage system, and determines the background task configuration of the storage system when the storage system runs the workload.
[0124] In this step, as shown in FIG8 , the determination module 421 can receive the business requirements of the user business set by the user from the configuration module 411 in the terminal device 410, determine the load configuration and background task configuration of the workload based on the business requirements, and then send the load configuration and background task configuration to the prediction module 422. Specifically, each business can include one or more workloads. As shown in FIG8 , business A can include one workload (load 1), and business B can include multiple workloads (load 2 and load 3). Business requirements can include workload load configuration, background task requirements, hardware configuration, etc.
[0125] In addition, the business requirements may also include access records of one or more workloads of the user business. When the business requirements include access records of one or more workloads, the determination module 421 can perform analysis and statistics based on the access records to determine the load configuration of each workload. The business requirements may also directly include the load configuration of one or more workloads of the user business, and the determination module 421 can extract the load configuration of each workload from the business requirements. The specific process of the analysis and statistics performed by the determination module 421 can be found in the specific introduction of the model training method shown in Figure 5 above, and will not be repeated here.
[0126] The background tasks of the storage system may include a first background task and a second background task. Therefore, the background task configuration may include the first background task configuration and the second background task configuration. The first background task refers to a background task whose configuration cannot be modified by the user, and the second background task refers to a background task whose configuration can be modified by the user. The background task requirements are used to determine the second background task configuration. The first background task may include, for example, garbage collection (GC), and the second background task may include, for example, remote replication tasks and / or data deduplication tasks.
[0127] For the first background task configuration, the determination module 421 can determine it based on the workload's load configuration. Taking garbage collection (GC) as an example, the GC configuration can include the GC's status and concurrency, with the GC's status including on or off. Specifically, the determination module 421 can determine the GC's status based on the runtime in the workload's load configuration. Specifically, when the runtime is greater than a pre-set threshold, the GC's status is determined to be on, otherwise it is off. When the GC's status is determined to be on, the determination module 421 can determine the capacity of the storage system occupied by IO based on the workload's load configuration, and then determine the GC's concurrency based on a predetermined relationship between the capacity of the storage system occupied by IO and the GC's concurrency. The process of determining the capacity of the storage system occupied by IO can include: determining the IO rate corresponding to the concurrency in the workload's load configuration based on the relationship between the IO concurrency and the IO rate; determining the IO's write capacity per unit time based on the IO rate and the product of the IO size and the read-write ratio in the workload's load configuration; and finally, obtaining the capacity of the storage system occupied by IO based on the write capacity per unit time and the runtime.
[0128] For the second background task configuration, the determination module 421 can be determined based on the background task requirements. Taking the remote replication task and the deduplication task as an example, the background task requirements may include the rate of the remote replication task and the rate of the deduplication task. The determination module 421 can determine the concurrency of the remote replication task and the concurrency of the deduplication task according to the rate of the remote replication task and the rate of the deduplication task. Specifically, the correspondence between the rate and the concurrency of the remote replication task and the correspondence between the rate and the concurrency of the deduplication task can be pre-stored in the determination module 421. In this way, the determination module 421 can determine the concurrency of the remote replication task and the concurrency of the deduplication task according to the rate of the remote replication task and the rate of the deduplication task set by the user, respectively, according to their respective correspondences.
[0129] In addition, in addition to determining the load configuration and background task configuration of the workload, the determination module 421 can also determine the hardware configuration of the storage system and the operating indicators of the storage system when running the workload, and send the hardware configuration and operating indicators to the prediction module 422.
[0130] The determination module 421 may obtain the hardware configuration of the storage system set by the user from the configuration module 411. The CPU configuration may include the CPU type and the number of CPUs, and the hard disk configuration may include the hard disk type and the number of hard disks.
[0131] As shown in FIG8 , the determination module 421 may run the workload on the target device in the storage system 420 according to the workload configuration and background task configuration, and obtain from the target device the operating indicators of the target device recorded by the target device during the workload operation. The operating indicators may include CPU utilization, hard disk utilization, network utilization, etc.
[0132] In S702 , the prediction module 422 determines input data, inputs the input data into a performance prediction model, and determines a performance prediction value output by the performance prediction model when the storage system runs the workload under the background task configuration.
[0133] In this step, the prediction module 422 can receive the load configuration and background task configuration of the workload from the determination module 421, determine the degree of influence of the background task configuration on the performance of the storage system based on the background task configuration, and then use the load configuration of the workload and the degree of influence of the background task configuration as input data of the performance prediction model, input the input data into the performance prediction model, and determine the performance prediction value of the storage system when running the workload output by the performance prediction model.
[0134] Wherein, determining the degree of influence of the background task configuration according to the background task configuration includes that it can be matched from a pre-obtained degree of influence set according to the background task configuration. The degree of influence set includes the degree of influence of different background task configurations obtained in the embodiment shown in Figure 4 above on the performance of the storage system, and a background task configuration may include configuration parameters of one or more background tasks. The configuration parameters of the background task may include the status and / or concurrency of the background task. Wherein, the input data may also include the background task configuration, and the hardware configuration and operating indicators of the storage system 420. Using the hardware configuration and operating indicators as the input of the model can further improve the accuracy of the performance prediction. Especially in the multi-load mixed operation scenario, the operating indicators can indirectly reflect the characteristics of the multi-load mixed operation, thereby supporting the performance prediction of the multi-load mixed operation scenario.
[0135] Specifically, taking a workload (Load 1) as an example, as shown in Figure 9a, prediction module 422 can input input data 1 corresponding to Load 1 into the performance prediction model to obtain model output data 1. Output data 1 is the performance prediction value 1 of the storage system when running Load 1. The input data 1 corresponding to Load 1 includes Load 1's load configuration, background task configuration, operating indicators, and hardware configuration.
[0136] Specifically, taking multiple workloads (load 1, load 2, and load 3) as an example, as shown in FIG9b , prediction module 422 may first input input data 1 corresponding to load 1, input data 2 corresponding to load 2, and input data 3 corresponding to load 3 into the performance prediction model to obtain output data 1, output data 2, and output data 3 of the model. Output data 1 is the performance prediction value 1 when the storage system runs load 1, output data 2 is the performance prediction value 2 when the storage system runs load 2, and output data 3 is the performance prediction value 3 when the storage system runs load 3.
[0137] Then, prediction module 422 determines the performance prediction value when the storage system simultaneously runs load 1, load 2, and load 3 based on the load configuration of load 1, the load configuration of load 2, and the load configuration of load 3, as well as the aforementioned performance prediction value 1, performance prediction value 2, and performance prediction value 3. The storage system simultaneously running load 1, load 2, and load 3 is a multi-load mixed operation scenario of the storage system.
[0138] Taking the performance prediction values including IOPS, latency and throughput as an example, the prediction module 422 can calculate the weighted sum based on the mixed proportion of load 1 and the latency in the performance prediction value 1, the mixed proportion of load 2 and the latency in the performance prediction value 2, and the mixed proportion of load 3 and the latency in the performance prediction value 3 to obtain the latency when the storage system runs load 1, load 2 and load 3 simultaneously.
[0139] After determining the latency when the storage system runs loads 1, 2, and 3 simultaneously, the prediction module 422 can determine the total concurrency of loads 1 minus load 3, and use the quotient of the total concurrency and the latency as the IOPS when the storage system runs loads 1, 2, and 3 simultaneously. The total concurrency of loads 1 minus load 3 is the sum of the concurrency of loads 1, 2, and 3.
[0140] After determining the IOPS of the storage system when running load 1, load 2, and load 3, the prediction module 422 can determine the average IO size of load 1-load 3, multiply the average IO size by the IOPS, and obtain the throughput of the storage system when running load 1, load 2, and load 3.
[0141] For example, if the input data includes workload configuration, background task configuration, hardware configuration, and operating indicators, prediction module 422 can process each input data into data of the corresponding data type before inputting it into the performance prediction model. Furthermore, the performance prediction value output by the performance prediction model can be processed into data of the corresponding data type before being sent to presentation module 412. The data types corresponding to the input data and the performance prediction values are shown in Table 1.
[0142] Table 1 Data attribute table
[0143] In this embodiment, after obtaining the performance prediction value, the prediction module 422 may send the performance prediction value to the display module 412 in the terminal device 410 , and the display module 412 may display the performance prediction value to the user.
[0144] In this embodiment, prediction module 422 can also use the performance prediction model to determine performance prediction values for the storage system when running a workload under other background task configurations that can be provided by storage system 420, thereby providing a multi-angle performance display to the user, thereby providing the user with different background task configurations for selection. Similarly, the performance prediction values corresponding to other background task configurations can be displayed to the user by presentation module 412.
[0145] In this embodiment, the prediction module 422 can also provide the user with an interface for setting the upper and / or lower limits of the QoS of the storage system through the display module 412, which is used to set the upper and lower limits of IOPS, the upper and lower limits of latency, and the upper and lower limits of throughput.
[0146] The method embodiment described in FIG7 takes into account the impact of the storage system's background tasks on the performance of the storage system's workload when predicting the performance of the storage system. This improves the accuracy of the performance prediction, allowing users to more accurately set the storage system's Quality of Service (QoS). Furthermore, the method embodiment shown in FIG7 supports performance prediction for both single-load and mixed-load scenarios, making it widely applicable.
[0147] Based on the performance prediction method shown in FIG7 , an embodiment of the present application also provides another performance prediction method.
[0148] Figure 10 is a flow chart of another performance prediction method provided by an embodiment of the present application. The method can be applied to a terminal device 410 and executed by a configuration module 411 and a presentation module 412. As shown in Figure 10, the method may include S1001 and S1002.
[0149] In S1001 , the configuration module 411 displays a first user interface, so that a user can set a business requirement of a workload through the first user interface.
[0150] The configuration module 411 may display a first user interface as shown in FIG11a to the user, so that the user can set business requirements for one or more businesses through the first user interface. Specifically, the business requirements may include workload configuration, background task requirements, and hardware configuration.
[0151] In the first user interface shown in FIG11a, the user can configure the load configuration of one or more workloads and can add the workload configuration by clicking the add icon in the first user interface. Add more load configurations. The workload configuration can include one or more of the workload's read and write IO size, read-write ratio, randomness, concurrency, and mix ratio.
[0152] In the first user interface shown in FIG11a, the user can also configure one or more background task requirements, and can add one or more background task requirements by clicking the add icon in the first user interface. Add more background task requirements.
[0153] In the first user interface shown in FIG11a, the user can also configure one or more hardware configurations of the storage system 420. The hardware may include but is not limited to a central processing unit (CPU) and a hard disk, and can also be configured by clicking the Add icon in the first user interface. Add more hardware configurations.
[0154] In S1002, the display module 412 displays a second user interface to display the performance prediction value of the storage system when running the workload predicted by the performance prediction model. The training process of the performance prediction model can refer to the description of the embodiment shown in FIG5 , and will not be repeated here. The process of the performance prediction model predicting the performance prediction value can refer to the description of S701 and S702 in the embodiment shown in FIG5 , and will not be repeated here. The second user interface may include the interface shown in FIG11b .
[0155] The second user interface shown in FIG11b includes the performance values of the storage system when running the workload under the current background task configuration, including IOPS, latency, and throughput. The current background task configuration is determined based on the user's business needs.
[0156] In the second user interface shown in Figure 11b, the IOPS, latency, and throughput corresponding to other background task configurations can also be displayed to provide users with the performance of the storage system when running the workload set by the user under other background task configurations that the storage system can provide.
[0157] In this embodiment, the display module 412 may also display a third user interface, as shown in FIG11c , which is used by the user to set the upper and / or lower limits of the QoS of the storage system. As shown in FIG11c , based on the IOPS, latency, and throughput displayed in the second user interface, the user may set the upper and lower limits of IOPS, the upper and lower limits of latency, and the upper and lower limits of throughput in the third user interface.
[0158] The performance prediction method of the embodiment of the present application supports performance prediction for single-load operation and multi-load mixed operation, and takes into account the performance impact of background tasks on user business loads; it uses AI methods to train performance prediction models, and uses data to deduce the relationship between load configuration, background task configuration, and performance, without relying on expert experience. At the same time, it provides the function of configuring workloads, background tasks, and displaying multi-dimensional performance values of the system on the user interface to help users make settings based on their own business characteristics. In addition, it provides the function of setting SLA QoS values on the user interface. Users can configure more appropriate and accurate SLA QoS based on the multi-dimensional performance values displayed on the user interface, which can reduce the SLA default rate.
[0159] Experiments were conducted for both single and mixed workloads, within the range of values shown for each load configuration. The background tasks in the experiments were configured with GC enabled, deduplication disabled, and remote replication disabled.
[0160] Table 3 Workload configuration table
[0161] The experimental results, shown in Tables 4-6, demonstrate that the performance evaluation technique of the present invention has a relatively low average prediction error and good prediction performance. It should be noted that for ease of presentation, the values in Tables 4-6 are rounded to two decimal places; the actual data values do not have only two decimal places. The prediction error in Tables 4-6 is the percentage of the quotient of the difference between the measured value and the predicted value divided by the measured value.
[0162] Table 4 Comparison of measured and predicted performance of a single workload
[0163] Table 5 Comparison of the measured and predicted performance of the two workloads during mixed runtime
[0164] Table 6 Comparison of measured and predicted performance of three workload hybrid runtimes
[0165] Based on the model training method shown in FIG5 , an embodiment of the present application further provides a training device. The training method can be applied to the storage system 420 shown in FIG4 , and is used to perform each step in the training method shown in FIG5 . The training device may include the determination module 421 and the update module 423 shown in FIG4 .
[0166] Based on the performance prediction method shown in FIG7 , an embodiment of the present application further provides a performance prediction device. The performance prediction device can be applied to the storage system 420 shown in FIG4 , and is used to perform each step in the performance prediction method shown in FIG7 . The performance prediction device can include the determination module 421 and the prediction module 422 shown in FIG4 .
[0167] Based on the performance prediction method shown in FIG10 , an embodiment of the present application further provides a performance prediction device. The performance prediction device can be applied to the terminal device 410 shown in FIG4 , and is used to perform each step of the performance prediction method shown in FIG10 . The performance prediction device can include the configuration module 411 and the presentation module 412 shown in FIG4 .
[0168] It should be noted that the aforementioned performance prediction device and training device, when executing the performance prediction method and training method, only uses the division of the aforementioned functional modules as an example. In actual applications, the functions of the aforementioned devices can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the performance prediction device and training device provided in the above embodiments are based on the same concept as the aforementioned performance prediction method embodiment and training method embodiment. The specific implementation process is detailed in the method embodiment and the invention content, and will not be repeated here.
[0169] In other embodiments, the output of the performance prediction model in the methods shown in Figures 5, 7, and 10 may also include performance prediction values for the storage system when running background tasks. Accordingly, the training samples of the performance prediction model may also include actual measured performance values for the storage system when running background tasks. This solution can improve the accuracy of background task performance prediction by considering the impact of the storage system's workload on the performance of background tasks.
[0170] In other embodiments, the methods shown in Figures 5, 7, and 10 can also be applied to network system scenarios to predict the performance of the network system when running user workloads and to train performance prediction models for the network system. The load configuration of the network system's workload can include the configuration of data packets transmitted by the network system, such as the size and type of the data packets. Background tasks of the network system can include data enhancement tasks, data consistency verification tasks, and the like.
[0171] 12 is a schematic diagram of the hardware structure of a computing device 1200 provided in an embodiment of the present application. The computing device 1200 may be a storage device in the storage system 420 or the terminal device 410 described above.
[0172] 12 , the computing device 1200 may include a processor 1201, a memory 1202, a communication interface 1203, and a bus 1204. The processor 1201, the memory 1202, and the communication interface 1203 are connected to each other via the bus 1204. The processor 1201, the memory 1202, and the communication interface 1203 may also be connected using other connection methods besides the bus 1204.
[0173] Processor 1201 may be a general-purpose processor that performs specific steps and / or operations by reading and executing content stored in a memory (e.g., memory 1202). For example, a general-purpose processor may be a central processing unit (CPU). Processor 1201 may include at least one circuit to perform all or part of the steps of the method shown in FIG. 5 , FIG. 7 , or FIG. 10 .
[0174] The memory 1202 may be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, optical storage, hard disk, etc.
[0175] When computing device 1200 is a storage device in storage system 420, memory 1202 can be specifically used to store the first program code. When processor 1201 executes this first program code, it implements the functions of determination module 421, prediction module 422, and update module 423 in FIG4 . Furthermore, it can be used to store access records of various workloads in storage system 420, background operation configurations, operation indicators, hardware configurations and measured performance values, as well as performance prediction models and their model parameters.
[0176] When the computing device 1200 is the terminal device 410, the memory 1202 can be used to store the second-dimensional program code. When the processor 1201 executes the second program code, the functions of the configuration module 411 and the display module 412 in FIG. 4 are implemented.
[0177] Communication interface 1203 includes input / output (I / O) interfaces, physical interfaces, and logical interfaces, which are used to interconnect components within computing device 1200, as well as interfaces for interconnecting computing device 1200 with other devices (e.g., other computing devices or user equipment). Physical interfaces can include Ethernet interfaces, fiber optic interfaces, ATM interfaces, and the like.
[0178] When the computing device 1200 is a storage device in the storage system 420 , the business requirements set by the user can be obtained from the terminal device 410 through the communication interface 1203 , and the performance prediction value can be sent to the terminal device 410 through the communication interface.
[0179] When the computing device 1200 is a terminal device 410 , the business requirements set by the user can be sent to the storage device in the storage system 420 through the communication interface 1203 , and the performance prediction value sent by the storage device can be received through the communication interface.
[0180] The bus 1204 may be any type of communication bus for interconnecting the processor 1201 , the memory 1202 , and the communication interface 1203 , such as a system bus.
[0181] The above-mentioned devices can be provided on separate chips, or at least partially or entirely on the same chip. Whether to provide each device independently on different chips or to integrate them on one or more chips often depends on the product design requirements. The embodiments of this application do not limit the specific implementation of the above-mentioned devices.
[0182] The computing device 1200 shown in FIG12 is merely exemplary. During implementation, the computing device 1200 may further include other components, which are not listed here one by one.
[0183] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).
[0184] It is understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not intended to limit the scope of the embodiments of the present application. It should be understood that in the embodiments of the present application, the order of the sequence numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0185] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this application should be included in the scope of protection of this application.
Claims
1. A method for training a performance prediction model of a storage system, characterized in that: The method comprises: Determining the degree of impact of background task configuration of the storage system on the performance of the storage system; Inputting input data into a performance prediction model, and outputting a performance prediction value of the storage system when running a historical workload under the background task configuration, the input data including the load configuration of the historical workload and the impact degree; The model parameters of the performance prediction model are updated according to the first measured performance value and the performance prediction value when the storage system runs the historical workload under the background task configuration.
2. The method according to claim 1, characterized in that Determining the degree of influence of the background task configuration of the storage system on the performance of the storage system includes: The impact degree is determined according to the first performance measured value and the second performance measured value when the storage system runs the historical workload without starting the background task.
3. The method according to claim 1 or 2, characterized in that: The input data also includes operating indicators and / or hardware configuration of the storage system when running the historical workload under the background task configuration.
4. The method according to any one of claims 1 to 3, characterized in that: The background task configuration includes: a first background task configuration that cannot be modified by a user of the storage system, and / or a second background task configuration that can be modified by the user.
5. A method for predicting performance of a storage system, characterized in that: The method comprises: Determine the workload profile of the storage system, and, Determining a background task configuration of the storage system when the storage system runs the workload; Input data is input into a performance prediction model to output a performance prediction value of the storage system when running the workload under the background task configuration, wherein the input data includes the influence of the load configuration and the background task configuration on the performance of the storage system.
6. The method according to claim 5, characterized in that The method further comprises: The impact degree is determined according to a first measured performance value of the storage system when running the historical workload under the background task configuration and a second measured performance value of the storage system when running the historical workload without starting the background task.
7. The method according to claim 5 or 6, characterized in that: The load configuration of determining the workload of the storage system includes: displaying a first user interface of the storage system, whereby a user of the storage system sets a load configuration of the workload through the first user interface; A load configuration for the workload is received from the first user interface.
8. The method according to any one of claims 5 to 7, characterized in that: The background task configuration includes a first background task configuration that cannot be modified by a user of the storage system, and / or a second background task configuration that can be modified by the user, The determining of the background task configuration of the storage system when the storage system runs the workload includes: determining the first background task configuration according to the load configuration of the workload, and / or, The second background task configuration is determined according to a requirement of the second background task, wherein the user sets the requirement of the second background task through a first user interface of the storage system.
9. The method according to any one of claims 5 to 8, characterized in that: The input data also includes operating indicators and / or hardware configuration of the storage system.
10. The method according to any one of claims 5 to 9, characterized in that: The method comprises: A second user interface is displayed, where the second user interface is used to display a performance prediction value when the storage system runs the workload under the background task configuration.
11. A training device for a performance prediction model of a storage system, characterized in that: The device comprises: A determination module, used to determine the degree of influence of the background task configuration of the storage system on the performance of the storage system; An update module is used to input input data into a performance prediction model, output a performance prediction value when the storage system runs a historical workload under the background task configuration, and update the model parameters of the performance prediction model according to a first measured performance value and the performance prediction value when the storage system runs the historical workload under the background task configuration, wherein the input data includes the load configuration of the historical workload and the degree of influence of the background task configuration on the performance of the storage system.
12. A performance prediction device for a storage system, characterized in that: The device comprises: A determination module, configured to determine a load configuration of a workload of a storage system, and to determine a background task configuration of the storage system when the storage system runs the workload; A prediction module is used to input input data into a performance prediction model to determine the performance prediction value of the storage system when running the workload output by the performance prediction model, wherein the input data includes the degree of influence of the load configuration and the background task configuration on the performance of the storage system.
13. A computing device, characterized in that: The computing device comprises: a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the method according to any one of claims 1 to 4, or to implement the method according to any one of claims 5 to 10.
14. A computer-readable storage medium, characterized in that: The method comprises instructions, which, when executed on a computer, enable the computer to execute the method according to any one of claims 1 to 4, or to execute the method according to any one of claims 5 to 10.
Citation Information
Patent Citations
Automated data forecasting using machine learning
US20220147895A1
Modeler for predicting storage metrics
US8620921B1
Evaluating proposed storage solutions
US9122739B1