An Automatic Parameter Tuning Method and System for a Distributed Storage System Based on Reinforcement Learning

By filtering the key parameters of the distributed storage system based on reinforcement learning, and combining GA and DDPG models for tuning, the problem of difficult parameter tuning in the existing technology is solved, fast and accurate system parameter adjustment is achieved, and the performance of the distributed storage system is improved.

CN116088761BActive Publication Date: 2025-07-29HUAZHONG UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310059553.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-18
Publication Date
2025-07-29
Estimated Expiration
2043-01-18

AI Technical Summary

Technical Problem

The prior art cannot reasonably and effectively optimize the parameters of distributed storage system, which leads to difficulty in screening parameters and the tuning model is too simple, making it difficult to recommend a reasonable optimal configuration in a short time.

Method used

The automatic parameter adjustment method of distributed storage system based on reinforcement learning is adopted. The system parameters with a greater impact are selected by statistically screening the performance indicators in the first sample pool, combining genetic algorithms and deep deterministic strategy gradient (DDPG) model for tuning, and pre-training is used to avoid cold start time, combining the advantages of GA and DDPG models, quickly training and achieving good tuning results.

Benefits of technology

It effectively reduces the complexity of optimization problems, improves the feasibility and accuracy of parameter adjustment, realizes rapid and accurate tuning of distributed storage system parameters, and improves system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116088761B_ABST
    Figure CN116088761B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic parameter tuning method and system for a distributed storage system based on reinforcement learning, belonging to the technical field of distributed storage. First, by statistically analyzing the performance indicators of the distributed storage system in the first sample pool, a plurality of system parameters that have a greater impact on the performance of the distributed storage system are screened out from the first sample pool as a set of tuning parameters for the distributed storage system, so as to reduce the complexity of the optimization problem and ensure the feasibility of subsequent parameter tuning based on reinforcement learning. Then, on the basis of parameter screening, the DDPG model of reinforcement learning is used for further tuning work. And in this process, considering the problem of the long cold start time of the DDPG model in the early stage, a genetic algorithm is adopted for the sample collection work in the early stage, and the collected samples are input to the DDPG model for pre-training, so as to avoid the problem of a large amount of time-consuming cold start in the early stage, and thus can reasonably and effectively tune the parameters of the distributed storage system accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of distributed storage, and more specifically, relates to an automatic parameter tuning method and system for a distributed storage system based on reinforcement learning. Background Technique

[0002] With the rapid development of the information age, a large amount of data needs to be stored. Therefore, as a reliable solution, the distributed file storage system has also developed rapidly, and at the same time, the performance requirements for it are increasing day by day. Parameter tuning, as an effective performance tuning method, has also received extensive attention.

[0003] In Chinese invention patent CN114564460A, a parameter tuning method for a distributed storage system is disclosed. This method first analyzes the read and write performance of the system using a log file to determine the business application model; according to different business application models, select the corresponding parameters to be adjusted; randomly assign values to these parameters and test the performance, and select the optimal parameter configuration according to the test results. In Chinese invention patent CN114415945A, a parameter tuning method for a distributed storage system is disclosed. This method first establishes independent Bayesian optimization models for multiple optimization targets, then performs parameter synthesis, converts the optimized results into prior knowledge to initialize the initial population of the multi-objective evolutionary algorithm; then uses the evolutionary algorithm to tune the parameters of the storage system.

[0004] In the above two parameter tuning methods for distributed storage systems, due to the large number of parameters in the storage system and the complex associations between parameters, it is very difficult to effectively and reasonably screen the parameters whether by using artificial experience or through a Bayesian optimization model. At the same time, the parameter tuning model is too simple, and it is very difficult to recommend a reasonable optimal configuration from the parameter space in a short time. Summary of the Invention

[0005] Aiming at the above defects or improvement requirements of the prior art, the present invention provides an automatic parameter tuning method and system for a distributed storage system based on reinforcement learning to solve the technical problem that the prior art cannot accurately tune the parameters of the distributed storage system reasonably and effectively.

[0006] To achieve the above object, in the first aspect, the present invention provides an automatic parameter tuning method for a distributed storage system based on reinforcement learning, including the following steps:

[0007] S1. Randomly select a set of configuration values of system parameters from the initial system parameter configuration set to tune the distributed storage system, obtain the corresponding performance indicators of the distributed storage system, and save the sample pair composed of the set of configuration values of the system parameters and the corresponding performance indicators of the distributed storage system to the first sample pool;

[0008] S2. Repeat step S1 for iteration until the number of sample pairs in the first sample pool reaches the first preset number;

[0009] S3. By statistically analyzing the performance metrics of the distributed storage system in the first sample pool, screen out m system parameters that have a greater impact on the performance of the distributed storage system from the first sample pool as a set of tuning parameters for the distributed storage system; 2 ≤ m ≤ M; M is the total number of a set of system parameters randomly selected in step S1;

[0010] S4. After obtaining the configuration values of each tuning parameter using the GA model, perform parameter tuning on the distributed storage system to obtain the corresponding performance metrics of the distributed storage system, and save the sample pair composed of the configuration values of this set of tuning parameters and the corresponding performance metrics of the distributed storage system to the second sample pool; Based on the performance metrics of the distributed storage system corresponding to the configuration values of this set of tuning parameters, calculate the fitness of the configuration values of this set of tuning parameters, and update the parameters in the GA model based on the fitness;

[0011] S5. Repeat step S4 for iteration until the number of sample pairs in the second sample pool reaches the second preset number or the GA model converges;

[0012] S6. Use the sample pairs in the second sample pool as the training set to pre-train the DDPG model;

[0013] S7. After obtaining the configuration values of each tuning parameter using the pre-trained DDPG model, perform parameter tuning on the distributed storage system, and obtain the corresponding performance metrics of the distributed storage system to calculate the corresponding reward value to perform online training on the DDPG model;

[0014] S8. Use the DDPG model after online training to obtain the configuration values of each tuning parameter to perform parameter tuning on the distributed storage system.

[0015] Further preferably, the above step S3 includes: respectively calculating the influence values of each system parameter in the first sample pool on the performance of the distributed storage system, and selecting the first m system parameters with larger influence values as a set of tuning parameters for the distributed storage system;

[0016] Among them, the influence value PI(p i ) of the i-th system parameter p on the performance of the distributed storage system is: i ) is:

[0017]

[0018] Among them, K i is the total number of the preset value range of the i-th system parameter p i ; is the average value of the performance metrics of the distributed storage system in ; is the average value of the performance metrics of the distributed storage system in ; is the variance of the performance metrics of the distributed storage system in ; is the variance of the performance metrics of the distributed storage system in ; The set is the set of sample pairs of the i-th system parameter p in the set S1 i within its k-th preset value range; The set is the set of sample pairs of the i-th system parameter p in the set S2 i within its k-th preset value range; The set S1 is the set composed of the top a% of the sample pairs after sorting the sample pairs in the set S in descending order; The set S2 is the set composed of the bottom b% of the sample pairs after sorting the sample pairs in the set S in descending order; Avg(S) is the average value of the performance metrics of the distributed storage system in the set S; STD(S) is the variance of the performance metrics of the distributed storage system in the set S; The set S is the set composed of all sample pairs in the first sample pool.

[0019] Further preferably, the above step S3 includes:

[0020] S31. Initialize the set P as the set composed of all system parameters in the first sample pool;

[0021] S32. After initializing the set S as the set composed of all sample pairs in the first sample pool, calculate the influence values of each system parameter in the set P on the performance of the distributed storage system respectively, and select the system parameter p with the largest influence value * and add it to the tuning parameter set;

[0022] S33. Remove the system parameter p * from the set P;

[0023] S34. Obtain multiple sets composed of the sample pairs of the system parameter p in the set S * respectively within its respective preset value ranges, and obtain the set set K * is the total number of preset value ranges of the system parameter p * ; The set set Assign the sets in it to set S respectively, and under each assignment, calculate the influence value of each system parameter in set P on the performance of the distributed storage system; calculate the sum of the influence values of each system parameter in set P on the performance of the distributed storage system under different assignments, and select the system parameter p with the largest sum of influence values * , and add it to the set of tuning parameters;

[0024] S34. Repeat steps S33 - S34 for iteration until the number of iterations reaches m times. At this time, the parameters in the set of tuning parameters are used as a set of tuning parameters for the distributed storage system;

[0025] Among them, the influence value PI(p i ) of the i-th system parameter p on the performance of the distributed storage system is: i ) is:

[0026]

[0027] K i is the total number of the preset value ranges of the i-th system parameter p i ; is the average value of the performance metrics of the distributed storage system in the set ; is the average value of the performance metrics of the distributed storage system in the set ; is the variance of the performance metrics of the distributed storage system in the set ; is the variance of the performance metrics of the distributed storage system in the set ; The set is the set composed of the samples of the i-th system parameter p i in its k-th preset value range in set S1; The set is the set composed of the samples of the i-th system parameter p i in its k-th preset value range in set S2; Set S1 is the set composed of the first a% of the sample pairs after sorting the sample pairs in set S from largest to smallest; Set S2 is the set composed of the last b% of the sample pairs after sorting the sample pairs in set S from largest to smallest; Avg(S) is the average value of the performance metrics of the distributed storage system in set S; STD(S) is the variance of the performance metrics of the distributed storage system in set S.

[0028] Further preferably, the above-mentioned performance metrics of the distributed storage system include: throughput or latency.

[0029] Further preferably, when the performance metric of the distributed storage system is throughput, the above-mentioned fitness f is:

[0030]

[0031]

[0032]

[0033] When the performance metric of the distributed storage system is latency, the above fitness f is:

[0034]

[0035]

[0036]

[0037] where T t is the throughput of the distributed storage system at time t; T0 is the throughput of the distributed storage system at the initial time; T t-1 is the throughput of the distributed storage system at time t - 1; L t is the latency of the distributed storage system at time t; L0 is the latency of the distributed storage system at the initial time; L t-1 is the latency of the distributed storage system at time t - 1.

[0038] Further preferably, when the performance metric of the distributed storage system is throughput, the above reward value r is:

[0039]

[0040]

[0041]

[0042] When the performance metric of the distributed storage system is latency, the above reward value r is:

[0043]

[0044]

[0045]

[0046] where T t is the throughput of the distributed storage system at time t; T0 is the throughput of the distributed storage system at the initial time; T t-1 is the throughput of the distributed storage system at time t - 1; L t is the latency of the distributed storage system at time t; L0 is the latency of the distributed storage system at the initial time; L t-1 is the latency of the distributed storage system at time t - 1.

[0047] Further preferably, the method for automatically tuning parameters of the above distributed storage system further includes: step S9 executed between step S2 and step S3;

[0048] Step S9 includes: sampling the sample pairs in the first sample pool, and only retaining the sampled sample pairs in the first sample pool.

[0049] In a second aspect, the present invention provides an automatic parameter tuning system for a distributed storage system based on reinforcement learning, including: a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it executes the method for automatically tuning parameters of the distributed storage system provided in the first aspect of the present invention.

[0050] In a fourth aspect, the present invention further provides a computer-readable storage medium, the computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the storage medium is located to execute the method for automatically tuning parameters of the distributed storage system provided in the first aspect of the present invention.

[0051] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0052] 1. The present invention provides a method for automatically tuning parameters of a distributed storage system based on reinforcement learning. First, by statistically analyzing the performance indicators of the distributed storage system in the first sample pool, a plurality of system parameters that have a greater impact on the performance of the distributed storage system are screened out from the first sample pool as a set of tuning parameters for the distributed storage system, so as to reduce the complexity of the optimization problem and ensure the feasibility of subsequent parameter tuning based on reinforcement learning; then, on the basis of parameter screening, the DDPG model of reinforcement learning is used for further tuning work, and in this process, considering the problem of long cold start time in the early stage of the DDPG model, a genetic algorithm is used for the early sample collection work, and the collected samples are input to the DDPG model for pre-training, so as to avoid the problem of a large amount of time-consuming cold start in the early stage, and thus can reasonably and effectively accurately tune the parameters of the distributed storage system.

[0053] 2. The method for automatically tuning parameters of the distributed storage system provided by the present invention constructs a system parameter importance measurement model based on mean and variance. By fixing the value of a certain system parameter and observing its impact on the distributed storage system based on mean and variance, the importance of each system parameter is obtained; by sorting the importance of system parameters and screening out some parameters that have a greater impact on performance for subsequent adjustment, the accuracy and efficiency of system optimization can be further improved.

[0054] 3. The automatic parameter tuning method for the distributed storage system provided by the present invention combines the advantages of the GA model and the DDPG model, which can not only train the model as quickly as possible but also achieve a better tuning effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a flowchart of the automatic parameter tuning method for the distributed storage system based on reinforcement learning provided in Embodiment 1 of the present invention;

[0056] Figure 2 It is the framework structure of the entire tuning system provided in Embodiment 1 of the present invention;

[0057] Figure 3 It is a flowchart of the parameter screening model provided in Embodiment 1 of the present invention;

[0058] Figure 4 It is a process diagram of the interaction and learning between the reinforcement learning model and the environment in the tuning model provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0060] Embodiment 1

[0061] An automatic parameter tuning method for a distributed storage system based on reinforcement learning, as Figure 1 shown, includes the following steps:

[0062] S1. Randomly select a set of configuration values of system parameters from the initial system parameter configuration set to tune the distributed storage system, obtain the corresponding performance indicators of the distributed storage system, and save the sample pair composed of the set of configuration values of system parameters and the corresponding performance indicators of the distributed storage system to the first sample pool; wherein, randomly selecting a set of system parameters includes: 56 system parameters such as the number of threads of the osd process osd_op_threads, the maximum number of bytes for cache read-ahead client_readahead_max_bytes, the number of worker threads Nb_Worker, and the function switch Performance_Opt; the performance indicators of the distributed storage system can be selected according to specific requirements and can be throughput, latency, etc.

[0063] S2. Repeat step S1 for iteration until the number of sample pairs in the first sample pool reaches the first preset number. In this embodiment, the first preset number is 500.

[0064] S3. By statistically analyzing the performance metrics of the distributed storage system in the first sample pool, screen out m system parameters that have a greater impact on the performance of the distributed storage system from the first sample pool as a set of tuning parameters for the distributed storage system; 2 ≤ m ≤ M; M is the total number of a set of system parameters randomly selected in step S1; in this embodiment, m is 10.

[0065] Specifically, the present invention constructs a parameter importance measurement model based on variance, ranks the parameters according to importance, screens out some parameters that have a greater impact on performance, and performs subsequent adjustments. After such processing, the complexity of the problem can be greatly reduced, ensuring the feasibility of subsequent parameter tuning using the reinforcement learning algorithm.

[0066] In an alternative embodiment, step S3 includes: respectively calculating the influence values of each system parameter in the first sample pool on the performance of the distributed storage system, and selecting the first m system parameters with larger influence values as a set of tuning parameters for the distributed storage system;

[0067] Among them, the influence value PI(p i ) of the i-th system parameter p on the performance of the distributed storage system is: i ) is:

[0068]

[0069] Among them, K i is the total number of the preset value range of the i-th system parameter p i ; is the average value of the performance metrics of the distributed storage system in the set ; is the average value of the performance metrics of the distributed storage system in the set ; is the variance of the performance metrics of the distributed storage system in the set ; is the variance of the performance metrics of the distributed storage system in the set ; The set is the set composed of sample pairs of the i-th system parameter p i in its k-th preset value range in set S1; The set is the set composed of sample pairs of the i-th system parameter p iA set composed of sample pairs within its k-th preset value range; set S1 is a set composed of the top a% of the sample pairs after sorting the sample pairs in set S in descending order; set S2 is a set composed of the bottom b% of the sample pairs after sorting the sample pairs in set S in descending order; Avg(S) is the average value of the distributed storage system performance metrics in set S; STD(S) is the variance of the distributed storage system performance metrics in set S; set S is a set composed of all sample pairs in the first sample pool. Among them, a% and b% can be 10%, 15%, etc., and are specifically selected according to needs.

[0070] In an alternative implementation, step S3 includes:

[0071] S31. Initialize set P as a set composed of all system parameters in the first sample pool;

[0072] S32. After initializing set S as a set composed of all sample pairs in the first sample pool, calculate the influence values of each system parameter in set P on the performance of the distributed storage system, and select the system parameter p with the largest influence value * and add it to the tuning parameter set;

[0073] S33. Remove the system parameter p * from set P;

[0074] S34. Obtain multiple sets composed of sample pairs of the system parameter p * in their respective preset value ranges in set S, and obtain a set set K * is the total number of preset value ranges of the system parameter p * ; Assign the sets in the set set to set S respectively, and under each assignment, calculate the influence values of each system parameter in set P on the performance of the distributed storage system; Calculate the sum of the influence values of each system parameter in set P on the performance of the distributed storage system under different assignments, and select the system parameter p * with the largest sum of influence values and add it to the tuning parameter set;

[0075] S34. Repeat steps S33 - S34 for iteration until the number of iterations reaches m times. At this time, the parameters in the tuning parameter set are used as a set of tuning parameters for the distributed storage system;

[0076] Among them, the influence value PI(p i ) of the i-th system parameter p on the performance of the distributed storage system is: i ) is:

[0077]

[0078] K i is the total number of preset value ranges for the i-th system parameter p i ; is the average value of the distributed storage system performance metrics in the set ; is the average value of the distributed storage system performance metrics in the set ; is the variance of the distributed storage system performance metrics in the set ; is the variance of the distributed storage system performance metrics in the set ; The set is the set of sample pairs of the i-th system parameter p in the set S1 i in its k-th preset value range; The set is the set of sample pairs of the i-th system parameter p in the set S2 i in its k-th preset value range; The set S1 is the set composed of the top a% of the sample pairs after sorting the sample pairs in the set S in descending order; The set S2 is the set composed of the bottom b% of the sample pairs after sorting the sample pairs in the set S in descending order; Avg(S) is the average value of the distributed storage system performance metrics in the set S; STD(S) is the variance of the distributed storage system performance metrics in the set S. Among them, a% and b% can be 10%, 15%, etc., and are specifically selected according to needs.

[0079] It should be noted that each system parameter has its corresponding multiple preset value ranges; for example, the conventional value range of the number of worker threads Nb_Worker is 60 - 140, which can be divided into four preset value ranges: 60 - 79, 80 - 99, 100 - 119, and 120 - 139; the conventional value range of the function switch Performance_Opt is false or true, which is divided into two preset value ranges: false and true.

[0080] S4. After obtaining the configuration values of each tuning parameter using the GA model, tune the distributed storage system, obtain the corresponding distributed storage system performance metrics, and save the sample pairs composed of this set of tuning parameter configuration values and the corresponding distributed storage system performance metrics to the second sample pool; calculate the fitness f of this set of tuning parameter configuration values based on the distributed storage system performance metrics corresponding to this set of tuning parameter configuration values, and update the parameters in the GA model based on the fitness;

[0081] Specifically, when the distributed storage system performance metric is throughput, the above fitness f is:

[0082]

[0083]

[0084]

[0085] Among them, T t is the throughput of the distributed storage system at time t; T0 is the throughput of the distributed storage system at the initial time; T t-1 is the throughput of the distributed storage system at time t - 1.

[0086] When the performance metric of the distributed storage system is latency, the above fitness f is:

[0087]

[0088]

[0089]

[0090] Among them, L t is the latency of the distributed storage system at time t; L0 is the latency of the distributed storage system at the initial time; L t-1 is the latency of the distributed storage system at time t - 1.

[0091] S5. Repeat step S4 for iteration until the number of sample pairs in the second sample pool reaches the second preset number or the GA model converges; in this embodiment, the second preset number is 300.

[0092] S6. Use the sample pairs in the second sample pool as the training set to pre-train the DDPG model;

[0093] Specifically, use the multiple sets of tuned parameter configuration values in the second sample pool as the input, and the corresponding distributed storage system performance metrics as the output to perform supervised training on the DDPG model.

[0094] S7. After obtaining the configuration values of each tuned parameter using the pre-trained DDPG model, adjust the parameters of the distributed storage system, and obtain the corresponding distributed storage system performance metrics to calculate the corresponding reward value for online training of the DDPG model;

[0095] Specifically, when the performance metric of the distributed storage system is throughput, the above reward value r is:

[0096]

[0097]

[0098]

[0099] Among them, T t is the throughput of the distributed storage system at time t; T0 is the throughput of the distributed storage system at the initial time; T t-1 is the throughput of the distributed storage system at time t - 1.

[0100] When the performance metric of the distributed storage system is latency, the above reward value r is:

[0101]

[0102]

[0103]

[0104] Among them, L t is the latency of the distributed storage system at time t; L0 is the latency of the distributed storage system at the initial time; L t-1 is the latency of the distributed storage system at time t - 1.

[0105] S8. Use the DDPG model after online training to obtain the configuration values of each tuning parameter to tune the parameters of the distributed storage system.

[0106] Based on parameter screening, the present invention uses the DDPG model of reinforcement learning for further tuning work. Considering the problem of the long cold start time of the DDPG model in the early stage, the genetic algorithm is first used for the sample collection work in the early stage, and the collected samples are input to the DDPG model for pre-training, so as to avoid the problem of a large amount of time consumption in the early cold start. The present invention proposes an efficient and reasonable tuning method, which can quickly adjust the parameters of the distributed storage system and improve the system performance. [[ID=�3]]

[0107] Furthermore, in an alternative embodiment, the above method for automatically tuning the parameters of the distributed storage system further includes: step S9 executed between step S2 and step S3;

[0108] Step S9 includes: sampling the sample pairs in the first sample pool and only retaining the sampled sample pairs in the first sample pool. Specifically, the Latin hypercube sampling method, importance sampling method, slice sampling method, etc. can be used to sample the sample pairs in the first sample pool.

[0109] To further illustrate the method for automatically tuning the parameters of the distributed storage system provided by the present invention, the following will be described in detail in combination with specific embodiments:

[0110] Figure 2It shows the framework structure of the entire tuning system. The entire system is mainly divided into four modules: the environment module, the configuration recommendation model module, the parameter configuration set module, and the sample pool module. The environment module consists of a workload generator, a storage system, and a performance monitor. The workload generator is used to simulate the IO load in real work. The storage system can be any mainstream distributed storage system. The performance monitor is used to monitor the performance metrics of the storage system to measure whether the current parameter configuration is effective and reasonable.

[0111] The parameter configuration recommendation model module includes: a random test model, the DDPG model in reinforcement learning, and the GA (Genetic Algorithm) model. These three models respectively correspond to the three sample pools in the sample pool module. The random test model recommends the configurations of each parameter according to a random policy, interacts with the environment to obtain performance feedback, and saves it in the random sample pool (the first sample pool). The GA model recommends the configurations of each parameter according to an evolutionary strategy based on the screened parameters, interacts with the environment to obtain performance feedback, and saves it in the GA sample pool (the second sample pool). The DDPG model first pre-trains using the samples in the GA sample pool, then continuously updates and trains the model by interacting with the environment, and at the same time puts the samples into the DDPG sample pool.

[0112] The parameter configuration module includes: an initial configuration and a running configuration. The initial configuration is a set of parameter configurations obtained after rough screening based on manual parameter tuning experience. The running configuration is based on the initial configuration, and uses a parameter screening model to measure the importance of each parameter (the impact value of system parameters on the performance of the distributed storage system) for the random samples, so as to screen out the final running configuration.

[0113] The specific operation process of the system includes the following processes:

[0114] (1) First, based on manual parameter tuning experience, rough screen the original set of parameter configurations to obtain an initial set of system parameter configurations.

[0115] (2) The random model in the configuration recommendation model module recommends the configurations of each parameter according to a random policy, interacts with the environment to obtain performance feedback, and saves it in the random sample pool;

[0116] (3) The parameter importance ranking model measures the importance of each parameter based on the random samples in the random sample pool, ranks them, and further finely screens the system parameters to obtain the final set of tuned parameters.

[0117] (4) The GA model in the configuration recommendation model module recommends the configurations of each tuned parameter according to an evolutionary strategy for the set of tuned parameters, interacts with the environment to obtain performance feedback, and saves it in the GA sample pool.

[0118] (5) The DDPG model in the configuration recommendation model module first uses the samples in the GA sample pool for pre-training to avoid the "cold start" problem in the early stage. Then, by interacting with the environment, the training model is continuously updated, and at the same time, the samples are put into the DDPG sample pool.

[0119] (6) The finally trained DDPG model will recommend the specific tuning parameter configuration suitable for the current IO scenario.

[0120] Figure 3 It is the algorithm description of the parameter screening model. First, determine the original parameter set P, and then take out a batch of samples S with a size of n from the random sample pool. This batch of samples consists of <system parameter configuration, performance>. Finally, determine the finish function: stop the process when 10 important parameters are currently selected. It includes the following steps:

[0121] Continuously collect random samples through the random model until n samples are collected, denoted as set S;

[0122] Perform Latin hypercube sampling on set S to obtain the set S after sampling * ;

[0123] For set S * Calculate the parameter importance, and select the most important parameter p * , add it to the selected set, and delete p from the P set * .

[0124] Further, Table 1 is the algorithm description of the genetic algorithm in the tuning model. Here, the genetic algorithm is mainly used to collect some high-quality training samples to quickly pre-train the DDPG model. Among them, first, the value of each parameter is normalized to the range of [0, 1] according to the corresponding value range, and then combined into a vector as the coding gene of the individual; the fitness evaluation function is set as the increase in the system performance compared to the default parameters of the current system. Finally, after setting the hyperparameters such as P c 、P m 、M、G, etc., according to the execution process, recommend appropriate parameters, then apply them to the storage system, and after evaluating the fitness of each individual, perform elimination, crossover, and mutation. Finally, after evolving through G generations, collect the samples in the process into the sample pool.

[0125] Table 1

[0126]

[0127] Figure 4It is a process diagram of the interaction and learning between the reinforcement learning model and the environment in the tuning model. The DDPG algorithm is an algorithm in reinforcement learning. Reinforcement learning (RL for short) is a field in machine learning that emphasizes how to act based on the environment to maximize the expected benefit. Its inspiration comes from the behaviorist theory in psychology, that is, how an organism gradually forms an expectation of a stimulus under the stimulation of rewards or punishments given by the environment, and generates a habitual behavior that can obtain the maximum benefit. Among them, there are two objects that can interact: the agent and the environment. The agent can perceive the state of the environment and, based on the feedback reward and the set policy, learn to select an appropriate action to maximize the long-term total reward. The environment will receive a series of actions executed by the agent, evaluate this series of actions and convert them into a quantifiable signal and feedback it to the agent. Here, we normalize the values of each parameter to the range [0, 1] according to the corresponding value range, and then combine them into a vector as the Action; combine the internal CPU, network and other information of the system as the State; according to the performance of the system and compare the increase with the performance in the default configuration as the Reward. However, due to the poor exploration performance and weak sense of direction of the DDPG model in the early stage, it cannot quickly find a better parameter configuration. Therefore, it is necessary to first use the high-quality samples in the GA sample pool to perform rapid pre-training on it, and then interact and train with the environment.

[0128] Example 2

[0129] A distributed storage system automatic parameter tuning system based on reinforcement learning, comprising: a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it executes the distributed storage system automatic parameter tuning method provided in Embodiment 1 of the present invention.

[0130] The related technical solutions are the same as those in Embodiment 1 and will not be elaborated here.

[0131] Example 3

[0132] A computer-readable storage medium, the computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the storage medium is located to execute the distributed storage system automatic parameter tuning method provided in Embodiment 1 of the present invention.

[0133] The related technical solutions are the same as those in Embodiment 1 and will not be elaborated here.

[0134] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An automatic parameter tuning method for a distributed storage system based on reinforcement learning, characterized in that The steps include the following: S1. Randomly select a set of configuration values of system parameters from the initial system parameter configuration set to tune the distributed storage system, obtain the corresponding performance metrics of the distributed storage system, and save the sample pair composed of the set of configuration values of the system parameters and the corresponding performance metrics of the distributed storage system to the first sample pool; S2. Repeat the step S1 for iteration until the number of sample pairs in the first sample pool reaches the first preset number; S3. By statistically analyzing the performance metrics of the distributed storage system in the first sample pool, select m system parameters that have a greater impact on the performance of the distributed storage system from the first sample pool as a set of tuning parameters for the distributed storage system; 2 ≤ m ≤ M; M is the total number of a set of system parameters randomly selected in the step S1; S4. After obtaining the configuration values of each tuning parameter by using the genetic algorithm GA model, tune the distributed storage system, obtain the corresponding performance metrics of the distributed storage system, and save the sample pair composed of the set of configuration values of the tuning parameters and the corresponding performance metrics of the distributed storage system to the second sample pool; Calculate the fitness of the set of configuration values of the tuning parameters based on the performance metrics of the distributed storage system corresponding to the set of configuration values of the tuning parameters, and update the parameters in the GA model based on the fitness; S5. Repeat the step S4 for iteration until the number of sample pairs in the second sample pool reaches the second preset number or the GA model converges; S6. Use the sample pairs in the second sample pool as a training set to pre-train the deep deterministic policy gradient (DDPG) reinforcement learning model; S7. After obtaining the configuration values of each tuning parameter by using the pre-trained DDPG model, tune the distributed storage system, and obtain the corresponding performance metrics of the distributed storage system to calculate the corresponding reward value for online training of the DDPG model; S8. Use the DDPG model after online training to obtain the configuration values of each tuning parameter to tune the distributed storage system; The performance metrics of the distributed storage system include: throughput or latency; When the performance metric of the distributed storage system is throughput, the fitness f is: When the performance metric of the distributed storage system is latency, the fitness f is: Among them, T t is the throughput of the distributed storage system at time t; T0 is the throughput of the distributed storage system at the initial time; T t-1 is the throughput of the distributed storage system at time t - 1; L t is the latency of the distributed storage system at time t; L0 is the latency of the distributed storage system at the initial time; L t-1 is the latency of the distributed storage system at time t - 1.

2. The automatic parameter tuning method for a distributed storage system according to claim 1, wherein The step S3 includes: respectively calculating the influence values of each system parameter in the first sample pool on the performance of the distributed storage system, and selecting the first m system parameters with larger influence values as a set of tuning parameters for the distributed storage system; The i-th system parameter p i The impact value PI(p i ) on the performance of the distributed storage system is as follows: Among them, K i is the total number of the preset value ranges of the i-th system parameter p i ; is the average value of the performance metrics of the distributed storage system in the set ; is the average value of the performance metrics of the distributed storage system in the set ; is the variance of the performance metrics of the distributed storage system in the set ; is the variance of the performance metrics of the distributed storage system in the set ; The set is the set composed of the sample pairs of the i-th system parameter p i in its k-th preset value range; The set is the set composed of the sample pairs of the i-th system parameter p i in its k-th preset value range; The set S1 is the set composed of the first a% of the sample pairs after sorting the sample pairs in the set S from largest to smallest; The set S2 is the set composed of the last b% of the sample pairs after sorting the sample pairs in the set S from largest to smallest; Avg(S) is the average value of the performance metrics of the distributed storage system in the set S; STD(S) is the variance of the performance metrics of the distributed storage system in the set S; The set S is the set composed of all the sample pairs in the first sample pool.

3. The automatic parameter tuning method for the distributed storage system according to claim 1, wherein The step S3 includes: S31. Initialize the set P as the set composed of all system parameters in the first sample pool; S32. After initializing the set S as the set composed of all sample pairs in the first sample pool, calculate the influence values of each system parameter in the set P on the performance of the distributed storage system respectively, and select the system parameter p with the largest influence value * and add it to the set of tuning parameters; S33. Remove the system parameter p * from the set P; S34. Obtain the system parameter p in the set S * A plurality of sets formed by sample pairs within their respective preset value ranges are obtained to form a set of sets K * is the total number of preset value ranges of the system parameter p * ; assign the sets in the set of sets to the set S respectively, and under each assignment, calculate the influence value of each system parameter in the set P on the performance of the distributed storage system; calculate the sum of the influence values of each system parameter in the set P on the performance of the distributed storage system under different assignments, and select the system parameter p with the largest sum of influence values * and add it to the tuning parameter set S34. Repeat the steps S33 - S34 for iteration until the number of iterations reaches m times. At this time, the parameters in the tuning parameter set are used as a set of tuning parameters for the distributed storage system; Among them, the i-th system parameter p i The influence value PI(p i ) on the performance of the distributed storage system is as follows: K i is the total number of preset value ranges of the i-th system parameter p i ; is the average value of the distributed storage system performance metrics in the set ; is the average value of the distributed storage system performance metrics in the set ; is the variance of the distributed storage system performance metrics in the said set ; is the variance of the distributed storage system performance metrics in the said set ; the said set is the set composed of sample pairs of the i-th system parameter p in its k-th preset value range in the set S1 i ; the said set is the set composed of sample pairs of the i-th system parameter p in its k-th preset value range in the set S2 i ; the set S1 is the set composed of the top a% of the sample pairs after sorting the sample pairs in the set S in descending order; the set S2 is the set composed of the bottom b% of the sample pairs after sorting the sample pairs in the set S in descending order; Avg(S) is the average value of the distributed storage system performance metrics in the set S; STD(S) is the variance of the distributed storage system performance metrics in the set S.

4. The automatic parameter tuning method for the distributed storage system according to claim 1, wherein When the performance metric of the distributed storage system is throughput, the reward value r is: When the performance metric of the distributed storage system is latency, the reward value r is: Among them, T t is the throughput of the distributed storage system at time t; T0 is the throughput of the distributed storage system at the initial time; T t-1 is the throughput of the distributed storage system at time t - 1; L t is the latency of the distributed storage system at time t; L0 is the latency of the distributed storage system at the initial time; L t-1 is the latency of the distributed storage system at time t - 1.

5. The automatic parameter tuning method for the distributed storage system according to claim 1, wherein It also includes: Step S9 performed between the said step S2 and the said step S3; The said step S9 includes: sampling the sample pairs in the first sample pool, and only retaining the sampled sample pairs in the first sample pool.

6. An automatic parameter tuning system for a distributed storage system based on reinforcement learning, characterized in that, Including: A memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it executes the automatic parameter tuning method for the distributed storage system according to any one of claims 1-5.

7. A computer-readable storage medium, characterized in that, The said computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the storage medium is located to execute the automatic parameter tuning method for the distributed storage system according to any one of claims 1-5.

Citation Information

Patent Citations

  • Parameter tuning method and device based on distributed storage system, equipment and medium

    CN114564460A

  • Parameter tuning method, system and device for distributed storage system

    CN113608677A

  • Parameter tuning method and system of distributed storage system, equipment and medium

    CN114415945A