Storage system parameter adjusting and optimizing method and device based on reinforcement learning

Through reinforcement learning combined with genetic algorithms and deep deterministic strategy gradient algorithms, the distributed storage system parameters are automatically tuned, which solves the problems of time-consuming and poor adaptability in the existing technology, and realizes efficient parameter configuration.

CN120371215APending Publication Date: 2025-07-25HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510506680.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The performance tuning of modern distributed storage systems is difficult, especially the large differences in the needs of different users and jobs, resulting in low resource utilization. The existing tuning methods are time-consuming and difficult to adapt to dynamic load changes. Ordinary users lack effective guidance.

Method used

The parameter tuning method of storage system based on reinforcement learning is adopted, combined with genetic algorithm and deep deterministic strategy gradient algorithm, and important parameters are screened through random sample sets, and the parameter optimization is used for genetic algorithm and deep deterministic strategy gradient algorithm to output the final key parameters.

Benefits of technology

It realizes automatic and accurate parameter adjustment of distributed storage systems, improves tuning efficiency, quickly finds the best parameter configuration, avoids the cold start problem of deep deterministic policy gradient algorithm, and ensures the accuracy and efficiency of parameter configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371215A_ABST
    Figure CN120371215A_ABST
Patent Text Reader

Abstract

The invention belongs to the related technical field of storage system parameter adjustment, and discloses a storage system parameter adjustment and optimization method and adjustment and optimization equipment based on reinforcement learning, and the adjustment and optimization method comprises the steps: carrying out the random configuration of global parameters in an initial parameter space, and constructing a random sample set; on the basis of the random sample set, performing importance sorting on the parameters in the initial parameter space, and obtaining important parameters which have the greatest influence on the performance index IOPS to form a new parameter space; taking {important parameter configuration, IOPS} as an individual, executing a genetic algorithm to optimize and configure parameters in a new parameter space, and obtaining a genetic sample set; taking configuration of important parameters as an action, taking IOPS of a storage system after the action as a state, firstly obtaining initial network parameters of a depth deterministic strategy gradient algorithm model based on a genetic sample set, then exploring configuration of the important parameters in a new parameter space, and outputting final key parameters. Through the method, the parameters of the storage system can be automatically and accurately adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field related to storage system parameter tuning, and more specifically, relates to a method for tuning storage system parameters based on reinforcement learning and a tuning device therefor. Background Art

[0002] For modern distributed storage systems, performance tuning is an important but difficult task. With the emergence of new storage devices and more complex workloads, the demands of different users and different jobs vary greatly. Default configuration parameters often result in low resource utilization and poor performance. It is necessary to adjust the parameters according to the specific situation.

[0003] On the other hand, parameter tuning in a modern storage environment is challenging. Even for human experts, tuning parameters is a time-consuming task. When tuning parameters, they often need to analyze the reasons for the bottleneck in system performance by debugging the underlying code and repeatedly experiment to summarize the relationship between parameters and performance. Worse still, the workload may change over time or vary periodically, and the optimal configuration for one workload may perform poorly for other workloads. Finding the optimal value for all possible workloads is costly and almost impossible. For ordinary users, tuning is even more daunting. Most users can only fall back on following an inflexible and uncustomized performance tuning guide. However, many systems even lack a functional description of parameters and tuning guidance, leaving users confused.

[0004] Therefore, there is an urgent need to propose a storage system parameter tuning technology to achieve automatic and accurate parameter tuning of the storage system. Summary of the Invention

[0005] In view of the above defects or improvement requirements of the prior art, the present invention provides a method for tuning storage system parameters based on reinforcement learning and a tuning device therefor, aiming to achieve automatic and accurate parameter tuning of the storage system.

[0006] To achieve the above object, according to the first aspect of the present invention, there is provided a method for tuning storage system parameters based on reinforcement learning, which includes:

[0007] S1. Randomly configure the global parameters in the initial parameter space and construct a random sample set S, where the samples are {global parameter configuration, IOPS}, and the IOPS in the samples is the IOPS performance metric of the storage system under the corresponding parameter configuration;

[0008] S2. Based on the random sample set S, perform importance ranking on each parameter in the initial parameter space and obtain the top m important parameters that have the greatest impact on the performance metric IOPS, forming a new parameter space;

[0009] S3. Take {important parameter configuration, IOPS} as an individual, and multiple individuals form a population. Use the improvement degree of the IOPS in the current individual compared to the IOPS of the storage system under the default configuration as the fitness function of the current individual. Execute the genetic algorithm to optimize the configuration of the parameters in the new parameter space, and obtain the individuals output by the genetic algorithm to form a genetic sample set;

[0010] S4. Take the configuration of important parameters as an action and the IOPS of the storage system after the action as the state. First, obtain the initial network parameters of the deep deterministic policy gradient algorithm model based on the genetic sample set, and then explore the configuration of the important parameters in the new parameter space based on the deterministic policy gradient algorithm model to output the final key parameters;

[0011] The importance ranking of each parameter in the initial parameter space includes: calculating the importance degree of each parameter and performing importance ranking based on the importance degree; the calculation process of calculating the importance degree of any parameter p includes:

[0012] Calculate the performance variance Var(IOPS) of the random sample set S;

[0013] Divide the values of parameter p in the random sample set S into intervals to form multiple subsets, calculate the variance weighted average of all subsets of this parameter p, and calculate the ratio of the performance variance Var(IOPS) to the variance weighted average of this parameter p as the importance degree I(p) of this parameter p.

[0014] Optionally, the calculation formula of the fitness function of the current individual is:

[0015]

[0016] In the formula, f(s i ) is the fitness of individual s i , IOPS cur and IOPS0 respectively represent the IOPS in the current individual s i and the IOPS of the storage system under the default configuration.

[0017] Optionally, the reward function set in the deep deterministic policy gradient algorithm model is:

[0018]

[0019] In the formula, r t-1 represents the reward obtained by executing the action at time t - 1, Δ t->0 represents the improvement degree of the IOPS in the state of the agent at time t compared to the IOPS of the system under the default configuration, Δ t->t-1Indicates the improvement degree of the IOPS in the state of the agent at time t compared to the IOPS in the state of the agent at time t-1.

[0020] Optionally, the improvement degree Δ t->0 and Δ t->t-1 are calculated by the following formula:

[0021]

[0022] In the formula, IOPS t is the IOPS in the state of the agent at time t, and IOPS t-1 is the IOPS in the state of the agent at time t-1, and IOPS0 is the IOPS of the system under the default configuration.

[0023] Optionally, the process of obtaining the initial network parameters of the deep deterministic policy gradient algorithm model based on the genetic sample set includes:

[0024] S401. Place the genetic sample set in the experience replay pool, initialize the network parameters of the deep deterministic policy gradient algorithm model, and fix its target network parameters;

[0025] S402. Randomly sample a batch of samples from the experience replay pool according to the uniform sampling method, and update the parameters of the backbone value network and the backbone policy network based on the sampled samples;

[0026] S403. Overwrite the parameters of the backbone value network to the target value network, and overwrite the parameters of the backbone policy network to the target policy network;

[0027] S404. Repeat steps S402 to S403 until the preset number of training times is reached.

[0028] Optionally, in S2 and S3, the non-important parameters of the storage system are randomly configured or configured by default.

[0029] According to the second aspect of the present invention, there is provided a storage system parameter tuning device based on reinforcement learning, which includes:

[0030] A random sample acquisition unit, configured to randomly configure global parameters in the initial parameter space and construct a random sample set S, where the samples are {global parameter configuration, IOPS}, and the IOPS in the samples is the IOPS performance index of the storage system under the corresponding parameter configuration;

[0031] A parameter importance selection unit, configured to perform importance ranking on each parameter in the initial parameter space based on the random sample set S and obtain the top m important parameters that have the greatest impact on the performance index IOPS, so as to form a new parameter space;

[0032] A genetic algorithm execution unit, which uses {important parameter configuration, IOPS} as an individual, multiple individuals form a population, uses the improvement degree of the IOPS in the current individual compared to the IOPS of the storage system under the default configuration as the fitness function of the current individual, executes the genetic algorithm to optimize the configuration of the parameters in the new parameter space, and obtains the individuals output by the genetic algorithm to form a genetic sample set;

[0033] A deep deterministic policy gradient algorithm execution unit, which uses the configuration of important parameters as an action and the IOPS of the storage system after the action as a state. First, based on the genetic sample set, it obtains the initial network parameters of the deep deterministic policy gradient algorithm model, and then explores the configuration of the important parameters in the new parameter space based on the deterministic policy gradient algorithm model, and outputs the final key parameters.

[0034] According to the third aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0035] According to the fourth aspect of the present invention, there is provided a computer program product, including a computer program or instruction, wherein when the computer program or instruction is executed by a processor, the steps of the method described in any one of the above are implemented.

[0036] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the present invention mainly has the following beneficial effects:

[0037] 1. The storage system parameter tuning method provided by the present invention combines the genetic algorithm and the deep deterministic policy gradient algorithm to achieve automatic tuning of the parameters of the distributed storage system. Compared with manual tuning, it can effectively improve the tuning efficiency. Moreover, both the genetic algorithm and the deep deterministic policy gradient algorithm are optimization algorithms. In the present invention, the IOPS performance index of the storage system is used as the optimization target to tune the parameters of the storage system. Through algorithm iteration, a parameter configuration that makes the IOPS performance index of the storage system better can be found, realizing accurate tuning of the parameters of the distributed storage system. In the present invention, first, a random sample is obtained through a random policy, and then the parameter importance evaluation is carried out to obtain important parameters. The important parameters are the key parameters affecting the system performance. Therefore, when the genetic algorithm and the deterministic policy gradient algorithm are subsequently executed, both explore the selected important parameters, and by shrinking the parameter space, the optimization speed can be improved and the overall configuration effect can be ensured. Moreover, before executing the deep deterministic policy gradient algorithm, the present invention first executes the genetic algorithm to obtain a genetic sample set, and based on the genetic sample set, the deep deterministic policy gradient algorithm model is quickly trained, which can avoid the cold start problem of the deep deterministic policy gradient algorithm, make the deep deterministic policy gradient algorithm quickly enter the effective exploration stage, and thus quickly output the parameter configuration result.

[0038] 2. Further, in some embodiments, the variance weighted average of all subsets of the parameter p is calculated, and the ratio of the performance variance Var(IOPS) to the variance weighted average of the parameter p is used as the importance I(p) of the parameter p. Since an influential parameter can significantly reduce the variance of the performance value after fixing its value, the larger the variance ratio of the total sample to the weighted average, the greater the importance of the parameter. Through this method, the importance of each parameter in the initial parameter space can be accurately evaluated.

[0039] 3. Further, in some embodiments, the design form of the reward function in the deep deterministic policy gradient algorithm model is provided. The reward function comprehensively considers the difference Δ t->0 between the performance value of this tuning and the default performance value and the difference Δ t->t-1 between the performance value of this tuning and the performance value of the previous tuning. If Δ t->0 is positive, the tuning trend is correct and the reward is positive; otherwise, the tuning trend is incorrect and the reward is negative. Δ t->0 determines the direction of the reward, while Δ t->t-1 affects the magnitude of the reward to a certain extent. Based on the above-set reward function, the parameter tuning of the model can be carried out in the correct direction, and at the same time, the optimal parameter configuration can be quickly searched. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1It is a flowchart of the steps of a storage system parameter tuning method in an embodiment of the present invention;

[0041] Figure 2 It is a schematic diagram of the steps of a storage system parameter tuning method in an embodiment of the present invention;

[0042] Figure 3 It is a flowchart of the DDPG algorithm in an embodiment of the present invention;

[0043] Figure 4 It is a structural block diagram of a storage system parameter tuning device in an embodiment of the present invention. Detailed implementation manners

[0044] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0045] Embodiment 1

[0046] The present invention proposes a storage system parameter tuning method based on reinforcement learning. As Figure 1 shown is a flowchart of the steps of a storage system parameter tuning method in an embodiment of the present invention. As Figure 2 shown is a schematic diagram of the steps of a storage system parameter tuning method in an embodiment of the present invention. The core steps are described in detail below.

[0047] S1. Randomly configure the global parameters in the initial parameter space to obtain a random sample set S, where the samples are {global parameter configuration, IOPS}, and the IOPS in the samples is the IOPS performance index of the storage system under the corresponding parameter configuration.

[0048] Specifically, first obtain random samples through a random policy to recommend parameter configurations, and form the sample set S.

[0049] Among them, the initial parameter space contains all parameters to be configured, denoted as P Z ={p 1 , p 2 , p 3 , ……, p z}, where z is the total number of parameters, and p i is the i-th parameter in this parameter space. The parameters include, for example, the number of shards describing the osd and the number of working threads for each shard, etc.

[0050] After each set of parameter configurations is generated, the workload generator generates a specific IO workload that conforms to the scenario according to the current parameter configuration and makes the storage system run this workload. Subsequently, performance monitoring is carried out, and stress testing is performed using the built-in tool Rados bench of the storage system to obtain and record the performance metric IOPS of the system in real time. The current parameter configuration and the corresponding system performance metric IOPS form a sample {parameter configuration, IOPS}.

[0051] S2. Based on the random sample set S, perform importance ranking on each parameter in the initial parameter space and obtain the top m parameters that have the greatest impact on the performance metric IOPS, forming a new parameter space.

[0052] Specifically, key parameters can be obtained according to the traditional principal component analysis method to construct a new parameter space.

[0053] In this embodiment, the following method is used to calculate the importance of each parameter.

[0054] S21. Calculate the performance variance Var(IOPS) of the random sample set S.

[0055] Among them, the calculation formula of the performance variance Var(IOPS) can be expressed as:

[0056]

[0057] In the formula, |S| represents the number of samples in the random sample set S, y i represents the performance value IOPS of the i-th sample, and μ represents the average performance value of the random sample set S.

[0058] S22. For each parameter p in the random sample set S, divide the value of the parameter p according to intervals to form multiple subsets, calculate the weighted average of the variances of all subsets of the parameter p, and calculate the ratio of the performance variance Var(IOPS) to the weighted average of the variances of the parameter p as the importance degree I(p) of the parameter p.

[0059] Among them, the calculation formula of the importance degree I(p) of the parameter p can be expressed as:

[0060]

[0061] Among them, M i is the i-th subset of the parameter P, |Mi| is the number of samples in the subset M i , n is the number of subsets; y j represents the performance value of the j-th sample in the subset M i ; μ i represents the average performance of the subset M i .

[0062] Through this formula, the variance-weighted average value of y for all subsets of each parameter can be calculated. j For an influential parameter, after fixing its value, the variance of the performance value can be significantly reduced. Therefore, the larger the ratio of the variance between the total sample and the weighted average value, the more important the parameter is. In this way, the importance of each parameter in the initial parameter space can be accurately evaluated.

[0063] At this time, m important parameters can be selected from the original complete parameter space to form a new parameter space, denoted as P = {p 1 , p 2 , p 3 , ……, p m}, where p i is the i-th parameter in this parameter space.

[0064] S3. Using {important parameter configuration, IOPS} as an individual, multiple individuals form a population. Taking the improvement degree of the IOPS in the current individual compared to the IOPS of the storage system under the default configuration as the fitness function of the current individual, execute the genetic algorithm to optimize the configuration of the parameters in the new parameter space, and obtain the individuals output by the genetic algorithm to form a genetic sample set.

[0065] Specifically, in order to apply the genetic algorithm (abbreviated as GA) to the parameter tuning of the distributed storage system Ceph, it is necessary to re-model the genetic algorithm for the distributed storage system. Based on the design principles of the genetic algorithm, define the sample s i as the binary tuple {P i , IOPS i}, as the m parameter configurations to be adjusted, IOPS i is the IOPS performance index of the system under the parameter configuration P i . Multiple individuals form a population. GA will calculate the fitness of each individual through the fitness function and generate new individuals through selection, crossover, and mutation strategies. This process will be iterated until the maximum iteration number or a satisfactory fitness level is reached.

[0066] The following will introduce the specific process of applying the genetic algorithm to this scenario.

[0067] S31. Obtain the initial population POP = {s1, s2, s3,..., s g}, s i is the i-th individual in the population, g is the number of individuals in the population, s i is the binary tuple {P i , IOPS i}, P i is the parameter configuration of the individual si The configuration found for m important parameters, IOPS i is the IOPS performance metric of the system under parameter configuration P i below.

[0068] S32. Calculate the fitness of each individual in the population, calculate the probability of each individual being selected according to the fitness, select the parental generation according to the probability, and perform hybridization and mutation operations based on the selected parental generation to generate offspring and obtain a new population.

[0069] Specifically, the fitness function can be expressed in the following form:

[0070]

[0071] In the formula, f(s i ) is the fitness of individual s i , IOPS cur and IOPS0 respectively represent the IOPS in the current individual s i and the IOPS of the storage system under the default configuration.

[0072] Subsequently, based on the obtained fitness values, calculate the probability of each individual being selected as the parental generation. The genetic algorithm selects individuals according to the principle of survival of the fittest, and individuals with higher fitness will have a higher probability of being selected.

[0073] The probability γ i that the sample s i is selected can be expressed in the following form:

[0074]

[0075] Subsequently, based on the obtained probabilities, select the parental generation for hybridization. To generate a new individual, the genetic algorithm will select the parental generation for hybridization with a selection strategy. Taking individuals s i , s j as the parental generation as an example, let represent a subset of P i with a cardinality of a, where Finally, define the new individual i generated by the hybridization of s j , s

[0076] Subsequently, perform the mutation operation. To widely explore the search space, the genetic algorithm changes each element of P k in s k with a certain probability to generate a new individual s ′ k , as the generated offspring.

[0077] Subsequently, after obtaining the offspring, a new population is formed.

[0078] It should be noted that for each individual s i , its parameter configuration only includes optimizing the configuration of m important parameters in the new parameter space. When testing the measurement system, the remaining non-critical parameter configurations can be randomly generated or directly configured as default values.

[0079] S33. Repeatedly execute S32 until the algorithm termination condition is reached, and output the final population.

[0080] Through the above genetic algorithm, a series of relatively optimal individuals can be obtained, which are used as new samples to form a genetic sample set and enter the subsequent operations.

[0081] S4. Taking the configuration of important parameters as the action and the IOPS of the storage system after the action as the state, first obtain the initial network parameters of the deep deterministic policy gradient algorithm model based on the genetic sample set, and then explore the configuration of important parameters in the new parameter space based on the deterministic policy gradient algorithm model, and output the final key parameters.

[0082] Specifically, in order to apply the deep deterministic policy gradient algorithm (hereinafter referred to as DDPG) to the parameter tuning of the distributed storage system, the present invention re-models DDPG according to the characteristics of the distributed storage system parameter tuning task, and defines the state, action, and reward to adapt to this parameter tuning task.

[0083] Action (actton): An action is a decision made by the agent in a given state. The choice of action will affect the interaction between the agent and the environment and cause a change in the state.

[0084] For the distributed storage system parameter tuning task, let a set of parameter configurations P t be the important parameter configuration obtained after the agent executes the action at time t, m represents the number of parameters to be tuned, and define the action a executed at time t t = P t , and the action space is the adjustable range space of the corresponding parameters of the distributed storage system.

[0085] State: A state is a description of the environment, which contains all relevant information of the environment where the agent is located, enabling the agent to make decisions based on this information.

[0086] When the environment is a distributed storage system, define the state s at time t t = {P t , IOPSt}, where TOPS t represents the parameter configuration P t is the performance value obtained after stress testing.

[0087] Reward: Reward is the immediate feedback obtained by the agent from the environment after performing an action. It is an evaluation of the quality of the agent's behavior and is usually a calculated numerical value.

[0088] The reward function is crucial in reinforcement learning and determines the learning direction of the network. To make DDPG suitable for parameter tuning of distributed storage systems, this embodiment hopes that the reward can simulate the tuning process of professional distributed storage system engineers in a real environment. At the same time, it is noted that the parameter tuning task of the distributed storage system not only hopes that the performance value after parameter tuning is better than the default parameter configuration, but also hopes that the performance value after parameter tuning can be as high as possible.

[0089] In summary, this embodiment defines the reward function r of DDPG at time t - 1 as follows t-1 :

[0090]

[0091] In the formula, Δ t->0 represents the degree of improvement of the IOPS in the state of the agent at time t compared to the IOPS of the system under the default configuration, and Δ t->t-1 represents the degree of improvement of the IOPS in the state of the agent at time t compared to the IOPS in the state of the agent at time t - 1. Among them, the state of the agent at time t is the state obtained after the agent performs an action at time t - 1.

[0092] Δ t->0 and, Δ t->t-1 The calculation formulas can be expressed in the following form:

[0093]

[0094] In the formula, IOPS t is the IOPS in the state of the agent at time t, and IOPS t-1 is the IOPS in the state of the agent at time t - 1, and IOPS0 is the IOPS of the system under the default configuration.

[0095] In this embodiment, the reward function comprehensively considers the difference Δ t->0 between the performance value of this tuning and the default performance value and the difference Δ t->t-1 between the performance value of this tuning and the performance value of the previous tuning. If Δ t->0 is positive, the tuning trend is correct and the reward is positive; otherwise, the tuning trend is incorrect and the reward is negative. Δ t->0Determines the direction of the reward, while Δ t->t-1 To a certain extent, it affects the magnitude of the reward. Based on the reward function set above, the parameter tuning of the model can be directed in the correct direction, and at the same time, the optimal parameter configuration can be quickly searched for.

[0096] In addition, the DDPG algorithm also involves the setting of two networks, namely the value network (Critic) and the policy network (Actor).

[0097] Value network (Critic): After defining the state s, action a, and reward function r, DDPG uses the deep neural network Critic to approximately estimate the value Q(s,a) of each action in each state. Q(s,a) is defined by the following formula:

[0098]

[0099] In the formula, θ t Is the parameter of the deep neural network at time t, s t+1 Represents the state at time t+1, μ(s t+1 ,θ t ) Represents the action output by the policy network μ when the state is s t+1 , and the parameter is θ t .

[0100] Q μ (s t+1 ,μ(s t+1 ,θ t )) Represents the action value in the state s t+1 , and the action is μ(s t+1 ,θ t ). γ is the attenuation factor, indicating the importance of future rewards relative to current rewards.

[0101] Let the parameter of Critic be w. Critic defines the action value under the policy μ as Q μ (s,a|w). When randomly sampling a sample (s t ,a t ,r t ,s t+1 ) from the experience replay pool, DDPG uses Q-learning to minimize the training objective:

[0102] min(L(w t )=E[(Q(s t ,a t |w t )-y(s,a|w)) 2 );

[0103]

[0104] Among them, A(s t+1 ) represents the entire action space in state s t+1 , and γ is the decay factor.

[0105] To minimize the error function, the Critic uses stochastic gradient ascent for value update:

[0106]

[0107] Among them, α w is a coefficient factor less than 1, and δ t is the TD error:

[0108] δ t = r t+1 + γQ[(s t+1 , μ(s t+1 , θ t ), w t ) - Q(s t , a t , w t )]

[0109] Policy network (Actor): DDPG uses a deep neural network Actor to output actions. Specifically, the Actor uses the action value estimated by the Critic for stochastic gradient ascent for policy update:

[0110]

[0111] Among them, α θ is a coefficient factor less than 1.

[0112] As Figure 3 shown is the flowchart of the DDPG algorithm in an embodiment of the present invention.

[0113] Since at the initial stage of the DDPG algorithm, the number of samples in the experience replay pool is small and the quality is poor, and the model parameters are still random initial parameters, the above situation will cause DDPG to be difficult to obtain high-quality samples in the initial stage of training, and the parameters are difficult to update in the correct direction, that is, the cold start problem will occur

[0114] To solve the above cold start problem, when obtaining the initial network parameters of the deep deterministic policy gradient algorithm model based on the genetic sample set in this step, genetic samples are used. The genetic samples carry IOPS information, so there is no need to perform stress testing on the Ceph system, and the samples are directly used to update the DDPG parameters, which is called fast training.

[0115] First, after supplementing the corresponding reward r in each genetic sample, it is placed in the experience replay pool, and then fast training is performed.

[0116] The process of rapid training includes the following steps:

[0117] S401. Initialize the Actor and Critic networks and their target networks, and fix the parameters of the target networks;

[0118] S402. Randomly sample a batch of samples from the experience replay pool in a uniform sampling manner and feed them into the Critic. Update the parameters of the main value network of the Critic according to the formula; then the Critic outputs a batch of Q(s,a,w main ) to the Actor, and update the parameters of the main policy network of the Actor according to the formula; w main is the parameter w of the main value network main ;

[0119] S403. Overwrite the parameters of the main networks of the Critic and Actor with those of their target networks;

[0120] S404. Continuously repeat steps S402 - S403 several times (set a number of rapid training times) to rapidly update the network parameters of DDPG.

[0121] In the present invention, first, the DDPG model is rapidly trained through a genetic sample set to overcome the cold start problem in the early stage of the DDPG algorithm, and then it enters the normal exploration process of DDPG to explore the configuration of m important parameters in the new parameter space.

[0122] The normal exploration process of DDPG includes the following steps:

[0123] S411. Fix the parameters of the target network;

[0124] S412. Obtain the initial state s0 by stress testing the default parameter configuration. The initial Actor outputs the action a0, and then the distributed storage system performs stress testing to obtain the next - moment state s1. Calculate the reward r0 based on s0 and s1, and save the generated sample (s0,a0,r0,s1) to the experience replay pool;

[0125] S413. Randomly sample a batch of samples from the experience replay pool in a uniform sampling manner and feed them into the Critic. Update the parameters of the main value network of the Critic according to the formula; then the Critic outputs a batch of Q(s,a,w main ) to the Actor, and update the parameters of the main policy network of the Actor according to the formula;

[0126] S414. Overwrite the parameters of the main networks of the Critic and Actor with those of their target networks;

[0127] S415. Use the recommended parameter a1 in the target network output state s1 of the Actor;

[0128] S416. Apply the parameter configuration a1 to the distributed storage system environment and perform a stress test to obtain the state s2. Calculate the reward r1 based on s0, s1, and s2, and save the generated sample (s1, a1, r1, s2) to the experience replay pool;

[0129] S417. Continuously repeat steps S412 to S416 until the DDPG searches for the optimal parameters and the tuning ends.

[0130] Combining the above steps S1 to S4 can achieve automatic and accurate parameter tuning for the storage system.

[0131] Generally speaking, the storage system parameter tuning method provided by the present invention combines the genetic algorithm and the deep deterministic policy gradient algorithm to achieve automatic tuning of the parameters of the distributed storage system. Compared with manual tuning, it can effectively improve the tuning efficiency. Moreover, both the genetic algorithm and the deep deterministic policy gradient algorithm are optimization algorithms. In the present invention, the IOPS performance index of the storage system is used as the optimization target to tune the parameters of the storage system. Through algorithm iteration, the parameter configuration that makes the IOPS performance index of the storage system better can be found, realizing accurate tuning of the parameters of the distributed storage system. In the present invention, random samples are first obtained through a random policy, and then the parameter importance is evaluated to obtain important parameters. The important parameters are the key parameters affecting the system performance. Therefore, when the genetic algorithm and the deterministic policy gradient algorithm are subsequently executed, the screened important parameters are explored, and the parameter space is shrunk, thereby improving the optimization speed and ensuring the overall configuration effect. Moreover, before executing the deep deterministic policy gradient algorithm, the genetic algorithm is first executed to obtain a genetic sample set, and the deep deterministic policy gradient algorithm model is quickly trained based on the genetic sample set, which can avoid the cold start problem of the deep deterministic policy gradient algorithm and enable the deep deterministic policy gradient algorithm to quickly enter the effective exploration stage, thereby quickly outputting the parameter configuration result.

[0132] Embodiment 2

[0133] The present invention also relates to a storage system parameter tuning device based on reinforcement learning. As Figure 4 shown in the structural block diagram of the storage system parameter tuning device in an embodiment of the present invention, it includes:

[0134] A random sample acquisition unit, configured to randomly configure the global parameters in the initial parameter space and construct a random sample set S, where the samples are {global parameter configuration, IOPS}, and the IOPS in the samples is the IOPS performance index of the storage system under the corresponding parameter configuration;

[0135] A parameter importance selection unit, which is used to rank the importance of each parameter in the initial parameter space based on the random sample set S, and obtain the top m important parameters that have the greatest impact on the performance metric IOPS, thereby constituting a new parameter space.

[0136] A genetic algorithm execution unit, which uses {important parameter configuration, IOPS} as an individual, and multiple individuals form a population. The improvement degree of the IOPS in the current individual compared to the IOPS of the storage system under the default configuration is used as the fitness function of the current individual. It executes the genetic algorithm to optimize the configuration of the parameters in the new parameter space, and obtains the individuals output by the genetic algorithm to form a genetic sample set.

[0137] A deep deterministic policy gradient algorithm execution unit, which uses the configuration of important parameters as an action and the IOPS of the storage system after the action as a state. First, it obtains the initial network parameters of the deep deterministic policy gradient algorithm model based on the genetic sample set, and then explores the configuration of the important parameters in the new parameter space based on the deterministic policy gradient algorithm model to output the final key parameters.

[0138] Specifically, the storage system parameter tuning device involves the interaction of five parts: the environment, the configuration recommendation model, the configuration, the sample pool, and the parameter importance ranking.

[0139] The environment consists of a distributed file system, a workload generator, and performance monitoring. The environment module receives the parameter configuration recommended by the configuration recommendation model. After setting the parameter configuration to the system, the workload generator will generate specific IO workloads that conform to the scenario and make the system run this workload. The performance monitoring will obtain and record the performance metrics of the system in real time.

[0140] The parameter recommendation model includes a total of three models: the random selection model, the genetic algorithm (GA) model, and the deep deterministic policy gradient algorithm (DDPG) model.

[0141] The random selection model recommends parameter configurations through a random policy and is used to obtain random test samples in the early stage.

[0142] The GA model explores the new parameter space through an evolutionary algorithm, recommends parameter configurations, and simultaneously collects samples during the exploration process as the sample set for the warm start of the DDPG.

[0143] The DDPG model first overcomes the cold start problem in the early stage through the sample set collected by the GA model, and then uses the reinforcement learning mechanism to explore the new parameter space, and finally generates the recommended parameter configuration.

[0144] The configuration corresponds to two parameter spaces, namely the initial parameter space with global parameters and the new parameter space composed of important parameters.

[0145] The parameter space includes parameter names, parameter types, maximum values of parameters, and minimum values of parameters.

[0146] The sample pool is divided into three parts: a random sample pool, a GA sample pool, and a DDPG sample pool.

[0147] The random sample pool is generated by a random selection model and is used to perform parameter importance ranking to generate a new parameter space.

[0148] The GA sample pool is generated by a GA model and is used to warm-start the DDPG model.

[0149] The DDPG sample pool is generated by a DDPG model and is used for the update training of the DDPG model.

[0150] Parameter importance ranking is to perform importance ranking on the parameters of the initial parameter space by analyzing the variances of the random test samples generated by the random selection model, so as to generate a new parameter space.

[0151] First, the random sample acquisition unit calls the random selection model in the configuration recommendation model to perform random parameter configuration, and transmits the parameter configuration to the environment to generate the performance metric IOPS of the storage system under the random configuration, thereby constructing a random sample set and placing it in the random sample pool in the sample pool;

[0152] Subsequently, the parameter importance selection unit starts parameter importance ranking, ranks the parameters based on the random sample set in the random sample pool, selects the top m important parameters with the greatest influence, and constitutes a new parameter space;

[0153] Subsequently, the genetic algorithm execution unit calls the GA model, explores the new parameter space through the genetic algorithm, recommends parameter configurations, and at the same time collects the genetic sample set output during the exploration process and places it in the GA sample pool in the sample pool;

[0154] Subsequently, the deep deterministic policy gradient algorithm execution unit calls the DDPG model, first performs fast training through the samples in the GA sample pool to warm-start the DDPG model, then explores the new parameter space using the reinforcement learning mechanism, and places the samples generated during the exploration in the DDPG sample pool. After the exploration ends, the final parameter configuration of the storage system is obtained.

[0155] Embodiment 3

[0156] The present invention also relates to a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0157] Specifically, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0158] Embodiment 4

[0159] An embodiment of the present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the method in the above embodiments of the present invention.

[0160] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification. It should be noted that the "in one embodiment", "for example", "again, for example", etc. in the present invention are intended to illustrate the present invention, rather than to limit the present invention.

[0161] The above embodiments only represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but it should not be understood as a limitation on the scope of the patent application. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A method for optimizing storage system parameters based on reinforcement learning, characterized in that Including: S1. Randomly configure the global parameters in the initial parameter space and construct a random sample set S, where the samples are {global parameter configuration, IOPS}, and the IOPS in the samples is the IOPS performance metric of the storage system under the corresponding parameter configuration; S2. Based on the random sample set S, perform importance ranking on each parameter in the initial parameter space and obtain the top m important parameters that have the greatest impact on the performance metric IOPS, thus forming a new parameter space; S3. Use {important parameter configuration, IOPS} as an individual, and multiple individuals form a population. Take the improvement degree of the IOPS in the current individual compared to the IOPS of the storage system under the default configuration as the fitness function of the current individual, execute the genetic algorithm to optimize the configuration of the parameters in the new parameter space, and obtain the individuals output by the genetic algorithm to form a genetic sample set; S4. Take the configuration of the important parameters as an action and the IOPS of the storage system after the action as the state. First, obtain the initial network parameters of the deep deterministic policy gradient algorithm model based on the genetic sample set, and then explore the configuration of the important parameters in the new parameter space based on the deterministic policy gradient algorithm model to output the final key parameters; Among them, the importance ranking of each parameter in the initial parameter space includes: calculating the importance degree of each parameter and performing importance ranking based on the importance degree; the calculation process of calculating the importance degree of any parameter p includes: Calculating the performance variance Var(IOPS) of the random sample set S; Dividing the values of parameter p in the random sample set S into intervals to form multiple subsets, calculating the variance weighted average of all subsets of this parameter p, and calculating the ratio of the performance variance Var(IOPS) to the variance weighted average of this parameter p as the importance degree I(p) of this parameter p.

2. The method for tuning storage system parameters based on reinforcement learning according to claim 1, wherein The calculation formula of the fitness function of the current individual is: where f(s i ) is the fitness of individual s i , IOPS cur and IOPS0 represent the IOPS in the current individual s i and the IOPS of the storage system under the default configuration, respectively.

3. The method for optimizing storage system parameters based on reinforcement learning according to claim 1, wherein The reward function set in the deep deterministic policy gradient algorithm model is: where r t-1 represents the reward obtained from the action performed at time t-1, and Δ t->0 represents the degree of improvement in IOPS in the state of the agent at time t compared to the IOPS of the system under the default configuration, and Δ t->t-1 represents the degree of improvement in IOPS in the state of the agent at time t compared to the IOPS in the state of the agent at time t-1.

4. The method for tuning storage system parameters based on reinforcement learning according to claim 3, wherein Degree of improvement Δ t->0 and Δ t->t-1 The calculation formula is as follows: where, IOPS t is the IOPS in the state of the agent at time t, and IOPS t-1 is the IOPS in the state of the agent at time t - 1, and IOPS0 is the IOPS of the system under the default configuration.

5. The method for tuning storage system parameters based on reinforcement learning according to claim 1, wherein The process of first obtaining the initial network parameters of the deep deterministic policy gradient algorithm model based on the genetic sample set includes: S401. Place the genetic sample set in the experience replay pool, initialize the network parameters of the deep deterministic policy gradient algorithm model, and fix its target network parameters; S402. According to the uniform sampling method, randomly sample a batch of samples from the experience replay pool, and update the parameters of the main value network and the main policy network based on the sampled samples; S403. Overwrite the parameters of the main value network to the target value network, and overwrite the parameters of the main policy network to the target policy network; S404. Repeat steps S402 - S403 until the preset number of training times is reached.

6. The method for tuning storage system parameters based on reinforcement learning according to claim 1, wherein In S2 and S3, the unimportant parameters of the storage system are randomly configured or default configured.

7. A storage system parameter tuning device based on reinforcement learning, characterized in that, Including: A random sample acquisition unit, used to randomly configure the global parameters in the initial parameter space and construct a random sample set S, where the samples are {global parameter configuration, IOPS}, and the IOPS in the samples is the IOPS performance metric of the storage system under the corresponding parameter configuration; A parameter importance selection unit, configured to perform importance ranking on each parameter in the initial parameter space based on the random sample set S, and obtain the top m important parameters that have the greatest impact on the performance metric IOPS, thereby constituting a new parameter space; A genetic algorithm execution unit, configured to use {important parameter configuration, IOPS} as an individual, and multiple individuals form a population. Using the improvement degree of the IOPS in the current individual compared to the IOPS of the storage system under the default configuration as the fitness function of the current individual, execute the genetic algorithm to optimize the configuration of the parameters in the new parameter space, and obtain the individuals output by the genetic algorithm to form a genetic sample set; A deep deterministic policy gradient algorithm execution unit, configured to use the configuration of important parameters as an action and the IOPS of the storage system after performing the action as a state. First, obtain the initial network parameters of the deep deterministic policy gradient algorithm model based on the genetic sample set, and then explore the configuration of the important parameters in the new parameter space based on the deterministic policy gradient algorithm model, and output the final key parameters.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.