Distributed storage system parameter tuning method based on experience parallel reinforcement learning
Through the method of parallel reinforcement learning based on experience, the distributed storage system parameters are automatically tuned, which solves the problems of low sampling efficiency and long cold start time in the existing technology, and achieves efficient parameter tuning and performance improvement.
Patent Information
- Application Number
- CN202510462252.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art parameter tuning method in distributed storage systems has problems such as low sampling efficiency, long cold start time, poor sample set uniformity and high resource consumption, resulting in low performance optimization efficiency.
The distributed storage system parameter tuning method based on experience parallel reinforcement learning is adopted, and the load is generated through the environment module. The parameter preprocessor filters the adjustable parameters, the important parameter filter collects samples, the system controller obtains performance indicators, the tuning agent performs reinforcement learning, and iteratively adjusts the parameter strategy to achieve automatic parameter tuning.
It significantly improves the performance of the storage cluster, reduces the time required to select the optimal parameter value, avoids system failure, alleviates cold start problems, and improves sampling efficiency.
Smart Images

Figure CN120295583A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of distributed storage, and particularly relates to a method for optimizing distributed storage system parameters based on experience parallel reinforcement learning. Background Art
[0002] A distributed storage system has a large number of parameters that can be adjusted by users themselves. However, the default values of these parameters often cannot enable platforms such as domesticated platforms to obtain optimal performance. Therefore, parameter tuning is a fast, effective, and economical way to optimize performance. Automatic parameter tuning is an optimization of the software layer of a distributed storage system. Under limited resource conditions, it automatically determines the optimal parameter value combination for the system under a given workload to maximize the performance optimization goal.
[0003] The automatic parameter tuning method of a distributed storage system usually consists of two sub-steps: important parameter screening and refined search for parameter values. Important parameter screening refers to sorting all parameters according to the degree of influence of each parameter on the performance optimization goal, reducing the difficulty of searching for the optimal parameter values; refined search for parameter values refers to searching for the parameter value combination that optimizes the performance optimization goal from the value space of important parameters. Chinese invention patent CN114489499A discloses an automatic parameter tuning system based on a performance simulation model. First, important parameters are screened based on the service scenario, then a simulation model of parameter values and performance is constructed by the SVR algorithm, and finally, the extreme point of the performance model is calculated in combination with the quasi-Newton algorithm to obtain the optimal parameter values. Chinese invention patent CN116088761A discloses an automatic parameter tuning method based on reinforcement learning. Under a specific workload, important parameters with the greatest impact on the performance of the distributed storage system are screened through sensitivity analysis based on performance variance, and the DDPG algorithm of reinforcement learning is used to further refine the search for parameter values. In addition, to alleviate the "cold start" time with poor effects in the early stage of DDPG algorithm tuning, the genetic algorithm is used to collect "parameter value - performance" samples in the early stage to reduce the overall tuning time.
[0004] In terms of important parameter screening, existing technologies all rely on a certain-scale labeled tuning experience set. Compared with the random sampling method, Latin hypercube sampling is a commonly used stratified sampling algorithm. Its basic idea is: first, split the value range of each parameter into N S equal-sized value intervals, and then collect N SSamples are selected such that each value range of each parameter contains only one sample. The sample sets collected by such algorithms are more evenly distributed and have stronger representativeness. In addition, since there is no performance emulator, the tuning samples can only be obtained through actual stress testing on the deployed distributed storage cluster. Some samples that do not satisfy the parameter constraint relationships will cause system failures after being actually applied to the storage cluster. At this time, the cluster can only be deleted and recreated, and the time required to rebuild the cluster is as long as 10 to 20 minutes, which is obviously unacceptable. However, when the existing constrained Latin hypercube sampling design is extended to high-dimensional spaces, there are problems of poor uniformity of the sample set and low sampling efficiency.
[0005] In terms of parameter configuration search, for the technology based on the performance simulation model, a simulation model with a relatively high prediction accuracy needs to be constructed. Therefore, a large number of tuning samples are required, and it is also required that the sample set covers all the performance inflection points of the system. However, the time cost of obtaining the tuning samples is extremely high. In addition, users also need to have prior knowledge of the storage system, cluster environment, workload, and simulation model. For the parameter value search based on reinforcement learning, although the final tuning effect is better than the former, there is generally a "cold start" time, that is, there is no performance improvement observed for up to 12 to 24 hours in the early stage of tuning. The main reason is that there are problems of low sampling efficiency and temporal dependence in the single-process serial training of the reinforcement learning algorithm. Summary of the Invention
[0006] (1) Technical Problems to be Solved
[0007] The technical problem to be solved by the present invention is to provide an automatic parameter tuning method for a distributed storage system, which is used to optimize the parameters at the software level of the distributed storage system under the condition that the preset hardware resource conditions and operating environment configurations are determined, and while significantly improving the performance of the storage cluster, reducing the time required to recommend the optimal parameter values.
[0008] (2) Technical Solutions
[0009] To solve the above technical problems, the present invention provides a distributed storage system parameter tuning system based on empirical parallel reinforcement learning. The system includes five components: an environment module, a parameter preprocessor, an important parameter filter, a system controller, and a tuning agent;
[0010] The environment module includes the distributed storage system and its workload, and the load generator simulates and generates the specific workload of the distributed storage system;
[0011] The input of the parameter preprocessor is the initial parameter information of the distributed storage system, and it can output information that can be automatically adjusted. The value ranges of these automatically adjustable parameters form the value space of the adjustable parameters;
[0012] The important parameter filter is used to take the value space of adjustable parameters as the sampling space, and collect a parameter value sample set. The system controller is used to apply the parameter value samples to the distributed storage system one by one to generate a performance metric sample set, and then combine them into a training data set {(parameter value, performance metric)} of the important parameter filter; the important parameter filter is also used to use the training data set to screen out the important parameters that significantly affect the distributed storage system, and the value ranges of these important parameters are used as the value space for subsequent fine-grained tuning of parameter values.
[0013] The system controller is used to update the parameter values of the distributed storage system by using the parameter update interface provided by the distributed storage system. After the parameter value update is completed, start the performance test and obtain the performance metric vector, and then use the status monitoring framework provided by the distributed storage system to obtain the status metric vector of the distributed storage system after the performance test is completed. Combine the parameter values, performance metric vector, status metric vector, and the state transition reward of the distributed storage system into tuning experience, and save it in the tuning experience cache pool of the tuning agent. The system controllers of each tuning agent are independent of each other.
[0014] The tuning agent is used to perform fine-grained tuning on the values of important parameters. During the tuning process, through online interaction with the environment module, iteratively adjust its parameter tuning strategy until stable performance feedback of the distributed storage system is obtained through the system controller or the tuning resource limit condition is reached, and then end the tuning process. Take the parameter value combination corresponding to the optimal performance during the tuning process as the optimal parameter value.
[0015] The present invention also provides a method for tuning parameters of a distributed storage system based on experience parallel reinforcement learning implemented based on the above system. The method includes the following steps:
[0016] (1) Determine the optimization goal. In the distributed storage system, the read-write performance metric indicators include throughput, latency, and IOPS. The optimization goal selects a single performance metric, or quantifies multiple performance metrics into a single integrated metric
[0017] (2) Obtain all the parameter information exposed by the distributed storage system that can be adjusted by users from the deployed distributed storage system. The parameter preprocessor screens the parameters that can be automatically adjusted and further processes these parameters, converting boolean-type and enumerated string-type parameters into numerical types. For integer and floating-point type parameters without specified value ranges, set their value ranges to [0.1×default(K i ), 10×default(K i ), where represents the parameter K iThe default value, where \(i\in[1,N]\) and \(N\) is a positive integer;
[0018] (3) The important parameter filter uses the value range of the parameters that can be automatically adjusted as the sampling space. Using the sampling algorithm, a parameter value sample set \(A = \{A_1,\ldots,A_M\}\) containing \(M\) samples is collected, where, M \(A_j=(Action_{j1},\ldots,Action_{jn})\) represents a combination of parameter values, t where \(Action_{jt}\) represents the Min - Max normalized value of the \(t\)-th parameter \(K_t\) value, where \(t = 1,2,\ldots,M\); N represents the parameter \(K_t\) i value's Min - Max normalized value, where \(t = 1,2,\ldots,M\);
[0019] (4) Through the system controller, the parameter values in the parameter value sample set \(A\) are applied to the distributed storage system one by one. Using the load generation tool, the performance of the distributed storage system is tested for a certain period of time to obtain the performance metric vector \(P=(Perf_1,\ldots,Perf_u)\), where the elements in the performance metric vector are performance metrics. Here, \(u\) represents the number of nodes in the distributed storage system. At the same time, the system controller uses the status monitoring framework to obtain the status metric vector \(S=(State_1,\ldots,State_v)\), where \(v\) represents the number of status metric values obtained through the status monitoring framework of the distributed storage system. The parameter values, performance metric vector, and status metric vector are combined to form the tuning experience sample set \(\{(A_j,P_j,S_j)\}\), where \(P_j\) and \(S_j\) respectively represent the performance metric vector and status metric vector after the distributed storage system applies the parameter value \(A_j\); t+1 \(P=(Perf_1,\ldots,Perf_u)\), u where the elements in the performance metric vector are performance metrics. Here, \(u\) represents the number of nodes in the distributed storage system. At the same time, the system controller uses the status monitoring framework to obtain the status metric vector \(S=(State_1,\ldots,State_v)\), t+1 where \(v\) represents the number of status metric values obtained through the status monitoring framework of the distributed storage system. The parameter values, performance metric vector, and status metric vector are combined to form the tuning experience sample set \(\{(A_j,P_j,S_j)\}\), v where \(P_j\) and \(S_j\) respectively represent the performance metric vector and status metric vector after the distributed storage system applies the parameter value \(A_j\); t ,P t+1 ,S t+1 )}, where, t+1 P t+1 and t S
[0020] (5) The important parameter filter uses the subset \(\{(A_j,P_j)\}\) of the tuning experience sample set to screen out the important parameters \(K_1',K_2',\ldots,K_m'\) that have a significant impact on the performance of the storage system, and forms the value range \(A'\) of the important parameters; t ,P t+1 )} to screen out the important parameters \(K_1',K_2',\ldots,K_m'\) that have a significant impact on the performance of the storage system, n and forms the value range \(A'\) of the important parameters;
[0021] (6) The tuning agent tunes the values of important parameters through reinforcement learning, which is implemented based on a neural network. The tuning agent uses the value space A' of the important parameters as the action space of the reinforcement learning algorithm, and obtains the state metrics S' representing the operating state of the distributed storage system from the state monitoring framework of the distributed storage system. The value range of these state metrics is used as the state space of the tuning algorithm. The tuning agent and the environment module continuously interact to generate tuning experiences (S t , A t , S t+1 , R t+1 ), where S t represents the state metric vector of the distributed storage system before the parameter value is updated to A t , and R t+1 is related to the difference ΔP t+1 between the performance metric vectors before and after the parameter value update, representing the state transition reward of the distributed storage system. After the number of tuning experiences in the tuning experience cache pool of the tuning agent reaches a preset threshold, the tuning agent collects a batch of tuning experiences {(S t , A t , S t+1 , R t+1 )} to update the neural network hyperparameters. Therefore, the tuning agent will continuously iterate and adjust its parameter tuning strategy until it obtains stable performance feedback of the distributed storage system through the system controller or reaches the preset tuning duration, at which point the training process ends, and the parameter value combination corresponding to the optimal performance during the training process is used as the optimal parameter value.
[0022] (III) Advantageous Effects
[0023] The present invention provides an automatic parameter tuning method for a distributed storage system based on experience parallel reinforcement learning, which is applicable to optimizing the software-level parameters of the distributed storage system under the condition that the preset hardware resource conditions and operating environment configurations are determined. This method realizes intelligent parameter tuning through the reinforcement learning algorithm and automatically completes all parameter processing work in the parameter tuning process, achieving a large improvement in read and write performance at a low cost. While significantly improving the performance of the storage cluster, it reduces the time required to recommend the optimal parameter values. Among them, the proposed SA-HCLHS algorithm makes the collected parameter value sample set satisfy all constraint relationships between parameters, avoiding frequent failures of the distributed storage system when collecting corresponding performance samples. Further, for the automatic parameter tuning task of the distributed storage system based on reinforcement learning under a fixed load, the EPDT technology based on experience parallel is proposed, effectively alleviating the cold start problem of the reinforcement learning algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is the design principle block diagram of the method of the present invention;
[0025] Figure 2 This is the flowchart of the automatic parameter tuning method of the present invention;
[0026] Figure 3 This is the schematic diagram of the dual-process EPDT;
[0027] Figure 4 This is the parallel architecture diagram of reinforcement learning;
[0028] Figure 5 It shows the recommended duration of the optimal parameter configuration for different combinations of Workers;
[0029] Figure 6 This is the tuning effect diagram of EPDT-C. Detailed implementation manners
[0030] To make the objectives, contents, and advantages of the present invention clearer, the following further describes the detailed implementation manners of the present invention in conjunction with the accompanying drawings and embodiments.
[0031] Aiming at the defects of the prior art and the improvement requirements, the present invention provides an automatic parameter tuning method for a distributed storage system based on experience parallel reinforcement learning, which is applicable to optimizing the parameters at the software level of the distributed storage system under the condition that the preset hardware resource conditions (such as CPU, memory, network bandwidth, storage devices, etc.) and the operating environment configuration (such as the operating system and its parameters, middleware, network topology, etc.) are determined, while significantly improving the performance of the storage cluster and reducing the time required to recommend the optimal parameter values.
[0032] To achieve the above objectives, in the first aspect, the present invention provides an automatic parameter tuning system for a distributed storage system based on multi-process reinforcement learning. Figure 1 This is the framework design diagram of the system, including five components: an environment module, a parameter preprocessor, an important parameter filter, a system controller, and a tuning agent.
[0033] The environment module includes a distributed storage cluster (i.e., the distributed storage system) and its workload, and a load generator simulates and generates the specific workload of the distributed storage cluster.
[0034] The input of the parameter preprocessor is the initial parameter information of the distributed storage system, and the output is the information of the parameters that can be automatically adjusted. The value ranges of these parameters that can be automatically adjusted (tunable parameters) form the value space of the tunable parameters.
[0035] The important parameter filter uses the value space of the adjustable parameters as the sampling space to collect a parameter value sample set. The system controller applies the parameter value samples to the distributed storage cluster one by one to generate a performance metric sample set, and then combines them into the training data set {(parameter value, performance metric)} of the important parameter filter. The important parameter filter uses the training data set to screen out the important parameters that significantly affect the distributed storage system, and the value ranges of these important parameters are used as the value space for the subsequent refined search of parameter values..
[0036] The system controller updates the parameter values of the distributed storage cluster using the parameter update interface provided by the distributed storage system. After the parameter value update is completed, it starts a performance test and obtains a performance metric vector, and then uses the status monitoring framework provided by the distributed storage system to obtain the status metric vector of the distributed storage cluster after the performance test is completed. The parameter value, performance metric vector, status metric vector, and the status transition reward of the distributed storage cluster are combined into tuning experience and saved in the tuning experience cache pool of the tuning agent. The system controllers of each tuning agent are independent of each other.
[0037] The tuning agent performs fine-grained tuning on the values of the important parameters. During the tuning process, it iteratively adjusts its automatic parameter tuning strategy through online interaction with the environment module until it obtains stable performance feedback of the distributed storage cluster through the system controller or reaches the tuning resource limit condition (preset tuning duration), at which point the training process ends, and the parameter value combination corresponding to the optimal performance during the training process is used as the optimal parameter value.
[0038] Since the system controller shields the differences between the underlying storage systems, the present invention is applicable to the automatic parameter tuning of various distributed storage systems. Specifically, the tuning agent takes the status metric values and the rewards of the status transition of the distributed storage system as inputs, which are obtained and converted by the system controller from the status monitoring framework of the underlying distributed storage system. Therefore, as long as a distributed storage system provides a parameter value update and a system feedback acquisition interface, it is applicable to the automatic parameter tuning method proposed by the present invention. Currently, mainstream distributed storage systems all support the status monitoring framework, including: GlusterFS, HDFS, Ceph, Lustre, Swift, etc.
[0039] Figure 2 is the workflow diagram of the technical solution of the present invention, mainly including the following steps:
[0040] (1) Determine the optimization goal. In the distributed storage cluster, the main read and write performance metric indicators are throughput, latency, and IOPS. The optimization goal can select a single performance metric, or multiple performance metrics can be quantified into a fusion metric Common linear scalarization functions are used for quantification, that is
[0041] (2) Obtain all the parameter information exposed by the distributed storage system that can be adjusted by users from the deployed distributed storage cluster. The parameter pre-processor filters the parameters that can be automatically adjusted, further processes these parameters, converts the boolean type and enumeration string type parameters into numerical types. For integer and floating-point type parameters without specified value ranges, set their value ranges to [0.1×default(K i ), 10×default(K i ), where represents the default value of parameter K i , and i ∈ [1, N].
[0042] (3) The important parameter filter uses the value space of the parameters that can be automatically adjusted as the sampling space, and uses the sampling algorithm designed by the present invention to collect a parameter value sample set A = {A1,..., A M} containing M samples. Among them, A t = (Action1,..., Action N ) represents a value combination of the parameters. represents the Min-Max normalization value of the value of parameter K i , where t = 1, 2,..., M.
[0043] (4) Through the system controller, apply the parameter values in the parameter value sample set A to the distributed storage cluster one by one, and use a load generation tool (such as Filebench, Vdbench, etc.) to perform a performance test on the distributed storage cluster for a certain period of time to obtain the performance metric vector P t+1 = <Perf1,..., Perf u >. The elements in the performance metric vector are performance metrics. Among them, u represents the number of nodes in the distributed storage cluster. At the same time, the system controller uses the status monitoring framework to obtain the status metric vector S t+1 = <State1,..., State v >, where v represents the number of status metric values obtained through the status monitoring framework of the distributed storage system. Combine the parameter value vector, the performance metric vector, and the status metric vector to form a tuning experience sample set {(A t , P t+1 , S t+1 ), where P t+1 and S t+1 respectively represent the performance metric vector and the status metric vector after the distributed storage cluster applies the parameter value A t .
[0044] (5) The important parameter filter utilizes a subset of the tuned experience sample set \(\{(A t ,P t+1 )\}\) to filter out important parameters \(K1', K2', \ldots, K'\) that have a significant impact on the performance of the storage cluster, n and form the value space \(A'\) of the important parameters.
[0045] (6) The tuning agent tunes the values of the important parameters through reinforcement learning, which is implemented based on a neural network. The tuning agent uses the value space \(A'\) of the important parameters as the action space of the reinforcement learning algorithm. From the state monitoring framework of the distributed storage system, it obtains the state metrics \(S'\) representing the operating state of the distributed storage cluster, and uses the value range of these state metrics as the state space of the tuning algorithm. The tuning agent and the environment module continuously interact to generate tuning experiences \((S t ,A t ,S t+1 ,R t+1 )\), where \(S t \) and \(S t+1 \) respectively represent the state metric vectors of the distributed storage system before and after the parameter values are updated to \(A t \) and \(A t \). \(R t+1 \) is related to the difference \(\Delta P\) between the performance metric vectors before and after the parameter value update t+1 \) and also represents the state transition reward of the distributed storage cluster. After the number of tuning experiences in the tuning experience cache pool of the tuning agent reaches a preset threshold, the tuning agent collects a batch of tuning experiences \((S t ,A t ,S t+1 ,R t+1 )\) to update the neural network hyperparameters; therefore, the tuning agent will continuously iterate and adjust its automatic parameter tuning strategy until it obtains stable performance feedback of the distributed storage cluster through the system controller or reaches the preset tuning duration, at which point the training process ends, and the parameter value combination corresponding to the optimal performance during the training process is used as the optimal parameter value.
[0046] Further, in step (4), the system controller obtains the state metric values through the state monitoring framework of the distributed storage cluster. However, in the distributed storage system, due to differences in data distribution and load conditions among nodes, the same state metric value collected from different nodes will be different. Therefore, it is necessary to determine the value for each state metric. For cumulative state metrics, the calculation method is to sum up the state metric values of each node. For mean state metrics, a linear weighted function is used to fuse these scattered state metric values for calculation.
[0047] Further, the screening function of the important parameter filter in step (5) is implemented based on machine learning. By learning the precise regression model between the independent variables (parameter values) and the dependent variable (performance metric), the influence degree of each parameter value on the performance metric is quantified using the coefficient of each parameter value, and the top K parameters with the greatest influence on the performance of the distributed storage system are selected as important parameters. Using a simple regression model (such as the Lasso linear regression model or the random forest regression model) can reduce the number of tuning experiences (A, P) required for model training. Specifically: First, divide the tuning sample set, perform log1p normalization on the values of each parameter in the sample set, so that the parameter value sample set A has the same order of magnitude, and convert the performance metric set {P} into an optimization target set (set of integrated metrics) as the labeled data for model training. Then, after the regression model training is completed, arrange the values of each parameter in descending order according to the regression coefficient, and select the top K parameters as important parameters for subsequent fine-tuning.
[0048] Further, in step (6), in each round of interaction between the tuning agent and the environment module, the policy network of reinforcement learning (the neural network includes the policy network and the value network) recommends new parameter values A t continuously according to the current state metric vector S of the distributed storage cluster t applied to the distributed storage cluster to obtain the feedback performance metric vector P t+1 and the state metric vector S t+1 After that, the tuning experience (S t , A t , S t+1 , R t+1 ) obtained in this round of interaction is saved in the tuning experience cache pool, and the reward R t+1 of state transition is calculated as where α i represents the linear weight of each performance metric, represents the reward value of a certain performance metric Perf i in the (t + 1)-th round of interaction with the environment module.
[0049] When calculating the reward value for pursuing positive benefits represented by IOPS:
[0050] Introduce a non-linear term to encourage the growth of IOPS, and at the same time use an absolute value term to maintain sensitivity to short-term fluctuations of IOPS;
[0051] where ΔIOPS0 represents the relative change rate of the IOPS of the distributed storage cluster after the parameter value is updated to A t compared with the IOPS when taking the default value, and the calculation formula is
[0052] ΔIOPS t Indicates that the parameter value is updated to A t After that, the IOPS of the distributed storage cluster is relatively taken as A t-1 The relative change rate of IOPS when taking A, and the calculation formula is
[0053] When calculating the reward value pursuing negative returns represented by latency:
[0054] Introduce a non - linear term to encourage the reduction of latency, and at the same time use an absolute value term to maintain sensitivity to short - term fluctuations of latency;
[0055] ΔLatency0 indicates that the parameter value is updated to A t After that, the relative change rate of the latency of the distributed storage cluster relative to the default value latency, and the calculation formula is
[0056] ΔLatency t Indicates that the parameter value is updated to A t After that, the Latency of the distributed storage cluster is relatively taken as A t-1 The relative change rate of Latency when taking A, and the calculation formula is
[0057] In the second aspect, the present invention provides a constrained hierarchical Latin hypercube sampling algorithm for the above - mentioned step (3). As shown in Algorithm 1, through the acceptance - rejection strategy, a certain number of sample sets that satisfy all constraint relationships can be obtained in a short time (lines 8 - 14 in Algorithm 1). However, actual tests found that when collecting 64 parameter value samples that satisfy 45 linear constraint relationships for 485 parameters of Ceph at one time, N cur Reaches the tens of millions level. Storing and judging whether these parameter configurations satisfy the constraint relationships requires nearly 100 GB of memory space. Therefore, to obtain a larger - scale sample set, combining the idea of hierarchical Latin hypercube sampling, combining small - scale sample sets (lines 2 - 4 in Algorithm 1), and optimizing the sample set through the simulated annealing method (Simulated Annealing, SA) (as shown in Algorithm 2) to improve the spatial uniformity of the combined sample set.
[0058]
[0059]
[0060] Specifically, in the SA-MM algorithm in Algorithm 1, the minimum distances between each parameter configuration sample point and other sample points are summed as the energy value of the overall sample set (lines 18 - 19 in Algorithm 2). By adding random perturbations to the parameter value vector of a certain parameter that does not involve constraint relationships, a new parameter configuration sample set is formed (lines 5 - 9 in Algorithm 2). If the sum of the minimum distances of the new sample set becomes larger, the new sample set is retained and the best sample set is updated; otherwise, the new sample set is accepted with the Boltzmann probability (lines 11 - 15 in Algorithm 2). After the iteration is completed, the best parameter configuration sample set is returned.
[0061]
[0062]
[0063]
[0064] In the third aspect, in the above step (6) of the present invention, an experience parallel reinforcement learning training acceleration technique is designed to optimize the values of important parameters by the agent. The reinforcement learning training acceleration technique is implemented using the DDPG algorithm of reinforcement learning, which is called the EPDT (Experience-based Parallel DDPG for Automatic Tuning, EPDT) algorithm. Figure 3 is a schematic diagram of EPDT. EPDT abstracts each single-process tuning agent and environment module as a Worker, and each Worker corresponds to a set of cloned environments. The tuning agent of the main process is called the global network (Learner), which collects batch experiences from the tuning experience cache pool shared by multiple Workers to update the value network and policy network of reinforcement learning.
[0065] The dual-process EPDT is shown in Algorithm 3. The main process randomly initializes the hyperparameters of the current network and the target network, and creates a pipeline for starting each slave process and the feedback of tuning experiences. In each round of interaction between the tuning agent and the environment module, first, it is verified whether the parameter configuration sample generated by the policy network satisfies the constraint relationship. If not, a new recommendation is made (lines 10 - 11) to avoid the risk of system failures. After receiving the feedback from its environment, the slave process puts the tuning experience into the tuning experience cache pool (lines 12 - 15). If the number of experiences in the pool is greater than the threshold, batch experiences are collected to update the global neural network of reinforcement learning (lines 16 - 21). The slave process and the main process collect tuning experiences in parallel. The difference between them is that after the slave process obtains the tuning experience, it only needs to put it into the local tuning experience cache pool and call child_conn.send(data) to send this experience to the main process (line 14).
[0066]
[0067]
[0068]
[0069] To verify the effectiveness of the present invention, a distributed storage cluster of Ceph Luminous version was built, and the automatic parameter tuning method of the present invention was further implemented. Specifically:
[0070] (1) The hierarchical Latin hypercube sampling method is improved by adopting an acceptance-rejection strategy, so that the sampled parameter value samples satisfy all the constraint relationships among the parameters, providing a high-quality training set for the screening of important parameters, avoiding the storage cluster failures caused by improper parameter values in step (4), and thus reducing the time cost of important parameter screening.
[0071] To evaluate the effectiveness of SA-HCLHS, 300 parameter value samples were collected using the original LHS and the SA-HCLHS designed in this paper respectively. Under the write-intensive load, these parameter value samples were applied to the system one by one. After the test was completed, the next set of parameter configurations was directly applied. During the test process, the status was captured every 30 s through the ceph-s command. Once keywords such as "100.000% pgs unknown", "HEALTH_ERR", and "osds down" appeared, it indicated that the system had failed. Record the parameter configuration at this time, and delete and recreate the cluster after that. The number of constraint failures caused by violating constraints during the collection of performance samples by the original LHS is shown in Table 1. No system failures occurred during the performance test of the performance samples collected by SA-HLHS. It can be seen that some of the constraint relationships among the parameters in Ceph are not strictly related, and not satisfying these constraints will lead to system failures, which also verifies the effectiveness of SA-HCLHS.
[0072] Table 1 Number of constraint failures in the LHS sample set
[0073]
[0074] (2) Based on the experience-parallel Ape-X architecture, the present invention proposes an experience-parallel reinforcement learning automatic parameter tuning technology, which reduces the "cold start" time of the tuning model while ensuring the optimality of the tuning performance, and thus shortens the overall tuning time.
[0075] The experience-parallel framework using multi-process parallel training can improve the quantity and quality of tuning experiences, which is an idea to solve the time series dependence problem of the single-process reinforcement learning algorithm. At present, the parallel computing of reinforcement learning is divided into two types: experience parallel and gradient parallel, both of which require multiple Workers (the environment with which a single-process agent interacts). The difference between the two lies in whether it is the experience or the gradient that is passed to the Learner (the global network), asFigure 4 As shown. In the case where the environmental feedback is very slow, the global network update utility is low, and at this time, the experience parallel architecture is more applicable.
[0076] The advantage of experience parallel is that except for the global network, it does not limit the tuning models adopted by other threads. To verify the impact of different Worker combination schemes on the tuning efficiency, the following dual-thread combination schemes are used for experiments: DDPG + DPPG, DDPG + SA-HCLHS, DDPG + GA, where the SA-HCLHS thread collects parameter configurations in advance and applies them to the environment one by one. Under write-intensive loads, 35h of automatic tuning is carried out, and the results are as Figure 5 shown. GA converges within 5 hours but finally falls into a local sub-optimal. Although DDPG + GA can provide more optimal experiences in the early stage, the time to recommend the optimal parameter configuration is the same as that of single-thread DDPG. DDPG + DDPG takes about 15 hours to recommend the optimal parameter configuration, while DDPG + LHS only takes 14 hours.
[0077] In the early stage of DDPG training, it is necessary to learn the mode of poor performance experiences through random exploration in order to avoid these action areas in the later stage. However, the samples generated by GA tend to be concentrated in local sub-optimal areas and do not help DDPG exploration substantially. LHS provides a wider set of performance samples through uniform sampling, which helps DDPG to contact experiences with various performances in the early stage of training, promotes the model's comprehensive understanding of the environment, and more efficiently realizes the automatic tuning process.
[0078] To further verify the effect of the present invention under different loads, tests are carried out under write-intensive loads, random read loads, and sequential read loads respectively, as Figure 6 shown. The experimental results show that: taking the performance of the storage cluster under the default parameter values as the benchmark, under write loads, the present invention can improve the performance by about 45%; under read loads, the present invention can improve the performance by about 100%.
[0079] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can still be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
Claims
1. A parameter tuning system for a distributed storage system based on empirical parallel reinforcement learning, characterized in that, The system includes five components: an environment module, a parameter pre-processor, an important parameter filter, a system controller, and a tuning agent; The environment module includes a distributed storage system and its workload, and the workload specific to the distributed storage system is simulated and generated by a load generator; The input of the parameter pre-processor is the initial parameter information of the distributed storage system, and it can output information that can be automatically adjusted. The value ranges of these adjustable parameters form the value space of the adjustable parameters; The important parameter filter is used to take the value space of the adjustable parameters as the sampling space to collect a parameter value sample set. The system controller is used to apply the parameter value samples to the distributed storage system one by one to generate a performance metric sample set, and then combine them into the training data set {(parameter value, performance metric)} of the important parameter filter. The important parameter filter is also used to use the training data set to screen out the important parameters that significantly affect the distributed storage system. The value ranges of these important parameters are used as the value space for fine-grained tuning of subsequent parameter values; The system controller is used to update the parameter values of the distributed storage system using the parameter update interface provided by the distributed storage system. After the parameter values are updated, it starts a performance test and obtains a performance metric vector, and then uses the state monitoring framework provided by the distributed storage system to obtain the state metric vector of the distributed storage system after the performance test is completed. The parameter values, the performance metric vector, the state metric vector, and the state transition reward of the distributed storage system are combined into tuning experience and saved in the tuning experience cache pool of the tuning agent. The system controllers of each tuning agent are independent of each other; The tuning agent is used to perform fine-grained tuning on the values of the important parameters. During the tuning process, through online interaction with the environment module, it iteratively adjusts its parameter tuning strategy until it obtains stable performance feedback of the distributed storage system through the system controller or reaches the tuning resource limit condition, at which point the tuning process ends, and the parameter value combination corresponding to the optimal performance during the tuning process is used as the optimal parameter value.
2. A method for tuning parameters of a distributed storage system based on experience parallel reinforcement learning implemented by the system as described in claim 1, characterized in that, The method includes the following steps: (1) Determine the optimization objective. In a distributed storage system, the read and write performance metrics include throughput, latency, and IOPS. The optimization objective can be to select a single performance metric or to quantify multiple performance metrics into a single integrated metric. (2) Obtain all the parameter information that can be adjusted by users exposed by the deployed distributed storage system. The parameter preprocessor filters the parameters that can be automatically adjusted, and further processes these parameters. Convert the boolean type and enumerated string type parameters into numerical types. For integer and floating-point type parameters without specified value ranges, set their value ranges to [0.1×default(K i ), 10×default(K i )], where represents the default value of parameter K i , i ∈ [1, N], and N is a positive integer; (3) The important parameter filter takes the value space of the parameters that can be automatically adjusted as the sampling space, and uses the sampling algorithm to collect a parameter value sample set A = {A1, …, A M}, where A t = (Action1, …, Action N ) represents a value combination of the parameters, represents the Min-Max normalization value of the value of parameter K i , where t = 1, 2, …, M; (4) Through the system controller, the parameter values in the parameter value sample set A are applied to the distributed storage system one by one. The load generation tool is used to perform a performance test on the distributed storage system for a certain period of time to obtain the performance metric vector P t+1 = <Perf1,…,Perf u >, where the elements in the performance metric vector are performance metrics. Here, u represents the number of nodes in the distributed storage system. At the same time, the system controller uses the status monitoring framework to obtain the status metric vector S t+1 = <State1,…,State v >, where v represents the number of status metric values obtained through the status monitoring framework of the distributed storage system; The parameter values, performance metric vector, and status metric vector are combined to form the tuning experience sample set {(A t ,P t+1 ,S t+1 )}, where P t+1 and S t+1 respectively represent the performance metric vector and the status metric vector of the distributed storage system after applying the parameter value A t ; (5) The important parameter filter utilizes a subset of the tuned experience sample set {(A t , P t+1 )} to filter out the important parameters K1′, K2′, …, K′ n that have a significant impact on the performance of the storage system, and forms the value space A’ of the important parameters; (6) The tuning agent tunes the values of important parameters through reinforcement learning, which is implemented based on a neural network. The tuning agent uses the value space A' of the important parameters as the action space of the reinforcement learning algorithm, and obtains the state metrics representing the operating state of the distributed storage system from the state monitoring framework of the distributed storage system. The value ranges of these state metrics are used as the state space of the tuning algorithm. The tuning agent and the environment module continuously interact to generate tuning experiences (S t , A t , S t+1 , R t+1 ), where S t represents the state metric vector of the distributed storage system before the parameter value is updated to A t , and R t+1 is related to the difference ΔP t+1 between the performance metric vectors before and after the parameter value update, representing the state transition reward of the distributed storage system. After the number of tuning experiences in the tuning experience cache pool of the tuning agent reaches a preset threshold, the tuning agent collects a batch of tuning experiences {(S t , A t , S t+1 , R t+1 )} to update the neural network hyperparameters. Therefore, the tuning agent will continuously iterate and adjust its parameter tuning strategy until it obtains stable performance feedback of the distributed storage system through the system controller or reaches the preset tuning duration, at which point the training process ends, and the parameter value combination corresponding to the optimal performance during the training process is used as the optimal parameter value.
3. The method according to claim 2, wherein In step (4), the system controller obtains the state metric values through the state monitoring framework of the distributed storage system. For cumulative state metrics, the calculation method is to accumulate the state metric values of each node. For mean-type state metrics, they are calculated by using a linear weighted function to fuse these scattered state metric values.
4. The method according to claim 2, wherein In step (4), the screening function of the important parameter filter in step (5) is implemented based on machine learning. By learning the regression model between the parameter values as independent variables and the performance metrics as dependent variables, the influence degree of each parameter value on the performance metrics is quantified using the coefficient of each parameter value, and the top K parameters with the greatest influence on the performance of the distributed storage system are selected as important parameters. The method for using the regression model to reduce the number of tuning experiences (A, P) required for model training is as follows: First, divide the tuning sample set, perform log1p normalization on the values of each parameter in the sample set, so that the parameter value sample set A has the same order of magnitude, and convert the performance metric set {P} into an optimization target set through the linear scalarization function g(·) in step (1) as the labeled data for model training; then, after the regression model training is completed, arrange the parameter values in descending order according to the regression coefficients, and take the top K parameters as important parameters.
5. The method according to claim 2, wherein In step (6), in each round of interaction between the tuning agent and the environment module, the policy network of reinforcement learning recommends new parameter values A t according to the current state metric vector S of the distributed storage system t and applies them to the distributed storage system to obtain the feedback performance metric vector P t+1 and the state metric vector S t+1 . After that, the tuning experience (S t , A t , S t+1 , R t+1 ) obtained in this round of interaction is saved in the tuning experience cache pool. The calculation formula for the reward R t+1 of state transition is where α i represents the linear weight of each performance metric, represents the reward value of a certain performance metric Perf i in the (t + 1)-th round of interaction with the environment module.
6. The method according to claim 5, wherein In step (6), the following formula is used to calculate the reward value for pursuing positive benefits represented by IOPS: Introducing a non - linear term encourages the growth of IOPS, while using an absolute - value term maintains sensitivity to short - term fluctuations of IOPS; where ΔIOPS0 represents that the parameter value is updated to A t The relative change rate of IOPS in the post - distributed storage system compared to when the default value is taken; ΔIOPS t represents that the parameter value is updated to A t The IOPS in the post - distributed storage system relative to taking A t-1 the relative change rate of IOPS at that time; The following formula is used to calculate the reward value for pursuing negative benefits represented by latency: Introducing a non - linear term encourages the reduction of latency, while using an absolute - value term maintains sensitivity to short - term fluctuations in latency; where ΔLatency0 represents the relative change rate of the latency of the distributed storage system after the parameter value is updated to A t The relative change rate of the latency of the distributed storage system after taking the default value of the latency; ΔLatency t represents the relative change rate of the latency of the distributed storage system after the parameter value is updated to A t The Latency of the distributed storage system after taking A t-1 The relative change rate of Latency when taking A 7. The method according to claim 5, wherein In step (3), the constrained hierarchical Latin hypercube sampling algorithm is used to collect a parameter value sample set containing M samples. In the constrained hierarchical Latin hypercube sampling algorithm, through the acceptance-rejection strategy, a certain number of sample sets that satisfy all constraint relationships are obtained in a short time, and combined with the idea of hierarchical Latin hypercube sampling, the sample set is optimized by the simulated annealing method to improve the spatial uniformity of the combined sample set.
8. The method according to claim 5, characterized in that, In step (6), an experience parallelized reinforcement learning training acceleration technique is designed to optimize the values of important parameters by the agent. The reinforcement learning training acceleration technique is implemented using the DDPG algorithm of reinforcement learning, which is called the EPDT algorithm. In the EPDT algorithm, each single-process tuning agent and the environment module are abstracted as a Worker, and each Worker corresponds to a set of cloned environments. The tuning agent in the main process is called the global network Learner, and batch experiences are collected from the tuning experience cache pool shared by multiple Workers to update the value network and policy network of reinforcement learning.
9. The method according to claim 8, wherein The experience parallelized reinforcement learning training acceleration technique is a two-process EPDT algorithm. Among them, the main process randomly initializes the hyperparameters of the current network and the target network, and creates a pipeline to start each slave process and the feedback of the tuning experience. In each round of interaction between the tuning agent and the environment module, first verify whether the parameter configuration sample generated by the policy network satisfies the constraint relationship. If not, recommend new parameter values again. After the slave process receives the feedback from the environment module, it puts the tuning experience into the tuning experience cache pool. If the number of tuning experiences in the tuning experience cache pool is greater than the threshold, batch experiences are collected to update the global neural network of reinforcement learning. The slave process and the main process collect tuning experiences in parallel. However, after the slave process obtains the tuning experience, it only puts it into the local tuning experience cache pool and sends this tuning experience to the main process.
10. The method according to claim 6, characterized in that, The distributed storage system is a distributed storage cluster of the Ceph Luminous version.
Citation Information
Patent Citations
Intelligent adjusting and optimizing system for performance parameters of distributed cloud storage platform
CN114489499A
Distributed storage system automatic parameter adjustment method and system based on reinforcement learning
CN116088761A