Parameter recommendation method, device and computer storage medium

By building a performance prediction model and iteratively updating the parameter set, the problem of the distributed storage system's parameter tuning depends on manual experience, and the parameter recommendation for automation and resource saving is realized, and the storage performance is improved.

CN115858660BActive Publication Date: 2025-08-26CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111115232.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-23
Publication Date
2025-08-26
Estimated Expiration
2041-09-23

AI Technical Summary

Technical Problem

In the prior art, the parameter tuning scheme of distributed storage systems relies too much on the knowledge base and practical experience of technicians, making it difficult to compatible with the needs of multiple storage scenarios, and it consumes a lot of resources and has low optimization efficiency.

Method used

By determining the parameter input set and system parameters of the distributed storage system, building a performance prediction model, conducting model training, and iteratively updating the parameter set to achieve automated parameter recommendations.

Benefits of technology

Reduce resource consumption, improve parameter optimization efficiency, and targeted recommendation of optimal parameters in different storage scenarios to avoid manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115858660B_ABST
    Figure CN115858660B_ABST
Patent Text Reader

Abstract

The present application discloses a parameter recommendation method, device, and computer storage medium. The method includes: determining a parameter input set of a distributed storage system and system parameters configured in at least one storage scenario; determining at least one performance indicator based on the parameter input set and the system parameters; performing model training based on the parameter input set, the system parameters, and the at least one performance indicator to determine at least one performance prediction model and at least one update parameter set; wherein each performance indicator corresponds to a performance prediction model and an update parameter set; and recommending parameters to the distributed storage system based on the at least one performance prediction model and the at least one update parameter set. In this way, since the system parameters in different storage scenarios are taken into account during model training, the optimal parameter recommendation in different storage scenarios can be achieved based on the trained performance prediction model, and resource consumption can also be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of distributed storage technology, and in particular to a parameter recommendation method, device, and computer storage medium. Background Art

[0002] Distributed storage systems, as the name suggests, store data in a distributed manner across multiple independent devices. Traditional network storage systems use centralized storage servers to store all data, which cannot meet the needs of large-scale storage applications. Distributed storage systems adopt a scalable system architecture, utilizing multiple storage servers to share the storage load and location servers to locate stored information. This not only improves system reliability, availability, and access efficiency, but also facilitates scalability.

[0003] In the related art, improving the performance of distributed storage systems typically involves considering both hardware and software optimization. Hardware optimization often imposes rigid hardware requirements, resulting in high optimization costs. Software optimization, for example, offers thousands of configurable parameters for a distributed storage system like Ceph. Configuring distributed storage systems relies heavily on the experience of storage engineers. However, even experienced storage engineers often struggle to guarantee that the parameters they set are optimal for achieving optimal performance.

[0004] At present, although there are many parameter tuning solutions for distributed storage systems, they all have defects such as over-reliance on the knowledge base and practical experience of technical personnel and difficulty in being compatible with the requirements of various storage scenarios. Summary of the Invention

[0005] The present application provides a parameter recommendation method, device and computer storage medium, which can make targeted parameter recommendations for distributed storage systems based on the performance indicators emphasized in different storage scenarios, thereby reducing resource consumption and improving parameter optimization efficiency.

[0006] The technical solution of this application is achieved as follows:

[0007] In a first aspect, an embodiment of the present application provides a parameter recommendation method, the method comprising:

[0008] Determining a parameter input set of a distributed storage system and system parameters configured in at least one storage scenario;

[0009] determining at least one performance indicator based on the parameter input set and the system parameters;

[0010] Performing model training based on the parameter input set, the system parameters, and the at least one performance indicator to determine at least one performance prediction model and at least one updated parameter set; wherein each performance indicator corresponds to a performance prediction model and an updated parameter set, and the updated parameter set is obtained after iterative updating of the parameter input set during the model training process;

[0011] Parameters are recommended to the distributed storage system based on the at least one performance prediction model and the at least one update parameter set.

[0012] In a second aspect, an embodiment of the present application provides a parameter recommendation device, which includes a determination unit, a training unit, and a recommendation unit, wherein:

[0013] The determining unit is configured to determine a parameter input set of the distributed storage system and system parameters configured in at least one storage scenario; and determine at least one performance indicator based on the parameter input set and the system parameters;

[0014] The training unit is configured to perform model training based on the parameter input set, the system parameters, and the at least one performance indicator, and determine at least one performance prediction model and at least one update parameter set; wherein each performance indicator corresponds to a performance prediction model and an update parameter set, and the update parameter set is obtained after iterative updating of the parameter input set during the model training process;

[0015] The recommendation unit is configured to recommend parameters to the distributed storage system based on the at least one performance prediction model and the at least one update parameter set.

[0016] In a third aspect, an embodiment of the present application further provides a parameter recommendation device, which includes a memory and a processor, wherein:

[0017] The memory is used to store a computer program that can be run on the processor;

[0018] The processor is configured to execute the parameter recommendation method as described in the first aspect when running the computer program.

[0019] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program, which implements the parameter recommendation method described in the first aspect when executed by at least one processor.

[0020] The present application provides a parameter recommendation method, device, and computer storage medium, which determine a parameter input set of a distributed storage system and system parameters configured in at least one storage scenario; determine at least one performance indicator based on the parameter input set and the system parameters; perform model training based on the parameter input set, the system parameters, and the at least one performance indicator to determine at least one performance prediction model and at least one update parameter set; wherein each performance indicator corresponds to a performance prediction model and an update parameter set, and the update parameter set is obtained after the parameter input set is iteratively updated during the model training process; and recommend parameters to the distributed storage system based on the at least one performance prediction model and the at least one update parameter set. In this way, when performing parameter recommendation, since the system parameters in different storage scenarios are taken into account in the model training, the performance prediction model and the update parameter set obtained by the training can be used to specifically recommend parameters to the distributed storage system, without having to analyze each parameter individually, and the entire process is automated, which can avoid manual intervention; in addition, adding the performance prediction model to the parameter optimization process not only improves the efficiency of parameter optimization, but also reduces resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A flow chart of a parameter recommendation method provided in an embodiment of the present application;

[0022] Figure 2 A flow chart of another parameter recommendation method provided in an embodiment of the present application;

[0023] Figure 3 A schematic diagram of the application architecture of a parameter recommendation system provided in an embodiment of the present application;

[0024] Figure 4 A schematic diagram of the structure of a parameter recommendation device provided in an embodiment of the present application;

[0025] Figure 5 A schematic diagram of the structure of another parameter recommendation device provided in an embodiment of the present application;

[0026] Figure 6 A schematic diagram of the specific hardware structure of a parameter recommendation device provided in an embodiment of the present application;

[0027] Figure 7 A schematic diagram of the composition structure of a parameter recommendation device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. It should be understood that the specific embodiments described herein are only used to explain the related applications and are not intended to limit the applications. It should also be noted that for ease of description, only the parts relevant to the related applications are shown in the drawings.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0030] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0031] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0032] It should be understood that a distributed storage system (also known as "distributed storage software," "distributed storage cluster," "distributed file system," etc.) can itself be viewed as a queuing model. The interaction between clients and the distributed storage system primarily involves queue processing of input / output (I / O)-related instructions. Therefore, distributed storage system optimization often focuses on the queue model, primarily aiming to increase parallelism and reduce service time. Both of these approaches are closely related to factors such as the server's hardware configuration and operating system settings, such as the processing power of the Central Processing Unit (CPU), the parallelism of the disks, the I / O path length in software design, memory size, memory recycling mechanisms, and network bandwidth.

[0033] Numerous parameter tuning schemes exist for distributed storage systems. Most begin by defining a parameter range and then make real-time adjustments based on performance monitoring during cluster operation. These tuning strategies rely heavily on the technical staff's knowledge and practical experience. Other tuning schemes apply statistical and machine learning techniques to parameter tuning. These schemes construct an input set by sampling the parameter range, determine evaluation metrics, and then try different parameter configurations to select the one with the best performance. This approach can theoretically find the optimal solution, especially for nonlinear scenarios. It offers significant tuning results and can be automated, requiring minimal operator intervention. However, the iterative process consumes significant system resources, and the results are not compatible with varying hardware and operating system configurations. Additionally, while some tuning schemes propose building a predictive model first and incorporating it into the parameter training process to reduce resource consumption and accelerate the search for the optimal parameter combination, these approaches are still highly coupled to the cluster's configuration.

[0034] Based on this, an embodiment of the present application provides a parameter recommendation method, the basic idea of ​​which is: determining a parameter input set of a distributed storage system and system parameters configured in at least one storage scenario; determining at least one performance indicator based on the parameter input set and the system parameters; performing model training based on the parameter input set, the system parameters and the at least one performance indicator to determine at least one performance prediction model and at least one update parameter set; wherein each performance indicator corresponds to a performance prediction model and an update parameter set, and the update parameter set is obtained after the parameter input set is iteratively updated during the model training process; and recommending parameters to the distributed storage system based on the at least one performance prediction model and the at least one update parameter set. In this way, when making parameter recommendations, since the system parameters in different storage scenarios are taken into account in the model training, the performance prediction model and the update parameter set obtained by the training can be used to make targeted parameter recommendations for the distributed storage system, without the need to analyze each parameter, and the entire process is automated to avoid manual intervention; in addition, adding the performance prediction model to the parameter optimization process not only improves the efficiency of parameter optimization, but also reduces resource consumption.

[0035] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0036] In one embodiment of the present application, see Figure 1 , which shows a flow chart of a parameter recommendation method provided by an embodiment of the present application. Figure 1 As shown, the method may include:

[0037] S101: Determine a parameter input set of a distributed storage system and system parameters configured in at least one storage scenario.

[0038] It should be noted that the parameter recommendation method provided in the embodiments of the present application is used to recommend a set of optimized parameters that optimize the performance indicators of the distributed storage system in combination with actual storage scenarios before creating the distributed storage system. The method can be applied to a device for performing parameter recommendation, or a device or system incorporating the device. Here, the electronic device can be a computer, a smartphone, a tablet computer, a laptop computer, a PDA, a personal digital assistant (PDA), a navigation device, etc., and the embodiments of the present application do not specifically limit this.

[0039] In an embodiment of the present application, the parameter input set of the distributed storage system refers to the configurable parameters from the distributed storage system itself. Taking the distributed storage system Ceph as an example, its own configurable parameters can reach thousands, for example, the maximum number of bytes written to the journal at one time: journal max write bytes, the maximum number of operations stored in the queue at any time: journal queue max ops, etc.

[0040] It should also be noted that in the application embodiment, the initial parameter set can be a parameter set optimized after screening a large number of configurable parameters of the distributed storage system itself, that is, not every configurable parameter of the distributed storage system is used as a parameter in the initial parameter set.

[0041] In addition, the system parameters here refer to some inherent configuration parameters of the system under different storage scenarios (such as different operating systems, hardware configurations, and storage strategies), such as: number of CPU cores, memory size, maximum thread number limit, etc. In different storage scenarios, even if the configuration parameters of the distributed storage system are the same, its performance indicators may have different performances. In the embodiments of the present application, for system parameters, the system parameters configured in different storage scenarios can be determined separately, such as different storage scenarios such as bare metal block storage, virtual machine block storage, object storage, and all-flash block storage.

[0042] S102: Determine at least one performance indicator according to the parameter input set and system parameters.

[0043] It should be noted that after the parameter input set and the system parameters are determined, the performance indicators of the distributed storage system under different parameter combinations can be determined based on the parameter input set and the system parameters.

[0044] In some embodiments, determining at least one performance indicator based on the parameter input set and the system parameter may include:

[0045] Based on the parameter input set and system parameters, build a test cluster corresponding to at least one storage scenario;

[0046] During operation of the test cluster corresponding to at least one storage scenario, at least one performance indicator is obtained using a performance testing tool;

[0047] The performance indicators include at least one of the following: IOPS (number of reads and writes per unit time), throughput, CPU utilization, memory utilization, swap memory SWAP utilization, and read and write latency.

[0048] It should be noted that the parameters of the distributed storage system and different values ​​of the system parameters may affect the performance of the distributed storage system. This step mainly determines the performance indicators of the distributed storage system under different parameter values ​​and combinations through the parameter input set and system parameters of the distributed storage system using performance testing tools such as perf and fio. Performance indicators mainly refer to read and write indicators, including but not limited to the number of read and write operations per second (IOPS), throughput, central processing unit (CPU) utilization, memory utilization, swap memory (SWAP) utilization, and read and write latency.

[0049] S103: Perform model training based on the parameter input set, system parameters, and at least one performance indicator to determine at least one performance prediction model and at least one update parameter set.

[0050] It should be noted that by using the parameter input set, system parameters and at least one performance indicator as the training set for model training, at least one performance prediction model and an updated parameter set can be obtained, wherein each performance indicator corresponds to a performance prediction model and an updated parameter set, and the updated parameter set is obtained after the parameter input set is iteratively updated during the model training process.

[0051] S104: Recommend parameters to the distributed storage system based on at least one performance prediction model and at least one updated parameter set.

[0052] It should be noted that after the aforementioned steps, at least one performance prediction model and one set of updated parameters corresponding to at least one performance indicator are obtained. Based on the at least one performance prediction model and the at least one set of updated parameters, an optimal set of parameters is determined, and parameter recommendations can be made to the distributed storage system.

[0053] It should also be noted that in order to improve the performance of distributed storage systems, two levels of optimization are generally considered: hardware optimization and software optimization. In one possible implementation, considering optimization from the hardware level, for hard disk types, the read and write performance of solid-state drives (SSDs) is much higher than that of mechanical hard disk drives (HDDs), especially SSDs based on the Non-Volatile Memory express (NVMe) protocol, which can fully utilize the advantages of CPU multi-cores and greatly increase concurrency. However, its cost is also considerable, so the scope of use of all-flash clusters is limited. In addition, if the cluster adopts an all-flash deployment, then disk IO will no longer be a bottleneck. The length of the IO path, especially the switching of the operating system kernel state, and the multi-node network communication involved in distributed operations will become a new bottleneck. Some manufacturers have introduced solutions such as Remote Direct Memory Access (RDMA) and the user state of the Storage Performance Development Kit (SPDK) to solve new performance bottlenecks. However, all of the above solutions to improve the performance of distributed storage systems undoubtedly have rigid requirements on hardware, high economic costs, and high transformation costs.

[0054] Another possible implementation involves optimizing at the software level. For example, Ceph, a common open-source distributed storage system, offers thousands of configurable parameters. These parameters impact performance in different storage scenarios (e.g., replication or erasure coding), when deployed on different devices (e.g., HDDs or SSDs), and when using different operating systems (e.g., x86 or ARM). Even experienced storage engineers cannot guarantee optimal performance in all scenarios by selecting the most appropriate parameters.

[0055] For example, Ceph's filestore has a parameter called filestore_op_threads, which represents the number of IO threads. Setting this parameter to a larger value can speed up IO processing, but if there are too many threads, frequent thread switching will also affect the performance of the distributed storage system. For another example, Ceph's bluestore design should try to avoid double writes, but metadata and some relatively small data (such as data less than 16KB on HDDs and less than 64KB on SSDs) will be written to the key-value (kv) database. The index feature of the database can speed up reading and writing, but at the same time, this part also brings the problem of double writes.

[0056] Based on this, in another possible implementation, the embodiment of the present application can recommend the most suitable parameter configuration for a distributed storage system based on hardware configuration, system settings, storage scenarios, etc. This application optimizes the performance of a distributed storage system from the software level, using automatic parameter tuning as a starting point, and provides a set of automatic parameter tuning methods and systems that can select the optimal parameter combination based on hardware configuration, storage strategy, and different usage scenarios to improve the storage performance of different distributed storage systems.

[0057] An embodiment of the present application provides a parameter recommendation method, which comprises determining a parameter input set of a distributed storage system and system parameters configured in at least one storage scenario; determining at least one performance indicator based on the parameter input set and the system parameters; performing model training based on the parameter input set, the system parameters, and the at least one performance indicator to determine at least one performance prediction model and at least one update parameter set; wherein each performance indicator corresponds to a performance prediction model and an update parameter set; and recommending parameters to the distributed storage system based on the at least one performance prediction model and the at least one update parameter set. In this way, since the system parameters in different storage scenarios are taken into account in the model training, the performance prediction model obtained by the training can achieve optimal parameter recommendations in different storage scenarios and can also reduce resource consumption.

[0058] In another embodiment of the present application, since the embodiment of the present application is mainly used in a distributed software-defined storage (SDS) system that provides storage resource services at the Infrastructure as a Service (IaaS) layer, for the distributed SDS system at the IaaS layer, the number of parameters involved in this type of distributed storage system is often large, and the value range of each parameter is very wide. In different usage scenarios, the parameters that have a decisive impact on performance are also different.

[0059] For example, the distributed storage system Ceph offers thousands of configurable parameters. However, not every parameter affects performance, and the parameters that affect performance vary across different storage scenarios. Therefore, it's necessary to filter these parameters to identify those that significantly impact or significantly impact the performance of the distributed storage system. Targeted training can then be performed on these parameters, ultimately generating the optimal set of parameters for creating the distributed storage system.

[0060] Therefore, in a specific example, the embodiment of the present application can screen the parameter types of the distributed storage system in a certain way to determine the parameter input set of the distributed storage system. Figure 2 , which shows a flow chart of another parameter recommendation method provided by an embodiment of the present application. Figure 2 As shown, the method may include:

[0061] S201: Obtain an initial parameter set of a distributed storage system, and randomly sample parameters in the initial parameter set within a preset parameter value range to determine a test sample.

[0062] S202: Determine at least one performance indicator based on the test sample and system parameters.

[0063] S203. Perform model training on the test sample, system parameters, and at least one performance indicator using a second preset algorithm to determine at least one intermediate performance prediction model and the contribution of each parameter in the test sample to the prediction accuracy of each intermediate performance prediction model.

[0064] S204: Perform contribution analysis on the determined prediction accuracy contribution, and select a parameter input set from the test sample based on the analysis result.

[0065] It should be noted that the initial parameter set of the distributed storage system includes all configurable parameters of the distributed storage system (or it can also be a parameter set obtained by experienced engineers after removing parameters that will inevitably have no impact on the performance of the distributed storage system). Each configurable parameter of the distributed storage system has its specific value range, that is, the preset parameter value range. Each parameter is randomly sampled within its preset parameter value range to obtain a test sample.

[0066] It should also be noted that the system parameters can be system parameters configured in a distributed storage system under at least one storage scenario; the at least one storage scenario here can be bare metal block storage, virtual machine block storage, object storage, all-flash block storage, etc., which is not specifically limited in the embodiments of this application. In this way, after determining the test samples and system parameters, based on the test samples and system parameters, multiple distributed storage system parameters and parameter combinations of system parameters can be obtained. Then, under each parameter combination, the performance of the distributed storage system under that parameter combination can be determined, that is, each parameter combination will correspond to a set of performance indicators.

[0067] Specifically, you can use performance testing tools such as perf and fio to build a distributed storage system based on test samples and system parameters, and then run the distributed storage system to obtain the correspondence between specific parameter combinations and performance indicators.

[0068] Here, performance indicators mainly refer to read and write indicators, and may include, but are not limited to, one or more of the following: the number of read and write operations per second (i.e., IOPS), throughput (expressed in mbps), CPU utilization, memory utilization, SWAP utilization, read and write latency, etc. Of these, IOPS, mbps, and read and write latency are of particular interest. However, the specific performance indicators to be obtained can be set based on actual conditions and are not specifically limited in this embodiment of the present application.

[0069] After obtaining at least one set of performance indicators, the test samples, system parameters and performance indicators can be trained using the second preset algorithm to obtain an intermediate performance prediction model. The intermediate performance prediction model here is the same type of model as the performance prediction model described in the aforementioned embodiment, but the intermediate performance prediction model cannot be used as the final performance prediction model. It serves as a partial basis for determining the parameter input set, which is equivalent to an intermediate process in determining the performance prediction model.

[0070] In an embodiment of the present application, the second preset algorithm selects the random forest algorithm, and the intermediate performance prediction model is obtained by training the random forest algorithm, with test samples and system parameters as input and performance indicators as output. Random Forest (RF) is a classifier containing multiple decision trees, each of which can be used for prediction. By establishing multiple decision trees and fusing them together, a more accurate and stable model is obtained. Since parameter values ​​and performance indicators are often not linearly related, the random forest algorithm can well solve the problem of nonlinear relationship between parameter values ​​and performance indicators.

[0071] It should also be noted that the set of intermediate performance prediction models obtained in the embodiments of the present application includes at least one intermediate performance prediction model corresponding to at least one performance indicator. That is, for each performance indicator, a corresponding intermediate performance prediction model can be obtained. Examples include an intermediate IOPS prediction model, an intermediate Mbps prediction model, and an intermediate read / write latency prediction model.

[0072] It should also be noted that, while obtaining the intermediate performance prediction model corresponding to each performance indicator, the contribution of each parameter to the prediction accuracy of each intermediate performance prediction model can also be obtained.

[0073] In some embodiments, before performing model training on the test sample, system parameters, and at least one performance indicator using the second preset algorithm to determine at least one intermediate performance prediction model and the contribution of each parameter in the test sample to the prediction accuracy of each intermediate performance prediction model, the method may further include:

[0074] Perform mutation point detection on each type of parameter and at least one performance indicator in the parameter input set;

[0075] A parameter type corresponding to a sudden change in at least one performance indicator is determined as a parameter type for model training.

[0076] It should be noted that before determining the contribution of each parameter to the prediction accuracy of each intermediate performance prediction model, in order to avoid calculating the contribution of unnecessary parameters to the prediction accuracy, the mutation point detection method can also be used to determine whether it is necessary to perform prediction accuracy contribution detection on the parameter, that is, to determine whether a certain parameter needs to be used when training the model.

[0077] Specifically, during the training of an intermediate performance prediction model, the impact of each change in the value of a parameter on performance may be significant or minimal. Calculating the parameter's contribution to prediction accuracy for each performance result is meaningless, as the performance change may be minimal. If a performance indicator undergoes a sudden change at a certain parameter value, then it is necessary to calculate the contribution to prediction accuracy. Mutation point detection involves finding the parameter type corresponding to the point at which the performance indicator undergoes a mutation. The parameter set corresponding to the parameter type with a mutation point affecting the performance indicator will be used as training data to train the intermediate performance prediction model and determine the contribution of these parameters to the prediction accuracy of the intermediate performance prediction model. Parameters that do not have a mutation point in their impact on the performance indicator will no longer be used as parameters for training the intermediate performance prediction model.

[0078] Here, the contribution of each parameter to the prediction accuracy of the intermediate performance prediction model corresponding to each performance indicator may be different. For example, for parameter 1, it has a great impact on performance A. A slight change in its value may cause performance A to change from extremely poor to extremely good. In this case, for performance A, parameter 1 is an important parameter; while for performance B, parameter 1 has little impact on it. Within the value range of parameter 1, no matter how parameter 1 is taken, performance B will not change or will change very little. In this case, for performance B, parameter 1 can be regarded as an "irrelevant" parameter. The impact of a parameter on a certain performance can be characterized by the contribution of the parameter to the prediction accuracy of the performance prediction model corresponding to the performance indicator. The greater the prediction accuracy contribution, the greater the impact of the parameter on this performance; conversely, the smaller the prediction accuracy contribution, the smaller the impact of the parameter on this performance.

[0079] Therefore, the contribution to the prediction accuracy can be analyzed, and a parameter input set can be selected from the initial parameter set based on the analysis results.

[0080] In some embodiments, performing contribution analysis on the determined prediction accuracy contribution and selecting the parameter input set from the test sample based on the analysis result may include:

[0081] A candidate prediction accuracy contribution degree whose prediction accuracy contribution degree is greater than a preset contribution degree threshold is selected from the determined prediction accuracy contribution degrees, and a first parameter set is determined using parameters corresponding to the candidate prediction accuracy contribution degrees.

[0082] It should be noted that in the embodiments of the present application, parameters are screened based on their contribution to the prediction accuracy of the intermediate performance prediction model. A preset contribution threshold can be set. When a parameter's contribution to the prediction accuracy of the intermediate performance prediction model exceeds the preset contribution threshold, it indicates that the parameter has a significant impact on performance and can be selected as an important parameter. For example, the preset contribution threshold can be 50%, 60%, etc., and this embodiment of the present application does not specifically limit this.

[0083] Specifically, first, a candidate prediction accuracy contribution whose prediction accuracy contribution is greater than a preset contribution threshold is selected from the determined prediction accuracy contribution, and then the parameters corresponding to the candidate prediction accuracy contribution are determined. These parameters are the important parameters that are screened out, and this important parameter constitutes the first parameter set. These parameters are randomly sampled within their preset parameter value range, and the process of screening the parameter input set of the distributed storage system is completed.

[0084] It should also be noted that different system parameters have different impacts on the performance of distributed storage systems. Therefore, while determining the contribution of each parameter in the test sample to the prediction accuracy of each performance prediction model, the contribution of each system parameter to the prediction accuracy of each performance prediction model can also be calculated. Then, using the same method, the system parameters whose contribution to the prediction accuracy of the performance prediction model corresponding to a certain performance indicator exceeds a contribution threshold are selected as the updated system parameters.

[0085] In this way, multiple iterations can be performed and the parameter input set can be continuously updated, that is, the first parameter input set is used to replace the initial parameter input set to continue model training.

[0086] Furthermore, in some embodiments, the method further comprises:

[0087] determining a number of iterations of the at least one performance prediction model;

[0088] When the number of iterations is less than a preset number of iterations, determining the first parameter set as the initial parameter set, incrementing the number of iterations by 1, and returning to the step of randomly sampling the parameters in the initial parameter set within a preset parameter value range to determine a test sample;

[0089] When the number of iterations reaches a preset number of iterations, the first parameter set is determined as the parameter input set, and the model obtained after the most recent iterative update is determined as the at least one performance prediction model.

[0090] It should be noted that the preset number of iterations can be determined based on factors such as the complexity of the actual model and the test environment, for example, it can be 10 times, 30 times, 100 times, etc., and the embodiments of the present application do not make specific limitations on this.

[0091] Among them, the process of obtaining at least one intermediate performance test model through each training is called an iteration. When the number of iterations is less than the preset number of iterations, it is necessary to continue screening parameters. After determining at least one performance prediction model and the corresponding first parameter set, the first parameter set is determined as the initial parameter input set, and the number of iterations is added by 1. Then, according to the first parameter set obtained after the iterative update (that is, the updated initial parameter input set), return to the step of randomly sampling the parameters in the initial parameter set within the preset parameter value range, determine the test sample, and continue training to obtain at least one new set of intermediate performance prediction models.

[0092] When the number of iterations reaches the preset number, there is no need to filter the parameters anymore. The first parameter set obtained at this time can be directly determined as the parameter input set, and the model obtained after the most recent iterative update can also be directly determined as the final at least one performance prediction model.

[0093] In addition, the basis for not iterating anymore may also be that when the absolute value of the difference between the performance index predicted by the performance prediction model and the actual performance index is less than the error threshold (or when the loss function value of the model is less than the preset loss threshold), it means that the prediction accuracy of the performance prediction model is already high, and then no iteration is performed.

[0094] That is to say, the initial parameter input set (which may also include system parameters) used to train the performance prediction model is continuously updated through the performance prediction model (the performance prediction model in the process of screening parameters is called an intermediate performance prediction model) itself until the final parameter input set is determined. This is an iterative process that can be completed by the Expectation-Maximization algorithm (EM).

[0095] In this way, by combining the process of training the performance prediction model with determining the parameter input set, a parameter input set is selected. The parameter input set serves as both sample parameters for training the performance prediction model and as parameters to be optimized. This approach has the advantage of making the resulting performance prediction model more accurate, as it is trained using the parameters that have the greatest impact on performance. It also eliminates the need to optimize parameters that have little or no impact on performance, thus reducing resource consumption.

[0096] In another specific example, the embodiment of the present application can also determine the parameter input set of the distributed storage system by means of IO stack analysis. Therefore, in some embodiments: determining the parameter input set of the distributed storage system may include:

[0097] Analyze the IO stack of the distributed storage system and determine the relevant parameters of the IO stack;

[0098] Random sampling is performed on relevant parameters of the IO stack within a preset IO stack parameter value range to obtain a relevant parameter set, and the relevant parameter set is determined as a parameter input set.

[0099] It should be noted that if certain types of parameters are related to the IO stack of a distributed storage system, then these types of parameters will inevitably affect the performance of the distributed storage system. Therefore, by analyzing the IO stack of the distributed storage system, relevant parameters of the IO stack can be determined. These parameters affect the response of the IO path and can serve as important parameters affecting the performance of the distributed storage system. By randomly sampling these important parameters (i.e., relevant parameters of the IO stack) within their parameter value range (i.e., the preset IO stack parameter value range), the parameter input set of the distributed storage system can be obtained.

[0100] However, when determining the parameter input set of the distributed storage system in this way, the R&D personnel are often required to have a relatively in-depth understanding of the implementation of the distributed storage system, which is relatively dependent on the experience of the R&D personnel.

[0101] After determining a parameter input set of the distributed storage system and system parameters configured in at least one storage scenario, at least one performance indicator may be determined based on the parameter input set and the system parameters.

[0102] It should be noted that, since in the embodiment of the present application, the process of determining the parameter input set is combined with the process of training the prediction model, the at least one performance prediction model and the at least one updated parameter set finally determined here are, in the aforementioned iterative process, when the number of iterations reaches the preset number, the at least one performance prediction model obtained after the most recent iterative update, and the updated parameter set corresponding to the at least one performance prediction model is the parameter set in the parameter input set after the most recent iterative update used to train the corresponding performance prediction model.

[0103] For example, assume that there are 100 configurable parameters for a distributed storage system, namely parameter 1, parameter 2, parameter 3...parameter 100. These 100 parameters are randomly sampled within their preset parameter value ranges to obtain a parameter input set, and then a performance test is performed together with the system parameters (the system parameters can be used as the filtered parameters, or the system parameters can be not filtered. This example assumes that the system parameters are not filtered, and only the configurable parameters of the distributed storage system itself are filtered), thereby obtaining multiple corresponding performance indicator values. Assume that only performance A, performance B, and performance C of the distributed storage system are tested here. Then, after the aforementioned iterative process (for example, 100 iterations), three prediction models corresponding to performance A, B, and C are finally obtained. Moreover, during the iterative process, based on the contribution of each parameter to the prediction accuracy of the performance prediction model, it is determined which parameters have a greater impact on specific performance indicators.

[0104] After 100 iterations, the final prediction model A corresponding to performance indicator A is obtained from the parameter input set consisting of parameters 1 to 15 and parameter 27. Therefore, the updated parameter set corresponding to prediction model A is parameters 1 to 15 and 27. Similarly, if the final prediction model B corresponding to performance indicator B is obtained from the parameter input set consisting of parameters 5 to 10 and parameters 70 to 75, the updated parameter set corresponding to prediction model B is parameters 5 to 10 and parameters 70 to 75. If the final prediction model C corresponding to performance indicator C is obtained from the parameter input set consisting of parameters 60 to 73, the updated parameter set corresponding to prediction model C is parameters 60 to 73. Parameters other than the execution parameters have little impact on performance and are not included in the parameter input set for training the performance prediction model.

[0105] It should also be noted that if in step S101, the parameter input set is obtained based on the IO stack analysis method, here, the model training can be directly performed based on the parameter input set, system parameters and at least one performance indicator to obtain at least one performance prediction model and a parameter set corresponding to each of the at least one performance prediction model (also called an updated parameter set).

[0106] Since, in the embodiment of the present application, the performance prediction model finally obtained is obtained based on the updated parameter set after multiple iterations, when the model is used for parameter recommendation, the optimal parameter set of parameters directly related to performance can be obtained. In addition, since, when training the performance prediction model, the embodiment of the present application also uses the system parameters as a parameter input set to determine the performance indicators and use them for training the performance prediction model, when performing parameter recommendation, a set of configuration parameters that cooperate with the system parameters of the storage scenario of the distributed storage system can be directly obtained, and the configuration parameters that make the distributed storage system perform best in a specific storage scenario can be obtained.

[0107] Furthermore, in some embodiments, during the model training process, the method may further include:

[0108] After obtaining the first parameter set, determining the prediction accuracy contribution of each parameter in the first parameter set corresponding to each performance indicator based on the prediction accuracy contribution of each parameter in the first parameter set to each performance prediction model and the corresponding relationship between the performance prediction model and the performance indicator;

[0109] Under each performance indicator, a candidate prediction accuracy contribution whose prediction accuracy contribution is greater than a preset contribution threshold is selected from the determined prediction accuracy contributions, and the parameters corresponding to the candidate prediction accuracy contribution are determined as an updated parameter set corresponding to each performance indicator.

[0110] It should be noted that in the process of training the performance prediction model, after obtaining the first parameter set, it is possible to determine the contribution of each parameter in the first parameter set to the prediction accuracy of each performance prediction model, as well as the correspondence between the performance prediction model and the performance indicator, for each performance indicator.

[0111] In this way, a first parameter set can be obtained for each performance indicator. Under each performance indicator, the parameters in the first parameter set whose contribution to the prediction accuracy of the performance indicator is greater than a preset contribution threshold are determined as the updated parameter set corresponding to the performance indicator. This allows the corresponding relationship between the parameters and the performance indicator to be determined.

[0112] Furthermore, after determining at least one performance prediction model and at least one updated parameter set, parameter recommendations can be made to the distributed storage system based on the at least one performance prediction model and at least one target parameter set. Specifically, in some embodiments, making parameter recommendations to the distributed storage system based on the at least one performance prediction model and at least one target parameter set can include:

[0113] Obtain target system parameters configured for the distributed storage system under a preset storage scenario;

[0114] According to the target performance indicator for parameter recommendation, a target performance prediction model corresponding to the target performance indicator is selected from at least one performance prediction model, and a target parameter set corresponding to the target performance indicator is selected from at least one update parameter set;

[0115] According to the target performance prediction model, the target parameter set and the target system parameters are optimized using a first preset algorithm to obtain an optimized parameter set corresponding to the target performance indicator, and the optimized parameter set is recommended to the distributed storage system.

[0116] It should be noted that the parameter recommendation method provided in the embodiment of the present application can recommend optimal parameters based on the current usage scenario. When recommending parameters to a distributed storage system based on at least one performance prediction model and at least one target parameter set, the target system parameters configured by the distributed storage system under a preset storage scenario are first obtained. The preset storage scenario is the specific storage scenario of the distributed storage system that needs to be created, for example, what storage strategy to adopt, how to configure the hardware, which operating system to choose, etc. The performance indicators emphasized by each storage scenario are different.

[0117] Therefore, based on the preset storage scenario, a target performance indicator for which parameter recommendation is required is obtained. A target performance prediction model corresponding to the target performance indicator is selected from at least one performance prediction model, and a target parameter set corresponding to the target performance indicator is selected from at least one update parameter set. The target performance model is then input into a first preset algorithm, and the target parameter set and target system parameters are optimized using the first preset algorithm to obtain an optimized parameter set corresponding to the target performance indicator. The optimized parameter set is then recommended to the distributed storage system.

[0118] It should be noted that in the embodiments of the present application, the first preset algorithm may be an EM algorithm. The target parameter set and target system parameters are input into the target performance prediction model. The performance prediction model is then input into the EM algorithm. Through iterative calculations of the EM algorithm, an optimized parameter set that optimizes the target performance indicator is obtained. Alternatively, the first preset algorithm may be a Bayesian algorithm, etc., which is not specifically limited in the embodiments of the present application.

[0119] In short, the parameter recommendation method provided by the embodiment of the present application first trains performance prediction models corresponding to different performance indicators. The parameter input set for training the performance prediction distributed storage system is based on IO stack analysis and / or, in the process of training the performance prediction model, the contribution of each parameter to the prediction accuracy of the performance prediction model corresponding to each performance indicator is determined. In this way, while obtaining the performance prediction model, the corresponding relationship between each parameter and the performance indicator can also be obtained. After obtaining the performance prediction model, it is possible to determine which type of performance indicator is more concerned in the current scenario based on the specific storage scenario, such as hardware configuration, storage strategy, etc., and then select the performance prediction model corresponding to the performance indicator, and input the parameter input set, system parameters and performance prediction model corresponding to the performance prediction model into the first preset algorithm for iterative calculation, and finally obtain the optimal parameter set in the current storage scenario, and recommend the optimal parameter set to the distributed storage system to create a distributed storage system based on the optimal parameter set.

[0120] That is to say, when making parameter recommendations, the embodiment of the present application can select a corresponding performance prediction model based on the performance indicators emphasized by the actual storage scenario. Based on the performance prediction model and the updated parameter set, it is determined which combination of parameters in the updated parameter set can achieve the best performance indicators. This set of parameter value combinations is recommended as the optimal parameter combination, and a distributed storage system is created based on the recommended optimal parameter combination.

[0121] This embodiment provides a parameter recommendation method, which includes determining a parameter input set of a distributed storage system and system parameters configured in at least one storage scenario; determining at least one performance indicator based on the parameter input set and the system parameters; performing model training based on the parameter input set, the system parameters, and the at least one performance indicator to determine at least one performance prediction model and at least one updated parameter set; wherein each performance indicator corresponds to a performance prediction model and an updated parameter set, and the updated parameter set is obtained after iteratively updating the parameter input set during the model training process; and recommending parameters to the distributed storage system based on the at least one performance prediction model and the at least one updated parameter set. In this way, a performance detection model corresponding to the performance indicator emphasized by the current storage scenario can be selected according to different storage scenarios, and the parameters with the greatest impact on the performance indicator can be determined. Thus, only these important parameters can be optimized in a targeted manner based on the performance prediction model, thereby reducing unnecessary resource consumption and improving the efficiency of parameter optimization. In addition, since the performance prediction model is incorporated into the parameter optimization process, while reducing resource consumption, the coupling of the system is also reduced. Because an accurate performance prediction model has been trained, there is no need to rebuild a distributed storage system to verify its performance indicators when performing parameter optimization. Moreover, the embodiment of the present application can obtain the correspondence between parameters and performance indicators during the training process of the performance prediction model, which helps R&D personnel to carry out targeted transformation of the IO stack and improve optimization efficiency.

[0122] In another embodiment of the present application, see Figure 3 , which shows a schematic diagram of the application architecture of a parameter recommendation system provided by an embodiment of the present application. Figure 3 As shown, the parameter recommendation system can include three major modules: parameter input set construction module 301, performance indicator acquisition module 302 and model training and parameter optimization module 303.

[0123] The parameter input set construction module 301 is used to construct the parameter input set of the distributed storage system. The parameter types can be determined based on the IO stack analysis method or the feature contribution analysis method, and then these types of parameters are randomly sampled to obtain the parameter input set of the distributed storage system.

[0124] The performance indicator acquisition module 302 is used to determine the performance indicator according to the parameter input set and system parameters of the distributed storage system.

[0125] The model training and parameter optimization module 303 is used to train the performance prediction model and optimize the parameters according to the performance prediction model. The model training and parameter optimization module 303 can be divided into two submodules, namely the model training submodule 303A and the parameter optimization submodule 303B. Among them, the model training submodule 303A is used to obtain the contribution of the input parameters to the prediction accuracy of the performance prediction model on the one hand, so as to update the parameter input set after analysis, and on the other hand to obtain the performance prediction model; the parameter optimization submodule 303B is used to optimize the hyperparameters of the second preset algorithm during the training of the performance prediction model.

[0126] The following will be combined Figure 3 The workflow of each module of a parameter recommendation system provided in an embodiment of the present application is explained in detail. It should be noted that the work between each module is not separated independently. For example, when explaining the workflow of the parameter input set construction module 301, it is also necessary to combine the model training and parameter optimization module 303 to specifically explain the workflow.

[0127] like Figure 3 As shown, the parameter input set construction module 301 may correspond to steps S3011 to S3022, specifically as follows:

[0128] S3011. Select parameter type.

[0129] S3012. Randomly sample parameters.

[0130] It should be noted that the embodiment of the present application provides a parameter recommendation method, which is applied to the distributed storage system of SSD that provides resource services as the IaaS layer, and recommends parameters for it. For the distributed storage system of SSD that provides resource services as the IaaS layer, it often involves a large number of configurable parameters. Taking Ceph as an example, it provides thousands of configurable parameters, and each parameter has a certain range of values. If every parameter combination is evaluated, a lot of resources will be wasted; and among these thousands of parameters, not every parameter will have a significant impact on the performance of the distributed storage system. Therefore, only the parameters that have a significant impact on the performance need to be trained as samples, which can have a good effect on parameter optimization; if the parameters that have no impact on the performance or have a very small impact on the performance are also trained, it will undoubtedly cause a waste of system resources and consume more time. Therefore, the embodiment of the present application will first screen out the parameter types that are necessary to be used as training samples.

[0131] Specifically, selecting the parameter type is to filter the parameters, which can be implemented in one or both of the following two ways.

[0132] Method 1: Based on IO stack analysis, by analyzing the IO stack of the distributed storage system, parameters related to the IO stack are identified. These parameters affect the responsiveness of the IO path and, in turn, the performance of the distributed storage system. This method requires developers to have a clear understanding of the distributed storage system architecture and know which parameters correspond to which IO processes. For example, Ceph's design involves dual writes, where the journal is written first and then flushed to disk. Parameters control how the journal is written, whether the cache is used, and when data is flushed, all of which affect the performance of the distributed storage system. Ceph developers can more easily identify parameters related to the IO stack, as these are more likely to affect the performance of the distributed storage system. Parameter types selected based on IO stack analysis can be considered IO stack-related.

[0133] However, distributed storage systems involve a large number of parameters, and their users are not all R&D personnel. This makes parameter selection difficult. If all parameters are included in the performance prediction model training process, the training process will be computationally intensive and require significant computing resources. In this case, method 2 can be used to select parameters.

[0134] Method 2: Based on feature contribution analysis. Feature contribution analysis means that during the training of the performance prediction model, each prediction result can be used to evaluate all parameters involved in the training. Some parameters contribute more to the accuracy of the prediction results, and these parameters are more likely to be selected and enter the next iteration process.

[0135] Specifically: first, all configurable parameters of the distributed storage system are determined as candidate parameters, and then all candidate parameters are randomly sampled within their parameter value ranges, so that the initial parameter input set of the distributed storage system is obtained.

[0136] Next, in the performance indicator acquisition module 302, performance indicators are obtained based on the initial parameter input set and the system parameter input set. The parameter types in the initial parameter input set that have a sudden impact on performance can also be screened out using a mutation point detection method and used as parameters for training the performance prediction model. Subsequently, in the model training and parameter optimization module 303, the performance prediction model is trained based on the initial parameter input set, the system parameter input set, and the performance indicators. This allows the performance prediction model corresponding to each performance indicator to be obtained, as well as the contribution of each parameter to the prediction accuracy of the performance prediction model (also known as the feature contribution). By analyzing the feature contribution, the more important parameter types, i.e., those with a higher contribution to the prediction accuracy of the performance prediction model, can be determined, and the more important parameter types can be selected. The selected parameter types are randomly sampled within the preset parameter value range to obtain a parameter input set, and then the performance indicators are determined based on the parameter input set and the system parameters. Then, the step of training the prediction model is continued to obtain the performance prediction model and the contribution of each parameter to the prediction accuracy of the performance prediction model again. The parameter types with higher contribution to the prediction accuracy of the performance prediction model are determined again, and the parameter input set is updated according to these parameters. Then, the above process is iterated until the number of iterations reaches the preset number of iterations.

[0137] It should be noted that the process of obtaining performance indicators and performing model training will be described in detail in the following steps and will not be repeated here.

[0138] like Figure 3 As shown, the performance indicator acquisition module 302 may correspond to steps S3021 to S3024, specifically as follows:

[0139] S3021. Determine a parameter input set of the distributed storage system.

[0140] S3022. Determine a system parameter input set.

[0141] The parameter input set of the distributed storage system is a parameter input set obtained by randomly sampling the aforementioned parameters of a certain type within their value ranges.

[0142] The system parameter input set may include customized system parameters, such as recommended system configuration parameters for different storage scenarios. Here, different storage scenarios may include bare metal block storage, virtual machine block storage, object storage, all-flash block storage, etc.

[0143] Exemplarily, the customized system parameters may include: maximum thread number limit (kernel.threads-max), maximum process number limit (kernel.pid_max), physical memory usage (vm.swappiness), number of CPUs, memory capacity, etc.

[0144] The parameter recommendation method provided in the embodiment of the present application can determine the optimal parameter combination based on hardware configuration, system settings, storage strategy, and storage usage scenarios, thereby improving the storage performance of the distributed storage system in different usage scenarios.

[0145] From a hardware configuration perspective, storage refers to whether the nodes deployed in the distributed storage system use mechanical hard drives, solid-state drives, or a hybrid deployment. For example, in a common object storage scenario, data disks are typically lower-cost mechanical hard drives, while solid-state drives are used as cache disks. During dual writes, journal writes to SSDs are faster than to HDDs.

[0146] From a system configuration perspective, the performance of the same distributed storage system on different operating systems may vary. For example, traditional x86 systems and ARM systems, as well as whether the operating system enables huge pages and the maximum number of processes, all require adaptation to the distributed storage system to maximize its performance. Typically, empirical values ​​are used, but the parameter recommendation method provided in the embodiments of this application incorporates system configuration into the training process.

[0147] From the perspective of storage scenarios, different storage scenarios have different requirements for performance and configuration. For example, object storage requires storage vendors to provide massive amounts of space, but has relatively low requirements for latency (i.e., read and write delays). Storage scenarios such as backup and migration have relatively loose requirements for data transmission time, while ultra-fast block storage scenarios that provide database services have very high latency requirements. Furthermore, the standard hardware configuration varies in different storage scenarios. Cost-sensitive scenarios will choose more and cheaper mechanical hard drives to provide greater storage space, while latency-sensitive storage scenarios will use higher-performance solid-state drives to minimize the impact of IO bottlenecks.

[0148] From the perspective of storage strategies, a key feature of distributed storage systems is data reliability. Each piece of data will have multiple copies, and this aspect also depends on the storage scenario. For example, object storage that aims to provide the largest possible storage space will generally choose an erasure coding storage strategy, while block storage generally uses a three-copy storage strategy. Furthermore, to ensure data security, the storage strategy also needs to consider how the copies are distributed. Replicas can be placed on different servers in the same rack, on different racks, or even distributed across different data centers. These differences in storage strategies will affect the performance of the distributed storage software, as different storage strategies will result in different communication costs between replicas.

[0149] To provide optimal distributed storage system parameter configuration recommendations based on these factors, we abstract them into parameter variables. For example, to determine whether the operating system uses x86 or ARM, we can write the parameters is_x86 = 1, is_arm = 0. Another example is the storage policy: for triple replication, copy = 3, for erasure coding, copy = 2. Through abstraction, hardware configuration, storage policy, and system configuration can be incorporated into the performance prediction model training process as specific parameters.

[0150] S3023. Run the distributed storage system.

[0151] S3024. Determine performance indicators.

[0152] By using performance testing tools (such as perf, fio, etc.) to simulate the distributed storage system based on the parameter input set of the distributed storage system and the combination of specific parameters in the system parameter input set, the performance indicators of the distributed storage system under different parameter combinations can be obtained.

[0153] Here, the performance metrics of a distributed storage system primarily refer to read and write metrics, and can include at least one or more of the following: IOPS, Mbps, CPU utilization, memory utilization, swap partition (SWAP) usage, and read and write latency. IOPS can include both sequential read and write IOPS and random read and write IOPS. In practical applications, IOPS performance, Mbps performance, and read and write latency performance are typically of particular interest.

[0154] like Figure 3 As shown, the model training and optimization module 303 may correspond to steps S3031 to S3032, specifically as follows:

[0155] S3031. Train a performance prediction model.

[0156] The parameter input set of the distributed storage system, the system parameter input set, and the performance indicators are used as training sets, and a performance prediction model is constructed using a random forest algorithm, thereby obtaining a performance prediction model. In an embodiment of the present application, the obtained performance prediction model may include multiple performance prediction models corresponding to each performance indicator.

[0157] Here, training the performance prediction model is combined with the parameter type selection process in the parameter input set construction module 301. While obtaining multiple performance prediction models corresponding to performance indicators, the contribution of each parameter to the prediction accuracy of the corresponding performance prediction model is determined through a mutation point detection method. Parameter types are then selected according to the aforementioned scheme until the number of iterations is reached. The resulting performance prediction model becomes the final performance prediction model, and the parameter input set determined at this point is the parameter that has a significant impact on the performance of the distributed storage system.

[0158] It should also be noted that, assuming there were originally 50 configurable parameters, after iterative updates, only 8 configurable parameters remain in the parameter input set: Parameter 1, Parameter 2, Parameter 3, ... Parameter 8. This ultimately results in two performance prediction models: prediction model A and prediction model B. Parameters 1, 2, 3, 4, and 5 were used to train prediction model A. Therefore, when recommending parameters using prediction model A, only parameters 1, 2, 3, 4, and 5 need to be optimized to obtain the optimized parameter set for recommendation.

[0159] Parameters 3, 4, 5, 6, 7, and 8 are used to train prediction model B. Therefore, when recommending parameters using prediction model B, only six parameters—3, 4, 5, 6, 7, and 8—need to be optimized to obtain the optimal parameter set for recommendation. This means that the parameters that significantly influence each performance indicator may be different.

[0160] In an embodiment of the present application, the trained performance prediction model can directly participate in the process of optimizing and recommending parameters, eliminating the step of obtaining performance indicators when constructing a new parameter set, which also reduces the tight coupling to a specific cluster to a certain extent.

[0161] This is because, if there is no performance prediction model, in order to implement parameter optimization and recommendation, the distributed storage system must be run every time a performance indicator is obtained, which is a huge workload. However, if the trained prediction model is used in the process of parameter optimization and recommendation, the performance indicator can be obtained directly through the performance prediction model without running the distributed storage system, greatly reducing the workload.

[0162] Furthermore, without a performance prediction model, each result is obtained on a fixed cluster. The optimized and recommended parameters are significantly affected by the cluster itself. While factors such as hardware configuration and storage policy are considered during the training process, the impact of these parameters on the performance of the distributed storage system is determined by running the model on different clusters. However, if a trained performance prediction model is already in place, and the impact of multiple hardware configurations and storage policies is already considered, incorporating this performance prediction model into the parameter optimization and recommendation process can mitigate the impact of the cluster itself on the results. This reduces the coupling between the final recommended parameters and the cluster. In other words, because the performance prediction model is involved in the process of obtaining the optimal recommended parameters, the results are not only accurate on the training cluster, but also universally applicable to other clusters.

[0163] In the embodiments of this application, since multiple performance prediction models are ultimately obtained for different performance indicators, it is possible that a parameter that is suitable for indicator A may be detrimental to indicator B. In this case, it is necessary to select the parameter type and specific model based on the performance indicator of interest. For example, if the model is trained for indicator A, then parameters that are beneficial for A will be selected. This is also why multiple models are trained. In this way, models trained for different indicators can be used to meet different storage requirements.

[0164] During the training of a performance prediction model, the random forest algorithm can directly determine the contribution of parameters to the model's prediction accuracy. Parameters whose prediction accuracy contribution exceeds a preset contribution threshold are selected as sample parameters. The preset contribution threshold can be customized, for example, 50%, depending on the number of iterations. Typically, a threshold setting of approximately 1 / 5 to 1 / 2 of the parameters can be retained in each iteration, indicating that this threshold is reasonable.

[0165] S3032. Perform hyperparameter optimization.

[0166] During the training of the performance prediction model, the hyperparameters of the random forest can also be optimized by a hyperparameter optimization algorithm. That is, in the embodiment of the present application, step S3031 and step S3032 can be performed simultaneously. For example, the optimization of the hyperparameters can be achieved by Bayesian optimization, for example:

[0167] Input: f, X

[0168] D←initSample(f,X)

[0169] for i←|D|to T do

[0170] f←buildModel(X)

[0171] x i ←argmax x∈X s(x,f)

[0172] y i ←f(x i )

[0173] D←D∪(x i ,y i )

[0174] end for

[0175] or

[0176] Input: f,X,S,M

[0177] D←initSample(f,X)

[0178] for i←|D|to T do

[0179] p(y|x,D)←buildModel(M,D)

[0180] xi←argmaxx∈Xs(x,p(y|x,D))

[0181] y i ←f(x i )

[0182] D←D∪(x i ,y i )

[0183] end for

[0184] Where X is all the parameters of random forest;

[0185] f represents the objective function, which is the mean AUC of the random forest model after five cross-validations.

[0186] Bayesian optimization of hyperparameters is a common technical method in this field and will not be described in detail here.

[0187] In this way, the final performance prediction model is obtained through the second preset algorithm (such as the random forest algorithm) and the hyperparameter optimization algorithm (such as the Bayesian optimization algorithm).

[0188] After obtaining the final performance prediction model, it is only necessary to input the performance prediction model and the system parameters of the storage scenario for which parameter recommendation is required into the first preset algorithm. By performing calculations through the first preset algorithm, the recommended parameters matching the current storage scenario can be obtained.

[0189] That is to say, through the parameter recommendation system provided in the embodiment of the present application, an optimized parameter set and a performance prediction model can be finally obtained. According to the performance prediction model, even if parameter recommendations are required for a new storage scenario, it is only necessary to input the system parameters, parameter input set and performance prediction model of the new storage scenario into the first preset algorithm to obtain the optimized parameter set.

[0190] In an embodiment of the present application, in a specific usage scenario, a corresponding specific performance prediction model is selected to recommend the optimal parameters. In this way, the configuration of the distributed storage system for different usage scenarios (including hardware configuration, system configuration, etc.) can be obtained, such as whether it is equipped with SSD, memory size, number of CPUs, number of process limits, etc. These parameters are fixed, and the required prediction model is selected, and the parameter optimization algorithm is run. The optimized parameter settings of the distributed storage system are the final recommended parameters. For example, for a usage scenario that is sensitive to latency, a performance prediction model trained with latency as the performance indicator is selected, and the parameter optimization algorithm is run. The parameter set corresponding to the obtained optimal latency result is the recommended applicable parameter setting for this scenario.

[0191] Exemplarily, the first preset algorithm is the EM algorithm, and the parameter recommendation process is as follows:

[0192] X: Initial parameter set

[0193] F: Performance prediction model

[0194] S: Adjust X value

[0195] M(X,y): The set of X that meets Y_exception

[0196] X_perfect: Recommended parameter values

[0197] Y_exception = [define]: performance indicator required to be achieved

[0198] for(n)

[0199] y=F(X)

[0200] X=S(X)

[0201] when y in Y_exception:

[0202] M.append(X,y)

[0203] X_perfect=argmaxM(X,y)

[0204] In this way, by running the above algorithm, the best parameter recommendation that suits the current scenario can be obtained.

[0205] In summary, the parameter recommendation method provided in the embodiment of the present application can be generally divided into three parts: constructing a parameter input set; obtaining performance indicators; model training and hyperparameter optimization.

[0206] Part 1: Constructing a parameter input set. Taking Ceph as an example, distributed storage systems have numerous configurable parameters, potentially reaching thousands, and each parameter has its own specific range of values. Evaluating every parameter combination would undoubtedly waste significant resources. Furthermore, not all parameters significantly impact the performance of a distributed storage system. Considering only those parameters that significantly impact performance can significantly improve optimization. Therefore, two approaches are possible for selecting a parameter input set. One approach involves analyzing the distributed storage system's I / O stack to identify I / O stack-related parameters. These parameters affect I / O path response and, consequently, performance. However, this approach requires developers to have a deep understanding of the distributed storage system's implementation. Another approach involves first randomly sampling all parameters of the distributed storage system as candidate parameters within a range of values ​​to construct a test sample. This then proceeds to Part 2 to obtain performance metrics, and then to Part 3 to construct a performance prediction model. While obtaining the performance prediction model, the contribution of each parameter to the model's prediction accuracy can be determined. By analyzing these contributions (i.e., feature contribution analysis), a more important parameter set can be selected.

[0207] The second part is to obtain performance indicators. This part mainly uses the parameter set obtained in the first part and customized system parameters (for example, the maximum number of threads (kernel.threads-max), the maximum number of processes (kernel.pid_max), how to use physical memory (vm.swappiness), etc.), and performance testing tools such as perf and fio to obtain performance indicators. The performance indicators are mainly read and write indicators, such as IOPS, Mbps, CPU utilization, memory utilization, SWAP usage, and read and write latency.

[0208] Part 3: Training and Optimization, this part can be further divided into two parts, training performance prediction model and parameter optimization. First of all, the purpose of training performance prediction model is not only to do feature contribution analysis, but also to reduce resource consumption. The performance prediction model trained by the sampling method can directly participate in the parameter optimization process, eliminating the step of obtaining performance indicators when constructing a new parameter set, which reduces the tight coupling to a certain extent. The prediction method can be selected as random forest, which can well solve the problem of nonlinear relationship between parameter values ​​and performance indicators. The parameter optimization / recommendation algorithm selected in the embodiment of the present application can obtain an approximate optimal solution at a very low evaluation cost, and the demand of actual application is that the training set is not rich enough, because obtaining performance evaluation indicators under parameter combinations is a relatively resource-consuming process. Whether adding the prediction model to the parameter optimization process or selecting the parameter optimization / recommendation algorithm is based on this demand.

[0209] In short, the parameter recommendation method provided in the embodiment of the present application can be implemented by the following steps:

[0210] Step 1: Construct the parameter input set: Randomly sample the parameter value range.

[0211] Step 2: System parameter input set: recommended system configuration parameters for several scenarios (bare metal block storage, virtual machine block storage, object storage, and all-flash block storage).

[0212] Step 3: Build a distributed storage cluster (i.e., distributed storage system) and obtain performance indicators: such as sequential read and write IOPS, random read and write IOPS, latency (i.e., read and write delay), and throughput.

[0213] Step 4: Use the parameter input set and performance indicator results as the training set to build a performance prediction model through the random forest algorithm. A performance prediction model will be obtained for each performance indicator, as well as the contribution of each parameter to the prediction accuracy of the corresponding model.

[0214] Step 5: Select the parameters whose contribution exceeds 50% in Step 4 and repeat Steps 1-4 to obtain a new performance prediction model and its contribution to prediction accuracy. Steps 1-5 can be iterated using the EM algorithm until the final performance prediction model is obtained. This process not only generates the performance prediction model used in the parameter optimization step, but also provides a correlation between parameters and performance indicators, helping R&D personnel to tailor the software I / O stack and improve optimization efficiency.

[0215] During the training of the performance prediction model, the hyperparameters of the random forest can also be optimized using the Bayesian algorithm.

[0216] Step 6: Parameter Recommendation: Select the corresponding performance prediction model based on the actual scenario, input the system parameters and performance prediction model into the first preset algorithm, and obtain the optimal parameter set through iterative calculation.

[0217] Compared with the solutions of related technologies, the parameter recommendation method provided in the embodiments of the present application has at least the following advantages: most of the existing parameter optimization solutions are tightly coupled with the system. In theory, a parameter optimization is required for each different system configuration, which consumes more resources. Some parameter optimization solutions also require the tuner to have a deeper understanding of the distributed storage system itself, and the learning cost is relatively high. The embodiments of the present application aim to automate the parameter optimization system from parameter set selection to model building to parameter recommendation as much as possible by reducing manual intervention. Moreover, since a non-black box algorithm is selected: such as the random forest algorithm in the embodiments of the present application, these algorithms with strong interpretability can obtain an in-depth understanding of each parameter through analysis of the calculation process, and can also be used for performance bottleneck analysis, providing reasonable ideas for the subsequent optimization of the distributed storage system.

[0218] In other words, the key points of this application proposal are:

[0219] (1) Provide recommendations for optimal distributed storage system parameter configurations based on different storage design usage scenarios and different configurations.

[0220] (2) Add the performance prediction model to parameter optimization to reduce resource consumption and reduce the coupling of the system.

[0221] (3) During the training process of the performance prediction model, the corresponding relationship between parameters and performance indicators can be obtained, which helps R&D personnel to carry out targeted transformation of the IO stack of the distributed storage system and improve optimization efficiency.

[0222] This embodiment provides a parameter recommendation method. The specific implementation of the aforementioned embodiment is elaborated in detail through the above embodiment. It can be seen that the parameter recommendation method proposed in this embodiment can select the corresponding performance prediction model according to the actual storage scenario, optimize the parameters that have a greater impact on the performance indicators corresponding to the performance model, and recommend parameters, thereby realizing targeted parameter recommendations; each link from parameter selection, model establishment, and parameter recommendation can be realized automatically, reducing the impact of manual intervention and no longer limited by the professional level of professionals themselves; since the prediction model is added to the parameter optimization process, it not only reduces the system's resource consumption, but also reduces the system's coupling. At the same time, according to the correspondence between the parameters and performance indicators obtained during the model training process, R&D personnel can carry out targeted transformation of the IO stack to further improve optimization efficiency.

[0223] In another embodiment of the present application, see Figure 4 , which shows a schematic diagram of the composition structure of a parameter recommendation device 40 provided in an embodiment of the present application. Figure 4 As shown, the parameter recommendation device 40 may include: a determination unit 401, a training unit 402 and a recommendation unit 403, wherein:

[0224] The determining unit 401 is configured to determine a parameter input set of a distributed storage system and system parameters configured in at least one storage scenario; and determine at least one performance indicator based on the parameter input set and the system parameters;

[0225] A training unit 402 is configured to perform model training based on the parameter input set, the system parameters, and the at least one performance indicator, and determine at least one performance prediction model and at least one update parameter set; wherein each performance indicator corresponds to a performance prediction model and an update parameter set, and the update parameter set is obtained after iterative updating of the parameter input set during the model training process;

[0226] The recommendation unit 403 is configured to recommend parameters to the distributed storage system based on the at least one performance prediction model and the at least one update parameter set.

[0227] In some embodiments, as Figure 5 As shown, the parameter recommendation device 40 may further include an acquisition unit 404 configured to acquire target system parameters configured for the distributed storage system in a preset storage scenario; and select, based on the target performance indicator for parameter recommendation, a target performance prediction model corresponding to the target performance indicator from the at least one performance prediction model, and select a target parameter set corresponding to the target performance indicator from the at least one update parameter set;

[0228] The recommendation unit 403 is further configured to optimize the target parameter set and the target system parameters using the first preset algorithm according to the target performance prediction model, obtain the optimized parameter set corresponding to the target performance indicator, and recommend the optimized parameter set to the distributed storage system.

[0229] In some embodiments, the determination unit 401 is further configured to perform input and output IO stack analysis on the distributed storage system to determine relevant parameters of the IO stack; and randomly sample the relevant parameters of the IO stack within a preset IO stack parameter value range to obtain a relevant parameter set, and determine the relevant parameter set as the parameter input set.

[0230] In some embodiments, the determination unit 401 is further configured to obtain an initial parameter set of the distributed storage system, and randomly sample the parameters in the initial parameter set within a preset parameter value range to determine a test sample; and determine at least one performance indicator based on the test sample and the system parameters; and use a second preset algorithm to perform model training on the test sample, the system parameters and the at least one performance indicator to determine at least one intermediate performance prediction model and the contribution of each parameter in the test sample to the prediction accuracy of each intermediate performance prediction model; and perform a contribution analysis on the determined prediction accuracy contribution, and select the parameter input set from the test sample based on the analysis results.

[0231] In some embodiments, as Figure 5 As shown, the parameter recommendation device 40 may further include an analyzing unit 405 configured to select a candidate prediction accuracy contribution whose prediction accuracy contribution is greater than a preset contribution threshold from the determined prediction accuracy contribution, and determine a first parameter set using parameters corresponding to the candidate prediction accuracy contribution;

[0232] The training unit 402 is further configured to determine the number of iterations of the at least one performance prediction model; and when the number of iterations is less than a preset number of iterations, determine the first parameter set as the initial parameter set, perform an increment of 1 on the number of iterations, and return to perform the step of randomly sampling the parameters in the initial parameter set within a preset parameter value range to determine the test sample; and when the number of iterations reaches a preset number of iterations, determine the first parameter set as the parameter input set, and determine the model obtained after the most recent iterative update as the at least one performance prediction model.

[0233] In some embodiments, the determination unit 401 is further configured to, after obtaining the first parameter set, determine the prediction accuracy contribution of each parameter in the first parameter set corresponding to each performance indicator based on the prediction accuracy contribution of each parameter in the first parameter set relative to each performance prediction model and the correspondence between the performance prediction model and the performance indicator; and under each performance indicator, select a candidate prediction accuracy contribution whose prediction accuracy contribution is greater than a preset contribution threshold from the determined prediction accuracy contributions, and determine the parameter corresponding to the candidate prediction accuracy contribution as the updated parameter set corresponding to each performance indicator.

[0234] In some embodiments, the determination unit 401 is further configured to build a test cluster corresponding to the at least one storage scenario based on the parameter input set and the system parameters; and during the operation of the test cluster corresponding to the at least one storage scenario, use a performance testing tool to obtain the at least one performance indicator; wherein the performance indicator includes at least some of the following: IOPS, throughput, CPU utilization, memory utilization, SWAP utilization, and read and write latency.

[0235] It is understood that in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular system. Furthermore, the various components in this embodiment can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The aforementioned integrated units can be implemented in the form of hardware or software functional modules.

[0236] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0237] Therefore, this embodiment provides a computer storage medium storing a computer program. When the computer program is executed by at least one processor, the parameter recommendation method of any one of the aforementioned embodiments is implemented.

[0238] Based on the above-mentioned composition of the parameter recommendation device 40 and the computer storage medium, see Figure 6 , which shows a specific hardware structure diagram of a parameter recommendation device 40 provided in an embodiment of the present application. Figure 6As shown, the parameter recommendation device 40 may include: a communication interface 601, a memory 602 and a processor 603; each component is coupled together via a bus system 604. It is understood that the bus system 604 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 604 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 6 Various buses are labeled as bus system 604. Among them, the communication interface 601 is used to receive and send signals in the process of sending and receiving information between other external network elements;

[0239] Memory 602, used to store computer programs that can be run on processor 603;

[0240] The processor 603 is configured to, when running the computer program, execute:

[0241] Determining a parameter input set of a distributed storage system and system parameters configured in at least one storage scenario;

[0242] determining at least one performance indicator based on the parameter input set and the system parameters;

[0243] Model training is performed based on the parameter input set, the system parameters, and the at least one performance indicator to determine at least one performance prediction model and at least one updated parameter set; wherein each performance indicator corresponds to a performance prediction model and an updated parameter set, and the updated parameter set is obtained after iterative updating of the parameter input set during the model training process;

[0244] Parameters are recommended to the distributed storage system based on the at least one performance prediction model and the at least one update parameter set.

[0245] It is understood that the memory 602 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DRRAM). The memory 602 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0246] Processor 603 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 603. The above processor 603 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 602, and processor 603 reads the information in memory 602 and, in conjunction with its hardware, completes the steps of the above method.

[0247] It is understood that the embodiments described herein may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or a combination thereof.

[0248] For software implementation, the techniques described herein can be implemented by modules (e.g., procedures, functions, etc.) that perform the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0249] Optionally, as another embodiment, the processor 603 is further configured to execute the steps of the method in any one of the aforementioned embodiments when running the computer program.

[0250] Based on the composition and hardware structure diagram of the parameter recommendation device 40, see Figure 7 , which shows a schematic diagram of the composition structure of a parameter recommendation device 70 provided in an embodiment of the present application. Figure 7 As shown, the parameter recommendation device 70 at least includes the parameter recommendation apparatus 40 according to any one of the aforementioned embodiments.

[0251] For the parameter recommendation system 70, since corresponding performance prediction models are selected for different storage scenarios, parameter recommendations are made to the distributed storage system based on the parameter set corresponding to the performance prediction model, thereby automatically recommending the optimal parameters for the distributed storage system under different storage scenarios, improving the efficiency of parameter recommendation, and reducing resource consumption.

[0252] The above description is merely a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application.

[0253] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0254] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0255] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0256] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0257] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0258] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A parameter recommendation method, characterized in that: The method comprises: Determining a parameter input set of a distributed storage system and system parameters configured in at least one storage scenario; determining at least one performance indicator based on the parameter input set and the system parameters; Performing model training based on the parameter input set, the system parameters, and the at least one performance indicator to determine at least one performance prediction model and at least one updated parameter set; wherein each performance indicator corresponds to a performance prediction model and an updated parameter set, and the updated parameter set is obtained after iterative updating of the parameter input set during the model training process; Obtain target system parameters configured for the distributed storage system under a preset storage scenario; According to the target performance indicator for parameter recommendation, a target performance prediction model corresponding to the target performance indicator is selected from the at least one performance prediction model, and a target parameter set corresponding to the target performance indicator is selected from the at least one update parameter set; According to the target performance prediction model, the target parameter set and the target system parameter are optimized using a first preset algorithm to obtain an optimized parameter set corresponding to the target performance indicator, and the optimized parameter set is recommended to the distributed storage system; Among them, the determination of the parameter input set of the distributed storage system includes: obtaining the initial parameter set of the distributed storage system, and randomly sampling the parameters in the initial parameter set within a preset parameter value range to determine a test sample; determining at least one performance indicator based on the test sample and the system parameters; using a second preset algorithm to perform model training on the test sample, the system parameters and the at least one performance indicator to determine at least one intermediate performance prediction model and the contribution of each parameter in the test sample to the prediction accuracy of each intermediate performance prediction model; performing a contribution analysis on the determined prediction accuracy contribution, and selecting the parameter input set from the test sample based on the analysis results.

2. The method according to claim 1, characterized in that The step of determining a parameter input set of the distributed storage system further includes: Performing an input / output (IO) stack analysis on the distributed storage system to determine relevant parameters of the IO stack; Random sampling is performed on relevant parameters of the IO stack within a preset IO stack parameter value range to obtain a relevant parameter set, and the relevant parameter set is determined as the parameter input set.

3. The method according to claim 1, characterized in that The performing contribution analysis on the determined prediction accuracy contribution, and selecting the parameter input set from the test sample according to the analysis result, includes: Selecting a candidate prediction accuracy contribution degree whose prediction accuracy contribution degree is greater than a preset contribution degree threshold from the determined prediction accuracy contribution degrees, and determining a first parameter set using parameters corresponding to the candidate prediction accuracy contribution degrees; determining a number of iterations of the at least one performance prediction model; When the number of iterations is less than a preset number of iterations, determining the first parameter set as the initial parameter set, incrementing the number of iterations by 1, and returning to the step of randomly sampling the parameters in the initial parameter set within a preset parameter value range to determine a test sample; When the number of iterations reaches a preset number of iterations, the first parameter set is determined as the parameter input set, and the model obtained after the most recent iterative update is determined as the at least one performance prediction model.

4. The method according to claim 3, characterized in that During the model training process, the method further includes: After obtaining the first parameter set, determining the prediction accuracy contribution of each parameter in the first parameter set corresponding to each performance indicator based on the prediction accuracy contribution of each parameter in the first parameter set to each performance prediction model and the corresponding relationship between the performance prediction model and the performance indicator; Under each performance indicator, a candidate prediction accuracy contribution whose prediction accuracy contribution is greater than a preset contribution threshold is selected from the determined prediction accuracy contribution, and the parameters corresponding to the candidate prediction accuracy contribution are determined as the update parameter set corresponding to each performance indicator.

5. The method according to any one of claims 1 to 3, characterized in that The determining of at least one performance indicator according to the parameter input set and the system parameter includes: Building a test cluster corresponding to the at least one storage scenario according to the parameter input set and the system parameters; During the operation of the test cluster corresponding to the at least one storage scenario, obtaining the at least one performance indicator using a performance testing tool; The performance indicators include at least one of the following: IOPS (number of reads and writes per unit time), throughput, CPU utilization, memory utilization, SWAP utilization, and read and write latency.

6. A parameter recommendation device, characterized in that: The parameter recommendation device includes a determination unit, a training unit, an acquisition unit and a recommendation unit, wherein: The determination unit is configured to determine a parameter input set of a distributed storage system and system parameters configured in at least one storage scenario; and determine at least one performance indicator based on the parameter input set and the system parameters; and is further configured to obtain an initial parameter set of the distributed storage system, and randomly sample the parameters in the initial parameter set within a preset parameter value range to determine a test sample; determine at least one performance indicator based on the test sample and the system parameters; perform model training on the test sample, the system parameters and the at least one performance indicator using a second preset algorithm to determine at least one intermediate performance prediction model and the contribution of each parameter in the test sample to the prediction accuracy of each intermediate performance prediction model; perform a contribution analysis on the determined prediction accuracy contribution, and select the parameter input set from the test sample based on the analysis result; The training unit is configured to perform model training based on the parameter input set, the system parameters, and the at least one performance indicator, and determine at least one performance prediction model and at least one update parameter set; wherein each performance indicator corresponds to a performance prediction model and an update parameter set, and the update parameter set is obtained after iterative updating of the parameter input set during the model training process; The acquisition unit is configured to acquire target system parameters configured for the distributed storage system under a preset storage scenario; and select, based on the target performance indicator for parameter recommendation, a target performance prediction model corresponding to the target performance indicator from the at least one performance prediction model, and select a target parameter set corresponding to the target performance indicator from the at least one update parameter set; The recommendation unit is configured to optimize the target parameter set and the target system parameters using a first preset algorithm based on the target performance prediction model, obtain an optimized parameter set corresponding to the target performance indicator, and recommend the optimized parameter set to the distributed storage system.

7. A parameter recommendation device, characterized in that: The parameter recommendation device includes a memory and a processor, wherein: The memory is used to store a computer program that can be run on the processor; The processor is configured to execute the parameter recommendation method according to any one of claims 1 to 5 when running the computer program.

8. A computer storage medium, characterized in that The computer storage medium stores a computer program, and when the computer program is executed by at least one processor, the parameter recommendation method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Model generation system and method and prediction system

    CN110032551A

  • Training method and device of information recommendation system

    CN110287420A