Storage system optimization method and device, nonvolatile storage medium and electronic equipment
By building a performance prediction model of Ceph storage system, and automatically adjusting parameters using an adaptive particle swarm optimization algorithm, the problems of low manual adjustment efficiency and waste of resources in Ceph storage system are solved, and the read and write performance of the system is improved.
Patent Information
- Application Number
- CN202510324585.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art manual adjustment of Ceph storage system parameters leads to problems such as low adjustment efficiency and serious waste of system resources.
By determining the configuration parameter type and value range of the target storage system, the data set is constructed, the first and second types of individuals are generated, and the long and short-term memory network and an adaptive particle swarm optimization algorithm based on cosine perturbation and reverse learning are used to iteratively determine the target configuration parameter set and automatically optimize the storage system.
It realizes automatic and efficient parameter configuration, improves the read and write performance of the storage system, and solves the problems of low adjustment efficiency and waste of resources.
Smart Images

Figure CN120335714A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of parameter tuning, and more particularly, to a method and apparatus for optimizing a storage system, a non-volatile storage medium, and an electronic device. Background Art
[0002] Ceph: Ceph is an open-source distributed storage system that provides various storage services such as object storage, block devices, and file systems. The design goal of Ceph is high reliability, high scalability, and high performance, which can meet the needs of large-scale data centers. However, the default parameters of Ceph cannot fully utilize the read and write performance of the system, and manual parameter adjustment will lead to low efficiency and waste a large amount of system resources.
[0003] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0004] Embodiments of the present application provide a method and apparatus for optimizing a storage system, a non-volatile storage medium, and an electronic device, so as to at least solve the technical problems of low adjustment efficiency and serious waste of system resources caused by manually adjusting storage system parameters in the prior art.
[0005] According to one aspect of the embodiments of the present application, a method for optimizing a storage system is provided, including: determining the types of configuration parameters of a target storage system, and the value ranges of the parameter values corresponding to each type of configuration parameter; determining a data set based on the types of configuration parameters and the value ranges of the parameter values, where the data set includes a plurality of configuration parameter sets, and each configuration parameter set includes the parameter values corresponding to each type of configuration parameter; determining each configuration parameter set as a first type of individual, and determining a second type of individual based on the first type of individual and the value ranges of the parameter values, where the second type of individual is the reverse individual of the first type of individual; determining the storage system performance prediction values corresponding to the first type of individual and the second type of individual respectively, and selecting initial individuals from the first type of individual and the second type of individual based on the storage system performance prediction values to form an initial population; iteratively determining a target configuration parameter set based on the initial population, and optimizing the target storage system according to the target configuration parameter set.
[0006] Optionally, determining the second type of individual based on the first type of individual and the value ranges of the parameter values includes: determining the median of the parameter values corresponding to each type of configuration parameter according to the value ranges of the parameter values; determining the second parameter values of each type of configuration parameter of the second type of individual corresponding to the first type of individual based on the first parameter values of each type of configuration parameter of the first type of individual and the median of the parameter values of each type of configuration parameter, where the average value of the first parameter value and the second parameter value of the same type of configuration parameter is the median of the parameter values.
[0007] Optionally, determining a set of target configuration parameters based on iteration of an initial population, and optimizing a target storage system according to the set of target configuration parameters includes: determining an initial position and a random initial velocity of each individual in the initial population, where the random initial velocity includes a moving direction and a moving distance of the individual in each dimension of the solution space, the dimensions correspond one-to-one with the types of configuration parameters, and the position of the individual in the solution space corresponds to the parameter value of the individual; iteratively updating the positions of each individual in the initial population according to the initial position and the random initial velocity, where the fitness index of the individual is the predicted value of the storage system performance corresponding to the individual; after reaching the maximum number of iterations, determining the set of target configuration parameters according to the globally optimal individual in the population, where the population is obtained by iteratively updating the positions of the individuals in the initial population, and the globally optimal individual is the individual with the largest fitness index in the population.
[0008] Optionally, iteratively updating the positions of each individual in the initial population according to the initial position and the random initial velocity includes: during the t-th iteration, determining a random velocity perturbation amount according to a sine function or a cosine function, where t is a positive integer not less than 2; determining the position of the globally optimal individual in the (t - 1)-th iteration; determining the velocities of each individual in the (t - 1)-th iteration and the historical optimal positions of each individual after the (t - 1)-th iteration, where the historical optimal position is the position corresponding to the historical maximum fitness index of the individual; determining the velocities of each individual during the t-th iteration according to the velocities of each individual in the (t - 1)-th iteration, the random velocity perturbation amount, the position of the globally optimal individual, and the historical optimal positions; determining the positions of the individuals after the t-th iteration according to the velocities of each individual.
[0009] Optionally, determining the velocities of each individual during the t-th iteration according to the random velocity perturbation amount, the position of the globally optimal individual, and the historical optimal positions includes: determining a velocity weight at the t-th iteration, where the velocity weight is used to determine the retention amount of the velocities of each individual in the (t - 1)-th iteration in the t-th iteration, and the value of the velocity weight decreases as the number of iterations increases; determining the velocities of each individual during the t-th iteration according to the velocity weight, the velocities of each individual in the (t - 1)-th iteration, the random velocity perturbation amount, the position of the globally optimal individual, and the historical optimal positions.
[0010] Optionally, determining a set of target configuration parameters based on iteration of an initial population, and optimizing a target storage system according to the set of target configuration parameters includes: determining the globally optimal individual in the iterated population, where the globally optimal individual is the individual with the largest fitness index, and the fitness index is the predicted value of the storage system performance corresponding to the individual; determining the set of configuration parameters corresponding to the globally optimal individual as the set of target configuration parameters; updating the configuration parameter values of the target storage system to the parameter values in the set of target configuration parameters.
[0011] Optionally, determining the storage system performance prediction values corresponding to the first type of individuals and the second type of individuals respectively includes: determining the application scenario information of the target storage system; determining the target performance type according to the application scenario information; determining the performance prediction model corresponding to the target performance type according to the target performance type, and determining the storage system performance prediction values corresponding to the first type of individuals and the second type of individuals respectively through the performance prediction model.
[0012] According to another aspect of the embodiments of the present application, there is also provided a storage system optimization device, including: a first processing module, configured to determine the configuration parameter types of the target storage system and the value ranges of the parameter values corresponding to each configuration parameter type; a second processing module, configured to determine a data set according to the configuration parameter types and the value ranges of the parameter values, where the data set includes multiple configuration parameter sets, and each configuration parameter set includes the parameter values corresponding to each configuration parameter type; a third processing module, configured to determine each configuration parameter set as the first type of individuals, and determine the second type of individuals according to the first type of individuals and the value ranges of the parameter values, where the second type of individuals are the reverse individuals of the first type of individuals; a fourth processing module, configured to determine the storage system performance prediction values corresponding to the first type of individuals and the second type of individuals respectively, and select initial individuals from the first type of individuals and the second type of individuals according to the storage system performance prediction values to form an initial population; a fifth processing module, configured to iteratively determine a target configuration parameter set according to the initial population, and optimize the target storage system according to the target configuration parameter set.
[0013] According to another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium, in which a program is stored, and when the program runs, it controls the device where the non-volatile storage medium is located to execute the storage system optimization method.
[0014] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, where the processor is configured to run the program stored in the memory, and when the program runs, it executes the storage system optimization method.
[0015] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the storage system optimization method.
[0016] In the embodiments of the present application, the types of configuration parameters of the target storage system are determined, as well as the value ranges of the parameter values corresponding to each type of configuration parameter; a data set is determined according to the type of configuration parameter and the value range of the parameter value, wherein the data set includes multiple configuration parameter sets, and each configuration parameter set includes the parameter values corresponding to each type of configuration parameter; each configuration parameter set is determined as a first type of individual, and a second type of individual is determined according to the first type of individual and the value range of the parameter value, wherein the second type of individual is the reverse individual of the first type of individual; the storage system performance prediction values corresponding to the first type of individual and the second type of individual are determined, and initial individuals are selected from the first type of individual and the second type of individual according to the storage system performance prediction values to form an initial population; the target configuration parameter set is iteratively determined according to the initial population, and the target storage system is optimized according to the target configuration parameter set. By combining the long short-term memory network with the adaptive particle swarm optimization algorithm based on sine-cosine perturbation and reverse learning, a performance prediction model of the storage system parameters is constructed and intelligently tuned, achieving the purpose of automatically and efficiently optimizing the parameter configuration, thereby realizing the technical effect of improving the read and write performance of the storage system, and further solving the technical problems of low adjustment efficiency and serious waste of system resources caused by manually adjusting the storage system parameters in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0018] Figure 1 is a schematic structural diagram of a computer terminal provided according to an embodiment of the present application;
[0019] Figure 2 is a schematic flowchart of a storage system optimization method provided according to an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of the influence of batch size on the model accuracy provided according to an embodiment of the present application;
[0021] Figure 4 is a schematic diagram of the comparison result between the predicted values and the true values of LSTM and RF provided according to an embodiment of the present application;
[0022] Figure 5 is a schematic flowchart of a storage system optimization method provided according to an embodiment of the present application;
[0023] Figure 6 is a schematic structural diagram of a storage system optimization device provided according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To enable those skilled in the art to better understand the solution of this application, the technical solution in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0026] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained as follows:
[0027] LSTM (Long Short-Term Memory, LSTM): It is a time recurrent neural network, which is specifically designed to solve the long-term dependence problem existing in general RNNs (recurrent neural networks). All RNNs have a chain form of repeating neural network modules.
[0028] In the related art, the performance optimization work of Ceph systems by domestic and foreign research scholars mainly focuses on three aspects: specific hardware environment optimization, application scenario-oriented optimization, and internal mechanism optimization. In terms of specific hardware environment optimization, with the emergence of NVDIMM (Non-Volatile Dual In-line Memory Module) products, byte-addressable non-volatile memory will provide IO performance similar to that of memory. The performance of using NVDIMM as the underlying medium in the Ceph system was simulated. For a single node, by mapping all content to NVDIMM, the throughput can be increased by more than 100%.
[0029] In the optimization for application scenarios, in the field of high-performance computing, Ceph is not the most suitable storage system. The files accessed by data-intensive applications in high-performance computing are classified into read-intensive, write-intensive, or read-write-intensive. The read-write characteristics of these files are used to set file placement decisions to balance high-performance computing workloads. In the field of cloud computing, data logging operations of data objects are carried out while maintaining write atomicity and reliability. Experimental results show that the capacity provided by the new storage engine is more than three times that of the original.
[0030] Although the above methods for specific hardware environment optimization and application scenario optimization have made certain progress in performance improvement, they do not consider the general environment and ignore the performance improvement space brought by the internal parameter tuning of Ceph. In the above research on internal mechanism optimization methods, they are generally applicable to performance improvement, but do not fully consider the non-linear relationship of parameters. Currently, the internal parameter tuning of Ceph mainly relies on manual adjustment based on experience, which is inefficient and wastes a large amount of resources.
[0031] To solve this problem, relevant solutions are provided in the embodiments of this application, which are described in detail below.
[0032] According to the embodiments of this application, a method embodiment of a storage system optimization method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0033] The method embodiments provided by the embodiments of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a storage system optimization method is shown. As Figure 1 shown, the computer terminal 10 (or mobile device 10) may include one or more (shown as 102a, 102b,..., 102n in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the computer terminal 10 may further include more than Figure 1more or fewer components as shown, or having a configuration different from that shown in Figure 1 that shown.
[0034] It should be noted that one or more of the above-mentioned processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0035] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the storage system optimization method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned storage system optimization method. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include memories remotely set relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0036] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0037] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 10 (or mobile device).
[0038] Under the above operating environment, the embodiments of the present application provide a storage system optimization method, as Figure 2 shown, the method includes the following steps:
[0039] Step S202: Determine the types of configuration parameters of the target storage system and the value ranges of the parameters corresponding to each configuration parameter type.
[0040] Optionally, for the Ceph storage system, it has a total of 8 configuration parameters, namely bluestore_cache_size_ssd, bluestore_cache_size_hdd, bluestore_cache_meta_ratio, bluestore_cache_kv_ratio, osd_max_write_size, osd_map_cache_size, rbd_cache_size, and rbd_cache_max_dirty; the types of bluestore_cache_size_ssd and bluestore_cache_size_hdd are integer, the types of bluestore_cache_meta_ratio and bluestore_cache_kv_ratio are float, and the types of osd_max_write_size, osd_map_cache_size, rbd_cache_size, and rbd_cache_max_dirty are integer. The 8 parameters of Ceph are shown in the following table.
[0041] Parameter Name Type Value Range bluestore_cache_size_ssd integer 1GB - 10GB bluestore_cache_size_hdd integer 1GB - 10GB bluestore_cache_meta_ratio float 0~1 bluestore_cache_kv_ratio float 0~1 osd_max_write_size integer 4~2000 osd_map_cache_size integer 64~1024 rbd_cache_size integer 1MB - 64MB rbd_cache_max_dirty integer 1MB - 64MB
[0042] Among them, the meanings of each parameter are as follows: bluestore_cache_size_ssd: Defines the size of the SSD cache in Ceph BlueStore, affecting read and write performance. bluestore_cache_size_hdd: Specifies the size of the HDD cache in Ceph BlueStore, optimizing storage efficiency. bluestore_cache_meta_ratio: Controls the proportion of metadata in the Ceph BlueStore cache, affecting data access speed. bluestore_cache_kv_ratio: Adjusts the proportion of key-value data in the Ceph BlueStore cache, optimizing cache usage. osd_max_write_size: Sets the maximum data block size for Ceph OSD write operations, affecting write efficiency and concurrency. osd_map_cache_size: Determines the size of the Ceph OSD metadata mapping cache, affecting metadata access speed. rbd_cache_size: Controls the size of the Ceph RBD cache, improving read performance. rbd_cache_max_dirty: Sets the maximum amount of dirty data allowed in the RBD cache that has not been synchronized to the disk, balancing read and write performance.
[0043] Step S204: Determine a data set according to the configuration parameter type and the parameter value range. The data set includes multiple configuration parameter sets, and each configuration parameter set includes parameter values corresponding to each configuration parameter type.
[0044] As an optional implementation, randomly assign values to 8 configuration parameters of Ceph within the adjustable range. Let the value range of the i-th parameter conf i be [lb i , ub i , conf i = random(lb i , ub i ), where i = 1, 2,..., 8. Take the configuration parameter set config = {conf1, conf2, conf3, conf4, conf5, conf6, conf7, conf8} as an item in the data set.
[0045] Step S206: Determine each configuration parameter set as the first type of individuals, and determine the second type of individuals according to the first type of individuals and the parameter value range. The second type of individuals are the reverse individuals of the first type of individuals.
[0046] In the technical solution provided in Step S206, determining the second type of individuals according to the first type of individuals and the parameter value range includes: determining the median of the parameter values corresponding to each configuration parameter type according to the parameter value range; determining the second parameter values of each configuration parameter type of the second type of individuals corresponding to the first type of individuals according to the first parameter values of each configuration parameter type of the first type of individuals and the median of the parameter values of each configuration parameter type. The average value of the first parameter value and the second parameter value of the same configuration parameter type is the median of the parameter values.
[0047] Optionally, determining the reverse individuals (i.e., the second type of individuals) includes: calculating the midpoint of the parameter values corresponding to each configuration parameter type (i.e., the median of the parameter values) based on the given parameter value range. This midpoint can be regarded as the center point or reference point in the parameter space. For example, for a parameter rbd_cache_size with a value range of 1MB to 64MB, its midpoint is 32.5MB, which provides a fixed coordinate reference for subsequent operations.
[0048] Next, based on the parameter values of the first type of individuals (the first parameter values) and the midpoint of the previously calculated parameter values, determine the parameter values of the second type of individuals corresponding to the first type of individuals (the second parameter values). This process actually involves a symmetric mapping, where the calculation of the second parameter values ensures that they are symmetric with respect to the midpoint of the parameter values with the first parameter values. The specific calculation method is to subtract the difference between the first parameter values and the midpoint from the midpoint to obtain the second parameter values. This means that if the parameter values of the first type of individuals are regarded as a point in the parameter space, the parameter values of the second type of individuals are the mirror image points of this point with respect to the midpoint of the parameter values. This operation ensures that for any parameter configuration of the first type of individuals, the parameter configuration of its corresponding reverse individuals (the second type of individuals) is numerically symmetric with it, and their average value is exactly the midpoint of the parameter values. For example, for a parameter rbd_cache_size with a value range of 1MB to 64MB, the midpoint is 32.5MB. When the parameter value of the first type of individuals is 32MB, the parameter value of the corresponding second type of individuals is 33MB.
[0049] By generating such a set of reverse individuals, not only the diversity of the population is increased, but also in the optimization process, by comparing the performance of the first type of individuals and the second type of individuals, select the individuals with better performance for subsequent optimization operations. This strategy is particularly important because it can avoid the parameter configuration of the population from being too concentrated in the early stage of the algorithm, thereby improving the global search efficiency of the algorithm, reducing the dependence on local optimal solutions, and further promoting the algorithm to find a parameter configuration combination closer to the global optimum faster.
[0050] In summary, the generation and utilization of reverse individuals are key steps in the parameter tuning process. It enhances the global optimization ability of the algorithm through symmetric mapping of parameter values.
[0051] Step S208, determine the predicted values of the storage system performance corresponding to the first type of individuals and the second type of individuals respectively, and select initial individuals from the first type of individuals and the second type of individuals according to the predicted values of the storage system performance to form an initial population.
[0052] In the technical solution provided in step S208, determining the predicted values of the storage system performance corresponding to the first type of individuals and the second type of individuals respectively includes: determining the application scenario information of the target storage system; determining the target performance type according to the application scenario information; determining the performance prediction model corresponding to the target performance type according to the target performance type, and determining the predicted values of the storage system performance corresponding to the first type of individuals and the second type of individuals respectively through the performance prediction model.
[0053] Optionally, determining the application scenario information of the target storage system is a key first step in optimizing storage performance. Different scenarios have very different requirements for storage systems. For example, in scenarios such as cloud gaming or online transaction processing, applications may focus more on system random read performance to quickly respond to user requests and provide a smooth experience. In big data analysis or database environments, random write IOPS (input and output operations per second) may become the focus because these applications need to frequently update and store large amounts of data. Understanding the core requirements of these application scenarios will help to select and build performance prediction models in a targeted manner.
[0054] As an optional implementation, if the application scenario information of the target storage system is big data analysis, the target performance type is determined to be random write IOPS based on the application scenario information, and the corresponding prediction model is determined to be a random write performance prediction model based on the target performance type.
[0055] Optionally, before determining the storage system performance prediction values corresponding to the first and second individuals, the performance prediction model is trained: randomly selecting values for the eight configuration parameters of Ceph within an adjustable range, assuming that the i-th parameter conf i The value range is [lb i ,ub i ],conf i =random(lb i ,ub i ), i = 1, 2, ..., 8, set the configuration parameter set config = {conf1, conf2, conf3, conf4, conf5, conf6, conf7, conf8}, and test the corresponding Ceph block storage system read and write performance; parameter combination config i And the corresponding ipos i Constitute a data item (config i ,ipos i ), and all the collected data items are used as the data set for building the Ceph performance prediction model.
[0056] Optionally, the IOPS value is divided into six indicators: random read IOPS, random write IOPS, sequential read IOPS, sequential write IOPS, mixed sequential read and write IOPS, and mixed random read and write IOPS; the method for collecting the IOPS value corresponding to the combined config is:
[0057] Step 1: Use random(lb i ,ub i ), i=1, 2, ..., 8 are 8 random parameter values;
[0058] Step 2: Use the cluster management tool Ansible to synchronize the modified configuration parameters to the entire Ceph cluster;
[0059] Step 3: Use the fio+rbd test tool to obtain the performance of the block storage system;
[0060] Step 4: Use the crontabs tool to execute the test tasks regularly, repeat steps 1-3, collect parameter combinations and corresponding IOPS values, and train to obtain performance prediction models corresponding to the corresponding IOPS performance metrics, namely: random read IOPS performance prediction model, random write IOPS performance prediction model, sequential read IOPS performance prediction model, sequential write IOPS performance prediction model, mixed sequential read and write IOPS performance prediction model, and mixed random read and write IOPS performance prediction model. The above models are applied to different scenarios for performance prediction of corresponding target performance types.
[0061] Optionally, the error model of the performance prediction model is as follows:
[0062]
[0063] where Error is the prediction error, Actual i is the true value of the Ceph block storage system performance, and Forecast i is the performance value predicted by the LSTM model, and n is the number of samples.
[0064] Optionally, in order to improve the accuracy of the Ceph performance prediction model and reduce the training duration, it is necessary to adjust the batchsize of the LSTM model. Batchsize is the number of samples selected for one training in the neural network. The size of this parameter affects the optimization degree and speed of the model and directly affects the use of GPU and memory. If the batchsize is too small, the gradient change will fluctuate greatly, which will cause the network to be not easy to converge. If the parameter is set too large, it will lead to too high memory capacity, inaccurate gradient, and longer time consumption. Figure 3 Shows the influence of batchsize on the model accuracy. As Figure 3 shown, the size of batchsize is determined by the experimental method.
[0065] According to Figure 3 it can be seen that as the batchsize increases, the accuracy of the model rises. When batchsize = 32, the model accuracy reaches the maximum. When batchsize is greater than 32, the model accuracy decreases. And the training duration will gradually increase after batchsize is greater than 32. According to the experimental results, selecting batchsize 32 can achieve the optimal training effect.
[0066] In Figure 4 , the comparison results between the predicted values and the true values of LSTM and RF (Random Forest) are shown. The abscissa represents different parameter configurations, and the ordinate represents the IOPS value. LSTM and RF respectively represent the predicted values of the performance models established by LSTM and RF. From the overall trend, the predicted values obtained by using LSTM and RF can both timely reflect the performance fluctuations caused by parameter changes. However, there are obvious differences between the predicted curve obtained by RF and the true value curve, and the deviation from the true value is relatively large at some moments. In order to intuitively compare the accuracy differences between the two models, the present invention uses Error to evaluate the accuracy of the models. Through comparative experiment calculations, the Error values of the RF and LSTM models are 0.56% and 0.28% respectively. Thus, it can be seen that the prediction accuracy of LSTM is better than that of the existing method RF.
[0067] Step S210: Iteratively determine the set of target configuration parameters based on the initial population, and optimize the target storage system according to the set of target configuration parameters.
[0068] In the technical solution provided in step S210, iteratively determining the set of target configuration parameters based on the initial population and optimizing the target storage system according to the set of target configuration parameters includes: determining the initial positions and random initial velocities of each individual in the initial population, where the random initial velocity includes the moving direction and moving distance of the individual in each dimension of the solution space, the dimension corresponds one-to-one with the type of configuration parameter, and the position of the individual in the solution space corresponds to the parameter value of the individual; iteratively updating the positions of each individual in the initial population according to the initial positions and random initial velocities, where the fitness index of the individual is the predicted value of the storage system performance corresponding to the individual; after reaching the maximum number of iterations, determining the set of target configuration parameters according to the globally optimal individual in the population, where the population is obtained by iteratively updating the positions of the individuals in the initial population, and the globally optimal individual is the individual with the largest fitness index in the population.
[0069] Optionally, the random initial velocity is randomly selected within a preset velocity range. The random initial velocity represents the moving direction and step size of the individual in the solution space, reflecting the comprehensive response of the individual to historical experience. The velocity is a multi-dimensional vector, and each dimension corresponds to the adjustment direction and amplitude of a parameter. The velocity directly affects the position of the individual through the position update formula. The position of the individual is determined by adding the updated velocity to the current position. The magnitude and direction of the velocity determine the next search direction of the individual.
[0070] Optionally, iteratively updating the positions of the individuals in the initial population according to the initial positions and the random initial velocities includes: during the t-th iteration, determining a random velocity perturbation amount according to a sine function or a cosine function, where t is a positive integer not less than 2; determining the position of the globally optimal individual in the (t - 1)-th iteration; determining the velocities of the individuals in the (t - 1)-th iteration and the historical optimal positions of the individuals after the (t - 1)-th iteration, where the historical optimal position is the position corresponding to the historical maximum fitness index of the individual; determining the velocities of the individuals during the t-th iteration according to the velocities of the individuals in the (t - 1)-th iteration, the random velocity perturbation amount, the position of the globally optimal individual, and the historical optimal positions; and determining the positions of the individuals after the t-th iteration according to the velocities of the individuals.
[0071] Optionally, determining the velocities of the individuals during the t-th iteration according to the random velocity perturbation amount, the position of the globally optimal individual, and the historical optimal positions includes: determining a velocity weight at the t-th iteration, where the velocity weight is used to determine the retention amount of the velocities of the individuals in the (t - 1)-th iteration in the t-th iteration, and the value of the velocity weight decreases as the number of iterations increases; and determining the velocities of the individuals during the t-th iteration according to the velocity weight, the velocities of the individuals in the (t - 1)-th iteration, the random velocity perturbation amount, the position of the globally optimal individual, and the historical optimal positions.
[0072] Optionally, for the adaptive particle swarm optimization algorithm based on sine-cosine perturbation and reverse learning, the reverse learning strategy is used for population initialization. During the population optimization process, sine-cosine perturbation is used to prevent the population from falling into local optimal solutions, and a weight update strategy with a downward-opening shape is used to enhance the global search ability of the population. Specifically, the method of using SCRLPSO (based on sine and cosine perturbation and reverse learning adaptive particle swarm optimization) for optimization is as follows: Let the population size be N and the maximum number of iterations be T. A set of parameter combinations config = {conf1, conf2, conf3, conf4, conf5, conf6, conf7, conf8} is used as the 8 dimensions in the particle, and each parameter represents a dimension of the particle. The reverse learning strategy is used to generate the reverse population of the initial population in step 1. Using Lose as the fitness index, compare the individuals in the initial population with their corresponding reverse individuals, and select the better individuals as the initial individuals of the population; Iteratively update, using Lose as the fitness index, update the global optimal solution of the population and the historical optimal solution of each individual, and use the hybrid sine-cosine particle swarm optimization algorithm to update the velocity and position of each individual in the population until the maximum number of iterations T is reached to obtain the global optimal solution of the population;
[0073] The pseudo-code of the SCRLPSO algorithm is as follows.
[0074] Input parameters: pms: upper bound of parameter value range
[0075] vms: lower bound of parameter value range
[0076] xbound: particle position range
[0077] vbound: particle velocity range
[0078] ωbound: weight range
[0079] T: number of generations to terminate evolution
[0080] M: population size
[0081] Output parameter: g: optimal solution
[0082]
[0083]
[0084] Among them, \(X(t)\) represents all individuals of the population at time \(t\), including the positions and velocities of the individuals. Lines 3 to 10 represent the population initialization operation. Line 13 updates the weights of the current generation population. Lines 14 to 16 update the positions and velocities of the individuals in the population. Lines 19 to 22 update the global optimal solution of the population and the historical optimal solutions of the individuals.
[0085] Specifically, in the initial stage of the algorithm, first set the current generation \(t = 1\), and then initialize the population, which involves generating \(M\) individuals, each representing a set of parameter configurations. The positions (parameter configuration values) and velocities (parameter adjustment amplitudes) of the individuals are randomly generated within the defined range to ensure the diversity of the population. Subsequently, through the opposition-based learning strategy, each individual is compared with its opposite individual, and the one with better performance is selected as an individual in the population. This process effectively increases the exploration ability of the population and avoids the trap of local optimum.
[0086] Entering the iterative optimization stage, at the beginning of each iteration, the weights of the population are updated according to the current iteration number to dynamically adjust the search strategy and enhance the global search ability. Next, the algorithm updates the positions and velocities of each individual, which is achieved by the hybrid sine-cosine particle swarm optimization algorithm to ensure that the individuals move in a non-linear manner in the parameter space and avoid premature convergence. The updated set of individuals \(X(t + 1)\) replaces the set of individuals \(X(t)\) of the previous generation to maintain the iterative evolution of the population. The algorithm continues this process until the iteration number reaches the preset termination evolution generation \(T\).
[0087] During the iteration process, the algorithm also continuously updates the global optimal solution of the population and the historical optimal solutions of each individual to record the best parameter configurations found so far. The update of the global optimal solution \(g\) ensures that the algorithm can track and retain the best parameter combination in terms of performance throughout the optimization process, while the update of the historical optimal solution \(p\) of the individual helps the algorithm remember the best configuration of the individual during the evolution process. The updates of both promote the algorithm to approach the optimal solution.
[0088] Finally, when the iteration termination condition is met, the algorithm outputs the optimal solution \(g\), and this set of parameter configurations is the optimal solution for the performance optimization of the Ceph system. The entire process demonstrates the powerful function of the SCRLPSO algorithm in parameter optimization, which can intelligently and efficiently explore the parameter space and finally find the parameter configuration that makes the system performance reach the optimal, thus significantly improving the read and write performance of Ceph.
[0089] Optionally, determining a set of target configuration parameters according to the iteration of the initial population, and optimizing the target storage system according to the set of target configuration parameters includes: determining the globally optimal individual in the population after iteration, where the globally optimal individual is the individual with the largest fitness index, and the fitness index is the predicted value of the storage system performance corresponding to the individual; determining the set of configuration parameters corresponding to the globally optimal individual as the set of target configuration parameters; updating the configuration parameter values of the target storage system to the parameter values in the set of target configuration parameters.
[0090] The embodiment of the present application provides a method for optimizing a storage system, as Figure 5 shown. The method includes the following steps: First, data collection is performed. The test workload is tested with random parameter configurations, and the Ceph block storage system is tested using the fio+rbd test tool. The read / write performance of the system is collected to obtain a training data set. The training data set includes multiple configuration combinations and corresponding performance metric values. A Ceph performance prediction model is constructed, and the LSTM model is trained through the training data set. After training the model, the population is initialized, a reverse population is generated based on the reverse learning strategy, the performance values of each individual are obtained through the performance prediction model, and then the optimal individual is retained. When the termination condition is met, the iteration ends and the optimal parameters are returned.
[0091] Optionally, in order to evaluate the effect of the method embodiment of the present application on the performance tuning of the Ceph system, the method embodiment of the present application is compared with the LSTM+GA (Genetic Algorithm) method and the Ceph automatic tuning method based on RF and GA. In order to obtain more accurate experimental results, the average value of 5 algorithm runs is taken as the final experimental result for each method. Through experiments, the read / write performance of the Ceph block storage system after parameter optimization is approximately 7184. There is no significant difference between the LSTM+GA method and the RF+GA method in the first 20 generations, and both methods reach a stable state at around 60 generations, but the RF+GA lags slightly behind the LSTM+GA. The LSTM+SCRLPSO method can reach a stable state at around 40 generations, indicating that the convergence speed of the LSTM+SCRLPSO is faster than that of the LSTM+GA and the RF+GA. Moreover, the optimal value obtained by the LSTM+SCRLPSO method is better than that of the LSTM+GA method and the RF+GA method, indicating that the convergence accuracy of the LSTM+SCRLPSO is higher. Substituting the obtained optimal parameter combination into the Ceph system in the real environment, the average value of the block storage system performance IOPS is measured to be 7184, which is not much different from the predicted value and is within the acceptable range. The performance of the default parameter configuration can only reach 3971, and the performance is about 1.8 times that of the default configuration.
[0092] Through the above steps, the types of configuration parameters of the target storage system are determined, as well as the value ranges of the parameter values corresponding to each type of configuration parameter; a data set is determined based on the type of configuration parameter and the value range of the parameter value, where the data set includes multiple configuration parameter sets, and each configuration parameter set includes the parameter values corresponding to each type of configuration parameter; each configuration parameter set is determined as a first type of individual, and a second type of individual is determined based on the first type of individual and the value range of the parameter value, where the second type of individual is the reverse individual of the first type of individual; the storage system performance prediction values corresponding to the first type of individual and the second type of individual are determined, and initial individuals are selected from the first type of individual and the second type of individual based on the storage system performance prediction values to form an initial population; the target configuration parameter set is iteratively determined based on the initial population, and the target storage system is optimized based on the target configuration parameter set. By combining the long short-term memory network with the adaptive particle swarm optimization algorithm based on sine-cosine perturbation and reverse learning, a performance prediction model of the storage system parameters is constructed and intelligently tuned, achieving the purpose of automatically and efficiently optimizing the parameter configuration, thus realizing the technical effect of improving the read and write performance of the storage system, and further solving the technical problems of low adjustment efficiency and serious waste of system resources caused by manually adjusting the storage system parameters in the prior art. Specifically, the method embodiment of the present application has the following advantages:
[0093] 1. SCRLPSO: Improved PSO algorithm
[0094] Traditional PSO algorithms have disadvantages such as being prone to falling into local optimal solutions and having a slow convergence speed. The proposed SCRLPSO uses a reverse learning strategy for population initialization. During the process of population optimization, sine-cosine perturbation is used to prevent the population from falling into local optimal solutions and enhance the global search ability of the population.
[0095] 2. Application of the LSTM + SCRLPSO parameter optimization method to the Ceph system
[0096] An accurate and reliable Ceph performance prediction model is constructed using LSTM. The prediction value of the performance prediction model is used as the fitness of the population individuals, and the optimal parameter configuration is found through SRCLPSO to optimize the system performance to the best.
[0097] The embodiment of the present application provides an optimization device for a storage system, Figure 6 which is a schematic structural diagram of the device, as Figure 6As shown, the device includes: a first processing module 60, configured to determine the types of configuration parameters of the target storage system and the value ranges of the parameter values corresponding to the respective types of configuration parameters; a second processing module 62, configured to determine a data set according to the types of configuration parameters and the value ranges of the parameter values, where the data set includes a plurality of configuration parameter sets, and each configuration parameter set includes the parameter values corresponding to the respective types of configuration parameters; a third processing module 64, configured to determine each configuration parameter set as a first type of individual, and determine a second type of individual according to the first type of individual and the value ranges of the parameter values, where the second type of individual is the reverse individual of the first type of individual; a fourth processing module 66, configured to determine the storage system performance prediction values corresponding to the first type of individual and the second type of individual respectively, and select initial individuals from the first type of individual and the second type of individual according to the storage system performance prediction values to form an initial population; and a fifth processing module 68, configured to iteratively determine a target configuration parameter set according to the initial population, and optimize the target storage system according to the target configuration parameter set.
[0098] In some embodiments of the present application, the third processing module 64 determines the second type of individual according to the first type of individual and the value ranges of the parameter values, including: determining the median of the parameter values corresponding to the respective types of configuration parameters according to the value ranges of the parameter values; and determining the second parameter values of the respective types of configuration parameters of the second type of individual corresponding to the first type of individual according to the first parameter values of the respective types of configuration parameters of the first type of individual and the median of the parameter values of the respective types of configuration parameters, where the average value of the first parameter value and the second parameter value of the same type of configuration parameter is the median of the parameter values.
[0099] In some embodiments of the present application, the fifth processing module 68 iteratively determines a target configuration parameter set according to the initial population, and optimizes the target storage system according to the target configuration parameter set, including: determining the initial positions and random initial velocities of the individuals in the initial population, where the random initial velocity includes the moving direction and moving distance of the individual in each dimension of the solution space, the dimensions correspond one-to-one to the types of configuration parameters, and the position of the individual in the solution space corresponds to the parameter value of the individual; iteratively updating the positions of the individuals in the initial population according to the initial positions and the random initial velocities, where the fitness index of the individual is the storage system performance prediction value corresponding to the individual; and after reaching the maximum number of iterations, determining the target configuration parameter set according to the global optimal individual in the population, where the population is obtained by iteratively updating the positions of the individuals in the initial population, and the global optimal individual is the individual with the largest fitness index in the population.
[0100] In some embodiments of the present application, the fifth processing module 68 iteratively updates the positions of the individuals in the initial population according to the initial position and the random initial velocity, including: during the t-th iteration, determining a random velocity perturbation amount according to a sine function or a cosine function, where t is a positive integer not less than 2; determining the position of the global optimal individual in the (t - 1)-th iteration; determining the velocities of the individuals in the (t - 1)-th iteration and the historical optimal positions of the individuals after the (t - 1)-th iteration, where the historical optimal position is the position corresponding to the historical maximum fitness index of the individual; determining the velocities of the individuals during the t-th iteration according to the velocities of the individuals in the (t - 1)-th iteration, the random velocity perturbation amount, the position of the global optimal individual, and the historical optimal positions; and determining the positions of the individuals after the t-th iteration according to the velocities of the individuals.
[0101] In some embodiments of the present application, when the fifth processing module 68 determines the velocities of the individuals during the t-th iteration according to the random velocity perturbation amount, the position of the global optimal individual, and the historical optimal positions, it includes: determining a velocity weight at the t-th iteration, where the velocity weight is used to determine the retention amount of the velocities of the individuals in the (t - 1)-th iteration in the t-th iteration, and the value of the velocity weight decreases as the number of iterations increases; and determining the velocities of the individuals during the t-th iteration according to the velocity weight, the velocities of the individuals in the (t - 1)-th iteration, the random velocity perturbation amount, the position of the global optimal individual, and the historical optimal positions.
[0102] In some embodiments of the present application, the fifth processing module 68 determines a set of target configuration parameters according to the iterative initial population and optimizes the target storage system according to the set of target configuration parameters, including: determining the global optimal individual in the population after iteration, where the global optimal individual is the individual with the largest fitness index, and the fitness index is the predicted value of the storage system performance corresponding to the individual; determining the set of configuration parameters corresponding to the global optimal individual as the set of target configuration parameters; and updating the configuration parameter values of the target storage system to the parameter values in the set of target configuration parameters.
[0103] In some embodiments of the present application, when the fourth processing module 66 determines the predicted values of the storage system performance corresponding to the first type of individuals and the second type of individuals respectively, it includes: determining the application scenario information of the target storage system; determining the target performance type according to the application scenario information; determining a performance prediction model corresponding to the target performance type according to the target performance type, and determining the predicted values of the storage system performance corresponding to the first type of individuals and the second type of individuals respectively through the performance prediction model.
[0104] It should be noted that each module in the above storage system optimization device can be a program module (for example, a set of program instructions that implements a specific function), or a hardware module. For the latter, it can be presented in the following forms, but not limited to: the manifestation of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.
[0105] An embodiment of the present application provides a non-volatile storage medium. A program is stored in the non-volatile storage medium. When the program runs, it controls the device where the non-volatile storage medium is located to execute the following storage system optimization method: determining the configuration parameter types of the target storage system and the value ranges of the parameter values corresponding to each configuration parameter type; determining a data set according to the configuration parameter types and the value ranges of the parameter values, where the data set includes multiple configuration parameter sets, and each configuration parameter set includes the parameter values corresponding to each configuration parameter type; determining each configuration parameter set as a first type of individual, and determining a second type of individual according to the first type of individual and the value range of the parameter values, where the second type of individual is the reverse individual of the first type of individual; determining the storage system performance prediction values corresponding to the first type of individual and the second type of individual respectively, and selecting initial individuals from the first type of individual and the second type of individual according to the storage system performance prediction values to form an initial population; iteratively determining a target configuration parameter set according to the initial population, and optimizing the target storage system according to the target configuration parameter set.
[0106] An embodiment of the present application provides an electronic device, including: a memory and a processor. The processor is used to run the program stored in the memory. When the program runs, it executes the following storage system optimization method: determining the configuration parameter types of the target storage system and the value ranges of the parameter values corresponding to each configuration parameter type; determining a data set according to the configuration parameter types and the value ranges of the parameter values, where the data set includes multiple configuration parameter sets, and each configuration parameter set includes the parameter values corresponding to each configuration parameter type; determining each configuration parameter set as a first type of individual, and determining a second type of individual according to the first type of individual and the value range of the parameter values, where the second type of individual is the reverse individual of the first type of individual; determining the storage system performance prediction values corresponding to the first type of individual and the second type of individual respectively, and selecting initial individuals from the first type of individual and the second type of individual according to the storage system performance prediction values to form an initial population; iteratively determining a target configuration parameter set according to the initial population, and optimizing the target storage system according to the target configuration parameter set.
[0107] An embodiment of the present application provides a computer program product, including a computer program, which implements the following storage system optimization method when executed by a processor: determining the types of configuration parameters of a target storage system and the value ranges of the parameter values corresponding to each type of configuration parameter; determining a data set based on the types of configuration parameters and the value ranges of the parameter values, where the data set includes multiple configuration parameter sets, and each configuration parameter set includes the parameter values corresponding to each type of configuration parameter; determining each configuration parameter set as a first type of individual, and determining a second type of individual based on the first type of individual and the value range of the parameter values, where the second type of individual is the reverse individual of the first type of individual; determining the storage system performance prediction values corresponding to the first type of individual and the second type of individual respectively, and selecting initial individuals from the first type of individual and the second type of individual based on the storage system performance prediction values to form an initial population; iteratively determining a target configuration parameter set based on the initial population, and optimizing the target storage system according to the target configuration parameter set.
[0108] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0109] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the units or modules can be in an electrical or other form.
[0110] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0111] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0112] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0113] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A method for optimizing a storage system, characterized in that, Including: Determine the types of configuration parameters of the target storage system and the value ranges of the parameter values corresponding to each of the types of configuration parameters; Determine a data set based on the types of configuration parameters and the value ranges of the parameter values, where the data set includes a plurality of configuration parameter sets, and each configuration parameter set includes the parameter values corresponding to each of the types of configuration parameters; Determine each of the configuration parameter sets as a first type of individual, and determine a second type of individual based on the first type of individual and the value ranges of the parameter values, where the second type of individual is the reverse individual of the first type of individual; Determine the storage system performance prediction values corresponding to the first type of individual and the second type of individual respectively, and select initial individuals from the first type of individual and the second type of individual based on the storage system performance prediction values to form an initial population; Iteratively determine a target configuration parameter set based on the initial population, and optimize the target storage system based on the target configuration parameter set.
2. The storage system optimization method according to claim 1, wherein Determining the second type of individual based on the first type of individual and the value ranges of the parameter values includes: Based on the value ranges of the parameter values, determine the median of the parameter values corresponding to each of the types of configuration parameters; Based on the first parameter values of each of the types of configuration parameters of the first type of individual and the median of the parameter values of each of the types of configuration parameters, determine the second parameter values of each of the types of configuration parameters of the second type of individual corresponding to the first type of individual, where the average value of the first parameter value and the second parameter value of the same type of configuration parameter is the median of the parameter values.
3. The storage system optimization method according to claim 1, characterized in that Iteratively determining a target configuration parameter set based on the initial population and optimizing the target storage system based on the target configuration parameter set includes: Determine the initial positions and random initial velocities of the individuals in the initial population, where the random initial velocity includes the moving direction and moving distance of the individual in each dimension of the solution space, the dimension corresponds one-to-one with the type of configuration parameter, and the position of the individual in the solution space corresponds to the parameter value of the individual; Iteratively update the positions of the individuals in the initial population based on the initial positions and the random initial velocities, where the fitness index of the individual is the storage system performance prediction value corresponding to the individual; After reaching the maximum number of iterations, determine the target configuration parameter set based on the global optimal individual in the population, where the population is obtained by iteratively updating the positions of the individuals in the initial population, and the global optimal individual is the individual with the largest fitness index in the population.
4. The storage system optimization method according to claim 3, wherein Iteratively updating the positions of the individuals in the initial population based on the initial positions and the random initial velocities includes: During the t-th round of iteration, determine a random velocity perturbation amount based on a sine function or a cosine function, where t is a positive integer not less than 2; Determine the position of the global optimal individual in the (t - 1)-th round of iteration; Determine the velocities of each individual in the (t - 1)-th iteration round, and the historical optimal positions of each individual after the (t - 1)-th iteration, where the historical optimal position is the position corresponding to the historical maximum fitness index of the individual; During the process of determining the velocities of each individual in the t-th iteration based on the velocities of each individual in the (t - 1)-th iteration round, the random velocity perturbation amount, the position of the global optimal individual, and the historical optimal position; Determine the positions of the individuals after the t-th iteration based on the velocities of each individual.
5. The storage system optimization method according to claim 4, wherein During the process of determining the velocities of each individual in the t-th iteration based on the random velocity perturbation amount, the position of the global optimal individual, and the historical optimal position, it includes: Determine the velocity weight at the t-th iteration, where the velocity weight is used to determine the retention amount of the velocities of each individual in the (t - 1)-th iteration round in the t-th iteration, and the value of the velocity weight decreases as the iteration round increases; During the process of determining the velocities of each individual in the t-th iteration based on the velocity weight, the velocities of each individual in the (t - 1)-th iteration round, the random velocity perturbation amount, the position of the global optimal individual, and the historical optimal position.
6. The storage system optimization method according to claim 1, wherein Iteratively determine the target configuration parameter set according to the initial population, and optimize the target storage system according to the target configuration parameter set, including: Determine the global optimal individual in the population after iteration, where the global optimal individual is the individual with the largest fitness index, and the fitness index is the predicted value of the storage system performance corresponding to the individual; Determine the configuration parameter set corresponding to the global optimal individual as the target configuration parameter set; Update the configuration parameter values of the target storage system to the parameter values in the target configuration parameter set.
7. The storage system optimization method according to claim 1, wherein Determine the predicted values of the storage system performance corresponding to the first type of individuals and the second type of individuals respectively, including: Determine the application scenario information of the target storage system; Determine the target performance type according to the application scenario information; Determine the performance prediction model corresponding to the target performance type according to the target performance type, and determine the predicted values of the storage system performance corresponding to the first type of individuals and the second type of individuals respectively through the performance prediction model.
8. An optimization device for a storage system, characterized in that Include: The first processing module is used to determine the configuration parameter types of the target storage system and the value range of the parameter values corresponding to each configuration parameter type; The second processing module is used to determine the data set according to the configuration parameter type and the value range of the parameter values, where the data set includes multiple configuration parameter sets, and each configuration parameter set includes the parameter values corresponding to each configuration parameter type; The third processing module is used to determine each configuration parameter set as the first type of individuals, and determine the second type of individuals according to the first type of individuals and the value range of the parameter values, where the second type of individuals is the reverse individual of the first type of individuals; A fourth processing module, configured to determine respective storage system performance prediction values corresponding to the first type of individuals and the second type of individuals, and select initial individuals from the first type of individuals and the second type of individuals according to the storage system performance prediction values to form an initial population; A fifth processing module, configured to iteratively determine a set of target configuration parameters according to the initial population, and optimize the target storage system according to the set of target configuration parameters.
9. A non-volatile storage medium, characterized in that, A program is stored in the non-volatile storage medium, wherein when the program runs, it controls the device where the non-volatile storage medium is located to execute the storage system optimization method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Comprising: A memory and a processor, the processor is configured to run a program stored in the memory, wherein when the program runs, it executes the storage system optimization method according to any one of claims 1 to 7.
11. A computer program product, characterized in that, Comprising a computer program, which implements the storage system optimization method according to any one of claims 1 to 7 when executed by a processor.