Parameter configuration method, device, computer equipment, storage medium and product
Through the dynamic parameter configuration method, the target system tuning model of multiple rounds of training is used to solve the problems of parameter configuration complexity and resource preemption of hyper-converged business systems, and the resource utilization rate and system economy are improved.
Patent Information
- Application Number
- CN202510020775.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-07
AI Technical Summary
The parameters of each module of the hyper-converged business system are complex and difficult to take into account, resulting in resource seizure and waste, affecting the system resource utilization rate.
A dynamic parameter configuration method is provided, by obtaining the target state information and target parameter information of the business system, using the target system tuning model obtained by multiple rounds of training, determining future parameter information, and performing parameter configuration. The model updates the training samples through incremental samples to improve the model accuracy.
It improves the resource utilization rate of hyper-converged business systems, reduces CPU and memory resource usage, and improves the economics of the system.
Smart Images

Figure CN119416070B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cloud computing technology, and in particular to a parameter configuration method, apparatus, computer equipment, storage medium and product. Background Art
[0002] For hyper-converged business systems, computing, storage, network, and management resources are deployed on the same physical machine as tenant systems, and the configuration parameters of each subsystem are very complex. In this scenario, how to optimize the configuration parameters of each module to ensure that the hyper-converged system occupies less central processing unit (CPU) and memory resources under certain pressure scenarios, so that tenants deployed in this scenario can obtain more resources and improve the economic efficiency of the hyper-converged system.
[0003] However, since it is difficult to balance the parameter configurations of various modules in a hyper-converged business system, and there is resource competition between modules, the default parameter configurations or the parameters configured based on expert experience are not entirely reasonable, which can easily lead to a waste of resources. Summary of the invention
[0004] Based on this, it is necessary to provide a parameter configuration method, device, computer equipment, storage medium and product to address the above technical problems, which can dynamically configure the parameters of the hyper-converged business system and improve the resource utilization of the hyper-converged business system.
[0005] In a first aspect, the present application provides a parameter configuration method, which is applied to a tuner of a business system, comprising:
[0006] Obtaining target state information and target parameter information of the business system in a target period;
[0007] Based on the target system tuning model of the business system, the future parameter information of the business system in the future period is determined according to the target state information and the target parameter information; wherein the target system tuning model is obtained through multiple rounds of training, and each round of training samples of the target system tuning model is obtained by updating the previous round of training samples with the current round of incremental samples, and the current round of incremental samples is determined according to the current round of state information and the current round of parameter information generated by the business system based on the previous round of predicted parameter information;
[0008] Parameter configuration of the business system is performed according to the future parameter information.
[0009] In one embodiment, the target system tuning model is trained in the following manner:
[0010] Obtaining current round status information and current round parameter information generated by the business system based on the previous round prediction parameter information;
[0011] Determining whether the target system tuning model obtained through the previous round of training meets the convergence condition according to the current round parameter information and the previous round prediction parameter information;
[0012] If not, determining the current round incremental samples according to the current round status information and the current round parameter information;
[0013] Use the incremental samples of this round to update the training samples of the previous round to obtain the training samples of this round;
[0014] The target system tuning model is trained using the current round of training samples.
[0015] In one embodiment, determining the current round incremental sample according to the current round state information and the current round parameter information includes:
[0016] Determine a state vector of this round according to the state information of this round;
[0017] Determine a current round parameter vector according to the current round parameter information;
[0018] Based on the state prediction model, according to the current round state vector and the current round parameter vector, determining the predicted state vector of the business system under the current round parameter vector;
[0019] The current round incremental samples are determined according to the current round state vector, the current round parameter vector, the predicted state vector and the initial state vector of the business system.
[0020] In one embodiment, determining the current round incremental sample according to the current round state vector, the current round parameter vector, the predicted state vector and the initial state vector of the business system includes:
[0021] Determine the reward value of this round according to the predicted state vector and the initial state vector of the business system;
[0022] Construct a current round incremental sample according to the current round state vector, the current round parameter vector, the predicted state vector and the current round reward value.
[0023] In one of the embodiments, the predicted state vector includes the predicted state value of the business system under the parameter vector of this round, and the initial state vector includes the initial state value of the business system in the initial state;
[0024] The determining the reward value of this round according to the predicted state vector and the initial state vector of the business system includes:
[0025] Taking the difference between the initial state value and the predicted state value as the intermediate state value;
[0026] The reward value of this round is determined according to the product between the intermediate state value and the preset reward coefficient.
[0027] In one embodiment, the obtaining of current round status information and current round parameter information generated by the business system based on the previous round prediction parameter information includes:
[0028] Sending the last round of prediction parameter information to the business system so that the business system updates the parameters of the business system according to the last round of prediction parameter information;
[0029] After the business system updates the parameters, a stress test is performed on the business system, and current round status information and current round parameter information of the business system after the stress test are obtained.
[0030] In a second aspect, the present application further provides a parameter configuration device, which is configured in a tuner of a business system, comprising:
[0031] An information acquisition module, used to acquire target state information and target parameter information of the business system in a target period;
[0032] An information prediction module, used to determine the future parameter information of the business system in the future period based on the target system tuning model of the business system, according to the target state information and the target parameter information; wherein the target system tuning model is obtained through multiple rounds of training, and each round of training samples of the target system tuning model is obtained by updating the previous round of training samples with the current round of incremental samples, and the current round of incremental samples is determined according to the current round of state information and the current round of parameter information generated by the business system based on the previous round of predicted parameter information;
[0033] A parameter configuration module is used to configure the parameters of the business system according to the future parameter information.
[0034] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0035] Obtaining target state information and target parameter information of the business system in a target period;
[0036] Based on the target system tuning model of the business system, the future parameter information of the business system in the future period is determined according to the target state information and the target parameter information; wherein the target system tuning model is obtained through multiple rounds of training, and each round of training samples of the target system tuning model is obtained by updating the previous round of training samples with the current round of incremental samples, and the current round of incremental samples is determined according to the current round of state information and the current round of parameter information generated by the business system based on the previous round of predicted parameter information;
[0037] Parameter configuration of the business system is performed according to the future parameter information.
[0038] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0039] Obtaining target state information and target parameter information of the business system in a target period;
[0040] Based on the target system tuning model of the business system, the future parameter information of the business system in the future period is determined according to the target state information and the target parameter information; wherein the target system tuning model is obtained through multiple rounds of training, and each round of training samples of the target system tuning model is obtained by updating the previous round of training samples with the current round of incremental samples, and the current round of incremental samples is determined according to the current round of state information and the current round of parameter information generated by the business system based on the previous round of predicted parameter information;
[0041] Parameter configuration of the business system is performed according to the future parameter information.
[0042] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:
[0043] Obtaining target state information and target parameter information of the business system in a target period;
[0044] Based on the target system tuning model of the business system, the future parameter information of the business system in the future period is determined according to the target state information and the target parameter information; wherein the target system tuning model is obtained through multiple rounds of training, and each round of training samples of the target system tuning model is obtained by updating the previous round of training samples with the current round of incremental samples, and the current round of incremental samples is determined according to the current round of state information and the current round of parameter information generated by the business system based on the previous round of predicted parameter information;
[0045] Parameter configuration of the business system is performed according to the future parameter information.
[0046] The above parameter configuration method, device, computer equipment, storage medium and product can use the current round incremental samples to update the previous round training samples to obtain each round of training samples during each round of training of the target system optimization model, and determine the current round incremental samples according to the current round state information and the current round parameter information generated by the business system based on the previous round prediction parameter information. The training method of the target system tuning model provided by the present application is based on the previous round training samples and the current round incremental samples in each round of training, so that even when there are fewer training samples, a high-precision target system optimization model can be trained; at the same time, since the current round incremental samples are determined based on the current round state information and the current round parameter information generated by the previous round prediction parameter information, it is equivalent to considering the impact of the previous round prediction parameter information on the business system, which further improves the accuracy of the target system optimization model. Furthermore, based on the target system optimization model obtained by such training, according to the target state information and target parameter information of the business system in the target time period, the future parameter information of the business system in the future time period can be dynamically and accurately determined, ensuring that the determined future parameter information can improve the resource utilization of the business system. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0048] Figure 1 An application environment diagram of a parameter configuration method in an embodiment;
[0049] Figure 2 A schematic diagram of a flow chart of a parameter configuration method in an embodiment;
[0050] Figure 3 A schematic diagram of a process for training a target system tuning model in one embodiment;
[0051] Figure 4 A schematic diagram of a process for determining incremental samples in this round in one embodiment;
[0052] Figure 5 A schematic diagram for explaining the specific reference of each parameter in an embodiment;
[0053] Figure 6 A schematic diagram of a process for determining incremental samples in this round in another embodiment;
[0054] Figure 7A schematic diagram of a process for determining a reward value for this round in one embodiment;
[0055] Fig. 8A is an architectural diagram of a hyper-converged system in one embodiment;
[0056] Figure 8B A system architecture diagram for training a target system tuning model in one embodiment;
[0057] Fig. 9 A schematic diagram of a process for training a target system tuning model in another embodiment;
[0058] Fig.10 It is a structural block diagram of a parameter configuration device in one embodiment;
[0059] Fig.11 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0061] The parameter configuration method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the business system 101 is a system that needs to be optimized for parameters. In the embodiment of the present application, the business system 101 can be a hyper-converged system; the tuner 102 is used to tune the parameters of the business system. In the embodiment of the present application, the tuner can be deployed on the server of the business system 101. Optionally, the tuner 102 obtains the target state information and target parameter information of the business system 101 in the target time period; the tuner 102 determines the future parameter information of the business system in the future time period based on the target system tuning model of the business system 101 according to the target state information and target parameter information; wherein, the target system tuning model is obtained through multiple rounds of training, and each round of training samples of the target system tuning model is obtained by updating the previous round of training samples with the current round of incremental samples, and the current round of incremental samples is determined according to the current round of state information and the current round of parameter information generated by the business system based on the previous round of predicted parameter information; the tuner 102 performs parameter configuration on the business system 101 according to the future parameter information.
[0062] In one embodiment, Figure 2 As shown, a parameter configuration method is provided, which is applied to Figure 1 The tuner 102 in FIG. 1 is taken as an example to illustrate, and specifically includes the following steps:
[0063] S201, obtaining target state information and target parameter information of a business system in a target period.
[0064] Among them, the business system can be any business system based on the hyper-converged system; the target period is a preset period, such as the current period; the state information is information that characterizes the operating state of the business system, including but not limited to memory state information, CPU state information, and system response delay; the target state information is the state information of the business system in the preset period. The parameter information is information that characterizes the operating parameters of the business system, for example, the parameter information includes but is not limited to host memory reservation information, database cache pool size information, etc.; the target parameter information is the parameter information of the business system in the preset period.
[0065] Optionally, the target state information and target parameter information of the business system in the target time period may be directly obtained from the business system.
[0066] Optionally, in order to ensure the accuracy and convenience of the acquired target state information and target parameter information, in the embodiment of the present application, a controller may be deployed on the same physical host of the business system, and the controller may collect the state information and parameter information of the business system and control the business system to perform parameter configuration. In this case, the target state information and target parameter information of the business system in the target period may be obtained through the acquisition module in the controller.
[0067] S202 , based on the target system tuning model of the business system, determine the future parameter information of the business system in the future period according to the target state information and the target parameter information.
[0068] Among them, the target system tuning model is a model used to tune the parameters of the business system; in the embodiment of the present application, the target system tuning model is obtained through multiple rounds of training, and each round of training samples of the target system tuning model is obtained by updating the previous round of training samples using the current round of incremental samples, and the current round of incremental samples is determined according to the current round of state information and current round of parameter information generated by the business system based on the previous round of predicted parameter information. Future parameter information is the parameter information of the business system in the future period.
[0069] In an embodiment of the present application, the target system tuning model may be a model constructed based on a deep deterministic policy gradient algorithm (DDPG).
[0070] It should be noted that in order to ensure the prediction accuracy of the target system tuning model, it is necessary to perform multiple rounds of training on the target system tuning model. After each round of training, it is necessary to verify whether the target system tuning model meets the convergence conditions. Furthermore, if the conditions are met, the multiple rounds of training of the target system tuning model are stopped; if not, it is necessary to continue to perform multiple rounds of training on the target system tuning model.
[0071] It should be further explained that, in the process of training the target system tuning model, the target system tuning model needs to be trained according to each round of training samples; and, each round of training samples is obtained by updating the previous round of training samples with this round of incremental samples, wherein this round of incremental samples is determined according to this round of status information and this round of parameter information generated by the business system based on the previous round of prediction parameter information.
[0072] Optionally, the first round of training is performed on the target system tuning model. At this time, since the prediction parameter information of the previous round is empty, the initial parameter information of the business system is used as the parameter information of this round, and the initial state information of the business system is used as the state information of this round; further, the incremental samples of this round are determined according to the parameter information of this round and the state information of this round. At the same time, if the training samples of the previous round are empty, the incremental samples of this round are directly used as the training samples of this round. The training samples of this round are further used to train the tuning model of the target system. And so on, until the optimization model of the target system meets the convergence conditions.
[0073] Optionally, the target state information and target parameter information can be input into the target system tuning model of the business system, so that the target system tuning model calculates the target state information and target parameter information based on the trained model parameters, and uses the output of the target system tuning model as the future parameter information of the business system in the future period.
[0074] S203: Parameter configuration of the business system according to the future parameter information.
[0075] Optionally, future parameter information may be sent to the business system so that the business system configures the parameters of the business system according to the future parameter information.
[0076] Optionally, if a controller is deployed in the business system, future parameter information can be sent to the controller. After receiving the future parameter information, the controller generates a new configuration file according to the future parameter information, further, overwrites the current configuration with the generated new configuration file, and restarts the corresponding process to make it effective.
[0077] In the above parameter configuration method, in each round of training of the target system optimization model, the current round incremental samples can be used to update the previous round training samples to obtain each round of training samples, and the current round incremental samples are determined according to the current round state information and the current round parameter information generated by the business system based on the previous round prediction parameter information. The training method of the target system tuning model provided by the present application is based on the previous round training samples and the current round incremental samples in each round of training, so that even when there are fewer training samples, a high-precision target system optimization model can be trained; at the same time, since the current round incremental samples are determined based on the current round state information and the current round parameter information generated by the previous round prediction parameter information, it is equivalent to considering the impact of the previous round prediction parameter information on the business system, further improving the accuracy of the target system optimization model. Furthermore, based on the target system optimization model obtained by such training, according to the target state information and target parameter information of the business system in the target time period, the future parameter information of the business system in the future time period can be dynamically and accurately determined, ensuring that the determined future parameter information can improve the resource utilization of the business system.
[0078] Optionally, in order to ensure the accuracy of the target system tuning model, the target system tuning model needs to be trained for multiple rounds. In one embodiment, a round of training of the target recognition model is used as an example. Figure 3 As shown, a training method for a target system tuning model is provided, which specifically includes the following steps:
[0079] S301, obtaining current round status information and current round parameter information generated by the business system based on the previous round prediction parameter information.
[0080] Optionally, if the current training round is the first training round, the prediction parameter information of the previous round is empty. At this time, the initial state information of the business system can be directly obtained as the current round state information, and the initial parameter information of the business system can be directly obtained as the current round parameter information.
[0081] Optionally, if the current training round is not the first training round, the tuner sends the last round of prediction parameter information to the business system so that the business system updates the parameters of the business system according to the last round of prediction parameter information; after the business system updates the parameters, the business system is stress tested, and the current round of state information and parameter information of the business system after the stress test are obtained. Optionally, after the last training round, the tuner sends the last round of prediction parameters output by the target system tuning model in the last training round to the business system. The business system will perform parameter overwriting configuration on the parameters of the business system according to the received last round of prediction parameters. Further, the stress script is started to stress the business system, and different stress test conditions can be added according to the business scenario, such as batch creation of virtual machines, batch deletion of virtual machines, batch creation of disks, and batch deletion of disks. After starting the stress for a period of time, the current round of state information and parameter information of the business system after the stress test are obtained.
[0082] S302: Determine whether the target system tuning model obtained by the previous round of training meets the convergence condition according to the parameter information of this round and the predicted parameter information of the previous round.
[0083] It should be noted that since the parameter information of this round is the parameter information generated by the business system based on the configuration of the previous round of predicted parameter information, the parameter information of this round can be regarded as the actual parameter information of the previous round of predicted parameter information.
[0084] Optionally, when it is necessary to determine whether the target system tuning model obtained by the previous round of training meets the convergence conditions, the parameter information of this round and the predicted parameter information of the previous round can be substituted into a pre-set loss function to calculate the value of the loss function of the target system tuning model obtained by the previous round of training, and based on the relationship between the value of the loss function and the preset loss threshold, it is judged whether the target system tuning model obtained by the previous round of training meets the convergence conditions.
[0085] Exemplarily, the loss function can be constructed based on the mean square error between the current round parameter information and the previous round prediction parameter information.
[0086] S303: If not, determine the incremental samples for this round according to the status information and parameter information for this round.
[0087] Optionally, if the target system tuning model obtained by the previous round of training does not meet the convergence condition, it is necessary to continue to perform multiple rounds of training operations on the target system tuning model. At this time, it is necessary to process and calculate the current round state information and current round parameter information according to the construction rules of the incremental samples to obtain the current round incremental samples.
[0088] S304, using the current round incremental samples to update the previous round training samples to obtain the current round training samples.
[0089] Optionally, the current round incremental samples and the previous round training samples can be used together as the current round training samples. Exemplarily, an experience return pool can be introduced to store the current round training samples, so that the current round incremental samples obtained during each round of training can be stored in the experience return pool.
[0090] S305, using the training samples of this round to train the target system tuning model.
[0091] Optionally, the target system tuning model may be trained using the training samples of this round, or some training samples may be randomly selected from the training samples of this round to train the target system tuning model.
[0092] For example, when the experience return pool is introduced, data in the experience return pool can be randomly extracted in batches to train the target system tuning model.
[0093] In this embodiment, by introducing the current round state information and the current round parameter information during the training process of the target system tuning model, and when the target system tuning model obtained by the previous round of training does not meet the convergence conditions, the target system tuning model continues to be trained according to the current round training samples, thereby ensuring that a high-precision target system tuning model is obtained through training.
[0094] Optionally, in order to ensure the accuracy of the determined incremental samples of this round, a state processing module may be introduced so that the state processing module determines the incremental samples of this round according to the state information of this round and the parameter information of this round. Figure 4 As shown, a method for determining the incremental samples of this round is provided to refine the above S203, which specifically includes the following steps:
[0095] S401, determining the current round state vector according to the current round state information.
[0096] The state vector of this round includes the state value of the business system.
[0097] Optionally, the state values of each state in the obtained state information of this round can be arranged in a preset order to obtain the state vector of this round. Exemplarily, the state values of different states can be represented by s0, s1, etc. In this case, the state vector of this round can be represented as (s0, s1, ...).
[0098] S402, determining a parameter vector for this round according to the parameter information for this round.
[0099] The parameter vector of this round includes the parameter values of the business system.
[0100] Optionally, the parameter values of each parameter in the obtained current round parameter information may be arranged in a preset order to obtain the current round parameter vector. Exemplarily, the parameter values of different parameters may be represented by p0, p1, etc. In this case, the current round parameter vector may be represented as (p0, p1, ...).
[0101] For example, taking the hyper-converged system composed of OpenStack (cloud computing management platform project) cloud platform and distributed storage as an example, based on historical experience, the parameter vectors (p0, p1, ... p18) that have a greater impact on system resources are selected. The specific reference explanations of each parameter are as follows: Figure 5 shown.
[0102] S403, based on the state prediction model, according to the current round state vector and the current round parameter vector, determine the predicted state vector of the business system under the current round parameter vector.
[0103] The state prediction model is a model for predicting the state of the business system; for example, the state prediction model can be a model built based on the DDPG reinforcement learning algorithm. The predicted state vector is the state vector of the business system predicted by the state prediction model under the current round of parameter vectors.
[0104] Optionally, the current round state vector and the current round parameter vector can be input into the state prediction model, so that the state prediction model analyzes and calculates the current round state vector and the current round parameter vector according to the model parameters, and outputs the predicted state vector of the business system under the current round parameter vector.
[0105] S404, determining the current round incremental samples according to the current round state vector, the current round parameter vector, the predicted state vector and the initial state vector of the business system.
[0106] Optionally, a sample vector may be constructed based on the current round state vector, the current round parameter vector, the predicted state vector, and the initial state vector of the business system, and the sample vector may be used as the current round incremental sample.
[0107] In this embodiment, in the process of determining the incremental samples of this round, by introducing the state prediction model, the predicted state vector of the business system under the parameter vector of this round can be accurately obtained. Furthermore, by introducing the state vector of this round and the parameter vector of this round, the accuracy of the determined incremental samples of this round is guaranteed.
[0108] Optionally, in order to further ensure the accuracy of the determined incremental samples of this round, based on the above embodiment, in one embodiment, Figure 6 As shown, a method for determining the incremental samples of this round is provided to refine the above S404, which specifically includes the following steps:
[0109] S601, determining the reward value of this round according to the predicted state vector and the initial state vector of the business system.
[0110] Among them, the reward value of this round is used to measure the degree of parameter optimization of the target system tuning model after the previous round of training.
[0111] Optionally, the reward value of this round can be determined according to the difference between the predicted state vector and the initial state vector of the business system. For example, if the predicted state vector is smaller than the initial state vector, the reward value of this round is set to a positive value, and vice versa, if the predicted state vector is larger than the initial state vector, the reward value of this round is set to a negative value to guide the target system tuning model to converge quickly.
[0112] S602, constructing the current round incremental sample according to the current round state vector, the current round parameter vector, the predicted state vector and the current round reward value.
[0113] Optionally, a vector may be constructed based on the current round state vector, the current round parameter vector, the predicted state vector, and the current round reward value, and the constructed vector may be used as the current round incremental sample.
[0114] In this embodiment, in the process of determining the incremental samples of this round, by introducing the reward value of this round, since the reward value of this round is determined based on the predicted state vector and the initial state vector of the business system, the reward value of this round can guide the optimization degree of the target system optimization model to a certain extent, thereby ensuring the availability and accuracy of the determined incremental samples of this round.
[0115] Optionally, in order to ensure the accuracy of the determined reward value of this round, in one embodiment, the predicted state vector includes the predicted state value of the business system under the parameter vector of this round, and the initial state vector includes the initial state value of the business system in the initial state; in this case, Figure 7 As shown, a method for determining the reward value of this round is provided to refine the above S601, which specifically includes the following steps:
[0116] S701, taking the difference between the initial state value and the predicted state value as the intermediate state value.
[0117] Optionally, the initial state value and the predicted state value may be subtracted, and the difference obtained may be used as the intermediate state value.
[0118] S702, determining the reward value of this round according to the product of the intermediate state value and the preset reward coefficient.
[0119] Among them, the preset reward coefficient is a pre-set reward value. In the embodiment of the present application, the preset reward coefficient can be determined according to the difference between the initial state value and the predicted state value; in the embodiment of the present application, the value range of the reward coefficient is (0,1].
[0120] Optionally, the product between the intermediate state value and the preset reward system can be used as the first product, and further, the ratio between the first product and the initial state value can be used as the reward value for this round.
[0121] It should be noted that if the business system has multiple predicted state values under the parameter vector of this round, then the reward value of this round is the sum of the reward values of this round corresponding to each predicted state value.
[0122] Exemplarily, the state of the business system may include but is not limited to the CPU state, memory state and system delay state. The state value of the CPU state is the number of occupied CPU cores, the state value of the memory state is the occupied memory size, and the state value of the system delay state is the response delay. The round reward value can be expressed by the following formula (1):
[0123] (1)
[0124] in, , and They are the reward coefficients for CPU status, memory status, and system delay status, and their value range is , , ; It is the initial state value of the CPU state; It is the predicted state value of the CPU state; is the initial state value of the memory state; is the predicted state value of the memory state; is the initial state value of the system delay state; It is the predicted state value of the system delay state.
[0125] In this embodiment, in the process of determining the reward value of this round, by introducing the initial state value and the predicted state value, it can be ensured that the intermediate state value reflects the accuracy of the predicted state value to a certain extent; further, by introducing the preset reward coefficient, the accuracy of the determined reward value of this round is guaranteed.
[0126] Optionally, in order to more intuitively describe the training process of the target system tuning model in the parameter configuration method provided in the embodiment of the present application, in one embodiment, Fig. 8AThe hyper-converged system shown is used as an example of a business system. Since the hyper-converged system deploys the host system, computing management, storage management, network management, management platform and tenant system on the same physical machine, the configuration parameters of each subsystem are very complex. In this scenario, how to optimize the configuration parameters of each module to ensure that under a certain pressure scenario, the hyper-converged system occupies less system CPU and system memory resources, so that the tenants deployed in this scenario can obtain more resources and improve the economic efficiency of the hyper-converged system. Therefore, the parameter configuration method provided in the embodiment of the present application can be applied to industrial and financial systems, such as Figure 8B As shown in the figure, a system architecture diagram for training the target system tuning model is provided. The controller is directly deployed on the hyper-converged system side and interacts directly with the hyper-converged system; the tuner is implemented based on the DDPG algorithm in reinforcement learning, which is suitable for complex high-dimensional continuous parameter tuning of the hyper-converged system.
[0127] The physical environment in which the hyper-converged system is deployed includes the hyper-converged system and physical resources (CPU, memory); the controller is divided into a system controller and a state collector. The system controller implements the application programming interface (API) for receiving parameter updates, the system configuration update program, and the system pressure program; the state collector implements the API interface for querying current parameters and status. The tuner obtains the state data and parameter data of the state collector, and inputs them into the state processing module as state information and parameter information respectively. The state processing module obtains the first The state vector of the training process , No. The parameter vector of the training process , the predicted state vector , and the first The reward value of the training process , generating Incremental samples of the training process ( , , , ) is stored in the experience replay pool. Random batches of data are taken from the experience replay pool, and the DDPG reinforcement learning algorithm is used to train the system tuning model, output new parameters, and update the parameters.
[0128] In one embodiment of the present application, a flowchart of the steps of a target system tuning model training method is provided. This embodiment is described by taking one round of training of the target system tuning model as an example. Fig. 9 The specific implementation process is as follows:
[0129] S901, sending the last round of prediction parameter information to the business system, so that the business system updates the parameters of the business system according to the last round of prediction parameter information.
[0130] S902, after the business system updates the parameters, perform a stress test on the business system, and obtain current round status information and current round parameter information of the business system after the stress test.
[0131] S903, based on the parameter information of this round and the predicted parameter information of the previous round, determine whether the target system tuning model obtained by the previous round of training meets the convergence condition; if so, execute S912; if not, execute S904.
[0132] S904, determining the current round state vector according to the current round state information.
[0133] S905, determining a parameter vector for this round according to the parameter information for this round.
[0134] S906, based on the state prediction model, according to the current round state vector and the current round parameter vector, determine the predicted state vector of the business system under the current round parameter vector.
[0135] S907: The difference between the initial state value and the predicted state value is used as the intermediate state value.
[0136] S908, determining the reward value of this round according to the product of the intermediate state value and the preset reward coefficient.
[0137] S909, construct the incremental sample of this round according to the state vector of this round, the parameter vector of this round, the predicted state vector and the reward value of this round.
[0138] S910, using the current round incremental samples to update the previous round training samples to obtain the current round training samples.
[0139] S911, use this round of training samples to train the target system tuning model.
[0140] S912, stop training the target system tuning model.
[0141] The specific process of S901-S912 above can refer to the description of the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated here.
[0142] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0143] Based on the same inventive concept, the embodiment of the present application also provides a parameter configuration device for implementing the parameter configuration method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more parameter configuration device embodiments provided below can refer to the limitations on the parameter configuration method above, and will not be repeated here.
[0144] In an exemplary embodiment, Fig.10 As shown, a parameter configuration device 1000 is provided, comprising: an information acquisition module 1010, an information prediction module 1020 and a parameter configuration module 1030, wherein:
[0145] The information acquisition module 1010 is used to acquire target state information and target parameter information of the business system in a target period.
[0146] The information prediction module 1020 is used to determine the future parameter information of the business system in the future period based on the target system tuning model of the business system according to the target state information and the target parameter information; wherein the target system tuning model is obtained through multiple rounds of training, and each round of training samples of the target system tuning model is obtained by updating the previous round of training samples with the current round of incremental samples, and the current round of incremental samples is determined according to the current round of state information and the current round of parameter information generated by the business system based on the previous round of predicted parameter information.
[0147] The parameter configuration module 1030 is used to configure the parameters of the business system according to the future parameter information.
[0148] The above parameter configuration device can use the current round incremental samples to update the previous round training samples to obtain each round of training samples during each round of training of the target system optimization model, and determine the current round incremental samples according to the current round state information and the current round parameter information generated by the business system based on the previous round prediction parameter information. The training method of the target system tuning model provided by the present application is based on the previous round training samples and the current round incremental samples in each round of training, so that even when there are fewer training samples, a high-precision target system optimization model can be trained; at the same time, since the current round incremental samples are determined based on the current round state information and the current round parameter information generated by the previous round prediction parameter information, it is equivalent to considering the impact of the previous round prediction parameter information on the business system, further improving the accuracy of the target system optimization model. Furthermore, based on the target system optimization model obtained by such training, according to the target state information and target parameter information of the business system in the target time period, the future parameter information of the business system in the future time period can be dynamically and accurately determined, ensuring that the determined future parameter information can improve the resource utilization of the business system.
[0149] In one embodiment, the parameter configuration device 1000 further includes:
[0150] The first acquisition module is used to acquire current round status information and current round parameter information generated by the business system based on the previous round prediction parameter information.
[0151] The condition judgment module is used to determine whether the target system tuning model obtained by the previous round of training meets the convergence condition based on the parameter information of this round and the predicted parameter information of the previous round.
[0152] The first determination module is used to determine the incremental samples of this round according to the current round state information and the current round parameter information if no.
[0153] The second determination module is used to update the previous round of training samples with the current round of incremental samples to obtain the current round of training samples.
[0154] The model training module is used to train the target system tuning model using the current round of training samples.
[0155] In one embodiment, the first acquisition module is specifically used for:
[0156] The previous round of prediction parameter information is sent to the business system so that the business system can update the parameters of the business system according to the previous round of prediction parameter information; after the business system updates the parameters, the business system is stress tested, and the current round status information and current round parameter information of the business system after the stress test are obtained.
[0157] In one embodiment, the first determining module includes:
[0158] The first determining unit is used to determine the current round state vector according to the current round state information.
[0159] The second determining unit is used to determine the current round parameter vector according to the current round parameter information.
[0160] The vector determination unit is used to determine the predicted state vector of the business system under the current round parameter vector based on the state prediction model according to the current round state vector and the current round parameter vector.
[0161] The sample determination unit is used to determine the incremental samples of this round according to the state vector of this round, the parameter vector of this round, the predicted state vector and the initial state vector of the business system.
[0162] In one embodiment, the sample determination unit includes:
[0163] The first determination subunit is used to determine the reward value of this round according to the predicted state vector and the initial state vector of the business system.
[0164] The second determination subunit is used to construct the current round incremental sample according to the current round state vector, the current round parameter vector, the predicted state vector and the current round reward value.
[0165] In one embodiment, the predicted state vector includes the predicted state value of the business system under the current round parameter vector, and the initial state vector includes the initial state value of the business system in the initial state; the first determining subunit is specifically used for:
[0166] The difference between the initial state value and the predicted state value is taken as the intermediate state value; the reward value of this round is determined according to the product between the intermediate state value and the preset reward coefficient.
[0167] Each module in the above parameter configuration device can be implemented in whole or in part by software, hardware and a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module above.
[0168] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Fig.11As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a parameter configuration method is implemented.
[0169] Those skilled in the art will understand that Fig.11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0170] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0171] Obtain target status information and target parameter information of the business system in the target period;
[0172] Based on the target system tuning model of the business system, the future parameter information of the business system in the future period is determined according to the target state information and the target parameter information; wherein the target system tuning model is obtained through multiple rounds of training, and each round of training samples of the target system tuning model is obtained by updating the previous round of training samples with the current round of incremental samples, and the current round of incremental samples is determined according to the current round of state information and the current round of parameter information generated by the business system based on the previous round of predicted parameter information;
[0173] Parameter configuration of the business system based on future parameter information.
[0174] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0175] Obtain the current round status information and current round parameter information generated by the business system based on the previous round prediction parameter information; determine whether the target system tuning model obtained by the previous round of training meets the convergence conditions based on the current round parameter information and the previous round prediction parameter information; if not, determine the current round incremental samples based on the current round status information and the current round parameter information; use the current round incremental samples to update the previous round training samples to obtain the current round training samples; use the current round training samples to train the target system tuning model.
[0176] In one embodiment, when the processor executes the computer program to determine the current round incremental sample according to the current round state information and the current round parameter information, the processor further implements the following steps:
[0177] According to the state information of this round, determine the state vector of this round; according to the parameter information of this round, determine the parameter vector of this round; based on the state prediction model, according to the state vector of this round and the parameter vector of this round, determine the predicted state vector of the business system under the parameter vector of this round; according to the state vector of this round, the parameter vector of this round, the predicted state vector and the initial state vector of the business system, determine the incremental sample of this round.
[0178] In one embodiment, when the processor executes the computer program to determine the current round incremental sample according to the current round state vector, the current round parameter vector, the predicted state vector and the initial state vector of the business system, the following steps are also implemented:
[0179] The reward value of this round is determined based on the predicted state vector and the initial state vector of the business system; the incremental sample of this round is constructed based on the state vector of this round, the parameter vector of this round, the predicted state vector and the reward value of this round.
[0180] In one embodiment, the predicted state vector includes the predicted state value of the business system under the parameter vector of this round, and the initial state vector includes the initial state value of the business system in the initial state; when the processor executes the computer program to determine the reward value of this round according to the predicted state vector and the initial state vector of the business system, the following steps are also implemented:
[0181] The difference between the initial state value and the predicted state value is taken as the intermediate state value; the reward value of this round is determined according to the product between the intermediate state value and the preset reward coefficient.
[0182] In one embodiment, when the processor executes the computer program to obtain the current round state information and current round parameter information generated by the business system based on the previous round prediction parameter information, the following steps are also implemented:
[0183] The previous round of prediction parameter information is sent to the business system so that the business system can update the parameters of the business system according to the previous round of prediction parameter information; after the business system updates the parameters, the business system is stress tested, and the current round status information and current round parameter information of the business system after the stress test are obtained.
[0184] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0185] Obtain target status information and target parameter information of the business system in the target period;
[0186] Based on the target system tuning model of the business system, the future parameter information of the business system in the future period is determined according to the target state information and the target parameter information; wherein the target system tuning model is obtained through multiple rounds of training, and each round of training samples of the target system tuning model is obtained by updating the previous round of training samples with the current round of incremental samples, and the current round of incremental samples is determined according to the current round of state information and the current round of parameter information generated by the business system based on the previous round of predicted parameter information;
[0187] Parameter configuration of the business system based on future parameter information.
[0188] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0189] Obtain the current round status information and current round parameter information generated by the business system based on the previous round prediction parameter information; determine whether the target system tuning model obtained by the previous round of training meets the convergence conditions based on the current round parameter information and the previous round prediction parameter information; if not, determine the current round incremental samples based on the current round status information and the current round parameter information; use the current round incremental samples to update the previous round training samples to obtain the current round training samples; use the current round training samples to train the target system tuning model.
[0190] In one embodiment, when the processor executes the computer program to determine the current round incremental sample according to the current round state information and the current round parameter information, the processor further implements the following steps:
[0191] According to the state information of this round, determine the state vector of this round; according to the parameter information of this round, determine the parameter vector of this round; based on the state prediction model, according to the state vector of this round and the parameter vector of this round, determine the predicted state vector of the business system under the parameter vector of this round; according to the state vector of this round, the parameter vector of this round, the predicted state vector and the initial state vector of the business system, determine the incremental sample of this round.
[0192] In one embodiment, when the processor executes the computer program to determine the current round incremental sample according to the current round state vector, the current round parameter vector, the predicted state vector and the initial state vector of the business system, the following steps are also implemented:
[0193] The reward value of this round is determined based on the predicted state vector and the initial state vector of the business system; the incremental sample of this round is constructed based on the state vector of this round, the parameter vector of this round, the predicted state vector and the reward value of this round.
[0194] In one embodiment, the predicted state vector includes the predicted state value of the business system under the parameter vector of this round, and the initial state vector includes the initial state value of the business system in the initial state; when the processor executes the computer program to determine the reward value of this round according to the predicted state vector and the initial state vector of the business system, the following steps are also implemented:
[0195] The difference between the initial state value and the predicted state value is taken as the intermediate state value; the reward value of this round is determined according to the product between the intermediate state value and the preset reward coefficient.
[0196] In one embodiment, when the processor executes the computer program to obtain the current round state information and current round parameter information generated by the business system based on the previous round prediction parameter information, the following steps are also implemented:
[0197] The previous round of prediction parameter information is sent to the business system so that the business system can update the parameters of the business system according to the previous round of prediction parameter information; after the business system updates the parameters, the business system is stress tested, and the current round status information and current round parameter information of the business system after the stress test are obtained.
[0198] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the following steps:
[0199] Obtain target status information and target parameter information of the business system in the target period;
[0200] Based on the target system tuning model of the business system, the future parameter information of the business system in the future period is determined according to the target state information and the target parameter information; wherein the target system tuning model is obtained through multiple rounds of training, and each round of training samples of the target system tuning model is obtained by updating the previous round of training samples with the current round of incremental samples, and the current round of incremental samples is determined according to the current round of state information and the current round of parameter information generated by the business system based on the previous round of predicted parameter information;
[0201] Parameter configuration of the business system based on future parameter information.
[0202] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0203] Obtain the current round status information and current round parameter information generated by the business system based on the previous round prediction parameter information; determine whether the target system tuning model obtained by the previous round of training meets the convergence conditions based on the current round parameter information and the previous round prediction parameter information; if not, determine the current round incremental samples based on the current round status information and the current round parameter information; use the current round incremental samples to update the previous round training samples to obtain the current round training samples; use the current round training samples to train the target system tuning model.
[0204] In one embodiment, when the processor executes the computer program to determine the current round incremental sample according to the current round state information and the current round parameter information, the processor further implements the following steps:
[0205] According to the state information of this round, determine the state vector of this round; according to the parameter information of this round, determine the parameter vector of this round; based on the state prediction model, according to the state vector of this round and the parameter vector of this round, determine the predicted state vector of the business system under the parameter vector of this round; according to the state vector of this round, the parameter vector of this round, the predicted state vector and the initial state vector of the business system, determine the incremental sample of this round.
[0206] In one embodiment, when the processor executes the computer program to determine the current round incremental sample according to the current round state vector, the current round parameter vector, the predicted state vector and the initial state vector of the business system, the following steps are also implemented:
[0207] The reward value of this round is determined based on the predicted state vector and the initial state vector of the business system; the incremental sample of this round is constructed based on the state vector of this round, the parameter vector of this round, the predicted state vector and the reward value of this round.
[0208] In one embodiment, the predicted state vector includes the predicted state value of the business system under the parameter vector of this round, and the initial state vector includes the initial state value of the business system in the initial state; when the processor executes the computer program to determine the reward value of this round according to the predicted state vector and the initial state vector of the business system, the following steps are also implemented:
[0209] The difference between the initial state value and the predicted state value is taken as the intermediate state value; the reward value of this round is determined according to the product between the intermediate state value and the preset reward coefficient.
[0210] In one embodiment, when the processor executes the computer program to obtain the current round state information and current round parameter information generated by the business system based on the previous round prediction parameter information, the following steps are also implemented:
[0211] The previous round of prediction parameter information is sent to the business system so that the business system can update the parameters of the business system according to the previous round of prediction parameter information; after the business system updates the parameters, the business system is stress tested, and the current round status information and current round parameter information of the business system after the stress test are obtained.
[0212] It should be noted that the data involved in this application (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0213] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0214] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0215] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A parameter configuration method, characterized in that: A tuner applied to a business system, the method comprising: Acquire target state information and target parameter information of the business system in the target period; wherein the target state information includes memory state information, central processing unit CPU state information and system response delay of the business system in the target period; Based on the target system tuning model of the business system, determining the future parameter information of the business system in the future period according to the target state information and the target parameter information; wherein the target system tuning model is obtained through multiple rounds of training; Parameter configuration of the business system according to the future parameter information; The target system tuning model is trained in the following way: Sending the last round of prediction parameter information to the business system so that the business system updates the parameters of the business system according to the last round of prediction parameter information; After the business system updates the parameters, the business system is stress tested, and current round status information and current round parameter information of the business system after the stress test are obtained; Determining whether the target system tuning model obtained through the previous round of training meets the convergence condition according to the current round parameter information and the previous round prediction parameter information; If not, determining the current round state vector according to the current round state information; Determine a current round parameter vector according to the current round parameter information; Based on the state prediction model, according to the current round state vector and the current round parameter vector, determining the predicted state vector of the business system under the current round parameter vector; Determine the current round incremental samples according to the current round status information and the current round parameter information; Use the incremental samples of this round to update the training samples of the previous round to obtain the training samples of this round; The target system tuning model is trained using the current round of training samples.
2. The method according to claim 1, characterized in that: The step of determining the current round incremental samples according to the current round state vector, the current round parameter vector, the predicted state vector, and the initial state vector of the business system includes: Determine the reward value of this round according to the predicted state vector and the initial state vector of the business system; Construct a current round incremental sample according to the current round state vector, the current round parameter vector, the predicted state vector and the current round reward value.
3. The method according to claim 2, characterized in that The predicted state vector includes the predicted state value of the business system under the current round parameter vector, and the initial state vector includes the initial state value of the business system in the initial state; The determining the reward value of this round according to the predicted state vector and the initial state vector of the business system includes: Taking the difference between the initial state value and the predicted state value as the intermediate state value; The reward value of this round is determined according to the product between the intermediate state value and the preset reward coefficient.
4. The method according to claim 1, characterized in that: The parameter configuration of the business system according to the future parameter information includes: The future parameter information is sent to the controller of the business system, so that the controller generates a new configuration file according to the future parameter information, and performs parameter configuration on the business system according to the new configuration file.
5. The method according to claim 1, characterized in that: The stress testing of the business system and obtaining current round status information and current round parameter information of the business system after the stress testing includes: Determine stress test conditions based on the business scenarios of the business system; Using a stress script, according to the stress test conditions, the business system is stress tested; Obtain current round status information and current round parameter information of the business system after the stress test.
6. The method according to claim 1, characterized in that The determining, based on the current round parameter information and the previous round prediction parameter information, whether the target system tuning model obtained through the previous round training meets the convergence condition includes: Obtain the preset loss function of the target system tuning model obtained through the previous round of training; Using the current round parameter information and the previous round prediction parameter information, the preset loss function is updated to obtain the loss value of the target system tuning model obtained through the previous round of training; According to the relationship between the loss value and the preset loss threshold, it is determined whether the target system tuning model obtained by the previous round of training meets the convergence condition.
7. A parameter configuration device, characterized in that: A tuner configured in a business system, the device comprising: An information acquisition module, used to acquire target state information and target parameter information of the business system in a target period; wherein the target state information includes memory state information, central processing unit CPU state information and system response delay of the business system in the target period; An information prediction module, used to determine the future parameter information of the business system in the future period based on the target system tuning model of the business system, according to the target state information and the target parameter information; wherein the target system tuning model is obtained through multiple rounds of training, and each round of training samples of the target system tuning model is obtained by updating the previous round of training samples with the current round of incremental samples, and the current round of incremental samples is determined according to the current round of state information and the current round of parameter information generated by the business system based on the previous round of predicted parameter information; A parameter configuration module, used for configuring parameters of the business system according to the future parameter information; The target system tuning model is trained in the following way: Sending the last round of prediction parameter information to the business system so that the business system updates the parameters of the business system according to the last round of prediction parameter information; After the business system updates the parameters, the business system is stress tested, and current round status information and current round parameter information of the business system after the stress test are obtained; Determining whether the target system tuning model obtained through the previous round of training meets the convergence condition according to the current round parameter information and the previous round prediction parameter information; If not, determining the current round state vector according to the current round state information; Determine a current round parameter vector according to the current round parameter information; Based on the state prediction model, according to the current round state vector and the current round parameter vector, determining the predicted state vector of the business system under the current round parameter vector; Determine the current round incremental samples according to the current round status information and the current round parameter information; Use the incremental samples of this round to update the training samples of the previous round to obtain the training samples of this round; The target system tuning model is trained using the current round of training samples.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Model incremental learning method and device, equipment and storage medium
CN116432780A
Parameter configuration method and system, storage medium and terminal equipment
CN118093039A