Model training method and model training system based on reinforcement learning

Through reinforcement learning-based methods, the performance evaluation model and perceptual response model are built, and the model training and evaluation are automated, which solves the problem of large workload and difficulty in exhaustive combination of manual parameter setting, and realizes efficient optimization of the model.

CN120338032APending Publication Date: 2025-07-18HANGZHOU WANLAN TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510426647.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, manual parameter setting workload is large and difficult to exhaustively combine parameters, resulting in the model not reaching the optimal state.

Method used

Using a reinforcement learning-based method, we automatically perform parameter combination and iterative optimization by building performance evaluation models and perceptual response models, and use the perceptual response mechanism of reinforcement learning to achieve parallel evaluation and iterative optimization of parameter combinations to avoid manual intervention.

Benefits of technology

The automated process of model training and evaluation is realized, the optimal parameter combination is found, and the optimal model is output, which improves training efficiency and accuracy, and avoids local optimal traps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338032A_ABST
    Figure CN120338032A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and a model training system based on reinforcement learning, and belongs to the technical field of intelligence. The method comprises the steps of performing fixed-area parameter combination on a parameter set of a to-be-trained model in a delimiting range of each parameter to obtain an initial parameter combination set; constructing a performance index-oriented performance evaluation model according to the data set and the initial parameter combination set; constructing a perception response model of reinforcement learning based on the performance evaluation model to carry out random parameter replacement iteration on each parameter in the reference parameter combination so as to determine an optimal performance parameter combination capable of generating a maximum reward value; and training the to-be-trained model by adopting the optimal performance parameter combination, and iterating the performance evaluation model to obtain an optimization target model. According to the model training method based on reinforcement learning, parameter combination setting can be carried out on models in batches, the process of completing model training and evaluation in an automatic mode is achieved, the optimal parameter combination is sought, and therefore an optimization model is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent technologies, and particularly to a model training method and a model training system based on reinforcement learning. Background Art

[0002] In the practical application of artificial intelligence, when training one or more models, it is often necessary to set the parameters of the models to different degrees, evaluate the impact of different parameter settings on the model performance, and finally select a set of parameters to make the model achieve better performance on the corresponding dataset.

[0003] In the process of traditional model parameter setting and model performance evaluation, two problems are likely to occur: one is that the setting process depends on manual coding, resulting in a large workload; the other is that it is difficult to effectively exhaust all parameter combinations manually, and it is easy to miss better parameter settings, resulting in the model not reaching the optimal state. Therefore, a more efficient and accurate model training method is needed. Summary of the Invention

[0004] In view of this, this paper proposes a model training method based on reinforcement learning, which can complete the process of model training and evaluation in an automated manner, seek the optimal parameter combination, and thus output an optimized training model.

[0005] To achieve the above object, the present invention provides a model training method based on reinforcement learning, including: performing fixed-region parameter combination on the parameter set of the model to be trained within the defined range of each parameter to obtain an initial parameter combination set; constructing a performance evaluation model for a performance index according to the dataset and the initial parameter combination set, wherein the performance evaluation model outputs the probability distribution of the performance index for any input parameter combination; constructing a perception response model of reinforcement learning based on the performance evaluation model, performing random parameter replacement iteration on each parameter in the benchmark parameter combination to determine the optimal performance parameter combination that can generate the maximum reward value; and training the model to be trained with the optimal performance parameter combination and iterating the performance evaluation model to obtain an optimized target model.

[0006] On the other hand, the present invention also provides a model training system based on reinforcement learning, including: a parameter combination device for performing zoned parameter combination on the parameter set of the model to be trained within the defined range of each parameter to obtain an initial parameter combination set; a first model construction device for constructing a performance evaluation model oriented to performance metrics according to the data set and the initial parameter combination set, wherein the performance evaluation model outputs the probability distribution of the performance metrics for any input parameter combination; a second model construction device for constructing a perception response model of reinforcement learning based on the performance evaluation model, performing random parameter replacement iteration on each parameter in the benchmark parameter combination to determine an optimal performance parameter combination that can generate the maximum reward value; and a model training device for training the model to be trained using the optimal performance parameter combination and iterating the performance evaluation model to obtain an optimized target model.

[0007] Through the above technical solution, the present invention proposes a model training method based on reinforcement learning. By constructing a performance evaluation model, the target model is trained by means of perception and response. Perception is reflected in collecting model performance data, and feedback is reflected in adjusting parameter settings through performance data, thereby completing the batch parameter combination setting of the model. The present invention can realize the process of model training and evaluation in an automated manner, seek the optimal parameter combination, and thus output an optimized model.

[0008] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. Brief Description of the Drawings

[0009] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification. Together with the following specific implementation, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:

[0010] Figure 1 is a schematic flowchart of a model training method based on reinforcement learning according to an embodiment of the present invention;

[0011] Figure 2 is a schematic flowchart of the technical route of a model training method based on reinforcement learning according to an embodiment of the present invention;

[0012] Figures 3 - 7 is a schematic diagram of the training process of an MLP model according to an embodiment of the present invention;

[0013] Figure 8 is a schematic structural diagram of a model training system based on reinforcement learning according to an embodiment of the present invention. Detailed Description of the Invention

[0014] The following will describe in detail the specific implementation manners of the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.

[0015] The present invention first provides a model training method 100 based on reinforcement learning. As Figure 1 shown, the model training method 100 may include:

[0016] Step S110: Perform fixed-region parameter combination on the parameter set of the model to be trained within the defined range of each parameter to obtain an initial parameter combination set.

[0017] First, a parameter set of the model to be trained (the target model to be established) can be established. For example, if the target model to be trained is M, then M can be expressed as the overall of multiple sub-models M = (m1, m2,..., m k ). Among them, the parameter set involved in M is: P = (p1, p2,..., p n ). In addition, the target model M of the present invention can also be multiple groups of models. In the case where each group of models affects each other, in order to avoid the situation where the local maximum value is not the overall optimal, it can be defined that:

[0018]

[0019] where l is the number of models. In this way, the present invention can be applied to more complex scenarios.

[0020] Among them, a parameter combination construction algorithm (such as the parameter fixed-region and fixed-boundary method) can be used to perform random parameter combination on the parameter set of the target model, so as to uniformize the initialization parameter settings. The advantage of such uniform initialization parameter settings is that it can shorten the path to find the optimal parameter combination and achieve a faster optimization process. That is, it can make the sampling of parameters cover the entire interval more evenly numerically, and make the parameter values (such as taking the interval average value) more representative.

[0021] Specifically, step S110 can be implemented according to the following steps:

[0022] Step S111: Limit the upper and lower bounds of the gathering interval for each parameter p n in the parameter set P = (p1, p2,..., p i ), and specify the number c i as the number of intervals to obtain the value interval and interval step size of each parameter p i .

[0023] Among them, the value of each parameter p i ∈P can be limited by the upper and lower bounds to obtain p iThe gathering interval is and a specified quantity c i is used as the number of intervals. Then the interval step of parameter p i is:

[0024]

[0025] After that, according to the interval step, within the gathering interval of p i , each interval or value interval of each p can be successively segmented. i

[0026] Step S112: Take the average of each value interval as the representative value to obtain the interval value set v i of each parameter p (i) .

[0027] Among them, the average of each interval can be taken as the representative value of this interval. Then the representative value of the j-th value interval of parameter p i is:

[0028]

[0029] Then the interval value set of parameter p i is:

[0030]

[0031] By setting multiple intervals through the above steps and performing value combinations within each interval, the various parameters can be evenly grouped. Specifically, setting the intervals is mainly to extract the central point within an interval, that is, using the mean value as the value of each parameter, which is by default more representative of the entire interval. Since these intervals are equivalent to the 8 vertices of a cube and some equilibrium points at the central positions, the performance prediction model trained with such parameter value combinations can better comprehensively display the prediction ability of the model, while the randomly selected parameter values later are equivalent to other points within the cube, and the framework constructed by the equilibrium points can better represent all the points of the entire cube.

[0032] Step S113: Determine the initial parameter combination set V i according to the interval value set v (i) of each parameter p init .

[0033] Therefore, the initial parameter combination set can be obtained based on the above multiple intervals. For example, the space V init = v (1) × v (2) ×…× v (n) = (v1, v2, …, v n ​) is the parameter value corresponding to the parameter P = (p1, p2,..., p n ), that is, the combination of all parameter interval values. Then V init can include a total of combination methods, expressed as:

[0034]

[0035] Step S120, construct a performance evaluation model for performance indicators according to the data set and the initial parameter combination set. Among them, the performance evaluation model outputs the probability distribution of performance indicators for any input parameter combination.

[0036] Specifically, this step S120 may include initializing model training, testing and performance evaluation, and a performance evaluation reward function, etc. In the embodiments of the present invention, the performance indicators may include one or more of the following: efficiency, accuracy, resource consumption, hardware requirements, etc. Among them, accuracy and performance are generally the most important indicators. Other indicators also include, for example, the cost of key indicators (corresponding to computing power requirements), the number of parameters (corresponding to miniaturization), and hardware requirements (corresponding to vendor dependence), etc. The present invention does not limit this.

[0037] Specifically, in the initialization step of determining the performance indicators, step S120 may include:

[0038] Step S121, divide the data set into a training set and a test set.

[0039] That is, divide the data set S, and divide it into a training set S train and a test set S test . It should be noted that different data sets have different requirements for the division method, and the present invention does not limit the specific division method. For example, the "proportional division" or "random division" method can be used.

[0040] Then, for the target model parameters to be trained: P = (p1, p2,..., p n ), randomly take values V i =(v ′ 1, v ′ 2,..., v ′ n ′ n ) within the defined range of each parameter p ′ | V′ .

[0041] Step S122, use the training set to train each parameter combination in the initial parameter combination set to obtain multiple target models.

[0042] Specifically, the data set Strain Train M ′ | V′ to obtain the model to be tested. Among them, c initial models (i.e., the target models above) can be obtained through init the combinations of parameters of V in step S110

[0043] Step S123: Use the test set to test multiple target models to obtain the performance metrics and reward values of each target model.

[0044] Then, the dataset S test can be used to test the model to be tested to obtain the model to be evaluated, so as to calculate the training performance value of its performance metrics. The following takes the calculation of the efficiency value and accuracy rate as an example:

[0045] 1) Method for calculating the efficiency value: The trimming mean method can be used. The time spent for each test data to run is sorted from small to large as Set the trimming rate to γ shrink %, delete the first γ shrink % of the data and the last γ shrink % of the data in the set, and calculate the average value to obtain t ′ , and take the efficiency value

[0046] 2) Method for calculating the accuracy rate: Let the total number of S test be n test , the number of accurate predictions be n ′ T , and the number of prediction errors be n ′ F , then the accuracy rate is: It can be understood that other efficiency value and / or accuracy rate calculation methods can also be used, and the present invention is not limited thereto.

[0047] Furthermore, the above c initial models can be trained using the dataset S train and tested using S test , and the measured efficiency value is: e init =(e1, e2,..., e c ), and the accuracy rate is: u init =(u1, u2,..., u c ). Furthermore, the reward values r init =(r1, r2,..., r c ) of the c initial models can also be calculated.

[0048] Among them, the reward value can be determined by the following steps:

[0049] 1) When there is only one performance metric, determine the reward value according to this performance metric;

[0050] 2) When there are multiple performance metrics, determine the reward value according to the performance weights of the multiple performance metrics and each performance metric.

[0051] Specifically, taking the efficiency value and accuracy rate as examples, first use the e ′ | V′ and u ′ of the target model M ′ as the efficiency benchmark value and accuracy benchmark value. Suppose the efficiency value and accuracy rate of the trained model are e (k) and u (k) respectively. Then, the improvement calculation of the efficiency value and accuracy rate can be performed through the following formula:

[0052]

[0053] For the case where the improvement value is negative, that is, negative optimization, the algorithm provided by the present invention can set a zeroing feedback mechanism in the follow-up to automatically discard the negative value combination.

[0054] Furthermore, the performance evaluation reward can be further defined so as to obtain the reward value r init = (r1, r2,..., r c ) corresponding to each initial model:

[0055]

[0056] where γ e and γ u are the performance weights of the efficiency value and accuracy rate respectively. It can be understood that the setting rules for the weights can be formulated according to actual needs, and the present invention does not limit this.

[0057] It can be seen that the reward function benchmark setting method of the present invention can comprehensively consider various performance metrics. In the text, only efficiency and accuracy are used as examples, but in fact, any item of metrics can be selected as the optimization target according to requirements. Therefore, the present invention has good applicability.

[0058] Step S124: According to each parameter combination and its corresponding performance metric, construct a fully connected neural network model for each performance metric as the performance evaluation model for this performance metric, where the performance evaluation model outputs the value and probability distribution of this performance metric for any input parameter combination.

[0059] Among them, based on the above parameters, a fully connected neural network model can also be constructed as the performance evaluation model: Me and M u , respectively for efficiency evaluation and accuracy evaluation. For example, (V init , e init ) can be used to split into and as the training set and the test set respectively, and train the efficiency value prediction model M e . The model output is the probability distribution of the efficiency value. It should be noted that the above probability distribution refers to the distribution of the output values. For example, if the output performance value is 8 with a probability of 95%; the performance value is 8.5 with a probability of 2%; the performance value is 9 with a probability of 3%, then the above data constitutes a probability distribution of the performance values. Similarly, (V init , u init ) can be used to split into and as the training set and the test set respectively, and train the accuracy prediction model M u . The model output is the probability distribution of the accuracy.

[0060] Step S130, based on the performance evaluation model, construct a perception response model for reinforcement learning, and perform random parameter replacement iteration on each parameter in the benchmark parameter combination to determine the optimal performance parameter combination that can generate the maximum reward value.

[0061] Among them, the random parameter replacement iteration of each parameter in the benchmark parameter combination in step S130 may include the following steps:

[0062] Step S131, randomly take values within the defined range of each parameter in the parameter set P=(p1, p2,..., p n ) to determine the benchmark parameter combination;

[0063] For example, for the target model parameters to be trained: P=(p1, p2,..., p n ), randomly take values V ′ =(v ′ 1, v ′ 2,..., v ′ n ) within the defined range of each parameter, and establish the target model M ′ | V′ .

[0064] Step S132, randomly extract k parameters from the parameter set P=(p1, p2,..., p n ), and randomly take values within the defined range of each of the k parameters to replace the original values to obtain the reward value of the parameter combination after random replacement; and

[0065] Specifically, first, from P=(p1, p2,..., pn ) Randomly extract k parameters, denoted as k = (1, 2, …, n). Then, from the value intervals (I r1 , I r2 , …, I rk ) corresponding to each of these k parameters, randomly select n r numbers (n r represents the value space of each parameter, and the number of values taken corresponds to the number of k values), to obtain a set of random values of the random parameter :

[0066]

[0067] Among them, the combination form of the random values: It can be seen that there are a total of n r to the power of k combination forms. It can be understood that this step is to randomly replace one or more parameters (k parameters) on the basis of the initial parameter set of the target model to be output. Among them, the set of random values can represent: all possible values of a certain parameter among the k parameters within its defined range For example, if each parameter has n r value spaces, then the product of the value spaces of the k parameters is the total combination form of the random value replacement, denoted as n r to the power of k.

[0068] The principle of the above step setting is: The planned search in the prior art usually has a direction. Due to the limitations of the planning algorithm itself, it is easy to guide the search behavior into a space without an optimal solution for search, or unable to jump out of this local space and regard the local optimal solution as the optimal solution. The advantage of the random search specified by the present invention is that there is no clear directivity, and all regions are equivalent. In this way, an initial parameter combination form with uniform values and an updated parameter combination form can be obtained.

[0069] It can be seen that the method for constructing a random parameter combination provided by the present invention can make full use of randomness and will not fall into the local space created by a specific plan. Therefore, it can avoid the search from entering the black area.

[0070] Step S133, perform random parameter replacement on the previous parameter combination in an iterative manner to obtain the reward value of the parameter combination after each random replacement.

[0071] It can be understood that when performing random replacement of parameters for the first time, the original parameters use the benchmark parameter combination, and then the parameter values are replaced using this algorithm. In the subsequent iterative process of parameter replacement, all are based on the previous parameter combination and then iteratively replace k newly extracted parameter sets.

[0072] In addition, the perception response model of reinforcement learning in this step S130 specifically includes: agent model M agent =(S, A, R, T). Among them, S=(s1, s2, …, s u ), is the state space, including the state information of all performance indicators that the agent model can perceive. For example, s u can include the efficiency value e, the accuracy u, the hardware requirements, etc.

[0073] A=(a1, a2, …, a m ), is the action space, representing the set of behaviors for randomly replacing parameters for each parameter combination. For example, it includes all actions that the agent can take in each state. Each action a i (i = 1, 2, …, m) represents one of all the parameter combinations in the initial parameter combination set obtained by the algorithm shown in the above step S110. Therefore and a i The specific implementation process is to form a new combination by replacing the corresponding parameter values in the current parameter combination with the random parameters and parameter values obtained by using the random parameter replacement method.

[0074] R is the reward function, representing the reward value generated by each action. For example R(s, a, s′) represents the reward value that the agent obtains from the environment when performing action a in state s and transitioning to state s′; its specific reward value is the new combination formed by replacing the corresponding parameter values in the current parameter combination with the random parameter combination obtained by action a; and taking the parameter combination of the new combination as the input value of the models M e and M u to perform operations, thereby obtaining the distributions of the efficiency value e and the accuracy value u. And the reward value is, for example, selecting the efficiency value e and the accuracy value u and evaluating them through the reward function in the above step S123. That is, the performance evaluation model can include the efficiency model M e (a) and the accuracy rate model M u (a), and determining the reward function R(s, a, s ′ ) with the following formula:

[0075] R(s, a, s ′ ) = γ e M e (a) + γ u M u (a)| s→s′

[0076] Among them, the efficiency model M e (a) and the accuracy rate model M u(a) It is constructed by the following formula:

[0077]

[0078] where n a is the first n e samplings among the total number of distributions obtained by running the input a through M a ; is the i-th output value in the distribution output by M e ; is the probability of the i-th output value; m a is the first m u samplings among the total number of distributions obtained by running the input a through M a ; is the i-th output value in the distribution output by M e ; is the probability of the i-th output value.

[0079] T is the state transition function, which characterizes the probability of each action occurring. Here, a performance model is used to predict the state change. Therefore, the parameters of the state transition function are the parameters of the prediction model, and the transition probability is the probability of the performance value output after the prediction model performs the calculation. For example, T: S × A × S → [0, 1], T(s, a, s ′ |θ) represents the probability of executing action a in state s and transitioning to state s′. The state transition function is determined by a series of parameters θ(θ1, θ2, …, θ m ), which is the result of the actions of models M e and M u . T = T(M e , M u ), and the parameters θ(θ1, θ2, …, θ m ) are the parameters obtained by training models M e and M u .

[0080] In summary, the entire agent model is equivalent to switching between two performances (states) using a parameter combination (action), and the switching probability in the middle is the performance probability output by the model.

[0081] After describing the basic architecture of the above agent model, the following steps can also be performed based on the above agent model to determine the optimal performance parameter combination that can generate the maximum reward value:

[0082] 1) Determine the target action a * that can generate the maximum reward value as the optimal policy for reinforcement learning;

[0083] Among them, the initial parameter combination form can be searched by the perception response model agent to find the target action that generates the maximum reward value. That is, through the above agent model, the most suitable behavior a * can be found to represent the optimal policy of reinforcement learning:

[0084] a * = argmax a (γ e M e (a) + γ u M u (a))| s→s′

[0085] Among them, the reward is positively correlated with the performance and can also be negative, but generally negative numbers are set to 0. That is to say, in the process of each decision-making of the agent, it is a process of searching for the behavior that maximizes the reward, that is, finding the parameter combination that predicts the optimal performance through the performance model. Therefore, the above optimization function can simply and intuitively represent the optimal policy.

[0086] 2) Execute the target action a * on the state s and use the parameter combination corresponding to the transferred state s′ as the optimal performance parameter combination. That is, determine the parameter combination form corresponding to the target action as the target parameter combination form, which is used as the optimal policy of reinforcement learning.

[0087] Step S140: Train the model to be trained using the optimal performance parameter combination and iterate the performance evaluation model to obtain the optimized target model.

[0088] First, after obtaining the optimal parameter value combination through the above steps, the model can be set and trained to obtain the candidate model M ′ . To better perform the model optimization work, M ′ can be further evaluated for performance, and the evaluation data is recorded and put into the candidate model set S t .

[0089] Then, after the training of n t candidate models is completed, the model with the highest performance evaluation (reward score) can be selected from S t , or other models with the highest single performance can be selected according to the actual scenario requirements.

[0090] Finally, the training data can be optimized, and the performance prediction model can be iterated, including the iteration of the performance prediction model and the candidate set S tIteration. As can be seen from the above steps in this article, the performance prediction model is trained by the initial parameter combination set and the performance prediction data set formed after the initial parameter combination set is trained and the performance values are measured. The initial parameter combination set has the characteristic of uniform parameter value distribution, but there may be a lack of characterization of the data space with non-uniform distribution. Therefore, the performance prediction model needs to be iteratively trained. That is, on the basis of the performance prediction data set formed by combining the initial parameter combination set and the performance prediction model, incorporate the parameter combination values and the measured performance values described in the above step "candidate model set S t ", and incorporate them into the performance prediction data set, and retrain the performance prediction model to update the performance prediction model to achieve a more comprehensive display.

[0091] That is to say, after the model to be trained has been trained, the performance value can be obtained through the model. Then, if there is a large deviation between the performance value and the predicted value, then these trained models need to perform a performance iteration on the performance prediction model. In one embodiment, the performance prediction model can be iterated periodically. That is to say, the currently constructed performance prediction model can be used within a certain period, and after a certain period, the performance prediction model needs to be iterated. It can be understood that the above iteration period can be determined according to the actual situation, and the present invention does not limit this.

[0092] In addition, if the models owned by S t are not sufficient to meet the training needs and there is a need for continuous training, then the current training data (uniformly preferably a part of the parameter combinations and model performance data with representativeness or other key significance) is transferred to the initial data set for the next round of training for re-training, that is, to perform one or more rounds of model training work involved in this case. Through multiple rounds of iteration, the S t set is gradually expanded, and finally the available models are selected from the total candidate set S t formed by the union of the candidate models generated in all rounds.

[0093] Through the above technical solutions, the key points of this article are:

[0094] 1) In the overall work process, the present invention can be used to batch-train models, find the optimal parameter settings, rather than manually adjusting the training parameters. Through the perception-response mechanism, parallel evaluation and iterative optimization of parameter combinations are realized, avoiding manual intervention. The Agent model feeds back the performance data to the parameter adjustment module to form a closed-loop optimization process;

[0095] 2) Through the parameter zoning and delimiting method in step S110, the initialization parameter settings can be homogenized. The advantage of homogenization is that the sampling of parameters numerically covers the entire interval more evenly, so that the sample of parameter values (taking the interval average) is more representative. This method can avoid the distribution deviation problem of traditional random search, ensure the full coverage of the parameter space, and reduce the local optimal trap;

[0096] 3) Through the initial set generation method and the performance evaluation model training method, a neural network or machine learning model can be used to predict the model performance of parameter combination settings, and the value-taking direction can be planned based on the predicted performance. For example, a performance evaluation model based on a fully connected neural network (MLP) can be constructed, and the probability distribution of multi-objective performance indicators (such as efficiency and accuracy) can be output after inputting the parameter combination;

[0097] 4) Through the reward function benchmark setting method in step S120, various performance indicators can be comprehensively considered. For example, based on the weighted reward function, conflicting indicators such as efficiency (such as training time) and accuracy can be optimized simultaneously to meet the requirements of complex scenarios. At the same time, the performance evaluation model supports extension to other indicators (such as hardware resource consumption), and only need to add corresponding neural network branches;

[0098] 5) Through the optimization function in step S130, it is simple and intuitive, and the process of model training and evaluation can be completed in an automated manner, thereby outputting an optimized model.

[0099] Through the above technical solutions, this paper proposes a model training method based on reinforcement learning. By using the model performance evaluation method, including the initial set generation method and the performance evaluation model training method (such as using a neural network / machine learning model), a performance evaluation model is constructed and the model performance of parameter combination settings is predicted. Among them, the present invention can train the target model in a perception and response manner. For example, the target model can be trained by setting an agent to perform the perception and response method. The value-taking direction is planned based on the predicted performance. Perception is reflected in collecting model performance data, and feedback is reflected in adjusting parameter settings through performance data, so as to complete the batch parameter combination setting of the model. Therefore, the present invention can batch-train the model through the above steps, and find the optimal parameter settings instead of manually adjusting the training parameters, thereby realizing the process of model training and evaluation in an automated manner, seeking the optimal parameter combination and outputting an optimized model.

[0100] In addition, the technical roadmap of the present invention can refer to Figure 2 , which shows the overall logic of the present invention. In the practical application of an embodiment, the above solution of the present invention can be executed according to the following steps:

[0101] Step S21: Establish a parameter set.

[0102] The model to be trained: M. Among them, M can be expressed as the whole of multiple models M = (m1, m2,..., m k ). The parameter set involved in M is: P = (p1, p2,..., p n ), when M is multiple models l is the number of models.

[0103] Step S22: Initial parameter combination set.

[0104] 1) Parameter value definition range

[0105] For each parameter p i ∈P, limit the upper and lower limits of its value. The value range of p i is And specify the number c i as the number of intervals. Then the interval step size of parameter p i is:

[0106]

[0107] Take the average value of the interval as the representative value of this interval. Then the representative value of the j-th value interval of parameter p i is:

[0108]

[0109] The interval value set of parameter p i is:

[0110]

[0111] 2) Obtaining the initial parameter combination set

[0112] Space V init = v (1) ×v (2) ×...×v (n) = (v1, v2,..., v n ) is the parameter value corresponding to parameter P = (p1, p2,..., p n ), that is, the combination of all parameter interval values. V init can include a total of combination methods:

[0113]

[0114] Step S23: Initialization.

[0115] Split the data set into a training set S train and a test set Stest ; For the target model parameters P = (p1, p2,..., p n ) to randomly take values V ′ = (v ′ 1, v ′ 2,..., v ′ n ) within the defined range of each parameter, and establish the target model M ′ | V′ .

[0116] Step S24: Model training, testing, performance evaluation, and performance evaluation reward function

[0117] First, the dataset S train can be used to train M ′ | V′ to obtain the model to be tested, and the dataset S test can be used to test M ′ | V′ to obtain the model to be evaluated, and calculate its performance value. The performance metrics can include efficiency, accuracy, resource consumption, hardware requirements, etc. Here, the efficiency value and accuracy are taken as examples:

[0118] 1) Method for calculating the efficiency value: Using the winsorized mean method, sort the time taken for each test data to run from smallest to largest as Set the winsorization rate to γ shrink %, delete the first γ shrink % of the data and the last γ shrink % of the data in the set, and calculate the average value to obtain t ′ , and take the efficiency value

[0119] 2) Method for calculating the accuracy: Let the total number of S test be n test , the number of accurately predicted be n ′ T , and the number of wrongly predicted be n ′ F . Then the accuracy is:

[0120] Then, the e ′ | V′ and u ′ of the initial model M ′ can be used as the efficiency benchmark value and accuracy benchmark value. Let the efficiency value and accuracy of the trained model be e (k) and u (k) . Based on this, perform the improvement calculation to obtain the efficiency value improvement as The accuracy improvement is

[0121] Finally, the performance evaluation reward can be set. where γ e and γ u are the performance weights of the efficiency value and the accuracy rate respectively.

[0122] Step S25: Train the initial model set and the performance reward prediction model with the parameter combination set.

[0123] First, c initial models can be trained using the parameter combinations of V init in step S22 with the data set S train . Test the above c initial models using S test , and measure the efficiency value as: e init =(e1, e2,..., e c ), and measure the accuracy rate as: u init =(u1, u2,..., u c ). Calculate the reward values r init =(r1, r2,..., r c ) of the c initial models.

[0124] A fully connected neural network model can be constructed as the performance evaluation model: M e and M u , respectively for efficiency evaluation and accuracy rate evaluation. Use (V init , e init ) to be divided into and as the training set and the test set respectively to train the efficiency value prediction model M e , and the model output is the probability distribution of the efficiency value; use (V init , u init ) to be divided into and as the training set and the test set respectively to train the accuracy rate prediction model M u , and the model output is the probability distribution of the accuracy.

[0125] Step S26: Parameter combination construction algorithm

[0126] Randomly extract k parameters from P=(p1, p2,..., p n ), denoted as k=(1, 2,..., n). Randomly select n r1 numbers from the value intervals (I r2 , I rk ) corresponding to these k parameters to obtain the random value set of the random parameter r : ​

[0127]

[0128] Combination forms of random values: There are a total of n r to the kth power of combination forms.

[0129] Step S27: The agent establishes

[0130] The perception response model of reinforcement learning is the agent: M agent =(S, A, R, T), including:

[0131] 1) State space: S = (s1, s2,..., s u ), which contains all the state information that the agent can perceive. s u can include the efficiency value e, the accuracy u, the hardware requirements, etc. For example, S = (e, u), where e is the efficiency value of the model and u is the accuracy performance value of the model.

[0132] 2) Action space: A = (a1, a2,..., a m ), which contains all the actions that the agent can take in each state. Each action a i (i = 1, 2,..., m) represents one of all the parameter combinations obtained through the algorithm shown in step S6. Therefore a i The specific implementation process is to form a new combination by replacing the corresponding parameter values in the current parameter combination with the obtained random parameters and parameter values.

[0133] 3) The reward function is: R(s, a, s′) represents the reward value that the agent obtains from the environment when performing action a in state s and transitioning to state s′; its specific reward value is a new combination formed by replacing the corresponding parameter values in the current parameter combination with the random parameter combination obtained through action a. The parameter combination of the new combination is used as the input value of the models M e and M u to perform operations, obtaining the distributions of the efficiency value e and the accuracy value u, and the reward value is evaluated through the reward function in step S24 by the efficiency value e and the accuracy value u, that is:

[0134] R(s, a, s ′ ) = γ e M e (a) + γ u M u (a)| s→s′

[0135] where

[0136]

[0137] where n a is the first n e samplings among the total number of distributions obtained by running input a through M a , is the i-th output value in the distribution output by M e , is the probability of the i-th output value, m a is the first m u samplings among the total number of distributions obtained by running input a through M a , is the i-th output value in the distribution output by M e , is the probability of the i-th output value.

[0138] 4) The state transition function is: T: S × A × S → [0, 1], T(s, a, s ′ |θ) represents the probability of executing action a on state s and transitioning to state s′. The state transition function is determined by a series of parameters θ(θ1, θ2, …, θ m ), that is, the result of the actions of models M e and M u . T = T(M e , M u ), and the parameters θ(θ1, θ2, …, θ m ) are the parameters obtained by training models M e and M u .

[0139] Step S28: Optimization function

[0140] represents the optimal policy of reinforcement learning by finding the most suitable action a * , that is:

[0141] a * = argmax a (γ e M e (a) + γ u M u )(a))|[[]END]] s→s′

[0142] Therefore, the process of each decision-making of the agent is a process of searching for the action that maximizes the reward, that is, finding the parameter combination that predicts the optimal performance through the performance model.

[0143] Step S29: Model training and output

[0144] ​​1) Set and train the model with the optimal parameter value combination obtained through the above steps to obtain the candidate model M ′ , and perform performance evaluation on M ′ , record the evaluation data, and put it into the candidate model set S t .

[0145] 2) After the training of the candidate models with the number to be completed being n t , select the model with the highest performance evaluation (reward points) from S t , or select the model with the highest single performance according to the actual scenario requirements.

[0146] 3) Follow-up processing, including the iteration of the performance prediction model and the iteration of the candidate set S t .

[0147] As can be seen from step S25 above, the performance prediction model is trained by the performance prediction data set formed after the initial parameter combination set and the initial parameter combination set are trained and the performance values are measured. The initial parameter combination set has the characteristic of uniform parameter value distribution, but there may be a lack of characterization of the characteristics of the data space with non-uniform distribution. Therefore, the performance prediction model needs to be iteratively trained. That is, on the basis of the performance prediction data set formed by combining the initial parameter combination set and the performance prediction model, the parameter combination values described in the above step "candidate model set S t " and the measured performance values are incorporated into the performance prediction data set, and the performance prediction model is retrained to update the performance prediction model to achieve a more comprehensive display.

[0148] In addition, if the models owned by S t are not sufficient to meet the training needs and there is a need for continuous training, then the current training data (uniformly select a part of the parameter combinations and model performance data with representativeness or other key significance) can be transferred to the initial data set of the next round of training for re-training, that is, perform one or more rounds of model training work involved in this case. Through multiple rounds of iteration, the S t set is gradually expanded, and finally an available model is selected from the total candidate set S t formed by the union of the candidate models generated in all rounds.

[0149] Through the above technical roadmap, the key points of the present invention include:

[0150] 1) Workflow: used to batch-train the model to find the optimal parameter settings;

[0151] 2) Parameter zoning and delimiting method (step S22): used to uniformly initialize the parameter settings;

[0152] 3) Model performance evaluation method: initial set generation method and performance evaluation model training method (step S24), using a neural network to predict the model performance of parameter combination settings;

[0153] 4) Reward function benchmark setting method (step S24): comprehensively considers various performance indicators;

[0154] 5) Random parameter combination construction method (step S26): makes full use of randomness to prevent the search from entering a black area;

[0155] 6) Optimization function (step S27): simple and intuitive.

[0156] Embodiment

[0157] To further explain the present invention, embodiments of specific applications are also provided. Taking the training of an MLP model (Multilayer Perceptron, a feedforward artificial neural network model) as an example, a method for evaluating the requirement development ability on a CCR-SSP system based on the MLP model is provided. The background of this embodiment: During the operation of the software system CRR-SSP, various new requirements will be generated to cope with new situations, and many of these requirements require the support of code development work. The CRR-SSP system incorporates a large number of software modules, and many modules are also used in other various systems. For example, the crr0266 module exists not only in CRR-SSP but also in the CRR-PES and CRR-AMT systems. Software engineers generally do not have work experience with all related modules of a system, but there are similarities and correlations among many modules. Therefore, if an engineer's work experience involves certain software modules, it is easier to understand some other related and similar modules, and then it will be helpful for the understanding of the entire system development work and it will be easier to be competent for this maintenance work.

[0158] Therefore, the goal of this embodiment is to judge whether a software engineer is competent for the development work of CRR-SSP by examining the development experience that the software engineer already has. The embodiment uses the MLP model to learn and train through the method provided by the present invention, and finally outputs a model for actual application. The specific process can refer to Figures 3 - 7 , including:

[0159] Step S31, prepare the data set and initialize its division into a training set and a test set:

[0160] Prepare the work experience of software engineers who have engaged in CRR-SSP development in the past before engaging in this work, and the final work evaluation (whether competent) as the data set.

[0161] Among them, such asFigure 3 As shown in the table, the columns crr0xxx (where 0xxx is the data code) all represent module numbers. The content 1 in the table indicates that the engineer fully understood the module before engaging in the development work of CRR-SSP, 0 indicates no relevant experience, and the value between 0 and 1 represents the degree of understanding. The column "Whether Competent" represents the post-work evaluation of the engineer, where 1 indicates competent and 0 indicates incompetent. This dataset records the work experience of the engineer before engaging in the development work of CRR-SSP and the post-work performance evaluation. Thus, it can be used to train a model for "predicting the work experience that is already known but not knowing whether one is competent for the development work of CRR-SSP".

[0162] Step S32: Establish an MLP model and construct a set of models to be trained using a parameter combination method;

[0163] First, establish a parameter set, as Figure 4 shown in the figure. The column parameters in the figure include: number of layers, number of neurons, activation function, solver, regularization coefficient, learning rate, initial learning rate, maximum number of iterations, shuffling, random seed, convergence threshold, etc., all of which are parameters of the MLP model.

[0164] Then, obtain the initial parameter combination set V in a way of defining boundaries and regions. init , and the following takes parameters of several different types of values as examples to illustrate the boundary definition of parameter values, including:

[0165] Activation function: integer type, with a value range of (1, 4), where 1 represents the identity function, 2 represents the logistic function, 3 represents the tanh function, 4 represents the relu function, and the defined region is divided into 4 intervals with a step size of 1;

[0166] Learning rate: floating-point type, with a value range of (0.0001, 0.005). In this example, the defined region is divided into 10 intervals, where the step size is 0.001;

[0167] Number of layers: integer type, with a value range of (2, 6). The model will generate parameter values according to the maximum number of layers. For example, if the maximum number of layers in this example is 6, then six parameters of number of layers 1, number of layers 2, number of layers 3, number of layers 4, number of layers 5, and number of layers 6 will be generated;

[0168] Number of neurons: array type, with a value range of (2, 5) for each layer. According to the different number of layers, when performing parameter combination, if the layer value is 2, then there are specific numbers of neurons in layers 1 - 2, and the number of neurons in layers 3 - 6 is 0. The boundary definition of the number of neurons for each layer is (2, 5), and the defined region is set manually with values.

[0169] The initial parameter combination set V formed by the above parameters init, c initial models are trained using the training set, and then fully connected neural network models for efficiency evaluation and accuracy evaluation can be constructed as performance evaluation models. For example, Figure 5 shows the parameter configuration of one of the models.

[0170] Step S33: Perform performance prediction on the model set through the method provided by the present invention, and preferably finally obtain the performance record of the optimized model through training; that is, the performance prediction of the optimized model can be obtained through the perception response model agent. As Figure 6 shows the column accuracy rate, error rate, etc., which all demonstrate the performance prediction of the performance model.

[0171] Step S34, Figure 7 is specific test data, which can be used for detailed analysis of the data;

[0172] Step S35: Through iterative training, optimize and finally output an optimal target model;

[0173] Step S36: Use the output optimal model as the final model, so that the model can be used to predict and judge whether a software engineer is competent for a development task.

[0174] Through the above technical solutions, this paper proposes a model training method based on reinforcement learning. By using model performance evaluation methods, including the initial set generation method and the performance evaluation model training method (such as using neural network / machine learning models), a performance evaluation model is constructed and the model performance of parameter combination settings is predicted. Among them, the present invention can train the target model in a perception and response manner. For example, the target model can be trained by setting the agent to perform the perception and response method. The value direction is planned through the predicted performance. Perception is reflected in collecting model performance data, and feedback is reflected in adjusting parameter settings through performance data, so as to complete the batch parameter combination setting of the model. Therefore, the present invention can batch train the model through the above steps, and find the optimal parameter settings instead of manually adjusting the training parameters, so as to realize the process of model training and evaluation in an automated manner, seek the optimal parameter combination and output the optimized model.

[0175] On the other hand, the present invention also provides a model training system 200 based on reinforcement learning, which can be implemented based on the above-mentioned model training method 100 based on reinforcement learning. As Figure 8 shown, the system 200 may include:

[0176] The parameter combination device 210 is used to perform zoning parameter combination on the parameter set of the model to be trained within the defined range of each parameter to obtain an initial parameter combination set;

[0177] The first model construction device 220 is configured to construct a performance evaluation model for performance metrics according to the data set and the set of initial parameter combinations, wherein the performance evaluation model outputs the probability distribution of the performance metrics for any input parameter combination;

[0178] The second model construction device 230 is configured to construct a perception response model for reinforcement learning based on the performance evaluation model, and perform random parameter replacement iteration on each parameter in the benchmark parameter combination to determine the optimal performance parameter combination that can generate the maximum reward value; and

[0179] The model training device 240 is configured to train the model to be trained by using the optimal performance parameter combination and iterate the performance evaluation model to obtain an optimized target model.

[0180] In summary, through the above technical solutions, this paper proposes a model training solution based on reinforcement learning. The target model is trained by constructing a performance evaluation model to perform perception and response. Perception is reflected in collecting model performance data, and feedback is reflected in adjusting parameter settings through performance data, so as to complete the parameter combination setting for the model in batches. The present invention can realize the process of model training and evaluation in an automated manner, seek the optimal parameter combination, and thus output an optimized model.

[0181] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.

[0182] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A model training method based on reinforcement learning, characterized in that The model training method includes: Performing zone parameter combination on the parameter set of the model to be trained within the defined range of each parameter to obtain an initial parameter combination set; Constructing a performance evaluation model oriented to performance indicators according to the data set and the initial parameter combination set, wherein the performance evaluation model outputs the probability distribution of the performance indicators for any input parameter combination; Constructing a perception response model for reinforcement learning based on the performance evaluation model, and performing random parameter replacement iteration on each parameter in the benchmark parameter combination to determine the optimal performance parameter combination that can generate the maximum reward value; and Training the model to be trained with the optimal performance parameter combination and iterating the performance evaluation model to obtain an optimized target model.

2. The model training method according to claim 1, wherein Performing zone parameter combination on the parameter set of the model to be trained within the defined range of each parameter includes: For each parameter p in the parameter set P = (p1, p2,..., p n ), perform upper and lower limit gathering interval limitation and specify the quantity c i as the number of intervals to obtain the value interval and interval step size for each of the said parameters p i as follows: i Take the average of each value interval as the representative value Obtain each of the parameters p i The set of interval values is as follows: and According to the set of interval values v i for each parameter p (i) , determine the set of initial parameter combinations V init = v (1) × v (2) × … × v (n) .

3. The model training method according to claim 1, wherein Constructing a performance evaluation model oriented to performance indicators according to the data set and the initial parameter combination set includes: Dividing the data set into a training set and a test set; Training each parameter combination in the initial parameter combination set with the training set to obtain multiple target models; Testing the multiple target models with the test set to obtain the performance indicators and reward values of each target model; and Constructing a fully connected neural network model oriented to each performance indicator as the performance evaluation model of the performance indicator according to each parameter combination and its corresponding performance indicator, wherein the performance evaluation model outputs the value and probability distribution of the performance indicator for any input parameter combination.

4. The model training method according to claim 1, wherein Performing random parameter replacement iteration on each parameter of the benchmark parameter combination includes: Randomly take values within the bounding ranges of the parameters in the parameter set P=(p1, p2,..., p n ), and determine the benchmark parameter combination; Randomly extract k parameters from the parameter set P = (p1, p2,..., p n ) and randomly take values within the delimiting range of each of the k parameters to replace the original values, so as to obtain the reward value of the parameter combination after random replacement; and Performing random parameter replacement on the previous parameter combination in an iterative manner to obtain the reward value of the parameter combination after each random replacement.

5. The model training method according to claim 1, wherein The perception response model for reinforcement learning includes: an agent model (S, A, R, T), wherein S is the state space, including the state information of all performance indicators that the agent model can perceive; A is the action space, representing the set of behaviors for performing random parameter replacement on each parameter combination; R is the reward function, representing the reward value generated by each action; T is the state transition function, representing the probability of each action occurring.

6. The model training method according to any one of claims 1-5, characterized in that, The performance indicators include one or more of the following: efficiency, accuracy, resource consumption, hardware requirements.

7. The model training method according to claim 5, wherein When the performance metrics include efficiency and accuracy, the performance evaluation model includes an efficiency model M e (a) and an accuracy model M u (a), and the reward function R(s,a,s ′ ) is determined by the following formula: R(s,a,s ′ ) = γ e M e (a) + γ u M u (a)| s→s′ Among them, R(s, a, s′) represents the reward value obtained when the agent model performs action a on state s and transfers to state s′; γ e is the performance weight for the said efficiency, γ u is the performance weight for the said accuracy rate.

8. The model training method according to claim 7, wherein The efficiency model M e (a) and the accuracy model M u (a) are constructed by the following formula: where n a is the first n e samplings in the total number of distributions obtained by running the input action a through M a ; is the i-th output value in the distribution output by M e ; is the probability of the i-th output value. m a The first m u samplings among the total number of distributions obtained by running the input a through M a samplings, where e is the i-th output value in the distribution output by M, and is the probability of the i-th output value.​ 9. The model training method according to claim 7 or 8, wherein The perception response model is used to perform the following steps to determine the optimal performance parameter combination that can generate the maximum reward value: Determine the target action a that can generate the maximum reward value * , as the optimal policy for reinforcement learning: a * = argmax a (γ e M e (a) + γ u M u (a))| s→s′ ; Execute the target action a on the state s * And use the parameter combination corresponding to the state s'transferred to as the optimal performance parameter combination.

10. A model training system based on reinforcement learning, characterized in that, The model training system includes: A parameter combination device for performing zone parameter combination on the parameter set of the model to be trained within the defined range of each parameter to obtain an initial parameter combination set; The first model construction device is used to construct a performance evaluation model for a performance metric according to a data set and the set of initial parameter combinations, wherein the performance evaluation model outputs a probability distribution of the performance metric for any input parameter combination; The second model construction device is used to construct a perception response model for reinforcement learning based on the performance evaluation model, and perform random parameter replacement iteration on each parameter in the benchmark parameter combination to determine an optimal performance parameter combination that can generate the maximum reward value; and The model training device is used to train the model to be trained with the optimal performance parameter combination and iterate the performance evaluation model to obtain an optimized target model.

Citation Information

Cited By

  • Model application method and system based on equivalence class division hierarchical dimension reduction

    CN120930715A

  • Deep learning model training method and device and electronic equipment

    CN121599025A