Automatic operator selection method based on deep reinforcement learning in evolutionary computation
By using an automatic operator selection model trained by deep reinforcement learning, the optimizer operator is dynamically selected, which solves the problem of inflexible operator selection in traditional methods, realizes an efficient and adaptive optimization process, and improves the performance of black-box optimization.
Patent Information
- Application Number
- CN202510702131.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-26
AI Technical Summary
Existing black-box optimization methods have difficulty dynamically adapting to changes in the optimization process when faced with complex and high-dimensional problems. Traditional operator selection methods lack real-time perception capabilities, resulting in insufficient optimization performance.
An automatic operator selection method based on deep reinforcement learning is adopted. By extracting feature parameters to train the automatic operator selection model, the most suitable optimizer operator is dynamically selected. The reward mechanism and policy gradient optimization are combined to achieve adaptive operator selection.
It improves the convergence speed and flexibility of the optimization process, can dynamically adjust strategies at different optimization stages, and significantly improves search efficiency and solution quality.
Smart Images

Figure CN120706460A_ABST
Abstract
Description
Technical Field
[0001] This invention is primarily applied to optimization problems, particularly operator selection in evolutionary computation. It relates to an optimization method for automatic operator selection based on deep reinforcement learning. By designing an efficient operator selection mechanism, it improves search efficiency and optimization performance in a black-box optimization environment. Background Art
[0002] Black-box optimization (BBO) is a method for performing optimization without knowing the mathematical formula of the objective function. It is of great significance in many fields, from automated machine learning to bioinformatics. Due to its black-box nature, BBO optimization requires an iterative search process to explore and evaluate candidate solutions. In this optimization framework, the explicit expression of the objective function or gradient information cannot be directly utilized. Instead, function value feedback can be used to infer and improve the solution.
[0003] However, black-box optimization problems are extremely challenging, especially when faced with complex and high-dimensional problems. Since the objective function is non-differentiable or computationally expensive, BBO relies on heuristic or experience-driven methods, such as evolutionary operators (EA), particle swarm optimization (PSO), differential evolution (DE), etc. These optimization operators usually have different search strategies, such as the balance between global exploration and local exploitation, and therefore perform differently on different optimization problems. According to the "No-Free-Lunch Theorem", there is no universal optimization operator that can perform best on all problems. Therefore, in practical applications, choosing an appropriate optimization operator based on the characteristics of the solution space of different problems becomes a key task.
[0004] Currently, research on operator selection in BBO focuses primarily on algorithm selection (AS) methods. AS methods analyze the characteristics of the problem and select the most appropriate optimization operator to improve optimization efficiency. Traditional AS methods typically use machine learning or heuristic methods to predict the optimal operator based on certain statistical characteristics of the problem, such as problem landscape analysis information and historical optimization trajectories, and then maintain this operator throughout the optimization process. Compared to manually adjusting the optimization operator, AS methods can reduce human intervention and improve operator adaptability, and therefore have received widespread attention in the BBO field.
[0005] While adaptive operator selection (AS) methods have significantly improved optimization performance, existing research still faces numerous challenges. First, static AS methods only select an operator at the beginning of optimization and maintain it throughout the optimization process, making them difficult to adapt to dynamic changes during the optimization process. Because different optimization stages may require different search strategies, static AS methods fail to fully exploit the performance complementarity between multiple operators. Second, traditional AS methods rely on manually designed feature extraction methods. For example, Jorge Maturana et al. (Adaptive Operator Selection and Management in Evolutionary Algorithms 2012) proposed a series of adaptive operator selection mechanisms, including ExCoDyMAB and Blacksmith. These methods dynamically manage operators by analyzing historical operator performance data, improving the flexibility of optimization strategies. However, these methods primarily rely on heuristic strategies driven by historical information and lack real-time awareness of the current optimization state, making it difficult to flexibly adjust selection strategies based on feature changes at different stages of the problem. Furthermore, these related work has not fully utilized modern methods such as deep learning to model complex feature spaces. Consequently, they often suffer from insufficient adaptability and generalization capabilities for high-dimensional and complex tasks. These features may not fully capture the complexity of the optimization problem, thus affecting the accuracy of operator selection. In addition, current AS research mainly focuses on the selection at the optimizer level, but lacks a dynamic adaptation mechanism for the operators inside the optimizer, which limits the flexibility and generalization ability of the operators.
[0006] In response to the above challenges, how to achieve dynamic operator selection in the optimization process and make adaptive decisions based on problem characteristics and optimization process is an important research direction in the current BBO field. Summary of the Invention
[0007] The technical problem that the present invention aims to solve is to provide a black box optimization method that can automatically select an operator according to the optimization effect.
[0008] The present invention is achieved through at least one of the following technical solutions.
[0009] The automatic operator selection method based on deep reinforcement learning in evolutionary computing includes the following steps:
[0010] Use the trained automatic operator selection model to select the most appropriate operator for its optimizer to adjust the optimization strategy in the optimizer;
[0011] The training of the automatic operator selection model includes the following steps: extracting feature parameters, and training the automatic operator selection model based on deep reinforcement learning based on the extracted feature parameters.
[0012] Furthermore, extracting characteristic parameters includes the following steps:
[0013] S1. Build an operator pool for automatic operator selection by parsing the configuration file. The operator contains the BBO optimizer that can perform the optimization task and the parameters required for the optimizer to run.
[0014] S2. Obtain the candidate solution population data of the optimizer and preprocess it to ensure that the population data format is consistent;
[0015] S3. Based on the obtained candidate solution population data, extract feature parameters for training.
[0016] Furthermore, the candidate population data includes data such as population mean, population variance and distance between individuals.
[0017] Furthermore, the characteristic parameters include distribution characteristics of the candidate solution population, fitness statistical information and time dynamic characteristics of the optimization process.
[0018] Furthermore, the automatic operator selection model is a multi-layer perceptron model or a graph neural network.
[0019] Furthermore, the loss value formula of the automatic operator selection model is:
[0020] Target(Θ)=E f~D [∑Reward(tra,F)|Θ];
[0021] Where Θ is the selected dynamic operator selector model parameter; D is the distribution of black box problems; f is the problem sampled from the distribution D; tra is the obtained dynamic operator selector operator trajectory, Reward is the reward; E is the expectation of the cumulative reward, and the reward expectation expressed in the above formula is maximized through the backpropagation operator iteration. The maximum expectation Target (Θ * ) corresponds to the automatic operator selection model parameter Θ * , which is the optimal automatic operator selection model.
[0022] Furthermore, based on the reward mechanism, the rewards of the automatic operator selection model are calculated and optimized in collaboration with the underlying optimizer. The reward mechanism includes two key indicators:
[0023] The first is the fitness ratio of the optimal solution of the current candidate population to the optimal solution of the initial population, which is used to measure the overall improvement of the optimization process;
[0024] The second is the time factor accumulated during the optimization process to evaluate the optimization efficiency, guide the operator selection model to evolve in a more optimal direction, and improve the stability and adaptability of the optimization algorithm.
[0025] Furthermore, based on the rewards of the automatic operator selection model and combined with the selected gradient optimization operator, policy gradient calculation is performed to optimize the operator selection strategy, so that the automatic operator selection model can adaptively learn and adjust during the optimization process to adapt to different optimization scenarios, thereby improving the overall optimization performance.
[0026] A computer device of the present invention includes a memory and a processor, wherein the memory is electrically connected to the processor and stores a computer program. The device is characterized in that when the computer program is executed by the processor, the processor implements the method described.
[0027] A computer-readable storage medium of the present invention stores a computer program, wherein when the computer program is executed by a processor, the processor implements the method described above.
[0028] The present invention has the following advantages and effects compared to the prior art:
[0029] (1) The black-box optimization method based on automatic operator selection disclosed in the present invention accurately preprocesses the initial candidate solution population data and extracts the characteristic parameters of the population landscape analysis, accurately captures the dynamic changes of the problem, improves the convergence speed of the optimization process, and avoids the slow convergence problem commonly seen in traditional methods.
[0030] (2) The black-box optimization method based on automatic operator selection disclosed in the present invention can flexibly adapt to the characteristics of different problems according to the current optimization state by dynamically selecting operators and adjusting the optimization strategy in real time. Compared with the traditional fixed operator selection method, it is more flexible and adaptable and can effectively cope with the changing optimization environment.
[0031] (3) The black-box optimization method based on automatic operator selection disclosed in this invention utilizes the most advanced BBO optimizer and automatic operator selection for co-evolution, which can significantly improve the exploration efficiency and performance of the selector. At the same time, it rewards the generator for its ability to outperform the most advanced BBO optimizer. Figure 1 Provided is a flowchart of an automatic operator selection method based on deep reinforcement learning in evolutionary computing for an embodiment; Figure 2 Detailed optimization flow chart of the automatic operator selection method based on deep reinforcement learning in the embodiment; Figure 3 This is a structural diagram of the automatic operator selection model in an embodiment. DETAILED DESCRIPTION
[0032] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0033] This embodiment uses drone path planning as an example. By constructing an automatic operator selection model, it acts as a "dispatcher" to intelligently deploy optimization strategies within the optimizer in optimization tasks such as drone path planning, significantly improving optimization efficiency and path quality, enabling fast and accurate path planning in complex environments. In drone path planning tasks, the aircraft must autonomously plan a safe and energy-efficient flight path within a complex environment. This task typically exhibits black-box characteristics such as high dimensionality, non-convexity, and discontinuity, making it difficult to solve using traditional optimization methods. The method of the present invention extracts the spatial distribution characteristics and fitness information of the current candidate path population and utilizes an automatic operator selection model composed of agents trained using a deep reinforcement learning model to dynamically select the operator most suitable for the current optimization stage, thereby achieving precise control over the path search process. For example, in the initial search phase, a global exploration operator is favored to identify potential feasible paths, while in the convergence phase, a local refined search operator is favored to optimize path quality. Compared with traditional fixed operator strategies, this significantly improves the efficiency and quality of path planning and reduces reliance on expert experience and parameter adjustment.
[0034] like Figures 1 to 3 As shown, this embodiment provides an automatic operator selection method based on deep reinforcement learning in evolutionary computing, a meta-learning optimizer that can automatically select appropriate operators for optimization in black-box problem optimization tasks, and the method includes the following steps:
[0035] S1. Determine the operator pool for automatic operator selection.
[0036] First, we build an operator pool for automatic operator selection. This pool contains the optimizers that can perform the optimization task, as well as the key parameters required for each optimizer to ensure diversity and adaptability in the optimization process. The constructed operator pool is stored in the model code as a configuration file. The following example shows an operator pool built for a drone path planning problem.
[0037] As an example, in this embodiment, the operators selected are JDE21, MadDE, and NL-SHADE-RSP. The parameters that each operator depends on are as follows:
[0038]
[0039] Among them, proτ1 and proτ2 are adaptive mutation and crossover probability parameters, AGELmt and eps control the number of operator stagnation, and myEqs determines the reinitialization of the operator; pro m and pro best Control operator mutation probability, pro qBX Then control the crossover probability; n A and pro A To control the size of historical archives and the probability of using archives.
[0040] S2. Acquisition and preprocessing of the initial candidate solution population.
[0041] In this step, the underlying optimizer is first initialized by randomly sampling the target black-box optimization problem to obtain the initial candidate solution population data for the underlying optimizer. This data is then preprocessed to ensure consistency in the population data format and prevent outliers from affecting the optimization process. Furthermore, the fitness of the candidate solution population is evaluated to provide a basis for adjusting the optimization strategy.
[0042] Candidate population data includes data such as population mean, population variance, and inter-individual distances. Furthermore, due to the missing problem structure of black-box optimization problems, the only way to obtain optimization information is by inputting the solution and obtaining the evaluation value: a solution-evaluation value pair for subsequent optimization. Therefore, by inputting the candidate solution into the target black-box optimization problem, we can obtain the corresponding population evaluation value data, providing a reference for the optimization process.
[0043] S3. Based on the obtained candidate solution population data, extract feature parameters for training.
[0044] Based on the acquired candidate solution population data, characteristic parameters are extracted for optimization process analysis, including but not limited to the distribution characteristics of the current candidate solution population, fitness statistics, and the temporal dynamics of the optimization process. These characteristic parameters can comprehensively characterize the search state during the optimization process and provide a decision basis for subsequent operator selection.
[0045] As an optional implementation, feature parameters for training are extracted by the following steps:
[0046] S31. Extracting distribution features
[0047] By calculating the average distances within the candidate solution population, between the global optimal individual, and between the local optimal individuals, we analyze the distribution pattern of the population in the search space and determine whether it is concentrated, dispersed, or evenly distributed. This information can be used to adjust the operator selection strategy to adapt it to different search environments.
[0048] S32. Fitness Statistical Analysis
[0049] The fitness of the candidate population is evaluated, and the average difference between its fitness and that of the global and local optima is calculated. The variance of the current population fitness is also calculated to quantify the fitness distribution of individuals within the population. The combination of the fitness gap and variance can be used to infer the current search stage, such as whether it is in a state of convergence or still has significant exploration potential.
[0050] S33. Evolutionary feasibility assessment
[0051] The algorithm's evolutionary capabilities are quantified by comparing the computational cost of population evolution with the pace of random search and analyzing the performance differences between the current population and different adopted populations. Furthermore, the improvement ratio between the best and worst individuals is calculated to assess the effectiveness of the current search strategy, thereby optimizing the operator selection strategy and improving search efficiency.
[0052] S34. Extracting temporal dynamic features
[0053] The time it takes for the population to reach the optimal solution is recorded, and the remaining time is calculated based on the maximum number of allowed iterations. Furthermore, the proportion of time with consistently good performance and whether the local optimal solution can be updated to the global optimal solution are analyzed to determine the current optimization process's progress. The time it takes for the population to reach the optimal solution can be used to dynamically adjust operator selection to match the search requirements at different stages.
[0054] S35. Operator Historical Behavior Analysis
[0055] By tracking and analyzing the impact of each operator on the position changes of the best and worst individuals in the population, the optimization history vector is recorded, and the long-term optimization trends and effects of the operators are quantified. This historical data can be used to form characteristic parameters of the operator's historical performance, providing a reliable reference for subsequent operator selection and ensuring intelligent and adaptive adjustment of the optimization strategy.
[0056] S4. Use the feature parameters to train an automatic operator selection model. The automatic operator selection model is a neural network model, such as a multi-layer perceptron model, a graph neural network, etc.
[0057] The loss formula of the automatic operator selection model is:
[0058] Target(Θ)=E f~D [∑Reward(tra,f)|Θ] (2)
[0059] The reward expectation expressed in the above formula is maximized through backpropagation operator iteration. Where Θ is the selected dynamic operator selector model parameter; D is the distribution of black-box problems; f is the problem sampled from distribution D; tra is the obtained dynamic operator selector operator trajectory, Reward is the reward; E represents the expected cumulative reward;
[0060] According to Target(Θ * ) The operator with the largest corresponding value selects the model parameter Θ * , which is the optimal operator selection model.
[0061] Based on the extracted feature parameters, an automatic operator selection model trained with deep reinforcement learning is used to dynamically select the most suitable operator from the operator pool to dynamically adapt to different stages of the optimization process to optimize the candidate solution population of the underlying optimizer.
[0062] The present invention uses an automatic operator selection model to optimize the candidate solution population. This model relies on the operator pool determined in step S1 and combines the extracted features. The model selects a neural network model (such as a multi-layer perceptron model, a graph neural network, etc.) to achieve automatic operator selection, dynamically adapting to different stages of the optimization process, improving search efficiency and solution quality.
[0063] S5. Develop a reward mechanism and collaborative optimization. This embodiment calculates the rewards of the automatic operator selection model based on the optimized candidate solution population and collaborates with the underlying optimizer to continuously optimize the operator selection strategy.
[0064] In practice, the reward calculation in step S5 is based on two key metrics: the fitness ratio of the current candidate population's optimal solution to the initial population's optimal solution, which measures the overall improvement in the optimization process; and the accumulated time factor during the optimization process, which assesses optimization efficiency. This reward mechanism can effectively guide the operator selection model toward a more optimal evolution, improving the stability and adaptability of the optimization algorithm.
[0065] S6. Iteratively update the operator selection model using policy gradient optimization. This process uses the reward calculated in step S5 and the selected gradient optimization operator to perform policy gradient calculations to optimize the operator selection strategy. This allows it to adaptively learn and adjust to different optimization scenarios during long-term optimization, thereby improving overall optimization performance. Based on the characteristic parameters, appropriate operators are dynamically selected from the operator pool and the candidate solution population of the underlying optimizer is optimized.
[0066] The following is combined with Figure 2 The black box optimizer based on automatic operator selection provided by the embodiment of the present invention is developed using Python language. The schematic diagram of the specific optimizer training process is as follows: Figure 1 shown.
[0067] First, determine the operator pool for automatic operator selection and the dependent parameters that support the normal optimization of the operators.
[0068] Determine a black box problem distribution D as the training data set, and f is the problem sampled from distribution D. By randomly generating multiple candidate solution individuals x i ∈X, i=1, 2, .... The fitness of the candidate solution y i is f(x i ).
[0069] Based on the candidate solution population X, a population landscape analysis and operator history characteristic parameters are generated. The distribution characteristics of the current candidate solution population are obtained by calculating the average distance within the candidate solution population, the global optimal individual, and the local optimal individual. Fitness statistics of the current candidate solution population are obtained by evaluating the average difference between the fitness of the candidate solution population and the fitness of the global optimal individual and the local optimal individual, as well as the variance of the current population fitness. The evolvability characteristics of the current candidate population are determined by evaluating the difference in cost changes caused by population evolution and the difference in random steps, the performance difference between the current population and different adopted populations, and the best-to-worst improvement ratio. The timestamp characteristics of the current optimization process are obtained by recording the time to the optimal solution, calculating the remaining time ratio based on the set maximum allowable iteration time, the proportion of time with continuous good performance, and whether the local optimal solution can be updated to the global optimal solution. By analyzing the impact of each operator on the position changes of the best and worst individuals in the population and recording the historical vectors of these changes, the operator history characteristics reflecting the optimization effect and behavioral characteristics of the candidate operators are obtained.
[0070] As an example, assume that the above feature calculation results in 9 candidate solution individuals s i , i = 1, 2, ..., 9, a total of 9 features as the population landscape analysis input of the operator selection model and the operator historical feature matrix m ij , i = 1, 2, ..., 2L, j = 1, 2, ..., D, as shown in the figure, assuming that the operator selection model uses a D × 1 feedforward layer to encode the above operator historical features, and then concatenates the encoding with the population landscape analysis, and assumes that a multi-layer perceptron (MLP) is used to generate the probability of each operator in the operator pool. The expression is as follows:
[0071]
[0072] DV=Tank(W dv (f LA ⊕Flatten(V AH ))+b dv ) (4)
[0073]
[0074] in Different neural network layers for the vector encoding network, and are different neural network layers of the actor network, f AH is the vector of characteristic parameters of the operator's historical performance, f LA(S31-S34) is a vector composed of the distribution characteristics of the current candidate solution population, fitness statistics, feasibility evaluation and time dynamic characteristics of the optimization process, σ is the ReLU activation function, π is the probability of each operator selection, V AH is the embedding vector output by the two-layer feedforward network, b actor 、b ve are the offset vectors of the actor network and the vector encoding network respectively, Flatten(.) is the residual link operation, and DV is the obtained decision vector.
[0075] The population is divided into different subpopulations based on probability, and each individual is assigned to the corresponding optimizer for guidance update according to the selected operator. After all individuals have completed the update, the optimized individuals are re-merged to form a new population, thereby promoting the iterative evolution of the optimization process and enhancing the diversity and adaptability of the search.
[0076] On this basis, the difference between the optimized population and the population before evolution is calculated, and the effect of the optimization process is evaluated based on this, and then the reward value r is calculated. t , in order to guide the optimization strategy towards a better direction.
[0077]
[0078] Where t is the current time step, To optimize the minimum cost to the tth time step, Represents the starting cost (as a normalization factor), Evaluations end Indicates the number of evaluations required to reach the termination condition, and MaxEvaluations is the maximum allowed evaluation test.
[0079] Based on the loss function form determined by formula (2), we assume that the proximal policy optimization (PPO) algorithm is used as the policy gradient optimization method for the operator selection model, and backpropagation is used for training. By continuously optimizing the policy parameters, we ultimately obtain the operator selection model with the best performance, thereby improving the adaptability and search efficiency of the optimization process.
[0080] In summary, the method of the present invention can dynamically select the optimal operator in the black box optimization task to achieve efficient optimization, while having low expert dependence and strong zero-sample generalization capabilities. The characteristics of the present invention are: based on the landscape characteristics of the candidate solution population and the historical characteristics between each operator, the operator selection strategy is dynamically adjusted to adapt to the search needs of different optimization stages. Through the deep reinforcement learning method, the present invention can continuously explore and utilize the optimal operator during the optimization process, thereby improving the search efficiency and the quality of the solution. Compared with the traditional static operator selection method, the present invention can adapt to the characteristics of different problems and realize a more intelligent and efficient optimization strategy. It is suitable for multiple fields such as automatic machine learning, engineering optimization, bioinformatics, etc., and has a wide range of application value.
[0081] The present invention has wide applicability in multiple specific black-box optimization tasks, and in particular, it exhibits powerful performance in practical complex problems such as drone path planning and protein structure splicing. In protein structure splicing optimization, the present invention also shows strong adaptability. Protein splicing design often involves the combination and conformational optimization of multiple structural fragments, with the goal of obtaining a three-dimensional structure with stable functions. This task is also a typical black-box optimization problem. By representing the protein conformation as an optimization population and using the energy scoring function as a fitness indicator, the automatic operator selection model of the present invention can dynamically select appropriate structural perturbations or reconstruction operators according to the characteristics of the current structural distribution, thereby efficiently searching for stable configurations, reducing experimental costs and improving prediction accuracy. In summary, the present invention not only provides an optimization method with efficient search and adaptive scheduling capabilities in theory, but also has good landing value in actual engineering and scientific fields, meeting the requirements of the specificity and applicability of the technical solution.
[0082] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, numerous modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can better understand and utilize the present invention.
Claims
1. An automatic operator selection method based on deep reinforcement learning in evolutionary computing, characterized in that: The following steps are involved: Use the trained automatic operator selection model to select the most appropriate operator for its optimizer to adjust the optimization strategy in the optimizer; The training of the automatic operator selection model includes the following steps: extracting feature parameters, and training the automatic operator selection model based on deep reinforcement learning based on the extracted feature parameters.
2. The method for automatic operator selection based on deep reinforcement learning in evolutionary computing according to claim 1, characterized in that: Extracting feature parameters includes the following steps: S1. Build an operator pool for automatic operator selection by parsing the configuration file. The operator contains the BBO optimizer that can perform the optimization task and the parameters required for the optimizer to run. S2. Obtain the candidate solution population data of the optimizer and preprocess it to ensure that the population data format is consistent; S3. Based on the obtained candidate solution population data, extract feature parameters for training.
3. The method for automatic operator selection based on deep reinforcement learning in evolutionary computing according to claim 2, characterized in that: The candidate population data include population mean, population variance and distance between individuals.
4. The method for automatic operator selection based on deep reinforcement learning in evolutionary computing according to claim 1, characterized in that: The characteristic parameters include the distribution characteristics of the candidate solution population, fitness statistical information and the time dynamic characteristics of the optimization process.
5. The method for automatic operator selection based on deep reinforcement learning in evolutionary computing according to claim 1, characterized in that: The automatic operator selection model is a multi-layer perceptron model or a graph neural network.
6. The method for automatic operator selection based on deep reinforcement learning in evolutionary computing according to claim 1, characterized in that: The loss formula of the automatic operator selection model is: Target(Θ)=E f~D [∑Reward(tra,f)∣Θ]; Where Θ is the selected dynamic operator selector model parameter; D is the distribution of black box problems; f is the problem sampled from the distribution D; tra is the obtained dynamic operator selector operator trajectory, Reward is the reward; E is the expectation of the cumulative reward, and the reward expectation expressed in the above formula is maximized through the backpropagation operator iteration. The maximum expectation Target (Θ * ) corresponds to the automatic operator selection model parameter Θ * , which is the optimal automatic operator selection model.
7. The method for automatic operator selection based on deep reinforcement learning in evolutionary computing according to claim 1, characterized in that: Based on the reward mechanism, the rewards of the automatic operator selection model are calculated and optimized in collaboration with the underlying optimizer. The reward mechanism includes two key indicators: The first is the fitness ratio of the optimal solution of the current candidate population to the optimal solution of the initial population, which is used to measure the overall improvement of the optimization process; The second is the time factor accumulated during the optimization process to evaluate the optimization efficiency, guide the automatic operator selection model to evolve in a more optimal direction, and improve the stability and adaptability of the optimization algorithm.
8. The method for automatic operator selection based on deep reinforcement learning in evolutionary computing according to claim 7, characterized in that: Based on the rewards of the automatic operator selection model and combined with the selected gradient optimization operator, policy gradient calculation is performed to optimize the operator selection strategy, enabling the automatic operator selection model to adaptively learn and adjust during the optimization process to adapt to different optimization scenarios, thereby improving the overall optimization performance.
9. A computer device comprising a memory and a processor, wherein the memory is electrically connected to the processor, and the memory stores a computer program, wherein: When the computer program is executed by the processor, the processor is enabled to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, wherein: When the computer program is executed by a processor, the processor implements the method according to any one of claims 1 to 8.