Multi-target deployment rapid optimization method based on deep reinforcement learning
Deep reinforcement learning with pre-training enhances the efficiency and convergence of multi-objective optimization by training agents to optimize deployment strategies, addressing the inefficiencies of traditional methods in complex constraint environments.
Patent Information
- Application Number
- CN202510461740.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-15
AI Technical Summary
Traditional optimization methods are difficult to find the global optimal solution to the multi-objective optimization problem under complex and changeable constraints, and are inefficient in computing, especially in case of multiple conflicting goals.
A multi-objective deployment rapid optimization method based on deep reinforcement learning is adopted. By simplifying the model and introducing deep reinforcement learning algorithms, training agents, combining pre-trained models and heuristic algorithm optimization, using TD3 algorithm and tanh activation function, a circular deployment area is designed to achieve multi-objective optimization.
The computational efficiency and convergence quality of multi-objective optimization problems have been significantly improved. The agent can quickly find the optimal deployment location, avoid local optimal trapping, and improve learning efficiency and strategy convergence.
Smart Images

Figure CN120317435A_ABST
Abstract
Description
Technical Field
[0001] The present invention is applicable to most engineering and technical fields that need to solve the multi-objective deployment optimization problem under constraints, and particularly relates to a multi-objective deployment rapid optimization method based on deep reinforcement learning. Background Art
[0002] The multi-objective optimization problem under constraints has always been a difficult point in the fields of scientific research and engineering. Traditional optimization methods often struggle to find the global optimal solution and have low computational efficiency when faced with complex and changing constraints. Especially when dealing with multiple conflicting objectives, the limitations of traditional methods are even more obvious. Therefore, there is an urgent need for a new method that can efficiently handle multi-objective optimization problems. As an emerging machine learning technology, deep reinforcement learning has powerful decision-making capabilities and adaptability, providing a new idea for solving this problem. By combining the advantages of deep learning and reinforcement learning, it is possible to achieve rapid optimization and deployment of multiple objectives in a complex environment, significantly improving the overall performance of the system. Summary of the Invention
[0003] The purpose of the present invention is to solve the problem of multi-objective rapid optimization and deployment using deep reinforcement learning under complex and changing constraints, and a multi-objective deployment rapid optimization method based on deep reinforcement learning is proposed.
[0004] The present invention is realized through the following technical solutions. The present invention proposes a multi-objective deployment rapid optimization method based on deep reinforcement learning, and the method includes the following steps:
[0005] Step 1, model simplification: Simplify and abstract the model of the multi-objective optimization and deployment target, and construct a simplified version of the model by extracting key features and ignoring secondary factors;
[0006] Step 2, framework construction of the multi-objective optimization problem under constraints: Assuming that all objectives are expected to be minimized, the discrete or continuous multi-objective optimization problem is described as:
[0007]
[0008] where x is an n-dimensional design variable, including discrete and continuous variables; x ∈ R n is the self-defined domain of the design variable, also known as the decision space; F(x) is an m-dimensional objective vector, constituting the problem domain mapped from the decision space; the mapping function f i (x) describes the real requirements and problem characteristics, and is called the objective function; g i (x) is the inequality constraint, p is the number of inequalities; h j (x) is the equality constraint, q is the number of equalities;
[0009] Step 3: Construction of a multi-objective optimization model based on a deep reinforcement learning algorithm: Use the deep reinforcement learning algorithm to train an agent to control all optimization deployment objectives, thereby finding the optimal deployment location;
[0010] Step 4: Construction of a multi-objective optimization model with pre-training introduced: Introduce a pre-training mode into the constructed multi-objective optimization model to form an agent pre-training model, and train the agent pre-training model, thereby obtaining the final multi-objective optimization deployment result according to the trained model.
[0011] Furthermore, in Step 2, for the multi-objective optimization problem, assuming the objective vectors u and v generated by two solutions, define that u dominates v as:
[0012]
[0013] If and only if:
[0014]
[0015] Given the solution set S, when any solution s ∈ S cannot be dominated by other solutions in the solution set, then the solution set S is called Pareto non-dominated, also known as the Pareto front PF; all non-dominated solutions in the decision space constitute the ideal Pareto front PF * ; To ensure that a reward function can be provided for the agent during subsequent training, different optimization terms are weighted and summed to obtain a unified optimization function:
[0016] F(x) = α1f1(x) + α2f2(x) +... + α m f m (x) (4)
[0017] where α1 / α2 / α m are all weight coefficients. By reasonably designing the weight coefficients, a balance can be achieved between different optimization objectives, thereby more effectively approaching the ideal Pareto front.
[0018] Furthermore, in Step 3, the deployment locations of all deployment objectives are the states where the current agent is located:
[0019] S = [x1, y1, x2, y2,... x 10 , y 10 , x 11 , y 11 (5)
[0020] Take the change in the deployment locations of all deployment objectives in each iteration as the action output by the agent, and the displacement is represented by v:
[0021]
[0022] After the agent takes an action, the environment returns the reward value of the current step and the state of the next step to the agent:
[0023] S t+1 = S t + a t (7)
[0024] The reward value is set to the value of the optimized objective function:
[0025] f(s) = f reward (s) - f punishment (s) (8)
[0026] where f reward is the sum of the weights of all optimization terms, and f punishment is the sum of the penalty function penalty values caused by not meeting the constraint conditions.
[0027] Furthermore, the deep reinforcement learning algorithm is the TD3 algorithm. Using the TD3 algorithm to train the agent to find the optimal deployment location is specifically as follows:
[0028] 1) The agent's Actor network selects the current action according to the current state S t :
[0029] a t = μ(S t ) (9)
[0030] 2) According to S t and a t , the deployment target changes the deployment location in the deployment environment to obtain the deployment location S t+1 at the next moment and the corresponding objective function value f t i.e., the reward value;
[0031] 3) Save the current state, action, reward value, and next state in the experience pool in the form of a tuple (S t , a t , f t , S t+1 );
[0032] 4) After the data in the experience pool accumulates to a certain amount, take out a certain number of tuples in batches to train the agent;
[0033] 5) During training, the agent's Target_Actor network receives S t+1 to obtain the target action of the next state:
[0034] a' t+1 = μ'(S t+1 ) (10)
[0035] 6) The Target_Critic network receives μ θ- (s i+1 ) = a' t+1 + N, the reward value f t and the next state deployment target deployment location S t+1 Estimate the update target of the Critic network. At this time, the Target_Critic network will output two valuations and select the smaller one:
[0036] y t = f t + γQ w- (S t+1 , μ θ- (S t+1 )) (11)
[0037] where N is Gaussian noise and γ is the reward decay coefficient;
[0038] 7) Obtain the loss function of the Critic network based on the update target valuation of the Target_Critic network and the valuation of the current action value function Q by the Critic network:
[0039]
[0040] 8) The Critic network updates its own parameters through gradient backpropagation according to the loss function L;
[0041] 9) The Actor network obtains the parameter update gradient based on the action value function Q provided by the Critic network:
[0042]
[0043] It represents that the goal of the Actor network is to maximize the function That is, to maximize the average initial state S0 action value function; where Q(s,a) is the valuation of the current action value function by the Critic network, and the Critic network will also estimate two Q values, and here the smaller one is selected;
[0044] 10) Use the Actor network and the Critic network to perform soft parameter updates on the Target_Actor and Target_Critic networks;
[0045] 11) Repeat steps 5) to 10) until the agent reward curve converges.
[0046] Further, in step four, first use the heuristic optimization algorithm to optimize the objective function, and use the optimized result as the target for the initial pre-training of the agent; set two experience pools, the first for the agent to store data for pre-training, and the other to record data that can be used for later self-exploration training during pre-training.
[0047] Further, the specific steps of step four are as follows:
[0048] 1) Use the heuristic algorithm to optimize the objective function to obtain the pre-training target S target ;
[0049] 2) Assume that the current state of the agent is S t , then the reward value is the distance of the current state relative to the pre-training target in the solution space:
[0050] r t =-|S t -S target | (14)
[0051] 3) During the exploration process of the agent, collect the pre-training reward value r t and the corresponding objective function value f t , and store the tuple (s t , a t , r t , s t+1 ) into the pre-training experience pool for the pre-training process of the agent, and store (s t , a t , f t , s t+1 ) into the self-exploration experience pool for the self-exploration training process after pre-training;
[0052] 4) Take out data from the pre-training experience pool in batches to pre-train the agent;
[0053] 5) Alternate network training and environment interaction. During environment interaction, continuously collect data tuples (s t , a t , f t , s t+1 ) for self-exploration training, and continuously update the experience pool;
[0054] 6) Pre-train until the reward value converges, save the agent model and the experience pool, return to the self-exploration process to start training, and finally obtain the optimized deployment result.
[0055] Further, use the tanh activation function instead of the Relu activation function in the multi-objective optimization process.
[0056] Further, in the multi-objective optimization process, borderless processing of the deployment area is carried out, specifically: a circular deployment area is designed, and when the deployment target of the agent exceeds the border, the target returns from the other side.
[0057] The present invention also provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the multi-objective deployment fast optimization method based on deep reinforcement learning are implemented.
[0058] The present invention also provides a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, the steps of the multi-objective deployment fast optimization method based on deep reinforcement learning are implemented.
[0059] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0060] The present invention proposes a multi-objective deployment fast optimization method based on deep reinforcement learning. The method first gives a model of the multi-objective optimization problem and transforms the problem into a single-objective optimization problem with a penalty term. Aiming at the problem that traditional optimization algorithms are prone to fall into local optima when dealing with multi-dimensional and high-complexity problems, the present invention innovatively proposes an idea of using the powerful environment understanding ability and generalization ability of the deep reinforcement learning algorithm to deal with complex optimization problems. The model constructed based on the TD3 algorithm enables the agent to learn and optimize the decision-making strategy by interacting with the environment. Preliminary experiments show that although the model can better understand the environment, the early learning efficiency is low and the strategy convergence is slow. Therefore, the present invention further proposes a reinforcement learning model based on pre-training, provides pre-training objectives through traditional optimization algorithms, accelerates the learning process and improves the convergence quality. Experiments prove that the pre-trained agent is superior to the non-pre-trained agent in both convergence speed and quality. Description of the Drawings
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0062] Figure 1 It is a flowchart of the multi-objective deployment fast optimization method based on deep reinforcement learning according to the present invention.
[0063] Figure 2 It is a schematic diagram of a multi-objective optimization model based on the TD3 algorithm.
[0064] Figure 3 It is a schematic diagram of the pre-training model of the multi-objective optimization agent.
[0065] Figure 4 Schematic diagram of the general situation of the covered target and the covered area.
[0066] Figure 5 Schematic diagram of the convergence curve of the radar target function.
[0067] Figure 6 Schematic diagram of the convergence curve of the reward value of the air defense weapon target function.
[0068] Figure 7 Schematic diagram of the convergence curve of the air defense weapon target function.
[0069] Figure 8 Schematic diagram of the TD3 deployment result.
[0070] Figure 9 Schematic diagram of the convergence curve of the pre-training reward value.
[0071] Figure 10 Schematic diagram of the convergence curve of the pre-training agent target function.
[0072] Figure 11 Schematic diagram of the deployment result of the pre-trained TD3 agent.
[0073] Figure 12 Schematic diagram of the convergence of each optimization item, where (a) coverage area, (b) overlapping area, (c) fire depth, (d) projection length of the key covered boundary, (e) uncovered area of the covered area, (f) number of covered layers that do not meet the standard for the covered target. Specific implementation manners
[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0075] In combination with Figures 1-12 , the present invention proposes a multi-objective deployment fast optimization method based on deep reinforcement learning, and the method includes the following steps:
[0076] Step 1. Model simplification: Simplify and abstract the model of the multi-objective optimization deployment target, and construct a simplified version of the model by extracting key features and ignoring secondary factors;
[0077] Generally, the models for multi-objective optimization deployment targets are relatively complex. For example, the killing area of air defense missiles, the route coverage rate, the geocentric coverage angle of the effective field of view of early warning satellite constellations, etc. This is not conducive to direct efficient calculation and optimization. To improve the calculation efficiency and simplify the processing process, the present invention first reasonably simplifies and abstracts the original model. By extracting key features and ignoring secondary factors, a simplified model is constructed, which can not only retain the core characteristics of the original model but also significantly reduce the computational complexity. This step ensures a significant increase in the calculation speed while maintaining the accuracy.
[0078] Step 2: Establish the framework of the multi-objective optimization problem under constraints.
[0079] Generally, in real engineering, the multi-objective optimization problem with constraints has various types of design variables and multiple conflicting objectives, and has characteristics such as high dimension, multi-scale, non-linearity, and non-convexity, belonging to typical NP-hard problems. For such multi-objective optimization problems, without loss of generality, assuming that all objectives are expected to be minimized, the discrete or continuous multi-objective optimization problem can be described as:
[0080]
[0081] where x is an n-dimensional design variable, including discrete and continuous variables; x ∈ R n is the self-defined domain of the design variable, also known as the decision space; F(x) is an m-dimensional objective vector, constituting the problem domain mapped from the decision space; the mapping function f i (x) describes the real requirements and problem characteristics and is called the objective function; g i (x) is the inequality constraint, and p is the number of inequalities; h j (x) is the equality constraint, and q is the number of equalities.
[0082] As can be seen from the above formula, establishing the framework of the multi-objective optimization problem mainly includes two parts of work: constructing the objective function and the constraint conditions. For multi-objective optimization problems, the superiority and inferiority of two solutions are mainly compared through the Pareto dominance relationship. Assuming that the objective vectors u and v generated by two solutions, define u dominates v as:
[0083]
[0084] If and only if:
[0085]
[0086] Given the solution set S, when any solution s ∈ S cannot be dominated by other solutions in the solution set, then the solution set S is called Pareto non-dominated, also known as the Pareto front PF; all non-dominated solutions in the decision space constitute the ideal Pareto front PF *; To ensure that a reward function can be provided to the agent during subsequent training, different optimization terms are weighted and summed to obtain a unified optimization function:
[0087] F(x) = α1f1(x) + α2f2(x) +... + α m f m (x) (4)
[0088] where α1 / α2 / α m are all weight coefficients. By reasonably designing the weight coefficients, a balance can be achieved among different optimization objectives, thereby more effectively approaching the ideal Pareto front. Specifically, the selection of weights should be based on the requirements of the actual problem and the importance of each optimization objective. For example, in a resource-constrained environment, resource utilization may require a higher weight, while in a performance-sensitive application, response time or processing speed may be the priority factors. By dynamically adjusting these weight coefficients, the model can exhibit stronger adaptability and flexibility in different optimization scenarios.
[0089] Step 3: Building a multi-objective optimization model based on the deep reinforcement learning algorithm: Use the deep reinforcement learning algorithm to train the agent so that the agent controls all optimization deployment objectives, thereby finding the optimal deployment location;
[0090] There are many types of deep reinforcement learning algorithms, and a suitable algorithm can be selected according to different specific application scenarios and optimization objectives. For example, for a dynamic environment that requires quick response, algorithms based on policy gradients such as PPO (Proximal Policy Optimization) or A3C (Asynchronous Advantage Actor-Critic) can be chosen; while for a static environment that requires high-precision solutions, algorithms based on value functions such as DQN (Deep Q-Network) or its variants can be considered. In addition, the advantages of multiple algorithms can be combined to design a hybrid deep reinforcement learning model to better adapt to complex and changing optimization problems. When selecting an algorithm, factors such as computing resources, training time, and model interpretability also need to be considered to ensure the efficiency of the optimization process and the reliability of the results.
[0091] In the present invention, the TD3 (Twin Delayed Deep Deterministic policy gradient algorithm) algorithm is taken as an example to train the agent so that the agent controls all optimization deployment objectives, and the deployment locations of all deployment objectives are the states where the current agent is located:
[0092] S = [x1, y1, x2, y2,... x 10 , y 10 , x11 , y 11 (5)
[0093] Take the change in the deployment positions of all deployment targets in each iteration as the action output by the agent. The displacement is represented by v:
[0094]
[0095] After the agent takes an action, the environment returns the reward value for the current step and the state of the next step to the agent:
[0096] S t+1 = S t + a t (7)
[0097] The reward value is set to the value of the optimization objective function:
[0098] f(s) = f reward (s) - f punishment (s) (8)
[0099] Among them, f reward is the weighted sum of all optimization terms, and f punishment is the sum of the penalty function penalty values due to non - satisfaction of the constraint conditions.
[0100] Since the number of steps in each round designed is fixed, the agent will automatically tend to find a strategy that maximizes the current state value. And the value of the state is the accumulation of all objective function values from the current state to the termination state. Therefore, the agent will tend to automatically and quickly find the position corresponding to a better objective function value to ensure the maximization of the state value, thus finding the optimal deployment position.
[0101] The deep reinforcement learning algorithm is the TD3 algorithm. Use the TD3 algorithm to train the agent to find the optimal deployment position. Specifically:
[0102] 1) The agent's Actor network selects the current action according to the current state S t :
[0103] a t = μ(S t ) (9)
[0104] 2) According to S t and a t , the deployment target changes its deployment position in the deployment environment, obtaining the deployment position (agent state) S t+1 at the next moment and the corresponding objective function value f t which is the reward value;
[0105] 3) Save the current state, action, reward value, and next state as a tuple (S t , a t , f t , S t+1 ) into the experience pool;
[0106] 4) After the data in the experience pool accumulates to a certain amount, take out a certain number of tuples in batches to train the agent;
[0107] 5) During training, the agent's Target_Actor network receives S t+1 and obtains the target action for the next state:
[0108] a' t+1 = μ'(S t+1 ) (10)
[0109] 6) The Target_Critic network receives μ θ- (s i+1 ) = a' t+1 + N, the reward value f t and the target deployment location S for the next state t+1 to estimate the update target of the Critic network. At this time, the Target_Critic network will output two estimates, and select the smaller value:
[0110] y t = f t + γQ w- (S t+1 , μ θ- (S t+1 )) (11)
[0111] where N is Gaussian noise and γ is the reward decay coefficient;
[0112] 7) Obtain the loss function of the Critic network according to the update target estimate of the Target_Critic network and the estimate of the current action value function Q of the Critic network:
[0113]
[0114] 8) The Critic network updates its own parameters through gradient backpropagation according to the loss function L;
[0115] 9) The Actor network obtains the parameter update gradient according to the action value function Q provided by the Critic network:
[0116]
[0117] It represents that the goal of the Actor network is to maximize the function That is, to maximize the average initial state S0 action value function; where Q(s,a) is the estimate of the current action value function by the Critic network. The Critic network will also estimate two Q values, and here the smaller one is selected.
[0118] 10) Use the Actor network and the Critic network to perform a soft update of the parameters of the Target_Actor and Target_Critic networks.
[0119] 11) Repeat steps 5) to 10) until the agent reward curve converges.
[0120] Step Four: Introduce a pre-trained multi-objective optimization model construction: Introduce a pre-training mode into the constructed multi-objective optimization model to form an agent pre-training model, and train the agent pre-training model, so as to obtain the final multi-objective optimization deployment result according to the trained model.
[0121] Since the state and action space dimensions involved in the multi-objective optimization problem with constraints in general engineering are relatively high and the search space is huge, which is not conducive to the rapid learning of the agent. Also, because the agent in step three does not have a good guidance in the initial exploration stage and the exploration is completely random, the quality of the experience collected in its initial stage is relatively low and also quite divergent. Therefore, the agent learns slowly.
[0122] To improve the learning speed of the agent and ensure that high-quality experience can be collected in the experience pool, a multi-objective optimization model introducing a pre-training mode is designed. Its general idea is: First, use a heuristic optimization algorithm (here, the double-population particle swarm algorithm) to optimize the objective function, and use the optimized result as the target for the initial pre-training of the agent; set up two experience pools, the first is for the agent to store data for pre-training, and the other records data that can be used for later autonomous exploration training during pre-training.
[0123] The specific steps of step four are as follows:
[0124] 1) Use a heuristic algorithm to optimize the objective function to obtain the pre-training target S target ;
[0125] 2) Assume that the current state of the agent is S t , then the reward value is the distance of the current state relative to the pre-training target in the solution space:
[0126] r t =-|S t -S target | (14)
[0127] 3) During the agent exploration process, collect the pre-training reward value r of the current state t and the corresponding objective function value f t , and store the tuple (s t , a t , r t , s t+1 ) into the pre-training experience pool for the agent pre-training process, and store (s t , a t , f t , s t+1 ) into the self-exploration experience pool for the self-exploration training process after the pre-training ends;
[0128] 4) Take out data from the pre-training experience pool in batches to pre-train the agent;
[0129] 5) Alternately perform network training and environment interaction. When interacting with the environment, continuously collect data tuples (s t , a t , f t , s t+1 ) for self-exploration training, and continuously update the experience pool;
[0130] 6) Pre-train until the reward value converges, save the agent model and the experience pool, return to the self-exploration process to start training, and finally obtain the optimized deployment result.
[0131] To ensure that the agent can effectively learn the multi-objective optimization strategy and avoid falling into ineffective learning and slow learning, the present invention adopts some effective design techniques in the process of designing the interaction between the agent and the environment. These techniques include the design of the agent network, the design of the interaction environment, and data processing, etc., which are helpful for the agent to learn quickly and effectively.
[0132] (1) Use the tanh activation function instead of the Relu activation function in the multi-objective optimization process.
[0133] In the design of the agent network, usually, the ReLU is selected as the activation function. However, the ReLU activation function has some inherent disadvantages. For example, there is a problem of gradient disappearance when the input is negative. At the same time, due to its asymmetry, the output data is not centered at 0, resulting in a certain bias in the data passed to the next layer, which is not conducive to the gradient descent of the neural network. The Tanh function can well avoid these problems. When the input quantity is not very far from the 0 point, the gradient of the tanh function will not disappear, and at the same time its output is symmetric with respect to the origin, which is beneficial to the training of the neural network.
[0134] (2) Normalize the input data
[0135] Although the Tanh function has the advantages of being able to handle negative values and having an unbiased output, it still has limitations on the input data. When the input data deviates too far from the 0 point, the gradient of the tanh function for it will approach 0, which is similar to the negative part of the ReLU function; this means that if the input quantity is too large or too small, the gradient of the tanh function for it will be too small, and the parameter update step size will be extremely small, which is not conducive to the learning of the neural network. Therefore, the present invention introduces a normalization layer into the neural network of the intelligent agent and performs normalization using the min-max normalization method:
[0136]
[0137] Normalize the input data so that the numerical ranges of different features are unified, thereby avoiding the excessive influence of certain features on the results during the training process of the machine learning model. Improve the stability and accuracy of the model, shorten the training time, and thus lay a solid foundation for subsequent model training and prediction.
[0138] (3) Deploy boundary-free processing during the multi-objective optimization process. At the initial stage of model design, if the actions of the intelligent agent cause the deployment target to exceed the range, the deployment position will be truncated at the boundary. However, this boundary restriction training may cause the intelligent agent to get stuck at the boundary and unable to move, collecting a large amount of invalid data. For this reason, design a circular deployment area: when the deployment target of the intelligent agent exceeds the boundary, the target returns from the other side. This design avoids the intelligent agent being restricted by the boundary, increases the opportunity to explore different states, helps the intelligent agent to globally grasp the environment, enriches experience, and promotes learning.
[0139] Embodiment
[0140] Taking the optimal deployment of the air defense system during the air defense operation as an example, use the method proposed by the present invention to train the intelligent agent to complete the multi-objective optimal deployment.
[0141] In this test, the area to be covered is set as a square area with a side length of 150 km. Taking the southwest corner as the origin, the north-south boundary as the y-axis, and the east-west boundary as the x-axis, a plane rectangular coordinate system is established; there are three targets that need to be key covered in this covered area, represented in the form of points, and the corresponding coordinates are (30, 60), (60, 120), and (130, 100); the key covered boundary of this area is its northern boundary. The terrain is set as plain. The specific form is as Figure 4 shown. Among them, the blue area is the covered area, the red dots are the covered targets, the black arrows represent the direction of the enemy's attack, and the boundary of the covered area facing the direction of the enemy's attack is the key covered boundary. The gray dotted line is the safety distance line (the enemy's bombing line). In this test scenario, our side has 3 types of air defense weapons and an air defense radar, and the specific combat parameters are shown in Tables 1 and 2:
[0142] Table 1 Anti-aircraft Missile Operation Data
[0143]
[0144] Table 2 Anti-aircraft Radar Operation Data
[0145]
[0146] First, perform the model simplification in Step 1: The kill zone of a typical multi-target channel anti-aircraft missile is relatively complex and is simplified to an annular area. Among them, the radius R of the large circle is the average maximum radius of the fire coverage area of the anti-aircraft missile, which can be represented by the attack radius of the anti-aircraft missile; is the average value of the horizontal projection of the near-range slant range and is also the inner ring radius within the fire coverage area. It is represented by r according to the relevant parameters of the vertical kill zone of the anti-aircraft weapon and the elevation angle of the kill zone:
[0147]
[0148] Among them, H max is the maximum height of the kill zone, H jj is the near-boundary height of the kill zone, and ε max is the maximum elevation angle of the kill zone.
[0149] Secondly, build the optimization objective function and constraint conditions in Step 2.
[0150] (I) Objective Function
[0151] The objective function mainly consists of two parts, the optimization item corresponding to the anti-aircraft missile and the optimization item corresponding to the anti-aircraft radar. The two parts are relatively independent, and each optimization item will be briefly described below.
[0152] (1) Fire Coverage Area
[0153] The fire coverage area refers to the scope of action of our anti-aircraft weapons in the deployment space, not limited to the cover area. When resources are limited, there may be a contradiction between fire coverage and overlapping area. The specific form is as follows:
[0154]
[0155] In the formula, S i is the fire coverage area of the i-th anti-aircraft weapon, and Z is the total number of anti-aircraft weapons.
[0156] (2) Fire Overlapping Area
[0157] The overlapping fire area is a key indicator for optimizing the deployment of the air defense system. It shows the likelihood of enemy units being jointly attacked by our air defense firepower. The more layers of overlapping firepower, the greater the density of air defense firepower at a certain point for the enemy unit. However, there is a conflict between the overlapping fire area and the fire coverage area. The specific form is as follows:
[0158]
[0159] In the formula, S Cji is the overlapping area between the i-th air defense weapon and the j-th air defense weapon.
[0160] (3) The sum of the fire depths of long-range air defense weapons
[0161] The fire depth is divided into two categories: one is for covering targets, starting from the target and extending towards the enemy direction, and the length covered by long-range air defense firepower beyond the safe distance; the other is when the covered area is on the boundary, starting from the boundary and extending towards the enemy, and the average covered length of long-range air defense firepower when exceeding the safe distance. The final function value is the sum of these two parts.
[0162] (4) The effective covering width of medium and short-range air defense missiles on the key covering boundary
[0163] The key covering boundary faces the enemy's incoming direction. The effective covering width of medium and short-range air defense weapons on the boundary shows their protection level against enemy air raids. The larger the covering width, the more enemy weapons the position of the air defense weapon can target. The specific form is as follows:
[0164]
[0165] (5) The detection coverage area of the early warning radar
[0166] The coverage area of the early warning radar represents the detection ability of the air defense system and is an important indicator for evaluating the early warning performance of the air defense system; the larger the detection coverage area of the early warning radar, the stronger the perception ability of the air defense system for the surrounding area, and the more conducive it is for the air defense system to react in advance to potential threats. Its specific form is similar to the fire coverage area:
[0167]
[0168] In the formula, D i is the detection area of the i-th air defense radar, and N is the total number of air defense weapons.
[0169] (6) The prediction accuracy of the landing point of the early warning radar
[0170] After the enemy missile shuts down, the early warning radar will perform multiple positioning and sampling on the missile to predict the impact point. The higher the positioning accuracy, the more accurate the sampling, the more accurate the prediction of the missile trajectory, and the higher the impact point prediction accuracy. Generally, the Geometric Dilution of Precision (GDOP) is used to evaluate the positioning accuracy of the radar.
[0171]
[0172] Among them are the root mean square errors of the positioning errors of the radar in the x, y, and z directions respectively.
[0173] Since the position of the enemy missile is unknown during deployment and it is considered that its target of attack is the cover target, the cover target point is used as the evaluation point. Let δ x = k x d r , δ y = k y d r , δ y = k y d r , where d r is the distance from the early warning radar to the cover target, and k x , k y , k z are the corresponding positioning parameters of the early warning radar. Then the impact point prediction accuracy becomes:
[0174]
[0175] (7) Probability of the early warning radar detecting the target
[0176] The probability of the radar detecting the target can be expressed as:
[0177]
[0178] In the formula, P is the probability of the radar detecting the target, σ is the effective radar cross section of the target, δ Z is the root mean square value of the noise amplitude, P T is the engine power, and d is the target distance.
[0179] Since the target is in motion and the distance d is not convenient for evaluation, the distance r E between the radar and the key cover boundary is selected to replace it, and the expression becomes:
[0180]
[0181] Among them, the objective functions of air defense early warning radars are all related to the distance between the early warning radar and the key protection boundary, as well as the distance between the early warning radar and the protected target, and these two are more easily expressed by independent variables (i.e., the radar position coordinates). Therefore, the last two objective functions of the early warning radar are replaced with:
[0182]
[0183] where N is the number of early warning radars, M is the number of protected targets, r E is the distance from the early warning radar to the key protection boundary, l ij is the distance from the i-th radar to the j-th target. minL is the minimum value of the denominator.
[0184] (2) Constraint conditions
[0185] There are three constraint conditions, which are respectively:
[0186] (1) Terrain constraint
[0187] Since the simulated terrain is a plain, it is considered that there is no terrain constraint.
[0188] (2) Coverage of the protected area by medium and long-range air defense weapons
[0189] This constraint condition requires that the medium and long-range air defense firepower should achieve full coverage of the protected area to ensure that there is no firepower vacuum inside the protected area.
[0190] (3) Coverage layer of medium and short-range air defense firepower for protected targets
[0191] This constraint condition requires that the medium and short-range air defense firepower must cover the protected targets, and each target must be covered by two or more medium and short-range air defense weapons simultaneously.
[0192] Since the processing of the optimization problem with constraint conditions is relatively cumbersome, the penalty function is used to introduce it into the objective function, and at the same time, each target is weighted, and this problem can be transformed into an unconstrained single-objective optimization problem. By adding a certain penalty to the solutions that do not meet the constraint conditions, the optimization algorithm or agent can be automatically inclined to find the solutions that meet the constraint conditions. Therefore, the problem can be expressed as two single-objective function optimization problems, and the optimization objective functions are respectively:
[0193] (1) The sum of the fire coverage area + the fire overlap area + the fire depth of the long-range air defense weapons + the effective cover width of the medium and short-range air defense missiles on the key protection boundary + the coverage of the protected area by the medium and long-range air defense weapons (penalty function) + the coverage layer of the medium and short-range air defense firepower for protected targets (penalty function).
[0194] (2) The detection coverage area of the early warning radar + the objective function f2.
[0195] Optimize and solve the two objective functions separately, and the final optimization result can be obtained by integrating the optimization results.
[0196] Then, perform the optimization of the agent's air defense deployment without pre-training in Step 3 to obtain the reward value convergence curve and the objective function value convergence curve. The radar objective function is relatively simple, so it converges faster. The air defense weapon objective function is more complex because it has more optimization terms and includes a penalty function part for restriction. The theoretically designed maximum value is 8, and Figure 7 the final convergence result in Figure 6 is 3.237, indicating that the deployment result is not yet perfect. Observing Figure 7 the reward value convergence curve, it can also be seen that
[0197] the result of Figure 10 is accidental because the reward value of the agent has not converged, which means that the agent has not learned a method to quickly deploy the air defense weapon to the optimal position. Therefore, for the air defense weapon deployment problem, the pre-training model in Step 4 is further introduced.
[0197] Comparing Figure 10 and Figure 7 it can be seen that the pre-training process greatly improves the convergence speed and convergence quality of the agent's strategy. At the beginning of exploration, due to the pre-training process, the agent has a relatively clear strategy orientation in the initial stage. Therefore, the return value of each round in the initial stage of exploration is significantly higher than that of the untrained agent. From Figure 10 it is known that the final convergence result is 5.306, which is better than the optimization result of the TD3 agent without pre-training.
[0198] Show the comparison chart of the convergence results of each optimization term before and after pre-training, as Figure 12 shown. It is easy to see from Figure 12 that the convergence results of the pre-trained agent are better than those of the untrained agent in all optimization terms except for the overlapping area. Analyzing the convergence results, the pre-trained agent converges faster in each optimization term and is more stable, indicating that when the agent has a good strategy orientation and experience for learning in the initial stage of training, it can find the optimal deployment position faster and has a deeper understanding of the environment during the learning process, avoiding a large number of ineffective explorations, thus improving the learning efficiency. In addition, the pre-trained agent performs slightly worse in the overlapping area term because only the ability of the agent to quickly enter the target deployment position is considered in the pre-training process, and the overlapping area is not restricted. Therefore, the convergence result in this term is slightly worse, but considering the significant improvement in other term indicators, this result is still acceptable.
[0199] The present invention also provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the multi-objective deployment fast optimization method based on deep reinforcement learning are implemented.
[0200] The present invention also provides a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, the steps of the multi-objective deployment fast optimization method based on deep reinforcement learning are implemented.
[0201] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). It should be noted that the memory of the method described in the present invention is intended to include but not limited to these and any other suitable types of memory.
[0202] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that contains one or more integrated available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as high-density digital video discs (DVDs)), or semiconductor media (such as solid state discs (SSDs)), etc.
[0203] In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by a combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0204] It should be noted that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0205] The above has introduced in detail the multi-objective deployment rapid optimization method based on deep reinforcement learning proposed by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A fast optimization method for multi-objective deployment based on deep reinforcement learning, characterized in that, The method includes the following steps: Step 1, model simplification: Simplify and abstract the model of the multi-objective optimization deployment target, and construct a simplified version of the model by extracting key features and ignoring secondary factors; Step 2, framework construction of the multi-objective optimization problem under constraints: Assuming that all objectives are expected to be minimized, the discrete or continuous multi-objective optimization problem is described as: where \(x\) is an \(n\)-dimensional design variable, including discrete and continuous variables; \(x\in R\) n is the self-defined domain of the design variable, also known as the decision space; \(F(x)\) is an \(m\)-dimensional objective vector, which constitutes the problem domain mapped from the decision space; the mapping function \(f\) i (x) describes the real-world requirements and problem characteristics and is called the objective function; \(g\) i (x) is the inequality constraint, and \(p\) is the number of inequalities; \(h\) j (x) is the equality constraint, and \(q\) is the number of equalities; Step 3, construction of the multi-objective optimization model based on the deep reinforcement learning algorithm: Use the deep reinforcement learning algorithm to train the agent to control all optimization deployment targets, so as to find the optimal deployment location; Step 4, construction of the pre-trained multi-objective optimization model: Introduce the pre-training mode into the constructed multi-objective optimization model to form the agent pre-training model, and train the agent pre-training model, so as to obtain the final multi-objective optimization deployment result according to the trained model.
2. The method according to claim 1, wherein In Step 2, for the multi-objective optimization problem, assuming the objective vectors u and v generated by two solutions, u is defined to dominate v as: If and only if: Given a solution set S, when any solution s ∈ S cannot be dominated by other solutions in the solution set, the solution set S is called Pareto non-dominated, also known as the Pareto front PF; all non-dominated solutions in the decision space constitute the ideal Pareto front PF * ; To ensure that a reward function can be provided for the agent during subsequent training of the agent, different optimization terms are weighted and summed to obtain a unified optimization function: F(x) = α1f1(x) + α2f2(x) +... + α m f m (x) (4) where α1 / α2 / α m are all weight coefficients. By reasonably designing the weight coefficients, a balance can be achieved among different optimization objectives, thereby more effectively approaching the ideal Pareto front.
3. The method according to claim 2, wherein In Step 3, the deployment locations of all deployment targets are the states where the current agent is located: S = [x1, y1, x2, y2,... x 10 , y 10 , x 11 , y 11 (5) Take the change in the deployment locations of all deployment targets in each iteration as the action output by the agent, and the displacement is represented by v: After the agent makes an action, the environment will return the reward value of the current step and the state of the next step to the agent: S t+1 = S t + a t (7) The reward value is set as the value of the optimization objective function: f(s) = f reward (s) - f punishment (s) (8) where, f reward is the sum of the weights of all optimization items, and f punishment is the sum of the penalty function penalty values caused by not satisfying the constraint conditions.
4. The method according to claim 3, wherein The deep reinforcement learning algorithm is the TD3 algorithm. Use the TD3 algorithm to train the agent to find the optimal deployment location. Specifically: 1) The Agent Actor network selects the current action based on the current state S t Select the current action: a t = μ(S t ) (9) 2) According to S t and a t , the deployment target changes its deployment location in the deployment environment to obtain the deployment location S t+1 at the next moment, as well as the corresponding objective function value f t i.e., the reward value; 3) Save the current state, action, reward value, and next state in the experience pool in the form of a tuple (S t , a t , f t , S t+1 ); 4) After the data in the experience pool accumulates to a certain amount, take out a certain number of tuples in batches to train the agent; 5) During training, the agent's Target_Actor network receives S t+1 , and obtains the next-state target action: a' t+1 = μ'(S t+1 ) (10) 6) The Target_Critic network receives μ θ- (s i+1 ) = a' t+1 + N, the reward value f t and the next state deployment target deployment location S t+1 Estimate the update target of the Critic network. At this time, the Target_Critic network will output two estimates and select the smaller value: y t = f t + γQ w- (S t+1 , μ θ- (S t+1 )) (11) Among them, N is Gaussian noise, and γ is the reward decay coefficient; 7) Obtain the loss function of the Critic network according to the updated target valuation of the Target_Critic network and the valuation of the current action value function Q of the Critic network: 8) The Critic network updates its own parameters through gradient backpropagation according to the loss function L; 9) The Actor network obtains the parameter update gradient according to the action value function Q provided by the Critic network: The goal of its representative Actor network is to maximize the function That is, to maximize the average initial state S0 action value function; where Q(s,a) is the estimate of the current action value function by the Critic network, and the Critic network will also estimate two Q values, and the smaller one is selected here; 10) Use the Actor network and the Critic network to perform soft parameter updates on the Target_Actor and Target_Critic networks; 11) Repeat processes 5) to 10) until the agent reward curve converges.
5. The method according to claim 4, wherein In Step 4, first use the heuristic optimization algorithm to optimize the objective function, and use the optimized result as the target for the initial pre-training of the agent; set two experience pools, the first is for the agent to store data for pre-training, and the other records data that can be used for later self-exploration training during pre-training.
6. The method according to claim 5, characterized in that, The specific content of Step 4 is: 1) Optimize the objective function using a heuristic algorithm to obtain the pre-training objective S target ; 2) Set the current state of the agent as S t , then the reward value is the distance of the current state relative to the pre-trained target in the solution space: r t = -|S t -S target | (14) 3) During the agent exploration process, collect the pre-training reward value r of the current state t and the corresponding objective function value f t , and store the tuple (s t , a t , r t , s t+1 ) into the pre-training experience pool for the agent's pre-training process, and store (s t , a t , f t , s t+1 ) into the self-exploration experience pool for the self-exploration training process after the pre-training ends; 4) Take out data from the pre-training experience pool in batches to pre-train the agent; 5) Network training and environment interaction are carried out alternately. During environment interaction, data tuples (s t , a t , f t , s t+1 ) for self-exploration training are continuously collected, and the experience pool is continuously updated; 6) Pre-train until the reward value converges, save the agent model and the experience pool, return to the self-exploration process to start training, and finally obtain the optimized deployment result.
7. The method according to claim 1, characterized in that Use the tanh activation function instead of the Relu activation function in the multi-objective optimization process.
8. The method according to claim 1, wherein During the multi-objective optimization process, borderless processing of the deployment area is carried out, specifically: a circular deployment area is designed, and when the deployment target of the agent exceeds the boundary, the target returns from the other side.
9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1-8 are implemented.
10. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, the steps of the method according to any one of claims 1-8 are implemented.
Citation Information
Cited By
MO-KTO legal model enhancement method, device and equipment and storage medium
CN121660020A