Ocean unmanned cluster hunting maneuvering decision-making method based on improved DDPG algorithm
By improving the DDPG algorithm, using the LSTM network to screen historical information and combining the pre-trained DDPG network, the problem of insufficient decision-making capabilities in the marine unmanned cluster was solved, the round-up success rate and decision-making capabilities were improved, and more efficient round-up maneuver confrontation was achieved.
Patent Information
- Application Number
- CN202510534670.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing DDPG algorithm cannot effectively utilize historical information in marine unmanned clusters, resulting in insufficient decision-making capabilities and low roundup success rate.
The DDPG algorithm is improved, and the state position of each aircraft is input into the LSTM network at the previous moment, and the effective historical information is obtained, and combined with the pre-trained DDPG network, the execution actions of each aircraft are determined. Reward values ensure that the vehicle learns the optimal strategy by building a kinematic model of the vehicle, combining individuals, teams and additional rewards.
It improves the decision-making ability and round-up success rate of the unmanned ocean cluster, balances competition and cooperation, avoids collisions between the aircraft beyond the map range and team, and improves round-up efficiency.
Smart Images

Figure CN120069018A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of marine unmanned cluster game decision-making, and particularly to a method for marine unmanned cluster encirclement maneuver decision-making based on an improved DDPG algorithm. Background Art
[0002] Marine unmanned clusters have the advantages of low cost, high efficiency, and strong flexibility, and show great application potential in the fields of marine environmental monitoring, target search and tracking, etc. Although marine unmanned clusters show great application prospects in maritime law enforcement, their technological development and application still face many challenges. For example, how to optimize the cooperative combat strategy of the cluster to improve its overall effectiveness in the maritime encirclement maneuver confrontation, etc.
[0003] In the field of artificial intelligence, deep reinforcement learning has become an important method for solving complex tasks. Among them, the Deep Deterministic Policy Gradient (hereinafter referred to as DDPG) algorithm has achieved efficient learning in the continuous action space and can perform policy learning and optimization in the continuous action space. However, during the training process of this algorithm, due to the fact that individuals in the marine unmanned cluster cannot effectively utilize historical information, there is a problem of insufficient decision-making ability, resulting in a low success rate of encirclement. Summary of the Invention
[0004] The main purpose of this application is to provide a method for marine unmanned cluster encirclement maneuver decision-making based on an improved DDPG algorithm, aiming to solve the problem of low success rate of encirclement existing in the existing methods.
[0005] To achieve the above object, the present application provides a marine unmanned cluster hunting maneuver decision method based on an improved DDPG algorithm for controlling multiple vehicles to hunt a target, including: obtaining the state position of each vehicle at the previous moment, respectively inputting the state position of each vehicle at the previous moment into a Long - Short Term Memory (LSTM) network, and correspondingly obtaining the historical state position of each vehicle; inputting the historical state position of each vehicle into the corresponding pre - trained DDPG network to determine the execution action of each vehicle, and obtaining a marine unmanned cluster hunting maneuver decision; wherein, during the training process of the pre - trained DDPG network, the method for determining the reward value includes: constructing a kinematic model of the vehicle; determining the next - moment position state of each vehicle according to the current execution action of each vehicle in combination with the kinematic model; determining the probability that the target is successfully hunted by itself according to the next - moment position state of each vehicle; determining an individual reward according to the probability that each target is successfully hunted by itself; determining a team reward according to the distance between the next - moment position states of all vehicles and the position of the target; determining an additional reward according to the next - moment position states of all vehicles; and taking the sum of the individual reward, the team reward, and the additional reward as the reward value.
[0006] Optionally, the method for determining the probability that the target is successfully hunted by itself is: determining the distance between the current vehicle position and the target position. When the distance is less than or equal to a preset threshold, the probability that the target is successfully hunted by itself is 1; when the distance is greater than the preset threshold, the probability that the target is successfully hunted by itself is 0.
[0007] Optionally, the method for determining the individual reward is: when the probability that the target is successfully hunted by itself is 1, the value of the individual reward is a first preset value, otherwise it is 0.
[0008] Optionally, the method for determining the team reward is as follows:
[0009] wherein, is the team reward, is the position coordinate of the i th vehicle, is the position of the target, N is the number of vehicles.
[0010] Optionally, the method for determining the additional reward includes: comparing the position of the current vehicle with a preset area and other objects respectively. When it is determined that the position of the vehicle exceeds the preset area, the value of the additional reward is a second preset value; Optionally, the determination method of the additional reward further includes: comparing the position of the current vehicle with other objects respectively. When it is determined that the position of the vehicle coincides with the position of other objects, the value of the additional reward is the second preset value; otherwise, it is 0. Among them, other objects include other vehicles, targets or obstacles.
[0011] Optionally, the DDPG network includes Q a network, a policy network, a target policy network, and a target Q network; the training process of the pre-trained DDPG network corresponding to each vehicle includes: obtaining the state position of the current vehicle at the previous moment, inputting the state position of the vehicle at the previous moment into the LSTM network to obtain the historical state position of the vehicle; inputting the historical state position of the vehicle into the policy network to determine the current execution action of the vehicle; determining the reward value and the position state of the vehicle at the next moment according to the current execution action of the vehicle; storing the historical state position, the current execution action, the reward value, and the position state at the next moment of the vehicle as a set of experience values into the experience sample pool; the experience sample pool includes the experience values of all vehicles; using the experience values to Q update the parameters of the network, the policy network, the target policy network, and the target Q network to obtain the updated DDPG network; repeating the above method until the DDPG network converges to obtain the pre-trained DDPG network.
[0012] Optionally, using the experience values to Q update the parameters of the network, the policy network, the target policy network, and the target Q network includes: randomly selecting multiple groups of experience values in the experience sample pool, inputting the position states of all vehicles at the next moment in the multiple groups of experience values into the target policy network to generate actions; determining the target Q value of the network according to the actions generated by the target policy network and the position state of the current vehicle at the next moment; inputting the current execution action and the historical state position of the vehicle into the Q network to determine the Q value; determining the loss function according to the Q value and the target Q value of the network, and updating the parameters of the Q network by minimizing the loss function; determining the target function of the policy network according to the historical state position of the vehicle, determining the gradient of the target function with respect to the parameters of the policy network, and updating the parameters of the policy network by using the gradient descent method for the gradient; using the soft update method to update the parameters of the target policy network and the target Q network.
[0013] Compared with the prior art, the beneficial effects of the present application are as follows: The marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present invention inputs the state position of each vehicle at the previous moment into the LSTM network. The LSTM is used to preprocess the position state of the vehicle at the previous moment, screen and obtain effective historical information, and avoid the problem of gradient disappearance when facing high-dimensional state information in reinforcement learning. Combining it with the DDPG algorithm endows each vehicle with the ability to process and save historical information, enabling it to select the best strategy based on the current state and historical information, improving the decision-making ability, and thus increasing the hunting success rate; the sum of the individual reward, the team reward and the additional reward is used as the reward value. The settings of the individual reward and the team reward help the vehicles learn the optimal strategy, and at the same time avoid the situation where some vehicles in the team do not work and gain benefits, and the situation where all vehicles only pursue the maximum reward and lead to the failure of the task; it allows each vehicle to learn and train better and balance the relationship between competition and cooperation; the setting of the additional reward can prevent the vehicle from exceeding the map range and effectively avoid collisions between teams and with obstacles; through the above reward value, the hunting success rate can be further increased. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a schematic flow chart of a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application; Figure 2 is a training flow chart of the DDPG network in a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application; Figure 3 is a hunting scenario diagram of a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application; Figure 4 is a parameter update flow chart of the DDPG network in a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application; Figure 5 is the hunting trajectory result of the marine unmanned cluster based on the marine unmanned cluster hunting maneuver decision-making method of the improved DDPG algorithm of the present application; Figure 6 is the average reward diagram of all vehicles in a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application; Figure 7 is the reward diagram obtained by each vehicle in a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application; Figure 8 is the hunting success rate diagram of a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application.
[0015] The realization of the purpose of this application, its functional characteristics and advantages will be further described in conjunction with embodiments with reference to the accompanying drawings. Specific Embodiments
[0016] To make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the accompanying drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all of them. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0017] The present invention provides a marine unmanned cluster hunting maneuver decision-making method based on an improved DDPG algorithm, which is used to control multiple vehicles to hunt a target, as Figure 1 shown, and specifically includes the following steps: Step S1, obtain the state positions of each vehicle at the previous moment, and input the state positions of each vehicle at the previous moment into the LSTM network respectively to obtain the historical state positions of each vehicle; the position state includes position and heading angle.
[0018] It should be noted that the hunting maneuver confrontation decision-making problem belongs to a sequential decision-making problem, and the vehicle needs to execute a long sequence of decisions. Therefore, when making a decision, the vehicle must consider historical state information in order to make a more accurate decision. In this embodiment, the historical state position information of each vehicle is pre-screened through the LSTM network, which can effectively add or delete historical information in the memory storage unit, thereby reducing the calculation of invalid information, and thus reducing the calculation amount of the DDPG network and the complexity of the model.
[0019] Step S2, input the historical state positions of each vehicle into the corresponding pre-trained DDPG network, determine the execution actions of each vehicle, and obtain the marine unmanned cluster hunting maneuver decision-making.
[0020] In this embodiment, each vehicle corresponds to a pre-trained DDPG network. The pre-trained DDPG network outputs the execution actions of the corresponding vehicle, and then the marine unmanned cluster hunting maneuver decision-making can be obtained. Each vehicle is controlled according to this decision to complete the hunting of the target.
[0021] Furthermore, the DDPG network includes Q a network, a policy network, a target policy network and a target Q network, as Figure 2 shown. The training process of the pre-trained DDPG network corresponding to each vehicle includes: Step S21: Obtain the state position of the current vehicle at the previous moment, input the state position of the vehicle at the previous moment into the LSTM network, and obtain the historical state position of the vehicle, that is, the hidden state at the current moment is:
[0022] In the formula, is the hidden state at the current moment.
[0023] It can be understood that the input of the LSTM network is the position state at the previous moment. The LSTM network will retain the information of the position state at the previous moment of each input, that is, the sequence , where k is the historical number of steps traced back from the current time step t . And filter this sequence to output the historical state position of the vehicle.
[0024] Input the historical state position of the vehicle into the policy network to determine the current execution action of the vehicle, where the execution action includes the movement speed and the yaw angular speed;
[0025] In the formula, is the random noise, is the parameter of the policy network, is the action taken by the policy network; Step S22: Determine the reward value and the position state of the vehicle at the next moment according to the current execution action of the vehicle; specifically, it includes the following steps: Step S221: Construct the kinematic model of the vehicle; according to the current execution action of each vehicle, combine the kinematic model to determine the position state of each vehicle at the next moment; that is, input the movement speed and the yaw angular speed of the current vehicle into the kinematic model, and the position state of the current vehicle at the next moment can be obtained.
[0026] Among them, the kinematic model of the vehicle on the horizontal plane is:
[0027] In the formula, is the first-order derivative of the position coordinate of the vehicle in the ground coordinate system, is the first-order derivative of the heading angle; , , r are respectively the longitudinal speed, the lateral speed and the yaw angular speed of the vehicle in its own vehicle coordinate system. The movement speed v is calculated by the formula:
[0028] Due to the limitations of dynamics on the movement of the vehicle, its movement speed v and yaw angular velocity r are within a limited range. In this embodiment, to make the movement of the vehicle more in line with the actual situation, the v range is set between 0 and 8 knots, r and the range of
[0029] is between -6° / s and 6° / s. It should be noted that after obtaining the position state of the vehicle at the next moment, the historical state position of the vehicle is updated using the position state of the vehicle at the next moment, and after the update, it is , which is used as the input of the LSTM network in the next iteration.
[0030] To more conveniently determine the reward value, this embodiment constructs an unmanned marine cluster hunting maneuver scenario, as shown in Figure 3 . In the figure, the position state of our side is the position state of the vehicle. The reward value is determined in this scenario, and the specific method for determining the reward value is as follows: Step S222: Determine the probability that the target is successfully hunted by itself according to the position state of each vehicle at the next moment; determine the individual reward according to the probability that each target is successfully hunted by itself; Specifically, the method for determining the probability that the target is successfully hunted by itself is: determine the distance between the current vehicle position and the target position. When the distance is less than or equal to the preset threshold L , the probability that the target is successfully hunted by itself is 1; when the distance is greater than the preset threshold L , the probability that the target is successfully hunted by itself is 0. The distance judgment formula is:
[0031] On this basis, the method for determining the individual reward is: when the probability that the target is successfully hunted by itself is 1, the value of the individual reward is the first preset value , otherwise it is 0, that is:
[0032] Preferably, take 10.
[0033] Step S223: Determine the team reward according to the distance between the position state of all vehicles at the next moment and the position of the target; the method for determining the team reward is as follows:
[0034] In the formula, is the position coordinate of the i th vehicle, is the position of the target,N is the number of vehicles.
[0035] Step S224: Determine the additional reward according to the position states of all vehicles at the next moment, specifically as follows: Compare the position of the current vehicle with the preset area and other objects respectively. When it is determined that the position of the vehicle exceeds the preset area (map range), the value of the additional reward is the second preset value ; Or, compare the position of the current vehicle with other objects respectively. When it is determined that the position of the vehicle coincides with the position of other objects, the value of the additional reward is the second preset value , otherwise it is 0; where other objects include other vehicles, targets or obstacles, that is:
[0036] Preferably, take -100.
[0037] Step S225: Take the sum of the individual reward, the team reward and the additional reward as the reward value, that is .
[0038] Step S23: Take the historical state position, the current executed action, the reward value and the position state at the next moment of the vehicle, that is , Save it as a set of experience values into the experience sample pool; the experience sample pool includes the experience values of all vehicles; Step S24: In the improved DDPG algorithm, each vehicle needs to use the experience values generated by other vehicles in the team to update its own neural network parameters. That is, randomly select multiple sets of experience values from the experience replay pool to update the parameters of the DDPG network, as Figure 4 shown, to obtain the updated DDPG network; In this embodiment, the experience sample pool is used to save the experience sample data obtained by the vehicle when interacting with the environment. During training, experience data can be extracted from the experience sample pool for training, thereby improving the utilization rate and training effect of the experience data. In addition, the addition of the experience sample pool helps to stabilize the training process of the vehicle. By randomly extracting experience data, the correlation between data is reduced, effectively preventing overfitting and making the training process more stable. Finally, the establishment of the experience sample pool solves the problem of how to balance exploration and exploitation of known training data in reinforcement learning. The vehicle can randomly select experience samples from the central experience sample pool for learning, which helps to alleviate this problem.
[0039] Step S241: Randomly select multiple groups of empirical values from the empirical sample pool, and input the next moment position states of all the vehicles in the multiple groups of empirical values into the target policy network of the current vehicle to generate actions; Exemplarily, m The group of empirical values is expressed as:
[0040]
[0041] In the formula, is the target policy, is the action generated by the target policy network, are the parameters of the target policy network; Step S242: Determine the target Q value of the network according to the action generated by the target policy network and the position state of the current vehicle at the next moment ;
[0042] In the formula, is the reward value, γ is the discount factor, is the target Q parameters of the network, is the target Q network function; Step S243: Input the current execution action and the historical state position of the vehicle into the Q network to determine the value;
[0043] In the formula, is the Q network function, is the Q parameters of the network; Step S244: Determine the loss function according to the value and the target Q value of the network, and update the parameters of the Q network by minimizing the loss function; The calculation formula of the loss function is as follows:
[0044] In the formula, is the expectation; Step S245: Determine the target function of the policy network according to the historical state position of the vehicle, determine the gradient of the target function with respect to the parameters of the policy network, and update the parameters of the policy network by using the gradient descent method for the gradient; The calculation formula of the target function of the policy network is as follows:
[0045] The gradient is obtained by using the gradient descent method for the gradient:
[0046] wherein, is the gradient operator of the subscript variable.
[0047] Step S246, update the parameters of the target policy network and the target Q network to obtain the updated DDPG network; specifically, use the soft update method to update the parameters of the target policy network and the target Q network, and the update formula is as follows:
[0048]
[0049] wherein, τ is the inertia update rate, is the network parameter of the updated target policy network, is the updated target Q network parameter of the network.
[0050] Step S25, repeat Step S21 - Step S24 until the DDPG network converges to obtain the pre-trained DDPG network.
[0051] Next, a specific example is used to introduce the decision-making method of the present invention.
[0052] The parameters of the experimental environment of this example are shown in Table 1.
[0053] Table 1 Experimental parameter setting values
[0054] The main objective of this example is to evaluate the convergence ability of the improved DDPG algorithm used by the unmanned marine cluster in the pursuit maneuver confrontation, and compare it with the traditional DDPG algorithm. To ensure the randomness of the results, the positions and headings of the vehicles, as well as the position of the target, are randomly initialized at the beginning of each simulation.
[0055] Control the movement of the vehicle using the pursuit maneuver decision-making method of the unmanned marine cluster of the present invention to obtain the pursuit trajectory map of the vehicle, as shown in Figure 5 , the black circular area in the figure represents randomly generated obstacles, the triangles represent the starting positions of our side and the target (i.e., the vehicles), and after the pursuit is successful, the circles within the pursuit circle represent the final positions of both sides. It can be seen that after the simulation starts, the pursuit vehicles quickly disperse to search for the target and finally achieve the pursuit target through cooperation.
[0056] During the training process of the improved DDPG network in this example, the average reward value of our side (i.e., the average reward value of all vehicles) in the traditional DDPG training process was compared, as shown in Figure 6 . It can be seen that both algorithms showed good convergence. However, the improved DDPG algorithm basically reached convergence after 25,000 episodes, and the reward value was higher than that of the traditional DDPG algorithm. This is mainly because the improved algorithm made better use of historical information. The results show that the improved DDPG algorithm performs better than the traditional algorithm in the pursuit maneuver confrontation task, improving the pursuit efficiency.
[0057] During the training process of the improved DDPG network in this example, the reward values of each of the 4 vehicles were compared with the reward values of each of the 4 vehicles during the training of the traditional DDPG network, as shown in Figure 7 . Figure 7 In (a), the solid line represents the change in the reward value of our vehicle 1 corresponding to the DDPG network training method of the present invention. Figure 7 In (b), the solid line represents the change in the reward value of our vehicle 2 corresponding to the DDPG network training method of the present invention. Figure 7 In (c), the solid line represents the change in the reward value of our vehicle 3 corresponding to the DDPG network training method of the present invention. Figure 7 In (d), the solid line represents the change in the reward value of our vehicle 4 corresponding to the DDPG network training method of the present invention. The dashed lines in the four figures represent the change in the reward value of the vehicles of the corresponding traditional DDPG algorithm. It can be seen that the individual reward values under both algorithms can converge and reach a relatively high level. In the reward graph of pursuing our vehicle 3, the difference in the reward values between the two algorithms is not significant, and sometimes the reward value of the traditional DDPG algorithm is higher, indicating that the pursuing vehicles pay more attention to the overall team reward rather than just the individual reward.
[0058] In the pursuit maneuver confrontation environment, in order to accurately calculate the pursuit success rate, the number of successful pursuits in the traditional DDPG and improved DDPG algorithms was recorded respectively to calculate the success rate. Considering the limited power of the actual vehicle, if the simulation step size exceeds 500 steps, it is regarded as a failed pursuit.
[0059] The calculation rule of the pursuit success rate is defined as follows: Each round of the above simulation is repeated 100 times, the number of successful pursuits is calculated, and the success rate of each round of simulation is obtained. This success rate was statistically analyzed for 500 rounds, and the results are shown in Figure 8 . As the number of simulation rounds increases, the pursuit success rate gradually increases. After 500 rounds of simulation, the success rate of the traditional DDPG algorithm reaches 75%, while that of the improved DDPG algorithm reaches 88%, which further proves the superiority of the improved algorithm.
[0060] The simulation results show that the proposed method for maneuver decision-making of unmanned marine swarms for encirclement and capture can effectively utilize the historical information of unmanned marine swarms. Compared with the traditional DDPG algorithm, the individual average reward value is higher, the convergence speed is faster, and the success rate of encirclement and capture is greatly improved. In addition, the encircling and capturing vehicles demonstrate excellent teamwork capabilities during the process, significantly enhancing the efficiency of maneuver confrontation for encirclement and capture.
[0061] The above are only the preferred embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of this application.
Claims
1. A marine unmanned swarm capture maneuvering decision-making method based on an improved DDPG algorithm, which is used to control multiple aircraft to capture a target, and is characterized by: include: Obtain the state position of each aircraft at a previous moment, input the state position of each aircraft at a previous moment into the LSTM network, and obtain the corresponding historical state position of each aircraft; Input the historical state position of each of the aircraft into the corresponding pre-trained DDPG network, determine the execution action of each aircraft, and obtain the marine unmanned swarm capture maneuver decision; Among them, during the training process of the pre-trained DDPG network, the method for determining the reward value includes: Constructing a kinematic model of the aircraft; determining the position state of each aircraft at the next moment in combination with the kinematic model according to the current execution action of each aircraft; Determine the probability of the target being successfully captured by itself according to the position state of each of the aircraft at the next moment; determine the individual reward according to the probability of the target being successfully captured by itself; Determine the team reward based on the distance between the next moment position state of all spacecraft and the target position; Determine the additional reward based on the next moment position status of all spacecraft; The sum of the individual reward, team reward and additional reward is taken as the reward value.
2. The marine unmanned swarm capture maneuvering decision-making method based on the improved DDPG algorithm according to claim 1 is characterized in that: The probability of the target being captured by itself is determined as follows: The distance between the current aircraft position and the target position is determined. When the distance is less than or equal to a preset threshold, the probability that the target is successfully captured by itself is 1; when the distance is greater than the preset threshold, the probability that the target is successfully captured by itself is 0.
3. The marine unmanned swarm capture maneuvering decision-making method based on the improved DDPG algorithm according to claim 1 is characterized in that: The individual reward is determined in the following manner: when the probability that the target is successfully captured by itself is 1, the value of the individual reward is a first preset value, otherwise it is 0.
4. The marine unmanned swarm capture maneuvering decision-making method based on the improved DDPG algorithm according to claim 1 is characterized in that: The team rewards are determined as follows: In the formula, Rewards for the team, For the i The position coordinates of the spacecraft, is the target location, N is the number of aircraft.
5. The marine unmanned swarm capture maneuvering decision-making method based on the improved DDPG algorithm according to claim 1 is characterized in that: The additional rewards are determined in the following ways: The current position of the aircraft is compared with a preset area and other objects respectively. When it is determined that the position of the aircraft exceeds the preset area, the value of the additional reward is a second preset value.
6. The marine unmanned swarm capture maneuvering decision-making method based on the improved DDPG algorithm according to claim 5 is characterized in that: The additional reward may also be determined by: Compare the current position of the aircraft with the other objects respectively, and when it is determined that the position of the aircraft coincides with the position of the other object, the value of the additional reward is a second preset value, otherwise it is 0; The other objects include other aircraft, targets or obstacles.
7. The marine unmanned swarm capture maneuvering decision-making method based on the improved DDPG algorithm according to claim 1 is characterized in that: The DDPG network includes Q Network, Policy Network, Target Policy Network and Target Q Network; The training process of the pre-trained DDPG network corresponding to each of the aircraft includes: Obtain the state position of the current aircraft at a previous moment, input the state position of the aircraft at the previous moment into the LSTM network, and obtain the historical state position of the aircraft; Inputting the historical state position of the aircraft into a policy network to determine the current execution action of the aircraft; Determine a reward value and a position state of the aircraft at a next moment according to the current execution action of the aircraft; The historical state position, current execution action, reward value and next moment position state of the aircraft are stored in an experience sample pool as a set of experience values; the experience sample pool includes the experience values of all aircraft; Using the experience value Q Network, Policy Network, Target Policy Network and Target Q The parameters of the network are updated to obtain the updated DDPG network; Repeat the above method until the DDPG network converges to obtain the pre-trained DDPG network.
8. The marine unmanned swarm capture maneuvering decision-making method based on the improved DDPG algorithm according to claim 7 is characterized in that: The use of the experience value Q Network, Policy Network, Target Policy Network and Target Q The network parameters are updated, including: Randomly selecting multiple sets of experience values from the experience sample pool, and inputting the next moment position states of all the aircraft in the multiple sets of experience values into the target strategy network to generate actions; According to the action generated by the target strategy network and the position state of the current aircraft at the next moment, the target is determined. Q The target value of the network; The current execution action and historical state position of the aircraft are input Q Network, OK Q value; According to the Q Value, target Q The target value of the network determines the loss function, and the loss function is minimized Q The parameters of the network are updated; Determine the policy network objective function according to the historical state position of the aircraft, determine the gradient of the objective function with respect to the policy network parameters, and update the policy network parameters using the gradient descent method; Use soft updates to target policy networks and targets Q The parameters of the network are updated.
Citation Information
Patent Citations
Robot path navigation method and system based on improved DDPG algorithm
CN113408782A
State prediction and DDPG combined multi-unmanned aerial vehicle hunting method
CN113625775A
Multi-unmanned aerial vehicle hunting strategy method based on CEL-MADDPG
CN115097861A
Unmanned ship cluster task scheduling and collaborative confrontation method based on MADDPG
CN116050795A