An Ocean Unmanned Cluster Encirclement Maneuvering Decision-Making Method Based on Improved DDPG Algorithm

By improving the DDPG algorithm, using the LSTM network to filter the historical information of the aircraft and combining the reward mechanism, the roundup decision-making of the unmanned marine cluster was optimized, which solved the problem of insufficient decision-making ability and improved the roundup success rate and efficiency.

CN120069018BActive Publication Date: 2025-07-04RES & DEV INST OF NORTHWESTERN POLYTECHNICAL UNIV IN SHENZHEN +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510534670.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-04
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing marine unmanned clusters lack decision-making capabilities in round-up maneuver confrontation, resulting in a low round-up success rate.

Method used

By improving the DDPG algorithm, the LSTM network is used to filter the historical status position information of the aircraft, and combined with individual rewards, team rewards and additional rewards, the reward value is constructed, the decision-making process of the aircraft is optimized, and the decision-making ability is improved.

Benefits of technology

It improves the decision-making ability and round-up success rate of the aircraft, enhances teamwork, avoids the aircraft from exceeding the map range and collision, and significantly improves the efficiency of round-up maneuver confrontation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069018B_ABST
    Figure CN120069018B_ABST
Patent Text Reader

Abstract

The present application discloses a method for marine unmanned cluster hunting maneuver decision-making based on an improved DDPG algorithm, specifically related to the field of marine unmanned cluster game decision-making. It includes: obtaining the state positions of each vehicle at the previous moment, respectively inputting the state positions of each vehicle at the previous moment into the LSTM network, and correspondingly obtaining the historical state positions of each vehicle; inputting the historical state positions of each vehicle into the corresponding pre-trained DDPG network to determine the execution actions of each vehicle, and obtaining the marine unmanned cluster hunting maneuver decision-making; wherein, during the training process of the pre-trained DDPG network, the reward value is the sum of the individual reward, the team reward and the additional reward. It can further improve the hunting success rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of marine unmanned swarm game decision-making, and particularly to a method for marine unmanned swarm pursuit maneuver decision-making based on an improved DDPG algorithm. Background Art

[0002] Marine unmanned swarms have advantages such as low cost, high efficiency, and strong flexibility, and show great application potential in fields such as marine environmental monitoring, target search and tracking. Although marine unmanned swarms show great application prospects in maritime law enforcement, their technological development and application still face many challenges. For example, how to optimize the cooperative combat strategy of the swarm to improve its overall effectiveness in the maritime pursuit maneuver confrontation.

[0003] In the field of artificial intelligence, deep reinforcement learning has become an important method for solving complex tasks. Among them, the Deep Deterministic Policy Gradient (hereinafter referred to as DDPG) algorithm has achieved efficient learning in the continuous action space and can perform policy learning and optimization in the continuous action space. However, during the training process of this algorithm, due to the fact that individuals in the marine unmanned swarm cannot effectively utilize historical information, there is a problem of insufficient decision-making ability, resulting in a low success rate of pursuit. Summary of the Invention

[0004] The main purpose of this application is to provide a method for marine unmanned swarm pursuit maneuver decision-making based on an improved DDPG algorithm, aiming to solve the problem of low success rate of pursuit existing in the existing methods.

[0005] To achieve the above object, the present application provides a method for maneuver decision-making of ocean unmanned clusters based on an improved DDPG algorithm, which is used to control multiple vehicles to hunt a target, including: obtaining the state position of each vehicle at the previous moment, respectively inputting the state position of each vehicle at the previous moment into a Long-Short Term Memory (LSTM) network, and correspondingly obtaining the historical state position of each vehicle; inputting the historical state position of each vehicle into the corresponding pre-trained DDPG network to determine the execution action of each vehicle, and obtaining the maneuver decision-making of ocean unmanned clusters for hunting; wherein, during the training process of the pre-trained DDPG network, the method for determining the reward value includes: constructing a kinematic model of the vehicle; according to the current execution action of each vehicle, combining the kinematic model to determine the position state of each vehicle at the next moment; according to the position state of each vehicle at the next moment, determining the probability that the target is successfully hunted by itself; according to the probability that each target is successfully hunted by itself, determining the individual reward; according to the distance between the position states of all vehicles at the next moment and the position of the target, determining the team reward; according to the position states of all vehicles at the next moment, determining the additional reward; taking the sum of the individual reward, the team reward and the additional reward as the reward value.

[0006] Optionally, the method for determining the probability that the target is successfully hunted by itself is: determining the distance between the current position of the vehicle and the position of the target, when the distance is less than or equal to a preset threshold, the probability that the target is successfully hunted by itself is 1; when the distance is greater than the preset threshold, the probability that the target is successfully hunted by itself is 0.

[0007] Optionally, the method for determining the individual reward is: when the probability that the target is successfully hunted by itself is 1, the value of the individual reward is the first preset value, otherwise it is 0.

[0008] Optionally, the method for determining the team reward is as follows:

[0009]

[0010] wherein, is the team reward, is the position coordinate of the i th vehicle, is the position of the target, N is the number of vehicles.

[0011] Optionally, the method for determining the additional reward includes: comparing the position of the current vehicle with a preset area and other objects respectively, when it is determined that the position of the vehicle exceeds the preset area, the value of the additional reward is the second preset value;

[0012] Optionally, the determination method of the additional reward further includes: comparing the position of the current vehicle with other objects respectively. When it is determined that the position of the vehicle coincides with the position of other objects, the value of the additional reward is the second preset value; otherwise, it is 0. Among them, the other objects include other vehicles, targets or obstacles.

[0013] Optionally, the DDPG network includes Q a network, a policy network, a target policy network, and a target Q network; the training process of the pre-trained DDPG network corresponding to each vehicle includes: obtaining the state position of the current vehicle at the previous moment, inputting the state position of the vehicle at the previous moment into the LSTM network to obtain the historical state position of the vehicle; inputting the historical state position of the vehicle into the policy network to determine the current execution action of the vehicle; determining the reward value and the next moment position state of the vehicle according to the current execution action of the vehicle; taking the historical state position, the current execution action, the reward value, and the next moment position state of the vehicle as a set of experience values and storing them in the experience sample pool; the experience sample pool includes the experience values of all vehicles; using the experience values to Q update the parameters of the network, the policy network, the target policy network, and the target Q network to obtain the updated DDPG network; repeating the above method until the DDPG network converges to obtain the pre-trained DDPG network.

[0014] Optionally, using the experience values to Q update the parameters of the network, the policy network, the target policy network, and the target Q network includes: randomly selecting multiple groups of experience values in the experience sample pool, inputting the next moment position states of all vehicles in the multiple groups of experience values into the target policy network to generate actions; determining the target Q value of the network according to the actions generated by the target policy network and the next moment position state of the current vehicle; inputting the current execution action and the historical state position of the vehicle into the Q network to determine the Q value; determining the loss function according to the Q value and the target Q value of the network, and updating the parameters of the Q network by minimizing the loss function; determining the target function of the policy network according to the historical state position of the vehicle, determining the gradient of the target function with respect to the parameters of the policy network, and updating the parameters of the policy network by using the gradient descent method for the gradient; using the soft update method to update the parameters of the target policy network and the target Q network.

[0015] Compared with the prior art, the beneficial effects of the present application are as follows:

[0016] The marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present invention inputs the state position of each vehicle at the previous moment into the LSTM network, uses the LSTM to pre-process the position state of the vehicle at the previous moment in advance, screens out effective historical information, and avoids the problem of gradient disappearance when facing high-dimensional state information in reinforcement learning. Combining it with the DDPG algorithm, it adds the ability to process and save historical information for each vehicle, enabling it to select the best strategy based on the current state and historical information, improving the decision-making ability, and thus increasing the hunting success rate; the sum of the individual reward, the team reward, and the additional reward is used as the reward value. The settings of the individual reward and the team reward help the vehicle learn the optimal strategy, and at the same time avoid the situation where some vehicles in the team do not work and gain benefits, and the situation where all vehicles only pursue the maximum reward and lead to the failure of the task; it can enable each vehicle to better learn and train, and balance the relationship between competition and cooperation; the setting of the additional reward can prevent the vehicle from exceeding the map range and effectively avoid collisions between teams and with obstacles; through the above reward value, the hunting success rate can be further increased. Description of the Drawings

[0017] Figure 1 It is a schematic flowchart of a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application;

[0018] Figure 2 It is a training flowchart of the DDPG network in a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application;

[0019] Figure 3 It is a hunting scenario diagram of a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application;

[0020] Figure 4 It is a parameter update flowchart of the DDPG network in a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application;

[0021] Figure 5 It is the hunting trajectory result of the marine unmanned cluster based on the marine unmanned cluster hunting maneuver decision-making method of the improved DDPG algorithm of the present application;

[0022] Figure 6 It is the average reward diagram of all vehicles of a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application;

[0023] Figure 7 It is the reward diagram obtained by each vehicle of a marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm of the present application;

[0024] Figure 8This is the graph of the capture success rate of a marine unmanned cluster capture maneuver decision-making method based on an improved DDPG algorithm in this application.

[0025] The realization, functional characteristics, and advantages of the purpose of this application will be further described with reference to the accompanying drawings in combination with embodiments. Specific embodiments

[0026] To make the purpose, technical solutions, and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below with reference to the accompanying drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0027] The present invention provides a marine unmanned cluster capture maneuver decision-making method based on an improved DDPG algorithm for controlling multiple vehicles to capture a target, such as Figure 1 shown, and specifically includes the following steps:

[0028] Step S1, obtain the state positions of each vehicle at the previous moment, and input the state positions of each vehicle at the previous moment into the LSTM network respectively to obtain the historical state positions of each vehicle; the position state includes the position and the heading angle.

[0029] It should be noted that the capture maneuver confrontation decision-making problem belongs to a sequential decision-making problem, and the vehicle needs to execute a long sequence of decisions. Therefore, when making a decision, the vehicle must consider the historical state information to make a more accurate decision. In this embodiment, the historical state position information of each vehicle is pre-screened through the LSTM network, which can effectively increase or delete the historical information in the memory storage unit, thereby reducing the calculation of invalid information, and thus reducing the calculation amount of the DDPG network and the complexity of the model.

[0030] Step S2, input the historical state positions of each vehicle into the corresponding pre-trained DDPG network, determine the execution actions of each vehicle, and obtain the capture maneuver decision of the marine unmanned cluster.

[0031] In this embodiment, each vehicle corresponds to a pre-trained DDPG network. The pre-trained DDPG network outputs the execution actions of the corresponding vehicle, and then the capture maneuver decision of the marine unmanned cluster can be obtained. Each vehicle is controlled according to this decision to complete the capture of the target.

[0032] Further, the DDPG network includes Q network, policy network, target policy network, and target Q network, such as Figure 2As shown, the training process of the pre-trained DDPG network corresponding to each vehicle includes:

[0033] Step S21, obtain the state position of the current vehicle at the previous moment, input the state position of the current vehicle at the previous moment into the LSTM network, and obtain the historical state position of the vehicle, that is, the hidden state at the current moment is:

[0034]

[0035] In the formula, is the hidden state at the current moment.

[0036] It can be understood that the input of the LSTM network is the position state at the previous moment, and the LSTM network will retain the information of the position state at the previous moment of each input, that is, the sequence , where k is the historical number of steps traced back from the current time step t , and this sequence is screened to output the historical state position of the vehicle.

[0037] Input the historical state position of the vehicle into the policy network to determine the current execution action of the vehicle, where the execution action includes the movement speed and the yaw angular speed;

[0038]

[0039] In the formula, is the random noise, is the parameter of the policy network, is the action taken by the policy network;

[0040] Step S22, determine the reward value and the position state of the vehicle at the next moment according to the current execution action of the vehicle; specifically, it includes the following steps:

[0041] Step S221, construct the kinematic model of the vehicle; according to the current execution action of each vehicle, combine the kinematic model to determine the position state of each vehicle at the next moment; that is, input the movement speed and the yaw angular speed of the current vehicle into the kinematic model, and the position state of the current vehicle at the next moment can be obtained.

[0042] Among them, the kinematic model of the vehicle on the horizontal plane is:

[0043]

[0044] In the formula, is the first derivative of the position coordinate of the vehicle in the ground coordinate system, is the first derivative of the heading angle; 、 , r They are respectively the longitudinal speed, lateral speed, and yaw angular velocity of the vehicle in its own vehicle coordinate system. The motion speed v The calculation formula is:

[0045]

[0046] Due to the motion of the vehicle being restricted by dynamics, its motion speed v and yaw angular velocity r are within a limited range. In this embodiment, to make the motion of the vehicle more in line with the actual situation, v is set in the range of 0 to 8 knots, r is in the range of -6° / s to 6° / s.

[0047] It should be noted that after obtaining the position state of the vehicle at the next moment, the historical state position of the vehicle is updated using the position state of the vehicle at the next moment, and after the update, it is , which is used as the input to the LSTM network during the next iteration.

[0048] To more conveniently determine the reward value, this embodiment constructs an unmanned marine cluster hunting maneuver scenario, as shown in Figure 3 . In the figure, the position state of our side is the position state of the vehicle. The reward value is determined in this scenario. The specific method for determining the reward value is as follows:

[0049] Step S222: Determine the probability that the target is successfully surrounded by itself according to the position state of each vehicle at the next moment; determine the individual reward according to the probability that each target is successfully surrounded by itself;

[0050] Specifically, the method for determining the probability that the target is successfully surrounded by itself is: Determine the distance between the current vehicle position and the target position. When the distance is less than or equal to the preset threshold L , the probability that the target is successfully surrounded by itself is 1; when the distance is greater than the preset threshold L , the probability that the target is successfully surrounded by itself is 0. The distance judgment formula is:

[0051]

[0052] On this basis, the method for determining the individual reward is: When the probability that the target is successfully surrounded by itself is 1, the value of the individual reward is the first preset value , otherwise it is 0, that is:

[0053]

[0054] Preferably, Take 10.

[0055] Step S223: Determine the team reward according to the distances between the next - moment position states of all vehicles and the position of the target. The method for determining the team reward is as follows:

[0056]

[0057] In the formula, is the position coordinate of the i -th vehicle, is the position of the target, N is the number of vehicles.

[0058] Step S224: Determine the additional reward according to the next - moment position states of all vehicles. Specifically, it is as follows:

[0059] Compare the position of the current vehicle with the preset area and other objects respectively. When it is determined that the position of the vehicle exceeds the preset area (map range), the value of the additional reward is the second preset value ;

[0060] Or, compare the position of the current vehicle with other objects respectively. When it is determined that the position of the vehicle coincides with the position of other objects, the value of the additional reward is the second preset value , otherwise it is 0. Among them, other objects include other vehicles, the target or obstacles, that is:

[0061]

[0062] Preferably, take - 100.

[0063] Step S225: Take the sum of the individual reward, the team reward and the additional reward as the reward value, that is .

[0064] Step S23: Take the historical state position, the current execution action, the reward value and the next - moment position state of the vehicle, that is , and store them as a set of experience values in the experience sample pool. The experience sample pool includes the experience values of all vehicles;

[0065] Step S24: In the improved DDPG algorithm, each vehicle needs to use the experience values generated by other vehicles in the team to update its own neural network parameters. That is, randomly select multiple sets of experience values in the experience replay pool to update the parameters of the DDPG network. As Figure 4 shown, obtain the updated DDPG network;

[0066] In this embodiment, the empirical sample pool is used to store the empirical sample data obtained when the vehicle interacts with the environment. During training, empirical data can be sampled from the empirical sample pool for training, thereby improving the utilization rate of empirical data and the training effect. In addition, the addition of the empirical sample pool helps to stabilize the training process of the vehicle. By randomly sampling empirical data, the correlation between data is reduced, effectively preventing overfitting and making the training process more stable. Finally, the establishment of the empirical sample pool solves the problem of how to balance exploration and exploitation of known training data in reinforcement learning. The vehicle can randomly select empirical samples from the central empirical sample pool for learning, which helps to alleviate this problem.

[0067] Step S241, randomly select multiple groups of empirical values in the empirical sample pool, and input the position states of all vehicles at the next moment in the multiple groups of empirical values into the target policy network of the current vehicle to generate actions; Exemplarily, m The group of empirical values is expressed as:

[0068]

[0069]

[0070] In the formula, is the target policy, is the action generated by the target policy network, are the parameters of the target policy network;

[0071] Step S242, determine the target Q value of the target network according to the action generated by the target policy network and the position state of the current vehicle at the next moment

[0072]

[0073] In the formula, is the reward value, γ is the discount factor, is the parameter of the target Q network, is the target Q network function;

[0074] Step S243, input the current execution action and the historical state position of the vehicle into the Q network to determine the value;

[0075]

[0076] In the formula, is the Q network function, is the QParameters of the network;

[0077] Step S244. According to value and the target Q value of the network, determine the loss function, and update the parameters of the Q network by minimizing the loss function; the calculation formula of the loss function is as follows:

[0078]

[0079] In the formula, is the expectation;

[0080] Step S245. Determine the target function of the policy network according to the historical state position of the vehicle, determine the gradient of the target function with respect to the parameters of the policy network, and update the parameters of the policy network by using the gradient descent method for the gradient; the calculation formula of the target function of the policy network is as follows:

[0081]

[0082] The gradient descent method for the gradient gives:

[0083]

[0084] In the formula, is the gradient operator of the subscript variable.

[0085] Step S246. Update the parameters of the target policy network and the target Q network to obtain the updated DDPG network; specifically, use the soft update method to update the parameters of the target policy network and the target Q network, and the update formula is as follows:

[0086]

[0087]

[0088] In the formula, τ is the inertia update rate, are the network parameters of the updated target policy network, are the network parameters of the updated target Q network.

[0089] Step S25. Repeat Steps S21 - S24 until the DDPG network converges to obtain the pre-trained DDPG network.

[0090] The decision-making method of the present invention will be introduced below with a specific example.

[0091] The parameters of the experimental environment of this example are shown in Table 1.

[0092] Table 1 Experimental parameter settings

[0093]

[0094] The main objective of this example is to evaluate the convergence ability of the improved DDPG algorithm used by the unmanned marine cluster in the pursuit maneuver confrontation and compare it with the traditional DDPG algorithm. To ensure the randomness of the results, the positions and head angles of the vehicles, as well as the position of the target, are randomly initialized at the beginning of each simulation.

[0095] Using the pursuit maneuver decision-making method of the unmanned marine cluster of the present invention to control the movement of the vehicle, a pursuit trajectory map of the vehicle is obtained, as shown in Figure 5 , where the black circular area in the figure represents randomly generated obstacles, the triangles represent the starting positions of our side (i.e., the vehicles) and the target. After the pursuit is successful, the circles within the pursuit circle represent the final positions of both sides. It can be seen that after the simulation starts, the pursuit vehicles quickly disperse to search for the target and finally achieve the pursuit target through cooperation.

[0096] Compare the average reward value of our side (i.e., the average reward value of all vehicles) in the training process of the improved DDPG network with that in the training process of the traditional DDPG in this example, as shown in Figure 6 , it can be seen that both algorithms show good convergence, but the improved DDPG algorithm basically reaches convergence after 25,000 episodes, and the reward value is higher than that of the traditional DDPG algorithm. This is mainly because the improved algorithm makes better use of historical information. The results show that the improved DDPG algorithm performs better than the traditional algorithm in the pursuit maneuver confrontation task and improves the pursuit efficiency.

[0097] Compare the reward values of each of the 4 vehicles in the training process of the improved DDPG network in this example with the reward values of each of the 4 vehicles when training the traditional DDPG network, as shown in Figure 7 shown, Figure 7 In (a), the solid line represents the change in the reward value of our vehicle 1 corresponding to the DDPG network training method of the present invention, Figure 7 In (b), the solid line represents the change in the reward value of our vehicle 2 corresponding to the DDPG network training method of the present invention, Figure 7 In (c), the solid line represents the change in the reward value of our vehicle 3 corresponding to the DDPG network training method of the present invention, Figure 7The solid line in (d) of the figure shows the change in the reward value of our vehicle 4 corresponding to the DDPG network training method of the present invention. The dashed lines in the four figures represent the change in the reward value of the vehicle using the traditional DDPG algorithm. It can be seen that the individual reward values under both algorithms can converge and reach a relatively high level. In the reward graph of the encirclement of our vehicle 3, the difference in the reward values of the two algorithms is not significant, and sometimes the reward value of the traditional DDPG algorithm is higher, which indicates that the encircling vehicle pays more attention to the overall team reward rather than just the individual reward.

[0098] In the encirclement maneuver confrontation environment, in order to accurately calculate the encirclement success rate, the number of successful encirclements in the traditional DDPG and improved DDPG algorithms is recorded respectively to calculate the success rate. Considering the limited power of the actual vehicle, if the simulation step size exceeds 500 steps, it is regarded as a failed encirclement.

[0099] The calculation rule of the encirclement success rate is defined as follows: Each round of the above simulation is repeated 100 times, the number of successful encirclements is calculated, and the success rate of each round of simulation is obtained. This success rate is statistically analyzed for 500 rounds, and the results are as Figure 8 shown. As the number of simulation rounds increases, the encirclement success rate gradually increases. After 500 rounds of simulation, the success rate of the traditional DDPG algorithm reaches 75%, while the improved DDPG algorithm reaches 88%, which further proves the superiority of the improved algorithm.

[0100] The simulation results show that the proposed ocean unmanned cluster encirclement maneuver decision-making method of the present invention can effectively utilize the historical information of the ocean unmanned cluster. Compared with the traditional DDPG algorithm, it has a higher individual average reward value, a faster convergence speed, and greatly improves the encirclement success rate. In addition, the encircling vehicle demonstrates excellent teamwork ability during the process, significantly improving the efficiency of the encirclement maneuver confrontation.

[0101] The above is only the preferred embodiment of this application, and it does not limit the patent scope of this application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied to other related technical fields, shall be included in the patent protection scope of this application by the same token.

Claims

1. An ocean unmanned cluster hunting maneuver decision-making method based on an improved DDPG algorithm, which is used to control multiple vehicles to hunt a target, and is characterized in that Including: Obtain the state positions of each vehicle at the previous moment, and respectively input the state positions of each vehicle at the previous moment into the LSTM network to correspondingly obtain the historical state positions of each vehicle; Input the historical state positions of each vehicle into the corresponding pre-trained DDPG network, determine the execution actions of each vehicle, and obtain the decision-making for the unmanned marine cluster to conduct encirclement maneuvers; Among them, during the training process of the pre-trained DDPG network, the method for determining the reward value includes: Construct the kinematic model of the vehicle; according to the current execution actions of each vehicle, combine the kinematic model to determine the position states of each vehicle at the next moment; According to the position states of each vehicle at the next moment, determine the probability that the target is successfully encircled by itself; according to the probability that each target is successfully encircled by itself, determine the individual reward; According to the distances between the position states of all vehicles at the next moment and the positions of the targets, determine the team reward; According to the position states of all vehicles at the next moment, determine the additional reward; Take the sum of the individual reward, team reward, and additional reward as the reward value; The DDPG network includes Q a network, a policy network, a target policy network, and a target Q network; The training process of the pre-trained DDPG network corresponding to each of the said vehicles includes: Obtain the state position of the current vehicle at the previous moment, input the state position of the vehicle at the previous moment into the LSTM network, and obtain the historical state position of the vehicle; Input the historical state position of the vehicle into the policy network to determine the current execution action of the vehicle; According to the current execution action of the vehicle, determine the reward value and the position state of the vehicle at the next moment; Take the historical state position, current execution action, reward value, and position state at the next moment of the vehicle as a set of experience values and store them in the experience sample pool; the experience sample pool includes the experience values of all vehicles; Using the empirical value pair Q the network, the policy network, the target policy network, and the target Q network parameters are updated to obtain the updated DDPG network; Repeat the above training process until the DDPG network converges to obtain the pre-trained DDPG network.

2. The method for maneuver decision-making of unmanned marine clusters for encirclement and capture based on the improved DDPG algorithm according to claim 1, wherein The method for determining the probability that the target is successfully encircled by itself is: Determine the distance between the current vehicle position and the target position. When the distance is less than or equal to the preset threshold, the probability that the target is successfully encircled by itself is 1; when the distance is greater than the preset threshold, the probability that the target is successfully encircled by itself is 0.

3. The method for marine unmanned cluster hunting maneuver decision-making based on the improved DDPG algorithm according to claim 1, wherein, The method for determining the individual reward is: when the probability that the target is successfully encircled by itself is 1, the value of the individual reward is the first preset value, otherwise it is 0.

4. The marine unmanned cluster hunting maneuver decision-making method based on the improved DDPG algorithm according to claim 1, characterized in that, The method for determining the team reward is as follows: In the formula, is the team reward, is the i th position coordinate of the vehicle, is the position of the target, N is the number of vehicles.

5. The method for maneuver decision-making of an unmanned marine cluster for surrounding and capturing based on the improved DDPG algorithm according to claim 1, wherein, The method for determining the additional reward includes: Compare the position of the current vehicle with the preset area and other objects respectively. When it is determined that the position of the vehicle exceeds the preset area, the value of the additional reward is the second preset value.

6. The method for marine unmanned cluster hunting maneuver decision-making based on the improved DDPG algorithm according to claim 5, wherein The method for determining the additional reward also includes: Compare the position of the current vehicle with other objects respectively. When it is determined that the position of the vehicle coincides with the position of other objects, the value of the additional reward is the second preset value, otherwise it is 0; Among them, the other objects include other vehicles, targets, or obstacles.

7. The method for maneuver decision-making of unmanned marine clusters for encirclement and capture based on the improved DDPG algorithm according to claim 1, characterized in that, Using the experience value pair Q The network, the policy network, the target policy network, and the target Q Update the parameters of the network, including: Randomly select multiple groups of experience values from the experience sample pool, and input the position states of all vehicles at the next moment in the multiple groups of experience values into the target policy network to generate actions; Determine the target value of the target network according to the actions generated by the target policy network and the position state of the current vehicle at the next moment. Q value of the network; Input the current execution action and historical state position of the vehicle Q into the network to determine Q a value; According to the said Q value, the target Q value of the network to determine the loss function, and update the parameters of the Q network by minimizing the loss function; Determine the policy network objective function based on the historical state positions of the vehicle, determine the gradient of the objective function with respect to the policy network parameters, and update the parameters of the policy network using the gradient descent method for the gradient. Update the parameters of the target policy network and the target Q network using soft updates.

Citation Information

Patent Citations

  • State prediction and DDPG combined multi-unmanned aerial vehicle hunting method

    CN113625775A

  • Multi-unmanned aerial vehicle hunting strategy method based on CEL-MADDPG

    CN115097861A