Marine unmanned cluster attack and defense game decision-making method based on improved MADDPG algorithm

Through the improved MADDPG algorithm, the future position status of the aircraft and targets is predicted, and combined with long and short-term memory networks and multiple reward values, the problem of insufficient decision adaptability of marine unmanned clusters in dynamic environments is solved, significantly improving the winning rate of the aircraft.

CN120124675AActive Publication Date: 2025-06-10RES & DEV INST OF NORTHWESTERN POLYTECHNICAL UNIV IN SHENZHEN +1

Patent Information

Application Number
CN202510546607.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-06-10
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing offensive and defensive game decision-making methods in the ocean unmanned cluster face dynamically changing strike environments, which shows insufficient adaptability and slow response speed, resulting in a low winning rate of the aircraft.

Method used

The improved MADDPG algorithm is used to obtain the historical position status of the aircraft and targets, predict their future position status, and use a long and short-term memory network to process historical data, and input the pre-trained MADDPG network to determine the execution action. The reward value includes the target's reward value, self-protection reward value and environmental adaptive reward value to guide decision-making.

Benefits of technology

It improves the decision adaptability and response speed of the aircraft in a dynamic environment, enhances the winning rate in offensive and defensive games, ensures that decisions are based on the latest information, and makes full use of historical data to make precise decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124675A_ABST
    Figure CN120124675A_ABST
Patent Text Reader

Abstract

The invention discloses a marine unmanned cluster attack and defense game decision-making method based on an improved MADDPG algorithm, and particularly relates to the field of marine unmanned cluster game decision-making technologies. Comprising the following steps: respectively acquiring the position states of all aircrafts at historical moments and the position state of a target at historical moments; according to the position state of each aircraft at the historical moment, predicting to obtain a predicted position state of each aircraft; according to the position state of each target at the historical moment, predicting to obtain a predicted position state of each target; inputting the position state of each aircraft at the previous moment into a long short-term memory network to obtain the historical position state of each aircraft; and inputting the historical position state of each aircraft, the position state at the current moment, the predicted position state and the predicted position state of the target into a corresponding pre-trained MADDPG network, determining the execution action of each aircraft, and obtaining a marine unmanned cluster attack and defense game decision. The problem that the win rate of an aircraft is low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of marine unmanned cluster game decision-making technology, and particularly to a marine unmanned cluster attack and defense game decision-making method based on an improved MADDPG algorithm. Background Art

[0002] In traditional marine unmanned cluster systems, decision-making usually relies on preset rules or a central control unit. In the face of a dynamically changing combat environment, this approach often exhibits problems such as insufficient adaptability and slow response speed. Especially in attack and defense games, the behavior of the target is highly uncertain, and traditional decision-making methods are difficult to handle complex game strategies. Therefore, there is an urgent need for a marine unmanned cluster attack and defense game decision-making method that can adapt to complex marine environments and has efficient cooperation capabilities to improve the autonomous decision-making ability and task execution efficiency of the cluster in combat scenarios.

[0003] The Multi-Agent Deep Deterministic Policy Gradient (hereinafter referred to as MADDPG) algorithm has shown significant advantages in marine unmanned cluster attack and defense game decision-making. The MADDPG algorithm is particularly suitable for multi-agent environments, where the behavior of each agent depends not only on the state of the environment but also on the strategies of other agents. This characteristic enables the MADDPG algorithm to handle the complex interactions between multiple agents in marine unmanned cluster attack and defense game decision-making well. However, in order to attack each other, both sides will continuously learn new strategies and find suitable actions to execute. As time goes by, the number of both sides will keep changing, resulting in a change in the attack and defense game decision-making environment. This will cause the attack and defense game decision-making system of the marine unmanned cluster not to satisfy the Markov property, and the historical information is not fully utilized, resulting in a low winning rate of the vehicle. Summary of the Invention

[0004] The main purpose of this application is to provide a marine unmanned cluster attack and defense game decision-making method based on an improved MADDPG algorithm, aiming to solve the problem of the low winning rate of the vehicle existing in the existing marine unmanned cluster attack and defense game decision-making method.

[0005] To achieve the above object, the present application provides a method for ocean unmanned cluster attack and defense game decision-making based on an improved MADDPG algorithm, including: respectively obtaining the position states of all vehicles at historical moments and the position states of targets at historical moments; wherein, the position state of a vehicle or a target at a historical moment is: within the time period when the vehicle detects the target, starting from the last moment, the position states of the vehicle or the target for a continuous preset number of moments forward; predicting the predicted position states of each vehicle according to the position states of each vehicle at historical moments; predicting the predicted position states of each target according to the position states of each target at historical moments; inputting the position state of each vehicle at the previous moment into a long short-term memory network to obtain the historical position states of each vehicle; obtaining the position state of each vehicle at the current moment, and inputting the historical position state, the position state at the current moment, the predicted position state of each vehicle, and the predicted position state of the target into the corresponding pre-trained MADDPG network to determine the execution actions of each vehicle, and obtaining the ocean unmanned cluster attack and defense game decision; wherein, during the training process of the pre-trained MADDPG network, the reward value includes the reward value of the target, the self-defense reward value, and the environmental adaptability reward value; the method for determining the reward value includes: constructing a kinematic model of the vehicle; determining the position state of the current vehicle at the next moment according to the execution action of the current vehicle in combination with the kinematic model; determining the strike success rate of the current vehicle and the strike success rate of the target against the current vehicle according to the position state of the current vehicle at the next moment and the predicted position state of the target; determining the reward value of the target according to the strike success rate of the current vehicle; and determining the self-defense reward value according to the strike success rate of the target.

[0006] Optionally, the reward value of the target is determined according to the following method:

[0007]

[0008]

[0009] Wherein, is the reward value of the target, is the self-reward value, is the team reward value.

[0010] Optionally, the self-defense reward value is determined according to the following method:

[0011] Wherein, l is the number of targets, is the position coordinate of vehicle i, is the position coordinate of target j, It is the self - protection reward value.

[0012] Optionally, the method for determining the environmental adaptability reward value includes: comparing the position state of the current vehicle at the next moment with the position states of other vehicles at the next moment, the predicted position state of the target, the positions of obstacles, and the preset area to determine the environmental adaptability reward value.

[0013] Optionally, the method for determining the strike success rate of the current vehicle includes: obtaining the strike range length, the strike - available range angle of the current vehicle, and its straight - line distance from the target; when it is determined that the strike range length of the current vehicle is less than the straight - line distance between the current vehicle and the target, the strike success rate of the current vehicle is 0; when it is determined that the strike range length of the current vehicle is greater than or equal to the straight - line distance between the current vehicle and the target, and the angle between the velocity vector of the current vehicle and the line connecting to the target is less than the strike - available range angle of the current vehicle, the strike success rate of the current vehicle is 1.

[0014] Optionally, predicting the predicted position state of each vehicle based on the position state of each vehicle at the historical moment includes: using the position state of the vehicle at the historical moment as the input and obtaining the predicted position state of the vehicle by using a sliding - state prediction model.

[0015] Optionally, the sliding - state prediction model is a polynomial fitting model based on the least - squares principle.

[0016] Optionally, the method for determining the predicted position state of the target is the same as that of the predicted position state of the vehicle.

[0017] Optionally, each vehicle corresponds to a pre-trained MADDPG network; the training process of the pre-trained MADDPG network includes: respectively obtaining the position states of the current vehicle at historical moments and the position states of all targets at historical moments; predicting the predicted position state of the current vehicle according to the position state of the current vehicle at historical moments; predicting the predicted position states of each target according to the position states of each target at historical moments; inputting the position state of the current vehicle at the previous moment into a long short-term memory network to obtain the historical position state of the current vehicle; obtaining the position state of the current vehicle at the current moment, and inputting the historical position state, the position state at the current moment, the predicted position state of the current vehicle, and the predicted position states of all targets into the MADDPG network to obtain the execution action of the current vehicle; determining the position state and reward value of the current vehicle at the next moment according to the execution action of the current vehicle; storing the position state, the position state at the next moment, the execution action, and the reward value at the current moment into an experience replay pool; wherein, the experience replay pool contains multiple groups of experience values, and each group of experience values includes the position state, the execution action, the position state at the next moment, and the reward value corresponding to all vehicles at each moment; randomly selecting several groups of experience values from the experience replay pool to update the parameters of the MADDPG network to obtain an updated MADDPG network; repeating the above process to iteratively update the MADDPG network until the MADDPG network converges to obtain a pre-trained MADDPG network.

[0018] Optionally, the MADDPG network includes a policy network, a target policy network, Q network, and a target Q network; randomly selecting several groups of experience values from the experience replay pool to update the parameters of the MADDPG network includes: randomly selecting several groups of experience values from the experience replay pool, and determining the target Q value of the current vehicle by using the reward value of the current vehicle at the current moment and the position states of all vehicles at the next moment in the several groups of experience values; using the position state of the current vehicle at the current moment, the execution actions of all vehicles at the current moment, and the target Q value in the several groups of experience values to update the Q network parameters; using the position state of the current vehicle at the current moment and the execution actions of all vehicles at the current moment in the several groups of experience values to update the policy network parameters; using a soft update method to update the parameters of the target policy network and the target Q network.

[0019] Compared with the prior art, the beneficial effects of the present application are as follows: The ocean unmanned cluster attack - defense game decision - making method based on the improved MADDPG algorithm of the present invention predicts the predicted position state of the target according to the position state of each target at historical moments. It enables the vehicle to pre - insight and predict the position state information of the target at future moments through the learning ability of environmental changes, so as to generate execution actions matching the changed attack - defense game decision - making environment when the attack - defense game decision - making environment changes, and improve the winning rate of the vehicle. Further, starting from the last moment when the vehicle detects the target, the position states of the vehicle or the target at a continuously preset number of moments forward are used as the position states at historical moments. As the vehicle continues to move, the collected position states at historical moments are updated in real - time, gradually replacing and covering the old sampling points, ensuring that the prediction is always analyzed based on the latest and most relevant information, and further ensuring the winning rate of the vehicle. Using the long - short - term memory network to process the position states at historical moments can make full use of the accumulated historical data, and then achieve more accurate and efficient decision - making, improving the overall task execution efficiency. The reward value includes the reward value of the target, the self - protection reward value, and the environmental adaptability reward value. Considering both the overall income of the vehicle team and the respective incomes of each vehicle, it can effectively avoid the situation where some vehicles do not work and only focus on the overall income of the team, further improving the winning rate of the ocean unmanned cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic flowchart of a method for ocean unmanned cluster attack - defense game decision - making based on the improved MADDPG algorithm of the present application; Figure 2 It is a schematic diagram of the selection principle of sliding state prediction data for a method for ocean unmanned cluster attack - defense game decision - making based on the improved MADDPG algorithm of the present application; Figure 3 It is a network structure diagram of a method for ocean unmanned cluster attack - defense game decision - making based on the improved MADDPG algorithm of the present application; Figure 4 It is a prediction flowchart of a method for ocean unmanned cluster attack - defense game decision - making based on the improved MADDPG algorithm of the present application; Figure 5 It is an adversarial scenario diagram of a method for ocean unmanned cluster attack - defense game decision - making based on the improved MADDPG algorithm of the present application; Figure 6 It is a map distribution of a method for ocean unmanned cluster attack - defense game decision - making based on the improved MADDPG algorithm of the present application; Figure 7 It is a comparison diagram of the average reward value of an embodiment of a method for ocean unmanned cluster attack - defense game decision - making based on the improved MADDPG algorithm of the present application and the traditional method; Figure 8A comparison chart of the reward values ​​of multiple aircraft in an embodiment of an unmanned marine swarm attack and defense game decision method based on an improved MADDPG algorithm of this application and a traditional method; Figure 9 This is a winning rate statistics chart of an embodiment of an ocean unmanned swarm attack and defense game decision-making method based on an improved MADDPG algorithm in this application.

[0021] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0023] The first embodiment of the present invention provides a marine unmanned cluster attack and defense game decision-making method based on an improved MADDPG algorithm, such as Figure 1 As shown, the specific steps include: Step S1, respectively obtaining the position states of all aircraft and targets at historical moments; wherein the position states of aircraft or targets at historical moments are: the position states of aircraft or targets at preset moments in the period of time when the aircraft detects the target, starting from the last moment; Step S2, predicting the predicted position state of each aircraft according to the position state of each aircraft at a historical moment; predicting the predicted position state of each target according to the position state of each target at a historical moment; It is worth noting that in the marine unmanned swarm attack and defense game scenario, the dynamic changes in the environment are not only subject to the internal mechanism of the state transfer of the aircraft itself, but also deeply affected by the change of the target state. This characteristic causes the attack and defense mobile strike decision system to be unable to strictly follow the Markov property. Therefore, this embodiment predicts the position dynamics of the aircraft and the target at the next moment. The prediction method is to use the position state of the aircraft at the historical moment as input and use the sliding state prediction model to obtain the predicted position state of the aircraft.

[0024] The method for determining the predicted position state of the target is the same as that for determining the predicted position state of the aircraft; for example, Figure 2 As shown, the number of consecutive preset moments forward can be 8. , or If within the time period when the vehicle detects a target, the number of samples from the last moment when the current vehicle detects the target to the first moment when the vehicle detects the target does not meet 8, then take the position states from the last moment forward to the position state detected at the very beginning. 。

[0025] Exemplarily, the sliding state prediction model is a polynomial fitting model based on the least squares principle. Taking the prediction of the vehicle as an example below, the prediction method will be specifically introduced.

[0026] Assume that the historical movement trajectory of the vehicle includes consecutive n coordinate points , , …, . Using the polynomial curve fitting method to fit the above coordinate points, the fitting formula is as follows:

[0027] In the formula, m is the order of the polynomial. By solving , the fitting curve can be obtained, and the solution method is as follows: Calculate the sum of the squares of the residuals Ɛ between the values of the fitted polynomial and the historical trajectory:

[0028] By determining the value of the parameter to make reach the minimum, let

[0029] Solve to obtain . The value of

[0030] Step S3, input the position state of each vehicle at the previous moment into the long short-term memory network to obtain the historical position state of each vehicle; it can be understood that the long short-term memory (LSTM) network is a recursive neural network (RNN) with excellent performance, which can learn and utilize long-term historical information. The LSTM network realizes the selective transmission of information through the cooperation of the forget gate, input gate, output gate and memory storage unit. This structure can effectively add or delete historical information in the memory storage unit. In this embodiment, the LSTM network is used to screen the position states of the vehicle at historical moments to obtain the historical position state of the vehicle. Specifically, input the position state of the vehicle at the previous moment and the output value of the LSTM network at the previous moment into the forget gate, and finally output the historical position state of the vehicle. Among them, initially, the output value at the previous moment is 0.

[0031] In this embodiment, by introducing the LSTM network into the MADDPG network, information can be filtered, the calculation of invalid information can be reduced, and thus the computational complexity of the MADDPG network and the complexity of the model can be lowered.

[0032] Step S4: Obtain the position state of each vehicle at the current moment, and input the historical position state, the position state at the current moment, and the predicted position state of each vehicle, as well as the predicted position state of the target, into the corresponding pre-trained MADDPG network to determine the execution actions of each vehicle and obtain the offensive and defensive game decision-making for the unmanned marine cluster. Among them, the execution actions include the movement speed and the yaw angular velocity.

[0033] It should be noted that the target in this embodiment is a moving target in the ocean. The number of targets can be one or multiple. Therefore, it is necessary to predict the predicted position states of all targets. The input of each MADDPG network includes the predicted position states of all targets. In addition, each vehicle corresponds to a pre-trained MADDPG network; the MADDPG network includes a policy network, a target policy network, Q network, and a target Q network; assume that there are k vehicles in the scenario, then and are respectively k the parameter sets of the policy network and the target policy network of and are respectively k the Q network and the target Q network parameter sets of

[0034] Specifically, as Figure 3 shown, the training process of the pre-trained MADDPG network includes: Step S41: As Figure 4 shown, according to the offensive and defensive interaction environment, obtain the position states of the current vehicle and the target at the historical moment respectively; Step S42: According to the position state of the current vehicle at the historical moment, predict the predicted position state of the current vehicle; according to the position state of each target at the historical moment, predict the predicted position state of each target; input the position state of the current vehicle at the previous moment into the long short-term memory network to obtain the historical position state of the current vehicle; Step S43: Input the historical position state i of the current vehicle , the predicted position state of the current vehicle, and the predicted position states , input the policy network to obtain the execution action of the current vehicle; the formula is as follows:

[0035] In the formula, vehicle i the action to be taken, that is, the execution action, is a deterministic policy, are the parameters of the policy network, is random noise.

[0036] Step S44, determine the position state of the current vehicle at the next moment according to the execution action of the current vehicle and the reward value ; In this embodiment, it focuses on the decision-making problem of unmanned marine clusters in the attack-defense game. The core lies in constructing a reasonable reward function system, aiming to guide the vehicle to effectively strike the target while avoiding target attacks. Among them, the reward value includes the reward value of the target , the self-protection reward value and the environmental adaptability reward value , and the calculation formula is: .

[0037] Specifically, in step S441, construct the kinematic model of the vehicle; according to the execution action of the current vehicle, combine the kinematic model to determine the position state of the current vehicle at the next moment; that is, input the movement speed and yaw angular velocity of the current vehicle into the kinematic model, and the position state of the current vehicle at the next moment can be obtained.

[0038] Among them, the kinematic model of the vehicle on the horizontal plane is:

[0039] In the formula, is the first derivative of the position coordinate of the vehicle in the ground coordinate system, is the first derivative of the heading angle; , , r are the longitudinal speed, lateral speed and yaw angular velocity of the vehicle in its own vehicle coordinate system respectively. The movement speed v The calculation formula of is:

[0040] Since the movement of the vehicle is restricted by dynamics, its movement speed v and yaw angular velocity r are within a limited range. In this embodiment, to make the movement of the vehicle more in line with the actual situation, vThe range is set between 0 and 8 sections. r The range is between -6° / s and 6° / s.

[0041] Step S442: Determine the strike success rate of the current vehicle and the strike success rate of the target against the current vehicle according to the position state of the current vehicle at the next moment and the predicted position state of the target. It can be understood that the strike success rate of the target against the current vehicle is the probability that the current vehicle is successfully struck by the target. The method for determining the strike success rate of the current vehicle and the strike success rate of the target against the current vehicle is the same. The following takes the strike success rate of the current vehicle as an example for specific introduction.

[0042] Such as Figure 5 shown, obtain the strike range length of the current vehicle and the strikeable range angle , the strike range length of the target and the strikeable range angle , the straight-line distance between the vehicle and the target L , the angle between the velocity vector of the current vehicle and the line connecting to the target α , the angle between the velocity vector of the target and the line connecting to the current vehicle β ; When determining that the strike range length of the current vehicle is less than the straight-line distance between the current vehicle and the target L , the strike success rate of the current vehicle is 0; When determining that the strike range length of the current vehicle is greater than or equal to the straight-line distance between the current vehicle and the target L , and the angle between the velocity vector of the current vehicle and the line connecting to the target α is less than or equal to the strikeable range angle of the current vehicle , the strike success rate of the current vehicle is 1.

[0043] Step S443: Determine the reward value of the target according to the strike success rate of the current vehicle; determine the self-protection reward value according to the strike success rate of the target.

[0044] Specifically, reward values are set for both the individual and the entire team. In terms of safety considerations, the greater the distance between the target and the vehicle, the higher the safety of the vehicle; however, to pursue higher rewards, the vehicle needs to actively initiate attacks. Therefore, the reward value of the target is determined in the following manner:

[0045]

[0046]

[0047] Among them, is the reward value for the target, is the self - reward value, is the team reward value.

[0048] Preferably, Take 100, Take 10. The hit success rate of the current vehicle is the probability that the current vehicle hits the target successfully, and the hit success rate of other vehicles is the probability that other vehicles in the team hit the target successfully.

[0049] Since when designing the reward function, it is necessary to ensure that the reward for self - protection is not higher than the reward for hitting the target. At the same time, the penalty for being hit also needs to be set within a reasonable range to avoid overly passive avoidance of hitting the target. Therefore, the self - protection reward value is determined according to the following method:

[0050] Among them, l is the number of targets, is the vehicle i 's position coordinates, is the target j 's position coordinates, is the self - protection reward value. The hit success rate of the target is the probability that the target hits the vehicle successfully. Preferably, Take - 100.

[0051] The method for determining the environmental adaptability reward value includes: Compare the position state of the current vehicle at the next moment with the position states of other vehicles at the next moment, the predicted position state of the target, the positions of obstacles, and the preset area (i.e., the map range) to determine the environmental adaptability reward value. Specifically, as Figure 6 shown, when the position state of the current vehicle at the next moment coincides with the predicted position state of the target or the position of the obstacle, it is determined that a collision occurs. When it is determined that it exceeds the preset area or a collision occurs, a relatively high negative reward value is set in order to achieve the effect of early warning and avoidance, specifically as follows:

[0052] Preferably, Take - 1000.

[0053] Step S45, obtain the position state of the current vehicle at the current moment, and store the position state of the current vehicle at the current moment, the position state at the next moment, the executed action, and the reward value into the experience replay pool; Among them, the experience replay pool contains multiple groups of experience values. Each group of experience values includes the position state, execution action, position state at the next moment, and reward value of all vehicles corresponding to each moment, that is , the MADDPG networks of multiple vehicles share an experience replay pool; Step S46, randomly select multiple groups of experience values in the experience replay pool to update the parameters of the MADDPG network, and obtain the updated MADDPG network; Specifically, randomly select q groups of experience values , and q in the groups of experience values, the reward value of the current vehicle at the current moment, the position states of all vehicles at the next moment , and the target policies of all vehicles to determine the target i th vehicle's target Q value ;

[0054] In the formula, γ is the discount factor, is the target Q network function, is the target policy of the i th vehicle; Using the position state of the current vehicle at the current moment, the execution actions of all vehicles at the current moment, and the target Q value in multiple groups of experience values to update the Q network parameters; specifically, the minimum loss function can be used to obtain the Q network parameters, and the formula is as follows:

[0055] In the formula, is Q network function.

[0056] Using the position state of the current vehicle at the current moment, and the execution actions of all vehicles at the current moment in multiple groups of experience values to update the policy network parameters corresponding to the current vehicle; specifically, use gradient descent to update the policy network parameters:

[0057] In the formula, is the gradient operator of the subscript variable.

[0058] Use soft update to update the target policy network and target Q network parameters:

[0059]

[0060] In the formula, τ is the inertial update rate, are the network parameters of the updated target policy network, is the updated target Q network's network parameters.

[0061] Step S47: Repeat steps S42 - S46 to iteratively update the MADDPG network until the MADDPG network converges, obtaining the pre - trained MADDPG network.

[0062] In this embodiment, on the basis of the traditional MADDPG network, a sliding prediction model and an LSTM network are introduced. The position state of the historical moment of the vehicle is predicted through the sliding prediction model, and the position state of the historical moment is processed by the LSTM network to obtain an improved MADDPG network. Using the improved MADDPG network can improve the winning rate of the unmanned marine cluster. Specifically, through the learning ability of environmental changes, the vehicle can pre - insight and predict the position state information of the target at future moments, so that when the offensive - defensive game decision - making environment changes, it can generate execution actions matching the changed offensive - defensive game decision - making environment, improving the winning rate of the vehicle; at the same time, as the vehicle continues to move, the collected historical trajectory points are updated in real - time, and the newly collected data points will gradually replace and cover the old sampling points, ensuring that the prediction model always analyzes based on the latest and most relevant information, further ensuring the winning rate of the vehicle; introducing the LSTM network can make full use of its accumulated historical data, and then achieve more accurate and efficient decision - making, improving the overall task execution efficiency; in the design of the reward value, not only the overall benefit of the vehicle team is considered, but also the individual benefits of each vehicle are considered, effectively avoiding the situation where some vehicles do not work and only focus on the overall benefit of the team.

[0063] Next, a specific example is used to introduce the decision - making method of the present invention.

[0064] In the experimental design, the number of vehicles and targets is both set to 3, the number of obstacles is set to 4, the attack range lengths and of the vehicle and the target are both 500 meters, the attack - able range angles and are both 60°, and the area range is set to 20 km ×20 km。The learning rate parameter is configured to 0.01, and the discount factor is set to 0.9 to ensure the importance of long-term rewards. In addition, the maximum storage capacity of the experience replay pool is set to 100,000. In the simulated strike scenario, the historical position states and predicted position states of each vehicle, as well as the predicted position states of all targets, are obtained according to steps S1 - S3, and are input into the pre-trained MADDPG network to obtain the execution actions of each vehicle, that is, the offensive and defensive game decision-making of the unmanned marine cluster, including the movement speed and yaw angular velocity. The vehicles are controlled to conduct offensive and defensive games according to this offensive and defensive game decision-making of the unmanned marine cluster.

[0065] In the strike experiment, the vehicles adopt the above-mentioned offensive and defensive game decision-making of the unmanned marine cluster, that is, the improved MADDPG algorithm, while the targets follow a random strategy. After a total of 200,000 interaction rounds, the average reward values of both sides in every 100 rounds are statistically analyzed to quantify the performance. The vehicle obtains an average reward curve, that is, the curve of the average reward value of our side, as Figure 7 shown. The improved MADDPG algorithm shows significant convergence characteristics. Compared with the traditional MADDPG algorithm, the improved one shows an obvious improvement in the obtained reward value. This result indicates that the proposed improved MADDPG algorithm is superior in performance.

[0066] Figure 8 shows the reward curves obtained by each of the three vehicles. From Figure 8 Figs. (a)-(c), it can be seen that the fluctuations of the reward values generated by the algorithms before and after improvement on the three vehicles (i.e., our vehicle 1, our vehicle 2, and our vehicle 3) are not significant. The root cause of this phenomenon is that in the offensive and defensive stage, the decision-making focus of the vehicle is no longer simply on its own survival state or the maximization of individual rewards, but has shifted to ensuring that the overall reward value of the entire strike system reaches the optimal level. Specifically, the strategy formulation of the vehicle focuses more on striking the target and effectively completing the offensive and defensive game tasks, so as to maximize the overall strike efficiency.

[0067] A win rate statistical strategy is adopted to statistically analyze the win rate after implementing the decision-making method of this embodiment. This strategy specifically calculates the win rate after 100 consecutive strike rounds, and a total of 200 rounds of simulation experiments are carried out, and a win rate change graph is drawn as Figure 9 shown. In this embodiment, the win rate is divided into three cases: the vehicle win rate, that is, the proportion of the number of rounds in which the vehicle successfully strikes the target in the total number of rounds within the given simulation steps; the target win rate, that is, the proportion of the number of rounds in which the target successfully strikes all vehicles under the same conditions; and the draw situation, that is, the proportion of the number of rounds in which neither side is completely struck within the simulation steps.

[0068] Figure 9In (a), it shows that when the traditional MADDPG algorithm is adopted, the winning rate of the vehicle (i.e., our unmanned marine cluster) is about 60%. Figure 9 In (b), it shows that when the improved MADDPG algorithm is adopted, the winning rate of the vehicle increases to about 80%. The results show that the adoption of the improved MADDPG algorithm can significantly improve the winning rate, showing significant superiority.

[0069] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A marine unmanned swarm attack and defense game decision-making method based on the improved MADDPG algorithm, characterized in that: include: The position states of all aircraft at historical moments and the position states of the targets at historical moments are obtained respectively; wherein the position states of the aircraft or targets at historical moments are: the position states of the aircraft or targets at preset moments in succession, starting from the last moment, within the time period when the aircraft detects the target; According to the position state of each of the aircraft at a historical moment, a predicted position state of each aircraft is predicted; according to the position state of each of the targets at a historical moment, a predicted position state of each target is predicted; Inputting the position state of each of the aircraft at the previous moment into the long short-term memory network to obtain the historical position state of each aircraft; Obtain the current position status of each aircraft, input the historical position status, current position status and predicted position status of each aircraft, and the predicted position status of the target into the corresponding pre-trained MADDPG network, determine the execution action of each aircraft, and obtain the attack and defense game decision of the marine unmanned swarm; In the training process of the pre-trained MADDPG network, the reward value includes the target reward value, the self-protection reward value and the environmental adaptability reward value, and the method for determining the reward value includes: Construct a kinematic model of the aircraft; determine the position state of the current aircraft at the next moment based on the current aircraft's execution action and the kinematic model; Determine the strike success rate of the current aircraft and the strike success rate of the target against the current aircraft based on the position state of the current aircraft at the next moment and the predicted position state of the target; Determining a reward value of a target according to the strike success rate of the current aircraft; The self-protection reward value is determined based on the success rate of hitting the target.

2. The marine unmanned swarm attack and defense game decision-making method based on the improved MADDPG algorithm according to claim 1 is characterized in that: The reward value of the target is determined as follows: in, is the reward value of the target, is its own reward value, Reward value for the team.

3. The marine unmanned swarm attack and defense game decision-making method based on the improved MADDPG algorithm according to claim 2 is characterized in that: The self-protection reward value is determined as follows: in, l is the number of targets, is the position coordinate of spacecraft i, is the position coordinate of target j, Reward value for self-protection.

4. The marine unmanned swarm attack and defense game decision-making method based on the improved MADDPG algorithm according to claim 3 is characterized in that: The method for determining the environmental adaptability reward value includes: The position state of the current aircraft at the next moment is compared with the position state of other aircraft at the next moment, the predicted position state of the target, the position of obstacles, and the preset area to determine the environmental adaptability reward value.

5. The marine unmanned swarm attack and defense game decision-making method based on the improved MADDPG algorithm according to claim 2 is characterized in that: The method for determining the strike success rate of the current aircraft includes: Obtain the strike range length, strike range angle, and straight-line distance between the current aircraft and the target; When it is determined that the strike range of the current aircraft is less than the straight-line distance between the current aircraft and the target, the strike success rate of the current aircraft is 0; When it is determined that the strike range length of the current aircraft is greater than or equal to the straight-line distance between the current aircraft and the target, and the angle between the current aircraft velocity vector and the target line is less than the strike range angle of the current aircraft, the strike success rate of the current aircraft is 1.

6. The marine unmanned swarm attack and defense game decision-making method based on the improved MADDPG algorithm according to claim 1 is characterized in that: The step of predicting the predicted position state of each aircraft according to the position state of each aircraft at a historical moment includes: The position state of the aircraft at a historical moment is taken as input, and the predicted position state of the aircraft is obtained by using a sliding state prediction model.

7. The marine unmanned swarm attack and defense game decision-making method based on the improved MADDPG algorithm according to claim 6 is characterized in that: The sliding state prediction model is a polynomial fitting model based on the least squares principle.

8. The marine unmanned swarm attack and defense game decision-making method based on the improved MADDPG algorithm according to claim 6 is characterized in that: The predicted position state of the target is determined in the same manner as the predicted position state of the aircraft.

9. The marine unmanned swarm attack and defense game decision-making method based on the improved MADDPG algorithm according to claim 1 is characterized in that: Each of the aircraft corresponds to a pre-trained MADDPG network; The training process of the pre-trained MADDPG network includes: Get the historical position status of the current spacecraft and the historical position status of all targets respectively; According to the position state of the current aircraft at a historical moment, predicting the predicted position state of the current aircraft; According to the position state of each target at a historical moment, a predicted position state of each target is predicted; Input the position state of the current spacecraft at the previous moment into the long short-term memory network to obtain the historical position state of the current spacecraft; Get the current position state of the current aircraft, input the historical position state, the current position state and the predicted position state of the current aircraft, and the predicted position state of all targets into the MADDPG network to obtain the execution action of the current aircraft; Determine the position state and reward value of the current aircraft at the next moment according to the execution action of the current aircraft; The current position state, the next position state, the executed action, and the reward value are stored in an experience replay pool; wherein the experience replay pool contains multiple sets of experience values, each set of experience values ​​includes the current position state, the executed action, the next position state, and the reward value corresponding to all aircraft at each moment; Randomly select several groups of experience values ​​in the experience replay pool to update the parameters of the MADDPG network to obtain an updated MADDPG network; Repeat the above process and iteratively update the MADDPG network until the MADDPG network converges to obtain the pre-trained MADDPG network.

10. The marine unmanned swarm attack and defense game decision-making method based on the improved MADDPG algorithm according to claim 9 is characterized in that: The MADDPG network includes a policy network, a target policy network, Q Network and Target Q Network; the randomly selected sets of experience values ​​in the experience replay pool are used to update the parameters of the MADDPG network, including: Randomly select several sets of experience values ​​from the experience replay pool, and use the reward value of the current aircraft at the current moment and the position status of all aircraft at the next moment in the several sets of experience values ​​to determine the target of the current aircraft Q value; Using several sets of experience values, the current position of the current aircraft, the execution actions of all aircraft at the current moment, and the target Q Value, Update Q Network parameters; Using the current position state of the current aircraft and the execution actions of all aircraft at the current moment in several sets of experience values, update the policy network parameters; Use soft update to update target policy network and target Q Network parameters.

Citation Information

Patent Citations

  • Multi-AUV dynamic maneuvering decision-making method based on interval information game

    CN112306070A

  • Unmanned cluster task collaboration method based on multi-agent reinforcement learning

    CN113589842A

  • Unmanned ship cluster task scheduling and collaborative confrontation method based on MADDPG

    CN116050795A

Cited By

  • Multi-unmanned ship cooperative hunting training system and training method

    CN120972936A

  • A multi-unmanned ship cooperative hunting training system and a training method

    CN120972936B