An Offensive-Defensive Game Decision-Making Method for Unmanned Marine Swarms Based on Improved MADDPG Algorithm
Through the improved MADDPG algorithm, sliding state prediction and long-term memory network are used to optimize offensive and defensive game decisions for marine unmanned clusters, solving the problem of low winning rate of the vehicle and achieving more efficient decision-making and task execution.
Patent Information
- Application Number
- CN202510546607.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The winning rate of the voyage in the offensive and defensive game is relatively low, making it difficult to cope with complex game strategies and changes in dynamic environments, resulting in insufficient adaptability to decisions and slow response speed.
The improved MADDPG algorithm is used to obtain the historical position status of the aircraft and the target, use the sliding state prediction model and long-term memory network to predict future position status, and combine the target's reward value, self-protection reward value and environmental adaptive reward value to optimize the decision process.
It improves the winning rate of the aircraft, ensures that decisions are based on the latest information, improves the task execution efficiency, avoids individuals from deviating from team goals, and achieves more efficient offensive and defensive game decisions.
Smart Images

Figure CN120124675B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of marine unmanned cluster game decision-making technology, and particularly to a marine unmanned cluster attack and defense game decision-making method based on an improved MADDPG algorithm. Background Art
[0002] In traditional marine unmanned cluster systems, decision-making usually relies on preset rules or a central control unit. In the face of a dynamically changing attack environment, this approach often exhibits problems such as insufficient adaptability and slow response speed. Especially in attack and defense games, the behavior of the target is highly uncertain, and traditional decision-making methods are difficult to handle complex game strategies. Therefore, there is an urgent need for a marine unmanned cluster attack and defense game decision-making method that can adapt to complex marine environments and has efficient coordination capabilities to improve the autonomous decision-making ability and task execution efficiency of the cluster in attack scenarios.
[0003] The Multi-Agent Deep Deterministic Policy Gradient (hereinafter referred to as MADDPG) algorithm has shown significant advantages in marine unmanned cluster attack and defense game decision-making. The MADDPG algorithm is particularly suitable for multi-agent environments, where the behavior of each agent depends not only on the state of the environment but also on the strategies of other agents. This characteristic enables the MADDPG algorithm to handle the complex interactions between multiple agents in marine unmanned cluster attack and defense game decision-making well. However, in order to attack each other, both sides will continuously learn new strategies and find suitable actions to execute. As time goes by, the number of both sides will keep changing, resulting in a change in the attack and defense game decision-making environment. This will cause the attack and defense game decision-making system of the marine unmanned cluster not to satisfy the Markov property, and the historical information is not fully utilized, leading to the problem of a low winning rate of the vehicle. Summary of the Invention
[0004] The main purpose of this application is to provide a marine unmanned cluster attack and defense game decision-making method based on an improved MADDPG algorithm, aiming to solve the problem of the low winning rate of the vehicle existing in the existing marine unmanned cluster attack and defense game decision-making method.
[0005] To achieve the above object, the present application provides a method for ocean unmanned cluster attack and defense game decision-making based on an improved MADDPG algorithm, including: respectively obtaining the position states of all vehicles at historical moments and the position states of targets at historical moments; wherein, the position state of a vehicle or a target at a historical moment is: within the time period when the vehicle detects the target, starting from the last moment, the position states of the vehicle or the target for a continuous preset number of moments forward; predicting the predicted position state of each vehicle according to the position state of each vehicle at historical moments; predicting the predicted position state of each target according to the position state of each target at historical moments; inputting the position state of each vehicle at the previous moment into a long short-term memory network to obtain the historical position state of each vehicle; obtaining the position state of each vehicle at the current moment, and inputting the historical position state, the position state at the current moment, the predicted position state of each vehicle, and the predicted position state of the target into the corresponding pre-trained MADDPG network to determine the execution action of each vehicle, and obtaining the ocean unmanned cluster attack and defense game decision; wherein, during the training process of the pre-trained MADDPG network, the reward value includes the reward value of the target, the self-protection reward value, and the environmental adaptability reward value; the method for determining the reward value includes: constructing a kinematic model of the vehicle; determining the position state of the current vehicle at the next moment according to the execution action of the current vehicle in combination with the kinematic model; determining the strike success rate of the current vehicle and the strike success rate of the target for the current vehicle according to the position state of the current vehicle at the next moment and the predicted position state of the target; determining the reward value of the target according to the strike success rate of the current vehicle; determining the self-protection reward value according to the strike success rate of the target.
[0006] Optionally, the reward value of the target is determined according to the following method:
[0007]
[0008]
[0009]
[0010] Wherein, is the reward value of the target, is the self-reward value, is the team reward value.
[0011] Optionally, the self-protection reward value is determined according to the following method:
[0012]
[0013] Wherein, l is the number of targets, is the position coordinate of vehicle i, is the position coordinate of target j. is the self - protection reward value.
[0014] Optionally, the method for determining the environmental adaptability reward value includes: comparing the position state of the current vehicle at the next moment with the position states of other vehicles at the next moment, the predicted position state of the target, the positions of obstacles, and a preset area to determine the environmental adaptability reward value.
[0015] Optionally, the method for determining the strike success rate of the current vehicle includes: obtaining the strike range length, the strikeable range angle of the current vehicle, and its straight - line distance from the target; when it is determined that the strike range length of the current vehicle is less than the straight - line distance between the current vehicle and the target, the strike success rate of the current vehicle is 0; when it is determined that the strike range length of the current vehicle is greater than or equal to the straight - line distance between the current vehicle and the target, and the angle between the velocity vector of the current vehicle and the line connecting to the target is less than the strikeable range angle of the current vehicle, the strike success rate of the current vehicle is 1.
[0016] Optionally, predicting the predicted position state of each vehicle based on the position state of each vehicle at historical moments includes: using the position state of the vehicle at historical moments as input and obtaining the predicted position state of the vehicle by using a sliding - state prediction model.
[0017] Optionally, the sliding - state prediction model is a polynomial fitting model based on the least - squares principle.
[0018] Optionally, the method for determining the predicted position state of the target is the same as that of the vehicle.
[0019] Optionally, each vehicle corresponds to a pre-trained MADDPG network; the training process of the pre-trained MADDPG network includes: respectively obtaining the position states of the current vehicle at historical moments and the position states of all targets at historical moments; predicting the predicted position state of the current vehicle according to the position state of the current vehicle at historical moments; predicting the predicted position states of each target according to the position states of each target at historical moments; inputting the position state of the current vehicle at the previous moment into a long short-term memory network to obtain the historical position state of the current vehicle; obtaining the position state of the current vehicle at the current moment, and inputting the historical position state, the position state at the current moment, the predicted position state of the current vehicle, and the predicted position states of all targets into the MADDPG network to obtain the execution action of the current vehicle; determining the position state and reward value of the current vehicle at the next moment according to the execution action of the current vehicle; storing the position state, the position state at the next moment, the execution action, and the reward value at the current moment into an experience replay pool; wherein, the experience replay pool contains multiple groups of experience values, and each group of experience values includes the position state, the execution action, the position state at the next moment, and the reward value corresponding to all vehicles at each moment; randomly selecting several groups of experience values from the experience replay pool to update the parameters of the MADDPG network to obtain an updated MADDPG network; repeating the above process to iteratively update the MADDPG network until the MADDPG network converges to obtain a pre-trained MADDPG network.
[0020] Optionally, the MADDPG network includes a policy network, a target policy network, Q network, and a target Q network; randomly selecting several groups of experience values from the experience replay pool to update the parameters of the MADDPG network includes: randomly selecting several groups of experience values from the experience replay pool, and determining the target Q value of the current vehicle by using the reward value of the current vehicle at the current moment and the position states of all vehicles at the next moment in the several groups of experience values; using the position state of the current vehicle at the current moment, the execution actions of all vehicles at the current moment, and the target Q value in the several groups of experience values to update the Q network parameters; using the position state of the current vehicle at the current moment and the execution actions of all vehicles at the current moment in the several groups of experience values to update the policy network parameters; using a soft update method to update the parameters of the target policy network and the target Q network.
[0021] Compared with the prior art, the beneficial effects of the present application are as follows:
[0022] The ocean unmanned cluster attack-defense game decision-making method based on the improved MADDPG algorithm of the present invention can predict the predicted position state of the target according to the position state of each target at the historical moment, enabling the vehicle to pre-insight and predict the position state information of the target at the future moment through the learning ability of environmental changes, so as to generate execution actions matching the changed attack-defense game decision-making environment when the attack-defense game decision-making environment changes, improving the winning rate of the vehicle; further, starting from the last moment when the vehicle detects the target, the position states of the vehicle or the target at a continuously preset moment forward are used as the position states at the historical moment, realizing real-time update of the collected position states at the historical moment as the vehicle continues to move, gradually replacing and covering the old sampling points, ensuring that the prediction is always analyzed based on the latest and most relevant information, and further ensuring the winning rate of the vehicle; using the long short-term memory network to process the position states at the historical moment can make full use of the accumulated historical data, and then realize more accurate and efficient decision-making, improving the overall task execution efficiency; the reward value includes the reward value of the target, the self-protection reward value and the environmental adaptability reward value, considering both the overall benefit of the vehicle team and the individual benefits of each vehicle, effectively avoiding the situation where some vehicles do not work and only focus on the overall benefit of the team, and further improving the winning rate of the ocean unmanned cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 FIG. is a schematic flow chart of a method for ocean unmanned cluster attack-defense game decision-making based on the improved MADDPG algorithm of the present application;
[0024] Figure 2 FIG. is a schematic diagram of the selection principle of sliding state prediction data for a method for ocean unmanned cluster attack-defense game decision-making based on the improved MADDPG algorithm of the present application;
[0025] Figure 3 FIG. is a network structure diagram of a method for ocean unmanned cluster attack-defense game decision-making based on the improved MADDPG algorithm of the present application;
[0026] Figure 4 FIG. is a prediction flow chart of a method for ocean unmanned cluster attack-defense game decision-making based on the improved MADDPG algorithm of the present application;
[0027] Figure 5 FIG. is an adversarial scenario diagram of a method for ocean unmanned cluster attack-defense game decision-making based on the improved MADDPG algorithm of the present application;
[0028] Figure 6 FIG. is a map distribution of a method for ocean unmanned cluster attack-defense game decision-making based on the improved MADDPG algorithm of the present application;
[0029] Figure 7A comparison graph of the average reward value of an embodiment of a method for ocean unmanned cluster attack and defense game decision-making based on an improved MADDPG algorithm in this application and a traditional method;
[0030] Figure 8 A comparison graph of the reward values of multiple vehicles in an embodiment of a method for ocean unmanned cluster attack and defense game decision-making based on an improved MADDPG algorithm in this application and a traditional method;
[0031] Figure 9 A statistical graph of the winning rate of an embodiment of a method for ocean unmanned cluster attack and defense game decision-making based on an improved MADDPG algorithm in this application.
[0032] The realization, functional features and advantages of the purpose of this application will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments
[0033] To make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below with reference to the accompanying drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts belong to the scope of protection of this application.
[0034] The first embodiment of the present invention provides a method for ocean unmanned cluster attack and defense game decision-making based on an improved MADDPG algorithm, as Figure 1 shown, and specifically includes the following steps:
[0035] Step S1, respectively obtain the position states of all vehicles and targets at historical moments; wherein, the position state of a vehicle or target at a historical moment is: within the time period when the vehicle detects the target, starting from the last moment, the position states of the vehicle or target for a continuous preset number of moments forward;
[0036] Step S2, based on the position states of each vehicle at historical moments, predict the predicted position state of each vehicle; based on the position states of each target at historical moments, predict the predicted position state of each target.
[0037] It should be noted that in the scenario of ocean unmanned cluster attack and defense game, the dynamic changes of the environment are not only restricted by the internal mechanism of the state transition of the vehicle itself, but also deeply affected by the changes in the target state. This characteristic causes the attack and defense maneuver strike decision-making system to not strictly follow the Markov property. Therefore, in this embodiment, the position dynamics of the vehicle and the target at the next moment are predicted. The prediction method is to use the position state of the vehicle at historical moments as the input and obtain the predicted position state of the vehicle using a sliding state prediction model.
[0038] The method for determining the predicted position state of the target is the same as that of the vehicle; for example, Figure 2 as shown, the continuous preset moments forward can be 8, , or . If the number of samples does not meet 8 from the last moment when the vehicle detects the target to the first moment when the vehicle detects the target during the period when the vehicle detects the target, then the position state detected from the last moment is taken forward to the position state detected at the very beginning. .
[0039] Exemplarily, the sliding state prediction model is a polynomial fitting model based on the least squares principle. Taking the prediction of the vehicle as an example, the prediction method will be specifically introduced below.
[0040] Assume that the historical movement trajectory of the vehicle includes consecutive n coordinate points , , …, . The above coordinate points are fitted using the polynomial curve fitting method, and the fitting formula is as follows:
[0041]
[0042] In the formula, m is the order of the polynomial. Solving for , the fitting curve can be obtained, and the solution method is as follows:
[0043] Calculate the sum of the squares of the residuals Ɛ between the values of the fitted polynomial and the historical trajectory:
[0044]
[0045] By determining the value of the parameter such that reaches the minimum, let
[0046]
[0047] Solve to obtain . The value of
[0048] Step S3: Input the position state of each vehicle at the previous moment into the long short-term memory network to obtain the historical position state of each vehicle. It can be understood that the long short-term memory (LSTM) network is a recurrent neural network (RNN) with excellent performance, capable of learning and utilizing long-term historical information. The LSTM network realizes the selective transmission of information through the cooperation of the forget gate, input gate, output gate, and memory storage unit. This structure can effectively add or delete historical information in the memory storage unit. In this embodiment, the LSTM network is used to screen the position state information of the vehicle at the historical moment to obtain the historical position state of the vehicle. Specifically, the position state of the vehicle at the previous moment and the output value of the LSTM network at the previous moment are input into the forget gate, and finally the historical position state of the vehicle is output. Among them, initially, the output value at the previous moment is 0.
[0049] In this embodiment, introducing the LSTM network into the MADDPG network can screen information, reduce the calculation of invalid information, and thus reduce the computational complexity of the MADDPG network and the complexity of the model.
[0050] Step S4: Obtain the position state of each vehicle at the current moment, and input the historical position state, the position state at the current moment, the predicted position state of each vehicle, and the predicted position state of the target into the corresponding pre-trained MADDPG network to determine the execution action of each vehicle and obtain the offensive and defensive game decision of the unmanned marine cluster. Among them, the execution action includes the movement speed and the yaw angular velocity.
[0051] It should be noted that the target in this embodiment is a moving target in the ocean, and the number of targets can be one or more. Therefore, it is necessary to predict the predicted position states of all targets. The input of each MADDPG network includes the predicted position states of all targets. In addition, each vehicle corresponds to a pre-trained MADDPG network; the MADDPG network includes a policy network, a target policy network, Q network, and target Q network; assuming that there are k vehicles in the scenario, then and are respectively k the parameter sets of the policy network and the target policy network of vehicles, and k are respectively Q the network and target Q network parameter sets of
[0052] specifically, as Figure 3 shown, the training process of the pre-trained MADDPG network includes:
[0053] Step S41, as Figure 4 shown, according to the attack - defense interaction environment, obtain the position states of the current vehicle and the target at historical moments respectively;
[0054] Step S42, based on the position state of the current vehicle at the historical moment, predict the predicted position state of the current vehicle; based on the position states of each target at the historical moment, predict the predicted position states of each target; input the position state of the current vehicle at the previous moment into the long - short - term memory network to obtain the historical position state of the current vehicle;
[0055] Step S43, input the i historical position state of the current vehicle , the predicted position state of the current vehicle , and the predicted position states of all targets into the policy network to obtain the execution action of the current vehicle; the formula is as follows:
[0056]
[0057] In the formula, the action that the vehicle i should take, that is, the execution action, is a deterministic policy, is the parameter of the policy network, is the random noise.
[0058] Step S44, determine the position state and reward value of the current vehicle at the next moment according to the execution action of the current vehicle ; In this embodiment, it focuses on the decision - making problem of unmanned marine clusters in the attack - defense game. The core lies in constructing a reasonable reward function system, aiming to guide the vehicle to effectively strike the target while avoiding the attack of the target. Among them, the reward value includes the reward value of the target , the self - protection reward value and the environmental adaptability reward value , and the calculation formula is: .
[0059] Specifically, in step S441, construct the kinematic model of the vehicle; according to the execution action of the current vehicle, combine the kinematic model to determine the position state of the current vehicle at the next moment; that is, input the movement speed and yaw angular velocity of the current vehicle into the kinematic model, and the position state of the current vehicle at the next moment can be obtained.
[0060] Among them, the kinematic model of the vehicle on the horizontal plane is:
[0061]
[0062] In the formula, is the first derivative of the position coordinates of the vehicle in the ground coordinate system, is the first derivative of the heading angle; , , r are respectively the longitudinal speed, lateral speed and yaw angular velocity of the vehicle in its own vehicle coordinate system. The motion speed v The calculation formula of is:
[0063]
[0064] Since the motion of the vehicle is restricted by dynamics, its motion speed v and the yaw angular velocity r are within a limited range. In this embodiment, to make the motion of the vehicle more in line with the actual situation, v is set in the range of 0 to 8 knots, r is in the range of -6° / s to 6° / s.
[0065] Step S442, determine the strike success rate of the current vehicle and the strike success rate of the target for the current vehicle according to the position state of the current vehicle at the next moment and the predicted position state of the target;
[0066] It can be understood that the strike success rate of the target for the current vehicle, that is, the probability that the current vehicle is successfully struck by the target. The methods for determining the strike success rate of the current vehicle and the strike success rate of the target for the current vehicle are the same. The following takes the strike success rate of the current vehicle as an example for specific introduction.
[0067] As Figure 5 shown, obtain the strike range length and the strikeable range angle of the current vehicle, the strike range length and the strikeable range angle of the target, the straight-line distance L between the vehicle and the target, the angle α between the velocity vector of the current vehicle and the line connecting to the target, and the angle β between the velocity vector of the target and the line connecting to the current vehicle;
[0068] When it is determined that the strike range length of the current vehicle is less than the straight-line distance L between the current vehicle and the target, the strike success rate of the current vehicle is 0;
[0069] When it is determined that the strike range length Greater than or equal to the straight-line distance between the current vehicle and the target L , the angle between the current vehicle's velocity vector and the line connecting to the target α is less than or equal to the strike range angle of the current vehicle , and the strike success rate of the current vehicle is 1
[0070] Step S443: Determine the reward value of the target according to the strike success rate of the current vehicle; determine the self-protection reward value according to the strike success rate of the target
[0071] Specifically, reward values are set for the individual and the entire team respectively. In terms of safety considerations, the greater the distance between the target and the vehicle, the higher the safety of the vehicle; however, to pursue higher rewards, the vehicle needs to actively initiate attacks. Therefore, the reward value of the target is determined as follows
[0072]
[0073]
[0074]
[0075] Among them is the reward value of the target is the self-reward value is the team reward value
[0076] Preferably Take 100 Take 10. The strike success rate of the current vehicle is the probability of the current vehicle successfully striking the target, and the strike success rate of other vehicles is the probability of other vehicles in the team successfully striking the target
[0077] Since when designing the reward function, it is necessary to ensure that the reward for self-protection is not higher than the reward for striking the target. At the same time, the punishment for being struck also needs to be set within a reasonable range to avoid overly passive evasion of striking the target. Therefore, the self-protection reward value is determined as follows
[0078]
[0079] Among them l is the number of targets is the vehicle i 's position coordinates is the target j 's position coordinates is the self-protection reward value. The strike success rate of the target is the probability of the target successfully striking the vehicle. Preferably Take -100
[0080] The method for determining the environmental adaptability reward value includes:
[0081] Compare the position state of the current vehicle at the next moment with the position states of other vehicles at the next moment, the predicted position state of the target, the positions of obstacles, and the preset area (i.e., the map range) to determine the environmental adaptability reward value. Specifically, as Figure 6 shown, when the position state of the current vehicle at the next moment coincides with the predicted position state of the target or the position of the obstacle, it is determined that a collision has occurred. When it is determined that it exceeds the preset area or a collision occurs, a relatively high negative reward value is set in order to achieve the effect of warning and avoidance, specifically as follows:
[0082]
[0083] Preferably, Take -1000.
[0084] Step S45, obtain the position state of the current vehicle at the current moment, and store the position state, the position state at the next moment, the executed action, and the reward value of the current vehicle at the current moment into the experience replay pool;
[0085] Among them, the experience replay pool contains multiple groups of experience values. Each group of experience values includes the position state, the executed action, the position state at the next moment, and the reward value corresponding to all vehicles at each moment, that is , and the MADDPG networks of multiple vehicles share one experience replay pool;
[0086] Step S46, randomly select multiple groups of experience values from the experience replay pool to update the parameters of the MADDPG network, and obtain the updated MADDPG network;
[0087] Specifically, randomly select q groups of experience values , and use q the reward value of the current vehicle at the current moment, the position states of all vehicles at the next moment , and the target policies of all vehicles to determine the target i (the Q th) value of the current vehicle ;
[0088]
[0089] In the formula, γ is the discount factor, is the target Q network function, is the target policy of the i th vehicle;
[0090] Using the position state of the current vehicle at the current moment among multiple sets of empirical values, the execution actions of all vehicles at the current moment, and the target Q value, update Q network parameters; specifically, the minimum loss function can be used to obtain Q network parameters, and the formula is as follows:
[0091]
[0092] In the formula, is Q the network function.
[0093] Using the position state of the current vehicle at the current moment among multiple sets of empirical values, and the execution actions of all vehicles at the current moment, update the policy network parameters corresponding to the current vehicle; specifically, use gradient descent to update the policy network parameters:
[0094]
[0095] In the formula, is the gradient operator of the subscript variable.
[0096] Use the soft update method to update the target policy network and the target Q network parameters:
[0097]
[0098]
[0099] In the formula, τ is the inertia update rate, is the network parameter of the updated target policy network, is the updated target Q network parameter of the network.
[0100] Step S47. Repeat steps S42 - S46 to iteratively update the MADDPG network until the MADDPG network converges to obtain the pre-trained MADDPG network.
[0101] In this embodiment, based on the traditional MADDPG network, a sliding prediction model and an LSTM network are introduced. The position state of the vehicle at historical moments is predicted through the sliding prediction model, and the position state at historical moments is processed by the LSTM network to obtain an improved MADDPG network. Using the improved MADDPG network can increase the winning rate of the unmanned marine cluster. Specifically, through the learning ability of environmental changes, the vehicle can pre-insight and predict the position state information of the target at future moments, so that when the attack and defense game decision-making environment changes, an execution action matching the changed attack and defense game decision-making environment is generated, increasing the winning rate of the vehicle; at the same time, as the vehicle continues to move, the collected historical trajectory points are updated in real time, and the newly collected data points will gradually replace and cover the old sampling points, ensuring that the prediction model always analyzes based on the latest and most relevant information, further ensuring the winning rate of the vehicle; introducing the LSTM network can make full use of its accumulated historical data, and then achieve more accurate and efficient decision-making, improving the overall task execution efficiency; in the design of the reward value, not only the overall income of the vehicle team is considered, but also the individual income of each vehicle is considered, effectively avoiding the situation where some vehicles do not work and only focus on the overall income of the team.
[0102] The following introduces the decision-making method of the present invention with a specific example.
[0103] In the experimental design, the number of vehicles and targets is both set to 3, the number of obstacles is set to 4, and the attack range lengths of the vehicles and targets and are both 500 meters, and the attackable range angles and are both 60°, and the area range is set to 20 km ×20 km . The learning rate parameter is configured to 0.01, and the discount factor is set to 0.9 to ensure the importance of long-term benefits. In addition, the maximum storage capacity of the experience replay pool is set to 100000. In the simulated strike scenario, the historical position states and predicted position states of each vehicle, and the predicted position states of all targets are obtained according to steps S1 - S3, and are input into the pre-trained MADDPG network to obtain the execution actions of each vehicle, that is, the attack and defense game decision-making of the unmanned marine cluster, including the movement speed and yaw angular velocity, and the vehicles are controlled to conduct the attack and defense game according to the attack and defense game decision-making of the unmanned marine cluster.
[0104] In the combat experiment, the vehicle adopts the above-mentioned ocean unmanned cluster attack-defense game decision-making, that is, the improved MADDPG algorithm, while the target follows a random strategy. After a total of 200,000 interaction rounds, the average reward values of both sides within every 100 rounds are statistically calculated to quantify the performance. The vehicle obtains the average reward curve, that is, the curve of the average reward value of our side, as Figure 7 shown. The improved MADDPG algorithm exhibits significant convergence characteristics. Compared with the traditional MADDPG algorithm, the improved one shows an obvious improvement in the obtained reward value. This result indicates that the proposed improved MADDPG algorithm is superior in performance.
[0105] Figure 8 shows the reward curves obtained by each of the three vehicles. From Figure 8 figures (a)-(c), it can be seen that the fluctuations in the reward values generated by the algorithms before and after improvement on the three vehicles (i.e., our vehicle 1, our vehicle 2, and our vehicle 3) are not significant. The root cause of this phenomenon is that in the attack-defense stage, the decision-making focus of the vehicle is no longer simply on its own survival status or the maximization of individual rewards, but has shifted to ensuring that the overall reward value of the entire combat system reaches the optimal level. Specifically, the strategy formulation of the vehicle focuses more on attacking the target and effectively completing the attack-defense game task, so as to maximize the overall combat effectiveness.
[0106] Adopt a win rate statistical strategy to statistically calculate the win rate after implementing the decision-making method of this embodiment. This strategy is specifically to calculate the win rate after 100 consecutive combat rounds, and a total of 200 rounds of simulation experiments are carried out, and the win rate change diagram is drawn as Figure 9 shown. In this embodiment, the win rate is divided into three cases: the win rate of the vehicle, that is, within the given simulation steps, the proportion of the number of rounds in which the vehicle successfully attacks the target to the total number of rounds; the win rate of the target, that is, under the same conditions, the proportion of the number of rounds in which the target successfully attacks all vehicles; and the case of a draw, that is, the proportion of the number of rounds in which neither side is completely attacked within the simulation steps.
[0107] Figure 9 figure (a) shows that when the traditional MADDPG algorithm is adopted, the win rate of the vehicle (i.e., our ocean unmanned cluster) is about 60%; Figure 9 figure (b) shows that when the improved MADDPG algorithm is adopted, the win rate of the vehicle increases to about 80%. The results show that adopting the improved MADDPG algorithm can significantly improve the win rate, showing significant superiority.
[0108] The above are only the preferred embodiments of the present application, which do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.
Claims
1. An offensive and defensive game decision-making method for unmanned marine swarms based on an improved MADDPG algorithm, characterized in that Including: Obtain the position states of all vehicles at historical moments and the position states of the targets at historical moments respectively; wherein, the position state of a vehicle or a target at a historical moment is: within the time period when the vehicle detects the target, starting from the last moment, the position states of the vehicle or the target for a continuous preset number of moments forward; Based on the position states of each vehicle at historical moments, predict the predicted position states of each vehicle; based on the position states of each target at historical moments, predict the predicted position states of each target; Input the position state of each vehicle at the previous moment into a long short-term memory network to obtain the historical position state of each vehicle; Obtain the position state of each vehicle at the current moment, and input the historical position state, the position state at the current moment, and the predicted position state of each vehicle, as well as the predicted position state of the target, into the corresponding pre-trained MADDPG network to determine the execution actions of each vehicle and obtain the offensive and defensive game decision-making for the unmanned marine cluster; Wherein, during the training process of the pre-trained MADDPG network, the reward value includes the reward value of the target, the self-protection reward value, and the environmental adaptability reward value, and the method for determining the reward value includes: Construct a kinematic model of the vehicle; based on the execution action of the current vehicle, combine the kinematic model to determine the position state of the current vehicle at the next moment; Based on the position state of the current vehicle at the next moment and the predicted position state of the target, determine the strike success rate of the current vehicle and the strike success rate of the target against the current vehicle; Based on the strike success rate of the current vehicle, determine the reward value of the target; Based on the strike success rate of the target, determine the self-protection reward value.
2. The method for ocean unmanned cluster attack and defense game decision-making based on the improved MADDPG algorithm according to claim 1, characterized in that, The reward value of the target is determined according to the following method: Among them, is the reward value for the target, is the self-reward value, is the team reward value.
3. The method for ocean unmanned cluster attack and defense game decision-making based on the improved MADDPG algorithm according to claim 2, wherein The self-protection reward value is determined according to the following method: Among them, l is the number of targets, is the position coordinates of vehicle i, is the position coordinates of target j, is the self-protection reward value.
4. The method for ocean unmanned cluster attack and defense game decision-making based on the improved MADDPG algorithm according to claim 3, characterized in that The method for determining the environmental adaptability reward value includes: Compare the position state of the current vehicle at the next moment with the position states of other vehicles at the next moment, the predicted position state of the target, the position of the obstacle, and the preset area to determine the environmental adaptability reward value.
5. The method for ocean unmanned cluster attack and defense game decision-making based on the improved MADDPG algorithm according to claim 2, characterized in that The method for determining the strike success rate of the current vehicle includes: Obtain the strike range length, the strikeable range angle of the current vehicle, and its straight-line distance from the target; When it is determined that the strike range length of the current vehicle is less than the straight-line distance between the current vehicle and the target, the strike success rate of the current vehicle is 0; When it is determined that the strike range length of the current vehicle is greater than or equal to the straight-line distance between the current vehicle and the target, and the angle between the velocity vector of the current vehicle and the connection line of the target is less than the strikeable range angle of the current vehicle, the strike success rate of the current vehicle is 1.
6. The method for ocean unmanned cluster attack and defense game decision-making based on the improved MADDPG algorithm according to claim 1, characterized in that The predicting the predicted position state of each vehicle based on the position state of each vehicle at historical moments includes: Use the position state of the vehicle at historical moments as the input, and use a sliding state prediction model to obtain the predicted position state of the vehicle.
7. The method for ocean unmanned cluster attack and defense game decision-making based on the improved MADDPG algorithm according to claim 6, characterized in that, The sliding state prediction model is a polynomial fitting model based on the least squares principle.
8. The method for ocean unmanned cluster attack and defense game decision-making based on the improved MADDPG algorithm according to claim 6, characterized in that The method for determining the predicted position state of the target is the same as that of the vehicle.
9. The method for ocean unmanned cluster attack and defense game decision-making based on the improved MADDPG algorithm according to claim 1, characterized in that, Each of the said vehicles corresponds to a pre-trained MADDPG network; The training process of the said pre-trained MADDPG network includes: Obtain the position states of the current vehicle at historical moments and the position states of all targets at historical moments respectively; Predict the predicted position state of the current vehicle according to the position state of the current vehicle at historical moments; Predict the predicted position states of each target according to the position states of each target at historical moments; Input the position state of the current vehicle at the previous moment into the long short-term memory network to obtain the historical position state of the current vehicle; Obtain the position state of the current vehicle at the current moment, and input the historical position state, the position state at the current moment, the predicted position state of the current vehicle, and the predicted position states of all targets into the MADDPG network to obtain the execution action of the current vehicle; Determine the position state and reward value of the current vehicle at the next moment according to the execution action of the current vehicle; Store the position state, the position state at the next moment, the execution action, and the reward value at the current moment into the experience replay pool; wherein, the experience replay pool contains multiple groups of experience values, and each group of experience values includes the position state, the execution action, the position state at the next moment, and the reward value corresponding to all vehicles at each moment; Randomly select several groups of experience values in the experience replay pool to update the parameters of the MADDPG network, and obtain the updated MADDPG network; Repeat the above process to iteratively update the MADDPG network until the MADDPG network converges to obtain the pre-trained MADDPG network.
10. The method for ocean unmanned cluster attack and defense game decision-making based on the improved MADDPG algorithm according to claim 9, characterized in that The MADDPG network includes a policy network, a target policy network, Q network, and a target Q network; updating the parameters of the MADDPG network by using several groups of experience values randomly selected from the experience replay pool, including: Randomly select several groups of experience values from the experience replay pool, and use the reward value of the current vehicle at the current moment and the position states of all vehicles at the next moment among the several groups of experience values to determine the target of the current vehicle Q value; Using the position state of the current vehicle at the current moment among several sets of empirical values, the execution actions of all vehicles at the current moment, and the target Q value, update Q network parameters; Update the policy network parameters by using the position state of the current vehicle at the current moment and the execution actions of all vehicles at the current moment among several groups of experience values; Update the target policy network and target Q network parameters using soft updates.
Citation Information
Patent Citations
Multi-AUV dynamic maneuvering decision-making method based on interval information game
CN112306070A
Unmanned cluster task collaboration method based on multi-agent reinforcement learning
CN113589842A