Multi-unmanned ship game confrontation method and system based on DRMRPG algorithm
By introducing multi-level experience playback strategies, dynamic soft update strategies and residual connection strategies to the MADDPG algorithm, an intelligent game framework is built, which solves the problems of difficulty convergence and poor stability of multi-unmanned boat game confrontation algorithms in the existing technology, and achieves efficient and stable multi-unmanned boat game confrontation effect.
Patent Information
- Application Number
- CN202510215765.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The existing multi-agent deep reinforcement learning algorithms are difficult to converge when facing high-dimensional state space and complex observation environments, and have poor training stability, low sampling efficiency, and lack effective simulation platforms, which limits the research and application of unmanned boat cluster game confrontation.
A multi-unmanned boat game confrontation method based on DRMRPG algorithm is proposed. By introducing multi-level experience replay strategy, dynamic soft update strategy and residual connection strategy based on the MADDPG algorithm, an intelligent game framework is built, and an intelligent decision-making model is trained to carry out multi-unmanned boat game confrontation.
The cumulative rewards of the agent are improved, sampling efficiency and stability are improved, the convergence performance of the algorithm is improved, and the multi-unmanned boat game confrontation effect with high reward value, high stability and high efficiency is achieved.
Smart Images

Figure CN120047006A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-agent game confrontation, and particularly relates to a multi-unmanned boat game confrontation method and system based on the DRMRPG algorithm. Background Technique
[0002] Unmanned boats have the ability to perform tasks all-weather, especially they can replace humans to perform dangerous, time-consuming and laborious operation tasks in harsh marine environments, and have broad application prospects in both military and civil fields. Intelligent decision-making is an important supporting technology for the application of unmanned boats and can be applied to unmanned boats to replace human drivers to perform dangerous tasks. Deep reinforcement learning combines the advantages of deep learning in perception and understanding and the excellent performance of reinforcement learning in decision-making to achieve precise decision-making of unmanned boats for real-time situations in different environments.
[0003] On the other hand, as an important branch of artificial intelligence technology, reinforcement learning currently has important application value in multi-agent game confrontation problems such as unmanned boats and drones. In 2021, Li Bo et al. proposed an improved multi-agent deep deterministic policy gradient algorithm, which was applied to the collaborative task research of multi-drones and could solve simple task decision-making problems. In 2023, Liu Jing et al. proposed a method for collaborative encirclement and capture of drone swarms by combining game theory and Q-Learning, and the results showed that this method could complete the effective encirclement and capture of a single target. In 2022, Zhan proposed a multi-agent proximal policy optimization algorithm for realizing distributed decision-making and collaborative task completion of heterogeneous drones. In contrast, there is relatively little research work on the game confrontation of unmanned boats at home and abroad, and it is still in the development stage. In 2022, Su Zhen et al. carried out research on the dynamic game confrontation of unmanned boat swarms, proposed to use the deep deterministic policy gradient algorithm to design a strategy solution method, and the trained agents could better complete the collaborative encirclement and capture tasks. In 2023, Xia Jiawei et al. used the MAPPO algorithm to complete the collaborative encirclement and capture tasks of a single unmanned boat. By combining the background of the encirclement and capture tasks, a scalable and permutation-invariant state space was established, and finally the training of the encirclement and capture strategy was completed using curriculum learning training techniques. However, the research work on the game confrontation of unmanned boat swarms is still in its infancy, there is still a large room for improvement, and there is currently a lack of authoritative and real simulation platforms. At the same time, for reinforcement learning algorithms, in the face of high-dimensional state spaces and complex observation environments, existing multi-agent deep reinforcement learning algorithms still face problems such as difficult algorithm convergence, poor training stability, and low sampling efficiency. Summary of the Invention
[0004] The present invention aims to solve the deficiencies of the prior art and provides the following solutions:
[0005] A multi-unmanned boat game confrontation method based on the DRMRPG algorithm, comprising the following steps:
[0006] Construct a multi-unmanned boat intelligent game framework based on the DRMRPG algorithm;
[0007] On the basis of the MADDPG algorithm, introduce a multi-level experience replay strategy, a dynamic soft update strategy and a residual connection strategy to obtain the DRMRPG algorithm;
[0008] Based on the DRMRPG algorithm, construct an initial multi-unmanned boat game decision model;
[0009] Use the multi-unmanned boat intelligent game framework to train the initial multi-unmanned boat game decision model to obtain an intelligent decision model;
[0010] Use the intelligent decision model to conduct multi-unmanned boat game confrontation to obtain game decisions.
[0011] Preferably, the method for constructing the multi-unmanned boat intelligent game framework includes:
[0012] Construct the state space of the multi-unmanned boat game environment based on the DRMRPG algorithm to obtain the state information of the red and blue unmanned boat agents during the game process;
[0013] Based on the state information, construct the action space of the multi-unmanned boat game environment based on the DRMRPG algorithm to obtain the set of actions that the red and blue unmanned boat agents can execute during the game process;
[0014] Based on the action set, construct a segmented reward function to guide the red and blue unmanned boat agents to continuously optimize their own strategies during the game process to complete the construction of the multi-unmanned boat intelligent game framework.
[0015] Preferably, the method for constructing the segmented reward function includes:
[0016] The reward function of the red unmanned boat agent is:
[0017] Y end = r adv + r pos + r col
[0018]
[0019]
[0020] Among them, r end represents the reward finally obtained by the red unmanned boat, pos adv represents the position of the opponent's unmanned boat, pos desIndicates the location of the target location, pos i Indicates the location of the i-th unmanned boat in the red unmanned boat cluster, range map Indicates the range of the map, r adv Indicates the negative reward of the opponent's target for the agent, r pos Indicates the positive reward obtained by the agent approaching the target landmark, r col Indicates the penalty caused by a collision, collide indicates a collision;
[0021] The reward function of the blue unmanned boat agent is:
[0022] r end = r blue + r col
[0023]
[0024] Among them, r blue Indicates the positive reward of the self for the agent.
[0025] Preferably, the method for obtaining the DRMRPG algorithm includes:
[0026] Introduce the multi-level experience replay strategy to replace the single-layer random experience replay pool in the MADDPG algorithm;
[0027] Introduce the dynamic soft update strategy in the MADDPG algorithm to guide the parameter update of the Actor network and the Critic network:
[0028]
[0029] Among them, τ represents the current soft update parameter, γ represents the weight factor, f(x) represents the attenuation function, τ 0 Represents the soft update parameter at the previous time point, τ e Represents the truncation parameter;
[0030] Introduce the residual connection strategy in the MADDPG algorithm to extract the state space features of the agent and ensure the gradient loss.
[0031] Preferably, the method for constructing the initial multi-unmanned boat game decision model includes:
[0032] Set the initial positions and heading angles of the red and blue unmanned boats;
[0033] Use the multi-level experience replay pool to hierarchically divide the experience data, and hierarchically divide the experience data according to a ratio to obtain the initial multi-unmanned boat game decision model.
[0034] Preferably, the method for obtaining the intelligent decision-making model includes:
[0035] Under the multi-unmanned boat intelligent game framework, the red and blue unmanned boats are made to confront in the initial scenario to generate confrontation data;
[0036] Based on the confrontation data, the data is split into time series sequences of a fixed length T, and through maximum-minimum normalization processing, it is input into the initial multi-unmanned boat game decision-making model to output the strategy and value of the state information at the current moment;
[0037] Calculate the loss value using the strategy and the value, and update the parameters of the initial multi-unmanned boat game decision-making model through gradient descent optimization to obtain the intelligent decision-making model.
[0038] The present invention also provides a multi-unmanned boat game confrontation system based on the DRMRPG algorithm. The system applies the method described in any one of the above, and includes: a framework construction module, an algorithm improvement module, a model construction module, a model training module, and a decision-making module;
[0039] The framework construction module is used to construct a multi-unmanned boat intelligent game framework based on the DRMRPG algorithm;
[0040] The algorithm improvement module is used to introduce a multi-level experience replay strategy, a dynamic soft update strategy, and a residual connection strategy on the basis of the MADDPG algorithm to obtain the DRMRPG algorithm;
[0041] The model construction module constructs an initial multi-unmanned boat game decision-making model based on the DRMRPG algorithm;
[0042] The model training module uses the multi-unmanned boat intelligent game framework to train the initial multi-unmanned boat game decision-making model to obtain an intelligent decision-making model;
[0043] The decision-making module uses the intelligent decision-making model to conduct multi-unmanned boat game confrontation to obtain game decisions.
[0044] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0045] The present invention uses the residual connection method to extract environmental features, improve the cumulative reward of the intelligent agent, uses the dynamic soft update mechanism to improve the sampling efficiency of the intelligent agent, improve the stability of the intelligent agent, and uses the multi-level experience replay pool to improve the utilization rate of historical experience and improve the algorithm convergence performance. Generally speaking, the present invention obtains high reward values, high stability, and high efficiency, and has a certain effectiveness. Description of the Drawings
[0046] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings required in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0047] Figure 1 Schematic diagram of the method flow of the embodiment of the present invention;
[0048] Figure 2 Block diagram of the method flow of the embodiment of the present invention;
[0049] Figure 3 Schematic diagram of the multi-unmanned boat game framework based on the DRMRPG algorithm in the embodiment of the present invention;
[0050] Figure 4 Architecture diagram of the simulation engine in the embodiment of the present invention;
[0051] Figure 5 Information interaction diagram of the multi-unmanned boat intelligent framework based on the Unity engine in the embodiment of the present invention;
[0052] Figure 6 Schematic diagram of the three-degree-of-freedom model of the unmanned boat in the embodiment of the present invention;
[0053] Figure 7 Schematic diagram of the MADDPG algorithm network architecture with residual connections in the embodiment of the present invention;
[0054] Figure 8 Schematic diagram of the structure of the multi-level experience replay pool in the embodiment of the present invention;
[0055] Figure 9 Schematic diagram of the simulation engine environment in the embodiment of the present invention;
[0056] Figure 10 Schematic diagram of the unmanned boat model in the embodiment of the present invention, where a is the schematic diagram of the red unmanned boat model, and b is the schematic diagram of the blue unmanned boat model;
[0057] Figure 11 Structure diagram of the sea battle decision training based on the DRMRPG algorithm in the embodiment of the present invention;
[0058] Figure 12 Training trajectory diagram of the red and blue sides in the sea battle in the embodiment of the present invention, where a is the combat trajectory diagram of the red and blue sides within 200 time steps when using the original MADDPG algorithm, b is the combat trajectory diagram of the red and blue sides within 200 time steps when using the MRDRPG algorithm, and c is the combat trajectory diagram of the red and blue sides within 400 time steps when using the MRDRPG algorithm;
[0059] Figure 13 The simulation reward change diagram of the red - side intelligent agent of the MADDPG and DRMRPG algorithms in the embodiments of the present invention;
[0060] Figure 14 The simulation reward change diagram of the blue - side intelligent agent of the MADDPG and DRMRPG algorithms in the embodiments of the present invention. Detailed implementation manners
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0062] To make the above - mentioned objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0063] Embodiment 1
[0064] In this embodiment, as shown in Figure 1 、 Figure 2 , a multi - unmanned - boat game confrontation method based on the DRMRPG algorithm includes the following steps:
[0065] S1. Construct a multi - unmanned - boat intelligent game framework based on the DRMRPG algorithm.
[0066] The method for constructing the multi - unmanned - boat intelligent game framework includes: constructing the state space of the multi - unmanned - boat game environment based on the DRMRPG algorithm to obtain the state information of the red - and blue - side unmanned - boat intelligent agents during the game process; based on the state information, constructing the action space of the multi - unmanned - boat game environment based on the DRMRPG algorithm to obtain the set of actions that the red - and blue - side unmanned - boat intelligent agents can execute during the game process; based on the set of actions, constructing a segmented reward function to guide the red - and blue - side unmanned - boat intelligent agents during the game process to continuously optimize their own strategies, and complete the construction of the multi - unmanned - boat intelligent game framework.
[0067] In this embodiment, as shown in Figure 3 , in the sea - battle game framework based on deep reinforcement learning, the states of the intelligent agents are integrated and calculated to describe the situation of the sea - battle environment. During the game process, the intelligent agents make decisions based on the reinforcement learning model. During the interaction process, the intelligent agents change the environmental state by executing actions and receive reward signals from the environment as feedback. The goal of the intelligent agents is to find a strategy that maximizes the total reward obtained after executing a series of actions.
[0068] The naval battle game framework consists of three basic elements: the state space, the action space, and the reward function. In this embodiment, the Unity engine is used as the simulation engine. In the Python reinforcement learning training code, the environment object is virtualized, and the TCP protocol family is used to implement the information interaction between the simulation engine and the virtual environment. The simulation engine generates the overall environmental observation value in real time, serializes it into a byte stream in Json format, and transmits it to the virtual environment at the reinforcement learning end through the TCP protocol for parsing. After the virtual environment parses the environmental observation value, it assigns the local observation value to the agent and calculates the state of the agent online, collects the instructions of the agent, and sends them back to the simulation engine to achieve the information interaction between the simulation engine and the virtual environment.
[0069] The architecture diagram of the simulation engine in this embodiment is as Figure 4 shown: Based on the Unity engine, a simulation environment architecture for the multi-unmanned boat game problem based on the DRMRPG algorithm is proposed. Generally speaking, the simulation environment consists of a rendering module, entity objects, an object manager pool, a TCP client, and an environmental observation object. Among them, the rendering module includes a skybox, a lighting model, cloud effects, and water body rendering effects. The materials used are from the Unity store and are open-source and free materials. The environmental rendering style and entity objects are as Figure 9 shown, and the unmanned boat models of both the red and blue sides used are as Figure 10 shown. The object manager pool maintains three objects, namely the message manager, the environment manager, and the status information manager. Among them, the message manager is responsible for parsing the entity control information received by the TCP client process and sending the control instructions to the entity objects to control the entities in the environment; the environment manager is responsible for parsing the environmental control information received by the TCP client process and controlling the basic functions such as the opening, closing, and acceleration of the overall environment; the status information manager is the serialization of the messages generated by the environmental observation singleton object. It encapsulates and serializes the overall environmental observation value generated by the environmental singleton object into a TCP data stream in Json format for parsing by the virtual environment in Python. The entity objects control all possible environmental entities in the environment, such as agents and obstacles. Each entity object is equipped with a dedicated control script to parse the instructions passed by the message manager and control the entity. The TCP client process is set at the bottom of the simulation environment and is responsible for the sending and receiving of all communication information with Python. It is implemented using the client-server architecture and the TCP / IP protocol family, and a dedicated private protocol is designed. The environmental observation singleton is a singleton object that is generated when the environment starts and is responsible for generating the overall observation information of the environment.
[0070] Based on this, the information interaction diagram of the multi-unmanned boat game framework based on the Unity engine in the embodiment of the present invention is as Figure 5Shown as follows: The implemented logic is to virtualize a reinforcement learning environment on the Python side. It is actually a server that is always on, waiting for the access of the simulation environment in a blocking manner. When an access request is received, it will start the reinforcement learning process, collect the observation information of the Unity simulation environment and parse it. The agent generates control information based on the current local observation, which is transmitted through the virtual environment to control the entities in the simulation environment and change the state of the agent.
[0071] ①Construct the naval battle state space
[0072] To better reflect the current combat situation, the state space should not only include its own state information, but also the state information of friendly forces, the state information of enemy forces, the relative relationship between the two sides, and the state of non-agent entities in the environment.
[0073] In this environment, the state space of the cooperative agent includes: its own position, its own heading angle, its own vector velocity, the position of the target location, and the distance of other agents from the target location. The adversarial agent does not include the position of the target location. The coordinated state space of a single agent is shown as follows. In this embodiment, the state space is normalized, which is very important in most cases.
[0074] The state space of the red unmanned boat agent is:
[0075] state colla = [obs self , obs other , obs ext , obs des
[0076] obs self = [Xpos self , Ypos self , Xvel self , Yvel self , yaw self
[0077] obbs ohter = [∑ i≠self Xpos i , ∑ i≠self Ypos i
[0078] obs ext = [∑ i dis(agt, des)]
[0079] obs des = [Xdes, Ydes]
[0080] The state space of the blue - side unmanned boat agent is as follows:
[0081] state adv =[obs self , obs other , obs ext
[0082] obs self =[Xpos self , Ypos self , Xvel self , Yvel self , yaw self
[0083] obs ohter =[∑ i≠self Xpos i , ∑ i≠self Ypos i
[0084] obs ext =[∑ i dis(agt, des)]
[0085] Among them, obs self represents the state information of the unmanned boat itself. Xpos self and Ypos self represent the spatial position coordinates of the unmanned boat itself. Xvel self and Yvel self represent the speed of the unmanned boat itself. yaw self represents the course of the unmanned boat itself. obs other represents the state information of other unmanned boats. Xpos i and Ypos i represent the spatial position coordinates of other unmanned boats. obs ext represents the relative position information between other unmanned boats and the target landmark. dis() represents the relative position information. agt represents other unmanned boats, and des represents the target landmark. obs des represents the position information of the target landmark entity in the environment. Xdes and Ydes represent the spatial position information of the target landmark.
[0086] ② Construct the naval battle action space
[0087] To simulate the control problem of warships in real situations, a responsive model is used to describe the motion problem of warships. In this problem, the ship motion is dynamically modeled using three degrees of freedom: forward, side drift, and yaw. Its motion parameters are as Figure 6 defined below.
[0088] Where x, y, and r are the x-axis position, y-axis position, and z-axis rotation speed in the navigation coordinate system; Ψ is the heading; u and v are the x-axis speed and y-axis speed in the body-fixed coordinate system, and the transformation matrix T(Ψ) is defined as follows:
[0089]
[0090] The conversion relationship between the speed vectors in the navigation coordinate system and the body-fixed coordinate system is:
[0091]
[0092] With the help of the physical engine of Unity Engine, the above-mentioned three-degree-of-freedom dynamic modeling of the unmanned boat can be easily realized, and it is controlled using a four-dimensional action space, as shown in Table 1.
[0093] Table 1
[0094]
[0095] It should be noted that in the specific implementation, the details of the action engine are shielded, and the motion control of the intelligent body is realized using the upper-layer speed calculation and steering torque calculation. Therefore, the actual control equation of the unmanned boat is as follows:
[0096] F acc = max(0.1, F - (f ship + f resis ))
[0097] vel cur = Lerp(vel cur , vel tar , l 1 )
[0098] rot cur = Lerp(rot cur , rot tar , l 2 )
[0099] rot cur (angle) → rot cur (rad)
[0100] rot axis = [u x , u y , u z
[0101]
[0102] vel rig = Q(rot cur , rotaxis )*vel cur
[0103] Among them, among them, F acc is the resultant force, F is the driving force, f ship and f resis are the hull resistance and the environmental resistance respectively, vel tar is the desired speed scalar, which is determined by F acc ; vel cur is the current speed scalar after Lerp interpolation calculation, and l is the interpolation parameter. rot tar is the desired rotation scalar, rot cur is described in the form of a quaternion, and the rot axis decomposed into vectors is used to calculate a new rotation quaternion, vel rig is the final rigid body velocity vector, which is jointly determined by vel cur and rot cur . It should be noted that in the interaction between the algorithm and the environment, in order to simulate the instruction synchronization in the real situation, the discrete action instructions of the algorithm will be implemented by the environment as a continuous control problem for the intelligent agent.
[0104] ③Construct the reward function
[0105] Generally speaking, this embodiment uses a reward function design based on proximity reward and collision penalty. The relative distance between the intelligent agent and the target landmark and the relative distance between the intelligent agent and other intelligent agents will be used as indicators to estimate the confrontation on the water surface. The same reward function is adopted by all intelligent agents in the environment. When the adversarial intelligent agent approaches the target landmark, the friendly intelligent agent will obtain a negative reward, while when the friendly intelligent agent approaches the target landmark, it will obtain a positive reward. In addition, a penalty term related to collision is designed for all intelligent agents. When two intelligent agents are too close, they will obtain a very large penalty. The design of the reward function is as follows: The reward function of the red side's unmanned boat intelligent agent is:
[0106] r end = r adv + r pos + r col
[0107]
[0108] Among them, r end represents the final reward obtained by the red side's unmanned boat, pos adv represents the position of the other side's unmanned boat, pos des represents the position of the target location, pos i represents the position of the i-th unmanned boat in the red side's unmanned boat cluster, range mapIndicates the range of the map, r adv Indicates the negative reward of the opponent's target for the agent, r pos Indicates the positive reward obtained by the agent approaching the target landmark, r col Indicates the penalty caused by a collision, collide indicates a collision; the reward function of the blue unmanned boat agent is:
[0109] r end = r adv + r col
[0110]
[0111] Among them, r blue Indicates the positive reward of the self for the agent.
[0112] S2. Based on the MADDPG algorithm, a multi-level experience replay strategy, a dynamic soft update strategy, and a residual connection strategy are introduced to obtain the DRMRPG algorithm.
[0113] The method to obtain the DRMRPG algorithm includes: introducing a multi-level experience replay strategy to replace the single-layer random experience replay pool in the MADDPG algorithm; introducing a dynamic soft update strategy in the MADDPG algorithm to guide the parameter update of the Actor network and the Critic network; introducing a residual connection strategy in the MADDPG algorithm to extract the state space features of the agent and ensure the gradient loss.
[0114] In this embodiment, the DRMRPG algorithm is proposed. By introducing a multi-level experience replay strategy on the basis of MADDPG, it guides the agent to learn and update; uses a dynamic soft update strategy to guide the parameter update of the Actor network and the Critic network; uses a residual connection strategy to extract the state space features of the agent cluster and ensure the gradient loss. Figure 11 It is a structural diagram of naval battle decision training based on the DRMRPG algorithm.
[0115] ① Residual connection method
[0116] This embodiment proposes a method for using residual connections to extract features for all agents during execution. In this method, the observation of each agent will first pass through a multi-layer perceptron. The first multi-layer perceptron encodes the local observation of agent i into Subsequently, the local observation of agent i will be added to through a residual connection. Therefore, the observation value obtained by the Actor network of agent i finally is is a tensor that has been feature-extracted once and added to the original state, It will be passed into the multi-layer perceptron where the Actor network is located as a new state feature, and the Actor network will obtain an action These and will be concatenated into a unified observation vector, and the Critic network will give a score Q i (s, a). Subsequently, agent i will update its own policy π i through the score Q i (s, a) of the Critic network. During the update process of the Actor and Critic networks, ReLU is used for optimization in all layers, and the loss function is defined by the mean square error function MSE. The network architecture of the MADDPG algorithm with residual connections is as Figure 7 shown.
[0117] ② Dynamic soft update strategy
[0118] After introducing residual connections for feature extraction of state information, although the problem of gradient vanishing / gradient explosion will be alleviated due to the use of residual connections, the deep neural network will still inevitably slow down the training speed due to the increase in the number of network layers, which is undesirable under limited computing resources. Therefore, a dynamic soft update strategy is introduced to increase the update frequency of the target network, accelerate the sampling efficiency, and achieve the purpose of accelerating the algorithm convergence.
[0119] In the update of the value network parameters, the ultimate goal of the update is to make Q ω (s, a) gradually approach r + max a′ Q ω (s′, a′). When using a single network, the training will be very unstable due to the update of the TD error. Therefore, the design of a double-layer network is introduced, that is, an online network and a target network are used simultaneously. During the training process, the target network is fixed first, and the parameters of the target network are updated asynchronously by updating the online network, which can solve the problem of unstable training to a certain extent. The soft update strategy further implements this purpose. In the algorithm using the soft update strategy, the update of the target network is based on the slow update method of the soft update parameter τ, and its formula is ω - ←τω+(1 - τ)ω - . Usually, τ is a small number. τ ensures the slow update of the network parameters through this strategy and maintains the stability of the training.
[0120] However, this strategy also has corresponding deficiencies. The traditional soft update strategy selects a relatively small update step during the training process. Although this helps to maintain the stability of the strategy, it also leads to a relatively slow learning process and requires more time to achieve the ideal performance. Therefore, in this embodiment, a dynamic and adaptive parameter update method is selected, and its formula is as follows:
[0121]
[0122] Among them, τ represents the current soft update parameter, γ represents the weight factor, τ 0 represents the soft update parameter at the previous time point, τ e represents the truncation parameter, represents the decay function, which will continuously increase during the training process, making g(x) = γ(τ 0 / f(x)) rapidly decline in the initial stage of training. To avoid the situation that the network does not update in the later stage of algorithm training, a truncation parameter τ e is set. It will make the policy network still update with a relatively small amplitude in the later stage to ensure the effectiveness of training. In summary, this embodiment dynamically adjusts the soft update strategy, enabling the policy network to update with a relatively large amplitude in the initial stage of training, achieving rapid convergence of the algorithm. At the same time, it updates with a relatively small amplitude in the later stage of training, maintaining the continuous update of the algorithm, which helps the iteration of the policy and the convergence of the algorithm. Experiments show that this dynamic update strategy is very effective in simple environments and also has a certain role in complex environments.
[0123] ③ Multi-level experience replay pool strategy
[0124] Although the convergence speed of the algorithm has been improved after adding the dynamic soft update strategy, when facing complex environments, this strategy may instead have a negative impact on the policy construction of the intelligent agent (this is usually due to the dependence of the intelligent agent on the initial policy under the soft update strategy). However, it is still hoped to further improve the sampling efficiency of the intelligent agent in the environment on this basis. Therefore, the widely used random experience replay method is redesigned, and a multi-level experience replay pool strategy is proposed. The multi-level experience replay pool is as Figure 8 shown.
[0125] In this set of experience replay pools, the basic experience replay pool is responsible for storing all historical experience data generated. At the end of each episode, the algorithm calculates the maximum cumulative experience sum of the current episode. When this cumulative experience sum is greater than that of the previous episode, the preferred experience replay pool will absorb the experience data of this episode; otherwise, the secondary experience replay pool will receive the experience data of this episode. This means that the historical experience of each episode is divided into two-quality experience replay pools through a multi-level experience replay pool. In the initial stage of training, the agent is made to learn historical experience from the basic experience replay pool containing all experience data and update the algorithm, which enables the agent to conduct a higher degree of exploration in the initial stage of training to avoid falling into local optimality prematurely. This is the first-level experience replay. When the data scale in the secondary experience pool reaches a certain level, random data screening will be carried out from the second-level experience replay pool, that is, the preferred experience replay pool and the secondary experience replay pool. The principles of experience screening are shown in Table 2:
[0126] Table 2
[0127]
[0128] S3. Based on the DRMRPG algorithm, construct an initial multi-unmanned boat game decision-making model.
[0129] The method for constructing the initial multi-unmanned boat game decision-making model includes: setting the initial positions and heading angles of the red and blue unmanned boats; using a multi-level experience replay pool to hierarchically divide the experience data and dividing the experience data according to a ratio to obtain the initial multi-unmanned boat game decision-making model.
[0130] S4. Use the multi-unmanned boat intelligent game framework to train the initial multi-unmanned boat game decision-making model to obtain an intelligent decision-making model.
[0131] The method for obtaining the intelligent decision-making model includes: under the multi-unmanned boat intelligent game framework, making the red and blue unmanned boats confront in the initial scenario to generate confrontation data; based on the confrontation data, splitting the data into time series sequences of a fixed length T, and through maximum-minimum normalization processing, inputting it into the initial multi-unmanned boat game decision-making model to output the strategies and values of the state information at the current moment; calculating the loss value using the strategies and values, and optimizing and updating the parameters of the initial multi-unmanned boat game decision-making model through gradient descent to obtain the intelligent decision-making model.
[0132] In this embodiment:
[0133] ① Initialize the model parameters
[0134] In each training session, model parameters are set for both the red and blue sides. The environmental parameters are set as shown in Table 3, and the hyperparameters are shown in Table 4. To ensure the generalization ability of the agents and improve the diversity of strategies, the initial positions and heading angles of the agents are randomly selected at the start of each game, and their initial parameter settings are shown in Table 5.
[0135] Table 3
[0136]
[0137] Table 4
[0138]
[0139] Table 5
[0140]
[0141] ② Generate naval battle data based on the naval battle game framework
[0142] Based on the naval battle game framework in the initial scenario, the red and blue unmanned boats use the proposed MADDPG algorithm network model to generate strategies for confrontation, and real-time simulation naval battle data is generated through the simulation environment and input into the naval battle strategy generation model.
[0143] ③ Use the DRMRPG algorithm to update the model to achieve autonomous generation of naval battle intention strategies
[0144] The sequence data is input into the proposed DRMRPG algorithm network model, and the state information S t of the current moment is output, along with the strategy π θ (a t |s t ) and the value V(s t ). The DRMRPG algorithm calculates the loss value using the strategy and value, and updates the model parameters θ through gradient descent optimization to obtain the intelligent decision-making model.
[0145] S5. Use the intelligent decision-making model for multi-unmanned boat game confrontation to obtain game decisions.
[0146] Example 2
[0147] In this example, to verify the effectiveness of the multi-unmanned boat game algorithm based on the DRMRPG algorithm proposed in the present invention, comparative simulation training experiments and ablation training experiments are carried out based on the DRMRPG algorithm proposed in the present invention. Figure 12 It is a training trajectory map of the naval battle between the red and blue sides. Figure 12It is a strategy analysis of agents trained using the MADDPG algorithm and the DRMRPG algorithm in the collaborative and adversarial environment of unmanned boats. In this environment, due to the physical properties of the unmanned boats themselves, they cannot move arbitrarily but must follow the physical laws of the real world. In Figure 12 (a), the strategy of the collaborative unmanned boats is as follows: In the initial stage of the game, unmanned boat A, whose spawn point is closer to the target landmark, will accelerate towards the target landmark, while unmanned boat B, whose spawn point is farther away, will swim slowly, resulting in a situation where one unmanned boat is in front and the other is behind. This makes it impossible for the adversarial unmanned boat C to identify the location of the target landmark. However, as the game progresses, unmanned boat A shows no tendency to decelerate and will still accelerate away after reaching near the target landmark. After A leaves the target landmark, unmanned boat B will accelerate forward towards the target landmark to obtain rewards for the team. This strategy causes unmanned boat C to approach the location where unmanned boat B is in the late stage of the game, thus reducing the rewards of the collaborative unmanned boats. This is not the desired strategy. In Figure 12 (b) and Figure 12 (c), there are some pleasantly surprising improvements in the strategy of the collaborative unmanned boats: In the short term, in the initial stage of the game, unmanned boat A, which is born closer, will still accelerate towards the target landmark. After approaching the target landmark, instead of blindly moving far away, it will perform a turning maneuver to induce unmanned boat C. Unmanned boat B faithfully sails towards the target location at a constant speed (slower than A). This strategy makes it impossible for unmanned boat C to identify the location of the target landmark and thus hover in place without being able to approach the target landmark. As a result, the cumulative rewards obtained by the collaborative agents have increased significantly. The differences in the strategies of the two groups of unmanned boats prove that the DRMRPG algorithm has a greater improvement in the agent strategy compared to the MADDPG algorithm. Figure 13 It is the simulation reward change graph of the red agents of the MADDPG and DRMRPG algorithms in the embodiments of the present invention; Figure 14 It is the simulation reward change graph of the blue agents of the MADDPG and DRMRPG algorithms in the embodiments of the present invention; In this environment, 5000 games were trained, and the training duration for every 5000 times was about 25 hours. It can be seen that in a relatively complex control environment, the average cumulative rewards of the agents trained by the DRMRPG algorithm are significantly higher than those of the agents trained by the MADDPG algorithm, and it also has a better performance in terms of algorithm stability. For the red agents, the agents trained by the DRMRPG algorithm are close to convergence at about 1500 steps, and their stability after convergence is significantly higher than that of the agents trained by the MADDPG algorithm; for the blue agents, the agents trained by the DRMRPG algorithm are also close to convergence at about 1500 steps, and their stability is also significantly higher than that of the agents trained by the MADDPG algorithm. This shows the effectiveness and superiority of the DRMRPG algorithm.
[0148] To further verify the effectiveness of each optimization module in the multi-unmanned boat game confrontation framework based on the DRMRPG algorithm proposed in the present invention, ablation experiments were conducted on the DRMRPG algorithm. These experiments were based on the MADDPG algorithm, and a series of different algorithm variants were constructed by gradually increasing or decreasing these three improvement strategies. Table 6 details the various algorithm variants used in the ablation experiments. The "√" symbol in the table indicates that the algorithm includes the corresponding improvement mechanism, while the "×" symbol indicates that it does not. Among them, DMRPG refers to an improved MADDPG method that includes a dynamic soft update strategy and a multi-level experience replay pool strategy, RMRPG refers to an improved MADDPG method that includes a residual connection method and a multi-level experience replay pool strategy, and DRPG refers to an improved MADDPG method that includes a dynamic soft update strategy and a residual connection method.
[0149] Table 6
[0150]
[0151] Specifically: (a) The RMRPG algorithm lacking the dynamic soft update strategy performed poorly in terms of convergence. The RMRPG had a slow convergence speed, and there were obvious fluctuations in the training of the algorithm, indicating the positive impact of the dynamic soft update strategy on the algorithm's convergence speed. (b) The stability of the DRPG algorithm lacking the multi-level experience replay pool strategy was unacceptable. When the data volume was large, the multi-level experience replay pool strategy had a more obvious effect on improving the convergence performance of the algorithm. (c) The cumulative reward of the DMRPG algorithm lacking the residual connection method decreased significantly, and its convergence performance on the adversary was poor, indicating the effectiveness of the residual connection method. The above experimental results verified the correctness and feasibility of the DRMRPG algorithm.
[0152] Through comparative experiments and ablation experiments, the average reward value and improvement rate of each algorithm at convergence (if not converged, it is the end of training) were finally obtained, as shown in Table 7.
[0153] Table 7
[0154]
[0155] The data shows that on the premise of maintaining the convergence of the algorithm, the DRMRPG algorithm has a very obvious improvement efficiency in the reward value of the agent. Compared with the MADDPG algorithm, the cumulative reward of the collaborator agent trained by the DRMRPG algorithm has increased by more than ten times.
[0156] Embodiment 3
[0157] In this embodiment, a multi-unmanned boat game confrontation system based on the DRMRPG algorithm includes: a framework construction module, an algorithm improvement module, a model construction module, a model training module, and a decision-making module.
[0158] The framework construction module is used to construct a multi-unmanned boat intelligent game framework based on the DRMRPG algorithm. The algorithm improvement module is used to introduce a multi-level experience replay strategy, a dynamic soft update strategy, and a residual connection strategy on the basis of the MADDPG algorithm to obtain the DRMRPG algorithm. The model construction module constructs an initial multi-unmanned boat game decision-making model based on the DRMRPG algorithm. The model training module uses the multi-unmanned boat intelligent game framework to train the initial multi-unmanned boat game decision-making model to obtain an intelligent decision-making model. The decision-making module uses the intelligent decision-making model to conduct multi-unmanned boat game confrontation to obtain game decisions.
[0159] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A multi-unmanned boat game confrontation method based on DRMRPG algorithm, characterized in that: The following steps are involved: Construct a multi-unmanned boat intelligent game framework based on the DRMRPG algorithm; On the basis of the MADDPG algorithm, a multi-level experience playback strategy, a dynamic soft update strategy and a residual connection strategy are introduced to obtain the DRMRPG algorithm; Based on the DRMRPG algorithm, an initial multi-unmanned boat game decision model is constructed; The initial multi-unmanned boat game decision model is trained by using the multi-unmanned boat intelligent game framework to obtain an intelligent decision model; The intelligent decision-making model is used to conduct a game confrontation among multiple unmanned boats to obtain a game decision.
2. According to claim 1, a multi-unmanned boat game confrontation method based on DRMRPG algorithm is characterized in that: The method for constructing the multi-unmanned boat intelligent game framework includes: Constructing the state space of the multi-unmanned boat game environment based on the DRMRPG algorithm, and obtaining the state information of the red and blue unmanned boat agents during the game; Based on the state information, construct an action space of a multi-unmanned boat game environment based on the DRMRPG algorithm, and obtain a set of actions that can be performed by the red and blue unmanned boat agents during the game; Based on the action set, a piecewise reward function is constructed to guide the red and blue unmanned boat agents to continuously optimize their own strategies during the game, thereby completing the construction of the multi-unmanned boat intelligent game framework.
3. According to claim 2, a multi-unmanned boat game confrontation method based on DRMRPG algorithm is characterized in that: Methods for constructing piecewise reward functions include: The reward function of the Red Team’s unmanned boat agent is: r end =r adv +r pos +r col Among them, r end Indicates the final reward obtained by the Red Team’s unmanned boat, pos adv Indicates the location of the other unmanned boat, pos des Indicates the location of the target location, pos i represents the location of the i-th unmanned boat in the red team’s unmanned boat cluster, range map Indicates the range of the map, r adv represents the negative reward of the other party’s goal to the agent, r pos represents the positive reward obtained by the agent when approaching the target landmark, r col Indicates the penalty caused by a collision, collide means a collision occurs; The reward function of the blue unmanned boat agent is: r end =r blue +r col Among them, r blue Represents the positive reward for the agent.
4. According to claim 1, a multi-unmanned boat game confrontation method based on DRMRPG algorithm is characterized in that: The method for obtaining the DRMRPG algorithm includes: Introducing the multi-level experience replay strategy to replace the single-layer random experience replay pool in the MADDPG algorithm; The dynamic soft update strategy is introduced into the MADDPG algorithm to guide the parameter update of the Actor network and the Critic network: Among them, τ represents the current soft update parameter, γ represents the weight factor, f(x) represents the attenuation function, τ0 represents the soft update parameter at the previous time point, τ e Indicates truncation parameters; The residual connection strategy is introduced into the MADDPG algorithm to extract the state space characteristics of the agent and ensure the gradient loss.
5. According to claim 1, a multi-unmanned boat game confrontation method based on DRMRPG algorithm is characterized in that: The method for constructing the initial multi-unmanned boat game decision model includes: Set the initial position and heading angle of the red and blue unmanned boats; The experience data is divided into layers using a multi-level experience replay pool, and the experience data is divided into layers according to a proportion to obtain the initial multi-unmanned boat game decision model.
6. The multi-unmanned boat game confrontation method based on the DRMRPG algorithm according to claim 1 is characterized in that: The method for obtaining the intelligent decision-making model includes: Under the multi-unmanned boat intelligent game framework, the red and blue unmanned boats are made to confront each other in an initial scenario to generate confrontation data; Based on the confrontation data, the data is split into a time series of fixed length T, and is input into the initial multi-unmanned boat game decision model through maximum-minimum normalization processing, and the strategy and value of the current state information are output; The loss value is calculated using the strategy and the value, and the parameters of the initial multi-unmanned boat game decision model are updated through gradient descent optimization to obtain the intelligent decision model.
7. A multi-unmanned boat game confrontation system based on DRMRPG algorithm, the system applies the method described in any one of claims 1-6, characterized in that: include: Framework building module, algorithm improvement module, model building module, model training module and decision-making module; The framework construction module is used to construct a multi-unmanned boat intelligent game framework based on the DRMRPG algorithm; The algorithm improvement module is used to introduce a multi-level experience playback strategy, a dynamic soft update strategy and a residual connection strategy on the basis of the MADDPG algorithm to obtain the DRMRPG algorithm; The model building module builds an initial multi-unmanned boat game decision model based on the DRMRPG algorithm; The model training module uses the multi-unmanned boat intelligent game framework to train the initial multi-unmanned boat game decision model to obtain an intelligent decision model; The decision-making module uses the intelligent decision-making model to conduct a multi-unmanned boat game confrontation to obtain a game decision.
Citation Information
Patent Citations
Multi-agent reinforcement learning method for collaborative decision-making of multiple combat units
CN114358141A
Unmanned aerial vehicle cluster cooperative combat game method and system based on deep reinforcement learning
CN115903903A
Multi-unmanned aerial vehicle cooperative field source search trajectory planning method based on deep reinforcement learning
CN117193372A
Stable fixed-wing unmanned aerial vehicle main wing aircraft clustering method, device and equipment
CN118394107A
Robot navigation and object tracking
US20190217476A1
Cited By
Multi-unmanned ship vector propeller game control method based on hierarchical multi-agent reinforcement learning
CN121386387A