Multi-agent near-end strategy optimization unmanned aerial vehicle game method of adaptive iteration resolution

By constructing a kinematic and control model with adaptive iterative resolution in a drone game environment and combining it with a deep reinforcement learning algorithm to dynamically adjust the resolution of strategy updates, the problems of low training efficiency and difficult strategy convergence of the MAPPO algorithm in complex scenarios are solved, and efficient drone cluster game is achieved.

CN120669520APending Publication Date: 2025-09-19NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510657435.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In complex scenarios, the MAPPO algorithm faces problems such as complex scenarios and a large number of parameters, which makes it impossible to train and difficult to converge the strategy.

Method used

A multi-agent proximal strategy optimization UAV game method with adaptive iterative resolution is proposed. By constructing a UAV kinematic and control model that introduces a resolution reduction factor parameter and combining it with a deep reinforcement learning algorithm, the resolution of the policy update is dynamically adjusted to improve training efficiency and convergence speed.

Benefits of technology

It significantly improves training efficiency and convergence speed, and can achieve good training results even under sparse rewards, enabling drone swarms to quickly adapt to mission requirements and promoting the application and development of deep reinforcement learning in drone swarm intelligent games.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669520A_ABST
    Figure CN120669520A_ABST
Patent Text Reader

Abstract

The invention discloses an adaptive iteration resolution multi-agent near-end strategy optimization unmanned aerial vehicle game method and device, a medium and equipment, and the method comprises the steps: constructing a kinematics model and a control model of an unmanned aerial vehicle, which introduces a resolution reduction multiple parameter, in an unmanned aerial vehicle game environment; based on interaction of an unmanned aerial vehicle cluster game deep reinforcement learning algorithm, the kinematic model and the control model, environment state information is obtained, and each reward function difference value is obtained; judging that the absolute value of each reward function difference value is smaller than a preset threshold value, and updating a resolution reduction multiple parameter by utilizing a preset resolution parameter updating algorithm so as to update the kinematic model and the control model; according to the method, the unmanned aerial vehicle cluster game deep reinforcement learning algorithm of the adaptive iteration resolution is deployed in each unmanned aerial vehicle agent participating in the game, and the unmanned aerial vehicle is guided to make the optimal decision in the game environment by using the agents, so that the training efficiency and the convergence speed are remarkably improved by dynamically adjusting the resolution updated by the strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of drone clusters, and in particular to a method, device, medium and equipment for optimizing drone games with a multi-agent proximal strategy and adaptive iterative resolution. Background Art

[0002] In recent years, fixed-wing drone technology has developed rapidly and has been widely used in a variety of fields, including agriculture, disaster relief, and environmental monitoring, bringing significant convenience and efficiency improvements to social production and daily life. However, traditional manual remote control methods are limited by communication distance, weather conditions, and electromagnetic interference, making them difficult to adapt to complex practical application scenarios. Therefore, improving the intelligence level of drones and giving them stronger autonomous decision-making capabilities and environmental adaptability will not only improve their effectiveness in missions but also promote their widespread application in various fields. Furthermore, the development of drone technology has the potential to promote progress in related disciplines such as artificial intelligence, communications technology, and materials science, providing important support for economic development and scientific and technological progress, and has important research significance and practical value.

[0003] Deep reinforcement learning has important research significance and value in the field of drone swarm games. As a machine learning technique, deep reinforcement learning autonomously learns action strategies through interaction with the environment, without relying on sample data, and can effectively deal with complex continuous decision-making problems that lack prior models. In drone swarm combat, deep reinforcement learning can provide drones with autonomous decision-making capabilities, enabling them to achieve efficient swarm games and task execution in dynamic and complex environments. The proximal policy optimization (PPO) algorithm has become one of the important methods in the field of reinforcement learning due to its outstanding performance in sample efficiency, stability, and simplicity. By limiting the amplitude of policy updates, it balances the relationship between exploration and exploitation, significantly improving the efficiency and stability of the algorithm, making it suitable for policy optimization problems in drone swarm games. MAPPO further improves the collaborative performance of multi-agent systems by introducing a more efficient policy update mechanism and adaptive reward design. It demonstrates unique advantages in solving complex multi-agent cooperation problems and provides strong technical support for drone swarm intelligent games.

[0004] However, in complex scenarios, the MAPPO algorithm faces problems such as complex scenarios and a large number of parameters, which lead to problems such as inability to train and difficulty in strategy convergence. Summary of the Invention

[0005] The main purpose of this application is to provide a multi-agent proximal strategy optimization drone game method, device, medium and equipment with adaptive iterative resolution, aiming to solve the technical problems faced by the MAPPO algorithm in complex scenarios, such as complex scenarios and large number of parameters, which lead to the inability to train and difficulty in strategy convergence.

[0006] To achieve the above objectives, the present application provides a multi-agent proximal strategy optimization method for drone games with adaptive iterative resolution, comprising: constructing a kinematic model and a control model of a drone that introduces a resolution reduction factor parameter in a drone game environment; constructing a deep reinforcement learning algorithm for drone swarm games, and obtaining environmental state information based on the interaction between the deep reinforcement learning algorithm for drone swarm games and the kinematic model and the control model; training the deep reinforcement learning algorithm for drone swarm games based on the environmental state information to obtain the actions of the drones; and using the actions of the drones to update the kinematic model, the control model and the preset reward function to obtain the reward function value of this round, and using the reward function value of this round to obtain the reward function value of this round. The round reward function value is subtracted from the previous round reward function values ​​to obtain the difference of each reward function; if the absolute value of the difference of each reward function is less than the preset threshold, it is judged that the drone swarm game deep reinforcement learning algorithm has converged, and the preset resolution parameter update algorithm is used to update the resolution reduction factor parameter to update the kinematic model and control model; the above process is iteratively executed until the drone swarm game deep reinforcement learning algorithm with adaptive iterative resolution is obtained, and the drone swarm game deep reinforcement learning algorithm with adaptive iterative resolution is deployed in each drone intelligent agent participating in the game, and the intelligent agent is used to guide the drone to make the optimal decision in the game environment.

[0007] Optionally, the construction of the kinematic model and control model of the UAV with the resolution reduction factor parameter introduced includes: constructing the kinematic model and control model of the UAV based on the linear velocity and angular velocity of the UAV, wherein the control model includes the linear acceleration of the UAV and the angular acceleration of the UAV; multiplying the upper limit of the linear velocity, the upper and lower limits of the angular velocity of the angular velocity, the linear acceleration of the UAV and the angular acceleration of the UAV with the resolution reduction factor to obtain the kinematic model and control model of the UAV with the resolution reduction factor parameter introduced.

[0008] Optionally, after the drone swarm game deep reinforcement learning algorithm is trained based on the environmental state information to obtain the drone's movements, and the kinematic model, control model and preset reward function are updated using the drone's movements, the method further includes: determining the relative situation between any two drones based on the drone relative situation model, wherein the relative situation includes the relative position vector between any two drones, the first angle between the velocity vector and the relative position vector of the target drone among the two drones, and the second angle between the velocity vector and the relative position vector of the drone; determining the maximum attack distance and the maximum attack angle of the drone based on the drone attack model; if the first angle / second angle of the target drone is less than the attack angle of the own drone, and the relative position vector of the target drone is less than the maximum attack distance of the own drone, it is determined that the target drone is destroyed by the own drone, otherwise it is determined that it is not destroyed by the own drone; obtaining a first number of target drones destroyed by the own drone and a second number of the own drones destroyed by the target drone, processing the first number and the second number based on the preset reward function, and obtaining the reward function value of this round.

[0009] Optionally, the method of training the UAV swarm game deep reinforcement learning algorithm based on the environmental state information to obtain the UAV action includes: each UAV obtains the environmental state information and the game reward function value from the kinematic model and the control model, and processes the environmental state information and the game reward function value based on its own Actor network to obtain the output action; the output action is input into the Critic network to obtain the feedback variable, and the Actor network is optimized based on the feedback variable; each UAV updates the kinematic model and the control model based on the output action.

[0010] Optionally, the preset reward function is:

[0011]

[0012] Among them, i represents the number of target drones destroyed by our drones; j represents the number of our drones destroyed by target drones, and r represents the resolution reduction factor.

[0013] Optionally, the update expression of the resolution reduction factor r is:

[0014]

[0015] Optionally, the updating of the resolution reduction factor parameter using a preset resolution parameter updating algorithm to update the kinematic model and the control model includes:

[0016] At the beginning of training, set the resolution parameter to the minimum resolution reduction factor;

[0017] During the training process, the minimum resolution reduction factor is continuously increased until the resolution parameter reaches the normal resolution.

[0018] In addition, to achieve the above objectives, the present application also provides a multi-agent proximal strategy optimization drone game device with adaptive iterative resolution, comprising:

[0019] The model building module is used to construct the kinematic model and control model of the drone in the drone game environment, which introduces the resolution reduction factor parameter;

[0020] The reward calculation module is used to construct a deep reinforcement learning algorithm for drone swarm games. The deep reinforcement learning algorithm for drone swarm games interacts with the kinematic model and control model to obtain environmental state information. The deep reinforcement learning algorithm for drone swarm games is trained based on the environmental state information to obtain the drone's actions. The kinematic model, control model, and preset reward function are then used to update the drone's actions to obtain the reward function value for this round. The reward function value for this round is then subtracted from the reward function values ​​for multiple previous rounds to obtain the difference between the reward functions.

[0021] The judgment module is used to determine if the absolute value of the difference between each reward function is less than a preset threshold, then determine that the drone swarm game deep reinforcement learning algorithm has converged, and use the preset resolution parameter update algorithm to update the resolution reduction factor parameter to update the kinematic model and control model;

[0022] The decision-making module is used to iteratively execute the above process until a deep reinforcement learning algorithm for drone swarm game with adaptive iterative resolution is obtained. The deep reinforcement learning algorithm for drone swarm game with adaptive iterative resolution is deployed in each drone agent participating in the game, and the agent is used to guide the drone to make the optimal decision in the game environment.

[0023] To achieve the above objectives, the present application also provides a computer-readable storage medium, which includes instructions that, when run on a computer, enable the computer to execute the multi-agent proximal strategy optimization drone game method with adaptive iterative resolution provided in the above embodiment.

[0024] To achieve the above-mentioned purpose, the present application also provides an electronic device, which includes: at least one processor, a memory and an input and output unit; wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the multi-agent proximal strategy optimization drone game method with adaptive iterative resolution provided by any of the aforementioned embodiments.

[0025] The embodiment of the present application proposes a method, device, medium and equipment for optimizing drone games with a multi-agent proximal strategy of adaptive iterative resolution. The method constructs a kinematic model and a control model of a drone that introduces a resolution reduction factor parameter in a drone game environment; constructs a deep reinforcement learning algorithm for drone cluster games, and obtains environmental state information based on the interaction between the deep reinforcement learning algorithm for drone cluster games and the kinematic model and the control model; trains the deep reinforcement learning algorithm for drone cluster games based on the environmental state information to obtain the actions of the drone; and uses the actions of the drone to update the kinematic model, the control model and the preset reward function to obtain the reward function value of this round, and subtracts the reward function value of this round from the reward function values ​​of the previous multiple rounds to obtain the difference of each reward function; and determines whether the absolute value of the difference of each reward function is less than a preset threshold. value, it is judged that the deep reinforcement learning algorithm for drone swarm game has converged, and the preset resolution parameter update algorithm is used to update the resolution reduction factor parameter to update the kinematic model and the control model; the above process is iteratively executed until the drone swarm game deep reinforcement learning algorithm with adaptive iterative resolution is obtained, and the drone swarm game deep reinforcement learning algorithm with adaptive iterative resolution is deployed in each drone intelligent agent participating in the game, and the intelligent agent is used to guide the drone to make the optimal decision in the game environment. The method of this application significantly improves the training efficiency and convergence speed by dynamically adjusting the resolution of the strategy update, and can also achieve good training results under sparse rewards, so that the drone group can quickly adapt to the task requirements in a complex dynamic environment, and promotes the application and development of deep reinforcement learning in drone swarm intelligent games. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a flow chart of a multi-agent proximal strategy optimization drone game method with adaptive iterative resolution according to the present invention;

[0027] Figure 2 It is an iterative function curve of the resolution reduction factor r of the present invention;

[0028] Figure 3 The present invention is based on the MAPPO deep reinforcement learning algorithm game framework with adaptive iterative resolution;

[0029] Figure 4 It is the kinematic model of a single UAV in two-dimensional space of the present invention;

[0030] Figure 5 It is the relative situation model of the two-dimensional UAV in the present invention.

[0031] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0032] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.

[0033] To address the shortcomings of existing technologies, this paper proposes an Adaptive Iterative Resolution (AIR) method, combined with the MAPPO algorithm to form the AIR-MAPPO algorithm, which effectively addresses this difficulty. By dynamically adjusting the resolution of policy updates, this method significantly improves training efficiency and convergence speed. It also achieves good training results under sparse rewards, enabling drone swarms to quickly adapt to mission requirements in complex dynamic environments, further promoting the application and development of deep reinforcement learning in drone swarm intelligent games.

[0034] Reference Figure 1 , Figure 1 This is a flowchart of a multi-agent proximal strategy optimization drone game method with adaptive iterative resolution provided in one embodiment of the present application. The multi-agent proximal strategy optimization drone game method with adaptive iterative resolution may include the following execution process:

[0035] S10. In the UAV game environment, construct the kinematic model and control model of the UAV by introducing the resolution reduction factor parameter.

[0036] In an embodiment of the present application, step S10 may include the following execution process:

[0037] S101. Construct a kinematic model and a control model of the UAV based on the linear velocity and the yaw angular velocity of the UAV, wherein the control model includes the linear acceleration of the UAV and the yaw angular acceleration of the UAV.

[0038] The kinematic and control model of the drone is used to describe the position changes and attitude adjustments of the drone itself in two-dimensional space. This invention focuses on drone trajectory planning rather than underlying flight control, so the drone modeling method can only represent its position and game-related attitude.

[0039] Figure 4 is the kinematic model of a single UAV in two-dimensional space. The kinematic model formula of a two-dimensional UAV is as follows:

[0040]

[0041] Where x and y represent the position coordinates of the drone in two-dimensional space, θ represents the deflection angle of the drone, and v represents the linear velocity of the drone. ω represents the deflection angular velocity of the drone, which controls the change in the deflection angle of the drone. Both ‖v‖ and ω are limited, and their upper limit is V max ={‖v‖ max ,ω m}, the lower limit is Vmin ={‖v‖ min ,-ω m}.

[0042] The UAV control model is used to control the UAV motion in two-dimensional space. Its formula is as follows:

[0043]

[0044] Where [a0, a1] are control variables, controlling the change in the drone's position and attitude. a0 represents the drone's linear acceleration, controlling the change in the drone's velocity. a1 is the drone's angular acceleration, controlling the change in the drone's angular velocity. The drone game scenario is discrete, and variables can be changed in single steps: increase, decrease, or remain unchanged. Both increases and decreases have fixed values.

[0045] S102: Multiply the upper limit of the linear velocity, the upper and lower limits of the deflection angular velocity, the linear acceleration of the UAV, and the deflection angular acceleration of the UAV by the resolution reduction factor to obtain the kinematic model and control model of the UAV with the resolution reduction factor parameter introduced.

[0046] The resolution is modified by introducing the kinematic model through the influence factor r. The influence factor r changes the scene resolution mainly by modifying the two-dimensional space UAV kinematic model. The UAV's acceleration a0, a1, and the UAV's linear velocity v upper limit ‖v‖ max , the upper and lower limits of the deflection angle angular velocity ω ±ω m Multiply them by the resolution reduction factor r to reduce the resolution by r times. max , V′ min , and its relationship with the original variables is shown below. It is worth noting that the resolution in this application can represent the resolution reduction factor:

[0047] a i ′=r·a i (i=1,2)

[0048] V′ max =(r·||v|| max , r·ω m )

[0049] V′ min =||v| min , -r·ω m )

[0050] S20. Construct a deep reinforcement learning algorithm for drone swarm game, and obtain environmental state information based on the interaction between the deep reinforcement learning algorithm for drone swarm game and the kinematic model and control model. Train the deep reinforcement learning algorithm for drone swarm game based on the environmental state information to obtain the actions of the drones, and use the actions of the drones to update the kinematic model, control model and preset reward function to obtain the reward function value of this round. Subtract the reward function value of this round from the reward function values ​​of previous rounds to obtain the difference values ​​of each reward function.

[0051] In one embodiment of the present application, after the UAV swarm game deep reinforcement learning algorithm is trained based on the environmental state information to obtain the UAV motions, and the UAV motions are used to update the kinematic model, the control model, and the preset reward function, the method further includes:

[0052] The relative situation between any two drones is determined based on the drone relative situation model, where the relative situation includes the relative position vector between the any two drones, a first angle between the velocity vector of the target drone among the two drones and the relative position vector, and a second angle between the velocity vector of the drone and the relative position vector.

[0053] in, Figure 5 is a two-dimensional UAV relative situation model, which is used to describe the relative situation between any two UAVs in two-dimensional space. Its formula is as follows:

[0054] D=[x b -x r ,y b -y r ]

[0055]

[0056] Among them, (x r ,y r ) is the position coordinate of the red drone, (x b ,y b ) is the position coordinate of the blue drone, and D is the position vector of the red drone relative to the blue drone. r 、v b Represent the velocity vectors of the red and blue drones respectively. r Represents the angle between the velocity vector of the red drone and its position vector D. b Represents the angle between the blue drone's velocity vector and its position vector -D.

[0057] It is understandable that in this application, a red drone may represent a friendly drone, a blue drone may represent an enemy drone, or vice versa, without making too many restrictions here.

[0058] Determine the maximum attack distance and maximum attack angle of the UAV based on the UAV attack model.

[0059] If the first angle / second angle of the target UAV is less than the attack angle of the friendly UAV, and the relative position vector of the target UAV is less than the maximum attack distance of the friendly UAV, it is determined that the target UAV is destroyed by the friendly UAV; otherwise, it is determined that the target UAV is not destroyed by the friendly UAV.

[0060] Among them, the two-dimensional air combat environment model between our two UAVs also includes the UAV attack model. The attack range of the UAV in two-dimensional space is a sector. d is the radius length of the sector, that is, the maximum attack distance, and q is the angle between the outer edge of the sector and the UAV velocity vector, that is, the maximum attack angle. When the opponent's UAV enters the sector range, it is determined to be destroyed, that is, q r <q and D < d

[0061] Obtain the first quantity of the friendly UAVs that destroy the target UAV and the second quantity of the friendly UAVs that are destroyed by the target UAV, and process the first quantity and the second quantity based on a preset reward function to obtain the value of the reward function for this round.

[0062] Among them, in an embodiment of the present application, the preset reward function can be:

[0063] <l

[0064] Among them, i represents the number of friendly UAVs that destroy the target UAV. j represents the number of friendly UAVs that are destroyed by the target UAV, and r represents the resolution reduction multiple.

[0065] It should be noted that the reward function in this application uses sparse rewards. The sparse rewards are mainly judged according to the survival status of the UAVs in the environment, that is, destroying enemy UAVs and friendly UAVs being destroyed. In addition, there is also the value size of the resolution reduction multiple r of the current scene. When the first enemy UAV is destroyed, the reward is +1, when the second enemy UAV is destroyed, the reward is +2, and so on. When the i-th enemy UAV is destroyed, the reward is +i. When the first friendly UAV is destroyed, the reward is -1, when the second friendly UAV is destroyed, the reward is -2, and so on. When the j-th friendly UAV is destroyed, the reward is -j. On this basis, the reward value is divided by the resolution reduction multiple r to obtain the final reward.

[0066] Therefore, to maximize the reward in the game, our UAVs must actively destroy enemy UAVs while trying to avoid losses to our own UAVs.

[0067] In an embodiment of the present application, the process of obtaining the actions of the UAV swarm through the game deep reinforcement learning algorithm based on the environmental state information may include the following execution process:

[0068] Each UAV obtains environmental state information and game reward function values ​​from the kinematic model and control model, and processes the environmental state information and game reward function values ​​based on its own Actor network to obtain output actions.

[0069] The output action is input into the Critic network to obtain the feedback variable, and the Actor network is optimized based on the feedback variable.

[0070] Figure 3 This is the MAPPO deep reinforcement learning algorithm game framework based on adaptive iterative resolution. There are n drones in the scene. At time t, each drone is based on the environment state s t Perform observations, and the i-th UAV obtains the observation value At the same time, the game reward value Under the influence of i π, calculated output action All drone actions The output acts on the environment to cause state changes, and at the same time it is input into the centralized critic network Criticπ, the feedback variable Optimize Actor network.

[0071] In the present invention, the observation value of the i-th UAV is where q is the angle between the position vector of the agent and the velocity vector of the agent, and q is the angle between the position vector of the agent and the velocity vector of the agent. r , the angle q between the position vector and the velocity vector of other agents b Drone Action It controls the linear acceleration a0 and the deflection angular acceleration a1, and then updates its own state based on the two-dimensional kinematic formula of the UAV to realize the simulation.

[0072] It should be noted that when obtaining the action output by Actor The agent then multiplies the action and action cap by the resolution reduction factor r, reducing the scene resolution by a factor of r. Furthermore, the agent receives feedback on changes in the reward value. When the reward value Rw changes slowly, training reaches convergence and the scene resolution is updated to continue training.

[0073] The process of determining whether to update the resolution and continue training is as follows:

[0074] S30. If the absolute values ​​of the differences between the reward functions are all less than a preset threshold, it is determined that the UAV cluster game deep reinforcement learning algorithm has converged, and the preset resolution parameter update algorithm is used to update the resolution reduction factor parameter to update the kinematic model and the control model.

[0075] In a specific game scenario, the current reward value Rw is compared with the previous m reward values ​​[Rw1, Rw2, ..., Rw m ] respectively make the difference and get [ΔRw1, ΔRw2,…, ΔRw m ], when [ΔRw1,ΔRw2,…,ΔRw m ] are all less than a set threshold ΔRw0, the overall upward trend of the reward value Rw stops, the training reaches convergence, and the updated resolution is reduced by a factor of r.

[0076] In one embodiment of the present application, updating the resolution reduction factor parameter using a preset resolution parameter updating algorithm to update the kinematic model and the control model may include the following execution process:

[0077] At the beginning of training, set the resolution parameter to the minimum resolution reduction factor.

[0078] During the training process, the minimum resolution reduction factor is continuously increased until the resolution parameter reaches the normal resolution.

[0079] Figure 2 It is the iterative function curve of the resolution reduction factor of the present invention. Figure 2 , the iterative function for updating the resolution reduction factor is The conditions for selecting the iterative function are: when the resolution reduction factor r takes the initial value r0 (r0>1), it is necessary to make r decrease monotonically from high to low when substituting it into the iterative function, and finally converge to r m = 1. The number of convergences during the iteration should be moderate, neither too many nor too few. The update amplitude of each iteration should be appropriate, and the update amplitude should be appropriately increased when r is large, and appropriately reduced when r is small (close to 1).

[0080] If you want the r value to converge from high to low to r m =1, requiring iterative function Passing through the point (1,1). However, in reality, if the iterative function just passes through the fixed point (1,1), when the iteration is close to 1, it will iterate several times in 1+ε(ε→0). In reality, it is hoped that the r iteration will be directly updated to 1. In this case, the optimization method is changed to when it is close to 1. (The present invention sets the condition r<1.1), when the value is greater than 1.1, avoid multiple small iterations, and when the function value is less than 1.1, take This avoids multiple small iterations.

[0081] Iteration Function The choice of linear function Starting from, its derivative function is First, the iterative function needs to satisfy the convergence condition of the function, that is, the absolute value of the derivative value at the (1,1) point Secondly, the iterative function must satisfy r to decrease monotonically from high to low, that is, the iterative function is monotonically increasing.

[0082] The size of k determines the amplitude of the iterative update. When r is large, the update amplitude can be increased appropriately. When r is small (close to 1), the update amplitude can be reduced appropriately. k can be set as a function k(r) that follows the change of r. Combined with the above convergence conditions With the monotonically increasing condition k(r) can be set as follows:

[0083]

[0084] Where r m = 1. The derivative k′(r) of k(r) is given by:

[0085]

[0086] It has been verified that when r≥1, k′(r)>0, the minimum value of k(r) is k(1)=0, and when r→+∞, lim r→+∞ k(r)=1, that is, k(r)∈(0,1), which satisfies the above conditions.

[0087] Since the slope k(r) changes with r, and the iterative function passes through the fixed point (1,1), then The vertical intercept b of is also a function b(r) that changes with r. Substitute You can get:

[0088]

[0089] Then the iterative function is:

[0090]

[0091] In particular, in actual situations, when x=r, we have:

[0092]

[0093] It has been verified that the iterative function Meet the above conditions. Function curve is as follows Figure 3 As shown, the red arrow represents that when r0>1, the iteration can eventually converge to rm =1.

[0094] S40. Iterate the above process until a deep reinforcement learning algorithm for drone swarm game with adaptive iterative resolution is obtained, deploy the deep reinforcement learning algorithm for drone swarm game with adaptive iterative resolution in each drone agent participating in the game, and use the agent to guide the drone to make the optimal decision in the game environment.

[0095] In summary, the present invention proposes a method of adaptive iterative resolution, which takes a low-resolution scene as the starting point, and adaptively and gradually improves the scene resolution during training according to the improvement of the training effect, and uses the strategy learned in the low-resolution scene as an experience pool to help train the higher-resolution scene, until finally learning the strategy to defeat the enemy in the normal-resolution scene. This overcomes the problem that ordinary deep reinforcement learning algorithms are difficult to converge on strategies and cannot be trained in drone games due to the large number of parameters, a lot of uncertain information, and sparse rewards. Compared with ordinary strategy optimization algorithms suitable for drone group games, the drone swarm game deep reinforcement learning algorithm with adaptive iterative resolution proposed in the present invention can find the strategy to defeat the enemy more quickly and avoid local optimality during the training stage, greatly improving the training efficiency, and the reward values ​​obtained from the training results are higher and less volatile.

[0096] Based on the above method embodiments, the present application also provides a multi-agent proximal strategy optimization drone game device with adaptive iterative resolution. The device may include a model construction module, a reward calculation module, a judgment module, and a decision module. The model construction module is used to construct a kinematic model and control model of a drone in a drone game environment that incorporates a resolution reduction factor parameter. The reward calculation module is used to construct a deep reinforcement learning algorithm for the drone swarm game. Based on the interaction between the deep reinforcement learning algorithm for the drone swarm game and the kinematic model and control model, the reward calculation module is used to obtain environmental state information. The deep reinforcement learning algorithm for the drone swarm game is trained based on the environmental state information to obtain drone actions. The kinematic model, control model, and preset reward function are then used to update the kinematic model, control model, and preset reward function to obtain the current round reward function value. The current round reward function value is then subtracted from the reward function values ​​of previous rounds to obtain the reward function difference values. The judgment module is used to determine if the absolute value of each reward function difference value is less than a preset threshold, thereby determining that the deep reinforcement learning algorithm for the drone swarm game has converged. The resolution reduction factor parameter is then updated using a preset resolution parameter update algorithm to update the kinematic model and control model. The decision-making module is used to iteratively execute the above process until a deep reinforcement learning algorithm for drone swarm game with adaptive iterative resolution is obtained. The deep reinforcement learning algorithm for drone swarm game with adaptive iterative resolution is deployed in each drone agent participating in the game, and the agent is used to guide the drone to make the optimal decision in the game environment.

[0097] Based on the above method embodiments, the present application also provides a computer-readable storage medium, characterized in that it includes instructions that, when run on a computer, enable the computer to execute the multi-agent proximal strategy optimization drone game method with adaptive iterative resolution described in any one of the previous method embodiments.

[0098] Based on the above method embodiments, the present application further provides an electronic device, characterized in that the electronic device includes: at least one processor, a memory, and an input / output unit. The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the multi-agent proximal strategy optimization drone game method with adaptive iterative resolution described in any of the above method embodiments.

[0099] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes a plurality of computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes multiple available media integration. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0100] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0101] Each embodiment in this specification is described in a related manner. Similar portions between the embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are described briefly because they are generally similar to the method embodiments. For related portions, reference can be made to the description of the method embodiments.

[0102] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A multi-agent proximal strategy optimization drone game method with adaptive iterative resolution, characterized by: include: In the UAV game environment, the kinematic model and control model of the UAV are constructed by introducing the resolution reduction factor parameter; Construct a deep reinforcement learning algorithm for drone swarm game, and obtain environmental state information based on the interaction between the deep reinforcement learning algorithm and the kinematic model and control model. Train the deep reinforcement learning algorithm for drone swarm game based on the environmental state information to obtain the drone's actions. Use the drone's actions to update the kinematic model, control model, and preset reward function to obtain the reward function value of this round. Subtract the reward function value of this round from the reward function values ​​of previous rounds to obtain the difference between each reward function. If the absolute value of the difference between each reward function is less than the preset threshold, the deep reinforcement learning algorithm for the drone swarm game is considered to have converged, and the preset resolution parameter update algorithm is used to update the resolution reduction factor parameter to update the kinematic model and control model; The above process is iterated until a deep reinforcement learning algorithm for drone swarm games with adaptive iterative resolution is obtained. The deep reinforcement learning algorithm for drone swarm games with adaptive iterative resolution is deployed in each drone agent participating in the game, and the agent is used to guide the drone to make the optimal decision in the game environment.

2. The multi-agent proximal strategy optimization drone game method with adaptive iterative resolution according to claim 1, characterized in that: The kinematic model and control model of the UAV that introduces the resolution reduction factor parameter are constructed, including: The kinematic model and control model of the UAV are constructed based on the linear velocity and yaw angular velocity of the UAV, wherein the control model includes the linear acceleration of the UAV and the yaw angular acceleration of the UAV; The upper limit of linear velocity, the upper and lower limits of angular velocity of deflection angle, the linear acceleration of UAV and the deflection angular acceleration of UAV are multiplied by the resolution reduction factor to obtain the kinematic model and control model of the UAV with the resolution reduction factor parameter introduced.

3. The multi-agent proximal strategy optimization drone game method with adaptive iterative resolution according to claim 1, characterized in that: After the UAV swarm game deep reinforcement learning algorithm is trained based on the environmental state information to obtain the UAV's motion, and the UAV's motion is used to update the kinematic model, the control model, and the preset reward function, the method further includes: Determining the relative situation between any two drones based on the drone relative situation model, wherein the relative situation includes a relative position vector between the any two drones, a first angle between a velocity vector of a target drone among the two drones and the relative position vector, and a second angle between the velocity vector of the drone and the relative position vector; Determine the maximum attack distance and maximum attack angle of the drone based on the drone attack model; If the target drone's first angle / second angle is smaller than the friendly drone's attack angle, and the target drone's relative position vector is smaller than the friendly drone's maximum attack range, the target drone is considered destroyed by the friendly drone; otherwise, it is considered not destroyed by the friendly drone. Obtain a first number of target drones destroyed by the own drone and a second number of the own drones destroyed by the target drone, process the first number and the second number based on a preset reward function, and obtain a reward function value for this round.

4. The multi-agent proximal strategy optimization drone game method with adaptive iterative resolution according to claim 1, characterized in that: The method of training the UAV swarm game deep reinforcement learning algorithm based on the environmental state information to obtain the UAV actions includes: Each UAV obtains environmental state information and game reward function values ​​from the kinematic model and control model, and processes the environmental state information and game reward function values ​​based on its own Actor network to obtain output actions; The output action is input into the Critic network to obtain the feedback variable, and the Actor network is optimized based on the feedback variable.

5. The multi-agent proximal strategy optimization drone game method with adaptive iterative resolution according to claim 1, characterized in that: The preset reward function is: Among them, i represents the number of target drones destroyed by our drones; j represents the number of our drones destroyed by target drones, and r represents the resolution reduction factor.

6. The multi-agent proximal strategy optimization drone game method with adaptive iterative resolution according to claim 1, characterized in that: The update expression of the resolution reduction factor r is:

7. The multi-agent proximal strategy optimization drone game method with adaptive iterative resolution according to claim 1, characterized in that: The method of updating the resolution reduction factor parameter by using a preset resolution parameter updating algorithm to update the kinematic model and the control model includes: At the beginning of training, set the resolution parameter to the minimum resolution reduction factor; During the training process, the minimum resolution reduction factor is continuously increased until the resolution parameter reaches the normal resolution.

8. A multi-agent proximal strategy optimization drone game device with adaptive iterative resolution, characterized by: include: A model building module is used to construct the kinematic model and control model of the drone in the drone game environment, which introduces the resolution reduction factor parameter; The reward calculation module is used to construct a deep reinforcement learning algorithm for drone swarm games, and obtain environmental state information based on the interaction between the deep reinforcement learning algorithm for drone swarm games and the kinematic model and control model. The deep reinforcement learning algorithm for drone swarm games is trained based on the environmental state information to obtain the drone's actions. The kinematic model, control model, and preset reward function are then used to update the drone's actions to obtain the reward function value of this round. The reward function value of this round is then subtracted from the reward function values ​​of the previous rounds to obtain the difference between the reward functions. A judgment module is used to determine if the absolute value of the difference between each reward function is less than a preset threshold, then determine that the drone swarm game deep reinforcement learning algorithm has converged, and use a preset resolution parameter update algorithm to update the resolution reduction factor parameter to update the kinematic model and control model; The decision-making module is used to iteratively execute the above process until a deep reinforcement learning algorithm for drone swarm game with adaptive iterative resolution is obtained. The deep reinforcement learning algorithm for drone swarm game with adaptive iterative resolution is deployed in each drone agent participating in the game, and the agent is used to guide the drone to make the optimal decision in the game environment.

9. A computer-readable storage medium, characterized in that The invention includes instructions, which, when run on a computer, enable the computer to execute the multi-agent proximal strategy optimization drone game method with adaptive iterative resolution as described in any one of claims 1 to 7.

10. An electronic device, characterized in that: The electronic device comprises: at least one processor, memory, and input-output unit; The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the multi-agent proximal strategy optimization drone game method with adaptive iterative resolution according to any one of claims 1 to 7.

Citation Information

Cited By

  • Training method based on dynamic adjustment reward mechanism

    CN120975268A