Training method based on dynamic adjustment reward mechanism
By using a three-stage training course with dynamically adjusted reward mechanisms and employing a multi-agent proximal policy optimization algorithm to adjust reward configurations, the training instability problem caused by fixed reward weights in existing technologies is solved, enabling rapid improvement of agent capabilities and shortening of training time.
Patent Information
- Application Number
- CN202511521637.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-10-23
AI Technical Summary
Existing course-based training methods use fixed reward weights, which cannot adjust the reward ratio according to task complexity or the learning progress of the agent. This leads to unstable policy updates and fails to provide effective guidance when the task difficulty increases, affecting the stability of the training process.
A training method based on a dynamic reward adjustment mechanism is adopted. The reward configuration, including attack, defense and balance reward parameters, is dynamically adjusted through a multi-agent proximal policy optimization algorithm to form a three-stage training course, which guides the agent to improve its comprehensive adversarial capabilities from basic attack capabilities to defensive skills.
By dynamically adjusting the reward mechanism, the stability and efficiency of the training process were improved, the training time was shortened, and the rapid development of the agent from basic attack to comprehensive adversarial capabilities was achieved.
Smart Images

Figure CN120975268B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) technology, and in particular to a training method based on a dynamically adjusted reward mechanism. Background Technology
[0002] Current training methods based on course learning use fixed reward weights, which cannot adjust the reward ratio according to task complexity or the learning progress of the agent. The core of course learning is to dynamically adjust the training difficulty, but fixed reward weights limit the adaptive adjustment capability of the reward. Furthermore, when the task difficulty increases, fixed reward weights cannot provide effective guidance, resulting in oscillations in policy updates and causing instability in the training process. Summary of the Invention
[0003] This invention provides a training method based on a dynamically adjusted reward mechanism to address the technical problems of existing course-based training methods that use fixed reward weights, which cannot adjust the reward ratio according to task complexity or the learning progress of the agent, and whose fixed reward weights cannot provide effective guidance when the task difficulty increases.
[0004] This invention provides a training method based on a dynamically adjusted reward mechanism, applicable to multi-UAV adversarial learning, the method comprising:
[0005] Acquire a first-party drone and a second-party drone; the first-party drone is any drone among the multiple drones; the second-party drone is any drone among the multiple drones other than the first-party drone;
[0006] When the first-party drone and the second-party drone are learning attack courses, the first parameter of the attack reward configuration in the first-party drone and the second-party drone is determined according to the preset multi-agent near-end policy optimization algorithm.
[0007] When the first-party drone and the second-party drone are learning defense courses, the second parameter of the defense reward configuration in the first-party drone and the second-party drone is determined according to the multi-agent near-end policy optimization algorithm.
[0008] When the first drone and the second drone are learning adversarial courses, the third parameter of the balanced reward configuration of the first drone is determined according to the multi-agent proximal policy optimization algorithm.
[0009] The target reward configuration parameters are determined based on the first parameter, the second parameter, and the third parameter; the target reward configuration parameters are used to learn and train the multiple UAVs.
[0010] In some implementations, determining the first parameter for the attack reward configuration of the first-party UAV and the second-party UAV according to a preset multi-agent proximal policy optimization algorithm includes:
[0011] The first reward parameter related to the attack capability of the first party drone and the second reward parameter related to the defense capability of the second party drone are determined according to the multi-agent proximal policy optimization algorithm.
[0012] The first parameter is determined based on the first reward parameter and the second reward parameter.
[0013] In some implementations, obtaining the value parameters of the target task includes:
[0014] Obtain the importance parameter of the target task and the success rate parameter of the first resource of the first UAV platform in executing the target task;
[0015] The importance parameter and the success rate parameter are used as the value parameter.
[0016] In some implementations, determining the first reward parameter related to the attack capability of the first-party drone and the second reward parameter related to the defense capability of the second-party drone based on the multi-agent proximal policy optimization algorithm includes:
[0017] The first policy parameters, first value network parameters, first course configuration parameters, and first environment configuration parameters of the multi-agent proximal policy optimization algorithm are obtained according to the multi-agent proximal policy optimization algorithm.
[0018] Based on the first strategy parameters, the first value network parameters, the first course configuration parameters, and the first environment configuration parameters, determine the attitude advantage reward parameters, altitude reward parameters, event triggering reward parameters, and launch penalty reward parameters of the first UAV.
[0019] The attitude advantage reward parameter, the altitude reward parameter, and the event triggering reward parameter are used as the first reward parameter, and the launch penalty reward parameter is used as the second reward parameter.
[0020] In some implementations, determining the first parameter based on the first reward parameter and the second reward parameter includes:
[0021] Obtain the first weighting coefficient of the attitude advantage reward parameter, the second weighting coefficient of the altitude reward parameter, the third weighting coefficient of the event triggering reward parameter, and the fourth weighting coefficient of the launch penalty reward parameter;
[0022] The first parameter is determined based on the first weighting coefficient, the second weighting coefficient, the third weighting coefficient, the fourth weighting coefficient, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter.
[0023] In some implementations, determining the second parameter of the defense reward configuration in the first and second party drones according to the multi-agent proximal policy optimization algorithm includes:
[0024] The third reward parameter related to the action space capability of the first drone and the fourth reward parameter related to the active pursuit attack of the second drone are determined according to the multi-agent proximal policy optimization algorithm.
[0025] The second parameter is determined based on the third reward parameter and the fourth reward parameter.
[0026] In some implementations, determining the third reward parameter related to the action space capability of the first UAV and the fourth reward parameter related to the active pursuit attack of the second UAV based on the multi-agent proximal policy optimization algorithm includes:
[0027] The second policy parameters, second value network parameters, second curriculum configuration parameters, and second environment configuration parameters of the multi-agent proximal policy optimization algorithm are obtained according to the multi-agent proximal policy optimization algorithm.
[0028] Based on the second strategy parameters, the second value network parameters, the second course configuration parameters, and the second environment configuration parameters, determine the missile evasion reward parameters, attitude advantage reward parameters, altitude reward parameters, event triggering reward parameters, and the launch penalty reward parameters of the first party's UAV.
[0029] The missile evasion reward parameter, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter are used as the third reward parameter, and the launch penalty reward parameter is used as the fourth reward parameter.
[0030] In some implementations, determining the second parameter based on the third reward parameter and the fourth reward parameter includes:
[0031] The fifth weighting coefficient of the missile evasion reward parameter, the sixth weighting coefficient of the attitude advantage reward parameter, the seventh weighting coefficient of the altitude reward parameter, the eighth weighting coefficient of the event triggering reward parameter, and the ninth weighting coefficient of the launch penalty reward parameter are obtained.
[0032] The second parameter is determined based on the fifth weighting coefficient, the sixth weighting coefficient, the seventh weighting coefficient, the eighth weighting coefficient, the ninth weighting coefficient, the missile evasion reward parameter, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter.
[0033] In some implementations, determining the third parameter of the balanced reward configuration of the first-party UAV based on the multi-agent proximal policy optimization algorithm includes:
[0034] The fifth, sixth, and seventh reward parameters related to the selection of attack or defense actions by the first party UAV are determined based on the multi-agent proximal policy optimization algorithm.
[0035] The third parameter is determined based on the fifth reward parameter, the sixth reward parameter, and the seventh reward parameter.
[0036] In some implementations, determining the fifth reward parameter related to the selection of attack or defense actions by the first-party UAV based on the multi-agent proximal policy optimization algorithm includes:
[0037] The third policy parameters, third value network parameters, third course configuration parameters, and third environment configuration parameters of the multiple UAVs are obtained according to the multi-agent proximal policy optimization algorithm.
[0038] Based on the third strategy parameters, the third value network parameters, the third course configuration parameters, and the third environment configuration parameters, the attitude advantage reward parameters, altitude reward parameters, and launch penalty reward of the first party UAV are determined.
[0039] Determine the first weighting factor of the attitude advantage reward parameter, the second weighting factor of the altitude reward parameter, and the third weighting factor of the launch penalty reward;
[0040] The fifth reward parameter is determined based on the attitude advantage reward parameter, the altitude reward parameter, the launch penalty reward, the first weighting factor, the second weighting factor, and the third weighting factor.
[0041] In some implementations, the sixth reward parameter includes a missile evasion reward parameter; the seventh reward parameter includes an event triggering reward parameter; determining the third parameter based on the fifth, sixth, and seventh reward parameters includes:
[0042] Obtain the first weight value of the fifth reward parameter, the second weight value of the missile evasion reward parameter, and the third weight value of the event triggering reward parameter;
[0043] The third parameter is determined based on the first weight value, the second weight value, the third weight value, the fifth reward parameter, the sixth reward parameter, and the seventh reward parameter.
[0044] This invention also provides a training device based on a dynamically adjusted reward mechanism, configured for use in adversarial learning with multiple unmanned aerial vehicles (UAVs). The device includes:
[0045] The acquisition unit is used to acquire a first-party drone and a second-party drone; the first-party drone is any drone among the multiple drones; the second-party drone is a drone among the multiple drones other than the first-party drone.
[0046] The first determining unit is used to determine the first parameter of the attack reward configuration in the first drone and the second drone according to a preset multi-agent near-end strategy optimization algorithm when the first drone and the second drone are learning an attack course.
[0047] The second determining unit is used to determine the second parameter of the defense reward configuration in the first drone and the second drone according to the multi-agent near-end policy optimization algorithm when the first drone and the second drone are learning a defense course.
[0048] The third determining unit is used to determine the third parameter of the balanced reward configuration of the first drone based on the multi-agent proximal policy optimization algorithm when the first drone and the second drone are learning adversarial courses.
[0049] The fourth determining unit is used to determine the target reward configuration parameters based on the first parameter, the second parameter, and the third parameter; the target reward configuration parameters are used to learn and train the multiple UAVs.
[0050] This invention provides an electronic device, the device comprising: a processor and a memory for storing a computer program capable of running on the processor, wherein, when the processor runs the computer program, it performs the steps of any of the methods described above.
[0051] This invention provides a storage medium storing a computer program; when the computer program is executed by a processor, it implements the steps of any of the methods described above.
[0052] This invention provides a training method based on a dynamically adjusted reward mechanism. The method includes: using multiple drones for adversarial learning; the method includes: acquiring a first-party drone and a second-party drone; the first-party drone being any drone among the multiple drones; the second-party drone being any drone among the multiple drones other than the first-party drone; when the first-party drone and the second-party drone are learning an attack course, determining a first parameter for the attack reward configuration of the first-party drone and the second-party drone according to a preset multi-agent proximal policy optimization algorithm; when the first-party drone and the second-party drone are learning a defense course, determining a second parameter for the defense reward configuration of the first-party drone and the second-party drone according to the multi-agent proximal policy optimization algorithm; when the first-party drone and the second-party drone are learning an adversarial course, determining a third parameter for the balanced reward configuration of the first-party drone according to the multi-agent proximal policy optimization algorithm; and determining a target reward configuration parameter based on the first parameter, the second parameter, and the third parameter; the target reward configuration parameter is used for learning and training the multiple drones. This involves a three-stage training course based on a dynamically adjusted reward mechanism, which guides the agent from basic attack capabilities to mastering defensive skills, and ultimately develops it into a comprehensive adversarial agent integrating offense and defense. At the same time, a parallel environment training mechanism is proposed to accelerate the convergence speed of multi-UAV adversarial training and significantly reduce training time. Attached Figure Description
[0053] Figure 1 A flowchart illustrating a training method based on a dynamically adjusted reward mechanism, provided in an embodiment of the present invention;
[0054] Figure 2 This is a schematic diagram of the parallel environment training architecture according to an embodiment of this application;
[0055] Figure 3 This is a schematic diagram illustrating the framework of the training process in an embodiment of this application;
[0056] Figure 4 A schematic diagram of a training device based on a dynamically adjusted reward mechanism provided in an embodiment of the present invention;
[0057] Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0059] The specific technical features described in the various embodiments in the detailed implementation can be combined in various ways without contradiction. For example, different implementation methods can be formed by combining different specific technical features. In order to avoid unnecessary repetition, the various possible combinations of the specific technical features in this invention will not be described separately.
[0060] It should also be noted that, in order to avoid obscuring the present invention with unnecessary details, only the structures and / or processing steps closely related to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.
[0061] Additionally, it should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. In the following description, the terms "first," "second," etc., are used merely to distinguish different objects and do not indicate any similarity or connection between them. It should be understood that the directional descriptions such as "above," "below," "inside," and "outside" refer to the orientation under normal use conditions.
[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the specific technical solutions of the invention will be further described in detail below with reference to the accompanying drawings of the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0063] Traditional drone decision-making methods primarily rely on manually preset rules, which are ill-suited to the rapidly changing adversarial situations in complex and dynamic drone combat environments. The development of artificial intelligence (AI) technology has brought new possibilities to drone adversarial decision-making. With its powerful adaptability and autonomous learning capabilities, AI, requiring no extensive specialized background knowledge, has become an effective approach to solving drone decision-making problems in complex adversarial environments. Currently, various drone adversarial decision-making methods exist, which can be categorized into three types based on their approach: game theory-based methods, optimization theory-based methods, and data-driven methods.
[0064] Game Theory-Based UAV Adversarial Decision-Making Methods: Game theory is a mathematical model that studies the interaction and strategy selection among multiple rational decision-making agents. It reveals the optimal decision path by analyzing the objective functions, strategy spaces, and information structures of each party. In UAV adversarial environments, UAVs and their opponents constitute a typical non-cooperative game relationship; therefore, modeling and solving UAV adversarial decision-making problems using game theory is an important research direction. One approach is to discretize UAV adversarial maneuvers into a maneuver library using matrix game theory. All possible maneuver combinations for both sides are represented by matrices, and then the aircraft motion equations are solved through numerical integration to obtain the optimal decision sequence. Another approach proposes a multi-stage influence graph game framework for modeling pilot maneuver decision-making in one-to-one adversarial situations. This framework integrates aircraft dynamics, pilot preferences, and uncertainties through a visualized influence graph, and uses a moving time-domain control method to solve the feedback Nash equilibrium of the dynamic game in stages, ultimately generating the optimal maneuver sequence. Finally, some researchers have proposed a more advanced automatic maneuver generation algorithm based on differential game theory. This algorithm follows a hierarchical decision structure and evaluates the UAV's air superiority by calculating the scoring function matrix and analyzing the aircraft's relative geometric position, relative distance, and speed.
[0065] UAV Adversarial Decision-Making Methods Based on Optimization Theory: Optimization theory refers to methods for finding the extrema of an objective function under constraints when solving decision-making problems in complex systems. Since the 1980s, intelligent optimization methods inspired by natural phenomena have been widely used. In UAV adversarial scenarios, maneuver decision-making methods based on optimization theory transform adversarial strategy selection into a multi-objective hybrid optimization problem, mainly including genetic algorithms and ant colony algorithms. Some researchers have developed a UAV adversarial maneuver rule learning method based on genetic algorithms, which discovers and optimizes UAV adversarial strategies through genetic algorithms. This method can automatically generate and evaluate new adversarial strategy rules, providing decision-making references for human experts. Research has verified that, under the condition of mutual learning, both aircraft can develop objectively interesting strategies. Rodlin et al. proposed a heuristic ant colony algorithm that integrates adversarial knowledge. By establishing a missile-target allocation optimization model with the goal of minimizing the expected residual threat of the enemy swarm, UAV adversarial decision-making is transformed into a combinatorial optimization problem. A local heuristic search strategy based on cooperative adversarial rules is designed and combined with the ant colony algorithm to improve global search efficiency.
[0066] Data-Driven UAV Adversarial Decision-Making Methods: With the improvement of computing power and the development of deep learning theory, artificial intelligence-based adversarial decision-making methods have shown great potential. These methods, by simulating human cognitive processes, learning from historical data, or continuously optimizing through interaction with the environment, can handle high-dimensional and complex state spaces, adapt to dynamically changing environments, and generate effective decision strategies even in the absence of precise models. Artificial intelligence methods in UAV adversarial decision-making mainly include three categories: expert systems, deep learning, and deep reinforcement learning. Some researchers have constructed decision-making models based on hierarchical expert systems and proposed decision-making methods based on adversarial strategy coordination, subdividing the adversarial situation into two cases: optimization-driven and adversarial strategy coordination, and establishing a hierarchical decision-making framework with a UAV collaborative adversarial knowledge base. Zhang Hongpeng et al. used neural networks to predict the future situation based on the current adversarial situation and select the optimal action for a given action, shortening the decision-making time. Some researchers have proposed a maneuver decision-making model that combines multi-objective optimization with reinforcement learning, solving the problems of difficulty in setting objective weights and insufficient game theory in traditional optimization methods.
[0067] Drone adversarial training faces severe learning challenges. Agents need to explore effective strategies in a high-dimensional state space while simultaneously dealing with complex adversary dynamic responses, often leading to inefficiencies when directly applying standard reinforcement learning methods. This phenomenon stems primarily from two fundamental problems: exploration difficulty and reward sparsity. The high-dimensional action-state space makes it difficult for random exploration to discover effective strategies; simultaneously, the unclear correlation between long-term rewards and current decisions in complex adversarial environments makes the policy optimization process prone to getting trapped in local optima.
[0068] Course-based learning offers a natural solution, its core concept being the simulation of a progressive educational model in human learning. This method breaks down complex tasks into a series of progressively more challenging learning stages, each focusing on developing specific skills and providing the agent with an adaptive learning environment and feedback mechanisms. Through a structured learning path that progresses from easy to difficult, the agent can quickly build foundational skills in simplified scenarios, then gradually tackle more challenging situations, ultimately mastering complex and comprehensive skills.
[0069] However, existing training methods based on course learning use fixed reward weights, which cannot adjust the reward ratio according to the task complexity or the learning progress of the agent. The core of course learning is to dynamically adjust the training difficulty, but fixed reward weights limit the adaptive adjustment capability of the reward. Moreover, when the task difficulty increases, fixed reward weights cannot provide effective guidance, resulting in oscillations in policy updates and causing instability in the training process.
[0070] In the training process of multi-UAV cooperative control system, there is a critical bottleneck that urgently needs to be addressed: when the complexity of the task environment and the scale of the agent cluster increase simultaneously, the computation time required for the system to interact with the environment grows exponentially. This non-linear growth relationship directly leads to a significant decrease in the model convergence speed, which severely limits the improvement of system training efficiency.
[0071] Based on this, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0072] This invention provides a training method based on a dynamically adjusted reward mechanism, such as... Figure 1 As shown, Figure 1 A flowchart illustrating a training method based on a dynamically adjusted reward mechanism, provided by an embodiment of the present invention; applied to adversarial learning using multiple unmanned aerial vehicles, the method includes:
[0073] Step S101: Obtain a first-party drone and a second-party drone; the first-party drone is any drone among the multiple drones; the second-party drone is any drone among the multiple drones other than the first-party drone.
[0074] Step S102: When the first drone and the second drone are learning the attack course, determine the first parameter of the attack reward configuration in the first drone and the second drone according to the preset multi-agent near-end policy optimization algorithm.
[0075] Step S103: When the first drone and the second drone are learning the defense course, determine the second parameter of the defense reward configuration in the first drone and the second drone according to the multi-agent near-end policy optimization algorithm.
[0076] Step S104: When the first drone and the second drone are learning adversarial courses, determine the third parameter of the balanced reward configuration of the first drone according to the multi-agent proximal policy optimization algorithm.
[0077] Step S105: Determine the target reward configuration parameters based on the first parameter, the second parameter, and the third parameter; the target reward configuration parameters are used to learn and train the multiple UAVs.
[0078] In this embodiment, the training method based on the dynamically adjusted reward mechanism can be determined according to the actual situation and is not limited here. As an example, the training method based on the dynamically adjusted reward mechanism can be a parallel training method for multi-UAV adversarial course learning based on the dynamically adjusted reward mechanism.
[0079] In step S101, both the first-party drone and the second-party drone can be determined according to the actual situation, and no limitation is made here. As an example, the first-party drone may include our drone, the red team's drone, etc.; the second-party drone may include the opposing team's drone, the blue team's drone, etc.
[0080] In step S102, the attack training course for the first and second drones can be understood as the first phase of basic attack training, focusing on developing the attack capabilities of the agents. The first drone can be our own drone; the second drone can be the opposing drone.
[0081] The specific determination process for the first parameter of the attack reward configuration in the first and second drones, determined according to the preset multi-agent proximal policy optimization algorithm, is determined based on actual circumstances and is not limited here. As an example, determining the first parameter of the attack reward configuration in the first and second drones according to the preset multi-agent proximal policy optimization algorithm may include: determining a first reward parameter related to the attack capability of the first drone and a second reward parameter related to the defense capability of the second drone according to the multi-agent proximal policy optimization algorithm; and determining the first parameter based on the first reward parameter and the second reward parameter. The first parameter can be denoted as... .
[0082] In practical applications, the first stage is a basic attack course, focusing on developing the agent's attack capabilities. In this stage, our drone engages in combat against an adversary drone employing a baseline maneuvering strategy; the adversary drone simply flies along a fixed path and lacks attack capabilities. The reward function in this stage selectively increases the weight of attack-related factors while minimizing the weight of defense-related factors.
[0083] In step S103, the defense training for the first and second drones can be understood as a second phase of defense enhancement training, focusing on developing the agent's defense capabilities. The first drone can be a red team drone; the second drone can be a blue team drone.
[0084] The specific determination process for the second parameter of the defense reward configuration in the first and second UAVs determined according to the multi-agent proximal policy optimization algorithm is determined based on actual circumstances and is not limited here. As an example, the determination of the second parameter of the defense reward configuration in the first and second UAVs according to the multi-agent proximal policy optimization algorithm may include: determining a third reward parameter related to the action space capability of the first UAV and a fourth reward parameter related to the active pursuit attack of the second UAV according to the multi-agent proximal policy optimization algorithm; and determining the second parameter based on the third and fourth reward parameters. The second parameter can be denoted as... .
[0085] In practical applications, the second phase is a defense enhancement course, focusing on developing the agent's defensive capabilities. In this phase, the red team's drones have limited weapon usage, only allowing for maneuvering and evasion. The blue team's drones, on the other hand, are equipped with full attack capabilities and employ an active pursuit and attack strategy. The reward function in this phase significantly increases the weighting of defense-related factors.
[0086] In step S104, the adversarial learning between the first and second drones can be understood as the third stage being a comprehensive adversarial course, focusing on cultivating the agent's ability to balance offense and defense. The first drone can be the red team drone; the second drone can be the blue team drone.
[0087] The specific determination process for the third parameter in determining the balanced reward configuration of the first party UAV according to the multi-agent proximal policy optimization algorithm is determined based on actual circumstances and is not limited here. As an example, determining the third parameter of the balanced reward configuration of the first party UAV according to the multi-agent proximal policy optimization algorithm may include: determining a fifth, sixth, and seventh reward parameter related to the first party UAV's selection of attack or defense actions according to the multi-agent proximal policy optimization algorithm; and determining the third parameter based on the fifth, sixth, and seventh reward parameters. The third parameter can be denoted as... .
[0088] In practical applications, the third stage is a comprehensive adversarial course, focusing on cultivating the agent's ability to balance offense and defense. In this stage, all action restrictions are removed, and the entire action space is opened, allowing the red team's drone to freely choose attack or defensive actions. The reward function in this stage restores a balanced configuration.
[0089] In step S105, the specific determination process for determining the target reward configuration parameters based on the first parameter, the second parameter, and the third parameter can be determined according to actual circumstances and is not limited here. The target reward configuration parameters can be understood as the optimized policy network, and can be denoted as... .
[0090] In practical applications, the first stage can be denoted as... The second stage can be denoted as The third stage can be denoted as The set of course stages can be denoted as ; Through a three-stage training course with dynamically adjusted reward mechanisms, the agent is guided from mastering basic attack capabilities to defensive skills, and finally develops into a comprehensive adversarial agent integrating offense and defense. At the same time, a parallel environment training mechanism is proposed to accelerate the convergence speed of multi-UAV adversarial training and significantly reduce training time.
[0091] In some embodiments, determining the first parameter of the attack reward configuration in the first party drone and the second party drone according to a preset multi-agent proximal policy optimization algorithm includes:
[0092] The first reward parameter related to the attack capability of the first party drone and the second reward parameter related to the defense capability of the second party drone are determined according to the multi-agent proximal policy optimization algorithm.
[0093] The first parameter is determined based on the first reward parameter and the second reward parameter.
[0094] In this embodiment, the specific determination process for determining the first reward parameter related to the attack capability of the first drone and the second reward parameter related to the defense capability of the second drone according to the multi-agent proximal policy optimization algorithm can be determined according to the actual situation and is not limited here. As an example, determining the first reward parameter related to the attack capability of the first drone and the second reward parameter related to the defense capability of the second drone according to the multi-agent proximal policy optimization algorithm may include obtaining the first policy parameter, first value network parameter, first course configuration parameter, and first environment configuration parameter of the multi-drone according to the multi-agent proximal policy optimization algorithm; determining the attitude advantage reward parameter, altitude reward parameter, event triggering reward parameter of the first drone and the launch penalty reward parameter of the second drone based on the first policy parameter, the first value network parameter, the first course configuration parameter, and the first environment configuration parameter; and using the attitude advantage reward parameter, the altitude reward parameter, and the event triggering reward parameter as the first reward parameter and the launch penalty reward parameter as the second reward parameter.
[0095] The specific determination process for determining the first parameter based on the first reward parameter and the second reward parameter can be determined according to the actual situation and is not limited here. As an example, determining the first parameter based on the first reward parameter and the second reward parameter may include obtaining a first weighting coefficient of the attitude advantage reward parameter, a second weighting coefficient of the altitude reward parameter, a third weighting coefficient of the event triggering reward parameter, and a fourth weighting coefficient of the launch penalty reward parameter; and determining the first parameter based on the first weighting coefficient, the second weighting coefficient, the third weighting coefficient, the fourth weighting coefficient, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter.
[0096] In some embodiments, determining the first reward parameter related to the attack capability of the first drone and the second reward parameter related to the defense capability of the second drone according to the multi-agent proximal policy optimization algorithm includes:
[0097] The first policy parameters, first value network parameters, first course configuration parameters, and first environment configuration parameters of the multi-agent proximal policy optimization algorithm are obtained according to the multi-agent proximal policy optimization algorithm.
[0098] Based on the first strategy parameters, the first value network parameters, the first course configuration parameters, and the first environment configuration parameters, determine the attitude advantage reward parameters, altitude reward parameters, event triggering reward parameters, and launch penalty reward parameters of the first UAV.
[0099] The attitude advantage reward parameter, the altitude reward parameter, and the event triggering reward parameter are used as the first reward parameter, and the launch penalty reward parameter is used as the second reward parameter.
[0100] It should be noted that Proximal Policy Optimization (PPO) is a deep reinforcement learning method based on the Actor-Critic framework, applicable to problems with continuous or discrete action spaces. This algorithm directly optimizes the policy function through the policy gradient method to maximize the expected reward obtained by the agent in the environment. The core objective function of the PPO algorithm is defined as:
[0101] MERGEFORMAT (1);
[0102] in, and These are the loss functions for the Actor network and the Critic network, respectively. These are network parameters. and To balance the coefficients of various influences, The entropy regularization term is used to facilitate policy exploration. At any moment The observation status of the Actor network.
[0103] Strategy Network Loss The key innovation of the PPO algorithm is that it ensures training stability by limiting the magnitude of policy updates:
[0104] MERGEFORMAT(2);
[0105] in, The policy probability ratio represents the ratio of the probability of choosing the same action under the current policy to the probability under the old policy. The clip function limits this ratio to... Within the scope, prevent a single update from causing excessive changes to the strategy. It is the advantage function estimate, which measures the additional gain of a particular action relative to the average state value. This pruning mechanism ensures the stability of policy updates and avoids policy collapse during training.
[0106] Value network loss Using mean square error form:
[0107] MERGEFORMAT (3);
[0108] in, It is the state value estimate output by the Critic network. It is the actual observed reward. Value networks provide a reliable advantage function benchmark for Actor networks by accurately estimating state values.
[0109] Advantage function estimation is a key factor in the performance of the PPO algorithm. This application employs the Generalized Advantage Estimation (GAE) technique:
[0110] MERGEFORMAT (4);
[0111] MERGEFORMAT (5);
[0112] in, As a discount factor, These are the GAE parameters. GAE balances the bias and variance of the estimation by weighted combining multi-step time difference errors, making the dominance estimate more accurate.
[0113] Compared to traditional policy optimization methods, PPO is characterized by high computational efficiency, simple implementation, and stable performance. Compared to Trust Region Policy Optimization (TRPO) algorithms, PPO avoids complex second-derivative calculations and constrained optimization problems, significantly reducing computational complexity. Furthermore, PPO supports parallel sample acquisition and multi-round utilization, significantly improving data efficiency. The PPO algorithm is suitable for single-UAV adversarial scenarios and forms the basis of multi-UAV adversarial algorithms.
[0114] Multi-UAV collaborative parallel training process:
[0115] (1) Multi-agent proximal policy optimization algorithm.
[0116] Extending the PPO algorithm from single-drone scenarios to multi-drone scenarios faces two problems that need to be addressed: First, environmental non-stationarity, that is, when multiple agents learn simultaneously, the dynamics of the environment change from the perspective of a single agent, which violates the basic assumptions of Markov decision processes; Second, the curse of dimensionality, that is, as the number of agents increases, the joint action space grows exponentially, making it difficult to effectively apply traditional single-agent reinforcement learning methods.
[0117] Multi-Agent Proximal Policy Optimization (MAPPO) is an extended version of PPO designed for multi-agent environments. It retains core components of PPO such as policy pruning, value function estimation, and generalized advantage estimation, while employing a centralized training-distributed execution framework and a parameter sharing mechanism to address the multi-agent collaboration problem. The parameter sharing technology introduced in MAPPO significantly reduces model complexity and improves sample utilization efficiency. For the unique collaborative challenges of multi-agent systems, MAPPO guides each agent towards optimizing towards the team goal through global value function design, overcoming the suboptimal policy problem inherent in independent learning.
[0118] The MAPPO algorithm employs a Centralized Training and Decentralized Execution (CTDE) framework, which is an effective method for addressing the non-stationarity problem in multi-agent reinforcement learning. Under the CTDE framework, different information acquisition strategies are used during the training and execution phases: global information is utilized to optimize the strategy during training, while decisions are made based solely on local observations during execution.
[0119] During the training phase, the system fully utilizes global information for policy optimization. The central trainer can acquire the state, actions, and reward information of all agents to construct a global value function. Assess the overall environmental status The value of this global perspective lies in its ability to capture the interactions between agents during training, leading to a coordinated policy. Dominance function estimation based on the global state provides more accurate reward signals for each agent, effectively guiding the policy towards global optimum.
[0120] During the execution phase, each agent relies only on its own local observations. Making a decision, that is It does not require obtaining global information. This meets the communication constraints and distributed decision-making requirements of practical applications, significantly reducing communication costs and computational complexity during execution. Moreover, although each agent makes independent decisions during execution, the distributed strategy still maintains good collaborative performance because the influence of global information has been considered during the training phase.
[0121] The mathematical description of CTDE is as follows:
[0122] During the training phase, the policy network and value network The optimization objective is:
[0123] MERGEFORMAT (6);
[0124] in, For global state The estimated value of the advantage function.
[0125] During the execution phase, the strategy for each agent is as follows:
[0126] MERGEFORMAT (7);
[0127] in, For intelligent agents Local observations.
[0128] Another challenge in multi-UAV cooperative control is that as the number of agents increases, the number of policy network parameters grows linearly, leading to a sharp increase in training difficulty and computational resource requirements. To address this issue, MAPPO employs a parameter-sharing mechanism, whereby all agents share the same set of policy network and value network parameters, effectively reducing the number of parameters to be optimized.
[0129] The core idea of the policy network parameter sharing mechanism is that when multiple agents perform similar tasks, their policy structures can be shared, resulting in differentiated behaviors only through different observation inputs. All agents share the same policy network. The formula is expressed as:
[0130] MERGEFORMAT (8);
[0131] Due to the observations of different intelligent agents They vary, even when using the same network parameters. The resulting action distributions will also differ, leading to different behavioral patterns. This mechanism decouples the number of parameters from the number of agents, keeping the model size constant and significantly improving computational efficiency and sample utilization.
[0132] Value network parameter sharing follows a similar principle. Within the CTDE framework, the central value network receives the global state. As input, evaluate the value of the current state. Since the global state already contains information about all agents, parameter sharing does not lead to a loss of expressive power. The value estimation formula is:
[0133] MERGEFORMAT (9);
[0134] (2) Parallel environment accelerates training.
[0135] The training efficiency of multi-UAV collaborative control faces a practical problem: as environmental complexity and the number of agents increase, the environmental interaction time significantly extends, severely restricting training efficiency. To address this issue, this application designs an environment interface that supports multi-process parallel sampling. By executing multiple environment instances in parallel through multiple processes, data sampling efficiency is significantly improved. The parallel environment training architecture is as follows: Figure 2 As shown; Figure 2 This is a schematic diagram of the parallel environment training architecture in an embodiment of this application.
[0136] The core design of the parallel environment architecture includes a main process and multiple child processes. The main process is responsible for policy optimization, parameter updates, action distribution, and training control, while the child processes are responsible for environment instance operation and data acquisition. The two communicate efficiently through a pipeline mechanism: the main process sends action instructions generated by the current policy to each child process; the child processes execute actions and return interactive data such as status, observations, rewards, and termination flags; the main process updates network parameters based on the collected batch data, forming a closed-loop optimization.
[0137] The course learning and training process based on a dynamically adjusted reward mechanism:
[0138] The training methodology framework based on curriculum learning: The core of curriculum learning lies in its progressive learning strategy, which breaks down complex adversarial tasks into multiple stages of increasing difficulty. A key principle in curriculum design is ensuring that the agent does not forget previously learned skills when transitioning to new stages. This application achieves this through reward function design—when the agent enters a new stage, the reward weight of the previous stage is appropriately reduced but not completely eliminated, while the reward weight of the new stage is correspondingly increased. This gradual reward adjustment ensures that the agent can develop new skills while retaining existing capabilities, and provides the agent with clearer learning guidance, forming a coherent capability system.
[0139] The framework of the training process is as follows Figure 3 As shown; Figure 3 This is a schematic diagram illustrating the training process in an embodiment of this application. The algorithm's execution flow follows a closed-loop structure of "environment initialization - course loading - adversarial training - policy evaluation - course update". The course learning module selects a course stage based on the current training progress and transmits the environmental parameters of the current stage to the adversarial environment, including the opponent's UAV strategy type and reward function weights. The system collects environmental observation information, inputs it into the decision network to generate action commands, and then the MAPPO algorithm updates the network parameters based on environmental feedback. When the agent reaches the performance index under the current course, the system automatically loads the next stage of training tasks, gradually increasing the adversarial complexity.
[0140] In this embodiment, the first strategy parameter, the first value network parameter, the first course configuration parameter, and the first environment configuration parameter can all be determined according to actual conditions, and are not limited here. As an example, the first strategy parameter may include an initialization strategy parameter, which can be denoted as... The first value network parameter can be simply referred to as the value network parameter, and can be denoted as... The first course configuration parameters may include the course initialization phase and the course configuration loading phase, which can be denoted as follows: and As an example, It can be 1; loading course configuration can include opponent strategy and reward weights, denoted as . and The first environment configuration parameter can also be called the configuration environment parameter. The configuration environment parameter can include attack permissions and defense capabilities, respectively denoted as... and .
[0141] The specific determination process for determining the attitude advantage reward parameters, altitude reward parameters, event-triggered reward parameters, and launch penalty reward parameters of the first UAV based on the first policy parameters, the first value network parameters, the first course configuration parameters, and the first environment configuration parameters can be determined according to actual circumstances and is not limited here. As an example, the determination of the attitude advantage reward parameters, altitude reward parameters, event-triggered reward parameters, and launch penalty reward parameters of the first UAV based on the first policy parameters, the first value network parameters, the first course configuration parameters, and the first environment configuration parameters can be based on the first policy parameters, the first value network parameters, the first course configuration parameters, and the first environment configuration parameters using a preset algorithm to determine the attitude advantage reward parameters, altitude reward parameters, event-triggered reward parameters, and launch penalty reward parameters of the second UAV; wherein, the preset algorithm can be determined according to actual circumstances and is not limited here. As an example, the preset algorithm can be the CL-MAPPO training algorithm.
[0142] The attitude advantage reward parameters, altitude reward parameters, and event triggering reward parameters of the first-party UAV, as well as the launch penalty reward parameters of the second-party UAV, can all be determined according to actual circumstances and are not limited here. As an example, the attitude advantage reward parameters can be denoted as... The height reward parameter can be denoted as: The event-triggered reward parameter can be denoted as... The launch penalty / reward parameters can be denoted as: .
[0143] Using the attitude advantage reward parameter, the altitude reward parameter, and the event triggering reward parameter as the first reward parameter can be understood as the first reward parameter including... , , , Using the launch penalty reward parameter as the second reward parameter can be understood as the second reward parameter including... .
[0144] In some embodiments, determining the first parameter based on the first reward parameter and the second reward parameter includes:
[0145] Obtain the first weighting coefficient of the attitude advantage reward parameter, the second weighting coefficient of the altitude reward parameter, the third weighting coefficient of the event triggering reward parameter, and the fourth weighting coefficient of the launch penalty reward parameter;
[0146] The first parameter is determined based on the first weighting coefficient, the second weighting coefficient, the third weighting coefficient, the fourth weighting coefficient, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter.
[0147] In this embodiment, the first weighting coefficient, the second weighting coefficient, the third weighting coefficient, and the fourth weighting coefficient can all be determined according to actual circumstances, and are not limited here. As an example, the first weighting coefficient can be denoted as... The second weighting coefficient can be denoted as... The third weighting coefficient can be denoted as... The fourth weighting coefficient can be denoted as... For example, the weight of posture advantage reward. Increased to 0.4 (standard value 0.2), relative height reward weight. Increase the event-triggered reward weight to 0.3 (standard value 0.15). Maintain at 0.5, while adjusting the launch penalty / reward weight. Reduced to 0.1 (standard value 0.2).
[0148] The specific determination process for determining the first parameter based on the first weighting coefficient, the second weighting coefficient, the third weighting coefficient, the fourth weighting coefficient, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter can be determined according to actual circumstances and is not limited here. As an example, determining the first parameter based on the first weighting coefficient, the second weighting coefficient, the third weighting coefficient, the fourth weighting coefficient, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter can be achieved by determining the first parameter using a preset algorithm based on the first weighting coefficient, the second weighting coefficient, the third weighting coefficient, the fourth weighting coefficient, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter; wherein, the preset algorithm can be determined according to actual circumstances and is not limited here. As an example, the preset algorithm can refer to... .
[0149] In practical applications, this process can be understood as the first stage, which is a basic attack course that focuses on developing the agent's attack capabilities. In this stage, our drone engages in combat with an adversary drone employing a baseline maneuvering strategy. The adversary drone only flies along a fixed path and lacks attack capabilities. The reward function in this stage selectively increases the weight of attack-related factors while minimizing the weight of defense-related factors. ;
[0150] Among them, the posture advantage reward weight Increased to 0.4 (standard value 0.2), relative height reward weight. Increase the event-triggered reward weight to 0.3 (standard value 0.15). Maintain at 0.5, while adjusting the launch penalty / reward weight. Reduced to 0.1 (standard value 0.2). Simultaneously, the missile evasion bonus will be... The weight of the reward is reduced to a minimum of 0.05, prompting the agent to focus on learning aggressive behaviors. Through this reward configuration, the agent can quickly master proactive aggressive behaviors in a simplified environment, establish aggressive awareness and basic skills, and lay the foundation for subsequent stages.
[0151] In some embodiments, determining the second parameter of the defense reward configuration in the first and second party drones according to the multi-agent proximal policy optimization algorithm includes:
[0152] The third reward parameter related to the action space capability of the first drone and the fourth reward parameter related to the active pursuit attack of the second drone are determined according to the multi-agent proximal policy optimization algorithm.
[0153] The second parameter is determined based on the third reward parameter and the fourth reward parameter.
[0154] In this embodiment, the determination of the third reward parameter related to the action space capability of the first UAV and the fourth reward parameter related to the active pursuit attack of the second UAV based on the multi-agent proximal policy optimization algorithm can be determined according to the actual situation and is not limited here. As an example, the determination of the third reward parameter related to the action space capability of the first UAV and the fourth reward parameter related to the active pursuit attack of the second UAV based on the multi-agent proximal policy optimization algorithm may include obtaining the second policy parameter, second value network parameter, second curriculum configuration parameter, and second environment configuration parameter of the multi-UAV based on the multi-agent proximal policy optimization algorithm; determining the missile evasion reward parameter, attitude advantage reward parameter, altitude reward parameter, event triggering reward parameter of the first UAV and the launch penalty reward parameter of the second UAV based on the second policy parameter, second value network parameter, second curriculum configuration parameter, and second environment configuration parameter; and using the missile evasion reward parameter, attitude advantage reward parameter, altitude reward parameter, and event triggering reward parameter as the third reward parameter and the launch penalty reward parameter as the fourth reward parameter.
[0155] The specific determination process for determining the second parameter based on the third and fourth reward parameters can be determined according to actual circumstances and is not limited here. As an example, determining the second parameter based on the third and fourth reward parameters may include: obtaining the fifth weighting coefficient of the missile evasion reward parameter, the sixth weighting coefficient of the attitude advantage reward parameter, the seventh weighting coefficient of the altitude reward parameter, the eighth weighting coefficient of the event triggering reward parameter, and the ninth weighting coefficient of the launch penalty reward parameter; and determining the second parameter based on the fifth, sixth, seventh, eighth, and ninth weighting coefficients, the missile evasion reward parameter, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter.
[0156] In some embodiments, determining the third reward parameter related to the action space capability of the first drone and the fourth reward parameter related to the active pursuit attack of the second drone according to the multi-agent proximal policy optimization algorithm includes:
[0157] The second policy parameters, second value network parameters, second curriculum configuration parameters, and second environment configuration parameters of the multi-agent proximal policy optimization algorithm are obtained according to the multi-agent proximal policy optimization algorithm.
[0158] Based on the second strategy parameters, the second value network parameters, the second course configuration parameters, and the second environment configuration parameters, determine the missile evasion reward parameters, attitude advantage reward parameters, altitude reward parameters, event triggering reward parameters, and the launch penalty reward parameters of the first party's UAV.
[0159] The missile evasion reward parameter, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter are used as the third reward parameter, and the launch penalty reward parameter is used as the fourth reward parameter.
[0160] In this embodiment, the second strategy parameter, the second value network parameter, the second course configuration parameter, and the second environment configuration parameter can all be determined according to actual conditions, and are not limited here. As an example, the second strategy parameter may include an initialization strategy parameter, which can be denoted as... The second value network parameter can be simply referred to as the value network parameter, and can be denoted as... The second course configuration parameters may include the course initialization phase and the course configuration loading phase, which can be denoted as follows: and As an example, It can be 2; loading course configuration can include opponent strategy and reward weights, denoted as . and The second environment configuration parameter can also be called the configuration environment parameter. The configuration environment parameter can include attack permissions and defense capabilities, denoted as follows: and .
[0161] The specific determination process for the missile evasion reward parameters, attitude advantage reward parameters, altitude reward parameters, event-triggered reward parameters, and launch penalty reward parameters of the first UAV based on the second strategy parameters, the second value network parameters, the second course configuration parameters, and the second environment configuration parameters can be determined according to actual circumstances and is not limited here. As an example, the determination of these parameters based on the second strategy parameters, the second value network parameters, the second course configuration parameters, and the second environment configuration parameters can be achieved by using a preset algorithm to determine these parameters. The preset algorithm can be determined according to actual circumstances and is not limited here. As an example, the preset algorithm can be the CL-MAPPO training algorithm.
[0162] The missile evasion reward parameters, attitude advantage reward parameters, altitude reward parameters, and event triggering reward parameters of the first-party UAV, as well as the launch penalty reward parameters of the second-party UAV, can all be determined according to the actual situation and are not limited here. As an example, the missile evasion reward parameters can be denoted as... The posture advantage reward parameter can be denoted as: The height reward parameter can be denoted as: The event-triggered reward parameter can be denoted as... The launch penalty / reward parameters can be denoted as: .
[0163] The missile evasion reward parameter, the attitude advantage reward parameter, the altitude reward parameter, and the event triggering reward parameter can be understood as the second reward parameter including the third reward parameter. , , , , Using the launch penalty reward parameter as the fourth reward parameter can be understood as the fourth reward parameter including... .
[0164] In some embodiments, determining the second parameter based on the third reward parameter and the fourth reward parameter includes:
[0165] The fifth weighting coefficient of the missile evasion reward parameter, the sixth weighting coefficient of the attitude advantage reward parameter, the seventh weighting coefficient of the altitude reward parameter, the eighth weighting coefficient of the event triggering reward parameter, and the ninth weighting coefficient of the launch penalty reward parameter are obtained.
[0166] The second parameter is determined based on the fifth weighting coefficient, the sixth weighting coefficient, the seventh weighting coefficient, the eighth weighting coefficient, the ninth weighting coefficient, the missile evasion reward parameter, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter.
[0167] In this embodiment, the fifth, sixth, seventh, eighth, and ninth weighting coefficients can all be determined according to actual circumstances, and are not limited here. As an example, the fifth weighting coefficient can be denoted as... The sixth weighting coefficient can be denoted as... The seventh weighting coefficient can be denoted as... The eighth weighting coefficient can be denoted as... For example, missile evasion reward weights. Increased to 0.45 (standard value 0.15), postural advantage reward weight. Adjust the configuration to focus on safe distance and increase it to 0.3 (standard value 0.2), with a relative height reward weight. Increase the weight of survival-related event trigger rewards to 0.25 (standard value 0.15). Increased to 0.4 (standard value 0.2).
[0168] The specific determination process for determining the second parameter based on the fifth, sixth, seventh, eighth, and ninth weighting coefficients, the missile evasion reward parameter, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter can be determined according to actual circumstances and is not limited here. As an example, determining the second parameter based on the fifth, sixth, seventh, eighth, and ninth weighting coefficients, the missile evasion reward parameter, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter can be achieved by using a preset algorithm based on the fifth, sixth, seventh, eighth, and ninth weighting coefficients, the missile evasion reward parameter, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter; wherein, the preset algorithm can be determined according to actual circumstances and is not limited here. As an example, the preset algorithm can refer to... .
[0169] In practical applications, this process can be understood as the second phase, a defense enhancement course that focuses on developing the agent's defensive capabilities. In this phase, the red team's drones have limited weapon usage, only allowing for maneuvering and evasion. The blue team's drones, on the other hand, are equipped with full attack capabilities and employ an active pursuit and attack strategy. The reward function in this phase significantly increases the weighting of defense-related factors. Among them, the missile evasion reward weight Increased to 0.45 (standard value 0.15), postural advantage reward weight. Adjust the configuration to focus on safe distance and increase it to 0.3 (standard value 0.2), with a relative height reward weight. Increase the weight of survival-related event trigger rewards to 0.25 (standard value 0.15). Increased to 0.4 (standard value 0.2). Meanwhile, due to motion space limitations, the launch penalty / reward... The weights are temporarily set to 0. However, this setting does not mean that the agent will forget the attack capabilities learned in the first phase, because the network parameters still retain these skills. The purpose of this phase is to allow the agent to strengthen its defensive skills while maintaining its offensive awareness, and to master the ability to effectively avoid threats.
[0170] In some embodiments, determining the third parameter of the balanced reward configuration of the first party UAV according to the multi-agent proximal policy optimization algorithm includes:
[0171] The fifth, sixth, and seventh reward parameters related to the selection of attack or defense actions by the first party UAV are determined based on the multi-agent proximal policy optimization algorithm.
[0172] The third parameter is determined based on the fifth reward parameter, the sixth reward parameter, and the seventh reward parameter.
[0173] In this embodiment, the process of determining the fifth, sixth, and seventh reward parameters related to the first-party UAV's choice of attack or defense action based on the multi-agent proximal policy optimization algorithm can be determined according to actual circumstances and is not limited here. As an example, determining the fifth reward parameter related to the first-party UAV's choice of attack or defense action based on the multi-agent proximal policy optimization algorithm may include obtaining the third policy parameters, third value network parameters, third curriculum configuration parameters, and third environment configuration parameters of the multiple UAVs according to the multi-agent proximal policy optimization algorithm; determining the first-party UAV's attitude advantage reward parameter, altitude reward parameter, and launch penalty reward based on the third policy parameters, the third value network parameters, the third curriculum configuration parameters, and the third environment configuration parameters; determining the first weighting factor of the attitude advantage reward parameter, the second weighting factor of the altitude reward parameter, and the third weighting factor of the launch penalty reward; and determining the fifth reward parameter based on the attitude advantage reward parameter, the altitude reward parameter, the launch penalty reward, the first weighting factor, the second weighting factor, and the third weighting factor.
[0174] The specific determination process for the third parameter based on the fifth, sixth, and seventh reward parameters can be determined according to actual circumstances and is not limited here. As an example, the sixth reward parameter includes a missile evasion reward parameter; the seventh reward parameter includes an event triggering reward parameter; and the determination of the third parameter based on the fifth, sixth, and seventh reward parameters includes:
[0175] Obtain the first weight value of the fifth reward parameter, the second weight value of the missile evasion reward parameter, and the third weight value of the event triggering reward parameter;
[0176] The third parameter is determined based on the first weight value, the second weight value, the third weight value, the fifth reward parameter, the sixth reward parameter, and the seventh reward parameter.
[0177] In some embodiments, determining the fifth reward parameter related to the selection of attack or defense actions by the first party UAV based on the multi-agent proximal policy optimization algorithm includes:
[0178] The third policy parameters, third value network parameters, third course configuration parameters, and third environment configuration parameters of the multiple UAVs are obtained according to the multi-agent proximal policy optimization algorithm.
[0179] Based on the third strategy parameters, the third value network parameters, the third course configuration parameters, and the third environment configuration parameters, the attitude advantage reward parameters, altitude reward parameters, and launch penalty reward of the first party UAV are determined.
[0180] Determine the first weighting factor of the attitude advantage reward parameter, the second weighting factor of the altitude reward parameter, and the third weighting factor of the launch penalty reward;
[0181] The fifth reward parameter is determined based on the attitude advantage reward parameter, the altitude reward parameter, the launch penalty reward, the first weighting factor, the second weighting factor, and the third weighting factor.
[0182] In this embodiment, the third strategy parameter, the third value network parameter, the third course configuration parameter, and the third environment configuration parameter can all be determined according to actual conditions, and are not limited here. As an example, the third strategy parameter may include an initialization strategy parameter, which can be denoted as... The third value network parameter can be simply referred to as the value network parameter, and can be denoted as... The third course configuration parameters may include the course initialization phase and the course configuration loading phase, which can be denoted as follows: and As an example, It can be 3; loading course configuration can include opponent strategy and reward weights, denoted as . and The third environment configuration parameter can also be called the configuration environment parameter. The configuration environment parameter can include attack permissions and defense capabilities, denoted as follows: and .
[0183] The specific determination process for the attitude advantage reward parameters, altitude reward parameters, and launch penalty reward of the first-party UAV based on the third policy parameters, the third value network parameters, the third course configuration parameters, and the third environment configuration parameters can be determined according to actual circumstances and is not limited here. As an example, the determination of the attitude advantage reward parameters, altitude reward parameters, and launch penalty reward of the first-party UAV based on the third policy parameters, the third value network parameters, the third course configuration parameters, and the third environment configuration parameters can be based on a preset algorithm to determine the attitude advantage reward parameters, altitude reward parameters, and launch penalty reward of the first-party UAV; wherein, the preset algorithm can be determined according to actual circumstances and is not limited here. As an example, the preset algorithm can be the CL-MAPPO training algorithm.
[0184] The attitude advantage reward parameters, altitude reward parameters, and launch penalty reward of the first-party UAV can all be determined according to the actual situation, and are not limited here. As an example, the attitude advantage reward parameters can be denoted as... The height reward parameter can be denoted as: The event-triggered reward parameter can be denoted as... The launch penalty / reward parameters can be denoted as: .
[0185] The specific determination process for the first weighting factor of the attitude advantage reward parameter, the second weighting factor of the altitude reward parameter, and the third weighting factor of the launch penalty reward can be determined according to actual circumstances and is not limited here. As an example, the first weighting factor can be denoted as... The second weighting factor can be denoted as... The third weighting factor can be denoted as... ;
[0186] The specific determination process for the fifth reward parameter based on the attitude advantage reward parameter, the altitude reward parameter, the launch penalty reward, the first weighting factor, the second weighting factor, and the third weighting factor can be determined according to actual circumstances and is not limited here. As an example, determining the fifth reward parameter based on the attitude advantage reward parameter, the altitude reward parameter, the launch penalty reward, the first weighting factor, the second weighting factor, and the third weighting factor can be achieved by using a preset algorithm to determine the fifth reward parameter; wherein, the fifth reward parameter can be denoted as R; the preset algorithm can be determined according to actual circumstances and is not limited here. As an example, the preset algorithm can refer to... .
[0187] In some embodiments, the sixth reward parameter includes a missile evasion reward parameter; the seventh reward parameter includes an event triggering reward parameter; determining the third parameter based on the fifth reward parameter, the sixth reward parameter, and the seventh reward parameter includes:
[0188] Obtain the first weight value of the fifth reward parameter, the second weight value of the missile evasion reward parameter, and the third weight value of the event triggering reward parameter;
[0189] The third parameter is determined based on the first weight value, the second weight value, the third weight value, the fifth reward parameter, the sixth reward parameter, and the seventh reward parameter.
[0190] In this embodiment, the sixth reward parameter includes a missile evasion reward parameter; the seventh reward parameter includes an event triggering reward parameter; the missile evasion reward parameter can be denoted as... The event-triggered reward parameter can be denoted as... .
[0191] The first weight value, the second weight value, and the third weight value can all be determined according to actual circumstances, and are not limited here. As an example, the first weight value can be denoted as... The second weight value can be denoted as The third weight value can be denoted as: .
[0192] The specific determination process for determining the third parameter based on the first weight value, the second weight value, the third weight value, the fifth reward parameter, the sixth reward parameter, and the seventh reward parameter can be determined according to actual circumstances and is not limited here. As an example, determining the third parameter based on the first weight value, the second weight value, the third weight value, the fifth reward parameter, the sixth reward parameter, and the seventh reward parameter can be achieved by determining the third parameter using a preset algorithm based on the first weight value, the second weight value, the third weight value, the fifth reward parameter, the sixth reward parameter, and the seventh reward parameter; wherein, the preset algorithm can be determined according to actual circumstances and is not limited here. As an example, the preset algorithm can refer to... .
[0193] In practical applications, this process can be understood as the third stage, a comprehensive adversarial course that focuses on cultivating the agent's ability to balance offense and defense. In this stage, all action restrictions are removed, opening up the entire action space, allowing the red team's drone to freely choose attack or defensive actions. The reward function in this stage restores a balanced configuration. .
[0194] In practical applications, a training method based on a dynamically adjusted reward mechanism can be specifically a parallel training method for multi-UAV adversarial course learning based on a dynamically adjusted reward mechanism, which consists of two parts: a multi-UAV collaborative parallel training method based on a multi-agent proximal policy optimization algorithm and a course learning training method based on a dynamically adjusted reward mechanism.
[0195] 1.1 A multi-UAV collaborative parallel training method based on multi-agent proximal policy optimization algorithm.
[0196] 1.1.1 Near-end strategy optimization algorithm.
[0197] Proximal Policy Optimization (PPO) is a deep reinforcement learning method based on the Actor-Critic framework, applicable to problems with continuous or discrete action spaces. This algorithm directly optimizes the policy function through the policy gradient method to maximize the expected reward obtained by the agent in the environment. The core objective function of the PPO algorithm is defined as:
[0198] (1);
[0199] in, and These are the loss functions for the Actor network and the Critic network, respectively. These are network parameters. and To balance the coefficients of various influences, The entropy regularization term is used to facilitate policy exploration. At any moment The observation status of the Actor network.
[0200] Strategy Network Loss The key innovation of the PPO algorithm is that it ensures training stability by limiting the magnitude of policy updates:
[0201] (2);
[0202] in, The policy probability ratio represents the ratio of the probability of choosing the same action under the current policy to the probability under the old policy. The clip function limits this ratio to... Within the scope, prevent a single update from causing excessive changes to the strategy. It is the advantage function estimate, which measures the additional gain of a particular action relative to the average state value. This pruning mechanism ensures the stability of policy updates and avoids policy collapse during training.
[0203] Value network loss Using mean square error form:
[0204] (3);
[0205] in, It is the state value estimate output by the Critic network. It is the actual observed reward. Value networks provide a reliable advantage function benchmark for Actor networks by accurately estimating state values.
[0206] Advantage function estimation is a key factor in the performance of the PPO algorithm. This application employs the Generalized Advantage Estimation (GAE) technique:
[0207] (4);
[0208] (5);
[0209] in, As a discount factor, These are the GAE parameters. GAE balances the bias and variance of the estimation by weighted combining multi-step time difference errors, making the dominance estimate more accurate.
[0210] Compared to traditional policy optimization methods, PPO is characterized by high computational efficiency, simple implementation, and stable performance. Compared to Trust Region Policy Optimization (TRPO) algorithms, PPO avoids complex second-derivative calculations and constrained optimization problems, significantly reducing computational complexity. Furthermore, PPO supports parallel sample acquisition and multi-round utilization, significantly improving data efficiency. The PPO algorithm is suitable for single-UAV adversarial scenarios and forms the basis of multi-UAV adversarial algorithms.
[0211] 1.1.2 Multi-UAV collaborative parallel training method.
[0212] (1) Multi-agent proximal policy optimization algorithm.
[0213] Extending the PPO algorithm from single-drone scenarios to multi-drone scenarios faces two problems that need to be addressed: First, environmental non-stationarity, that is, when multiple agents learn simultaneously, the dynamics of the environment change from the perspective of a single agent, which violates the basic assumptions of Markov decision processes; Second, the curse of dimensionality, that is, as the number of agents increases, the joint action space grows exponentially, making it difficult to effectively apply traditional single-agent reinforcement learning methods.
[0214] Multi-Agent Proximal Policy Optimization (MAPPO) is an extended version of PPO designed for multi-agent environments. It retains core components of PPO such as policy pruning, value function estimation, and generalized advantage estimation, while employing a centralized training-distributed execution framework and a parameter sharing mechanism to address the multi-agent collaboration problem. The parameter sharing technology introduced in MAPPO significantly reduces model complexity and improves sample utilization efficiency. For the unique collaborative challenges of multi-agent systems, MAPPO guides each agent towards optimizing towards the team goal through global value function design, overcoming the suboptimal policy problem inherent in independent learning.
[0215] The MAPPO algorithm employs a Centralized Training and Decentralized Execution (CTDE) framework, which is an effective method for addressing the non-stationarity problem in multi-agent reinforcement learning. Under the CTDE framework, different information acquisition strategies are used during the training and execution phases: global information is utilized to optimize the strategy during training, while decisions are made based solely on local observations during execution.
[0216] During the training phase, the system fully utilizes global information for policy optimization. The central trainer can acquire the state, actions, and reward information of all agents to construct a global value function. Assess the overall environmental status The value of this global perspective lies in its ability to capture the interactions between agents during training, leading to a coordinated policy. Dominance function estimation based on the global state provides more accurate reward signals for each agent, effectively guiding the policy towards global optimum.
[0217] During the execution phase, each agent relies only on its own local observations. Making a decision, that is It does not require obtaining global information. This meets the communication constraints and distributed decision-making requirements of practical applications, significantly reducing communication costs and computational complexity during execution. Moreover, although each agent makes independent decisions during execution, the distributed strategy still maintains good collaborative performance because the influence of global information has been considered during the training phase.
[0218] The mathematical description of CTDE is as follows:
[0219] During the training phase, the policy network and value network The optimization objective is:
[0220] (6);
[0221] in, For global state The estimated value of the advantage function.
[0222] During the execution phase, the strategy for each agent is as follows:
[0223] (7);
[0224] in, For intelligent agents Local observations.
[0225] Another challenge in multi-UAV cooperative control is that as the number of agents increases, the number of policy network parameters grows linearly, leading to a sharp increase in training difficulty and computational resource requirements. To address this issue, MAPPO employs a parameter-sharing mechanism, whereby all agents share the same set of policy network and value network parameters, effectively reducing the number of parameters to be optimized.
[0226] The core idea of the policy network parameter sharing mechanism is that when multiple agents perform similar tasks, their policy structures can be shared, resulting in differentiated behaviors only through different observation inputs. All agents share the same policy network. The formula is expressed as:
[0227] (8);
[0228] Due to the observations of different intelligent agents They vary, even when using the same network parameters. The resulting action distributions will also differ, leading to different behavioral patterns. This mechanism decouples the number of parameters from the number of agents, keeping the model size constant and significantly improving computational efficiency and sample utilization.
[0229] Value network parameter sharing follows a similar principle. Within the CTDE framework, the central value network receives the global state. As input, evaluate the value of the current state. Since the global state already contains information about all agents, parameter sharing does not lead to a loss of expressive power. The value estimation formula is:
[0230] (9);
[0231] (2) Parallel environment accelerates training.
[0232] The training efficiency of multi-UAV collaborative control faces a practical problem: as environmental complexity and the number of agents increase, the environmental interaction time significantly extends, severely restricting training efficiency. To address this issue, this application designs an environment interface that supports multi-process parallel sampling. By executing multiple environment instances in parallel through multiple processes, data sampling efficiency is significantly improved. The parallel environment training architecture is as follows: Figure 2 As shown.
[0233] The core design of the parallel environment architecture includes a main process and multiple child processes. The main process is responsible for policy optimization, parameter updates, action distribution, and training control, while the child processes are responsible for environment instance operation and data acquisition. The two communicate efficiently through a pipeline mechanism: the main process sends action instructions generated by the current policy to each child process; the child processes execute actions and return interactive data such as status, observations, rewards, and termination flags; the main process updates network parameters based on the collected batch data, forming a closed-loop optimization.
[0234] 1.2 Course learning and training methods based on dynamic reward adjustment mechanism.
[0235] 1.2.1 Training methodology framework based on course learning.
[0236] The core of the course lies in its progressive learning strategy, breaking down complex adversarial tasks into multiple stages of increasing difficulty. A key principle in designing the course is ensuring that the agent does not forget previously learned skills when transitioning to new stages. This application achieves this through a reward function design—when the agent enters a new stage, the reward weight of the previous stage is appropriately reduced but not completely eliminated, while the reward weight of the new stage is correspondingly increased. This gradual adjustment of rewards ensures that the agent can develop new skills while retaining existing abilities, and provides the agent with clearer learning guidance, forming a coherent capability system.
[0237] Training method framework such as Figure 3 As shown, the algorithm's execution flow follows a closed-loop structure of "environment initialization - course loading - adversarial training - policy evaluation - course update". The course learning module selects a course stage based on the current training progress and transmits the environmental parameters of the current stage to the adversarial environment, including the opponent's drone strategy type and reward function weights. The system collects environmental observation information, inputs it into the decision network to generate action commands, and then the MAPPO algorithm updates the network parameters based on environmental feedback. When the agent reaches the performance index under the current course, the system automatically loads the next stage of training tasks, gradually increasing the adversarial complexity.
[0238] 1.2.2 Phased curriculum and dynamic reward design.
[0239] This application divides multi-UAV combat training into three progressive stages. Each stage not only focuses on cultivating different combat capabilities but also achieves precise learning guidance through dynamic adjustment of the reward function weights. The dynamic adjustment of the reward function follows the principle of "emphasizing key points and gradual transition," that is, emphasizing the current learning focus at each stage while retaining rewards for capabilities from previous stages, albeit with gradually decreasing weights, to ensure the integration and connection between new and old capabilities.
[0240] The first phase is a basic attack course, focusing on developing the agent's attack capabilities. In this phase, our drone engages in combat against an adversary drone employing a baseline maneuvering strategy. The adversary drone simply flies along a fixed path and lacks attack capabilities. The reward function in this phase selectively increases the weight of attack-related factors while minimizing the weight of defense-related factors.
[0241] (10);
[0242] Among them, the posture advantage reward weight Increased to 0.4 (standard value 0.2), relative height reward weight. Increase the event-triggered reward weight to 0.3 (standard value 0.15). Maintain at 0.5, while adjusting the launch penalty / reward weight. Reduced to 0.1 (standard value 0.2). Simultaneously, the missile evasion bonus will be... The weight of the reward is reduced to a minimum of 0.05, prompting the agent to focus on learning aggressive behaviors. Through this reward configuration, the agent can quickly master proactive aggressive behaviors in a simplified environment, establish aggressive awareness and basic skills, and lay the foundation for subsequent stages.
[0243] The second phase is a defense enhancement course, focusing on developing the agent's defensive capabilities. In this phase, the red team's drones' weapon usage is restricted, allowing only evasive maneuvers. The blue team's drones, on the other hand, are equipped with full attack capabilities and employ an active pursuit and attack strategy. The reward function in this phase significantly increases the weighting of defense-related factors.
[0244] (11);
[0245] Among them, the missile evasion reward weight Increased to 0.45 (standard value 0.15), postural advantage reward weight. Adjust the configuration to focus on safe distance and increase it to 0.3 (standard value 0.2), with a relative height reward weight. Increase the weight of survival-related event trigger rewards to 0.25 (standard value 0.15). Increased to 0.4 (standard value 0.2). Meanwhile, due to motion space limitations, the launch penalty / reward... The weights are temporarily set to 0. However, this setting does not mean that the agent will forget the attack capabilities learned in the first phase, because the network parameters still retain these skills. The purpose of this phase is to allow the agent to strengthen its defensive skills while maintaining its offensive awareness, and to master the ability to effectively avoid threats.
[0246] The third stage is a comprehensive adversarial course, focusing on cultivating the agent's ability to balance offense and defense. In this stage, all action restrictions are removed, opening up the entire action space, allowing the red team's drones to freely choose attack or defensive actions. The reward function in this stage restores a balanced configuration.
[0247] (12);
[0248] Table 1. Algorithm 1: CL-MAPPO Training Algorithm
[0249]
[0250]
[0251] Among them, attack-related reward weights Set to 0.35, defense-related reward weight. Set to 0.35, event reward weight Set to 0.3. The weights of each specific reward indicator are also restored to their standard values (posture advantage). =0.2, relative height =0.15, launch penalty =0.2, missile evasion weight is 0.15), prompting the agent to develop comprehensive adversarial strategy capabilities under a balance of offense and defense. At this stage, specific capabilities are no longer guided by reward functions, but rather the agent is encouraged to autonomously explore the optimal combination of adversarial strategies to form a comprehensive adversarial strategy system that combines offense and defense.
[0252] By employing a progressive and dynamically adjusted reward mechanism, the agent can focus on developing different abilities at different stages, while ensuring the continuity of learned abilities. Compared to training methods with fixed reward weights, this dynamic adjustment mechanism more effectively addresses the problem of immediate reward preference, helping the agent overcome local optima traps and form truly effective composite adversarial strategy behaviors.
[0253] The algorithm flow is shown in Algorithm 1, as shown in Table 1. During the loading of each course stage, the algorithm configures the corresponding reward weight vector according to the current stage. The agent adjusts its opponent's strategy and environmental parameters accordingly. This ensures that the agent receives the most suitable reward signal for the current learning objective at each stage, thereby efficiently completing the stage-by-stage learning task.
[0254] This invention proposes a training method for multi-UAV adversarial learning based on a dynamically adjusted reward mechanism. This method guides the agent from basic attack capabilities to mastery of defensive skills through a carefully designed three-stage training course based on the dynamically adjusted reward mechanism, and finally develops into a comprehensive adversarial agent integrating offense and defense. At the same time, it proposes a parallel environment training mechanism to accelerate the convergence speed of multi-UAV adversarial training and significantly reduce training time.
[0255] Based on the same inventive concept as described above Figure 4 This is a schematic diagram of a training device based on a dynamically adjusted reward mechanism, provided in an embodiment of the present invention. Figure 4 As shown, a multi-drone setup is provided for adversarial learning, the device 400 comprising:
[0256] Acquisition unit 401 is used to acquire a first-party drone and a second-party drone; the first-party drone is any drone among the multiple drones; the second-party drone is a drone among the multiple drones other than the first-party drone;
[0257] The first determining unit 402 is used to determine the first parameter of the attack reward configuration in the first drone and the second drone according to a preset multi-agent near-end strategy optimization algorithm when the first drone and the second drone are learning an attack course.
[0258] The second determining unit 403 is used to determine the second parameter of the defense reward configuration in the first drone and the second drone according to the multi-agent near-end policy optimization algorithm when the first drone and the second drone are learning a defense course.
[0259] The third determining unit 404 is used to determine the third parameter of the balanced reward configuration of the first drone according to the multi-agent proximal policy optimization algorithm when the first drone and the second drone are learning adversarial courses.
[0260] The fourth determining unit 405 is used to determine the target reward configuration parameters based on the first parameter, the second parameter, and the third parameter; the target reward configuration parameters are used to learn and train the multiple UAVs.
[0261] In some embodiments, the first determining unit 402 is further configured to determine a first reward parameter related to the attack capability of the first drone and a second reward parameter related to the defense capability of the second drone according to the multi-agent near-end policy optimization algorithm; and determine the first parameter based on the first reward parameter and the second reward parameter.
[0262] In some embodiments, the first determining unit 402 is further configured to obtain the first policy parameters, first value network parameters, first course configuration parameters, and first environment configuration parameters of the multiple UAVs according to the multi-agent proximal policy optimization algorithm; determine the attitude advantage reward parameters, altitude reward parameters, event triggering reward parameters, and launch penalty reward parameters of the first UAV based on the first policy parameters, first value network parameters, first course configuration parameters, and first environment configuration parameters; and use the attitude advantage reward parameters, altitude reward parameters, and event triggering reward parameters as the first reward parameters and the launch penalty reward parameters as the second reward parameters.
[0263] In some embodiments, the first determining unit 402 is further configured to obtain a first weighting coefficient of the attitude advantage reward parameter, a second weighting coefficient of the altitude reward parameter, a third weighting coefficient of the event triggering reward parameter, and a fourth weighting coefficient of the launch penalty reward parameter; and determine the first parameter based on the first weighting coefficient, the second weighting coefficient, the third weighting coefficient, the fourth weighting coefficient, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter.
[0264] In some embodiments, the second determining unit 403 is further configured to determine a third reward parameter related to the action space capability of the first drone and a fourth reward parameter related to the active pursuit attack of the second drone according to the multi-agent near-end policy optimization algorithm; and determine the second parameter based on the third reward parameter and the fourth reward parameter.
[0265] In some embodiments, the second determining unit 403 is further configured to obtain the second policy parameters, second value network parameters, second curriculum configuration parameters, and second environment configuration parameters of the multiple UAVs according to the multi-agent near-end policy optimization algorithm; determine the missile evasion reward parameters, attitude advantage reward parameters, altitude reward parameters, event triggering reward parameters, and launch penalty reward parameters of the first UAV based on the second policy parameters, second value network parameters, second curriculum configuration parameters, and second environment configuration parameters; and use the missile evasion reward parameters, attitude advantage reward parameters, altitude reward parameters, and event triggering reward parameters as the third reward parameters and the launch penalty reward parameters as the fourth reward parameters.
[0266] In some embodiments, the second determining unit 403 is further configured to acquire the fifth weighting coefficient of the missile evasion reward parameter, the sixth weighting coefficient of the attitude advantage reward parameter, the seventh weighting coefficient of the altitude reward parameter, the eighth weighting coefficient of the event triggering reward parameter, and the ninth weighting coefficient of the launch penalty reward parameter; and determine the second parameter based on the fifth weighting coefficient, the sixth weighting coefficient, the seventh weighting coefficient, the eighth weighting coefficient, the ninth weighting coefficient, the missile evasion reward parameter, the attitude advantage reward parameter, the altitude reward parameter, the event triggering reward parameter, and the launch penalty reward parameter.
[0267] In some embodiments, the third determining unit 404 is further configured to determine the fifth reward parameter, the sixth reward parameter, and the seventh reward parameter related to the first party UAV's selection of attack or defense actions according to the multi-agent proximal policy optimization algorithm; and to determine the third parameter based on the fifth reward parameter, the sixth reward parameter, and the seventh reward parameter.
[0268] In some embodiments, the third determining unit 404 is further configured to obtain the third policy parameters, third value network parameters, third curriculum configuration parameters, and third environment configuration parameters of the multiple UAVs according to the multi-agent proximal policy optimization algorithm; determine the attitude advantage reward parameters, altitude reward parameters, and launch penalty reward of the first UAV based on the third policy parameters, the third value network parameters, the third curriculum configuration parameters, and the third environment configuration parameters; determine the first weighting factor of the attitude advantage reward parameters, the second weighting factor of the altitude reward parameters, and the third weighting factor of the launch penalty reward; and determine the fifth reward parameter based on the attitude advantage reward parameters, the altitude reward parameters, the launch penalty reward, the first weighting factor, the second weighting factor, and the third weighting factor.
[0269] In some embodiments, the third determining unit 404 is further configured to obtain a first weight value of the fifth reward parameter, a second weight value of the missile evasion reward parameter, and a third weight value of the event triggering reward parameter; and determine the third parameter based on the first weight value, the second weight value, the third weight value, the fifth reward parameter, the sixth reward parameter, and the seventh reward parameter.
[0270] It should be noted that the training device based on the dynamic adjustment reward mechanism provided in the embodiments of the present invention and the configuration method provided in the aforementioned embodiments of the present invention belong to the same inventive concept. The meanings of the terms appearing here have been explained in detail above and will not be repeated here.
[0271] This invention also provides a storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0272] This invention also provides an electronic device, comprising: a processor and a memory for storing a computer program capable of running on the processor, wherein when the processor runs the computer program, it executes the steps of the method embodiments described above stored in the memory.
[0273] Figure 5 This is a schematic diagram of a hardware structure of an electronic device according to an embodiment of the present invention. The electronic device 500 includes at least one processor 501 and a memory 502. Optionally, the electronic device 500 may further include at least one communication interface 503. The various components in the electronic device 500 are coupled together through a bus system 504. It can be understood that the bus system 504 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 504 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 5 The general designated all buses as Bus System 504.
[0274] It is understood that memory 502 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 502 described in this embodiment of the invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0275] The memory 502 in this embodiment of the invention is used to store various types of data to support the operation of the electronic device 500. Examples of such data include any computer program for operation on the electronic device 500, and programs implementing the methods of this embodiment of the invention may be included in the memory 502.
[0276] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 501. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in a memory. The processor reads information from the memory and, in conjunction with its hardware, completes the steps of the aforementioned method.
[0277] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the methods described above.
[0278] In the several embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units; some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. In addition, all functional units in the various embodiments of this invention can be integrated into one processing module, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated units can be implemented in hardware or in the form of hardware plus software functional units.
[0279] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.
Claims
1. A training method based on a dynamically adjusted reward mechanism, characterized in that, The method is applied to a multi-unmanned aerial vehicle for adversarial curriculum learning, and the method comprises the following steps: obtaining a first-party unmanned aerial vehicle and a second-party unmanned aerial vehicle; the first-party unmanned aerial vehicle is any unmanned aerial vehicle in the multi-unmanned aerial vehicle; the second-party unmanned aerial vehicle is any unmanned aerial vehicle in the multi-unmanned aerial vehicle except the first-party unmanned aerial vehicle; in the case that the first-party unmanned aerial vehicle and the second-party unmanned aerial vehicle perform attack curriculum learning, determining, according to a preset multi-agent proximal policy optimization algorithm, a first reward parameter related to attack capability of the first-party unmanned aerial vehicle and a second reward parameter related to defense capability of the second-party unmanned aerial vehicle; determining, based on the first reward parameter and the second reward parameter, a first parameter of attack reward configuration in the first-party unmanned aerial vehicle and the second-party unmanned aerial vehicle; in the case that the first-party unmanned aerial vehicle and the second-party unmanned aerial vehicle perform defense curriculum learning, determining, according to the multi-agent proximal policy optimization algorithm, a third reward parameter related to action space capability of the first-party unmanned aerial vehicle and a fourth reward parameter related to active pursuit attack of the second-party unmanned aerial vehicle; determining, based on the third reward parameter and the fourth reward parameter, a second parameter of defense reward configuration in the first-party unmanned aerial vehicle and the second-party unmanned aerial vehicle; in the case that the first-party unmanned aerial vehicle and the second-party unmanned aerial vehicle perform adversarial curriculum learning, determining, according to the multi-agent proximal policy optimization algorithm, a fifth reward parameter, a sixth reward parameter and a seventh reward parameter related to selection of attack or defense action of the first-party unmanned aerial vehicle; determining, based on the fifth reward parameter, the sixth reward parameter and the seventh reward parameter, a third parameter of balance reward configuration of the first-party unmanned aerial vehicle; determining a target reward configuration parameter based on the first parameter, the second parameter and the third parameter; the target reward configuration parameter is used for learning training of the multi-unmanned aerial vehicle.
2. The training method of claim 1, wherein, The method comprises the following steps: obtaining a first strategy parameter, a first value network parameter, a first curriculum configuration parameter and a first environment configuration parameter of the multi-unmanned aerial vehicle according to the multi-agent proximal policy optimization algorithm; determining, based on the first strategy parameter, the first value network parameter, the first curriculum configuration parameter and the first environment configuration parameter, an attitude advantage reward parameter, a height reward parameter and an event trigger reward parameter of the first-party unmanned aerial vehicle and a launch penalty reward parameter of the second-party unmanned aerial vehicle; taking the attitude advantage reward parameter, the height reward parameter and the event trigger reward parameter as the first reward parameter and taking the launch penalty reward parameter as the second reward parameter.
3. The training method of claim 2, wherein, The method comprises the following steps: obtaining a first weight coefficient of the attitude advantage reward parameter, a second weight coefficient of the height reward parameter, a third weight coefficient of the event trigger reward parameter and a fourth weight coefficient of the launch penalty reward parameter; determine the first parameter based on the first weight coefficient, the second weight coefficient, the third weight coefficient, the fourth weight coefficient, the attitude advantage reward parameter, the height reward parameter, the event trigger reward parameter and the launch penalty reward parameter.
4. The training method of claim 1, wherein, The third reward parameter related to the action space capability of the first unmanned vehicle and the fourth reward parameter related to the active pursuit attack of the second unmanned vehicle are determined according to the multi-agent proximal policy optimization algorithm, and the third reward parameter and the fourth reward parameter include: The second strategy parameter, the second value network parameter, the second curriculum configuration parameter and the second environment configuration parameter of the multiple unmanned vehicles are obtained according to the multi-agent proximal policy optimization algorithm. The missile evasion reward parameter, the attitude advantage reward parameter, the height reward parameter, the event trigger reward parameter of the first unmanned vehicle and the launch penalty reward parameter of the second unmanned vehicle are determined based on the second strategy parameter, the second value network parameter, the second curriculum configuration parameter and the second environment configuration parameter. The missile evasion reward parameter, the attitude advantage reward parameter, the height reward parameter and the event trigger reward parameter are taken as the third reward parameter, and the launch penalty reward parameter is taken as the fourth reward parameter.
5. The training method of claim 4, wherein, The second parameter is determined based on the third reward parameter and the fourth reward parameter, and the second parameter includes: The fifth weight coefficient of the missile evasion reward parameter, the sixth weight coefficient of the attitude advantage reward parameter, the seventh weight coefficient of the height reward parameter, the eighth weight coefficient of the event trigger reward parameter and the ninth weight coefficient of the launch penalty reward parameter are obtained. The second parameter is determined based on the fifth weight coefficient, the sixth weight coefficient, the seventh weight coefficient, the eighth weight coefficient, the ninth weight coefficient, the missile evasion reward parameter, the attitude advantage reward parameter, the height reward parameter, the event trigger reward parameter and the launch penalty reward parameter.
6. The training method of claim 1, wherein, The fifth reward parameter related to the selection of attack or defense action of the first unmanned vehicle is determined according to the multi-agent proximal policy optimization algorithm, and the fifth reward parameter includes: The third strategy parameter, the third value network parameter, the third curriculum configuration parameter and the third environment configuration parameter of the multiple unmanned vehicles are obtained according to the multi-agent proximal policy optimization algorithm. The attitude advantage reward parameter, the height reward parameter and the launch penalty reward of the first unmanned vehicle are determined based on the third strategy parameter, the third value network parameter, the third curriculum configuration parameter and the third environment configuration parameter. The first weight factor of the attitude advantage reward parameter, the second weight factor of the height reward parameter and the third weight factor of the launch penalty reward are determined. The fifth reward parameter is determined based on the attitude advantage reward parameter, the height reward parameter, the launch penalty reward, the first weight factor, the second weight factor and the third weight factor.
7. The training method of claim 6, wherein The sixth reward parameter includes a missile evasion reward parameter, and the seventh reward parameter includes an event trigger reward parameter. The third parameter is determined based on the fifth reward parameter, the sixth reward parameter and the seventh reward parameter, including: obtaining a first weight value of the fifth reward parameter, a second weight value of the missile avoidance reward parameter and a third weight value of the event trigger reward parameter; determining the third parameter based on the first weight value, the second weight value, the third weight value, the fifth reward parameter, the sixth reward parameter and the seventh reward parameter.
Citation Information
Patent Citations
Data security defense system based on multi-agent reinforcement learning
CN115859283A
Device and method for adjusting game difficulty based on analysis of user game ability based on artificial intelligence
KR102789064B1