Complex task-oriented agent training method and system

Through the hierarchical decision-making framework and local joint optimization algorithm, the learning efficiency and performance problems of agents in complex task scenarios are solved, and efficient agent training and learning effects are achieved.

CN120354968APending Publication Date: 2025-07-22NANKAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510439890.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing technology has insufficient overall learning efficiency and performance performance of agents in complex task scenarios, and the computational overhead of stratified global joint optimization is large, while completely independent optimization leads to insufficient coordination of sub-tasks and affects overall performance.

Method used

A hierarchical decision-making framework is adopted to independently train agents at all levels, and use local joint optimization algorithms and the 'rules-imitation-reinforcement' training paradigm, combined with reinforcement learning algorithms to jointly optimize some agents to coordinate strategies at all levels.

Benefits of technology

It improves the training efficiency and learning effect of the agent in complex tasks, improves exploration ability and strategy performance, coordinates strategies between various levels, and ensures overall learning efficiency and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354968A_ABST
    Figure CN120354968A_ABST
Patent Text Reader

Abstract

The invention provides a complex task-oriented agent training method and system, and relates to the technical field of reinforcement learning, and the method comprises the steps: constructing a hierarchical decision framework according to a complex task in an unmanned aerial vehicle air combat confrontation scene, so as to enable the complex task to be decomposed into a plurality of subtasks, and enabling a plurality of agents to execute the subtasks respectively; independently training each layer of agent under the hierarchical decision framework; and after the independent training is completed, part of agents are extracted from different frameworks according to task characteristics to form a joint optimization combination, and joint optimization is performed on the agents in the combination by using a reinforcement learning algorithm. According to the method, in complex scenes such as high-fidelity air battle games and the like, the performance of the intelligent agent in complex game tasks and the autonomous game confrontation ability of the intelligent agent are improved by deeply researching the layered local joint optimization technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of reinforcement learning, and in particular, to an agent training method and system for complex tasks. Background Art

[0002] Hierarchical reinforcement learning is a reinforcement learning method that decomposes complex tasks into multiple subtasks or levels. Each level is responsible for a specific sub-goal. Through this hierarchical structure, it improves the learning efficiency and generalization ability. Hierarchical reinforcement learning reduces the dimensionality of the state space, enabling the agent to converge to the optimal policy faster. At the same time, the knowledge of subtasks can be reused in different contexts, enhancing the generalization ability. In addition, hierarchical reinforcement learning simplifies the task description. The reward function for each subtask is relatively simple, and the overall task is clearer and more intuitive. The flexibility of the hierarchical structure enables the agent to dynamically adjust the relationship between subtasks, adapt to environmental changes, and improve robustness. Therefore, it shows significant advantages in dealing with complex and dynamic tasks. In the process of implementing the present invention, the applicant found that: previous research mainly focused on designing agent update methods for hierarchical architectures to improve their performance. Many methods achieve this goal through hierarchical global joint optimization or completely independent optimization. The hierarchical global joint optimization method aims to consider the policies and states of multiple levels simultaneously and optimize from a global perspective to achieve better overall performance. This method can usually ensure the coordination among various subtasks and has a high performance ceiling, but it involves complex optimization algorithms and large computational overhead and has fallen into local optima. On the other hand, completely independent optimization trains agents separately at each level. Each agent only focuses on the subtasks it is responsible for and does not consider the influence of other levels. This strategy simplifies the optimization process and improves the training efficiency, but it may lead to a lack of effective coordination between subtasks in multi-level collaborative tasks, affecting the overall performance.

[0003] Therefore, how to improve the overall learning efficiency and performance of agents in complex task scenarios has become a technical problem to be solved. Summary of the Invention

[0004] The present invention aims to at least solve one of the technical problems existing in the prior art or related technologies, and discloses an agent training method and system for complex tasks, introducing a hierarchical decision-making framework and proposing a local joint optimization algorithm, which can effectively coordinate the policies among various levels while ensuring the overall learning efficiency of the agent.

[0005] The first aspect of the present invention discloses an intelligent agent training method for complex tasks, including: constructing a hierarchical decision-making framework according to the complex tasks in the UAV air combat confrontation scenario, so as to decompose the complex tasks into multiple subtasks, and multiple intelligent agents respectively execute the subtasks. Among them, the hierarchical decision-making framework specifically includes: a bottom layer framework composed of UAV controller intelligent agents, a middle layer framework composed of attack tactic intelligent agents and defense tactic intelligent agents, and an upper layer framework composed of tactic regulation intelligent agents; independently training the intelligent agents at each level of the hierarchical decision-making framework; after the independent training is completed, extract some intelligent agents from different frameworks according to the task characteristics to form a joint optimization combination, and use the reinforcement learning algorithm to jointly optimize the intelligent agents in the combination.

[0006] In this technical solution, the intelligent agents in each level framework have different training objectives, and each is responsible for exploring part of the decision-making space under complex tasks. A single intelligent agent alone cannot complete the task, and all intelligent agents need to be combined together to solve complex tasks. Based on the task attributes and the design objectives of each level of intelligent agents, some intelligent agents in the hierarchical framework can be extracted and combined, and they can be jointly optimized using reinforcement learning, rather than jointly optimizing the entire hierarchical framework. This can effectively coordinate the strategies between different levels while ensuring the overall learning efficiency of the intelligent agents. It solves the problems such as low exploration efficiency and low sample utilization rate caused by the overly large action space.

[0007] According to the intelligent agent training method for complex tasks disclosed by the present invention, preferably, the independent training is carried out according to the "rule - imitation - reinforcement" training paradigm. The "rule - imitation - reinforcement" training paradigm specifically includes: constructing rules: constructing expert rules for each level based on an expert system, and using the constructed expert rules for sampling to construct an expert demonstration database; imitation learning: using the expert demonstration database for imitation learning to endow each layer of intelligent agents with initial intelligence; reinforcement learning: constructing an expert-guided reinforcement loss to achieve the secondary update of the intelligent agents.

[0008] According to the intelligent agent training method for complex tasks disclosed by the present invention, preferably, the decision-making algorithm based on a behavior tree is used to construct the expert rules to generate the mapping from the current situation to the intelligent decision; when using the expert rules for sampling to construct the expert demonstration database, the sampling scenarios should be randomly initialized to ensure the randomness and diversity of the samples.

[0009] According to the intelligent agent training method for complex tasks disclosed by the present invention, preferably, the loss L I for endowing the intelligent agents with initial intelligence using imitation learning is calculated as follows:

[0010]

[0011] where is based on the policy π BThe mean of the random variable X during sampling, π B is the expert rule strategy, indicating the advantage function of action a compared to other actions under the current strategy π B and state s, π I is the agent currently being trained, and β is used to regulate the influence degree of the advantage function.

[0012] According to the agent training method for complex tasks disclosed in the present invention, preferably, in the reinforcement learning stage, the constructed expert-guided reinforcement loss calculation formula is as follows:

[0013] L = L reinforce + αL imitation

[0014] L reinforce = E t (min(μ t A t , clip(μ t , 1 - ∈, 1 + ∈)A t ))

[0015] L imitation = D KL (π R (a|s), π I (a|s))

[0016] where α is the weight factor to regulate the influence degree of expert knowledge on training; L reinforce is the reinforcement learning loss, calculated using the proximal policy gradient optimization algorithm to promote the agent to optimize in the direction of obtaining higher rewards, E t (X) is the mean of the random variable X, is the probability factor, π R and respectively represent the policies before and after update, A t is the advantage function, ∈ is the gradient clipping factor; L imitation is the expert-guided loss to encourage the agent to explore near the expert policy, thereby preventing the agent training from deviating from the reasonable direction, D KL (P, Q) represents the KL divergence between the probability distribution P and the probability distribution Q, used to evaluate the difference size between the two probability distributions, π I is the policy distribution obtained through imitation learning, π R then represents the policy distribution obtained through reinforcement learning.

[0017] According to the intelligent agent training method for complex tasks disclosed by the present invention, preferably, the task objectives of the UAV controller intelligent agent include: generating the throttle control signal and rudder angle control signal of the aircraft according to the target altitude, target speed, and target heading, so that the aircraft can perform maneuvers according to the throttle control signal and rudder angle control signal.

[0018] According to the intelligent agent training method for complex tasks disclosed by the present invention, preferably, the task objectives of the attack tactic intelligent agent include: maximizing the game advantage to shoot down the opponent, and the attack tactics at least include the attack maneuver instruction strategy and the attack weapon instruction strategy.

[0019] According to the intelligent agent training method for complex tasks disclosed by the present invention, preferably, the task objectives of the defense tactic intelligent agent include: improving the aircraft's own missile evasion ability, and the defense tactics at least include the defense maneuver instruction strategy and the defense weapon instruction strategy.

[0020] According to the intelligent agent training method for complex tasks disclosed by the present invention, preferably, the task objectives of the tactical regulation intelligent agent include: regulating the offensive and defensive tactics to be selected currently according to the real-time situation of the confrontation scenario, and the offensive and defensive tactics are the tactics obtained based on the training of the intelligent agents in the middle-level framework.

[0021] The second aspect of the present invention discloses an intelligent agent training system for complex tasks, including: a memory for storing program instructions; a processor for calling the program instructions stored in the memory to implement the intelligent agent training method for complex tasks according to any one of the above technical solutions.

[0022] The beneficial effects of the present invention at least include: by introducing a hierarchical decision-making framework, the present invention effectively improves the training efficiency and learning effect of intelligent agents when facing complex tasks. By using expert rules to provide training samples at the initial stage of training, intelligent agents can quickly obtain initial intelligence through imitation learning; as training progresses, the reinforcement learning algorithm is used to enhance the intelligent agents to further improve their decision-making ability; finally, some core intelligent agents in the hierarchical framework are extracted and combined according to the task characteristics, and they are jointly optimized using reinforcement learning. This method can effectively coordinate the strategies between different levels under the hierarchical framework, and can also improve the overall learning efficiency and performance, providing an efficient and comprehensive solution for the learning of intelligent agents that need to cope with complex environments in various fields. Description of the Drawings

[0023] Figure 1 Shows a schematic flow chart of an intelligent agent training method for complex tasks according to an embodiment of the present invention.

[0024] Figure 2 Shows a schematic block diagram of an intelligent agent training system for complex tasks according to an embodiment of the present invention. Detailed implementation manners

[0025] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0026] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the present invention is not limited to the limitations of the specific embodiments disclosed below.

[0027] As Figure 1 shown, according to an embodiment of the present invention, a method for training an agent for complex tasks in a high-fidelity 1v1 air combat game scenario is disclosed, including:

[0028] S1. Construct a hierarchical decision-making framework for complex tasks:

[0029] Design a three-layer framework for the air combat game. The bottom layer is used as a UAV controller, which can provide the mapping from the target altitude, target speed and target heading of the UAV to the throttle and three rudder angles of the UAV; based on the UAV controller, the middle layer designs attack tactics and defense tactics. The attack tactics aim to maximize the game advantage and shoot down the opponent, and the defense tactics aim to improve the UAV's own missile evasion ability. Both tactics include two strategies, which are respectively used to control the maneuver and weapon commands of the UAV; based on the middle-layer tactics obtained through training, the upper layer adjusts the current attack and defense tactics to be selected in real time according to the situation.

[0030] S2. Construct a "rule-imitation-reinforcement" training paradigm to realize the independent optimization of agents at each level:

[0031] First, construct expert rules based on the behavior tree, and then use the expert rules to sample [s t , a t , r t , s t+1 in a randomly initialized scenario, where s t is the state at the current moment t, a t represents the action, r t represents the reward, and s t+1 represents the state at the next moment. Construct an expert demonstration database to ensure the randomness and diversity of the samples.

[0032] Subsequently, endow the agent with initial intelligence using imitation learning. The loss calculation used in imitation learning is as follows:

[0033]

[0034] Among them, is based on the policy πB The mean of the random variable X during sampling, π B is the expert rule strategy, indicating that under the current strategy π B and in the state s, the advantage function of action a compared to other actions, π I is the currently training agent, and β is used to regulate the influence degree of the advantage function.

[0035] Next, in the reinforcement learning stage, the constructed expert-guided reinforcement loss calculation formula is:

[0036] L = L reinforce + αL imitation

[0037] where α is the weight factor used to regulate the influence degree of expert knowledge on training, L reinforce is the reinforcement learning loss to promote the agent to optimize in the direction of obtaining higher rewards, and L imitation is the expert-guided loss to encourage the agent to explore near the expert strategy, thereby preventing the agent training from deviating from the reasonable direction. In this embodiment, the proximal policy gradient optimization algorithm is used to calculate the loss:

[0038] L reinforce = E t (min(μ t A t , clip(μ t , 1 - ∈, 1 + ∈)A t ))

[0039] L imitation = D KL (π R (a|s), π I (a|s))

[0040] where E t (X) is the mean of the random variable X, is the probability factor, π R and respectively represent the policies before and after update, A t is the advantage function, and ∈ is the gradient clipping factor; D KL (P, Q) represents the KL divergence between the probability distributions P and Q, used to evaluate the difference size between two probability distributions, π I is the policy distribution obtained through imitation learning, and π R represents the policy distribution obtained through reinforcement learning.

[0041] All three designed decision-making frameworks adopt the above process for independent optimization training.

[0042] S3. Use the local joint optimization algorithm to jointly update the agents:

[0043] Based on the independently optimized three-layer agents, a local joint optimization algorithm is designed to achieve the secondary joint optimization of the agents. Specifically, considering that the control of the underlying drones and the maneuvering strategies in the middle-layer tactics are tightly coupled, in this embodiment, the weapon control strategy in the middle-layer tactics and the upper-layer tactical regulation agents are extracted and combined, and the loss function L of the joint optimization is joint as follows:

[0044]

[0045] where is the policy loss of the weapon control strategy in the middle-layer tactics, and L top represents the policy loss of the upper-layer tactical regulation agents. In this embodiment, both are calculated by the proximal policy gradient optimization algorithm. Specifically, first, taking the victory or defeat of the game as the reward function, calculate and update to obtain the joint optimization value function, and then calculate the advantage function based on this value function, and call the proximal policy gradient optimization algorithm to simultaneously optimize the weapon control strategy in the middle-layer tactics and the upper-layer tactical regulation strategy until the training requirements are met (the training stop requirement for each stage in this embodiment is 10,000 training iterations), so as to achieve the joint optimization of the hierarchical framework.

[0046] As Figure 2 shown, according to another embodiment of the present invention, an agent training system 200 for complex tasks is also disclosed, including: a memory 201 for storing program instructions; a processor 202 for calling the program instructions stored in the memory to implement the agent training method for complex tasks as described in the above embodiment.

[0047] In summary, the present invention provides an agent training method for complex tasks based on hierarchical local joint optimization technology. By introducing a hierarchical decision-making framework, the training efficiency and learning effect of the agent in the face of complex tasks are effectively improved; by constructing a "rule - imitation - reinforcement" training paradigm, the exploration ability and policy performance of the agent are effectively improved; by extracting and combining some core agents in the hierarchical framework, a joint optimization algorithm is constructed to efficiently coordinate the strategies between different levels under the hierarchical framework, while ensuring the overall learning efficiency and performance, providing an efficient and comprehensive solution for agent learning that needs to cope with complex environments in various fields.

[0048] All or part of the steps in the various methods of the above embodiments can be completed by controlling related hardware through a program, and the program can be stored in a readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, magnetic tape memories, or any other readable medium capable of carrying or storing data.

[0049] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent agent training method for complex tasks, characterized in that Including: Construct a hierarchical decision-making framework according to the complex tasks in the UAV air combat confrontation scenario, so as to decompose the complex tasks into multiple subtasks, and multiple agents execute the subtasks respectively. Among them, the hierarchical decision-making framework specifically includes: a bottom layer framework composed of UAV controller agents, a middle layer framework composed of attack tactic agents and defense tactic agents, and an upper layer framework composed of tactic regulation agents; Independently train each level of agents under the hierarchical decision-making framework; After the independent training is completed, extract some agents from different frameworks according to the task characteristics to form a joint optimization combination, and use the reinforcement learning algorithm to jointly optimize the agents in the combination.

2. The intelligent agent training method for complex tasks according to claim 1, wherein Perform the independent training according to the training paradigm of "rule - imitation - reinforcement", and the training paradigm of "rule - imitation - reinforcement" specifically includes: Construct rules: Construct expert rules for each level based on the expert system, and use the constructed expert rules for sampling to construct an expert demonstration database; Imitation learning: Use the expert demonstration database for imitation learning to endow each level of agents with initial intelligence; Reinforcement learning: Construct an expert-guided reinforcement loss to realize the secondary update of the agents.

3. The intelligent agent training method for complex tasks according to claim 2, wherein Use a decision-making algorithm based on the behavior tree to construct expert rules to generate a mapping from the current situation to intelligent decisions; when using expert rules for sampling to construct an expert demonstration database, the sampling scenarios should be randomly initialized to ensure the randomness and diversity of the samples.

4. The intelligent agent training method for complex tasks according to claim 2, characterized in that Loss L for endowing an agent with initial intelligence using imitation learning I The calculation formula is as follows: Among them, is the mean of the random variable X when sampling based on the policy π, B where π B is the expert rule policy, represents the advantage function of the action a compared to other actions under the current policy π B and the state s, where π I is the agent being currently trained, and β is used to regulate the influence degree of the advantage function.

5. The intelligent agent training method for complex tasks according to claim 2, wherein In the reinforcement learning stage, the calculation formula of the expert-guided reinforcement loss constructed is as follows: L = L reinforce + αL imitation L reinforce = E t (min(μ t A t , clip(μ t , 1 - ∈, 1 + ∈)A t )) L imitation = D KL (π R (a|s), π I (a|s)) Among them, α is a weight factor to regulate the influence degree of expert knowledge on training; L reinforce is the reinforcement learning loss, calculated by using the proximal policy gradient optimization algorithm to promote the optimization of the agent in the direction of obtaining higher rewards, E t () represents the mean of the random variable in the parentheses, is the probability factor, π R and respectively represent the policies before and after update, A t is the advantage function, ∈ is the gradient clipping factor; L imitation is the expert-guided loss to encourage the agent to explore near the expert policy, thereby preventing the agent from deviating from the reasonable direction during training, D KL () represents the KL divergence between two probability distributions, used to evaluate the difference size between two probability distributions, π I is the policy distribution obtained through imitation learning, π R then represents the policy distribution obtained through reinforcement learning.

6. The intelligent agent training method for complex tasks according to claim 1, wherein The task objectives of the UAV controller agents include: Generate the throttle control signal and rudder angle control signal of the own aircraft according to the target altitude, target speed, and target heading, so that the own aircraft maneuvers according to the throttle control signal and the rudder angle control signal.

7. The intelligent agent training method for complex tasks according to claim 1, wherein The task objectives of the attack tactic agents include: Maximize the game advantage to shoot down the opponent, and the attack tactics at least include an attack maneuver instruction strategy and an attack weapon instruction strategy.

8. The intelligent agent training method for complex tasks according to claim 1, wherein The task objectives of the defense tactic agents include: Improve the own aircraft's ability to avoid missiles, and the defense tactics at least include a defense maneuver instruction strategy and a defense weapon instruction strategy.

9. The intelligent agent training method for complex tasks according to claim 1, wherein The task objectives of the tactic regulation agents include: Regulate the offensive and defensive tactics to be selected currently according to the real-time situation of the confrontation scenario, and the offensive and defensive tactics are the tactics obtained from the training of the agents in the middle layer framework.

10. An intelligent agent training system for complex tasks, characterized in that, Including: A memory for storing program instructions; A processor for calling the program instructions stored in the memory to implement the intelligent agent training method for complex tasks as described in any one of claims 1 to 9.