Diffusion-teaching learning multi-agent game algorithm for emergency scene
By combining teaching learning and diffusion models, the problem of poor adaptability of agents in multi-agent games in emergencies is solved, and efficient decision-making and rapid adaptation of agents in complex dynamic environments is achieved.
Patent Information
- Application Number
- CN202510241201.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-13
AI Technical Summary
In the field of multi-agent game, it is difficult for agents to maintain stability and adaptability in complex and dynamic burst scenarios. Traditional methods rely on a large amount of environmental interaction and expert data, and the training efficiency is low and difficult to adapt to emergencies.
Combining teaching learning and diffusion model, through the combination of teaching data and synthetic data generated by diffusion model, we will improve the diversity of data and the generalization ability of models, and enhance the agent's emergency response ability in sudden scenarios.
It significantly improves the decision-making ability and adaptability of the agent in emergencies, improves the generalization ability of the model, reduces data bias and overfitting, and ensures that the agent can quickly adapt to complex dynamic environments and make efficient decisions.
Smart Images

Figure CN120146093A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-agent games, and specifically relates to a diffusion-teaching learning multi-agent game algorithm for emergency scenarios. Background Art
[0002] In the field of multi-agent games, how to maintain the stability and generalization ability of agents in complex and dynamic environments is a key challenge in practical deployment. Especially in emergency scenarios, when agents face sudden changes in the environment or enemy strategies, they often cannot effectively adapt without sufficient training. Therefore, how to improve the emergency response ability of agents in complex and emergency environments is an important research direction in this field.
[0003] Traditional multi-agent game algorithms, such as reinforcement learning-based methods, rely on a large number of environmental interactions to collect data and train. Reinforcement learning algorithms optimize the strategies of agents through interactions with the environment, but this method usually requires a large number of trials and errors in the environment, with low training efficiency and long training time. In addition, in many high-risk or high-cost environments, it may be very difficult to collect sufficient training data. Especially in complex and dynamic adversarial environments, agents need to be able to react quickly rather than relying on a large number of environmental interactions.
[0004] To improve training efficiency and address this challenge, teaching learning has gradually become an effective solution. Teaching learning helps agents learn by using expert demonstration data, avoiding the process of exploration from scratch. Through the demonstration data provided by human experts, agents can adapt to specific tasks more quickly, especially in high-risk or complex environments. However, a key problem with teaching learning lies in the diversity of data. Traditional teaching data usually only covers limited scenarios and cannot comprehensively reflect various emergency situations. Therefore, how to expand the diversity of the dataset, especially for training in emergency scenarios, has become a difficult point in teaching learning.
[0005] To solve this problem, diffusion models, as a new type of generative model, have made significant progress in the field of data generation in recent years. Diffusion models generate synthetic data similar to real data by gradually adding noise to the data and removing the noise through the inverse process. This ability to generate data makes diffusion models an important tool in reinforcement learning. Especially in cases where data is insufficient or the environment is complex, diffusion models can generate high-quality synthetic data, enrich the training dataset, and improve the generalization ability of the model.
[0006] In multi-agent games, agents not only need to learn how to make decisions in regular scenarios but also maintain decision stability in complex and dynamic emergency scenarios. Existing algorithms often rely on a large number of environmental interactions and expert data for training. However, in the face of emergencies, traditional methods often show low adaptability. To improve the adaptability of agents in emergency scenarios, this study proposes a multi-agent game algorithm that combines learning from demonstration and diffusion models. By combining demonstration data and synthetic data generated by diffusion models, not only the diversity of data is enhanced, but also the emergency response ability of the model in emergency scenarios is strengthened.
[0007] In summary, the main problems in the current multi-agent game field include how to effectively handle emergency scenarios, how to expand and optimize training data, and how to improve the decision stability and adaptability of agents. This study provides innovative ideas for solving these problems by combining learning from demonstration and diffusion models. Especially in dealing with emergency scenarios, the proposed algorithm has great application potential and practical value. Summary of the Invention
[0008] The present invention proposes a diffusion-demonstration learning multi-agent game algorithm for emergency scenarios, aiming to improve the stability and adaptability of multi-agent game systems in complex environments. By combining learning from demonstration and diffusion models, the decision-making ability of agents in emergency scenarios can be effectively enhanced.
[0009] To solve the above technical problems, the technical solution adopted by the present invention is: a diffusion-demonstration learning multi-agent game algorithm for emergency scenarios, and the algorithm includes the following steps:
[0010] S1. Model the multi-agent game problem based on learning from demonstration, formalize it into a Markov decision process as the basic structure of the dataset; through the learning from demonstration method, use the demonstration data of human experts to fine-tune the agents to help the agents quickly adapt to the complex environment and make reasonable decisions in emergency scenarios;
[0011] S2. Collect diverse demonstration data through behavior tree strategies and human operations, covering different tactical decisions to improve the quality of demonstration data; construct a demonstration dataset based on behavior tree strategies and human operation data collection, enabling agents to learn efficiently in the face of uncertain environments;
[0012] S3. Use a diffusion model to perform empirical augmentation on the demonstration data. Since the acquisition cost of the demonstration data is high, and in the face of complex and sudden scenarios, insufficient samples may affect the learning effect of the intelligent agent, the present invention uses a diffusion model to perform empirical augmentation on the demonstration data. The diffusion model generates synthetic data through a step-by-step denoising process, increasing the diversity and representativeness of the data, and optimizing the data quality by combining perceptual loss and similarity comparison loss, thereby improving the generalization ability of the model, enhancing the training effect of the model, and avoiding data bias and overfitting.
[0013] S4. Build a demonstration decision attention network. To further optimize the decision-making ability of the intelligent agent, the present invention designs and builds a decision attention network based on the demonstration data. The demonstration decision attention network is based on the decision attention network. The demonstration decision attention network receives the demonstration data of the sudden scenario and the offline data of the ordinary scenario to update the strategy, improving the real-time emergency decision-making ability of the intelligent agent in the sudden scenario.
[0014] S5. Loss calculation and multi-agent parameter update. To optimize the cooperation and confrontation ability of the multi-agent system, the loss calculation uses the mean square error loss of the upper-level macro strategy and strategy target selection. Through joint training, combining the advantages of the demonstration data of the sudden scenario and the offline data of the ordinary scenario, not only improves the decision-making ability of individual intelligent agents, but also enhances the cooperation between intelligent agents; in the process of parameter update, the strategy of the intelligent agent is optimized based on the gradient descent method, and a heterogeneous multi-agent framework is adopted, and each intelligent agent uses different network parameters to ensure that in the multi-agent system, each intelligent agent can quickly adapt to environmental changes and make efficient decisions.
[0015] In step S2 of the present invention, by integrating the behavior tree strategy with the human operation data, a demonstration data set is constructed and dynamically enhanced to further improve the learning ability and adaptability of the intelligent agent in case of emergencies. The behavior tree strategy helps to ensure that the intelligent agent can maintain efficient decision-making in a complex tactical environment and can dynamically adjust the strategy to cope with different types of emergencies.
[0016] In step S3 of the present invention, the diffusion model generates synthetic data through a step-by-step denoising process, and combines perceptual loss and similarity comparison loss to enhance the quality and diversity of the generated data. By removing noise, more diverse and real-data-like synthetic samples are generated to ensure that the generated data can fully reflect the variability in complex scenarios.
[0017] In step S4 of the present invention, based on the decision attention network, ordinary scenario offline data and augmented generated emergency scenario teaching data are received as training samples to construct a teaching decision attention network. The teaching decision attention network optimizes the decision-making ability of the agent in a complex dynamic environment by receiving ordinary scenario offline data and augmented generated emergency scenario teaching data. The teaching decision attention network can improve the emergency response ability of the agent in emergency scenarios, ensuring that the agent can quickly adjust its strategy and make efficient decision responses.
[0018] In step S5 of the present invention, the emergency scenario teaching data generated by diffusion model augmentation and ordinary scenario offline data are jointly used to train the agent, enhancing the generalization ability of the model during the training process and reducing the overfitting problem. Especially when the training samples of emergency scenarios are relatively few, the synthetic data provides an effective supplement to ensure that the model can continue to operate stably in the face of uncertainty.
[0019] In step S5 of the present invention, the decision attention network updates and optimizes the strategy through the mean square error loss function, comprehensively considering the current situation and historical data, further improving the decision stability and strategy adaptability of the agent in emergency scenarios, ensuring that the agent can accurately judge and execute the optimal strategy in a complex dynamic environment.
[0020] During the training process of step S4 of the present invention, the teaching decision attention uses the method of time series modeling and the Attention mechanism in the Transformer network to encode the state, action and reward, optimizing the decision-making ability of the agent in a complex environment.
[0021] The beneficial effects of adopting the above technical solutions are as follows: The present invention innovatively solves the adaptability problem of agents in emergency scenarios in multi-agent games by combining teaching learning and diffusion models. Its main advantages are accelerating the learning process of agents through teaching learning and generating diverse synthetic data using the diffusion model, thereby effectively expanding the training samples and improving the generalization ability of the model. In addition, the introduction of the teaching decision attention network enables the agent to quickly adapt and make efficient decisions when facing a complex and dynamically changing environment. This algorithm has significant advantages and application value as it outperforms traditional methods in emergency scenarios by enhancing the decision stability and emergency response ability of the agent, and is particularly suitable for complex and dynamic adversarial environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flowchart of a diffusion-teaching learning multi-agent game algorithm for emergency scenarios of the present invention;
[0023] Figure 2Collect diverse teaching data behavior tree logic diagram for implementation example;
[0024] Figure 3 The training flow chart of the empirical augmentation algorithm based on the diffusion model of the embodiment;
[0025] Figure 4 The flowchart of training the attention algorithm for decision making in emergency scenarios of the embodiment;
[0026] Figure 5 Schematic diagram of the trained intelligent agent in the embodiment in the confrontation between luring the enemy deep into the territory and improvisation. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only embodiments of a part of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present application. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0028] A typical embodiment of the present invention is a diffusion-teaching learning multi-agent game agent training method for emergency scenarios, such as Figure 1 As shown, the following steps are included:
[0029] S1. Modeling the multi-agent game problem based on teaching learning can be formalized as a Markov decision process M = (S, A, P, R, μ), where:
[0030] -S: state space;
[0031] -A: action space;
[0032] -P(s ′ |s,a): state transition probability;
[0033] -R(s,a): reward function;
[0034] -μ: initial state distribution.
[0035] In the teaching learning environment, the goal of the agent is to demonstrate the dataset D = {(s 0 , a 0 , r 0 ,…,s H , a H , r H )} Learn the optimal strategy π * , to maximize long-term returns:
[0036]
[0037] Among them, γ is the discount factor and π is the policy function.
[0038] S2. Collect diverse demonstration data through the behavior tree strategy and human operations. In this embodiment, we design a behavior tree strategy adapted to the characteristics of air combat missions to collect demonstration data. The behavior tree is a hierarchical task modeling method widely used in dynamic decision-making systems. In the air combat scenario, we design a multi-level behavior tree structure according to the mission requirements, and the specific decision logic is as Figure 2 shown.
[0039] The core decision-making process of the behavior tree includes:
[0040] 1. Root node startup: The behavior tree is started through the root node, and depth-first traversal is adopted to ensure that the intelligent agent preferentially executes key tasks.
[0041] 2. Left sequence node: First, judge whether the remaining number of weapons N weapon exceeds the set threshold θ threshold :
[0042] - If there are sufficient weapons, enter the selection node to judge whether the enemy target is locked. If the target is not locked, perform level flight to maintain situational awareness; if the target is locked, trigger the attack decision sub-node to decide whether to launch a missile or switch the attack target.
[0043] - If the number of weapons is insufficient, turn to the right sequence node and perform operations to avoid and escape from the battlefield.
[0044] In addition, this study also combines the collection of demonstration data from human operators. Through real-time interaction collection, the decision-making processes of multiple volunteers on the simulation platform are recorded.
[0045] The demonstration data collected by the behavior tree strategy and human operators is in the following form:
[0046] D = {(s 0 , a 0 , r 0 , …, s H , a H , r H )}
[0047] Among them, s t represents the state at time t, a t is the executed action, and r t is the obtained reward.
[0048] These data will be used for training together to provide diverse teaching samples and further improve the generalization ability of the model.
[0049] S3. Use a diffusion model to perform empirical augmentation on the teaching data. The diffusion model generates synthetic data through the processes of forward diffusion and reverse denoising. The forward diffusion process is as follows:
[0050]
[0051] The reverse denoising process is:
[0052]
[0053] For the data generated by this diffusion model, we further calculate different loss functions to optimize the quality of the generated samples:
[0054] - Mean squared error loss:
[0055]
[0056] - Perceptual loss:
[0057]
[0058] Among them, F is a pre-trained feature extraction network used to capture the high-level features of the image.
[0059] - Similarity contrast loss:
[0060] L contr = -E[cosine_sim(F(X gen )) - F(X real ))]
[0061] Combining these loss functions, we obtain the total loss:
[0062] L total = λ MSE L MSE + λ perc L perc + λ contr L contr
[0063] Figure 3 Figure shows the augmentation training process based on the diffusion model. First, set the real data ratio r ∈ [0, 1], and initialize the real data replay buffer pool D real and the synthetic data replay buffer pool D synthetic . Next, generate the policy π and its corresponding model, and update and generate samples through the diffusion model M. In each training step, collect real data from the environment and update the diffusion model, while generating synthetic samples and calculating the total loss Ltotal , until the end of training.
[0064] S4. The present invention introduces a decision attention network to improve the decision-making ability of the agent. The decision attention network learns trajectories in an autoregressive manner, inputs state, action, and reward data, and predicts future actions. Its core idea is to transform the reinforcement learning problem into a sequence modeling problem and model the sequences of state, action, and reward through a Transformer.
[0065] The representation of the trajectory in the decision attention network is as follows:
[0066]
[0067] where represents the cumulative return from time step t to the end of the game, s t is the state, and a t is the action.
[0068] To effectively capture the impact of time on decision-making, the decision attention network introduces time embedding. The time embedding is added to the state, action, and reward embeddings. The specific formula is as follows:
[0069] s t = s t + t t
[0070] a t = a t + t t
[0071]
[0072] This way enables the model to make full use of time information during the learning process and then predict future actions.
[0073] The demonstration decision attention network is an extension of the decision attention network, which combines the demonstration data τ prompt of the emergency scenario and the offline data τ offline of the normal scenario to optimize the decision-making ability of the model. The input data is represented as:
[0074] τ input =(τ prompt , τ offline )
[0075] S5. In a multi-agent system, we make decisions for each agent separately and calculate the loss of its policy optimization. For each agent, we define two loss functions:
[0076] - The loss of the macro-policy action:
[0077]
[0078] - Loss of the target selection strategy:
[0079]
[0080] The total loss is the sum of the two parts of the loss:
[0081] L total = L action + L target
[0082] The specific training process of the emergency scenario teaching decision attention network is as Figure 4 shown: Sample trajectories from the dataset and generate corresponding prompt trajectories τ ★i,m . Then, calculate the total loss L MSE , and update the model parameters through backpropagation. This process is repeated until all training iterations are completed. In this embodiment, the training stop requirement is 20,000 training iterations. Figure 5 It is the confrontation graph of the agents in the scenarios of luring the enemy in deep and making a sudden move opportunistically. In the scenario of luring the enemy in deep, in the face of the behavior of the blue side retreating to the base, the red side agent did not get entangled in the battle and saw through the ruse of the blue side to lure the enemy in deep; in the scenario of making a sudden move opportunistically, 4 enemy planes appeared from the side, and the red side agent chose to retreat and avoid the battle, readjust the formation and then conduct the confrontation.
[0083] In summary, the diffusion-teaching learning multi-agent game algorithm for emergency scenarios proposed by the present invention innovatively combines teaching learning and diffusion models, significantly improving the adaptability and decision-making ability of agents in complex and dynamic environments. By generating diverse synthetic data through the diffusion model, the coverage of training samples is expanded, thus effectively enhancing the generalization ability of the model. At the same time, the introduction of the decision attention network enables the agent to more quickly adapt to emergency scenarios and make efficient decisions, especially performing well in a dynamically changing confrontation environment. This algorithm has high practical value, especially applicable to high-risk and high-complexity confrontation environments, and can significantly improve the stability and emergency response ability of multi-agent systems.
[0084] The above embodiments are only used to illustrate rather than limit the technical solutions of the present invention. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand: It is still possible to modify the present invention or make equivalent replacements, and any modification or partial replacement without departing from the spirit and scope of the present invention shall be covered by the scope of the claims of the present invention.
Claims
1. A diffusion-teaching learning multi-agent game algorithm for emergency scenarios, characterized by: The algorithm comprises the following steps: S1. Model the multi-agent game problem based on teaching learning and formalize it into a Markov decision process as the basic structure of the data set; S2. Collect diverse teaching data through behavior tree strategies and human operations, covering different tactical decisions to improve the quality of teaching data; S3, using a diffusion model to empirically augment the teaching data, the diffusion model generates synthetic data through a step-by-step denoising process; S4. Build a teaching decision attention network. The teaching decision attention network updates the strategy by receiving teaching data of emergency scenarios and offline data of normal scenarios. S5. Loss calculation and multi-agent parameter update. The loss calculation adopts the mean square error loss selected by the upper-level macro strategy and strategy target. During the parameter update process, the agent's strategy is optimized based on the gradient descent method, and a heterogeneous multi-agent framework is adopted. Each agent uses different network parameters to ensure that in the multi-agent system, each agent can quickly adapt to environmental changes and make efficient decisions.
2. According to claim 1, a diffusion-teaching learning multi-agent game algorithm for emergency scenarios is characterized in that: In step S2, a teaching data set is constructed by integrating the behavior tree strategy with human operation data.
3. According to claim 1, a diffusion-teaching learning multi-agent game algorithm for emergency scenarios is characterized in that: In step S3, the diffusion model generates synthetic data through a step-by-step denoising process, and combines perceptual loss with similarity contrast loss to enhance the quality and diversity of the generated data, thereby generating synthetic samples that are more diverse and close to real data by removing noise.
4. A diffusion-teaching learning multi-agent game algorithm for emergency scenarios according to any one of claims 1 to 3, characterized in that: In step S4, based on the decision attention network, the normal scene offline data and the sudden scene teaching data generated by augmentation are received as training samples to construct a teaching decision attention network.
5. A diffusion-teaching learning multi-agent game algorithm for emergency scenarios according to any one of claims 1 to 3, characterized in that: In step S5, the emergency scenario teaching data generated by diffusion model augmentation and the common scenario offline data are used together to train the intelligent agent. When the training samples of emergency scenarios are relatively small, the synthetic data provides effective data supplement.
6. A diffusion-teaching learning multi-agent game algorithm for emergency scenarios according to any one of claims 1 to 3, characterized in that: In step S5, the decision attention network performs strategy update and optimization through the mean square error loss function.
7. The diffusion-teaching learning multi-agent game algorithm for emergency scenarios according to any one of claims 1 to 3, characterized in that: In the step S4, during the training process, the teaching decision attention network utilizes the time series modeling method and the Attention mechanism in the Transformer network to encode the state, action and reward.