Multi-agent collaborative interaction decision and control method in interaction scene

Through the multi-agent SAC reinforcement learning algorithm combined with LSTM timing processing and dual-empirical playback mechanism, the decision accuracy and learning speed of agents in complex interactive scenarios in traditional methods are solved, and the agents can make efficient decisions and rapid adaptation in dynamic environments are achieved.

CN120386386AActive Publication Date: 2025-07-29NORTHWESTERN POLYTECHNICAL UNIV

Patent Information

Application Number
CN202510880327.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Traditional multi-agent reinforcement learning schemes cannot effectively utilize the timing correlation of environmental states in complex interactive scenarios, resulting in agents being unable to accurately predict the movement trajectory of dynamic objects, lack of prospectiveness and accuracy in decision-making, slow learning speed and poor adaptability.

Method used

The multi-agent SAC reinforcement learning algorithm is adopted, combined with LSTM timing processing and dual experience playback mechanism, and a reasonable reward function is designed to predict and feature extraction of timing state information through the LSTM network, and a discrete incremental PID controller is used for the agent control.

Benefits of technology

It improves the accuracy and timeliness of decision-making in complex dynamic environments, improves the training speed and interactive victory rate, and ensures that the agent can quickly adapt to environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386386A_ABST
    Figure CN120386386A_ABST
Patent Text Reader

Abstract

The invention provides a multi-agent collaborative interaction decision and control method in an interaction scene. The method comprises the following steps: firstly, acquiring time sequence state information of a confrontation scene by each agent; in a decision-making system of each agent, inputting obtained confrontation scene time sequence state information into an LSTM network for prediction and feature extraction to obtain a hidden state, inputting the hidden state into an SAC reinforcement learning model, and training the LSTM network and the SAC reinforcement learning model through a designed reward function to obtain a confrontation scene time sequence state; finally, the maneuvering decision action vector of each agent is obtained; and each agent takes a maneuvering decision action vector as a control target value, and the agents are controlled according to a discrete incremental PID controller. According to the method, on the basis of a multi-agent SAC reinforcement learning algorithm, LSTM time sequence processing and a double-experience playback mechanism are combined, and a reasonable reward function mechanism is designed, so that in an environment with a complex dynamic object, the motion trail of the dynamic object can be quickly predicted, and the continuous action amount of the dynamic object can be accurately decided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of agent decision-making and control, and specifically to a multi-agent collaborative interaction decision-making and control method in an interaction scenario. Background Art

[0002] An agent is an AI entity that can perceive the environment and make autonomous decisions. A multi-agent system (MAS) breaks through the limitations of individual capabilities through cooperation, and improves the system efficiency by means of information sharing and task sharing, and is widely used in various fields. Due to problems such as partially observable environment, conflicts of individual and group interests, etc., the cooperation of MAS requires dynamic coordination rather than simple behavior superposition. Multi-agent reinforcement learning (MARL) introduces the trial-and-error optimization mechanism of reinforcement learning into MAS, and takes the Markov decision process (MDP) as the framework, and trains the group cooperation strategy through environmental feedback. Since Littman proposed the basic framework, it has now become the core technology in the fields of distributed decision-making, intelligent robots, etc.

[0003] In the context of the wide application of multi-agent systems in complex interaction scenarios (such as autonomous flight in response to air targets, autonomous driving in response to complex road conditions, etc.), agents face a dynamic and uncertain environment. In such scenarios, key information such as the position and speed of obstacles or opponents changes rapidly over time, which poses extremely high requirements for the decision-making and control capabilities of agents.

[0004] Traditional multi-agent reinforcement learning schemes have significant technical bottlenecks when dealing with the decision-making and control tasks of agents in complex interaction scenarios. Traditional schemes usually make decisions only based on the current state, ignoring the continuity and temporal correlation of environmental state changes. In an unknown dynamic obstacle environment, this limitation is particularly prominent. Due to the lack of effective utilization of past state information, agents cannot accurately capture the movement trajectories of dynamic objects and are difficult to predict their future positions, resulting in decisions lacking foresight and accuracy. For example, in the autonomous driving scenario, the vehicle may not be able to make reasonable decisions according to the changing trend of the road conditions, increasing the probability of traffic accidents. At the same time, in the training process of traditional schemes, the experience replay mechanism often adopts a single sampling strategy, such as uniform sampling or prioritized experience replay. In the case of uniform sampling, key experiences are difficult to be fully emphasized, and the learning speed is slow. Agents need to spend a lot of time to master effective decision-making strategies; while prioritized experience replay can accelerate the convergence speed, but it overemphasizes some high-value experiences and sacrifices the comprehensive exploration of the environment, making agents less adaptable and robust when facing new environments or unexpected situations and unable to flexibly handle various complex situations. Summary of the Invention

[0005] To ensure that the agent can accurately predict dynamic environment information, better adapt to the environment, and make effective decisions, the present invention proposes a multi-agent collaborative interaction decision and control method in an interaction scenario. This method is based on the multi-agent SAC reinforcement learning algorithm, combines LSTM time series processing and double experience replay mechanism, and through a reasonable reward function mechanism design, can quickly predict the movement trajectory of dynamic objects in an environment with complex dynamic objects (such as other agents and their kinetic bodies), and accurately make decisions on its own continuous action amount. Compared with traditional methods, the present invention can not only accelerate the training speed, but also greatly improve the winning rate of interactive confrontation.

[0006] The technical solution of the present invention is as follows:

[0007] A multi-agent collaborative interaction decision and control method in an interaction scenario, comprising the following steps:

[0008] Step 1: Each agent obtains the time series state information of the confrontation scenario;

[0009] Step 2: In the decision-making system of each agent, the obtained time series state information of the confrontation scenario is input into the LSTM network for prediction and feature extraction to obtain a hidden state vector , which is used as the input for subsequent reinforcement learning;

[0010] Step 3: In the decision-making system of each agent, the hidden state vector obtained in Step 2 is input into the SAC reinforcement learning model, and the LSTM network and the SAC reinforcement learning model are trained through the designed reward function, and finally the maneuver decision action vector of each agent is obtained;

[0011] Step 4: After each agent obtains the maneuver decision action vector through Step 3, using the maneuver decision action vector as the control target value, the agent is controlled according to the discrete incremental PID controller.

[0012] Further, in Step 1, the time series state information of the confrontation scenario obtained by the agent includes:

[0013] (1) The three-dimensional coordinates of the agent itself and the other agents on its own side relative to the agent itself at the current time step and historical time steps;

[0014] (2) The three-dimensional coordinates of the other agent relative to the agent itself at the current time step and historical time steps;

[0015] (3) The three-axis velocity components of the agent itself and the other agents on its own side relative to the agent itself at the current time step and historical time steps;

[0016] (4) The three-axis velocity components of the opponent agent relative to the agent itself at the current time step and historical time steps;

[0017] (5) The attitude angles of the agent itself at the current time step and historical time steps;

[0018] (6) The three-dimensional coordinates of the kinetic energy body carried by the opponent agent relative to the agent itself at the current time step and historical time steps;

[0019] (7) The three-axis velocity components of the kinetic energy body carried by the opponent agent relative to the agent itself at the current time step and historical time steps;

[0020] (8) The absolute value of the current velocity of the kinetic energy body carried by the opponent agent relative to the agent itself.

[0021] Further, in step 2, the hidden state vector has a lower dimension than the dimension of the temporal state information of the adversarial scenario obtained by the agent, and in the hidden state vector explicitly includes the predicted values of the three-dimensional coordinates of the opponent agent relative to the agent itself in the next time step.

[0022] Further, in step 3, the SAC reinforcement learning model includes an Actor network and a Critic network;

[0023] The predicted and dimension-reduced hidden state vector obtained through step 2 is input into the Actor network to generate a maneuver decision action vector a = [ pitch,roll,rudder , engine ] , and each element in the vector represents the pitch decision target value, roll decision target value, yaw decision target value, and thrust decision target value of the agent in sequence;

[0024] The hidden state vector and the generated maneuver decision action vector are input into the Critic network to output the action value .

[0025] Further, in step 3, during the training process, the designed reward function includes angle advantage reward, altitude advantage reward, speed advantage reward, and win-loss reward, and a trajectory prediction reward is also added. The final reward function is the sum of each reward function with random weights added.

[0026] Further, each reward function is:

[0027] Angle advantage reward :

[0028]

[0029] is the line of sight angle;

[0030] Height advantage reward :

[0031]

[0032] Among them, is the set ideal height; is the height of the agent itself, is the height safety interval;

[0033] Speed advantage reward :

[0034]

[0035] Among them, is the speed of the agent itself, is the effective speed interval;

[0036] Winning and losing reward :

[0037] If the opponent is defeated, the winning and losing reward takes a positive value. If the agent is defeated by the opponent, the winning and losing reward takes a negative value;

[0038] Trajectory prediction reward :

[0039]

[0040] Among them, is the position of the opponent agent n predicted by the LSTM network at the next time step, is the actual position of the opponent agent n at the next time step;

[0041] The finally obtained total reward function is:

[0042]

[0043] Among them , represents the weights of each part.

[0044] Further, during the training process of step 3, a dual experience replay mechanism is adopted. In the initial stage of training, a random experience pool is used for uniform sampling to ensure that the agent can fully explore the environment. In the middle and late stages of training, a prioritized experience pool is used for priority sampling based on the reward value to accelerate the learning of high-value experiences.

[0045] Further, a basic action library is also stored in the decision-making system of the agent. A variety of typical maneuvering actions of the agent are set in the basic action library. Under certain set conditions, the agent is not controlled by the reinforcement learning model but executes typical maneuvering actions according to the settings of the basic action library. When reaching other set conditions, the reinforcement learning model is used for control again.

[0046] In addition, the present invention also proposes an electronic device and a readable storage medium:

[0047] An electronic device includes a processor and a memory, and the memory is used to store one or more programs;

[0048] When the one or more programs are executed by the processor, the above method is implemented.

[0049] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0050] Beneficial effects:

[0051] The multi-agent collaborative interaction decision-making and control method in the interaction scenario proposed by the present invention has the following advantages:

[0052] 1. Temporal information is introduced into the state space, so that the state information obtained by the agent is no longer limited to the current moment but includes rich data of the current and historical time steps, thus being able to effectively capture the long-term dependence relationship in the data;

[0053] 2. Compared with the traditional reinforcement learning process method that directly inputs the state vector into the Actor network, in the present invention, the high-dimensional state vector containing temporal information is first predicted and dimension-reduced by an LSTM unit, and the temporal information is converted into a hidden state vector for reinforcement learning. This not only reduces the dimension and avoids the phenomenon of dimensional explosion when directly inputting the high-dimensional state vector into the Actor network, but also includes the prediction of the other party's next time step, making the predicted value gradually approach the true relative position, so as to be able to more accurately predict the future movement trajectory of the dynamic object, plan countermeasures in advance, significantly improve the accuracy and timeliness of decision-making, and provide a more comprehensive and accurate basis for the agent's decision-making;

[0054] 3. The present invention uses a double experience replay mechanism to assist reinforcement learning. In the initial stage of training, this mechanism adopts a random sampling strategy, enabling the agent to fully explore the environment, obtain diverse state and action combinations, discover more potential effective strategies, and avoid falling into local optimal solutions. As the training progresses, in the middle and late stages, it switches to a priority sampling strategy based on the reward value, preferentially learning those experiences that are more crucial for improving the system performance, accelerating the learning process, and enhancing the training efficiency. By adopting different sampling strategies at different training stages in this way, it effectively balances exploration and exploitation in reinforcement learning, ensuring that the agent can quickly learn efficient decision-making strategies while fully understanding the environment.

[0055] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0057] Figure 1 is a logical flowchart of the interaction between the agent adopting the method of the present invention and the built-in AI of simulation platforms at all levels in the embodiments of the present invention;

[0058] Figure 2 is a change curve of the sum of the cumulative reward values of the agent adopting the method of the present invention in the embodiments of the present invention.

[0059] Figure 3 is a change curve of the sum of the cumulative reward values of the agent in the traditional SAC algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The embodiments of the present invention will be described in detail below. The embodiments are exemplary and are intended to explain the present invention, but should not be construed as a limitation to the present invention.

[0061] In this embodiment, taking the 4V4 agent confrontation interaction as an example, through the multi-agent collaborative interaction decision-making and control method in the interaction scenario proposed by the present invention, the built-in decision-making model of the agent is trained and interacted with the built-in AI agents of the existing interaction simulation platform to verify the decision-making effectiveness of the method of the present invention.

[0062] Specifically, it includes the following steps:

[0063] Step 1: The agent obtains the sequential state information of the confrontation scenario;

[0064] This example is a 4V4 intelligent agent interactive confrontation scenario, including four intelligent agents of Party A and four intelligent agents of Party B. Each intelligent agent also has two kinetic energy bodies, which can be emitted by the intelligent agent and move according to a preset control rate.

[0065] Taking intelligent agent No. 1 of Party A as an example, the obtained confrontation scenario time series state information includes:

[0066] (1) In the world coordinate system, the three-dimensional coordinates of intelligent agent No. 1 of Party A and the other three intelligent agents of Party A relative to intelligent agent No. 1 of Party A at the current time step and historical time steps (such as the previous 10 steps), where the coordinate information of the current time step is expressed as:

[0067] [ x 1(i) , y 1(i) , z 1(i) , x 2(i) , y 2(i) , z 2(i) , x 3(i) , y 3(i) , z 3(i) , x 4(i) , y 4(i) , z 4(i) ]

[0068] where is the three-dimensional coordinate time series matrix of intelligent agent No. 1 of Party A in the world coordinate system, composed of the three-dimensional coordinates of intelligent agent No. 1 of Party A at the current time step and historical time steps, 、 、 are the three-dimensional coordinate time series matrices of the other three intelligent agents of Party A relative to intelligent agent No. 1 of Party A in the world coordinate system, respectively, composed of the relative three-dimensional coordinates of the other three intelligent agents of Party A relative to intelligent agent No. 1 of Party A at the current time step and historical time steps.

[0069] (2) In the world coordinate system, the three-dimensional coordinates of the four intelligent agents of Party B relative to intelligent agent No. 1 of Party A at the current time step and historical time steps:

[0070] [ x 5(i) , y 5(i) , z 5(i) , x 6(i) , y 6(i) , z 6(i) , x 7(i) , y 7(i) , z 7(i) , x 8(i) , y 8(i) , z 8(i) ]

[0071] where 、 , , They are the three - dimensional coordinate time - series matrices of the four agents of Party B relative to Agent 1 of Party A in the world coordinate system, which are composed of the relative three - dimensional coordinates of the four agents of Party B relative to Agent 1 of Party A at the current time step and historical time steps.

[0072] (3) In the world coordinate system, the three - axis velocity components of Agent 1 of Party A and the other three agents of Party A relative to Agent 1 of Party A at the current time step and historical time steps:

[0073] [ v 1x i , v 1y i , v 1z i , v 2x i , v 2y i , v 2z i , v 3x i , v 3y i , v 3z i , v 4x i , v 4y i , v 4z i ]

[0074] Among them is the three - dimensional velocity time - series matrix of Agent 1 of Party A in the world coordinate system, which is composed of the three - axis velocity components of Agent 1 of Party A at the current time step and historical time steps, , , are the three - dimensional velocity time - series matrices of the other three agents of Party A relative to Agent 1 of Party A in the world coordinate system, which are composed of the relative three - axis velocity components of the other three agents of Party A relative to Agent 1 of Party A at the current time step and historical time steps.

[0075] (4) In the world coordinate system, the three - axis velocity components of the four agents of Party B relative to Agent 1 of Party A at the current time step and historical time steps:

[0076] [ v 5x(i) , v 5y(i) , v 5z(i) , v 6x(i) , v 6y(i) , v 6z(i) , v 7x(i) , v 7y(i) , v 7z(i) , v 8x(i) , v 8y(i) , v 8z(i) ]

[0077] Among them , , , They are the three-dimensional velocity time series matrices of the four agents on Party B with respect to Agent No. 1 on Party A in the world coordinate system, which are composed of the relative three-axis velocity components of the four agents on Party B with respect to Agent No. 1 on Party A at the current time step and historical time steps.

[0078] (5) The attitude angles of the four agents on Party A at the current time step and historical time steps, including:

[0079] Pitch angle:

[0080]

[0081] Roll angle:

[0082]

[0083] Yaw angle:

[0084] ;

[0085] Among them , , , They are the attitude angle matrices of the four agents on Party A in sequence, which are composed of the attitude angles of the four agents on Party A at the current time step and historical time steps.

[0086] (6) Each of the four agents on Party B carries two releasable kinetic energy bodies. Therefore, the time series state information of the confrontation scenario also includes the three-dimensional coordinates of the eight kinetic energy bodies on Party B with respect to Agent No. 1 on Party A at the current time step and historical time steps in the world coordinate system:

[0087] , , , , , , ,

[0088] They are the three-dimensional coordinate time series matrices of the eight kinetic energy bodies on Party B with respect to Agent No. 1 on Party A in sequence, which are composed of the relative three-dimensional coordinates of the eight kinetic energy bodies on Party B with respect to Agent No. 1 on Party A at the current time step and historical time steps.

[0089] (7) The three-axis velocity components of the eight kinetic energy bodies on Party B with respect to Agent No. 1 on Party A at the current time step and historical time steps in the world coordinate system:

[0090] , , , , , , ,

[0091] They are respectively the three-dimensional velocity time series matrices of the eight kinetic energy bodies of Party B relative to Agent 1 of Party A in the world coordinate system, which are composed of the relative three-dimensional velocity components of the eight kinetic energy bodies of Party B relative to Agent 1 of Party A at the current time step and historical time steps.

[0092] (8) The absolute value of the current velocity of the eight kinetic energy bodies relative to Agent 1 of Party A in the world coordinate system , , , , , , , .

[0093] It can be seen that the adversarial scenario time series state information obtained by a single agent is a 1196-dimensional high-dimensional state vector. The decision-making systems of each agent obtain the adversarial scenario time series state information at set time intervals. The time interval in this embodiment is set to .

[0094] Step 2: In the decision-making systems of each agent, the obtained adversarial scenario time series state information is input into the LSTM network for prediction and feature extraction to obtain a hidden state vector, which is used as the input for subsequent reinforcement learning.

[0095] Still taking Agent 1 of Party A as an example, the decision-making system of Agent 1 of Party A inputs the 1196-dimensional high-dimensional time series state information corresponding to this agent into the LSTM network of its own decision-making system for prediction and feature extraction to obtain a 256-dimensional hidden state vector . It contains the key dependencies in the timing information, providing more accurate and comprehensive information for Agent 1 of Party A. Moreover, the last 12 dimensions of the 256-dimensional data are explicitly configured as the predicted three-dimensional coordinate values of the four agents of Party B relative to Agent 1 of Party A in the next time step. This information is also predicted through the LSTM network, which is the key feature distinguishing it from the prior art where the LSTM network only performs encoding. By predicting in advance the predicted three-dimensional coordinate values of the four agents of Party B relative to Agent 1 of Party A in the next time step, it can more accurately guide the decision-making of Agent 1 of Party A; the remaining 244 dimensions are implicit decision features, and the LSTM network gating mechanism automatically extracts the timing dependencies, mainly including the following information: (1) the coordinate trends, speed stability, and attitude coordination of the four agents of Party A; (2) threat features such as the speed change information of the agents of Party B and the risk of kinetic energy body impact; (3) the effectiveness of historical actions and the memory of threat events (such as successful obstacle avoidance actions).

[0096] Here is a theoretical explanation of the LSTM network:

[0097] The LSTM network processes the input sequence data through a series of gating mechanisms and can effectively capture the long-term dependencies in the data. Its core formula is as follows:

[0098] Input gate: , the input gate determines how much information in the current input should be added to the cell state. is the sigmoid activation function, which maps the input value to the interval (0, 1), enabling the input gate to output a value between 0 and 1 to control how much information in the current input should be added to the cell state. is the weight matrix of the input gate, which determines the weight distribution of the input information and the hidden state vector at the previous moment in the input gate, and is continuously adjusted during the training process to optimize the network performance. is the bias vector, providing an additional learnable parameter for the calculation of the input gate, which helps the model better fit the data. represents concatenating the hidden state vector at the previous moment and the current input . This concatenation operation enables the input gate to comprehensively consider the current information and historical information, thereby more accurately determining which information should be incorporated into the cell state.

[0099] Forget gate: , the forget gate determines how much information in the cell state at the previous moment should be forgotten. is the weight matrix of the forget gate, which controls the influence degree of the previous hidden state vector and the current input on the forget gate decision. During the training process, the function of the forget gate is optimized by adjusting the weights. is the bias vector, which adds flexibility to the calculation of the forget gate. Through the sigmoid activation function , the forget gate outputs a value between 0 and 1. If the value is close to 1, it means that the system tends to retain most of the past information; if it is close to 0, it means that the system tends to forget most of the past information. In this way, the forget gate can dynamically adjust the retention degree of historical information according to the current input and the previous state, enabling the cell state to better adapt to environmental changes.

[0100] Cell state update: , is a candidate cell state, which is obtained by processing through the hyperbolic tangent activation function tanh. Tanh can highlight important information and suppress unimportant information, enabling the candidate cell state to effectively capture key features in the current input and historical state. is the weight matrix, is the bias vector, and they are continuously optimized during the training process to adjust the generation of the candidate cell state. The latter formula shows the update process of the cell state. Among them, represents retaining the part needed in the previous cell state according to the output of the forget gate , ⊙ represents element-wise multiplication, that is, only when the corresponding elements in are close to 1, the information at the corresponding position in will be retained. represents adding the appropriate part in the candidate cell state to the current cell state according to the output of the input gate. In this way, the cell state can not only retain important historical information but also update the current information in a timely manner, thus better reflecting the dynamic changes of the environment.

[0101] Output gate: , and are the weight matrix and the bias vector respectively, and they are optimized during the training to determine how much information in the current cell state should be output to the hidden state vector . The function maps the output value to the interval (0,1) to control the output ratio. represents selecting the appropriate information from the processed cell state as the output of the hidden state vector according to the control of the output gate. The finally obtained hidden state vector It includes the short-term features of the current time step and the long-term dependencies of historical information, providing key information for the decision-making of agents in the four-agent interactive confrontation scenario of this embodiment, and helping the agents predict the actions of opponents and plan their own actions.

[0102] Step 3: In the decision-making system of each agent, input the hidden state vector obtained in Step 2 into the SAC reinforcement learning model, and train the LSTM network and the SAC reinforcement learning model through the designed reward function, and finally obtain the maneuver decision action vector of each agent.

[0103] The SAC reinforcement learning model includes an Actor network and a Critic network.

[0104] The predicted and dimension-reduced hidden state vector obtained through Step 2 is input into the Actor network to generate a maneuver decision action vector a = [ pitch,roll,rudder , engine ] , and each element in the vector represents the pitch decision target value, roll decision target value, yaw decision target value, and thrust decision target value of the agent in turn.

[0105] Input the hidden state vector and the generated maneuver decision action vector into the Critic network to output the action value .

[0106] Compared with the traditional reinforcement learning process method that directly inputs the state vector into the Actor network, the present invention first performs prediction and dimension reduction processing on the high-dimensional state vector containing temporal information through the LSTM unit, and converts the temporal information into a hidden state vector for reinforcement learning, which not only reduces the dimension and avoids the phenomenon of dimensional explosion when directly inputting the high-dimensional state vector into the Actor network, but also includes the prediction of the opponent's next time step, making the predicted value gradually approach the true relative position, so as to be able to more accurately predict the future motion trajectory of dynamic objects, plan countermeasures in advance, and significantly improve the accuracy and timeliness of decision-making.

[0107] During the training process, the designed reward function includes angle advantage reward, height advantage reward, speed advantage reward, and win-loss reward, and a trajectory prediction reward is also added to help the agent better adapt to the dynamic environment. Finally, the reward function is the sum of each reward function with random weights added.

[0108] (1) Angle advantage reward

[0109] Physical meaning: line-of-sight angle (The angle between our velocity vector and the line connecting us and the enemy) The smaller it is, the stronger the ability to lock on to the target and the higher the reward.

[0110] Formula:

[0111]

[0112] (2) Altitude advantage reward

[0113] Physical meaning: the altitude of our agent Needs to be maintained within a safe range , to avoid falling to the ground or losing the range advantage.

[0114] Formula:

[0115]

[0116] Among them, is the set ideal altitude;

[0117] (3) Speed advantage reward

[0118] Physical meaning: the speed of our agent Needs to be maintained within an effective range , to avoid stalling or overloading .

[0119] Formula:

[0120]

[0121] Among them, the reward within the safe range is linearly positively correlated with the speed, encouraging maintaining medium and high speeds to improve mobility;

[0122] (4)Winning and losing reward

[0123] Physical meaning: It represents whether our agent defeats the other party or is defeated by the other party in the interactive confrontation. In this embodiment, if the other party is defeated, the winning and losing reward takes 100, and if defeated by the other party, the winning and losing reward takes -100; in this embodiment, the basis for judging whether to defeat is that when the attacked party is in the inescapable area of the attacker's kinetic body, it is considered that the attacked party is defeated.

[0124] (5)Trajectory prediction reward

[0125]

[0126] Among them, is the position of Agent B n predicted by the LSTM network at the next time step, and is the actual position of Agent B n at the next time step.

[0127] The finally obtained total reward function is:

[0128]

[0129] where , represents the weights of each part.

[0130] In the reward function, the trajectory prediction reward not only guides the LSTM network to train the prediction of the three-dimensional coordinates of the four Agent Bs in the next time step relative to Agent A No. 1 in the hidden state vector , enabling the LSTM network to preferentially retain the recent coordinate data of the enemy by adjusting the weights of the forget gate and input gate, and making the predicted value gradually approach the true relative position; but also jointly optimizes with the angle reward, speed reward, etc., dynamically adjusting the pitch, roll and other action decisions of our agents to achieve coordinated actions and path planning.

[0131] And to improve the training efficiency, a double experience replay mechanism is also introduced in this embodiment. Different sampling strategies are adopted at different stages of training. In the initial stage of training, uniform sampling is used with a random experience pool to ensure that the agents can fully explore the environment. In the middle and late stages of training (for example, after 15000000 steps), a prioritized experience pool is used for priority sampling based on the reward value to accelerate the learning of high-value experiences.

[0132] Step 4: After each agent obtains the maneuver decision action vector through Step 3, using the maneuver decision action vector as the control target value, the agent is controlled according to the discrete incremental PID controller. The discrete incremental PID control is a conventional means in the control field.

[0133] In addition, in this embodiment, a basic action library is also stored in the decision-making system of the agent. A variety of typical maneuver actions of the agent are set in the basic action library, such as target tracking, rapid climb, etc. The agent is not controlled by the reinforcement learning model under certain conditions, but executes typical maneuver actions according to the settings of the basic action library. For example, when the target enters the recognition range of our agent, our agent executes the target tracking action in the basic action library, and when the target leaves the recognition range or our agent is threatened by the kinetic energy body of the other party, the reinforcement learning model is re-adopted for control, so as to avoid problems such as too long intelligent decision sequences, unequal decision sequences, and overly sparse rewards.

[0134] In this embodiment, the four agents on Party A's side are agents that make decisions and controls using the method proposed in the present invention, and the four agents on Party B's side are built-in AI agents of the interactive simulation platform. The hyperparameter settings for the training of the maneuver decision-making network are shown in Table 1.

[0135] Table 1 Hyperparameter settings for the training of the maneuver decision-making network

[0136]

[0137] Finally, the training model with 20,000 Epochs is taken as the test network, and it interacts with the built-in AI of the simulation platforms at all levels. The interaction logic is as Figure 1 shown. The comparison of the test results is shown in Table 2. The winning rate of the present invention is significantly higher than that of the traditional SAC, where the traditional SAC refers to the SAC reinforcement learning model that does not use the LSTM network.

[0138] Table 2 Comparison of test results of the four-agent interactive decision-making network

[0139]

[0140] From Figure 2 , Figure 3 the change in the cumulative reward of each agent per game, it can be seen that the sum of the cumulative reward values of all agents of the present invention based on the SAC reinforcement learning algorithm combined with the LSTM time series processing and the double experience replay mechanism is basically stable above 400 after the 25th round and is basically maintained at about 500. Compared with the traditional SAC, this method can make the cumulative reward reach the ideal value faster and more stably.

[0141] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and purposes of the present invention.

Claims

1. A multi-agent collaborative interaction decision-making and control method in an interaction scenario, characterized in that: It includes the following steps: Step 1: Each agent obtains the temporal state information of the adversarial scenario; Step 2: In the decision-making system of each agent, the obtained adversarial scenario time-series state information is input into the LSTM network for prediction and feature extraction to obtain a hidden state vector , which is used as the input for subsequent reinforcement learning; Step 3: In the decision-making system of each agent, input the hidden state vector obtained in Step 2 into the SAC reinforcement learning model, and train the LSTM network and the SAC reinforcement learning model through the designed reward function, and finally obtain the maneuver decision action vector of each agent; Step 4: Each agent obtains the maneuver decision action vector through Step 3 After that, using the maneuver decision action vector as the control target value, the agent is controlled according to the discrete incremental PID controller.

2. The multi-agent collaborative interaction decision-making and control method in an interaction scenario according to claim 1, wherein: In Step 1, the temporal state information of the adversarial scenario obtained by the agent includes: (1) The three-dimensional coordinates of the agent itself and the other friendly agents relative to the agent itself at the current time step and historical time steps; (2) The three-dimensional coordinates of the opposing agents relative to the agent itself at the current time step and historical time steps; (3) The three-axis velocity components of the agent itself and the other friendly agents relative to the agent itself at the current time step and historical time steps; (4) The three-axis velocity components of the opposing agents relative to the agent itself at the current time step and historical time steps; (5) The attitude angles of the friendly agents at the current time step and historical time steps; (6) The three-dimensional coordinates of the kinetic energy body carried by the opposing agent relative to the agent itself at the current time step and historical time steps; (7) The three-axis velocity components of the kinetic energy body carried by the opposing agent relative to the agent itself at the current time step and historical time steps; (8) The absolute value of the current velocity of the kinetic energy body carried by the opposing agent relative to the agent itself.

3. The multi-agent collaborative interaction decision-making and control method in an interaction scenario according to claim 1, characterized in that: In step 2, the hidden state vector has a lower dimension than the dimension of the temporal state information of the adversarial scenario obtained by the agent, and in the hidden state vector , the predicted value of the three-dimensional coordinates of the other agent relative to the agent itself in the next time step is explicitly included.

4. The multi-agent collaborative interaction decision-making and control method in an interaction scenario according to claim 1 or 3, characterized in that: In Step 3, the SAC reinforcement learning model includes an Actor network and a Critic network; The predicted and dimension-reduced hidden state vector obtained through Step 2 Input it into the Actor network to generate a maneuver decision action vector , and each element in the vector respectively represents the pitch decision target value, roll decision target value, yaw decision target value, and thrust decision target value of the agent in sequence; Input the hidden state vector and the generated maneuver decision action vector into the Critic network, and output the action value .

5. The multi-agent collaborative interaction decision-making and control method in an interaction scenario according to claim 4, characterized in that: In Step 3, during the training process, the designed reward function includes angular advantage reward, altitude advantage reward, speed advantage reward and win-loss reward, and a trajectory prediction reward is also added. Finally, the reward function is the sum of each reward function with random weights added.

6. The multi-agent collaborative interaction decision-making and control method in an interaction scenario according to claim 5, characterized in that: Each reward function is: Angle Advantage Reward : is the line-of-sight angle; Height Advantage Reward : wherein, is the set ideal height; is the height of the agent itself, is the height safety interval; Speed Advantage Reward : Among them, is the speed of the agent itself, is the effective speed range; Winning and Losing Rewards : If you defeat the opponent, the win-loss reward takes a positive value. If you are defeated by the opponent, the win-loss reward takes a negative value; Trajectory Prediction Reward : Among them, is the position of the opponent agent n predicted by the LSTM network at the next time step, is the actual position of the opponent agent n at the next time step; The final total reward function obtained is: Among them , represents the weights of each part.

7. The multi-agent collaborative interaction decision-making and control method in an interaction scenario according to claim 1, wherein: In the training process of Step 3, a double experience replay mechanism is adopted. In the initial stage of training, uniform sampling is performed using a random experience pool to ensure that the agent can fully explore the environment. In the middle and late stages of training, a prioritized experience pool is used for prioritized sampling based on the reward value to accelerate the learning of high-value experiences.

8. The multi-agent collaborative interaction decision-making and control method in an interaction scenario according to claim 1, characterized in that: In the decision-making system of the agent, a basic action library is also stored. A variety of typical maneuver actions of the agent are set in the basic action library. Under certain set conditions, the agent is not controlled by the reinforcement learning model but executes the typical maneuver actions according to the settings of the basic action library. When another set condition is reached, the reinforcement learning model is used for control again.

9. An electronic device, comprising a processor and a memory, the memory being configured to store one or more programs; characterized in that: When the one or more programs are executed by the processor, the method according to any one of claims 1 to 8 is implemented.

10. A readable storage medium stores a computer program, characterized in that: When the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Reinforcement learning decision-making method for heterogeneous multi-agent simulation confrontation environment

    CN115542777A

  • Stand-alone air combat decision-making method based on curriculum type reinforcement learning

    CN116415646A

  • Multi-agent collaborative confrontation decision-making method and device based on reinforcement learning

    CN117273057A

  • Intersection decision-making method based on multi-agent deep reinforcement learning

    CN117496707A

  • Robot agent reinforcement learning training method and system in complex scene

    CN119129642A

Cited By

  • Vehicle lane changing decision control method and device, equipment and storage medium

    CN120863640A

  • Unmanned aerial vehicle flight task allocation method and system based on artificial intelligence

    CN120973073A