Method and device for implementing theory of mind model based on multi-agent reinforcement learning

By establishing and training the target joint mental model network, predicting the intention information of multi-agents and performing explicit learning in the multi-agent reinforcement learning algorithm, the problem of poor synergy effect in the multi-agent scenario is solved, and efficient collaboration of multi-agents is achieved.

CN115081617BActive Publication Date: 2025-09-05BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210635877.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-06
Publication Date
2025-09-05
Estimated Expiration
2042-06-06

AI Technical Summary

Technical Problem

In the multi-agent scenario, in the existing technology, when the number of agents is large, the coordination effect is poor, the network training pressure is high, and it is difficult to achieve effective coordination.

Method used

By establishing the original joint mental model network, predict the intention feature information of the friendly agents of multiple independent agents, and combining the task scenarios of multi-intelligence reinforcement learning, the main target and its sub-objectives are modeled in a hierarchical manner, and the converged main target implementation algorithm and regularized sub-objective implementation algorithm are used to train the target joint mental model network, and the mind theory model is introduced to conduct explicit learning in the multi-intelligence reinforcement learning algorithm.

Benefits of technology

The synergistic effect of multi-agents has been improved, and the final synergistic effect of multi-agents has been achieved through the combination of explicit learning and centralized training frameworks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115081617B_ABST
    Figure CN115081617B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for implementing a theory of mind model based on multi-agent reinforcement learning. The method includes: establishing an original joint mind model network based on a theory of mind model to predict the intention feature information of friendly agents of multiple self-agent agents; establishing a task scenario for multi-agent reinforcement learning in combination with the intention feature information, hierarchically modeling the main goal and its sub-goals of the scenario task; collecting data to be used through the main goal realization algorithm after the convergence of the main goal and the regularized sub-goal realization algorithm of the sub-goals to train the original joint mind model network, predicting the intention information of the current self-agent agent through the target joint mind model network and adding it to the input information of the multi-agent algorithm to achieve the collaboration of the self-agent agent. The method for implementing a theory of mind model based on multi-agent reinforcement learning provided in the embodiment of the present application combines multi-agent reinforcement learning, theory of mind models and task scenarios to improve the collaborative effect of multiple agents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of mental models and multi-agent control, and in particular to a method and device for implementing a theory of mind model based on multi-agent reinforcement learning. Background Art

[0002] Currently, most methods combining reinforcement learning with theory of mind use a single-agent algorithm combined with a theory of mind model. Furthermore, the number of agents in a task scenario is relatively small, and different agents require separate theory of mind modeling. Directly applying this approach to multi-agent scenarios would place significant pressure on network training, leading to poor multi-agent collaboration. Summary of the Invention

[0003] This application provides a method and device for implementing a theory of mind model based on multi-agent reinforcement learning, aiming to improve the collaborative effect of multiple agents.

[0004] In a first aspect, the present application provides a method for implementing a theory of mind model based on multi-agent reinforcement learning, comprising:

[0005] Establishing an original joint mental model network based on the theory of mind model, and predicting the intention feature information of multiple friendly agents of the own agent through the original joint mental model network;

[0006] Establishing a multi-agent reinforcement learning task scenario and combining the intention feature information to hierarchically model the main goal and sub-goals of the scenario task;

[0007] Training the main goal to obtain a converged main goal implementation algorithm, and obtaining a regularized sub-goal implementation algorithm for the sub-goals based on the underlying platform rules;

[0008] Collecting data to be used by the converged main goal realization algorithm and the regularized sub-goal realization algorithm to train the original joint mental model network to obtain a target joint mental model network;

[0009] The target joint mental model network is used to predict the intention information of the current self-agent agent, and the intention information is added to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to achieve the collaboration of the self-agent agent.

[0010] In one embodiment, predicting the intention information of the current self-agent agent through the target joint mental model network and adding the intention information to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to achieve the collaboration of the self-agent agent includes:

[0011] Predicting the intention information of the current self-agent agent through the target joint mind model network, and adding the intention information to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to obtain a multi-agent reinforcement learning algorithm based on mind theory;

[0012] Controlling one's own agent using the theory of mind-based multi-agent reinforcement learning algorithm, controlling the enemy agent using the reinforcement learning algorithm, adjusting the parameters of the scenario task to preset parameters, setting rewards for the own agent and the enemy agent, conducting combat training at a preset number of rounds and a preset round time, and recording changes in the first overall radar coverage index of the own agent and the enemy agent in each round during training;

[0013] The first overall radar coverage index change is compared and verified with the second overall radar coverage index change obtained through multi-agent reinforcement learning algorithm training alone to achieve collaboration between one's own intelligent agents.

[0014] The method of establishing an original joint mental model network based on the theory of mind model includes:

[0015] Determining global observation information of the plurality of friendly agents, wherein the global observation information includes friendly agent information and enemy agent information observable by the friendly agent;

[0016] The theory of mind model is trained using the own agent information of the multiple own agents and the enemy agent information that can be observed by the own agent to obtain the original joint mind model network.

[0017] The predicting of intention feature information of friendly agents of multiple own agents through the original joint mental model network includes:

[0018] Predicting the intention probability distribution of each of the friendly agents through the original joint mental model network to obtain surface intention information of each of the friendly agents;

[0019] Predicting the probability distribution of each friendly agent through the original joint mental model network to obtain the deep intention information of each friendly agent;

[0020] The surface intention information and deep intention information of each of the friendly intelligent agents are determined as the intention feature information of each of the friendly intelligent agents.

[0021] The step of establishing a multi-agent reinforcement learning task scenario and combining the intention feature information to hierarchically model the main goal and sub-goals of the scenario task includes:

[0022] Determine a multi-agent reinforcement learning mission scenario, where the mission scenario layout includes the scenario size, initial position information of the combatants, mission objectives, and final mission evaluation indicators;

[0023] The scene size, the initial position information of the combat parties, the mission objectives and the final mission evaluation indicators are combined with the intention feature information to hierarchically model the main objectives and sub-objectives of the scene mission.

[0024] The training of the main goal to obtain a converged main goal implementation algorithm, and obtaining a regularized sub-goal implementation algorithm of the sub-goal based on the underlying rules of the platform, includes:

[0025] The main goal is trained by a multi-agent reinforcement learning algorithm with the own agent information and the enemy agent information observable by the own agent as input and the coverage target selected by the own agent as output, thereby obtaining the converged main goal realization algorithm;

[0026] Pursue the target selected by the own intelligent agent and obtain the regularized sub-goal implementation algorithm based on the underlying rules of the platform.

[0027] The method of collecting data to be used by the converged main goal realization algorithm and the regularized sub-goal realization algorithm to train the original joint mental model network to obtain a target joint mental model network includes:

[0028] Conducting a multi-target coverage task battle using the converged main goal realization algorithm, the regularized sub-goal realization algorithm, and the enemy strategy realization algorithm, and collecting data to be used required for training the original joint mental model network during operation, wherein the data to be used includes training data, labeled data, and test data;

[0029] The original joint mental model network is supervisedly trained using the training data and the label data, and the intention prediction accuracy of the original joint mental model network is tested using the test data to obtain a target joint mental model network with a preset accuracy.

[0030] In a second aspect, the present application provides a device for implementing a theory of mind model based on multi-agent reinforcement learning, comprising:

[0031] Establishing a prediction module for establishing an original joint mental model network based on the theory of mind model, and predicting the intention feature information of the friendly agents of the plurality of own agents through the original joint mental model network;

[0032] A hierarchical modeling module is used to establish a multi-agent reinforcement learning task scenario and, in combination with the intention feature information, hierarchically model the main goal and sub-goals of the scenario task;

[0033] A training module is used to train the main goal to obtain a converged main goal implementation algorithm, and to obtain a regularized sub-goal implementation algorithm for the sub-goals based on the underlying rules of the platform;

[0034] a collection and training module, configured to collect data to be used by the converged main goal realization algorithm and the regularized sub-goal realization algorithm to train the original joint mental model network, thereby obtaining a target joint mental model network;

[0035] A prediction implementation module is used to predict the intention information of the current self-agent agent through the target joint mental model network, and add the intention information to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to achieve the collaboration of the self-agent agent.

[0036] In a third aspect, the present application also provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the method for implementing a theory of mind model based on multi-agent reinforcement learning as described in the first aspect is implemented.

[0037] In a fourth aspect, the present application also provides a non-transitory computer-readable storage medium, which includes a computer program. When the computer program is executed by the processor, it implements the method for implementing the theory of mind model based on multi-agent reinforcement learning described in the first aspect.

[0038] In a fifth aspect, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by the processor, it implements the method for implementing the theory of mind model based on multi-agent reinforcement learning described in the first aspect.

[0039] The method and apparatus for implementing a theory-of-mind model based on multi-agent reinforcement learning, provided in this application, combine multi-agent reinforcement learning, theory-of-mind models, and task scenarios. This method uses a theory-of-mind model to effectively capture the intentions of one's own agents in collaborative task scenarios, and explicitly learns this information within the multi-agent reinforcement learning algorithm, thereby enhancing the ultimate collaborative effect of the multi-agents. Furthermore, by leveraging the advantages of the centralized training and distributed execution framework within the multi-agent reinforcement learning algorithm, a joint mind model network structure is proposed that integrates well with this framework, achieving complementary advantages and enhancing the ultimate collaborative effect of the multi-agents. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solution of the present application, a brief introduction is given below to the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0041] Figure 1 This is a flowchart of a method for implementing a theory of mind model based on multi-agent reinforcement learning provided by this application;

[0042] Figure 2 This is a schematic diagram of the multi-agent reinforcement learning algorithm based on the theory of mind of this application;

[0043] Figure 3 This is a schematic diagram of the structure of the device for implementing the theory of mind model based on multi-agent reinforcement learning provided by this application;

[0044] Figure 4 It is a structural diagram of the electronic device provided in this application. DETAILED DESCRIPTION

[0045] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0046] Combine Figures 1 to 4 This application describes the method and device for implementing a theory of mind model based on multi-agent reinforcement learning. Figure 1 This is a flowchart of a method for implementing a theory of mind model based on multi-agent reinforcement learning provided by this application; Figure 2 This is a schematic diagram of the multi-agent reinforcement learning algorithm based on the theory of mind of this application; Figure 3 This is a schematic diagram of the structure of the device for implementing the theory of mind model based on multi-agent reinforcement learning provided by this application; Figure 4 It is a structural diagram of the electronic device provided in this application.

[0047] The embodiments of the present application provide an embodiment of a method for implementing a theory of mind model based on multi-agent reinforcement learning. It should be noted that although a logical order is shown in the flowchart, under certain data, the steps shown or described may be completed in an order different from that shown here.

[0048] The embodiments of the present application take electronic devices as examples of execution entities. The electronic devices in the embodiments of the present application include but are not limited to terminals, computers, and devices.

[0049] Reference Figure 1 , Figure 1 This is a flow chart of a method for implementing a theory of mind model based on multi-agent reinforcement learning provided by this application. The method for implementing a theory of mind model based on multi-agent reinforcement learning provided by this embodiment of the application includes:

[0050] Step S10: establishing an original joint mental model network based on the theory of mind model, and predicting the intention feature information of multiple friendly intelligent agents of the own intelligent agents through the original joint mental model network.

[0051] According to the task requirements, the joint mental model network can adopt any simple multi-layer perceptron network structure (Multilayer Perceptron, MLP). The network structure adopted in the embodiment of the present application specifically includes two hidden layers and one output layer, both of which are fully connected layers containing 32 hidden nodes, and the activation function of the two hidden layers adopts the Relu activation function. The theory of mind model (multi-layer perceptron MLP network) is further trained through the global observation of the own agent to obtain the original joint mental model network ToM. Furthermore, the intention feature information of the friendly agents of multiple own agents is predicted through the original joint mental model network, as described in steps S101 to S105.

[0052] Furthermore, steps S101 to S105 are described as follows:

[0053] Step S101, determining global observation information of the plurality of own agents, wherein the global observation information includes own agent information and enemy agent information observable by the own agent;

[0054] Step S102, training the theory of mind model using the self-agent information of the multiple self-agents and the enemy agent information observable by the self-agent to obtain the original joint mind model network;

[0055] Step S103, predicting the intention probability distribution of each of the friendly agents through the original joint mental model network to obtain the surface intention information of each of the friendly agents;

[0056] Step S104, predicting the probability distribution of each friendly agent through the original joint mental model network to obtain the deep intention information of each friendly agent;

[0057] Step S105: Determine the surface intention information and deep intention information of each of the friendly intelligent agents as the intention feature information of each of the friendly intelligent agents.

[0058] Specifically, the input information of the MLP network (theory of mind model) requires the global observation information of the friendly agent, where the global observation information includes the friendly agent information (i.e., the agent information of the friendly agent) and the enemy agent information that the friendly agent can observe. The MLP network is based on the concept of the mental model and needs to predict the intention information of the friendly agent. The predicted intention information will be modeled as the enemy coverage target selected by the friendly drone in the specific scenario. The output of the MLP network is the intention of the current predicted agent. Each time a prediction is made, a piece of friendly agent information and the observable enemy overall agent information are input separately, so as to predict the intention of each friendly agent in turn. The specific formula can be expressed as

[0059]

[0060] in, Indicates the predicted intention of the i-th agent, which is the probability distribution of the intention after the network output features are calculated by a softmax layer (intent i1 ,…,intent iN ), each element represents the probability of the current agent i selecting the jth intention, and the overall sum is 1; s i Represents the basic information of the i-th agent, s e Represents the overall information of the observed enemy agent, MLP θ Represents an MLP network with θ as a parameter. Since all intelligent agents share a network model for intention prediction and use global observations for centralized training, the original joint mental model network ToM is obtained by training the MLP network (theory of mind model).

[0061] Furthermore, the original mental model network is used to predict the intentions of each friendly agent. During the intention prediction process, not only the friendly agent's own probability distribution is predicted, but also the friendly agent's intention probability distribution needs to be predicted. Specifically, the original joint mental model network is used to predict the intention probability distribution of each friendly agent to obtain the surface intention information of each friendly agent. At the same time, the original joint mental model network is used to predict the probability distribution of each friendly agent to obtain the deep intention information of each friendly agent. Furthermore, the surface intention information and deep intention information of each friendly agent are determined as the intention feature information of each friendly agent. The above multi-level intention prediction is the nested belief prediction mechanism of the mental model network.

[0062] The embodiment of the present application introduces a theory of mind model to effectively capture the intention information of one's own intelligent agent in a collaborative task scenario, and explicitly learns it in the multi-agent reinforcement learning algorithm, thereby improving the final collaborative effect of the multi-agent.

[0063] Step S20: Establish a task scenario for multi-agent reinforcement learning and, in combination with the intention feature information, hierarchically model the main goal and sub-goals of the scenario task.

[0064] Furthermore, a multi-agent reinforcement learning mission scenario is established, and its layout is set. This includes the scenario size, initial position information for all parties involved, the mission objective, and final mission evaluation metrics. Furthermore, the scenario size, initial position information for all parties involved, the mission objective, and final mission evaluation metrics are combined with the intentional feature information of each friendly agent output by the original mental model network to hierarchically model the primary objective and sub-goals of the scenario mission, as described in steps S201 and S202.

[0065] Furthermore, the description of step S201 to step S202 is as follows:

[0066] Step S201: Determine a multi-agent reinforcement learning mission scenario, where the mission scenario layout includes the scenario size, initial position information of multiple combat parties, mission objectives, and final mission evaluation indicators;

[0067] Step S202: Combine the scene size, the initial position information of the combat parties, the mission objectives and the final mission evaluation indicators with the intention feature information to hierarchically model the main objectives and sub-objectives of the scene mission.

[0068] Specifically, the scene size, initial position information of multiple combat parties, task objectives and final task evaluation indicators of the task scenario of multi-agent reinforcement learning are determined. In one embodiment, the task scenario used is a multi-target coverage task scenario based on the multi-UAV air combat simulation platform Xsim. The simulated battlefield range (scene size) is set to a 300,000*300,000 meter square battlefield. The combat units (combat parties) are blue UAVs and red UAVs, with a maximum of four on each side. With the center of the battlefield as the origin, the initial position information of the red UAV and the initial position information center of the blue UAV are set to (-60,000, 0) and (40,000, 0) respectively. The distance between the two sides is 100,000 meters and the distance between the UAVs on the same side is 5,000 meters, and the altitude is 9,500 meters. For the UAV parameter setting, the maximum speed is set to 300 meters / second, the detection radar azimuth range is [-30 degrees, 30 degrees], the pitch range is [-10 degrees, 10 degrees], and the detection distance range is 60,000 meters.

[0069] Furthermore, the multi-target coverage scenario is conducted in a round-based system, with each round consisting of 600 time steps. After each round, the scenario automatically resets and enters the next round. The objective of the scenario is that within each round, the Red Team's drones achieve radar coverage of as many Blue Team drones as possible in as many time steps as possible. Therefore, two test metrics are set for this mission: the total target coverage ratio per round and the total target coverage time per round. The formula for calculating the total coverage ratio per round can be expressed as Ratio.

[0070]

[0071] Among them, N e Represents the total number of blue team drones; t i represents the total time that the i-th blue drone is covered in this round; T episode Represents the total number of time steps per round. The target full coverage time per round indicator represents the total duration that the Red UAV achieves full radar coverage of the Blue UAV per round.

[0072] Furthermore, the drone information, drone action and drone reward interface are set, specifically: the multi-target coverage scenario task is based on the POMDP state sequence and is a partially observable task. When modeling the scenario, it is necessary to consider the drone information acquisition method and the content of the information obtained; the drone control method provided by the Xsim platform is mainly instruction control commands, which include but are not limited to initializing entities (make_entityinitinfo), line patrol (make_linepatrolparam), area patrol (make_areapatrolparam), maneuver parameter adjustment (make_motioncmdparam), following (make_followparam) and attacking targets (make_attackparam). When realizing the control of drones on the battlefield, several instructions need to be used flexibly; and when designing the rewards for drones, it is necessary to consider the characteristics of the algorithm used and the mission objectives.

[0073] Furthermore, the environmental situation information (obs) provided by the Xsim scenario includes, but is not limited to, the simulation time step, the Red team situation, and the Blue team situation. The Red and Blue team situation includes, but is not limited to, "platforminfos" weapon platform information (friendly), "trackinfos" intelligence information (enemy), and "missileinfos" missile information. Enemy targets known to both the Red and Blue teams are shared by all friendly aircraft, and battlefield situation information can be acquired at a customizable frequency, with the platform-given situation update frequency set at a maximum of 1 second. Based on the multi-target coverage mission requirements and algorithmic needs, drone position and orientation information from the weapon platform information is selected as friendly information state input. Enemy information acquisition requires radar coverage; otherwise, enemy drone information cannot be obtained, and the acquired information remains the drone's position and orientation.

[0074] Furthermore, the UAV's motion control needs to be combined through six instruction control commands. According to the algorithm training and modeling requirements, the control of the UAV is divided into chasing designated targets and free control; among them, chasing designated targets is achieved by calling the follow (make_followparam) instruction. By entering the pursuit UAV ID, the pursuit control of the designated UAV can be achieved through the platform's underlying rules. The motion control is mainly used for the red team's main target strategy training; the free motion control calls the line patrol (make_linepatrolparam) command, which realizes full-speed forward motion control in the horizontal direction of up, down, left, right, upper left, upper right, lower left, lower right, and vertical direction of ascent and descent freedom according to the current UAV coordinate position. It is mainly used for the blue team's strategy training to avoid the red team's radar coverage.

[0075] Furthermore, the training reward setting for the drone is divided into the reward setting for the red drone and the reward setting for the blue drone, and each drone is calculated separately; the reward setting for the red drone is the basic reward and the radar coverage reward for the blue drone, where the basic reward is a reward loss of -0.3 for each time step, which is used to encourage the drone to explore the action of increasing the reward. The calculation formula for the radar coverage reward for the blue drone can be expressed as (per time step) R cover .

[0076]

[0077] Among them, N e Represents the total number of blue team drones; N i Indicates how many red drones have radar coverage on the i-th blue drone (N i ≥1), C i Indicates whether the current red drone has radar coverage for the current i-th blue drone. If C i=1; otherwise C i =0, so the above design encourages the Red UAV to achieve radar coverage of a Blue UAV that is not covered; the reward setting for the Blue UAV is mainly to encourage it to evade the opponent's radar coverage, that is, when the Blue UAV is covered by the Red UAV radar, it will receive a -0.2 reward for each time step covered, in addition to a basic reward of -0.1 for each time step. red Set to a predefined coverage reward R cover , the blue drone's R blue The reward can be expressed as follows:

[0078]

[0079] This means that the blue team's drone will receive a -0.3 reward when it is covered, and a -0.1 reward otherwise.

[0080] Furthermore, based on the intention feature information output by the original joint mind model network, the primary and sub-goals of the scenario task are hierarchically modeled, introducing the concept of intent information through the primary goal. Using the Xsim platform, a multi-target coverage combat scenario is deployed as the mission scenario, involving the control of both Red and Blue drones. The Red drone is treated as the friendly agent, and the Blue drone as the enemy agent. The process of the Red drone selecting different Blue drones as coverage targets is modeled as the Red drone's primary goal, and the Red drone's pursuit of the current target Blue drone is modeled as a sub-goal based on the primary goal. Thus, the primary goal corresponds to the concept of intent in the theory of mind. The next step in intent recognition is to identify the Blue drone's coverage targets selected by different friendly drones.

[0081] The embodiment of the present application combines the advantages of the centralized training distributed execution framework in the multi-agent reinforcement learning algorithm, and proposes a joint mental model network structure that can be well integrated with the framework, thereby achieving complementary advantages and improving the ultimate synergistic effect of multiple agents.

[0082] Step S30: training the main goal to obtain a converged main goal implementation algorithm, and obtaining a regularized sub-goal implementation algorithm of the sub-goal based on the underlying rules of the platform.

[0083] Furthermore, a multi-agent reinforcement learning algorithm is used, using the self-agent information and observable enemy agent information as input, and the self-agent selected coverage target as output to train the main goal, resulting in a converged main goal achievement algorithm. Simultaneously, the self-agent selected target is pursued, and based on the platform's underlying rules, a regularized sub-goal achievement algorithm for the sub-goal is derived, as described in steps S301 and S302.

[0084] Furthermore, the description of step S301 to step S302 is as follows:

[0085] Step S301: Using a multi-agent reinforcement learning algorithm with the own agent information and the enemy agent information observable by the own agent as input and the coverage target selected by the own agent as output, the main goal is trained to obtain a converged main goal realization algorithm;

[0086] Step S302: pursue the target selected by the own intelligent agent and obtain the regularized sub-goal implementation algorithm based on the underlying rules of the platform.

[0087] Specifically, the main objective training is implemented through the multi-agent reinforcement learning algorithm MAPPO, which takes the basic state information of the own drone and the observable enemy drone information as input, and the blue coverage target selected by the current aircraft as output. The sub-objective task is to pursue the target after the drone selects it, which is directly implemented based on the underlying rules of the Xsim platform. In addition, the enemy drone's mission objective is only to avoid radar coverage, which is mainly achieved through the multi-agent reinforcement learning algorithm IDQN. The training output of the main objective task can be expressed as the following process:

[0088] A i =π θ (s i ,s intent ,s e )

[0089] Among them, A i Indicates the enemy target currently selected for pursuit, π θ represents the action selection strategy trained by multi-agent reinforcement learning, θ represents the network parameters; s i Indicates the basic information of the current UAV. intent Represents the overall intention information of the friendly party obtained through mental model recognition, s e Indicates the overall basic information of the currently observed enemy drone.

[0090] Furthermore, based on the selected main target task, the sub-target implementation of the red team drone calls the Xsim underlying rule command follow (make_followparam) instruction to implement it. The red team drone that executes this command will chase the selected blue team drone to cover the target with the optimal path and maximum speed.

[0091] Furthermore, the multi-agent reinforcement learning algorithm MAPPO algorithm is used to control the main goal of the self-agent, wherein each agent i in the multi-agent reinforcement learning algorithm MAPPO algorithm is based on the local observation o iand a shared strategy (the shared strategy here is for the case where the agents are of the same type. For agents of different types, they can have their own independent actor and critic network) π θ (a i ∣o i ) to generate an action a i To maximize the discount cumulative reward: Learn a centralized value function V based on the global state s φ (s). Among them, the Actor network optimization goal is:

[0092]

[0093] Among them, the advantage function The GAE method is used, S represents the entropy of the strategy, and σ is a hyperparameter that controls the entropy coefficient. The optimization goal of the critic network is:

[0094]

[0095] in, is the discounted reward; B represents the batch_size, and n represents the number of agents; finally, the structural design of both the Actor and Critic networks adopts a neural network structure consisting of two hidden layers and one output layer, with 64 hidden nodes in each layer. The output layer activation function of the Actor network uses Tanh, and the output layer activation function of the Critic network uses Softmax.

[0096] Furthermore, the reinforcement learning algorithm IDQN algorithm is used to control the enemy intelligent agent to avoid radar coverage targets. In the reinforcement learning algorithm IDQN algorithm, each drone is independently deployed with a reinforcement learning algorithm DQN algorithm. The main optimization formula is as follows:

[0097] L i (θ i )=E s [(y i -Q(s,a;θ i )) 2 ]

[0098] Where i is the time period, L i Represents the calculated loss value, θ i Represents the use of the current real-time updated network parameters, that is, iterative network parameters, E s represents the expected value, s and a represent the state and action respectively, so Q(s,a;θ i ) represents the Q estimate calculated using the iterative network, and the y in the formulai The Q target value calculated based on the target network in the previous iteration cycle is different from the network parameters used in the current iteration.

[0099] Furthermore, in terms of input and output, the input of the multi-agent reinforcement learning algorithm MAPPO is the joint observation of one's own agent X = (s i ,s intent ,s e ), which is the joint input feature mentioned above, and the output is a=(a1,…,a N ) represents the action a selected by each agent i i , which is the main target A in the multi-target task scenario i The input of the reinforcement learning algorithm IDQN includes the basic information of the current enemy agent and the observable global information of the own agent. The output is the action of the current enemy agent, which is one of the predefined 10 degrees of freedom discrete actions.

[0100] After setting up the multi-agent reinforcement learning algorithm MAPPO algorithm and the reinforcement learning algorithm IDQN algorithm, a certain number of rounds of multi-target coverage training of the red drone and the blue drone are carried out until the strategies of both algorithms basically converge.

[0101] The embodiment of the present application combines the advantages of the centralized training distributed execution framework in the multi-agent reinforcement learning algorithm, and proposes a joint mental model network structure that can be well integrated with the framework, thereby achieving complementary advantages and improving the ultimate synergistic effect of multiple agents.

[0102] Step S40 , collecting data to be used by the converged main goal realization algorithm and the regularized sub-goal realization algorithm to train the original joint mental model network to obtain a target joint mental model network.

[0103] Furthermore, a multi-target coverage task is conducted using the converged MAPPO algorithm (the converged main goal achievement algorithm), the regularized sub-goal achievement algorithm, and the enemy strategy implementation algorithm IDQN. At this point, the algorithm strategy is frozen, and during operation, the data required for training the original joint mental model network (to-be-used data) is collected. Furthermore, the original joint mental model network is trained using the to-be-used data to obtain the target joint mental model network, as described in steps S401 and S402.

[0104] Furthermore, the description of step S401 to step S402 is as follows:

[0105] Step S401: Conduct a multi-target coverage task battle using the converged main goal implementation algorithm, the regularized sub-goal implementation algorithm, and the enemy strategy implementation algorithm, and collect data to be used required for training the original joint mental model network during the operation, wherein the data to be used includes training data, labeled data, and test data;

[0106] Step S402 : Supervise the original joint mental model network using the training data and the label data, and test the intent prediction accuracy of the original joint mental model network using the test data to obtain a target joint mental model network with a preset accuracy.

[0107] Specifically, the converged main goal implementation algorithm MAPPO, the regularized sub-goal implementation algorithm and the enemy strategy implementation algorithm IDQN are used to carry out multi-target coverage task battles. At this time, the algorithm strategy is frozen, and the data required for the original joint mental model network training is collected during operation. The data to be used includes training data, label data and test data.

[0108] Furthermore, the training data includes the overall orientation and position information of the red drone in each round, as well as the overall orientation and position information of the blue drone observed at the corresponding moment. About 100,000 pieces of training data are collected each time and stored as JSON format files; the label data is the blue attack target selected by the red drone at the corresponding moment of the training data, expressed in one-hot encoding format, and the number is consistent with the training data; then about 2,000 test data are collected in the same way.

[0109] Furthermore, the original joint mental model network ToM is supervised and trained using the above-collected training data, label data, and test data. At the same time, the intention prediction accuracy of the original joint mental model network ToM is tested using the above-collected test data to obtain a target joint mental model network ToM with a preset accuracy. The loss function formula of the original joint mental model network ToM is as follows:

[0110]

[0111] Among them, C represents the total number of enemy agents, y i Indicates label data, g i It represents the probability predicted by the network that the current agent selects the enemy's i-th intention.

[0112] Furthermore, when training and testing the original joint mental model network (ToM), information about each agent in the collected data is randomly input. The original joint mental model network (ToM) is trained using training data, labeled data, and test data. The network's prediction accuracy is then compared on the test set. When the prediction accuracy reaches 95% or higher, the original joint mental model network (ToM) meets the training requirements.

[0113] The embodiment of the present application combines the advantages of the centralized training distributed execution framework in the multi-agent reinforcement learning algorithm, and proposes a joint mental model network structure that can be well integrated with the framework, thereby achieving complementary advantages and improving the ultimate synergistic effect of multiple agents.

[0114] Step S50: predict the intention information of the current own agent through the target joint mental model network, and add the intention information to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to achieve the collaboration of the own agent.

[0115] Furthermore, the target joint mind model network predicts the intention information of the current self-agent agent. During the training process of the multi-agent algorithm, this intention information is added to the input information of the multi-agent algorithm, thereby obtaining a multi-agent reinforcement learning algorithm based on theory of mind. Furthermore, the self-agent agent is controlled by the multi-agent reinforcement learning algorithm based on theory of mind, and the enemy agent is controlled by the reinforcement learning algorithm. The parameters of the scenario task are adjusted to preset parameters, and the rewards for the self-agent and enemy agents are set. Combat training is conducted for a preset number of rounds and a preset round time. Changes in the first overall radar coverage index of the self-agent and enemy agents are recorded during each round of training. Furthermore, the changes in the first overall radar coverage index are compared and verified with changes in the second overall radar coverage index obtained through training with the multi-agent reinforcement learning algorithm alone, thereby achieving coordination among the self-agent agents, as described in steps S501 to S503.

[0116] Furthermore, steps S501 to S503 are described as follows:

[0117] Step S501: Predicting the intention information of the current self-agent agent through the target joint mind model network, and adding the intention information to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to obtain a multi-agent reinforcement learning algorithm based on mind theory;

[0118] Step S502: Control the own agent using the theory of mind-based multi-agent reinforcement learning algorithm, control the enemy agent using the reinforcement learning algorithm, adjust the parameters of the scenario task to preset parameters, set rewards for the own agent and the enemy agent, conduct combat training for a preset number of rounds and a preset round time, and record changes in the first overall radar coverage index of the own agent and the enemy agent in each round during the training period;

[0119] Step S503 : Compare and verify the change in the first overall radar coverage index with the change in the second overall radar coverage index obtained through training of the multi-agent reinforcement learning algorithm alone to achieve collaboration between the own intelligent agents.

[0120] Specifically, the target joint mental model network TOM is combined with the multi-agent algorithm MAPPO. During the training process of the multi-agent algorithm, the target joint mental model network is used to predict the current intention information of the own agent and add it to the input information of the multi-agent algorithm to obtain the multi-agent reinforcement learning algorithm ToM-MAPPO based on the theory of mind. Figure 2 , Figure 2 Schematic diagram of the multi-agent reinforcement learning algorithm based on theory of mind in this application.

[0121] Further, in combination with the present application Figure 2 The ToM-MAPPO multi-agent reinforcement learning algorithm based on the theory of mind is analyzed as follows: The combination of the target joint mental model network and the multi-agent reinforcement learning algorithm framework adopts a docking additive combination on the input and output, and puts the target joint mental model network ToM into the centralized training part of the centralized training distributed execution. The global observation is reused on the input, and the current intention of each drone of the red team, that is, the main target selection, is predicted in turn, and the predicted intention probability vector is used. This is spliced ​​into the input of the centrally trained critic network to achieve explicit modeling and learning of friendly intention information. At this point, the modeling completes the new theory-of-mind-based multi-agent reinforcement learning algorithm ToM-MAPPO. The updated algorithm's joint critic network training objective function is as follows (highlighting the differences and being compatible with other multi-agent algorithms):

[0122]

[0123] in, In the above parameters, φ represents the network parameters, is the discounted reward; B represents the batch_size, n represents the number of agents; x represents the joint observation, x ToM Represents the joint intention probability vector obtained by concatenating the intention vectors predicted by the target joint mental model network TOM, satisfying xToM =(I i1 ,…,I ij ;I ji ), where I i1 ,…,I ij Represents the current agent's prediction of the probability distribution vector of other agents' intentions, and I ji It represents the prediction of the intention probability distribution vector of the current agent by other agents. Due to the use of the target joint mental model network TOM, the prediction results are consistent. This method not only predicts the probability distribution of the friendly intention, that is, the surface intention, but also learns and predicts the friendly prediction of its own probability distribution, that is, the deep intention. The method of combining the multi-level predicted intentions with the input is the nested belief prediction mechanism of the mental model network. Represents the current joint action; the update of the Actor network is consistent with the original MAPPO algorithm;

[0124] Furthermore, after adjusting the parameters of the mission scenario, the ToM-MAPPO algorithm, a multi-agent reinforcement learning algorithm based on the theory of mind, is used to control the red drone, and the IDQN algorithm, an enemy strategy implementation algorithm, is used to control the blue drone. After setting the rewards for both sides, combat training is carried out under a preset number of rounds and a preset round time. The changes in the first overall radar coverage rate indicator of the red drone over the blue drone are recorded in each round during the training as the final evaluation indicator. Among them, the preset number of rounds and the preset round time are set according to actual conditions.

[0125] Furthermore, the change in the second overall radar coverage index obtained by using the multi-agent reinforcement learning algorithm MAPPO algorithm for combat training alone is determined.

[0126] Furthermore, the change in the first overall radar coverage index is compared with the change in the second overall radar coverage index obtained through training of the multi-agent reinforcement learning algorithm alone, verifying that the multi-agent reinforcement learning algorithm ToM-MAPPO algorithm based on the theory of mind can improve the collaborative effect of multiple agents and realize the collaboration of one's own agents.

[0127] Embodiments of the present application

[0128] The embodiment of this application provides a method for implementing a theory-of-mind model based on multi-agent reinforcement learning. This method combines multi-agent reinforcement learning, a theory-of-mind model, and task scenarios. The theory-of-mind model is introduced to effectively capture the intention information of one's own agent in collaborative task scenarios, and is explicitly learned within the multi-agent reinforcement learning algorithm, thereby improving the ultimate collaborative effect of the multi-agents. Furthermore, by combining the advantages of the centralized training and distributed execution framework of the multi-agent reinforcement learning algorithm, a joint mind model network structure is proposed that can be well integrated with the framework, thereby achieving complementary advantages and improving the ultimate collaborative effect of the multi-agents.

[0129] Furthermore, the present application describes a device for implementing a theory of mind model based on multi-agent reinforcement learning, and the device for implementing a theory of mind model based on multi-agent reinforcement learning and a method for implementing a theory of mind model based on multi-agent reinforcement learning correspond to each other.

[0130] like Figure 3 As shown, Figure 3 This is a schematic diagram of the structure of a device for implementing a theory of mind model based on multi-agent reinforcement learning provided by this application. The device for implementing a theory of mind model based on multi-agent reinforcement learning includes:

[0131] Establishing a prediction module 301 for establishing an original joint mental model network based on the theory of mind model, and predicting the intention feature information of the friendly agents of multiple self-agents through the original joint mental model network;

[0132] A hierarchical modeling module 302 is configured to establish a multi-agent reinforcement learning task scenario and, in combination with the intention feature information, hierarchically model the main goal and sub-goals of the scenario task;

[0133] The training module 303 is used to train the main goal to obtain a converged main goal implementation algorithm, and to obtain a regularized sub-goal implementation algorithm for the sub-goals based on the underlying rules of the platform;

[0134] A collection and training module 304 is configured to collect data to be used by the converged main goal realization algorithm and the regularized sub-goal realization algorithm to train the original joint mental model network to obtain a target joint mental model network;

[0135] The prediction implementation module 305 is used to predict the intention information of the current self-agent agent through the target joint mental model network, and add the intention information to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to achieve the collaboration of the self-agent agent.

[0136] Furthermore, the prediction implementation module 305 is further configured to:

[0137] Predicting the intention information of the current self-agent agent through the target joint mind model network, and adding the intention information to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to obtain a multi-agent reinforcement learning algorithm based on mind theory;

[0138] Controlling one's own agent using the theory of mind-based multi-agent reinforcement learning algorithm, controlling the enemy agent using the reinforcement learning algorithm, adjusting the parameters of the scenario task to preset parameters, setting rewards for the own agent and the enemy agent, conducting combat training at a preset number of rounds and a preset round time, and recording changes in the first overall radar coverage index of the own agent and the enemy agent in each round during training;

[0139] The first overall radar coverage index change is compared and verified with the second overall radar coverage index change obtained through multi-agent reinforcement learning algorithm training alone to achieve collaboration between one's own intelligent agents.

[0140] Furthermore, the prediction module 301 is also used to:

[0141] Determining global observation information of the plurality of friendly agents, wherein the global observation information includes friendly agent information and enemy agent information observable by the friendly agent;

[0142] The theory of mind model is trained using the own agent information of the multiple own agents and the enemy agent information that can be observed by the own agent to obtain the original joint mind model network.

[0143] Furthermore, the prediction module 301 is also used to:

[0144] Predicting the intention probability distribution of each of the friendly agents through the original joint mental model network to obtain surface intention information of each of the friendly agents;

[0145] Predicting the probability distribution of each friendly agent through the original joint mental model network to obtain the deep intention information of each friendly agent;

[0146] The surface intention information and deep intention information of each of the friendly intelligent agents are determined as the intention feature information of each of the friendly intelligent agents.

[0147] Furthermore, the hierarchical modeling module 302 is further configured to:

[0148] Determine a multi-agent reinforcement learning mission scenario, where the mission scenario layout includes the scenario size, initial position information of the combatants, mission objectives, and final mission evaluation indicators;

[0149] The scene size, the initial position information of the combat parties, the mission objectives and the final mission evaluation indicators are combined with the intention feature information to hierarchically model the main objectives and sub-objectives of the scene mission.

[0150] Furthermore, the training module 303 is further configured to:

[0151] The main goal is trained by a multi-agent reinforcement learning algorithm with the own agent information and the enemy agent information observable by the own agent as input and the coverage target selected by the own agent as output, thereby obtaining the converged main goal realization algorithm;

[0152] Pursue the target selected by the own intelligent agent and obtain the regularized sub-goal implementation algorithm based on the underlying rules of the platform.

[0153] Furthermore, the collection and training module 304 is further configured to:

[0154] Conducting a multi-target coverage task battle using the converged main goal realization algorithm, the regularized sub-goal realization algorithm, and the enemy strategy realization algorithm, and collecting data to be used required for training the original joint mental model network during operation, wherein the data to be used includes training data, labeled data, and test data;

[0155] The original joint mental model network is supervisedly trained using the training data and the label data, and the intention prediction accuracy of the original joint mental model network is tested using the test data to obtain a target joint mental model network with a preset accuracy.

[0156] The specific embodiments of the apparatus for implementing a theory of mind model based on multi-agent reinforcement learning provided in this application are substantially the same as the embodiments of the method for implementing a theory of mind model based on multi-agent reinforcement learning described above, and are not described in detail here.

[0157] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute a method for implementing a theory of mind model based on multi-agent reinforcement learning, which includes:

[0158] Establishing an original joint mental model network based on the theory of mind model, and predicting the intention feature information of multiple friendly agents of the own agent through the original joint mental model network;

[0159] Establishing a multi-agent reinforcement learning task scenario and combining the intention feature information to hierarchically model the main goal and sub-goals of the scenario task;

[0160] Training the main goal to obtain a converged main goal implementation algorithm, and obtaining a regularized sub-goal implementation algorithm for the sub-goals based on the underlying platform rules;

[0161] Collecting data to be used by the converged main goal realization algorithm and the regularized sub-goal realization algorithm to train the original joint mental model network to obtain a target joint mental model network;

[0162] The target joint mental model network is used to predict the intention information of the current self-agent agent, and the intention information is added to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to achieve the collaboration of the self-agent agent.

[0163] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0164] On the other hand, the present application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can perform the method for implementing a theory of mind model based on multi-agent reinforcement learning provided by the above methods, which includes:

[0165] Establishing an original joint mental model network based on the theory of mind model, and predicting the intention feature information of multiple friendly agents of the own agent through the original joint mental model network;

[0166] Establishing a multi-agent reinforcement learning task scenario and combining the intention feature information to hierarchically model the main goal and sub-goals of the scenario task;

[0167] Training the main goal to obtain a converged main goal implementation algorithm, and obtaining a regularized sub-goal implementation algorithm for the sub-goals based on the underlying platform rules;

[0168] Collecting data to be used by the converged main goal realization algorithm and the regularized sub-goal realization algorithm to train the original joint mental model network to obtain a target joint mental model network;

[0169] The target joint mental model network is used to predict the intention information of the current self-agent agent, and the intention information is added to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to achieve the collaboration of the self-agent agent.

[0170] In another aspect, the present application further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for implementing a theory of mind model based on multi-agent reinforcement learning provided above is implemented, the method comprising:

[0171] Establishing an original joint mental model network based on the theory of mind model, and predicting the intention feature information of multiple friendly agents of the own agent through the original joint mental model network;

[0172] Establishing a multi-agent reinforcement learning task scenario and combining the intention feature information to hierarchically model the main goal and sub-goals of the scenario task;

[0173] Training the main goal to obtain a converged main goal implementation algorithm, and obtaining a regularized sub-goal implementation algorithm for the sub-goals based on the underlying platform rules;

[0174] Collecting data to be used by the converged main goal realization algorithm and the regularized sub-goal realization algorithm to train the original joint mental model network to obtain a target joint mental model network;

[0175] The target joint mental model network is used to predict the intention information of the current self-agent agent, and the intention information is added to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to achieve the collaboration of the self-agent agent.

[0176] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0177] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0178] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for implementing a theory of mind model based on multi-agent reinforcement learning, characterized in that: include: Establishing an original joint mental model network based on the theory of mind model, and predicting the intention feature information of multiple friendly agents of the own agent through the original joint mental model network; Establishing a multi-agent reinforcement learning task scenario and combining the intention feature information to hierarchically model the main goal and sub-goals of the task scenario; Training the main goal to obtain a converged main goal implementation algorithm, and obtaining a regularized sub-goal implementation algorithm for the sub-goals based on the underlying platform rules; Collecting data to be used by the converged main goal realization algorithm and the regularized sub-goal realization algorithm to train the original joint mental model network to obtain a target joint mental model network; Predicting the intention information of the current self-agent agent through the target joint mental model network, and adding the intention information to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to achieve collaboration of the self-agent agent; Among them, the intelligent agent is a drone, and the mission scenario involved is a multi-target coverage mission scenario based on the multi-drone air combat simulation platform Xsim; The method of establishing an original joint mental model network based on the theory of mind model includes: Determining global observation information of the plurality of friendly agents, wherein the global observation information includes friendly agent information and enemy agent information observable by the friendly agent; Training the theory of mind model using the own agent information of the multiple own agents and the enemy agent information observable by the own agent to obtain the original joint mind model network; The theory of mind model is used to predict the intention of each friendly agent. The specific formula is: in, Indicates the predicted intention of the i-th agent, which is the probability distribution of the intention after the network output features are calculated by a softmax layer (intent i1 ,…,intent iN ), each element represents the probability of the current agent i selecting the jth intention, and the overall sum is 1; s i Represents the basic information of the i-th agent, s e Represents the overall information of the observed enemy agent, MLP θ represents an MLP network with θ as a parameter; The predicting of intention feature information of friendly agents of multiple own agents through the original joint mental model network includes: Predicting the intention probability distribution of each of the friendly agents through the original joint mental model network to obtain surface intention information of each of the friendly agents; Predicting the probability distribution of each friendly agent through the original joint mental model network to obtain the deep intention information of each friendly agent; Determining the surface intention information and deep intention information of each of the friendly intelligent agents as intention feature information of each of the friendly intelligent agents; The step of establishing a multi-agent reinforcement learning task scenario and combining the intention feature information to hierarchically model the main goal and sub-goals of the task scenario includes: Determine a multi-agent reinforcement learning mission scenario, where the mission scenario layout includes the scenario size, initial position information of the combatants, mission objectives, and final mission evaluation indicators; Combining the scene size, the initial position information of the combat parties, the mission objectives, and the final mission evaluation index with the intention feature information to hierarchically model the main objectives and sub-objectives of the mission scene; The hierarchical modeling of the main goal and sub-goals of the task scenario includes: The process of the friendly agent selecting different enemy agents as coverage targets is modeled as the main goal of the friendly agent; Model the friendly agent chasing the current target enemy agent as a sub-goal of the main goal; The training of the main goal to obtain a converged main goal implementation algorithm, and obtaining a regularized sub-goal implementation algorithm of the sub-goal based on the underlying rules of the platform, includes: The main goal is trained by a multi-agent reinforcement learning algorithm with the own agent information and the enemy agent information observable by the own agent as input and the coverage target selected by the own agent as output, thereby obtaining the converged main goal realization algorithm; Pursue the target selected by the own intelligent agent and obtain the regularized sub-goal implementation algorithm based on the underlying rules of the platform.

2. The method for implementing a theory of mind model based on multi-agent reinforcement learning according to claim 1, characterized in that: The method of predicting the intention information of the current self-agent agent through the target joint mental model network and adding the intention information to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to achieve the collaboration of the self-agent agent includes: Predicting the intention information of the current self-agent agent through the target joint mind model network, and adding the intention information to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to obtain a multi-agent reinforcement learning algorithm based on mind theory; Controlling the own agent using the theory of mind-based multi-agent reinforcement learning algorithm, controlling the enemy agent using the reinforcement learning algorithm, adjusting the parameters of the mission scenario to preset parameters, setting rewards for the own agent and the enemy agent, conducting combat training at a preset number of rounds and a preset round time, and recording changes in the first overall radar coverage index of the own agent and the enemy agent in each round during the training period; The first overall radar coverage index change is compared and verified with the second overall radar coverage index change obtained through multi-agent reinforcement learning algorithm training alone to achieve collaboration between one's own intelligent agents.

3. The method for implementing a theory of mind model based on multi-agent reinforcement learning according to claim 1, characterized in that: The method of collecting data to be used by the converged main goal realization algorithm and the regularized sub-goal realization algorithm to train the original joint mental model network to obtain a target joint mental model network includes: Conducting a multi-target coverage task battle using the converged main goal realization algorithm, the regularized sub-goal realization algorithm, and the enemy strategy realization algorithm, and collecting data to be used required for training the original joint mental model network during operation, wherein the data to be used includes training data, labeled data, and test data; The original joint mental model network is supervisedly trained using the training data and the label data, and the intention prediction accuracy of the original joint mental model network is tested using the test data to obtain a target joint mental model network with a preset accuracy.

4. A device for implementing a theory of mind model based on multi-agent reinforcement learning, characterized in that: include: Establishing a prediction module for establishing an original joint mental model network based on the theory of mind model, and predicting the intention feature information of the friendly agents of the plurality of own agents through the original joint mental model network; A hierarchical modeling module is used to establish a multi-agent reinforcement learning task scenario and, in combination with the intention feature information, hierarchically model the main goal and sub-goals of the task scenario; A training module is used to train the main goal to obtain a converged main goal implementation algorithm, and to obtain a regularized sub-goal implementation algorithm for the sub-goals based on the underlying rules of the platform; a collection and training module, configured to collect data to be used by the converged main goal realization algorithm and the regularized sub-goal realization algorithm to train the original joint mental model network, thereby obtaining a target joint mental model network; A prediction implementation module, configured to predict the intention information of the current self-agent agent through the target joint mental model network, and add the intention information to the input information of the multi-agent algorithm during the training process of the multi-agent algorithm to achieve collaboration of the self-agent agent; Among them, the intelligent agent is a drone, and the mission scenario involved is a multi-target coverage mission scenario based on the multi-drone air combat simulation platform Xsim; The method of establishing an original joint mental model network based on the theory of mind model includes: Determining global observation information of the plurality of friendly agents, wherein the global observation information includes friendly agent information and enemy agent information observable by the friendly agent; Training the theory of mind model using the own agent information of the multiple own agents and the enemy agent information observable by the own agent to obtain the original joint mind model network; The theory of mind model is used to predict the intention of each friendly agent. The specific formula is: in, Indicates the predicted intention of the i-th agent, which is the probability distribution of the intention after the network output features are calculated by a softmax layer (intent i1 ,…,intent iN ), each element represents the probability of the current agent i selecting the jth intention, and the overall sum is 1; s i Represents the basic information of the i-th agent, s e Represents the overall information of the observed enemy agent, MLP θ represents an MLP network with θ as a parameter; The predicting of intention feature information of friendly agents of multiple own agents through the original joint mental model network includes: Predicting the intention probability distribution of each of the friendly agents through the original joint mental model network to obtain surface intention information of each of the friendly agents; Predicting the probability distribution of each friendly agent through the original joint mental model network to obtain the deep intention information of each friendly agent; Determining the surface intention information and deep intention information of each of the friendly intelligent agents as intention feature information of each of the friendly intelligent agents; The step of establishing a multi-agent reinforcement learning task scenario and combining the intention feature information to hierarchically model the main goal and sub-goals of the task scenario includes: Determine a multi-agent reinforcement learning mission scenario, where the mission scenario layout includes the scenario size, initial position information of the combatants, mission objectives, and final mission evaluation indicators; Combining the scene size, the initial position information of the combat parties, the mission objectives, and the final mission evaluation index with the intention feature information to hierarchically model the main objectives and sub-objectives of the mission scene; The hierarchical modeling of the main goal and sub-goals of the task scenario includes: The process of the friendly agent selecting different enemy agents as coverage targets is modeled as the main goal of the friendly agent; Model the friendly agent chasing the current target enemy agent as a sub-goal of the main goal; The training of the main goal to obtain a converged main goal implementation algorithm, and obtaining a regularized sub-goal implementation algorithm of the sub-goal based on the underlying rules of the platform, includes: The main goal is trained by a multi-agent reinforcement learning algorithm with the own agent information and the enemy agent information observable by the own agent as input and the coverage target selected by the own agent as output, thereby obtaining the converged main goal realization algorithm; Pursue the target selected by the own intelligent agent and obtain the regularized sub-goal implementation algorithm based on the underlying rules of the platform.

5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the method for implementing a theory of mind model based on multi-agent reinforcement learning as described in any one of claims 1 to 3.

6. A non-transitory computer-readable storage medium comprising a computer program, characterized in that: When the computer program is executed by a processor, the method for implementing a theory of mind model based on multi-agent reinforcement learning according to any one of claims 1 to 3 is implemented.