Multi-agent pursuit decision-making method and system based on hybrid imitation learning

By combining two imitation learning methods, MT-GAIL and TD-BC, we process multimodal and singlemodal expert trajectory data, and generate a hybrid pursuit decision model, solving the shortcomings of multiagent pursuit decision-making methods in the existing technology in complex environments, and improving the training efficiency and generalization ability of the model.

CN120046646APending Publication Date: 2025-05-27CHINA SHIPBUILDING ZHIHAI INNOVATION RES INST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411948286.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing multi-agent pursuit decision-making methods are difficult to effectively respond in complex environments, especially the multi-modal expert trajectory data is difficult to accurately distinguish, resulting in insufficient generalization capabilities of the model.

Method used

The hybrid imitation learning method is adopted to generate two imitation learning methods: Adversarial imitation learning (MT-GAIL) and timing differential error behavior cloning (TD-BC) through multi-expert trajectory, and the expert trajectory data of multimodal and singlemodal are processed respectively to generate a hybrid hunting decision model.

Benefits of technology

It significantly improves the training efficiency and adaptability of the model, can better process multimodal data in complex environments, and improves the generalization ability and application flexibility of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046646A_ABST
    Figure CN120046646A_ABST
Patent Text Reader

Abstract

The invention provides a multi-agent pursuit decision-making method and system based on hybrid imitation learning, and the method comprises the steps: training expert track data through employing a multi-expert track generation adversarial imitation learning method when the type of the expert track data is multi-modal, so as to obtain a first decision-making model; when the type of the expert trajectory data is a single mode, training the expert trajectory data by adopting a time sequence difference error behavior cloning method to obtain a second decision model; the first decision-making model and the second decision-making model are endowed with intelligent agents, the first decision-making model and the second decision-making model are deduced through the intelligent agents, and a mixed pursuit decision-making model is obtained; the agent performs decision processing on the pursuit scene containing the moving and static target through the pursuit decision model to obtain a corresponding pursuit strategy; according to the method, time sequence difference error behavior cloning and multi-expert track generation adversarial imitation learning are effectively combined, so that the decision and cooperation capability of a multi-agent system in a complex and dynamic environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multi-agent systems, and particularly relates to a multi-agent pursuit decision-making method and system based on hybrid imitation learning. Background Art

[0002] Currently, for the multi-agent pursuit problem, traditional non-learning methods are difficult to effectively meet the real-time decision-making requirements in complex environments. In contrast, reinforcement learning methods have gradually become the mainstream methods for solving such problems due to their powerful high-dimensional information perception, understanding, and non-linear processing capabilities. However, pure reinforcement learning methods have defects such as long training time and slow convergence speed. To address this problem, most current research adopts imitation learning to quickly learn behavioral strategies from expert experience, thereby significantly improving the model training efficiency. Existing imitation learning methods are mostly based on single-modal expert data, but such expert trajectories do not have diversity, and the trained models are difficult to handle relatively complex scenarios. Therefore, most scholars at the present stage have started to study multi-modal imitation learning. The related research on multi-modal imitation learning can be divided into two categories. One category is that the multi-modal expert trajectory information does not carry labels, and the modal label information is distinguished through unsupervised learning. For example, in 2017, the DeepMind team introduced variational autoencoders to judge modal labels. However, using unsupervised learning methods for such problems sometimes cannot accurately distinguish the expert trajectory data of each modality, resulting in difficulty in learning a multi-modal strategy with good performance. The other category is that the research content of multi-modal imitation learning is to learn from multi-modal expert trajectories with modal labels. For example, the DeepMind team input modal labels into the generator and discriminator based on the GAIL method and used modal labels to guide the training. Lin et al. introduced an auxiliary classifier on the basis of the GAIL framework to classify sample data according to the information of the modality to which they belong. Raunak et al. studied multi-agent imitation learning methods and introduced imitation learning methods into the multi-agent field. How to mine effective information from multi-modal expert trajectories and improve the generalization ability of the model is the focus and difficulty of current research in this field.

[0003] Therefore, how to provide a multi-agent pursuit decision-making method and system based on hybrid imitation learning has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-agent pursuit decision-making method and system based on hybrid imitation learning.

[0005] According to the first aspect of the present invention, a multi-agent pursuit decision-making method based on hybrid imitation learning is provided. The method includes,

[0006] Step S1: Select a pursuit scenario for imitation learning and determine the expert model of the pursuit scenario;

[0007] Step S2: Generate expert trajectory data by interacting the expert model of the pursuit scenario with the environment, and determine the type of the expert trajectory data;

[0008] Step S3: When the type of the expert trajectory data is multimodal, train the expert trajectory data by using the multi-expert trajectory generation adversarial imitation learning method to obtain a first decision-making model;

[0009] Step S4: When the type of the expert trajectory data is unimodal, train the expert trajectory data by using the temporal difference error behavior cloning method to obtain a second decision-making model;

[0010] Step S5: Endow the first decision-making model and the second decision-making model to an agent, and deduce the first decision-making model and the second decision-making model through the agent to obtain a hybrid pursuit decision-making model;

[0011] Step S6: The agent makes decision processing on the pursuit scenario containing static and dynamic targets through the pursuit decision-making model to obtain corresponding pursuit strategies.

[0012] Optionally, in the step S1, determine the expert model of the pursuit scenario by using a decision tree or a reinforcement learning algorithm.

[0013] Optionally, in the step S2, when pursuing a static target with a fixed position in the pursuit scenario, determine that the type of the expert trajectory data is unimodal;

[0014] When pursuing a dynamic target with a constantly changing position in the pursuit scenario, determine that the type of the expert trajectory data is multimodal.

[0015] Optionally, in the step S3, the process of training the expert trajectory data by using the multi-expert trajectory generation adversarial imitation learning method is as follows:

[0016] Step S31: Input each expert trajectory in the expert trajectory data and the generated trajectory of the policy generator into a discriminator at the same time to train the discriminator network, and update the policy network of the policy generator according to the binary cross-entropy loss of the discriminator;

[0017] Step S32: Calculate the reward values of each expert trajectory and each generated trajectory by using the accuracy of the discriminator, store all the calculated reward values in an experience pool, and update the parameters of the policy network of the policy generator by using a policy gradient optimization method so that the generated trajectory of the policy generator gradually approaches the expert trajectory.

[0018] Optionally, before the step S31, it further includes:

[0019] Step S30: Use the discriminator to discriminate each expert trajectory to obtain the reliability coefficient of each expert trajectory, so as to determine the quality of each expert trajectory.

[0020] Optionally, in step S32, the expert accuracy E_acc and the student accuracy L_acc are respectively expressed as:

[0021]

[0022] where D w (s,a) is the output of the discriminator, s represents the state, and a represents the action; T i represents the trajectory data generated by the policy generator, T E represents the trajectory data generated by expert experience, represents the expectation that the discriminator output is greater than 0.5 under the condition of using the expert trajectory data as the input, represents the expectation that the discriminator output is less than 0.5 under the condition of using the student trajectory data as the input;

[0023] The reward value is defined as:

[0024]

[0025] where E(j)_w is the weight of the j-th discriminator.

[0026] Optionally, in step S4, the process of training the expert trajectory data by using the temporal difference error behavior cloning method is as follows:

[0027] Step S41: Define that the student model includes a policy network and a value network. Use the gradient descent strategy to update the parameters of the policy network of the student model, and introduce the temporal difference error TD error to guide the parameter update of the value network of the student model. Among them, the update target of the gradient descent strategy is to minimize the state-action matching error in the expert trajectory data;

[0028] Step S42: Continuously repeat step S42 for iterative update to optimize the parameters of the policy network and the value network of the student model, so that the policy of the student model gradually approaches the policy of the expert model.

[0029] According to the second aspect of the present invention, there is provided a multi-agent pursuit decision-making system based on hybrid imitation learning, and the system includes:

[0030] The first processing module is configured to select a pursuit scenario for imitation learning and determine the expert model of the pursuit scenario;

[0031] A second processing module, configured to generate expert trajectory data by interacting with the environment through the expert model of the pursuit scenario and determine the type of the patent trajectory data;

[0032] A third processing module, configured to, when the type of the patent trajectory data is multimodal, train the expert trajectory data by using a multi-expert trajectory generative adversarial imitation learning method to obtain a first decision-making model;

[0033] A fourth processing module, configured to, when the type of the expert trajectory data is unimodal, train the expert trajectory data by using a temporal difference error behavior cloning method to obtain a second decision-making model;

[0034] A fifth processing module, configured to endow the first decision-making model and the second decision-making model to an intelligent agent, and deduce the first decision-making model and the second decision-making model through the intelligent agent to obtain a hybrid pursuit decision-making model;

[0035] A sixth processing module, configured to the intelligent agent makes decision processing on a pursuit scenario containing moving and static targets through the pursuit decision-making model to obtain a corresponding pursuit strategy.

[0036] According to a third aspect of the present invention, there is provided an electronic device, where the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps in any one of the first aspects of the present invention in a multi-agent pursuit decision-making method based on hybrid imitation learning are implemented.

[0037] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps in any one of the first aspects of the present invention in a multi-agent pursuit decision-making method based on hybrid imitation learning are implemented.

[0038] The beneficial effects brought by the present invention are as follows:

[0039] As can be seen from the above solution, the embodiments of the present invention provide a multi-agent pursuit decision-making method and system based on hybrid imitation learning, which have the following beneficial effects:

[0040] 1) The training efficiency of the model is improved: By combining two imitation learning methods of TD-BC and MT-GAIL, useful information in various quality expert data can be effectively absorbed, thereby significantly reducing the training time of the model and improving the convergence speed.

[0041] 2) The adaptability and generalization ability of the model are enhanced: The hybrid imitation learning model can handle multi-modal expert trajectories, improve the model's performance in complex environments, and avoid the negative impact of low-quality expert data on model training.

[0042] 3) Strong application flexibility: The generated model can be directly applied to subsequent reinforcement learning training. With only minor adjustments and optimizations, a directly usable reinforcement learning model based on expert experience can be trained. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic flow chart of a multi-agent pursuit decision-making method based on hybrid imitation learning provided according to an embodiment;

[0044] Figure 2 It is a classification framework diagram of a task model provided according to an embodiment;

[0045] Figure 3 It is a framework diagram of a hybrid imitation learning method provided according to an embodiment;

[0046] Figure 4 It is a framework diagram of the MT-GAIL algorithm provided according to an embodiment;

[0047] Figure 5 It is a schematic structural diagram of a multi-agent pursuit decision-making system based on hybrid imitation learning provided according to an embodiment;

[0048] Figure 6 It is a schematic diagram of an electronic device provided according to an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] According to the first aspect of the present invention, a multi-agent pursuit decision-making method based on hybrid imitation learning is provided. Referring to Figure 1 as shown, the method includes,

[0051] Step S1: Select a pursuit scenario for imitation learning and determine the expert model of the pursuit scenario;

[0052] Step S2: Generate expert trajectory data by the expert model of the pursuit scenario interacting with the environment and determine the type of the expert trajectory data;

[0053] Step S3: When the type of the expert trajectory data is multimodal, use the multi-expert trajectory generative adversarial imitation learning method to train the expert trajectory data to obtain a first decision-making model;

[0054] Step S4: When the type of the expert trajectory data is unimodal, use the temporal difference error behavior cloning method to train the expert trajectory data to obtain a second decision-making model;

[0055] Step S5: Assign the first decision-making model and the second decision-making model to the agent, and let the agent deduce the first decision-making model and the second decision-making model to obtain a hybrid pursuit decision-making model;

[0056] Step S6: The agent makes decision processing on the pursuit scenario containing static and dynamic targets through the pursuit decision-making model to obtain corresponding pursuit strategies.

[0057] Optionally, in step S1 of the multi-agent pursuit decision-making method based on hybrid imitation learning according to the embodiment of the present invention, the expert model of the pursuit scenario is determined by a decision tree or a reinforcement learning algorithm.

[0058] Specifically, in this embodiment, first select a pursuit scenario available for imitation learning, and the expert model of this pursuit scenario has been obtained through a decision tree or other reinforcement learning methods. For example, an expert model in the form of a decision tree is constructed using expert experience, or a neural network model obtained through reinforcement learning.

[0059] Optionally, in step S2 of the multi-agent pursuit decision-making method based on hybrid imitation learning according to the embodiment of the present invention, when the static target with a fixed position in the pursuit scenario is determined, the type of the expert trajectory data is unimodal; when the dynamic target with a constantly changing position in the pursuit scenario is determined, the type of the expert trajectory data is multimodal.

[0060] For a complex pursuit decision-making problem, the expert experiences worth imitating can be of various types, some are unimodal and some are multimodal. Multimodal expert trajectories refer to expert experiences for a specific task, but the strategies adopted are different. For unimodal expert trajectories with similar modal types of expert trajectories, the Temporal Difference Behavioral Cloning (TD-BC) method is used for training (such as pursuing a static target with a fixed position). For the case of multimodal expert trajectories with diverse modal types of expert trajectories (such as pursuing a dynamic target with a constantly changing position), the Multi-Task Generative Adversarial Imitation Learning (MT-GAIL) method is used for training. In this case, although the task of the agent is the same dynamic target, the expert trajectories may vary greatly. Therefore, for a relatively complex pursuit scenario, in some cases, only unimodal expert experience needs to be learned, and in some cases, a better and more applicable expert experience needs to be obtained based on multimodal expert experience.

[0061] In this embodiment, the imitation learning method adopted is divided according to the type of expert trajectory data (i.e., task type), as Figure 2 shown, the hybrid imitation learning method is used to train a complex pursuit scenario containing static and dynamic targets. The TD-BC method is used to learn the strategy of the agent to pursue the static target, and the MT-GAIL method is used to learn the strategy of the agent to pursue the dynamic target. Expert data for TD-BC training is obtained by sequentially selecting agent models from the agents pursuing the static target, and expert data for MT-GAIL training is obtained by selecting three models from the agents pursuing the dynamic target.

[0062] Optionally, in step S3 of the multi-agent pursuit decision-making method based on hybrid imitation learning according to the embodiment of the present invention, the process of training the expert trajectory data by using the multi-expert trajectory generative adversarial imitation learning method is as follows:

[0063] Step S31: Input each expert trajectory in the expert trajectory data and the generated trajectory of the policy generator into the discriminator at the same time to train the discriminator network, and update the policy network of the policy generator according to the binary cross-entropy loss of the discriminator;

[0064] Step S32: Calculate the reward values of each expert trajectory and each generated trajectory by using the accuracy of the discriminator, store all the calculated reward values in the experience pool, and update the parameters of the policy network of the policy generator through the policy gradient optimization method, so that the generated trajectory of the policy generator gradually approaches the expert trajectory.

[0065] In this embodiment, the multi-expert trajectory generation adversarial imitation learning (MT-GAIL) method is used to learn multi-modal expert trajectories: multiple expert trajectories are input into the generator and discriminator, and the optimization direction of the policy network is determined through the output of the discriminator. The discriminator loss function is defined as the binary cross-entropy loss, and through multiple rounds of iterative training, it is ensured that the discriminator can effectively distinguish expert data from generated data, and finally the model is trained.

[0066] Optionally, before step S31, the multi-agent pursuit decision-making method based on hybrid imitation learning according to an embodiment of the present invention further includes:

[0067] Step S30: The discriminator discriminates each expert trajectory to obtain the reliability coefficient of each expert trajectory, so as to determine the quality of each expert trajectory.

[0068] Optionally, in step S32 of the multi-agent pursuit decision-making method based on hybrid imitation learning according to an embodiment of the present invention, the expert accuracy E_acc and the student accuracy L_acc are respectively expressed as:

[0069]

[0070] where D w (s,a) is the output of the discriminator, s represents the state, and a represents the action; T i represents the trajectory data generated by the policy generator, T E represents the trajectory data generated by expert experience, represents the expectation that the discriminator output is greater than 0.5 under the condition of taking the expert trajectory data as the input, represents the expectation that the discriminator output is less than 0.5 under the condition of taking the student trajectory data as the input;

[0071] The reward value is defined as:

[0072]

[0073] where E(j)_w is the weight of the j-th discriminator.

[0074] Specifically,

[0075] Optionally, in step S4 of the multi-agent pursuit decision-making method based on hybrid imitation learning according to an embodiment of the present invention, the process of training the expert trajectory data by using the temporal difference error behavior cloning method is:

[0076] Step S41: Define that the student model consists of a policy network and a value network. Use the gradient descent strategy to update the parameters of the policy network of the student model, and introduce the temporal difference error TD error to guide the parameter update of the value network of the student model. Among them, the update objective of the gradient descent strategy is to minimize the state-action matching error in the expert trajectory data;

[0077] Step S42: Through continuous repetition of step S42 for iterative update, optimize the parameters of the policy network and the value network of the student model, so that the policy of the student model gradually approaches the policy of the expert model.

[0078] In this embodiment, expert data is generated by the interaction between the expert model and the environment. The temporal difference error is used for value network update, and the mean square error loss function is used to optimize the parameters of the policy network. Through the TD error of the value network and the policy loss function of the policy network, the parameters of the student model are jointly optimized to ensure that the model can approximate the expert policy.

[0079] The following uses a specific embodiment to elaborate in detail on the multi-agent pursuit decision-making method based on hybrid imitation learning in the embodiments of the present invention:

[0080] (1) Select a pursuit scenario available for imitation learning, and an expert model has been obtained for this scenario through decision trees or other reinforcement learning methods.

[0081] (2) Combine the TD-BC method and the MT-GAIL method to form a hybrid imitation learning method. The method framework diagram is as Figure 3 shown. For cases where the expert trajectory modalities are similar, the TD-BC method is used for training. For cases where the expert trajectory modalities are diverse, the MT-GAIL method is used for training.

[0082] (3) Specific implementation steps of the temporal difference error behavior cloning (TD-BC) method:

[0083] a) Generate an expert model: Use expert experience to construct an expert model in the form of a decision tree, or a neural network model obtained through reinforcement learning.

[0084] b) Generate expert data: According to the trajectory data generated by the expert model, represent the expert data in the form of triples X = {(s 1 , a 1 , s 1 '), (s 2 , a 2 , s 2 '),...,(s n , a n , s n ')}. Among them, s j represents the state at time j, and aj Represents the corresponding action, s j ' Represents the state after action a is executed j is completed.

[0085] c) Training the student model: Define that the student model consists of a policy network and a value network. Use the gradient descent strategy to update the parameters of the policy network, and the update objective is to minimize the state-action matching error in the expert data. For the update of the value network, introduce the temporal difference error TD error and use the TD error to guide the parameter update of the value network. Through continuous iteration, optimize the parameters of the policy network and the value network to make the student model gradually approximate the policy of the expert model.

[0086] (4) Specific implementation steps of the multi-expert trajectory generation adversarial imitation learning (MT-GAIL) method:

[0087] a) Acquisition and quality discrimination of expert data: The algorithm framework is as Figure 4 shown. Use multiple expert models to interact with the environment to generate multiple expert trajectories with different qualities. Use the discriminator to discriminate each expert trajectory, and determine the reliability coefficient of the trajectory according to the output of the discriminator to form a complete expert data set.

[0088] b) Update of the policy generator and the discriminator: Input the expert data and the trajectory data generated by the policy generator into the discriminator at the same time, and train the discriminator network to enable it to effectively distinguish expert data and generated data. Update the policy network of the generator according to the binary cross-entropy loss of the discriminator.

[0089] c) Calculation of the reward value and optimization of the policy network: Calculate the reward value of each expert trajectory using the accuracy of the discriminator, and store the reward value of the generated data in the experience pool. Update the parameters of the policy network through the policy gradient optimization method to make the generated data gradually approximate the expert data. Among them, the expert accuracy is represented by E_acc, and the student accuracy is represented by L_acc. The corresponding calculation method is:

[0090]

[0091] Take the output of the discriminator as part of the reward value, and the definition method of the reward value is:

[0092]

[0093] Among them, D w (s,a) is the output of the discriminator, s represents the state, and a represents the action; T i represents the trajectory data generated by the policy generator, T E represents the trajectory data generated by expert experience, Denote the expectation that the discriminator output is greater than 0.5 under the condition of using expert trajectory data as input, and denote the expectation that the discriminator output is less than 0.5 under the condition of using student trajectory data as input;

[0094] During the training process, according to the quality of different expert trajectories, the value network and policy network of the TD-BC model, as well as the policy generator and discriminator of the MT-GAIL model, are updated respectively, so that the model can adapt to the multi-agent decision-making requirements in different complex environments.

[0095] (5) Assign the learned model to the corresponding agent for deduction. Finally, assign the trained model to the corresponding agent, so as to quickly learn the experience of pursuing dynamic and static combined targets.

[0096] According to the second aspect of the present invention, there is provided a multi-agent pursuit decision-making system based on hybrid imitation learning. Refer to Figure 5 as shown, the system 500 includes:

[0097] A first processing module 501, configured to select a pursuit scenario for imitation learning and determine an expert model for the pursuit scenario;

[0098] A second processing module 502, configured to generate expert trajectory data by interacting with the environment through the expert model of the pursuit scenario and determine the type of patent trajectory data;

[0099] A third processing module 503, configured to, when the type of patent trajectory data is multimodal, train the expert trajectory data by using the multi-expert trajectory generation adversarial imitation learning method to obtain a first decision model;

[0100] A fourth processing module 504, configured to, when the type of expert trajectory data is unimodal, train the expert trajectory data by using the temporal difference error behavior cloning method to obtain a second decision model;

[0101] A fifth processing module 505, configured to assign the first decision model and the second decision model to the agent, and deduce the first decision model and the second decision model through the agent to obtain a hybrid pursuit decision model;

[0102] A sixth processing module 506, configured to enable the agent to perform decision-making processing on the pursuit scenario containing dynamic and static targets through the pursuit decision model to obtain corresponding pursuit strategies.

[0103] In summary, the multi-agent pursuit decision-making scheme based on hybrid imitation learning provided by the embodiments of the present invention improves the training efficiency of the model. By combining two imitation learning methods, TD-BC and MT-GAIL, it can effectively extract useful information from various quality expert data, thus significantly reducing the training time of the model and improving the convergence speed. The adaptability and generalization ability of the model are enhanced. The hybrid imitation learning model can process multi-modal expert trajectories, improve the performance of the model in complex environments, and avoid the negative impact of low-quality expert data on model training. It has strong application flexibility. The generated model can be directly applied to subsequent reinforcement learning training. After only minor adjustments and optimizations, a directly usable reinforcement learning model based on expert experience can be trained.

[0104] According to a third aspect of the present invention, there is provided an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in any one of the methods for multi-agent pursuit decision-making based on hybrid imitation learning in the first aspect of the present invention are implemented.

[0105] Figure 6 FIG. is a structural diagram of an electronic device according to an embodiment of the present invention. As Figure 6 shown, the electronic device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, near-field communication (NFC), or other technologies. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the electronic device, or an external keyboard, touchpad, or mouse, etc.

[0106] Those skilled in the art can understand that Figure 6 the structure shown in is only a structural diagram of a part related to the technical solution of the present disclosure, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.

[0107] The fourth aspect of the present invention discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps in a method for wireless energy-harvesting assisted relay of an unmanned aerial vehicle according to any one of the first aspects disclosed in the embodiments of the present invention are implemented.

[0108] According to the fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in a multi-agent pursuit decision-making method based on hybrid imitation learning according to any one of the first aspects of the present invention are implemented.

[0109] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A multi-agent pursuit decision-making method based on hybrid imitation learning, characterized in that: The method comprises: Step S1: selecting a pursuit scene for imitation learning and determining an expert model of the pursuit scene; Step S2: generating expert trajectory data by interacting with the expert model of the pursuit scene and the environment, and determining the type of the expert trajectory data; Step S3: when the type of the expert trajectory data is multimodal, a multi-expert trajectory generation adversarial imitation learning method is used to train the expert trajectory data to obtain a first decision model; Step S4: when the type of the expert trajectory data is single-mode, the expert trajectory data is trained using a temporal difference error behavior cloning method to obtain a second decision model; Step S5: assigning the first decision model and the second decision model to an intelligent agent, and using the intelligent agent to deduce the first decision model and the second decision model to obtain a hybrid pursuit decision model; Step S6: The intelligent agent makes a decision on the pursuit scene containing dynamic and static targets through the pursuit decision model to obtain a corresponding pursuit strategy.

2. The multi-agent pursuit decision-making method based on hybrid imitation learning according to claim 1 is characterized in that: In the step S1, the expert model of the pursuit scenario is determined by a decision tree or a reinforcement learning algorithm.

3. The multi-agent pursuit decision-making method based on hybrid imitation learning according to claim 1 is characterized in that: In the step S2, when a static target with a fixed position in the pursuit scene is being captured, the type of the expert trajectory data is determined to be single-mode; When a dynamic target whose position changes all the time in the pursuit scene is being captured, the type of the expert trajectory data is determined to be multimodal.

4. The multi-agent pursuit decision-making method based on hybrid imitation learning according to claim 1 is characterized in that: In step S3, the process of training the expert trajectory data using the multi-expert trajectory generation adversarial imitation learning method is as follows: Step S31: inputting each expert trajectory in the expert trajectory data and the generated trajectory of the strategy generator into the discriminator at the same time to train the discriminator network, and updating the strategy network of the strategy generator according to the binary cross entropy loss of the discriminator; Step S32: Calculate the reward value of each expert trajectory and each generated trajectory using the accuracy of the discriminator, store all the calculated reward values ​​in the experience pool, and update the parameters of the policy network of the policy generator through the policy gradient optimization method, so that the generated trajectory of the policy generator gradually approaches the expert trajectory.

5. The multi-agent pursuit decision-making method based on hybrid imitation learning according to claim 4 is characterized in that: Before step S31, the method further includes: Step S30: using the discriminator to discriminate each expert trajectory, and obtaining a reliability coefficient of each expert trajectory to determine the quality of each expert trajectory.

6. The multi-agent pursuit decision-making method based on hybrid imitation learning according to claim 4 is characterized in that: In step S32, the expert accuracy E_acc and the student accuracy L_acc are respectively expressed as: Among them, D w (s,a) is the output of the discriminator, s represents the state, a represents the action; T i represents the trajectory data generated by the strategy generator, T E represents the trajectory data generated by expert experience, It represents the expectation that the discriminator output is greater than 0.5 when the expert trajectory data is used as input. It indicates the expectation that the discriminator output is less than 0.5 when the student trajectory data is used as input; The reward value is defined as: Among them, E(j)_w is the weight of the j-th discriminator.

7. The multi-agent pursuit decision-making method based on hybrid imitation learning according to claim 1 is characterized in that: In step S4, the process of training the expert trajectory data using the temporal difference error behavior cloning method is as follows: Step S41: Define the student model to include a policy network and a value network, and use the gradient descent strategy to update The parameters of the policy network of the student model are updated by introducing a temporal difference error (TD) to guide the parameter update of the value network of the student model, wherein the update target of the gradient descent strategy is to minimize the state-action matching error in the expert trajectory data; Step S42: Iterative updating is performed by continuously repeating step S42 to optimize the parameters of the strategy network and the value network of the student model so that the strategy of the student model gradually approaches the strategy of the expert model.

8. A multi-agent pursuit decision system based on hybrid imitation learning, characterized in that: The system comprises: A first processing module is configured to select a pursuit scene for imitation learning and determine an expert model of the pursuit scene; a second processing module configured to generate expert trajectory data through interaction between the expert model of the pursuit scenario and the environment, and determine the type of the patent trajectory data; A third processing module is configured to, when the type of the patent trajectory data is multimodal, train the expert trajectory data using a multi-expert trajectory generation adversarial imitation learning method to obtain a first decision model; a fourth processing module, configured to, when the type of the expert trajectory data is single-modal, train the expert trajectory data using a temporal difference error behavior cloning method to obtain a second decision model; A fifth processing module is configured to assign the first decision model and the second decision model to an intelligent agent, and deduce the first decision model and the second decision model through the intelligent agent to obtain a hybrid pursuit decision model; The sixth processing module is configured such that the intelligent agent performs decision processing on a pursuit scene containing dynamic and static targets through the pursuit decision model to obtain a corresponding pursuit strategy.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps in a multi-agent pursuit decision method based on hybrid imitation learning described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps in the multi-agent pursuit decision method based on hybrid imitation learning described in any one of claims 1 to 7 are implemented.