Multi-agent cooperative guidance method and device

By introducing a collaborative guidance method into a multi-UAV collaborative adversarial decision-making system, and using a base model and a collaborative evaluation network to calculate collaborative scores, the problem of insufficient utilization of collaborative information is solved, training quality and deployment stability are improved, and collaborative quality and team benefits are enhanced.

CN122431405APending Publication Date: 2026-07-21HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
Filing Date
2026-03-26
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In multi-drone collaborative adversarial decision-making systems, existing technologies struggle to effectively utilize collaborative information, resulting in insufficient utilization of collaborative actions, low training sample quality, and unstable team gains. Furthermore, existing collaborative guidance methods are highly invasive, have high adaptation costs, and are complex to deploy.

Method used

By introducing a collaborative guidance method during the training phase, the collaborative score is calculated using the comparable preference scalar output of the pedestal model and the collaborative evaluation network. Combined with scale alignment and gating fusion, actions are selected, and collaborative guidance is disabled during the testing and deployment phases to maintain the stability of the pedestal model and simplify deployment.

Benefits of technology

It improved collaboration quality and training stability, reduced deployment complexity, maintained the stability of adversarial learning, and enhanced team benefits and the smoothness of the training curve.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431405A_ABST
    Figure CN122431405A_ABST
Patent Text Reader

Abstract

The application provides a kind of multi-agent cooperation guidance method and device, it is related to artificial intelligence technical field.The method includes obtaining unmanned aerial vehicle information, for each candidate action in legal action set, the base station model is output to obtain base preference score by comparable preference scalar;Role determination is carried out to unmanned aerial vehicle to be decided, belong to cooperative formation, then enable cooperative guidance, construct cooperative evaluation input to candidate action, obtain cooperative score by cooperative evaluation network, carry out scale alignment, calculate fusion preference, select training phase action;Belong to confrontation formation, then disable cooperative guidance;Training data is generated by interacting with environment and updating base station model;In test and deployment phase, disable cooperative guidance, retain the output unmanned aerial vehicle action decision result of updated base station model.The application can improve cooperation quality and team benefit under the premise of keeping confrontation learning stable, and enhance training curve smoothness, while reducing uncertainty and additional overhead in deployment phase.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and intelligent decision-making technology, and in particular to a method and apparatus for guiding multi-agent collaboration. Background Technology

[0002] Multi-agent cooperative decision-making is widely used in scenarios such as multi-UAV cooperative combat, intelligent equipment swarm control, adversarial simulation training, and game-theoretic decision optimization. Taking a multi-UAV cooperative combat decision-making system as an example, different UAV formations make decisions in an adversarial environment around tasks such as reconnaissance, containment, penetration, target allocation, and path planning. UAVs in the same formation need to complete action coordination and timing coordination under the premise of consistent mission objectives, while different formations have competitive or adversarial relationships.

[0003] In the reinforcement learning training process of multi-UAV cooperative adversarial decision-making systems, the action sampling or action selection mechanism during the training phase directly affects the data distribution of interaction trajectories and further influences the direction of policy updates. If the training phase relies solely on each UAV independently selecting actions based on its own local inputs, it is difficult to fully utilize the collaborative information between UAVs in the same formation, easily leading to problems such as insufficient utilization of collaborative actions, low training sample quality, and unstable team gains. Especially in tasks with significant cross-temporal collaboration, the maneuvering actions, mode switching actions, or target assignment actions performed by a certain UAV at a given moment will affect the action space and gains of other UAVs in the same formation. Therefore, it is necessary to guide collaborative actions during the training phase.

[0004] Existing technologies typically enhance collaborative capabilities through explicit communication mechanisms, centralized training structures, joint value learning, collaborative reward shaping, or mutual information objectives. However, these methods generally suffer from the following drawbacks: First, they are structurally intrusive, requiring modifications to the original training objectives, network structure, or training process, resulting in high adaptation costs. Second, the guidance intensity is difficult to control, easily causing excessive shifts in action distribution in adversarial tasks, thereby disrupting the stable learning of the base policy. Third, some solutions still require additional collaborative modules during the testing or deployment phases, increasing the complexity of the execution end and hindering engineering deployment.

[0005] Therefore, there is a need for a collaborative guidance method and device that can be applied to the training phase of a multi-UAV cooperative adversarial decision-making system. This method and device can provide controlled collaborative guidance for candidate actions of UAVs in the same formation without significantly changing the training objectives of the base and the decision-making structure of the execution end, thereby improving the quality of collaboration, the effectiveness of training samples and the stability of training, and reducing the additional complexity of the deployment phase. Summary of the Invention

[0006] To address the problems of insufficient utilization of collaborative information, high invasiveness of collaborative guidance, and poor consistency between training and deployment in existing technologies for multi-UAV cooperative adversarial decision-making training, embodiments of the present invention provide a multi-agent cooperative guidance method and apparatus. The technical solution is as follows: On the one hand, a multi-agent cooperative guidance method is provided, which is applied to a multi-UAV cooperative adversarial decision-making system. This method is implemented by a multi-agent cooperative guidance device in the training platform of the multi-UAV cooperative adversarial decision-making system. The system includes one or more cooperative formations and one or more adversarial formations, wherein the UAVs in the cooperative formations are cooperative roles, and the UAVs in the adversarial formations are adversarial roles. The method includes: S1. In the reinforcement learning training phase of the multi-UAV cooperative adversarial decision-making system, the local state information, legal action set, historical public information, recent action information of other UAVs in the same formation, and role identification information of the UAV to be decided are obtained. For each candidate action in the legal action set, a comparable preference scalar is output through the base model. The base preference score of each candidate action is obtained based on the comparable preference scalar.

[0007] S2. Determine the role of the drone to be decided. If the drone is a cooperative role, enable cooperative guidance, construct cooperative evaluation input for each candidate action, output cooperative relevance evaluation results through the cooperative evaluation network and calculate the cooperative score, scale align the cooperative score, calculate the fusion preference based on the base preference score and the scale-aligned cooperative score, and select actions for the training phase based on the fusion preference. If the drone is an adversarial role, disable cooperative guidance and select actions for the training phase based on the base preference score.

[0008] S3. Through the action control of the selected training phase, the current decision-making UAV interacts with the multi-UAV cooperative adversarial environment to generate training data and update the base model.

[0009] S4. Disable collaborative guidance during the testing and deployment phases of reinforcement learning, and only retain the drone action decision results output by the updated pedestal model.

[0010] Optionally, the base model in S1 is a policy network or a value network, which is used to output a comparable preference scalar for each candidate action in the set of legal actions corresponding to the current drone to be decided, and the comparable preference scalar is comparable among the candidate actions.

[0011] Optionally, the collaborative evaluation inputs in S2 include: historical window features, recent action features of other UAVs in the same formation, and candidate action features; wherein, historical window features correspond to historical public information in the multi-UAV collaborative confrontation decision-making system, recent action features of other UAVs in the same formation correspond to recent action information of other UAVs in the same formation, and candidate action features correspond to candidate action information in the set of legal actions of the UAV to be decided.

[0012] Optionally, the collaboration score in S2 is calculated from the collaboration relevance evaluation result output by the collaboration evaluation network; when the collaboration relevance evaluation result is a binary classification probability, the calculation formula for the collaboration score is as follows (1): (1) In the formula, Indicates collaboration rating. This represents the probability of binary classification.

[0013] Optionally, the collaboration rating in S2 is scale-aligned, including: Calculate the scaling factor and align the collaboration score to the scale based on the scaling factor; The scaling factor is calculated using the following formula (2): (2) In the formula, Indicates the scaling factor. This represents a preset constant. Indicates the base preference rating. Represents the set of legal actions. Indicates collaboration rating. .

[0014] Optionally, the actions in S2 selected during the training phase based on fusion preferences include: The action of the UAV to be decided during the reinforcement learning training phase is selected based on fusion preferences and exploration strategies.

[0015] On the other hand, a multi-agent cooperative guidance device is provided, which is applied to a multi-agent cooperative guidance method. The device includes: The pedestal preference score calculation module is used in the reinforcement learning training phase of the multi-UAV cooperative adversarial decision-making system to obtain the local state information, legal action set, historical public information, recent action information of other UAVs in the same formation, and role identification information of the UAV to be decided. For each candidate action in the legal action set, the pedestal model outputs a comparable preference scalar, and the pedestal preference score of each candidate action is obtained based on the comparable preference scalar.

[0016] The fusion and sampling module is used to determine the role of the current drone to be decided. If the drone is a cooperative role, cooperative guidance is enabled. Cooperative evaluation input is constructed for each candidate action. The cooperative evaluation network outputs the cooperative relevance evaluation result and calculates the cooperative score. The cooperative score is scale-aligned. The fusion preference is calculated based on the base preference score and the scale-aligned cooperative score. The action in the training phase is selected based on the fusion preference. If the drone is an adversarial role, cooperative guidance is disabled. The action in the training phase is selected based on the base preference score.

[0017] The training update module is used to generate training data and update the base model by controlling the current decision-making UAV to interact with the multi-UAV cooperative adversarial environment through the action control of the selected training phase.

[0018] The output module is used to disable collaborative guidance during the testing and deployment phases of reinforcement learning, retaining only the updated pedestal model output of the drone action decision results.

[0019] Optionally, the base model is a policy network or a value network, which is used to output a comparable preference scalar for each candidate action in the set of legal actions corresponding to the current drone to be decided, and the comparable preference scalar is comparable among the candidate actions.

[0020] Optionally, the collaborative evaluation inputs include: historical window features, recent action features of other UAVs in the same formation, and candidate action features; wherein, historical window features correspond to historical public information in the multi-UAV collaborative confrontation decision-making system, recent action features of other UAVs in the same formation correspond to recent action information of other UAVs in the same formation, and candidate action features correspond to candidate action information in the set of legal actions of the UAV to be decided.

[0021] Optionally, the collaboration score is calculated from the collaboration relevance assessment result output by the collaboration assessment network; when the collaboration relevance assessment result is a binary classification probability, the calculation formula for the collaboration score is as follows (1): (1) In the formula, Indicates collaboration rating. This represents the probability of binary classification.

[0022] Optionally, the fusion and sampling module is further used for: Calculate the scaling factor and align the collaboration score to the scale based on the scaling factor; The scaling factor is calculated using the following formula (2): (2) In the formula, Indicates the scaling factor. This represents a preset constant. Indicates the base preference rating. Represents the set of legal actions. Indicates collaboration rating. .

[0023] The optional fusion and sampling module is further used for: The action of the UAV to be decided during the reinforcement learning training phase is selected based on fusion preferences and exploration strategies.

[0024] On the other hand, a multi-agent cooperative guidance device is provided, the multi-agent cooperative guidance device comprising: a processor; a memory, the memory storing computer-readable instructions, which, when executed by the processor, implement any of the methods described above for multi-agent cooperative guidance.

[0025] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement any of the above-described multi-agent cooperative guidance methods.

[0026] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this invention, a multi-UAV collaborative combat decision-making system is used as a carrier to provide collaborative guidance for the candidate actions of the collaborative formation of UAVs during the training phase, which can improve the quality of collaboration and team benefits.

[0027] By introducing collaborative scoring on top of pedestal preference scoring and aligning it with scales, collaborative guidance can act on the action selection process in a controlled manner, which helps maintain the stability of adversarial learning and enhances the smoothness of the training curve.

[0028] Disabling collaborative guidance during the testing and deployment phases, and retaining only the updated base model for drone action decisions, reduces uncertainty and additional overhead during the deployment phase. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart of a multi-agent collaborative guidance method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the scale difference between the base preference score and the collaboration score provided in an embodiment of the present invention; Figure 3 This is a flowchart of the collaborative gating guidance method for the training phase provided in an embodiment of the present invention; Figure 4 This is a block diagram of a multi-agent cooperative guidance device provided in an embodiment of the present invention; Figure 5 This is a structural block diagram of a multi-agent cooperative guidance device provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of a multi-agent collaborative guidance device provided in an embodiment of the present invention. Detailed Implementation

[0031] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0032] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0033] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0034] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0035] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0036] This invention provides a multi-agent cooperative guidance method, which can be implemented by a multi-agent cooperative guidance device deployed in a multi-UAV cooperative adversarial decision-making system training platform. The multi-agent cooperative guidance device can be a terminal or a server. Figure 1 The flowchart shown illustrates a multi-agent collaborative guidance method. The processing flow of this method may include the following steps: S1. In the reinforcement learning training phase of the multi-UAV cooperative adversarial decision-making system, the local state information, legal action set, historical public information, recent action information of other UAVs in the same formation, and role identification information of the UAV to be decided are obtained. For each candidate action in the legal action set, a comparable preference scalar is output through the base model. The base preference score of each candidate action is obtained based on the comparable preference scalar.

[0037] In one feasible implementation, the local state information includes one or more of the following: the current position information, speed information, heading information, altitude information, remaining task resource information, local environmental perception results, relative state information of nearby targets, and current task stage information of the UAV to be decided.

[0038] The set of legal actions is determined by the system state, task constraints, kinematic constraints, dynamic constraints, collision avoidance constraints, and control rules of the UAV currently making a decision at the current decision-making moment. Each candidate action in the set of legal actions corresponds to an executable action option for the UAV currently making a decision in a multi-UAV cooperative adversarial decision-making system.

[0039] Specifically, the present invention obtains a set of legal actions during the training phase. The base model outputs a base preference score. . It can come from value function classes (such as...) ) or strategy type (preferred) ),in, Indicates an action, It represents a state. However, it must be a comparable scalar for candidate actions or be monotonically transformable into a comparable scalar.

[0040] Among them, multi-agent reinforcement learning is a method system in which multiple agents learn strategies through interaction in a shared environment to maximize cumulative rewards.

[0041] Multi-UAV Cooperative Countermeasure Decision System: There are at least two formations, one or more of which consist of multiple UAVs and need to coordinate, while another formation or multiple formations confront it; the decision can be synchronous or sequential.

[0042] The base model is a policy network or value network, which outputs a comparable preference scalar for each candidate action in the set of legal actions corresponding to the current drone to be decided, and the comparable preference scalar is comparable among the candidate actions.

[0043] Base Preference Rating The base model outputs a comparable preference scalar (or can be monotonically transformed into a comparable scalar) for each candidate action in the set of legal actions, reflecting the adversarial capability preference learned by the base model based on training samples; the value function class can be... Strategy-based preferred selection .

[0044] S2. Determine the role of the drone to be decided. If the drone is a cooperative role, enable cooperative guidance, construct cooperative evaluation input for each candidate action, output cooperative relevance evaluation results through the cooperative evaluation network and calculate the cooperative score, scale align the cooperative score, calculate the fusion preference based on the base preference score and the scale-aligned cooperative score, and select actions for the training phase based on the fusion preference. If the drone is an adversarial role, disable cooperative guidance and select actions for the training phase based on the base preference score.

[0045] In one feasible implementation, role determination and gating are performed. It is determined whether the drone to be decided is in a cooperative or adversarial role. If it is a cooperative role, cooperative guidance is enabled; if it is an adversarial role, cooperative guidance is disabled, and the decision is made directly based on the base preference score. Choose an action.

[0046] Specifically, collaborative guidance includes: Constructing Collaborative Evaluation Input (Action-by-Action): For each candidate action, construct collaborative evaluation input, which includes at least: historical window features, teammate's most recent action features, and candidate action features. The historical window features correspond to historical common information in the multi-UAV collaborative adversarial decision-making system; the teammate's most recent action features correspond to the most recent action information of other UAVs in the same formation; and the candidate action features correspond to candidate action information in the set of legal actions for the UAV currently awaiting decision. The historical window can be maintained using explicit queue maintenance or by reusing base state coding.

[0047] Furthermore, the collaborative evaluation network outputs binary classification probabilities. It is used to estimate the correlation between candidate actions and team collaboration.

[0048] When the collaboration relevance assessment result is a binary probability, the collaboration score is... (Log-odds score) is based on binary probability. After logarithmic probability transformation, we obtain: (1) Among them, the collaborative evaluation network structure can adopt a multilayer perceptron, a lightweight network, or other network structures that can extract features and jointly represent historical public information, recent action information of other UAVs in the same formation, and candidate action information (without affecting the log probability output format).

[0049] Furthermore, scale alignment: based on pedestal preference scores Collaboration rating Numerical scaling factor calculation Used for scoring collaboration Scaling (and cropping if necessary) is performed to control the guidance perturbation and avoid disrupting adversarial learning, ensuring that collaborative guidance does not dominate the action sampling of the current decision-making UAV during the training phase with uncontrolled amplitude.

[0050] One preferred embodiment is: (2) In the formula, As a preset constant, Prevent division by zero.

[0051] at the same time, Other methods for estimating the scale include range, quantile range, and standard deviation.

[0052] Explanation of the necessity of scale alignment: Collaboration scoring With base preference The numerical ranges of data from different networks and targets can vary significantly. If fused directly without scale alignment, cooperative terms may dominate the action selection of the UAV during the training sampling phase, leading to an uncontrollable shift in the training data distribution and thus disrupting the stable learning of the base's adversarial capabilities. For example... Figure 2 As shown, there is a significant scale difference between the pedestal preference rating distribution and the collaboration rating distribution. Therefore, this invention uses a scaling factor... right Scale alignment is performed to make collaborative guidance a controlled perturbation (further pruning is required) to ensure training stability.

[0053] Furthermore, through gating fusion, it is determined whether to combine the collaboration score and the pedestal preference score for action sampling / selection during the training phase, based on the role and conditions.

[0054] Specifically, fusion preference is calculated under the team member role: (3) In the formula, Indicates fusion preference, Indicates the base preference rating. Indicates the scaling factor. This indicates a collaborative rating.

[0055] In addition, the fusion method can also adopt additive fusion, candidate subset restricted fusion (Top-M rearrangement) or pruning fusion.

[0056] Furthermore, based on the calculated fusion preference, and combined with exploration strategies (such as ε-greedy algorithm or soft sampling based on fusion score), the current decision-making UAV selects its actions during the training phase; under adversarial roles, the pedestal preference score is directly used. Choose an action.

[0057] This invention introduces cooperative guidance during the training phase without altering the core structure of the base model or the training objective; the intensity of cooperative guidance is controllable, avoiding inconsistent cooperative signal scales that could lead to the current UAV being dominated by cooperative terms during the training sampling phase; it distinguishes between cooperative and adversarial roles, and cooperative guidance only applies to UAVs within the cooperative formation.

[0058] S3. Through the action control of the selected training phase, the current decision-making UAV interacts with the multi-UAV cooperative adversarial environment to generate training data and update the base model.

[0059] In one feasible implementation, training data is generated by selecting actions and interacting with the environment, and the pedestal model is updated according to the pedestal training method.

[0060] S4. Disable collaborative guidance during the testing and deployment phases of reinforcement learning, and only retain the drone action decision results output by the updated pedestal model.

[0061] In one feasible implementation, the present invention employs an execution phase bypass: the collaborative guidance module is shut down during the testing and deployment phases, and only the drone action decision results are output using the trained base model, in order to reduce deployment uncertainty.

[0062] like Figure 3 As shown, the collaborative gating guidance method for the training phase of this invention includes steps 101–107: Step 101: Obtain the local state information and legal action set of the UAV to be decided. And calculate the base preference score Step 102 involves role determination and gating; under the collaborative role branch, steps 103–106 are executed sequentially to construct collaborative input and calculate collaborative scores. Step 107: Align and merge scales to obtain the action selection for the training phase; Step 107: Perform training update.

[0063] This invention proposes a collaborative guidance scheme suitable for the training phase of a multi-UAV cooperative adversarial decision-making system. It involves action sampling / selection guidance technology during the training phase of multi-agent reinforcement learning in a multi-UAV cooperative adversarial environment, particularly a training phase guidance method based on binary classification log-probability cooperative scoring, scale alignment, and gating fusion. It is applicable to scenarios such as adversarial simulation, UAV cooperative adversarial control, and intelligent equipment swarm control. In specific application scenarios, the current UAV to be decided, the set of legal actions, and the data generated during the training phase can be specifically defined according to the business constraints of the target system.

[0064] In one implementation scenario, the adversarial simulation or group adversarial control scenario can be a multi-UAV cooperative adversarial decision-making system. At this time, the agent corresponds to each individual UAV; the set of legal actions refers to the set of allowed actions determined by the system state, task constraints, kinematic constraints, dynamic constraints, collision avoidance constraints and control rules at the current decision time. Legal actions may include one or more of discrete maneuver actions, discrete control mode switching actions, target allocation actions or path selection actions. The cooperative evaluation samples generated during the training phase include at least: (1) historical public information, preferably shared observations, public environment states, historical trajectory information, historical action sequences or task state information; (2) the action information of the most recent action performed by other UAVs in the same formation, preferably the most recent control command, maneuver action, target switching action or discrete control mode; (3) the candidate action information or selected action information of the UAV to be decided at the present time. The action information comes from the current set of legal actions. The above information is encoded and used as the input of the cooperative evaluation network to calculate the cooperative score of the current action.

[0065] The method of this invention can also be extended to other multi-agent cooperative adversarial decision-making scenarios.

[0066] Therefore, the 'legal action set', 'historical public information', 'recent action information of other UAVs in the same formation', and 'candidate action information or selected action information of the UAV to be decided' in this invention all correspond to data items that can be collected, encoded, and used for training in the multi-UAV cooperative adversarial decision-making system, thereby combining the cooperative guidance method with specific technical application scenarios, rather than just staying at the level of abstract algorithms.

[0067] To address the problems of existing CTDE / centralized critic, which employ centralized critic, credit allocation, and value decomposition, resulting in structure / target coupling, high reuse costs, and difficulty in controlling guidance strength, this invention adopts pluggable guidance during training without altering the base training objective. Scale alignment limits perturbations.

[0068] For existing value decomposition / union The joint value learning approach has problems such as specific training structures and high transfer adaptation costs. This invention adopts a method of decoupling from the base, guiding only on the sampling / selection side.

[0069] To address the problems of existing mutual information collaboration methods that incorporate mutual information targets into training, such as the need to incorporate loss and training coupling, this invention uses log-probability collaborative scoring as a sampling guidance signal, eliminating the need to modify the training target.

[0070] To address the problem that existing reward-based shaping / intrinsic reward methods, which employ additional reward shaping, suffer from scale inconsistencies that can easily disrupt stability, this invention uses scale alignment to make cooperative terms controlled perturbations.

[0071] To address the issues of significant modifications and complex deployment associated with existing communication / intent inference methods that employ the belief / communication module, this invention adopts an execution-phase bypass strategy, where only the base module is retained during deployment.

[0072] In the training and sampling phase, this invention introduces a collaborative score based on binary log odds for the collaborative roles in the collaborative formation, evaluates candidate actions one by one, and superimposes the collaborative score onto the base preference score in a controlled perturbation manner through scale alignment and gating fusion, thereby guiding the sampling distribution to form a better collaborative formation. In the testing and deployment phases, the collaborative guidance is bypassed, and only the base model outputs the UAV action decision results to ensure deployment stability.

[0073] In this embodiment of the invention, the following beneficial effects are achieved: Low-intrusion pluggable: Collaborative guidance is implemented through training sampling side modules, without changing the base training objectives and structure, thus reducing adaptation costs.

[0074] Maintaining stability in adversarial learning: By limiting collaborative guidance to controlled perturbations through scale alignment, the risk of adversarial degradation caused by dominant sampling of collaborative items is reduced.

[0075] Improved profitability and stability: Compared to the fixed scaling factor scheme, the adaptive scaling factor scheme adopted in this invention improves team profitability and provides a smoother training curve; compared to the optimal fixed scaling factor of 0.05, the WP (Winning Percentage) of the adaptive scaling factor in this invention increases from 0.27 to 0.30, and the ADP (Average Difference in Points) increases from -1.7 to -1.5.

[0076] Stable deployment: The bypass collaboration guidance module during the execution phase reduces deployment uncertainty and additional overhead.

[0077] Figure 4 This is a block diagram of a multi-agent cooperative guidance device according to an exemplary embodiment, which is used to implement the above-described multi-agent cooperative guidance method. (Refer to...) Figure 4 The device includes a base preference rating calculation module 310, a fusion and sampling module 320, a training and update module 330, and an output module 340. Among them: The base preference score calculation module 310 is used to acquire the local state information, legal action set, historical public information, recent action information of other UAVs in the same formation, and role identification information of the UAV to be decided during the reinforcement learning training phase of the multi-UAV cooperative adversarial decision system. For each candidate action in the legal action set, the base model outputs a comparable preference scalar and the base preference score of each candidate action is obtained based on the comparable preference scalar.

[0078] The fusion and sampling module 320 is used to determine the role of the current drone to be decided. If the current drone to be decided is a cooperative role, cooperative guidance is enabled, cooperative evaluation input is constructed for each candidate action, cooperative correlation evaluation results are output through the cooperative evaluation network and cooperative score is calculated, scale alignment is performed on the cooperative score, fusion preference is calculated based on the base preference score and the scale-aligned cooperative score, and actions in the training phase are selected based on the fusion preference. If the current drone to be decided is an adversarial role, cooperative guidance is disabled, and actions in the training phase are selected based on the base preference score.

[0079] The training update module 330 is used to generate training data and update the base model by interacting with the multi-drone cooperative adversarial environment through the action control of the selected training phase.

[0080] Output module 340 is used to disable collaborative guidance during the testing and deployment phases of reinforcement learning, retaining only the updated pedestal model output of the UAV action decision results.

[0081] like Figure 5 As shown, the device of the present invention specifically includes: a legal action acquisition module, a base preference score calculation module, a role determination and gating module, a cooperative evaluation input construction module, a cooperative evaluation module, a cooperative score calculation module, a scale alignment module, a fusion and sampling module, a training update module, and an execution phase bypass module. The role determination and gating module is used to control the cooperative guidance branch (cooperative evaluation input construction module, cooperative evaluation module, cooperative score calculation module, scale alignment module, fusion and sampling module) to only be effective for cooperative roles; the execution phase bypass module is used to close the cooperative guidance branch during the testing and deployment phases, retaining only the updated base model output of the UAV action decision results.

[0082] In this embodiment of the invention, the following beneficial effects are achieved: Low-intrusion pluggable: Collaborative guidance is implemented through training sampling side modules, without changing the base training objectives and structure, thus reducing adaptation costs.

[0083] Maintaining stability in adversarial learning: By limiting collaborative guidance to controlled perturbations through scale alignment, the risk of adversarial degradation caused by dominant sampling of collaborative items is reduced.

[0084] Improved profitability and stability: Compared with the fixed scaling factor scheme, the adaptive scaling factor scheme adopted in this invention improves the team profitability index and the training curve is more stable.

[0085] Stable deployment: The bypass collaboration guidance module during the execution phase reduces deployment uncertainty and additional overhead.

[0086] Figure 6 This is a schematic diagram of the structure of a multi-agent collaborative guidance device provided in an embodiment of the present invention, as shown below. Figure 6 As shown, the multi-agent collaborative guidance device may include the above-mentioned Figure 4 The multi-agent collaborative guidance device shown.

[0087] Optionally, the multi-agent collaborative guidance device 410 may include a first processor 2001.

[0088] Optionally, the multi-agent collaborative guidance device 410 may also include a memory 2002 and a transceiver 2003.

[0089] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.

[0090] The following is combined with Figure 6 A detailed description is provided of each component of the adversarial team game multi-agent reinforcement learning training collaborative gating guidance device 410: The first processor 2001 is the control center of the adversarial team game multi-agent reinforcement learning training collaborative gating guidance device 410. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).

[0091] Optionally, the first processor 2001 can perform various functions of the multi-agent cooperative guidance device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0092] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 6 CPU0 and CPU1 are shown in the diagram.

[0093] In a specific implementation, as one example, the adversarial team game multi-agent reinforcement learning training collaborative gating guidance device 410 may also include multiple processors, such as... Figure 6 The first processor 2001 and the second processor 2004 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).

[0094] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.

[0095] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently, and may be connected to the interface circuit of the adversarial team game multi-agent reinforcement learning training collaborative gating guidance device 410. Figure 6 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.

[0096] The transceiver 2003 is used to communicate with network devices or with terminal devices.

[0097] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 6(Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.

[0098] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently, and can be trained through the interface circuit of the collaborative gating guidance device 410 via adversarial team game multi-agent reinforcement learning. Figure 6 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.

[0099] It should be noted that, Figure 6 The structure of the adversarial team game multi-agent reinforcement learning training collaborative gating guidance device 410 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown, or combine some components, or have different component arrangements.

[0100] Furthermore, the technical effects of the multi-agent collaborative guidance device 410 can be referred to the technical effects of the multi-agent collaborative guidance method described in the above method embodiments, and will not be repeated here.

[0101] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0102] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0103] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0104] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0105] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0106] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0107] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0108] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0109] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0110] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0111] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0112] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0113] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A multi-agent collaborative guidance method, characterized in that, The method is applied to a multi-UAV cooperative adversarial decision-making system and is implemented by a multi-agent cooperative guidance device in the training platform of the multi-UAV cooperative adversarial decision-making system. The system includes one or more cooperative formations and one or more adversarial formations, wherein the UAVs in the cooperative formations are cooperative roles and the UAVs in the adversarial formations are adversarial roles. The method includes: S1. In the reinforcement learning training phase of the multi-UAV cooperative adversarial decision-making system, the local state information, legal action set, historical public information, recent action information of other UAVs in the same formation, and role identification information of the UAV to be decided are obtained. For each candidate action in the legal action set, a comparable preference scalar is output through the base model. The base preference score of each candidate action is obtained based on the comparable preference scalar. S2. Determine the role of the drone to be decided. If the drone is a cooperative role, enable cooperative guidance, construct cooperative evaluation input for each candidate action, output cooperative relevance evaluation results through the cooperative evaluation network and calculate the cooperative score, perform scale alignment on the cooperative score, calculate the fusion preference based on the base preference score and the scale-aligned cooperative score, and select actions for the training phase based on the fusion preference. If the drone is an adversarial role, disable cooperative guidance and select actions for the training phase based on the base preference score. S3. Through the action control of the selected training phase, the current decision-making UAV interacts with the multi-UAV cooperative adversarial environment to generate training data and update the base model. S4. Disable collaborative guidance during the testing and deployment phases of reinforcement learning, and only retain the drone action decision results output by the updated pedestal model.

2. The multi-agent cooperative guidance method according to claim 1, characterized in that, The base model in S1 is a policy network or a value network, which is used to output a comparable preference scalar for each candidate action in the set of legal actions corresponding to the current UAV to be decided, and the comparable preference scalar is comparable among the candidate actions.

3. The multi-agent cooperative guidance method according to claim 1, characterized in that, The collaborative evaluation input in S2 includes: historical window features, recent action features of other UAVs in the same formation, and candidate action features; wherein, the historical window features correspond to historical public information in the multi-UAV collaborative confrontation decision-making system, the recent action features of other UAVs in the same formation correspond to the recent action information of other UAVs in the same formation, and the candidate action features correspond to the candidate action information in the set of legal actions of the UAV to be decided.

4. The multi-agent cooperative guidance method according to claim 1, characterized in that, The collaboration score in S2 is calculated from the collaboration relevance evaluation result output by the collaboration evaluation network; when the collaboration relevance evaluation result is a binary classification probability, the calculation formula of the collaboration score is as follows (1): (1) In the formula, Indicates collaboration rating. This represents the probability of binary classification.

5. The multi-agent cooperative guidance method according to claim 1, characterized in that, The scale alignment of the collaboration score in S2 includes: Calculate the scaling factor and align the collaboration score to scale based on the scaling factor; The scaling factor is calculated using the following formula (2): (2) In the formula, Indicates the scaling factor. This represents a preset constant. Indicates the base preference rating. Represents the set of legal actions. Indicates collaboration rating. .

6. The multi-agent cooperative guidance method according to claim 1, characterized in that, The action in the training phase selected according to fusion preference in S2 includes: The action of the UAV to be decided during the reinforcement learning training phase is selected based on fusion preferences and exploration strategies.

7. A multi-agent cooperative guidance device, wherein the multi-agent cooperative guidance device is used to implement the multi-agent cooperative guidance method as described in any one of claims 1-6, characterized in that, The device includes: The base preference score calculation module is used to acquire the local state information, legal action set, historical common information, recent action information of other drones in the same formation, and role identification information of the drone to be decided during the reinforcement learning training phase of the multi-drone cooperative adversarial decision system. For each candidate action in the legal action set, the base model outputs a comparable preference scalar and obtains the base preference score for each candidate action based on the comparable preference scalar. The fusion and sampling module is used to determine the role of the drone to be decided. If the drone is a cooperative role, cooperative guidance is enabled. Cooperative evaluation input is constructed for each candidate action. The cooperative evaluation network outputs the cooperative relevance evaluation result and calculates the cooperative score. The cooperative score is scale-aligned. The fusion preference is calculated based on the pedestal preference score and the scale-aligned cooperative score. The action is selected for the training phase based on the fusion preference. If the drone is an adversarial role, cooperative guidance is disabled. The action is selected for the training phase based on the pedestal preference score. The training update module is used to generate training data and update the base model by interacting with the multi-drone cooperative adversarial environment through the action control of the selected training phase. The output module is used to disable collaborative guidance during the testing and deployment phases of reinforcement learning, retaining only the updated pedestal model output of the drone action decision results.

8. The multi-agent cooperative guidance device according to claim 7, characterized in that, The base model is a policy network or a value network, which is used to output a comparable preference scalar for each candidate action in the set of legal actions corresponding to the current UAV to be decided, and the comparable preference scalar is comparable among the candidate actions.

9. A multi-agent collaborative guidance device, characterized in that, The multi-agent collaborative guidance device includes: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the multi-agent cooperative guidance method as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code, which can be invoked by a processor to execute the multi-agent cooperative guidance method as described in any one of claims 1 to 6.