Multi-agent combat mission cooperation method of structure entropy guided graph neural network

By guiding the dynamic grouping and information aggregation of graph neural networks through structural entropy, the information bottleneck and strategy stability problems of multi-agent systems in complex battlefield environments are solved, achieving efficient and adaptive multi-agent collaboration and improving battlefield situational awareness and mission coordination capabilities.

CN121503969APending Publication Date: 2026-02-10BEIHANG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511486604.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing multi-agent reinforcement learning methods struggle to effectively capture the nonlinear evolution and multiple interaction patterns of battlefield situations in complex battlefield environments. They suffer from insufficient policy stability and generalization ability, poor model adaptability in small sample areas, low transfer efficiency, and difficulty in achieving dynamic task allocation and real-time collaboration.

Method used

By employing a structural entropy-guided graph neural network, a hierarchical community structure is constructed through dynamic grouping and information aggregation, utilizing the principle of structural entropy minimization. This enables decentralized execution and centralized training, enhancing the system's scalability and learning efficiency.

Benefits of technology

It achieves efficient adaptive policy learning and collaborative optimization, enhances the information modeling and policy generalization capabilities of multi-agent systems, and adapts to complex and ever-changing battlefield environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121503969A_ABST
    Figure CN121503969A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent combat task cooperation method for a structure entropy guided graph neural network, and the method comprises the steps: S10, each combat agent interacts with an environment according to an action generated by a strategy network, the environment comprises environment information, task parameters and a preset task target, and the strategy of each combat agent is completely executed in a decentralized manner; collecting complete empirical trajectory data; s20, using the collected data for centralized training; performing value evaluation on the global state of each time step by using a value network; s30, calculating strategy loss and value loss by using a multi-agent near-end strategy optimization algorithm in combination with the output of the strategy network and the value estimation of the output of the value network; updating parameters of the strategy network and the value network by using a gradient descent method; and S40, performing loop iteration. The problems that in a traditional method, the battlefield game dynamic structure sensing ability is insufficient, the hierarchical strategy learning and generalization ability is limited, the adaptability of a model in a small sample area is poor, and the migration efficiency is low are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent warfare technology, and in particular relates to a multi-agent combat mission cooperation method using a structural entropy-guided graph neural network. Background Technology

[0002] In modern intelligent warfare systems, multi-agent collaborative operations are a core capability for achieving efficient mission execution and rapid situational response. Operational mission organization and collaboration methods aim to plan optimal mission allocation and collaboration strategies for multi-service, multi-platform, and multi-mission combat entities in complex, dynamic, and highly adversarial battlefield environments to achieve operational objectives such as rapid assembly, precision strikes, and efficient withdrawal. Compared to traditional single-operation planning, mission collaboration in intelligent warfare environments needs to penetrate to the tactical and combat group levels, dynamically adjust decisions based on real-time situational awareness information, and achieve dynamic trade-offs and optimal allocation among multiple objectives such as security, operational effectiveness, stealth, and mobility.

[0003] With the large-scale deployment of multi-agent systems, how to achieve data-driven multi-agent collaboration mechanisms to continuously optimize the overall combat system performance, improve battlefield adaptability, and effectively cope with complex and ever-changing enemy and friendly situations is the core focus of current research and practical applications. However, existing multi-agent reinforcement learning still faces multiple bottlenecks in decision-making and collaboration in complex battlefield environments. First, in terms of battlefield situation modeling, although traditional directed graph modeling methods can describe the direct interaction relationships between different combat units, they generally suffer from insufficient characterization of high-dimensional complex terrain, multi-task convergence, multi-directional troop maneuvers, and target transfers. This single node relationship has weak expressive power and is difficult to effectively capture the nonlinear evolution and multiple interaction patterns of the battlefield situation, resulting in limited accuracy in predicting future situations and enemy and friendly actions, thereby weakening the robustness and timeliness of task coordination strategies.

[0004] While graph neural network (GNN) methods have inherent advantages in modeling agent interactions and decentralized coordination, existing methods primarily emphasize communication mechanisms, scalability, and expressive power, failing to effectively address the problem of representing high-order structures in multi-agent systems within highly dynamic battlefield environments. In complex situations, the actions of combat units exhibit significant non-determinism, and interaction relationships constantly change with battlefield evolution. Traditional graph structures struggle to support large-scale, multi-dimensional dynamic relationship modeling and reasoning.

[0005] Regarding the stability and generalization capabilities of strategies, existing multi-agent methods still have significant shortcomings in high-dimensional battlefield task decomposition, long-term goal planning, and dynamic task reorganization, relying excessively on manually preset tactical rules and grouping methods. Although methods such as reinforcement learning and game theory have made some progress in some intelligent combat scenarios, their strategy stability and generalization capabilities still urgently need improvement when dealing with complex combat tasks involving multiple levels and multiple objectives. Especially in large-scale collaborative systems, how to achieve dynamic task allocation, real-time coordination within and between groups, and effectively evaluate and optimize joint action plans without relying on fixed grouping and preset task structures is a key challenge currently facing multi-agent collaborative operations.

[0006] The generalization and adaptability of models in different theaters and complex combat environments also face severe challenges. Intelligent combat systems often need to be deployed rapidly in environments with significant differences in terrain, enemy situation, and mission objectives. These differences lead to an imbalance in data richness and situational sampling granularity, which in turn affects the accuracy of operational decisions in low-data areas. At the same time, existing methods are clearly insufficient in their strategy transfer and rapid response capabilities when dealing with real-time mission changes (such as sudden enemy situations, target maneuvers, and dynamic adjustments to mission objectives). If models lack effective knowledge transfer and online adaptation mechanisms, it will greatly increase training and deployment costs and weaken the real-time response capability and overall combat effectiveness of intelligent combat systems. Summary of the Invention

[0007] To address the aforementioned issues, this invention proposes a multi-agent combat mission collaboration method guided by structural entropy using graph neural networks. This method aims to solve core problems in traditional approaches, such as insufficient dynamic structure perception in battlefield games, limited hierarchical policy learning and generalization capabilities, poor model adaptability in small sample areas, and low transfer efficiency. It also aims to address the challenge of achieving scalable and effective multi-agent collaboration in environments with dynamic interaction patterns. This framework dynamically groups agents using the principle of structural entropy and guides information aggregation within the graph neural network architecture. This achieves robust separation between decentralized execution and centralized training, thereby enhancing the system's scalability and learning efficiency.

[0008] To achieve the above objectives, the technical solution adopted by this invention is: a multi-agent combat mission cooperation method using a structure entropy-guided graph neural network, comprising the following steps:

[0009] S10, Environmental Interaction and Data Acquisition: Each combat agent interacts with the environment based on actions generated by the policy network. The environment includes environmental information, task parameters, and preset task objectives. The policy of each combat agent is executed in a completely decentralized manner; complete experience trajectory data is collected.

[0010] S20, State Value Assessment: The collected data is used for centralized training; the global state at each time step is assessed using a value network.

[0011] S30, Network parameter update: Using the multi-agent proximal policy optimization algorithm, the policy loss and value loss are calculated by combining the output of the policy network and the value estimate of the value network output in step S20; then the gradient descent method is used to update the parameters of the policy network and the value network.

[0012] S40, Iterative Loop: Repeat steps S10 to S30 until the agent's policy converges or the preset termination condition is met.

[0013] Furthermore, in step S10, environmental interaction and data acquisition include: collecting complete empirical trajectory data, including local observations from each combat agent. (i) Actions taken (a) (i) The reward R(S) obtained t A t and global state S t Local observations of each combat agent (i) =O(s) (i) Local observation includes the position and velocity of the combat agent itself, as well as the relative position of its target; the joint action A consists of the individual individual actions a of each combat agent i. (i) ∈A constitutes.

[0014] Furthermore, the state value assessment includes the following steps:

[0015] S21, Perform multi-agent environment modeling and enhance node features:

[0016] S22, Graph Information Encoding: Input the graph with enhanced features into the graph neural network (GNN) encoder, process the relationships between nodes through the message passing mechanism, and learn the contextualized feature representation of each node under a given global state and local graph structure;

[0017] S23, Hierarchical graph pooling based on minimizing structural entropy: Pooling the feature representations of each node output by the graph neural network encoder to generate a single global graph feature representation. The pooling process is guided by the principle of minimizing structural entropy.

[0018] S24, Value Output: Evaluate the joint actions of all agents using a centralized value network.

[0019] Furthermore, in step S21, the multi-agent environment modeling and node feature enhancement include the following steps:

[0020] S211, the environment is dynamically represented as a local agent entity graph g for each agent i. (i) = (V, E); In the graph, each node v∈V corresponds to a specific entity, while the edge e∈E exists between the agent and any entity within the perception radius p; the edges between agents are bidirectional, modeling communication; the edges between agents and non-agent entities are unidirectional, modeling perception.

[0021] S212, construct a graph G = (V, E) representing the current global state; where each node i ∈ V represents an independent agent, and each edge (i, j) ∈ E represents a communication or potential interaction link between agents i and j.

[0022] S213, perform detailed feature representation for each node j in the graph; in addition to node features, each edge in the graph is precisely labeled using the Euclidean distance between i and j.

[0023] Furthermore, a detailed feature representation is performed for each node j in the graph, resulting in a feature vector x. j for:

[0024]

[0025] in, These represent the relative position and relative velocity information of j relative to i, respectively. entity type(j) is the type of entity j, used to distinguish different categories of entities.

[0026] The fields vary depending on the entity type:

[0027] If entity j is another agent, it represents the position of the target relative to agent i;

[0028] If entity j is an obstacle or a target, it represents the entity's position in the environment relative to agent i.

[0029] Furthermore, in step S22, the image information encoding includes the following steps:

[0030] Embedding of agent i in the l-th layer of GNN It is calculated by aggregating information from its neighbors and combining it with its own embedding in the previous layer. Its mathematical expression is:

[0031]

[0032] in, Let represent the embedding vector of agent i in the l-th layer of the GNN, and represent the high-level feature representation it has learned; is the embedding of agent i in the previous layer, which contains the information accumulated in its previous iterations; || represents vector concatenation, which connects the current agent's embedding with the information aggregated from its neighbors; W is a permutation-invariant aggregation function used to summarize the embeddings of all neighbors j from agent i at layer (l-1); (l) It is a learnable weight matrix used to perform linear transformations and feature mapping on the concatenated features; σ is a non-linear activation function that introduces non-linearity to enhance the expressive power of the model.

[0033] After iterative information aggregation and feature transformation across multiple layers of GNNs, the agent's strategy is ultimately based on its embedding in the last layer. Generate an action; the process of selecting this action is represented as: ;

[0034] Where, π i The policy network representing agent i is determined based on the final embedding. Output action a i The probability distribution.

[0035] Furthermore, in step S23, the hierarchical graph pooling based on minimizing structural entropy includes the following steps:

[0036] S231, Optimal coding tree construction: Using all nodes as leaf nodes, find a coding tree that minimizes the structural entropy of the entire graph by iteratively executing preset structural adjustment operations; structural entropy H(T) is a measure of the structural complexity of graph G under coding tree T. The lower the value of H(T), the clearer and more natural the community structure revealed by coding tree T.

[0037] S232, Generation of Hierarchical Partition Matrix: Based on the constructed optimal coding tree, extract the node partitioning relationships of each internal level; represent the partitioning relationship of each level as a clustering assignment matrix S. i ; Series matrix {S1,S2,...,S L This describes how to aggregate lower-level nodes layer by layer into higher-level supernodes, up to the root node;

[0038] S233, Graph Pooling Execution: Using the hierarchical partitioning matrix described above, the node feature matrix X output by the graph neural network encoder and the adjacency matrix A of the graph are aggregated layer by layer.

[0039] The aggregation operation at each level is represented as:

[0040]

[0041] in, Assign clustering matrix S i The transpose of Ai Let be the graph adjacency matrix of the i-th layer, representing the connection relationships between nodes in the current layer;

[0042] This process coarsens the node features and structural information of the graph layer by layer, ultimately yielding a single global graph feature representation vector that contains multi-level structural information.

[0043] Furthermore, the value network utilizes a global graph embedding h G Global embedding is achieved by embedding all agents in the last layer of the graph neural network. Generated through pooling operations;

[0044] The value network then utilizes this global context information h G The joint action value function Q(o,a) is estimated by considering the joint action a performed by all agents. This value function quantifies the expected cumulative reward of performing a particular joint action in a given state.

[0045] Furthermore, the value network calculation method is as follows:

[0046]

[0047] READOUT is a pooling function that embeds and aggregates a group of nodes into a single, global representation that can represent the entire graph.

[0048] Furthermore, the value function is expressed as:

[0049] Q(o,a)=f critic (h G ,a);

[0050] Among them, f critic It is the mapping function of the value network, responsible for mapping global embeddings and joint actions to Q-values.

[0051] The beneficial effects of adopting this technical solution are:

[0052] For operational mission organization and planning, this invention proposes a multi-agent cooperation method based on structural entropy-guided graph neural networks, which addresses the multi-agent coordination problem in complex environments and achieves efficient and adaptive policy learning and collaborative optimization.

[0053] This invention achieves lossless graph structure modeling of global information: through an innovative node feature enhancement mechanism, the value network, while deeply mining the topological structure information between agents using graph neural networks, ensures that each computing unit can access complete global state information, fundamentally solving the information bottleneck problem of traditional methods.

[0054] This invention enables adaptive discovery of intelligent agent community structures: based on a pooling method that minimizes structural entropy, it can adaptively and unsupervisedly discover the optimal hierarchical community structure according to the inherent data characteristics of the graph, without pre-setting the number of clusters or compression ratio, making the information aggregation process more intelligent and more in line with the inherent logic of the task scenario.

[0055] This invention achieves a globally optimal hierarchical representation: Unlike traditional methods that perform pooling layer by layer in a greedy manner, structural entropy minimization is a global optimization process that aims to find the overall optimal coding tree. This effectively avoids getting trapped in local optima, thereby generating a more globally valuable and optimal hierarchical graph representation. Attached Figure Description

[0056] Figure 1 This is a schematic diagram of a multi-agent combat mission collaboration method using a structural entropy-guided graph neural network according to the present invention. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described below with reference to the accompanying drawings.

[0058] In this embodiment, see Figure 1 As shown, this invention proposes a multi-agent combat mission cooperation method using a structure entropy-guided graph neural network, comprising the following steps:

[0059] S10, Environmental Interaction and Data Acquisition: Each combat agent (vehicle, aircraft, soldier formation, etc.) interacts with the environment based on actions generated by the policy network. The environment includes environmental information, mission parameters, and preset mission objectives. The policy of each combat agent is executed in a completely decentralized manner; complete experience trajectory data is collected.

[0060] S20, State Value Assessment: The collected data is used for centralized training; the global state at each time step is assessed using a value network.

[0061] S30, Network parameter update: Using the multi-agent proximal policy optimization algorithm, the policy loss and value loss are calculated by combining the output of the policy network and the value estimate of the value network output in step S20; then the gradient descent method is used to update the parameters of the policy network and the value network.

[0062] S40, Iterative Loop: Repeat steps S10 to S30 until the agent's policy converges or the preset termination condition is met.

[0063] As an optimization of the above embodiment, in step S10, environmental interaction and data acquisition include: collecting complete experience trajectory data, including local observations from each combat agent. (i) Actions taken (a)(i) The reward R(S) obtained t A t and global state S t Local observations of each combat agent (i) =O(s) (i) Local observation includes the position and velocity of the combat agent itself, as well as the relative position of its target; the joint action A consists of the individual individual actions a of each combat agent i. (i) ∈A constitutes.

[0064] As an optimization of the above embodiments, the state value assessment includes the following steps:

[0065] S21, Perform multi-agent environment modeling and enhance node features:

[0066] S22, Graph Information Encoding: Input the graph with enhanced features into the graph neural network (GNN) encoder, process the relationships between nodes through the message passing mechanism, and learn the contextualized feature representation of each node under a given global state and local graph structure;

[0067] S23, Hierarchical graph pooling based on minimizing structural entropy: Pooling the feature representations of each node output by the graph neural network encoder to generate a single global graph feature representation. The pooling process is guided by the principle of minimizing structural entropy.

[0068] This invention optimizes agent partitioning and improves collaboration efficiency by minimizing the structural entropy of the graph. The invention uses structural entropy as an indicator to quantify the amount of information and uncertainty in the graph structure, and its minimization process enables the framework to dynamically form highly adaptable agent groups.

[0069] S24, Value Output: Evaluate the joint actions of all agents using a centralized value network.

[0070] Specifically, in step S21, multi-agent environment modeling and node feature enhancement include the following steps:

[0071] S211, the environment is dynamically represented as a local agent entity graph g for each agent i. (i) = (V, E); In the graph, each node v∈V corresponds to a specific entity, while the edge e∈E exists between the agent and any entity within the perception radius p; the edges between agents are bidirectional, modeling communication; the edges between agents and non-agent entities are unidirectional, modeling perception.

[0072] S212, construct a graph G = (V, E) representing the current global state; where each node i ∈ V represents an independent agent, and each edge (i, j) ∈ E represents a communication or potential interaction link between agents i and j.

[0073] S213, perform detailed feature representation for each node j in the graph; in addition to node features, each edge in the graph is precisely labeled using the Euclidean distance between i and j.

[0074] For each node j in the graph, a detailed feature representation is performed, resulting in a feature vector x. j for:

[0075]

[0076] in, These represent the relative position and relative velocity information of j relative to i, respectively. entity type(j) is the type of entity j, used to distinguish different categories of entities.

[0077] The fields vary depending on the entity type:

[0078] If entity j is another agent, it represents the position of the target relative to agent i;

[0079] If entity j is an obstacle or a target, it represents the entity's position in the environment relative to agent i.

[0080] This edge annotation further enriches the spatial relationship information in the graph structure. Through this refined node feature and edge annotation, the local observations of all agents are stitched together to form a global state observation vector. This step ensures that each node holds complete global information in subsequent processing.

[0081] Specifically, in step S22, the image information encoding includes the following steps:

[0082] Embedding of agent i in the l-th layer of GNN It is calculated by aggregating information from its neighbors and combining it with its own embedding in the previous layer. Its mathematical expression is:

[0083]

[0084] in, Let represent the embedding vector of agent i in the l-th layer of the GNN, and represent the high-level feature representation it has learned; is the embedding of agent i in the previous layer, which contains the information accumulated in its previous iterations; || represents vector concatenation, which connects the current agent's embedding with the information aggregated from its neighbors; W is a permutation-invariant aggregation function used to summarize the embeddings of all neighbors j from agent i at layer (l-1); (l) It is a learnable weight matrix used to perform linear transformations and feature mapping on the concatenated features; σ is a non-linear activation function that introduces non-linearity to enhance the expressive power of the model.

[0085] After iterative information aggregation and feature transformation across multiple layers of GNNs, the agent's strategy is ultimately based on its embedding in the last layer. Generate an action; the process of selecting this action is represented as:

[0086] Where, π i The policy network representing agent i is determined based on the final embedding. Output action a i The probability distribution.

[0087] Specifically, in step S23, the hierarchical graph pooling based on minimizing structural entropy includes the following steps:

[0088] S231, Optimal Coding Tree Construction: Using all nodes as leaf nodes, an iterative process of performing pre-defined structural adjustment operations is used to find a coding tree that minimizes the structural entropy of the entire graph. Structural entropy H(T) is a measure of the structural complexity of graph G under coding tree T; a lower H(T) value indicates a clearer and more natural community structure revealed by coding tree T. This optimization process aims to discover the optimal hierarchical community structure in the graph. This dynamic grouping mechanism based on minimizing structural entropy enables the entire multi-agent collaborative structure to adjust flexibly and adaptively to respond to constantly changing environmental contexts and diverse task requirements.

[0089] S232, Generation of Hierarchical Partition Matrix: Based on the constructed optimal coding tree, extract the node partitioning relationships of each internal level; represent the partitioning relationship of each level as a clustering assignment matrix S. i ; Series matrix {S1,S2,...,S L This describes how to aggregate lower-level nodes layer by layer into higher-level supernodes, up to the root node;

[0090] S233, Graph Pooling Execution: Using the hierarchical partitioning matrix described above, the node feature matrix X output by the graph neural network encoder and the adjacency matrix A of the graph are aggregated layer by layer.

[0091] The aggregation operation at each level is represented as:

[0092]

[0093] in, Assign clustering matrix Si The transpose of A i Let be the graph adjacency matrix of the i-th layer, representing the connection relationships between nodes in the current layer;

[0094] This process coarsens the node features and structural information of the graph layer by layer, ultimately yielding a single global graph feature representation vector that contains multi-level structural information.

[0095] As an optimization of the above embodiments, the value network utilizes a global graph embedding h G Global embedding is achieved by embedding all agents in the last layer of the graph neural network. Generated through pooling operations; this global pooling operation aims to capture the overall state and potential cooperative patterns of the entire multi-agent system.

[0096] The value network calculation method is as follows:

[0097]

[0098] Here, READOUT is a pooling function, such as global average pooling or global max pooling, which embeds and aggregates a set of nodes into a single, global representation that can represent the entire graph.

[0099] The value network then utilizes this global context information h G The joint action value function Q(o,a) is estimated by considering the joint action a performed by all agents. This value function quantifies the expected cumulative reward of performing a particular joint action in a given state.

[0100] The value function is expressed as:

[0101] Q(o,a)=f critic (h G ,a);

[0102] Among them, f critic It is the mapping function of the value network, responsible for mapping global embeddings and joint actions to Q-values.

[0103] This centralized value design allows it to fully leverage the global structural embeddings from graph neural networks, enabling more expressive and effective evaluation of joint actions, significantly improving policy optimization and training stability. Through this synergistic mechanism of policy and value, this framework cleverly utilizes the global information advantage brought by centralized training while ensuring the flexibility of decentralized execution by individual agents, thereby greatly improving the learning efficiency and overall collaborative performance of multi-agent systems.

[0104] This invention proposes a multi-agent combat mission collaboration method guided by structural entropy graph neural networks. It aims to address the key bottlenecks of traditional multi-agent reinforcement learning methods in complex dynamic environments through innovative structurally aware graph neural network modeling, a structural entropy-guided hierarchical policy learning and optimization mechanism, and a scalable cross-scenario collaborative policy transfer paradigm. First, this invention introduces the principle of structural entropy to construct structural awareness capabilities, mapping high-dimensional raw states to abstract representations capable of capturing complex spatial dependencies, effectively overcoming the limitations of traditional graph modeling and improving the accuracy of multi-agent state prediction. Based on this, a multi-layered policy optimization mechanism is constructed using the structural entropy minimization criterion to achieve spatiotemporal consistency optimization of long-term policies and enhance the model's adaptability to dynamic environmental changes. Finally, a scalable cross-scenario collaborative policy transfer paradigm improves the generalization ability of policies in new scenarios, reduces dependence on large-scale data, and improves the deployment efficiency of the model in different regions and complex environments.

[0105] At the dynamic environment perception level, this invention overcomes the limitations of traditional methods in characterizing high-order spatial dependencies in complex combat scenarios. Existing technologies often overlook complex spatial interaction features such as multi-path convergence and divergence in combat flows, resulting in the abstract state space failing to accurately represent the dynamic characteristics of the environment, thus causing strategies to oscillate under noise interference. This solution, by dynamically forming agent groups and using structural entropy to guide information aggregation, can better capture dynamic interaction patterns and high-order structural features in multi-agent systems, providing a more accurate dynamic perception foundation for agent collaboration.

[0106] At the multi-agent policy optimization level, addressing the disconnect between coarse-grained and fine-grained task decomposition and long-term goal planning in existing reinforcement learning methods, this invention proposes a hierarchical reinforcement learning path planning mechanism guided by structural entropy. By integrating the principle of minimizing structural entropy into the Actor-Value architecture, an adaptive hierarchical reinforcement learning framework with structure awareness is constructed. Based on the criterion of minimizing structural entropy of states and actions, effective decomposition and path interpretation of long-term goals are achieved. This mechanism eliminates the need for manually designing predefined skills and task partitions, dynamically generating high-level states and actions by minimizing structural entropy, significantly reducing the dimensionality of the learning space and improving the model's responsiveness to environmental changes, thus meeting the real-time requirements of the system.

[0107] At the level of model generalization and adaptability, this invention proposes a scalable cross-scenario collaborative strategy transfer paradigm to address the challenges of low model accuracy and insufficient model transferability in small sample areas due to differences in data richness and sampling granularity across different regions. Abstract policy and skill modules trained in the source environment can be transferred to the target environment via abstract representations guided by structural entropy. Policy parameters can be quickly fine-tuned with only a small number of interaction samples to adapt to the dynamics and transfer patterns of new tasks. This mechanism fully utilizes abstract representations guided by structural information, significantly reducing sample requirements and improving strategy transfer efficiency, making it suitable for multi-agent path planning tasks with sparse operational flow data and dynamically changing environments.

[0108] The core objective of this invention is to construct a complete technology system encompassing high-level perception in multi-agent systems, hierarchical decision-making, and cross-scenario policy transfer, overcoming the bottlenecks of traditional methods in terms of environmental adaptability, policy coherence, and system scalability. Through the reconstruction of structural information theory, it achieves a synergistic improvement in noise suppression, long-term goal optimization, and cross-scenario adaptability, providing a highly robust next-generation decision-making solution for complex and dynamic scenarios such as combat.

[0109] The objectives of this invention include the following aspects: 1. To avoid the loss of global state information while utilizing graph structures to model the relationships between agents. 2. To adaptively and unsupervisedly discover the hierarchical community structure formed by a group of agents in the current state. 3. To aggregate information based on the above structure, thereby achieving a more accurate and profound evaluation of the value of the global state and providing high-quality guiding signals for policy learning.

[0110] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A multi-agent combat mission cooperation method using a structural entropy-guided graph neural network, characterized in that, Including the following steps: S10, Environmental Interaction and Data Acquisition: Each combat agent interacts with the environment based on actions generated by the policy network. The environment includes environmental information, task parameters, and preset task objectives. The policy of each combat agent is executed in a completely decentralized manner; complete experience trajectory data is collected. S20, State Value Assessment: The collected data is used for centralized training; the global state at each time step is assessed using a value network. S30, Network Parameter Update: Using the multi-agent proximal policy optimization algorithm, the policy loss and value loss are calculated by combining the output of the policy network and the value estimate of the value network output in step S20. Then, gradient descent is used to update the parameters of the policy network and the value network; S40, Iterative Loop: Repeat steps S10 to S30 until the agent's policy converges or the preset termination condition is met.

2. The multi-agent combat mission cooperation method of a structural entropy-guided graph neural network according to claim 1, characterized in that, In step S10, environmental interaction and data acquisition include: collecting complete empirical trajectory data, including local observations from each combat agent. (i) Actions taken (a) (i) The reward R(S) obtained t A t and global state S t Local observations of each combat agent (i) =O(s) (i) Local observation includes the position and velocity of the combat agent itself, as well as the relative position of its target; the joint action A consists of the individual individual actions a of each combat agent i. (i) ∈A constitutes.

3. The multi-agent combat mission cooperation method of a structure entropy-guided graph neural network according to claim 1, characterized in that, The state value assessment includes the following steps: S21, Perform multi-agent environment modeling and enhance node features: S22, Graph Information Encoding: Input the graph with enhanced features into the graph neural network (GNN) encoder, process the relationships between nodes through the message passing mechanism, and learn the contextualized feature representation of each node under a given global state and local graph structure; S23, Hierarchical graph pooling based on minimizing structural entropy: Pooling the feature representations of each node output by the graph neural network encoder to generate a single global graph feature representation. The pooling process is guided by the principle of minimizing structural entropy. S24, Value Output: Evaluate the joint actions of all agents using a centralized value network.

4. The multi-agent combat mission cooperation method of a structure entropy-guided graph neural network according to claim 3, characterized in that, In step S21, multi-agent environment modeling and node feature enhancement include the following steps: S211, the environment is dynamically represented as a local agent entity graph g for each agent i. (i) = (V, E); In the graph, each node v∈V corresponds to a specific entity, while the edge e∈E exists between the agent and any entity within the perception radius p; the edges between agents are bidirectional, modeling communication; the edges between agents and non-agent entities are unidirectional, modeling perception. S212, construct a graph G = (V, E) representing the current global state; where each node i ∈ V represents an independent agent, and each edge (i, j) ∈ E represents a communication or potential interaction link between agents i and j. S213, perform detailed feature representation for each node j in the graph; in addition to node features, each edge in the graph is precisely labeled by the Euclidean distance between i and j.

5. The multi-agent combat mission cooperation method of a structural entropy-guided graph neural network according to claim 4, characterized in that, For each node j in the graph, a detailed feature representation is performed, resulting in a feature vector x. j for: in, These represent the relative position and relative velocity information of j relative to i, respectively. entity type(j) is the type of entity j, used to distinguish different categories of entities. The fields vary depending on the entity type: If entity j is another agent, it represents the position of the target relative to agent i; If entity j is an obstacle or a target, it represents the entity's position in the environment relative to agent i.

6. The multi-agent combat mission cooperation method of a structure entropy-guided graph neural network according to claim 3, characterized in that, In step S22, the image information encoding includes the following steps: Embedding of agent i in the l-th layer of GNN It is calculated by aggregating information from its neighbors and combining it with its own embedding in the previous layer. Its mathematical expression is: in, Let represent the embedding vector of agent i in the l-th layer of the GNN, and represent the high-level feature representation it has learned; is the embedding of agent i in the previous layer, which contains the information accumulated in its previous iterations; || represents vector concatenation, which connects the current agent's embedding with the information aggregated from its neighbors; W is a permutation-invariant aggregation function used to summarize the embeddings of all neighbors j from agent i at layer (l-1); (l) It is a learnable weight matrix used to perform linear transformations and feature mapping on the concatenated features; σ is a non-linear activation function that introduces non-linearity to enhance the expressive power of the model. After iterative information aggregation and feature transformation across multiple layers of GNNs, the agent's strategy is ultimately based on its embedding in the last layer. Generate an action; the process of selecting this action is represented as: Where, π i The policy network representing agent i is determined based on the final embedding. Output action a i The probability distribution.

7. The multi-agent combat mission cooperation method of a structural entropy-guided graph neural network according to claim 3, characterized in that, In step S23, the hierarchical graph pooling based on minimizing structural entropy includes the following steps: S231, Optimal coding tree construction: Using all nodes as leaf nodes, find a coding tree that minimizes the structural entropy of the entire graph by iteratively executing preset structural adjustment operations. Structural entropy H(T) is a measure of the structural complexity of graph G under coding tree T. The lower the value of H(T), the clearer and more natural the community structure revealed by coding tree T. S232, Generation of hierarchical partitioning matrix: Based on the constructed optimal coding tree, extract the node partitioning relationship of each level; The partitioning relationship of each level is represented as a clustering assignment matrix S. i ; Series matrix {S1,S2,...,S L This describes how to aggregate lower-level nodes layer by layer into higher-level supernodes, up to the root node; S233, Graph Pooling Execution: Using the hierarchical partitioning matrix described above, the node feature matrix X output by the graph neural network encoder and the adjacency matrix A of the graph are aggregated layer by layer. The aggregation operation at each level is represented as: in, Assign clustering matrix S i The transpose of A i Let be the graph adjacency matrix of the i-th layer, representing the connection relationships between nodes in the current layer; This process coarsens the node features and structural information of the graph layer by layer, ultimately yielding a single global graph feature representation vector that contains multi-level structural information.

8. The multi-agent combat mission cooperation method of a structural entropy-guided graph neural network according to claim 3, characterized in that, The value network utilizes a global graph embedding h G Global embedding is achieved by embedding all agents in the last layer of the graph neural network. Generated through pooling operations; The value network then utilizes this global context information h G The joint action value function Q(o,a) is estimated by considering the joint action a performed by all agents. This value function quantifies the expected cumulative reward of performing a particular joint action in a given state.

9. A multi-agent combat mission cooperation method using a structural entropy-guided graph neural network according to claim 8, characterized in that, The value network calculation method is as follows: READOUT is a pooling function that embeds and aggregates a group of nodes into a single, global representation that can represent the entire graph.

10. A multi-agent combat mission cooperation method using a structural entropy-guided graph neural network according to claim 8, characterized in that, The value function is expressed as: Q(o,a)=f critic (h G ,a); Among them, f critic It is the mapping function of the value network, responsible for mapping global embeddings and joint actions to Q-values.

Citation Information

Cited By

  • Multi-robot patrol method, device and equipment

    CN121900492A

  • CFD sparse linear system iteration strategy and precondition intelligent selection method based on structure perception graph embedding

    CN121920100A

  • Layered multi-agent reinforcement learning-based army unmanned combat cluster confrontation command method, apparatus and device, and storage medium

    CN121960558A