Multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory

Through a hierarchical reinforcement learning method based on structural information theory, the direction-sensitive sparse state map and coding tree are dynamically generated, which solves the problems of role division limitations and insufficient strategy stability in traditional methods, and realizes efficient adaptive decision-making and collaboration of multi-agent collaboration.

CN120335305APending Publication Date: 2025-07-18BEIHANG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510526441.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional hierarchical reinforcement learning has inherent limitations in role division, frequent action conflicts, parameter sensitivity, difficulty in balancing exploration efficiency and collaboration accuracy in complex dynamic environments of multi-agent collaboration, and cannot effectively characterize direction-sensitive state transfer characteristics, resulting in insufficient strategy stability and decreased generalization ability.

Method used

A hierarchical reinforcement learning method based on structural information theory is adopted, through dynamic hierarchical abstraction, multi-scale skill discovery and group collaboration mechanism, deep autoencoder and directed structural entropy optimization technology, a direction-sensitive sparse state map and coding tree are constructed, combined with directed graph modeling and spectral clustering, a role-specific action subspace is dynamically generated to realize autonomous collaboration of the agent.

Benefits of technology

It improves the accuracy and efficiency of multi-agent collaboration, improves adaptive decision-making capabilities in complex dynamic environments, reduces policy conflict rates and noise interference, and realizes stable and efficient collaboration of agents in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335305A_ABST
    Figure CN120335305A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent cooperation method based on hierarchical reinforcement learning of a structural information theory, and the method comprises the steps: carrying out the hierarchical representation of environment observation through a directed structural entropy optimization technology, mapping a high-dimensional original state into a low-dimensional abstract space, eliminating noise interference, and capturing a direction-sensitive state transition mode; on the basis, a multi-scale skill tree is constructed in combination with time sequence characteristics and a structural information theory, and time-space consistency optimization of a long-range strategy is achieved. And finally, expanding to a multi-agent scene, and improving the group decision-making efficiency through dynamic role division and a cooperation framework. The invention aims to solve the key bottleneck of traditional hierarchical reinforcement learning in a multi-agent cooperation complex dynamic environment through dynamic hierarchical abstraction, multi-scale skill discovery and a group cooperation mechanism, and improve the precision and efficiency of multi-agent cooperation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of agent control, and particularly relates to a multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory. Background Art

[0002] The multi-agent cooperation field faces inherent limitations in role division. Mainstream methods need to pre-define the role action space, resulting in frequent role conflicts in dynamic tasks; the action decomposition scheme based on fixed metrics is sensitive to parameter settings and it is difficult to balance exploration efficiency and cooperation accuracy. A deeper constraint stems from the lack of dimension in structural information modeling - traditional theories are modeled based on undirected graphs and cannot effectively represent the direction-sensitive state transition characteristics, and existing directed graph methods have obvious performance bottlenecks in key state recognition and sparse observation processing. These defects jointly restrict the practical application efficiency of hierarchical reinforcement learning in complex dynamic scenarios.

[0003] Current hierarchical reinforcement learning faces multiple bottlenecks in complex environment decision-making. Traditional methods rely too much on manual preset skill definitions and state decomposition rules. For example, although some frameworks can automatically generate skill hierarchies, they still require manual intervention in feature selection; generative methods are limited by expert demonstration data and their generalization ability decreases significantly in unknown scenarios. Existing hierarchical abstraction mechanisms generally adopt a fixed-scale skill set and it is difficult to dynamically adapt to changes in environmental complexity, resulting in insufficient policy stability. The lack of coordination between temporal feature modeling and skill discovery is prominent. Methods based on preset termination conditions restrict the effective execution of long-term policies, and existing feature extraction methods have significant estimation biases in sparse reward scenarios. Summary of the Invention

[0004] To solve the above problems, the present invention proposes a multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory, aiming to solve the key bottlenecks of traditional hierarchical reinforcement learning in the complex dynamic environment of multi-agent cooperation through dynamic hierarchical abstraction, multi-scale skill discovery and group cooperation mechanisms, and improve the accuracy and efficiency of multi-agent cooperation.

[0005] To achieve the above object, the technical solution adopted by the present invention is:

[0006] A multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory, comprising the steps of:

[0007] S10, establishing a dynamic hierarchical abstraction mechanism, including:

[0008] S101, non-linearly reducing the dimension of the original state through a deep autoencoder; for the low-dimensional feature vector generated by the encoder, constructing a weighted graph G representing the similarity relationship between states s ;

[0009] S102, Optimize the graph structure using the dynamic k-nearest neighbor kNN graph filtering algorithm, extract the core structure from the weighted graph G s and eliminate redundant or noisy connections to generate a sparse state graph with the minimum structural information entropy

[0010] S103, Based on the sparse state graph Construct and optimize a multi-layer coding tree structure to reveal the inherent multi-level and multi-granularity community structure in the state space. Achieve dynamic hierarchical adjustment through the improved hierarchical structure entropy minimization algorithm HSCE to obtain the hierarchical coding tree

[0011] S104, Based on the optimal coding tree Calculate the embedding representation of each node in the tree representing an abstract level or cluster, and define how to map the original state to these abstract representations;

[0012] S20, Perform directed structure entropy optimization, including steps:

[0013] S201, After obtaining the static hierarchical abstraction of the state, introduce directionality, model and optimize the transition dynamics between the abstract states, and output the processed strongly connected and weight-normalized weighted directed graph G' d ;

[0014] S202, For the directed graph G' d , Define a directed hierarchical structure entropy H D (T) as the optimization index for hierarchical partitioning, which is defined as an extension of the undirected structure entropy and captures the asymmetric transfer pattern by separating in-degree and out-degree information; The optimization process needs to reflect directionality, and the optimization objective is defined as minimizing and the direction-sensitive optimization principle;

[0015] S203, Find the optimal directed coding tree that minimizes the directed structure entropy through the greedy algorithm

[0016] S30, Multi-scale skill discovery and optimization, including steps:

[0017] S301, Based on the optimal directed coding tree and combined with the properties of the directed graph, discover skills with different time scales, and use the reward function to guide skill learning, and output a set of skills K at one or more levels, where each skill K i is defined by its initiation set, policy, and termination conditions, and is associated with the nodes in ;

[0018] S302, Based on the skill K iEstablish the intrinsic reward r int to guide the learning of the skill strategy so that it can effectively navigate within the corresponding state subspace or complete sub-goals;

[0019] S303. Apply the abstract state, skill set, and intrinsic reward obtained from dynamic abstraction to a hierarchical reinforcement learning framework, and train to construct a two-layer policy structure including a high-level policy and a low-level policy;

[0020] S40. Character-driven multi-agent collaboration. Utilize the action similarity of the agent group to automatically discover characters, and define an exclusive action subspace for each character; By constructing a group action graph and optimizing its encoding tree obtain the character partition {R1, R2,..., R K}}, each character R k corresponds to one or more nodes in it, and the nodes cover a group of similar low-dimensional action representations

[0021] Map the character R k back to the original action space to obtain the corresponding action subspace of this character which contains all the original actions corresponding to the action representations in A' k The optional actions of the agents assigned to the character R k are restricted within ; Obtain the character set Ψ = {R k} and the corresponding action subspace for each character

[0022] Based on the group action graph, use the joint optimization algorithm of spectral clustering and structural entropy to partition characters, and utilize Ψ and the action subspace Realize dynamic character assignment optimization and collaboration conflict resolution through the character utility equilibrium mechanism and policy similarity measurement; Through the policy soft synchronization mechanism, broadcast the new character parameters to relevant agents to achieve a smooth transition of the policy during the character switching process and avoid collaboration oscillations caused by sudden behavior changes.

[0023] Furthermore, the sparse graph construction includes:

[0024] Perform non-linear dimensionality reduction on the original state through a deep autoencoder, and the deep autoencoder adopts a convolutional structure;

[0025] The encoder training adopts a dual-constraint mechanism, including: reconstruction loss: reconstructing the original input through a fully connected decoder to constrain the features to retain complete information; causal preservation loss, enabling the encoder to capture the causal chain of action-state transitions:

[0026] For the low-dimensional feature vectors generated by the encoder, a weighted graph G representing the similarity relationship between states is constructed s , including:

[0027] Calculate the Pearson correlation coefficient between state pairs (s i , s j );

[0028] Take the absolute value of the correlation coefficient |C| as the weight W(i, j) of the edge connecting nodes s i and s j , and construct a weighted graph G representing the similarity relationship between states s .

[0029] Furthermore, a dynamic k-nearest neighbor graph filtering algorithm is used to optimize the graph structure, and the structural uncertainty is minimized by optimizing the graph connectivity, so as to find an optimal number of neighbors k * , including the steps:

[0030] First, perform a dynamic k-value search, traverse k ∈ {3, 5,..., 15}, for each candidate k value, retain the strongest k neighbor connections for each node according to the edge weights of the original graph, so as to construct a corresponding kNN graph and calculate each kNN graph G k ;

[0031] Subsequently, calculate the one-dimensional structural entropy H k of this G 1 (G k ), and this entropy value quantifies the average information amount or uncertainty of a single-step random walk in this graph structure

[0032] By comparing the one-dimensional structural entropies of Gk generated by different k values, select the k that makes H 1 (G k ) reach the minimum value * , and generate a minimum entropy sparse graph

[0033] Furthermore, the structural entropy optimization coding tree includes:

[0034] The initial coding tree T s takes each state node as a leaf node, the root node λ covers all states, and the volume is calculated as

[0035] The optimization process minimizes the structural entropy of the coding tree by iteratively performing stretching and compression operations.

[0036] Furthermore, multi-granularity feature aggregation includes:

[0037] First, calculate the encoding tree representation h of each node α in α ; the representation of the leaf node ν is its corresponding original low-dimensional state representation j ν = s'; for non-leaf node α, the representation j α is aggregated through the representations j of its child nodes αi , and the aggregation weight is based on the structural entropy contribution of the child nodes:

[0038]

[0039] where and represent the structural entropy contributions of child node α i and child node α j relative to their parent node α; L α represents the largest child node;

[0040] After obtaining the representations of all nodes in the encoding tree, define the abstract state set Z s as the set of representations of the direct children of the root node λ Each abstract state corresponds to a subset S in the original state space i ; when a new original state s t is received, first obtain its low-dimensional representation s' t through the encoder, and then map it to the most similar abstract state

[0041] Furthermore, directed graph modeling and weight assignment include

[0042] using the abstract state set Z s as nodes, and based on the historical trajectory data generated by the interaction between the agent and the environment, based on the interaction trajectory data of the agent, define a directed graph G d =(V, E d , W d ), where the node v ∈ V corresponds to the abstract state z s ∈ Z s , and the weight W u→v ∈ E d of the directed edge e d (u → v) represents the transition probability from state u to v;

[0043] To dynamically adapt to environmental changes, a dual decay strategy is adopted for weight update: long-term stability is achieved through an exponential decay factor; while in the event of sudden environmental mutations, a short-term rapid update mode is enabled;

[0044] Direction-sensitive modeling is further achieved through the calculation of in-degree and out-degree separation: for each node u, the sum of out-edge weights and the sum of in-edge weights are respectively counted, and an asymmetric adjacency matrix is constructed;

[0045] During the hierarchical optimization process, nodes with a high out-degree ratio are preferentially retained as community centers to ensure that the partitioning result conforms to the mainstream direction of state transfer.

[0046] Furthermore, the definition and optimization objectives of hierarchical structure entropy include:

[0047] For the directed graph G' d , a directed hierarchical structure entropy H D (T) that can capture directional information is defined as an optimization index for hierarchical partitioning, which is defined as an extension of the undirected structure entropy and captures the asymmetric transfer pattern by separating in-degree and out-degree information; for any non-root node α in the coding tree T, its entropy contribution is composed of the negative entropy of the cross-community flow of the out-edges and the ratio of the hierarchical volume;

[0048] The calculation formula is:

[0049]

[0050] where, represents the cross-community flow of the out-edges between node α and its parent node α- (reflecting the intensity of external transfer at this level), is the out-degree volume of node α (characterizing the activity of the state at this level in the global transfer), vol out (G d ) = ∑ u∈V deg out (u) is the total out-degree volume of the entire graph;

[0051] The optimization objective is to find a coding tree to minimize driving the hierarchical partitioning of the coding tree to conform to the mainstream direction of state transfer;

[0052] The optimization process needs to reflect directionality. During the stretching operation, nodes α with significant differences in the in-degree or out-degree ratio of internal nodes are preferentially split, and the direction separation degree Δ dir is calculated;

[0053] Only when Δ dir is greater than the threshold is the split accepted; during the compression operation, an in-degree similarity constraint is added, that is, only nodes with similar in-degree patterns are allowed to be merged.

[0054] Furthermore, the greedy algorithm includes:

[0055] Taking a directed graph G' d and an initial coding tree T dir as inputs;

[0056] First, in the current tree T dir find a node split that can minimize H D (T dir ) and satisfies the directionality constraint; if such an operation is found and the entropy reduction ΔH fwd > θ fwd , then perform this operation, update the tree structure T dir , and increase the tree height

[0057] Second, perform reverse compression or combination. In the current tree T dir find a node merge that can reduce H D (T dir ) and satisfies the directionality constraint; if such an operation is found and the entropy reduction ΔH rev < θ rev , then perform this operation, update the tree structure T dir , and reduce the tree height;

[0058] Alternate between stretching and compressing or dynamically adjust the strategy according to the decreasing rate of H D (T dir ) until H D (T dir ) converges or reaches the maximum number of iterations or the tree height limit.

[0059] Furthermore, the construction of skill levels based on the directed abstract state transition graph includes:

[0060] First, directly use the hierarchical structure to define skills. Traverse each non-leaf node α of the coding tree, extract the subset of states V α covered by it, and calculate the difference in transition patterns inside and outside the node: for the cross-community traffic and internal transfer traffic of the outgoing edges of node α, define the skill boundary strength, and select the nodes with a skill boundary strength greater than the threshold 2.0 as candidate skill anchor points; for nodes at different levels, generate differentiated skill expressions: for high-level nodes, define coarse-grained skills as macro actions spanning multiple state communities; for low-level nodes, extract fine-grained skills as atomic action sequences;

[0061] For each selected skill anchor point α i , define a corresponding skill κ i ; where the initiation set is the skill κi The set of abstract states at the start of execution; typically States within except for the termination state; skill strategy For skill κ i When activated, according to the current abstract state The strategy for selecting the next action; termination strategy For according to the current abstract state Determine skill κ i The probability of whether to terminate;

[0062] In the skill intrinsic reward function, the transfer entropy reward and / or diversity penalty are also combined. The intrinsic reward function is a weighted combination of these terms;

[0063] The two - layer policy structure includes a high - level policy and a low - level policy;

[0064] High - level policy Π h : Responsible for selecting skills;

[0065] Low - level policy Corresponding to each skill κ i The policy

[0066] Furthermore, a joint optimization algorithm of spectral clustering and structural entropy is used to partition roles:

[0067] Spectral embedding dimensionality reduction: Calculate the normalized Laplacian matrix and extract the first K eigenvectors to form a low - dimensional embedding space where K is the preset maximum number of roles;

[0068] Structural entropy - constrained clustering: On the embedding space Optimize with the goal of minimizing the intra - role structural entropy and the inter - role overlap;

[0069] Use the equilibrium mechanism and policy similarity metric to achieve dynamic role assignment optimization and cooperation conflict resolution, including: First, define the marginal contribution utility of roles;

[0070] When utility imbalance between roles or environmental mutation is detected, trigger the role reorganization process: Role unbinding: Remove the binding relationship of low - utility roles and release their member agents to the free pool;

[0071] Policy similarity clustering: For agents in the free pool, calculate the KL divergence between their policy distributions and existing roles;

[0072] Dynamic re - assignment: Assign agents to the smallest role.

[0073] The beneficial effects of adopting this technical solution:

[0074] First, the present invention hierarchically represents environmental observations through the directed structural entropy optimization technique, maps the high-dimensional original state to a low-dimensional abstract space, eliminates noise interference, and captures the direction-sensitive state transition patterns. On this basis, a multi-scale skill tree is constructed by combining temporal features and structural information theory to achieve the spatio-temporal consistency optimization of long-range strategies. Finally, it is extended to the multi-agent scenario, and the group decision-making efficiency is improved through dynamic role division and cooperation framework.

[0075] At the level of dynamic environment perception, the present invention breaks through the limitation of traditional methods that rely on manually preset state decomposition rules. The existing technology uses a fixed abstraction level, which is difficult to adapt to the direction-sensitive state transition characteristics in complex scenarios, resulting in policy oscillation under noise interference. This solution autonomously extracts the multi-granularity representation of environmental features by constructing a direction-sensitive sparse state transition graph and combining the dynamic coding tree generation technology. The bottom-layer nodes of the coding tree represent fine-grained immediate interactions, and the upper-layer nodes capture the long-range state evolution law, enabling the intelligent agent to autonomously select the best abstraction level according to the task complexity and achieve noise-robust decision-making ability. Aiming at the problem that traditional methods rely on manually preset skills and fixed abstraction granularity, the present invention realizes the autonomous hierarchical representation of environmental features through the directed structural entropy optimization model, combined with the sparse state graph construction and coding tree dynamic generation technology, and improves the adaptive decision-making ability in complex scenarios.

[0076] At the level of temporal strategy optimization, aiming at the fragmentation problem between temporal modeling and strategy generation in the existing skill discovery methods, the present invention proposes a multi-scale skill discovery framework driven by structural information. The temporal causal relationship is modeled through a directed abstract state transition graph, and the skill units across time scales are extracted by combining the hierarchical entropy minimization algorithm. The bottom-layer skills represent short-term action sequences, the upper-layer skills encapsulate long-range strategy segments, and an inherent reward function is designed based on the steady-state distribution to optimize the credit assignment between the nodes of the skill tree. This method breaks through the termination condition limitation of the traditional feature option framework and significantly improves the tactical coherence and execution success rate in complex scenarios such as StarCraft II. Aiming at the problem that the existing skill discovery methods are difficult to integrate temporal dynamic features, the present invention proposes a multi-scale skill discovery mechanism based on the directed abstract state transition graph, and realizes the cross-time-scale strategy association and long-range task optimization through the steady-state distribution-driven reward function and the hierarchical entropy minimization algorithm.

[0077] At the level of group collaborative decision-making, aiming at the problems of rigid multi-agent role division and action conflicts, the present invention proposes a role automatic division mechanism driven by action similarity graphs. By constructing a weighted directed graph to model the action interaction relationships among agents, and combining the structure entropy minimization algorithm to dynamically cluster action patterns with similar functions, a role-specific subspace is generated. The upper-level strategy coordinates the global goals in the abstract role space, and the lower-level strategy executes specific operations in the original action space with structural constraints, realizing the autonomous collaboration and elastic expansion of heterogeneous agents. This framework supports the dynamic adjustment of the role scale according to the task complexity, significantly reducing the policy conflict rate in scenarios such as robot cluster collaboration and enhancing the real-time response ability under sudden tasks. Regarding the limitation of the traditional role-based method that requires predefined action spaces, the present invention extends the structural information theory to multi-agent scenarios. Through minimizing the structure entropy of the action similarity graph, a dynamically adjustable role-specific action subspace is automatically generated, significantly improving the collaboration efficiency of heterogeneous agents.

[0078] The core objective of the present invention is to construct a full-link technology system from single-agent perception and decision-making to multi-agent collaboration, breaking through the bottlenecks of traditional methods in terms of environmental adaptability, policy coherence, and system scalability. Through the reconstruction of the structural information theory, the collaborative improvement of noise suppression, long-range policy optimization, and group collaboration ability is realized, providing a new generation of highly robust decision-making solutions for complex dynamic scenarios such as autonomous driving and intelligent cluster control. Brief Description of the Drawings

[0079] Figure 1 It is a schematic flowchart of a multi-agent collaboration method based on hierarchical reinforcement learning of structural information theory of the present invention. Detailed Embodiment

[0080] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings.

[0081] In this embodiment, referring to Figure 1 as shown, the present invention proposes a multi-agent collaboration method based on hierarchical reinforcement learning of structural information theory, including the steps:

[0082] S10, Establish a dynamic hierarchical abstraction mechanism, including:

[0083] S101, Non-linearly reduce the original state through a deep autoencoder; for the low-dimensional feature vectors generated by the encoder, construct a weighted graph G representing the similarity relationship between states s ; It contains all potential structural relationships and will be used as the input for the subsequent graph structure sparsification step;

[0084] S102, Optimize the graph structure using the dynamic k-nearest neighbor (kNN) graph filtering algorithm, starting from the weighted graph G sExtract the core structure and eliminate redundant or noisy connections to generate a sparse state graph with the minimum structural information entropy. This optimized sparse graph more clearly reflects the core correlation characteristics between states and will serve as the key input data for the next step of constructing and optimizing the hierarchical coding tree.

[0085] S103, Based on the sparse state graph Construct and optimize a multi-layer coding tree structure to reveal the inherent multi-level and multi-granularity community structure in the state space. Achieve dynamic hierarchical adjustment through the improved hierarchical structure entropy minimization algorithm HSCE, and obtain the hierarchical coding tree This optimal coding tree not only deeply reveals the inherent structure of the state (or action) space, but also serves as the basis for subsequent generation of specific abstract state (or abstract action / role) sets and definition of the transition relationships between them.

[0086] S104, Based on the optimal coding tree Calculate the embedding representation of each node in the tree representing an abstract level or cluster, and define how to map the original state to these abstract representations;

[0087] S20, Perform directed structure entropy optimization, including steps:

[0088] S201, After obtaining the static hierarchical abstraction of the state, introduce directionality, model and optimize the transition dynamics between abstract states, and output the processed strongly connected, weight-normalized weighted directed graph G' d ; This graph represents the transition probability or frequency between abstract states;

[0089] S202, For the directed graph G' d , Define a directed hierarchical structure entropy H D (T) as the optimization metric for hierarchical partitioning, defined as an extension of the undirected structure entropy, capturing asymmetric transition patterns by separating in-degree information; The optimization process needs to reflect directionality, and define the optimization goal as minimizing and the direction-sensitive optimization principle;

[0090] S203, Use the greedy algorithm to find the optimal directed coding tree that minimizes the directed structure entropy Output the optimal directed coding tree optimized based on the directed graph G' Output the optimal directed coding tree optimized based on the directed graph G' d The structure of this tree reflects the directional and hierarchical transition patterns between abstract states; This tree's structure reflects the directional, hierarchical transition patterns between abstract states;

[0091] S30, Multi-scale skill discovery and optimization, including steps:

[0092] S301, Based on the optimal directed encoding tree Combined with the properties of the directed graph, skills with different time scales are discovered, and a reward function is used to guide skill learning, outputting a set of skills K at one or more levels, where each skill K i is defined by its initiation set, policy, and termination conditions, and is associated with the nodes in;

[0093] S302, Based on skill K i Establish the intrinsic reward r int to guide the learning of the skill policy so that it can effectively navigate within the corresponding state subspace or complete sub-goals;

[0094] S303, Apply the abstract state, skill set, and intrinsic reward obtained from dynamic abstraction to a hierarchical reinforcement learning framework, and train to construct a two-layer policy structure including a high-level policy and a low-level policy;

[0095] S40, Role-driven multi-agent collaboration, using the action similarity of the agent group to automatically discover roles and define exclusive action subspaces for each role; by constructing a group action graph and optimizing its encoding tree to obtain the role partition {R1, R2,..., R K}}, where each role R k corresponds to one or more nodes in, and the nodes cover a set of similar low-dimensional action representations

[0096] Map the role R k back to the original action space to obtain the corresponding action subspace of the role containing all the original actions corresponding to the action representations in A' k , and the optional actions of the agents assigned to the role R k are restricted within ; Obtain the role set Ψ = {R k} and the corresponding action subspace of each role This step realizes the automatic partition and dynamic binding of group roles by constructing a multi-agent joint state transition graph.

[0097] Based on the group action graph, use the spectral clustering and structural entropy joint optimization algorithm to partition roles, and utilize Ψ and the action subspace Optimize dynamic role allocation and resolve collaboration conflicts through the role-utility equilibrium mechanism and strategy similarity measurement; broadcast new role parameters to relevant agents through the strategy soft synchronization mechanism to achieve a smooth transition of strategies during role switching and avoid collaboration oscillations caused by behavioral mutations. Dynamically adjust the role allocation of agents according to the task progress and environmental changes, and solve potential collaboration conflicts.

[0098] As an optimized solution to the above embodiment, in step S10, the sparse graph construction includes:

[0099] Perform non-linear dimensionality reduction on the original state through a deep autoencoder, and the deep autoencoder Adopts a convolutional structure;

[0100] The mathematical expression is:

[0101]

[0102] Among them, ConvBlock contains convolutional kernels, ReLU activation, and max-pooling operations, and outputs the feature vector s' t , corresponding to the spatial coordinates, motion speed, and key state parameters, and s t Is the original state.

[0103] The encoder training adopts a dual-constraint mechanism, including: reconstruction loss: reconstruct the original input through a fully connected decoder to ensure that the features retain complete information; causal preservation loss, enabling the encoder to capture the causal chain of action-state transitions.

[0104] The causal preservation loss is:

[0105] Among them, λ = 0.5 balances the weights of action prediction and state prediction.

[0106] For the low-dimensional feature vectors generated by the encoder, construct a weighted graph G representing the similarity relationship between states s , including:

[0107] Calculate the Pearson correlation coefficient between state pairs (s i , s j );

[0108] The Pearson correlation coefficient is:

[0109]

[0110] Among them, μ represents the mean of the feature vectors;

[0111] Take the absolute value of the correlation coefficient ∣C∣ as the connection between nodes s i And s jThe weight W(i, j) of the edge is used to construct a weighted graph G representing the similarity relationship between states. s 。

[0112] In step S10, the dynamic k-nearest neighbor graph filtering algorithm is used to optimize the graph structure. By optimizing the connectivity of the graph, its structural uncertainty is minimized to find an optimal number of neighbors k. * , including the steps of:

[0113] First, perform a dynamic k-value search, traversing k ∈ {3, 5,..., 15}. For each candidate k value, keep the strongest k neighbor connections for each node according to the edge weights of the original graph, thereby constructing a corresponding kNN graph to calculate each kNN graph G. k ;

[0114] Subsequently, calculate the one-dimensional structural entropy H k of this G 1 (G k ). This entropy value quantifies the average amount of information or uncertainty of a single-step random walk in this graph structure.

[0115] Its calculation formula is:

[0116]

[0117] Among them, d v is the degree of node v (the sum of the weights of all connected edges), vol(G k ) is the total volume of graph G k (the sum of the degrees of all nodes), and V is the total number of nodes.

[0118] By comparing the one-dimensional structural entropies of Gk generated by different k values, select the k that makes H 1 (G k ) reach the minimum value, that is, perform optimal sparsification, select k * such that k * = argmin k H 1 (G k ), generating a minimum entropy sparse graph.

[0119] In step S10, the structural entropy optimization encoding tree includes:

[0120] The initial encoding tree T s has each state node as a leaf node, and the root node λ covers all states. The volume is calculated as

[0121] The optimization process minimizes the structural entropy of the encoding tree by iteratively performing stretching and compression operations.

[0122]

[0123] Among them, g α represents the sum of the cross-community edge weights between node α and its parent node α-. is the volume of node α.

[0124] Stretching operation: Perform binary splitting on non-leaf node α. First, use the spectral clustering algorithm to divide α into {α1, α2}, and calculate the entropy change after splitting Only when ΔH st < -0.1·H K (T s ) is the split accepted, and the tree height h T increases to h T+1 .

[0125] Compression operation: Simplify the abstraction granularity by merging nodes at the same level, and calculate the merging gain of the node pair (α i , α j )

[0126] When △H cp < 0 and the cross-node edge weights satisfy merging is performed.

[0127] The dynamic adjustment trigger mechanism monitors the change of the environmental complexity in real time, and defines the state space coverage and the transfer entropy volatility When the change rate of Coverage exceeds 15% or ΔH env > 20%, a full-tree reconstruction is triggered.

[0128] By executing the above structural entropy optimization process, an optimal hierarchical coding tree T that can adaptively reflect the multi-granularity community structure of the state is finally obtained s* . This optimal coding tree is not only a profound revelation of the internal structure of the state space, but also the basis for subsequent generation of the set of specific abstract states (or abstract actions / roles) and definition of the transition relationships between them.

[0129] In step S10, the multi-granularity feature aggregation includes:

[0130] First, calculate the representation h of each node α in the coding tree α ; the representation of the leaf node ν is its corresponding original low-dimensional state representation h ν = s'; for the non-leaf node α, the representation h α is aggregated through the representations h of its child nodes αi , and the aggregation weight is based on the structural entropy contribution of the child nodes:

[0131]

[0132] Among them, and represent the structural entropy contributions of child node α i and child node α j relative to its parent node α; L α represents the largest child node;

[0133] After obtaining the representations of all nodes in the encoding tree, define the abstract state set Z s as the set of representations of the direct children of the root node λ Each abstract state corresponds to a subset S in the original state space i ; when a new original state s t is received, first obtain its low-dimensional representation s' t through the encoder, and then map it to the most similar abstract state

[0134] The similarity calculation usually uses the vector inner product or cosine similarity:

[0135]

[0136] Among them, <·,·> represents the similarity calculation.

[0137] In addition, the hierarchical structure of the encoding tree allows for cross-level policy coordination. The agent can, based on the task requirements and the current state, dynamically select an appropriate abstract level h for decision-making based on the structural entropy at a specific level.

[0138] The selection probability p(h) is designed to be related to the level entropy:

[0139]

[0140] where β is a hyperparameter, and is the tree height.

[0141] As an optimized solution to the above embodiment, in step S20, directed graph modeling and weight assignment include

[0142] using the abstract state set Z s as nodes, and based on the historical trajectory data generated by the interaction between the agent and the environment, based on the interaction trajectory data of the agent, define the directed graph G d =(V, E d , W d ), where the node v ∈ V corresponds to the abstract state z s ∈ Z s , and the directed edge eu→v ∈E d The weight W d (u→v) represents the transition probability from state u to v;

[0143] Calculate the fusion of temporal difference TD learning and importance sampling:

[0144]

[0145] where γ = 0.95 is the discount factor, and t reach represents the time step from the initial state to state u, ∈ = 10 -5 Prevent division-by-zero errors;

[0146] To dynamically adapt to environmental changes, the weight update adopts a dual decay strategy: long-term stability is achieved through the exponential decay factor η long = 0.98;

[0147]

[0148] When sudden environmental mutations occur, enable the short-term rapid update mode η short = 0.5, β = 3.0, and accelerate the weight decay of invalid paths through the enhancement coefficient β;

[0149] Direction-sensitive modeling is further achieved through in-degree and out-degree separation calculation: for each node u, separately calculate the sum of out-edge weights and deg out (u)=∑ v∈V W d (u→v) and the sum of in-edge weights and deg in (u)=∑ v∈V W d (v→u), and construct an asymmetric adjacency matrix;

[0150] During the hierarchical optimization process, preferentially retain nodes with a high out-degree ratio as community centers to ensure that the partitioning result fits the mainstream direction of state transition.

[0151] In step S20, the definition of hierarchical structure entropy and the optimization objective include:

[0152] For the directed graph G' d , define a directed hierarchical structure entropy H D (T) as the optimization index for hierarchical partitioning, defined as an extension of the undirected structure entropy, capturing the asymmetric transfer pattern by separating in-degree and out-degree information; for any non-root node α in the encoding tree T, its entropy contribution consists of the negative entropy of the out-edge cross-community flow and the hierarchical volume ratio;

[0153] The calculation formula is:

[0154]

[0155] Among them, represents the cross - community traffic of the out - edge between node α and its parent node α− (reflecting the intensity of external transfer at this level), is the out - degree volume of node α (characterizing the activity of the state at this level in the global transfer), vol out (G d ) = ∑ u∈V deg out (u) is the total out - degree volume of the whole graph.

[0156] The optimization goal is to find an encoding tree to minimize and drive the hierarchical division of the encoding tree to fit the main direction of state transfer;

[0157] The optimization process needs to reflect directionality. In the stretching operation, preferentially split the node α with a significant difference in the in - degree or out - degree ratio of its internal nodes, and calculate the direction separation degree Δ dir ;

[0158] The calculation formula is:

[0159]

[0160] Only accept the split when Δ dir is greater than the threshold; in the compression operation, add the in - degree similarity constraint, that is, only the nodes with similar in - degree patterns are allowed to merge. This step defines the optimization goal (minimize and the direction - sensitive optimization principle.

[0161] In step S20, the greedy algorithm includes:

[0162] Take a directed graph G' d and an initial encoding tree T dir as the input;

[0163] First, in the current tree T dir find a node split that can most significantly reduce H D (T dir ) and satisfies the directionality constraint; if such an operation is found and the entropy reduction amount ΔH fwd > θ fwd , then execute this operation, update the tree structure T dir , and increase the tree height

[0164] Secondly, perform reverse compression or combination. In the current tree T dir find a combination that can reduce H D (T dir) And node merging that satisfies the directional constraint; if such an operation is found and the entropy reduction amount ΔH rev < θ rev , then execute this operation, update the tree structure T dir , and reduce the tree height;

[0165] Alternate between stretching and compression or dynamically adjust the strategy according to the descent rate of H D (T dir ), until H D (T dir ) converges or reaches the maximum number of iterations or the tree height limit.

[0166] As an optimized solution to the above embodiment, in step S30, the construction of the skill hierarchy based on the directed abstract state transition graph includes:

[0167] First, directly use the hierarchical structure to define skills. Traverse each non-leaf node α of the encoding tree, extract the subset of states V α covered by it, and calculate the difference in transition patterns inside and outside the node: for the cross-community traffic of the outgoing edges of node α and the internal transition traffic Define the skill boundary strength Filter the nodes with a skill boundary strength greater than the threshold of 2.0 as candidate skill anchor points; for nodes at different levels, generate differentiated skill expressions: for high-level nodes (h≥3), define coarse-grained skills as macro actions that span multiple state communities; for low-level nodes (h≤2), extract fine-grained skills as atomic action sequences;

[0168] For each selected skill anchor point α i (in the node set at level h ), define a corresponding skill κ i , in the form of where the initiation set is the set of abstract states at which skill κ i starts to execute; usually the states in except the termination state; the skill policy is the policy for selecting the next action according to the current abstract state i when skill κ is activated; the termination policy is the probability of deciding whether skill κ terminates according to the current abstract state i ;

[0169] In step S30, in the skill intrinsic reward function, the transfer entropy reward and / or diversity penalty are also combined, and the intrinsic reward function is a weighted combination of these terms.

[0170] Transfer entropy reward Encourage choosing paths with a clearer transfer mode (related to the skill boundary strength ρ α ).

[0171] Diversity penalty Suppress ineffective wandering within the same state cluster.

[0172] The intrinsic reward function is a weighted combination of these terms:

[0173]

[0174] This reward r int is related to the skill κ i .

[0175] When is dynamically updated, Π s , etc. will also be updated, and the intrinsic reward is adaptively adjusted accordingly.

[0176] Based on this intrinsic reward, the termination condition i and the initiation set of the skill κ can be specified. This step outputs the intrinsic reward function r i defined for each skill κ int as well as the skill initiation set and termination condition

[0177] In step S30, the two - layer policy structure includes a high - level policy and a low - level policy;

[0178] The high - level policy Π h : is responsible for selecting skills;

[0179] The input is the current abstract state (from Section 1.4), and the output is a distribution over the current set of available skills (from Section 3.1), from which a skill κ i is selected.

[0180] The low - level policy corresponding to each skill κ i has a policy

[0181] The input is the current raw state s t and the identifier of the selected skill κ i , and the output is a distribution over the raw action a t .

[0182]

[0183] Low-level policy In the initiation set of skill κ i Execute until the termination condition is met Its training objective is to maximize the intrinsic reward r corresponding to this skill int

[0184] As an optimized solution to the above embodiment, in step S40, a spectral clustering and structural entropy joint optimization algorithm is used to divide roles:

[0185] Spectral embedding dimensionality reduction: Calculate the normalized Laplacian matrix Extract the first K eigenvectors to form a low-dimensional embedding space where K is the preset maximum number of roles;

[0186] Structural entropy constrained clustering: On the embedding space Optimize with the goal of minimizing the intra-role structural entropy and the inter-role overlap.

[0187] The specific formula is:

[0188]

[0189] Among them, is the directed structural entropy corresponding to role λ r = 0.3 is the overlap penalty coefficient. The optimization process is implemented through a greedy merge-split algorithm. Initially, each agent is regarded as an independent role, and candidate role pairs with reduced structural entropy and overlap rate lower than the threshold θ overlap = 0.1 are iteratively merged until convergence. This step outputs the role set Ψ = {R k} and the action subspace corresponding to each role

[0190] In step S40, a balanced mechanism and policy similarity metric are used to achieve dynamic role assignment optimization and cooperation conflict resolution, including: First, define the marginal contribution utility of the role;

[0191] The calculation formula is:

[0192]

[0193] Among them represents the joint action after disabling all member actions of role Q joint is the global state-action value function.

[0194] ​​When utility imbalance between roles or environmental mutations are detected, trigger the role reorganization process: Role unbinding: Remove the binding relationships of low-utility roles and release their member agents to the free pool;

[0195] Policy similarity clustering: For the agents in the free pool, calculate the KL divergence between their policy distributions and the existing roles;

[0196] Dynamic reassignment: Assign agents to the smallest role.

[0197] This patent reconstructs the technical paradigm of hierarchical reinforcement learning through three major technological innovations: The dynamic hierarchical abstraction mechanism breaks through the dimension limitations of manual presetting, the multi-scale skill discovery establishes policy associations across time levels, and the group collaboration framework realizes elastic and scalable role collaboration. In typical scenarios such as four-room navigation, StarCraftII tactical decision-making, and heterogeneous robot collaboration, the decision-making efficiency, policy reliability, and system scalability have all achieved order-of-magnitude improvements, providing a new generation of solutions for intelligent decision-making systems in complex dynamic environments.

[0198] Dynamic Hierarchical Abstraction Mechanism Based on Directed Structural Entropy Optimization

[0199] Traditional hierarchical reinforcement learning methods face core challenges in complex environment decision-making - the manually preset abstraction rules are difficult to capture direction-sensitive features in dynamic environments, and the fixed-granularity hierarchical division leads to limited policy generalization ability. Existing technologies usually model state transition relationships based on undirected graphs, ignoring the key impact of the directionality of action execution on state evolution, resulting in the abstract state space being unable to accurately represent the dynamic characteristics of the environment. This patent breaks through this limitation and proposes a hierarchical abstraction optimization framework with directed structural entropy as the core metric. By constructing a direction-sensitive sparse state transition graph and combining dynamic coding tree generation technology, it realizes the adaptive hierarchical representation of environmental features.

[0200] The core of this mechanism lies in integrating Pearson correlation analysis and kNN graph filtering technology to extract direction-sensitive key state dimensions from high-dimensional raw observations and eliminate redundant noise interference. Based on the modified sparse state graph, design a hierarchical structure entropy optimization algorithm to dynamically adjust the abstraction level by iteratively merging coding tree nodes, forming a multi-granularity environmental representation. Each level of the coding tree corresponds to a state transition pattern with a specific time span, and the bottom-level nodes represent fine-grained immediate interactions, while the upper-level nodes capture long-range state evolution rules. This hierarchical abstraction mechanism enables agents to autonomously select the best abstraction level according to the task complexity, while maintaining policy stability while ensuring decision-making efficiency.

[0201] Compared with traditional static abstraction methods, this mechanism models through a direction-constrained graph structure, significantly enhancing the ability to capture key state transition paths. The dynamic generation process of the encoding tree effectively solves the problem of policy oscillation caused by fixed abstraction granularity, enabling the agent to achieve noise-robust environmental perception and decision-making in complex dynamic scenarios.

[0202] Multi-scale Skill Discovery Method Driven by Structural Information

[0203] Existing skill discovery methods generally suffer from the disconnection problem between temporal feature modeling and skill generation, making it difficult to establish policy associations across time scales. Traditional feature option frameworks rely on manually preset termination conditions, resulting in low collaborative efficiency among skills; skill discovery methods based on spectral clustering are limited by the static partitioning of the original state space and cannot adapt to dynamic environmental changes. This patent proposes a multi-scale skill discovery framework based on a Directed Abstract State Transition Graph (Abstract State Transition Graph), which realizes the deep integration of temporal features and skill generation through a hierarchical optimization mechanism driven by structural entropy.

[0204] The innovation of this method is reflected in three aspects: First, construct an ASG graph with temporal direction constraints, and ensure the integrity of the state transition path through strong connectivity correction technology, providing a direction-sensitive basic structure for skill discovery. Second, the improved hierarchical structural entropy optimization algorithm synchronously optimizes local and global entropy values on the K-layer encoding tree, extracting skill units with time spans - the underlying skills represent short-term action sequences, and the upper-level skills encapsulate long-range policy fragments. Finally, design an inherent reward function based on the steady-state distribution, and optimize the credit assignment between skill tree nodes through cross-level policy gradients, breaking through the sub-optimality problem of skill collaboration in traditional methods.

[0205] Through the autonomous construction of a multi-scale skill tree, the agent can dynamically select the skill granularity according to task requirements. In scenarios of sudden environmental changes, the underlying skills ensure immediate response capabilities, while the upper-level skills maintain long-range policy consistency, forming a decision-making system that balances flexibility and stability.

[0206] Multi-Agent Role Collaboration Framework for Group Collaboration

[0207] The field of multi-agent collaboration has long been restricted by the rigid constraints of role division. Traditional role-based methods need to predefine fixed action subspaces, resulting in policy conflicts and low collaboration efficiency among heterogeneous agents in dynamic tasks. Existing technologies use fixed distance metrics for action clustering, which are difficult to adapt to changes in complex interaction patterns and lack flexibility in adjusting the role scale. This patent innovatively extends the structural information theory to the group collaboration scenario and proposes an automatic role division mechanism driven by an Action Affinity Graph.

[0208] This framework models the action interaction relationships among agents by constructing a weighted directed graph. Nodes represent the behavioral characteristics of agents, and edge weights quantify the action coordination effects. Based on an improved minimum structure entropy algorithm, it dynamically clusters action patterns with similar functions to generate role-specific subspaces. The core breakthrough lies in the design of a two-layer optimization architecture: the upper-layer policy coordinates global goals in the abstract role space, and the lower-layer policy executes specific operations in the original action space with structural constraints, realizing the decoupled optimization of the decision-making layer and the execution layer.

[0209] The role division mechanism has the ability of dynamic adjustment - when the task complexity changes, it autonomously expands or contracts the role scale by adjusting the hierarchical depth of the coding tree. In the heterogeneous agent cooperation scenario, this framework can effectively identify complementary action patterns and avoid redundant behavior conflicts. The dynamic generation process of role subspaces breaks through the rigid constraints of the predefined role system, enabling the group system to autonomously reconstruct the cooperation mode according to environmental changes and significantly improving the adaptability to large-scale collaborative tasks.

[0210] The objectives of the present invention include the following aspects:

[0211] Achieve autonomous hierarchical abstraction in a dynamic environment: Aiming at the problems that traditional methods rely on manually preset state decomposition rules and are difficult to capture direction-sensitive transfer characteristics, the present invention proposes a hierarchical abstraction mechanism based on directed structure entropy. Through sparse graph modeling with direction constraints and dynamic coding tree optimization, it autonomously extracts the key state features of the environment, eliminates the interference of high-dimensional observation noise, and forms a multi-granularity abstract representation space. This mechanism enables agents to dynamically adjust the abstraction level according to the environmental complexity and achieve noise-robust decision-making capabilities in scenarios such as visual navigation and robot control.

[0212] Establish a cross-time-scale policy coordination mechanism: To solve the fragmentation problem of temporal feature modeling and policy generation in existing skill discovery methods, the present invention designs a multi-scale skill discovery framework driven by structural information. By modeling the temporal causal relationship with a directed state transition graph, combining a hierarchical entropy minimization algorithm to extract cross-level skill units, and designing an inherent reward function based on the steady-state distribution. This method breaks through the rigid constraints of traditional skill termination conditions and realizes the dynamic coordination of bottom-level action sequences and high-level policy segments in long-range task execution.

[0213] Construct an elastic and scalable group cooperation paradigm: Aiming at the bottlenecks of rigid role division and frequent action conflicts in multi-agent systems, the present invention proposes a role cooperation framework inspired by structural similarity. It generates role-specific subspaces through dynamic clustering of action interaction graphs and realizes the autonomous cooperation of heterogeneous agents by combining a two-layer optimization strategy (upper-layer role coordination and lower-layer action execution). This framework supports the elastic adjustment of the role scale according to the task complexity, breaks through the limitations of the predefined role system in scenarios such as robot swarm cooperation, significantly reduces the policy conflict rate, and improves the dynamic task adaptability.

[0214] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will also have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory, characterized in that, Including steps: S10. Establish a dynamic hierarchical abstraction mechanism, including: S101, Non-linearly reduce the dimension of the original state through a deep autoencoder; for the low-dimensional feature vectors generated by the encoder, construct a weighted graph G representing the similarity relationship between states s ; S102. Optimize the graph structure using the dynamic k-nearest neighbor (kNN) graph filtering algorithm, extract the core structure from the weighted graph G s and eliminate redundant or noisy connections to generate a sparse state graph with the minimum structural information entropy S103, Based on the sparse state diagram Construct and optimize a multi-layer coding tree structure, reveal the inherent multi-level and multi-granularity community structure in the state space, and implement dynamic level adjustment through the improved hierarchical structure entropy minimization algorithm HSCE to obtain a hierarchical coding tree S104, Based on the optimal coding tree Calculate the embedding representation of each node in the tree, where each node represents an abstract level or cluster, and define how to map the original state to these abstract representations; S20. Conduct directed structural entropy optimization, including steps: S201, After obtaining the static hierarchical abstraction of the state, introduce directionality, model and optimize the transition dynamics between the abstract states, and output the processed strongly connected and weight-normalized weighted directed graph G'. d ; S202. For the directed graph G' d , define a directed hierarchical structure entropy H D (T) as an optimization metric for hierarchical partitioning. Define it as an extension of the undirected structure entropy, which captures asymmetric transfer patterns by separating in-degree and out-degree information; the optimization process needs to reflect directionality, and define the optimization goal as minimizing and the direction-sensitive optimization principle; S203, Find the optimal directed encoding tree that minimizes the directed structural entropy through the greedy algorithm for the S30. Multi-scale skill discovery and optimization, including steps: S301, Based on the optimal directed encoding tree Combined with the properties of the directed graph, skills with different time scales are discovered, and the reward function is used to guide skill learning, outputting a set of skills K at one or more levels, where each skill K i is defined by its initiation set, policy, and termination conditions, and is associated with the nodes in; S302, Based on skill K i Establish an inherent reward r int to guide the learning of the skill strategy so that it can effectively navigate within the corresponding state subspace or complete sub-goals; S303. Apply the abstract states, skill sets, and intrinsic rewards obtained through dynamic abstraction to a hierarchical reinforcement learning framework, and train and construct a two-layer policy structure including a high-level policy and a low-level policy; S40, Role-driven multi-agent collaboration, which automatically discovers roles by leveraging the action similarity of the agent group and defines an exclusive action subspace for each role; by constructing a group action graph and optimizing its encoding tree to obtain the role partition {R1, R2,..., R K}, where each role R k corresponds to one or more nodes in it, and the nodes cover a set of similar low-dimensional action representations Map the character R k back to the original action space to obtain the action subspace corresponding to this character which contains A' k All the original actions corresponding to the action representations in are assigned to the character R k The available actions of the agent of are restricted to inside; Obtain the character set Ψ = {R k} and the action subspace corresponding to each character Based on the group action graph, a combined optimization algorithm of spectral clustering and structural entropy is used to divide roles, and Ψ and the action subspace are utilized. The dynamic role assignment optimization and collaboration conflict resolution are achieved through the role utility equilibrium mechanism and the policy similarity metric; through the policy soft synchronization mechanism, the new role parameters are broadcast to relevant agents to achieve a smooth transition of policies during the role switching process, avoiding collaboration oscillations caused by behavioral mutations.

2. The multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory according to claim 1, characterized in that Sparse graph construction includes: Nonlinear dimensionality reduction of the original state is performed by a deep autoencoder, and the deep autoencoder adopts a convolutional structure; The encoder training adopts a dual-constraint mechanism, including: reconstruction loss: reconstruct the original input through a fully connected decoder to ensure that the features retain complete information; causal preservation loss, enabling the encoder to capture the causal chain of action-state transitions: Construct a weighted graph G representing the similarity relationship between states for the low-dimensional feature vectors generated by the encoder s , including: Calculate the Pearson correlation coefficient between the computational state pairs (s i , s j ); Take the absolute value ∣C∣ of the correlation coefficient as the weight W(i, j) of the edge connecting nodes s i and s j to construct a weighted graph G representing the similarity relationship between states s .

3. A multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory according to claim 2, characterized in that, Optimize the graph structure using the dynamic k-nearest neighbor graph filtering algorithm, and minimize the structural uncertainty of the graph by optimizing its connectivity, thereby finding an optimal number of neighbors k * , including the steps: First, perform dynamic k-value search, traversing k ∈ {3, 5,..., 15}. For each candidate k-value, retain the strongest k neighbor connections for each node according to the edge weights of the original graph, thereby constructing a corresponding kNN graph to calculate each kNN graph G k ; Subsequently, calculate the one-dimensional structural entropy H k of this G 1 (G k ). This entropy value quantifies the average amount of information or uncertainty in a single-step random walk under this graph structure. By comparing the one-dimensional structural entropy of Gk generated by different k values, select the k value that makes H 1 (G k ) reach the minimum value, and generate a minimum entropy sparse graph * ​ 4. A multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory according to claim 3, characterized in that Structural entropy optimization of the coding tree includes: Initial coding tree T s Taking each state node as a leaf node, the root node λ covers all states, and the volume is calculated as The optimization process minimizes the structural entropy of the coding tree by iteratively performing stretching and compression operations.

5. A multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory according to claim 4, characterized in that, Multi-granularity feature aggregation includes: First, calculate the encoding tree representation h of each node α in α ; the representation of the leaf node ν is its corresponding original low-dimensional state representation h ν = s'; for the non-leaf node α, the representation h α is aggregated through its child nodes representation h αi with the aggregation weights based on the structural entropy contributions of the child nodes: Among them, and represent the structural entropy contributions of child node α i and child node α j with respect to their parent node α; L α represents the largest child node; After obtaining the representations of all nodes in the encoding tree, define the set of abstract states $Z$ s as the set of representations of the direct children of the root node $\lambda$ Each abstract state corresponds to a subset $S$ in the original state space i ; when a new original state $s$ is received t , first obtain its low-dimensional representation $s'$ through the encoder t , and then map it to the most similar abstract state 6. A multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory according to claim 1, characterized in that Directed graph modeling and weight assignment, including using the abstract state set Z s as nodes, and defining a directed graph G d =(V, E d , W d ), where the node v ∈ V corresponds to the abstract state z s ∈ Z s , the weight W u→v ∈ E d of the directed edge e d (u → v) represents the transition probability from state u to v; To dynamically adapt to environmental changes, the weight update adopts a dual decay strategy: long-term stability is achieved through an exponential decay factor; while in the event of sudden environmental mutations, a short-term rapid update mode is enabled; Direction-sensitive modeling is further achieved through in-degree and out-degree separation calculation: for each node u, the sum of out-edge weights and the sum of in-edge weights are respectively counted, and an asymmetric adjacency matrix is constructed; During the hierarchical optimization process, nodes with a high out-degree ratio are preferentially retained as community centers to ensure that the partitioning result conforms to the mainstream direction of state transitions.

7. A multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory according to claim 6, characterized in that Definition and optimization objectives of hierarchical structural entropy, including: For the directed graph G' d , a directed hierarchical structure entropy H D (T) that can capture directional information is defined as an optimization metric for hierarchical partitioning. It is defined by extending the undirected structural entropy and captures asymmetric transfer patterns by separating in-degree and out-degree information; for any non-root node α in the coding tree T, its entropy contribution consists of the negative entropy of the out-edge cross-community flow ratio to the layer volume; The calculation formula is: Among them, represents the cross - community traffic of the outgoing edge between node α and its parent node α− (reflecting the intensity of external transfer at this level), is the out - degree volume of node α (characterizing the activity of the state of this layer in the global transfer), vol out (G d ) = ∑ u∈V deg out (u) is the total out - degree volume of the whole graph; The optimization objective is to find an encoding tree Minimize Drive the hierarchical division of the encoding tree to fit the main direction of state transition; The optimization process needs to reflect directionality. In the stretching operation, nodes α with significant differences in the in-degree or out-degree of internal nodes are preferentially split, and the direction separation degree Δ is calculated dir ; Accept splitting only when Δ dir is greater than the threshold; in the compression operation, add an in-degree similarity constraint, that is, only nodes with similar in-degree patterns are allowed to merge.

8. A multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory according to claim 7, characterized in that, The greedy algorithm includes: With a directed graph G' d and an initial coding tree T dir as input; First, in the current tree T dir find a node split that can minimize H D (T dir ) and satisfies the directional constraint; if such an operation is found and the entropy reduction amount ΔH fwd > θ fwd , then execute this operation, update the tree structure T dir , and increase the tree height Secondly, perform reverse compression or combination. In the current tree T dir find nodes that can reduce H D (T dir ) and satisfy the directional constraint for merging; if such an operation is found and the entropy reduction amount ΔH rev < θ rev , then execute this operation, update the tree structure T dir and reduce the tree height. Alternating between stretching and compression or dynamically adjusting the strategy according to the rate of decline of H D (T dir ), until H D (T dir ) converges or reaches the maximum number of iterations or the tree height limit.

9. A multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory according to claim 1, characterized in that Skill hierarchy construction based on a directed abstract state transition graph, including: First, directly utilize 's hierarchical structure to define skills. Traverse each non-leaf node α of the encoding tree, extract the subset of states V α it covers, and calculate the difference in transition patterns inside and outside the node: for the cross-community traffic and internal transfer traffic of the outgoing edges of node α, define the skill boundary strength, and select the nodes with a skill boundary strength greater than the threshold of 2.0 as candidate skill anchor points; for nodes at different levels, generate differentiated skill expressions: for high-level nodes, define coarse-grained skills as macro actions that span multiple state communities; for low-level nodes, extract fine-grained skills as atomic action sequences; For each selected skill anchor α i , define a corresponding skill κ i ; where the initiation set is the set of abstract states at which skill κ i starts to execute; typically the states within except for the termination state; the skill policy is the policy for selecting the next action according to the current abstract state i when skill κ is activated; the termination policy is the probability of deciding whether to terminate skill κ according to the current abstract state i ; In the skill intrinsic reward function, transfer entropy reward and / or diversity penalty are also combined, and the intrinsic reward function is a weighted combination of these terms; The two-layer policy structure includes a high-level policy and a low-level policy; High-level strategy Π h : Responsible for selecting skills; Low-level policy For each skill κ i policy 10. A multi-agent cooperation method based on hierarchical reinforcement learning of structural information theory according to claim 1, characterized in that, Adopt a joint optimization algorithm of spectral clustering and structural entropy to partition roles: Spectral embedding dimensionality reduction: Calculate the normalized Laplacian matrix and extract the first K eigenvectors to form a low-dimensional embedding space ε m , where K is the preset maximum number of roles; Structural entropy-constrained clustering: Optimize on the embedding space ε m to minimize the intra-role structural entropy and the inter-role overlap; Use an equilibrium mechanism and policy similarity metric to achieve dynamic role assignment optimization and cooperation conflict resolution, including: first define the marginal contribution utility of roles; When utility imbalance between roles or environmental mutations are detected, trigger the role reorganization process: role unbinding: remove the binding relationship of low-utility roles and release their member agents to the free pool; Policy similarity clustering: for agents in the free pool, calculate the KL divergence between their policy distributions and existing roles; Dynamic reassignment: Assign agents to the smallest role.

Citation Information

Cited By

  • Multi-agent internal biochemical model training method and system

    CN120654765A

  • Body intelligent control method and system based on physical entropy reduction and manifold remodeling

    CN121572321A

  • Multi-robot collaborative carrying method and system

    CN122414727A