Unmanned aerial vehicle group collaborative optimization method based on alliance mark and structure information
By adopting alliance marking technology and structural information principles in multi-UAV collaborative systems, the problem of collaboration chaos is solved, the collaboration efficiency and stability are improved, and the efficient collaboration of drone swarms in complex tasks is achieved.
Patent Information
- Application Number
- CN202510181821.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-30
AI Technical Summary
The multi-UAV collaborative system has a chaotic collaboration problem in the execution of complex tasks, resulting in unreasonable task allocation, poor information interaction, and incoordinated actions, which in turn affects the efficiency and smooth completion of the task.
The drone cluster collaborative optimization method based on alliance marking and structural information is adopted, and the drone cluster is divided and marked through alliance marking technology, the observation vector is integrated to clarify the collaboration relationship and task division; the structural information principle is used to abstract the environment and task information, enhance the collaborative targetedness, and dynamically adjust the collaboration strategy with the help of structural entropy.
It improves the efficiency and stability of drone cluster collaboration, clarifies the role and tasks of drones in collaboration, reduces the problems of unreasonable task allocation and incoordinated actions, and enhances the dynamic adaptability of the system and the accuracy of information interaction.
Smart Images

Figure CN120066069A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of agent control, and particularly relates to a collaborative optimization method for a swarm of unmanned aerial vehicles (UAVs) based on coalition marking and structural information. Background Art
[0002] With the rapid development of multi-agent technology, multi-agent collaborative systems have shown significant advantages in the execution of complex tasks and are gradually being applied in various fields, bringing higher efficiency and adaptability to the execution of different tasks. Through the collaboration among agents, these systems can efficiently handle complex tasks, quickly adapt to environmental changes, and achieve goals.
[0003] However, the problem of chaotic collaboration in multi-agent collaborative systems is particularly prominent in the complex task of multi-UAV collaborative exploration. The so-called chaotic collaboration means that in the process of collaboration, the UAV swarm has unreasonable task allocation, poor information interaction, and uncoordinated actions. In the multi-UAV collaborative exploration task, this may lead to low exploration efficiency, such as overlapping search areas, delayed or even incorrect information transmission, thus affecting the smooth completion of the task.
[0004] The main reasons for this problem are that the UAV swarm lacks a clear understanding of the roles it plays during collaboration and the dynamic adaptability to environmental changes and task requirements. When facing a complex and changing exploration environment and diverse task requirements, it is difficult for UAVs to form an efficient and stable collaborative relationship. Currently, common collaborative optimization methods are mostly based on fixed rules or simple feedback mechanisms, which are difficult to handle complex dynamic scenarios and may even exacerbate the instability of collaboration. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a collaborative optimization method for a swarm of UAVs based on coalition marking and structural information, aiming to solve the problem of chaotic collaboration in the UAV swarm and improve the collaboration efficiency and stability.
[0006] To achieve the above object, the technical solution adopted by the present invention is: a collaborative optimization method for a swarm of UAVs based on coalition marking and structural information, including the steps of:
[0007] First, divide and mark the UAV swarm through coalition marking technology, and incorporate the observation vector so that the UAV swarm can clarify the collaborative relationship and task division at the initial stage;
[0008] Next, use the principle of structural information to abstractly process the environmental and task information to assist the UAV swarm in refining roles, reasonably allocating tasks, and enhancing the collaboration pertinence; during the collaboration process, use structural entropy to monitor the collaboration status in real time. Once an abnormality is found, dynamically adjust the collaboration strategy in combination with coalition marking and abstract information.
[0009] Furthermore, the drone swarm is divided and labeled through coalition marking technology, and the observation vector is incorporated to clarify the collaboration relationship and task division among the drones at the initial stage, including the steps:
[0010] Data collection: Collect the hardware performance parameters of each drone in the drone swarm; and collect task data, including the goals and requirements of the drone tasks;
[0011] Coalition marking: First, the COLLAB-MARL framework enriches the observation space of the agents by introducing coalition marking, making it contain coalition-specific meta-information and allowing policy adjustment, thus enhancing the basic multi-agent reinforcement learning algorithm; Second, structural entropy quantifies the coordination among agents by analyzing the interaction graph that changes over time, providing a clear measure of collaboration effectiveness.
[0012] Furthermore, the main goal of the COLLAB-MARL framework is to enhance the coordination and cooperation among agents in the same coalition in a heterogeneous reward-sharing environment; the COLLAB-MARL framework includes:
[0013] Coalition marking technology to achieve enhanced observation;
[0014] Through the policy framework, scene changes are made;
[0015] Then model-agnostic learning is utilized.
[0016] Furthermore, the coalition marking technology attaches a coalition-specific identifier to the observation vector of each agent, including:
[0017] Given a multi-agent system: S = (V, E);
[0018] where V = {v 1 , v 2 ,..., v n} represents the set of agents, E represents the set of agent interactions, and C = {C 1 , C 2 ,..., C k} is the set of coalitions in S; each agent v ∈ V belongs to a certain coalition C i ∈ C;
[0019] The integer coalition label of agent v is defined as:
[0020]
[0021] where the scalar label represents the coalition to which agent v belongs and is directly attached to the observation vector o v to form the enhanced observation
[0022] By including in the agent's observations The policy network identifies coalition membership as part of the agent state representation, enabling the shared policy to learn coalition-specific behaviors and improve coordination within the same coalition;
[0023] The policy of agent v depends on the augmented observation where a v represents the action taken by agent v, and θ represents the shared policy parameters.
[0024] Furthermore, the policy framework uses different architectures depending on the specific algorithm.
[0025] Furthermore, the policy framework processes the augmented observation using a shared network combining the original observation o v with coalition information; the policy will map to an action distribution, parameterized by a mean μ v and a variance σ v :
[0026]
[0027] where, is a normal distribution;
[0028] The policy network is implemented as a multi-layer perceptron MLP with L hidden layers:
[0029]
[0030] where φ is the activation function; W (l) is the weight matrix of the l-th layer, is the output of the (l - 1)-th layer, is the output of the l-th layer, and b (l) is the bias vector of the l-th layer;
[0031] The value function is modeled by a critic network, which can be centralized or decentralized; a centralized critic is used for the actor-critic method, and a decentralized critic is used for the value-based method. Training uses batches of collected observations, actions, and rewards; for the actor-critic method, proximal policy optimization is applied, and generalized advantage estimation is used for advantage calculation; for the value-based method, a tailored loss function is minimized to improve value prediction.
[0032] Furthermore, the model-agnostic learning process includes: considering a task distribution p(j), which includes multiple related task distributions j d ; each task distribution j dCorresponding to a specific coalition mark;
[0033] For each task, update it in a meta - learning manner. Before the agent encounters the new task distribution, it first obtains a good initial parameter configuration by training on related tasks.
[0034] After obtaining the updated parameters, further dynamically adjust the coalition mark of the agent; fine - tune for the new target task.
[0035] After completing the dynamic adjustment of the coalition mark, update the global model parameters by optimizing the performance of each task.
[0036] Furthermore, adopt a hierarchical decision - making framework SIDM based on the principle of structural information. Use the principle of structural information to abstract the environmental and task information, assist the UAV swarm to refine roles and reasonably allocate tasks, and enhance the pertinence of cooperation; during the cooperation process, use structural entropy to monitor the cooperation status in real - time. Once an abnormality is found, dynamically adjust the cooperation strategy by combining the coalition mark and abstract information.
[0037] Furthermore, the hierarchical decision - making framework SIDM based on the principle of structural information includes: an adaptive abstraction mechanism, directed structural entropy, and role - based learning.
[0038] Furthermore, the adaptive abstraction mechanism includes:
[0039] 1.1 Data sampling, at each time step t, extract a batch of data of size n from the replay buffer B, which covers the current state S t , action A t , subsequent state S t+1 and reward R t ;
[0040] 1.2 Encoding conversion, adopt an encoder - decoder architecture to convert the original high - dimensional variables into low - dimensional representations;
[0041] The formula is S' t = f s (S t ), S' t+1 = f s (S t+1 ), A' t = f a (A t );
[0042] f s and f a are the state encoder and action encoder respectively. At the same time, design two decoders. The state decoder d s uses cross - entropy inverse target to predict the action between adjacent states, and the action decoder d aReconstruct the next state based on the action and state representation;
[0043] 1.3 Graph Construction, for a pair of different states s in state S i ∈S and s j ∈S, using Pearson correlation analysis to measure its feature similarity C(s i ,s j ), C(s i ,s j )The higher the absolute value, the stronger the state similarity; each state is regarded as a vertex and connected according to the similarity to form a complete weighted undirected state graph G s ;
[0044] 1.4 Optimal Partitioning and Aggregation, Determining the Sparse State Graph G s The key to the optimal partitioning structure;
[0045] Directed structural entropy, including:
[0046] 2.1 Graph structure adjustment: Given a directed graph G dir =(V,E dir ,W dir ), adjust its structure so that there is a directed path between any pair of vertices, and normalize the weights so that the sum of the weighted out-degree of each vertex is 1; the adjusted graph G' dir There exists a unique stationary distribution π s , which is equivalent to the adjacency matrix A' dir The eigenvector corresponding to the maximum eigenvalue 1;
[0047] 2.2 Definition of directed structural entropy: In G' dir Calculate the vertex stationary distribution π s , define one-dimensional directed structural entropy:
[0048] H 1 (G' dir )=-∑ v∈V π s (v)·logπ s (v);
[0049] Based on π s Adjust the relevant items and define the coding tree T dir The distribution entropy of non-root node α in
[0050] Then define G' dir The K-dimensional directed structural entropy of:
[0051]
[0052] 2.3 Entropy optimization process: Merging η based on deDoc algorithm mg and combination ηcb Operator, optimizing the directed structure entropy Determine the optimal tree structure coding strategy
[0053] Role-based learning, including:
[0054] 3.1 Role set definition: Define the abstract action set Z a as the role set Ψ, i.e., each role ρ i ∈ Ψ corresponds to an action subspace
[0055] 3.2 Learning process: At each time step t, for each agent n i , the high-level policy selects a role ρ from the role set Ψ j and its corresponding action subspace A j , completing the role assignment; here τ i is the state-action history of agent n i , and the high-level policy selects a suitable role according to the agent's historical experience; then, the low-level role policy outputs an action according to the global state s t ; the individual actions of all agents together form the joint action a , after acting on the environment, the environment returns the joint reward r t and the subsequent global state s t +1; when training the low-level role policy, a multi-agent reinforcement learning algorithm is adopted, and agents with the same role use a common policy network; the relevant training loss is denoted as L t , and the low-level role policy is optimized by minimizing this loss; marl
[0056] 3.3 Update process: At the regular update interval t up , sample a batch of data from the replay buffer B to construct the variables S t , A t , S t+1 and R t ; use the action abstraction mechanism to redefine the sets Z s and Ψ; the specific operations include constructing an action graph, filtering edges, and generating an optimal coding tree to update the abstract action set Z a , i.e., the role set Ψ; then, calculate the losses L de and L marl ; L de is used to optimize the relevant parameters in the action abstraction process, and L marl is used to optimize the hierarchical policy; by minimizing these two losses, the parameters of the hierarchical policy are continuously adjusted.
[0057] Beneficial effects of adopting this technical solution:
[0058] In the scenario of multi-UAV collaborative exploration, this method enables UAVs to clearly define their roles and tasks in collaboration. Through the coalition marking technology, UAVs can quickly identify collaborative objects and task division of labor. At the same time, by using the principle of structural information, they can abstractly analyze the environment and tasks, optimize the information interaction and action coordination mechanism, improve the task execution efficiency, enhance the system's dynamic adaptability, and achieve real-time monitoring and collaborative adjustment, so as to realize efficient and stable collaboration among multi-agents and build a more reliable and intelligent multi-agent collaborative system.
[0059] This invention can improve the task execution efficiency. With the help of the coalition marking technology, intelligent agents can quickly clarify their own roles and task division of labor, effectively avoid unreasonable task allocation, reduce problems such as overlapping search areas, and thus improve the task execution efficiency. In the management of smart cities, this technology can be used to optimize tasks such as traffic flow and environmental monitoring to ensure the efficient use of resources.
[0060] This invention can enhance the dynamic adaptability. Based on the coalition marking and the principle of structural information, intelligent agents can make dynamic responses to environmental changes and task requirements and flexibly adjust the collaboration strategy. When new obstacles are encountered or the task objectives change during the collaborative work of multi-agents, with the help of model-free learning technology, intelligent agents can re-plan tasks and actions faster, maintain the high efficiency and stability of collaboration, and enhance the system's adaptability to complex environments and diverse tasks.
[0061] This invention can optimize information interaction and action coordination. By abstractly analyzing the environment and tasks using the principle of structural information and combining coalition marking to build an efficient information interaction and action coordination mechanism, it solves the problems of poor information interaction and uncoordinated actions among intelligent agents. In the application of smart cities, this mechanism ensures that in the fields of intelligent transportation, public safety, etc., information can be transmitted between intelligent agents in a timely and accurate manner, avoiding information transmission delays or errors, enabling the work of each intelligent agent to cooperate with each other, and improving the collaborative effect.
[0062] When facing a complex exploration environment and changing task requirements, UAVs can dynamically adjust the collaboration strategy according to the coalition marking and structural information, avoiding unreasonable task allocation and uncoordinated actions. In the environmental monitoring of smart cities, for example, UAVs can quickly respond in air quality detection, identify key monitoring areas, and collaboratively adjust the flight path to improve the comprehensiveness and accuracy of monitoring.
[0063] During the exploration process, the drones can also monitor the collaboration effect in real time. Once signs of chaos in the collaboration are detected, such as delayed information transmission or conflicting actions, they can promptly adjust the collaboration method. This mechanism greatly improves the efficiency and stability of multi-drone collaborative exploration, enhances the system's adaptability to complex environments and tasks, and thus constructs a more efficient and reliable multi-drone collaborative exploration system. With this technology, multi-drones can achieve rapid response and precise execution in collaborative exploration tasks, while minimizing the efficiency loss caused by collaborative chaos to the greatest extent, and demonstrating great value in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 Schematic flow diagram of a method for collaborative optimization of a drone swarm based on coalition marking and structural information according to the present invention;
[0065] Figure 2 Schematic structural diagram of the COLLAB-MARL framework in an embodiment of the present invention;
[0066] Figure 3 Schematic structural diagram of the hierarchical decision-making framework SIDM framework based on the principle of structural information in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings.
[0068] In this embodiment, as shown in Figure 1 a method for collaborative optimization of a drone swarm based on coalition marking and structural information is proposed according to the present invention, including the steps:
[0069] First, the drone swarm is divided and marked through coalition marking technology, and the observation vector is incorporated, so that the drone swarm can clarify the collaboration relationship and task division at the initial stage;
[0070] Next, the environmental and task information is abstracted by using the principle of structural information to assist the drone swarm in refining roles and reasonably allocating tasks, enhancing the collaboration pertinence; during the collaboration process, the collaboration state is monitored in real time by means of structural entropy. Once an abnormality is detected, the collaboration strategy is dynamically adjusted in combination with coalition marking and abstract information.
[0071] As an optimized solution of the above embodiment, the drone swarm is divided and marked through coalition marking technology, and the observation vector is incorporated, so that the drone swarm can clarify the collaboration relationship and task division at the initial stage, including the steps:
[0072] Data collection: Collect the hardware performance parameters of each drone in the drone swarm; in the collaborative tasks of the drone swarm, the hardware performance parameters of the drones are crucial data. For example, the hardware performance such as the endurance and flight speed of the drone affects the efficiency of task execution, while the sensor accuracy is related to the accuracy of the collected data. Accurately collecting the performance parameters of each agent can better plan and allocate tasks.
[0073] And collect task data, including the goals and requirements of the drone tasks; in the collaborative tasks of the drone swarm, the detailed information of the task goals is crucial. Information such as the location and type of the goal determines the exploration and processing methods that the agents need to adopt. The priority of the goal helps the agents reasonably arrange the task sequence under limited resources and time conditions.
[0074] Alliance marking: First, the COLLAB-MARL framework enhances the basic multi-agent reinforcement learning algorithm by introducing alliance marking, enriching the observation space of the agents to include alliance-specific meta-information and allowing policy adjustment; second, structural entropy quantifies the coordination among agents by analyzing the interaction graph that changes over time, providing a clear measure of collaboration effectiveness.
[0075] Among them, as Figure 2 shown, the main goal of the COLLAB-MARL framework is to enhance the coordination and cooperation among agents in the same alliance in a heterogeneous reward-sharing environment. By incorporating alliance-specific information into the state representation of each agent, COLLAB-MARL can achieve more complex internal alliance collaboration, thereby improving the collective performance.
[0076] The COLLAB-MARL framework includes:
[0077] Alliance marking technology to achieve enhanced observation;
[0078] Through the policy framework, perform scenario changes;
[0079] Then utilize model-free learning.
[0080] Specifically, in a multi-agent system, the agents operate in the same environment and share a common policy network in scenarios where cooperation among subsets of agents (alliances) is required to optimize a heterogeneous reward function. However, the agents cannot inherently perceive their alliance membership from the environmental observations because the alliance structure is usually not embedded in the environment. To solve this problem, the present invention introduces alliance marking technology to attach an alliance-specific identifier to the observation vector of each agent, including:
[0081] Given a multi-agent system: S=(V, E);
[0082] where, V={v 1, v 2 ,..., v n}, which represents the set of agents, E represents the set of agent interactions, and C = {C 1 , C 2 ,..., C k} is the set of coalitions in S; each agent v ∈ V belongs to a certain coalition C i ∈ C;
[0083] The integer coalition label of agent v is defined as:
[0084]
[0085] where the scalar label represents the coalition to which agent v belongs and is directly attached to the observation vector o v of the agent to form the augmented observation
[0086] By including in the agent's observation, the policy network identifies coalition membership as part of the agent state representation, enabling the shared policy to learn coalition-specific behaviors and improve coordination within the same coalition.
[0087] This approach eliminates the need for agents to perceive coalition membership through environmental cues, thus avoiding modifications to the environment or sensor models.
[0088] The policy of agent v depends on the augmented observation where a v represents the action taken by agent v and θ represents the shared policy parameters.
[0089] Specifically, the policy framework uses different architectures according to different algorithms.
[0090] The policy framework of the present invention is compatible with various benchmark methods, such as MAPPO and MASAC, etc. Different architectures (e.g., shared or separate policies) are used according to different algorithms to adapt to each method.
[0091] The policy framework uses a shared network to process the augmented observation combining the original observation o v with coalition information; the policy will map to an action distribution, parameterized by the mean μ v and variance σ v :
[0092]
[0093] where is a normal distribution;
[0094] The policy network is implemented as a multi-layer perceptron MLP with L hidden layers:
[0095]
[0096] where φ is the activation function; W (l) is the weight matrix of the l-th layer, is the output of the (l - 1)-th layer, is the output of the l-th layer, and b (l) is the bias vector of the l-th layer;
[0097] The value function is modeled by a critic network, including centralized or decentralized ones; a centralized critic is used for the actor-critic method, and a decentralized critic is used for the value-based method. Training uses batches of collected observations, actions, and rewards; for the actor-critic method, proximal policy optimization is applied, and generalized advantage estimation is used for advantage calculation; for the value-based method, a tailored loss function is minimized to improve value prediction;
[0098] This framework achieves coordinated learning by sharing parameters and supports multiple multi-agent reinforcement learning algorithms, enhancing scalability and generality in cooperative tasks.
[0099] Specifically, in a multi-agent cooperation system, the coalition marking technology provides clear role recognition and task division for agent cooperation. This technology assigns specific coalition marks to each agent, enabling it to quickly identify cooperation objects and task division, thus effectively reducing problems such as unreasonable task assignment and uncoordinated actions. In the face of dynamically changing task distributions and environments, a more flexible adjustment mechanism is needed. Here, a model-agnostic learning-based strategy is introduced, aiming to dynamically adjust coalition marks through a meta-learning mechanism to improve the adaptability and cooperation efficiency of agents. Model-agnostic learning refers to a learning strategy where the structure and parameter update rules of the model do not depend on a specific task. This method makes the learning process more flexible and can adapt to various different task distributions and environmental changes.
[0100] The model-agnostic learning process includes: considering a task distribution p(j), which includes multiple related task distributions j d ; each task distribution j d corresponds to a specific coalition mark; this process lays the foundation for subsequent meta-learning updates. By sampling multiple task distributions, agents can obtain richer experiences, thereby quickly finding appropriate adjustment strategies in unknown task distributions.
[0101] For each task, update is performed in a meta - learning manner. Before the agent encounters the new task distribution, it first obtains a good initial parameter configuration by training on related tasks;
[0102] By introducing model - agnostic learning, under the new task distribution j d adjust the coalition markers to enhance the adaptability of the agent. The core of meta - learning lies in improving learning efficiency through shared experiences between tasks. This enables the agent to quickly utilize existing knowledge for adjustment when facing new tasks.
[0103] After obtaining the updated parameters, further dynamically adjust the coalition markers of the agent; perform fine - tuning for the new target task;
[0104] At this stage, by fine - tuning its parameters, the agent can quickly adapt to the new task requirements. The process of dynamic adjustment is not limited to parameter updates. The coalition markers themselves can also be re - evaluated and adjusted according to the current task requirements. To achieve a more efficient adjustment process, the agent can incorporate feedback information from the environment during fine - tuning. This means that while the agent is performing a task, it can monitor the effects of its actions in real - time and then make adaptive adjustments to the coalition markers. This mechanism provides the agent with greater flexibility, enabling it to maintain efficient cooperation in complex and ever - changing environments.
[0105] After completing the dynamic adjustment of the coalition markers, update the global model parameters by optimizing the performance of each task. In this way, the agent can not only quickly adapt to the new task distribution but also gradually improve its cooperation performance in multiple rounds of iteration. In this process, the training objective of the agent is to maximize its performance in the new task while ensuring its transfer ability between different task distributions. This goal - setting not only promotes the agent's adaptability to new tasks but also enhances its flexibility when facing unknown tasks.
[0106] The objective of the present invention is to enable the agent to quickly adjust its coalition markers when encountering a new task distribution, ensuring effective cooperation in the new environment.
[0107] As an optimized solution of the above - mentioned embodiment, adopt a hierarchical decision - making framework SIDM based on the principle of structural information, use the principle of structural information to abstractly process environmental and task information, assist the UAV swarm in refining roles and reasonably allocating tasks, and enhance the pertinence of cooperation; during the cooperation process, use structural entropy to monitor the cooperation status in real - time. Once an abnormality is detected, dynamically adjust the cooperation strategy by combining coalition markers and abstract information.
[0108] Such as Figure 3As shown, the hierarchical decision-making framework SIDM framework based on the principle of structural information includes: an adaptive abstraction mechanism, directed structural entropy, and role-based learning.
[0109] Specifically, to effectively process high-dimensional and noisy environmental information, an adaptive abstraction mechanism based on the principle of structural information is proposed. The core of this mechanism is to classify similar states and actions to generate an abstract representation. The adaptive abstraction mechanism includes:
[0110] 1.1 Data sampling, at each time step t, a batch of data of size n is drawn from the replay buffer B, which covers the current state S t , action A t , subsequent state S t+1 and reward R t ;
[0111] 1.2 Encoding conversion, using an encoder-decoder architecture to convert the original high-dimensional variables into low-dimensional representations;
[0112] The formula is S' t = f s (S t ), S' t+1 = f s (S t+1 ), A' t = f a (A t );
[0113] f s and f a are the state encoder and action encoder respectively. At the same time, two decoders are designed. The state decoder d s uses cross-entropy inverse target to predict the action between adjacent states, and the action decoder d a reconstructs the next state based on the action and state representations;
[0114] 1.3 Graph construction, for a pair of different states s i ∈ S and s j ∈ S in the state S, use Pearson correlation analysis to measure their feature similarity C(s i , s j ), the higher the absolute value of C(s i , s j ), the stronger the state similarity; each state is regarded as a vertex and connected according to the similarity to form a complete weighted undirected state graph G s ;
[0115]
[0116] where, μ si , μ sj is the mean, σsi , σ sj is the variance. C(s i , s j ) The higher the absolute value, the stronger the state similarity.
[0117] 1.4 Optimal partitioning and aggregation to determine the key to the optimal partitioning structure of the sparse state graph G s .
[0118] Introduce the stretching (n st ) and compression (n cp ) operators of the HSCE algorithm to optimize the initial coding tree T s , reduce the structural entropy and increase the height of the tree. Each iteration traverses the set of tree nodes at the same layer, and selects the set of nodes that can reduce the structural entropy to the greatest extent for the "stretching-compression" operation. When the tree height reaches or no set of nodes satisfies , end the iteration and output the optimal coding tree
[0119] In , design an aggregation function according to the node uncertainty to deduce the node representation. For the leaf node If V ν = s, its representation is h ν = s; for the non-leaf node α, calculate the representation after normalizing the weights of the child nodes with the softmax function:
[0120]
[0121] The representation h λ of the child nodes of the root node λ is defined as the abstract state The original environmental state s t is mapped to the abstract state through the dot product with the abstract representation in Z s
[0122]
[0123] Similarly, the abstract action variable Z a can be defined.
[0124] To overcome the limitation of the undirected constraint of the current structural information principle, the present invention proposes a definition and optimization method for the high-dimensional structural entropy of a directed graph to accurately capture the directed transitions between abstract states in reinforcement learning.
[0125] Specifically, the directed structural entropy includes:
[0126] 2.1 Graph structure adjustment: Given a directed graph G dir = (V, E dir , W dir) Adjust its structure so that there is a directed path between any pair of vertices, and normalize the weights so that the weighted out-degree sum of each vertex is 1; the adjusted graph G' dir There exists a unique stationary distribution π s , which is equivalent to the adjacency matrix A' dir The eigenvector corresponding to the largest eigenvalue 1;
[0127] 2.2 Definition of directed structure entropy: Calculate the vertex stationary distribution π in G' dir , and define the one-dimensional directed structure entropy: s
[0128] H 1 (G' dir ) = -∑ v∈V π s (v)·logπ s (v);
[0129] Based on π s Adjust the relevant terms and define the allocation entropy of the non-root node α in the coding tree T dir
[0130] Furthermore, define the K-dimensional directed structure entropy of G' dir :
[0131]
[0132] 2.3 Entropy optimization process: Based on the merging η mg and combination η cb operators of the deDoc algorithm, optimize the directed structure entropy Determine the optimal tree structure coding strategy
[0133] In a fully cooperative multi-agent scenario with a discrete action space, using the previous abstraction mechanism, a two-layer role-based learning method is proposed to promote effective cooperation among multi-agents.
[0134] Specifically, role-based learning includes:
[0135] 3.1 Definition of role set: Define the abstract action set Z a as the role set Ψ, that is, each role ρ i ∈Ψ corresponds to an action subspace Through this definition method, the joint action space of multi-agents is partitioned, and each role is responsible for a specific action subspace, which helps to simplify the complexity of multi-agent cooperation.
[0136] 3.2 Learning process: At each time step t, for each agent n i , the high-level policy Select a role ρ from the role set Ψ j and its corresponding action subspace A j , and complete the role assignment; where τ i is the state-action history of agent n i . The high-level policy selects a suitable role according to the historical experience of the agent; then, the low-level role policy outputs an action according to the global state s t ; The individual actions of all agents together form the joint action a . After acting on the environment, the environment returns the joint reward r t and the subsequent global state s t +1; When training the low-level role policy, a multi-agent reinforcement learning algorithm is adopted, and agents with the same role use a common policy network; this can improve the learning efficiency and promote the cooperation consistency among agents with the same role. The relevant training loss is denoted as L t , and the low-level role policy is optimized by minimizing this loss. marl
[0137] 3.3 Update process: At the regular update interval t up , sample a batch of data from the replay buffer B to construct the variables S t , A t , S t+1 and R t ; Use the action abstraction mechanism to redefine the sets Z s and Ψ; The specific operations include constructing an action graph, filtering edges, and generating an optimal coding tree to update the abstract action set Z a , that is, the role set Ψ; After that, calculate the losses L de and L marl ; L de is used to optimize the relevant parameters in the action abstraction process, and L marl is used to optimize the hierarchical policy; By minimizing these two losses, the parameters of the hierarchical policy are continuously adjusted to enable the multi-agent system to better adapt to the environment and improve the cooperation performance.
[0138] The multi-agent collaborative optimization method based on coalition marking and structural information proposed by the present invention aims at the problem of chaotic cooperation faced by multi-agent collaborative systems in complex task execution, especially in the scenarios of multi-UAV collaborative exploration and smart city management, and has outstanding optimization efficiency. Through the coalition marking technology, agents can accurately identify their own roles, clearly define the task division of labor, and greatly improve the task execution efficiency. By using the principle of structural information, not only can the environment and tasks be abstractly analyzed, the information interaction and action coordination mechanism be optimized, but also the roles of agents can be mined and defined.
[0139] Meanwhile, by combining the alliance mark and structural information, agents can better coordinate their roles with each other. In the face of diverse task requirements, each agent can clarify its role in collaboration based on its own characteristics and task requirements, and flexibly switch roles according to the actual situation, enhancing the dynamic adaptability of the system, constructing a more efficient and reliable multi-agent collaborative system, and is expected to promote the wide application and development of multi-agent technology in more fields, especially playing an important role in the intelligent management and services of smart cities.
[0140] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A UAV swarm collaborative optimization method based on alliance labeling and structural information, characterized in that: Includes steps: First, the UAV swarm is divided and marked through alliance marking technology, and observation vectors are integrated to make the UAV swarm clear about the collaborative relationship and task division at the initial stage; Then, the structural information principle is used to abstract the environment and task information, assisting the drone swarm to refine roles, reasonably allocate tasks, and enhance the targetedness of collaboration; During the collaboration process, the collaboration status is monitored in real time with the help of structural entropy. Once an abnormality is found, the collaboration strategy is dynamically adjusted in combination with alliance tags and abstract information.
2. The method for collaborative optimization of drone swarms based on alliance labeling and structural information according to claim 1 is characterized in that: The UAV swarm is divided and marked through alliance marking technology, and the observation vector is integrated, so that the UAV swarm can clarify the cooperation relationship and task division at the initial stage, including the following steps: Data collection: Collect hardware performance parameters of each drone in the drone group; and collect mission data, including the goals and requirements of the drone mission; Coalition Labeling: First, the COLLAB-MARL framework enhances the basic multi-agent reinforcement learning algorithm by introducing coalition labeling, enriching the observation space of agents to include coalition-specific meta-information and allowing policy adjustment; second, structural entropy provides an explicit measure of collaborative effectiveness by analyzing the interaction graph over time to quantify the coordination between agents.
3. The method for collaborative optimization of drone swarms based on alliance labeling and structural information according to claim 2 is characterized in that: The main goal of the COLLAB-MARL framework is to enhance the coordination and cooperation among agents in the same coalition in a heterogeneous reward sharing environment; the COLLAB-MARL framework includes: Alliance tagging technology to achieve enhanced observation; Make scenario changes through the strategic framework; Reuse model-independent learning.
4. The method for collaborative optimization of drone swarms based on alliance labeling and structural information according to claim 3 is characterized in that: The coalition labeling technique appends a coalition-specific identifier to each agent’s observation vector, consisting of: Given a multi-agent system: S = (V, E); Where V = {v1,v2,...,v n } represents the set of agents, E represents the set of agent interactions, C = {C1, C2, ..., C k } is the set of alliances in S; each agent v∈V belongs to a certain alliance C i ∈C; The integer coalition label of agent v is defined as: Among them, the scalar label represents the coalition to which agent v belongs and is directly attached to the agent's observation vector o v On top, form enhanced observation By including in the agent's observations The policy network recognizes coalition membership as part of the agent state representation, enabling a shared policy to learn behaviors specific to each coalition and improve coordination within the same coalition. Agent v’s strategy Depends on enhanced observations where a v represents the action taken by the agent v and θ represents the shared policy parameters.
5. The method for collaborative optimization of drone swarms based on alliance labeling and structural information according to claim 4 is characterized in that: The strategy framework uses different architectures depending on the specific algorithm.
6. The method for collaborative optimization of drone swarms based on alliance labeling and structural information according to claim 5 is characterized in that: Strategic framework uses shared network processing to enhance observation The original observation o v Combined with alliance information; strategy will Mapped to an action distribution, parameterized by mean μ v and variance σ v : in, is a normal distribution; The policy network is implemented as a multilayer perceptron MLP with L hidden layers: Among them, φ is the activation function; W (l) is the weight matrix of the lth layer, is the output of the l-1th layer, is the output of layer l, b (l) is the bias vector of the lth layer; The value function is modeled by a critic network, either centralized or decentralized; a centralized critic is used for actor-critic methods and a decentralized critic is used for value-based methods, and training is done using collected batches of observations, actions, and rewards; for actor-critic methods, proximal policy optimization is applied, and advantage computation is done using generalized advantage estimation; for value-based methods, a tailored loss function is minimized to improve value predictions.
7. The method for collaborative optimization of drone swarms based on alliance labeling and structural information according to claim 3 is characterized in that: The model-independent learning process includes: considering a task distribution p(j), which includes multiple related task distributions j d ; Each task distribution j d Corresponds to a specific alliance mark; For each task, meta-learning is used to update the agent. Before the agent is exposed to a new task distribution, it first obtains a good initial parameter configuration by training on related tasks. After obtaining the updated parameters, the alliance tags of the agents are further dynamically adjusted; new target tasks are fine-tuned; After completing the dynamic adjustment of the alliance labels, the global model parameters are updated by optimizing the performance of each task.
8. The method for collaborative optimization of drone swarms based on alliance labeling and structural information according to claim 1 is characterized in that: The hierarchical decision-making framework SIDM based on the structural information principle is adopted to abstract the environment and task information using the structural information principle, assist the drone group to refine roles, reasonably allocate tasks, and enhance the targetedness of collaboration; During the collaboration process, the collaboration status is monitored in real time with the help of structural entropy. Once an abnormality is found, the collaboration strategy is dynamically adjusted in combination with alliance tags and abstract information.
9. The method for collaborative optimization of drone swarms based on alliance labeling and structural information according to claim 8, characterized in that: The hierarchical decision-making framework SIDM framework based on the structural information principle includes: adaptive abstraction mechanism, directed structural entropy and role-based learning.
10. The method for collaborative optimization of drone swarms based on alliance labeling and structural information according to claim 9, characterized in that: Adaptive abstraction mechanisms include: 1.1 Data sampling, at each time step t, a batch of data of size n is extracted from the playback buffer B, which covers the current state S t 、Action A t , subsequent state S t+1 and reward R t ; 1.2 Encoding conversion, using the encoder-decoder architecture to convert the original high-dimensional variables into low-dimensional representations; The formula is S' t =f s (S t ),S' t+1 =f s (S t+1 ),A' t =f a (A t ); f s and f a They are state encoder and action encoder respectively, and two decoders are designed at the same time, the state decoder d s Use cross entropy inverse target to predict actions between adjacent states, action decoder d a Reconstruct the next state based on the action and state representation; 1.3 Graph Construction, for a pair of different states s in state S i ∈S and s j ∈S, using Pearson correlation analysis to measure its feature similarity C(s i ,s j ), C(s i ,s j )The higher the absolute value, the stronger the state similarity; each state is regarded as a vertex and connected according to the similarity to form a complete weighted undirected state graph G s ; 1.4 Optimal Partitioning and Aggregation, Determining the Sparse State Graph G s The key to the optimal partitioning structure; Directed structural entropy, including: 2.1 Graph structure adjustment: Given a directed graph G dir =(V,E dir ,W dir ), adjust its structure so that there is a directed path between any pair of vertices, and normalize the weights so that the sum of the weighted out-degree of each vertex is 1; the adjusted graph G' dir There exists a unique stationary distribution π s , which is equivalent to the adjacency matrix A' dir The eigenvector corresponding to the maximum eigenvalue 1; 2.2 Definition of directed structural entropy: In G' dir Calculate the vertex stationary distribution π s , define one-dimensional directed structural entropy: H 1 (G' dir )=-∑ v∈V π s (v) logπ s (v); Based on π s Adjust the relevant items and define the coding tree T dir The distribution entropy of non-root node α in Then define G' dir The K-dimensional directed structural entropy of: 2.3 Entropy optimization process: Merging η based on the deDoc algorithm mg and combination η cb Operator, optimizes the directed structural entropy Determining the optimal tree encoding strategy Role-based learning, including: 3.1 Role set definition: define the abstract action set Z a is the role set Ψ, that is, each role ρ i ∈Ψ corresponds to an action subspace 3.2 Learning process: At each time step t, for each agent n i , high-level strategy Select a role ρ from the role set Ψ j and its corresponding action subspace A j , complete the role assignment; here τ i is an agent n i The state-action history of the agent, the high-level strategy selects the appropriate role based on the historical experience of the agent; then, the low-level role strategy According to the global state s t Output an action The individual actions of all agents together form a joint action a t , after acting on the environment, the environment returns the joint reward r t and the subsequent global state s t +1; when training low-level role strategies, a multi-agent reinforcement learning algorithm is used, and agents of the same role use a common strategy network; the relevant training loss is recorded as L marl , optimize the low-level role policy by minimizing this loss; 3.3 Update process: At regular update intervals t up , sample a batch of data from the playback buffer B and construct the variable S t , A t , S t+1 and R t ; Redefine the set Z using the action abstraction mechanism s and Ψ; the specific operations include constructing an action graph, filtering edges, and generating an optimal coding tree to update the abstract action set Z a , that is, the role set Ψ; then, calculate the loss L de and L marl ; L de Used to optimize the relevant parameters in the action abstraction process, L marl Used to optimize tiering strategies; By minimizing these two losses, the parameters of the stratification strategy are continuously adjusted.