A distributed-based self-healing method and device for a centerless cluster network, equipment and medium

CN122742013APending Publication Date: 2026-09-11AEROSPACE YISHU (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610972052.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0007]通过采用上述技术方案,对无人机集群中各无人机节点的邻居节点进行心跳信息和监测包交付率的监听来确定可疑邻居节点,并对可疑邻居节点的已知邻居节点进行交叉验证以判定故障节点,相较于现有技术中单一节点独立判断的方式,能够有效避免因局部通信干扰导致的误判,提高了故障检测的准确性;当故障状态导致无人机集群拓扑破裂形成目标子群时,通过收集目标子群中各健康无人机节点根据自身物理状态参数计算得到的综合评分,并利用分布式一致性算法选举出子群领航者,解决了现有技术中集群分裂后缺乏协调节点导致修复效率低下的问题,在子群领航者的协调下,通过故障节点的邻居节点基于刚性图理论执行局部拓扑修复来恢复目标子群的内部连通性,避免了全局拓扑重构带来的巨大通信开销,同时由子群领航者通过自适应扩展发现其他目标子群并触发合并得到重构后的集群,有效解决了现有技术中多个子群无法自主重组的问题;当故障节点为关键角色时,基于分布式共识算法从重构后的集群中确定目标健康无人机节点作为接替节点来执行关键角色的任务,保证了集群功能的连续性;通过基于分级递归自愈策略引导重构后的集群从网络一致性恢复至物理一致性以恢复预设的编队构型,实现了从通信拓扑到物理编队的完整自愈闭环,相较于现有技术仅关注网络连通性恢复的方案,能够在保持集群任务执行能力的前提下完成完整的自愈过程,提升了无人机集群在复杂环境下的自主性和任务持续执行能力

Benefits of technology

1、对无人机集群中各无人机节点的邻居节点进行心跳信息和监测包交付率的监听来确定可疑邻居节点,并对可疑邻居节点的已知邻居节点进行交叉验证以判定故障节点,相较于现有技术中单一节点独立判断的方式,能够有效避免因局部通信干扰导致的误判,提高了故障检测的准确性;当故障状态导致无人机集群拓扑破裂形成目标子群时,通过收集目标子群中各健康无人机节点根据自身物理状态参数计算得到的综合评分,并利用分布式一致性算法选举出子群领航者,解决了现有技术中集群分裂后缺乏协调节点导致修复效率低下的问题,在子群领航者的协调下,通过故障节点的邻居节点基于刚性图理论执行局部拓扑修复来恢复目标子群的内部连通性,避免了全局拓扑重构带来的巨大通信开销,同时由子群领航者通过自适应扩展发现其他目标子群并触发合并得到重构后的集群,有效解决了现有技术中多个子群无法自主重组的问题;当故障节点为关键角色时,基于分布式共识算法从重构后的集群中确定目标健康无人机节点作为接替节点来执行关键角色的任务,保证了集群功能的连续性;通过基于分级递归自愈策略引导重构后的集群从网络一致性恢复至物理一致性以恢复预设的编队构型,实现了从通信拓扑到物理编队的完整自愈闭环,相较于现有技术仅关注网络连通性恢复的方案,能够在保持集群任务执行能力的前提下完成完整的自愈过程,提升了无人机集群在复杂环境下的自主性和任务持续执行能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122742013A_ABST
    Figure CN122742013A_ABST
Patent Text Reader

Abstract

A kind of centerless cluster network self-healing method, device, equipment and medium based on distribution, it is related to unmanned aerial vehicle network communication technical field.The neighbor node of each unmanned aerial vehicle node in unmanned aerial vehicle cluster is verified, and the neighbor node in fault state is determined as fault node;When unmanned aerial vehicle cluster topology breaks into target subgroup, the comprehensive score of each healthy unmanned aerial vehicle node in target subgroup is collected, and the comprehensive score is elected to determine subgroup leader;Under the coordination of subgroup leader, local topology repair and merge are carried out by the neighbor node of fault node, and the reconstructed cluster is obtained;When fault node is a key role, determine target healthy unmanned aerial vehicle node from the reconstructed cluster as replacement node to execute the task of key role;After confirming that network topology reconstruction and role replacement are completed, guide the reconstructed cluster to recover the preset formation configuration.Implement the method, improve the autonomy and task continuous execution capability of unmanned aerial vehicle cluster in complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of unmanned aerial vehicle (UAV) network communication technology, specifically to a self-healing method, apparatus, device, and medium based on a distributed, decentralized cluster network. Background Technology

[0002] With the rapid development of drone technology, drone swarms have demonstrated enormous application potential in complex mission scenarios such as collaborative reconnaissance, area search, and emergency rescue. Their high degree of coordination and robustness relies on real-time and reliable communication network connectivity between nodes within the swarm. During mission execution, swarm networks often face threats such as severe electromagnetic interference or physical damage, which can easily lead to communication link interruptions and network topology disruptions, causing the swarm to split into multiple isolated subnets that cannot communicate with each other. Ensuring that the remaining healthy nodes can autonomously and quickly rebuild network connectivity and restore swarm coordination after a network topology disruption is crucial for guaranteeing the security of the swarm system and the continuity of mission execution.

[0003] To address the self-healing problem following network topology breaks, existing technologies often employ distributed topology repair methods. For example, they utilize periodic heartbeat signals between neighboring nodes to monitor communication status, and when a link interruption is detected, they adaptively adjust the communication radius or deploy relay nodes to enhance the network's connectivity robustness.

[0004] However, in practical tasks, nodes that lose connection due to damage or interference often play critical roles, such as navigators. The absence of these roles directly leads to the interruption of collaborative tasks. Existing topology repair methods only focus on re-establishing link connectivity at the communication network level, neglecting the issue of autonomous succession after the loss of critical roles. This results in the cluster remaining in a disordered state and unable to continue collaboratively executing the intended tasks even after topology reconstruction at the network level due to the lack of role succession. Summary of the Invention

[0005] This application provides a self-healing method, apparatus, device, and medium for a distributed, decentralized cluster network. Compared with existing technologies that only focus on network connectivity restoration, this method can complete the entire self-healing process while maintaining the cluster's task execution capability, significantly improving the autonomy and continuous task execution capability of UAV clusters in complex environments.

[0006] Firstly, this application provides a distributed, decentralized cluster network self-healing method applied to a drone swarm. The method includes: monitoring the heartbeat information and monitoring packet delivery rate of each drone node's neighbor nodes in the drone swarm; identifying suspicious neighbor nodes based on the heartbeat information and monitoring packet delivery rate; performing cross-validation on the known neighbor nodes of the suspicious neighbor nodes; determining whether the suspicious neighbor nodes are in a faulty state based on the validation results; and identifying the suspicious neighbor nodes determined to be in a faulty state as faulty nodes; when the faulty state causes the drone swarm topology to break, forming at least one target subgroup, collecting the comprehensive scores calculated by each healthy drone node in the target subgroup based on its own physical state parameters, and using a distributed consensus algorithm to evaluate the scores of multiple comprehensive scores. The system uses a combined score to elect the healthy drone node with the highest overall score as the subgroup leader. Under the coordination of the subgroup leader, the neighboring nodes of the failed node perform local topology repair based on rigid graph theory to restore the internal connectivity of the target subgroup. The subgroup leader then adaptively expands to discover other target subgroups and triggers merging, resulting in a reconstructed cluster. When the failed node plays a critical role, a target healthy drone node is selected from the reconstructed cluster as the successor node based on a distributed consensus algorithm. This healthy drone node then takes over the tasks of the critical role. After confirming that the network topology reconstruction and role succession are complete, a hierarchical recursive self-healing strategy guides the reconstructed cluster to recover from network consistency to physical consistency, thereby restoring the preset formation configuration.

[0007] By employing the above technical solution, suspicious neighbor nodes are identified by monitoring the heartbeat information and packet delivery rate of the neighbor nodes of each drone node in the drone swarm. Cross-validation of the known neighbor nodes of the suspicious neighbor nodes is then performed to determine the faulty node. Compared to the existing method of independent judgment by a single node, this effectively avoids misjudgments caused by local communication interference, improving the accuracy of fault detection. When a fault condition causes the drone swarm topology to break, forming a target subgroup, a distributed consensus algorithm is used to elect a subgroup leader by collecting the comprehensive score calculated by each healthy drone node in the target subgroup based on its own physical state parameters. This solves the problem of low repair efficiency caused by the lack of a coordinating node after swarm splitting in existing technologies. Under the coordination of the subgroup leader, the internal structure of the target subgroup is restored by the neighbor nodes of the faulty node performing local topology repair based on rigid graph theory. Connectivity is improved, avoiding the huge communication overhead caused by global topology reconstruction. Meanwhile, the subgroup leader discovers other target subgroups through adaptive expansion and triggers merging to obtain the reconstructed cluster, effectively solving the problem that multiple subgroups cannot autonomously reorganize in existing technologies. When the faulty node plays a critical role, a target healthy UAV node is selected from the reconstructed cluster based on a distributed consensus algorithm to take over the critical role's tasks, ensuring the continuity of cluster functions. By guiding the reconstructed cluster from network consistency to physical consistency based on a hierarchical recursive self-healing strategy to restore the preset formation configuration, a complete self-healing closed loop from communication topology to physical formation is achieved. Compared with existing technologies that only focus on network connectivity restoration, this approach can complete the entire self-healing process while maintaining the cluster's task execution capability, improving the autonomy and continuous task execution capability of the UAV cluster in complex environments.

[0008] Optionally, the fault state includes physical damage state and local interference state. Cross-validation is performed on the known neighbor nodes of the suspected neighbor node, and the failure state of the suspected neighbor node is determined based on the validation results. Specifically, this includes: obtaining observation nodes from the UAV swarm; if a suspected neighbor node exists among the neighbor nodes of the observation node, a cross-validation request is sent to the known neighbor nodes of the suspected neighbor node; collecting the observation dataset for the suspected neighbor node fed back by the known neighbor nodes, wherein the observation dataset includes the historical communication quality, current heartbeat reception status, and collaborative sensing data between the known neighbor nodes and the suspected neighbor node; if all current heartbeat reception statuses in the observation dataset are not received, and no collaborative sensing data from the suspected neighbor node is received within a preset waiting time, the suspected neighbor node is determined to be in a physical damage state; if some known neighbor nodes in the observation dataset can receive the heartbeat information of the suspected neighbor node, or the physical motion trajectory in the collaborative sensing data conforms to the preset formation motion constraints and the historical communication quality is lower than the preset communication threshold, the suspected neighbor node is determined to be in a local interference state.

[0009] By adopting the above technical solution, cross-validation requests are sent from observation nodes in the UAV swarm to the known neighbor nodes of suspected neighbor nodes. This avoids misjudgments caused by a single node's own communication failure or local environmental interference. The observation dataset fed back by known neighbor nodes is collected. When all current heartbeat reception states in the observation dataset are not received and no collaborative sensing data from suspected neighbor nodes is received within a preset waiting time, it is determined to be in a physical damage state. The determination method based on multi-node consistency observation can accurately identify permanent failures of nodes. When some known neighbor nodes in the observation dataset can receive heartbeat information from suspected neighbor nodes, or when the physical movement trajectory in the collaborative sensing data conforms to preset formation movement constraints and the historical communication quality is lower than a preset communication threshold, it is determined to be in a local interference state. This determination can accurately identify temporary communication interruptions caused by electromagnetic interference or obstacle obstruction. By distinguishing between physical damage state and local interference state, differentiated handling measures can be taken for different fault types. Nodes in the physical damage state can be directly replaced and topology reconstructed, while nodes in the local interference state can restore communication by adjusting communication parameters, avoiding resource waste and unnecessary topology adjustments.

[0010] Optionally, a comprehensive score is collected from each healthy UAV node in the target subgroup, calculated based on its own physical state parameters. Specifically, this includes: obtaining the physical state parameters corresponding to the healthy nodes from the target subgroup, including current remaining energy, node degree, average relative distance to other healthy UAV nodes in the target subgroup, and average signal-to-noise ratio; normalizing the physical state parameters using a preset normalization function to obtain the corresponding energy score, topology score, position score, and channel score, and constructing a score feature vector; calculating the local information entropy of each score feature based on the score feature vectors of all healthy UAV nodes in the target subgroup, and determining the initial weight of each score feature based on the local information entropy; performing a nonlinear mapping on the energy score, topology score, position score, channel score, and initial weights of the healthy nodes based on a preset state-variable-weighting activation function to generate dynamic weights that match the current physical state of the healthy nodes; and using the dynamic weights to perform a weighted summation of the energy score, topology score, position score, and channel score to obtain the comprehensive score of the healthy UAV node.

[0011] By adopting the above technical solution, physical state parameters corresponding to healthy nodes in the target subgroup are obtained, including current remaining energy, node degree, average relative distance to other healthy UAV nodes in the target subgroup, and average signal-to-noise ratio. These physical state parameters are normalized using a preset normalization function to obtain energy score, topology score, location score, and channel score, and a score feature vector is constructed. This eliminates the influence of differences in physical dimensions. Based on the score feature vectors of all healthy UAV nodes in the target subgroup, the local information entropy of each score feature is calculated, and initial weights are determined. A preset state-weighted excitation function is used to perform a nonlinear mapping of the various scores and initial weights of the healthy nodes to generate dynamic weights. The dynamic weights are then used to weighted sum the various scores to obtain the comprehensive score of the healthy UAV node. This ensures that the elected subgroup leader is in an optimal state in multiple dimensions such as energy reserves, topology location, and communication quality, improving the efficiency and success rate of subsequent topology repair and subgroup merging. This effectively solves the problem of slow convergence in the repair process caused by the selection of a non-optimal coordinating node in existing technologies.

[0012] Optionally, a distributed consensus algorithm is used to elect the healthy drone node with the highest comprehensive score from multiple comprehensive scores, which is then selected as the subgroup leader. Specifically, this involves: each healthy drone node broadcasting an election message containing its own identifier and comprehensive score within the target subgroup; each healthy drone node collecting the election messages broadcast by other healthy drone nodes within the target subgroup and constructing a score mapping table locally; each healthy drone node performing multiple rounds of interactive verification on the data in the score mapping table based on a distributed consensus protocol until all healthy drone nodes in the target subgroup reach a consensus on the score mapping; and each healthy drone node determining the healthy drone node with the highest comprehensive score as the subgroup leader based on the consensus-reached score mapping table.

[0013] By adopting the above technical solution, each healthy drone node broadcasts an election message containing its own identifier and comprehensive score within the target subgroup, realizing a fully distributed information disclosure mechanism. Each healthy drone node collects the election messages broadcast by other healthy drone nodes within the target subgroup and builds a score mapping table locally, so that each node has a complete view of candidate information. Through multiple rounds of interactive verification of the data in the score mapping table by each healthy drone node based on a distributed consensus protocol, until all healthy drone nodes in the target subgroup reach a consensus on the score mapping table, the data inconsistency problem caused by message loss in a distributed environment is effectively solved. Each healthy drone node determines the healthy drone node with the highest comprehensive score as the subgroup leader based on the score mapping table that has reached a consensus. The entire election process does not require any centralized coordination and is achieved entirely through peer-to-peer communication and distributed consensus between nodes. This improves the autonomy and real-time performance of the drone cluster in quickly rebuilding the command system after topology failure, and effectively solves the problem of low repair efficiency caused by the lack of a unified coordination mechanism after cluster split in existing technologies.

[0014] Optionally, under the coordination of the subgroup leader, the neighboring nodes of the failed node perform local topology repair based on rigid graph theory to restore the internal connectivity of the target subgroup. The subgroup leader then adaptively expands to discover other target subgroups and triggers merging, resulting in a reconstructed cluster. Specifically, this includes: under the coordination of the subgroup leader, obtaining the set of neighboring nodes within the local topology association range of the failed node, and extracting a local subgraph composed of the neighboring node set; constructing a local minimum rigid graph in the local subgraph based on rigid graph theory, and merging the local minimum rigid graph into the global topology of the target subgroup to complete the local topology restoration of the subgroup's internal connectivity. Topology repair; the subgroup navigator periodically broadcasts subgroup probe messages. If no feedback is received from other subgroups within a preset probe time, the navigator controls the nodes in the target subgroup to adaptively increase their communication radius to scan for the existence of other target subgroups; or the subgroup navigator determines relay candidate nodes in the target subgroup, controls the relay candidate nodes to move to a preset adjacent subgroup location to be deployed as communication relays, and actively establishes cross-subgroup communication links; after establishing communication connections with other target subgroups, the subgroup navigator controls the exchange of member lists and status information with the navigators of other target subgroups, and performs subgroup merging based on distributed joint consensus to obtain the reconstructed cluster.

[0015] By adopting the above technical solution, under the coordination of the subgroup leader, a set of neighboring nodes within the local topology association range of the faulty node is obtained, and a local subgraph composed of the neighboring node set is extracted. Based on rigid graph theory, a local minimum rigid graph is constructed in the local subgraph and merged into the global topology of the target subgroup to complete the local topology repair of the connectivity within the subgroup. The subgroup leader periodically broadcasts subgroup probe messages. If no feedback is received from other subgroups within a preset probe time, the nodes in the target subgroup are controlled to adaptively increase their communication radius to scan for the existence of other target subgroups. Alternatively, the subgroup leader determines relay candidate nodes in the target subgroup and moves them to preset adjacent subgroup locations to deploy them as communication relays and actively establish cross-subgroup communication links. After establishing communication connections with other target subgroups, the subgroup leader exchanges member lists and status information with the leaders of other target subgroups and performs subgroup merging based on distributed joint consensus to obtain the reconstructed cluster. This distributed joint consensus method avoids the single point of failure risk and communication bottleneck in centralized merging, ensuring the reliability and efficiency of the merging process.

[0016] Optionally, when the faulty node plays a critical role, a target healthy drone node is determined from the reconstructed cluster based on a distributed consensus algorithm as the replacement node. The target healthy drone node is then controlled to take over the tasks of the critical role. Specifically, when the faulty node plays a critical role, the subgroup leader broadcasts the role replacement request within the subgroup. The role replacement request carries the hardware payload requirements and task attribute parameters of the critical role. Based on the current state parameters of each healthy drone node in the reconstructed cluster and the role replacement request, the corresponding bidding score is calculated, and a bidding response message carrying the bidding score is sent to the subgroup leader. The subgroup leader, based on distributed consensus, determines the target healthy drone node with the highest bidding score from all received bidding response messages as the replacement node and controls the target healthy drone node to take over the tasks of the critical role.

[0017] By adopting the above technical solution, when the faulty node plays a critical role, the sub-group leader broadcasts a role replacement request within the sub-group, carrying the hardware payload requirements and task attribute parameters of the critical role. This ensures that the replacement node possesses the necessary hardware conditions and capabilities to execute specific critical tasks, avoiding task failures or performance degradation due to mismatched capabilities of the replacement node. Based on the current state parameters and role replacement requirements of each healthy drone node in the reconstructed cluster, the corresponding bidding score is calculated, and a bidding response message carrying the bidding score is sent to the sub-group leader. The sub-group leader, based on distributed consensus, determines the target healthy drone node with the highest bidding score from all received bidding response messages as the replacement node and controls the target healthy drone node to take over the critical role's task. This deterministic selection mechanism based on distributed consensus not only ensures that the globally optimal node is selected but also ensures that all nodes in the cluster agree on the role replacement decision through the consensus process, avoiding chaos and resource waste caused by decision conflicts. It effectively solves the problem of poor functional continuity after the failure of a critical node in existing technologies, ensuring that the drone cluster can quickly recover its task execution capability when encountering a critical role node failure.

[0018] Optionally, based on the current state parameters and role succession requirements of each healthy drone node in the reconstructed cluster, the corresponding bidding score is calculated. Specifically, this includes: matching and verifying the hardware configuration of the healthy drone node to be bid on according to the hardware load requirements in the role succession requirements; the healthy drone node to be bid on can be any healthy drone node in the reconstructed cluster; if the matching verification fails, the bidding score of the healthy drone node to be bid on is determined to be zero; if the matching verification succeeds, the current state parameters of the healthy drone node to be bid on are obtained, and combined with the task attribute parameters in the role succession requirements, the target input variables are calculated. The target input variable parameters include remaining energy, relative proximity based on the task execution location, and the expected task load after superimposing the task attribute parameters. The process involves: 1) Mapping remaining energy, relative proximity, and expected task load to corresponding fuzzy sets based on a preset membership function to obtain the membership of the target input variable under different fuzzy states; 2) Performing fuzzy inference on each membership based on a preset fuzzy inference library to calculate the rule strength of each triggered rule. The preset fuzzy rule library contains multiple rules constructed based on conditional statements, and the rule strength is the minimum value of each membership in the corresponding rule condition; 3) Trunculating the output fuzzy set of the corresponding rule conclusion based on the rule strength, and performing a union operation on all truncated output fuzzy sets to obtain an aggregated output fuzzy set; 4) Defuzzifying the aggregated output fuzzy set using the centroid method, and using the abscissa of the geometric centroid of the aggregated output fuzzy set as the bidding score corresponding to the health drone node to be bid on.

[0019] By adopting the above technical solution, the hardware configuration of the bidding health drone node is matched and verified according to the hardware load requirements in the role succession requirements. If the matching verification fails, the bidding score of the bidding health drone node is directly determined to be zero. When the matching verification is successful, the current state parameters of the bidding health drone node are obtained and combined with the task attribute parameters in the role succession requirements to calculate the target input variable. Based on the preset membership function, the target input variable is mapped to the corresponding fuzzy set to obtain the membership degree of the target input variable in different fuzzy states. Based on the preset fuzzy inference library, fuzzy inference is performed on each membership degree to calculate the rule strength of each triggered rule. The fuzzy logic-based processing method is more advanced than traditional methods. The system's precise threshold judgment method can effectively handle the uncertainty and fuzziness of state parameters in the drone swarm environment, making the calculation of bidding scores more in line with the flexible decision-making needs in actual application scenarios. According to the rule strength, the output fuzzy set of the corresponding rule conclusion is truncated, and all the truncated output fuzzy sets are combined to obtain the aggregated output fuzzy set. Then, the centroid method is used to defuzzify the aggregated output fuzzy set, and the abscissa of the geometric centroid is used as the bidding score corresponding to the healthy drone node to be bid. This avoids the problem that extreme cases may be ignored due to simple linear weighting, so that the selected replacement node achieves optimal balance in multiple aspects, improving the stability and continuity of task execution after role replacement.

[0020] The second aspect of this application provides a distributed, decentralized cluster network self-healing device. The device is located within a drone swarm and further includes an acquisition unit, a fault detection unit, a reconstruction unit, a role succession unit, and a recovery unit. The acquisition unit listens to the heartbeat information and monitoring packet delivery rate of each drone node's neighbor nodes in the drone swarm, and identifies suspicious neighbor nodes based on these information. The fault detection unit performs cross-validation on the known neighbor nodes of the suspicious neighbor nodes, determines whether the suspicious neighbor node is in a fault state based on the validation results, and identifies the suspicious neighbor node in a fault state as a fault node. The reconstruction unit, when a fault state causes the drone swarm topology to break, forming at least one target subgroup, collects the comprehensive data calculated by each healthy drone node in the target subgroup based on its own physical state parameters. The system scores multiple drones and elects the healthiest drone node as the subgroup leader using a distributed consensus algorithm. Under the coordination of the subgroup leader, the neighboring nodes of the failed node perform local topology repair based on rigid graph theory to restore the internal connectivity of the target subgroup. The subgroup leader then adaptively expands to discover other target subgroups and triggers merging, resulting in a reconstructed cluster. A role succession unit, when the failed node is a key role, determines a target healthy drone node from the reconstructed cluster as a successor node based on a distributed consensus algorithm, and controls the target healthy drone node to take over the tasks of the key role. A recovery unit, after confirming the completion of network topology reconstruction and role succession, guides the reconstructed cluster to recover from network consistency to physical consistency based on a hierarchical recursive self-healing strategy, restoring the preset formation configuration.

[0021] In a third aspect, this application provides an electronic device including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory, causing the electronic device to perform any of the methods described above in this application.

[0022] In a fourth aspect, this application provides a computer-readable storage medium storing instructions that, when executed, perform any of the methods described above in this application.

[0023] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. By monitoring the heartbeat information and packet delivery rate of each drone node in the drone swarm, suspicious neighbor nodes are identified. Cross-validation of the known neighbors of these suspicious neighbors is then performed to determine the faulty node. Compared to the existing method of independent judgment by a single node, this effectively avoids misjudgments caused by local communication interference, improving the accuracy of fault detection. When a fault condition causes the drone swarm topology to break, forming a target subgroup, a comprehensive score calculated by each healthy drone node in the target subgroup based on its own physical state parameters is collected. A distributed consensus algorithm is then used to elect a subgroup leader. This solves the problem of low repair efficiency due to the lack of a coordinating node after swarm splitting in existing technologies. Under the coordination of the subgroup leader, local topology repair based on rigid graph theory is performed by the neighbor nodes of the faulty node to restore the internal connectivity of the target subgroup, avoiding... This approach avoids the significant communication overhead associated with global topology reconstruction. Furthermore, the subgroup leader adaptively expands to discover other target subgroups and triggers merging to obtain the reconstructed cluster, effectively solving the problem of multiple subgroups being unable to autonomously reorganize in existing technologies. When a faulty node plays a critical role, a healthy UAV node is selected from the reconstructed cluster based on a distributed consensus algorithm to serve as the replacement node and perform the critical role's tasks, ensuring the continuity of cluster functionality. By guiding the reconstructed cluster from network consistency to physical consistency through a hierarchical recursive self-healing strategy to restore the preset formation configuration, a complete self-healing closed loop from communication topology to physical formation is achieved. Compared to existing technologies that only focus on network connectivity restoration, this approach can complete the entire self-healing process while maintaining the cluster's task execution capabilities, enhancing the autonomy and continuous task execution capabilities of the UAV cluster in complex environments. Attached Figure Description

[0024] Figure 1 This is a flowchart illustrating a distributed, decentralized cluster network self-healing method provided in an embodiment of this application. Figure 2 This is a schematic diagram of the structure of a distributed, decentralized cluster network self-healing device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application.

[0025] Explanation of reference numerals in the attached drawings: 201, Acquisition unit; 202, Fault detection unit; 203, Reconstruction unit; 204, Role succession unit; 205, Recovery unit; 300, Electronic device; 301, Processor; 302, Memory; 303, User interface; 304, Network interface; 305, Communication bus. Detailed Implementation

[0026] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0027] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0028] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0029] Therefore, how to change the existing technology's focus solely on network connectivity restoration, which leads to the inability of remaining nodes to continue performing tasks due to a lack of role replacement, is a pressing problem that needs to be solved. This application provides a distributed, decentralized cluster network self-healing method, applied to unmanned aerial vehicle (UAV) swarms. UAV swarms are used in scenarios such as military reconnaissance, coordinated strikes, disaster relief, and logistics delivery. Figure 1 This is a flowchart illustrating a distributed, decentralized cluster network self-healing method provided in an embodiment of this application. (Refer to...) Figure 1 The method includes the following steps S101-S106.

[0030] S101: Monitor the heartbeat information and monitoring packet delivery rate of each drone node in the drone cluster, and identify suspicious neighbor nodes based on the heartbeat information and monitoring packet delivery rate.

[0031] In the above S101, during the flight of the UAV swarm, factors such as enemy air defense fire, electromagnetic interference, and bad weather may cause some nodes to be shot down or communication links to be interrupted. Traditional methods rely solely on heartbeat timeout detection, which cannot accurately distinguish whether a node is truly physically damaged or merely experiencing communication interference. This can easily lead to misjudgment of faults and result in incorrect selection of subsequent self-healing strategies.

[0032] A fault detection layer is deployed in the onboard communication module of each UAV to continuously perform monitoring tasks. Each UAV node broadcasts a heartbeat message to its one-hop neighbor at a fixed period T. This period is typically set to 100 to 500 milliseconds and can be dynamically adjusted according to the current communication load and network congestion to achieve an optimal balance between detection sensitivity and communication overhead. The heartbeat message uses a lightweight data packet format, including the node ID for uniquely identifying the sender, the current status code to indicate the node's working status, location coordinates to assist in subsequent topology reconstruction, a timestamp for synchronizing the clock and detecting message delays, and a message sequence number for detecting packet loss and duplicate reception. Simultaneously, each node maintains a heartbeat timer for all one-hop neighbor nodes. When node A listens for the heartbeat information of neighbor node B, if it does not receive a heartbeat message from B within N consecutive heartbeat periods, it marks B as a suspicious neighbor node. N is typically set to 3 to 5 to avoid misjudgments due to occasional packet loss. The reason for determining an anomaly by not receiving a heartbeat for several consecutive cycles rather than by not receiving it once is to avoid misjudgment caused by occasional wireless channel fading or data packet collisions, and to improve the reliability of detection.

[0033] Meanwhile, Node A continuously monitors the Packet Delivery Rate (PDR) with each neighboring node. This PDR is calculated by comparing the number of successfully received packets within a given time window to the total number of packets that should have been received. Node A maintains a sliding window for each neighboring node B, recording packet reception over a recent period and calculating the PDR-AB in real time. When PDR-AB falls below a preset PDR-T threshold (typically set between 0.5 and 0.7), it indicates a severe degradation in communication quality between Node A and Node B. Monitoring the PDR provides a more granular assessment of communication quality than heartbeat timeouts, because even if heartbeat messages are occasionally delivered, a large number of lost packets still indicate a serious link problem. More importantly, the trend in PDR can help distinguish different types of failures. If the PDR suddenly drops to zero and heartbeats are completely interrupted, it likely indicates physical damage to the node. Conversely, if the PDR fluctuates at a low level but heartbeats are still occasionally received, it is more likely that the node is in an environment with electromagnetic interference or is experiencing communication limitations due to obstruction.

[0034] When determining suspicious neighbor nodes based on a combination of heartbeat information and packet delivery rate, node A uses the following logic: If node A detects that neighbor node B has not sent a heartbeat message for N consecutive heartbeat cycles, or if it occasionally receives heartbeats but its packet delivery rate (PDR-AB) remains consistently below the threshold (PDR-T), then node B is marked as a suspicious neighbor node, and the timestamp of the start of B's ​​suspicious state is recorded. At this point, node A does not immediately determine that node B has failed, but instead adds B to the list of suspicious nodes for further cross-validation, avoiding misjudgments caused by local communication anomalies. This dual-indicator-based determination mechanism, compared to a single heartbeat timeout detection method, can more accurately identify nodes with real problems. Furthermore, continuous monitoring of packet delivery rate can detect communication quality degradation trends early, providing early warning information for preventative topology adjustments.

[0035] S102: Perform cross-validation on the known neighbor nodes of the suspected neighbor node, determine whether the suspected neighbor node is in a fault state based on the validation results, and identify the suspected neighbor node that is in a fault state as a fault node.

[0036] In step S102 above, after the initial identification of suspicious neighbor nodes, it is necessary to further confirm whether these suspicious nodes are physically damaged or only experiencing temporary communication interference, as these two types of failures require completely different self-healing strategies. If a node is physically damaged, a topology reconfiguration and role takeover process can be initiated. However, if a node is only experiencing localized communication interference, the connection can be restored by adjusting communication parameters or waiting for the interference to subside, avoiding unnecessary resource consumption caused by topology reconfiguration. Cross-validation is performed on the known neighbor nodes of the suspected neighbor node. Based on the validation results, it is determined whether the suspected neighbor node is in a faulty state. Specifically, this includes: obtaining observation nodes from the UAV swarm; if a suspected neighbor node exists among the neighbor nodes of the observation node, a cross-validation request is sent to the known neighbor nodes of the suspected neighbor node; collecting the observation dataset for the suspected neighbor node fed back by the known neighbor nodes, where the observation dataset includes the historical communication quality, current heartbeat reception status, and collaborative sensing data between the known neighbor nodes and the suspected neighbor node; if all current heartbeat reception statuses in the observation dataset are not received, and no collaborative sensing data from the suspected neighbor node is received within a preset waiting time, the suspected neighbor node is determined to be in a physically damaged state; if some known neighbor nodes in the observation dataset can receive the heartbeat information of the suspected neighbor node, or the physical motion trajectory in the collaborative sensing data conforms to preset formation motion constraints and the historical communication quality is lower than a preset communication threshold, the suspected neighbor node is determined to be in a local interference state.

[0037] Specifically, observation nodes are obtained from the drone swarm. An observation node is any healthy drone node that detects a suspicious neighbor node. It's important to note that observation nodes are not simply nodes marked as suspicious neighbors, but rather healthy nodes that are themselves in normal condition but have suspicious neighbors in their neighbor list. Upon discovering a neighbor node marked as suspicious, the observation node proactively initiates a cross-validation process to confirm the true type of failure. When a suspicious neighbor node exists among the observation node's neighbors, the observation node needs to send cross-validation requests to the known neighbors of the suspicious neighbor node. Here, known neighbors refer to the set of one-hop neighbors recorded in the cluster topology graph before the topology break occurred. Each drone node maintains and periodically updates its neighbor list during normal operation. This list records not only the node IDs of direct neighbors but also their neighbor information, forming a two-hop topology view. The observation node retrieves the list of known neighbors of the suspicious neighbor node by querying the locally stored topology information, then constructs a cross-validation request message and sends it to each node in the list. The cross-validation request message contains the node ID of the suspected neighbor node, the node ID of the observing node itself, the timestamp of the request being sent, and the verification request sequence number. This message requires the known neighbor node that receives the request to report its current observations of the suspected neighbor node.

[0038] After sending a cross-validation request, the observation node sets a cross-validation timeout timer Tv. This timeout is configured based on the cluster size and communication conditions, with a typical value set between 500 milliseconds and 2 seconds. This ensures sufficient response time for known neighbor nodes while maintaining real-time fault detection. During the timeout period, the observation node continuously collects observation datasets from known neighbor nodes regarding suspected neighbor nodes. The observation dataset contains three key types of information: first, the historical communication quality between known and suspected neighbor nodes, characterized by parameters such as packet delivery rate, signal strength, and link latency over a past time window, reflecting the long-term stability of the communication link between the two nodes; second, the current heartbeat reception status, indicating whether the known neighbor node is currently receiving heartbeat messages from the suspected neighbor node. This status is a Boolean value, marked as received if a heartbeat is received and marked as not received if no heartbeat is received; and third, collaborative perception data, which comes from the collaborative perception system of the UAV swarm, including information such as the physical movement trajectory, flight attitude, and position coordinates of suspected neighbor nodes observed by known neighbor nodes through visual sensors, radar, or other sensing devices. Collaborative sensing data can provide physical-level verification independent of communication links. Even if communication is interfered with and data packets cannot be transmitted, if a suspicious node is still flying normally, its physical trajectory should conform to the formation's preset movement pattern.

[0039] After collecting observation datasets from all known neighbor nodes, or after collecting feedback from some known neighbor nodes when the timeout period Tv arrives, the observation node begins a comprehensive analysis of the observation dataset to determine the failure type of the suspected neighbor node. When all known neighbor nodes in the observation dataset report no heartbeat reception, it indicates that the suspected neighbor node has stopped sending heartbeat messages, a strong signal of node failure. However, heartbeat interruption alone is insufficient to determine physical damage, as even brief communication interruptions can lead to heartbeat loss. The system continues to check the collaborative sensing data. If no collaborative sensing data about the suspected neighbor node is received from any known neighbor nodes within a preset waiting time, or if the received collaborative sensing data shows that the suspected node has deviated from its normal flight path and is stationary, it can be confirmed that the node has crashed or completely lost power, classifying the suspected neighbor node as physically damaged. This judgment logic combines evidence from both the communication-level heartbeat interruption and the physical-level cessation of movement, significantly reducing the false positive rate and ensuring that only truly damaged nodes are classified as physically damaged.

[0040] Corresponding to the physical damage state, if some known neighbor nodes in the observation dataset report that their current heartbeat reception status is "received," it indicates that the suspicious neighbor node is still sending heartbeat messages. The problem is that the observation node cannot receive the heartbeat due to locality of the communication link. This situation typically occurs when there is electromagnetic interference, obstruction, or the node is at the edge of the communication radius between the observation node and the suspicious node. In this case, examining the physical motion trajectory in the collaborative sensing data reveals that if the trajectory conforms to preset formation motion constraints—that is, the suspicious node's position, velocity, and acceleration changes are consistent with the expected movement pattern of the formation—it indicates that the node is still performing formation flight normally, and the communication problem is limited to the local link with the observation node. Further examining the historical communication quality reveals that if the historical communication quality is below a preset communication threshold, such as an average packet delivery rate below 60% over a period of time, it indicates that the communication link between the observation node and the suspicious node has been unstable for a long time. The current heartbeat interruption is likely a continuation of communication quality degradation rather than node damage. Combining this information, the suspicious neighbor node is determined to be in a state of local interference. This determination indicates that the node itself is functioning normally, and the problem lies in the communication link. The connection can be restored by adjusting communication parameters, switching communication channels, or deploying relay nodes.

[0041] After determining the fault type, the observation node records the result in its local fault status table. This table maintains the current status of all nodes in the cluster, including healthy status, suspicious status, physically damaged status, and local interference status. For a suspicious neighbor node determined to be in a physically damaged state, the observation node is identified as a faulty node and broadcasts a fault notification message within the subgroup. This message includes the faulty node's node ID, the fault type as physical damage, the determination timestamp, and a list of known neighbor nodes participating in cross-validation. After the fault notification message propagates within the subgroup, all healthy nodes that receive the message update their local topology graph, remove the faulty node from the topology, and trigger the subsequent topology reconstruction process. For a suspicious neighbor node determined to be in a local interference state, the observation node updates its status to local interference and shares this information within the subgroup. However, it does not trigger a complete topology reconstruction process; instead, it initiates communication enhancement measures, such as requesting other nodes to assist in forwarding data packets or temporarily increasing transmission power to penetrate interference.

[0042] S103: When a fault condition causes the topology of the drone cluster to break and form at least one target subgroup, collect the comprehensive score calculated by each healthy drone node in the target subgroup based on its own physical state parameters, and use a distributed consensus algorithm to elect the healthy drone node with the highest comprehensive score as the subgroup leader.

[0043] In S103 above, when the fault detection layer confirms that a node is physically damaged and that this damage has caused a topology break, each healthy drone node analyzes its local topology graph to identify whether the current cluster has split into multiple target subgroups. The nodes use a connected component analysis algorithm to group nodes in the local topology graph that can reach each other via communication links into the same connected component, with each connected component corresponding to a target subgroup. When a node finds that its connected component has lost connection with other connected components, it determines that the topology is broken and triggers the subgroup leader election process. Within the target subgroup, each healthy drone node maintains a random election timeout timer. The timeout period of this timer is randomly generated within a preset election timeout interval, typically between 150 and 300 milliseconds. This randomization avoids election conflicts caused by multiple nodes simultaneously initiating elections. When the election timeout timer of a healthy node reaches its timeout period, the node transitions to a candidate state and begins collecting comprehensive scores from all healthy drone nodes in the target subgroup to prepare for initiating an election voting request.

[0044] Before collecting the comprehensive score, each healthy UAV node needs to calculate its own comprehensive score based on its physical state parameters. The collection of comprehensive scores calculated by each healthy UAV node in the target subgroup based on its own physical state parameters includes: obtaining the physical state parameters corresponding to the healthy nodes from the target subgroup, including current remaining energy, node degree, average relative distance to other healthy UAV nodes in the target subgroup, and average signal-to-noise ratio; normalizing the physical state parameters using a preset normalization function to obtain the corresponding energy score, topology score, position score, and channel score, and constructing a score feature vector; calculating the local information entropy of each score feature based on the score feature vectors of all healthy UAV nodes in the target subgroup, and determining the initial weights of each score feature based on the local information entropy; performing a nonlinear mapping on the energy score, topology score, position score, channel score, and initial weights of the healthy nodes based on a preset state-variable-weighting activation function to generate dynamic weights that match the current physical state of the healthy nodes; and using the dynamic weights to perform a weighted summation of the energy score, topology score, position score, and channel score to obtain the comprehensive score of the healthy UAV node.

[0045] Specifically, the physical state parameters of each healthy node in the target subgroup are obtained. These parameters include current remaining energy, node degree, average relative distance to other healthy UAV nodes in the target subgroup, and average signal-to-noise ratio (SNR). Current remaining energy reflects the node's continuous working capacity; a higher remaining energy allows the node to act as the subgroup leader for a longer period without failing due to energy depletion. Node degree refers to the number of neighbors a node has in the target subgroup's topology. A higher degree indicates that the node has established direct communication links with more other nodes, placing it at the core of the topology and enabling more efficient information exchange with other nodes within the subgroup as the subgroup leader. Average relative distance is the average Euclidean distance between the node and other healthy UAV nodes in the target subgroup. A smaller average relative distance indicates that the node is located at the geometric center of the subgroup, facilitating coordination of flight formation and topology reconfiguration. Average SNR is the average SNR of the communication links between the node and its neighbors; a higher average SNR indicates better communication quality and more reliable transmission of control commands and status information.

[0046] After obtaining the physical state parameters, a preset normalization function is used to normalize these parameters, mapping physical state parameters with different dimensions and numerical ranges to a unified zero-to-one interval, facilitating subsequent comprehensive scoring calculations. For the current remaining energy, the normalization function divides the current remaining energy by the node's maximum battery capacity to obtain an energy score; the closer the energy score is to one, the more abundant the remaining energy. For the node degree, the normalization function divides the node's degree by the maximum node degree of all nodes in the target subgroup to obtain a topology degree score; the closer the topology degree score is to one, the stronger the node's connectivity in the topology. For the average relative distance, since smaller distances are better, the normalization function uses reciprocal normalization, dividing the minimum average relative distance of all nodes in the target subgroup by the node's average relative distance to obtain a location score; the closer the location score is to one, the closer the node is to the geometric center of the subgroup. For the average signal-to-noise ratio (SNR), the normalization function divides the node's average SNR by the maximum average SNR of all nodes in the target subgroup to obtain a channel score; the closer the channel score is to one, the better the node's communication quality. After normalization, the energy score, topology score, location score, and channel score are combined to construct a four-dimensional score feature vector, which comprehensively describes the node's performance across multiple dimensions.

[0047] After obtaining the score feature vectors of each healthy node, it is necessary to determine the weight of each score feature in the comprehensive score calculation. Traditional fixed-weight methods pre-set the weights of each score feature and keep them constant throughout the operation. However, this method does not consider the differences in the relative importance of each score feature under different subgroup scenarios. For example, in subgroups with generally abundant energy, the energy score has low discriminative power, while the topology score and channel score may be more decisive. Conversely, in subgroups with generally scarce energy, the energy score should receive a higher weight to ensure that the selected subgroup leader has sufficient endurance. Therefore, the initial weights of each rating feature are determined based on the local information entropy. Specifically, this includes: calculating the normalized weight of all healthy drone nodes in the target subgroup under the current rating feature, where the current rating feature is any rating feature; calculating the information entropy value of the current rating feature based on the normalized weight, and calculating the information utility value of the corresponding rating feature based on the information entropy value, where the information entropy value and the information utility value are negatively correlated; and normalizing the information utility values ​​of all rating features to obtain the initial weights of each rating feature, so that the rating feature with higher dispersion in the target subgroup receives a higher initial weight.

[0048] Specifically, for any given rating feature, the distribution of values ​​for all healthy drone nodes in the target subgroup under that rating feature is statistically analyzed. The normalized weight of each node under that rating feature is calculated, which is equal to the node's rating feature value divided by the sum of all nodes' rating feature values. Based on these normalized weights, the information entropy value of the rating feature is calculated. The formula for the information entropy value is the negative of the sum of the products of each node's normalized weight and its logarithm. The information entropy value reflects the dispersion of the rating feature in the target subgroup. A higher information entropy value indicates a more dispersed distribution of values ​​among nodes on that rating feature, meaning the feature has greater discriminative power among nodes and should receive a higher weight in the overall score. The information utility value of the corresponding rating feature is calculated based on the information entropy value. The information utility value is equal to one minus the information entropy value. Information entropy and information utility values ​​are negatively correlated. In practice, to reflect the logic that high entropy corresponds to high discriminative power and should receive a higher weight, the information utility value is defined as one minus the normalized information entropy value, so that rating features with higher information entropy values ​​receive higher information utility values. The information utility values ​​of all rating features are normalized so that the sum of the initial weights of all rating features equals one. The normalization method involves dividing the information utility value of a given rating feature by the sum of the information utility values ​​of all rating features to obtain the initial weight of that feature. Through this information entropy-based weight calculation method, rating features with higher dispersion in the target subgroup automatically receive higher initial weights. This allows the comprehensive score calculation to more accurately reflect the differences between nodes in key dimensions, improving the rationality of subgroup leader election.

[0049] However, using only initial weights for linear weighting still has limitations because this method does not consider the extreme performance of individual nodes on certain scoring features. For example, if a node's energy score is very low, close to zero, even if other scoring features are excellent, this node is not suitable as a subgroup leader because it may fail due to energy depletion at any time. Based on a preset state-variable incentive function, a nonlinear mapping is performed on the energy score, topology score, location score, channel score, and initial weights of healthy nodes to obtain dynamic weights. Specifically, this involves: constructing a preset state-variable incentive function, which includes a penalty interval and an incentive interval; when the target score of a healthy node is lower than the penalty threshold, the weight corresponding to the target score is reduced through the penalty interval to penalize the target score (which can be any score); when the target score of a healthy node is greater than the incentive threshold, the weight corresponding to the target score is increased through the incentive interval to incentivize the target score; and the processed weights are normalized and reconstructed to obtain the dynamic weights.

[0050] Specifically, a pre-defined state-weighted incentive function is constructed, comprising a penalty interval and an incentive interval. Different processing intervals are defined by setting penalty and incentive thresholds. When a healthy node's target score is below the penalty threshold, its performance on that score feature is considered insufficient. A non-linear mapping function within the penalty interval reduces the weight corresponding to this target score, achieving a punitive weight reduction. The mapping function for the penalty interval is designed so that the weight decreases at an accelerated rate as the target score decreases, ensuring that the contribution of poorly performing score features to the overall score is significantly weakened. When a healthy node's target score is above the incentive threshold, its performance on that score feature is considered excellent. A non-linear mapping function within the incentive interval increases the weight corresponding to this target score, achieving an incentive-based weight increase. The mapping function for the incentive interval is designed so that the weight increases at an accelerated rate as the target score increases, ensuring that high-performing score features gain a higher influence in the overall score. For target scores between the penalty and incentive thresholds, the weight of the score feature is kept at its initial weight without adjustment, forming a transition region. Here, the target score refers to any one of the following: energy score, topology score, location score, or channel score.

[0051] The preset state-adjusted incentive function can be expressed as a piecewise function. For the i-th rating feature of a healthy node, let the normalized score be Si, the initial weight be Wi, the penalty threshold be Tp, and the incentive threshold be Te. When Si is less than Tp, the adjusted weight Wi1 is calculated using the penalty function, which can adopt an exponential decay form, such as Wi1 = Wi * exp((-kp) * Tp - Si), where kp is the penalty intensity coefficient. This formula makes the weight decrease faster as the score is lower than the penalty threshold. When Si is greater than Te, the adjusted weight Wi1 is calculated using the incentive function, which can adopt an exponential growth form, such as Wi1 = Wi * exp(ke * Si - Te), where ke is the incentive intensity coefficient. This formula makes the weight increase faster as the score is higher than the incentive threshold. When Si is between Tp and Te, the adjusted weight Wi1 = Wi remains unchanged. After adjusting the weights of each scoring feature by penalties or incentives, all adjusted weights are normalized and reconstructed to ensure that the sum of the dynamic weights of each scoring feature remains equal to one. The normalization method is to divide the adjusted weight of a scoring feature by the sum of the adjusted weights of all scoring features to obtain the dynamic weight of that scoring feature. Through this state-variable incentive mechanism, the calculation of the comprehensive score can adaptively respond non-linearly to the actual performance of nodes on each scoring feature, avoiding the selection of nodes with serious weaknesses in certain key dimensions as subgroup leaders, while prioritizing the selection of nodes that perform well in multiple dimensions. The preset state-variable incentive function can be a piecewise exponential function. When the score Si is lower than the penalty threshold Tp, the new weight is Wi*exp(-kp(Tp-Si)); when the score Si is higher than the incentive threshold Te, the new weight is Wi*exp(ke(Si-Te)), where kp and ke are preset penalty and incentive coefficients.

[0052] After obtaining the energy score, topology score, location score, channel score, and corresponding dynamic weights of each healthy node, the comprehensive score of the healthy drone node is calculated by weighted summing of these scores using the dynamic weights. The comprehensive score is calculated as follows: energy score multiplied by the dynamic weight of the energy score, topology score multiplied by the dynamic weight of the topology score, location score multiplied by the dynamic weight of the location score, and channel score multiplied by the dynamic weight of the channel score. The result is a value between zero and one. A higher comprehensive score indicates that the node has stronger overall capabilities across multiple physical state dimensions and is more suitable to serve as a subswarm leader. After calculating the comprehensive score locally, each healthy drone node stores the comprehensive score value in its local state table, ready for use in subsequent election processes.

[0053] For example, the target subgroup contains five healthy drone nodes, numbered N1 to N5. The physical state parameters of each node are obtained: node N1 has 80% of its maximum capacity remaining energy, a node degree of 3, an average relative distance of 25 meters, and an average signal-to-noise ratio of 18 dB; node N2 has 95% of its maximum capacity remaining energy, a node degree of 2, an average relative distance of 30 meters, and an average signal-to-noise ratio of 20 dB; other node parameters are obtained similarly. The physical state parameters are then normalized. The energy score of node N1 is 0.80 divided by 1.00, which equals 0.80; the topology degree score is 3 divided by the maximum node degree of the subgroup (4), which equals 0.75; the position score is reciprocally normalized to the minimum average relative distance of the subgroup (20 meters divided by 25 meters), which equals 0.80; and the channel score is 18 dB divided by the maximum signal-to-noise ratio of the subgroup (20 dB), which equals 0.90. The score feature vector of node N1 is (0.80, 0.75, 0.80, 0.90). Similarly, calculate the rating feature vectors for other nodes. Calculate the information entropy of each rating feature. Taking energy rating as an example, the normalized weights of the energy ratings for the five nodes are 0.18, 0.21, 0.19, 0.22, and 0.20, respectively. The calculated information entropy value is approximately 1.60, and the information utility value is 2.32 minus 1.60, which equals 0.72. After normalizing the information utility values ​​of the four rating features, the initial weights are obtained. Assume the initial weight for energy rating is 0.25, topology rating is 0.30, location rating is 0.20, and channel rating is 0.25. Apply the state-varying incentive function. Set the penalty threshold to 0.30 and the incentive threshold to 0.85. Node N1's channel rating of 0.90 is greater than the incentive threshold, and its weight increases from 0.25 to 0.32 through the incentive function. Node N1's topology rating of 0.75 is in the transition region, and its weight remains unchanged at 0.30. After normalizing and reconstructing the adjusted weights, the dynamic weights are 0.23, 0.28, 0.19, and 0.30. The overall score of node N1 is 0.80*0.23 + 0.75*0.28 + 0.80*0.19 + 0.90*0.30 = 0.817. Similarly, the overall scores of other nodes are calculated. Assuming node N2 has an overall score of 0.845, then N2 will receive a higher voting priority in the election.

[0054] Furthermore, after all healthy nodes in the target subgroup have completed their comprehensive score calculations, the nodes that have transitioned to the candidate state initiate the election process. A distributed consensus algorithm is used to elect the healthy drone node with the highest comprehensive score as the subgroup leader. Specifically, this involves: each healthy drone node broadcasting an election message containing its own identifier and comprehensive score within the target subgroup; each healthy drone node collecting the election messages broadcast by other healthy drone nodes in the target subgroup and building a score mapping table locally; each healthy drone node performing multiple rounds of interactive verification on the data in the score mapping table based on the distributed consensus protocol until all healthy drone nodes in the target subgroup reach a consensus on the score mapping; and each healthy drone node determining the healthy drone node with the highest comprehensive score as the subgroup leader based on the consensus-reached score mapping table.

[0055] Specifically, each healthy drone node within the target subgroup enters the leader election phase. Each node maintains a randomly generated election timeout timer, the timeout duration of which is randomly selected within a preset range, for example, between 150 milliseconds and 300 milliseconds. This randomization design borrows from the Raft consensus algorithm, aiming to avoid vote scattering and election failure caused by multiple nodes simultaneously launching elections. When a node's election timeout timer expires, the node changes its state from follower to candidate and immediately broadcasts the election message to all neighboring nodes within the target subgroup.

[0056] The election message is encapsulated in a structured data format, containing the following key fields: unique node identifier, current election round number, node comprehensive score, message timestamp, and message sequence number. The unique node identifier distinguishes different candidates, the election round number indicates the current election stage, the comprehensive score is a competitiveness indicator calculated by the node based on its physical state, the message timestamp synchronizes the time base of each node, and the message sequence number detects message loss and duplicate reception. Taking node N3 as an example, assuming its comprehensive score is 0.827 after the topology break, when its election timeout expires after 180 milliseconds, node N3 constructs an election message {ID: N3, Term: 1, Score: 0.827, Timestamp: T1, SeqNum: 001} and broadcasts it to its neighboring nodes N1, N2, N4, and N5 within the subgroup via the wireless communication module. To ensure reliable transmission of the election message, an application-layer confirmation mechanism is used. Nodes receiving the election message must reply with a confirmation message to the sending node within a set time window. If the sending node does not receive an acknowledgment from a neighboring node within the timeout period, it determines that the message transmission has failed and will retransmit the message to that neighboring node up to three times. If no acknowledgment is received after three retransmissions, the sending node will mark the neighboring node as temporarily unreachable and exclude it from voting in the subsequent election process.

[0057] Each healthy drone node within the target subgroup, upon receiving election messages broadcast by other nodes, begins constructing a local score mapping table. This table is organized using a hash table data structure, with the node identifier as the key and the node's overall score as the value. It also records the message's timestamp and sequence number for subsequent consistency checks. Taking node N1 as an example, assuming node N1 receives election messages from nodes N2, N3, N4, and N5 during the election phase, node N1 first verifies the integrity and validity of each message, checking if the message format conforms to the predefined data structure, if the timestamp is within a reasonable range, and if the sequence number is consecutive. After successful verification, node N1 extracts the node identifier and overall score from the election messages and inserts them into its local score mapping table. During the construction of the score mapping table, nodes need to handle various exceptions. If a node receives multiple election messages from the same sender, it determines whether the message is a duplicate or an updated message based on the message sequence number. If the sequence numbers are the same, it is determined to be a duplicate message. The first received message is retained, and subsequent duplicate messages are discarded. If the sequence numbers are incremented, it is determined that the sending node has updated its own score, and the new message overwrites the old record in the score mapping table. If a node detects an interruption in the heartbeat signal of a candidate node during the construction of the score mapping table, it immediately removes the node from the score mapping table and broadcasts a notification message that the node is suspected of being invalid to other nodes in the subgroup, triggering the neighbor cross-validation process to confirm the true status of the node.

[0058] After each node completes its local score mapping table, it does not immediately conduct leader election but instead enters the consistency verification phase of the score mapping table. Due to the unreliability of wireless communication and the dynamic changes in node locations, the score mapping tables built by different nodes at the same time may differ. For example, node N1 may fail to receive node N5's election message due to its distance, resulting in a missing record for N5 in its score mapping table; while node N2 may have received N5's election message, and its score mapping table contains N5's record. If nodes conduct leader election based on inconsistent score mapping tables, different nodes may elect different leaders, disrupting the coordination and consistency of the subgroup. To solve the consistency problem of the score mapping table, a multi-round interactive verification method based on a distributed consensus protocol is adopted. The interactive verification method draws on the log replication and consistency guarantee ideas in the Raft algorithm, and is adapted to the actual constraints of the drone swarm. The core objective of the protocol design is to ensure that all healthy drone nodes in the target subgroup eventually reach a completely consistent consensus on the score mapping table, that is, all nodes' score mapping tables contain the same set of nodes and corresponding comprehensive score values.

[0059] The protocol execution consists of multiple interaction rounds, each round comprising two phases: information exchange and local verification. In the first round of interaction, each node encapsulates the complete content of its local scoring map into a synchronization message and broadcasts it to all neighboring nodes within the subgroup. The synchronization message includes the node's own identifier, the current round number, the hash value of the scoring map, and the complete scoring map data. The hash value is used to quickly determine whether the scoring maps of different nodes are consistent, avoiding the communication overhead of transmitting and comparing complete table data. Taking node N2 as an example, assuming that N2's local scoring map contains records from five nodes {N1: 0.802, N2: 0.845, N3: 0.827, N4: 0.791, N5: 0.813} in the first round of interaction, N2 calculates the SHA-256 hash value of the table to obtain Hash-N2, then constructs a synchronization message {ID: N2, Round: 1, Hash: Hash_N2, Table: {...}}, and broadcasts it to its neighboring nodes. After receiving synchronization messages from other nodes, each node enters the local verification phase. The node first compares the hash value of the received rating mapping table with the hash value of its local rating mapping table. If the hash values ​​match exactly, it means that the node's rating mapping table has reached an agreement with the sending node, and no further processing is needed for that node in this round of interaction. If the hash values ​​do not match, the node compares the received rating mapping table with its local rating mapping table item by item to identify discrepancies. Discrepancies may include three types: records present in the local table but missing in the remote table, records present in the remote table but missing in the local table, and records present in both tables but with different rating values.

[0060] For the first type of discrepancy, a node needs to determine the validity of its local record. The node checks whether the node corresponding to the record is still active within the subgroup, confirming this by querying the most recent heartbeat signal and communication status. If the node is out of contact or confirmed to be corrupted, the record is deleted from the local score mapping table; if the node is still active, the record is appended to the update message sent to the remote node, prompting the remote node to replenish the record. For the second type of discrepancy, the node proactively sends a query message to the node that exists in the remote table but is missing in the local table, requesting that node to resend its election message. If a response is received from the node within the timeout period, the node's record is added to the local score mapping table; if no response is received, a challenge message is sent to the remote node, suggesting that the remote node delete the record. For the third type of discrepancy, the node compares the message timestamps corresponding to the two score values, retaining the score value with the newer timestamp and discarding the record with the older timestamp.

[0061] After the first round of interactive verification, each node updates its local scoring mapping table and recalculates the hash value. In the second round of interaction, each node broadcasts the updated scoring mapping table and its hash value again. Theoretically, after a finite number of rounds of interaction, the scoring mapping tables of all nodes will converge to a consistent state. To prevent convergence failure due to network latency or message loss, a maximum number of interaction rounds is set, such as three or five rounds. If some nodes still fail to reach a consensus on their scoring mapping tables within the maximum number of rounds, a majority rule is used for adjudication: the hash value distribution of each node's scoring mapping table is statistically analyzed, and the scoring mapping table corresponding to the hash value with the highest frequency is selected as the standard mapping table for the subgroup. All nodes are then required to update to this standard mapping table. This adjudication mechanism ensures that even when communication is restricted for some nodes, the subgroup can still complete the leader election process, avoiding the system entering an indefinite waiting state. Once all healthy drone nodes in the target subgroup reach a consensus on the scoring mapping table, each node determines the leader based on the comprehensive score value in the agreed-upon scoring mapping table. The determination rule adopts the maximum value priority principle, that is, the healthy drone node with the highest comprehensive score is determined as the subgroup leader. Each node traverses the scoring mapping table to find the maximum value of the overall score and its corresponding node identifier. Assuming the agreed-upon scoring mapping table is {N1: 0.802, N2: 0.845, N3: 0.827, N4: 0.791, N5: 0.813}, after calculation, each node finds that N2's overall score of 0.845 is the highest, therefore unanimously determining N2 as the subgroup leader.

[0062] To handle the extreme case of identical overall scores, if multiple nodes have the same overall score within the numerical precision range, secondary indicators are compared sequentially: remaining energy percentage, communication capability score, and node identifier value. Specifically, among nodes with the same overall score, the node with the highest remaining energy percentage is selected; if the remaining energy is still the same, the node with the highest communication capability score is selected; if the communication capability scores are also the same, the node with the smallest node identifier value is selected. This ensures that the subgroup leader is uniquely determined under any circumstances, avoiding conflicts caused by multiple nodes simultaneously claiming to be the leader. After confirming its role, the selected subgroup leader immediately broadcasts a leader declaration message to all nodes in the subgroup. The declaration message includes the leader node's identifier, election round number, leader's overall score, and term identifier. The term identifier is a monotonically increasing integer used to identify the leader's term, incremented by one after each new leader election. Upon receiving the leader declaration message, other nodes verify that the node identifier and overall score in the declaration message match their local score mapping table. Once the verification is successful, each node will change its state from candidate or follower to follower and record the current navigator identifier and term identifier. All subsequent operations that require subgroup coordination will be performed through this navigator node.

[0063] After the leader node completes its role transition, it assumes the core coordination responsibility within the subgroup. The leader periodically sends heartbeat messages to all nodes in the subgroup to maintain its leader status; the typical heartbeat period is 100 to 500 milliseconds. The heartbeat message contains the leader's status information, the subgroup member list, and task allocation instructions. Upon receiving a heartbeat message, a follower node resets its local election timeout timer. As long as it continues to receive a heartbeat from the leader, the follower node will not initiate a new election. If a follower node does not receive a heartbeat message from the leader for several consecutive heartbeat periods, it determines that the leader may have failed, automatically restarts its election timeout timer, and prepares to initiate a new round of leader election.

[0064] S104: Under the coordination of the subgroup leader, the neighboring nodes of the faulty node perform local topology repair based on rigid graph theory to restore the internal connectivity of the target subgroup. The subgroup leader then discovers other target subgroups through adaptive expansion and triggers merging to obtain the reconstructed cluster.

[0065] In S104 above, after the drone swarm completes the election of sub-swarm leaders, there may still be insufficient local connectivity within each sub-swarm due to the absence of faulty nodes, and different sub-swarms are in a completely isolated state. If relying solely on traditional global topology reconstruction methods, all nodes need to participate in the computation and exchange complete topology information, which will generate huge communication overhead and computational burden in large-scale swarms, and is difficult to implement in adversarial environments with limited communication.

[0066] Under the coordination of the subgroup leader, the neighboring nodes of the failed node perform local topology repair based on rigid graph theory to restore the internal connectivity of the target subgroup. The subgroup leader then adaptively expands to discover other target subgroups and triggers merging, resulting in a reconstructed cluster. Specifically, this involves: under the coordination of the subgroup leader, obtaining the set of neighboring nodes that are within the local topology association range of the failed node, and extracting a local subgraph composed of the neighboring node set; constructing a local minimum rigid graph in the local subgraph based on rigid graph theory, and merging the local minimum rigid graph into the global topology of the target subgroup to complete the local topology of the subgroup's internal connectivity. The fix involves the subgroup navigator periodically broadcasting subgroup probe messages. If no feedback is received from other subgroups within a preset probe time, the navigator controls the nodes in the target subgroup to adaptively increase their communication radius to scan for the existence of other target subgroups. Alternatively, the subgroup navigator determines relay candidate nodes in the target subgroup, controls these candidate nodes to move to a preset adjacent subgroup location to be deployed as communication relays, and actively establishes cross-subgroup communication links. Once a communication connection is established with other target subgroups, the subgroup navigator exchanges member lists and status information with the navigators of the other target subgroups, and performs subgroup merging based on distributed joint consensus to obtain the reconstructed cluster.

[0067] Specifically, after the subgroup leader is elected and confirms its role, it broadcasts a topology status query message to all nodes in the subgroup. This message requests each node to report its current neighbor connectivity and known faulty node information. Upon receiving the query message, each node encapsulates its locally maintained neighbor list, neighbor connectivity quality parameters, and information on disconnected nodes marked by the fault detection layer into a topology status report message, and sends it back to the subgroup leader via a single-hop or multi-hop path. Taking node N2 as the subgroup leader as an example, assuming that the subgroup contains healthy nodes N1, N2, N3, N4, and N5 after a topology breakdown, and the faulty node is N6. Node N1 reports its neighbors as N2 and N6 (disconnected), node N3 reports its neighbors as N2, N4, and N6 (disconnected), node N4 reports its neighbors as N3 and N5, and node N5 reports its neighbor as N4.

[0068] After collecting all topology status reports, the subgroup leader N2 constructs a global topology graph representation of the subgroup locally. This topology graph is stored using an undirected graph data structure, with node identifiers as vertices and communication links as edges, the weight of which represents link quality or communication distance. By performing connectivity analysis on the topology graph, the leader identifies which nodes' connectivity is affected by the absence of the faulty node N6. The specific analysis method uses a depth-first search or breadth-first search algorithm to traverse the topology graph starting from any healthy node, checking if all healthy nodes are reachable. In this example, due to the absence of N6, nodes N1 and N3, which were originally connected by N6, lose their direct communication path. If the physical distance between N1 and N3 exceeds the single-hop communication range, a connectivity break occurs within the subgroup. To determine the range of nodes that need to participate in topology repair, the concept of a local topology association range is introduced. The local topology association range is defined as the set of one-hop neighbor nodes of the faulty node before the topology break. This definition is based on the locality principle in rigid graph theory: in a planar rigid graph, the removal of a single vertex only affects the local rigidity of its immediate neighbors, without directly affecting distant nodes more than two hops away. The subgroup leader extracts all nodes that had a neighbor relationship with N6 before the failure, forming a neighbor set, based on the collected topology state reports. In this example, the neighbor set is {N1, N3}, because both N1 and N3 reported a connection to N6 before the failure.

[0069] The subgroup leader sends a topology repair task instruction to each node in its neighboring node set. This instruction message includes the faulty node identifier, a list of neighboring node sets, the repair coordinator identifier (usually the leader itself), and a repair task timeout. Upon receiving the instruction, nodes N1 and N3 enter a topology repair preparation state and begin collecting the local topology information required for repair. Each neighboring node not only records its own connections to other healthy nodes but also needs to acquire the connections between neighboring nodes to construct a complete local subgraph. Node N1 queries its neighbor N2 and requests N2 to provide its connection status with N3; node N3 similarly queries its neighbors N2 and N4, collecting the connections between them and with N1. After collecting the necessary topology information, each node in the neighboring node set independently extracts its local subgraph. The local subgraph is a derived subgraph of the global topology graph. The vertex set of the local subgraph includes all nodes in the neighboring node set and their direct neighbors, and the edge set includes all existing communication links between these vertices. Taking node N1 as an example, the vertex set of the local subgraph is {N1, N2, N3}, and the edge set is {(N1, N2), (N2, N3)}. This is because N1 is directly connected to N2, N2 is directly connected to N3, and the direct connection between N1 and N3 is lost due to the absence of N6. The vertex set of the local subgraph extracted from node N3 is {N2, N3, N4}, and the edge set is {(N2, N3), (N3, N4)}.

[0070] After the local subgraph is extracted, each node sends its local subgraph data to the sub-swarm leader via an encrypted communication channel. Upon receiving the local subgraphs from all neighboring nodes, the leader performs a local subgraph merging operation. This merging operation uses graph union to combine the vertex and edge sets of the local subgraphs submitted by different nodes, eliminating duplicate vertices and edges to form a merged local subgraph containing complete information about the neighborhood of the faulty node. In the example above, the merged local subgraph has vertex set {N1, N2, N3, N4} and edge set {(N1, N2), (N2, N3), (N3, N4)}. Based on the merged local subgraph, the sub-swarm leader initiates the algorithm for constructing a local minimum rigid graph. Rigid graph theory states that in a two-dimensional plane, a minimum rigid graph with n vertices requires exactly 2n-3 independent edges, while in three-dimensional space, it requires 3n-6 edges. For the topological scenario of a UAV swarm in three-dimensional space, the system adopts the construction criteria for a three-dimensional rigid graph. The core idea of ​​the algorithm is to add a minimum number of new edges to the existing merged local subgraph so that the repaired local topology meets the edge count requirements of a three-dimensional rigid graph, while ensuring connectivity between all neighboring nodes.

[0071] The algorithm's execution is divided into three stages. The first stage is the connectivity repair stage. The algorithm checks the connectivity of the merged local subgraph. If disconnected node pairs exist, edges are added to restore connectivity. In the example above, nodes N1 and N3 lack a direct connection, but a two-hop path exists through N2, so the merged local subgraph is logically connected. The second stage is the rigidity evaluation stage. The algorithm calculates the rigidity metrics of the current local subgraph. Commonly used metrics include the graph's algebraic connectivity (Fiedler value) and the rank of the rigidity matrix. Algebraic connectivity reflects the graph's connectivity robustness; a higher value indicates that the graph is less likely to be destroyed by the removal of a small number of nodes or edges. The rank of the rigidity matrix directly reflects the graph's rigidity; a full-rank rigidity matrix corresponds to a rigid graph. By calculating these metrics, the system determines whether the current local subgraph already possesses sufficient rigidity. The third stage is the edge addition optimization stage. If the current local subgraph does not meet the requirements of a minimum rigidity graph, the algorithm selects the optimal edge from the candidate edge set for addition based on the nodes' physical location and communication capabilities. The candidate edge set includes all node pairs within the physical communication range but for which no link has been established yet. Edge selection follows several optimization criteria: the principle of closest physical distance, prioritizing edges between nodes with shorter physical distances to reduce communication power consumption and improve link reliability; the principle of optimal communication quality, prioritizing links with high signal-to-noise ratios and low interference; and the principle of topology balancing, avoiding nodes becoming communication bottlenecks by connecting too many edges, prioritizing edges for nodes with lower degrees. In the example above, assuming nodes N1 and N3 are 18 meters apart, within a 20-meter communication radius, and without significant obstacles, the system decides to add a direct communication link between N1 and N3, forming edge (N1, N3). After adding this edge, the edge set of the local subgraph is updated to {(N1, N2), (N2, N3), (N3, N4), (N1, N3)}.

[0072] To verify whether the local subgraph after adding edges satisfies the requirements of a rigid graph, the algorithm recalculates the rank of the rigidity matrix. For a 3D rigid graph containing four nodes, theoretically 3*4-6=6 edges are needed. The current local subgraph only contains 4 edges, which does not yet meet the minimum edge count requirement for a rigid graph. The algorithm continues searching the candidate edge set. Assuming the physical distance between nodes N2 and N4 is 16 meters, within communication range, the algorithm adds the edge (N2, N4). At this point, the edge set is {(N1, N2), (N2, N3), (N3, N4), (N1, N3), (N2, N4)}, containing 5 edges. Further evaluation is performed. Assuming the physical distance between nodes N1 and N4 is 25 meters, exceeding the standard communication radius, but a link can be established by temporarily increasing the transmission power, the algorithm weighs the communication cost and decides not to add this edge. Instead, it optimizes the topology performance by adjusting the weights of other edges. After multiple rounds of iterative calculations, the system confirmed that the current topology of the 5 edges is close to the optimal under actual physical constraints. Although the number of edges has not reached the requirement of the theoretical minimum rigid graph, the topology can still maintain connectivity in the event of a single link failure by introducing a fault-tolerant redundancy mechanism.

[0073] After constructing the local minimum rigid graph, the subgroup leader needs to merge it into the global topology of the target subgroup, ensuring seamless integration between the repaired local topology and other unaffected parts of the subgroup. The key to this merging operation is ensuring that newly added edges do not conflict with the existing topology, while verifying that the merged global topology maintains its overall connectivity and rigidity. The subgroup leader distributes the completed local minimum rigid graph data to each node in its neighboring node set. The distribution message contains a complete description of the local topology, including a vertex list, an edge list, and communication parameters for each edge (such as frequency, power, modulation scheme, etc.). Nodes N1 and N3, upon receiving the distribution message, establish new communication links locally according to the instructions in the message. Taking node N1 as an example, this node needs to establish a direct communication link with N3. Node N1's communication module configures the transceiver's operating parameters according to the allocated communication parameters and sends a link establishment request message to N3. Upon receiving the request, node N3 performs a link handshake and parameter negotiation; both parties confirm the link establishment through a three-way handshake protocol. After the link is established, N1 and N3 each send a link confirmation message to the subgroup leader, reporting the actual communication quality parameters of the link, including signal strength, bit error rate, round-trip time, etc.

[0074] After collecting confirmation messages for all newly added links, the subgroup leader updates its locally maintained global topology graph. The update operation adds the new edges from the local minimum rigid graph to the edge set of the global topology graph, and simultaneously updates the neighbor lists and routing tables of the relevant nodes. To verify the correctness of the merge operation, the leader performs an integrity check on the updated global topology graph. This check includes: connectivity verification, confirming that there is at least one communication path between all healthy nodes in the subgroup using a graph traversal algorithm; loop detection, identifying redundant loops in the topology (redundant loops can improve the topology's fault tolerance, but too many loops increase the complexity of routing protocols and communication overhead); and degree distribution analysis, counting the number of neighbor connections for each node to avoid hotspot nodes with excessively high degrees or isolated nodes with excessively low degrees. After completing the global topology update, the subgroup leader broadcasts a topology update notification message to all nodes in the subgroup. This message contains a snapshot of the updated global topology. Upon receiving the notification, each node synchronously updates its local topology view and routing table. To ensure the consistency of topology updates, the system employs a version number mechanism. After each topology update, the global topology version number is incremented, and each node records the current topology version number locally. When a node sends a data packet or control message, it includes its local topology version number in the message header. The receiver compares the version number to determine if the sender's topology view is consistent with its own. If the version numbers do not match, the receiver requests the latest topology update message from the sender or the subgroup leader to maintain global topology consistency.

[0075] After local topology repair, connectivity within the subgroup is restored. Nodes N1 and N3, which were previously unable to communicate directly due to the absence of the faulty node N6, now achieve one-hop communication through the newly established direct link. The communication latency has been reduced from two hops via N2 to a single hop, significantly improving data transmission efficiency. The entire local repair process involves only the neighboring nodes N1 and N3 of the faulty node, as well as the coordinator N2, far less than the communication overhead required for global topology reconstruction, which necessitates the participation of all nodes. According to rigid graph theory, the repaired local topology, supported by the two newly added edges (N1, N3) and (N2, N4), ensures that even if one link is interrupted again in the future, the subgroup can still maintain connectivity through other paths, enhancing the robustness of the topology.

[0076] After completing local topology repair within a subgroup, the subgroup leader initiates the inter-subgroup connectivity restoration process. Since topology breakdowns can cause the cluster to split into multiple isolated subgroups, each subgroup leader needs to actively probe for the existence of other subgroups and establish cross-subgroup communication links, ultimately merging all subgroups and restoring the cluster's integrity. The subgroup leader broadcasts subgroup probe messages at fixed periodic intervals. The typical range for the probe period Td is 1 to 5 seconds, and the specific value can be dynamically adjusted based on task urgency and communication resource constraints. Subgroup probe messages are encapsulated using a specific message format. The message header includes a message type identifier, the sender's subgroup leader identifier, the subgroup identifier, the number of subgroup members, and the message sequence number. The message body includes a subgroup member list, the subgroup center coordinates, the subgroup's average velocity vector, and a summary of the subgroup's current task status. The broadcasting of probe messages employs a multi-level power expansion strategy. In the first round of broadcasting, the subgroup leader uses standard communication power to send probe messages to the surrounding area, covering a standard communication radius, assumed to be 20 meters. If no response is received from other subgroups within a detection period Td, the navigator will increase its transmission power to 1.5 times the standard power in the second round of broadcasting, extending the coverage range to approximately 30 meters. If no response is received within three consecutive detection periods (3*Td), it is determined that no other subgroups exist within the current communication radius, triggering the adaptive communication radius extension mechanism.

[0077] To avoid the subgroup leader bearing the energy consumption of high-power broadcasting alone, a distributed probing strategy is adopted. The subgroup leader issues probing task instructions to all healthy nodes within the subgroup, assigning each node a probing sector. A probing sector is a cone-shaped region divided into three dimensions, with the subgroup center as the origin. Each node is responsible for broadcasting probing messages to its assigned sector. For example, in a subgroup with five nodes, the system evenly divides the horizontal 360 degrees into five 72-degree sectors, with each node responsible for probing one sector. Node N1 is responsible for sectors from 0 to 72 degrees, node N2 for sectors from 72 to 144 degrees, and so on. This distributed probing strategy expands the probing range and balances the energy consumption of each node, preventing the leader from becoming an energy bottleneck. When a node in a subgroup receives a probing message from another subgroup, that node immediately reports the discovery to its subgroup's leader. The report message includes the leader identifier of the discovered external subgroup, the subgroup identifier, subgroup information carried in the probe message, and the Received Signal Strength Indicator (RSSI). Upon receiving the discovery report, the subgroup leader records the information of the external subgroup and immediately sends a response message to the leader of the external subgroup. The response message contains complete information about its own subgroup, including a member list, leader identifier, and subgroup center location, used to establish a two-way inter-subgroup communication link.

[0078] Assume node N2 is the leader of subgroup A and node N8 is the leader of subgroup B. When node N5 in subgroup A is performing a distributed probing task, it receives a probing message broadcast by node N7 in subgroup B. Node N5 sends a discovery report to its own subgroup leader N2, and simultaneously sends a temporary response message to node N7, informing node N7 that it has discovered the existence of subgroup B. Node N7 forwards N5's response message to the leader of subgroup B, N8. At this point, both the leader of subgroup A, N2, and the leader of subgroup B, N8, are aware of the existence of the other's subgroup, and they enter the subgroup merging negotiation phase. If, within a preset probing time, a subgroup still fails to discover other subgroups through periodic broadcasting of probing messages, an adaptive communication radius expansion mechanism is activated. The trigger condition for this mechanism is that no response from any external subgroup is received within a consecutive number of probing cycles (e.g., three to five cycles). Upon triggering, the subgroup leader calculates the topological connectivity metric of the current subgroup. A commonly used metric is the algebraic connectivity of the graph, i.e., the Fiedler value λ². The Fiedler value is the second smallest eigenvalue of the Laplacian matrix of the topological graph. A larger value indicates stronger connectivity and less susceptibility to partitioning. When the Fiedler value falls below a preset safety threshold, it indicates that the current subgroup's topology is relatively fragile, requiring the addition of redundant links by expanding the communication radius to improve robustness.

[0079] The subgroup leader broadcasts a communication radius extension command to all nodes within the subgroup. The command specifies the new communication radius Rc, which is typically set to 1.2 to 1.5 times the standard communication radius Rcom. For example, with a standard communication radius of 20 meters, the extended communication radius can be set to 25 to 30 meters. The communication radius extension is achieved by adjusting the transmit power and receive sensitivity of the wireless transceiver. The increase in transmit power follows the inverse square law of electromagnetic wave propagation; extending the communication distance by 1.5 times requires an increase in transmit power of approximately 2.25 times. Considering the energy constraints of the UAV, the duration of the communication radius extension is limited; the extension state is maintained for a limited time window, typically 10 to 30 seconds. If other subgroups are discovered within the extension time window, the system establishes communication links between subgroups, and each node reverts to its standard communication radius. If no other subgroups are discovered by the end of the extension time window, the extension strategy is temporarily abandoned, and a relay node deployment mechanism is initiated. During the communication radius extension state, each node rescans the surrounding wireless channels, attempting to establish communication links with nodes at greater distances. The node's communication module executes the neighbor discovery protocol, broadcasting a neighbor discovery message within the extended communication radius. This message includes the node's identifier, its subgroup identifier, and its current location coordinates. External nodes receiving the neighbor discovery message, if belonging to a different subgroup, reply with a neighbor response message, reporting their own subgroup affiliation information. Through this proactive scanning and response mechanism, the system can discover other subgroups that were previously inaccessible due to distance within the extended communication radius.

[0080] Assuming the center of subgroup A is 28 meters away from the center of subgroup C, a direct communication link cannot be established between the edge nodes of the two subgroups within a standard communication radius of 20 meters. When subgroup A initiates adaptive communication radius expansion, extending the communication radius to 30 meters, the distance between edge node N5 of subgroup A and edge node N11 of subgroup C decreases to within the expanded communication radius. Node N5's neighbor discovery message is successfully received by N11, and N11 replies with a response message, confirming the feasibility of a cross-subgroup communication link. Node N5 reports the discovered information about external subgroup C to its subgroup leader N2, while node N11 reports the information about subgroup A to the leader of subgroup C, N10. After both leaders are aware of each other's existence, they enter the subgroup merge negotiation process.

[0081] If the adaptive communication radius expansion mechanism fails to discover other subgroups within a specified time, it is inferred that there may be subgroups that are far apart. Simply increasing the communication radius is insufficient to establish a connection, and communication relay nodes need to be deployed to establish cross-subgroup communication links. The subgroup leader initiates a selection process for relay candidate nodes within the subgroup. The selection of relay candidate nodes is based on multi-dimensional evaluation indicators, comprehensively considering factors such as the node's remaining energy, current task load, communication capability, and relative location. Remaining energy is the primary consideration, as relay nodes need to maintain communication links for a relatively long period. Insufficient energy may lead to premature relay link failure; therefore, nodes with a remaining energy percentage higher than 80% are preferred. Current task load reflects whether the node plays a critical role. If a node is performing high-priority tasks such as reconnaissance or attack, it is not suitable to be reassigned to relay tasks. Communication capability evaluates the performance of the node's wireless transceiver, including maximum transmit power, receive sensitivity, and supported communication protocol types. Nodes with strong communication capabilities can establish longer-distance and more stable relay links. Relative position refers to the spatial location of a node within a subgroup. Nodes located at the edge of a subgroup and facing a preset adjacent subgroup are more likely to move closer to other subgroups.

[0082] The subgroup leader calculates the relay candidate score for each node within the subgroup based on the aforementioned evaluation metrics. The scoring formula uses a weighted summation, with the weights of each metric configured according to the task scenario. Assume the remaining energy weight is 0.4, the task load weight is 0.2, the communication capability weight is 0.2, and the relative position weight is 0.2. The subgroup leader sorts the scores of each node and selects one or more nodes with the highest scores as relay candidate nodes. In subgroup A containing five nodes, assuming node N5 has 90% remaining energy, is not currently undertaking a critical task, has a communication capability score of 0.85 (normalized), and is located on the southern edge of the subgroup, while based on task planning and previous cluster topology information, it is inferred that other possible subgroups may exist in the southern direction, and node N5's relative position score is 0.90. The calculated overall score for N5 is 0.90*0.4 + 0.80*0.2 + 0.85*0.2 + 0.90*0.2 = 0.87, ranking highest within the subgroup and thus selected as a relay candidate node. The subgroup navigator sends a relay deployment command to relay candidate node N5. The command includes the target movement direction, expected movement distance, communication scan parameters, and mission timeout limit. Upon receiving the command, node N5 switches its state to relay mode, suspends its current regular flight mission, and plans a southward movement trajectory according to the command. During the movement, node N5 maintains communication connections with other nodes within subgroup A, periodically reporting its current position and communication scan results to the subgroup navigator N2. Node N5's communication module continuously performs external subgroup scans, broadcasting high-power probe messages in the movement direction, attempting to establish initial contact with any potential external subgroups.

[0083] For example, after node N5 moves 50 meters south, its probe message is received by edge node N11 of subgroup C. Node N11 replies with a response message, confirming the existence and location of subgroup C. Node N5 immediately reports the discovery of subgroup C to the navigator N2 of subgroup A, and simultaneously negotiates with node N11 to establish a stable relay link. After the relay link is established, node N5 acts as a communication relay between subgroup A and subgroup C, and all messages that need to be transmitted between the two subgroups are forwarded via N5. To improve the reliability of the relay link, the system can also deploy multiple relay candidate nodes to simultaneously search for external subgroups in different directions, increasing the success rate of discovery and connection.

[0084] After two subgroups establish a preliminary communication connection through periodic probing, communication radius expansion, or relay node deployment, the leaders of each subgroup begin exchanging detailed member lists and status information to prepare for subsequent subgroup merging. After confirming the establishment of a communication link with subgroup C, leader N2 of subgroup A sends a subgroup information request message to leader N10 of subgroup C. This message indicates that N2 wishes to negotiate a subgroup merge with N10 and requests complete information about subgroup C. Upon receiving the request, node N10 packages and encapsulates information such as the member list, topology, task status, and resource statistics of subgroup C into a subgroup information response message and sends it to N2 via relay node N5 or a direct communication link. The member list contains detailed information such as the identifier, location coordinates, remaining energy, and role of all healthy nodes in subgroup C. The topology describes the connection relationships between nodes within subgroup C, represented in the form of an adjacency matrix or edge list. The task status describes the type, progress, and priority of the task currently being executed by subgroup C. Resource statistics summarize the overall resource situation of subgroup C, including total remaining energy, available communication bandwidth, sensor configuration, etc.

[0085] After receiving information from subgroup C, navigator N2 of subgroup A performs a local feasibility assessment of the merge. The assessment includes: a topology compatibility check to verify whether the topologies of the two subgroups can be merged without conflict, such as checking for node identifier conflicts or communication frequency conflicts; a resource complementarity analysis to determine whether the merged cluster resource configuration is more balanced and whether it can improve overall task execution capabilities; and a task coordination assessment to confirm whether the task objectives of the two subgroups are consistent or coordinateable, avoiding task conflicts after the merge. If the assessment results indicate that the merge is feasible, node N2 sends a merge intention confirmation message to node N10, and both parties enter the formal subgroup merge process. Information exchange between subgroup navigators employs encryption and authentication mechanisms to prevent malicious nodes from forging subgroup information or tampering with exchanged data. Each subgroup navigator is pre-configured with a public and private key pair. Information exchange messages are digitally signed using the sender's private key, and the receiver verifies the validity of the signature using the sender's public key. If signature verification fails, the receiver rejects the message and requests a retransmission from the sender. This encryption and authentication mechanism can resist man-in-the-middle attacks and information forgery attacks, ensuring the security of the subgroup merging process.

[0086] After the subgroup leaders complete information exchange and confirm the merge intention, the subgroup merge process based on a distributed joint consensus mechanism is initiated. This mechanism borrows the idea of ​​joint consensus from the Raft protocol to ensure consistency in member change operations between the two subgroups during the subgroup merge process, avoiding inconsistencies where some nodes believe the merge is complete while others remain in the pre-merge state. The core of the joint consensus mechanism is the introduction of a transitional state in which members of both subgroups participate in decision-making simultaneously. Any changes involving cluster configuration require majority agreement in both subgroups to take effect. The leaders N2 of subgroup A and N10 of subgroup C negotiate and determine a joint configuration, Cold and Cnew, which includes the member sets of both subgroup A and subgroup C. In the joint configuration state, any proposal must obtain more than half of the votes in Cold (members of subgroup A) and Cnew (members of subgroup C). The leaders N2 of subgroup A and N10 of subgroup C broadcast the joint configuration message to all members within their respective subgroups. The message contains a detailed description of the joint configuration, including the merged member list, the new topology, and the merged leader election scheme. Upon receiving the message, members of subgroup A update their local configuration state, record the joint configuration information, and reply with a configuration confirmation message to leader N2. Members of subgroup C perform the same operation, replying with a confirmation message to leader N10. Once N2 and N10 have both collected confirmations from more than half of the members in their respective subgroups, the joint configuration takes effect simultaneously in both subgroups.

[0087] In a joint configuration, the navigator roles of the two subgroups are coordinated. Since the merged cluster can only have one global navigator, a navigator competition mechanism is used to determine the final navigator. This competition mechanism is still based on a comprehensive score. Navigator N2 of subgroup A and navigator N10 of subgroup C each calculate their own navigator score. The scoring formula is the same as the formula used during the navigator election within the subgroup, comprehensively considering factors such as node energy, communication capabilities, and subgroup size. Assuming N2's navigator score is 0.87 and N10's is 0.82, N2's score is higher. According to the competition rules, N2 becomes the global navigator of the merged cluster, and N10 transitions to a follower role. To ensure a smooth transition of the navigator role, a phased transition strategy is adopted. In the initial stage of the joint configuration, both N2 and N10 retain the navigator role, each responsible for coordination tasks within their respective atomic groups. Global decisions involving cross-subgroups require joint negotiation between N2 and N10. After a transition period, a new round of leader election voting is initiated. All members of the merged cluster participate in the vote, and a unique global leader is determined based on a comprehensive score. Assuming the voting confirms N2 as the global leader, N10 broadcasts a leader change notification to all members of subgroup C, informing them that N2 is the new global leader. Members of subgroup C update their local leader records, and all subsequent operations requiring leader coordination are submitted to N2. After confirming the global leader, the cluster exits the joint configuration state and enters the merged stable configuration state Cnew. The new configuration only includes the set of members of the merged cluster, and the decision-making mechanism reverts to the majority-agreement rule of a single cluster; any proposal only needs to obtain more than half of the agreement in Cnew to take effect. The global leader N2 broadcasts a stable configuration message to all members of the merged cluster, and each member updates its local configuration, completing the subgroup merge process.

[0088] After the merge, the previously isolated subgroups A and C form a unified reconstructed cluster. All nodes within the cluster can communicate via direct links or multi-hop paths, and the topology connectivity is fully restored. The global navigator N2 coordinates the task allocation, resource scheduling, and flight control of the reconstructed cluster, significantly enhancing the cluster's collaborative capabilities. If other unmerged subgroups exist in the system, the global navigator continues to initiate subgroup detection and merging processes, repeating the above steps until all subgroups are merged, forming a complete reconstructed cluster. The core objective of self-healing from cluster topology fractures is thus achieved.

[0089] S105: When the faulty node plays a critical role, a target healthy drone node is determined from the reconstructed cluster based on a distributed consensus algorithm as the replacement node, and the target healthy drone node is controlled to take over the task of the critical role.

[0090] In S105 above, after the UAV swarm completes topology reconstruction and sub-swarm merging, although network connectivity is restored, the downed and faulty nodes may have originally played critical roles. The absence of these roles will directly affect the swarm's mission execution capability and overall combat effectiveness. Traditional role redistribution methods usually rely on a central node for global scheduling or simply replace nodes according to a predetermined backup order. These methods are difficult to implement in a decentralized distributed architecture and cannot dynamically optimize role allocation based on the real-time status of nodes. When a faulty node plays a critical role, a target healthy drone node is selected from the reconstructed cluster as the replacement node based on a distributed consensus algorithm. This target healthy drone node then takes over the critical role's tasks. Specifically, when a faulty node plays a critical role, the subgroup leader broadcasts a role replacement request within the subgroup. This request carries the critical role's hardware payload requirements and task attribute parameters. Based on the current state parameters of each healthy drone node in the reconstructed cluster and the role replacement request, a corresponding bidding score is calculated, and a bidding response message carrying the bidding score is sent to the subgroup leader. The subgroup leader, based on distributed consensus, selects the target healthy drone node with the highest bidding score from all received bidding response messages as the replacement node and takes over the critical role's tasks.

[0091] Specifically, in a drone swarm system, different nodes are assigned different roles based on their hardware configuration, software functions, and mission responsibilities. Five key roles are predefined: navigator, communication relay node, reconnaissance node, attack node, and backup node. The navigator is responsible for flight path planning and formation control of the sub-swarm or the entire swarm, serving as the core node for swarm coordination. Communication relay nodes are deployed in specific spatial locations, responsible for forwarding messages between nodes or sub-swarms with long communication distances, maintaining swarm communication connectivity. Reconnaissance nodes are equipped with high-precision optical sensors, infrared sensors, or radar, responsible for intelligence gathering and situational awareness of target areas. Attack nodes carry weapon payloads, responsible for precision strikes against enemy targets. Backup nodes are reserved redundant resources that can quickly fill gaps when other key roles fail, improving the system's fault tolerance.

[0092] Each UAV node is assigned one or more role tags during system initialization or dynamic task allocation. Role tags are stored in the node's local configuration database and contain information such as role type identifier, role priority value, role capability requirements description, and role status flags. Role type identifiers are represented by enumerated values; for example, the navigator is identified as RLE, the relay node as RRE, the reconnaissance node as RSC, the attack node as RAT, and the backup node as RBA. The role priority value reflects the importance of the role to the cluster's task execution; a higher value indicates a higher priority. When multiple roles are simultaneously unavailable, the system processes role replacement requests in priority order. Based on the analysis of typical military mission scenarios, the navigator's priority is set to the highest level of 5, the relay and reconnaissance nodes' priorities are set to 4, the attack node's priority is set to 3, and the backup node's priority is set to 2.

[0093] Once the fault detection layer confirms that a node has been physically damaged through neighbor cross-validation, it further determines whether the faulty node plays a critical role. When a neighboring node that detected the fault reports the fault information to the subgroup leader, it includes the faulty node's identifier in the fault notification message. The subgroup leader maintains a complete subgroup member information table, which records the basic attributes of each member node, including node identifier, node type, hardware configuration, currently executing role, and node status. Upon receiving a fault notification message, the subgroup leader queries the member information table for the node's role label based on the faulty node's identifier. If the query result shows that the faulty node's role label list contains any of the five types of critical roles mentioned above, the subgroup leader determines that the fault event is a critical role missing event and needs to initiate a role succession process.

[0094] For example, in the reconstructed cluster, node N6, before the topology breakdown, assumed the role of a reconnaissance node. This node was equipped with a high-resolution visible light camera and an infrared thermal imager, responsible for real-time monitoring and image acquisition of ground targets in a designated area. When node N6 was shot down by enemy air defense fire, the neighbor cross-validation mechanism of the fault detection layer confirmed that N6 had been physically damaged. Node N3, a neighbor of node N6, sent a fault notification message to the subgroup leader N2. The message included the faulty node identifier IDN6 and the fault type identifier PDE. Subgroup leader N2 looked up the record corresponding to IDN6 in its local member information table and found that N6's role label was RSC, its role priority was 4, and its hardware payload included a high-resolution camera and an infrared thermal imager. According to the critical role definition rules, N2 determined that the damage to N6 was a critical role loss event and immediately triggered the takeover process for the reconnaissance node role.

[0095] After confirming the absence of a key role, the subgroup leader needs to broadcast the role replacement request to all healthy nodes within the subgroup, initiating the role bidding and consensus election process. The role replacement request message must fully describe the capability requirements and task constraints of the missing role, enabling each healthy node to accurately assess its suitability for the role. The role replacement request message is encapsulated in a structured data format. The message header includes a message type identifier (RRR), a sender identifier (i.e., the subgroup leader identifier), a message sequence number, and a message timestamp. The message body contains several key fields that detail the specific requirements for role replacement. The missing role type field indicates whether the role to be replaced is the leader, communication relay node, reconnaissance node, attack node, or backup node. The role priority field indicates the priority value of the role, used for coordinating the processing order when multiple roles are simultaneously missing. The hardware payload requirement field lists the hardware equipment necessary for the role to perform its tasks; for example, a reconnaissance node needs an optical sensor, an attack node needs weapon mounting capabilities, and a communication relay node needs a high-power communication module. The task attribute parameter field describes the task characteristics related to the role, including task type, target area coordinates, expected working duration, communication frequency requirements, etc. The soft constraint preference field gives non-mandatory preferred conditions, such as the expectation that the current position of the replacement node is close to the original working area of ​​the missing node, or the expectation that the remaining energy of the replacement node is relatively sufficient. These soft constraints are considered as bonus points in the bidding evaluation.

[0096] Continuing with the scenario where reconnaissance node N6 is missing, when subgroup leader N2 constructs the role replacement request message, it sets the missing role type to RSC and the role priority to 4. The hardware payload requirement field requires the replacement node to be equipped with at least one visual sensor, which can be any one of a visible light camera, infrared camera, or synthetic aperture radar. The sensor resolution must meet minimum requirements; for example, a visible light camera must have at least 12 megapixels, and an infrared camera must have a thermal sensitivity of at least 50 millikrvin. The mission attribute parameter field describes that the reconnaissance node needs to continuously monitor the ground area at coordinates 34.5 degrees North latitude and 112.3 degrees East longitude, with an expected working time of at least 30 minutes. The data transmission communication bandwidth requirement is at least 5 Mbps to support real-time image transmission. The soft constraint preference field indicates that nodes whose current location is closer to the reconnaissance target area are preferred, as are nodes with a remaining energy percentage higher than 60%. This shortens the time it takes for the node to travel to the reconnaissance location and ensures sufficient energy to complete the mission.

[0097] After constructing the role replacement request message, the subgroup navigator N2 broadcasts it to all healthy nodes through the subgroup's communication network. The broadcast uses a flooding mechanism; the message originates from the navigator node and is forwarded hop-by-hop to every node in the subgroup. Each node, upon receiving the broadcast message, forwards it to its neighbors until all healthy nodes have received it. To avoid infinitely looping and wasting communication resources, the system includes a Time-to-Live (TTL) field in the broadcast message. The TTL value is decremented by 1 with each forward, and the message is discarded when the TTL reaches 0. Simultaneously, each node maintains a cached list of received messages, determining whether a message has already been processed based on its sequence number. Duplicate messages are discarded to avoid duplicate processing and forwarding. Upon receiving the role replacement request message, each healthy drone node in the subgroup calculates its bidding score based on its current state parameters and the role requirements described in the message. The bidding score is a comprehensive evaluation indicator reflecting the node's overall ability and suitability to replace a missing role; a higher score indicates a more suitable node for the role.

[0098] Furthermore, based on the current state parameters and role succession requirements of each healthy drone node in the reconstructed cluster, the corresponding bidding score is calculated. Specifically, this includes: matching and verifying the hardware configuration of the healthy drone node to be bid on according to the hardware load requirements in the role succession requirements; the healthy drone node to be bid on can be any healthy drone node in the reconstructed cluster; if the matching verification fails, the bidding score of the healthy drone node to be bid on is determined to be zero; if the matching verification succeeds, the current state parameters of the healthy drone node to be bid on are obtained, and combined with the task attribute parameters in the role succession requirements, the target input variables are calculated. The target input variable parameters include remaining energy, relative proximity based on the task execution location, and the expected task load after superimposing the task attribute parameters. The process involves: 1) Mapping remaining energy, relative proximity, and expected task load to corresponding fuzzy sets based on a preset membership function to obtain the membership of the target input variable under different fuzzy states; 2) Performing fuzzy inference on each membership based on a preset fuzzy inference library to calculate the rule strength of each triggered rule. The preset fuzzy rule library contains multiple rules constructed based on conditional statements, and the rule strength is the minimum value of each membership in the corresponding rule condition; 3) Trunculating the output fuzzy set of the corresponding rule conclusion based on the rule strength, and performing a union operation on all truncated output fuzzy sets to obtain an aggregated output fuzzy set; 4) Defuzzifying the aggregated output fuzzy set using the centroid method, and using the abscissa of the geometric centroid of the aggregated output fuzzy set as the bidding score corresponding to the health drone node to be bid on.

[0099] Specifically, the hardware configuration of the bidding health drone nodes is matched and verified to ensure that the node possesses the necessary hardware equipment and capabilities to perform the target role. Hardware matching verification, as a hard constraint, employs a "one-vote veto" mechanism. If the node's hardware configuration does not meet the role requirements, its bidding score is directly determined to be zero, eliminating the need for subsequent fuzzy inference calculations, thereby improving computational efficiency and preventing unqualified nodes from participating in the bidding. The hardware payload requirement field in the role succession request message describes the various hardware devices required for the target role in the form of a structured list. The hardware payload requirement list contains multiple hardware items, each consisting of three attributes: hardware type, performance requirements, and mandatory flags. The hardware type identifies the category of the equipment, such as optical sensors, infrared sensors, synthetic aperture radar, weapon mounting systems, high-power communication modules, etc. The performance requirements describe the minimum performance indicators that this type of hardware must meet, listing the thresholds for each performance parameter in key-value pairs. For example, the performance requirements for optical sensors might include a minimum pixel resolution of 12 megapixels, a minimum field of view of 60 degrees, and a minimum frame rate of 15 frames per second. The "Required" flag uses a Boolean value to indicate whether the hardware is a mandatory requirement for the role. A true flag indicates that the hardware is required and the node must be equipped with it to pass the verification. A false flag indicates that the hardware is optional and the node can get bonus points for equipping the hardware, but not equipping it will not affect the verification passing.

[0100] The bidding health drone nodes maintain a complete hardware list in their local configuration database. The hardware list records all hardware devices mounted on the node and their detailed configuration parameters, including device type, model, performance parameters, and current operating status. The hardware matching and verification process iterates through each hardware item in the hardware payload requirement list, searches the local hardware list for matching devices, and compares the device performance parameters item by item. Let's continue with a detailed explanation using the scenario of a reconnaissance node role change as an example. The hardware payload requirement list in the role change request message contains two hardware items. Hardware item 1 is an optical sensor, with performance requirements including a minimum pixel resolution of 12 megapixels, a minimum field of view of 60 degrees, and a minimum frame rate of 15 frames per second. The mandatory flag is set to true, indicating that the optical sensor is a mandatory hardware requirement for the reconnaissance node. Hardware item 2 is an infrared sensor, with performance requirements including a minimum thermal sensitivity of 50 milliklvin and a minimum field of view of 50 degrees. The mandatory flag is set to false, indicating that the infrared sensor is optional hardware; equipping it can improve reconnaissance capabilities but is not mandatory.

[0101] After bidding node N5 receives the role replacement request message, it retrieves the hardware list information from its local configuration database. Node N5's hardware list shows it is equipped with a CAM-1300 visible light camera and an IR-THERMAL-40 infrared thermal imager. The visible light camera's performance parameters are 13 megapixel resolution, 75-degree field of view, and 30 frames per second; the infrared thermal imager's performance parameters are 40 milliklvin thermal sensitivity and 55-degree field of view. Node N5 initiates the hardware matching and verification process, iterating through the first hardware item in the hardware payload requirement list, namely the optical sensor item. The verification program searches for devices of type optical sensor in the local hardware list and finds the CAM-1300 visible light camera. The verification program compares the performance parameters item by item: the visible light camera's 13 megapixel resolution is greater than the required 12 megapixels, meeting the requirements; the 75-degree field of view is greater than the required 60 degrees, meeting the requirements; and the 30 frames per second frame rate is greater than the required 15 frames per second, meeting the requirements. All performance parameters of the optical sensor have been verified, and the mandatory flag for this hardware item is true. The verification program marks the optical sensor item as a successful match.

[0102] The verification program continues to iterate through the second hardware item, the infrared sensor item. It searches the local hardware list for devices of type infrared sensor and finds the infrared thermal imager IR-THERMAL-40. Comparing performance parameters, the infrared thermal imager's thermal sensitivity of 40 mKelvin is better than the required 50 mKelvin, because a lower thermal sensitivity value indicates a more sensitive sensor, thus meeting the requirement; the field of view of 55 degrees is greater than the required 50 degrees, also meeting the requirement. All performance parameters of the infrared sensor pass the verification, and the mandatory flag for this hardware item is false. The verification program marks the infrared sensor item as a successful match and records it as an optional hardware bonus item. After completing the iteration of all hardware items, the verification program tallies the matching results. All mandatory hardware items for node N5 are successfully matched, and the optional hardware items are also successfully matched. The final result of the hardware matching verification is determined to be passed.

[0103] Nodes that pass hardware matching verification enter the fuzzy inference calculation stage for their bidding scores. The fuzzy inference system requires three key target input variables: remaining energy, relative proximity, and expected task load. These three variables comprehensively reflect the node's capability reserves to take over the target role, its spatial location advantage, and current task pressure, and are the core factors affecting the bidding score. The remaining energy variable directly relates to whether the node can continuously perform its role. The node's battery management system monitors parameters such as battery pack voltage, current, and temperature in real time, calculates the current remaining charge using coulomb counting or voltage-capacity curve estimation, and converts it into a remaining charge percentage E. The remaining charge percentage is defined as the ratio of the current remaining charge to the battery's full capacity, ranging from 0% to 100%. The node reading the real-time value of E from the battery management system's status register serves as the first input variable for fuzzy inference.

[0104] The relative proximity variable assesses the spatial proximity between a node's current location and its assigned task execution location. The task attribute parameter field in the role replacement request message contains the coordinate information of the task execution location. For a reconnaissance node, the task execution location is the center coordinates of the reconnaissance target area; for a communication relay node, it is the ideal relay location between two subgroups; and for an attack node, it is the coordinates of the attack target or the attack initiation location. The bidding node extracts the three-dimensional coordinates (x1, y1, z1) of the task execution location from the role replacement request message, reads the three-dimensional coordinates (x2, y2, z2) of its current location provided by its own navigation system, and calculates the Euclidean distance D between the two points. This distance value is then converted into an index reflecting proximity, defined as relative proximity equal to a preset parameter distance constant divided by the sum of the Euclidean distance and the preset parameter distance constant. The preset reference distance constant is set to 100 meters. For example... The expected task load variable comprehensively considers the task load currently being executed by the node and the task load added after taking over the target role. The node's task management module maintains a task queue, recording all currently executing or pending tasks. Each task entry includes attributes such as task type, task priority, expected task duration, and required CPU utilization. The node traverses the task queue and calculates the current task load L, which is equal to the weighted sum of the CPU utilization of all tasks, with the weighting coefficient being the normalized value of the task priority. The task attribute parameter field in the role replacement request message contains a description of the target role's task load, including the expected CPU utilization (Load) and task priority (Priority) of that role's tasks. The calculation of the expected task load Le needs to comprehensively consider the current task load and the task load of the newly added role, and the calculation formula is Le = L + the normalized value of Load * Priority. The expected task load ranges from 0% to 200%, where 0% indicates the node is completely idle, 100% indicates the node is running at full load, and more than 100% indicates the node is overloaded.

[0105] For example, taking node N5 as an example, its battery management system reports a current remaining charge of 2880 mAh, with a full charge capacity of 4000 mAh. The remaining charge percentage E is calculated as 2880 divided by 4000 multiplied by 100%, resulting in 72%. The center coordinates of the reconnaissance target area in the role replacement request message are 34.5 degrees North latitude, 112.3 degrees East longitude, and 150 meters above sea level. The geographic coordinates are converted to a Cartesian coordinate system with a fixed reference point as the origin. The converted task execution location coordinates are (5000 meters, 3000 meters, 150 meters). The current location coordinates of node N5 are (4950 meters, 3065 meters, 120 meters). The calculated Euclidean distance D is equal to 87.3 meters. The preset reference distance is set to 100 meters, and the relative proximity is calculated to be 0.534. There is currently a low-priority routine patrol task in node N5's task queue, with a CPU utilization of 15% and a task priority of 2. Since there is only one task, the current task load L is calculated as a normalized value of 15% multiplied by 2. Assuming the priority normalization range is 1 to 5, the normalized value of 2 is 2 divided by 5, which equals 0.4. Therefore, L equals 15% multiplied by 0.4, which equals 6%. In the role replacement request message, the expected CPU utilization (Load) of the reconnaissance node role is 20%, the task priority is 4, and the normalized value is 4 divided by 5, which equals 0.8. The expected task load Le equals 22%.

[0106] The core of a fuzzy inference system lies in converting precise numerical inputs into fuzzy linguistic descriptions, achieved through a membership function. The membership function defines the degree to which an input variable's value belongs to a certain fuzzy set, ranging from 0 to 1. 0 indicates no membership in the fuzzy set, 1 indicates full membership, and values ​​between 0 and 1 represent partial membership. For the remaining energy input variable, three fuzzy sets are defined, named low, medium, and high. These three fuzzy sets describe the different levels of a node's remaining energy in a linguistic way. A low fuzzy set indicates that the node has little remaining energy, potentially insufficient to support long-term task execution; a medium fuzzy set indicates that the node's remaining energy is at a moderate level, sufficient to support task execution for a certain duration; and a high fuzzy set indicates that the node has sufficient remaining energy, enabling continuous task execution without worry.

[0107] The membership function of the remaining energy adopts a piecewise linear function form of triangle or trapezoid. The membership function μLow(x) of the low fuzzy set is defined as follows: when the remaining energy x is less than or equal to 0%, the membership is 1; when x is between 0% and 40%, the membership decreases linearly, and the calculation formula is μLow(x) equal to (40 minus x) divided by 40; when x is greater than or equal to 40%, the membership is 0. The membership function μMedium(x) of a medium fuzzy set is defined as follows: when x is less than 20%, the membership is 0; when x is between 20% and 50%, the membership increases linearly, calculated as μMedium(x) = (x - 20) / 30; when x is between 50% and 80%, the membership remains 1; when x is between 80% and 100%, the membership decreases linearly, calculated as μMedium(x) = (100 - x) / 20; and when x is greater than 100%, the membership is 0. The membership function μHigh(x) of a high fuzzy set is defined as follows: when x is less than 60%, the membership is 0; when x is between 60% and 100%, the membership increases linearly, calculated as μHigh(x) = (x - 60) / 40; and when x is greater than or equal to 100%, the membership is 1.

[0108] For the remaining energy of node N5, which is 72%, its membership degree on the three fuzzy sets is calculated. The membership degree μLow(72%) of the low fuzzy set is 0 by definition, since 72% is greater than 40%. The membership degree μMedium(72%) of the medium fuzzy set is 1 by definition, since 72% is between 50% and 80%. The membership degree μHigh(72%) of the high fuzzy set is 0 by definition, since 72% is between 60% and 100%. The calculated μHigh(72%) equals (72 - 60) divided by 40, which equals 12 divided by 40, which equals 0.3. The remaining energy of node N5 has a membership degree of 0 on the low fuzzy set, 1 on the medium fuzzy set, and 0.3 on the high fuzzy set. The membership degree calculation results indicate that the remaining energy of node N5 belongs entirely to the medium level, partially to the high level, and not to the low level.

[0109] For the relative proximity input variable, three fuzzy sets were defined, named far, medium, and near. The far fuzzy set indicates that the node is far from the task execution location and requires a longer time to reach; the medium fuzzy set indicates that the node is at a moderate distance; and the near fuzzy set indicates that the node is very close to the task execution location and can be reached quickly. The membership function of relative proximity also adopts a piecewise linear form. The membership function μFar(x) of the far fuzzy set is defined as follows: when the proximity x is less than or equal to 0, the membership degree is 1; when x is between 0 and 0.3, the membership degree decreases linearly, calculated as μFar(x) equal to (0.3 minus x) divided by 0.3; and when x is greater than or equal to 0.3, the membership degree is 0. The membership function μMedium(x) of a medium fuzzy set is defined as follows: when x is less than 0.2, the membership is 0; when x is between 0.2 and 0.5, the membership increases linearly, and the formula is μMedium(x) = (x - 0.2) divided by 0.3; when x is between 0.5 and 0.7, the membership remains 1; when x is between 0.7 and 0.9, the membership decreases linearly, and the formula is μMedium(x) = (0.9 - x) divided by 0.2; when x is greater than 0.9, the membership is 0. The membership function μClose(x) of a near fuzzy set is defined as follows: when x is less than 0.6, the membership is 0; when x is between 0.6 and 1.0, the membership increases linearly, and the formula is μClose(x) = (x - 0.6) divided by 0.4; when x equals 1.0, the membership is 1.

[0110] For node N5, with a relative proximity value of 0.534, its membership degree on the three fuzzy sets is calculated. The membership degree μFar (0.534) in the far fuzzy set is 0 by definition, since 0.534 is greater than 0.3. The membership degree μMedium (0.534) in the medium fuzzy set is 1 by definition, as 0.534 falls between 0.5 and 0.7. The membership degree μClose (0.534) in the near fuzzy set is 0 by definition, since 0.534 is less than 0.6. The relative proximity of node N5 is 0 in the far fuzzy set, 1 in the medium fuzzy set, and 0 in the near fuzzy set, indicating that the node's position is at a medium proximity level.

[0111] For the expected task load input variable, the system defines three fuzzy sets, named low, medium, and high. The low fuzzy set indicates that the node's current task load is relatively light, with sufficient computing resources to take over the new role; the medium fuzzy set indicates that the node's task load is at a moderate level; and the high fuzzy set indicates that the node's task load is relatively heavy, and taking over the new role may lead to resource strain. The membership function of the expected task load is defined as follows: The membership function μLow(x) of the low fuzzy set is defined as follows: when the load x is less than or equal to 0%, the membership degree is 1, meaning the node is completely idle; when x is between 0% and 20%, the membership degree remains 1; when x is between 20% and 50%, the membership degree decreases linearly, calculated as μLow(x) equal to (50 minus x) divided by 30; when x is greater than or equal to 50%, the membership degree is 0. The membership function μMedium(x) of a medium fuzzy set is defined as follows: when x is less than 30%, the membership is 0; when x is between 30% and 50%, the membership increases linearly, calculated as μMedium(x) = (x - 30) / 20; when x is between 50% and 80%, the membership remains 1; when x is between 80% and 100%, the membership decreases linearly, calculated as μMedium(x) = (100 - x) / 20; and when x is greater than 100%, the membership is 0. The membership function μHigh(x) of a high fuzzy set is defined as follows: when x is less than 70%, the membership is 0; when x is between 70% and 100%, the membership increases linearly, calculated as μHigh(x) = (x - 70) / 30; and when x is greater than or equal to 100%, the membership remains 1, indicating that the node is overloaded.

[0112] For node N5, with an expected workload of 22%, its membership degrees on the three fuzzy sets are calculated. The membership degree μLow(22%) of the low fuzzy set, by definition, is between 20% and 50%, so μLow(22%) equals (50 - 22) divided by 30, which is 28 divided by 30, approximately 0.933. The membership degree μMedium(22%) of the medium fuzzy set, by definition, is 0 since 22% is less than 30%. The membership degree μHigh(22%) of the high fuzzy set, by definition, is 0 since 22% is less than 70%. The expected workload of node N5 has a membership degree of 0.933 on the low fuzzy set, 0 on the medium fuzzy set, and 0 on the high fuzzy set, indicating that the node's workload is at a low level. Through the calculation of the membership function, the three precise input variables of node N5 are converted into nine membership values: the membership of the remaining energy on the low, medium, and high fuzzy sets (0, 1, 0.3), the membership of the relative proximity on the far, medium, and near fuzzy sets (0, 1, 0), and the membership of the expected task load on the low, medium, and high fuzzy sets (0.933, 0, 0). These membership values ​​serve as the basis for fuzzy inference and will be used for inference calculation through the fuzzy rule base in the next step.

[0113] The fuzzy rule base incorporates expert experience and domain knowledge, describing the mapping relationship between input and output variables in the form of conditional statements. Each fuzzy rule consists of an IF-THEN structure, where the IF part is the antecedent or condition of the rule, describing the state combination of the input variables on various fuzzy sets, and the THEN part is the consequent or conclusion of the rule, describing the fuzzy set to which the output variable should belong. For the bidding score calculation task, the output variable is defined as fitness, representing the overall degree of fit of a node in replacing the target role. The universe of discourse for the fitness variable ranges from 0 to 1, where 0 represents no fit and 1 represents a perfect fit. The fitness variable defines five output fuzzy sets, named poor, average, good, excellent, and outstanding. These five fuzzy sets describe different levels of node fitness in a linguistic way. The output fuzzy sets of the fitness variable are also defined using membership functions. The membership function of poor fuzzy sets covers the fitness range of 0 to 0.3, general fuzzy sets cover 0.2 to 0.5, good fuzzy sets cover 0.4 to 0.7, excellent fuzzy sets cover 0.6 to 0.9, and superb fuzzy sets cover 0.8 to 1.0. The membership function of each fuzzy set adopts a triangular or trapezoidal shape to ensure that any point in the entire universe of discourse belongs to at least one fuzzy set. The fuzzy rule base contains multiple rules covering the decision logic of input variables under different combinations of fuzzy sets. The construction of rules follows expert experience and common-sense reasoning. For example, when a node has high remaining energy, is geographically proximate, and has low task load, the node is most suitable to replace the target role, and the corresponding fitness should be superb; when a node has low remaining energy, regardless of its location and load, it is not suitable to replace the role, and the corresponding fitness should be poor.

[0114] The fuzzy rule base includes the following rules: Rule 1 states: IF remaining energy is high AND relative proximity is near AND expected task load is low THEN fit is excellent. Rule 2 states: IF remaining energy is high AND relative proximity is medium AND expected task load is low THEN fit is good. Rule 3 states: IF remaining energy is medium AND relative proximity is near AND expected task load is low THEN fit is good. Rule 4 states: IF remaining energy is low THEN fit is poor. This rule uses a veto logic; if the remaining energy is low, regardless of other conditions, the fit is directly judged as poor. Rule 5 states: IF expected task load is high THEN fit is average, indicating that when the task load is too high, the node is not suitable to take over the new role. Based on the combination of three fuzzy sets for each of the three input variables, theoretically 27 rules can be constructed to cover all possible input states. In practical applications, a representative and meaningful subset of rules is selected to construct the rule base, typically with 15 to 20 rules.

[0115] The fuzzy inference process employs the Mamdani inference method, which is widely used in industrial control and decision-making systems. The inference process includes four steps: rule matching, rule strength calculation, output fuzzy set truncation, and aggregation. In the rule matching stage, it is determined whether the antecedent conditions of each rule are met, i.e., whether the membership degrees of the input variables trigger the fuzzy set combination described in the rule. In the rule strength calculation stage, the activation degree of the triggered rules is calculated. Rule strength is defined as the minimum value of the membership degrees of each input variable in the rule's antecedents, using a minimum value operation with AND logic. In the output fuzzy set truncation stage, the output fuzzy set in the rule conclusion is truncated according to the rule strength. The truncated fuzzy set is vertically restricted to the height of the rule strength. In the aggregation stage, the truncated output fuzzy sets of all triggered rules are unioned to obtain the final aggregated output fuzzy set.

[0116] Taking the membership degree of node N5 as an example, a detailed calculation of fuzzy inference is performed. The membership degree of node N5's remaining energy is low (0), medium (1), and high (0.3); its membership degree of relative proximity is far (0), medium (1), and near (0); and its membership degree of expected task load is low (0.933), medium (0), and high (0). During the rule matching phase, all rules in the rule base are traversed. The antecedent of rule 1 is "remaining energy is high AND relative proximity is near AND expected task load is low," with corresponding membership degrees of 0.3, 0, and 0.933, respectively. The rule strength is calculated as min(0.3, 0, 0.933), which equals 0. A rule strength of 0 indicates that the rule has not been effectively triggered and will not participate in subsequent inference. Rule 2's antecedents are high remaining energy, medium relative proximity, and low expected task load, with corresponding membership degrees of 0.3, 1, and 0.933 respectively. The rule strength is calculated as min(0.3, 1, 0.933) = 0.3, so the rule is triggered with a strength of 0.3, and the rule conclusion is excellent fit. Rule 3's antecedents are medium remaining energy, near relative proximity, and low expected task load, with corresponding membership degrees of 1, 0, and 0.933 respectively. The rule strength is calculated as min(1, 0, 0.933) = 0, so the rule is not triggered. Rule 4's antecedent is low remaining energy, with a corresponding membership degree of 0, and a rule strength of 0, so the rule is not triggered. Continuing to iterate through other rules, suppose there is rule 6 which is expressed as: IF Remaining energy is medium AND Relative position proximity is medium AND Expected task load is low THEN Adaptability is good. The antecedents of this rule have membership degrees of 1, 1, and 0.933, respectively. The rule strength is calculated as min(1, 1, 0.933) equals 0.933. The rule is triggered, the rule strength is 0.933, and the rule conclusion is that the adaptability is good.

[0117] After rule matching and rule strength calculation, two rules were triggered during the inference process of node N5: Rule 2, with a rule strength of 0.3, concluded that the fit was excellent; Rule 6, with a rule strength of 0.933, concluded that the fit was good. The output fuzzy set truncation stage processes the fuzzy sets of these two rules. Rule 2's conclusion of excellent fit is defined within the fit range of 0.6 to 0.9, assuming the peak value of the triangular membership function is 0.75 and the peak membership is 1. Based on the rule strength of 0.3, truncating the excellent fuzzy set results in a trapezoidal region with a vertical height of 0.3, where the base corresponds to the fit range. Rule 6's conclusion of good fit is also addressed by the good fuzzy set with a peak membership function of 0.55, truncated with a rule strength of 0.933, resulting in a vertical height of 0.933 for the good fuzzy set.

[0118] The aggregation phase performs a union operation on the two truncated output fuzzy sets. The union operation takes the maximum membership degree of the two fuzzy sets at each fitness value point, forming a new aggregated output fuzzy set. The aggregated output fuzzy set is dominated by the good fuzzy set in the fitness range of 0.4 to 0.7, with a membership degree height of 0.933. In the fitness range of 0.6 to 0.9, it is partially influenced by the excellent fuzzy set, but because the good fuzzy set has higher rule strength, the aggregation result mainly reflects the characteristics of a good level. The aggregated output fuzzy set is a fuzzy regional distribution and cannot be directly used for precise comparison and ranking of bidding scores. It needs to be converted into a definite numerical value through a defuzzification method. There are various defuzzification methods, including the centroid method, the maximum membership degree method, and the weighted average method. This embodiment uses the centroid method. The centroid method treats the aggregated output fuzzy set as a geometric figure on a two-dimensional plane, with the horizontal axis representing the fitness variable value and the vertical axis representing the membership degree. The figure is represented by the area region enclosed by the membership function curve and the horizontal axis. The centroid method is used to calculate the position of the centroid of the geometric figure. The x-coordinate of the centroid is used as the output value after defuzzification, which is the bid score BidV. The integral ranges from 0 to 1 across the entire universe of discourse of the fitness variable, and u(x) is the membership degree of the aggregated output fuzzy set at the fitness value x. Since the membership function is usually in piecewise linear form, the integral can be numerically calculated by piecewise summation. The fitness universe is divided into several discrete sampling points, for example, one sampling point is taken every 0.01, for a total of 101 sampling points. At each sampling point xi, the membership degree u(xi) of the aggregated output fuzzy set is calculated, the product of u(xi) and xi is accumulated to obtain the numerator summation term, the product of u(xi) and xi is accumulated to obtain the denominator summation term, and finally the numerator is calculated by dividing the denominator to obtain the barycentric abscissa. For the aggregated output fuzzy set of node N5, assuming that the numerator summation term is 32.15, the denominator summation term is 54.8, and the barycentric abscissa BidV is calculated as 32.15 divided by 54.8, which is approximately equal to 0.587. The bidding score of node N5 is determined to be 0.587.

[0119] After the bidding node completes the fuzzy inference calculation of the bidding score, the calculation result needs to be validated to ensure that the score is within a reasonable range and that the calculation process is without anomalies. The validation of the bidding score includes two aspects: numerical range checking and logical consistency checking. Numerical range checking verifies whether the bidding score is within the expected universe of discourse. Since the universe of discourse for the adaptability variable is defined as 0 to 1, the bidding score obtained from the defuzzification calculation should theoretically fall between 0 and 1. It checks whether the calculation result satisfies 0 ≤ BidV ≤ 1. If the value exceeds the range, it indicates an anomaly in the calculation process, such as an incorrect membership function definition, improper fuzzy rule base configuration, or insufficient numerical integration precision. The system records the error log and sets the bidding score to the default value of 0 to prevent abnormal data from participating in subsequent bidding decisions. Logical consistency checking verifies whether the qualitative relationship between the bidding score and the input variables conforms to common sense. A set of consistency rules is defined, such as a bid score of less than 0.3 when remaining energy is less than 20% of the threshold; a bid score of less than 0.4 when the expected task load is greater than 80% of the threshold; and a bid score of less than 0.5 when relative proximity is less than 0.2 and remaining energy is less than 50%. The system determines which consistency rules should be met based on the input variable values ​​and then checks whether the bid score violates these rules. If a logical inconsistency is found, the system records a warning message and can choose to adjust the bid score or keep the original value but mark it as a suspicious result. After successful verification, the calculation result is encapsulated into a bid response message and sent to the subgroup leader. The design of the bid response message needs to ensure that the message can be delivered to the leader in a timely and accurate manner, while avoiding communication conflicts caused by multiple nodes sending messages simultaneously.

[0120] The data structure of the bidding response message includes a message header and a message body. The message header contains a message type identifier, a sender identifier (i.e., the identifier of the bidding node), a target receiver identifier (i.e., the identifier of the subgroup leader), a message sequence number, and a message timestamp. The message body contains a bidding score field and a detailed capability description field, listing detailed information such as the node's hardware configuration, remaining energy, current location, and current task status, for the leader to perform secondary verification and decision-making reference. To avoid communication channel congestion and message collisions caused by multiple nodes sending bidding response messages simultaneously, a random delay mechanism is introduced. After calculating the bidding score, each node does not immediately send a bidding response message but starts a random delay timer. The random delay time is evenly distributed within a preset time range, for example, randomly selecting a value between 0 and 500 milliseconds. The node waits for the random delay time to expire before sending the bidding response message to the subgroup leader. This random delay mechanism effectively disperses the message sending times, reduces the probability of communication collisions, and improves the message transmission success rate.

[0121] Bidding response messages are sent to the subgroup leader via the routing protocol within the subgroup. If a direct communication link exists between the bidding node and the leader, the message is transmitted to the leader in a single hop. If there is no direct link between the bidding node and the leader, the message is forwarded via multi-hop routing. A topology-based shortest path routing algorithm is used. Based on the global topology graph maintained by the topology reconstruction layer, the shortest path from the bidding node to the leader is calculated, and the bidding response message is forwarded hop-by-hop along this path. To ensure the reliability of message transmission, a reliable transmission mechanism based on acknowledgment and retransmission is implemented at the link layer. After sending a message, the sending node starts a timeout timer, waiting for the receiver's acknowledgment message (ACK). If no ACK is received within the timeout period, a retransmission is performed, with a maximum of 3 retransmissions. If the transmission still fails after 3 retransmissions, a transmission failure is reported to the upper layer.

[0122] After broadcasting the role succession request message, the subgroup leader initiates the bidding response collection phase. The leader sets a bidding collection timeout, typically ranging from 1 to 5 seconds, with the specific value configured based on the subgroup size and communication latency. Within the timeout, the leader continuously receives bidding response messages from each healthy node and stores them in a bidding response cache list. When the timeout expires, the leader stops receiving new bidding responses and analyzes and makes decisions on all collected responses. The bidding collection timeout needs to strike a balance between decision-making speed and participation breadth. The leader's processing of bidding response messages includes two steps: validity verification and bidding score sorting. Validity verification checks the completeness and legality of the bidding response messages, including: whether the message format conforms to the protocol specifications, whether the sender's identifier belongs to a legitimate member within the subgroup, whether the bidding score is within a reasonable range, and whether the hardware configuration in the detailed capability description meets the mandatory requirements of the role succession requirement. If a bid response message fails the validity verification, the message is marked as invalid and removed from the cache list, and will not participate in subsequent bidding decisions.

[0123] Under a distributed consensus mechanism, the final decision on role succession cannot be made solely by the subgroup leader; consensus must be reached within the subgroup to avoid decision-making errors due to leader node failure or misjudgment. The leader encapsulates the bidding score ranking results into a succession candidate proposal message. This message contains the identifier of the recommended successor node, its bidding score, and a detailed description of its capabilities. The leader broadcasts the succession candidate proposal message to all healthy nodes in the subgroup, requesting a vote. The voting on the succession candidate proposal follows a simple majority rule. Upon receiving the proposal message, each healthy node in the subgroup independently assesses the information within it. If it deems the recommended successor node suitable for the scout node role, it votes in favor; otherwise, it votes against or abstains. Each node encapsulates its voting results into a voting message and sends it to the subgroup leader. The leader collects all voting messages and counts the number of votes in favor. If the number of votes in favor exceeds half of the total number of healthy nodes in the subgroup, the proposal passes, and the recommended successor node is officially confirmed as the successor node for the scout node role; if the number of votes in favor does not reach half, the proposal is rejected, and the leader needs to restart the bidding process or select the node with the second highest bidding score as a new candidate proposal for voting.

[0124] After the subgroup leader determines the successor node through distributed consensus, it sends a role handover instruction message to that node, officially initiating the role transition process. The role handover instruction message contains complete configuration information for the new role. The successor node adjusts its hardware and software operating modes according to this configuration information, receives the task context data of the missing role, and begins executing the tasks of the new role. The data structure of the role handover instruction message includes a message header and a message body. The message header contains a message type identifier, a sender identifier (i.e., the subgroup leader identifier), a target receiver identifier (i.e., the successor node identifier), a message sequence number, and a message timestamp. The message body contains several key configuration fields. The new role type field specifies the role the successor node will assume; for example, if the key role is a reconnaissance node role. The role priority field is set to 4, indicating the importance of the role. The task context data field contains detailed parameters for the reconnaissance mission, including the geographical coordinates of the reconnaissance target area, the characteristic description of the reconnaissance target, the required data acquisition frequency, the required image resolution, and the target node or ground station address for data return. The flight control parameter field contains control commands such as flight path planning, flight altitude, flight speed, and hovering position for the successor node when performing the reconnaissance mission. The communication configuration parameters field includes communication protocol configurations such as the communication frequency, modulation method, and encryption key used for data return. The time constraint field specifies the deadline for role handover and the start time for task execution.

[0125] For example, subgroup navigator N2 sends a role handover instruction message to the assuming node N5. The message body details the reconnaissance mission requirements: the target area is a circular area with a radius of 500 meters, located at 34.5 degrees North latitude and 112.3 degrees East longitude; it is necessary to identify the type, quantity, and location of ground fixed facilities and mobile vehicles; take a 13-megapixel JPEG photo every 5 seconds with 90% compression quality; transmit data back to the ground command station via UDP protocol; fly at an altitude of 150 meters, a speed of 8 meters per second, hover upon arrival, and maintain a horizontal position error of less than 5 meters; communication uses 2.4 GHz band OFDM modulation and a pre-configured encryption key; preparation must be completed within 60 seconds, and arrival at the target area must be within 120 seconds. After receiving the instruction, node N5's mission management module adds the reconnaissance node role to its local role list, updating the priority to 4. The flight control module plans the flight path from the current position to the target area, calculates waypoint coordinates and velocity curves, and configures the autopilot. The sensor management module activates the visible light and infrared cameras, sets the continuous shooting mode to 5-second intervals, a resolution of 13 megapixels, and configures a real-time compressed transmission channel. The communication module switches to the 2.4GHz frequency, configures OFDM modulation, loads the encryption key, and establishes an encrypted link with the ground command station. After preparation, N5 sends a role handover confirmation message to N2. Upon receiving the confirmation, N2 broadcasts a role update notification to the subgroup, and each node updates its member information table, marking N5 as a reconnaissance node. N5 then initiates its flight mission, with the navigation module continuously monitoring position deviation and adjusting in real time. After approximately 15 seconds of flight, N5 reaches a hovering position above the target, at coordinates 34.5003 degrees North latitude and 112.2998 degrees East longitude, with a horizontal deviation of 3.2 meters, meeting the accuracy requirements. The flight control module activates the position-holding algorithm, compensating for wind interference by adjusting the rotor speed. After hovering, the sensor management module initiates the reconnaissance mission execution.

[0126] In certain scenarios, missing key roles may involve cross-subgroup collaboration. For example, a global communication relay node might be responsible for forwarding messages between multiple subgroups, or a global navigator might be responsible for coordinating the collaborative flight of multiple subgroups. The succession of such cross-subgroup roles cannot be completed within a single subgroup; it requires coordination from multiple subgroup navigators to determine the optimal successor node through global consensus. The trigger condition for cross-subgroup role succession is that the scope of responsibility of the missing role involves multiple subgroups. When constructing a role succession request message, the subgroup navigator queries the role attribute database to determine whether the role is a cross-subgroup role. If the scope in the role attribute is marked as GLOBAL (global), the navigator determines that a cross-subgroup role succession process needs to be initiated. The navigator not only broadcasts the role succession request message within its own subgroup but also forwards it to navigators in other subgroups via inter-subgroup communication links.

[0127] In scenarios involving large-scale node failures, multiple critical roles may be simultaneously absent. To improve the efficiency of role succession, the system supports parallel processing of multiple role succession requests. Each bidding process is initiated according to role priority, but they do not block each other during execution, allowing for simultaneous bidding and consensus decision-making for multiple roles. The trigger condition for multi-role succession is that the fault detection layer confirms the failure of multiple critical role nodes within a short period of time. The subgroup leader maintains a role absence queue locally, sorting all detected missing roles from high to low priority and triggering the role succession process sequentially. A role batch identifier field is added to the role succession request message to distinguish different role succession processes and avoid message confusion between processes. The bidding collection, consensus decision-making, and role handover stages of the two role succession processes are executed in parallel without blocking each other. Nodes can participate in bidding for multiple roles simultaneously, calculating different bidding scores for different roles based on their own capabilities and resource constraints. In the distributed consensus voting phase, each node votes on the reconnaissance node succession proposal and the attack node succession proposal respectively based on the role batch identifier.

[0128] After the role handover is completed, the handover effect is verified to ensure that the replacement node can perform the duties of the new role normally. The operational status of the replacement node is continuously monitored to promptly detect and handle any anomalies. Verification of the role handover effect includes two aspects: functional verification and performance verification. Functional verification checks whether the replacement node can perform all the functions required by the new role. For example, whether a reconnaissance node can normally capture and transmit images, whether an attack node can lock onto targets and execute attack commands, and whether a communication relay node can normally forward messages. After the replacement node sends a role handover confirmation message, the subgroup leader sends a functional test command to the replacement node, requiring it to perform a role-related test operation and report the test results to the leader. Performance verification checks the performance indicators of the replacement node when performing role tasks, such as whether the image acquisition frequency meets the requirements, whether the data transmission latency is within an acceptable range, and whether the message forwarding success rate of the communication relay meets the threshold. The subgroup leader collects the operation logs and performance statistics of the replacement node, calculates the actual values ​​of each performance indicator, and compares them with the performance thresholds required for the role. If the actual performance indicators meet or exceed the threshold requirements, the performance verification is successful; if some indicators fail to meet the requirements, the navigator analyzes the reasons, which may be due to insufficient hardware capabilities of the replacement node or a decline in the quality of the communication link, requiring adjustments or re-initiating the role replacement process.

[0129] S106: After confirming that the network topology reconstruction and role succession are completed, based on the hierarchical recursive self-healing strategy, guide the reconstructed cluster to recover from network consistency to physical consistency, so as to restore the preset formation configuration.

[0130] In S106 above, after completing network topology reconstruction and key role handover, although the drone swarm logically restores communication connectivity, the positions, speeds, and flight attitudes of each node in physical space remain chaotic. Network consistency only ensures information transmission and consensus, but does not translate into coordinated movement in physical space. During topology disruption, nodes may deviate from their original formation positions to maintain communication or avoid obstacles, and some nodes move to new positions to perform role handover tasks, breaking the formation configuration.

[0131] The level generation process starts automatically after network topology reconfiguration. The subgroup leader node sets its own level to Level = 1 and broadcasts a level initialization message to its one-hop neighbors. This message includes the node identifier, level value, and timestamp. Upon receiving the message, neighboring nodes, if they haven't yet set a level, set their own level to the sender's level plus 1 and continue broadcasting to their neighbors. Level information spreads hop-by-hop from the leader, forming a hierarchical distribution structure. When multiple source nodes compete, a minimum level priority mechanism is used. When a node receives multiple level messages, it selects the message source with the smallest level value as its superior and calculates its own level accordingly. If messages of the same level are received, arbitration is conducted using the node identifier. The level diffusion terminates when all nodes have been assigned a level and no updates occur within the time window.

[0132] Taking a 10-node subgroup as an example, the leader N2 sets its level to 1 and broadcasts it to its neighbors N5, N6, and N7. Upon receiving the message, N5, N6, and N7 set their levels to 2 and propagate the message to their respective neighbors. N8 receives N5's message and sets its level to 3, N9 receives N6's message and sets its level to 3, and N10 receives N7's message and sets its level to 3. The level 3 nodes continue propagating outwards, and N1, N3, and N4 eventually set their levels to 4. The entire process is completed in approximately 150 milliseconds, forming a clear four-layer hierarchy.

[0133] After the hierarchical distribution is established, a tiered repair path is constructed based on the hierarchical information. The preset formation configuration uses the geometric center as the reference origin, with higher-level nodes assigned to core positions. Taking a V-formation as an example, the level 1 navigator N2 is located at the apex of the V, level 2 nodes N5, N6, and N7 are located on either side of the second row, and level 3 and level 4 nodes are arranged sequentially in subsequent positions. The repair path adopts a hierarchical recursive strategy, expanding sequentially from higher-level nodes. Navigator N2 calculates the deviation vector between its current position and the target formation center, planning a flight trajectory that satisfies dynamic constraints. A fifth-order polynomial trajectory generation method is used to ensure that the velocity acceleration at the start and end points is zero, guaranteeing smooth start and stop. While initiating trajectory tracking, N2 broadcasts formation recovery instructions to lower-level nodes N5, N6, and N7. After receiving the instructions, level 2 nodes calculate their own target positions and plan their trajectories. A virtual spring-damping model is introduced into the trajectory planning, modeling the desired relative positional relationship with higher-level nodes as a virtual spring connection, generating an attractive force to ensure that the relative position remains within a reasonable range. Level 2 nodes initiate trajectory tracking and converge to the target position after approximately 15 seconds of flight, forming the core framework of the formation. Level 3 nodes are triggered based on the stable state of Level 2 nodes. The system sets convergence criteria: when a node's position deviation is less than 2 meters and its speed is close to its cruising speed, convergence is determined, and a command is broadcast to the next level. Level 3 nodes plan their trajectories with reference to the positions of their superior nodes, introducing an artificial potential field method to avoid collisions with peer nodes. Level 4 nodes are initiated after Level 3 convergence, completing the outermost position adjustment. The entire graded repair process exhibits a hierarchical characteristic, with a total duration of approximately 50 to 60 seconds.

[0134] After a node approaches its target location, a distributed clustering method is used to achieve precise formation and speed synchronization. This method includes three basic rules: The aggregation rule is implemented based on the relative position information of the node and its neighbors. Node Ni maintains a neighbor list containing the identifiers and location information of all nodes within its communication range. Node Ni calculates the geometric center position of its neighbors using the formula Pc = (1 / n) * the sum of the positions Pj of all neighbors j, where n is the number of neighbors. The aggregation force Fc points towards the geometric center. The magnitude of Pc is proportional to the distance from node Ni's current position Pi to the geometric center, calculated using the formula Fc = kc * (Pc - Pi), where kc is the aggregation force gain coefficient, typically ranging from 0.5 to 2.0.

[0135] The alignment rule is implemented based on the speed information of a node and its neighbors. Node N_i calculates the average speed of its neighbors using the formula Va = (1 / n) * the sum of the speeds Vj of all neighbors j. The alignment force Fa points in the direction of the difference between the average speed Va and the current speed Vi, and its magnitude is proportional to the speed difference. The formula is Fa = ka * (Va - Vi), where ka is the alignment force gain coefficient, typically ranging from 1.0 to 3.0. The purpose of the alignment rule is to eliminate speed differences between nodes, ensuring that the entire cluster moves in the same direction at the same speed, achieving global speed consistency.

[0136] The separation rule is implemented based on the distance information between a node and its neighbors. Node Ni traverses the neighbor list. For each neighbor j, the distance dij between them is calculated as equal to the Euclidean norm of Pi minus Pj. If dij is less than the safe distance threshold dsafe, a repulsive force Fsj is generated. The direction of Fsj is a unit vector pointing from neighbor j to node Ni, and its magnitude is proportional to the reciprocal of the distance. The formula is Fsj = ks multiplied by (dsafe minus dij) divided by the square of dij, multiplied by the unit vector (Pi minus Pj) divided by dij, where ks is the separation force gain coefficient, typically ranging from 5.0 to 10.0. Node Ni sums the repulsive forces of all neighbor j that generated the repulsive force, obtaining the total separation force Fs, which is equal to the sum of all Fsj. The purpose of the separation rule is to prevent collisions caused by nodes being too close together and to maintain a safe distance between nodes in the cluster.

[0137] The total control force Ftotal of a node is obtained by a weighted sum of three rules, with the weight coefficients dynamically adjusted according to the formation phase. Initially, a higher aggregation weight encourages nodes to quickly converge; as they approach their configuration, the alignment weight increases to enhance speed synchronization; and in dense areas, the separation weight increases to avoid collisions. After saturation limiting, the total control force is converted into thrust and attitude commands, updated at a frequency of 10 to 50 Hz. Taking node N5 as an example, the geometric centers and average velocities of its neighbors N2, N8, and N6 are calculated to generate aggregation and alignment forces. The distance to neighbors is checked; if it exceeds the safe distance of 10 meters, there is no repulsive force. The weighted sum yields the total control force, which is then converted into flight control commands. Within 5 seconds, N5's position deviation decreases from 1.5 meters to 0.3 meters, and its velocity direction and magnitude gradually align with its neighbors. All nodes execute the distributed cluster algorithm in parallel, making decisions based solely on local neighbor information without requiring global coordination. This distributed implementation reduces communication burden and computational overhead, exhibits good scalability and robustness, and demonstrates self-organizing emergent characteristics.

[0138] After the distributed clustering method has been in effect for a period of time, the cluster gradually recovers from a chaotic state to an ordered formation configuration. However, it is necessary to verify the effectiveness of the formation recovery through quantitative consistency indicators to determine whether a convergence state of physical consistency has been reached. The formation consistency verification is coordinated and executed by the subgroup leader node. By collecting the state information of each node, it calculates the global formation error index and determines whether to end the formation recovery process based on whether the error meets the preset threshold.

[0139] Formation consistency assessment includes two aspects: position consistency and velocity consistency. Position consistency measures the deviation of each node's current position from the target formation position, while velocity consistency measures the dispersion of each node's velocity. The position consistency index is defined as the root mean square of the position deviations of all nodes, calculated as Ep = √[(1 divided by N) multiplied by the square of the norm of the position deviation vector ei of all nodes i], where N is the total number of nodes in the cluster, ei equals Pi minus Pti, where Pi is the current position of node i, and Pti is the target formation position of node i. The velocity consistency index is defined as the root mean square of the velocity deviations of all nodes from the average velocity, calculated as Ev = √[(1 divided by N) multiplied by the square of the norm of the velocity deviation vector vi of all nodes i minus Va], where Va is the average velocity of all nodes. The sub-swarm leader N2 periodically broadcasts status query messages to all nodes in the cluster, requesting each node to report its current position, velocity, and flight status. The status query period is Tq, typically 2 to 5 seconds. The choice of period needs to balance the timeliness of information updates and the consumption of communication bandwidth. After receiving the status query message, each node reads the current position and velocity information provided by the navigation system, packages it into a status report message, and sends it to the navigator. The status report message includes the node identifier, position coordinates (3D), velocity vector (3D), timestamp, and message sequence number.

[0140] After sending a status query message, the Navigator N2 starts a receive timer, waiting for status reports from each node. The timer timeout is set to Tt, typically 1 to 2 seconds, and needs to cover the round-trip communication latency and processing latency of the furthest node. All status report messages collected by the Navigator before the timeout are stored in a local buffer. After the timeout, the receive window is closed, and consistency computation begins. If some nodes fail to report their status before the timeout, the Navigator marks these nodes as unresponsive, excludes their data from the consistency computation, but records the number of unresponsive nodes. If the number of unresponsive nodes exceeds a threshold (e.g., more than 10% of nodes), the consistency verification is considered to have failed, requiring a re-execution of the status query or the initiation of a fault detection process. The Navigator broadcasts a formation recovery completion notification message to all nodes in the cluster, containing the consistency metric values, the convergence determination results, and subsequent task instructions. After receiving the notification message, each node switches its local state machine from formation recovery mode to normal cruise mode and begins to execute the predetermined task process. For example, the reconnaissance node starts target search and tracking, the communication relay node establishes a data forwarding link, and the attack node enters standby mode to wait for attack instructions.

[0141] If the consistency metric fails to meet the threshold requirement, the navigator determines that formation recovery has not yet converged and continues to execute the distributed cluster method, waiting for the next state query cycle for reassessment. If the consistency verification fails multiple times consecutively (e.g., 5 times consecutively), and the improvement rate of the consistency metric is slow or oscillating, the navigator determines that formation recovery is encountering difficulties, possibly due to improper control parameter configuration of some nodes or excessive external environmental interference. The navigator initiates a diagnostic process, collecting detailed control logs and sensor data from each node, analyzing the reasons for the lack of convergence in the consistency metric, and adjusting control parameters or replanning the formation configuration based on the analysis results. The formation consistency verification mechanism provides quantitative convergence judgment criteria, objectively evaluating the effectiveness of formation recovery through error indicators in two dimensions: position and velocity. Periodic state queries and consistency calculations enable the system to monitor the formation status in real time, promptly detect and handle anomalies, and ensure the reliable completion of the formation recovery process.

[0142] This application also provides a self-healing device based on a distributed, decentralized cluster network. Figure 2 This is a schematic diagram of the structure of a self-healing device based on a distributed, decentralized cluster network provided in an embodiment of this application, with reference to... Figure 2 The device is located within a drone swarm and also includes an acquisition unit 201, a fault detection unit 202, a reconstruction unit 203, a role replacement unit 204, and a recovery unit 205. Acquisition unit 201 monitors the heartbeat information and monitoring packet delivery rate of the neighboring nodes of each drone node in the drone cluster, and identifies suspicious neighboring nodes based on the heartbeat information and monitoring packet delivery rate; fault The detection unit 202 performs cross-validation on the known neighbor nodes of the suspected neighbor node, determines whether the suspected neighbor node is in a fault state based on the validation results, and identifies the suspected neighbor node that is in a fault state as a fault node. Reconstruction unit 203, when a fault condition causes the topology of the UAV cluster to break and form at least one target subgroup, collects the comprehensive score calculated by each healthy UAV node in the target subgroup based on its own physical state parameters, and elects the healthy UAV node with the highest comprehensive score as the subgroup leader through a distributed consensus algorithm; under the coordination of the subgroup leader, the neighboring nodes of the faulty node perform local topology repair based on rigid graph theory to restore the internal connectivity of the target subgroup, and the subgroup leader discovers other target subgroups and triggers merging through adaptive expansion to obtain the reconstructed cluster; Role succession unit 204: When the faulty node is a critical role, it determines the target healthy drone node as the successor node from the reconstructed cluster based on the distributed consensus algorithm, and controls the target healthy drone node to take over the task of the critical role. Recovery unit 205, after confirming that the network topology reconstruction and role handover are completed, guides the reconstructed cluster to recover from network consistency to physical consistency based on a hierarchical recursive self-healing strategy, so as to restore the preset formation configuration.

[0143] In one possible implementation, the acquisition unit 201 is used to acquire observation nodes from the UAV cluster. If there is a suspicious neighbor node among the neighbor nodes of the observation node, a cross-validation request is sent to the known neighbor nodes of the suspicious neighbor node. The fault detection unit 202 is used to collect the observation dataset for the suspicious neighbor node fed back by the known neighbor nodes. The observation dataset includes the historical communication quality, current heartbeat reception status, and collaborative sensing data between the known neighbor nodes and the suspicious neighbor nodes. If all current heartbeat reception statuses in the observation dataset are not received, and no collaborative sensing data from the suspicious neighbor node is received within a preset waiting time, the suspicious neighbor node is determined to be in a physically damaged state. If some known neighbor nodes in the observation dataset can receive the heartbeat information of the suspicious neighbor node, or the physical motion trajectory in the collaborative sensing data conforms to the preset formation motion constraints and the historical communication quality is lower than the preset communication threshold, the suspicious neighbor node is determined to be in a local interference state.

[0144] In one possible implementation, the acquisition unit 201 is used to acquire the physical state parameters corresponding to healthy nodes from the target subgroup. The physical state parameters include the current remaining energy, node degree, average relative distance to other healthy UAV nodes in the target subgroup, and average signal-to-noise ratio. The reconstruction unit 203 is used to normalize the physical state parameters using a preset normalization function to obtain the corresponding energy score, topology score, position score, and channel score, and construct a score feature vector. Based on the score feature vectors of all healthy UAV nodes in the target subgroup, the local information entropy of each score feature is calculated, and the initial weight of each score feature is determined based on the local information entropy. Based on a preset state-weighted excitation function, the energy score, topology score, position score, channel score, and initial weight of the healthy node are nonlinearly mapped to generate dynamic weights that match the current physical state of the healthy node. The energy score, topology score, position score, and channel score are weighted and summed using the dynamic weights to obtain the comprehensive score of the healthy UAV node.

[0145] In one possible implementation, the reconstruction unit 203 is used for each healthy drone node to broadcast an election message containing its own identifier and comprehensive score within the target subgroup; each healthy drone node collects the election messages broadcast by other healthy drone nodes within the target subgroup and constructs a score mapping table locally; each healthy drone node performs multiple rounds of interactive verification on the data in the score mapping table based on a distributed consensus protocol until all healthy drone nodes within the target subgroup reach a consensus on the score mapping; each healthy drone node determines the healthy drone node with the highest comprehensive score as the subgroup leader based on the consensus-reached score mapping table.

[0146] In one possible implementation, the acquisition unit 201, under the coordination of the subgroup leader, acquires a set of neighboring nodes within the local topology association range of the faulty node and extracts a local subgraph composed of the neighboring node set; the reconstruction unit 203 is used to construct a local minimum rigid graph in the local subgraph based on rigid graph theory and merge the local minimum rigid graph into the global topology of the target subgroup to complete the local topology repair of the connectivity within the subgroup; the subgroup leader periodically broadcasts subgroup probe messages, and if no feedback is received from other subgroups within a preset probe time, it controls the nodes in the target subgroup to adaptively increase the communication radius to scan for the existence of other target subgroups; or the subgroup leader determines relay candidate nodes in the target subgroup, controls the relay candidate nodes to move to a preset adjacent subgroup location to be deployed as communication relays, and actively establishes cross-subgroup communication links; after establishing communication connections with other target subgroups, it controls the subgroup leader to exchange member lists and status information with the leaders of other target subgroups, and performs subgroup merging based on distributed joint consensus to obtain the reconstructed cluster.

[0147] In one possible implementation, the role succession unit 204 is used to broadcast a role succession request within the subgroup when the faulty node is a critical role. The role succession request carries the hardware payload requirements and task attribute parameters of the critical role. Based on the current state parameters of each healthy drone node in the reconstructed cluster and the role succession request, the corresponding bidding score is calculated, and a bidding response message carrying the bidding score is sent to the subgroup leader. The subgroup leader, based on distributed consensus, determines the target healthy drone node with the highest bidding score from all received bidding response messages as the successor node, and controls the target healthy drone node to take over the task of the critical role.

[0148] In one possible implementation, the role succession unit 204 is used to match and verify the hardware configuration of the bidding healthy drone node according to the hardware load requirements in the role succession requirements. The bidding healthy drone node is any healthy drone node in the reconstructed cluster. If the matching verification fails, the bidding score of the bidding healthy drone node is determined to be zero. If the matching verification succeeds, the current state parameters of the bidding healthy drone node are obtained, and combined with the task attribute parameters in the role succession requirements, the target input variables are calculated. The target input variable parameters include remaining energy, relative proximity based on the task execution location, and expected task load after superimposing the task attribute parameters. Based on a preset membership function, the remaining energy is used to determine the target input variable. Energy, relative proximity, and expected task load are mapped to corresponding fuzzy sets to obtain the membership degree of the target input variable under different fuzzy states. Fuzzy inference is performed on each membership degree according to a preset fuzzy inference library to calculate the rule strength of each triggered rule. The preset fuzzy rule library contains multiple rules constructed based on conditional statements, and the rule strength is the minimum value of each membership degree in the corresponding rule condition. The output fuzzy set of the corresponding rule conclusion is truncated according to the rule strength, and all truncated output fuzzy sets are merged to obtain an aggregated output fuzzy set. The centroid method is used to defuzzify the aggregated output fuzzy set, and the abscissa of the geometric centroid of the aggregated output fuzzy set is used as the bidding score corresponding to the health drone node to be bid on.

[0149] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0150] This application also discloses an electronic device. (See reference...) Figure 3 , Figure 3 This application provides a schematic diagram of the structure of an electronic device. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 302, and at least one communication bus 305.

[0151] The communication bus 305 is used to enable communication between these components.

[0152] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.

[0153] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0154] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 302, and by calling data stored in memory 302. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and application requests; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.

[0155] The memory 302 may include random access memory (RAM) or read-only memory. Optionally, the memory 302 may include a non-transitory computer-readable storage medium. The memory 302 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 302 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), instructions for implementing the various method embodiments described above, etc. The data storage area may store data involved in the various method embodiments described above. Optionally, the memory 302 may also be at least one storage device located remotely from the aforementioned processor 301.

[0156] like Figure 3 As shown, the memory 302, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program based on a distributed, decentralized cluster network for self-healing.

[0157] exist Figure 3 In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for users to input data and obtain user input data; while the processor 301 can be used to call the application program based on a distributed, decentralized cluster network self-healing stored in the memory 302. When executed by one or more processors, the electronic device performs one or more of the methods described in the above embodiments.

[0158] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0159] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0160] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some service interfaces; indirect couplings or communication connections between devices or units may be electrical or other forms.

[0161] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0162] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0163] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0164] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and practical application disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure.

Claims

1. A self-healing method for distributed, decentralized cluster networks, characterized in that, When applied to drone swarms, the method includes: The heartbeat information and monitoring packet delivery rate of each drone node in the drone cluster are monitored, and suspicious neighbor nodes are identified based on the heartbeat information and the monitoring packet delivery rate. Cross-validation is performed on the known neighbor nodes of the suspected neighbor node. Based on the validation results, it is determined whether the suspected neighbor node is in a faulty state, and the suspected neighbor node that is determined to be in the faulty state is identified as a faulty node. When the fault state causes the topology of the drone cluster to break and form at least one target subgroup, the comprehensive score calculated by each healthy drone node in the target subgroup based on its own physical state parameters is collected, and the healthy drone node with the highest comprehensive score is elected as the subgroup leader through a distributed consensus algorithm. Under the coordination of the subgroup leader, the neighboring nodes of the faulty node perform local topology repair based on rigid graph theory to restore the internal connectivity of the target subgroup. The subgroup leader then adaptively expands to discover other target subgroups and triggers merging, resulting in a reconstructed cluster. When the faulty node plays a critical role, a target healthy drone node is determined from the reconstructed cluster based on a distributed consensus algorithm as a replacement node, and the target healthy drone node is controlled to take over the task of the critical role. After confirming that the network topology reconstruction and role succession are completed, the reconstructed cluster is guided to recover from network consistency to physical consistency based on a hierarchical recursive self-healing strategy, so as to restore the preset formation configuration.

2. The method according to claim 1, characterized in that, The fault states include physical damage states and local interference states. Cross-validation is performed on the known neighbor nodes of the suspected neighbor node, and the determination of whether the suspected neighbor node is in a fault state is based on the validation results. Specifically, this includes: Observation nodes are obtained from the drone cluster. If a suspicious neighbor node exists among the neighbor nodes of the observation node, a cross-validation request is sent to the known neighbor nodes of the suspicious neighbor node. Collect observation datasets of the known neighbor nodes for the suspected neighbor nodes, wherein the observation datasets include historical communication quality, current heartbeat reception status, and cooperative sensing data between the known neighbor nodes and the suspected neighbor nodes; If all current heartbeat reception states in the observation dataset are not received, and no collaborative sensing data from the suspicious neighbor node is received within a preset waiting time, then the suspicious neighbor node is determined to be in the physical damage state. If some known neighbor nodes in the observation dataset can receive the heartbeat information of the suspected neighbor node, or if the physical motion trajectory in the collaborative sensing data conforms to the preset formation motion constraints and the historical communication quality is lower than the preset communication threshold, then the suspected neighbor node is determined to be in the local interference state.

3. The method according to claim 1, characterized in that, The comprehensive score calculated by each healthy UAV node in the target subgroup based on its own physical state parameters specifically includes: Obtain the physical state parameters corresponding to the healthy nodes from the target subgroup. The physical state parameters include the current remaining energy, node degree, average relative distance to other healthy UAV nodes in the target subgroup, and average signal-to-noise ratio. The physical state parameters are normalized using a preset normalization function to obtain the corresponding energy score, topology score, location score, and channel score, and a score feature vector is constructed. Based on the rating feature vectors of all healthy drone nodes in the target subgroup, calculate the local information entropy of each rating feature, and determine the initial weight of each rating feature based on the local information entropy. Based on a preset state-variable incentive function, the energy score, topology score, location score, channel score, and initial weight of the healthy node are nonlinearly mapped to generate dynamic weights that match the current physical state of the healthy node. The energy score, topology score, location score, and channel score are weighted and summed using the dynamic weights to obtain the comprehensive score of the healthy drone node.

4. The method according to claim 3, characterized in that, The process of electing the healthiest drone node as the subgroup leader by using a distributed consensus algorithm to evaluate multiple comprehensive scores specifically includes: Each of the aforementioned healthy unmanned nodes broadcasts an election message containing its own identifier and the comprehensive score within the target subgroup; Each healthy drone node collects campaign messages broadcast by other healthy drone nodes within the target subgroup and builds a scoring mapping table locally; Each healthy drone node performs multiple rounds of interactive verification on the data in the scoring mapping table based on a distributed consensus protocol until all healthy drone nodes in the target subgroup reach a consensus on the scoring mapping expression. Each healthy drone node determines the healthy drone node with the highest comprehensive score as the subgroup leader based on a consensus-based scoring mapping table.

5. The method according to claim 1, characterized in that, Under the coordination of the subgroup leader, the neighboring nodes of the faulty node perform local topology repair based on rigid graph theory to restore the internal connectivity of the target subgroup. The subgroup leader then adaptively expands to discover other target subgroups and triggers merging, resulting in a reconstructed cluster. Specifically, this includes: Under the coordination of the subgroup leader, a set of neighboring nodes that are within the local topological association range of the faulty node is obtained, and a local subgraph composed of the set of neighboring nodes is extracted. Based on rigid graph theory, a local minimum rigid graph is constructed in the local subgraph, and the local minimum rigid graph is merged into the global topology of the target subgroup to complete the local topology repair of the connectivity within the subgroup; The subgroup leader periodically broadcasts subgroup probe messages. If no feedback is received from other subgroups within a preset probe time, it controls the nodes in the target subgroup to adaptively increase their communication radius in order to scan for the existence of other target subgroups. Alternatively, the subgroup leader may determine relay candidate nodes in the target subgroup, control the relay candidate nodes to move to a preset adjacent subgroup location to be deployed as communication relays, and actively establish cross-subgroup communication links; Once a communication connection is established with the other target subgroups, the subgroup leader is controlled to exchange member lists and status information with the leaders of the other target subgroups, and subgroup merging is performed based on distributed joint consensus to obtain the reconstructed cluster.

6. The method according to claim 1, characterized in that, When the faulty node plays a critical role, a target healthy drone node is determined from the reconstructed cluster based on a distributed consensus algorithm as a replacement node. This target healthy drone node is then controlled to take over the tasks of the critical role. Specifically, this includes: When the faulty node is the critical role, the subgroup leader broadcasts a role replacement request within the subgroup, wherein the role replacement request carries the hardware load requirements and task attribute parameters of the critical role. Based on the current status parameters of each healthy drone node in the reconstructed cluster and the role succession requirements, the corresponding bidding score is calculated, and a bidding response message carrying the bidding score is sent to the sub-group navigator. Based on distributed consensus, the subgroup leader determines the target healthy drone node with the highest bidding score from all received bidding response messages as the successor node, and controls the target healthy drone node to take over the task of the key role.

7. The method according to claim 6, characterized in that, The step of calculating the corresponding bidding score based on the current state parameters of each healthy drone node in the reconstructed cluster and the role succession requirements specifically includes: Based on the hardware load requirements in the role succession requirements, the hardware configuration of the health drone node to be bid on is matched and verified. The health drone node to be bid on is any one of the health drone nodes in the reconstructed cluster. If the matching verification fails, the bidding score of the health drone node to be bid for is determined to be zero. If the matching verification is successful, the current status parameters of the health drone node to be bid are obtained, and the target input variables are calculated by combining them with the task attribute parameters in the role succession requirements. The target input variable parameters include remaining energy, relative position proximity based on the task execution position, and expected task load after superimposing the task attribute parameters. Based on a preset membership function, the remaining energy, the relative position proximity, and the expected task load are mapped to corresponding fuzzy sets to obtain the membership of the target input variable under different fuzzy states. Fuzzy reasoning is performed on each membership degree according to a preset fuzzy reasoning library to calculate the rule strength of each triggered rule. The preset fuzzy rule library contains multiple rules constructed based on conditional statements, and the rule strength is the minimum value of each membership degree in the corresponding rule condition. Based on the rule strength, the output fuzzy set of the corresponding rule conclusion is truncated, and all the truncated output fuzzy sets are combined to obtain the aggregated output fuzzy set. The centroid method is used to defuzzify the aggregated output fuzzy set, and the abscissa of the geometric centroid of the aggregated output fuzzy set is used as the bidding score corresponding to the healthy drone node to be bid on.

8. A self-healing device based on a distributed, decentralized cluster network, characterized in that, The device is located within a drone swarm and further includes an acquisition unit, a fault detection unit, a reconfiguration unit, a role takeover unit, and a recovery unit. The acquisition unit listens to the heartbeat information and monitoring packet delivery rate of the neighboring nodes of each drone node in the drone cluster, and determines suspicious neighboring nodes based on the heartbeat information and the monitoring packet delivery rate. The fault detection unit performs cross-validation on the known neighbor nodes of the suspected neighbor node, determines whether the suspected neighbor node is in a fault state based on the validation results, and identifies the suspected neighbor node that is in the fault state as a fault node. When the fault state causes the topology of the UAV cluster to break and form at least one target subgroup, the reconstruction unit collects the comprehensive scores calculated by each healthy UAV node in the target subgroup based on its own physical state parameters, and elects the healthy UAV node with the highest comprehensive score as the subgroup leader through a distributed consensus algorithm. Under the coordination of the subgroup leader, the neighboring nodes of the faulty node perform local topology repair based on rigid graph theory to restore the internal connectivity of the target subgroup. The subgroup leader then adaptively expands to discover other target subgroups and triggers merging, resulting in the reconstructed cluster. When the faulty node is a critical role, the role succession unit determines a target healthy drone node from the reconstructed cluster as a successor node based on a distributed consensus algorithm, and controls the target healthy drone node to take over the task of the critical role. The recovery unit, after confirming that the network topology reconstruction and role handover are completed, guides the reconstructed cluster to recover from network consistency to physical consistency based on a hierarchical recursive self-healing strategy, so as to restore the preset formation configuration.

9. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.