Emergency cooperative scheduling method based on swarm intelligence
By adopting an emergency collaborative scheduling method based on swarm intelligence, the problems of insufficient data integration and sharing, computing resource allocation and scheduling, system resilience and collaborative capabilities in the emergency management system are solved, and efficient and reliable emergency response and resource optimization are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING QUNXIN SPACE-TIME INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-01-13
- Publication Date
- 2026-04-28
AI Technical Summary
Existing emergency management systems have significant shortcomings in data integration and sharing, computing resource allocation and scheduling, and system resilience and coordination capabilities, resulting in low disaster response efficiency and resource waste.
An emergency collaborative scheduling method based on swarm intelligence is adopted. This method acquires multi-source data for preprocessing, generates collaborative scheduling strategies using a lightweight distributed ledger and federated deep Q-network model, and decomposes and distributes the strategies at edge nodes to achieve localized data utilization and global coordination.
It improves the efficiency of perception and response in emergency environments, enhances the robustness of the system in the event of failures or network fluctuations, ensures policy consistency and reliability, and improves the collaborative scheduling capability in complex disaster scenarios.
Smart Images

Figure CN121936828A_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to the field of emergency management, and more specifically, to an emergency collaborative scheduling method based on swarm intelligence. Background Technology
[0002] In current emergency management practices, especially in disaster relief and accident response scenarios, existing technology systems face significant challenges in achieving efficient and collaborative decision-making and resource allocation, and urgently require breakthrough architectures and methodologies to support them.
[0003] First, there are structural obstacles to the integration and sharing of emergency data. Data from diverse platforms such as satellites, ground monitoring stations, and unmanned vehicles differ significantly in terms of spatiotemporal references and format protocols. This heterogeneity makes it difficult to effectively integrate cross-platform information and form a unified and timely global situational awareness. In particular, when facing dynamic and complex disaster environments, traditional data calibration and fusion methods are often inefficient and lack real-time performance, severely restricting the ability to accurately perceive and analyze disaster situations.
[0004] Secondly, the allocation and scheduling of computing resources are insufficient to meet the dynamic and time-sensitive needs of emergency response. Reliance on centralized cloud computing or traditional cloud-edge collaboration models has inherent limitations: on the one hand, centralized processing struggles to efficiently respond to the explosive and dynamically changing computing demands of frontline mobile terminals; on the other hand, existing resource allocation mechanisms lack proactive prediction and intelligent coordination capabilities, resulting in uneven distribution of computing resources across regions, with valuable time windows wasted on task queuing or resource waiting. Simultaneously, the potential of edge computing resources carried by numerous mobile intelligent devices deployed on-site (such as drones and robots) has not been fully activated and effectively utilized, leading to both redundancy and waste of overall computing resources.
[0005] A more critical weakness lies in the system's overall resilience and coordination capabilities. Existing emergency decision-making systems are mostly vertically designed, with command flows often being unidirectional, leading to a disconnect between decision-making and execution, and causing delays or even breaks in the response process. The system relies excessively on fixed infrastructure (such as communication base stations), making it highly susceptible to failure under extreme disaster conditions (such as earthquakes and floods), disrupting information transmission links.
[0006] The core of the above problems lies in the fact that traditional architectures lack effective distributed coordination mechanisms at the data flow, computing flow, and control flow levels. Information silos, resource barriers, and broken collaboration directly lead to lagging situation assessment, low decision-making efficiency, and loose action of rescue forces. Summary of the Invention
[0007] According to the present invention, an emergency collaborative scheduling scheme based on swarm intelligence is provided. This scheme addresses the technical problems in the prior art, such as structural obstacles in the integration and sharing of emergency data in emergency scenarios, the difficulty in allocating and scheduling computing resources to meet the dynamic and time-sensitive requirements of emergency response, and insufficient overall resilience and collaborative capabilities of the system.
[0008] In a first aspect, the present invention provides an emergency collaborative scheduling method based on swarm intelligence. The method includes: acquiring multi-source data from various agents; preprocessing the multi-source data and feeding it back to the corresponding agents; generating a preliminary scheduling strategy based on the received multi-source data, and updating and maintaining the strategy distribution snapshots of edge nodes in the preliminary scheduling strategy using a lightweight distributed ledger; acquiring the evolution path of secondary disasters, inputting the evolution path of secondary disasters, the preprocessed multi-source data, and the strategy distribution snapshots of edge nodes into a federated deep Q-network model to generate a collaborative scheduling strategy; and decomposing the collaborative scheduling strategy into specific tasks, initially allocating them to various agents, and performing collaborative scheduling.
[0009] Compared with the prior art, the present invention has the following beneficial technical effects: This invention acquires and preprocesses multi-source data from various agents and feeds it back to the corresponding agents, achieving localized data utilization and global coordination. This provides real-time and accurate decision-making basis for subsequent strategy generation, effectively improving the system's perception and response efficiency in emergency environments. After generating initial scheduling strategies, each agent uses a lightweight distributed ledger to update and maintain the strategy distribution snapshot, ensuring that all edge nodes reach a consensus on the system state. This mechanism not only enhances the tamper-proof and traceability of strategy synchronization but also improves the overall robustness of the system in the event of partial node failures or network fluctuations, ensuring the consistency and reliability of distributed strategies. By inputting the secondary disaster evolution path, preprocessed multi-source data, and strategy distribution snapshots into a federated deep Q-network model, the system can integrate environmental evolution trends and multi-party strategy states to generate a globally optimized collaborative scheduling strategy while protecting the data privacy of each agent. Finally, this strategy is decomposed and distributed to each agent for execution, realizing a complete closed loop from intelligent decision-making to task implementation. This significantly enhances the overall integrity, adaptability, and intelligence level of the emergency dispatch system and improves the collaborative scheduling capability in complex disaster scenarios.
[0010] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0011] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A flowchart of an emergency collaborative dispatch method based on swarm intelligence according to an embodiment of the present invention is shown. Detailed Implementation
[0012] Figure 1 A flowchart of an emergency collaborative scheduling method based on swarm intelligence, according to an embodiment of the present invention, is shown.
[0013] The method includes: S101. Obtain multi-source data from each intelligent agent, preprocess the multi-source data, and then feed it back to the corresponding intelligent agent.
[0014] In this embodiment, the intelligent agents include a swarm of drones, a swarm of robot dogs, and an emergency vehicle.
[0015] Specifically, the drone swarm adopts an adaptive clustering protocol to construct a LoRa communication relay layer to form a topological sensing network.
[0016] In one embodiment of this invention, the UAV is equipped with a dual-spectrum gimbal for visible light and infrared light to collect visible light images and infrared temperature data, enabling monitoring of temperature distribution at the fire site. The swarm constructs a LoRaMesh network at an altitude of 50-100 meters, employing a TDMA time slot allocation mechanism to avoid signal collisions. The UAV swarm constructs a LoRa communication relay layer to form a topology sensing network. An adaptive clustering protocol is used, with each UAV clustering based on link quality. Electing a cluster head. The cluster head node schedules the transmission of child nodes through TDMA, enabling the formation of a self-organizing network within a 10km² area. The topology-sensing network dynamically adjusts relay paths through link quality index, improving the reliability of communication coverage at disaster sites.
[0017] Specifically, the robot dog swarm is equipped with UWB (Ultra-Wideband) tags, which are used to work in conjunction with UWB base stations to construct a ground grid coordinate system. The UWB base stations are deployed at intervals of ≤50m, achieving centimeter-level positioning (error <10cm) through the TDOA (Time Difference of Arrival) algorithm. The grid coordinate system is divided into 1m×1m units using GeoHash encoding, storing attributes such as ground bearing capacity and obstacle distribution. This scheme improves path planning accuracy in complex terrain.
[0018] In one embodiment of this invention, the robotic dog is equipped with a vibration sensor and a gas detector to collect evidence of geological loosening (e.g., using the vibration sensor to collect micro-vibration signals, acoustic emission signals, environmental vibration signals, and infrasound signals) and toxic gas concentrations. A swarm of robotic dogs constructs a ground grid coordinate system using UWB positioning, forming a path planning network. UWB positioning base stations are deployed on the roof of the rescue vehicle, combining with IMU inertial navigation to construct a dynamic grid map with centimeter-level accuracy. The path planning network uses a potential field algorithm to generate a gradient navigation field, enabling autonomous obstacle avoidance navigation in complex ruin environments.
[0019] The UAV relay layer is clustered according to Delaunay triangulation, with the cluster head node responsible for cross-layer data forwarding. The ground UWB grid uses the UAV as the origin of the coordinate system, and the grid resolution dynamically adapts to the terrain complexity (5 meters / grid in plains, 2 meters / grid in mountains). This enhances coverage capabilities in complex terrains.
[0020] Specifically, the emergency vehicle serves as a mobile communication anchor point and energy dispatch center. Its core tasks are to autonomously determine the optimal location to optimize network coverage and intelligently manage onboard energy. The emergency vehicle forms mobility decisions based on a dynamic priority scoring model, and autonomously determines its location as a communication anchor point for communicating with other intelligent agents.
[0021] The process of forming a movement decision based on a dynamic priority scoring model includes: A comprehensive priority score for the current location is output using a dynamic priority scoring model, and a movement decision is generated using this comprehensive priority score; the dynamic priority scoring model is as follows: in, The overall priority score for the target location; The signal strength factor (measured by an onboard spectrum analyzer, characterizing the communication quality at that location); The path smoothness factor (derived from the global passability matrix) (This indicates the feasibility of traveling to that location). Energy consumption penalty item (related to the vehicle's remaining battery power) negative correlation The lower the value, the higher the penalty. This is the first adjustable weighting coefficient; This is the second adjustable weighting coefficient.
[0022] The emergency vehicle periodically calculates potential target points in the surrounding area. The system calculates the network coverage value and moves to the position with the highest score, thereby achieving self-optimization of network coverage.
[0023] Specifically, the logic for generating movement decisions using the dynamic priority scoring model is as follows: 1. Candidate point generation: The emergency vehicle's decision-making system periodically (e.g., every 60 seconds) generates candidate points based on the onboard map and the global traffic coefficient matrix. Within a certain range (e.g., 500 meters) of the current location, generate a set of N reachable "potential stopping candidate points" (e.g., road intersections, open areas with higher elevations). Simultaneously, the vehicle's "current location" is also considered as one of the candidate points (i.e., the candidate point). ) are included in the calculation.
[0024] 2. Traversal Evaluation: The system traverses these N candidate points (including candidate points) ), and calculate the comprehensive priority score P for each candidate point.
[0025] For candidate points The system needs to evaluate its signal strength factor. Path smoothness factor and energy consumption penalty items .
[0026] Signal strength: By combining an onboard spectrum analyzer with a signal propagation prediction model, the vehicle's movement to the candidate point is estimated. The expected communication quality afterwards.
[0027] Path accessibility: determined by querying the global accessibility matrix. Get the distance from the current position to the candidate point The smoothness of the path (or the cost of passage).
[0028] Energy consumption penalty: This item is mainly based on the vehicle's current remaining battery power. Calculation. The energy consumption penalty term for all candidate points within the same decision loop. The value is the same, serving as a global "inhibition factor" when the remaining battery power... When it is very low, the energy consumption penalty term The value is very high, resulting in all candidate points (including candidate points) The overall priority score P value of the target position was lowered, tending to "remain unchanged".
[0029] 3. Move Decision Generation: The system compares all candidate points (including candidate points) Find the candidate point with the highest score by calculating the p-value. .
[0030] Generating movement decisions involves two scenarios: Scenario 1 (Retention): If the candidate point has the highest score That is, "current position" (candidate point) ),or Score Score with current position If the difference is less than a preset "movement incentive threshold" (to avoid frequent invalid moves), the movement decision is "keep the current position".
[0031] Scenario 2 (Movement): If the candidate point has the highest score It is a new position, and its score Significantly higher than the current position score The move decision is "move to..." ", and call the route planning module to generate the destination The path will be used as the new communication anchor point. In one embodiment of this invention, the emergency vehicle integrates a panoramic camera and a sonar array to construct a three-dimensional terrain point cloud. The emergency vehicle (emergency rescue vehicle) serves as a mobile decision-making anchor point, dynamically executing: a) Objective function: Score = 0.7 × Signal Strength + 0.3 × Path Accessibility. The signal strength factor is dynamically monitored by a spectrum analyzer to measure electromagnetic interference intensity, while path accessibility is calculated by fusing visual SLAM and laser point cloud data. The weighted scoring model uses a fuzzy logic controller to adjust coefficients in real time, optimizing vehicle maneuvering decisions in environments with strong interference.
[0032] b) Energy consumption priority: Prio=1 for life support equipment, Prio=0.8-0.1×remaining power for other equipment. Life support equipment (such as oxygen generators) is connected to an independent power bus and equipped with hardware-level power failure protection. Conventional equipment operates at a reduced frequency based on the state of charge (SOC) of the battery management system (BMS) to ensure critical missions are carried out under limited energy conditions.
[0033] In this embodiment, the multi-source data includes visible light image data (collected by drone), infrared temperature data (collected by drone for monitoring temperature distribution at the fire site), vibration signal data (collected by robot dog, including micro-vibration, acoustic emission, environmental vibration, and infrasound signals), toxic gas concentration data (collected by robot dog), panoramic image data (collected by emergency vehicle), and sonar point cloud data (collected by emergency vehicle for constructing three-dimensional terrain).
[0034] In this embodiment, multi-source data needs to be preprocessed before being fed back to the corresponding intelligent agent. The preprocessing process includes: aligning the multi-source data using a spatiotemporal registration algorithm (based on GPS / IMU fusion positioning) and using wavelet transform to eliminate sensor noise. Through the above preprocessing, the comprehensiveness and accuracy of environmental perception in complex disaster scenarios are significantly improved. The preprocessed data is then returned to the intelligent agent that sent the data.
[0035] As an extended implementation of this embodiment, the onboard equipment of the emergency vehicle can also be managed hierarchically, and onboard energy scheduling can be carried out using dynamic adjustment rules for energy consumption priorities. Specifically, this includes: assigning task priorities to onboard equipment and monitoring the operating status of the onboard equipment in real time; the onboard equipment includes critical and non-critical equipment. Critical equipment includes life support equipment, core communication equipment, and vehicle power and control systems. Non-critical equipment includes: data backup storage, environmental monitoring sensors (non-safety related), auxiliary lighting, etc.
[0036] The dynamic adjustment rules for energy consumption priority include: In response to the activation signal of life support equipment, a highest-priority interrupt request is sent to the task scheduler, triggering the task scheduler to suspend all currently non-highest-priority device tasks and reallocate system resources to the life support equipment. Life support equipment is always the highest priority: life support equipment (such as oxygen generators and heart monitors) is directly connected to the high-priority power bus via a dedicated independent power supply circuit (such as a dual-relay interlock circuit), ensuring power supply priority at the hardware level. At the software level, the priority arbitrator monitors the equipment's operating status in real time. Once a life support equipment activation signal is detected, a highest-priority interrupt request is immediately sent to the task scheduler. The task scheduler responds to the interrupt, suspends all currently non-highest-priority device tasks, and immediately reallocates system resources (such as power quotas and calculation cycles) to the life support equipment.
[0037] When the remaining power of life support equipment falls below a first threshold, the priority of non-critical equipment is downgraded. This means setting a first threshold. When the device has remaining power In this case, the priority of non-critical equipment is downgraded. The downgrade process for non-critical equipment is triggered by the Battery Management System (BMS) and executed using a predefined progressive strategy. The downgrade of non-critical equipment includes: First, non-critical equipment is downgraded at the task level. This downgrade includes suspending background tasks marked as background or low_priority (such as log uploads and non-emergency data synchronization) and continuously monitoring the downgraded non-critical equipment. These tasks are characterized by not being essential for core emergency rescue instructions and immediate operations (such as life support, vehicle movement, and real-time communication). Suspending them will not immediately affect the execution of critical rescue work, so they are suspended first to save power when energy is scarce.
[0038] If the rate of battery drain per second exceeds the second threshold (BMS detects that the rate of battery drain per second exceeds the threshold), the non-critical device will be functionally downgraded, and the downgraded non-critical device will be continuously monitored; the functional downgrade includes shutting down the non-core functional modules of the device (for example, downgrading the 1080P recording function of the camera to 480P, or shutting down the LiDAR and keeping only the ultrasonic sensor).
[0039] If the power supply remains critical, a soft shutdown command is sent to non-critical devices to shut them down. "Remaining critical" refers to a situation where the power supply situation has not improved or stabilized after implementing both task-level and function-level energy-saving measures.
[0040] The specific execution process for downgrading the priority of non-critical equipment is as follows: 1) Initial Critical State: When the device has remaining power... Below the preset minimum threshold At this point, the system enters an initial critical state and triggers "task-level degradation" (such as pausing background tasks). For example, the battery threshold. The battery is set to 20% of the rated capacity. When the battery management system (BMS) detects that the battery level has reached the threshold, it sends a downgrade command to the task scheduler.
[0041] 2) Continuous Critical State: If the Battery Management System (BMS) detects that "the rate of power loss per second exceeds the threshold" after the task-level degradation, it means that the power is still being consumed rapidly, and the system will trigger "functional-level degradation" (such as shutting down non-core functional modules of the device).
[0042] 3) "Still Critical" Status: If the power problem persists after functional downgrading (the document does not provide a new specific numerical threshold, but logically it means that the first two measures are ineffective), the system will determine that the power is "still critical" and execute the last resort - sending a soft shutdown command to the non-critical device.
[0043] This mechanism significantly extends the continuous operational capability of critical equipment in power outage scenarios through pre-defined equipment capability description files (defining the core / non-core functional modules of each device) and real-time power monitoring.
[0044] S102. The agent generates a preliminary scheduling strategy based on the received multi-source data, and uses a lightweight distributed ledger to update and maintain the strategy distribution snapshot of the edge nodes in the preliminary scheduling strategy.
[0045] In this embodiment, the agent makes autonomous decisions based on local environment perception; the agent generates a preliminary scheduling strategy based on received multi-source data, including: 1. Status assessment: The intelligent agent (such as a drone or robot dog) first integrates the preprocessed data fed back from S101 (such as the fire point temperature identified by the drone and the concentration of toxic gas detected by the robot dog) with its own internal status (such as the remaining power and the current task).
[0046] 2. Local Decision Making: The agent runs a lightweight local decision-making model (e.g., a pre-trained lightweight neural network, decision tree, or a rule-based finite state machine FSM).
[0047] 3. Strategy Generation: The model outputs one or more "preliminary scheduling strategies". These strategies are local and immediate.
[0048] Example: Drone: If the input is 'high priority fire point' and 'self-battery > 30%', the generated initial strategy is {Action:"Monitor",Target:[Coord_X,Coord_Y],Priority:0.9}.
[0049] Robot Dog: If the input is 'high concentration of toxic gas' and 'path ahead', the initial strategy is {Action:"Avoid",Target:[Safe_Coord_Z],Priority:1.0}.
[0050] Emergency vehicle: If the reported communication quality Ss is lower than the threshold, the initial strategy is {Action:"Relocate",Priority:0.8} to trigger subsequent mobility decisions.
[0051] 4. Snapshot Preparation: This generated JSON format or vectorized preliminary policy constitutes the agent's "policy intent" at the current moment. This "policy intent" is then used to update the "policy distribution snapshot" in subsequent steps.
[0052] In this embodiment, to ensure global consistency, a lightweight distributed ledger is used to maintain policy-distributed snapshots of edge nodes. In the improved Hyperledger Fabric architecture: Each edge node acts as a peer node, comprising an agent and / or an edge computing gateway connected to the agent. These edge nodes store a Merkle tree hash chain of policy distribution snapshots, rather than the complete transaction history, significantly reducing storage burden. The policy distribution snapshot is a lightweight hash digest of the agent's current decision state (e.g., resource allocation, target location, task queue).
[0053] The communication protocol between peer nodes uses the Gossip protocol for data synchronization, replacing the complex channel mechanism in the meta-architecture and reducing communication latency.
[0054] The ordering service includes receiving policy snapshot update requests from various peer nodes, employing a simplified Byzantine fault-tolerant consensus mechanism to reach a consensus on the update order, replacing the original distributed consensus algorithm (RAFT) or Kafka to tolerate more node failures. The consensus mechanism is the core of the ordering service. Its main task is to receive policy snapshot update requests from various peer nodes (i.e., edge agents) and reach a network-wide consensus on the order of these updates. Even if some nodes fail or send malicious information, this mechanism ensures that all honest nodes agree on the final order of policy updates. The ordered and consensus-agreed policy snapshots are then used to update the Merkle tree hash chain stored on each peer node.
[0055] By ensuring the consistency of the entire network state, it provides a reliable data foundation for the communication protocol (Gossip protocol), making the data synchronized between peer nodes verified and consensus-based, thereby avoiding policy conflicts and ultimately achieving global state synchronization of the entire system.
[0056] In this embodiment, the simplified Byzantine fault-tolerant consensus mechanism includes: Pre-preparation phase: Same as standard Byzantine Fault Tolerance (PBFT). In response to a policy distribution snapshot update request, the master node assigns a sequence number to the request, packages it, and broadcasts a pre-preparation message to all backup nodes. Preparation phase: Once any backup node receives and verifies the pre-preparation message, it no longer broadcasts a separate preparation message to all other nodes. Instead, it uses the current backup node's private key share to sign the digest of the pre-preparation message, generates a signature share, and broadcasts the signature share. Submission Phase: If any node (including the master node) collects more than a first threshold number of signatures for the same message (e.g., 2f+1, where f is the number of faulty nodes that can be tolerated), then the signatures collected by that node for the same message are aggregated into a single threshold signature of constant size. This final threshold signature is equivalent to the "submission certificate" in the standard Byzantine Fault Tolerance (PBFT) mechanism that has collected 2f+1 "submission" messages. It is the final proof of consensus and the threshold signature is broadcast.
[0057] If any node generates or receives the signature threshold, it indicates that all nodes have reached a consensus on the prepared message, and the policy distribution snapshot update request is written to the local ledger and the policy distribution snapshot update is executed. Once a node generates or receives the final threshold signature, it means that the entire network has reached a consensus on the message, and it can be executed and written to the local ledger.
[0058] S103. Obtain the evolution path of secondary disasters, input the evolution path of secondary disasters, preprocessed multi-source data and policy distribution snapshots of edge nodes into the federated deep Q network model to generate a collaborative scheduling policy.
[0059] In this embodiment, the evolution path of the secondary disaster is obtained through a digital twin engine. The digital twin engine predicts the evolution of the disaster using simulation techniques such as computational fluid dynamics (CFD) or finite element analysis. Based on finite element analysis, the digital twin engine generates a disaster evolution probability graph (e.g., a debris flow diffusion path represented as G(V,E,P), where V is the node set, E is the edge set, and P is the path activation probability). The probability graph is then converted into a state transition matrix of a Q-network. And used in the following update formula: in, In the state Take action below The expected cumulative reward value. The learning rate controls the extent to which new information covers older information. This serves as a discount factor, balancing the importance of current rewards versus future rewards. After performing action a in state s, the environment transitions to state s. The probability (provided by the digital twin); In the next state Below, all possible actions The maximum that can be obtained in the middle value; These represent the current state, the action taken, and the next state, respectively. It is evident that the federated deep Q-network model can anticipate risks in areas prone to aftershocks, significantly improving the safety of rescue routes.
[0060] The evolution path of secondary disasters is output in the form of a "state transition probability matrix", which is used as an input parameter for the state transition equation of the Q network to optimize and guide the generation of strategies.
[0061] In this embodiment, the federated deep Q-network model is deployed and run on front-end computing devices close to the data source, rather than on remote cloud servers. These edge nodes can be computing gateways mounted on emergency vehicles or other field devices. This approach ensures the localized processing of sensitive data, thus balancing real-time decision-making with privacy and security. The federated deep Q-network model is used to generate collaborative scheduling strategies. The state space S of the federated deep Q-network model is defined as tuples representing key dimensions such as <resource availability, disaster spread, and device status>.
[0062] The input data received by the federated deep Q-network model during local training and decision-making includes: 1) Local perception data: Multi-mode data collected and processed in real time by intelligent agents such as drones and robot dogs in step S101.
[0063] 2) Global state snapshot: In step S102, a policy distribution snapshot of other edge nodes synchronized through the lightweight distributed ledger.
[0064] 3) Evolution path of secondary disasters: This is data that is predicted by a digital twin engine and then input into the federated deep Q network model (FDQN).
[0065] In this embodiment, the federated deep Q-network model adopts a federated learning framework, which consists of two parts: 1) Multiple local deep Q-network (DQN) models are deployed and run on edge nodes; 2) A federated evolution center deployed in the cloud is responsible for integrating the model parameters of each edge node to generate a global evolution strategy model.
[0066] In this embodiment, the step of inputting the evolution path of the secondary disaster, the preprocessed multi-source data, and the policy distribution snapshot of the edge nodes into the federated deep Q-network model to generate a cooperative scheduling policy includes: Each edge node independently trains its local DQN model based on preprocessed local multi-source data and synchronized global state data; During training, the evolution path of the secondary disaster is embedded into the update formula of the DQN model in the form of a state transition probability matrix for iteration; this allows the model to anticipate risks and improve the safety of decision-making; for example, during DQN model training, its core Q-value update formula directly uses the state transition probability matrix P(s'|s,a) provided by the digital twin.
[0067] The local DQN model periodically (e.g., every 5 minutes) uploads the gradient parameters output during training to the cloud. The cloud-based federated evolution center uses a differential privacy gradient aggregation algorithm to integrate all uploaded gradient parameters and generate a better global model, namely the global evolution strategy model. The performance and generalization ability of the global evolution strategy model are superior to any single local DQN model.
[0068] This improved global model is sent back to the edge nodes to update their local DQN models. Finally, the edge nodes use this updated local model, combined with current real-time data, to generate and output the final collaborative scheduling strategy.
[0069] Edge nodes use the updated local DQN model to output a collaborative scheduling strategy.
[0070] Specifically, the federated deep Q-network model needs to be continuously trained. The training data consists of locally sensed data (from step S101) and a global state snapshot (from step S102). The local DQN model on each edge node receives the locally sensed data from step S101 and the global state snapshot from step S102 as input for training and decision-making. Simultaneously, the disaster path predicted by the digital twin is also used as input to guide policy optimization. This training process, referred to as "local training," is performed independently on each edge node.
[0071] Upload and Distribution: After training, the local model periodically uploads its gradient parameters to the Federated Evolution Center in the cloud. The Federated Evolution Center aggregates these parameters to generate a better global model, which is then distributed to each edge node to update its local model.
[0072] Gradient parameters are intermediate products calculated using the backpropagation algorithm during model training. The gradient represents the direction and magnitude of the adjustment that the model's internal parameters should undergo to reduce error.
[0073] To provide decision-making support for the Federation Deep Q-Network (FDQN) model, this system employs multi-source fusion technology (UAV thermal imaging, robot dog sonar, and emergency vehicle UWB radar) to achieve precise positioning. The positioning results are input into the FDQN as part of the state information. The system automatically performs injury grading (ISS standard) by analyzing visible light images using CNN and outputs injury weights. This weight is used to construct the reward function of FDQN.
[0074] The reward function of the federated deep Q-network model is designed as follows: in, This is the total reward value, used to guide the model's learning direction. These are the third and fourth adjustable weighting coefficients, used to balance the importance of rescue value and resource consumption. The calculation formula is as follows: . This refers to the number of people trapped. The injury weighting coefficient is (minor injury: 0.3; serious injury: 1.0; critical injury: 2.0). The cost of resource consumption is calculated using the following formula: .in, For the first Dynamic weights of resource classes (calculated using the entropy weighting method); For the first Real-time consumption rate of resource types. For the first The upper limit of the consumption rate of a resource type. On the other hand, when outputting actions through a collaborative scheduling strategy, a global access coefficient matrix is also required. Considering path cost, the global traversal coefficient matrix It provides the basis for calculating path costs (e.g., the route to point A requires passing through a certain number of road segments). Its high weighting means that the value of this decision will be reduced.
[0075] S104. Decompose the collaborative scheduling strategy into specific tasks, initially allocate them to each intelligent agent, and perform collaborative scheduling.
[0076] In this embodiment, the step of decomposing the cooperative scheduling strategy into specific tasks and initially assigning them to each intelligent agent includes: The collaborative scheduling strategy is decomposed into tasks to generate a task list; According to the collaborative scheduling strategy, a pheromone field is constructed for each task in the task list. The agent is treated as an individual in the ant colony, and the attractiveness of the task to the agent is calculated through the task attraction function. The agent moves to the task point with the highest attractiveness, thus forming a biological group intelligent scheduling mechanism.
[0077] The task attraction function is: in, The task urgency is the quantified value output by S103, i.e., the FDQN model in S103 for each task (i.e., action) in the "task list". The Q-value calculated by ) ), equivalent to the formula in ); The distance from the agent to the task point is calculated using the global accessibility matrix. Calculated; Environmental visibility factor (obtained from meteorological sensors or time information); , , These are the fifth, sixth, and seventh adjustable weight coefficients, which are dynamically configured according to the disaster scenario.
[0078] Intelligent agents are like ants, prioritizing attraction within the field. The highest-ranking task point is moved, thus achieving initial rapid task distribution. During this movement, the global traversal coefficient matrix needs to be considered. It considers real-time, dynamic environmental accessibility information. When the ant colony algorithm calculates the task's attractiveness... Distance in At that time, this distance should not be the physical straight-line distance, but should be based on The shortest weighted path distance calculated by the matrix. A more circuitous route with a lower traffic weight may be "shorter" than a straight route with a higher traffic weight.
[0079] In this embodiment, the biological swarm intelligent scheduling mechanism achieves dynamic task distribution and adaptive role allocation through biomimetic algorithms. During execution, a two-layer decision-making architecture is employed: first, a bee swarm mechanism is used for real-time role adjustment (microsecond-level response); then, at intervals, policy distribution data is pulled from the global state synchronizer, and a Hungarian algorithm is used for global rebalancing, outputting the global reassignment scheme to the execution agent layer. The bee swarm mechanism achieves microsecond-level role switching locally at edge nodes, broadcasting state changes through a lightweight message queue. During the global rebalancing phase, the central coordinator periodically pulls node state snapshots, uses the Hungarian algorithm to solve for the globally optimal allocation scheme, and then pushes incremental update instructions, balancing real-time response with system-level resource optimization goals.
[0080] As another implementation of this embodiment, ant colony pheromone-driven: Using an ant colony pheromone mechanism, intelligent agents are attracted to the task point, thereby achieving initial rapid task distribution. A pheromone gradient field is established in the task region. In this case, the aforementioned task attraction function can be further updated to an exponential form, making it more closely aligned with the biomimicry of the "pheromone gradient field." in: Due to the urgency of the task, This represents the distance from the agent to the task point. As a reference distance constant, Visibility factor. The agent prioritizes highly attractive tasks. By establishing a dynamic pheromone concentration field in the digital map, the urgency of the mission is calculated in real time based on the vital signs monitoring data of the trapped personnel, and the distance factor... The shortest travel time is integrated into the path planning network, and visibility parameters are dynamically adjusted based on weather sensors. The attraction function employs an exponential decay mechanism to simulate the pheromone evaporation process, enabling intelligent priority identification and rapid response for high-threat tasks.
[0081] As one implementation method of this embodiment, after the tasks are initially assigned to each intelligent agent, there are also two roles, a and b, that can be switched.
[0082] First, a capability vector is constructed for each agent, and a task requirement vector is constructed for each task, resulting in a cost matrix; next: a. Calculate the capability matching degree between each agent and the initially assigned task. When the capability matching degree is less than the capability matching threshold, it is determined that the current agent is not suitable for the current role, and the role switch is triggered.
[0083] Calculating the fit (fitness) is to verify the rationality of the initial assignment and determine whether a "role switch" needs to be triggered for reallocation. The role switch decision is based on a quantified ability fit function: in, The calculated capability matching degree (or fitness) has a value between -1 and 1 (or between 0 and 1 if all vector components are non-negative). The closer the value is to 1, the better the agent's capabilities match the task requirements. This is the demand vector for the current task assigned to the agent. The capability vector of the agent performing the task.
[0084] For example, when the calculated ability matching degree F between the agent and the task is less than a threshold (e.g., 0.7), the system determines that the agent is not suitable for the current role, thus triggering a role switch to find a more suitable task. The system will assign a new task type that matches its ability vector better (e.g., a drone that has switched from a "transportation role" to a "surveillance role" will be reassigned to perform relay communication tasks; other tasks include search and rescue, transportation, demolition, surveillance, and relay). This mechanism can determine whether a specific agent is capable of performing a specific task, that is, measure the matching degree between the individual and the task, thereby dynamically optimizing the efficiency of global task allocation.
[0085] b. After dynamically assigning tasks to each agent, the following steps are also included: monitoring the backlog rate of the assigned tasks. When the backlog rate of a task exceeds the backlog rate threshold and there are agents in an idle state, the execution role switch is triggered.
[0086] For example, when the task backlog rate is >20%, the role fitness formula applies. ( For the first The system dynamically converts idle agents into scarce roles based on the current backlog of tasks for a particular role. Triggering conditions and actions: When a backlog of tasks (such as demolition) occurs while a specific role (such as a transport robot) is idle, the system can trigger the online learning module, enabling idle agents to quickly learn new skills (e.g., increasing the demolition skill matching rate from 0.3 to 0.8), thereby fulfilling the role switching conditions. They were then reassigned to roles in short supply.
[0087] The task backlog rate is calculated in real time using the ratio of task queue length to processing rate. When the system detects that the task backlog for a specific role exceeds a threshold, it automatically triggers an idle device capability reconfiguration protocol (such as improving skill matching through an online learning module), effectively resolving response latency issues caused by local resource shortages.
[0088] Case b determines whether a certain type of role (such as a "demolition" task) urgently needs reinforcements. It measures the congestion level (backlog rate) of the task queue.
[0089] The purpose of introducing a swarm role switching mechanism is to solve the local resource bottleneck that may occur after task distribution, where "no one can perform certain tasks".
[0090] These two determination methods together constitute a more robust scheduling system: case a can quickly correct erroneous individual task allocations, and case b can solve the problem of uneven resource availability from a global perspective.
[0091] As one implementation of this embodiment, a capability profile matching algorithm is also included: constructing a five-dimensional capability vector for each agent. (Action accuracy, response speed, payload capacity, energy efficiency ratio, reliability), using the Hungarian algorithm to match task requirements. The capability profiling matching algorithm can accurately and standardizedly quantify the agent's capabilities, i.e., "capability profiling." The "Hungarian algorithm" is used to process this quantified data, calculating the Euclidean distance between task requirements and agent capabilities to find the optimal "task-agent" allocation scheme that minimizes the total cost.
[0092] The five-dimensional capability vector is generated through historical equipment operation logs and real-time performance diagnostics. Perception accuracy is correlated with sensor calibration parameters, and the failure rate is based on predictions from the equipment health management system. The Hungarian algorithm introduces a capability difference weight matrix during the matching phase to ensure optimal adaptation between heterogeneous equipment and complex task requirements.
[0093] Definition of a five-dimensional vector: Action accuracy (e.g., robotic arm positioning error ≤ 2cm); Response speed (command latency <100ms); Load capacity (kg); Energy efficiency ratio (task power consumption rate); Reliability (MTBF hours); Based on this profile, the Hungarian Algorithm is used to find the optimal match: a. Constructing the cost matrix (a) Calculate the Euclidean distance between the task requirement vector and the agent capability vector; b) Solve the matrix to obtain the task-agent allocation scheme with the minimum total cost.
[0094] Optionally, after constructing the pheromone field, the pheromone field can be dynamically updated by updating the pheromone concentration. The pheromone concentration is: in, The pheromone volatility coefficient ( ); For the agent in the path The increment of pheromones released is denoted as _____. Q is the task value coefficient (similar to the attraction function mentioned earlier). (The corresponding Q [task urgency]). For path The task's appeal; Let be the pheromone concentration on path ij at time t; This represents the current time step.
[0095] By dynamically updating the pheromone field, the efficiency of task allocation during the golden period after an earthquake is significantly improved.
[0096] Optionally, the attraction function The coefficients in the system can be dynamically adjusted according to the environmental context to achieve adaptive optimization: for example, the visibility factor is increased at night or in low-visibility environments. The weighting of emergency missions such as medical rescue increases the mission urgency factor. The weight.
[0097] The swarm role switching mechanism and ability profile matching scheme provide specific algorithmic support for achieving efficient and accurate task allocation.
[0098] Optionally, after decomposing the cooperative scheduling strategy into specific tasks and initially assigning them to various agents, the process further includes: converting the tasks into redundant byte stream instructions executable by the agents using a compiler. Specifically, this includes: High-level abstraction: Compiles the task description into an agent-independent intermediate representation layer, usually in JSON format; "agent-independent" means that this is a general, standardized task description format that only defines "what to do" without caring about "who does it" or "how to do it specifically".
[0099] Low-level adaptation: Based on the type of intelligent agent (such as drone or robot dog), the corresponding intelligent agent-specific backend is invoked to translate the intermediate representation layer into a bytecode instruction stream of the control protocol specific to that intelligent agent (for example, generating MAVLink instructions for drones and ROS messages for ROS-based robot dogs). Redundancy reinforcement: Hamming code encoding is applied to the bytecode instruction stream to achieve real-time detection and correction of single-byte errors, ensuring the reliability of instruction transmission in unreliable communication environments.
[0100] On the other hand, the waypoints included in the instructions issued to mobile intelligent agents such as emergency vehicles and drones must be obtained through the global accessibility matrix. The planned safe and feasible path. The role of the compiler is to translate the high-level task description (including which path points to go to) into low-level bytecode instructions (such as MAVLink instructions or ROS messages) that can be understood and executed by specific devices (such as drones or robot dogs).
[0101] In the above embodiment, the global access coefficient matrix The global traffic coefficient matrix is obtained through dynamic road network reconstruction, specifically by predicting changes in road capacity using an LSTM (Long Short-Term Memory) network and updating it in real time. The LSTM network is chosen to learn temporal patterns from historical disaster data, thereby predicting road conditions over a future period. (Global Traffic Coefficient Matrix) It has become the unified, real-time data foundation for all intelligent agents in the entire system to perform path planning.
[0102] The LSTM network takes as input historical traffic flow data sensed and processed by S101, ground-penetrating radar echoes (or humidity values), and meteorological information (such as precipitation probability) the result of its output road segment attenuation rate over a future period. ( (A higher attenuation rate indicates poorer traffic capacity). This fusion allows the prediction to consider not only traffic but also the risk of secondary disasters caused by weather and geological changes. The system calculates based on the attenuation rate. The value of the global access coefficient matrix is dynamically updated. The weight values for the corresponding road segments are assigned, and the following traffic statuses are set: Passable status: Weight Set as the standard value.
[0103] Alert / Restricted Access Status: Weight As the ratio increases, the path planning algorithm will prioritize avoiding such road sections.
[0104] Impassable status: Weight Mark it as infinity and exclude it from the feasible path.
[0105] (Note: Thresholds 0.3 and 0.6 can be dynamically adjusted based on actual road surface type and experience, for example, 0.4 and 0.7, or 0.2 and 0.5.) This mechanism provides real-time and reliable environmental constraints for path planning of all agents.
[0106] The hidden layer of the LSTM network structure contains 128 neurons (i.e., 128 hidden units), and the output... (0 represents full passage, 1 represents complete blockage). For example, in flood disaster testing, if the national highway's traffic capacity is predicted to drop to 0.4 30 minutes in advance, an emergency detour plan will be triggered.
[0107] Instead of using fixed road closure standards, the system sets a dynamically adjustable passage threshold θ. This threshold adaptively adjusts based on road surface type (e.g., asphalt, gravel, mud), significantly improving the robustness of the planning. Passage Threshold Dynamically set rules and adjust them according to road surface type We can assume the threshold adaptation rule is as follows: • Asphalt pavement: =0.3 (High load-bearing requirements) • Gravel road surface: =0.6 • Muddy road surface: =0.8 (Low traffic expectations) When the drone remote sensing identifies the road surface material as "collapseable loess," it automatically switches to... To a minimum of 0.7, to avoid the vehicle falling into danger.
[0108] when The detour topology generation is triggered and mapped to the twin guidance strategy optimization unit.
[0109] When the predicted attenuation rate of a certain road section Exceeding its corresponding dynamic threshold When needed, the system will automatically trigger detour topology generation and use an improved Dijkstra algorithm to plan the optimal safe path, including: a) Graph construction: Modeling the road network as a graph structure, where intersections are nodes and road segments are edges, and the traffic attenuation rate of the road segment is calculated. Assign the weight to the corresponding edge.
[0110] b) Algorithm Execution: Using the current location of the rescue vehicle as the source node, execute Dijkstra's algorithm, but the search target is not a single destination, but all paths whose cumulative path weights (i.e., the sum of traffic decay rates) are less than a threshold. Reachable nodes.
[0111] c) Path generation: The algorithm termination condition is modified to: when the cumulative weight of a certain path is explored. If this happens, the path is pruned. Ultimately, the path with the lowest cumulative weight among all feasible paths from the source node to the target node is the optimal detour path. The generated optimal detour path is automatically imported into the digital twin engine, feeding back into the collaborative decision-making and task allocation process. This provides real-time environmental constraints for higher-dimensional decision-making and drives the virtual rescue vehicle to rehearse the passage process.
[0112] In some optional implementations of this embodiment, the following method is also included: the cloud-based federated evolution center aggregates the model parameters of all edge nodes using a differential privacy gradient aggregation algorithm, uses this aggregation result to update the central global model parameters, optimizes the global model, obtains a global evolution strategy model, and distributes it to each edge node. In cross-regional joint rescue drills, this design can successfully resist gradient poisoning attacks while ensuring the privacy and security of patient locations. Specifically, it includes the following three core steps: 1. Gradient Adversarial Filtering: First, the system calculates the mean of the gradients uploaded by all edge nodes. with standard deviation , for satisfying The gradient vectors are marked as anomalies and removed (where...) For the first (Gradient vectors uploaded by each node). After filtering, to fill in missing data and maintain convergence, Gibbs sampling is used to reconstruct a smooth gradient distribution from the posterior distribution to defend against model poisoning attacks by malicious nodes and improve the robustness of the federated learning system.
[0113] 2. Differential privacy perturbation: Next, Laplace noise is added to the gradient distribution processed in the previous step. to satisfy - Differential privacy. Aggregated gradients. The calculation formula is: .in The scale parameter is represented as The Laplace distribution, For the global sensitivity of the function, For privacy budgets. The noise addition mechanism is designed as follows: sensitivity Maximum change in model parameters (take) Privacy Budget In practice, a bucketing strategy is used when adding noise—smaller noise is added to key parameters (such as rescue path coordinates). Add significant noise to secondary parameters (such as equipment status). While ensuring data privacy (such as rescue route coordinates), the model's generalization ability is improved. Noise is added to ensure that the original data of any individual edge node cannot be deduced from the aggregation result during this aggregation process, thereby protecting privacy.
[0114] 3. Evolution of Game Strategies: Finally, based on the strategy distribution snapshot obtained in S102, the proportion of each strategy is calculated. (i.e., adopting a strategy) The proportion of nodes in the total number of nodes. The global model generated by the above steps is optimized by Nash equilibrium solving algorithms (such as fictitious play) so that it can achieve cooperative equilibrium in a resource competition environment.
[0115] The calculation of the proportion of each strategy includes: First, construct the strategy return matrix: ;in, Representation Strategy strategy The benefits; Secondly, the proportion of iterative update strategies: in For strategy Expected returns; To select strength parameters; For strategy The proportion; For learning rate or strength selection; For strategy The average return; Finally, after multiple iterations, the policy proportions reach a convergent state, outputting a Pareto optimal policy set, thus completing the optimization of the global model.
[0116] It is worth mentioning that the "global model" refers to the basic model generated through gradient aggregation. The "global evolution strategy model" refers to the more complete version of the "global model" after it has been optimized through the "game strategy evolution" step.
[0117] In some embodiments, device failure events (robot dog disconnection, drone collision) can also be injected into the disaster path prediction process of the digital twin to calculate the probability of strategy failure. Multi-level failure events are injected into the digital twin environment: a robot dog losing contact simulates a communication link interruption, and a drone colliding with an obstacle sets random motion deviations. The strategy failure probability is calculated using the Monte Carlo method to statistically analyze the frequency of decision failures in multiple disaster scenarios, providing pre-validation of the reliability of key strategies.
[0118] In some embodiments, when the policy failure probability generated by the master model (FDQN) in S103 is detected... If the value is too high (e.g., causing frequent obstacle collisions for the drone in simulation), the system will automatically switch to a backup decision model. For example, when... The system automatically switches to a backup decision model. When the failure probability exceeds a preset threshold, the system uses memory mapping technology to directly load the backup model (such as a pre-trained deep reinforcement learning network) into the GPU memory. The switching process maintains the input and output interfaces of the original model unchanged, ensuring the continuity and stability of the emergency response process.
[0119] In some embodiments, when assigning tasks, tasks can be decomposed first, i.e., task fission, and task assignment efficiency can be improved by "cluster networking" and "cross-domain resource borrowing".
[0120] For example, in the case of multiple emergency rescue vehicles operating in coordination: Mission fission phase: The lead rescue vehicle decomposes the complex mission into a set of atomic tasks, with the following constraints: .
[0121] The complex task is decomposed into a set of atomic tasks using a directed acyclic graph (DAG), with atomicity constraints ensuring no state dependencies between tasks. The decomposition algorithm employs critical path analysis to identify parallelizable subtasks, maximizing the parallel processing capability of the rescue operation.
[0122] Cluster network allocation: Constructing a two-layer network topology for vehicles and devices: a. Each vehicle acts as a cluster head node, and the drone / robot dog it carries acts as a node within the cluster; b. Each cluster head exchanges task progress through a distributed consensus protocol; Vehicles act as cluster heads to construct an 802.11s wireless mesh backbone, and their devices access the cluster via ZigBee short-range communication. The distributed consensus protocol uses an improved Raft algorithm to exchange task progress, and heartbeat messages contain resource load fingerprint data, achieving transparency of resource status across vehicles.
[0123] Cross-domain resource borrowing: When resources are scarce within a cluster, cross-cluster scheduling of devices is triggered. ,in This serves as a resource-sharing factor in disaster scenarios.
[0124] Resource shortages are quantified by the length of the in-cluster device task queue, while neighbor availability is based on real-time inventory data exchanged via a consensus protocol. The sharing factor λ is dynamically adjusted according to the disaster severity level (Level III disaster λ=0.5 / Level I disaster λ=0.2), constructing a resilient cross-domain resource sharing mechanism.
[0125] In this embodiment, biomimetic neural network recombination is also included to achieve collaborative regeneration after equipment failure: 1) Establish node topology relationships based on each agent to obtain a device connection strength map; the node topology relationships include neuron nodes and synaptic weights.
[0126] In this embodiment, establishing the node topology relationship based on each agent includes: First, acquire multi-source data from each intelligent agent; Secondly, each agent is abstracted as a neuron node, which contains a state vector; and the communication links between agents are modeled as synaptic weights.
[0127] Specifically, the intelligent agent is abstracted as a neuron node (e.g., a drone is labeled N1, and a robot dog is labeled N2), and each intelligent agent, as a neuron, contains a state vector. Communication link modeling as synaptic weights .
[0128] As one implementation of this embodiment, the synaptic weight .in The distance is in Euclidean form (unit: km). Terrain attenuation coefficient (mountainous area) Plains ), This represents the signal attenuation rate (weather-dependent).
[0129] As another preferred implementation of this embodiment, the synaptic weight calculation adopts a hybrid attenuation model: the distance attenuation component is calibrated in real time through UWB ranging, and the signal attenuation component is dynamically updated based on the channel probe packet loss rate. The weight calculation adopts a hybrid attenuation model: in, Synaptic weights; This is a reference distance baseline value; Environmental adaptability coefficient; The signal attenuation index; This represents the distance between neuron nodes.
[0130] As can be seen from the hybrid attenuation model, This constitutes the spatial distance attenuation factor, which is the distance between nodes. After processing with a Fresnel zone diffraction model, the node spacing is calculated in real time based on UWB positioning data. The non-line-of-sight propagation error is compensated by using a Fresnel zone diffraction model.
[0131] This constitutes the electromagnetic environment attenuation factor, which is the signal attenuation index. The result, after processing by the historical environment database correction model, dynamically calibrates the signal attenuation index based on the signal attenuation rate. The model parameters are corrected by combining historical environmental databases.
[0132] This embodiment establishes a quantifiable device connection strength map, providing a dynamic data foundation for topology reconfiguration. When electromagnetic interference intensifies at a disaster site or equipment shifts, the weight values will respond in real-time to environmental changes. Communication link weights are constructed based on a signal propagation model, the spatial distance attenuation factor employs a Fresnel zone attenuation compensation algorithm, and the signal attenuation rate is measured in real-time using channel sounding packets. The node topology graph is stored in an adjacency matrix format, enabling quantitative modeling of the connection strength between devices.
[0133] In this embodiment, the multi-source data includes at least the agent's location, battery level, and task status.
[0134] 2) When the failure of a neuron node in the device connection strength map is detected, the synaptic weight redistribution is triggered to realize the reorganization of the biomimetic neural network.
[0135] Furthermore, when a neuron node malfunctions in the biomimetic neural network, it triggers synaptic weight redistribution and virtual relay wake-up.
[0136] In this embodiment, when If so, it is considered that the neuron node is faulty.
[0137] As one implementation method, node failure detection can adopt a basic heartbeat mechanism (e.g., broadcasting a HELLO message every 2 seconds, and determining node failure after 3 timeouts).
[0138] As an alternative implementation, for more robust implementation, node failure detection employs a triple heartbeat mechanism (data packet / physical layer carrier / hardware watchdog), triggering a fault flag when there are three consecutive no responses.
[0139] In this embodiment, synaptic weight redistribution can be implemented in two ways: The first method: Synaptic weight redistribution based on Dijkstra's algorithm: Remove the associated edges of the faulty node; calculate the synaptic weights of the alternative path. .in, As alternative path weights; For nodes To the intermediate node Synaptic weights, representing the node To the target node The new communication cost after rerouting following a failure; intermediate node To the target node Synaptic weights; As an intermediate node, that is, in the search from arrive When choosing an alternative path, the non-faulty third-party nodes traversed.
[0140] As an alternative implementation, the message passing mechanism based on graph neural networks is reassigned: in, Synaptic weights for alternative paths; To replace the original synaptic weights; These are the weight update coefficients, used to control the degree to which neighborhood information affects the current weight adjustment; For neighboring nodes The remaining resource index, which combines battery capacity and computing load, reflects the "capability" or "health" of the neighboring node as a relay, based on the node's state vector. (Calculated by combining "battery level" and "task status"); For set One of the neighboring nodes; For neighboring nodes With the target node The synaptic weights between neighbors represent the connection strength between the neighbor and the target. The maximum reference value for the resource index is used to normalize the resource index and ensure that the ratio is within a reasonable range. For nodes The set of neighboring nodes.
[0141] In this embodiment, if the synaptic weight of the alternative path does not reach a preset threshold after synaptic weight redistribution, a new virtual relay node is added. The addition of a new virtual relay node includes: If the synaptic weight of the alternative path After reallocation, it satisfies Then, it will wake up dormant devices (such as backup drones) to take over. This is the preset minimum acceptable connection strength.
[0142] Virtual relay nodes are dynamically activated by the software-defined radio (SDR) module of nearby devices, and wake-up commands for dormant devices are sent through out-of-band management channels, forming a collaborative fault-tolerance capability for device-level failures. For example, in a tunnel collapse scenario, this mechanism reduces the communication network recovery time to within a few seconds.
[0143] As an optional implementation of this embodiment, to ensure the sustainability of the reconfiguration process, the reconstruction process is constrained by energy consumption to ensure that the self-healing behavior itself does not cause the system to collapse due to energy depletion. Therefore, after the bionic neural network is reconfigured, if the reconfiguration energy consumption is less than the maximum energy consumption and / or the total reconfiguration energy consumption is not greater than a preset incremental threshold, the device connection strength map is reconfigured; otherwise, the reconfiguration is repeated.
[0144] As one implementation method, the energy consumption constraint of the reorganization process meets the requirements. Reorganization energy consumption Includes device wake-up power consumption Energy consumption of route reconstruction , This can be seen as a constraint on the total cost of the restructuring process. Upper limit. ( (Total available energy for the network). The dynamic optimization algorithm prioritizes low-power paths. This solution effectively ensures the sustainability of the restructuring process in areas experiencing power shortages and disasters.
[0145] As another implementation method, the incremental energy consumption of reorganization is limited. , ( Energy consumption coefficient per unit link reconfiguration for ). This can be viewed as a constraint on the amount of network topology change caused by the reconfiguration. The reconfiguration energy consumption model considers the communication overhead of link switching and the power consumption of device state switching, and the unit energy consumption coefficient ε is calibrated through device energy efficiency benchmark tests. The constraint conditions trigger the iterative termination conditions of the topology optimization algorithm, ensuring the sustainable operation of the system in energy-constrained environments.
[0146] As another implementation method, it is necessary to satisfy both of the above energy consumption constraints at the same time. If the energy consumption constraints are satisfied, then the energy consumption constraints are satisfied; otherwise, the equipment connection strength map needs to be reorganized.
[0147] By employing a biomimetic neural network-based dynamic perception-decision-reconstruction closed loop, this system achieves the synergistic goals of topology self-healing, load rebalancing, and uninterrupted service during equipment failures, while ensuring energy sustainability. This significantly enhances system resilience in complex disaster environments. This mechanism overcomes the bottleneck of traditional emergency systems where a single point of failure leads to global paralysis, providing a bio-inspired fault-tolerant paradigm for large-scale intelligent equipment clusters.
Claims
1. An emergency collaborative scheduling method based on swarm intelligence, characterized in that, include: Obtain the collaborative scheduling strategy, decompose the collaborative scheduling strategy into tasks, and generate a task list; A pheromone field is constructed for each task in the task list according to the aforementioned collaborative scheduling strategy; Treat the agent as an individual in the ant colony, and calculate the attractiveness of the task to the agent using the task attraction function, so that the agent moves to the task point with the highest attractiveness.
2. The method according to claim 1, characterized in that, Also includes: After constructing the pheromone field, the pheromone field is dynamically updated by updating the pheromone concentration; The pheromone concentration is: in, The pheromone evaporation coefficient; For the agent in the path The increase in pheromone released; Let be the pheromone concentration on path ij at time t; This represents the current time step.
3. The method according to claim 1, characterized in that, After the agent moves to the task point with the highest attractiveness, it also includes: A capability vector is constructed for each agent, and a task requirement vector is constructed for each task, resulting in a cost matrix; Based on the cost matrix, the capability matching degree between each agent and the assigned task is calculated. When the capability matching degree is less than the capability matching threshold, it is determined that the current agent is not suitable for the current role, and a role switch is triggered; or The backlog rate of assigned tasks is monitored. When the backlog rate of a task exceeds the backlog rate threshold and there is an agent in an idle state, the execution role switch is triggered.
4. The method according to claim 1, characterized in that, After task decomposition, it also includes: converting the task into redundant byte stream instructions that can be executed by the agent; The process of converting the task into a redundant byte stream of instructions executable by the target intelligent agent includes: The task description is compiled into an agent-independent intermediate representation layer; Depending on the type of intelligent agent, the intermediate representation layer is translated into a bytecode instruction stream of the control protocol specific to that intelligent agent; Hamming encoding is applied to the bytecode instruction stream to achieve real-time detection and correction of single-byte errors.
5. The method according to claim 1, characterized in that, The acquisition of the collaborative scheduling strategy includes: Acquire multi-source data from each intelligent agent, preprocess the multi-source data, and then feed it back to the corresponding intelligent agent; The agent generates a preliminary scheduling strategy based on the received multi-source data, and uses a lightweight distributed ledger to update and maintain the policy distribution snapshot of the edge nodes in the preliminary scheduling strategy. The evolution path of secondary disasters is obtained, and the evolution path of secondary disasters, preprocessed multi-source data, and policy distribution snapshots of edge nodes are input into the federated deep Q network model to generate a collaborative scheduling policy.
6. The method according to claim 5, characterized in that, The intelligent agents include drone swarms, robot dog swarms, and emergency vehicles; The drone swarm adopts an adaptive clustering protocol to construct a LoRa communication relay layer to form a topology sensing network. The robot dog swarm is equipped with UWB tags, which are used to work in conjunction with UWB base stations to construct a ground grid coordinate system. The emergency vehicle makes a movement decision based on a dynamic priority scoring model, and autonomously decides its location to stay as a communication anchor point for communication with other intelligent agents. The process of forming a movement decision based on a dynamic priority scoring model includes: The dynamic priority scoring model is used to output a comprehensive priority score for the current position, and the comprehensive priority score is used to generate a movement decision. The dynamic priority scoring model is as follows: in, The overall priority score for the target location; The signal strength factor (measured by an onboard spectrum analyzer, characterizing the communication quality at that location); The path smoothness factor (derived from the global passability matrix) (This indicates the feasibility of traveling to that location). Energy consumption penalty item (related to the vehicle's remaining battery power) negative correlation The lower the value, the higher the penalty. This is the first adjustable weighting coefficient; This is the second adjustable weighting coefficient.
7. The method according to claim 6, characterized in that, Also includes: The on-board equipment of the emergency vehicle is managed in a hierarchical manner, and on-board energy is dispatched using dynamic adjustment rules for energy consumption priority. The hierarchical management of the onboard equipment of the emergency vehicle, and the use of energy consumption priority dynamic adjustment rules for onboard energy scheduling, include: Task priorities are assigned to the vehicle-mounted equipment, and the operating status of the vehicle-mounted equipment is monitored in real time; the vehicle-mounted equipment includes critical equipment and non-critical equipment, and the critical equipment includes life support equipment. The dynamic adjustment rules for energy consumption priority include: In response to the activation signal of the life support device, a highest priority interrupt request is sent to the task scheduler, triggering the task scheduler to suspend all current non-highest priority device tasks and reallocate system resources to the life support device. When the remaining power of life support equipment falls below the first threshold, the priority of non-critical equipment is downgraded.
8. The method according to claim 5, characterized in that, The lightweight distributed ledger uses an improved Hyperledger Fabric architecture; In the improved Hyperledger Fabric architecture: Peer nodes are edge nodes, which include agents and / or edge computing gateways connected to agents, and are used to store Merkle tree hash chains of policy distribution snapshots; The communication protocol between peer nodes uses the Gossip protocol for data synchronization. The ordering service includes: receiving policy distribution snapshot update requests from each peer node, and reaching a consensus on the update order using a simplified Byzantine fault-tolerant consensus mechanism; The simplified Byzantine fault-tolerant consensus mechanism includes: In response to the policy distribution snapshot update request, the master node assigns a sequence number to the policy distribution snapshot update request and then broadcasts a pre-preparation message to all backup nodes. Once any backup node receives and verifies the pre-preparation message, it uses its current backup node's private key share to sign the digest of the pre-preparation message, generates a signature share, and broadcasts the signature share. If any node collects more than a first quantity threshold for the same message, then the collected signatures for the same message are aggregated into a threshold signature, and the threshold signature is broadcast. If any node generates or receives the signature threshold, it indicates that all nodes have reached a consensus on the pre-prepared message, and write the policy distribution snapshot update request into the local ledger and update the policy distribution snapshot.
9. The method according to claim 5, characterized in that, The process of inputting the evolution path of the secondary disaster, preprocessed multi-source data, and policy distribution snapshots of edge nodes into the federated deep Q-network model to generate a collaborative scheduling policy includes: Each edge node trains a local DQN model based on preprocessed local multi-source data and synchronized global state data; During training, the evolution path of the secondary disaster is embedded into the DQN model in the form of a state transition probability matrix for iteration; The local DQN model periodically integrates all gradient parameters output during training using the differential privacy gradient aggregation algorithm to generate an updated global model, which then updates the local DQN model. Edge nodes use the updated local DQN model to output a collaborative scheduling strategy.
10. The method according to claim 9, characterized in that, Also includes: The global model is optimized by aggregating the model parameters of all edge nodes and then distributed to each edge node. The process of generating a global evolutionary strategy model by aggregating the model parameters of all edge nodes includes: Calculate the mean and standard deviation of the gradients of all edge nodes, identify and remove outlier gradient vectors, and reconstruct the gradient distribution from the posterior distribution using Gibbs sampling. Laplace noise is added to the reconstructed gradient distribution, and the proportion of each policy in the policy distribution snapshot is calculated. A global evolutionary policy model is generated by the Nash equilibrium solution algorithm.