A space-time grid closed-loop multi-machine obstacle avoidance anchor point allocation system and method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-08-11
AI Technical Summary
在多机计划于极短时间差内到达同一锚点时,该方法难以提前识别锚点占用冲突,高密度流量下易产生悬停对峙
[0021] This invention introduces four-dimensional spatiotemporal grid coding, expanding anchor point resource management from a static management method based solely on spatial location to a spatiotemporally integrated management method that incorporates both spatial and temporal dimensions. As a result, the system can identify potential conflicts in advance based on the estimated arrival time of candidate anchor point reservation requests, even before the drone arrives at the anchor point. This allows for proactive prediction of conflicts arising from occupancy of the same anchor point at different times, reducing the problem of missed timing conflicts due to the lack of a temporal dimension.
Smart Images

Figure CN122347887B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-altitude unmanned aerial vehicle (UAV) traffic management technology, specifically to a multi-UAV obstacle avoidance anchor point collaborative allocation system and method based on spatiotemporal grid closed-loop perception in a structured low-altitude corridor network. Background Technology
[0002] With the rapid development of the low-altitude economy, the operational density of drones in urban and suburban airspace continues to increase, and scenarios where multiple drones simultaneously encounter sudden obstacles (such as sudden weather events, temporary no-fly zones, and foreign objects in the air) are becoming increasingly frequent. In such high-density and high-dynamic scenarios, multiple drones need to compete for, occupy, and restore safe obstacle avoidance anchor points (temporary hovering or waypoints) in a very short time, which places stringent demands on the collaborative management and scheduling capabilities of airspace resources.
[0003] Currently, existing technologies typically employ static spatial point allocation mechanisms and serial scheduling architectures when dealing with multi-machine sudden obstacle avoidance scenarios. However, in practical engineering applications, these technologies still have the following limitations: First, the coarse-grained nature of airspace resource management makes it difficult to predict timing conflicts in advance. Existing anchor point management methods mostly use static spatial points as the basic unit. For example, Chinese patent application CN120656340B discloses a three-dimensional scheduling method and system applicable to low-altitude multi-aircraft mixed take-off and landing fields. It constructs a four-dimensional spatiotemporal grid based on BeiDou grid codes and is mainly used for route conflict prediction in global path planning of take-off and landing fields, rather than for predicting timing conflicts related to the reservation and occupancy of multi-aircraft obstacle avoidance anchor points. When multiple aircraft plan to arrive at the same anchor point within a very short time difference, this method struggles to identify anchor point occupancy conflicts in advance, easily leading to hovering standoffs under high-density traffic.
[0004] Second, conflict detection and allocation decisions are sequential and fragmented, resulting in a delayed scheduling response. Existing obstacle avoidance scheduling methods typically design detection and resolution as sequential steps. For example, Chinese patent application CN121325577A discloses a conflict resolution method for air-to-ground unmanned swarms based on multi-agent reinforcement learning, but it still relies on path replanning after a conflict occurs as its core logic, making conflict detection and decision-making actually sequential steps. The decision-making system cannot provide real-time feedback on the predicted future conflict probability to drive anchor point pre-allocation, leading to a significant lag in scheduling response in highly dynamic scenarios involving multiple aircraft experiencing sudden obstacle avoidance.
[0005] Third, energy constraints rely on passive responses and lack full-range energy awareness. Existing energy management methods largely depend on comparing the current battery level with a fixed threshold to trigger protection. For example, Chinese patent application CN115867459A discloses a system and method for battery capacity management in a UAV fleet. It sets a fixed capacity threshold for each UAV in the fleet, compares the current remaining battery level with this threshold, and triggers battery capacity management when the level falls below the threshold. This method only relies on a static comparison between the current battery level and the fixed threshold, failing to incorporate random dynamic energy consumption caused by temporary waiting and route replanning in multi-aircraft obstacle avoidance scenarios into the full-range assessment, and lacks the ability to dynamically predict accumulated energy consumption over future flight segments. When multiple aircraft are queuing for obstacle avoidance, the fixed threshold mechanism is still prone to energy depletion due to accumulated waiting energy consumption.
[0006] Fourth, existing technologies lack effective anchor point collaborative release mechanisms, leading to secondary congestion during the resumption of traffic flow after obstacles are removed. Current obstacle avoidance scheduling methods primarily focus on anchor point allocation and conflict resolution, lacking systematic collaborative management of how to orderly reintegrate the backlog of drones into the main corridor after obstacles are removed. For example, Chinese patent application CN110888453A discloses a method for autonomous drone flight based on LiDAR data to construct a 3D real-world scene, but it mainly addresses single-drone trajectory correction and does not address traffic control in multi-drone collaborative recovery scenarios. Existing academic research (Computer Science, Vol. 51, No. 6, 2024) proposes fast trajectory recovery algorithms for obstacle avoidance scenarios, but their recovery logic still focuses on single-drone trajectory optimization, without considering the collaborative release order and inlet traffic constraints among multiple anchor points. When obstacles are removed, a large number of backlogged drones simultaneously flood into the main corridor. Existing technologies lack structural control measures for release order, rate, and inlet traffic constraints, easily creating instantaneous traffic peaks at the main corridor entrance and causing secondary congestion. This disorderly release not only reduces the operational efficiency of the low-altitude corridor, but also weakens the system's resilience and long-term stability under continuous disturbances.
[0007] The four contradictions mentioned above are interdependent and jointly restrict the emergency carrying capacity of low-altitude corridors in multi-aircraft sudden obstacle avoidance scenarios. Current technologies lack a systematic solution that can simultaneously resolve all four contradictions. Summary of the Invention
[0008] The purpose of this invention is to provide a spatiotemporal grid closed-loop multi-aircraft obstacle avoidance anchor point allocation system and method. By introducing a four-dimensional spatiotemporal grid coding mechanism, a real-time closed-loop mechanism for conflict detection and allocation decision, a full-range energy closed-loop management mechanism, and a virtual anchor point chain cascade allocation and symmetrical release mechanism, the invention simultaneously solves the shortcomings of the prior art at the mechanism level.
[0009] This invention provides a spatiotemporal grid closed-loop multi-machine obstacle avoidance anchor point allocation system, including a corridor network construction module and an emergency obstacle avoidance collaborative scheduling module; The corridor network construction module is used to construct a structured low-altitude corridor network and determine the location of obstacle avoidance anchor points. The emergency obstacle avoidance and collaborative scheduling module includes: The four-dimensional spatiotemporal grid management unit is used to perform three-dimensional segmentation of the target airspace based on the Beidou grid code and introduce the time dimension to form a four-dimensional spatiotemporal grid code and a four-dimensional spatiotemporal grid unit, and to predict spatiotemporal conflicts in UAV anchor point reservation requests. The graph neural network conflict prediction unit is used to construct a spatiotemporal graph structure with the four-dimensional spatiotemporal grid units as nodes and the temporal correlation of the UAV flight path as edges, and input the spatiotemporal graph structure into the spatiotemporal graph convolutional network to output the conflict probability of each four-dimensional spatiotemporal grid unit in the future time window. A multi-agent deep reinforcement learning allocation unit is used to dynamically adjust the anchor point allocation decision by taking the conflict probability as the input of the reward function. The full-range energy closed-loop management unit is used to predict remaining energy consumption, calculate the energy safety margin of candidate anchor points, and trigger energy priority preemption. The virtual anchor chain management unit is used to construct primary and secondary anchor chain, and to cascade and allocate anchors when the primary anchor is saturated. The resumption release control unit is used to control the release based on the release priority score and symmetrical release sequence, and to adjust the release rate so that the instantaneous flow at the main corridor entrance does not exceed the capacity limit.
[0010] As a preferred embodiment of the present invention, the full-range energy closed-loop management unit is configured with an energy priority interruption mechanism; When the energy safety margin of the candidate anchor point corresponding to the drone is lower than the preset margin threshold, or the predicted remaining energy of the drone is lower than the dynamic emergency threshold, the energy priority interruption mechanism is triggered to raise the allocation priority of the drone to the highest level and allow it to preempt the anchor point time slot that has been reserved but not yet occupied. The dynamic emergency threshold is calculated in real time based on the predicted flight energy consumption from the UAV's current location to the nearest alternate landing site, the current hovering waiting energy consumption, and the safety redundancy power.
[0011] As a preferred embodiment of the present invention, the virtual anchor chain management unit dynamically determines the cascading allocation order based on the anchor occupancy status, anchor capacity, and the reachability cost between anchors. When the capacity of the main anchor point reaches the preset saturation threshold, the UAV is cascaded and allocated from the main anchor point to the secondary anchor point according to the cascade allocation sequence. After the obstacle is removed, the multi-agent deep reinforcement learning allocation unit generates a symmetric release order by combining the conflict probability, the UAV energy safety margin, and the task priority under the reverse constraint of the cascade allocation order. The resumption release control unit calculates a release priority score based on the spatial distance from each anchor point to the main corridor entrance, the average energy safety margin of the UAVs within the anchor point, and the average task priority. It then adjusts the exit order of UAVs within the same release batch according to the release priority score and adjusts the release rate of each anchor point according to the instantaneous flow constraint at the main corridor entrance to ensure that the instantaneous flow at the main corridor entrance does not exceed the carrying capacity limit.
[0012] As a preferred embodiment of the present invention, the graph neural network conflict prediction unit predicts the conflict probability in the following manner: The candidate anchor points and their expected arrival times of each UAV are mapped together to the corresponding four-dimensional spatiotemporal grid cells. A spatiotemporal graph structure is constructed with each spatiotemporal grid cell as a node and the temporal correlation of the UAV flight path as an edge. The spatiotemporal graph structure is propagated and predicted by a spatiotemporal graph convolutional network, and the conflict probability of each spatiotemporal grid cell in the future time window is output. When the conflict probability exceeds a preset threshold, it is determined that the corresponding four-dimensional spatiotemporal grid unit has a conflict risk, and the multi-agent deep reinforcement learning allocation unit is triggered to dynamically adjust the anchor point allocation decision according to the conflict probability. The conflict probability prediction is completed before the UAV reaches the anchor point to achieve predictive conflict avoidance.
[0013] As a preferred embodiment of the present invention, the multi-agent deep reinforcement learning allocation unit adopts a centralized training and distributed execution paradigm; During the training phase, a centralized value network is trained using global state information that includes all UAV states, anchor point states, conflict probabilities of four-dimensional spatiotemporal grid cells, and energy safety margins. The policy network parameters are then updated based on the advantage function estimate output by the centralized value network. During the execution phase, each UAV independently generates anchor point allocation decisions based solely on its local observation information, without the need for real-time centralized coordination. The reward function of the multi-agent deep reinforcement learning allocation unit includes four components: safety reward, efficiency reward, energy reward, and fairness reward. The weights of the four components are dynamically configured according to the task scenario type to adapt to the differentiated needs of optimization objectives in urban logistics, medical delivery, and emergency rescue scenarios.
[0014] As a preferred embodiment of the present invention, the multi-agent deep reinforcement learning allocation unit in the emergency obstacle avoidance collaborative scheduling module adopts a multi-source information fusion network with a multi-head attention mechanism. The multi-source information fusion network uses the UAV feature vector as a query and the anchor point feature vector as a key and value. It calculates the matching score between the UAV and each candidate anchor point through a multi-head attention mechanism. The matching score comprehensively reflects the coupling relationship between task priority, arrival timeliness, remaining battery urgency, flight direction matching degree and anchor point resource availability. The optimal anchor point allocation scheme is output based on the matching score through PointerNetworks (Ptr-Net).
[0015] Compared to traditional linear weighted scoring methods, multi-head attention mechanisms can adaptively capture the non-linear coupling relationship between multi-dimensional indicators without the need for manual pre-setting of fixed weights for each indicator.
[0016] As a preferred embodiment of the present invention, the full-range energy closed-loop management unit further includes a three-level emergency rescue protocol; When the drone's energy status enters the emergency zone, a rescue response of the corresponding level is triggered based on the energy safety margin range corresponding to the ratio of remaining energy to the distance to the nearest alternate landing site. The first-level response corresponds to an energy safety margin that is in the upper part of the preset emergency zone, and the execution priority is increased and the anchor point is preempted. The Level 2 response corresponds to an energy safety margin in the middle of the preset emergency zone, and executes guidance to the nearest alternate airport and flight replanning. The Level 3 response corresponds to an energy safety margin in the lower part of the preset emergency zone, and involves designating an emergency landing area and clearing the surrounding airspace. The trigger threshold ranges of each level of response correspond one-to-one with the numerical range of the energy safety margin, and the switching between each level of response is driven by the real-time calculation results of the energy safety margin.
[0017] As a preferred embodiment of the present invention, the method for constructing the virtual anchor chain in the virtual anchor chain management unit includes: The location, capacity, direction of arrival, and historical usage characteristics of each anchor point are encoded, and the functional similarity between anchor points is calculated. Anchor points are clustered based on functional similarity and spatial proximity. Within each cluster, the main anchor point is determined based on a weighted comprehensive score of capacity, accessibility, and historical usage frequency. Establish a directed topology from primary anchor points to secondary anchor points, where the edge weights of the directed topology reflect the spatial reachability cost between anchor points; The directed topology of the virtual anchor chain is dynamically updated as the anchor occupancy status changes.
[0018] As a preferred embodiment of the present invention, the corridor network construction module includes a risk map construction unit and a modular corridor unit library; The risk map construction unit integrates geographic information, building information, population distribution information, noise-sensitive area information, meteorological information, and airspace occupancy information of the target airspace to construct a three-dimensional risk map; The modular corridor unit library predefines standardized three-dimensional corridor units that match the drone's cruise, turning, takeoff, and node switching actions. Based on the three-dimensional risk map, standardized three-dimensional corridor units are selected and combined from the modular corridor unit library to generate a structured low-altitude corridor network; The location of obstacle avoidance anchor points is determined by the risk distribution in the three-dimensional risk map and the directed topology of the structured low-altitude corridor network.
[0019] Another objective of this invention is to provide a collaborative allocation method for a spatiotemporal grid closed-loop multi-machine obstacle avoidance anchor point allocation system, comprising the following steps: Step S1: By integrating the geographic information, building information, population distribution information, noise-sensitive area information, meteorological information and airspace occupancy information of the target airspace through the corridor network construction module, a three-dimensional risk map is constructed. Based on the three-dimensional risk map, standardized three-dimensional corridor units are selected and combined from the modular corridor unit library to generate a structured low-altitude corridor network and determine the location of obstacle avoidance anchor points. Step S2: The target airspace is divided into three dimensions and a time dimension is introduced by the four-dimensional spatiotemporal grid management unit based on the Beidou grid code to form a four-dimensional spatiotemporal grid code; the anchor point reservation requests of each UAV are mapped to the corresponding four-dimensional spatiotemporal grid unit, and reservation conflicts are predicted in the spatiotemporal dimension. Step S3: Construct a spatiotemporal graph structure using each four-dimensional spatiotemporal grid unit as a node and the temporal correlation of the UAV flight path as an edge through the graph neural network conflict prediction unit. Output the conflict probability of each four-dimensional spatiotemporal grid unit in the future time window through the spatiotemporal graph convolutional network. When the conflict probability exceeds the preset threshold, it is determined that a conflict has occurred, and the collaborative allocation process is triggered. Step S4: The conflict probability is used as the input component of the reward function by the multi-agent deep reinforcement learning allocation unit. The weights of the four components, namely safety reward, efficiency reward, energy reward and fairness reward, are combined to drive the dynamic adjustment of the anchor point allocation decision in real time. Each UAV independently generates an allocation decision based on its local observation information. The matching score between the UAV and each candidate anchor point is calculated by the multi-source information fusion network with multi-head attention mechanism. The optimal anchor point allocation scheme is output through the pointer network. Step S5: The full-range energy closed-loop management unit dynamically predicts the remaining energy consumption of each UAV at each decision step and calculates the energy safety margin corresponding to each candidate anchor point. When the energy safety margin of a UAV is lower than the dynamic emergency threshold calculated in real time based on its distance to the nearest alternate landing site and expected energy consumption, the energy priority preemption mechanism is triggered to raise the allocation priority of the UAV to the highest level and allow it to preempt the anchor point time slot that has been reserved but not yet occupied. Step S6: After the obstacle is removed, the recovery release order is generated by the multi-agent deep reinforcement learning allocation unit. The release priority score is calculated based on the spatial distance from each anchor point to the main corridor entrance, the average energy safety margin of the UAVs within the anchor point, and the average task priority. The UAVs are controlled to exit from the secondary anchor point to the main anchor point in a symmetrical release order from high to low release priority score. At the same time, the release rate of each anchor point is adjusted to ensure that the instantaneous flow at the main corridor entrance does not exceed the capacity limit. The recovery path for each UAV to access the main corridor or backup corridor is replanned. Beneficial effects
[0020] Effect 1: Improves the ability to identify timing conflicts in advance.
[0021] This invention introduces four-dimensional spatiotemporal grid coding, expanding anchor point resource management from a static management method based solely on spatial location to a spatiotemporally integrated management method that incorporates both spatial and temporal dimensions. As a result, the system can identify potential conflicts in advance based on the estimated arrival time of candidate anchor point reservation requests, even before the drone arrives at the anchor point. This allows for proactive prediction of conflicts arising from occupancy of the same anchor point at different times, reducing the problem of missed timing conflicts due to the lack of a temporal dimension.
[0022] Effect 2: Achieve closed-loop linkage between conflict detection and allocation decision-making.
[0023] This invention directly uses the conflict probability output by the graph neural network as the input component of the reward function of a multi-agent deep reinforcement learning framework, enabling the conflict detection results to be fed back to the anchor point allocation decision process in real time. This breaks through the serial processing mode in existing technologies where conflict detection and allocation decision are separated. Through the above closed-loop linkage mechanism, the system can proactively correct the allocation strategy before conflict risk arises, realizing a shift from passive response to proactive avoidance, thereby improving the real-time performance and intelligence level of allocation decision-making.
[0024] Effect 3: Improve the efficiency of anchor point resource utilization.
[0025] This invention maps candidate anchor point reservation requests to a four-dimensional spatiotemporal grid and coordinates allocation by comprehensively considering factors such as conflict probability, arrival timeliness, task priority, flight direction matching degree, and anchor point resource availability. This allows the system to prioritize anchor point resources with lower conflict risk and higher matching degree. Compared with traditional static allocation or linear weighted allocation methods, this mechanism can reduce duplicate reservations, invalid occupancy, and inefficient idleness of anchor points, thereby improving the overall utilization rate and scheduling efficiency of anchor point resources.
[0026] Effect 4: Enhanced energy supply capabilities throughout the entire flight.
[0027] This invention constructs a closed-loop energy management mechanism for the entire flight path. At each decision step, it dynamically predicts the remaining energy consumption of the UAV throughout its flight and calculates the energy safety margin corresponding to each candidate anchor point, using this margin as a crucial constraint for anchor point allocation decisions. This mechanism comprehensively considers factors such as flight distance, hovering energy consumption, and alternate landing requirements, transforming energy management from post-flight threshold triggering to dynamic perception throughout the entire process. This reduces the risk of energy depletion due to UAVs waiting, detouring, or path changes, thereby improving the system's energy security capabilities.
[0028] Effect 5: Improves scheduling flexibility in scenarios where the main anchor point is saturated.
[0029] This invention introduces a virtual anchor chain cascading allocation mechanism, organizing spatially adjacent and functionally complementary anchors into a chain structure with a primary-secondary hierarchical relationship. When the capacity of the primary anchor is saturated, the system can automatically trigger cascading allocation to secondary anchors, thereby avoiding the concentration of allocation requests on a single anchor. This mechanism enables hierarchical access and orderly expansion of anchor resources, enhancing the system's scheduling flexibility and carrying capacity in high-density, multi-request scenarios.
[0030] Effect 6: Reduce the risk of secondary congestion during the resumption of traffic flow.
[0031] This invention employs a symmetrical release sequence from secondary anchor points to primary anchor points during the recovery phase after obstacle removal, and dynamically controls the release rate by combining the release priority scores of each anchor point and the instantaneous flow constraints at the main corridor entrance. This approach allows backlogged drones to resume operation smoothly and gradually, preventing a concentrated influx of drones into the main corridor entrance after obstacle removal, thus avoiding new flow peaks. Structurally, this reduces the risk of secondary congestion and improves the smoothness and controllability of the recovery process.
[0032] Effect 7: Improves the overall stability, robustness, and adaptability of the system.
[0033] This invention integrates corridor network construction, conflict prediction, anchor point allocation, energy management, and resumption control into a unified and collaborative design, forming a complete closed-loop mechanism covering front-end modeling, mid-end scheduling, and back-end recovery. This mechanism can adapt to low-altitude operation scenarios with high density and strong disturbances involving multiple aircraft, and can dynamically adjust decision-making strategies according to different airspace environments, mission types, and traffic densities, thereby significantly improving the overall operational stability, robustness, and emergency response capabilities of the system. Attached Figure Description
[0034] Figure 1 This is a schematic diagram of the overall system structure of the present invention; Figure 2 A schematic diagram of a four-dimensional spatiotemporal grid coding and collision detection closed-loop mechanism; Figure 3 A schematic diagram of a multi-agent deep reinforcement learning framework; Figure 4 This is a schematic diagram illustrating the closed-loop relationship between conflict probability prediction and the reward function. Figure 5 A flowchart of the closed-loop energy management mechanism for the entire flight; Figure 6 A schematic diagram illustrating the cascaded allocation and symmetrical release sequence of virtual anchor points; Figure 7 This is a flowchart of the overall process of the method of the present invention. Detailed Implementation
[0035] Example 1
[0036] System overall structure, like Figure 1 As shown, the spatiotemporal grid closed-loop multi-aircraft obstacle avoidance anchor point allocation system provided by the present invention includes a corridor network construction module and an emergency obstacle avoidance collaborative scheduling module. The corridor network construction module includes a risk map construction unit and a modular corridor unit library. The emergency obstacle avoidance collaborative scheduling module includes a four-dimensional spatiotemporal grid management unit, a graph neural network conflict prediction unit, a multi-agent deep reinforcement learning allocation unit, a full-range energy closed-loop management unit, a virtual anchor point chain management unit, and a recovery release control unit.
[0037] Corridor network building module, The risk map construction unit integrates geographic information, building information, population distribution information, noise-sensitive area information, meteorological information, and airspace occupancy information of the target airspace to construct a three-dimensional risk map. This map is used for any spatial point within the airspace. Its risk level Defined as: ; in, For the first Risk factors at spatial points Normalized score at the location, For the corresponding weight coefficients, satisfying , This represents the total number of risk factors, which include building obstruction rate, population density, noise sensitivity, weather hazard index, and historical frequency of airspace conflicts. The placement of obstacle avoidance anchor points is determined by the risk distribution in the 3D risk map and the directed topology of the structured low-altitude corridor network. Priority is given to placing them at corridor intersections, entrances to high-risk areas, and intermediate nodes of long-distance corridors, with specific risk levels required for each anchor point's location. Below the preset safety threshold .
[0038] Example 2 Four-dimensional spatiotemporal grid coding and collision detection closed-loop mechanism This embodiment details the four-dimensional spatiotemporal grid coding mechanism and its closed-loop integration with a multi-agent deep reinforcement learning framework. The multi-agent deep reinforcement learning allocation unit integrates the multi-agent deep reinforcement learning framework and is used to perform anchor point allocation policy learning, reward function calculation, and policy network updates. The aforementioned four-dimensional spatiotemporal grid coding mechanism and closed-loop integration method are the core technical solutions for resolving the first and second systemic contradictions in the background technology.
[0039] (a) Four-dimensional spatiotemporal grid coding, like Figure 2 As shown, the four-dimensional spatiotemporal grid management unit performs three-dimensional segmentation of the target airspace based on the BeiDou grid code, generating spatial grid units uniquely identified by longitude, latitude, and altitude layer indices. On this basis, the time axis is discretized at fixed time intervals, with each time interval corresponding to a time grid unit. The spatial grid code and the time grid code are combined to form the four-dimensional spatiotemporal grid code. ; in, This represents a four-dimensional spacetime grid cell; This indicates the longitude index corresponding to the four-dimensional spatiotemporal grid cell; This indicates the dimensional index corresponding to the four-dimensional spatiotemporal grid cell; This indicates the height layer index corresponding to the four-dimensional spatiotemporal grid cell; This indicates the time grid index corresponding to the four-dimensional spatiotemporal grid cell.
[0040] Time grid index From the current moment With time resolution (fixed time interval) Jointly determined: ; in, Indicates a time grid index; Indicates the current time to be encoded; Indicates the system initialization time; Indicates time resolution, i.e., a fixed time interval; Indicates floor function; time resolution The system dynamically adjusts the value based on airspace traffic density: a smaller value (preset value) is used in high-density scenarios to improve detection accuracy, while a larger value (preset value) is used in low-density scenarios to reduce computational overhead; thus maintaining a balance between accuracy and efficiency.
[0041] For drones Submitted for anchor point The system will anchor the reservation request. Spatial coordinates and drones The estimated arrival times are jointly mapped to the corresponding four-dimensional spatiotemporal grid cells: ; in, Indicates drone For anchor points The four-dimensional spatiotemporal grid unit corresponding to the reservation request; Indicates the drone's serial number; Indicates the anchor point number; Indicates anchor point The longitude index of the spatial grid in which it is located; Indicates anchor point The latitude index of the spatial grid in which it is located; Indicates anchor point The height layer index of the spatial grid; Indicates drone Expected arrival at anchor point Time; Indicates the system initialization time; Indicates time resolution.
[0042] The system writes the reservation request into the reservation queue of the four-dimensional spatiotemporal grid cell. If another drone's reservation request is mapped to the same or adjacent four-dimensional spatiotemporal grid cell, the system can identify the potential conflict at the current decision moment without waiting for the two drones to actually approach the anchor point before triggering the detection.
[0043] Two four-dimensional spatiotemporal grid units and The spatiotemporal distance between them is defined as: ; in, and These represent two four-dimensional spatiotemporal grid cells to be compared; express and The combined spatiotemporal distance between them; express and The spatial Euclidean distance between them; express and The time interval between them; Normalized weight coefficients representing spatial dimensions; This represents the normalized weight coefficients for the time dimension, and and All are integers. (Settings are used to...) and The system can adjust the relative influence of spatial proximity and temporal proximity in conflict determination according to different scenarios.
[0044] when Below the safety interval threshold At that time, it is determined that there is a risk of spatiotemporal conflict between the two reservation requests, triggering the conflict probability prediction process. Among these, This represents a preset safety interval threshold, used to determine whether there is a potential spatiotemporal conflict between reservation requests corresponding to two four-dimensional spatiotemporal grid units.
[0045] (ii) Graph Neural Network Conflict Probability Prediction The graph neural network conflict prediction unit models the four-dimensional spatiotemporal grid unit as a spatiotemporal graph structure. .in, This represents a spacetime graph structure composed of four-dimensional spacetime grid cells; Represents a set of nodes, a set of nodes Each node in Corresponding to a four-dimensional spacetime grid unit ; Let edge set be an edge set. Each edge in Connect two spatially adjacent or temporally continuous four-dimensional spatiotemporal grid cells. and .in, Indicates the connection node With nodes The edge. edge weight Defined as: ; in, Represents a node With nodes The edge weights are used to characterize the correlation strength between two four-dimensional spatiotemporal grid cells; This represents an exponential function with the natural constant e as its base. Represents a four-dimensional spacetime grid cell and The combined spatiotemporal distance between them; This represents the scale parameter that controls the decay of edge weights as the spatiotemporal distance increases. The larger the value, the slower the edge weight decays with increasing distance. The smaller the value, the faster the edge weight decays with increasing distance. Based on the above definition of edge weight, two four-dimensional spatiotemporal grid cells that are spatially closer or have a shorter time interval have a greater correlation weight.
[0046] Graph neural networks employ a spatiotemporal graph convolutional network structure, the first... Layer nodes The hidden state update rule is as follows: ; in, For nodes The set of neighboring nodes, and They are nodes and The degree; Represents a node In the Hidden state vectors in a layered graph convolutional network; Represents a node In the Hidden state vectors in a layered graph convolutional network; Represents a node In the Hidden state vectors in a layered graph convolutional network; Represents a node Belongs to node , neighboring nodes; This represents the normalization factor used to normalize the neighborhood aggregation terms. Indicates the first The spatial graph convolution weight matrix of the layer is used to extract the spatial correlation features between adjacent spatiotemporal grid cells; Indicates the first The time autoregressive weight matrix of the layer is used to preserve the state evolution characteristics of the node itself in the time dimension; Indicates the first Layer bias vector; This represents a non-linear activation function.
[0047] It should be noted that this article adopts This represents the nonlinear activation function in a graph convolutional network, to distinguish it from the scale parameter in the edge weight formula. To avoid confusion of symbols.
[0048] go through After layer graph convolution, the hidden state of each node is mapped to the collision probability of the corresponding four-dimensional spatiotemporal grid cell through the output layer: ; in, Represents a four-dimensional spacetime grid cell The probability of a conflict occurring within a future time window; This refers to the Sigmoid function, used to map the output to... interval; This represents the output layer weight matrix; Represents a node go through The final hidden state vector obtained after layer graph convolution; Indicates the output layer bias term; This represents the total number of layers in the graph convolutional network. When... Exceeding the preset conflict probability threshold When the four-dimensional spatiotemporal grid cell is identified as having a conflict risk, a collaborative allocation process is triggered. This represents the preset collision probability safety threshold.
[0049] (III) Closed-loop integration of conflict detection and allocation decision-making. like Figure 4 As shown, the collision probability The reward function of the multi-agent deep reinforcement learning framework is continuously fed back, forming a real-time closed loop. Specifically, for drones... Assigned to anchor point The system selects candidate solutions based on the corresponding four-dimensional spatiotemporal grid units. The conflict probability is used to calculate the safety reward component.
[0050] The safe reward component in the reward function is defined as follows: ; in, This represents the safety bonus received by the drone under the current anchor point allocation scheme. This represents the positive reward coefficient given under safe conditions; This indicates the penalty coefficient applied when a collision occurs; This represents the weighting coefficient used to penalize based on the probability of conflict. Indicates drone Minimum distance between the nearest obstacle or other drone This is the safe distance threshold; Indicates drone Reservation Anchor Point Time corresponds to four-dimensional spacetime grid unit The probability of conflict; This indicates the preset collision probability safety threshold.
[0051] Compared to existing technologies where security rewards rely solely on current spatial distance, this embodiment considers the probability of collisions within a future time window. By introducing the calculation of safety rewards, the policy network not only avoids current spatial collision risks when optimizing allocation schemes, but also proactively avoids spatiotemporal conflict risks within future time windows. When it is higher, even at the current moment Still greater than the safe distance threshold Security rewards are also reduced due to conflict probability penalties, prompting the policy network to actively choose candidate anchors with lower conflict probabilities.
[0052] In addition to the safety reward component, this embodiment further sets up an efficiency reward component, an energy reward component, and a fairness reward component, and weights and combines each reward component to form a total reward function.
[0053] Efficiency reward components are defined as follows: ; in Indicates drone The efficiency bonus obtained under the current anchor point allocation scheme; Indicates drone The estimated waiting time at the anchor point; Indicates drone The estimated flight time required to fly from the current location to the assigned anchor point or target path node; This represents the waiting time penalty weighting coefficient; This represents the flight time penalty weighting coefficient. Since waiting time typically causes anchor point resource consumption, energy depletion, and subsequent scheduling delays, therefore... This is to demonstrate that waiting time has a greater negative impact on scheduling efficiency compared to flight time.
[0054] The fair reward component is defined as: ; in, This indicates the fair reward amount corresponding to the current allocation scheme; This represents the waiting time vector of all drones currently participating in the scheduling. This represents a function that calculates the standard deviation over the input latency vector. Specifically, it is expressed as: ; in, This indicates the total number of drones currently participating in the scheduling process; These represent the expected waiting time for each drone. The larger the standard deviation of the waiting time, the greater the difference in waiting time between different drones, and the worse the fairness of the allocation. Therefore, the fair reward component is negative to encourage the system to reduce the difference in waiting time between different drones.
[0055] The total reward function is: ; in, This represents the total reward obtained by the multi-agent deep reinforcement learning framework at the current decision step; Indicates the amount of safety reward; Indicates the efficiency reward component; Indicates the amount of energy reward; Indicates the fair amount of reward; , , , Let these represent the weighting coefficients for safety rewards, efficiency rewards, energy rewards, and fairness rewards, respectively, and satisfy the following conditions: ; in, , , , All values are non-negative. The weighting coefficients mentioned above are dynamically configured according to the task scenario. For example, in emergency rescue scenarios, the weights of safety rewards and energy rewards are increased, while in urban logistics scenarios, the weights of efficiency rewards and fairness rewards are increased, thus adapting to the differences in optimization objectives under different task scenarios.
[0056] The multi-agent deep reinforcement learning framework is trained using the MAPPO algorithm, and the objective function for updating the policy network is: ; in, ; in, Let represent the objective function of the policy network to be optimized; Indicates the current policy network parameters; This represents the network parameters of the old policy before the update; Indicates the training sampling time step Take the expected value from the data above; Indicates time step The environmental conditions; This indicates the state of the agent in the environment. The action to be performed; This indicates the current state of the policy network in the environment. Select action The probability of; This indicates the old policy network in the environmental state. Select action The probability of; This represents the probability ratio between the current strategy and the old strategy; Indicates time step The estimated value of the dominance function; Represents the clipping function; This represents the PPO pruning parameter, used to limit the magnitude of policy updates; This means taking the smaller value of the two input terms. This pruning mechanism avoids excessively large single updates to the policy network, thus improving the stability of the training process.
[0057] The advantage function is calculated using generalized advantage estimation (GAE): ; in, ; in, Indicates time step The estimated value of the dominance function; Indicates the time step from the current time. The time offset for backward recursion; This represents the discount factor, used to measure the impact of future rewards on current decisions; This represents the GAE smoothing parameter, used to balance the bias and variance. Indicates time step The timing difference error; Indicates time step The timing difference error; Indicates time step The instant reward received; Indicates time step The environmental conditions; Indicates the value network in environmental state Estimation of the state value of the output; Indicates the value network in environmental state Estimation of the state value of the output; This represents the parameters of the value network.
[0058] Example 3 A closed-loop energy management mechanism for the entire flight. This embodiment details the closed-loop energy management mechanism for the entire flight, which is the core technical solution to resolve the third systemic contradiction in the background technology.
[0059] (I) Dynamic energy prediction throughout the entire flight. like Figure 5 As shown, the full-range energy closed-loop management unit adopts a deep neural network energy prediction model to dynamically predict the remaining energy consumption of each UAV throughout its entire flight at each decision step.
[0060] For any flight segment to be predicted, the input feature vector of the prediction model is defined as: ; in, This represents the energy prediction input feature vector corresponding to a single flight segment; This represents the current flight distance. This represents the change in altitude; This indicates the real-time wind speed in the airspace corresponding to the current flight segment; This indicates the wind direction angle in the airspace corresponding to the current flight segment; Ambient temperature; Indicates the current payload weight of the drone; This is the model code for the drone.
[0061] Predicted energy consumption for each flight segment Output from the deep neural network energy prediction model: ; in, This indicates the predicted energy consumption for a single flight segment; This represents a deep neural network energy prediction model; This represents the network parameters of a deep neural network energy prediction model; This represents the segment feature vector input to the deep neural network energy prediction model.
[0062] For the entire flight path from the current location through candidate anchor points to the destination, the path is divided into... If there are consecutive flight segments, the total energy consumption for the entire flight is predicted to be [number] segments. The sum of predicted energy consumption for each flight segment: ; in, This represents the predicted total energy consumption for the entire flight from the drone's current location to its destination via candidate anchor points; Indicates the segment number; This indicates the total number of segments included in the entire flight path; Indicates the first Predicted energy consumption for each flight segment.
[0063] In this way, the system can dynamically update the predicted remaining energy consumption for the entire flight at each decision step based on the current state of the UAV, environmental conditions, payload status, and changes in candidate paths, providing a basis for subsequent energy safety margin calculations and anchor point allocation decisions.
[0064] (ii) Calculation of energy safety margin, Based on the full-range energy prediction results, the full-range energy closed-loop management unit targets UAVs. Each candidate anchor point Calculate the corresponding energy requirement. The energy requirement is defined as follows: ; in, Indicates drone via candidate anchor points The predicted total energy required to complete subsequent flight missions; Indicates drone Fly from current position to candidate anchor Predicted energy consumption; Indicates drone hovering power; Indicates drone At candidate anchor points The longest estimated waiting time at the location; Indicates drone From candidate anchor points Predicted energy consumption for flying to the destination or subsequent target path nodes; Indicates drone Reserved safety redundancy power.
[0065] Correspondingly, drones Select candidate anchor points The energy safety margin is defined as follows: ; in, Indicates drone Corresponding candidate anchor points Energy safety margin; Indicates drone Current remaining battery power; (This refers to a drone) via candidate anchor points The predicted total energy required to complete subsequent flight missions.
[0066] when hour The range of values is .when At that time, it indicates that the drone The current remaining power is sufficient to cover the candidate anchor points. Forecast total energy demand; when At that time, it indicates that the drone The current remaining power is insufficient to complete the journey via the candidate anchor point. For subsequent flight missions, this candidate anchor point From drones Remove from the set of candidate anchor points.
[0067] Energy safety margin This serves as a constraint on the anchor point allocation decision and also as the basis for calculating the energy reward component in the reward function. Let's assume that in the current allocation scheme, the drone... Assigned to anchor point The energy reward component is then defined as: ; in, Indicates drone The energy reward amount obtained under the current anchor point allocation scheme; Indicates the number of drones in the current allocation scheme. The actual anchor point number assigned; Indicates drone Corresponding to its assigned anchor point Energy safety margin; This indicates the threshold for determining sufficient energy. This represents the energy stress threshold, and satisfies... ; , , All are positive reward coefficients.
[0068] The energy reward function described above indicates that when the drone The energy safety margin of the assigned anchor points is higher than the energy sufficiency threshold. When the energy safety margin falls below the energy shortage threshold, the system provides a positive energy reward; when the energy safety margin falls below the energy shortage threshold, the system provides a positive energy reward. When the energy is at a certain level, the system imposes an energy stress penalty; when the energy safety margin is within a certain range... and During this period, the system provides continuous rewards based on the size of the energy safety margin, causing the anchor point allocation strategy to tend to select anchor points with higher energy safety margins.
[0069] (III) Dynamic emergency threshold and energy priority interruption, Dynamic emergency threshold According to drones The distance to the nearest alternate landing site, the predicted energy consumption for flying to the nearest alternate landing site, the safety redundancy power, and the energy consumption for current hovering are calculated in real time. It is defined as follows: ; in, Indicates drone Dynamic emergency energy threshold; Indicates drone The corresponding safety factor, and satisfying ; Indicates drone Predicted energy consumption for flying from the current location to the nearest alternate airport; Indicates drone The hovering energy consumption corresponding to the current hovering wait time; Indicates drone Safety redundancy power.
[0070] Among them, the predicted energy consumption for flying to the nearest alternate airport. According to drones The energy consumption per unit of flight segment is calculated from the spatial distance from the current location to the nearest alternate landing site and the expected energy consumption per unit of flight segment, or it can be predicted by the aforementioned deep neural network energy prediction model based on characteristics such as flight segment distance, altitude change, wind speed, wind direction, ambient temperature, payload weight, and UAV type code. Current hovering waiting energy consumption. Defined as: ; in, Indicates drone hovering power; Indicates drone The current cumulative hovering wait time.
[0071] Safety factor The system is dynamically adjusted based on meteorological conditions and airspace complexity. This adjustment is particularly relevant in scenarios with adverse weather conditions, high airspace traffic density, or complex airspace structure. Take the larger value to trigger the energy protection mechanism in advance; in scenarios with stable weather conditions, low airspace traffic density, or sufficient alternate landing resources, Choose a smaller value to reduce unnecessary interception operations.
[0072] When drones Predicted remaining energy Below At this time, the energy priority preemption mechanism is triggered. After triggering, the system will... The priority of anchor point allocation is raised to the highest level, allowing it to preempt anchor point slots that have been reserved by other drones but not yet occupied. The drones whose anchor point slots have been preempted enter the reallocation process.
[0073] For the reassigned drones Its new set of candidate anchor points is defined as: ; in, Indicates drone The set of candidate anchor points in the reallocation process; Indicates the candidate anchor point number; Indicates drone Select candidate anchor points Energy safety margin at that time; Indicates drone The current remaining power is sufficient to support its passage through candidate anchor points Complete subsequent flight missions; Indicates drone The current position coordinates; Indicates candidate anchor points Position coordinates; Indicates drone With candidate anchors Spatial distance between them; Indicates drone The maximum achievable distance under the current remaining battery power constraint; Indicates candidate anchor points The available status indicator, when Time indicates candidate anchor point It is in a state of being available for booking or receiving, when Time indicates candidate anchor point Not available.
[0074] Through the aforementioned dynamic emergency threshold and energy priority interruption mechanism, the system can trigger protective actions in advance before the drone enters an energy danger state, transforming energy constraints from post-event threshold alarms to dynamic perception and proactive intervention throughout the entire process, thereby reducing the risk of the drone running out of energy due to waiting, detouring, or path changes.
[0075] (iv) Level 3 Emergency Response Agreement The triggering condition for a Level 3 emergency response protocol is determined by the continuous change of the energy safety margin. Assuming that in the current anchor point allocation scheme, the drone... The actual anchor points assigned are Then drone The current corresponding energy safety margin is .in, Indicates the number of drones in the current allocation scheme. The actual anchor point number assigned.
[0076] The Level 3 emergency response protocol triggers different levels of response based on different energy safety margin ranges, as detailed below: First-level response triggering conditions: When the above conditions are met, it indicates that the drone is The energy state entered a mild emergency zone, and the system executed priority escalation and anchor point preemption operations, deploying the drone. The priority of anchor point allocation is increased, and it is allowed to preempt anchor point time slots that have been reserved but not yet actually occupied; Second-level response triggering conditions: When the above conditions are met, it indicates that the drone is The energy state has entered a moderate emergency zone. If the drone... If available charging anchor points are available nearby, the system guides the drone. The drone flies to an available charging anchor point and replans its subsequent flight path; if no available charging anchor point is found nearby, the system immediately grants the highest priority clearance, allowing the drone to proceed. Prioritize leaving the waiting state and entering a safer flight path with lower energy consumption; Level 3 response triggering conditions: When the above conditions are met, it indicates that the drone is The remaining energy is insufficient to safely complete the subsequent flight mission via the currently assigned anchor point. The system immediately triggers the emergency landing procedure and calculates the nearest safe landing point using a dynamic energy map. Emergency landing target. The calculation method is as follows: ; in, Represented as drone Selected emergency landing target point; This represents the set of available safe landing points within the current airspace; Indicates drone The current location; Represents the set of available safe landing points Any candidate landing point in the list; Indicates drone From current location Fly to candidate landing point The system predicts energy consumption from a set of available safe landing points. The landing point with the lowest predicted energy consumption was selected as the emergency landing target.
[0077] in, This indicates the threshold for triggering the first-level response. This represents the threshold for triggering the second-level response, and satisfies: ; The switching between different levels of response is determined by the energy safety margin. Driven by real-time calculation results, the system progressively enhances rescue measures in the order of Level 1, Level 2, and Level 3 response as the energy safety margin continues to decline. When the energy status recovers, the system can exit the corresponding level of emergency rescue state based on the real-time energy safety margin and re-enter the regular anchor point allocation process.
[0078] Example 4 Virtual anchor chain cascading allocation and symmetric release mechanism This embodiment details the construction method of the virtual anchor chain and its cascading allocation and symmetric release mechanism.
[0079] (a) Construction of the virtual anchor chain, like Figure 6 As shown, the virtual anchor chain management unit first encodes the high-dimensional features of each obstacle avoidance anchor in the structured low-altitude corridor network. Anchor points The eigenvectors are defined as follows: ; in Indicates anchor point High-dimensional feature vectors, Indicates anchor point spatial coordinate vector, Indicates anchor point capacity, Indicates anchor point The direction vector of arrival, Indicates anchor point Historical usage frequency Indicates anchor point The surrounding traffic flow density.
[0080] To reduce feature dimensionality and extract latent functional relationships between anchor points, an autoencoder is used to process high-dimensional feature vectors. Compressed encoding is performed to obtain a low-dimensional latent feature representation. An autoencoder consists of an encoder and a decoder. The encoder is used to process high-dimensional feature vectors. Mapping to low-dimensional latent feature representation The decoder is used to represent low-dimensional latent features. Reconstruct the original feature vector. The training objective is to minimize the reconstruction error. ; in, This represents the reconstruction loss of the autoencoder; Indicates the decoder's connection to the anchor point. The eigenvector reconstruction output; This represents the L2 norm.
[0081] Any two anchor points and The functional similarity between them is measured using cosine similarity: ; in, Indicates anchor point With anchor point Functional similarity between them; and These represent anchor points. With anchor point The low-dimensional latent feature representation.
[0082] After obtaining the functional similarity between anchor points, the virtual anchor point chain management unit comprehensively considers both functional similarity and spatial proximity to cluster the anchor points. Specifically, anchor points that meet the criteria of functional similarity higher than a preset similarity threshold and spatial distance less than a preset proximity threshold are grouped into the same cluster: ; in, Indicates the first Each anchor point cluster; Indicates the functional similarity threshold; Indicates the spatial proximity threshold; Indicates anchor point With anchor point The spatial distance between them.
[0083] Within each anchor cluster, the primary anchor is determined based on a weighted composite score of capacity, accessibility, and historical usage frequency. Anchor The main anchor point score is defined as: ; in, Indicates anchor point Main anchor score; Indicates anchor point The capacity; This represents the maximum capacity of anchor points within the cluster. Indicates anchor point The reachability score (availability status indicator) has a range of values. ; Indicates anchor point Historical usage frequency; This represents the maximum historical frequency of use of anchor points within the cluster. , , Let represent the weighting coefficients corresponding to capacity, accessibility, and historical usage frequency, respectively, and satisfy . Within each cluster, the anchor point with the highest weighted composite score is determined as the primary anchor point, and the remaining anchor points are determined as secondary anchor points.
[0084] After determining the primary and secondary anchor points, a directed topology is established, pointing from the primary anchor point to the secondary anchor points. For the primary anchor point... With secondary anchor points The cascade reachability weight of its directed edges is defined as: ; in, Indicates from the main anchor point To secondary anchor point Cascade reachable weights; Indicates the main anchor point With secondary anchor points Spatial distance between them; To prevent extremely small positive numbers with a denominator of zero; Indicates secondary anchor point The current occupancy rate, with a value range of 100%. .
[0085] From the above formula, it can be seen that when the secondary anchor point Distance from main anchor point The closer and the lower the current occupancy rate The larger the value, the higher the selection priority of the secondary anchor point in the cascading allocation after the primary anchor point's capacity is saturated. The directed topology of the virtual anchor point chain is dynamically updated based on the anchor point's occupancy status, reachability status, and traffic flow density.
[0086] (ii) Cascading allocation mechanism, When the main anchor point The occupancy rate exceeds the saturation threshold At this time, the virtual anchor chain management unit triggers the cascading allocation mechanism. For newly arrived UAV allocation requests, the system first checks the main anchors in descending order of edge weights in the directed topology of the virtual anchor chain. The corresponding set of secondary anchor points The availability of each secondary anchor point is determined, and a set of available secondary anchor points that satisfy the occupancy constraint, availability state constraint, and energy safety margin constraint is selected: ; in, Indicates the main anchor point The corresponding set of available secondary anchor points; Indicates the main anchor point The set of secondary anchors in the virtual anchor chain; Indicates secondary anchor point Current occupancy rate; Indicates the anchor point saturation threshold; Indicates secondary anchor point The availability status; Indicates drone The energy safety margin corresponding to flight through candidate secondary anchor points.
[0087] The fitness score of cascaded assignment is determined by a multi-agent graph neural network collaborative decision-making mechanism. calculate: ; in For drones State feature vector, secondary anchor point eigenvectors (low-dimensional latent feature representations). The similarity score is cosine similarity. This fit score comprehensively reflects the degree of feature matching between the UAV and the secondary anchor point, the energy safety margin, and the remaining available capacity of the anchor point.
[0088] The system starts from the set of available secondary anchor points. The secondary anchor point with the highest fit score is selected as the cascaded assignment target. : ; If a set of secondary anchor points is available If the result is empty, the system continues to expand the search outward along the next level of the virtual anchor chain topology; if no anchor that meets the conditions is found within the preset maximum cascading depth, the reallocation process or the backup corridor avoidance process is triggered.
[0089] (iii) Symmetrical release mechanism, After the obstacle is cleared, the release control unit releases the backlog of drones sequentially in a symmetrical order from secondary anchor points to primary anchor points. To avoid a large influx of drones into the main corridor entrance simultaneously after the obstacle is cleared, the system calculates a release priority score for each anchor point based on the spatial distance from each anchor point to the main corridor entrance, the average energy safety margin of drones within each anchor point, and the average task priority. The release priority score for each anchor point is defined as follows: ; in, Indicates anchor point Release priority score; Indicates anchor point To the main corridor entrance Normalized spatial distance; Indicates anchor point The average energy safety margin of all drones in the system; Indicates anchor point The average priority of internal drone missions; To prevent extremely small positive numbers with a denominator of zero; , , These are the corresponding weight coefficients. And they satisfy: ; As shown in the formula above, the farther the anchor point is from the main corridor entrance, the lower the energy safety margin of the drone within the anchor point, and the higher the task priority, the higher the release priority score of that anchor point. The resumption of the release control unit follows... In descending order, the drones within each anchor point are controlled to sequentially exit the obstacle avoidance anchor point and reconnect to the main corridor or backup corridor.
[0090] To avoid creating new instantaneous flow peaks during the reopening phase, the instantaneous flow constraint at the main corridor entrance is defined as follows: ; in for A set of anchor points that are always in a released state. anchor point exist The drone release rate at any given moment The maximum capacity flow at the main corridor entrance is determined by the instantaneous flow constraint. The resumption control unit dynamically adjusts the release rate of each anchor point based on the instantaneous flow constraint to ensure that the total flow at the main corridor entrance never exceeds the capacity limit during the resumption process. This prevents drones from flooding into the main corridor entrance after the obstacle is removed, reducing the risk of secondary congestion.
[0091] Example 5 Multi-source information fusion networks with multi-head attention mechanisms Anchor point allocation employs a multi-source information fusion network with a multi-head attention mechanism to achieve end-to-end learning-based anchor point allocation.
[0092] This multi-source information fusion network takes UAV observation feature sequences, candidate anchor point state features, and task constraint features as inputs. It extracts multi-dimensional dynamic state features of the UAV through a CNN-LSTM fusion network, calculates the matching relationship between the UAV and candidate anchor points using a multi-head attention mechanism, and finally outputs the selection probability distribution of candidate anchor points through a pointer network.
[0093] For drones Observational feature sequence The CNN branch extracts local coupling features between multidimensional observation metrics through one-dimensional convolution operations: ; in, For one-dimensional convolution kernel parameters, For convolution bias terms, This represents the convolution operation. This represents the local coupling features of the CNN branch output.
[0094] The LSTM branch captures the dynamic trends of various indicators over time through a bidirectional LSTM layer, and the forward and backward hidden states are concatenated to obtain: ; in, Indicates the forward LSTM at time... The hidden state, Indicates the backward LSTM at time... The hidden state, This represents the temporal dynamic characteristics of the UAV observation sequence.
[0095] After concatenating the CNN output and the LSTM output, the result is mapped to a drone feature vector of uniform dimension through a fully connected layer. ; in, Indicates drone The fused feature vector, and These are the weight matrix and bias term of the fully connected layer, respectively.
[0096] for Candidate anchor set Set candidate anchor points The feature vector is The multi-head attention layer will integrate the drone feature vectors. As the query vector, the candidate anchor feature vector As key and value vectors, the attention weights between the UAV and each candidate anchor point are calculated. Attention weights corresponding to each attention head for: ; in, and The first The query mapping matrix and key mapping matrix in the attention head The dimension of the key vector is used to scale the dot product to prevent gradient vanishing. Operations on the set of candidate anchor points Normalization is performed on the above.
[0097] No. The output of each attention head for: ; in, For the first The value mapping matrix in the attention head.
[0098] Will After concatenating the outputs of each attention head, a multi-head attention fusion feature is obtained through linear transformation: ; in, Output the mapping matrix for multi-head attention. Indicates drone The fusion matching features between the candidate anchor set and the target anchor set.
[0099] The pointer network output layer is based on the fused matching features and candidate anchor features Calculation of drones Select candidate anchor points Matching score : ; in, , , For learnable parameters, Indicates drone With candidate anchors Match score.
[0100] Furthermore, the selection probability distribution of candidate anchor points is obtained through the Softmax function. : ; in, To find the dummy element, we represent traversing the set of candidate anchor points. Each anchor point in the middle, Indicates drone With candidate anchors Match score.
[0101] Finally, a greedy strategy is used to select the anchor point with the highest probability as the allocation result: ; in, Indicates drone The optimal anchor point allocation result.
[0102] Through the multi-source information fusion network with the aforementioned multi-head attention mechanism, the system can adaptively learn the nonlinear coupling relationship between UAV status, candidate anchor point status, task priority, arrival timeliness, remaining battery urgency, flight direction matching degree, and anchor point resource availability. This avoids the problem of traditional linear weighting methods relying on manually preset fixed weights, thereby improving the accuracy and adaptability of anchor point allocation results.
[0103] Example 6 The overall workflow of the system like Figure 7 As shown, the overall workflow of the method of the present invention is as follows: Step S1: The corridor network construction module constructs a structured low-altitude corridor network based on the environmental information of the target airspace, no-fly zone information, obstacle distribution information, and mission route information. In the structured low-altitude corridor network, the system presets obstacle avoidance anchor points at key nodes, branch nodes, merging nodes, and obstacle avoidance buffer areas, and encodes each obstacle avoidance anchor point using spatiotemporal grid units as the basic management granularity, and initializes the state information of each spatiotemporal grid unit.
[0104] Step S2: When multiple drones detect a sudden obstacle, each drone determines its own set of candidate obstacle avoidance anchor points. The selection criteria include: spatial reachability. The system checks the following parameters: orientation matching (deflection angle not exceeding 90°), and time availability (anchor points have available capacity within their expected arrival time windows). Each candidate anchor point and its corresponding expected arrival time are mapped to a four-dimensional spatiotemporal grid, and the reservation queue for the corresponding spatiotemporal grid cell is updated.
[0105] Step S3: The graph neural network conflict prediction unit performs state prediction on the four-dimensional spatiotemporal grid and outputs the conflict probability of each spatiotemporal grid unit within a future time window. ;Will Real-time feedback to the reward function of the multi-agent deep reinforcement learning framework enables conflict detection and allocation decisions to form a closed loop.
[0106] Step S4: The multi-agent deep reinforcement learning allocation unit, combined with the multi-source information fusion network of the multi-head attention mechanism, comprehensively considers multi-dimensional constraints such as conflict probability, energy safety margin, and task priority, and collaboratively generates anchor point allocation schemes for each UAV, establishing anchor point occupancy reservations for each UAV.
[0107] Step S5: The full-range energy closed-loop management unit continuously monitors the energy safety margin of each UAV. When drones of The energy priority preemption mechanism is triggered, and the anchor point allocation priority is adjusted accordingly; The continuous changes trigger the corresponding Level 3 emergency response protocol, with the specific triggering conditions as follows: when The first-level response is triggered when priority is increased and anchor point preemption is executed; when When a second-level response is triggered, the drone is guided to the charging anchor point or given the highest priority clearance; when The third-level response is triggered, guiding the drone to the nearest safe landing point along the path of least energy consumption. .
[0108] Step S6: After the obstacle is removed, the release control unit generates a release order through the deep reinforcement learning policy network of the multi-agent deep reinforcement learning allocation unit. The release priority score of each anchor point is calculated according to the following formula: ; according to Release the drones at each anchor point in descending order of height, while ensuring that the instantaneous flow rate at the main corridor entrance is always met: ; When the release of a certain anchor point causes the inlet flow to approach... At that time, the release rate of the anchor point will be automatically reduced. The normal release rate will be restored only after the inlet flow drops to a safe range, thus maintaining a smooth transition of the main corridor inlet flow throughout the recovery process and avoiding secondary congestion.
[0109] Example 7 Training methods for multi-agent deep reinforcement learning frameworks This embodiment details the training method of the multi-agent deep reinforcement learning framework, providing complete training support for the implementation of the closed-loop mechanism in Embodiment 2.
[0110] (a) Environmental modeling A multi-UAV obstacle avoidance simulation environment is constructed, and the state space consists of two parts: local observation state and global state.
[0111] drones The local observation state is defined as: ; in For drones The position vector, For drones The velocity vector, For drones The remaining battery power, drones The priority of the task is _____. For drones The set of candidate anchor points; for For the candidate anchor set Candidate anchor points The set of four-dimensional spatiotemporal grid units corresponding to the reservation process; This is the set of conflict probabilities corresponding to the set of four-dimensional spatiotemporal grid cells.
[0112] Global state Used only by the centralized value network during the training phase, defined as: ; in, Indicates the number of drones participating in the scheduling; For airspace traffic flow, This represents the average occupancy rate of anchor points. The probability of collision in the current spatiotemporal grid exceeds a threshold. The grid cell density.
[0113] Action space Discrete action space: ; in to These represent the sets of candidate anchor points. Candidate anchor points corresponding to the sorting positions in the middle. This indicates that the vehicle is hovering in place and waiting. This indicates a return to the upstream anchor point. Candidate anchor point set. The ranking of each candidate anchor point is determined by the anchor point reachability, anchor point availability, conflict probability, and energy safety margin.
[0114] The state transition updates the environmental state based on the UAV dynamics model and anchor point occupancy rules. (UAV) In performing the action The position after is updated to: ; in, Indicates drone At any moment The position vector; Indicates drone At any moment The velocity vector; Indicates drone At any moment The acceleration vector; This represents the time interval between two adjacent decision steps.
[0115] Corresponding energy consumption update for: ; in, Indicates drone At any moment The remaining battery power; The flight power of the drone is determined by the drone's speed. Load With ambient temperature A joint decision.
[0116] When drones Select candidate anchor points Subsequently, the four-dimensional spatiotemporal grid management unit, based on the UAV... The current location, flight speed, candidate anchor point locations, and estimated arrival time are used to map the reservation request to the corresponding four-dimensional spatiotemporal grid cell. The graph neural network conflict prediction unit further outputs the conflict probability of this four-dimensional spatiotemporal grid unit. This information is then fed back into the reward function of the multi-agent deep reinforcement learning allocation unit, enabling the policy network to learn allocation strategies that avoid high-conflict-probability anchors during training.
[0117] (II) MAPPO Algorithm Training Process Initialize policy network With value network Iterative training is performed according to the following process: Initialize the shared policy network for all drones With centralized value networks ,in, Indicates the policy network parameters, Represents the parameters of the value network. Policy network. Used to output anchor point assignment actions based on the local observation status of the UAV, centralized value network Used to estimate state value based on global state during the training phase.
[0118] During the training phase, a centralized training and distributed execution paradigm is adopted. Each UAV interacts with the simulation environment and collects trajectory data to form an experience dataset. ; in, Represents an empirical dataset; Indicates the number of drones participating in the training; Indicates the time length of a single-round trajectory sampling; Indicates time step The global state; Indicates drone At time step The local observation status; Indicates drone At time step The anchor point assignment action performed; Indicates drone At time step The rewards received; Indicates the global state at the next time step; Indicates drone The local observation status at the next time step.
[0119] The first step is to gather experience. Each drone, based on the current strategy network, [does something]. and its own local observation state Independent action selection: ; in, The policy network represents the situation of a given drone. Local observation status The output action probability distribution. After an action is executed, the environment updates the system state based on the anchor point allocation result, the four-dimensional spatiotemporal grid conflict probability, the energy safety margin, and the anchor point occupancy status, and returns the reward. .
[0120] The second step is to calculate the advantage function. Generalized advantage estimation (GAE) is used to calculate the advantage function for the UAV. At time step Advantage function: ; in, Indicates drone At time step The estimated value of the dominance function; Indicates the discount factor; This represents the attenuation coefficient of the generalized advantage estimation, i.e., the GAE smoothing parameter; Indicates time step The temporal difference residual.
[0121] The time-series differential residual is defined as: ; in, This indicates the global state of a centralized value network. The state value estimate of the output.
[0122] The third step is to update the policy network. The policy network is updated using the PPO (Pruning Objective Function) method. For drones At time step The probability ratio of an action is defined as: ; in, This indicates the difference between the current policy network and the old policy network in terms of actions. The probability ratio above; Indicates the current policy network; This refers to the old policy network used when sampling empirical data.
[0123] PPO pruning objective function of policy network Defined as: ; This is used to limit the magnitude of policy updates, preventing the policy network from changing too much in a single update. By maximizing... Update strategy network parameters .
[0124] The fourth step is to update the value network. Centralized value network. Update by minimizing the value loss function, the value loss function Defined as: ; in, This represents the target value. The target value is obtained by combining the advantage function estimate and the current state value estimate: ; in, The advantage functions of each UAV can be used We get the following by averaging or weighted summation: ; Alternatively, the advantage functions of each UAV can be weighted and combined based on mission priority, energy urgency, and conflict risk.
[0125] The fifth step is repeated training. The processes of experience collection, advantage function calculation, policy network update, and value network update are repeated until the policy network converges. The convergence criterion is that the change in total reward is less than a preset threshold over several consecutive training rounds, or key performance indicators such as conflict rate, average waiting time, and energy safety margin default rate tend to stabilize.
[0126] After training, retain the policy network. Used for the execution phase. During the execution phase, each UAV relies solely on its own local observation status. Anchor point assignment actions are generated independently without real-time access to the global state. It also eliminates the need for real-time centralized coordination; a centralized value network Use only during the training phase.
[0127] (III) Training methods for CNN-LSTM fusion networks The CNN-LSTM fusion network is trained using a combination of supervised learning and reinforcement learning. In the supervised learning phase, pre-training is performed using historical anchor point-assigned data, and the loss function... Designed as follows: ; in, Represents classification loss, Indicates regression loss, Indicates the sorting loss; and These are the loss weighting coefficients, used to adjust the contribution of different loss terms to the total loss function.
[0128] Classification The cross-entropy loss function is used to measure the accuracy of the allocation success rate prediction: ; in, Indicates the first The true label of the class, The network prediction yields the first... Class probability.
[0129] Regression loss The mean squared error loss function is used to measure the accuracy of waiting time and energy consumption predictions: ; in, Indicates the number of training samples; Indicates the first Prediction waiting time for each sample; Indicates the first The actual waiting time for each sample; Indicates the first Predicted energy consumption for each sample; Indicates the first The actual energy consumption of each sample.
[0130] Ranking loss The ListNet loss function is used to measure the accuracy of anchor sorting: ; in and The scores calculated based on the actual scores and predicted scores are respectively the first... The probability distribution of anchor point selection is obtained by normalizing using Softmax: ; ; After supervised pre-training, the CNN-LSTM fusion network is embedded into the MAPPO framework and further fine-tuned through reinforcement learning, so that the network can continuously optimize the allocation strategy in real multi-machine interaction scenarios.
[0131] By introducing a four-dimensional spatiotemporal grid coding mechanism into the structured low-altitude corridor network, the management granularity of anchor point resources is elevated from static spatial points to dynamic spatiotemporal units, fundamentally eliminating the problem of missed temporal conflict detection caused by the lack of a time dimension in existing technologies. By directly embedding the conflict probability prediction results of graph neural networks into the reward function of a multi-agent deep reinforcement learning framework, a real-time closed loop is formed between conflict detection and allocation decisions, realizing a mechanism shift from reactive processing to predictive avoidance. By constructing a full-fledged closed-loop management mechanism centered on energy safety margins, energy constraints are transformed from post-flight threshold triggering to full-process dynamic perception, and graded responses to energy risks are achieved through dynamic emergency thresholds and a three-level emergency rescue protocol. By introducing a virtual anchor point chain cascade allocation and symmetrical release mechanism, the risk of secondary congestion during the resumption of passage is structurally eliminated. The above four mechanisms work together to resolve systemic contradictions in existing technologies, such as coarse granularity of spatial domain resource management during multi-aircraft emergency obstacle avoidance, fragmented conflict detection and allocation decisions, passive response to energy constraints, and disordered resumption of passage. They can significantly improve the emergency carrying capacity, decision-making intelligence level, and operational stability of low-altitude corridors in multi-aircraft scenarios.
[0132] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A spatiotemporal grid closed-loop multi-machine obstacle avoidance anchor point allocation system, characterized in that, This includes a corridor network construction module and an emergency obstacle avoidance and collaborative scheduling module; The corridor network construction module is used to construct a structured low-altitude corridor network and determine the location of obstacle avoidance anchor points. The emergency obstacle avoidance and collaborative scheduling module includes: The four-dimensional spatiotemporal grid management unit is used to perform three-dimensional segmentation of the target airspace based on the Beidou grid code and introduce the time dimension to form a four-dimensional spatiotemporal grid code and a four-dimensional spatiotemporal grid unit, and to predict spatiotemporal conflicts in UAV anchor point reservation requests. The graph neural network conflict prediction unit is used to construct a spatiotemporal graph structure with the four-dimensional spatiotemporal grid units as nodes and the temporal correlation of the UAV flight path as edges, and input the spatiotemporal graph structure into the spatiotemporal graph convolutional network to output the conflict probability of each four-dimensional spatiotemporal grid unit in the future time window. A multi-agent deep reinforcement learning allocation unit is used to dynamically adjust the anchor point allocation decision by taking the conflict probability as the input of the reward function. The full-range energy closed-loop management unit is used to predict remaining energy consumption, calculate the energy safety margin of candidate anchor points, and trigger energy priority preemption. The virtual anchor chain management unit is used to construct primary and secondary anchor chain, and to cascade and allocate anchors when the primary anchor is saturated. The resumption release control unit is used to control the release based on the release priority score and symmetrical release sequence, and to adjust the release rate so that the instantaneous flow at the main corridor entrance does not exceed the capacity limit.
2. The spatiotemporal grid closed-loop multi-machine obstacle avoidance anchor point allocation system according to claim 1, characterized in that, The full-range energy closed-loop management unit is equipped with an energy priority interruption mechanism; When the energy safety margin of the candidate anchor point corresponding to the drone is lower than the preset margin threshold, or the predicted remaining energy of the drone is lower than the dynamic emergency threshold, the energy priority interruption mechanism is triggered to raise the allocation priority of the drone to the highest level and allow it to preempt the anchor point time slot that has been reserved but not yet occupied. The dynamic emergency threshold is calculated in real time based on the predicted flight energy consumption from the UAV's current location to the nearest alternate landing site, the current hovering waiting energy consumption, and the safety redundancy power.
3. The spatiotemporal grid closed-loop multi-machine obstacle avoidance anchor point allocation system according to claim 1, characterized in that, The virtual anchor chain management unit dynamically determines the cascading allocation order based on the anchor occupancy status, anchor capacity, and the reachability cost between anchors. When the capacity of the main anchor point reaches the preset saturation threshold, the UAV is cascaded and allocated from the main anchor point to the secondary anchor point according to the cascade allocation sequence. After the obstacle is removed, the multi-agent deep reinforcement learning allocation unit generates a symmetric release order by combining the conflict probability, the UAV energy safety margin, and the task priority under the reverse constraint of the cascade allocation order. The resumption release control unit calculates a release priority score based on the spatial distance from each anchor point to the main corridor entrance, the average energy safety margin of the UAVs within the anchor point, and the average task priority. It then adjusts the exit order of UAVs within the same release batch according to the release priority score and adjusts the release rate of each anchor point according to the instantaneous flow constraint at the main corridor entrance to ensure that the instantaneous flow at the main corridor entrance does not exceed the carrying capacity limit.
4. The spatiotemporal grid closed-loop multi-machine obstacle avoidance anchor point allocation system according to claim 1, characterized in that, The graph neural network conflict prediction unit predicts the conflict probability in the following way: The candidate anchor points and their expected arrival times of each UAV are mapped together to the corresponding four-dimensional spatiotemporal grid cells. A spatiotemporal graph structure is constructed with each spatiotemporal grid cell as a node and the temporal correlation of the UAV flight path as an edge. The spatiotemporal graph structure is propagated and predicted by a spatiotemporal graph convolutional network, and the conflict probability of each spatiotemporal grid cell in the future time window is output. When the conflict probability exceeds a preset threshold, it is determined that the corresponding four-dimensional spatiotemporal grid unit has a conflict risk, and the multi-agent deep reinforcement learning allocation unit is triggered to dynamically adjust the anchor point allocation decision according to the conflict probability. The conflict probability prediction is completed before the UAV reaches the anchor point to achieve predictive conflict avoidance.
5. The spatiotemporal grid closed-loop multi-machine obstacle avoidance anchor point allocation system according to claim 1, characterized in that, The multi-agent deep reinforcement learning allocation unit adopts a centralized training and distributed execution paradigm; During the training phase, a centralized value network is trained using global state information that includes all UAV states, anchor point states, conflict probabilities of four-dimensional spatiotemporal grid cells, and energy safety margins. The policy network parameters are then updated based on the advantage function estimate output by the centralized value network. During the execution phase, each UAV independently generates anchor point allocation decisions based solely on its local observation information, without the need for real-time centralized coordination. The reward function of the multi-agent deep reinforcement learning allocation unit includes four components: safety reward, efficiency reward, energy reward, and fairness reward. The weights of the four components are dynamically configured according to the task scenario type to adapt to the differentiated needs of optimization objectives in urban logistics, medical delivery, and emergency rescue scenarios.
6. The spatiotemporal grid closed-loop multi-machine obstacle avoidance anchor point allocation system according to claim 1, characterized in that, The multi-agent deep reinforcement learning allocation unit in the emergency obstacle avoidance collaborative scheduling module adopts a multi-source information fusion network with a multi-head attention mechanism. The multi-source information fusion network uses the UAV feature vector as a query and the anchor point feature vector as a key and value. It calculates the matching score between the UAV and each candidate anchor point through a multi-head attention mechanism. The matching score comprehensively reflects the coupling relationship between task priority, arrival timeliness, remaining battery urgency, flight direction matching degree and anchor point resource availability. The optimal anchor point allocation scheme is output through the pointer network based on the matching score.
7. A spatiotemporal grid closed-loop multi-machine obstacle avoidance anchor point allocation system according to claim 2, characterized in that, The full-range energy closed-loop management unit also includes a three-level emergency rescue protocol; When the drone's energy status enters the emergency zone, a corresponding level of rescue response is triggered based on the energy safety margin range corresponding to the ratio of remaining energy to the distance to the nearest alternate landing site. The first-level response corresponds to an energy safety margin that is in the upper part of the preset emergency zone, and the execution priority is increased and the anchor point is preempted. The Level 2 response corresponds to an energy safety margin in the middle of the preset emergency zone, and executes guidance to the nearest alternate airport and flight replanning. The Level 3 response corresponds to an energy safety margin in the lower part of the preset emergency zone, and involves designating an emergency landing area and clearing the surrounding airspace. The trigger threshold ranges of each level of response correspond one-to-one with the numerical range of the energy safety margin, and the switching between each level of response is driven by the real-time calculation results of the energy safety margin.
8. A spatiotemporal grid closed-loop multi-machine obstacle avoidance anchor point allocation system according to claim 3, characterized in that, The method for constructing the virtual anchor chain in the virtual anchor chain management unit includes: The location, capacity, direction of arrival, and historical usage characteristics of each anchor point are encoded, and the functional similarity between anchor points is calculated. Anchor points are clustered based on functional similarity and spatial proximity. Within each cluster, the main anchor point is determined based on a weighted comprehensive score of capacity, accessibility, and historical usage frequency. Establish a directed topology from the primary anchor point to the secondary anchor point, wherein the edge weights of the directed topology reflect the spatial reachability cost between the anchor points; The directed topology of the virtual anchor chain is dynamically updated as the anchor occupancy status changes.
9. A spatiotemporal grid closed-loop multi-machine obstacle avoidance anchor point allocation system according to claim 1, characterized in that, The corridor network construction module includes a risk map construction unit and a modular corridor unit library; The risk map construction unit integrates geographic information, building information, population distribution information, noise-sensitive area information, meteorological information, and airspace occupancy information of the target airspace to construct a three-dimensional risk map; The modular corridor unit library predefines standardized three-dimensional corridor units that match the drone's cruise, turning, takeoff, and node switching actions. Based on the three-dimensional risk map, standardized three-dimensional corridor units are selected and combined from the modular corridor unit library to generate a structured low-altitude corridor network; The location of obstacle avoidance anchor points is determined by the risk distribution in the three-dimensional risk map and the directed topology of the structured low-altitude corridor network.
10. The method for a spatiotemporal grid closed-loop multi-machine obstacle avoidance anchor point allocation system according to claim 1, characterized in that, Includes the following steps: Step S1: By integrating the geographic information, building information, population distribution information, noise-sensitive area information, meteorological information and airspace occupancy information of the target airspace through the corridor network construction module, a three-dimensional risk map is constructed. Based on the three-dimensional risk map, standardized three-dimensional corridor units are selected and combined from the modular corridor unit library to generate a structured low-altitude corridor network and determine the location of obstacle avoidance anchor points. Step S2: The target airspace is divided into three dimensions and a time dimension is introduced by the four-dimensional spatiotemporal grid management unit based on the Beidou grid code to form a four-dimensional spatiotemporal grid code; the anchor point reservation requests of each UAV are mapped to the corresponding four-dimensional spatiotemporal grid unit, and reservation conflicts are predicted in the spatiotemporal dimension. Step S3: Construct a spatiotemporal graph structure using each four-dimensional spatiotemporal grid unit as a node and the temporal correlation of the UAV flight path as an edge through the graph neural network conflict prediction unit. Output the conflict probability of each four-dimensional spatiotemporal grid unit in the future time window through the spatiotemporal graph convolutional network. When the conflict probability exceeds the preset threshold, it is determined that a conflict has occurred, and the collaborative allocation process is triggered. Step S4: The conflict probability is used as the input component of the reward function by the multi-agent deep reinforcement learning allocation unit. The weights of the four components, namely safety reward, efficiency reward, energy reward and fairness reward, are combined to drive the dynamic adjustment of the anchor point allocation decision in real time. Each UAV independently generates an allocation decision based on its local observation information. The matching score between the UAV and each candidate anchor point is calculated by the multi-source information fusion network with multi-head attention mechanism. The optimal anchor point allocation scheme is output through the pointer network. Step S5: Dynamically predict the remaining energy consumption of each UAV at each decision step using the full-range energy closed-loop management unit, and calculate the energy safety margin corresponding to each candidate anchor point; When the energy safety margin of a certain drone is lower than the dynamic emergency threshold calculated in real time based on its distance to the nearest alternate landing site and expected energy consumption, the energy priority interruption mechanism is triggered, which raises the allocation priority of the drone to the highest level and allows it to preempt the anchor point time slot that has been reserved but not yet occupied. Step S6: After the obstacle is removed, the recovery release order is generated by the multi-agent deep reinforcement learning allocation unit. The release priority score is calculated based on the spatial distance from each anchor point to the main corridor entrance, the average energy safety margin of the UAVs within the anchor point, and the average task priority. The UAVs are controlled to exit from the secondary anchor point to the main anchor point in a symmetrical release order from high to low release priority score. At the same time, the release rate of each anchor point is adjusted to ensure that the instantaneous flow at the main corridor entrance does not exceed the capacity limit. The recovery path for each UAV to access the main corridor or backup corridor is replanned.
Citation Information
Patent Citations
Unmanned aerial vehicle autonomous flight method for constructing three-dimensional real scene based on LiDAR data
CN110888453A
System and method for battery capacity management in UAV from
CN115867459A
Three-dimensional scheduling method and system suitable for low-altitude multi-aircraft hybrid take-off and landing field
CN120656340B
Air-ground unmanned cluster conflict resolution method based on multi-agent reinforcement learning
CN121325577A
Distribution line first-aid repair operation risk assessment method, system and equipment based on multi-source sensing data, and medium
CN121504182A