Unmanned aerial vehicle on-board edge computing method and system integrating perception-planning-control
By using a cloud-edge-airborne collaborative intelligent flight management system, combined with blockchain and digital twin technologies, the shortcomings of UAV air traffic management systems in real-time situational awareness, data consistency, and security verification have been addressed. This enables efficient and reliable airspace decision-making and continuous optimization, and allows for multi-aircraft collaborative obstacle avoidance in complex airspace environments.
Patent Information
- Application Number
- CN202610426274.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-02
- Publication Date
- 2026-06-26
AI Technical Summary
Existing UAV air traffic management systems are inadequate in terms of real-time situational awareness and decision-making, data consistency, security verification, and continuous autonomous evolution capabilities, making it difficult to cope with multi-aircraft collaborative avoidance and dynamic threats in complex airspace environments.
The intelligent flight management system adopts a cloud-edge-airborne collaborative approach. By combining a cloud brain decision center, an edge intelligent collaborative network, and an airborne autonomous safety entity, it utilizes a decentralized and trusted data lake powered by blockchain for data interaction and policy synchronization. Combined with a simulation and deduction environment driven by digital twins, it performs closed-loop iteration and autonomous evolution to achieve real-time perception, trusted decision-making, and continuous optimization.
It significantly reduces the response delay to sudden threats, ensures macro-optimal decision-making, improves the system's fault tolerance when some nodes fail, has the ability to respond to unknown threats in sub-second time, and can continuously adapt to dynamically changing airspace environments and mission requirements.
Smart Images

Figure CN122293694A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) flight control and air traffic management technology, specifically relating to a cloud-edge-airborne collaborative intelligent flight management system and its management method that combines cloud computing, edge computing, blockchain and artificial intelligence technologies. Background Technology
[0002] In recent years, with the explosive growth in the number of drones and the gradual opening of airspace for mixed manned / unmanned operations, traditional air traffic management systems are facing unprecedented challenges. Existing systems mostly employ centralized or hierarchical static planning models, which have the following inherent drawbacks:
[0003] First, the real-time nature of situational awareness and decision-making is insufficient. While cloud-based centers can process global information and perform macro-level planning, they suffer from significant delays in responding to sudden, localized dynamic threats (such as severe weather or unauthorized aircraft intrusion). On the other hand, relying entirely on onboard autonomous decision-making is limited by its limited sensing and computing power, making it difficult to address the challenges of multi-aircraft coordinated avoidance in complex airspace environments.
[0004] Second, there are challenges in system collaboration and data consistency. Data interaction between the cloud, edge nodes, and aircraft often adopts the traditional client-server model, resulting in significant issues such as asynchronous data updates and version inconsistencies. There is a lack of mechanisms to ensure the reliability and traceability of the overall system state. When some nodes fail or communication is interrupted, the overall system performance drops sharply.
[0005] Third, the ability to verify security and respond to unknown threats is weak. Existing rule-based or traditional machine learning-based decision-making systems have security boundaries that are difficult to rigorously prove. When faced with long-tail threats or coordinated attacks outside of the training data, they may make unpredictable or even dangerous decisions. Airborne control systems lack a safety reflection mechanism that does not rely on complex upper-level calculations in extreme emergency situations.
[0006] Fourth, the system lacks continuous autonomous evolution capabilities. Model and policy updates often rely on offline, batch retraining, making it impossible to utilize real-time flight performance data during system operation for online, personalized adaptive optimization, and thus difficult to adapt to dynamic changes in airspace structure and mission requirements.
[0007] Therefore, there is an urgent need in this field for a new generation of intelligent flight management system that can deeply integrate cloud-based global intelligence, edge group collaboration and airborne trusted execution capabilities, and has the characteristics of data consistency assurance, security verifiability and continuous autonomous evolution. Summary of the Invention
[0008] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a cloud-edge-airborne collaborative intelligent flight management system and method. This system and method aim to build a new generation of air traffic management system with real-time and accurate response, reliable collaborative decision-making, secure and verifiable execution, and continuous autonomous evolution capabilities through architectural innovation and the integration of cutting-edge technologies.
[0009] In a first aspect, embodiments of this application provide an integrated perception-planning-control UAV onboard edge computing system, the method comprising:
[0010] The cloud brain decision center, deployed on a cloud server cluster, serves as the strategy generation and evolution hub of the system, used to generate and continuously optimize global flight strategies.
[0011] The edge intelligent collaborative network consists of multiple heterogeneous edge computing nodes deployed at key nodes in the airspace. It serves as the real-time perception and computing network of the system and is used for local airspace threat perception and multi-machine collaborative conflict resolution computing.
[0012] An airborne autonomous safety agent, deployed on the UAV platform, serves as the execution verification and feedback terminal of the system, used to perform flight missions and transmit performance data back.
[0013] The cloud brain decision center, the edge intelligent collaborative network, and the airborne autonomous safety entity interact and synchronize data through a decentralized trusted data lake powered by blockchain. The system performs closed-loop iteration and autonomous evolution of global strategies, edge computing models, and airborne decision logic based on a simulation and deduction environment driven by digital twins.
[0014] Secondly, embodiments of this application provide an intelligent flight management method based on digital twins and collaborative evolution, applied to the UAV airborne edge computing system integrating perception-planning-control as described in the first aspect, the method comprising:
[0015] S1. Digital Twin Environment Construction and Initialization Steps: Based on high-precision geographic information, airspace rules, and equipment models, construct a full-scale digital twin airspace environment synchronized with the physical system;
[0016] S2. Cloud-based global policy co-evolution steps: In the twin environment, a hierarchical federated meta-reinforcement learning model is used to train the global policy model; based on feedback from the physical world and the twin environment, an evolutionary algorithm is used to continuously optimize the model architecture and parameters to generate a policy knowledge graph;
[0017] S3. Edge Distributed Perception and Collaborative Decision-Making Steps: Edge nodes utilize local multimodal data to recognize threats through causal reasoning; when facing group conflicts, they generate collaborative resolution strategies online through distributed consensus algorithms and game equilibrium solutions; and adapt cloud-based general models to local personalized models through online knowledge distillation.
[0018] S4. Airborne Trusted Execution and Reflection Control Steps: The airborne terminal performs runtime verification of the received policy within the trusted execution environment; it executes tasks using a layered control architecture and handles extreme contingencies through an independent reflection layer; it collects and transmits flight causal data back.
[0019] S5. Cross-layer data synchronization and system evolution steps: Ensure the consistency and reliability of all interactive data and model updates through a blockchain data lake; use the returned causal feedback data to drive the update of the digital twin environment and iterate from S2 to S4 to achieve the overall collaborative evolution of the system.
[0020] Thirdly, embodiments of this application provide an electronic device, including:
[0021] processor;
[0022] Memory used to store processor-executable instructions;
[0023] The processor is configured to implement the UAV airborne edge computing method with integrated perception-planning-control as described in the second aspect when executing the instructions.
[0024] Fourthly, embodiments of this application provide a computer-readable storage medium storing a program that instructs a device to execute the integrated perception-planning-control UAV onboard edge computing method as described in the second aspect.
[0025] Compared to existing technologies, this invention offers the following significant advantages: Distributed collaborative computing at edge nodes drastically reduces response latency to sudden threats; combined with cloud-based global optimization, it ensures macro-optimal decision-making. A decentralized data lake built using blockchain technology guarantees real-time synchronization, tamper-proofing, and traceability of cross-layer data, enhancing the system's fault tolerance in the event of partial node failures. Formal verification and a trusted execution environment provide rigorous security proofs for decision-making; a unique reflection layer design enables the system to instinctively avoid unknown sudden threats at sub-second levels. Simulation and online learning mechanisms based on a digital twin environment allow the system to continuously optimize strategies and models using real-time flight data, adapting to dynamically changing airspace environments and mission requirements. Multi-agent collaborative algorithms optimize airspace resource allocation and conflict resolution schemes, improving overall airspace capacity and operational efficiency. Attached Figure Description
[0026] Figure 1 This is a schematic flowchart of an integrated perception-planning-control UAV airborne edge computing method provided in an embodiment of this application.
[0027] Figure 2 This application provides an architecture diagram of an integrated perception-planning-control UAV airborne edge computing system.
[0028] Figure 3 A schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.
[0030] It should be noted that in the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.
[0031] Based on the embodiments described in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0032] Example 1
[0033] Figure 1 This is a schematic diagram of an integrated perception-planning-control UAV airborne edge computing system module provided in one embodiment of this application. Figure 1 As shown, the system includes:
[0034] The Cloud Brain Decision Center 110, deployed on a cloud server cluster, serves as the central hub for strategy generation and evolution within the system, generating and continuously optimizing global flight strategies. As the system's top-level intelligent hub, it generates, iterates, and distributes a globally optimal flight strategy library based on multi-source historical and real-time data through machine learning and optimization algorithms. It coordinates the overall evolution of the system, ensuring the macro-optimal nature and consistency of the strategies.
[0035] The edge intelligent collaborative network 120, composed of multiple heterogeneous edge computing nodes deployed at key airspace nodes, serves as the system's real-time perception and computing network, used for local airspace threat perception and multi-machine collaborative conflict resolution computation. Deployed at key airspace nodes, it forms a real-time response network responsible for dynamic threat perception and multi-machine collaborative computation in local airspace. It utilizes local computing power to quickly handle emergencies (such as conflict resolution) and collaborates with neighboring nodes to achieve efficient, distributed real-time decision-making.
[0036] The airborne autonomous safety agent 130, deployed on the UAV platform, serves as the system's execution verification and feedback terminal, responsible for performing flight missions and transmitting performance data. As the final execution unit, it receives upper-layer policies and controls the UAV's flight under strict safety constraints. It integrates a safety verification mechanism to ensure command reliability and collects and transmits flight performance data in real time, forming a closed-loop system optimization mechanism.
[0037] The cloud brain decision center, the edge intelligent collaborative network, and the airborne autonomous security body interact and synchronize data through a decentralized trusted data lake 140 powered by blockchain; the system performs closed-loop iteration and autonomous evolution of global strategies, edge computing models, and airborne decision logic based on a simulation and deduction environment 150 driven by digital twins.
[0038] The blockchain-enabled decentralized trusted data lake 140 serves as the system's data hub. Through blockchain technology, it ensures the immutability, traceability, and real-time synchronization of all interactive data (such as policies, states, and feedback) between the cloud, edge, and onboard terminals, resolving cross-layer data consistency and trust issues. The digital twin-driven simulation environment 150 serves as the system's security sandbox, constructing a virtual space environment synchronized with the physical world. This environment is used for risk-free testing, verification, and optimization of new policies, models, and control logic, driving the system to achieve secure, efficient, and autonomous continuous evolution.
[0039] Specifically, in this embodiment, the cloud brain decision center 110 includes a global policy generator, an adaptive policy distribution and orchestration unit, and a global model evolution engine, wherein:
[0040] The global policy generator employs a hierarchical federated meta-reinforcement learning framework to aggregate heterogeneous data from multiple sources, generating and continuously optimizing a benchmark flight policy knowledge graph covering various airspace structures and mission scenarios.
[0041] In large-scale urban logistics network planning, this module first works collaboratively through a hierarchical federated meta-reinforcement learning framework. Hierarchical refers to dividing the decision-making process into a strategic layer (e.g., regional distribution center planning), a tactical layer (e.g., parcel route allocation), and a maneuver layer (e.g., real-time obstacle avoidance for drones); federated refers to training each logistics company's local data (e.g., historical delivery routes, traffic flow) locally, uploading only model updates rather than raw data to protect privacy; meta-reinforcement learning is a method that enables models to learn quickly, allowing them to adapt rapidly to small amounts of new urban data. This framework aggregates multi-source data such as city maps, traffic flow, weather, and building heights to generate a massive flight strategy knowledge graph. This graph is stored in a graph structure, with nodes representing different task scenarios (e.g., peak-hour business district delivery, nighttime medical supply transportation), and edges representing transferable relationships between strategies (e.g., rainy day delivery strategies can be transferred to foggy day delivery). The graph is continuously optimized based on new data; for example, when the construction of new high-rise buildings in a certain area increases signal interference, the graph automatically updates the strategy associations and obstacle avoidance logic of relevant nodes.
[0042] The adaptive policy distributor and orchestrator dynamically orchestrates and distributes differentiated policy fragments to corresponding edge collaboration nodes based on real-time airspace load, network status, and task urgency. Specifically, in the scenario of UAV emergency material delivery, the orchestrator continuously monitors the real-time load of each airspace (such as UAV density and task queues). When the communication network status of a disaster area edge node shows insufficient bandwidth, the orchestrator extracts key information from the global policy knowledge graph, converts detailed path planning policy fragments into compressed versions containing only key waypoints and emergency protocols (differentiated policies), and distributes them to that node. Simultaneously, for emergency medical supplies delivery with high task urgency, the orchestrator temporarily increases its policy priority and proactively pushes relevant policy fragments to all edge nodes along the route to ensure relay support. During the distribution process, the orchestrator is also linked to the future functions of the onboard autonomous safety agent to ensure that the distributed policy fragments can be correctly identified and processed in the safety verification module. These distribution records and subsequent execution effects are synchronized to the trusted data lake as feedback data for subsequent optimization of the global model evolution engine.
[0043] The global model evolution engine utilizes feedback data from the edge intelligent collaborative network and the airborne autonomous safety agent, combined with simulation results, to perform structural and parametric co-evolution of the global policy model through evolutionary algorithms and neural network architecture search techniques. This engine drives the overall evolution of the system. It first receives feedback from the edge intelligent collaborative network, such as a report from an airport edge node that the efficiency of a traditional queuing landing strategy decreases by 30% under crosswind conditions, and data from the airborne autonomous safety agent, such as a UAV recording that its current maneuvering strategy leads to an abnormal increase in energy consumption when encountering turbulence. The engine reproduces these real-world problems in a digital twin simulation environment, constructing a high-fidelity test scenario. Then, the engine initiates a two-stage co-evolution:
[0044] Parametric evolution: Using evolutionary algorithms that simulate the selection, crossover, and mutation processes of biological populations, the engine creates multiple parametric variants of a policy model (like multiple fine-tuned versions of a single policy), which compete for survival in a digital twin environment. Poor-performing variants are eliminated, and the parameters of superior variants are combined and fine-tuned, iteratively optimizing model performance.
[0045] Structural Evolution: This employs neural network architecture search technology, an automated method for designing neural network structures. The engine attempts to adjust the internal logical structure of the policy model. For example, it automatically tests and discovers that adding a dedicated decision-making submodule for handling complex weather cooperative avoidance improves the model's overall performance under severe weather conditions. The evolved next-generation global policy model will be used by the global policy generator to update its policy knowledge graph and distributed to edge nodes by an adaptive policy distributor and orchestrator. Simultaneously, successful structural or parameter adjustments discovered during the evolution process can serve as prior knowledge, feeding back into the initial design of the hierarchical federated meta-reinforcement learning framework, forming a closed loop from application to basic research.
[0046] Specifically, in this embodiment, each heterogeneous edge computing node in the edge intelligent collaborative network 120 includes a multi-source threat cognition engine, a group collaborative decision-making unit, and a personalized model distillation unit, wherein:
[0047] The multi-source threat recognition engine employs a combination of multimodal pre-trained models and spatiotemporal causal reasoning, fusing multimodal data from images, radar, and text to achieve causal prediction and interpretable attribution of dynamic threats. At an edge computing node in the airspace of a coastal airport, the threat recognition engine receives three sets of real-time data: a weather radar image showing rapidly developing convective clouds in a certain direction (radar mode); an onboard camera of an approaching flight capturing the initial shape of a funnel cloud beneath a cloud cluster ahead (image mode); and simultaneously, a text message automatically issued by the flight information service indicating a thunderstorm warning for the area (text mode). The engine first inputs these three different types of data into a large multimodal pre-trained model. This model has learned deep correlations between different modalities through massive amounts of images, radar charts, and weather reports, enabling it to preliminarily determine the potential threat of tornado formation. Next, the engine activates its spatiotemporal causal reasoning module. This module not only analyzes the current data but also traces the meteorological evolution sequence over the past half-hour. Combining this with local geographic databases (such as topography and land-sea distribution), it infers the cause of the threat: warm, moist air flows in from the sea, converging with cold air over land at the coastline. Under orographic lifting, this forms strong convection, which evolves into a tornado under specific wind shear conditions. The engine's final output prediction is: within the next 8-12 minutes, the tornado will move northeastward, affecting the approach path of Runway 05. This is accompanied by a visualized source tracing report containing the aforementioned causal chain and reference radar echo sequences for controllers and subsequent decision-making units.
[0048] Supporting the aforementioned explainable threat prediction is a recursive Bayesian filtering framework that integrates a causal graph model, as proposed in this invention. This framework formalizes threat cognition as a state estimation problem, and its core mathematical model is as follows:
[0049] 1. Definition of State Space and Observation Model. Threat Causal State: Let express The threat state vector at any given time includes not only physical attributes (location, velocity, type) but also implicit causal factors (such as convection intensity and wind shear index). Causal dynamic model: State evolution is driven by a parameterized causal graph. Driver, i.e. ,in To encode the nonlinear state transition function that encodes causal relationships, Let be the process noise covariance. Multimodal likelihood model: [The remaining text appears to be incomplete and requires further context.] Observations of each modality By fusion model Generation likelihood: ,in For the observation function, To observe the noise covariance.
[0050] 2. Online filtering process (prediction-update loop). The system runs the following Bayesian filtering steps in real time to solve for the posterior distribution. :
[0051] (predict):
[0052] (renew): ,
[0053] in, Given the posterior probability distribution and the state conditional distribution after all observations, The modal likelihood function is the probability of observing data in a given state, and measures the degree of matching between a single piece of sensor evidence and the assumed state. It represents the causal state transition probability, the state evolution probability based on causal mechanisms, and the interpretable prediction of the next stage of a threat's development based on physical causal laws. The likelihood function quantifies the probability that the current threat state is... So, the actual data observed , How likely is it that...? The total likelihood is the product of the modal likelihoods. ∫ denotes the integral.
[0054] 3. Implementation and Evolution. In practical systems, the posterior distribution is efficiently approximated using a particle filter or an extended Kalman filter. Causal graph parameters. With fusion model parameters By leveraging local data streams from edge nodes, continuous updates can be performed through online variational inference, enabling the threat perception model to autonomously evolve and adapt to new threat patterns. The final output is a maximum a posteriori estimate. and its causal tracing path This provides interpretable input for collaborative decision-making.
[0055] The group collaborative decision-making unit employs a distributed collaborative algorithm based on graph neural networks and game theory. When multi-aircraft conflict or complex airspace congestion is detected, it generates a distributed resolution solution that satisfies the group's optimality without relying on central coordination among nodes. Specifically, within the jurisdiction of a certain edge node in an urban air traffic (UAM) corridor, a sudden public event leads to temporary airspace control, requiring eight eVTOLs (electric vertical takeoff and landing aircraft) to quickly adjust their routes and converge on an alternate landing site, resulting in a complex multi-aircraft convergence conflict. The group collaborative decision-making units of this node and its adjacent nodes are activated. Each unit first abstracts each eVTOL within its jurisdiction and its flight plan (position, speed, destination) as a node in the graph, and abstracts the relative positions and speeds between aircraft, as well as potential conflict risks (such as the expected minimum separation being lower than the safety standard), as edges connecting these nodes, jointly constructing a dynamic conflict relationship graph. A graph neural network is applied to this graph, which can efficiently process such irregular relational data, learn the complex interactions between node (aircraft) states and edges (conflict relationships), and quickly assess the overall conflict situation. Subsequently, each unit modeled the problem as a cooperative game based on game theory. Each eVTOL is a participant, whose strategy is to adjust speed, altitude, or make a slight yaw. The objective function is to minimize overall delay, optimize energy consumption, and meet safety intervals. Each edge node only exchanges encrypted local intent and cost information with other nodes within its communication range. Through multiple rounds of consensus-based iterative computation, they negotiate a scheme in a distributed manner: for example, eVTOLs A and B slightly accelerate first, C and D climb slightly and pass overhead, and E through H decelerate and increase their spacing in sequence. This scheme does not require cloud-based central adjudication and is accepted by all participants, achieving group optimality. After the scheme is formed, it will undergo rapid feasibility verification through a local model optimized by a personalized model distillation unit and is ready to be deployed to the airborne autonomous safety body for execution.
[0056] The personalized model distillation unit receives a general model from the cloud brain decision center and uses historical data and real-time feedback from its local area to generate a lightweight, high-precision personalized edge model through online knowledge distillation technology. Specifically, an edge node deployed in a mountainous power line inspection area is tasked with controlling a drone to autonomously inspect high-voltage power lines. The cloud-based general drone obstacle avoidance and path tracking model experiences reduced control accuracy when facing the region's unique, strong, and variable canyon winds, leading to increased drone swaying and blurry images. At this point, the personalized model distillation unit within the node begins operation. It continuously receives two types of local data: first, real-time monitoring and short-term prediction data of the local wind field from a multi-source threat cognition engine (as input features); second, data from the flight controller in the onboard autonomous safety body, such as the deviation between the actual flight trajectory and the ideal trajectory, and motor power consumption data (as the model's target or feedback). The unit employs online knowledge distillation technology: the powerful general model in the cloud serves as the teacher model, its complex network possessing strong generalization capabilities; the unit's goal is to train a student model with few parameters and fast inference speed to mimic the teacher. In the continuous data stream, the student model not only learns the obstacle avoidance strategies of the teacher for general obstacles (such as towers and trees), but more importantly, it utilizes the local, frequently occurring canyon wind field data to focus on learning and internalizing control strategies to counteract this specific interference. Ultimately, the unit generates a lightweight flight control model optimized specifically for this inspection corridor. This model is deployed on the node to generate more accurate and stable control commands in real time and serves as local knowledge accumulation. The model's effectiveness data (such as the percentage improvement in inspection image clarity) is uploaded via a trusted data lake as feedback, providing a reference for the global model evolution engine to generate more adaptable general models for complex terrains in the future, forming a collaborative optimization loop from global to local and back to global.
[0057] Specifically, in this embodiment, the airborne autonomous safety body 130 includes a trusted fusion and security shield module, a hierarchical-reflective adaptive controller, and an efficiency acquisition and causal feedback unit, wherein:
[0058] The Trusted Fusion and Security Shield module integrates runtime verification and a trusted execution environment to perform millisecond-level formal verification of external strategies and local decisions. It employs an adaptive confidence propagation network to dynamically evaluate the credibility of each information source. A drone is performing a patrol mission, and its onboard security shield module simultaneously receives suggested flight path adjustments around thunderstorm areas from an edge intelligent collaborative network and optimal energy-consuming flight paths autonomously planned by local sensors. These strategies are first decrypted and analyzed in an isolated trusted execution environment (a hardware-level secure zone resistant to software attacks). Subsequently, the module initiates runtime verification, using formal methods to transform flight safety rules (such as maintaining a minimum terrain clearance of 100 meters) into machine-verifiable logical propositions, providing millisecond-level security proofs for both paths. Simultaneously, the adaptive confidence propagation network is activated, an algorithm that dynamically adjusts the credibility of information sources based on their historical performance and real-time consistency. Network evaluation data: the edge network strategy is based on weather radar and multiple historical successful detours (high confidence), while the local sensor's strategy suffers from signal interference due to current heavy rainfall (confidence dynamically lowered). Finally, the module integrates the verification results and confidence weights within the TEE, selects and securely signs the edge policy, rejects risky autonomous paths locally, and stores the verification records as an audit chain in the airborne black box for analysis by the performance acquisition unit.
[0059] The hierarchical-reflective adaptive controller comprises a task planning layer, a behavior layer, and a reflection layer. The task planning layer executes high-level task sequences. The behavior layer generates smooth trajectories based on model predictive control. The reflection layer is a high-speed reaction loop independent of the upper layers, based on a predefined safety primitive library, to make sub-second-level instinctive avoidance responses to sudden and unforeseen threats.
[0060] To ensure the overall stability and safety of this hybrid control system containing independent high-speed reflection loops, this invention establishes a unified stability analysis framework based on the Lyapunov method. The closed-loop dynamics of this system can be expressed as:
[0061] ,
[0062] This equation describes the drone's state. Over time rate of change How is it generated? It consists of three parts added together. Among them, The nonlinear function (including gravity and basic aerodynamic effects) is used to describe the nominal dynamics of the UAV. Extend the state vector (such as position, velocity, attitude) of the UAV. This represents the rate of change of the state vector over time. This represents the role of the behavioral layer in controlling input. This represents the function of the reflective layer in controlling the input. and These are the action matrices for the control inputs of the behavior layer and the reflection layer, respectively. Their elements are determined by the configuration and efficiency of the UAV's actuators. The two can be the same (sharing actuators) or different (equipped with dedicated emergency mechanisms). and These are the control quantities output by the two layers of controllers. The function is used to characterize the control quantity of the reflective layer at the moment of sudden threat triggering. The instantaneous pulse application.
[0063] The core theoretical contribution of this invention lies in proposing the following stability criterion:
[0064] , making ,
[0065] in, It is a Lyapunov function. The nominal exponentially stable rate, guaranteed by the behavioral layer, is represented by the second term on the right-hand side of the inequality. The instantaneous disturbance caused by the reflection action is quantified, and its effect must be expressed over time at a rate of... Exponential decay. This criterion provides a mathematical guideline for the design and formal verification of each reflective action in the safety primitive library: any primitive must, after being triggered, guarantee that the system state satisfies the above inequality, thereby ensuring that the system can smoothly recover to the controllable range of the behavioral layer while achieving rapid risk avoidance. This is fundamentally different from simple rule-based reflection or dual-mode switching control, which lack strict stability guarantees.
[0066] In a complex urban aerial logistics scenario, a drone receives a task sequence from the cloud-based decision-making center: take off from warehouse A → pass through waypoints B and C → land on rooftop D. The task planning layer breaks this down into executable flight segments. The behavior layer uses model predictive control, a rolling optimization control method that calculates a smooth, efficient flight trajectory that meets constraints (such as maximum turn rate) in real time for the next few seconds based on the drone's current state, dynamics model, and environmental predictions. While the drone is flying normally towards waypoint C, a small consumer-grade drone, without any prior sensor warning, suddenly crosses its path at high speed from a building blind spot (a sudden threat). At this moment, the upper-level behavior and planning layers cannot respond in time due to computational and communication delays (hundreds of milliseconds). The independent reflective layer is activated instantaneously. This layer employs independent low-latency sensors (such as event cameras) and a processor. Without complex planning, it directly selects and triggers the most suitable emergency jump primitive within 100 milliseconds based on the relative position and speed of the threat from its safety primitive library (a rigorously validated and biomimetic-optimized library of emergency actions, such as emergency jump, emergency stop hovering, and small-angle sideslip). This causes the drone to instinctively climb vertically, successfully avoiding a collision. Detailed triggering data of this reflective action and its preceding and following flight states are fully recorded by the performance acquisition unit.
[0067] The performance acquisition and causal feedback unit collects flight status data and constructs a flight decision-outcome causal graph, analyzes the root causes of strategy advantages and disadvantages, and generates a structured performance report for feedback. After the task is completed, the performance acquisition unit begins its work. It comprehensively collects data throughout the entire flight: decision records from the trusted fusion and safety shield module (why the detour path was chosen), control commands and actual execution status from each layer of the hierarchical controller (including emergency jump events triggered by the reflection layer), and raw data from sensors (position, energy consumption, turbulence level). The unit uses causal inference algorithms to analyze this time-series data. For example, the analysis found that although the strategy of detouring around thunderstorm areas (decision) increased the flight distance by 15%, it avoided severe turbulence and energy consumption surges caused by turbulence (outcome), resulting in a positive net benefit. Especially for emergency jump events, the unit deeply analyzes the root cause: it is due to a coverage gap in the threat perception of the edge intelligent collaborative network in the building blind spot (cause), and constructs a local decision-outcome causal graph based on this. Finally, the unit generates a structured performance report, including a strategy effectiveness assessment, causal analysis of security incidents, and recommendations for system weaknesses (such as suggesting the addition of area awareness capabilities). This report is uploaded via a secure channel to a blockchain-enabled decentralized trusted data lake, where it is used by the global model evolution engine of the cloud brain decision center for model optimization. It can also be used by the edge collaborative network to adjust its threat perception range or optimize its collaborative strategies, forming a closed-loop learning and improvement process from individual flight to system evolution.
[0068] Furthermore, the group collaborative decision-making unit adopts a credit allocation and reputation mechanism: each edge node obtains corresponding credit points after contributing to the conflict resolution solution, and the points affect its voice in subsequent collaborative decision-making; at the same time, the reputation values of other nodes are maintained to identify and isolate nodes that may provide false information or engage in malicious behavior.
[0069] Specifically, in dense urban airspace, the drone swarms under the jurisdiction of edge nodes A, B, and C face congestion conflicts at crossroads. The collaborative decision-making unit of the three nodes initiates distributed collaborative computing.
[0070] In the first round of collaboration, Node A, leveraging its advanced multi-source threat recognition engine, proposed an efficient diversion scheme (e.g., the eastern drone swarm collectively climbs 50 meters, while the western drone swarm maintains its original altitude). This scheme was subsequently proven to be close to swarm optimum in graph neural network and game theory deductions. Therefore, the system automatically recorded +10 credit points for Node A through a credit allocation mechanism. Meanwhile, Nodes B and C also participated in refining the scheme. Node B contributed fine-tuned speed adjustment parameters (+5 points), while Node C's data was incorrect, leading to its suggestion being rejected (no points awarded).
[0071] The second round of collaboration: An emergency medical drone priority passage incident occurs in another area. At this time, the credit scores of each node directly affect its influence. The proposal to open a temporary fast track by node A (15 points) with a high credit score is given a higher initial weight, the auxiliary adjustment suggestions of node B (5 points) with a medium credit score are moderately adopted, while node C with a low credit score and a history of bad behavior is restricted to participating in collaboration within a limited scope.
[0072] The system continuously tracks node behavior. Node C's reputation value steadily declines due to repeatedly providing conflicting data (such as false airspace occupancy reports). On one occasion, Node C suddenly broadcast an emergency alert about an illegally intruding drone in the northbound airspace, requesting all nearby drones to evade it. However, because Node C's reputation value had fallen below the system's malicious behavior isolation threshold, other nodes (A and B) did not immediately execute its instructions. Instead, they first cross-verified the information using their own multi-source threat recognition engine and neighboring nodes. The verification revealed the information to be a false alarm or malicious interference. Therefore, the system not only rejected Node C's instructions but also penalized it by significantly reducing its credit score and marking its reputation value as untrustworthy. In subsequent collaborative decision-making, Node C's proposals will be automatically downgraded or even ignored by the system, and its transmitted information will be carefully reviewed by other nodes, effectively isolating potentially erroneous or malicious nodes and ensuring the overall robustness and security of the collaborative network.
[0073] Throughout the entire collaboration process, changes in credit scores and reputation values, as well as the final solutions and execution effects of each collaboration, are meticulously recorded. These records are synchronized through a blockchain-enabled decentralized trusted data lake, ensuring their immutability. Ultimately, this historical data, reflecting the collaborative capabilities and trustworthiness of nodes, can serve as crucial feature inputs, feeding back into the global model evolution engine of the cloud brain decision center. This data optimizes parameters for future collaboration algorithms (such as credit allocation formulas and reputation decay models), and even guides the adaptive strategy distributor and orchestrator to prioritize nodes with high reputation and high credit as lead nodes for regional collaboration. This creates a tightly linked, self-reinforcing closed loop from individual behavior to group trust, and finally to the system's global optimization strategy.
[0074] The above mechanism is implemented by the following rigorous mathematical model, and deeply embedded with a collaborative optimization algorithm:
[0075] 1. Credit-reputation driven dynamic weight optimization. Each node The optimization problem for finding a solution to local conflict resolution is as follows:
[0076] ,
[0077] ,
[0078] in, For decision-making actions, This is the local cost function. It is the set of neighboring nodes.
[0079] and These are, respectively, one's own credit and one's neighbor's reputation. Credit The higher the level, the greater the range of innovation allowed for the node. The larger the value, the more firmly it adheres to its own direction (penalty weight decreases); reputation value The higher the node The more inclined one is to trust and align with nodes Neighbor proposal vector . ξ is an adjustable parameter. For nominal actions, in the absence of conflict, nodes The originally planned actions (such as the original flight path). This is a credit-related trust radius function. Nodes are defined. Permitted actions Deviating from its nominal actions The maximum range.
[0080] 2. Dynamics of Credit and Reputation Updates. After each round of collaboration, credit and reputation are iteratively updated based on the results:
[0081] ,
[0082] ,
[0083] in, ∈(0,1) is the learning rate. This is a performance function based on the final execution effect of the solution (such as the reduction in average delay). To measure the proposal With final consensus A consistent function.
[0084] 3. Algorithm convergence and robustness.
[0085] This model constructs a self-reinforcing positive feedback loop by transforming credit (C) and reputation (R) from traditional external metrics into internally dynamically adjusted factors during distributed optimization iterations. Specifically, a node's high performance enhances its influence, guiding the entire network to find better solutions. The realization of these better solutions further solidifies the node's high performance, forming a virtuous cycle. Mathematically, this design ensures that the collaborative network maintains robustness (i.e., the ability to resist interference and operate stably) and self-organizing optimization capabilities (i.e., the ability of nodes to autonomously achieve global optimization through local interaction without central control) even in the presence of faulty or malicious nodes. In short, through the dynamic adjustment of credit and reputation, this model enables the network to work efficiently and stably, continuously optimizing even in complex environments with destructive factors.
[0086] Furthermore, the primitives in the security primitive library of the reflection layer are generated and optimized in the digital twin environment through adversarial reinforcement learning, and their triggering conditions and response actions are encoded into verifiable formal rules.
[0087] In urban low-altitude logistics scenarios, to address the specific and sudden threat of falling flowerpots from windowsills that drones may encounter, the system generates dedicated safety primitives through the following closely interconnected steps: First, the global policy generator in the cloud brain decision center identifies falling objects from heights as a high-incidence threat pattern from the city management database and accident reports, and issues this requirement. The multi-source threat recognition engine in the edge intelligent collaborative network supplements real-time data: based on historical camera data, it statistically identifies a higher risk of falling objects in specific older residential areas during specific time periods. This information is synchronized to a digital twin-driven simulation environment, constructing a virtual training scenario that includes falling object trajectories at different heights, angles, and speeds.
[0088] Adversarial reinforcement learning is initiated in a digital twin environment. Specifically, the defensive agent learns emergency avoidance strategies for the drone, aiming to maximize the success rate of avoidance while minimizing trajectory deviation and energy consumption. The adversarial agent learns to generate the most threatening falling object trajectories (such as objects thrown from a blind spot with a specific initial velocity), aiming to maximize the probability of the drone failing to avoid them. During adversarial training, the defensive agent gradually learns to recognize early characteristics of falling objects (such as shadow changes and initial motion) and masters efficient combined maneuvers (such as emergency lateral movement + descent). The training process also receives constraints from the collaborative decision-making unit of the edge node group: avoidance actions must not endanger other drones in the adjacent airspace.
[0089] The primitive P-01, emergency avoidance of falling objects from a window sill, was extracted from the successfully trained strategy. The trigger condition and response action of this primitive were formally encoded as follows: Trigger condition: IF [The lidar detects an object within the window projection area at time t0, and its vertical displacement exceeds the threshold δ within time Δt]; Response action: DO [Execute the control sequence within time t1 (t1 < 100 ms): increase the roll angle by α degrees while applying a downward acceleration β]. This formal rule was submitted to the trusted fusion and safety shield module of the airborne autonomous safety body for verification, ensuring that the action sequence does not violate the physical and dynamic constraints of the UAV and does not conflict with known no-fly zones or buildings.
[0090] The verified primitive P-01 was injected into the reflective layer security primitive library of drones in the relevant area. When a drone was performing a delivery mission in the old city, the reflective layer successfully identified and triggered P-01, avoiding falling objects. All data from this event was recorded by the performance acquisition and causal feedback unit. Analysis showed that the triggering timing could be advanced by 50ms to reduce energy consumption. This feedback was synchronized through a blockchain-enabled decentralized trusted data lake.
[0091] Based on feedback from real-world events, the global model evolution engine of the cloud brain decision center initiated targeted optimizations: enhancing the training weights for such scenarios within the digital twin environment. Simultaneously, the personalized model distillation unit of the edge intelligent collaborative network utilized locally collected data from multiple falling object incidents to locally fine-tune the primitive P-01, generating a lightweight version more suitable for the building characteristics of the local area. This optimized primitive version is shared through a trusted data lake, and edge nodes in other areas can selectively adopt it as needed.
[0092] Example 2
[0093] like Figure 2 As shown, this application provides a flowchart of an intelligent flight management method based on digital twins and collaborative evolution, which is applied to the intelligent flight management system described in Embodiment 1, including...
[0094] S1. Digital Twin Environment Construction and Initialization Steps: Based on high-precision geographic information, airspace rules, and device models, construct a full-scale digital twin airspace environment synchronized with the physical system. Map the initial state of the real world (such as UAV location and weather conditions) to the twin environment, establishing a starting point for virtual-real synchronization. This provides a safe, controllable, and repeatable digital sandbox for all subsequent learning, verification, and simulation, avoiding the high risks and costs associated with trial and error directly in the physical world.
[0095] S2. Cloud-based Global Policy Co-evolution Steps: In the twin environment, a hierarchical federated meta-reinforcement learning model is trained using hierarchical federated meta-reinforcement learning. Based on feedback from the physical world and the twin environment, evolutionary algorithms are used to continuously optimize the model architecture and parameters, generating a policy knowledge graph. This step generates and continuously optimizes the system's brain and knowledge base. Through hierarchical federated meta-reinforcement learning, strategic (task allocation), tactical (path planning), and execution (maneuver control) policies are learned simultaneously through hierarchical layering. Through federation, local training results from edge nodes in various locations are aggregated without centralizing the original data, protecting privacy. Meta-reinforcement learning enables the model to learn quickly, allowing it to adapt rapidly to a small amount of new scenario data. Through policy knowledge graph construction, the learned policies are stored in a graph structure, with nodes representing scenarios or policies and edges representing transferable relationships, forming a queryable and reasonable policy knowledge system. Through model evolution, based on real feedback and twin inference, evolutionary algorithms are used to optimize model parameters, and neural architecture search is used to optimize the model structure, achieving continuous improvement in model capabilities. This forms the system's highest-level intelligent decision-making capability, providing an optimal policy benchmark that adapts to different scenarios and continuously evolves.
[0096] S3. Edge Distributed Sensing and Collaborative Decision-Making Steps: Edge nodes utilize local multimodal data to recognize threats through causal reasoning; when facing group conflicts, they generate collaborative resolution strategies online through distributed consensus algorithms and game equilibrium solutions; and they adapt cloud-based general models to local personalized models through online knowledge distillation. This provides low-latency, highly reliable real-time intelligence in the last mile near the threat, alleviating the latency pressure of centralized cloud processing and adapting to local specificities.
[0097] Furthermore, the distributed consensus algorithm in S3 is an air traffic consensus protocol based on an improved practical Byzantine fault-tolerant algorithm. This protocol takes the relief scheme as a proposal, reaches consensus through multiple rounds of voting and proof, and tolerates no more than one-third of the nodes failing or engaging in malicious behavior.
[0098] The algorithm used in this embodiment is an improvement on the Practical Byzantine Fault-Tolerant Algorithm (PBFT). PBFT is a classic consensus algorithm that allows distributed systems to reach consensus even when a minority of nodes act maliciously or fail. Its core is to ensure that honest nodes reach consensus on the same decision through a three-phase protocol (pre-ready, ready, commit) and multiple rounds of communication.
[0099] For air traffic scenarios, the protocol is improved as follows: Proposal content: Disconnection schemes (such as a combination of commands for speed adjustment, altitude layer allocation, and heading fine-tuning). Consensus goal: All normal nodes will eventually execute the same, effective disconnection scheme. Fault tolerance: With a total number of nodes N=3, it can tolerate no more than f=floor((N-1) / 3)=1 node failures or malicious (Byzantine) behavior.
[0100] Detailed steps (taking node A as the proposal initiator as an example):
[0101] Phase 1: Pre-preparation Phase. In the current view (a logical round), node A is elected as the master node (leader). Node A submits its locally generated relief solution Pa (containing solution details, timestamp, and current view number) as a proposal, attaches its own digital signature, and broadcasts it to all backup nodes (nodes B and C). Upon receiving the proposal, nodes B and C verify it: checking if the proposal format is correct and checking if the digital signature is valid.
[0102] The logic of the local trusted fusion and security shield module is invoked to perform preliminary and rapid formal security verification on solution Pa (such as checking for violations of minimum interval rules). This is a key improvement for air traffic scenarios, ensuring that the consensus object is primarily a safe and feasible solution.
[0103] Phase Two: Preparation Phase.
[0104] If nodes B and C pass verification, they each enter the preparation phase, and a...<PREPARE,v,n,d(A),i> The message (where v is the view number, n is the sequence number, d(A) is the digest of proposal Pa, and i is its own node number) is signed and broadcast to all other nodes (including the primary node A and other backup nodes). This is equivalent to casting a vote of approval.
[0105] Each node (including A, B, and C) is collecting PREPARE messages sent by other nodes.
[0106] Phase 3: Submission Phase.
[0107] For a proposal (e.g., Pa), when node i (e.g., node B) collects at least 2f (2 in this example, from A and C) valid PREPARE messages from different nodes, and these messages are all for the same (v,n,d) (i.e., the same view, sequence number, and proposal summary), node B considers that proposal Pa has obtained proof of readiness and can enter the submission phase.
[0108] Node B broadcasts a message<COMMIT,v,n,d(A),i> Signature message.
[0109] When node B collects at least 2f+1 (3 in this example, from A, C, and itself) valid COMMIT messages for the same (v,n,d), node B finally commits the proposal Pa. Committing means that node B formally recognizes the proposal Pa as the consensus result and is ready to execute or forward it.
[0110] The fault tolerance and view change handling process is as follows: Scenario 1: Master node A fails (no response): Within the timeout period, nodes B and C do not receive a valid proposal from the master node. At this time, nodes B and C will trigger the view change protocol, jointly elect a new master node (such as node B), and re-initiate consensus under the new view. Because the number of normal nodes (2) is greater than two-thirds of the total number of nodes (2>3*2 / 3), the system can continue to run.
[0111] Scenario 2: Malicious Backup Node C (Sending Conflicting Messages): Suppose node C is a malicious node. It sends a PREPARE message to node A in favor of Pa, but sends a PREPARE message to node B opposing or regarding other proposals. According to the nature of the PBFT protocol, as long as the number of malicious nodes does not exceed f (1 in this example), honest nodes A and B will eventually be able to collect enough (2f+1) consistent COMMIT messages to reach a consensus on Pa. Node C's contradictory behavior will be recorded and affect its score in the credit and reputation mechanism.
[0112] The consensus-reached resolution plan serves as the fundamental input for solving the game equilibrium, ensuring that collaborative decision-making is refined and optimized based on a recognized candidate solution. All signed messages in the consensus process constitute important auditable data, which can be synchronized to a blockchain-enabled decentralized trusted data lake for post-event analysis and credit allocation. The three nodes ultimately reach an agreement on the execution plan Pa. Subsequently, based on the consensus result, each node directs the drones under its jurisdiction to execute the corresponding part of plan Pa, completing the collaborative conflict resolution.
[0113] S4. Airborne Trusted Execution and Reflection Control Steps: The airborne terminal performs runtime verification of the received policy within a trusted execution environment; it executes tasks using a layered control architecture and handles extreme contingencies through an independent reflection layer; it collects and transmits flight causal data. As the final gateway for the system's interaction with the physical world, it translates intelligent decisions into safe physical actions and ensures a safety buffer even in extreme situations.
[0114] Furthermore, the runtime verification in S4 specifically involves: encoding the flight safety protocol into a time-series logic formula, converting the flight strategy or control command to be verified into a hybrid automaton model, completing the reachability analysis within a limited time using model verification technology, and outputting a verification certificate or counterexample path.
[0115] A drone performing power line inspection received a new flight path suggestion from an edge intelligent collaborative network. This suggested route, to shorten the distance, required flying close to a hillside. Within the trusted execution environment of the drone's trusted fusion and security shield module, runtime verification of this flight path suggestion was initiated to ensure absolute safety. First, the safety rules in the flight manual were translated into mathematical statements that the machine could rigorously process, using temporal logic. This is a mathematical language specifically designed to describe how the system state should change over time.
[0116] 1. Mathematical Expression of Safety Rules. Terrain Separation Rule: Translated into the requirement that the distance between the UAV and the nearest terrain point must always be greater than 50 meters throughout the entire flight. Battery Safety Rule: Translated into the requirement that once the battery level drops below 20%, the UAV must be able to reach the nearest alternate landing site within 15 minutes. Airspace Restriction Rule: Translated into the requirement that the flight altitude must be strictly maintained within the specified range when crossing temporary control zones. These are no longer vague natural language rules, but precise and unambiguous mathematical propositions.
[0117] 2. Constructing a state machine model of the flight strategy. Next, the flight strategy to be verified (a new route close to the hillside), along with the UAV's physical characteristics (such as maximum climb rate, turning radius, and battery consumption model) and external environmental factors (such as the 3D model of the mountain and predicted wind speed), are collectively constructed into a hybrid automaton model. A hybrid automaton can be understood as an intelligent state machine that can describe both the different modes the UAV is in (such as discrete states like cruise, climb, and evasion) and the continuous changes in its position, speed, and battery power within each mode. It abstracts the complex flight process into a set of rules and equations that a computer can systematically analyze.
[0118] 3. Perform exhaustive state search verification. Then, use model verification techniques to analyze the constructed hybrid automata model. Model verification: The core idea is to use the computing power of the computer to systematically and exhaustively simulate all possible flight scenarios from takeoff (including various reasonable wind speed changes, minor sensor errors, command execution delays, etc.) to check if there is even one scenario that would cause the UAV to violate the safety regulations set in step 1. Reachability analysis: This is the key operation of model verification, which involves analyzing all possible flight paths (state sequences) to see if there is a path that can reach a dangerous state (such as less than 50 meters from a mountain). This process will be completed within a set maximum task time window to ensure that the verification can provide conclusions within a limited time.
[0119] 4. Generate executable verification conclusions. After the verification calculation is completed, a clear conclusion is generated:
[0120] Scenario 1: Validation Passed. The model verifier confirms that the UAV will not violate safety regulations when executing this strategy under all simulated and reasonable conditions. At this point, it generates a validation certificate. This certificate is a digitally signed, structured security proof indicating that this route is safe under the current model. The controller will confidently execute the route based on this certificate.
[0121] Scenario 2: Validation Failure. The model validator detects at least one possible scenario (e.g., encountering a persistent crosswind in a specific direction during the third segment of the flight path) that would cause the UAV to continuously correct its course to maintain heading, eventually causing it to drift too close to the hillside due to inertial drift at a certain point, violating the terrain spacing convention. In this case, the validator will output a counterexample path—a detailed description of the complete story chain from takeoff, including the states and events that ultimately led to the danger.
[0122] The safety shield module immediately rejected the proposed route. Simultaneously, the counterexample path was sent to the performance acquisition and causal feedback unit. This unit conducted an in-depth analysis of the counterexample and discovered that the core reason was that the route was on the leeward side of a hillside, making it particularly sensitive to crosswinds. This analysis result generated a structured report, which was fed back to the edge intelligent collaborative network via the blockchain data lake. In subsequent personalized model distillation, the edge network can enhance its learning of similar terrain wind field effects; during collaborative decision-making, it will also design more conservative safety margins for routes in similar areas. This counterexample may also be uploaded to the cloud brain decision center to enrich the training scenarios in the digital twin environment, drive the evolution of the global policy model, and avoid generating similar risky policies in the future.
[0123] S5. Cross-layer data synchronization and system evolution steps: A blockchain data lake ensures the consistency and reliability of all interactive data and model updates; causal feedback data is used to drive updates to the digital twin environment, iterating from S2 to S4 to achieve overall collaborative evolution of the system. As the system's central nervous system and metabolic system, it ensures the reliability and smooth flow of information, and drives the entire system to form a self-evaluating, self-improving, and continuously evolving collaborative closed loop.
[0124] Figure 3 This is an electronic device provided in one embodiment of this application. For example... Figure 3 As shown, the electronic device includes at least the following components: processor 301 and memory 300, communication interface 303, and bus 302.
[0125] In this embodiment of the application, memory 300 is used to store executable instructions of processor 301, which, when configured to execute instructions, implements the method as described in the first aspect.
[0126] In embodiments of this application, a computer-readable storage medium includes instructions that instruct a device to perform the method as described in the first aspect. For example, the instructions instruct the device to perform... Figure 1 The method is shown in the process steps.
[0127] In one embodiment of this application, the program operating in the electronic device may be a program that controls a central processing unit (CPU) or similar device to achieve the functions of the above-described embodiments of the present invention (a program that enables the computer to function). Information processed by these systems is then temporarily stored in random access memory (RAM) during processing, and subsequently stored in various ROMs such as read-only memory (FlashROM) and hard disk drives (HDDs), and read, corrected, and written by the CPU as needed.
[0128] It should be noted that a portion of the electronic device described above can also be implemented using a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and the program recorded on the recording medium can be read into the computer and executed.
[0129] It should be noted that the computer mentioned here refers to a computer built into an electronic device, employing hardware including an operating system and peripheral devices. Furthermore, computer-readable recording media refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage systems such as hard drives built into the computer.
[0130] Furthermore, computer-readable recording media can include: media that dynamically stores programs for short periods of time, such as communication lines used when transmitting programs via networks like the Internet or communication lines like telephone lines; and media that store programs for fixed periods of time, such as volatile memory inside a computer that serves as a server or client in this case. In addition, the aforementioned program can be a program used to implement the above-mentioned functions, or it can be a program that can implement the above-mentioned functions by combining them with programs already recorded in the computer.
[0131] Furthermore, the electronic device in the above embodiments can also be implemented as an assembly (system group) composed of multiple systems. Each system constituting the system group can possess some or all of the functions or functional blocks of the electronic device in the above embodiments. As a system group, it is sufficient to have all the functions or functional blocks of the electronic device.
[0132] Those skilled in the art should recognize that the above embodiments are only used to illustrate this application and are not intended to limit this application. Any appropriate changes and variations made to the above embodiments within the essential spirit and scope of this application fall within the scope of protection claimed in this application.
Claims
1. An integrated perception-planning-control UAV airborne edge computing system, characterized in that, Includes the following steps: The cloud brain decision center, deployed on a cloud server cluster, serves as the strategy generation and evolution hub of the system, used to generate and continuously optimize global flight strategies. The edge intelligent collaborative network consists of multiple heterogeneous edge computing nodes deployed at key nodes in the airspace. It serves as the real-time perception and computing network of the system and is used for local airspace threat perception and multi-machine collaborative conflict resolution computing. An airborne autonomous safety agent, deployed on the UAV platform, serves as the execution verification and feedback terminal of the system, used to perform flight missions and transmit performance data back. The cloud brain decision center, the edge intelligent collaborative network, and the airborne autonomous safety entity interact and synchronize data through a decentralized trusted data lake powered by blockchain. The system performs closed-loop iteration and autonomous evolution of global strategies, edge computing models, and airborne decision logic based on a simulation and deduction environment driven by digital twins.
2. The system according to claim 1, characterized in that, The cloud brain decision-making center includes: The global policy generator adopts a hierarchical federated meta-reinforcement learning framework to aggregate multi-source heterogeneous data, generate and continuously optimize a benchmark flight policy knowledge graph covering various airspace structures and mission scenarios; An adaptive policy distributor and orchestrator dynamically orchestrates and distributes differentiated policy fragments to corresponding edge collaboration nodes based on real-time spatial load, network status, and task urgency. The global model evolution engine utilizes feedback data from the edge intelligent collaborative network and the airborne autonomous safety agent, combined with simulation results, to perform structural and parametric co-evolution of the global policy model through evolutionary algorithms and neural network architecture search technology.
3. The system according to claim 1, characterized in that, Each heterogeneous edge computing node in the edge intelligent collaborative network includes: The multi-source threat cognition engine employs a combination of multimodal pre-trained models and spatiotemporal causal reasoning, integrating multimodal data from images, radar, and text to achieve causal prediction and interpretable attribution of dynamic threats. The group collaborative decision-making unit adopts a distributed collaborative algorithm based on graph neural network and game theory. When multi-machine conflict or complex airspace congestion is detected, it generates a distributed relief solution that satisfies the group's optimality without relying on central coordination among nodes. The personalized model distillation unit receives a general model from the cloud brain decision center and uses historical data and real-time feedback from its local area to generate a lightweight, high-precision personalized edge model through online knowledge distillation technology.
4. The system according to claim 1, characterized in that, The airborne autonomous safety system includes: The Trusted Fusion and Security Shield module integrates runtime verification and a trusted execution environment to perform millisecond-level formal verification of external policies and local decisions; it also employs an adaptive confidence propagation network to dynamically evaluate the trustworthiness of each information source. The hierarchical-reflective adaptive controller comprises a task planning layer, a behavior layer, and a reflection layer. The task planning layer executes high-level task sequences. The behavior layer generates smooth trajectories based on model predictive control. The reflection layer is a high-speed reaction loop independent of the upper layers, based on a predefined safety primitive library, to make sub-second-level instinctive avoidance responses to sudden and unforeseen threats. The performance acquisition and causal feedback unit collects flight status data and constructs a flight decision-outcome causal graph, analyzes the root causes of strategy advantages and disadvantages, and generates a structured performance report for feedback.
5. The system according to claim 3, characterized in that, The group collaborative decision-making unit adopts a credit allocation and reputation mechanism: each edge node obtains corresponding credit points after contributing to the conflict resolution solution, and the points affect its voice in subsequent collaborative decision-making; at the same time, the reputation value of other nodes is maintained to identify and isolate nodes that may provide false information or engage in malicious behavior.
6. The system according to claim 4, characterized in that, The primitives in the security primitive library of the reflection layer are generated and optimized in the digital twin environment through adversarial reinforcement learning, and their triggering conditions and response actions are encoded into verifiable formal rules.
7. An intelligent flight management method based on digital twins and collaborative evolution, characterized in that, Performed by the intelligent flight management system according to any one of claims 1-6, the method includes the following steps: S1. Digital Twin Environment Construction and Initialization Steps: Based on high-precision geographic information, airspace rules, and equipment models, construct a full-scale digital twin airspace environment synchronized with the physical system; S2. Cloud-based global policy co-evolution steps: In the twin environment, a hierarchical federated meta-reinforcement learning model is used to train the global policy model; based on feedback from the physical world and the twin environment, an evolutionary algorithm is used to continuously optimize the model architecture and parameters to generate a policy knowledge graph; S3. Edge Distributed Perception and Collaborative Decision-Making Steps: Edge nodes utilize local multimodal data to recognize threats through causal reasoning; when facing group conflicts, they generate collaborative resolution strategies online through distributed consensus algorithms and game equilibrium solutions; and adapt cloud-based general models to local personalized models through online knowledge distillation. S4. Airborne Trusted Execution and Reflection Control Steps: The airborne terminal performs runtime verification of the received policy within the trusted execution environment; it executes tasks using a layered control architecture and handles extreme contingencies through an independent reflection layer; it collects and transmits flight causal data back. S5. Cross-layer data synchronization and system evolution steps: Ensure the consistency and reliability of all interactive data and model updates through a blockchain data lake; use the returned causal feedback data to drive the update of the digital twin environment and iterate from S2 to S4 to achieve the overall collaborative evolution of the system.
8. The method according to claim 7, characterized in that, The distributed consensus algorithm in S3 is an air traffic consensus protocol based on an improved practical Byzantine fault-tolerant algorithm. This protocol uses a relief scheme as a proposal, reaches consensus through multiple rounds of voting and proof, and tolerates no more than one-third of node failures or malicious behavior.
9. The method according to claim 7, characterized in that, The runtime verification in S4 specifically involves: encoding the flight safety protocol into a time-series logic formula, converting the flight strategy or control command to be verified into a hybrid automaton model, completing the reachability analysis within a limited time using model verification technology, and outputting a verification certificate or counterexample path.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the intelligent flight management method based on digital twins and co-evolution as described in any one of claims 7-9.