Low-altitude traffic flow cooperative perception and conflict prediction method and system based on multi-modal large model
By fusing multi-source data into a multimodal large model and incorporating aircraft dynamics and airspace rules, combined with an embodied agent large model and game theory, the problem of multi-source data fusion and prediction of complex interactive behaviors in low-altitude traffic flow was solved, achieving high-precision and robust conflict early warning and avoidance strategy generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-15
- Publication Date
- 2026-06-16
AI Technical Summary
Existing low-altitude traffic flow conflict prediction methods struggle to effectively integrate multi-source heterogeneous data, understand complex interactive behaviors, and navigate highly uncertain environments. This results in inaccurate predictions and insufficient robustness, failing to meet the safety requirements of high-density, highly dynamic low-altitude traffic flows.
Data fusion is performed using a multimodal large model, incorporating aircraft dynamics constraints and airspace rules, fine-tuning instructions through an embodied intelligent agent large model, and risk assessment is conducted using incomplete information game theory to generate a collaborative avoidance strategy.
It achieves high-precision and robust low-altitude traffic flow collaborative perception and conflict prediction, reduces false alarm rate, improves the accuracy and timeliness of early warning, and supports real-time collision avoidance decision-making.
Smart Images

Figure CN121725677B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary fields of intelligent air traffic management, artificial intelligence and operations research, and specifically relates to a method and system for collaborative perception and conflict prediction of low-altitude traffic flow based on a multimodal large model. Background Technology
[0002] With the rapid development of the low-altitude economy, including urban air mobility (UAM) and logistics drones, low-altitude airspace is becoming increasingly congested and complex. The surge in the number of aircraft, their diverse operating modes (such as vertical takeoff and landing, high-speed cruise), the dynamic changes in the airspace environment (such as buildings and weather interference), and the mixing of humans and aircraft pose serious challenges to low-altitude traffic flow safety. Traditional traffic conflict prediction methods mainly rely on rule-based systems or simple statistical models, which struggle to effectively handle the fusion of multi-source heterogeneous data, the understanding of complex interactive behaviors, and real-time risk assessment under high uncertainty environments.
[0003] Currently, deep learning-based trajectory prediction methods have improved prediction accuracy to some extent, but they still have significant limitations: First, existing methods mostly rely on single or limited data sources, lacking effective fusion and collaborative utilization of multimodal data such as visual, radar, flight dynamics, meteorological, and flight planning data, making it difficult to build a comprehensive and unified perception of traffic conditions. Second, most models lack an inherent understanding of aircraft physical dynamic constraints and airspace operation rules, leading to predicted trajectories that often violate physical common sense or airspace regulations, affecting the rationality and credibility of the prediction results. Third, existing risk assessments are mostly based on simple spatiotemporal proximity indicators, failing to fully consider strategic interactions and game-theoretic behaviors between aircraft, easily resulting in false alarms or missed alarms in dense, dynamic scenarios, and the accuracy and timeliness of early warnings need to be improved. Furthermore, existing models generally lack robustness in the face of sensor noise, data contamination, and even malicious adversarial interference.
[0004] At a deeper level, the existing technological system suffers from three major technical bottlenecks: First, the contradiction between the fixed nature of perception representation and the dynamic nature of the scene. Existing fusion methods mostly use spatiotemporal grids with fixed granularity or rules, which wastes computational resources when traffic is sparse, and fails to capture subtle risks in conflict hotspots due to insufficient accuracy, making it impossible to achieve an adaptive balance between accuracy and efficiency. Second, the contradiction between the openness of model prediction and the closed nature of domain common sense. Mainstream prediction models learn statistical laws from data, lacking a fundamental understanding of the physical performance of aircraft (such as maneuver envelope) and airspace management rules (such as geofencing), resulting in prediction results that often violate physical common sense or operational norms, requiring complex post-processing for patching corrections, a lengthy process with poor reliability. Third, the contradiction between the static nature of risk assessment and the dynamic nature of behavioral interaction. Existing assessment methods are essentially static calculations based on geometry or probability, treating aircraft as unconscious moving points, completely ignoring their intelligent game-theoretic behavior of actively avoiding conflicts based on their own intentions and capabilities. This has led to the system facing the core challenges of high false alarm rates and poor interpretability in its early warning of conflicts involving high-density, highly interactive low-altitude traffic flows.
[0005] Therefore, there is an urgent need for a new method for collaborative perception and conflict prediction of low-altitude traffic flow that can deeply integrate multi-source information, internalize physical and rule-based common sense, simulate intelligent agent game interaction, and maintain stability under complex interference, so as to support the safe and efficient operation of high-density and high-dynamic low-altitude traffic flow in the future. Summary of the Invention
[0006] To overcome the shortcomings of existing technologies, this invention provides a method and system for collaborative perception and conflict prediction of low-altitude traffic flow based on a multimodal large model, aiming to achieve intelligent prediction and early warning of conflict risks with high accuracy, high robustness and in accordance with physical and rule common sense.
[0007] In a first aspect, embodiments of this application provide a method for collaborative perception and conflict prediction of low-altitude traffic flow based on a multimodal large model, the method comprising:
[0008] S1. Collect multi-source heterogeneous sensing data and flight plan data of aircraft in low-altitude environment, and based on real-time traffic situation, fuse and map the data into a spatiotemporal grid with dynamically adjustable granularity to generate unified spatiotemporal fusion features and complete the spatiotemporal adaptive representation of multimodal data.
[0009] S2. Based on a pre-trained embodied agent large model as the base model, fine-tuning is performed using a low-altitude traffic flow dataset containing real-world scenarios, adversarial disturbance samples, and anomalous samples. During the fine-tuning process, aircraft dynamics constraints and airspace rules are introduced as model priors, and a loss function containing adversarial regularization terms is used for optimization to obtain a dedicated prediction model suitable for low-altitude traffic flow scenarios. The embodied agent large model is a visual language action model based on VIMA or RT2 architecture.
[0010] S3. Input the spatiotemporal fusion features into the dedicated prediction model; the dedicated prediction model performs integrated reasoning on the input features and simultaneously outputs the aircraft's behavioral intention recognition results and its future multi-step probability trajectory prediction; based on the probability trajectory prediction, construct an incomplete information game inference model to simulate the strategy interaction between aircraft, and then calculate a comprehensive collision risk index that integrates trajectory uncertainty, relative motion state and avoidance strategy derived from game inference.
[0011] S4. Based on the comprehensive collision risk index and combined with the preset dynamic safety threshold, generate graded conflict warning information and corresponding cooperative avoidance strategy suggestions.
[0012] S5. Distribute the graded conflict warning information and cooperative avoidance strategy suggestions to the relevant aircraft or air traffic management systems through the communication network to support real-time collision avoidance decisions.
[0013] Secondly, embodiments of this application provide a low-altitude traffic flow collaborative perception and conflict prediction system based on a multimodal large model, applied to the method described in the first aspect, the system comprising:
[0014] The multi-source data fusion processing module is used to collect multi-source heterogeneous sensing data and flight plan data of aircraft in the low-altitude environment, and based on the real-time traffic situation, fuse the data and map it into a spatiotemporal grid with dynamically adjustable granularity to generate unified spatiotemporal fusion features and complete the spatiotemporal adaptive representation of multimodal data.
[0015] A dedicated prediction model module is used to fine-tune a low-altitude traffic flow dataset containing real-world scenarios, adversarial disturbance samples, and anomalous samples, based on a pre-trained embodied agent large model. During fine-tuning, aircraft dynamics constraints and airspace rules are introduced as model priors, and a loss function with adversarial regularization is used for optimization, resulting in a dedicated prediction model suitable for low-altitude traffic flow scenarios. The embodied agent large model is a visual-language-action model based on VIMA or RT2 architecture.
[0016] An integrated reasoning and risk assessment module is used to input the spatiotemporal fusion features into the dedicated prediction model; the dedicated prediction model performs integrated reasoning on the input features and simultaneously outputs the aircraft's behavioral intent recognition results and its future multi-step probability trajectory prediction; based on the probability trajectory prediction, an incomplete information game inference model is constructed to simulate the strategy interaction between aircraft, and then calculates a comprehensive collision risk index that integrates trajectory uncertainty, relative motion state and avoidance strategy derived from game inference;
[0017] The intelligent early warning and strategy generation module is used to generate graded conflict early warning information and corresponding cooperative avoidance strategy suggestions based on the comprehensive collision risk index and a preset dynamic safety threshold.
[0018] The early warning information distribution and communication module is used to distribute the graded conflict early warning information and cooperative avoidance strategy suggestions to relevant aircraft or air traffic management systems through a communication network to support real-time collision avoidance decisions.
[0019] Thirdly, embodiments of this application provide an electronic device, including:
[0020] processor;
[0021] Memory used to store processor-executable instructions;
[0022] The processor is configured to implement the low-altitude traffic flow cooperative perception and conflict prediction method based on a multimodal large model as described in the first aspect when executing the instructions.
[0023] Fourthly, embodiments of this application provide a computer-readable storage medium storing a program that instructs a device to execute the low-altitude traffic flow cooperative perception and conflict prediction method based on a multimodal large model as described in the first aspect.
[0024] This invention addresses the aforementioned technical bottlenecks by proposing a systematic solution. Its core innovations are reflected in three levels: First, it designs an information-driven, resource-constrained spatiotemporal adaptive representation mechanism, resolving the contradiction between the difficulty of balancing perception accuracy and computational efficiency. Second, it pioneers a method for constructing a domain-knowledge-based intrinsic security model based on differentiable fusion, deeply encoding physical laws and operational rules into the model's inference layer, surpassing traditional post-processing correction methods. Finally, it constructs a game-theoretic-driven dynamic intelligent risk assessment framework, treating the aircraft as a rational intelligent agent and assessing residual risks by simulating its future interaction strategies, achieving a leap from static collision detection to dynamic conflict prediction.
[0025] Compared with existing technologies, this invention has the following significant advantages: Through spatiotemporal adaptive fusion representation driven by maximizing the information entropy of multimodal data, it achieves comprehensive, unified, and computationally optimal perception of low-altitude traffic flow patterns with limited computing resources. Compared with fixed-grid methods, this method significantly reduces overall system resource consumption while ensuring high accuracy in key areas, laying the foundation for large-scale real-time processing. Utilizing embodied agent large-model fine-tuning technology, through differentiable physical constraint injection and rule knowledge distillation, domain common sense is internalized into model parameters, giving the prediction model physical intuition and rule awareness. This fundamentally avoids the problem of predicted trajectories violating dynamics or airspace regulations, achieving intrinsically safe predictions without relying on unreliable post-processing corrections. Introducing incomplete information game theory to simulate the strategic interactions between aircraft, this makes risk assessment not only based on the probability overlap of trajectories but also on the proactive deduction of agent avoidance strategies. This ensures that the warning signal reflects the residual risk after considering intelligent avoidance behavior, rather than a simple worst-case scenario, thereby greatly reducing the false alarm rate in dense dynamic scenarios and significantly improving the accuracy, timeliness, and decision interpretability of warnings. The generated avoidance strategy is both collaborative and executable. By distributing the strategy through a communication network and monitoring execution feedback, a complete intelligent closed loop is achieved, encompassing perception, prediction, evaluation, decision-making, execution, and continuous model optimization, significantly enhancing the system's adaptability and long-term practical value. Attached Figure Description
[0026] Figure 1 This is a schematic diagram of a low-altitude traffic flow collaborative perception and conflict prediction method based on a multimodal large model, provided as an embodiment of this application.
[0027] Figure 2 The system architecture diagram of low-altitude traffic flow collaborative perception and conflict prediction based on a multimodal large model provided in this application.
[0028] Figure 3 A schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.
[0030] It should be noted that in the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.
[0031] Based on the embodiments described in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0032] Example 1
[0033] Figure 1 This is a schematic flowchart illustrating a low-altitude traffic flow collaborative sensing and conflict prediction method based on a multimodal large model, provided as an embodiment of this application. Figure 1 As shown, a low-altitude traffic flow collaborative sensing and conflict prediction method based on a multimodal large model includes:
[0034] S1. Collect multi-source heterogeneous sensing data and flight plan data of aircraft in the low-altitude environment, and based on real-time traffic conditions, fuse and map the data into a spatiotemporal grid with dynamically adjustable granularity to generate unified spatiotemporal fusion features, completing the spatiotemporal adaptive representation of multimodal data. This transforms the raw, chaotic multi-source data into machine-understandable, spatiotemporally aligned, unified structured features. This is the sensing and preprocessing stage of the entire system, solving the problems of data heterogeneity and spatiotemporal inconsistency, and providing high-quality input for subsequent intelligent analysis.
[0035] Specifically, in this embodiment, step S1 includes:
[0036] The system collects multimodal data in real time from UAVs, manned aircraft, ground radar or visual sensors, meteorological monitoring stations, and flight planning systems. This data includes at least aircraft trajectory, speed, images, point clouds, meteorological information, and structured flight intentions. Specifically, the system collects multimodal data in real time through a distributed node network: acquiring ADS-B broadcast signals and airborne sensor data from UAVs and manned aircraft; collecting aircraft trajectory, velocity vectors, and optical images through ground radar arrays and visual sensor networks; connecting to meteorological monitoring stations to obtain real-time meteorological parameters such as wind speed and visibility; and accessing the flight planning system to parse structured flight intentions. All data is timestamped and aligned with a coordinate system to form a multi-dimensional data stream containing trajectory, speed, images, point clouds, meteorological text, and flight plans.
[0037] Based on real-time monitoring of traffic flow density and conflict hotspot distribution, the granularity of an adaptive spatiotemporal grid is dynamically adjusted. A fine-grained grid is used in identified high-risk areas or conflict hotspots, while a coarse-grained grid is used in sparsely trafficked areas. In the data processing unit, the system continuously analyzes traffic flow density distribution and identifies potential hotspot areas using a real-time conflict detection algorithm. Based on the analysis results, the system dynamically adjusts the adaptive spatiotemporal grid partitioning strategy: a fine-grained grid of 5m × 5m × 1 second is used in high-risk areas such as airport landing corridors and areas surrounding high-rise buildings; while a coarse-grained grid of 20m × 20m × 3 seconds is used in sparsely trafficked areas such as suburban airspace. This dynamic adjustment mechanism allows the system to effectively reduce the overall computational load while maintaining high accuracy in key areas.
[0038] The collected multimodal data are uniformly mapped and fused into the corresponding cells of the adaptive dynamic spatiotemporal grid according to their spatiotemporal attributes, generating a spatiotemporally aligned multimodal feature map that balances representation accuracy and computational load, serving as the spatiotemporal fusion feature. Various preprocessed data types are mapped to the corresponding spatiotemporal cells of the dynamic grid: radar point cloud data is converted into occupancy probabilities for grid cells; feature vectors extracted from visual images are filled into corresponding spatial cells; flight trajectory data is embedded in the time dimension as state vectors; and meteorological information is integrated into the entire grid as environmental parameters. Through a feature-level fusion algorithm, a 128-dimensional multimodal feature map is finally generated. This feature map retains the physical meaning of each data source and forms a unified spatiotemporally aligned representation, which can be directly input into downstream neural network models.
[0039] In actual deployment, the above three steps are executed cyclically at a frequency of 10Hz: the data acquisition module continuously receives streaming data from each sensor, the grid management unit re-evaluates and updates the grid partitioning strategy every 30 seconds, and the feature fusion engine generates feature maps in real time in a pipeline manner. This architecture enables the system to simultaneously handle collaborative perception tasks of more than 200 flying targets, providing a highly timely environmental representation foundation for subsequent prediction and decision-making.
[0040] Specifically, an optimization model for multimodal spatiotemporal adaptive representation can be constructed, modeling dynamic mesh partitioning as a constrained multi-objective optimization problem to maximize the information entropy of the scene representation under limited computational resources. Let the entire spatiotemporal region be... Discretize it into a grid set Each grid ( (For indexing) with variable spatiotemporal resolution Define the information density function. and calculate cost function :
[0041] ;
[0042] ;
[0043] The optimization objective is to find the optimal resolution field. To maximize overall information gains and minimize costs:
[0044] ;
[0045] The constraints are: , ,in, For resolution field, It is a mapping function from resolution to information capture efficiency (usually a decreasing function). Indicates position In time Traffic flow density; The variance of the velocity field characterizes the degree of motion disorder; Indicator function (0 or 1) representing conflict hotspot areas; , The weighting coefficients for each item are obtained through domain knowledge or learning. This indicates the data communication energy consumption of this grid cell; A parameter representing the trade-off between information benefits and computational costs; Represents the total computing / communication budget; Indicates the upper and lower limits of the resolution. This represents the coordinate change. The real-time solution to this optimization problem drives the dynamic adaptive adjustment of the mesh granularity.
[0046] S2. Based on a pre-trained embodied intelligent agent large model as the base model, fine-tuning is performed using a low-altitude traffic flow dataset containing real-world scenarios, adversarial disturbance samples, and anomalous samples. During fine-tuning, aircraft dynamics constraints and airspace rules are introduced as model priors, and a loss function including adversarial regularization is used for optimization, resulting in a specialized prediction model suitable for low-altitude traffic flow scenarios. The embodied intelligent agent large model is a visual-language-action model based on VIMA or RT2 architecture. A general artificial intelligence model is customized into an expert model specifically for low-altitude traffic flow scenarios—one that understands the rules and is resistant to interference—through domain knowledge injection and robust training. This is the process of forging the system's brain or decision-making core.
[0047] Specifically, in this embodiment, step S2 includes:
[0048] S2.1 Obtain a large-scale embodied agent model pre-trained on a general multimodal task as the base model. This base model possesses basic understanding and reasoning capabilities regarding the state of the physical world. In practice, a VIMA architecture or RT2 architecture embodied agent model pre-trained on a general robot operation and navigation task can be selected as the base model. This model has been trained on massive amounts of simulation and real data, including visual observation, language commands, and action sequences. It possesses basic understanding capabilities regarding the physical properties of objects, spatial relationships, and causal effects. Its multimodal encoder can simultaneously process image features and text commands, forming an ideal base adapted to low-altitude traffic flow scenarios.
[0049] S2.2 In the fine-tuning phase, a simplified dynamic model of the aircraft is injected as a priori constraints into the base model, enabling the model to simulate and comply with the physical limitations and maneuverability of the aircraft during its internal reasoning process. In the fine-tuning preparation phase, a simplified dynamic model containing limits on maximum acceleration, turning radius, and rate of climb is established based on typical parameters of quadcopter UAVs and small electric vertical takeoff and landing (eVTOL) aircraft. This model is transformed into a differentiable constraint module and fused with the model's original decision network through an attention mechanism. For example, by designing a differentiable projection layer or using it as a regularization term in the loss function, the trajectory predicted by the model satisfies the dynamic constraints. This ensures that the implicit state space of the model naturally satisfies physical feasibility conditions when predicting trajectories, such as automatically excluding sharp turn predictions exceeding the maximum overload.
[0050] The simplified dynamics model is encoded as a set of differentiable state transition inequalities. At the model architecture level, we achieve fusion by introducing a physical constraint adaptation layer: this layer receives feature representations from the intermediate layers of the model and outputs a correction vector designed to project the predicted next-time state into a dynamically feasible state space. Specifically, the output of the FFN (feedforward network) layer used to predict actions in the base model's Transformer decoder is added to the correction vector generated by the physical constraint adaptation layer, and then input into a Softmax layer to produce the final action distribution. The physical constraint adaptation layer itself is a small neural network whose input is the current state features, and its weights are updated along with the base model during fine-tuning.
[0051] A Bayesian framework for injecting physical constraints into a large-scale embodied intelligent agent model is constructed. Aircraft dynamics constraints are treated as prior knowledge and injected into the pre-trained large-scale model through Bayesian inference. Let the observed data be... The model parameters are Standard fine-tuning is maximizing likelihood. Introducing a physical consistency prior Simplified dynamic model encoding:
[0052] ,
[0053] in It is the trajectory prediction function of the model. It is to the feasible trajectory space The projection operator for the space defined by differential constraints:
[0054] ,
[0055] The goal of model optimization is to find the optimal parameters. This parameter can simultaneously fit the data, obey physical laws, and internalize the rules. The new posterior optimization objective is:
[0056] ,
[0057] This is equivalent to maximizing the posterior probability of the parameters. In this context, physics and rules serve as prior knowledge. It is the weight and bias parameter vector of the neural network. This is the training dataset used for fine-tuning. It is the first Multimodal observation inputs for each sample (such as images, radar point clouds, historical trajectories, etc.). These are the corresponding labels for the actual future trajectory. It encodes the prior probability distribution of the aircraft's dynamic feasibility; it is not a fixed function, but rather a distribution with respect to parameters. Probability assessment. If a set of parameters If the model can generate predictions that mostly conform to physical laws, then... The value is higher. It is a priori spatial rule, achieved through knowledge distillation. It is an aircraft In time The observation, These are the maximum acceleration, velocity, and curvature constraints. These are the strength coefficients of physical priors and rule priors. This is the set of state constraints.
[0058] S2.3. Using a small sample dataset containing real-world scenarios, adversarial perturbation samples, and anomalous samples, the model with injected physical constraints is fine-tuned. This fine-tuning process is achieved by minimizing a composite loss function, which includes: a task loss to optimize trajectory prediction accuracy, a knowledge distillation loss to inject geofencing and air corridor rules, and an adversarial regularization loss term specifically designed to improve the model's robustness against perturbations. For example, a small sample dataset is constructed containing 5000 sets of real-world urban low-altitude flight trajectories and 2000 sets of perturbation samples generated through adversarial attacks (such as adding sensor noise or simulating GPS spoofing). A three-stage learning strategy is adopted: first, pre-training on simulation data without adversarial perturbation samples and anomalous samples to achieve stable convergence; then, an adversarial regularization term is introduced, defined by calculating the norm of the difference between the model's output on clean samples and adversarial perturbation and anomalous samples, to enhance model robustness; finally, the task loss composed of mean squared error and the knowledge distillation loss supervised by soft labels output by the teacher model (with encoded geofencing rules) are jointly optimized. The entire training process lasted 72 hours on 8 GPUs, and the resulting dedicated prediction model significantly improved robustness against common adversarial perturbations while maintaining high-precision trajectory prediction capabilities.
[0059] Furthermore, in step S2.3, the process of constructing the small sample dataset and fine-tuning the model further includes:
[0060] Based on the simplified dynamics model and preset airspace rules, a large amount of synthetic trajectory data covering various conflict scenarios, extreme weather, and sensor failures is generated through simulation. For example, based on the quadcopter UAV dynamics model and urban airspace management rules, a synthetic dataset containing 100,000 flight trajectories was constructed in the simulation platform. This dataset specifically simulates various conflict scenarios such as head-on approach at intersecting airways, entry at merging points, and hovering at geofence boundaries, and incorporates extreme weather conditions such as strong crosswinds and low visibility, as well as failure modes such as GPS signal drift and temporary visual sensor malfunctions, ensuring that the data covers the main risk dimensions of low-altitude operations.
[0061] A multi-stage course learning fine-tuning process is designed. First, the synthetic data is used to pre-train the model, enabling it to initially grasp the general patterns and rules of low-altitude traffic flow. Then, a small sample dataset containing real, adversarial disturbance, and anomalous samples is used for fine-tuning, focusing on optimizing the model's performance under real noise distributions and its robustness to adversarial and anomalous samples. Specifically, a two-stage progressive training strategy is implemented: In the first stage, the aforementioned synthetic data is used to pre-train the base model for 48 hours. During this stage, the model quickly grasps the basic motion patterns of aircraft, airspace structural constraints, and typical conflict patterns, effectively improving trajectory prediction accuracy on the synthetic test set. In the second stage, the model switches to a refined dataset containing 3500 sets of real-world trajectories and 1500 sets of adversarial and anomalous samples. During training, the weights of the loss function are dynamically adjusted, increasing the adversarial regularization loss weight to three times that of the initial stage, forcing the model to enhance its ability to distinguish data disturbances and malicious interference. After two stages of training, the model's overall performance on the real-world test set is significantly improved compared to the baseline model trained directly with real data.
[0062] S3. Input the spatiotemporal fusion features into the dedicated prediction model; the dedicated prediction model performs integrated reasoning on the input features, and simultaneously outputs the aircraft's behavioral intent recognition results and its future multi-step probabilistic trajectory predictions; based on the probabilistic trajectory predictions, an incomplete information game theory model is constructed to simulate the strategic interaction between aircraft, and then calculates a comprehensive collision risk index that integrates trajectory uncertainty, relative motion state, and avoidance strategies derived from game theory. Using a customized expert model, a deep understanding of the traffic situation is achieved, future projections are made, and potential collision risks are intelligently assessed. The behavioral intent recognition results can be output as a probability distribution, such as {Maintain course: 0.7, Prepare for left turn: 0.25, Emergency climb: 0.05}. The future multi-step probabilistic trajectory is output as a set of weighted trajectory samples or a parameterized probability distribution (such as a multivariate Gaussian distribution). The incomplete information game inference model treats each aircraft as a rational agent with a strategy set containing a limited set of typical evasive maneuvers (such as acceleration / deceleration, left / right turns, climb / descent), an information set including its incomplete observations of the intentions and states of other aircraft in the vicinity, and a payoff function that integrates safety, efficiency, and rule compliance.
[0063] Specifically, in this embodiment, step S3, where the dedicated prediction model completes the task in an end-to-end manner, includes:
[0064] Based on its internally integrated aircraft dynamics model, the system identifies the target's current maneuvering behavior and infers its short-term flight intentions under physical constraints, with a focus on identifying unconventional or undeclared non-cooperative flight behaviors. A dedicated predictive model analyzes the target's kinematic parameters in real time through its internally integrated dynamics priors. When the model detects a drone continuously climbing at an acceleration exceeding its maximum rate of climb, or erratically loitering at the edge of a no-fly zone, it marks it as a potential non-cooperative target and combines this with flight plan data for intent matching, generating high-risk flags for undeclared course changes or altitude intrusions.
[0065] When predicting the future multi-step trajectories of aircraft, the model explicitly models the interactions between aircraft and their interactions with environmental constraints, outputting the probabilistic trajectory predictions. During the prediction phase, the model employs an attention-based interactive encoder to explicitly calculate the relative position and velocity vectors of the target aircraft with all its neighbors, as well as its distance to the nearest geofence boundary. For example, when predicting the trajectory of a drone approaching an airport boundary, the model simultaneously considers the potential avoidance effects between the drone and another aircraft taking off or landing, as well as the constraints imposed by airport boundary rules, ultimately outputting multiple probabilistic trajectory distributions at 100-millisecond intervals over the next 30 seconds.
[0066] Based on the predicted probabilistic trajectories, the game theory-driven risk assessment is performed. Based on these probabilistic trajectories, the system constructs an incomplete information game theory model with each drone as an agent. When assessing potential collisions between two drones, the model not only calculates the probability of their trajectories overlapping but also infers the avoidance strategies (such as left turn, climb, deceleration) and their corresponding probabilities, while considering incomplete information factors such as communication delays and sensor field of view. The final risk indicator is a comprehensive function of the spatiotemporal overlap probability, the probability of successful avoidance by each party, and the relative velocity vector hazard coefficient. In a test scenario, this method significantly reduces the false alarm rate of conflict warnings compared to the traditional nearest-distance threshold method.
[0067] Specifically, in this embodiment, the specific process of calculating the comprehensive collision risk index in step S3 includes:
[0068] For any pair of aircraft, based on their predicted probability trajectory distributions, the joint probability of spatial overlap at each future time step is calculated as the basic spatiotemporal conflict potential energy. In the embodiment, for two aircraft (aircraft A and B) predicted to intersect within the next 10 seconds, the system extracts their predicted position probability distributions (represented by a Gaussian mixture model) at each 100-millisecond interval and calculates the joint probability of their spatial volume overlapping at each time step. By integrating the cumulative amount of these instantaneous overlap probabilities over the time dimension, a scalar value is obtained as the basic conflict potential energy. For example, if the calculation results show that the overlap probability reaches a peak of 0.15 at t=5.2 seconds, the basic potential energy may accumulate to a value of 3.7.
[0069] An incomplete information game simulation model is constructed to simulate the possible avoidance actions and probabilities of each party based on its physical capabilities, observational information, and pre-set strategies in this conflict scenario. Based on the simulation results, a game correction factor is calculated to dynamically adjust the basic conflict potential energy. For this potential conflict, the system constructs a two-stage incomplete information game: the first stage simulates the updating of each party's beliefs about each other's intentions, and the second stage simulates avoidance strategies. The model considers the maximum turning capabilities of A and B, the currently observable history of the opponent's heading angle changes, and the pre-set avoidance rule priority (e.g., right-hand priority). After simulation, if the calculated probability of A actively turning right to avoid the collision is 0.8, and the probability of B maintaining its heading is 0.7, then based on the degree of reduction in the actual collision probability under this strategy combination, the game correction factor is calculated to be 0.25 (indicating that through reasonable game theory, the original risk can be reduced to one-quarter).
[0070] The incomplete information game deduction model employs a multi-agent policy search algorithm based on deep recursive belief updates for online deduction. The payoff function for each agent (aircraft) is as follows: Designed as follows: ,in These are preset positive weighting coefficients, corresponding to preferences for safety, energy consumption, rule compliance, and task progress, respectively. This is the sum of the self-collision risk and the other-collision risk calculated based on the aforementioned probability trajectory; Related to the magnitude of maneuvering actions (such as changes in acceleration and steering angle); It represents the degree of deviation from the planned flight path or entry into the no-fly zone, and is calculated based on the deviation of the aircraft's status from the geofence and the planned flight path; Encourage the aircraft to proceed toward the target point.
[0071] The system combines the basic conflict potential energy with the game-theoretic correction factor, and further weights and fuses the uncertainty represented by the covariance of the predicted trajectory and the urgency represented by the relative speed and heading angle of the two aircraft to generate the final comprehensive collision risk index. Finally, the system multiplies the basic conflict potential energy (3.7) with the game-theoretic correction factor (0.25) to obtain a preliminary adjustment value (0.925). Subsequently, an uncertainty weight (the determinant of the trajectory prediction covariance, normalized to 1.2) and an urgency weight (a composite function of the cosine of the difference between the relative speed and the heading angle, calculated to be 1.5) are introduced. Through weighted fusion, a comprehensive collision risk index value located in the standardized interval [0, 100] is finally generated. In this case, the comprehensive index calculated through this process is 41.5, exceeding the set intermediate risk threshold of 30, and the system generates a corresponding warning signal accordingly.
[0072] Specifically, in this embodiment, the output of the dedicated prediction model in step S3 and the calculation results of the comprehensive collision risk index are organized into a structured risk assessment map; the risk assessment map includes at least:
[0073] A spatiotemporal risk heatmap is used to visualize the spatial distribution of the comprehensive collision risk index over multiple future time steps on the adaptive dynamic spatiotemporal grid. In this embodiment, based on the dynamically updated adaptive spatiotemporal grid, the system renders the comprehensive collision risk index value of each grid cell at each moment within the next 5 seconds into a continuously gradient pseudo-color layer using an interpolation algorithm. This sequence of layers over time constitutes a dynamic heatmap, visually displaying the spatiotemporal evolution of high-risk areas. For example, a high-risk hotspot can be clearly seen moving rapidly from the southeast side of the intersection area towards the northwest and gradually intensifying, its color transitioning from yellow to red, corresponding to a risk index increase from 30 to 75.
[0074] The conflict event list lists all predicted potential conflict pairs and associates them with the corresponding comprehensive collision risk index value, the predicted nearest point in time and space, and the main avoidance strategy options derived from the game theory model. Simultaneously, for each potential conflict pair whose risk index exceeds a threshold (e.g., set to 20), the system generates a structured record and adds it to the event list. Each record includes the aircraft IDs of both conflicting aircraft, the current risk index value (e.g., aircraft A and B: 42.5), the predicted nearest approach point (latitude and longitude coordinates, altitude, and timestamp, e.g., predicted to occur at t+3.6 seconds at location [116.4°E, 39.9°N, altitude 150 meters]), and the main avoidance strategy options recommended by the game theory model (e.g., recommending A to turn right and climb, and B to decelerate and maintain its course).
[0075] This structured map (heatmap + event list) is pushed to the air traffic controller's situational awareness interface in real time. Administrators can click on high-risk areas on the heatmap to focus on them, while the event list can be sorted and filtered by risk value, estimated time, or avoidance strategy type. This helps them quickly assess the overall situation, locate key events, and make decisions based on system recommendations. In a real-world stress test, administrators using this map demonstrated faster and more efficient average identification and response times for potential conflicts compared to those using traditional radar trajectory maps.
[0076] S4. Based on the comprehensive collision risk index and a preset dynamic safety threshold, generate graded conflict warning information and corresponding cooperative avoidance strategy suggestions. This transforms abstract collision risk values into intuitive, actionable warning levels and specific avoidance action plans. This is the system's decision output stage, converting the conclusions of intelligent analysis into executable instructions.
[0077] Specifically, in this embodiment, step S4 includes:
[0078] Based on the overall operational status of the current airspace, meteorological conditions, and the current granularity of the adaptive dynamic spatiotemporal grid, multi-level safety thresholds are dynamically set. The comprehensive collision risk index is compared with these multi-level safety thresholds to generate conflict warning levels of varying urgency. In this embodiment, when dynamically setting multi-level safety thresholds, the system comprehensively considers real-time traffic flow density, historical airspace conflict statistics, current visibility and wind speed levels, and risk level directives issued by regulatory agencies. The system incorporates a threshold mapping function. ,in For local flow density, For visibility, For wind speed, This function is designed to monitor risk levels. It is itself a lightweight regression model trained using historical safe operation data and expert rules, capable of outputting high, medium, and low risk thresholds adapted to the current scenario in real time. The early warning system analyzes factors such as airspace traffic density (e.g., number of aircraft per cubic kilometer), visibility, wind speed, and the granularity of the current dynamic spatiotemporal grid (e.g., grid size of 20m / square or 50m / square). Based on these factors, the system dynamically adjusts the safety thresholds: in high-density areas with low visibility and fine-grained grid monitoring, the high-risk warning threshold is set to 60; conversely, in sparse areas with good weather and coarse-grained monitoring, this threshold can be relaxed to 80. After comparing the real-time calculated comprehensive collision risk index (e.g., a risk value of 75 for a conflict pair) with the corresponding threshold, the system automatically maps it to a specific warning level, such as generating an orange warning (corresponding to a high-risk level), and associates this level information with the conflict pair ID.
[0079] For conflicts triggering warnings, at least one set of cooperative avoidance strategy suggestions is generated based on the deduction results of the incomplete information game theory model. These strategy suggestions include at least: adjustments to the heading, speed, and altitude of each relevant aircraft, as well as the priority or timing of these adjustments. For the aforementioned orange warning conflict, the system calls the output of the incomplete information game theory model that pre-simulated the conflict. Based on the optimal or high-probability Nash equilibrium solution from the game theory deduction, a specific cooperative avoidance strategy is generated. For example, for aircraft A, it is suggested that it immediately turn right by 15° and climb 10 meters within 2 seconds; for aircraft B, it is suggested that it maintain its current heading but reduce its speed to 15 meters per second. Simultaneously, the strategy clearly defines priorities and timing, such as indicating the strategy execution priority: A (immediately) > B (after A's action takes effect), thereby ensuring orderly coordination of avoidance actions and preventing new conflicts arising from simultaneous actions.
[0080] S5. The graded conflict warning information and cooperative avoidance strategy suggestions are distributed to relevant aircraft or air traffic management systems via a communication network to support real-time collision avoidance decision-making. Decision instructions are safely, reliably, and promptly transmitted to the entities (aircraft or management personnel) that need to execute the instructions. This is the system's control and communication link, completing the final step from decision-making to execution and connecting to the actual traffic system.
[0081] Specifically, in this embodiment, step S5 includes:
[0082] Based on the conflict warning level, the communication capabilities of the relevant aircraft, and the avoidance responsibility weight inferred from the game theory model, the system adaptively selects communication links and protocols, and determines the priority and timeliness requirements for information distribution. In an embodiment, for a red-level (highest) warning, the system identifies a conflict involving an eVTOL equipped with a 5G ATG communication module and a conventional UAV that only supports narrowband data links. Based on the responsibility weight inferred from the game theory model (the eVTOL is determined to be the avoidance responsible party), the system prioritizes sending a complete instruction containing detailed policies to the eVTOL via the high-bandwidth, low-latency 5G network, while simultaneously sending a concise "maintain current state" advisory message to the UAV via the LoRa wide area network. The system sets the instruction distribution to the eVTOL to the highest priority and a 100-millisecond end-to-end latency requirement, while setting the message distribution to the UAV to a standard priority.
[0083] The tiered conflict warning information and cooperative avoidance strategy suggestions are encapsulated into standardized instructions or consultation messages conforming to the target system interface specifications, and securely transmitted through a communication network with authentication and encryption mechanisms. The system encapsulates the generated warning information (including conflict ID and risk level) and avoidance strategy (a 15-degree left turn and 20% deceleration instruction for eVTOLs) into a binary message format conforming to ASTM F3411 Remote ID and UAS service-specific domain standards. Before transmission, this message is signed with a PKI-based digital certificate and transmitted via a TLS 1.3 encrypted channel to ensure the integrity of the instructions, the trustworthiness of the source, and the confidentiality of transmission.
[0084] After information distribution, the trajectory response status of the relevant aircraft is continuously monitored. The actual response is compared with the proposed strategy, and the comparison result is used as feedback data to continuously optimize the parameters of the dedicated prediction model or the payoff function and strategy prior of the game theory model. After the command is issued, the system continuously monitors the ADS-B and reported status information of the two aircraft at a frequency of 5Hz, extracting their actual heading and speed changes. For example, the system detects that the eVTOL completed a 12-degree heading adjustment and an 18% deceleration within 3 seconds, which is highly consistent with the strategy recommendation; while the UAV's status remains stable. The actual response data of the aircraft in this avoidance event, the environmental conditions (wind speed and traffic density at the time), and the original game theory process are packaged together into an experience tuple and injected into a priority experience replay buffer. During periods of low system computational load, this new experience data will be used to perform small-batch incremental fine-tuning of the dedicated prediction model (especially its trajectory prediction module) and the game theory model (its payoff function), thereby achieving continuous online optimization of model performance.
[0085] Example 2
[0086] like Figure 2 As shown, this application provides an architecture diagram of a low-altitude traffic flow collaborative perception and conflict prediction system based on a multimodal large model, which is applied to the low-altitude traffic flow collaborative perception and conflict prediction system based on a multimodal large model as described in Embodiment 1. It includes: a multi-source data fusion processing module 210, a dedicated prediction model module 220, an integrated reasoning and risk assessment module 230, an intelligent early warning and strategy generation module 240, and an early warning information distribution and communication module 250.
[0087] The multi-source data fusion processing module 210 is used to collect multi-source heterogeneous sensing data and flight plan data of aircraft in the low-altitude environment, and based on the real-time traffic situation, fuse the data and map it into a spatiotemporal grid with dynamically adjustable granularity to generate unified spatiotemporal fusion features and complete the spatiotemporal adaptive representation of multimodal data.
[0088] A dedicated prediction model module 220 is used to fine-tune a low-altitude traffic flow dataset containing real-world scenarios, adversarial disturbance samples, and anomalous samples, based on a pre-trained embodied agent large model as the base model. During the fine-tuning process, aircraft dynamics constraints and airspace rules are introduced as model priors, and a loss function containing adversarial regularization terms is used for optimization to obtain a dedicated prediction model suitable for low-altitude traffic flow scenarios. The embodied agent large model is a visual language action model based on VIMA or RT2 architecture.
[0089] The integrated reasoning and risk assessment module 230 is used to input the spatiotemporal fusion features into the dedicated prediction model; the dedicated prediction model performs integrated reasoning on the input features and simultaneously outputs the behavioral intent recognition results of the aircraft and its future multi-step probability trajectory prediction; based on the probability trajectory prediction, an incomplete information game inference model is constructed to simulate the strategy interaction between aircraft, and then a comprehensive collision risk index that integrates trajectory uncertainty, relative motion state and avoidance strategy derived from game inference is calculated.
[0090] The intelligent early warning and strategy generation module 240 is used to generate graded conflict early warning information and corresponding collaborative avoidance strategy suggestions based on the comprehensive collision risk index and a preset dynamic safety threshold.
[0091] The early warning information distribution and communication module 250 is used to distribute the graded conflict early warning information and cooperative avoidance strategy suggestions to relevant aircraft or air traffic management systems through a communication network to support real-time collision avoidance decisions.
[0092] Figure 3 This is an electronic device provided in one embodiment of this application. For example... Figure 3 As shown, the electronic device includes at least the following components: processor 301 and memory 300, communication interface 303, and bus 302.
[0093] In this embodiment of the application, memory 300 is used to store executable instructions of processor 301, which, when configured to execute instructions, implements the method as described in the first aspect.
[0094] In embodiments of this application, a computer-readable storage medium includes instructions that instruct a device to perform the method as described in the first aspect. For example, the instructions instruct the device to perform... Figure 1 The method is shown in the process steps.
[0095] In one embodiment of this application, the program operating in the electronic device may be a program that controls a central processing unit (CPU) or similar device to achieve the functions of the above-described embodiments of the present invention (a program that enables the computer to function). Information processed by these systems is then temporarily stored in random access memory (RAM) during processing, and subsequently stored in various ROMs such as read-only memory (FlashROM) and hard disk drives (HDDs), and read, corrected, and written by the CPU as needed.
[0096] It should be noted that a portion of the electronic device described above can also be implemented using a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and the program recorded on the recording medium can be read into the computer and executed.
[0097] It should be noted that the computer mentioned here refers to a computer built into an electronic device, employing hardware including an operating system and peripheral devices. Furthermore, computer-readable recording media refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage systems such as hard drives built into the computer.
[0098] Furthermore, computer-readable recording media can include: media that dynamically stores programs for short periods of time, such as communication lines used when transmitting programs via networks like the Internet or communication lines like telephone lines; and media that store programs for fixed periods of time, such as volatile memory inside a computer that serves as a server or client in this case. In addition, the aforementioned program can be a program used to implement the above-mentioned functions, or it can be a program that can implement the above-mentioned functions by combining them with programs already recorded in the computer.
[0099] Furthermore, the electronic device in the above embodiments can also be implemented as an assembly (system group) composed of multiple systems. Each system constituting the system group can possess some or all of the functions or functional blocks of the electronic device in the above embodiments. As a system group, it is sufficient to have all the functions or functional blocks of the electronic device.
[0100] Those skilled in the art should recognize that the above embodiments are only used to illustrate this application and are not intended to limit this application. Any appropriate changes and variations made to the above embodiments within the essential spirit and scope of this application fall within the scope of protection claimed in this application.
Claims
1. A method for collaborative perception and conflict prediction of low-altitude traffic flow based on a multimodal large model, characterized in that, Includes the following steps: S1. Collect multi-source heterogeneous sensing data and flight plan data of aircraft in low-altitude environment, and based on real-time traffic situation, fuse and map the data into a spatiotemporal grid with dynamically adjustable granularity to generate unified spatiotemporal fusion features and complete the spatiotemporal adaptive representation of multimodal data. S2. Based on a pre-trained embodied agent large model as the base model, fine-tuning is performed using a low-altitude traffic flow dataset containing real-world scenarios, adversarial disturbance samples, and anomalous samples. During the fine-tuning process, aircraft dynamics constraints and airspace rules are introduced as model priors, and a loss function containing adversarial regularization terms is used for optimization to obtain a dedicated prediction model suitable for low-altitude traffic flow scenarios. The embodied agent large model is a visual language action model based on VIMA or RT2 architecture. S3. Input the spatiotemporal fusion features into the dedicated prediction model; the dedicated prediction model performs integrated reasoning on the input features and simultaneously outputs the aircraft's behavioral intention recognition results and its future multi-step probability trajectory prediction; based on the probability trajectory prediction, construct an incomplete information game inference model to simulate the strategy interaction between aircraft, and then calculate a comprehensive collision risk index that integrates trajectory uncertainty, relative motion state and avoidance strategy derived from game inference. S4. Based on the comprehensive collision risk index and combined with the preset dynamic safety threshold, generate graded conflict warning information and corresponding cooperative avoidance strategy suggestions. S5. Distribute the graded conflict warning information and cooperative avoidance strategy suggestions to the relevant aircraft or air traffic management systems through the communication network to support real-time collision avoidance decisions.
2. The method according to claim 1, characterized in that, Step S1 specifically includes: Real-time acquisition of multimodal data from UAVs, manned aircraft, ground radar or visual sensors, meteorological monitoring stations, and flight planning systems, including at least aircraft trajectory, speed, images, point clouds, meteorological information, and structured flight intentions; Based on real-time monitoring of traffic flow density and conflict hotspot distribution, the granularity of an adaptive dynamic spatiotemporal grid is dynamically adjusted. Fine-grained grids are used in identified high-risk areas or conflict hotspots, while coarse-grained grids are used in sparsely trafficked areas. The collected multimodal data is uniformly mapped and fused into the corresponding cells of the adaptive dynamic spatiotemporal grid according to its spatiotemporal attributes, generating a spatiotemporally aligned multimodal feature map that balances representation accuracy and computational load, which serves as the spatiotemporal fusion feature.
3. The method according to claim 1, characterized in that, Step S2 specifically includes: S2.1 Obtain a large embodied intelligent agent model that has been pre-trained on a general multimodal task as a base model. The base model has the basic understanding and reasoning ability of the physical world state. S2.2 In the fine-tuning stage, a simplified dynamic model of the aircraft is injected as a priori constraint into the basic model, so that the model can simulate and comply with the physical limitations and maneuverability of the aircraft during its internal reasoning process. S2.
3. Using a small sample dataset containing real-world scenarios, adversarial perturbation samples, and anomalous samples, fine-tune the model with injected physical constraints. The fine-tuning process is achieved by minimizing a composite loss function, which includes: a task loss for optimizing trajectory prediction accuracy, a knowledge distillation loss for injecting geofence and aerial corridor rules, and an adversarial regularization loss term specifically designed to improve the model's robustness against perturbations.
4. The method according to claim 3, characterized in that, In step S2.3, the construction of the small sample dataset and the fine-tuning of the model further include: Based on the simplified dynamics model and preset airspace rules, a large amount of synthetic trajectory data covering various conflict scenarios, extreme weather, and sensor failure conditions is generated through simulation. A multi-stage course learning fine-tuning process is designed. First, the synthetic data is used to pre-adapt the model, enabling it to initially grasp the general patterns and rules of low-altitude traffic flow. Then, the small sample dataset containing real, adversarial, and anomalous samples is used for fine-tuning, focusing on optimizing the model's performance under real noise distribution and its robustness to adversarial and anomalous samples.
5. The method according to claim 1, characterized in that, In step S3, the tasks completed by the dedicated prediction model in an end-to-end manner specifically include: Based on its internally integrated aircraft dynamics model, it identifies the target's current maneuvering behavior and infers its short-term flight intentions under physical constraints, with a focus on identifying non-cooperative flight behaviors that deviate from the norm or are not declared. When predicting the future multi-step trajectory of an aircraft, the interactions between aircraft and their interaction with environmental constraints are explicitly modeled, and the probabilistic trajectory prediction is output. Based on the predicted probability trajectory, the game theory-driven risk assessment is performed.
6. The method according to claim 5, characterized in that, The specific process for calculating the comprehensive collision risk index in step S3 includes: For any pair of aircraft, based on their predicted probability trajectory distribution, calculate the joint probability of spatial overlap at each future time step, which serves as the underlying spatiotemporal conflict potential energy. Construct the incomplete information game simulation model to simulate the avoidance actions and probabilities that each party may take based on its physical capabilities, observation information, and preset strategies in the conflict scenario; calculate a game correction factor based on the simulation results to dynamically adjust the basic conflict potential energy. The basic conflict potential energy is combined with the game correction factor, and the uncertainty represented by the covariance of the predicted trajectory and the urgency represented by the relative speed and heading angle of the two aircraft are further weighted and integrated to generate the final comprehensive collision risk index.
7. The method according to claim 5 or 6, characterized in that, The output of the dedicated prediction model in step S3, and the calculation results of the comprehensive collision risk index, are organized into a structured risk assessment map. The risk assessment map includes at least: A spatiotemporal risk heatmap is used to visualize the spatial distribution of the comprehensive collision risk index at multiple future time steps on the adaptive dynamic spatiotemporal grid. The conflict event list is used to list all predicted potential conflict pairs and associate them with the corresponding comprehensive collision risk index value, the predicted nearest point spatiotemporal location, and the main avoidance strategy options derived from the game theory model.
8. The method according to claim 1, characterized in that, Step S4 specifically includes: Based on the overall operational status of the current airspace, meteorological conditions, and the current granularity of the adaptive dynamic spatiotemporal grid, multi-level safety thresholds are dynamically set; the comprehensive collision risk index is compared with the multi-level safety thresholds to generate conflict warning levels of different urgency. For conflicts that trigger warnings, at least one set of cooperative avoidance strategy recommendations is generated based on the incomplete information game inference model. The strategy recommendations include at least the following: the recommended adjustments to the heading, speed, and altitude of each relevant aircraft, as well as the priority or timing of the adjustments.
9. The method according to claim 1, characterized in that, Step S5 specifically includes: Based on the conflict warning level, the communication capabilities of the relevant aircraft, and the avoidance responsibility weight inferred from the game theory model, the communication link and protocol are adaptively selected, and the priority and timeliness requirements for information distribution are determined. The hierarchical conflict warning information and cooperative avoidance strategy suggestions are encapsulated into standardized instructions or consultation messages that conform to the target system interface specifications, and are securely transmitted through a communication network with authentication and encryption mechanisms. After information is distributed, the trajectory response status of the relevant aircraft is continuously monitored; the actual response is compared with the strategy suggestions, and the comparison results are used as feedback data to continuously optimize the dedicated prediction model or the game inference model.
10. A low-altitude traffic flow collaborative sensing and conflict prediction system based on a multimodal large model, applied to the method described in any one of claims 1 to 9, characterized in that, The system includes: The multi-source data fusion processing module is used to collect multi-source heterogeneous sensing data and flight plan data of aircraft in the low-altitude environment, and based on the real-time traffic situation, fuse the data and map it into a spatiotemporal grid with dynamically adjustable granularity to generate unified spatiotemporal fusion features and complete the spatiotemporal adaptive representation of multimodal data. A dedicated prediction model module is used to fine-tune a low-altitude traffic flow dataset containing real-world scenarios, adversarial disturbance samples, and anomalous samples, based on a pre-trained embodied agent large model. During fine-tuning, aircraft dynamics constraints and airspace rules are introduced as model priors, and a loss function with adversarial regularization is used for optimization, resulting in a dedicated prediction model suitable for low-altitude traffic flow scenarios. The embodied agent large model is a visual-language-action model based on VIMA or RT2 architecture. An integrated reasoning and risk assessment module is used to input the spatiotemporal fusion features into the dedicated prediction model; the dedicated prediction model performs integrated reasoning on the input features and simultaneously outputs the aircraft's behavioral intent recognition results and its future multi-step probability trajectory prediction; based on the probability trajectory prediction, an incomplete information game inference model is constructed to simulate the strategy interaction between aircraft, and then calculates a comprehensive collision risk index that integrates trajectory uncertainty, relative motion state and avoidance strategy derived from game inference; The intelligent early warning and strategy generation module is used to generate graded conflict early warning information and corresponding cooperative avoidance strategy suggestions based on the comprehensive collision risk index and a preset dynamic safety threshold. The early warning information distribution and communication module is used to distribute the graded conflict early warning information and cooperative avoidance strategy suggestions to relevant aircraft or air traffic management systems through a communication network to support real-time collision avoidance decisions.
Citation Information
Patent Citations
Low-altitude airspace management method and system based on data analysis
CN119918739A
Multi-agent collision-free path planning method based on fusion DQN algorithm
CN120949778A