Airspace communication perception fusion method for low-altitude Internet of Things

By integrating millimeter-wave radar, multi-band radio frequency transceiver modules, and embedded processing units onto low-altitude aircraft, high-precision, low-latency airspace communication and perception fusion was achieved. This solved the problems of synchronization and anti-interference between perception and communication in low-altitude intelligent networks, and improved the accuracy of obstacle detection and the real-time performance of the system.

CN121924451APending Publication Date: 2026-04-24CHONGQING COLLEGE OF ELECTRONICS ENG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING COLLEGE OF ELECTRONICS ENG
Filing Date
2026-01-30
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies for communication and sensing in low-altitude environments suffer from insufficient accuracy, real-time performance, anti-interference capabilities, and resource efficiency, making it difficult to meet the actual operational requirements of low-altitude intelligent networks. In particular, insufficient time synchronization and motion compensation accuracy in high-speed flight scenarios leads to spatial offset between point clouds and CSI, affecting the accuracy and safety of obstacle detection.

Method used

A fusion sensing terminal using millimeter-wave radar and multi-band radio frequency transceiver modules achieves joint feature extraction of point clouds and CSI through high-stability clock source synchronization, real-time data alignment of embedded processing units, and lightweight neural network models. Motion compensation is performed in conjunction with IMU, and a frequency hopping spread spectrum mechanism is used to reduce broadcast conflicts. Ground base stations perform cross-view data fusion to construct a high-precision airspace situation map.

Benefits of technology

It achieves high-fidelity modeling of complex low-altitude environments, eliminates point cloud distortion and Doppler blur caused by aircraft self-motion, meets the perception refresh rate requirements of high-speed low-altitude flight scenarios, reduces the probability of broadcast collisions, and improves the geometric accuracy of obstacle detection and the stability of communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121924451A_ABST
    Figure CN121924451A_ABST
Patent Text Reader

Abstract

According to the airspace communication sensing fusion method for the low-altitude Internet of Things, a fusion sensing terminal integrated with a millimeter-wave radar, a multi-band radio frequency transceiver module and an embedded processing unit is deployed on a low-altitude aircraft, and the fusion sensing terminal establishes a bidirectional communication link with a ground base station and an adjacent aircraft through the multi-band radio frequency transceiver module; the embedded processing unit receives original point cloud data from the millimeter wave radar and channel state information (CSI) from the multi-band radio frequency transceiver module in real time. High-fidelity modeling of a complex low-altitude environment is realized through fusion of millimeter-wave radar and C-waveband radio frequency sensing, the millimeter-wave radar provides centimeter-level distance resolution and Doppler velocity information and is good at capturing a geometric contour, and a C-waveband sensing sub-module uses an orthogonal time-frequency-space pilot frequency and polarization switching mechanism to realize high-fidelity modeling of a complex low-altitude environment. Therefore, the static obstacle is no longer misjudged as a dynamic target, and the geometric accuracy and physical consistency of obstacle detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent network communication and sensing technology, specifically a method for airspace communication and sensing fusion for low-altitude intelligent networks. Background Technology

[0002] With the rapid rise of urban air traffic (UAM), logistics drones, and low-altitude emergency inspection applications, low-altitude airspace is evolving from "sparse and isolated" to "dense, collaborative, and intelligent." Against this backdrop, building highly reliable, high-precision, and low-latency airspace perception and communication capabilities has become a core infrastructure requirement to support the safe operation of large-scale low-altitude aircraft. However, existing technologies for achieving communication and perception functions in low-altitude environments generally employ a discrete architecture, leading to challenges in accuracy, real-time performance, interference resistance, and resource efficiency, making it difficult to meet the actual operational requirements of future low-altitude intelligent networks.

[0003] First, current mainstream solutions mostly rely on a single sensor. For example, some drones are only equipped with visual or lidar, which are easily affected by environmental factors such as lighting, rain, fog, and smoke. Other systems, although incorporating millimeter-wave radar, have weak reflection from non-metallic obstacles (such as carbon fiber rotors, thin wires, and transparent curtain walls), resulting in a large number of blind spots. While wireless communication systems can infer environmental scattering characteristics through channel state information (CSI), traditional communication modules are not optimized for perception tasks. Their simple pilot structures and fixed polarization methods lead to a severe deficiency in the ability to distinguish static obstacles and non-cooperative targets.

[0004] Secondly, regarding time synchronization and motion compensation, existing sensor fusion solutions mostly rely on software-level timestamps or network protocols (such as PTP) for data alignment, with synchronization accuracy only reaching the millisecond level. In high-speed flight scenarios (>20m / s), millisecond-level time deviations can cause point clouds and CSI to shift spatially by several meters, resulting in severe distortion of the fusion results. Simultaneously, most systems do not integrate high-dynamic IMUs, or although they have IMUs, they are not used for real-time motion compensation of radar point clouds. This causes the aircraft's own motion to introduce additional Doppler frequency shifts into the radar echo, misjudging stationary buildings and utility poles as high-speed moving targets, leading to false triggering or missed alarms in obstacle avoidance systems, seriously threatening flight safety.

[0005] Furthermore, in terms of computation and real-time performance, existing fusion algorithms mostly rely on general-purpose CPUs or GPUs to run complex neural networks. The models are not quantized or pruned, and inference latency often exceeds 50 milliseconds, which cannot meet the requirements of low-altitude aircraft for sensing refresh rates (≥30Hz, ideally ≥60Hz). At the same time, the lack of dedicated hardware acceleration support results in high power consumption and large size, making it difficult to deploy on small electric vertical takeoff and landing (eVTOL) aircraft or lightweight logistics drone platforms. Summary of the Invention

[0006] In order to overcome the shortcomings of the prior art, the present invention provides a method for airspace communication and sensing fusion for low-altitude intelligent networks, so as to at least partially solve the above-mentioned technical problems.

[0007] The technical solution adopted in this invention is as follows: This invention proposes a spatial communication sensing fusion method for low-altitude intelligent networks, comprising the following steps: Step 1: Deploy a fusion sensing terminal on a low-altitude aircraft that integrates millimeter-wave radar, multi-band radio frequency transceiver modules, and embedded processing units. Step 2: The fusion sensing terminal establishes a two-way communication link with the ground base station and the nearby aircraft through the multi-band radio frequency transceiver module; Step 3: The embedded processing unit receives raw point cloud data from the millimeter-wave radar and channel state information (CSI) from the multi-band radio frequency transceiver module in real time. Step 4: After aligning the original point cloud data with the Channel State Information (CSI) at timestamps, unify the spatial coordinate system to form fused input data. Based on a preset lightweight neural network model, perform joint feature extraction on the fused input data and output fused perception results containing obstacle positions, velocity vectors, and the orientation of communication interference sources. Step 5: Broadcast the fused sensing results to other nodes in the low-altitude intelligent network through the multi-band radio frequency transceiver module.

[0008] In one embodiment of the present invention, the millimeter-wave radar and the multi-band radio frequency transceiver module are placed on the same physical carrier, the phase center distance between their antenna arrays does not exceed 5 cm, and they are provided with a synchronization trigger signal through the same high-stability clock source.

[0009] In one embodiment of the present invention, the multi-band radio frequency transceiver module includes a Sub-6GHz communication submodule and a C-band sensing submodule. The Sub-6GHz communication submodule is used to interact with the ground base station for control commands and status information. The C-band sensing submodule is independently configured with a dedicated transmission pilot sequence. The pilot sequence adopts an orthogonal time-frequency space coding structure and alternately switches the polarization mode in each sensing cycle to enhance the ability to distinguish the scattering response of metallic obstacles and non-cooperative targets.

[0010] In one embodiment of the present invention, the embedded processing unit is provided with a dual-buffered circular queue. The first buffer is dedicated to caching the original point cloud data frames marked with timestamps, and the second buffer is dedicated to caching CSI data frames with the same timestamp. When the embedded processing unit detects that there are data frame pairs in the two buffers with a timestamp difference of less than a preset threshold Δt, it triggers a coordinate system transformation module to map the point cloud coordinates in the millimeter-wave radar coordinate system to the body coordinate system with the aircraft's center of mass as the origin through a pre-calibrated extrinsic parameter matrix, and converts the angle arrival at AoA information in the CSI data into a spatial vector in the same body coordinate system.

[0011] In one embodiment of the present invention, the lightweight neural network model comprises a feature stitching layer, a three-dimensional convolutional encoder, an attention weight allocation module, and a multi-task decoder connected sequentially. The feature stitching layer stitches the voxelized three-dimensional tensor of the point cloud with the CSI amplitude-phase two-dimensional map along the channel dimension. The three-dimensional convolutional encoder employs a depth-separable convolutional kernel structure, downsampling layer by layer to generate multi-scale spatial features. The attention weight allocation module dynamically adjusts the fusion weight coefficients of radar features and CSI features based on the aircraft's current altitude. The multi-task decoder outputs obstacle bounding box parameters, velocity vector components, and cosine values ​​of the interference source direction.

[0012] In one embodiment of the present invention, the fusion sensing terminal further includes an inertial measurement unit (IMU), which is directly connected to the embedded processing unit via an SPI bus to output the attitude angle, angular velocity, and acceleration data of the aircraft in real time. The embedded processing unit uses the IMU data to perform motion compensation on the point cloud acquired by the millimeter-wave radar, specifically including: rotating the point cloud coordinates according to the current attitude angle, and then applying a reverse translation offset to each point cloud point according to the linear velocity vector to eliminate Doppler blurring caused by the aircraft's own motion.

[0013] In one embodiment of the present invention, when the fusion sensing results are broadcast through a multi-band radio frequency transceiver module, a frequency hopping spread spectrum mechanism based on the aircraft identity ID is adopted. That is, each aircraft selects a working frequency point in a time-slot rotation from a pre-allocated set of frequency hopping sequences, and modulates the fusion sensing result data packet on the selected frequency point using a direct sequence spread spectrum method. The spread spectrum code is generated by the aircraft ID and the current UTC seconds, ensuring that adjacent aircraft use different spread spectrum codes in the same time slot, thereby reducing the probability of broadcast collisions.

[0014] In one embodiment of the present invention, the ground base station is configured with a ground fusion node that is symmetrical to the structure of the fusion sensing terminal. The ground fusion node also integrates millimeter-wave radar, multi-band radio frequency transceiver module and embedded processing unit, and is connected to the regional air traffic control server through an optical fiber link. After receiving the fusion sensing results broadcast by the aircraft in the air, the ground fusion node performs cross-view registration with the local sensing data. The registration process adopts the nearest neighbor matching algorithm based on quadtree spatial index to construct a three-dimensional airspace situation map with a coverage radius of not less than 3 kilometers under a unified geographic coordinate system.

[0015] In one embodiment of the present invention, the timestamp alignment adopts a hardware-level synchronization mechanism, specifically: a field-programmable gate array (FPGA) is set inside the fusion sensing terminal. The FPGA is simultaneously connected to the trigger port of the millimeter-wave radar, the baseband processor interrupt pin of the multi-band radio frequency transceiver module, and the PPS pulse output terminal of the global navigation satellite system (GNSS) receiver; after receiving the rising edge of the PPS pulse per second from the GNSS, the FPGA immediately sends a synchronization start signal to the radar and radio frequency module, and adds a microsecond-level time tag based on the PPS count to each frame of data subsequently acquired.

[0016] In one embodiment of the present invention, the lightweight neural network model is deployed on a neural network acceleration coprocessor within an embedded processing unit. The coprocessor adopts a fixed-point quantization architecture, supports INT8 precision operations, and has a built-in dedicated tensor slicing engine. During the model inference stage, the coprocessor loads the weight parameters of the three-dimensional convolutional encoder into on-chip SRAM in channel groups, and simultaneously uses a DMA controller to read point cloud voxel blocks and CSI map blocks from main memory in parallel to achieve pipelined feature fusion calculation, with a single inference latency of no more than 15 milliseconds.

[0017] The beneficial effects of the technical solution of this invention are as follows: This invention achieves high-fidelity modeling of complex low-altitude environments by fusing millimeter-wave radar and C-band radio frequency sensing. The millimeter-wave radar provides centimeter-level range resolution and Doppler velocity information, excelling at capturing geometric contours. The C-band sensing submodule utilizes an orthogonal time-frequency-space (OTFS) pilot and polarization switching mechanism, exhibiting unique sensitivity to metallic structures, non-cooperative targets, and even small obstacles (such as power lines). Both are co-located on the same carrier plate with a phase center distance ≤ 5 cm, reducing viewing angle differences. Combined with IMU-driven motion compensation (attitude rotation and velocity reverse translation), it effectively eliminates point cloud distortion and Doppler blurring caused by aircraft self-motion, preventing static obstacles from being misjudged as dynamic targets and improving the geometric accuracy and physical consistency of obstacle detection.

[0018] This invention ensures controllable end-to-end latency through hardware-level time synchronization and a high-efficiency inference pipeline. The FPGA uses GNSSPPS pulses as a reference to synchronize the sampling times of the radar and RF modules at the nanosecond level, and adds microsecond-level time stamps to each frame of data. The embedded processing unit employs a double-buffered circular queue to achieve frame-level alignment between the point cloud and CSI, triggering fusion only when the time difference is less than Δt, thus avoiding invalid computation. Simultaneously, through INT8 quantization, channel group loading, tensor slicing, and parallel DMA transmission, pipelined feature fusion is achieved, with a single inference latency ≤15 milliseconds, meeting the requirements of sensing refresh rates above 60Hz and fully adapting to high-speed, low-altitude flight scenarios.

[0019] This invention deeply couples the broadcasting of sensing results with a frequency-hopping spread spectrum mechanism. Each aircraft dynamically generates a spreading code based on a unique ID and UTC time, and rotates its operating frequency in a pre-allocated frequency-hopping sequence. This ensures that even in dense formations or high-density scenarios like urban canyons, adjacent nodes can achieve code domain orthogonality within the same time slice, reducing the probability of broadcast collisions. Simultaneously, the Sub-6GHz submodule ensures control link stability, while the C-band submodule focuses on sensing. This clear division of spectrum resources avoids mutual interference between communication and sensing, forming a virtuous cycle of "sensing without interfering with communication, and communication assisting sensing."

[0020] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0021] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a method framework diagram of the airspace communication and sensing fusion method for low-altitude intelligent networks proposed in an embodiment of the present invention; Figure 2 This is a first functional framework diagram of the airspace communication and sensing fusion method for low-altitude intelligent networks proposed in an embodiment of the present invention; Figure 3 This is a second functional framework diagram of the airspace communication and sensing fusion method for low-altitude intelligent networks proposed in an embodiment of the present invention; Figure 4 This is the third functional framework diagram of the airspace communication and sensing fusion method for low-altitude intelligent networks proposed in this embodiment of the invention. Detailed Implementation

[0022] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0023] The following describes an embodiment of the present invention, with reference to the accompanying drawings, a method for airspace communication and sensing fusion for low-altitude intelligent networks.

[0024] like Figures 1 to 4 As shown, this embodiment of the invention provides a spatial communication sensing fusion method for low-altitude intelligent networks, including the following steps: Step 1: Deploy a fusion sensing terminal on a low-altitude aircraft that integrates millimeter-wave radar, multi-band radio frequency transceiver modules, and embedded processing units. Step 2: The fusion sensing terminal establishes a two-way communication link with the ground base station and nearby aircraft through the multi-band radio frequency transceiver module; Step 3: The embedded processing unit receives raw point cloud data from the millimeter-wave radar and Channel State Information (CSI) from the multi-band radio frequency transceiver module in real time. Step 4: After aligning the original point cloud data with the Channel State Information (CSI) at timestamps, unify the spatial coordinate system to form fused input data. Based on a preset lightweight neural network model, perform joint feature extraction on the fused input data and output fused perception results containing obstacle positions, velocity vectors, and the orientation of communication interference sources. Step 5: Broadcast the fusion sensing results to other nodes in the low-altitude intelligent network via a multi-band radio frequency transceiver module.

[0025] In practical applications, the fusion sensing terminal integrates millimeter-wave radar, a multi-band radio frequency transceiver module, and an embedded processing unit, sharing a unified high-stability clock source and power management module to ensure synchronization of each subsystem in both time and space dimensions. Once the aircraft takes off, the fusion sensing terminal immediately activates, actively scanning preset frequency bands through its built-in multi-band radio frequency transceiver module to establish a downlink control link with the ground base station. Simultaneously, it utilizes the broadcast channel to construct a point-to-point or point-to-multipoint self-organizing communication network with nearby aircraft, forming a dynamically reconfigurable low-altitude communication topology.

[0026] Meanwhile, the millimeter-wave radar continuously transmits frequency-modulated continuous wave signals and receives echoes from the surrounding environment, generating raw point cloud data containing range, azimuth, and velocity information. The multi-band RF transceiver module, while performing routine communication tasks, also continuously collects Channel State Information (CSI). CSI not only reflects the quality of the communication link but also implicitly reveals the disturbance characteristics of scatterers and reflectors in the environment on the wireless signal propagation path. The embedded processing unit receives these two types of heterogeneous sensing data through a double-buffering mechanism and performs frame-level alignment based on the high-precision timestamps provided by the hardware synchronization mechanism. Specifically, the system uses an FPGA synchronization controller based on GNSS second pulse triggering, inserting microsecond-level time stamps at the start of each radar sampling and CSI acquisition frame to ensure that the time deviation between the two types of data is controlled within 10 nanoseconds.

[0027] After time alignment is completed, the system enters the spatial coordinate unification stage. Since millimeter-wave radars typically establish a local polar coordinate system with the phase center of their own antenna as the origin, and the Angular Arrival (AoA) information in CSI is calculated based on the geometric layout of the radio frequency antenna array, there is an inherent coordinate system offset between the two. To address this, the system has undergone extrinsic parameter calibration before leaving the factory, obtaining the rotation matrix and translation vector from the radar coordinate system to the aircraft body coordinate system. During runtime, the embedded processing unit uses the extrinsic parameter matrix to transform the point cloud data to the body coordinate system with the aircraft's center of mass as the origin, and simultaneously maps the spatial angle information resolved from CSI to the same coordinate system, thereby constructing a spatiotemporally consistent fused input tensor. The data is then fed into a pre-defined lightweight neural network model for joint feature extraction: the model first converts the point cloud into a three-dimensional mesh and encodes the CSI amplitude-phase map into a two-dimensional feature map. The two are then concatenated along the channel dimension and input into a three-dimensional convolutional encoder. The encoder extracts multi-scale spatial semantics layer by layer through depthwise separable convolution, and the attention module dynamically weights the contribution of radar features and CSI features according to the aircraft's current altitude and speed context information. Finally, the multi-task decoder outputs the three-dimensional bounding box of the obstacle, the velocity vector (including direction and magnitude), and the spatial azimuth angle of potential communication interference sources in parallel.

[0028] The generated fusion sensing results are broadcast in real time to other nodes in the low-altitude intelligent network via a multi-band radio frequency transceiver module. The broadcasting process employs a frequency hopping spread spectrum mechanism based on an identity ID: each aircraft has a unique ID, which the system uses to select the operating frequency for the current time slice from a pre-allocated frequency hopping sequence library and dynamically generate a spreading code in conjunction with UTC time. This ensures that even if multiple aircraft broadcast simultaneously in the same area, code division multiplexing can effectively avoid collisions. After receiving the fusion sensing results, ground base stations or other aircraft can perform cross-view fusion with local sensing data. For example, ground fusion nodes can use their own radar and airborne node data to construct a three-dimensional airspace situation map covering a range of several kilometers, thereby achieving full-domain collaborative sensing and resource scheduling.

[0029] In one specific implementation, the millimeter-wave radar and the multi-band radio frequency transceiver module are co-located on the same physical carrier. The phase center distance between their antenna arrays does not exceed 5 cm, and they are provided with a synchronization trigger signal through the same high-stability clock source. The multi-band radio frequency transceiver module includes a Sub-6GHz communication submodule and a C-band sensing submodule. The Sub-6GHz communication submodule is used to exchange control commands and status information with the ground base station. The C-band sensing submodule is independently configured with a dedicated transmission pilot sequence. The pilot sequence adopts an orthogonal time-frequency spatial coding structure and alternately switches the polarization mode in each sensing cycle to enhance the ability to distinguish the scattering response of metallic obstacles and non-cooperative targets.

[0030] In practical applications, the integrated front-end of this invention is uniformly driven by a single high-stability clock source. This clock source typically employs a temperature-compensated crystal oscillator (TCXO) or an oven-controlled crystal oscillator (OCXO), with a frequency stability better than ±0.1ppm. The clock signal is simultaneously fed into the waveform generator of the millimeter-wave radar, the baseband modulator of the C-band sensing submodule, and the sampling controller of the Sub-6GHz communication submodule, enabling synchronized triggering of all three on a nanosecond timescale. Specifically, at the start of each sensing cycle, the high-stability clock generates a synchronization pulse, simultaneously initiating the millimeter-wave radar to transmit a frequency-modulated continuous wave, the C-band submodule to transmit a dedicated pilot sequence, and the Sub-6GHz submodule to enter receiving and listening mode.

[0031] In terms of functional division, the multi-band RF transceiver module is internally divided into two subsystems that are logically and physically partially isolated but share RF front-end resources: the Sub-6GHz communication submodule mainly undertakes the bidirectional data interaction tasks with the ground base station, including flight command issuance, status reporting, identity authentication, and network access control. Its operating frequency band covers the mainstream 5G private network frequency band of 3.3-3.8GHz, and it adopts OFDM modulation and standard MAC protocol stack to ensure the reliability and interoperability of communication; while the C-band sensing submodule focuses on environmental sensing functions and operates in the 5.725-5.850GHz unlicensed ISM frequency band. Its core is to transmit a carefully crafted dedicated pilot sequence. The sequence adopts the orthogonal time-frequency space (OTFS) coding structure, which maps information symbols to the delay-Doppler domain. It not only has natural robustness to high-speed moving targets, but also can simultaneously obtain the channel delay spread and Doppler frequency shift characteristics in a single transmission, thereby inverting the spatial distribution and dynamic properties of scatterers in the environment.

[0032] Furthermore, the C-band sensing submodule actively switches antenna polarization modes within each sensing cycle, for example, alternately transmitting pilot signals between horizontal polarization (H-pol) and vertical polarization (V-pol), and recording the corresponding received CSI. Since targets of different materials (such as metal UAV frames, carbon fiber rotors, glass curtain walls, and tree branches) exhibit different reflection and scattering characteristics to waves of different polarizations, this effectively enhances the system's ability to distinguish non-cooperative targets (such as obstacles without transponders) and weak scatterers (such as thin wires and transparent obstacles). For example, metallic objects typically respond weakly to cross-polarization but strongly to the same polarization, while vegetation exhibits a strong depolarization effect. By comparing the amplitude and phase changes of CSI at the same spatial location under different polarizations, the system can extract polarization feature vectors to help determine the material properties of obstacles, thereby improving the semantic richness and decision reliability of the sensing results.

[0033] In one specific implementation, the embedded processing unit internally includes a dual-buffered circular queue. The first buffer is dedicated to caching raw point cloud data frames marked with timestamps, and the second buffer is dedicated to caching CSI data frames with the same timestamp. When the embedded processing unit detects a pair of data frames in the two buffers whose timestamp difference is less than a preset threshold Δt, it triggers a coordinate system transformation module. This module maps the point cloud coordinates in the millimeter-wave radar coordinate system to the body coordinate system with the aircraft's center of mass as the origin using a pre-calibrated extrinsic parameter matrix, and converts the angle arrival at altitude (AoA) information in the CSI data to the same body coordinate system. The lightweight neural network model for spatial vectors consists of a feature stitching layer, a 3D convolutional encoder, an attention weight allocation module, and a multi-task decoder connected sequentially. The feature stitching layer stitches the voxelized 3D tensor of the point cloud with the CSI amplitude-phase 2D map along the channel dimension. The 3D convolutional encoder uses a depth-separable convolutional kernel structure to downsample layer by layer, generating multi-scale spatial features. The attention weight allocation module dynamically adjusts the fusion weight coefficients of radar features and CSI features based on the aircraft's current altitude. The multi-task decoder outputs obstacle bounding box parameters, velocity vector components, and cosine values ​​of the interference source directions.

[0034] In practical applications, the embedded processing unit maintains two independent but collaborative circular buffers in parallel: the first buffer is dedicated to receiving and temporarily storing raw point cloud data frames output by millimeter-wave radar, with each frame being stamped with a microsecond-level timestamp generated by a high-stability clock upon writing; the second buffer is dedicated to receiving CSI data frames acquired by the multi-band RF transceiver module, also with aligned timestamps. Due to the difference between the radar sampling rate and the CSI acquisition frequency (e.g., radar frames every 20ms, CSI frames every 10ms), the data frames in the two buffers are not naturally aligned on the time axis. Therefore, the embedded processing unit continuously polls the head element of both buffers, calculating the timestamp difference between any pair of point cloud frames and CSI frames in real time. Once the absolute value of the time difference between a pair of data frames is detected to be less than a preset threshold Δt (usually set to 5-10 milliseconds to balance the dynamics of moving targets and the real-time performance of the system), the system immediately determines the data pair as a "valid synchronization frame pair" and triggers the subsequent fusion processing flow.

[0035] Upon triggering, the coordinate system transformation module intervenes first. This module loads the extrinsic parameter matrix obtained during the ground calibration phase. The matrix describes the rotation and translation relationship of the millimeter-wave radar antenna phase center relative to the origin (usually the centroid) of the aircraft's body coordinate system. Based on this, the system maps each 3D point (x_radar, y_radar, z_radar) in the point cloud to the body coordinate system through rigid body transformation, obtaining a unified spatial representation (x_body, y_body, z_body). Simultaneously, the angle of arrival (AoA) information (usually azimuth θ and pitch φ) calculated from the CSI data using MUSIC or compressed sensing algorithms is also converted into unit direction vectors (cosφ·cosθ, cosφ·sinθ, sinφ) in the body coordinate system. This unifies the two types of heterogeneous sensing data, originally belonging to different observation perspectives and coordinate systems, into a single physical space reference system.

[0036] After coordinate unification, the data enters a lightweight neural network model for end-to-end joint feature extraction. The model adopts an encoder-decoder architecture, but is deeply optimized to address the computing power and power consumption constraints of the embedded platform. First, the original point cloud is voxelized: the three-dimensional space in the body coordinate system is divided into a fixed-size voxel grid (e.g., 64×64×16). The number of points or the average reflection intensity within each voxel is encoded as channel values, forming a three-dimensional tensor. The CSI data is organized into an amplitude-phase two-dimensional map, where the horizontal axis is the subcarrier index and the vertical axis is the receiving antenna number. Amplitude and phase each constitute two channels. The feature stitching layer directly stitches these two heterogeneous tensors along the channel dimensions to form a fused input tensor, preserving the structural characteristics of the original mode without introducing additional projection loss.

[0037] Subsequently, the fusion tensor is fed into a 3D convolutional encoder. The encoder abandons the traditional dense convolution and instead adopts a depthwise separable 3D convolutional kernel structure: first, it performs spatial convolution independently on each input channel, and then performs pointwise convolution through 1×1×1 convolution. The encoder downsamples layer by layer, gradually expanding the receptive field and generating a multi-scale feature pyramid from fine-grained to coarse-grained, which captures the fine contours of nearby obstacles and also perceives the macroscopic distribution of distant scatterers.

[0038] Based on this, the attention weight allocation module dynamically adjusts the fusion weights of radar features and CSI features according to the current altitude of the aircraft. For example, when flying at very low altitudes (<50 meters), there are dense static obstacles such as buildings and trees, and the geometric accuracy of millimeter-wave radar is superior, so the system automatically increases the weight of point cloud voxel features; while in mid-to-high altitude scenarios (100-300 meters), there are mainly dynamic targets such as metal UAVs and birds, and C-band CSI is sensitive to metal scatterers, so the system enhances the contribution ratio of CSI spectral features.

[0039] Ultimately, the multi-task decoder outputs three types of perception results in parallel: first, the three-dimensional bounding box parameters of obstacles (center coordinates, length, width, height, and yaw angle), used for obstacle avoidance planning; second, the velocity vector components (vx, vy, vz) of each obstacle, jointly estimated by point cloud Doppler information and CSI time-varying characteristics, supporting trajectory prediction; and third, the direction cosine values ​​(l, m, n) of potential communication interference sources, directly derived from the weighted AoA spatial vector, used for beamforming or spectrum avoidance.

[0040] In one specific implementation, the fusion sensing terminal also includes an inertial measurement unit (IMU). The IMU is directly connected to the embedded processing unit via an SPI bus and outputs the aircraft's attitude angle, angular velocity, and acceleration data in real time. The embedded processing unit uses the IMU data to perform motion compensation on the point cloud acquired by the millimeter-wave radar. Specifically, this includes rotating the point cloud coordinates according to the current attitude angle and then applying a reverse translation offset to each point cloud point according to the linear velocity vector to eliminate Doppler blurring caused by the aircraft's own motion. When the fusion sensing results are broadcast through the multi-band radio frequency transceiver module, a frequency hopping spread spectrum mechanism based on the aircraft's identity ID is adopted. That is, each aircraft selects a working frequency point in a time-slot rotation from the pre-allocated set of frequency hopping sequences and modulates the fusion sensing result data packet on the selected frequency point using a direct sequence spread spectrum method. The spread spectrum code is generated by the aircraft ID and the current UTC seconds, ensuring that adjacent aircraft use different spread spectrum codes in the same time slot, reducing the probability of broadcast collisions.

[0041] In practical applications, the IMU continuously outputs the aircraft's three-axis angular velocity, three-axis acceleration, and attitude angles (including roll, pitch, and yaw) calculated from these at a sampling frequency of no less than 200Hz. This data is sent to the embedded processing unit in real time via a low-latency SPI interface and aligned temporally with the raw point cloud data acquired by the millimeter-wave radar. Because the IMU possesses microsecond-level timestamping capabilities and shares the same highly stable clock source with the radar, the time synchronization error between the two can be controlled to the sub-millisecond level. Once a frame of point cloud data is acquired, the embedded processing unit immediately calls upon the latest valid IMU attitude and velocity information to initiate the motion compensation process.

[0042] Motion compensation consists of two consecutive steps: First, a three-dimensional rotation matrix is ​​constructed based on the current attitude angle to rotate all point cloud points collected in the radar coordinate system to a stable body coordinate system with the aircraft's center of mass as the origin and aligned with geographic north, eliminating the spatial distortion of the point cloud caused by the aircraft's roll or pitch. Second, by integrating the acceleration signal and combining it with the instantaneous linear velocity vector (vx, vy, vz) estimated by Kalman filtering, a translational offset of opposite magnitude to the direction of the aircraft's motion is applied to each point cloud point. This reverse translational offset effectively cancels the additional Doppler frequency shift introduced by the aircraft's own high-speed movement in the radar echo, thereby reducing the "self-motion ambiguity" effect. Without this compensation, stationary obstacles (such as utility poles and building edges) will be incorrectly identified as dynamic targets with radial velocity in the point cloud due to the platform's motion, seriously interfering with subsequent obstacle classification and trajectory prediction.

[0043] Motion-compensated point cloud data exhibits higher geometric fidelity and physical consistency, serving as a reliable foundation for subsequent fusion with CSI data. The fusion perception results generated on this basis (including obstacle positions, velocities, and interference source locations) need to be broadcast to other nodes in the low-altitude intelligent network via a multi-band radio frequency transceiver module.

[0044] Each aircraft is assigned a globally unique ID upon network registration and simultaneously acquires a predefined set of frequency hopping sequences (e.g., a pseudo-random arrangement of 16 available C-band frequencies). The system divides time into fixed-length time slices (e.g., 100ms per slice). At the beginning of each time slice, the aircraft takes the modulo of its current UTC time based on its ID and selects the corresponding operating frequency from the frequency hopping sequence. After selecting the frequency, the aircraft does not directly send the raw data packets. Instead, it first encapsulates the fused sensing results into standard data frames and then uses a pseudo-random spreading code generated by combining the aircraft ID and the current UTC integer seconds to perform direct sequence spread spectrum modulation on the frames. Since the UTC seconds are updated every second, even if multiple aircraft select the same frequency in the same area and the same time slice, as long as their IDs are different, the spreading codes they use will be almost orthogonal. The receiver can effectively separate the signals through correlation despreading, greatly suppressing multiple access interference.

[0045] In one specific implementation, the ground base station is configured with a ground fusion node symmetrical to the structure of the fusion sensing terminal. The ground fusion node also integrates millimeter-wave radar, multi-band radio frequency transceiver module and embedded processing unit, and is connected to the regional air traffic control server through a fiber optic link. After receiving the fusion sensing results broadcast by the airborne aircraft, the ground fusion node performs cross-view registration with the local sensing data. The registration process adopts the nearest neighbor matching algorithm based on quadtree spatial index to construct a three-dimensional airspace situation map with a coverage radius of not less than 3 kilometers under a unified geographic coordinate system. The timestamp alignment adopts a hardware-level synchronization mechanism, specifically: a field-programmable gate array (FPGA) is set inside the fusion sensing terminal. The FPGA is simultaneously connected to the trigger port of the millimeter-wave radar, the baseband processor interrupt pin of the multi-band radio frequency transceiver module and the PPS pulse output terminal of the Global Navigation Satellite System (GNSS) receiver. After receiving the rising edge of the PPS pulse per second from the GNSS, the FPGA immediately sends a synchronization start signal to the radar and radio frequency module, and adds a microsecond-level time tag based on PPS count to each subsequent frame of data.

[0046] In practical applications, ground-based fusion nodes are typically deployed near urban high points, communication towers, or air traffic control infrastructure, with an antenna line-of-sight coverage radius of no less than 3 kilometers. When an airborne aircraft broadcasts its fused perception results (including obstacle positions, velocity vectors, and interference source azimuths) via a multi-band radio frequency transceiver module, the ground-based fusion node synchronously receives the data stream and processes it jointly with the local point cloud data collected in real time by its own millimeter-wave radar and the CSI data acquired by the radio frequency sensing module. Because the angles, distances, and obstruction relationships of airborne and ground-based nodes observing the same airspace targets differ significantly (for example, ground-based radar is easily obstructed by buildings but sensitive to low-altitude targets, while the airborne node's overhead view can penetrate some obstructions but has high requirements for ground clutter suppression), perception results from a single perspective often have blind spots or false detections. Therefore, the obstacle coordinates reported by the airborne node are first transformed from its aircraft coordinate system to a unified geographic coordinate system (such as WGS-84 or ENU local coordinate system). This transformation relies on the high-precision GNSS position and attitude information accompanying the aircraft's broadcast. Simultaneously, the ground-based node also maps its locally perceived targets to the same geographic coordinate system.

[0047] Within this unified spatial framework, the system employs a nearest neighbor matching algorithm based on quadtree spatial indexing to achieve efficient association. Specifically, ground nodes construct a quadtree structure on a two-dimensional horizontal plane (XY) for all target points they perceive, with each leaf node storing the target ID and its three-dimensional coordinates within the area it falls into. When an obstacle point is reported by an airborne node, the system first locates its corresponding region in the quadtree, and then searches only the region and its adjacent sub-regions for the nearest local target point. If the Euclidean distance between the two is less than a preset threshold (e.g., 2 meters), and the angle between their velocity vector directions is less than 30 degrees, they are determined to be the same physical target, triggering data fusion: a weighted average or Kalman filter strategy is used to fuse observations from the air and the ground to generate position and velocity estimates. For targets observed only by one party (such as drones obscured by buildings), the original observation results are retained and the confidence level is marked. Through this multi-source heterogeneous and multi-view complementary fusion mechanism, the system finally constructs a three-dimensional airspace situation map with a coverage radius of not less than 3 kilometers and an update frequency of more than 10 Hz. The map not only includes static obstacles (such as tower cranes and high-voltage lines), but also dynamically tracks the trajectories of all cooperative and non-cooperative aircraft, providing high-fidelity input for airspace management, conflict early warning and path planning.

[0048] Both the airborne fusion sensing terminal and the ground fusion node incorporate a Field Programmable Gate Array (FPGA) as a hardware-level synchronization controller, a frame trigger port for the millimeter-wave radar, an interrupt pin for the baseband processor of the multi-band RF transceiver module, and a pulses per second (PPS) output for the high-precision Global Navigation Satellite System (GNSS) receiver. Whenever the GNSS receiver outputs a rising edge of a PPS (representing an integer second in UTC), the FPGA immediately captures the event and sends a synchronization start signal to the radar and RF module within a nanosecond delay, forcing both to begin a new round of data acquisition simultaneously. Simultaneously, the FPGA internally maintains a microsecond-level timer based on PPS counting. For example, at the k-th microsecond after the Nth PPS, all acquired data frames are time-stamped "Nk". Because the GNSS PPS signal has nanosecond-level consistency globally (after differential correction), all airborne and ground nodes deployed within the area effectively share the same global time reference.

[0049] In one specific implementation, the lightweight neural network model is deployed on a neural network acceleration coprocessor within an embedded processing unit. The coprocessor adopts a fixed-point quantization architecture, supports INT8 precision operations, and has a built-in dedicated tensor slicing engine. During the model inference stage, the coprocessor loads the weight parameters of the 3D convolutional encoder into on-chip SRAM in channel groups, and simultaneously uses a DMA controller to read point cloud voxel blocks and CSI map blocks from main memory in parallel, realizing pipelined feature fusion calculation with a single inference latency of no more than 15 milliseconds.

[0050] In practical applications, the pre-trained lightweight neural network (including a feature concatenation layer, a 3D convolutional encoder, an attention weight allocation module, and a multi-task decoder) undergoes offline quantization processing. This converts floating-point weights and activation values ​​into 8-bit integers (INT8) format. Quantization not only compresses the model size (typically reducing it by more than 75%) but also ensures that subsequent calculations are fully adapted to the coprocessor's fixed-point arithmetic units, avoiding the high power consumption and long latency associated with floating-point operations. The quantized model parameters are divided into multiple logical blocks and grouped according to the channel dimension of the 3D convolutional encoder. For example, if a certain layer has 64 convolutional kernels, it is split into 4 groups, each with 16 channels. This grouping strategy matches the on-chip storage structure of the coprocessor, ensuring that the weight subset loaded each time exactly fills its high-speed SRAM cache, minimizing the frequency of access to external main memory.

[0051] Upon entering the real-time inference phase, the system first uses the main processor to complete point cloud voxelization and CSI map construction, temporarily storing the generated fused input tensor in main memory. At this point, the neural network acceleration coprocessor is activated, initiating a pipelined computation process. Its internal dedicated tensor slicing engine first parses the current input data structure, automatically dividing the 3D point cloud voxel blocks (e.g., 64×64×16×C) and 2D CSI map blocks (e.g., 64×32×2) into several sub-tiles according to spatial dimensions. The size of each sub-tile is strictly aligned with the throughput capacity of the coprocessor's computing array (e.g., an 8×8×8 voxel block). Simultaneously, the DMA (Direct Memory Access) controller reads these data sub-tiles from main memory in parallel without consuming main CPU resources and sends them to the coprocessor's input buffer via a high-bandwidth on-chip interconnect bus.

[0052] Almost simultaneously, the coprocessor's weight scheduling unit preloads the next set of INT8 weights from off-chip flash memory or main memory to on-chip SRAM based on the channel grouping information of the current inference layer. Since the SRAM access latency is only a few nanoseconds, far lower than the hundreds of nanoseconds of external DRAM, the convolution computation unit can prepare the weights for the next layer in advance while performing the current layer's operations, achieving overlap between computation and memory access. The 3D convolution computation unit uses a systolic array or reconfigurable SIMD architecture to perform intensive multiply-accumulate operations on the input sub-blocks and weight sub-blocks. All intermediate results are temporarily stored in on-chip register files in INT8 or INT16 format, avoiding precision overflow while maintaining high throughput.

[0053] Ultimately, the obstacle bounding boxes, velocity vectors, and interference source direction cosine values ​​output by the multi-task decoder are written back to main memory in one go for subsequent broadcast module use. Thanks to the aforementioned hardware and software co-optimization, including INT8 fixed-point quantization, channel grouping weight loading, tensor slicing, DMA parallel transfer, and deep pipeline execution, the entire model inference process can be completed within 15 milliseconds, meeting the stringent requirements of perception update rate (≥60Hz) for low-altitude aircraft in high-speed maneuvering scenarios. Meanwhile, the typical power consumption of the coprocessor is controlled below 2 watts, allowing it to be directly integrated into the embedded motherboard of a small UAV or eVTOL platform without additional heat dissipation or power supply modifications.

[0054] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0055] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A method for airspace communication and sensing fusion for low-altitude intelligent networks, characterized in that, Includes the following steps: Step 1: Deploy a fusion sensing terminal on a low-altitude aircraft that integrates millimeter-wave radar, multi-band radio frequency transceiver modules, and embedded processing units. Step 2: The fusion sensing terminal establishes a two-way communication link with the ground base station and the nearby aircraft through the multi-band radio frequency transceiver module; Step 3: The embedded processing unit receives raw point cloud data from the millimeter-wave radar and channel state information (CSI) from the multi-band radio frequency transceiver module in real time. Step 4: After aligning the original point cloud data with the Channel State Information (CSI) at timestamps, unify the spatial coordinate system to form fused input data. Based on a preset lightweight neural network model, perform joint feature extraction on the fused input data and output fused perception results including obstacle positions, velocity vectors, and the orientation of communication interference sources. Step 5: Broadcast the fused sensing results to other nodes in the low-altitude intelligent network through the multi-band radio frequency transceiver module.

2. The airspace communication sensing fusion method for low-altitude intelligent networks according to claim 1, characterized in that, The millimeter-wave radar and the multi-band radio frequency transceiver module are placed on the same physical carrier board, with the phase center distance between their antenna arrays not exceeding 5 centimeters, and are provided with a synchronization trigger signal through the same high-stability clock source.

3. The airspace communication sensing fusion method for low-altitude intelligent networks according to claim 1, characterized in that, The multi-band radio frequency transceiver module includes a Sub-6GHz communication submodule and a C-band sensing submodule. The Sub-6GHz communication submodule is used to exchange control commands and status information with the ground base station. The C-band sensing submodule is independently configured with a dedicated transmission pilot sequence. The pilot sequence adopts an orthogonal time-frequency space coding structure and alternately switches the polarization mode in each sensing cycle to enhance the ability to distinguish the scattering response of metallic obstacles and non-cooperative targets.

4. The airspace communication sensing fusion method for low-altitude intelligent networks according to claim 1, characterized in that, The embedded processing unit is equipped with a dual-buffered circular queue. The first buffer is dedicated to caching the original point cloud data frames marked with timestamps, and the second buffer is dedicated to caching CSI data frames with the same timestamp. When the embedded processing unit detects a pair of data frames in two buffers with a timestamp difference less than a preset threshold Δt, it triggers a coordinate system transformation module to map the point cloud coordinates in the millimeter-wave radar coordinate system to the body coordinate system with the aircraft's center of mass as the origin through a pre-calibrated extrinsic parameter matrix, and converts the angle arrival at AoA information in the CSI data into a spatial vector in the same body coordinate system.

5. The airspace communication and sensing fusion method for low-altitude intelligent networks according to claim 1, characterized in that, The lightweight neural network model consists of a feature stitching layer, a 3D convolutional encoder, an attention weight allocation module, and a multi-task decoder connected sequentially. The feature stitching layer stitches the voxelized 3D tensor of the point cloud with the CSI amplitude-phase 2D map along the channel dimension. The 3D convolutional encoder uses a depth-separable convolutional kernel structure to downsample layer by layer, generating multi-scale spatial features. The attention weight allocation module dynamically adjusts the fusion weight coefficients of radar features and CSI features based on the aircraft's current altitude. The multi-task decoder outputs obstacle bounding box parameters, velocity vector components, and cosine values ​​of the interference source directions.

6. The airspace communication sensing fusion method for low-altitude intelligent networks according to claim 1, characterized in that, The fusion sensing terminal also includes an inertial measurement unit (IMU), which is directly connected to the embedded processing unit via an SPI bus to output the aircraft's attitude angle, angular velocity, and acceleration data in real time. The embedded processing unit uses the IMU data to perform motion compensation on the point cloud acquired by the millimeter-wave radar. Specifically, it rotates the point cloud coordinates according to the current attitude angle, and then applies a reverse translation offset to each point cloud point according to the linear velocity vector to eliminate Doppler blurring caused by the aircraft's own motion.

7. The airspace communication sensing fusion method for low-altitude intelligent networks according to claim 1, characterized in that, When the fusion sensing results are broadcast through the multi-band radio transceiver module, a frequency hopping spread spectrum mechanism based on the aircraft's identity ID is adopted. That is, each aircraft selects a working frequency point in a time-slot rotation from the pre-allocated set of frequency hopping sequences, and modulates the fusion sensing result data packet on the selected frequency point using direct sequence spread spectrum. The spread spectrum code is generated by the aircraft ID and the current UTC seconds, ensuring that adjacent aircraft use different spread spectrum codes in the same time slot, thereby reducing the probability of broadcast collisions.

8. The airspace communication sensing fusion method for low-altitude intelligent networks according to claim 1, characterized in that, The ground base station is equipped with a ground fusion node that is symmetrical to the structure of the fusion sensing terminal. The ground fusion node also integrates millimeter-wave radar, multi-band radio frequency transceiver module and embedded processing unit, and is connected to the regional air traffic control server through fiber optic link. After receiving the fusion perception results broadcast by the airborne aircraft, the ground fusion node performs cross-view registration with the local perception data. The registration process adopts the nearest neighbor matching algorithm based on quadtree spatial index to construct a three-dimensional airspace situation map with a coverage radius of not less than 3 kilometers under a unified geographic coordinate system.

9. The airspace communication sensing fusion method for low-altitude intelligent networks according to claim 1, characterized in that, The timestamp alignment adopts a hardware-level synchronization mechanism, specifically: a field-programmable gate array (FPGA) is set up inside the fusion sensing terminal. The FPGA is simultaneously connected to the trigger port of the millimeter-wave radar, the baseband processor interrupt pin of the multi-band radio frequency transceiver module, and the PPS pulse output terminal of the global navigation satellite system (GNSS) receiver. Upon receiving the rising edge of the pulse per second (PPS) from the GNSS, the FPGA immediately sends a synchronization start signal to the radar and radio frequency modules, and adds a microsecond-level time stamp based on the PPS count to each subsequent frame of data.

10. The airspace communication sensing fusion method for low-altitude intelligent networks according to claim 1, characterized in that, The lightweight neural network model is deployed on a neural network acceleration coprocessor within an embedded processing unit. The coprocessor adopts a fixed-point quantization architecture, supports INT8 precision operations, and has a built-in dedicated tensor slicing engine. During the model inference stage, the coprocessor loads the weight parameters of the 3D convolutional encoder into on-chip SRAM in channel groups, and simultaneously uses a DMA controller to read point cloud voxel blocks and CSI map blocks from main memory in parallel, realizing pipelined feature fusion calculation with a single inference latency of no more than 15 milliseconds.