A distributed multi-modal unmanned aerial vehicle angle perception and prediction method

By deploying distributed cameras and base stations in urban areas, and fusing visual and wireless echo information for state prediction and beamforming, the stability issues of UAV perception and communication in complex urban environments have been resolved, enabling omnidirectional positioning and efficient communication for UAVs.

CN122237634APending Publication Date: 2026-06-19NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-26
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

In complex urban environments, single-base station sensing modes face signal obstruction and multipath reflection interference, leading to decreased UAV sensing performance and unstable communication links. Existing multimodal fusion methods are not robust enough in urban environments, making it difficult to achieve high-precision prediction and completely eliminate sensing blind spots.

Method used

By employing a distributed multimodal perception method, integrating sensing and communication base stations and multiple distributed cameras in urban areas, and fusing visual and wireless echo information, the system utilizes neural networks for state prediction and beamforming to achieve omnidirectional positioning and communication optimization for unmanned aerial vehicles (UAVs).

Benefits of technology

It enables omnidirectional, continuous, robust positioning and efficient communication of UAVs in complex urban environments, improving the stability and transmission efficiency of communication links.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122237634A_ABST
    Figure CN122237634A_ABST
Patent Text Reader

Abstract

This invention discloses a distributed multimodal UAV angle perception and prediction method, comprising the following steps: S1, determining the system model; S2, determining the visual signal and radio frequency signal models; S3, fusing visual information from multiple cameras to estimate the UAV's orientation relative to the base station at the current moment; S4, fusing the visual estimation results from historical time slots with the wireless echo perception results to predict the UAV's state at the next moment; S5, based on the UAV's state at the next moment predicted in step S4, calculating and generating the optimal transmission beam for the base station antenna in advance. This invention eliminates the blind spots of a single base station's perception through a distributed camera network, achieving omnidirectional, continuous, and robust positioning of UAVs in complex urban environments; by fusing visual and wireless echo information for high-precision motion prediction, the communication system can pre-align with the UAV, transforming passive tracking into proactive service, thereby significantly improving the stability and transmission efficiency of low-altitude communication links.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) perception and communication technology, and in particular to a distributed multimodal UAV angle perception and prediction method. Background Technology

[0002] With the rapid development of the low-altitude economy, drones have demonstrated broad application potential in smart cities, emergency response, and environmental monitoring. Ensuring the safe and efficient operation of drones in complex urban airspace, achieving accurate and continuous positioning and status tracking, and maintaining reliable communication links have become key requirements for low-altitude network construction. Integrated sensing and communication technologies are considered an effective solution, combining wireless communication with target perception capabilities by sharing hardware, spectrum, and signal processing resources, aiming to achieve mutual performance enhancement.

[0003] In a typical ISAC system, the base station usually serves as the core node for sensing and communication. It utilizes a line-of-sight wireless link with the UAV for data transmission while simultaneously analyzing the strong line-of-sight echo signals from the UAV to extract its position, speed, and other state information. However, in the complex urban low-altitude areas with numerous buildings, this base station-centric sensing model faces severe challenges. On the one hand, dense buildings cause severe signal obstruction, leading to sensing link interruptions and creating "sensing blind spots." On the other hand, multipath reflections from other objects (such as building walls and moving vehicles) generate strong environmental clutter, interfering with the effective extraction of target echoes. These factors collectively lead to a significant deterioration in sensing performance based on a single base station echo, resulting in fluctuations or even interruptions in the quality of the communication link, which relies on accurate state information for beam alignment. This makes it difficult to meet the stringent requirements of UAV missions for highly reliable communication.

[0004] To address these challenges, existing research attempts to incorporate multimodal information fusion to aid communication. For example, some studies utilize omnidirectional radar for initial access and motion estimation, while others employ a combined wide-beam and fine-beam sensing strategy to optimize beamforming. In recent years, perception-assisted communication based on multimodal fusion of visual and wireless signals has shown promise. Related research in vehicle-to-vehicle or drone scenarios has integrated camera visual information with GPS data to predict the optimal communication beam. Furthermore, work utilizing multi-device, multimodal information to improve perception performance demonstrates a significant advantage in beam prediction accuracy compared to single-modal methods. However, these existing methods still lack robustness in addressing the persistent loss or degradation of sensor signals caused by physical obstruction in urban environments. A comprehensive framework has yet to be systematically constructed to completely eliminate sensor blind spots and fully utilize complementary sensing information for high-precision prediction.

[0005] Therefore, it is necessary to develop a distributed multimodal UAV angle perception and prediction method to solve the above problems. Summary of the Invention

[0006] The purpose of this invention is to design a distributed multimodal UAV angle perception and prediction method to solve the above problems.

[0007] The present invention achieves the above objectives through the following technical solutions:

[0008] A distributed multimodal UAV angle perception and prediction method, comprising the following steps:

[0009] S1. Determine the system model;

[0010] S2. Determine the models for visual and radio frequency signals;

[0011] S3. By fusing visual information from multiple cameras, the direction of the drone relative to the base station at the current moment is estimated.

[0012] S4. By fusing the visual estimation results from historical time slots with the wireless echo sensing results, the state of the drone at the next moment can be predicted.

[0013] S5. Based on the UAV state predicted in step S4, calculate and generate the optimal transmission beam for the base station antenna in advance.

[0014] Furthermore, S1 specifically includes: deploying an integrated sensing and communication base station and K distributed cameras within the target low-altitude urban area, where K ≥ 2. The base station is equipped with a... A uniform planar array antenna with 1000 elements is used for simultaneous communication and echo-based sensing; K cameras are distributed in space to provide all-round, blind-spot-free visual coverage of the UAV's flight airspace. The base station and all cameras achieve time synchronization through a high-precision synchronization mechanism and complete the initial system access.

[0015] Furthermore, S2 specifically includes the following steps:

[0016] The system establishes a unified spatial coordinate system; in any discrete time slot The position of the target drone relative to the base station is determined by the azimuth angle. and elevation angle Indicates; the One camera The acquired images are denoted as The first camera It is usually located at the same site as the base station and is defined as the main camera;

[0017] The communication and sensing links between the base station and the drone share spectrum resources; the base station transmits signals and receives reflected echoes from the drone. The signal-to-noise ratio of the echo signals is affected by environmental obstruction and clutter. The base station generates a beamforming vector. To achieve beamforming, the array steering vector corresponding to the direction of the UAV... Closely related, guide vector The expression is: ;

[0018] in, The spacing between antenna elements. For carrier wavelength, and These represent the row index and column index of the array, respectively. The beamforming vector is typically set as the guide vector in the prediction direction, i.e.:

[0019] in( () represents the predicted angle of the UAV in the next time slot;

[0020] The system's ultimate performance metric is the achievable communication rate. Its relationship with the received signal-to-noise ratio Related, the expression is:

[0021] ;

[0022] in,

[0023] ;

[0024] here, To transmit SNR, For the complex coefficients of the LOS path, For time slots The true angle, This represents the conjugate transpose of the guiding vector;

[0025] In each work slot The system performs two types of data acquisition in parallel:

[0026] 1. The base station transmits detection signals and receives reflected echoes from the drone;

[0027] II. All Multiple distributed cameras are simultaneously triggered to capture images. All data is accompanied by precise timestamps and aligned by the system to form a time-slot aligned multimodal sensing dataset. And the corresponding echo signal.

[0028] Furthermore, S3 specifically includes:

[0029] S31: In the images captured by each camera, the drone is automatically identified and its center is located; then, a local image area containing the drone is cropped based on this center point.

[0030] S32: Process the images cropped from each camera and the location information detected by each camera, extract features, and use a neural network to uniformly transform these features into the observation view of the base station;

[0031] S33: Based on the clarity and completeness of the image content provided by each camera, automatically calculate and assign different fusion weights, and perform weighted merging of the converted features from all cameras;

[0032] S34: Based on the fused features, the current azimuth and elevation angles of the UAV are finally calculated.

[0033] Furthermore, S4 specifically includes:

[0034] S41: For each historical moment, evaluate the reliability of visual perception and echo perception separately, and select the more reliable perception result at that moment as the final state information adopted at that moment;

[0035] S42: Arrange the final state information of a series of historical moments in chronological order and input it into a temporal neural network; the network analyzes the trend and pattern of state changes over time;

[0036] S43: Using the output of a temporal neural network, predict the change in the state of the UAV from the current moment to the next moment, and then calculate the predicted azimuth and elevation angles for the next moment.

[0037] The beneficial effects of this invention are:

[0038] This invention eliminates the blind spots of a single base station by using a distributed camera network, enabling omnidirectional, continuous, and robust positioning of UAVs in complex urban environments. By fusing visual and wireless echo information for high-precision motion prediction, the communication system can be pre-aimed at the UAV, transforming passive tracking into proactive service, thereby significantly improving the stability and transmission efficiency of low-altitude communication links. Attached Figure Description

[0039] Figure 1 This is a flowchart of the present invention;

[0040] Figure 2 This is a structural diagram of the state estimation model and state prediction model of the present invention. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0042] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0043] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0044] In the description of this invention, it should be understood that the terms "upper," "lower," "inner," "outer," "left," "right," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are only used to facilitate the description of this invention and to simplify the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0045] Furthermore, the terms "first," "second," etc., are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.

[0046] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, terms such as "set" and "connection" should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0047] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0048] like Figure 1 As shown, a distributed multimodal UAV angle perception and prediction method includes the following steps:

[0049] S1. Determine the system model; specifically including: deploying an Integrated Sensing and Communication (ISAC) base station and K distributed cameras within the target low-altitude urban area, where K≥2. The base station is equipped with a… A uniform planar array antenna with 1000 elements is used to communicate with the UAV and sense its status based on echo signals. K cameras are distributed in space to provide omnidirectional, blind-spot-free visual coverage of the UAV's flight airspace. The base station and all cameras are synchronized in time through a high-precision synchronization mechanism to complete the initial system access, and the image acquisition and echo sensing data are aligned in time slots. The system also includes virtual modules: a status prediction module, which fuses historical visual perception information and echo sensing information to predict the UAV's status in the next time slot; and a beamforming module, which generates beamforming vectors based on the predicted UAV status to improve the quality of the communication link.

[0050] S2. Determine the models for visual and radio frequency signals;

[0051] The system establishes a unified spatial coordinate system. This applies to arbitrary discrete time slots. The position of the target drone relative to the base station is determined by the azimuth angle. and elevation angle Indicates the k-th camera. The acquired images are denoted as The first camera It is usually located at the same site as the base station and is defined as the main camera.

[0052] The communication and sensing link between the base station and the drone shares spectrum resources. The base station transmits signals and receives reflected echoes from the drone; the signal-to-noise ratio (SNR) of the echo signals is affected by environmental obstruction and clutter. To achieve beamforming, the base station needs to generate a beamforming vector. Its array steering vector corresponding to the direction of the UAV Closely related. Guiding vector The expression is:

[0053] ;

[0054] Where D is the spacing between antenna array elements. Where is the carrier wavelength, and m and n represent the row and column indices of the array, respectively. Beamforming vectors are typically set as guide vectors in the prediction direction, i.e.

[0055] ;

[0056] in( ( ) represents the predicted angle of the drone in the next time slot.

[0057] The system's final performance metric is the achievable communication rate C, which is related to the received signal-to-noise ratio G, and is expressed as:

[0058] ;

[0059] in,

[0060] ;

[0061] here, To transmit SNR, For the complex coefficients of the LOS path, The actual angle of time slot L. This represents the conjugate transpose of the guiding vector.

[0062] In each working time slot i, the system performs two types of data acquisition in parallel:

[0063] 1) The base station transmits detection signals and receives reflected echoes from the drone;

[0064] 2) All K distributed cameras are triggered synchronously to acquire images. All data (echo signals and image sequences) are accompanied by precise timestamps and aligned by the system to form a time-slot aligned multimodal sensing dataset. And the corresponding echo signal.

[0065] S3. Fusing visual information from multiple cameras, estimate the drone's orientation relative to the base station at the current moment; this step uses a state estimation model (SEM) to fuse visual information from multiple cameras and outputs the estimated azimuth angle of the drone relative to the base station in the current time slot i. and elevation angle . Figure 2 The structure of the state estimation model is described.

[0066] This step includes the following sub-steps:

[0067] S31: In the images captured by each camera, the drone is automatically identified and its center is located; then, a local image area containing the drone is cropped based on this center point.

[0068] Specifically, for each image The system uses target detection algorithms (such as the YOLO series) to identify drones and outputs their normalized center coordinates. Centered on this coordinate, a fixed-size image region is cropped out according to a preset ratio r. The cropping area is:

[0069]

[0070] Where H and W are the original image dimensions.

[0071] S32: Process the images cropped from each camera and the location information detected by each camera, extract features, and use a neural network to uniformly transform these features into the observation view of the base station;

[0072] Specifically, this includes visual feature extraction and orientation alignment for cropped images. Visual features are extracted using convolutional networks. At the same time, coordinates Encoded via linear layer After concatenating the two, a linear layer is used to map them to a latent space vector aligned with the base station's observation viewpoint. ;

[0073] S33: Based on the clarity and completeness of the image content provided by each camera, automatically calculate and assign different fusion weights, and perform weighted merging of the converted features from all cameras;

[0074] Specifically: S33: Generate fusion weights based on image features from the main camera (k=1).

[0075] ;

[0076] in, This represents a linear layer. The weights of the remaining cameras are distributed proportionally based on their features:

[0077] ;

[0078] The fused features are obtained by weighted summation of all aligned latent space vectors.

[0079] ;

[0080] S34: Based on the fused features, the current azimuth and elevation angles of the UAV are finally calculated.

[0081] Specifically, this means: integrating features Input a linear regression layer, and it directly outputs the estimated angle:

[0082] ;

[0083] S4. Fusing the visual estimation results from historical time slots with the wireless echo sensing results, predict the state of the UAV in the next time slot; This step uses a State Prediction Model (SPM) to fuse the visual estimation results from historical time slots with the echo sensing results to predict the state of the UAV in the next time slot L. . Figure 2 The structure of the state prediction model is described.

[0084] This step specifically includes the following sub-steps:

[0085] S41: For each historical moment, evaluate the reliability of visual perception and echo perception separately, and select the more reliable perception result at that moment as the final state information adopted at that moment;

[0086] Specifically, for each historical time slot i, calculate the confidence score of visual perception. Set a threshold. Generate decision gating variables The fusion state of this time slot is then:

[0087] ;

[0088] ;

[0089] in The angle is estimated based on the echo (such as the MUSIC algorithm).

[0090] S42: Arrange the final state information of a series of historical moments in chronological order and input it into a temporal neural network; the network analyzes the trend and pattern of state changes over time;

[0091] Specifically, this involves fusing the state sequences of L historical time slots. As input, add position encoding to the state embedding vector for each time slot. , to obtain the input vector

[0092] ;

[0093] will sequence Input a Transformer encoder and use its self-attention mechanism to model long-term temporal dependencies.

[0094] S43: Using the output of a temporal neural network, predict the change in the state of the UAV from the current moment to the next moment, and then calculate the predicted azimuth and elevation angles for the next moment.

[0095] Specifically, it involves retrieving the hidden representation corresponding to the last time slot output by the Transformer encoder. The state increment is predicted using a multilayer perceptron regression head. The final predicted state is:

[0096] ;

[0097] S5. Based on the UAV's next time slot state predicted in step S4, calculate and generate the optimal transmission beam for the base station antenna in advance. This includes active beamforming and communication performance enhancement, based on the UAV's next time slot state predicted in step S4. Calculate the corresponding array steering vector Based on this, beamforming vectors are constructed. Applying this beamforming vector to the downlink transmit signal of the base station allows the main lobe of the beam to be aligned in advance with the predicted position of the UAV in the next time slot, thereby maximizing the received signal-to-noise ratio G and ultimately improving the achievable communication rate C of the system.

[0098] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A distributed multimodal unmanned aerial vehicle (UAV) angle perception and prediction method, characterized in that, Including the following steps: S1. Determine the system model; S2. Determine the models for visual and radio frequency signals; S3. By fusing visual information from multiple cameras, the direction of the drone relative to the base station at the current moment is estimated. S4. By fusing the visual estimation results from historical time slots with the wireless echo sensing results, the state of the drone at the next moment can be predicted. S5. Based on the UAV state predicted in step S4, calculate and generate the optimal transmission beam for the base station antenna in advance.

2. The distributed multimodal UAV angle perception and prediction method according to claim 1, characterized in that, S1 specifically includes: deploying an integrated sensing and communication base station and K distributed cameras within the target low-altitude urban area, where K ≥ 2. The base station is equipped with a… A uniform planar array antenna with 1000 elements is used for simultaneous communication and echo-based sensing; K cameras are distributed in space to provide all-round, blind-spot-free visual coverage of the UAV's flight airspace. The base station and all cameras achieve time synchronization through a high-precision synchronization mechanism and complete the initial system access.

3. The distributed multimodal UAV angle perception and prediction method according to claim 2, characterized in that, S2 specifically includes the following steps: The system establishes a unified spatial coordinate system; in any discrete time slot The position of the target drone relative to the base station is determined by the azimuth angle. and elevation angle Indicates; the One camera The acquired images are denoted as The first camera It is usually located at the same site as the base station and is defined as the main camera; The communication and sensing links between the base station and the drone share spectrum resources; the base station transmits signals and receives reflected echoes from the drone. The signal-to-noise ratio of the echo signals is affected by environmental obstruction and clutter. The base station generates a beamforming vector. To achieve beamforming, the array steering vector corresponding to the direction of the UAV... Closely related, guide vector The expression is: ; Where D is the spacing between antenna array elements. For carrier wavelength, and These represent the row index and column index of the array, respectively. The beamforming vector is typically set as the guide vector in the prediction direction, i.e. ; in( () represents the predicted angle of the UAV in the next time slot; The system's final performance metric is the achievable communication rate C, which is related to the received signal-to-noise ratio G, and is expressed as: ; in, ; here, To transmit SNR, For the complex coefficients of the LOS path, The actual angle of time slot L. This represents the conjugate transpose of the guiding vector; In each work slot The system performs two types of data acquisition in parallel:

1. The base station transmits detection signals and receives reflected echoes from the drone; 2. All K distributed cameras are triggered synchronously to acquire images. All data is accompanied by precise timestamps and aligned by the system to form a time-slot aligned multimodal sensing dataset. And the corresponding echo signal.

4. The distributed multimodal UAV angle perception and prediction method according to claim 3, characterized in that, S3 specifically includes: S31: In the images captured by each camera, the drone is automatically identified and its center is located; then, a local image area containing the drone is cropped based on this center point. S32: Process the images cropped from each camera and the location information detected by each camera, extract features, and use a neural network to uniformly transform these features into the observation view of the base station; S33: Based on the clarity and completeness of the image content provided by each camera, automatically calculate and assign different fusion weights, and perform weighted merging of the converted features from all cameras; S34: Based on the fused features, the current azimuth and elevation angles of the UAV are finally calculated.

5. The distributed multimodal UAV angle perception and prediction method according to claim 4, characterized in that, S4 specifically includes: S41: For each historical moment, evaluate the reliability of visual perception and echo perception separately, and select the more reliable perception result at that moment as the final state information adopted at that moment; S42: Arrange the final state information of a series of historical moments in chronological order and input it into a temporal neural network; the network analyzes the trend and pattern of state changes over time; S43: Using the output of a temporal neural network, predict the change in the state of the UAV from the current moment to the next moment, and then calculate the predicted azimuth and elevation angles for the next moment.