Unmanned aerial vehicle multi-mode dynamic target decoupling system

By using a multimodal synchronization interface and dynamic Transformer perception technology, the synchronization and decoupling problems in UAV multimodal perception are solved, achieving high-precision data alignment and occlusion prediction, and improving the perception accuracy and robustness of UAVs in complex dynamic scenarios.

CN121878713APending Publication Date: 2026-04-17WUHAN VOCATIONAL COLLEGE OF SOFTWARE & ENG (WUHAN OPEN UNIV)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN VOCATIONAL COLLEGE OF SOFTWARE & ENG (WUHAN OPEN UNIV)
Filing Date
2026-01-06
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing UAV multimodal perception technologies suffer from problems such as insufficient multimodal data synchronization and alignment accuracy, limited dynamic target motion coupling processing capability, and lack of occlusion prediction and compensation mechanisms in complex dynamic scenarios, resulting in poor perception accuracy and robustness.

Method used

Employing a multimodal synchronization interface, an RGB-LiDAR-IMU synchronization module, a cross-modal spatiotemporal alignment unit, a dynamic Transformer sensing chip, a dynamic feature decoupling module, and an occlusion prediction unit, this system achieves high-precision synchronization, causal decoupling, and occlusion prediction of multimodal data through event-triggered interface compatibility, neuromorphic spiking network synchronization, causal cross-modal alignment, a dynamic Transformer architecture, and a diffusion generation model.

Benefits of technology

It significantly improves the synchronization and alignment accuracy of multimodal data, reduces the recognition error rate, enhances the perception accuracy and robustness of UAVs in complex dynamic scenarios, and strengthens the system's adaptability and real-time response capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121878713A_ABST
    Figure CN121878713A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle multi-modal dynamic target decoupling system, and relates to the technical field of unmanned aerial vehicle multi-modal sensing, and the system comprises a multi-modal synchronization interface, an RGB-LiDAR-IMU synchronization module, a cross-modal space-time alignment unit, a dynamic Transform sensing chip, a dynamic feature decoupling module, and a shielding prediction unit. The multi-modal synchronous interface is connected with the RGB camera, the LiDAR sensor and the IMU and is used for integrating multi-modal data input and executing event-triggered interface compatibility, and asynchronous data aggregation is realized by detecting key events in data streams. And the synchronization and alignment precision of the multi-modal data is obviously improved. The system detects data stream key events in a multi-modal synchronous interface to realize asynchronous aggregation, simulates biological neuron discharge in an RGB-LiDAR-IMU synchronous module to adjust modal delay, and realizes sub-microsecond time alignment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal perception technology for unmanned aerial vehicles (UAVs), specifically to a multimodal dynamic target decoupling system for UAVs. Background Technology

[0002] As drone technology evolves towards greater autonomy and intelligence, higher demands are being placed on dynamic target perception. Real-time tracking, adaptability, and robustness in complex dynamic scenarios have become a focus of industry attention. For example, low-altitude security requires accurate identification of multi-target intrusion trajectories, traffic monitoring needs to separate vehicle and pedestrian movements to analyze traffic flow, and logistics obstacle avoidance requires proactive avoidance of dynamic obstacles. Currently, multimodal perception technology is widely used in drone vision systems.

[0003] However, existing UAV target perception technologies face the following key technical challenges in meeting these requirements: Multimodal data synchronization and alignment accuracy is insufficient. Traditional synchronization methods mostly rely on software timestamps or simple hardware triggers, such as GPS clock-based synchronization mechanisms. The time alignment error is on the order of milliseconds, which makes it difficult to handle data delays and heterogeneity under high-speed motion, resulting in inconsistent fusion features and affecting the overall perception accuracy. The dynamic target motion coupling processing capability is limited. Existing perception algorithms such as standard CNN or RNN usually treat multiple targets as a whole and cannot effectively separate topological coupling. For example, in crowded scenes, the motion interference of vehicles and pedestrians can lead to an identification error rate of more than 20%, which limits the reliability of independent tracking. The lack of occlusion prediction and compensation mechanisms means that traditional methods rely on historical frame interpolation or simple filtering, such as trajectory prediction based on Kalman filters, which has an interruption rate as high as 15% under complex occlusion. The lack of simulation and pre-compensation for future scenes results in poor robustness of the system in dynamic environments.

[0004] Therefore, a multimodal dynamic target decoupling system for UAVs is needed to solve the above problems. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a multimodal dynamic target decoupling system for unmanned aerial vehicles (UAVs), which solves the problems mentioned in the background section.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solution: a multimodal dynamic target decoupling system for unmanned aerial vehicles, including a multimodal synchronization interface, an RGB-LiDAR-IMU synchronization module, a cross-modal spatiotemporal alignment unit, a dynamic Transformer perception chip, a dynamic feature decoupling module, and an occlusion prediction unit; The multimodal synchronization interface connects the RGB camera, LiDAR sensor, and IMU inertial measurement unit, and is used to integrate multimodal data input and perform event-triggered interface compatibility, achieving asynchronous data aggregation by detecting key events in the data stream; The RGB-LiDAR-IMU synchronization module is connected to the multimodal synchronization interface and is used to perform biomimetic synchronization of RGB image data, LiDAR point cloud data and IMU attitude data by applying a neuromorphic pulse network, simulating the firing mechanism of biological neurons to achieve sub-microsecond time alignment. The cross-modal spatiotemporal alignment unit is connected to the RGB-LiDAR-IMU synchronization module. It adopts a causal cross-modal alignment structure to construct an intermodal causal graph model, perform causal inference fusion on the synchronized multimodal data, and generate a dynamically adaptive spatiotemporal feature map. The dynamic Transformer perception chip is connected to the cross-modal spatiotemporal alignment unit and is used to deploy a meta-learning-enhanced Transformer architecture. It processes the spatiotemporally aligned feature data through an online meta-optimization mechanism and adaptively adjusts the attention weights to capture the nonlinear spatiotemporal dependencies of multiple targets. The dynamic feature decoupling module is connected to the dynamic Transformer sensing chip and is used to introduce a spectral graph convolutional network to separate the motion features of multiple targets. It decomposes the topological coupling between targets through frequency domain graph signal processing to achieve causal decoupling of independent motion vectors. The occlusion prediction unit is connected to the dynamic feature decoupling module and is used to predict occlusion events based on the diffusion generation model, simulate future multimodal scenarios through the conditional diffusion process, and perform pre-compensation fusion.

[0007] Preferably, the multimodal synchronization interface includes an event detector and an asynchronous aggregation buffer. The event detector is used to identify change event thresholds in RGB image data, LiDAR point cloud data, and IMU attitude data. The asynchronous aggregation buffer is used to dynamically aggregate data streams based on event priorities and transmit them to the RGB-LiDAR-IMU synchronization module through data interaction.

[0008] Preferably, the neuromorphic pulse network of the RGB-LiDAR-IMU synchronization module includes a biomimetic synaptic connection layer and a pulse timing encoder. The biomimetic synaptic connection layer is used to simulate plastic connections to adjust modal delays, and the pulse timing encoder is used to convert continuous data into pulse sequences to achieve synchronization, and to provide the synchronization data to the cross-modal spatiotemporal alignment unit through instruction transmission.

[0009] Preferably, the causal cross-modal alignment structure of the cross-modal spatiotemporal alignment unit includes a causal discovery network and an intervention fusion module. The causal discovery network is used to mine causal relationship graphs from multimodal data, and the intervention fusion module is used to generate a unified spatiotemporal feature representation through causal intervention operations and transmit the spatiotemporal feature graph to the dynamic Transformer sensing chip through data interaction.

[0010] Preferably, the meta-learning enhanced Transformer architecture of the dynamic Transformer perception chip includes an inner attention optimizer and an outer meta-gradient updater. The inner attention optimizer is used to fine-tune the self-attention matrix in real time, and the outer meta-gradient updater is used to learn general weight adaptation based on task distribution and send the processed feature data to the dynamic feature decoupling module through instruction transmission.

[0011] Preferably, the spectral convolutional network of the dynamic feature decoupling module includes a frequency domain graph filter and a coupling decomposer. The frequency domain graph filter is used to convert spatiotemporal features into spectral domain signals, and the coupling decomposer is used to separate independent motion components through graph Laplacian eigenvalue decomposition and transmit the decoupled motion vector to the occlusion prediction unit through data interaction.

[0012] Preferably, the diffusion generation model of the occlusion prediction unit includes a conditional noise injector and an inverse diffusion decoder. The conditional noise injector is used to inject current features as conditions to generate potential occlusion noise, and the inverse diffusion decoder is used to gradually recover the prediction scene for compensation.

[0013] Preferably, the system further includes a self-evolving feedback loop, which connects the occlusion prediction unit and the cross-modal spatiotemporal alignment unit. The self-evolving feedback loop is used to dynamically optimize the causal graph model parameters through a reinforcement learning agent and to realize a collaborative mechanism between the alignment unit and the prediction unit through a communication connection.

[0014] Preferably, the dynamic feature decoupling module integrates a causal spectral clustering algorithm to handle the dynamic coupling of vehicle and pedestrian targets, and achieves real-time decoupling in collaboration with the coupling decomposer through spectral domain causal intervention operations.

[0015] Preferably, the occlusion prediction unit is combined with a multimodal diffusion path simulator to generate a predicted trajectory field of dynamic obstacles, and provides obstacle avoidance decision input through instruction transmission between the inverse diffusion decoder and the simulator.

[0016] Beneficial effects This invention provides a multimodal dynamic target decoupling system for unmanned aerial vehicles (UAVs). It has the following beneficial effects: 1. This invention significantly improves the synchronization and alignment accuracy of multimodal data by introducing an event-triggered interface compatible with a synchronization mechanism based on neuromorphic spiking networks. The system detects key events in the data stream in the multimodal synchronization interface to achieve asynchronous aggregation, and simulates the firing of biological neurons in the RGB-LiDAR-IMU synchronization module to adjust modal delays, achieving sub-microsecond time alignment. This high-precision synchronization effectively handles data heterogeneity under high-speed motion, ensuring the consistency of fused features, thereby improving overall perception accuracy, reducing tracking errors caused by latency in complex dynamic scenarios, and enhancing the real-time response capability of UAVs.

[0017] 2. This invention employs a decoupled structure of a dynamic Transformer sensing chip and a spectral graph convolutional network, effectively enhancing the processing capability of dynamic target motion coupling. The meta-learning-enhanced Transformer architecture captures the nonlinear spatiotemporal dependencies of multiple targets, and frequency domain graph signal processing is introduced into the dynamic feature decoupling module to decompose topological coupling, achieving causal decoupling of independent motion vectors. This decoupling mechanism can separate vehicle and pedestrian motion interference in congested scenes, reducing the recognition error rate and achieving reliable independent tracking, thereby improving the system's adaptability and multi-target management efficiency in applications such as traffic monitoring.

[0018] 3. This invention integrates an occlusion prediction unit based on a diffusion generation model, providing an advanced occlusion prediction and compensation mechanism. By simulating future multimodal scenarios through a conditional diffusion process and performing pre-compensation fusion, this unit, combined with a multimodal diffusion path simulator, generates a dynamic obstacle prediction trajectory field, overcoming the shortcomings of traditional methods that rely on historical frames. This mechanism significantly reduces the tracking interruption rate under complex occlusion, improves the system's robustness in dynamic environments, and enables safer navigation planning in scenarios such as logistics obstacle avoidance. Attached Figure Description

[0019] Figure 1 This is a system framework diagram of the present invention; Figure 2 This is a system flowchart of the present invention; Figure 3 This is a structural diagram of the occlusion prediction unit of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Specific Implementation Example 1: like Figures 1 to 3As shown, the present invention provides a multimodal dynamic target decoupling system for unmanned aerial vehicles, including a multimodal synchronization interface, an RGB-LiDAR-IMU synchronization module, a cross-modal spatiotemporal alignment unit, a dynamic Transformer perception chip, a dynamic feature decoupling module, and an occlusion prediction unit; The multimodal synchronization interface connects an RGB camera, a LiDAR sensor, and an IMU inertial measurement unit to integrate multimodal data input and perform event-triggered interface compatibility. It achieves asynchronous data aggregation by detecting key events in the data stream. The RGB-LiDAR-IMU synchronization module connects to a multimodal synchronization interface, which is used to perform biomimetic synchronization of RGB image data, LiDAR point cloud data and IMU pose data using a neuromorphic pulse network, simulating the firing mechanism of biological neurons to achieve sub-microsecond time alignment. The cross-modal spatiotemporal alignment unit is connected to the RGB-LiDAR-IMU synchronization module. It adopts a causal cross-modal alignment structure to construct an intermodal causal graph model, perform causal inference fusion on the synchronized multimodal data, and generate a dynamically adaptive spatiotemporal feature map. The dynamic Transformer perception chip connects to a cross-modal spatiotemporal alignment unit to deploy a meta-learning-enhanced Transformer architecture. It processes the spatiotemporally aligned feature data through an online meta-optimization mechanism and adaptively adjusts attention weights to capture the nonlinear spatiotemporal dependencies of multiple targets. The dynamic feature decoupling module is connected to the dynamic Transformer sensing chip. It is used to introduce a spectral graph convolutional network to separate the motion features of multiple targets. Through frequency domain graph signal processing, it decomposes the topological coupling between targets and realizes the causal decoupling of independent motion vectors. The occlusion prediction unit is connected to the dynamic feature decoupling module and is used to predict occlusion events based on the diffusion generation model. It simulates future multimodal scenarios through the conditional diffusion process and performs pre-compensation fusion.

[0022] The multimodal synchronization interface includes an event detector and an asynchronous aggregation buffer. The event detector is used to identify change event thresholds in RGB image data, LiDAR point cloud data, and IMU attitude data. The asynchronous aggregation buffer is used to dynamically aggregate data streams based on event priorities and transmit them to the RGB-LiDAR-IMU synchronization module through data interaction.

[0023] The neuromorphic pulse network of the RGB-LiDAR-IMU synchronization module includes a biomimetic synaptic connection layer and a pulse timing encoder. The biomimetic synaptic connection layer is used to simulate plastic connections to adjust modal delays, and the pulse timing encoder is used to convert continuous data into pulse sequences to achieve synchronization. The synchronized data is then provided to the cross-modal spatiotemporal alignment unit via instruction transmission.

[0024] The causal cross-modal alignment structure of the cross-modal spatiotemporal alignment unit includes a causal discovery network and an intervention fusion module. The causal discovery network is used to mine causal relationship graphs from multimodal data, and the intervention fusion module is used to generate a unified spatiotemporal feature representation through causal intervention operations and transmit the spatiotemporal feature map to the dynamic Transformer sensing chip through data interaction.

[0025] The meta-learning enhanced Transformer architecture of the dynamic Transformer perception chip includes an inner attention optimizer and an outer meta-gradient updater. The inner attention optimizer is used to fine-tune the self-attention matrix in real time, and the outer meta-gradient updater is used to learn general weight adaptation based on task distribution. The processed feature data is sent to the dynamic feature decoupling module through instruction transmission.

[0026] The spectral convolutional network of the dynamic feature decoupling module includes a frequency domain graph filter and a coupling decomposer. The frequency domain graph filter is used to convert spatiotemporal features into spectral domain signals, and the coupling decomposer is used to separate independent motion components through graph Laplacian eigenvalue decomposition and transmit the decoupled motion vector to the occlusion prediction unit through data interaction.

[0027] The diffusion generation model of the occlusion prediction unit includes a conditional noise injector and an inverse diffusion decoder. The conditional noise injector is used to inject the current features as conditions to generate potential occlusion noise, and the inverse diffusion decoder is used to gradually recover the prediction scene for compensation.

[0028] The system further includes a self-evolving feedback loop, which connects the occlusion prediction unit and the cross-modal spatiotemporal alignment unit. This loop is used to dynamically optimize the causal graph model parameters through a reinforcement learning agent and to realize a collaborative mechanism between the alignment unit and the prediction unit through a communication connection.

[0029] The dynamic feature decoupling module integrates a causal spectral clustering algorithm to handle the dynamic coupling of vehicle and pedestrian targets. It achieves real-time decoupling through spectral domain causal intervention operations in collaboration with the coupling decomposer.

[0030] The occlusion prediction unit, combined with a multimodal diffusion path simulator, is used to generate a predicted trajectory field for dynamic obstacles. It provides obstacle avoidance decision input through instruction transmission between the inverse diffusion decoder and the simulator. Specific Implementation Example 2: like Figures 1 to 3As shown, this system is suitable for complex and dynamic scenarios such as low-altitude security patrols, traffic monitoring, or obstacle avoidance for logistics drones, enabling independent motion decoupling perception of multiple targets. The system integrates multimodal data input from an RGB camera, LiDAR sensor, and IMU inertial measurement unit, performing a series of processing steps including synchronization, alignment, perception, decoupling, and prediction to address recognition and tracking errors caused by target occlusion and motion coupling in complex dynamic scenarios. The following sections provide a detailed explanation of the system's structure, the working principles of each component, connection methods, data interaction processes, and optimization mechanisms.

[0032] The system's overall architecture includes a multimodal synchronization interface, an RGB-LiDAR-IMU synchronization module, a cross-modal spatiotemporal alignment unit, a dynamic Transformer sensing chip, a dynamic feature decoupling module, and an occlusion prediction unit. These components are sequentially connected via a high-speed data bus or a dedicated communication interface, forming an ordered processing pipeline. Specifically, the multimodal synchronization interface is directly electrically connected to the RGB camera, LiDAR sensor, and IMU inertial measurement unit for raw data acquisition; the RGB-LiDAR-IMU synchronization module is connected to the multimodal synchronization interface via a data interaction port; the cross-modal spatiotemporal alignment unit is connected to the RGB-LiDAR-IMU synchronization module via a high-speed data bus; the dynamic Transformer sensing chip is connected to the cross-modal spatiotemporal alignment unit via a dedicated chip interface; the dynamic feature decoupling module is connected to the dynamic Transformer sensing chip via an internal data bus; and the occlusion prediction unit is connected to the dynamic feature decoupling module via a command transmission channel. Furthermore, the system includes a self-evolving feedback loop, which connects the occlusion prediction unit and the cross-modal spatiotemporal alignment unit via a communication connection to achieve closed-loop optimization.

[0033] At the initial stage of the data processing flow, the multimodal synchronization interface, serving as the system's entry module, is responsible for integrating multimodal data inputs from the RGB camera, LiDAR sensor, and IMU (Inertial Measurement Unit). This interface first receives 2D image data (containing color and texture information, transmitted as a frame sequence) from the RGB camera, 3D point cloud data (containing distance and depth information, transmitted as a point cloud) from the LiDAR sensor, and attitude data (containing acceleration, angular velocity, and orientation information, transmitted as a high-frequency sequence) from the IMU via electrical connections. Then, the interface executes an event-triggered interface-compatible mechanism: an internally integrated event detector monitors key change events in each data stream in real time, such as significant pixel changes in the RGB image, abrupt changes in point density in the LiDAR point cloud, or significant changes in motion intensity in the IMU attitude data. Once these key events are detected, the event detector generates a trigger signal, activating an asynchronous aggregation buffer. This asynchronous aggregation buffer employs a priority queue structure, dynamically aggregating data streams according to the importance of events: prioritizing motion-related high-priority event data, placing it at the front of the buffer, and maintaining the temporal order of data segments. Finally, the aggregated asynchronous data stream is sent to the RGB-LiDAR-IMU synchronization module via data interaction. This event-triggered design avoids the limitations of traditional fixed clock synchronization and improves the system's ability to respond quickly to dynamic scenes.

[0034] Next, the RGB-LiDAR-IMU synchronization module receives asynchronous data streams from the multimodal synchronization interface and applies a neuromorphic spiking network for biomimetic synchronization. The module's internal structure includes a biomimetic synaptic connection layer and a pulse timing encoder. First, the biomimetic synaptic connection layer simulates the plasticity of connections in biological nervous systems, dynamically adjusting the intermodal connection strength to compensate for delay differences between different sensors, making the data more temporally consistent. Second, the pulse timing encoder converts continuous sensor data into a pulse sequence: RGB image data is converted into pulse frequency based on brightness changes, LiDAR point cloud data into pulse phase based on depth changes, and IMU attitude data into pulse intervals based on motion intensity. By simulating the firing mechanism of biological neurons, this encoder achieves high-precision alignment of all modal data along the time axis. After synchronization, the module provides the synchronized data to the cross-modal spatiotemporal alignment unit via command transmission, including a time-consistent RGB frame sequence, LiDAR point cloud sequence, and IMU attitude sequence.

[0035] Subsequently, the cross-modal spatiotemporal alignment unit receives the synchronized data and processes it using a causal cross-modal alignment structure. The core of this structure is the construction of an intermodal causal relationship model. First, a causal discovery network is used to mine causal relationships between different modalities from multimodal data, forming a causal relationship graph. Then, the intervention fusion module performs interventional fusion operations based on this causal relationship graph, simulating the impact of one modality change on other modalities, thereby generating a unified spatiotemporal feature representation. Specifically, RGB features are mapped to a spatiotemporal grid structure, LiDAR point clouds are voxelized and combined with IMU pose for coordinate transformation, and the features of each modality are weighted and fused through causal intervention to generate a dynamically adaptive spatiotemporal feature map. Finally, this spatiotemporal feature map is transmitted to the dynamic Transformer sensing chip through data interaction. This causal reasoning fusion method effectively addresses the heterogeneity problem of multimodal data.

[0036] The dynamic Transformer perception chip receives spatiotemporal feature maps and processes them using a meta-learning-enhanced Transformer architecture. This architecture includes an inner attention optimizer and an outer meta-gradient updater. First, the inner attention optimizer fine-tunes the self-attention mechanism in real time, making the attention distribution more adaptable to the specific features of the current scene. Second, the outer meta-gradient updater learns a general weight adaptation strategy based on various task distributions, enabling the model to quickly adjust in different dynamic environments. Through this online meta-optimization mechanism, the chip adaptively adjusts attention weights, deeply capturing the nonlinear spatiotemporal dependencies between multiple targets, such as highlighting the feature associations of fast-moving targets. The processed feature data is sent to the dynamic feature decoupling module via instruction transmission.

[0037] The dynamic feature decoupling module receives the processed feature data and introduces a spectral graph convolutional network to separate the motion features of multiple targets. This network includes a frequency domain graph filter and a coupling decomposer. First, the frequency domain graph filter constructs a graph structure between targets and converts the spatiotemporal features into frequency domain signals. Second, the coupling decomposer separates the low-frequency independent motion component and the high-frequency coupling interference component based on graph Laplacian eigenvalue decomposition, thereby achieving causal decoupling of the independent motion vectors of each target. In addition, this module integrates a causal spectral clustering algorithm to handle the dynamic coupling of vehicle and pedestrian targets. It works in conjunction with the coupling decomposer through spectral domain causal intervention to achieve real-time decoupling. The decoupled motion vectors are transmitted to the occlusion prediction unit through data interaction.

[0038] The occlusion prediction unit receives decoupled motion vectors and predicts occlusion events based on a diffusion generation model. This model includes a conditional noise injector and an inverse diffusion decoder. First, the conditional noise injector progressively adds noise based on the current decoupled features, generating a sequence of potential occlusion perturbations. Second, the inverse diffusion decoder progressively reconstructs the future multimodal scene through a multi-step inverse denoising process, achieving simulation and pre-compensation fusion of potential occlusions. Furthermore, this unit combines a multimodal diffusion path simulator to generate a predicted trajectory field of dynamic obstacles, providing probabilistic trajectory information for obstacle avoidance decisions through instruction transmission between the inverse diffusion decoder and the simulator.

[0039] Finally, the system further includes a self-evolving feedback loop connecting the occlusion prediction unit and the cross-modal spatiotemporal alignment unit, used to dynamically optimize the causal relationship model parameters through a reinforcement learning agent. Specifically, the feedback loop collects error information from the prediction results of the occlusion prediction unit and inputs it as a reward signal into the reinforcement learning agent; the agent generates optimization instructions based on these signals and adjusts the causal graph structure in the cross-modal spatiotemporal alignment unit through communication connections, achieving closed-loop collaborative optimization between the alignment unit and the prediction unit. This self-evolving mechanism significantly improves the system's long-term adaptability in complex environments. Specific Implementation Example 3: like Figures 1 to 3 As shown, the module components mentioned above will be described in detail below: Multimodal synchronization interface: The multimodal synchronization interface, serving as the system's data entry point and initial integration module, is primarily responsible for receiving and initially processing multimodal data inputs from the RGB camera, LiDAR sensor, and IMU (Inertial Measurement Unit). The interface's hardware comprises an FPGA-based integrated circuit board equipped with multiple input ports: a USB interface for connecting the RGB camera, an Ethernet interface for the LiDAR sensor, and a serial peripheral interface for the IMU. Additionally, it integrates event detector hardware circuitry and an asynchronous aggregation buffer. Hardware specifications: The FPGA chip on the integrated circuit board provides reprogrammable logic, supports high-speed data throughput, and ensures interface compatibility with different sensor protocols; the event detector circuitry uses a threshold comparator to identify key events in the data stream in real time, such as pixel changes or sudden changes in point cloud density; the asynchronous aggregation buffer uses priority queue logic circuitry to dynamically manage data stream priorities, preventing data loss. The RGB camera is a high-resolution visible light sensor employing a 2 / 3-inch CMOS image sensor with 20 million effective pixels, an 84° field of view, an equivalent focal length of 24 mm, and an aperture range from f / 2.8 to f / 11. It supports focusing distances from 1 m to infinity and provides high frame rate image capture to capture dynamic scene details. The LiDAR sensor is a lidar device that uses a laser pulse emitter and receiver to generate 3D point cloud data by measuring the time it takes for the laser to travel from emission to return. Typical scanning ranges reach hundreds of meters, with point cloud densities of millions of points per second, supporting high-precision distance measurement and environmental modeling. The IMU (Inertial Measurement Unit) is a multi-axis motion sensor integrating a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer. The gyroscope operates within a range of ±550° / s, and the acceleration measurement range is ±16g. It features high-precision self-calibration and temperature robustness, operating from -40°C to +71°C, with power consumption less than 10 W.

[0041] The module takes into account a 2D image data stream from an RGB camera, a 3D point cloud data stream from a LiDAR sensor, and an attitude data stream from an IMU. The output is a pre-aggregated asynchronous multimodal data stream, including timestamped event-triggered data packets. Data transmission path: Sensor data is directly input to the FPGA port via electrical connection. The event detector circuit scans the input signal to generate trigger pulses, activating the asynchronous aggregation buffer for priority sorting and aggregation. The aggregated data is transmitted to the RGB-LiDAR-IMU synchronization module via an internal high-speed bus to ensure continuity with subsequent modules and avoid independent operation.

[0042] RGB-LiDAR-IMU synchronization module: The RGB-LiDAR-IMU synchronization module connects to a multimodal synchronization interface for high-precision time synchronization of initially aggregated data. The module's hardware components include a neuromorphic pulse network circuit board, integrating a biomimetic synaptic connection layer and a pulse timing encoder. It also includes an auxiliary clock synchronization crystal oscillator and a power management module. Hardware description: The neuromorphic chip in this circuit board simulates the firing of biological neurons, adjusting connection strength through a time-dependent plasticity mechanism to compensate for modal delay; the pulse timing encoder converts continuous signals into pulse sequences, supporting sub-microsecond alignment accuracy; the crystal oscillator provides a unified reference clock to reduce jitter. The RGB camera is a high-resolution visible light sensor, employing a 2 / 3-inch CMOS image sensor with 20 million effective pixels, an 84° field of view, an equivalent focal length of 24 mm, an aperture range from f / 2.8 to f / 11, supports focusing distances from 1m to infinity, and provides high frame rate image capture to capture dynamic scene details. LiDAR sensors are lidar devices that use laser pulse emitters and receivers to generate 3D point cloud data by measuring the time it takes for the laser to travel from emission to return. Typical scanning ranges can reach hundreds of meters, with point cloud densities of millions of points per second, supporting high-precision distance measurement and environmental modeling. An IMU (Inertial Measurement Unit) is a multi-axis motion sensor integrating a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer. The gyroscope operates within a range of ±550° / s, and the acceleration measurement range is ±16 g. It features high-precision self-calibration and temperature robustness, operating from -40°C to +71°C, with power consumption less than 10 W.

[0043] The input is an asynchronous data stream from the multimodal synchronization interface; the output is a time-aligned multimodal data sequence, including pulse-coded synchronization frames. Data transmission path: Asynchronous data enters the neuromorphic circuit board through the data interaction port. The bionic synaptic connection layer first adjusts the delay, then the pulse timing encoder converts the signal into pulse form and aligns the timestamp. Synchronous data is provided to the cross-modal spatiotemporal alignment unit through the instruction transmission channel, ensuring seamless connection between the data stream and upstream and downstream units, forming a system-level continuous processing chain.

[0044] Cross-modal spatiotemporal alignment unit: The cross-modal spatiotemporal alignment unit connects to the RGB-LiDAR-IMU synchronization module, responsible for building causal relationship models and fusing synchronized data. The unit's hardware components include a dedicated AI accelerator board equipped with a causal discovery network processor and an intervention fusion module. Additionally, it includes high-bandwidth memory and a network interface controller. Hardware specifications: The accelerator board supports causal graph mining algorithms, passing conditional independence tests through parallel kernel processing; the digital signal processor performs intervention operations, achieving feature projection and fusion; high-bandwidth memory ensures low-latency access and supports real-time spatiotemporal grid construction. The RGB camera is a high-resolution visible light sensor, employing a 2 / 3-inch CMOS image sensor with 20 million effective pixels, an 84° field of view, an equivalent focal length of 24 mm, an aperture range from f / 2.8 to f / 11, supports focusing distances from 1 m to infinity, and provides high frame rate image capture to capture dynamic scene details. LiDAR sensors are lidar devices that use laser pulse emitters and receivers to generate 3D point cloud data by measuring the time it takes for the laser to travel from emission to return. Typical scanning ranges can reach hundreds of meters, with point cloud densities of millions of points per second, supporting high-precision distance measurement and environmental modeling. An IMU (Inertial Measurement Unit) is a multi-axis motion sensor integrating a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer. The gyroscope operates within a range of ±550° / s, and the acceleration measurement range is ±16 g. It features high-precision self-calibration and temperature robustness, operating from -40°C to +71°C, with power consumption less than 10 W.

[0045] The input is a pulse-coded data sequence provided by the synchronization module; the output is a unified dynamic adaptive spatiotemporal feature map. Data transmission path: Synchronization data enters the AI ​​accelerator board via a high-speed data bus. The causal discovery network processor first mines the relationship graph, and then intervenes in the fusion module to perform weighted fusion and coordinate transformation. The generated feature map is transmitted to the dynamic Transformer sensing chip through data interaction, maintaining a close data dependency with the preceding and following modules, ensuring that the fusion result directly affects downstream sensing.

[0046] Dynamic Transformer sensing chip: The dynamic Transformer perception chip connects to a cross-modal spatiotemporal alignment unit to process spatiotemporal features and capture target dependencies. The chip's hardware components include an edge AI processor and a meta-learning enhancement module. Specifically, it includes an inner-layer attention optimizer and an outer-layer meta-gradient updater. Furthermore, it integrates low-power memory and a thermal management system. Hardware description: The neural processing unit in the processor optimizes Transformer self-attention computation, supporting real-time weight fine-tuning; a dedicated accelerator implements the meta-learning loop, adapting weights based on task distribution; memory provides high-speed caching to reduce bottlenecks. The RGB camera is a high-resolution visible light sensor using a 2 / 3-inch CMOS image sensor with 20 million effective pixels, an 84° field of view, an equivalent focal length of 24 mm, an aperture range from f / 2.8 to f / 11, supports focusing distances from 1 m to infinity, and provides high frame rate image capture to capture dynamic scene details. LiDAR sensors are lidar devices that use laser pulse emitters and receivers to generate 3D point cloud data by measuring the time it takes for the laser to travel from emission to return. Typical scanning ranges can reach hundreds of meters, with point cloud densities of millions of points per second, supporting high-precision distance measurement and environmental modeling. An IMU (Inertial Measurement Unit) is a multi-axis motion sensor integrating a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer. The gyroscope operates within a range of ±550° / s, and the acceleration measurement range is ±16 g. It features high-precision self-calibration and temperature robustness, operating from -40°C to +71°C, with power consumption less than 10 W.

[0047] The input is the spatiotemporal feature map of the alignment unit; the output is the enhanced feature data embedding vector, including adjusted attention weights. Data transmission path: the feature map enters the processor through a dedicated chip interface, the inner optimizer fine-tunes the attention matrix, and the outer updater calculates the general adaptation; the processed data is sent to the dynamic feature decoupling module via instruction transmission, ensuring that the perception result serves as the direct input for decoupling, forming a continuous AI processing pipeline.

[0048] Dynamic feature decoupling module: The dynamic feature decoupling module connects to the dynamic Transformer sensing chip for separating multi-target motion features. The module's hardware components include an FPGA chip and a spectral graph convolutional network accelerator. Specifically, it includes a frequency domain graph filter and a coupling decomposer. Furthermore, it integrates causal spectral clustering algorithm circuitry and auxiliary storage. Hardware description: The FPGA chip in this module provides flexible graph signal processing, supporting frequency domain transformation and eigenvalue decomposition; the AI ​​engine accelerates spectral clustering, ensuring real-time decoupling; and a lookup table enables causal intervention operations. The RGB camera is a high-resolution visible light sensor using a 2 / 3-inch CMOS image sensor with 20 million effective pixels, an 84° field of view, an equivalent focal length of 24 mm, and an aperture range from f / 2.8 to f / 11. It supports focusing distances from 1 m to infinity and provides high frame rate image capture to capture dynamic scene details. The LiDAR sensor is a lidar device that uses a laser pulse emitter and receiver to generate 3D point cloud data by measuring the time it takes for the laser to travel from emission to return. Typical scanning ranges can reach hundreds of meters, with point cloud densities of millions of points per second, supporting high-precision distance measurement and environmental modeling. An IMU (Inertial Measurement Unit) is a multi-axis motion sensor that integrates a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer. The gyroscope operates within a range of ±550° / s, and the acceleration measurement range is ±16 g. It features high-precision self-calibration and temperature robustness, with an operating temperature range of -40°C to +71°C and a power consumption of less than 10 W.

[0049] The input is enhanced feature data from the sensing chip; the output is independent motion vectors, including decoupled trajectory components. Data transmission path: Feature data enters the FPGA via an internal data bus, a frequency domain graph filter constructs a graph structure and converts the signal, a coupling decomposer decomposes the signal and works in conjunction with the spectral clustering circuit; the decoupled vector is transmitted to the occlusion prediction unit via data interaction, emphasizing the integration dependency with upstream sensing and downstream prediction.

[0050] Occlusion prediction unit: The occlusion prediction unit, connected to the dynamic feature decoupling module, is used to predict occlusion events based on a diffusion model. The unit's hardware consists of a dedicated chip and a diffusion generation model processor, specifically a conditional noise injector and an inverse diffusion decoder. It also incorporates a multimodal diffusion path simulator and an output interface. Hardware description: The dedicated chip in this unit optimizes the diffusion process, supporting conditional noise generation and multi-step recovery; the simulator handles trajectory field prediction; and the graphics processing submodule accelerates sampling. The RGB camera is a high-resolution visible light sensor employing a 2 / 3-inch CMOS image sensor with 20 million effective pixels, an 84° field of view, an equivalent focal length of 24 mm, and an aperture range from f / 2.8 to f / 11. It supports focusing distances from 1 m to infinity and provides high frame rate image capture to capture dynamic scene details. The LiDAR sensor is a lidar device that uses a laser pulse emitter and receiver to generate 3D point cloud data by measuring the time it takes for the laser to travel from emission to return. Typical scanning ranges reach hundreds of meters, with point cloud densities of millions of points per second, supporting high-precision distance measurement and environmental modeling. An IMU (Inertial Measurement Unit) is a multi-axis motion sensor that integrates a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer. The gyroscope operates within a range of ±550° / s, and the acceleration measurement range is ±16 g. It features high-precision self-calibration and temperature robustness, with an operating temperature range of -40°C to +71°C and a power consumption of less than 10 W.

[0051] The input consists of independent motion vectors from the decoupled modules; the output is the predicted occlusion scene and the compensation fusion result, including the probability trajectory field. Data transmission path: Motion vectors enter the chip via the instruction transmission channel; a noise injector generates a perturbation sequence; the inverse diffusion decoder iteratively recovers the scene and collaborates with the simulator; the prediction result is fed back to the self-evolving feedback loop via a communication connection, providing obstacle avoidance decisions and ensuring closed-loop interaction with other modules in the system.

[0052] Self-evolving feedback loop: The self-evolving feedback loop connects the occlusion prediction unit and the cross-modal spatiotemporal alignment unit to optimize system parameters. The loop's hardware components include a reinforcement learning agent processor and a communication module. Additionally, it includes an error acquisition sensor and a parameter adjustment interface. Hardware description: The agent processor runs the agent algorithm to generate optimization instructions; the wireless module ensures low-latency feedback; the converter acquires prediction errors as reward signals. The RGB camera is a high-resolution visible light sensor using a 2 / 3-inch CMOS image sensor with 20 million effective pixels, an 84° field of view, an equivalent focal length of 24 mm, and an aperture range from f / 2.8 to f / 11. It supports focusing distances from 1 m to infinity and provides high frame rate image capture to capture dynamic scene details. The LiDAR sensor is a lidar device that uses a laser pulse emitter and receiver to generate 3D point cloud data by measuring the time it takes for the laser to travel from emission to return. Typical scanning ranges can reach hundreds of meters, with point cloud densities of millions of points per second, supporting high-precision distance measurement and environmental modeling. An IMU (Inertial Measurement Unit) is a multi-axis motion sensor that integrates a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer. The gyroscope operates within a range of ±550° / s, and the acceleration measurement range is ±16 g. It features high-precision self-calibration and temperature robustness, with an operating temperature range of -40°C to +71°C and a power consumption of less than 10 W.

[0053] The input is the error information of the occlusion prediction unit; the output is the optimization instruction, which adjusts the causal graph parameters. Data transmission path: The error enters the agent processor through the communication connection, and after calculation and optimization, it is transmitted to the alignment unit through the bus to achieve closed-loop collaboration; this path loops the prediction results back upstream to ensure that the entire system modules operate in a non-independent and interdependent manner. Specific Implementation Example 4: like Figures 1 to 3 As shown below, the algorithm architecture used in this system is described in detail: Neuromorphic spiking networks: Neuromorphic spiking networks are the core algorithm architecture of the RGB-LiDAR-IMU synchronization module, used to achieve biomimetic synchronization of multimodal data.

[0055] The input data is an asynchronous data stream from the multimodal synchronization interface, including RGB image data, LiDAR point cloud data, and IMU attitude data. This data is passed in as a raw sequence and may have time delay differences.

[0056] The output is a time-aligned multimodal data sequence, including synchronization frames converted to pulse form, where the data for each modality is precisely aligned on the time axis.

[0057] In the specific application of the system, synchronization is achieved by simulating the firing mechanism of biological neurons. First, the bionic synaptic connection layer detects and adjusts the time delay differences between modes to make the data streams more consistent. Then, the pulse timing encoder converts the continuous data signal into a discrete pulse sequence and encodes information according to the timing and frequency of the pulses. Finally, these pulse sequences are aligned to a unified time reference point, thereby providing consistent input data for the downstream cross-modal spatiotemporal alignment unit and ensuring that the entire system maintains data consistency in complex dynamic scenarios.

[0058] Causal cross-modal alignment structure: The causal cross-modal alignment structure is the core algorithm architecture of the cross-modal spatiotemporal alignment unit, used to fuse synchronized multimodal data.

[0059] The input data is a pulse-coded data sequence provided by the RGB-LiDAR-IMU synchronization module, including time-aligned RGB image frames, LiDAR point cloud sequences, and IMU attitude sequences.

[0060] The output is a unified dynamic adaptive spatiotemporal feature map, which is a multidimensional feature representation that integrates information from various modalities to reflect spatiotemporal relationships.

[0061] In the system, the specific application involves fusing modal causal relationship models. First, the causal discovery network analyzes the dependencies between modal variables from the input data, identifying the causal chain of how one modality affects other modalities. Then, the intervention fusion module simulates the impact of changes based on these causal relationships, maps RGB features to a spatiotemporal grid, voxels the LiDAR point cloud and adjusts the coordinates by combining IMU pose. Finally, a comprehensive feature map is generated through weighted fusion, thereby providing the dynamic Transformer sensing chip with input adapted to dynamic scenes and improving the system's ability to capture target dependencies.

[0062] Meta-learning-enhanced Transformer architecture: Meta-learning-enhanced Transformer architecture is the core algorithm architecture of dynamic Transformer perception chip, used to process spatiotemporal features and capture multi-object dependencies.

[0063] The input data is a spatiotemporal feature map generated by the cross-modal spatiotemporal alignment unit, including the fused multimodal feature representation.

[0064] The output is an enhanced feature data embedding vector, including adaptively adjusted attention weights, which capture the nonlinear spatiotemporal relationships of multiple targets.

[0065] In the system, the specific application is to perform feature processing through an online meta-optimization mechanism. First, the inner attention optimizer analyzes the input feature map and fine-tunes the self-attention distribution in real time to highlight the key parts of the current scene. Then, the outer meta-gradient updater learns a general adaptation strategy based on different task distributions and iteratively updates the model weights to adapt to the new environment. Finally, the processed embedding vector is passed to the dynamic feature decoupling module, thereby improving the system's perception accuracy of multi-target motion in complex dynamic scenes.

[0066] Spectral Convolutional Networks: Spectral graph convolutional networks are the core algorithm architecture of the dynamic feature decoupling module, used to separate motion features of multiple targets.

[0067] The input data is the enhanced feature data output by the dynamic Transformer sensing chip, including the captured spatiotemporally dependent embedding vectors.

[0068] The output consists of independent motion vectors, including decoupled trajectory components for each target, which separate the coupling interference between targets.

[0069] In the specific application of the system, decoupling is achieved through frequency domain graph signal processing. First, the frequency domain graph filter constructs a graph structure between targets, converting the input spatiotemporal features into frequency domain signals to analyze frequency components. Then, the coupling decomposer separates the low-frequency independent motion part and the high-frequency coupling interference part. The unique motion information of each target is extracted through graph Laplace decomposition. In addition, the causal spectrum clustering algorithm is integrated to process the dynamic coupling of vehicle and pedestrian targets. The target groups are clustered in collaboration with the decomposer through spectral domain intervention. Finally, the decoupling vector is transmitted to the occlusion prediction unit to ensure that the system can achieve real-time independent tracking in scenarios such as traffic monitoring.

[0070] Diffusion generation model: The diffusion generation model is the core algorithm architecture of the occlusion prediction unit, used to predict occlusion events.

[0071] The input data consists of independent motion vectors provided by the dynamic feature decoupling module, including the decoupled target trajectory component.

[0072] The output consists of the predicted occlusion scene and the compensation fusion result, including the simulated future multimodal scene and pre-compensated features.

[0073] In the specific application of the system, prediction is performed through a conditional diffusion process. First, the conditional noise injector gradually adds noise based on the current motion vector to generate a potential occlusion perturbation sequence. Then, the inverse diffusion decoder gradually recovers the future scene through multi-step inverse denoising and simulates possible occlusion paths. In addition, the multimodal diffusion path simulator generates a predicted trajectory field of dynamic obstacles, which provides obstacle avoidance decision input through collaborative transmission with the decoder, thereby improving the system's robustness to occlusion in scenarios such as obstacle avoidance for logistics drones.

[0074] Reinforcement learning agent: Reinforcement learning agents are the core algorithmic architecture of self-evolving feedback loops, used to dynamically optimize system parameters.

[0075] The input data consists of error information from the occlusion prediction unit, including the deviation index between the predicted scene and the actual scene.

[0076] The output is an optimization instruction, including the adjusted causal graph model parameters, used to update the cross-modal spatiotemporal alignment unit.

[0077] In the specific application of the system, optimization is achieved through a closed-loop collaborative mechanism. First, the agent processor collects errors as reward signals, analyzes deviation patterns to generate policy adjustments, then calculates optimization steps based on these signals, generates update instructions for the causal graph, and finally transmits the instructions back to the cross-modal spatiotemporal alignment unit through a communication connection to achieve collaboration between alignment and prediction. This improves the long-term adaptability of the entire system and ensures high-precision performance in applications such as low-altitude security patrol. Specific Implementation Example 5: like Figures 1 to 3 As shown, the following are specific use cases: Use Case 1: Low-altitude security patrol scenario In low-altitude security patrol applications, the UAV multimodal dynamic target decoupling system can be used to monitor sensitive areas, such as industrial parks or border areas, to identify and track potential intruders or abnormal activities. The system first integrates visible light images captured by an RGB camera, depth point clouds provided by a LiDAR sensor, and attitude data from an IMU (Inertial Measurement Unit) through a multimodal synchronization interface. This data is asynchronously aggregated under an event-triggered compatibility mechanism, ensuring that only critical change events are processed. Subsequently, the RGB-LiDAR-IMU synchronization module applies a neuromorphic spiking network for biomimetic synchronization, simulating the firing delay of biological neurons to achieve high-precision time alignment. The cross-modal spatiotemporal alignment unit uses a causal cross-modal alignment structure to construct an intermodal causal graph model, performing inference and fusion of synchronized data to generate a spatiotemporal feature map adapted to the dynamic environment. The dynamic Transformer perception chip deploys a meta-learning-enhanced Transformer architecture, adaptively adjusting attention weights through online meta-optimization to capture the nonlinear spatiotemporal dependencies of multiple targets, such as pedestrians or vehicles. The dynamic feature decoupling module introduces a spectral graph convolutional network to separate motion features, decomposing topological coupling through frequency domain processing to achieve causal decoupling of independent motion vectors. The occlusion prediction unit predicts potential occlusion events, such as targets being occluded by trees or buildings, based on a diffusion generation model. It simulates future scenarios through conditional diffusion and performs pre-compensation fusion. A self-evolving feedback loop optimizes the causal graph parameters through a reinforcement learning agent, enabling collaboration between the alignment unit and the prediction unit. In this scenario, the system helps the UAV decouple the movement of multiple intruding targets in real time, maintaining continuous tracking and triggering alarms even when occlusion occurs in complex terrain, thus improving security efficiency.

[0079] Use Case 2: Traffic Monitoring Scenarios In traffic monitoring applications, the UAV multimodal dynamic target decoupling system is suitable for real-time traffic analysis on urban roads or highways to detect congestion or abnormal behaviors, such as illegal lane changes or pedestrian crossings. The system begins with a multimodal synchronization interface connecting an RGB camera, LiDAR sensor, and IMU inertial measurement unit, implementing event-triggered interface compatibility and achieving asynchronous aggregation by detecting key events in the data stream. Subsequently, the neuromorphic spiking network of the RGB-LiDAR-IMU synchronization module performs biomimetic synchronization, adjusting modal delays and converting data into spiking sequences to ensure time alignment. The causal cross-modal alignment unit mines causal relationship graphs through a causal cross-modal alignment structure, generating a unified spatiotemporal feature map through intervention and fusion. The meta-learning enhancement of the Transformer architecture in the dynamic Transformer perception chip processes feature data, adaptively adjusting attention to capture the spatiotemporal dependencies of vehicles and pedestrians. The spectral convolutional network of the dynamic feature decoupling module separates motion features, integrates a causal spectral clustering algorithm to handle dynamic coupling, and achieves real-time decoupling through spectral domain intervention and decomposer collaboration. The occlusion prediction unit injects noise into its diffusion generation model to simulate occlusion, and recovers the predicted scene through inverse diffusion, providing compensation fusion. The feedback loop optimizes parameters through a reinforcement learning agent to ensure system coordination. In this scenario, the system decouples multi-target motion, such as separating vehicle trajectories in congestion. Even if the target is occluded by other objects, it can accurately track and generate traffic reports, supporting traffic management decisions.

[0080] Use Case 3: Obstacle Avoidance Scenarios for Logistics Drones In obstacle avoidance applications for logistics drones, this system is used for cargo transportation path planning to avoid dynamic obstacles such as flying birds or ground vehicles, ensuring safe navigation. The system integrates multimodal data input through a multimodal synchronization interface, performing event-triggered asynchronous aggregation. The RGB-LiDAR-IMU synchronization module uses a neuromorphic pulse network to synchronize data, simulating a discharge mechanism to align time. A cross-modal spatiotemporal alignment unit constructs a causal graph model, performing inference fusion to generate a spatiotemporal feature map. The dynamic Transformer perception chip processes features through meta-optimization, adaptively capturing target dependencies. The dynamic feature decoupling module separates motion features, achieving decoupling through frequency domain processing. The occlusion prediction unit predicts occlusion based on a diffusion model, combining it with a path simulator to generate a trajectory field, providing input for obstacle avoidance decisions. A feedback loop optimizes parameters to achieve coordination. In this scenario, the system decouples obstacle motion, predicting paths and adjusting drone routes even under occlusion conditions, improving the reliability of logistics delivery.

[0081] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising a reference structure" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0082] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

[0083] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multi-modal dynamic target decoupling system for unmanned aerial vehicles, characterized in that, It includes a multimodal synchronization interface, an RGB-LiDAR-IMU synchronization module, a cross-modal spatiotemporal alignment unit, a dynamic Transformer sensing chip, a dynamic feature decoupling module, and an occlusion prediction unit; The multimodal synchronization interface connects the RGB camera, LiDAR sensor, and IMU inertial measurement unit, and is used to integrate multimodal data input and perform event-triggered interface compatibility, achieving asynchronous data aggregation by detecting key events in the data stream; The RGB-LiDAR-IMU synchronization module is connected to the multimodal synchronization interface and is used to perform biomimetic synchronization of RGB image data, LiDAR point cloud data and IMU attitude data by applying a neuromorphic pulse network, simulating the firing mechanism of biological neurons to achieve sub-microsecond time alignment. The cross-modal spatiotemporal alignment unit is connected to the RGB-LiDAR-IMU synchronization module. It adopts a causal cross-modal alignment structure to construct an intermodal causal graph model, perform causal inference fusion on the synchronized multimodal data, and generate a dynamically adaptive spatiotemporal feature map. The dynamic Transformer perception chip is connected to the cross-modal spatiotemporal alignment unit and is used to deploy a meta-learning-enhanced Transformer architecture. It processes the spatiotemporally aligned feature data through an online meta-optimization mechanism and adaptively adjusts the attention weights to capture the nonlinear spatiotemporal dependencies of multiple targets. The dynamic feature decoupling module is connected to the dynamic Transformer sensing chip and is used to introduce a spectral graph convolutional network to separate the motion features of multiple targets. It decomposes the topological coupling between targets through frequency domain graph signal processing to achieve causal decoupling of independent motion vectors. The occlusion prediction unit is connected to the dynamic feature decoupling module and is used to predict occlusion events based on the diffusion generation model, simulate future multimodal scenarios through the conditional diffusion process, and perform pre-compensation fusion.

2. The UAV multimodal dynamic target decoupling system according to claim 1, characterized in that, The multimodal synchronization interface includes an event detector and an asynchronous aggregation buffer. The event detector is used to identify change event thresholds in RGB image data, LiDAR point cloud data, and IMU attitude data. The asynchronous aggregation buffer is used to dynamically aggregate data streams based on event priorities and transmit them to the RGB-LiDAR-IMU synchronization module through data interaction.

3. The UAV multimodal dynamic target decoupling system according to claim 1, characterized in that, The neuromorphic pulse network of the RGB-LiDAR-IMU synchronization module includes a biomimetic synaptic connection layer and a pulse timing encoder. The biomimetic synaptic connection layer is used to simulate plastic connections to adjust modal delays, and the pulse timing encoder is used to convert continuous data into pulse sequences to achieve synchronization, and to provide the synchronization data to the cross-modal spatiotemporal alignment unit through instruction transmission.

4. The UAV multimodal dynamic target decoupling system according to claim 1, characterized in that, The causal cross-modal alignment structure of the cross-modal spatiotemporal alignment unit includes a causal discovery network and an intervention fusion module. The causal discovery network is used to mine causal relationship graphs from multimodal data. The intervention fusion module is used to generate a unified spatiotemporal feature representation through causal intervention operations and transmit the spatiotemporal feature graph to the dynamic Transformer sensing chip through data interaction.

5. The UAV multimodal dynamic target decoupling system according to claim 1, characterized in that, The meta-learning enhanced Transformer architecture of the dynamic Transformer perception chip includes an inner attention optimizer and an outer meta-gradient updater. The inner attention optimizer is used to fine-tune the self-attention matrix in real time, and the outer meta-gradient updater is used to learn general weight adaptation based on task distribution and send the processed feature data to the dynamic feature decoupling module through instruction transmission.

6. The UAV multimodal dynamic target decoupling system according to claim 1, characterized in that, The spectral convolutional network of the dynamic feature decoupling module includes a frequency domain graph filter and a coupling decomposer. The frequency domain graph filter is used to convert spatiotemporal features into spectral domain signals, and the coupling decomposer is used to separate independent motion components through graph Laplacian eigenvalue decomposition and transmit the decoupled motion vector to the occlusion prediction unit through data interaction.

7. The UAV multimodal dynamic target decoupling system according to claim 1, characterized in that, The diffusion generation model of the occlusion prediction unit includes a conditional noise injector and an inverse diffusion decoder. The conditional noise injector is used to inject the current features as conditions to generate potential occlusion noise, and the inverse diffusion decoder is used to gradually recover the prediction scene for compensation.

8. The UAV multimodal dynamic target decoupling system according to claim 1, characterized in that, The system further includes a self-evolving feedback loop, which connects the occlusion prediction unit and the cross-modal spatiotemporal alignment unit. The self-evolving feedback loop is used to dynamically optimize the causal graph model parameters through a reinforcement learning agent and realize the collaborative mechanism between the alignment unit and the prediction unit through a communication connection.

9. The UAV multimodal dynamic target decoupling system according to claim 6, characterized in that, The dynamic feature decoupling module integrates a causal spectral clustering algorithm to handle the dynamic coupling of vehicle and pedestrian targets. It achieves real-time decoupling through spectral domain causal intervention operations in collaboration with the coupling decomposer.

10. The UAV multimodal dynamic target decoupling system according to claim 7, characterized in that, The occlusion prediction unit, combined with a multimodal diffusion path simulator, is used to generate a predicted trajectory field of dynamic obstacles and provides obstacle avoidance decision input through instruction transmission between the inverse diffusion decoder and the simulator.