Cooperative positioning method and device based on UWB radar and vision, and electronic equipment

By using a UWB radar and vision-based collaborative localization method, the limitations of single-sensor localization systems in complex dynamic environments are overcome. This method achieves deep fusion of multimodal information and real-time localization, improving localization accuracy and robustness, and is applicable to autonomous driving and intelligent robots.

CN121810787APending Publication Date: 2026-04-07INSPUR WORLDWIDE SERVICES LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In complex and dynamic environments, existing technologies have limitations for single-sensor positioning systems. Most solutions lack deep feature interaction and spatiotemporal alignment mechanisms, resulting in insufficient utilization of intermodal information. Traditional fusion strategies are poorly adaptable to environmental changes and are difficult to achieve continuous and real-time dynamic positioning.

Method used

A UWB radar and vision co-localization method is adopted. A multimodal co-localization dataset is constructed through joint calibration of radar point clouds and visual images. A dual-branch feature extraction network is used to extract radar spatiotemporal features and visual semantic features respectively. Combined with a dynamic temporal attention mechanism and a meta-learning framework, cross-modal feature fusion is achieved, and the localization results are processed in real time on an edge computing platform.

Benefits of technology

It achieves deep fusion of multimodal information, has strong environmental adaptability and high real-time performance, and significantly improves positioning accuracy and reliability in complex dynamic scenarios, making it suitable for applications such as autonomous driving and intelligent robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121810787A_ABST
    Figure CN121810787A_ABST
Patent Text Reader

Abstract

The invention provides a cooperative positioning method and device based on UWB radar and vision, and electronic equipment, and belongs to the technical field of positioning. Comprising the steps of collecting radar point cloud data and a visual image sequence, performing joint calibration, and generating target object information based on the radar point cloud data and the visual image sequence; performing double-branch feature extraction on the radar point cloud data and the visual image sequence in the multi-mode cooperative positioning data set; a cross-modal feature fusion model is constructed based on a dynamic time sequence attention mechanism, the cross-modal feature fusion model is trained and deployed on an edge computing platform, and real-time acquisition, processing and positioning result output of radar point cloud data and a visual image sequence are realized by using a multi-thread mechanism. According to the invention, deep feature fusion can be realized, and UWB radar-vision cooperative positioning which has environment adaptive capability and supports real-time deployment can be realized, so that complex and changeable practical application scenes can be coped with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of positioning technology, specifically relating to a collaborative positioning method, device, and electronic device based on UWB radar and vision. Background Technology

[0002] In cutting-edge fields such as autonomous driving, intelligent robotics, industrial automation, and smart security, high-precision and robust positioning technology is the core foundation for ensuring the safe and reliable operation of the system. However, single sensors all have inherent limitations in complex dynamic environments: Global Navigation Satellite System (GNSS) signals are easily blocked indoors or in urban canyons; visual systems experience a sharp decline in performance under conditions such as low light, fog, and dynamic obstruction; and while UWB radar has strong penetration and anti-interference capabilities, it is prone to positioning drift in scenarios with severe multipath effects or dense targets.

[0003] Although existing research has attempted to fuse radar and visual data, current methods still have significant shortcomings. Most solutions remain at the level of data-level stitching or decision-level fusion, lacking deep feature interaction and spatiotemporal alignment mechanisms, resulting in insufficient utilization of intermodal information; some systems rely on fixed nodes or posterior matching, making it difficult to achieve continuous, real-time dynamic positioning; in addition, traditional fusion strategies have poor adaptability to environmental changes, lacking adaptive adjustment capabilities when a sensor fails or its performance degrades, affecting the overall system stability. Summary of the Invention

[0004] To address at least one of the problems existing in the background technology, this application provides a UWB radar-vision collaborative localization method that can achieve deep feature fusion, has environmental adaptability, and supports real-time deployment of UWB radar-vision collaborative localization, in order to cope with complex and ever-changing real-world application scenarios.

[0005] The second aspect of this application provides a cooperative positioning device based on UWB radar and vision.

[0006] The technical solution adopted in this application is as follows: The first aspect of this application provides a cooperative localization method based on UWB radar and vision, including: Radar point cloud data and visual image sequences are collected and jointly calibrated. Target object information is generated based on the radar point cloud data and the visual image sequences. The 2D bounding box, 3D bounding box, and pose information of the target object are labeled to construct a multimodal collaborative localization dataset containing diverse dynamic scenes. A two-branch feature extraction is performed on the radar point cloud data and visual image sequences in the multimodal cooperative localization dataset to extract radar spatiotemporal features that characterize the target position, motion state and spatial structure, as well as visual semantic features that include object category, texture and contextual relationship. A cross-modal feature fusion model is constructed based on a dynamic temporal attention mechanism. The cross-modal feature fusion model includes a spatiotemporal feature alignment layer, a temporal dependency modeling layer, and an adaptive joint representation layer. Based on the cross-modal feature fusion model, a unified and adaptive joint feature representation is generated from the radar spatiotemporal features and the visual semantic features. The cross-modal feature fusion model is trained based on the task prioritization mechanism and dynamic weighting strategy in the meta-learning framework, and the training weights are adjusted based on modal reliability. The trained cross-modal feature fusion model is deployed on an edge computing platform, and a multi-threaded mechanism is used to realize the real-time acquisition, processing and localization output of the radar point cloud data and the visual image sequence.

[0007] According to the collaborative localization method based on UWB radar and vision provided in the first aspect of this application, firstly, by simultaneously acquiring radar point clouds and visual images and performing joint calibration, and combining real annotations and simulation data augmentation, a high-quality multimodal dataset covering diverse dynamic working conditions is constructed, providing a rich and representative data foundation for model training and improving the model's generalization ability to the complexity of the real world; secondly, a dual-branch network is used to extract the spatiotemporal features of radar and the semantic features of vision respectively, making full use of the stability of UWB radar in distance and motion state perception, and the advantages of vision in object category recognition and environmental context understanding, to achieve a structured expression of heterogeneous modalities at the semantic level; furthermore, a cross-modal feature fusion model is constructed based on a dynamic temporal attention mechanism, establishing fine-grained spatial associations between radar points and image regions through a spatiotemporal feature alignment layer, and utilizing a temporal dependency modeling layer (such as bidirectional...) LSTM captures the continuity of target motion, and an environment-aware gating mechanism is introduced through an adaptive joint representation layer to achieve dynamic adjustment of modal weights, thereby generating a unified and robust joint feature representation. This significantly enhances the system's localization stability under non-ideal conditions such as occlusion, lighting changes, or signal interference. Based on this, a meta-learning framework and task priority mechanism are introduced to enable the model to quickly adapt to new scenarios. Simultaneously, a dynamic weighting strategy for modal reliability assessment is combined to automatically adjust the loss weights of each branch according to the environmental state during training, improving the model's convergence efficiency and robustness under modal performance fluctuations. Finally, the optimized model is deployed on an edge computing platform, leveraging TensorRT acceleration and a multi-threaded double-buffering mechanism to achieve real-time acquisition, processing, and localization output of sensor data. Anomaly detection and degradation operation strategies are integrated to ensure the system's continuous availability under extreme conditions. In summary, this method not only achieves deep fusion of multimodal information from the data layer to the semantic layer, but also possesses strong environmental adaptability, high real-time performance, and system robustness. It significantly improves the positioning accuracy and reliability in complex dynamic scenarios, providing solid technical support for application scenarios with extremely high requirements for safety and response speed, such as autonomous driving and intelligent robots.

[0008] According to one embodiment of this application, the process of collecting radar point cloud data and visual image sequences and performing joint calibration, generating target object information based on the radar point cloud data and the visual image sequences, and labeling the 2D bounding box, 3D bounding box, and pose information of the target object to construct a multimodal cooperative localization dataset containing diverse dynamic scenes, specifically: UWB radar and vision sensors are deployed synchronously in the target scene. Hardware-level time synchronization is achieved using FPGA trigger signals to collect radar spatiotemporal point cloud data and color-depth image sequences with precise timestamps, respectively. The spatial transformation matrix between the radar coordinate system and the camera coordinate system is solved by performing joint calibration using the Kalibr toolbox. LabelImg is used to annotate the target in the visual image with 2D bounding boxes. CloudCompare is used to annotate the radar point cloud with 3D position and contour. The annotation results are then fused to generate pose information that includes the target's unique identifier, three-dimensional spatial coordinates and attitude angles.

[0009] According to one embodiment of this application, the step of performing dual-branch feature extraction on the radar point cloud data and visual image sequences in the multimodal cooperative localization dataset to extract radar spatiotemporal features characterizing target location, motion state, and spatial structure, as well as visual semantic features including object category, texture, and contextual relationships, specifically involves: The raw radar point cloud data is smoothed by a Savitzky-Golay filter, and the filtered point cloud is dynamically separated by the DBSCAN clustering algorithm. The retained dynamic point cloud is input into the PointNet++ network for high-dimensional feature learning. Through its hierarchical sampling and grouping strategy, spatiotemporal features with spatial structure perception capability are extracted, and a 256-dimensional radar feature vector is output. Distortion correction is performed on the visual image. The OpenCV cv2.undistort function is used in combination with the camera intrinsic matrix K and distortion coefficient D to eliminate lens geometric distortion. The YOLO series of object detection models are used to locate the region of interest in the image and generate object detection boxes with confidence. After cropping the detected target region, it is input into the ResNet-50 convolutional neural network for deep semantic feature extraction. Through its residual structure, the appearance, texture and context information of the target are mined, and a 512-dimensional visual feature vector is output to represent the target's category attributes and visual semantic information.

[0010] According to one embodiment of this application, the cross-modal feature fusion model constructed based on a dynamic temporal attention mechanism includes a spatiotemporal feature alignment layer, a temporal dependency modeling layer, and an adaptive joint representation layer. Based on the cross-modal feature fusion model, a unified and adaptive joint feature representation is generated from the radar spatiotemporal features and the visual semantic features. Specifically: Based on the 256-dimensional radar feature vector and the 512-dimensional visual feature vector, a spatiotemporal attention mechanism is constructed to establish fine-grained associations between heterogeneous modalities. The cross-modal attention weights between each radar point and the image region are calculated through a learnable weight matrix. The original features are weighted and fused using the attention weights. A bidirectional LSTM is introduced to perform temporal modeling on the aligned feature sequences of consecutive frames, generating forward and backward hidden states respectively, capturing the temporal continuity of target motion, and outputting an enhanced feature representation containing contextual temporal information. An environmental interference factor is introduced to evaluate each mode in real time, and the evaluation results are embedded into a gating function to generate dynamic weights. The sigmoid activation function and element-wise multiplication operation are used to achieve adjustable fusion of contributions between modes, and a unified joint feature representation is generated through nonlinear transformation. A temporal consistency constraint is introduced to minimize the L2 error between the predicted position change between adjacent frames and the displacement calculated based on the velocity of the previous frame.

[0011] According to one embodiment of this application, the cross-modal feature fusion model is trained based on the task prioritization mechanism and dynamic weighting strategy in the meta-learning framework, and the training weights are adjusted based on modal reliability, specifically as follows: A model-agnostic meta-learning framework is used to meta-train the cross-modal feature fusion model. Multiple scene subsets with different environmental characteristics are divided in the multimodal dataset. A task priority mechanism with scene complexity weighting is introduced on the basis of the original MAML. The training scenes are divided into different priority levels according to dynamic occlusion density and illumination change amplitude. A modal reliability assessment module is constructed to calculate the confidence scores of the radar and vision branches in real time. The visual semantic features are used to generate environmental semantic vectors, and the modal confidence is output by combining learnable parameters and the sigmoid function. Based on the confidence level, a dynamic weighted loss function is designed to automatically reduce the loss weight and enhance the radar branch contribution when the vision is obstructed or the illumination is degraded, and to increase the visual modal weight when the radar is subjected to multipath or electromagnetic interference.

[0012] According to one embodiment of this application, the step of deploying the trained cross-modal feature fusion model on an edge computing platform and using a multi-threaded mechanism to achieve real-time acquisition, processing, and localization output of the radar point cloud data and the visual image sequence specifically involves: Based on a dual-buffered pipeline architecture, the data acquisition and model inference tasks are separated through a multi-threaded mechanism. One thread is responsible for real-time reading and preprocessing of sensor data, while the other thread performs fusion localization inference. A hardware-level time synchronization system based on PTP is constructed to achieve microsecond-level clock alignment between UWB radar and visual sensors. A data buffer queue with timestamp matching is constructed to pair asynchronously arriving radar point clouds and image sequences. Specifically, when data loss or delay is detected, a motion model-based prediction compensation mechanism is activated to interpolate missing observations.

[0013] According to one embodiment of this application, the step of deploying the trained cross-modal feature fusion model on an edge computing platform and using a multi-threading mechanism to realize the real-time acquisition, processing, and localization result output of the radar point cloud data and the visual image sequence further includes: Perform online health monitoring of sensors, verify the integrity of input data, and judge the reasonableness of positioning results; When the visual sensor fails, it switches to a degraded positioning mode that relies primarily on radar. When radar signals are interfered with, increase the weight of the visual modality; In extreme cases where both the visual sensor and the radar signal are abnormal, the kinematic prediction module based on IMU or historical trajectory is activated to maintain short-term positioning output.

[0014] A second aspect of this application provides a cooperative positioning device based on UWB radar and vision, comprising: The data acquisition module is suitable for acquiring radar point cloud data and visual image sequences and performing joint calibration. Based on the radar point cloud data and the visual image sequences, it generates target object information and labels the 2D bounding box, 3D bounding box and pose information of the target object to construct a multimodal collaborative localization dataset containing diverse dynamic scenes. The feature extraction module is adapted to perform dual-branch feature extraction on radar point cloud data and visual image sequences in the multimodal cooperative localization dataset, so as to extract radar spatiotemporal features that characterize the target position, motion state and spatial structure, as well as visual semantic features that include object category, texture and contextual relationship. The model building module is suitable for building a cross-modal feature fusion model based on a dynamic temporal attention mechanism. The cross-modal feature fusion model includes a spatiotemporal feature alignment layer, a temporal dependency modeling layer, and an adaptive joint representation layer. Based on the cross-modal feature fusion model, a unified and adaptive joint feature representation is generated from the radar spatiotemporal features and the visual semantic features. The model training module is suitable for training the cross-modal feature fusion model based on the task priority mechanism and dynamic weighting strategy in the meta-learning framework, and adjusting the training weights based on modal reliability. The result output module is suitable for deploying the trained cross-modal feature fusion model on an edge computing platform and using a multi-threading mechanism to realize the real-time acquisition, processing and positioning results output of the radar point cloud data and the visual image sequence.

[0015] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the UWB radar and vision-based collaborative positioning method described in any of the first aspects above.

[0016] This application also provides a non-volatile computer storage medium storing computer-executable instructions thereon, wherein the computer program, when executed by a processor, implements the UWB radar and vision-based cooperative positioning method as described in any embodiment of the first aspect above. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating the collaborative localization method based on UWB radar and vision provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of the cross-modal feature fusion model provided in the embodiments of this application; Figure 3 This is a flowchart illustrating the actual deployment process of the cross-modal feature fusion model provided in this application embodiment. Detailed Implementation

[0018] To more clearly illustrate the overall concept of this application, a detailed explanation is provided below with reference to the accompanying drawings.

[0019] Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below. It should be noted that, unless otherwise specified, the embodiments of this application and the features thereof can be combined with each other.

[0020] In this application, unless otherwise expressly specified and limited, the "above" or "below" of the second feature can mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. In the description of this specification, references to terms such as "an embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples.

[0021] like Figures 1 to 3 As shown, the first aspect of this application provides a cooperative localization method based on UWB radar and vision, including: Step 100: Collect radar point cloud data and visual image sequences and perform joint calibration. Generate target object information based on radar point cloud data and visual image sequences, and label the 2D bounding box, 3D bounding box and pose information of the target object to construct a multimodal collaborative localization dataset containing diverse dynamic scenes.

[0022] Step 200: Perform dual-branch feature extraction on the radar point cloud data and visual image sequences in the multimodal cooperative localization dataset to extract radar spatiotemporal features that characterize the target's location, motion state, and spatial structure, as well as visual semantic features that include object category, texture, and contextual relationships.

[0023] Step 300: Construct a cross-modal feature fusion model based on a dynamic temporal attention mechanism. The cross-modal feature fusion model includes a spatiotemporal feature alignment layer, a temporal dependency modeling layer, and an adaptive joint representation layer. Based on the cross-modal feature fusion model, a unified and adaptive joint feature representation is generated from radar spatiotemporal features and visual semantic features.

[0024] Step 400: Train a cross-modal feature fusion model based on the task priority mechanism and dynamic weighting strategy in the meta-learning framework, and adjust the training weights based on modal reliability.

[0025] Step 500: Deploy the trained cross-modal feature fusion model on an edge computing platform and use a multi-threading mechanism to achieve real-time acquisition, processing and localization output of radar point cloud data and visual image sequences.

[0026] In step 100, "collecting radar point cloud data and visual image sequences and performing joint calibration" refers to simultaneously deploying a UWB radar and a high-resolution RGB-D camera, achieving hardware-level time synchronization under FPGA trigger signal control, ensuring strict alignment of radar and visual data in timestamps. "Joint calibration" utilizes multi-sensor calibration tools such as Kalibr to solve for the spatial transformation matrix (rotation and translation parameters) between the radar coordinate system and the camera coordinate system, achieving spatial alignment of heterogeneous data under a unified geometric reference. Based on this, "generating target object information" involves labeling the target's 2D / 3D bounding boxes and pose information (position + attitude angle) using tools such as LabelImg and CloudCompare, constructing a multimodal collaborative localization dataset covering indoor and outdoor environments, and different lighting and occlusion conditions.

[0027] This step addresses the most fundamental "spatiotemporal consistency" problem in multimodal sensing systems, preventing subsequent fusion deviations caused by sensor asynchrony or coordinate misalignment. By fusing and expanding real-world data with physical simulation data, the diversity and representativeness of the dataset are enhanced.

[0028] In step 200, "dual-branch feature extraction of multimodal data" refers to constructing independent UWB radar and vision branch network structures. In the radar branch, the "signal filtering layer" uses a Savitzky-Golay filter to suppress noise in the original point cloud, and the "spatiotemporal feature extraction layer" uses point cloud neural networks such as PointNet++ to extract high-dimensional spatiotemporal features representing the target's position, velocity, trajectory, and spatial distribution from the dynamic target point cloud. In the vision branch, the "image preprocessing layer" eliminates lens distortion using cv2.undistort, and the "semantic feature extraction layer" uses YOLO to detect the target and then uses ResNet-50 to extract deep semantic features containing object category, texture, shape, and contextual relationships.

[0029] The dual-branch design follows the "divide first, then combine" principle, optimizing the physical characteristics of radar and visual modalities separately to avoid information confusion caused by direct fusion. Point cloud networks excel at handling unordered sparse data, while CNNs are adept at extracting local and global semantics from images.

[0030] In step 300, “constructing a cross-modal feature fusion model based on a dynamic temporal attention mechanism” refers to calculating the attention weights between radar points and image regions through a “spatiotemporal feature alignment layer”, establishing a cross-modal spatial mapping relationship, and solving the feature misalignment caused by perspective differences; the “temporal dependency modeling layer” uses bidirectional LSTM to model the feature sequence of continuous frames, capturing the temporal dependency of target motion and enhancing the consistency of features in dynamic scenes; the “adaptive joint representation layer” designs a gated fusion mechanism, introduces environmental interference factors (such as image blurring and radar multipath intensity) to dynamically adjust the fusion weights of radar and visual features, and generates a unified and robust joint feature representation.

[0031] This model breaks through the traditional simple splicing or weighted averaging fusion method. It achieves fine-grained feature association through attention mechanism, enhances temporal continuity through LSTM, and achieves environmental adaptation through gating mechanism.

[0032] In step 400, “training the model based on the task priority mechanism and dynamic weighting strategy in the meta-learning framework” refers to using the MAML (Model-Agnostic Meta-Learning) framework to perform meta-training on multiple subsets of scenes with different complexities, enabling the model to learn common initial parameters, thus requiring only a small number of samples to quickly fine-tune and adapt in new scenes; the “task priority mechanism” assigns higher weights to more difficult tasks based on scene complexity (such as occlusion density and illumination changes), improving the model’s generalization ability to challenging environments; and the “dynamic weighting strategy” adjusts the training weights of each modality in the loss function based on the real-time evaluated modal reliability (such as visual clarity and radar signal-to-noise ratio), achieving adaptive optimization of the training process.

[0033] Meta-learning empowers models with the ability to "learn to learn," while dynamic weighting simulates the attention regulation mechanism in human perception, prioritizing more reliable perceptual channels.

[0034] In step 500, "deploying the trained model to an edge computing platform" refers to performing INT8 quantization and graph optimization on the optimized fusion model using TensorRT and deploying it to embedded devices such as Jetson AGX Xavier; "utilizing a multi-threading mechanism" enables the parallel operation of the data acquisition thread and the model inference thread, and adopts a double-buffered queue to avoid data blocking; at the same time, it combines the PTP protocol to achieve microsecond-level time synchronization, builds an anomaly detection and degradation operation mechanism (such as switching to radar-dominated mode when visual failure occurs), and outputs real-time positioning results to the outside world through ROS or HTTP interface.

[0035] Edge deployment ensures low latency and high real-time performance, multi-threaded architecture improves resource utilization, and fault tolerance mechanisms enhance system availability.

[0036] According to the collaborative localization method based on UWB radar and vision provided in the first aspect of this application, firstly, by simultaneously acquiring radar point clouds and visual images and performing joint calibration, and combining real annotations and simulation data augmentation, a high-quality multimodal dataset covering diverse dynamic working conditions is constructed, providing a rich and representative data foundation for model training and improving the model's generalization ability to the complexity of the real world; secondly, a dual-branch network is used to extract the spatiotemporal features of radar and the semantic features of vision respectively, making full use of the stability of UWB radar in distance and motion state perception, and the advantages of vision in object category recognition and environmental context understanding, to achieve a structured expression of heterogeneous modalities at the semantic level; furthermore, a cross-modal feature fusion model is constructed based on a dynamic temporal attention mechanism, establishing fine-grained spatial associations between radar points and image regions through a spatiotemporal feature alignment layer, and utilizing a temporal dependency modeling layer (such as bidirectional...) LSTM captures the continuity of target motion, and an environment-aware gating mechanism is introduced through an adaptive joint representation layer to achieve dynamic adjustment of modal weights, thereby generating a unified and robust joint feature representation. This significantly enhances the system's localization stability under non-ideal conditions such as occlusion, lighting changes, or signal interference. Based on this, a meta-learning framework and task priority mechanism are introduced to enable the model to quickly adapt to new scenarios. Simultaneously, a dynamic weighting strategy for modal reliability assessment is combined to automatically adjust the loss weights of each branch according to the environmental state during training, improving the model's convergence efficiency and robustness under modal performance fluctuations. Finally, the optimized model is deployed on an edge computing platform, leveraging TensorRT acceleration and a multi-threaded double-buffering mechanism to achieve real-time acquisition, processing, and localization output of sensor data. Anomaly detection and degradation operation strategies are integrated to ensure the system's continuous availability under extreme conditions. In summary, this method not only achieves deep fusion of multimodal information from the data layer to the semantic layer, but also possesses strong environmental adaptability, high real-time performance, and system robustness. It significantly improves the positioning accuracy and reliability in complex dynamic scenarios, providing solid technical support for application scenarios with extremely high requirements for safety and response speed, such as autonomous driving and intelligent robots.

[0037] In some embodiments of this application, radar point cloud data and visual image sequences are collected and jointly calibrated. Target object information is generated based on the radar point cloud data and visual image sequences, and the 2D bounding box, 3D bounding box, and pose information of the target object are labeled to construct a multimodal cooperative localization dataset containing diverse dynamic scenes. Specifically: UWB radar and vision sensors are deployed synchronously in the target scene. Hardware-level time synchronization is achieved using FPGA trigger signals to collect radar spatiotemporal point cloud data and color-depth image sequences with precise timestamps, respectively. The spatial transformation matrix between the radar coordinate system and the camera coordinate system is solved by performing joint calibration using the Kalibr toolbox. LabelImg is used to annotate the target in the visual image with 2D bounding boxes. CloudCompare is used to annotate the radar point cloud with 3D position and contour. The annotation results are then fused to generate pose information that includes the target's unique identifier, three-dimensional spatial coordinates and attitude angles.

[0038] In the target scene, UWB radar (such as a millimeter-wave radar development board) and a high-resolution RGB-D vision sensor (such as the Intel RealSense D455) are deployed. These sensors respectively perceive the target's distance, velocity, angle information, and rich texture, color, and depth images. To ensure strict data alignment in the temporal dimension, an FPGA (Field-Programmable Gate Array) is used to generate a high-precision synchronization trigger pulse signal, serving as the common start command for both sensors. This mechanism achieves hardware-level time synchronization, avoiding clock drift and latency jitter caused by software triggering. It ensures that each frame of radar point cloud has a precise timestamp match with its corresponding visual image, providing a reliable timing foundation for subsequent cross-modal feature association.

[0039] Because UWB radar and visual sensors differ in their perception principles, coordinate system origins, and orientations, spatial coordinate system unification is essential. This solution employs advanced multi-sensor calibration toolkits such as Kalibr, deploying calibration boards (e.g., AprilTag or checkerboard patterns) in the scene while simultaneously acquiring radar echo signals and visual images. An optimization algorithm is then used to solve for the rigid body transformation matrix (rotation matrix R and translation vector T) between the radar coordinate system and the camera coordinate system. This matrix describes how to transform any point in the radar point cloud to the camera coordinate system, thereby achieving geometric alignment of heterogeneous data in three-dimensional space, a prerequisite for subsequent cross-modal fusion.

[0040] Cross-modal joint annotation and pose information generation: Visual side annotation: Using image annotation tools such as LabelImg, 2D bounding boxes are annotated for targets in visual images to record the position, size and category information of the targets in the image plane, which are used for subsequent semantic feature extraction and detection verification.

[0041] Radar side annotation: Using point cloud processing software such as CloudCompare, the 3D point cloud collected by the radar is manually or semi-automatically annotated to extract the 3D spatial contour and center position of the target and form a 3D bounding box.

[0042] Pose fusion generation: Visual 2D annotations and radar 3D annotations are matched using timestamps and spatial transformation matrices to assign a unique ID to each target. The 3D position coordinates (x, y, z) and attitude angles (such as yaw) are then fused to form a complete 6DoF (six degrees of freedom) pose label. This process achieves semantic-level alignment of multimodal information and constructs high-quality ground truth that can be used for supervised learning.

[0043] To further enhance the generalization ability of the dataset, virtual scene data covering different lighting conditions (day / night), weather conditions (rain / fog / strong light), occlusion (partial occlusion / dense occlusion), and dynamic disturbances (multi-target cross motion) are generated by combining physical simulation platforms (such as CARLA or Gazebo). This data is then mixed with real-world collected data in proportion to construct a diverse dynamic multimodal dataset covering complex indoor and outdoor working conditions.

[0044] By using FPGA hardware synchronization and joint calibration with Kalibr, the most critical "spatiotemporal misalignment" problem in multi-sensor systems is solved, avoiding fusion errors caused by time delays or spatial deviations and significantly improving subsequent localization accuracy. A cross-modal joint annotation strategy is employed to generate ground truth data containing 2D / 3D bounding boxes and complete pose information, providing accurate supervision signals for deep learning models and improving the convergence speed and accuracy of model training. Through a real-world + simulation data fusion strategy, a dataset covering various complex scenarios is constructed, enabling the model to cope with challenges such as illumination changes, occlusion, and dynamic interference during the training phase, significantly improving its robustness in the real world. This dataset can not only be used for supervised training of object detection and localization tasks but also for joint optimization of sub-modules such as cross-modal alignment, feature fusion, and trajectory prediction, providing data support for building a complete cooperative localization system.

[0045] In some embodiments of this application, dual-branch feature extraction is performed on radar point cloud data and visual image sequences in a multimodal cooperative localization dataset to extract radar spatiotemporal features characterizing target location, motion state, and spatial structure, as well as visual semantic features including object category, texture, and contextual relationships. Specifically: The raw radar point cloud data is smoothed by a Savitzky-Golay filter, and the filtered point cloud is dynamically separated by the DBSCAN clustering algorithm. The retained dynamic point cloud is input into the PointNet++ network for high-dimensional feature learning. Through its hierarchical sampling and grouping strategy, spatiotemporal features with spatial structure perception capability are extracted, and a 256-dimensional radar feature vector is output. Distortion correction is performed on the visual image. The OpenCV cv2.undistort function is used in combination with the camera intrinsic matrix K and distortion coefficient D to eliminate lens geometric distortion. The YOLO series of object detection models are used to locate the region of interest in the image and generate object detection boxes with confidence. After cropping the detected target region, it is input into the ResNet-50 convolutional neural network for deep semantic feature extraction. Through its residual structure, the appearance, texture and context information of the target are mined, and a 512-dimensional visual feature vector is output to represent the target's category attributes and visual semantic information.

[0046] UWB radar is susceptible to noise and multipath interference in complex electromagnetic environments, leading to jitter or false targets in point cloud data. A Savitzky-Golay filter is employed to smooth the time-series signal of the original point cloud. This filter, based on local polynomial fitting, effectively suppresses random noise while preserving high-frequency details (such as target edges), outperforming traditional low-pass filters and is particularly suitable for denoising non-stationary radar signals.

[0047] In static environments, most point clouds originate from the background (walls, ground, etc.), and moving targets need to be separated from them. The DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering algorithm is employed to automatically identify and segment dynamic target clusters based on the spatial density characteristics of the point cloud. Its advantages include not requiring a pre-defined number of clusters, effectively identifying irregularly shaped targets, and marking isolated noise points as outliers for removal, thus improving the accuracy of target extraction.

[0048] The preserved dynamic point cloud is input into the PointNet++ deep neural network. This network captures the local structure and global distribution characteristics of the point cloud at different scales through hierarchical sampling (setabstraction) and grouping strategies. Its symmetric functions (such as max pooling) ensure the invariance of the point cloud's disorder, while a multilayer perceptron (MLP) extracts high-dimensional features that can characterize the target's position, velocity, trajectory, spatial contour, and structural changes. Finally, a compact 256-dimensional feature vector is output as the "spatiotemporal fingerprint" of the radar mode.

[0049] Camera lenses exhibit radial and tangential distortion, causing bending and deformation in image edge regions and affecting target localization accuracy. Using the `cv2.undistort` function in OpenCV, combined with pre-calibrated camera intrinsic parameter matrix K (focal length, principal point coordinates) and distortion coefficients D, geometric correction is performed on the original image to restore the linear structure of the real scene, providing geometrically accurate input for subsequent target detection and feature extraction.

[0050] The YOLO (You Only Look Once) series of object detection algorithms (such as YOLOv5 / v8) are used to quickly locate the target of interest in the entire image and output a 2D bounding box with class label and confidence score. YOLO's single-stage detection architecture has a good balance between high inference speed and accuracy, making it suitable for real-time system applications. The detection results are used to crop out the target region to avoid the influence of background interference on feature extraction.

[0051] The target image patches detected by YOLO are input into a ResNet-50 convolutional neural network. This network solves the gradient vanishing problem in deep networks through residual connections (skip connections), supporting deeper feature abstraction. After multiple convolutional layers, the network can extract the target's color, texture, shape, component structure, and contextual semantic information (such as whether pedestrians are carrying items, vehicle type, etc.). Finally, it outputs a 512-dimensional visual feature vector as the target's "visual semantic encoding".

[0052] The radar and vision branches adopt a parallel and decoupled design architecture, each optimizing for its own modal characteristics to avoid information confusion or dominant modality suppression caused by direct fusion. This "separate-then-integrate" strategy ensures that the advantages of each sensor are fully utilized, providing high-quality, structured input features for subsequent cross-modal fusion.

[0053] The radar branch, through a process of filtering, clustering, and PointNet++, achieves efficient transformation from raw noisy point clouds to structured spatiotemporal features, enhancing the perception of dynamic target motion states. The vision branch, through distortion correction, detection, and ResNet extraction, obtains semantic features rich in category and appearance information, significantly improving target recognition and discrimination capabilities. Savitzky-Golay filtering and DBSCAN effectively suppress radar noise and static clutter, improving the stability of target extraction; image distortion correction ensures visual geometric accuracy, and YOLO's robust detection mechanism maintains high recall even under partial occlusion and illumination changes. The output 256-dimensional radar features and 512-dimensional visual features are both compact and highly discriminative vector representations, facilitating subsequent alignment and fusion through attention mechanisms, gating fusion, etc., avoiding dimensionality mismatch and semantic gap problems caused by directly splicing raw data. The dual-branch structure can be computed in parallel, fully utilizing multi-core CPU or GPU resources; both YOLO and ResNet are lightweight designs, capable of running efficiently on edge devices, ensuring low-latency response of the overall system.

[0054] Specifically, in one embodiment, a digital filtering algorithm is used to smooth the signal to address noise interference in the original radar signal. A clustering analysis algorithm is used to separate the dynamic target point cloud and remove environmental clutter interference. An advanced point cloud neural network architecture is used to extract the spatiotemporal features of the radar signal, converting the original point cloud data into a high-dimensional feature representation while preserving the target's motion characteristics and spatial distribution information.

[0055] To address the noise and multipath interference issues in radar signals, a Savitzky-Golay filter is first used for signal smoothing.

[0056] Among them, Represents the original radar signal. These are the filter coefficients. n The width is half the width of the window. This is time information.

[0057] Then, the dynamic target point cloud is separated using the DBSCAN clustering algorithm:

[0058] in, Indicates radar point, As cluster center, is the neighborhood radius.

[0059] Finally, the PointNet++ network was used to extract spatiotemporal features:

[0060] in, This refers to the PointNet++ encoder. For a point cloud matrix containing coordinates and intensity, This is the output of 256-dimensional radar features.

[0061] Optical distortion correction is performed on the raw images captured by the camera to eliminate geometric distortion caused by the lens. A deep learning object detection algorithm is used to locate the target of interest in the image. Deep semantic features are extracted from the image based on a convolutional neural network architecture, and high-dimensional feature representations containing information about the target's appearance, texture, and structure are extracted from the corrected image region.

[0062] To address image distortion and lighting variations, OpenCV's cv2.undistort function is used for lens distortion correction.

[0063] in, For the original image, K For the camera intrinsic parameter matrix,D The distortion coefficient is denoted as .

[0064] Then, the target is detected using the YOLO model (different series of YOLO algorithms can be selected according to the current computing power conditions):

[0065] in, Indicates the detection box. c , where is the confidence level.

[0066] Finally, ResNet-50 is used to extract image features:

[0067] in, According to The target area to be cropped. This is the output of 512-dimensional visual features.

[0068] In some embodiments of this application, a cross-modal feature fusion model is constructed based on a dynamic temporal attention mechanism. This model includes a spatiotemporal feature alignment layer, a temporal dependency modeling layer, and an adaptive joint representation layer. Based on this model, a unified and adaptive joint feature representation is generated from radar spatiotemporal features and visual semantic features. Specifically: Based on 256-dimensional radar feature vectors and 512-dimensional visual feature vectors, a spatiotemporal attention mechanism is constructed to establish fine-grained associations between heterogeneous modalities. Cross-modal attention weights between each radar point and the image region are calculated through a learnable weight matrix. The attention weights are used to perform weighted fusion of the original features. A bidirectional LSTM is introduced to perform temporal modeling on the aligned feature sequences of consecutive frames, generating forward and backward hidden states respectively, capturing the temporal continuity of target motion, and outputting an enhanced feature representation containing contextual temporal information. An environmental interference factor is introduced to evaluate each mode in real time, and the evaluation results are embedded into a gating function to generate dynamic weights. The sigmoid activation function and element-wise multiplication operation are used to achieve adjustable fusion of contributions between modes, and a unified joint feature representation is generated through nonlinear transformation. A temporal consistency constraint is introduced to minimize the L2 error between the predicted position change between adjacent frames and the displacement calculated based on the velocity of the previous frame.

[0069] In dynamic scenes, the position and state of a target evolve over time, and relying solely on single-frame information can easily lead to positioning jitter or jumps. Therefore, a bidirectional long short-term memory (Bi-LSTM) network is introduced to model the aligned feature sequence of consecutive frames: The fused features of the current frame and several frames before and after are input into the Bi-LSTM in chronological order. The forward LSTM captures the state evolution from the past to the present, and the backward LSTM captures the information backtracking from the future to the present. Together, they output the hidden state containing contextual information.

[0070] This mechanism can effectively model the target's motion trend, acceleration changes, and trajectory smoothness, enhancing the system's predictive ability under conditions such as brief occlusion and signal loss, and improving the temporal consistency of positioning results.

[0071] The reliability of radar and vision systems varies under different environmental conditions (e.g., visual degradation at night, radar performance deterioration in rain and fog). To address this, a gating fusion mechanism is designed to adaptively adjust the modal weights. Introducing environmental interference factors, dynamically generated by real-time monitoring of sensor signal quality: For visual modalities: Confidence is assessed based on indicators such as image sharpness (e.g., Laplacian gradient variance), illumination intensity, and target blur. For radar modes: their reliability is evaluated based on parameters such as signal-to-noise ratio (SNR), multipath interference intensity, and point cloud density.

[0072] The spatiotemporal attention mechanism overcomes the limitations of coarse-grained stitching in traditional fusion methods, achieving semantic-level matching between radar points and image regions, significantly improving the accuracy and robustness of heterogeneous feature fusion. Bi-LSTM models the temporal dependence of target motion, giving the system "memory" capabilities, maintaining stable output even under brief occlusion, noise interference, or sensor failure, thus enhancing the overall system's continuity and reliability. The gated fusion mechanism dynamically adjusts modal weights based on real-time environmental conditions, simulating the attention switching mechanism in human perception, enabling the system to maintain high performance under complex conditions such as sudden changes in illumination, rain, fog, and strong electromagnetic interference, significantly improving robustness. Temporal consistency constraints ensure that the predicted trajectory conforms to the laws of physical motion, reducing jitter and jumps, and outputting more natural and reliable localization results, suitable for applications with high safety requirements (such as autonomous driving and robot navigation). The entire fusion model can be jointly trained with detection heads, regression heads, and other task heads, uniformly optimizing all parameters through backpropagation to achieve a complete mapping from multimodal input to precise localization output, improving the overall system performance ceiling.

[0073] Specifically, a spatiotemporal attention mechanism is designed to establish the correlation between radar point clouds and visual pixels. Attention weights between different modal features are calculated using a learnable weight matrix, reflecting the spatial correspondence between radar points and image regions. Attention weighting is used to align the feature space, resolving the feature mismatch problem caused by differences in sensor perspective.

[0074] To address the spatial misalignment between radar point clouds and visual pixels, a spatiotemporal attention mechanism is designed:

[0075] in, , For learnable weight matrix, Indicates the first i The radar point and the first j The association weights of each image region Let i be the feature vector of the i-th radar point. Let be the visual feature vector of the j-th image region.

[0076] Feature alignment is then achieved through attention weighting:

[0077] in, The features are aligned.

[0078] A Long Short-Term Memory (LSTM) network is introduced to perform temporal modeling of continuous radar-visual feature sequences, capturing the temporal continuity of target motion. Specifically, a bidirectional LSTM is used to model the feature sequences. Encode to generate feature representations containing temporal information. and This enhances feature consistency in dynamic scenarios.

[0079] A gating fusion mechanism is employed to integrate feature representations from radar and vision. A learnable gating function is designed to dynamically adjust the contribution of each modality's features, achieving adaptive feature fusion. A unified joint feature representation is generated through nonlinear combination, fully preserving the advantageous characteristics of each modality.

[0080] A gating fusion mechanism is adopted to integrate multi-source features, and an environmental interference factor is introduced into the original gating function. This factor is dynamically generated by real-time monitoring of sensor signal quality (such as radar multipath intensity and visual image sharpness).

[0081] in, For the gated weight matrix, This is the sigmoid function, where ⊙ represents element-wise multiplication.

[0082] A multi-task learning framework is designed to simultaneously optimize object detection and location estimation tasks. A self-supervised learning strategy is introduced to enhance the model's robustness to modality loss. End-to-end training optimizes the overall model parameters, achieving a complete mapping from multimodal input to precise localization output.

[0083] Design a multi-task learning framework to jointly optimize localization and detection, incorporating a temporal consistency constraint. By comparing the smoothness of localization results across consecutive frames, abrupt changes are penalized.

[0084] in, To predict the location, For the actual location, To detect the loss, , , This is the balance coefficient. The overall loss is minimized through end-to-end training.

[0085]

[0086] in, These are the frame numbers in the time series, representing different moments in time. The current frame Predicted location; The previous frame Predicted location; The previous frame Predicted speed; Indicates the time interval between two adjacent frames; This represents the L2 norm, used to calculate the Euclidean distance of a vector. The formula ensures the positioning results conform to kinematic laws by calculating the error between the positional change in the current frame and the displacement of the velocity in the previous frame within the time interval.

[0087] In some embodiments of this application, a cross-modal feature fusion model is trained based on the task prioritization mechanism and dynamic weighting strategy in the meta-learning framework, and the training weights are adjusted based on modal reliability, specifically as follows: A model-agnostic meta-learning framework is adopted to perform meta-training on the cross-modal feature fusion model. Multiple scene subsets with different environmental characteristics are divided in the multimodal dataset. A task priority mechanism with scene complexity weighting is introduced on the basis of the original MAML, and the training scene is divided into different priority levels according to dynamic occlusion density and illumination change amplitude. A modal reliability assessment module is constructed to calculate the confidence scores of the radar and vision branches in real time. The visual semantic features are used to generate environmental semantic vectors, and the modal confidence is output by combining learnable parameters and the sigmoid function. Based on confidence, a dynamic weighted loss function is designed to automatically reduce the loss weight and enhance the radar branch contribution when the visual mode is obstructed or the illumination is degraded, and to increase the visual mode weight when the radar is subjected to multipath or electromagnetic interference.

[0088] Different scenarios present varying degrees of challenge to the localization system. To improve the model's adaptability to challenging scenarios, a task priority mechanism weighted by scenario complexity is introduced on top of the original MAML: Define complexity metrics such as dynamic occlusion density (the proportion of a target that is occluded per unit time), illumination variation amplitude (the rate of change of the standard deviation of image brightness), and radar multipath intensity. Based on the above metrics, training tasks are divided into three priority levels: low, medium, and high. Higher complexity tasks are assigned higher weights to the outer loop loss. wi >1. The meta-learning framework enables models to obtain universal initialization parameters, possessing the ability to "quickly adapt to new environments," significantly reducing the amount of labeled data and training time required for deployment in new scenarios, and is suitable for diverse real-world application environments. A task prioritization mechanism guides the model to focus on high-difficulty scenarios (such as dense occlusion and extreme lighting) during the training phase, improving its robustness and performance ceiling under the most challenging conditions. The dynamic weighted loss function automatically adjusts training weights based on real-time modality reliability, avoiding the negative impact of low-quality data on model optimization and improving training stability and convergence efficiency. In cases of partial modality degradation or temporary failure, the model can automatically reduce its supervision intensity, relying on more reliable modalities to continue learning, simulating the redundancy and compensation mechanisms of the human perception system, enhancing the overall reliability of the system. The entire training strategy can be jointly optimized with feature extraction and fusion modules, forming a unified end-to-end training process; and the dynamic weighting mechanism during training and the gating fusion mechanism during inference are consistent, ensuring that model behavior remains consistent between the training and deployment phases.

[0089] Specifically, a model-agnostic meta-learning method is employed to enable the model to quickly adapt to new scenarios. By simulating multi-scenario training tasks, model initialization parameters with good generalization ability are learned. When applied to new environments, only a small number of samples are needed to quickly adjust the model parameters, significantly reducing the requirement for large amounts of data for new scenarios.

[0090] To address the need for rapid adaptation to new scenarios, a meta-learning framework based on MAML is designed. Building upon the existing MAML framework, a task prioritization mechanism weighted by scene complexity is added: the training scene is divided into different priority levels according to its complexity (e.g., dynamic occlusion density, illumination variation amplitude). Higher complexity, higher priority The larger the subset, the more likely it is to be divided into multiple scene subsets during the meta-training phase. Each of them This represents a dataset for a specific scenario. The model initialization parameters are learned through a two-layer optimization process.

[0091] in These are the initial parameters for the model. Indicates in Parameter update operations on the training set The test set loss is used. This framework enables the model to generalize across different scenarios, allowing it to quickly adapt to new environments with only a small number of samples for fine-tuning.

[0092] A modal reliability assessment mechanism is established to calculate the confidence level of each sensor's data in real time. The weights of different modalities in the loss function are dynamically adjusted according to environmental changes. When vision is limited, the contribution of radar information is enhanced; when radar is interfered with, the weight of visual information is increased, achieving adaptive importance allocation. Simultaneously, an environmental semantic vector is constructed using semantic features extracted from the visual branch. Incorporate it into the confidence calculation: To address the dynamic change problem of multimodal reliability, a modal adaptive weighting mechanism is proposed. First, the confidence score for each mode is calculated:

[0093] in Representing modes m Features (radar or vision). and For learnable parameters, This is the sigmoid function.

[0094] Then dynamically adjust the loss weights:

[0095] in, and As the benchmark weight, and These represent the real-time confidence scores for the radar and visual modalities, respectively. This design enables the model to automatically enhance the contribution of the radar branch when visual occlusion occurs, and to emphasize visual information when the radar is jammed.

[0096] Simulated environmental disturbances are injected during training to enhance the model's robustness. Adversarial perturbations are added to visual data to simulate image quality degradation, and synthetic noise is added to radar data to simulate signal interference. Adversarial training improves the model's stability and reliability in complex environments.

[0097] First, adversarial examples are injected during training. For the visual modality, adversarial perturbations are generated using the Fast Gradient Sign Method (FGSM):

[0098] in For the input image, For real labels, Controlling disturbance intensity. For radar modes, a physical model-driven interference simulation is introduced. In complex electromagnetic environments, traditional Gaussian noise models struggle to accurately characterize real-world interference characteristics. To enhance the model's adaptability to real-world scenarios, differentiated physical model-driven interference simulation strategies are designed for different modal data:

[0099] in, The interference signal representing the generated radar modal adversarial sample is the final result calculated by the entire formula, i.e., the signal used to interfere with the radar system. For the first The weighting coefficients corresponding to each reflection path are used to adjust the contribution of each reflection path signal to the final interference signal; This indicates that the mean is 0 and the covariance matrix is... Gaussian noise; This is a ray tracing function that uses the environment map. , with the first The starting point or related parameters of the reflection path As input, the calculation yields the first... The signal with a reflection path is made more realistic by introducing real environmental information to make the interference more closely resemble the actual scenario.

[0100] In some embodiments of this application, the trained cross-modal feature fusion model is deployed on an edge computing platform, and a multi-threaded mechanism is used to realize the real-time acquisition, processing, and localization output of radar point cloud data and visual image sequences, specifically: Based on a dual-buffered pipeline architecture, the data acquisition and model inference tasks are separated through a multi-threaded mechanism. One thread is responsible for real-time reading and preprocessing of sensor data, while the other thread performs fusion localization inference. A hardware-level time synchronization system based on PTP is constructed to achieve microsecond-level clock alignment between UWB radar and visual sensors. A data buffer queue with timestamp matching is constructed to pair asynchronously arriving radar point clouds and image sequences. Specifically, when data loss or delay is detected, a motion model-based prediction compensation mechanism is activated to interpolate missing observations.

[0101] In practical applications, sensor data acquisition, preprocessing, and deep learning inference have different computational characteristics and latency requirements. To ensure system real-time performance, a double-buffered pipeline architecture combined with a multi-threading mechanism is adopted to achieve task decoupling and efficient scheduling. Thread 1: Data Acquisition and Preprocessing Thread It is responsible for synchronously reading raw data from UWB radar and vision sensor drivers, performing lightweight preprocessing operations such as timestamp verification, noise reduction filtering (e.g., Savitzky-Golay), image distortion correction (cv2.undistort), and target clustering (DBSCAN), and writing the processed data into input buffer A or B.

[0102] Thread 2: Model Inference and Fusion Localization Thread Read aligned data pairs from the currently ready buffer, call the TensorRT-accelerated cross-modal fusion model to perform feature extraction, attention alignment, temporal modeling and joint localization inference, output the target's 6DoF pose information and write it to the output queue.

[0103] The double buffering mechanism avoids resource contention and blocking between producers (collection) and consumers (inference) by alternately reading and writing to two memory buffers, ensuring that the system runs stably at a constant frame rate (e.g., 10–30 FPS).

[0104] Radar and vision sensors typically use independent clock sources, resulting in clock drift ranging from microseconds to milliseconds, causing cross-modal data to be misaligned in the time dimension. To address this, a Precision Time Protocol (PTP, IEEE 1588) synchronization system is constructed. Deploy a PTP master clock on an edge computing platform (such as Jetson AGX Xavier) to send high-precision time synchronization messages to UWB radar and vision cameras via Ethernet or a dedicated synchronization line; The sensor-side hardware module receives the PTP timestamp and calibrates the local clock to achieve microsecond-level (<1μs) clock alignment; All collected data are tagged with a unified global timestamp to provide a high-precision time reference for subsequent cross-modal pairing.

[0105] Although hardware synchronization can significantly reduce clock skew, radar and vision data may still arrive asynchronously due to transmission delays, differences in interrupt response, and other factors. Therefore, a timestamp-aligned buffer queue is designed: Maintain two first-in-first-out (FIFO) queues to cache radar point clouds and visual images with timestamps, respectively; Set a reasonable time window (e.g. ±10ms) and find the radar-visual data pair with the closest time within the window; Once a match is found, the data pair is fed into the preprocessing pipeline to ensure that the data input to the model is highly consistent in time and space.

[0106] In extreme environments (such as strong electromagnetic interference or brief visual obstruction), data loss or severe delay may occur from a single sensor. To maintain continuous system operation, a motion model-based prediction and compensation mechanism is introduced: Utilize historical states (position, velocity, acceleration) to construct Kalman filter or LSTM trajectory prediction models; When a frame of visual or radar data fails to arrive on time, the system automatically switches to prediction mode and calculates the current target pose based on the previous state. Once the data is recovered, the new observations are immediately integrated to smoothly transition back to normal mode; This mechanism effectively prevents location jumps or system crashes caused by brief signal interruptions.

[0107] The multi-threaded, double-buffered architecture parallelizes data acquisition and model inference, fully utilizing the multi-core CPU and GPU resources of edge devices to significantly reduce end-to-end processing latency and meet the stringent response speed requirements of applications such as autonomous driving and robot navigation. The PTP hardware synchronization and timestamp matching mechanism ensures precise alignment of radar and visual data in the time dimension, avoiding feature mismatch and fusion errors caused by temporal misalignment, a prerequisite for achieving high-precision collaborative positioning. The prediction compensation mechanism enables the system to maintain short-term stable output even when some sensors fail or data is abnormal, providing "degraded operation" capability and significantly improving the overall system availability and security.

[0108] Specifically, the optimized collaborative localization model was deployed to the Jetson AGX Xavier edge computing platform. TensorRT was used to quantize and accelerate the model, converting the floating-point model to INT8 precision to reduce computational resource consumption. A double-buffering mechanism was designed to process the sensor data stream, with one thread handling data acquisition and another thread performing model inference to ensure real-time performance. Sensor-driven processing, data fusion, and localization result publishing were integrated using ROS and frameworks.

[0109] A hardware-triggered synchronization system was constructed, employing the PTP protocol to achieve time synchronization between the UWB radar and the vision sensor, achieving a synchronization accuracy at the microsecond level. A data buffer queue was designed to handle the timestamp alignment of sensor data. When data loss or delay is detected, a predictive compensation mechanism is automatically activated to maintain continuous system operation.

[0110] A multi-level anomaly detection mechanism is implemented, including sensor health monitoring, data validity checks, and verification of the rationality of positioning results. When a visual sensor failure is detected, the system automatically switches to pure radar positioning mode; when the radar signal is interfered with, the visual positioning weight is enhanced; when both are abnormal, motion model-based predictive state maintenance is activated to ensure the system's basic functionality under extreme conditions.

[0111] Develop a visual monitoring interface to display sensor data, location results, and system status in real time. Design a tiered alarm mechanism to provide differentiated alerts for anomalies of varying degrees, enabling operators to quickly locate and resolve issues. Provide an API interface for upper-layer applications to call the location service, supporting multiple communication protocols such as ROS messages and HTTP.

[0112] In some embodiments of this application, the trained cross-modal feature fusion model is deployed on an edge computing platform, and a multi-threaded mechanism is used to realize the real-time acquisition, processing, and localization output of radar point cloud data and visual image sequences. The model also includes: Perform online health monitoring of sensors, verify the integrity of input data, and judge the reasonableness of positioning results; When the visual sensor fails, it switches to a degraded positioning mode that relies primarily on radar. When radar signals are interfered with, increase the weight of the visual modality; In extreme cases where both visual sensor and radar signals are abnormal, the kinematic prediction module based on IMU or historical trajectory is activated to maintain short-term positioning output.

[0113] To achieve comprehensive monitoring of system health status, a three-level anomaly detection system is constructed: Sensor-based online health monitoring: The system reads the hardware status information of the UWB radar and vision sensors in real time (such as communication connection status, power supply voltage, temperature, frame rate, and packet loss rate), and uses a heartbeat mechanism to determine whether the device is online. If no data is received for several consecutive frames or communication is abnormal, it is determined to be a sensor failure.

[0114] Input data integrity verification: Each frame of input data undergoes a quality assessment. For example, it detects whether the image is blurry, abnormally exposed, or completely black (visual failure); and determines whether the radar point cloud is too sparse, has a low signal-to-noise ratio, or contains a large number of abnormal echoes (signal interference). Automatic identification can be achieved through preset thresholds or lightweight anomaly detection networks (such as Autoencoder reconstruction error).

[0115] Judgment of the reasonableness of the positioning results: Logical verification is performed on the pose results output by the model, such as detecting position jumps (sudden increases in displacement between adjacent frames), abnormal velocity (exceeding physical limits), and attitude angle jitter. The reliability of the output is determined by combining motion continuity constraints and contextual information (such as map priors).

[0116] Based on the above anomaly detection results, the system has intelligent decision-making capabilities and can automatically switch to the optimal working mode according to different fault scenarios: When the visual sensor fails: Switch to radar-dominated mode When the visual channel fails to provide effective data due to insufficient lighting, lens obstruction, or hardware malfunction, the system automatically shuts down the visual branch and fully shifts the fusion weights to the radar mode. In this mode, target detection and tracking rely solely on UWB radar point cloud data, combined with historical trajectory prediction to maintain basic positioning capabilities. While this mode sacrifices some semantic recognition capabilities, it still ensures stable tracking of moving targets.

[0117] When radar signals are interfered with: Enhance visual modality weights In metallic environments, scenarios with severe multipath reflections, or strong electromagnetic interference, radar may experience false alarms, drift, or target fragmentation. The system identifies these situations through signal-to-noise ratio analysis or point cloud stability detection, dynamically increasing the fusion weight of the visual branch. This makes the localization results rely more on high-confidence visual semantic information, thereby suppressing the negative impact of radar noise.

[0118] When both modes are abnormal: Enable the kinematic prediction module to maintain short-term output. In extreme cases (such as complete visual obstruction + strong radar interference), the system enters a "black box" state. At this time, the kinematic prediction module based on IMU or historical trajectories is activated: If the system integrates an IMU (Inertial Measurement Unit), then accelerometer and gyroscope data are used for dead reckoning, combined with Kalman filtering to estimate the current pose; If there is no IMU, an LSTM or multinomial extrapolation model is built based on the motion states (position, velocity, acceleration) of multiple consecutive historical frames to predict the target state at the next moment. The prediction results are used to maintain short-term (e.g., 1–3 seconds) positioning output to prevent the system from "disconnecting" and to provide operators with emergency response time.

[0119] To prevent location jumps or jitters during mode switching, a gradual weight transition strategy is introduced: Instead of hard switching, the fusion weights are continuously adjusted based on the sensor confidence scores; For example, when the visual confidence level drops from 0.9 to 0.3, the visual weight decreases linearly within 0.5 seconds, and the radar weight increases accordingly to ensure a smooth and continuous output trajectory.

[0120] The multi-level anomaly detection mechanism enables the system to have "self-diagnosis" capabilities, allowing it to identify and respond promptly in the early stages of a fault, thus preventing erroneous data from causing localization failures.

[0121] Through a degradation-based operation strategy, the system can still provide basic positioning services even in the event of single-sensor failure or dual-mode anomalies, exhibiting high availability with a "fault-free" characteristic and meeting the safety requirements of industrial applications. The system can dynamically adjust its fusion strategy based on real-time perception quality, simulating human compensatory behavior when sensory perception is limited, demonstrating the autonomy and adaptability of an intelligent perception system. The kinematic prediction module effectively prevents sudden changes or interruptions in positioning results, avoiding misjudgments or control instability caused by momentary target loss, making it particularly suitable for high-dynamic scenarios such as autonomous driving and drone obstacle avoidance. The automated anomaly handling mechanism reduces reliance on manual intervention, making it suitable for unattended scenarios such as remote inspections, underground operations, and nighttime security, thus lowering maintenance costs.

[0122] Based on a practically deployed UWB radar-vision collaborative localization method and system, a 15-day continuous test was conducted in various indoor and outdoor scenarios, covering typical working conditions such as single-target localization, multi-target tracking, dynamic occlusion, and changes in lighting. Data acquisition was performed simultaneously using a TI IWR6843ISK radar (100Hz sampling rate) and an Intel RealSense D455 camera (1280×720@30fps), with baseline values ​​obtained through the OptiTrack optical motion capture system (accuracy ±0.5mm). The dataset contains 20 scene variations, divided into training, validation, and test sets in a 7:2:1 ratio.

[0123] 1. Cross-modal positioning performance The verification was conducted from three dimensions: positioning accuracy, environmental adaptability, and modal loss robustness. Evaluation indicators: Absolute position error (APE): 3D Euclidean distance error; Relative accuracy (RP): The percentage reduction in error compared to a single-modal reference; Modal Failure Recovery Rate (MFR): The success rate of localization in the event of a single-mode failure.

[0124] Table 1 Experimental results of cross-modal positioning performance

[0125] 2. Multimodal fusion performance Specific verification was conducted on the feature fusion module: Evaluation indicators: Feature alignment error (FAE): distance across modal feature space; Gated Decision Consistency (GDC): The rationality of modal weight allocation; Latency: The processing time for a single frame.

[0126] Table 2 Experimental results of multimodal fusion performance

[0127] Experimental results show that the cross-modal positioning accuracy is significantly better than the single-modal scheme, especially in visual degradation scenarios. Furthermore, the gating fusion mechanism effectively improves modal failure recovery capability. Typical scenario test data demonstrates that even in scenarios with simultaneous visual occlusion and radar multipath interference, the system can still maintain a positioning accuracy within 0.15m, meeting industrial-grade application standards.

[0128] A second aspect of this application provides a cooperative positioning device based on UWB radar and vision, comprising: The data acquisition module is suitable for acquiring radar point cloud data and visual image sequences and performing joint calibration. Based on the radar point cloud data and visual image sequences, it generates target object information and annotates the 2D bounding box, 3D bounding box and pose information of the target object to construct a multimodal collaborative localization dataset containing diverse dynamic scenes. The feature extraction module is suitable for performing dual-branch feature extraction on radar point cloud data and visual image sequences in multimodal cooperative localization datasets, so as to extract radar spatiotemporal features that characterize the target's position, motion state and spatial structure, as well as visual semantic features that include object category, texture and contextual relationship. The model building module is suitable for building cross-modal feature fusion models based on dynamic temporal attention mechanisms. The cross-modal feature fusion model includes a spatiotemporal feature alignment layer, a temporal dependency modeling layer, and an adaptive joint representation layer. Based on the cross-modal feature fusion model, a unified and adaptive joint feature representation is generated from radar spatiotemporal features and visual semantic features. The model training module is suitable for training cross-modal feature fusion models based on the task priority mechanism and dynamic weighting strategy in the meta-learning framework, and for adjusting the training weights based on modal reliability. The output module is suitable for deploying the trained cross-modal feature fusion model on an edge computing platform and using a multi-threaded mechanism to realize the real-time acquisition, processing and localization output of radar point cloud data and visual image sequences.

[0129] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the UWB radar and vision-based collaborative positioning method described in any of the first aspects above.

[0130] Finally, the present invention also provides a non-volatile computer storage medium storing computer-executable instructions thereon, which, when executed by a processor, implement the cooperative localization method based on UWB radar and vision provided by the above methods.

[0131] For any parts not mentioned in this application, existing technologies may be used or referenced.

[0132] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0133] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A cooperative localization method based on UWB radar and vision, characterized in that, include: Radar point cloud data and visual image sequences are collected and jointly calibrated. Target object information is generated based on the radar point cloud data and the visual image sequences. The 2D bounding box, 3D bounding box, and pose information of the target object are labeled to construct a multimodal collaborative localization dataset containing diverse dynamic scenes. A two-branch feature extraction is performed on the radar point cloud data and visual image sequences in the multimodal cooperative localization dataset to extract radar spatiotemporal features that characterize the target position, motion state and spatial structure, as well as visual semantic features that include object category, texture and contextual relationship. A cross-modal feature fusion model is constructed based on a dynamic temporal attention mechanism. The cross-modal feature fusion model includes a spatiotemporal feature alignment layer, a temporal dependency modeling layer, and an adaptive joint representation layer. Based on the cross-modal feature fusion model, a unified and adaptive joint feature representation is generated from the radar spatiotemporal features and the visual semantic features. The cross-modal feature fusion model is trained based on the task prioritization mechanism and dynamic weighting strategy in the meta-learning framework, and the training weights are adjusted based on modal reliability. The trained cross-modal feature fusion model is deployed on an edge computing platform, and a multi-threaded mechanism is used to realize the real-time acquisition, processing and localization output of the radar point cloud data and the visual image sequence.

2. The cooperative localization method based on UWB radar and vision according to claim 1, characterized in that, The process involves collecting radar point cloud data and visual image sequences, performing joint calibration, generating target object information based on the radar point cloud data and visual image sequences, and labeling the 2D bounding box, 3D bounding box, and pose information of the target object to construct a multimodal collaborative localization dataset containing diverse dynamic scenes. Specifically: UWB radar and vision sensors are deployed synchronously in the target scene. Hardware-level time synchronization is achieved using FPGA trigger signals to collect radar spatiotemporal point cloud data and color-depth image sequences with precise timestamps, respectively. The spatial transformation matrix between the radar coordinate system and the camera coordinate system is solved by performing joint calibration using the Kalibr toolbox. LabelImg is used to annotate the target in the visual image with 2D bounding boxes. CloudCompare is used to annotate the radar point cloud with 3D position and contour. The annotation results are then fused to generate pose information that includes the target's unique identifier, three-dimensional spatial coordinates and attitude angles.

3. The cooperative localization method based on UWB radar and vision according to claim 1, characterized in that, The process involves performing a two-branch feature extraction on the radar point cloud data and visual image sequences in the multimodal collaborative localization dataset to extract radar spatiotemporal features representing target location, motion state, and spatial structure, as well as visual semantic features including object category, texture, and contextual relationships. Specifically: The raw radar point cloud data is smoothed by a Savitzky-Golay filter, and the filtered point cloud is dynamically separated by the DBSCAN clustering algorithm. The retained dynamic point cloud is input into the PointNet++ network for high-dimensional feature learning. Through its hierarchical sampling and grouping strategy, spatiotemporal features with spatial structure perception capability are extracted, and a 256-dimensional radar feature vector is output. Distortion correction is performed on the visual image. The OpenCV cv2.undistort function is used in combination with the camera intrinsic matrix K and distortion coefficient D to eliminate lens geometric distortion. The YOLO series of object detection models are used to locate the region of interest in the image and generate object detection boxes with confidence. After cropping the detected target region, it is input into the ResNet-50 convolutional neural network for deep semantic feature extraction. Through its residual structure, the appearance, texture and context information of the target are mined, and a 512-dimensional visual feature vector is output to represent the target's category attributes and visual semantic information.

4. The cooperative localization method based on UWB radar and vision according to claim 3, characterized in that, The cross-modal feature fusion model constructed based on the dynamic temporal attention mechanism includes a spatiotemporal feature alignment layer, a temporal dependency modeling layer, and an adaptive joint representation layer. Based on this model, a unified and adaptive joint feature representation is generated from the radar spatiotemporal features and the visual semantic features. Specifically: Based on the 256-dimensional radar feature vector and the 512-dimensional visual feature vector, a spatiotemporal attention mechanism is constructed to establish fine-grained associations between heterogeneous modalities. The cross-modal attention weights between each radar point and the image region are calculated through a learnable weight matrix. The original features are weighted and fused using the attention weights. A bidirectional LSTM is introduced to perform temporal modeling on the aligned feature sequences of consecutive frames, generating forward and backward hidden states respectively, capturing the temporal continuity of target motion, and outputting an enhanced feature representation containing contextual temporal information. An environmental interference factor is introduced to evaluate each mode in real time, and the evaluation results are embedded into a gating function to generate dynamic weights. The sigmoid activation function and element-wise multiplication operation are used to achieve adjustable fusion of contributions between modes, and a unified joint feature representation is generated through nonlinear transformation. A temporal consistency constraint is introduced to minimize the L2 error between the predicted position change between adjacent frames and the displacement calculated based on the velocity of the previous frame.

5. The cooperative localization method based on UWB radar and vision according to claim 1, characterized in that, The cross-modal feature fusion model is trained based on the task prioritization mechanism and dynamic weighting strategy in the meta-learning framework, and the training weights are adjusted based on modal reliability, specifically as follows: A model-agnostic meta-learning framework is used to meta-train the cross-modal feature fusion model. Multiple scene subsets with different environmental characteristics are divided in the multimodal dataset. A task priority mechanism with scene complexity weighting is introduced on the basis of the original MAML. The training scenes are divided into different priority levels according to dynamic occlusion density and illumination change amplitude. A modal reliability assessment module is constructed to calculate the confidence scores of the radar and vision branches in real time. The visual semantic features are used to generate environmental semantic vectors, and the modal confidence is output by combining learnable parameters and the sigmoid function. Based on the confidence level, a dynamic weighted loss function is designed to automatically reduce the loss weight and enhance the radar branch contribution when the vision is obstructed or the illumination is degraded, and to increase the visual modal weight when the radar is subjected to multipath or electromagnetic interference.

6. The cooperative localization method based on UWB radar and vision according to claim 1, characterized in that, The trained cross-modal feature fusion model is deployed on an edge computing platform, and a multi-threaded mechanism is used to achieve real-time acquisition, processing, and localization output of the radar point cloud data and the visual image sequence. Specifically: Based on a dual-buffered pipeline architecture, the data acquisition and model inference tasks are separated through a multi-threaded mechanism. One thread is responsible for real-time reading and preprocessing of sensor data, while the other thread performs fusion localization inference. A hardware-level time synchronization system based on PTP is constructed to achieve microsecond-level clock alignment between UWB radar and visual sensors. A data buffer queue with timestamp matching is constructed to pair asynchronously arriving radar point clouds and image sequences. Specifically, when data loss or delay is detected, a motion model-based prediction compensation mechanism is activated to interpolate missing observations.

7. The cooperative localization method based on UWB radar and vision according to claim 6, characterized in that, The step of deploying the trained cross-modal feature fusion model on an edge computing platform and using a multi-threaded mechanism to achieve real-time acquisition, processing, and localization output of the radar point cloud data and the visual image sequence also includes: Perform online health monitoring of sensors, verify the integrity of input data, and judge the reasonableness of positioning results; When the visual sensor fails, it switches to a degraded positioning mode that relies primarily on radar. When radar signals are interfered with, increase the weight of the visual modality; In extreme cases where both the visual sensor and the radar signal are abnormal, the kinematic prediction module based on IMU or historical trajectory is activated to maintain short-term positioning output.

8. A cooperative positioning device based on UWB radar and vision, characterized in that, include: The data acquisition module is suitable for acquiring radar point cloud data and visual image sequences and performing joint calibration. Based on the radar point cloud data and the visual image sequences, it generates target object information and labels the 2D bounding box, 3D bounding box and pose information of the target object to construct a multimodal collaborative localization dataset containing diverse dynamic scenes. The feature extraction module is adapted to perform dual-branch feature extraction on radar point cloud data and visual image sequences in the multimodal cooperative localization dataset, so as to extract radar spatiotemporal features that characterize the target position, motion state and spatial structure, as well as visual semantic features that include object category, texture and contextual relationship. The model building module is suitable for building a cross-modal feature fusion model based on a dynamic temporal attention mechanism. The cross-modal feature fusion model includes a spatiotemporal feature alignment layer, a temporal dependency modeling layer, and an adaptive joint representation layer. Based on the cross-modal feature fusion model, a unified and adaptive joint feature representation is generated from the radar spatiotemporal features and the visual semantic features. The model training module is suitable for training the cross-modal feature fusion model based on the task priority mechanism and dynamic weighting strategy in the meta-learning framework, and adjusting the training weights based on modal reliability. The result output module is suitable for deploying the trained cross-modal feature fusion model on an edge computing platform and using a multi-threading mechanism to realize the real-time acquisition, processing and positioning results output of the radar point cloud data and the visual image sequence.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the collaborative localization method based on UWB radar and vision as described in any one of claims 1 to 7.

10. A non-volatile computer storage medium storing computer-executable instructions thereon, characterized in that, When the computer program is executed by the processor, it implements the collaborative localization method based on UWB radar and vision as described in any one of claims 1 to 7.