Multi-modal data fusion-based multi-target tracking method for legged robot
Through multimodal data fusion and the improved DeepSORT algorithm, the accuracy and stability problems of multi-target tracking of legged robots in complex dynamic environments were solved, efficient and accurate multi-target tracking was achieved, and the ability of autonomous navigation and task execution was improved.
Patent Information
- Application Number
- CN202510564286.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-12
AI Technical Summary
Existing legged robots have low multimodal target tracking accuracy in complex dynamic environments, difficulty in data fusion, target occlusion and mismatching, making it difficult to achieve stable multi-target tracking.
A multimodal data fusion method is adopted to achieve stable tracking of multiple targets through timestamp alignment, feature extraction, BEV pooling, multi-task detection head and feature sharing module of sensor data such as vision, lidar and IMU, combined with Kalman filtering for target tracking, and cascade matching and trajectory generation using the improved DeepSORT algorithm.
The legged robot has improved its target detection and tracking accuracy and robustness in complex and highly dynamic scenes, and is able to continuously and stably identify and track multiple targets in a dynamic environment, enhancing its autonomous navigation and task execution capabilities.
Smart Images

Figure CN120635140A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot target tracking, and in particular to a multi-target tracking method based on multimodal data fusion for a legged robot. Background Art
[0002] In legged robot applications, multimodal target detection and tracking for highly dynamic scenes is a technical challenge. First, existing target tracking methods often rely on single-modal sensors, which are difficult to cope with dynamic changes in complex environments, resulting in low target detection and tracking accuracy. In particular, in situations of insufficient lighting, the presence of obstructions, or rapid scene changes, the data from a single sensor is easily interfered with or fails, affecting the stability and robustness of the system. Second, in the multimodal data fusion process, how to effectively integrate information from different sensors (such as vision, lidar, IMU, etc.) and achieve efficient and accurate data association and trajectory prediction is a problem that needs to be solved urgently. Third, existing multimodal target tracking methods have not yet achieved ideal performance in legged robot applications, especially for problems such as target occlusion, loss, and re-identification, for which existing technologies cannot provide effective solutions. Finally, how to implement this multimodal data fusion and multi-target tracking system on an embedded platform to ensure the system's real-time performance, lightweightness, and efficient operation is also a technical challenge that the present invention needs to overcome.
[0003] Legged robots, as highly flexible and autonomous mobile robots, have been widely researched and applied in recent years in fields such as environmental perception, target tracking, and intelligent navigation. Compared with traditional wheeled robots, legged robots have enhanced obstacle traversal capabilities and the ability to adapt to complex terrain, enabling more efficient locomotion on uneven surfaces. However, legged robots still face many technical challenges when tracking multiple targets in dynamic environments, particularly in multimodal perception data fusion, target recognition, and tracking accuracy.
[0004] First, the perception capabilities of traditional single-modal legged robots are limited. Traditional target tracking technologies are mostly based on single-modal data input, such as vision, lidar, or infrared sensors. However, these single-modal perception systems often expose their limitations in complex environments. For example, vision sensors are susceptible to interference in low-light conditions or highly complex environments; lidar may fail in the presence of obstructions or in rapidly changing scenes. Furthermore, data from a single sensor is susceptible to noise and errors, resulting in reduced target tracking accuracy. To address this issue, multimodal perception systems have gradually become a research focus. By fusing data from different sensors (such as vision, lidar, and IMU), the system's interference tolerance and adaptability can be effectively improved. When one sensor is interfered with, other sensors can compensate for the missing information, ensuring continuous target detection and tracking, and improving the target tracking capabilities of legged robots, especially in dynamic environments. Furthermore, different sensors have different adaptability to different target types. Multimodal fusion technology can integrate multiple sensors to adapt to a variety of target types and improve the versatility of target detection. While extensive research has been conducted in the autonomous driving field on multimodal fusion perception, this research in robotics is relatively limited. This is primarily due to differences in carriers, sensor types, scenarios, and target objects, which prevent direct data and algorithm reuse. Furthermore, robots must operate in multi-degree-of-freedom, dynamic environments, and the sensor perspectives, coordinate systems, and modeling methods differ significantly from those used in autonomous driving, further increasing the complexity of multimodal perception systems. Therefore, addressing these specific challenges and achieving efficient multimodal perception and target tracking has become a pressing challenge for legged robots operating in highly dynamic environments.
[0005] Secondly, when performing autonomous navigation and mission execution, legged robots need to continuously track multiple targets and adjust their strategies in a timely manner in complex dynamic scenarios. However, from the perspective of a legged robot, traditional methods struggle to achieve stable and continuous target tracking when faced with sudden changes in target appearance, severe occlusion of the target area, or target disappearance and reappearance. This makes it difficult to meet the diverse target perception requirements in complex and highly dynamic environments. In recent years, with the advancement of deep learning, image processing, and machine learning, researchers have proposed several target tracking methods based on multimodal data fusion, which have achieved significant progress in both theory and application. For example, multimodal tracking methods based on deep convolutional neural networks (CNNs) and recurrent neural networks (RNNs) can effectively extract and fuse features from multiple sensors, such as visual, infrared, and lidar, thereby improving the robustness and accuracy of target tracking. Furthermore, to address the challenges of multi-target tracking, researchers have proposed techniques based on trajectory similarity measurement and data association optimization to improve the accuracy and efficiency of multi-target tracking.
[0006] Third, while existing multimodal target tracking methods have achieved promising results in certain scenarios, combining target detection, data association, and trajectory prediction to achieve efficient and accurate multi-target tracking in complex environments remains an unresolved challenge for legged robot applications. In particular, dealing with sensor noise, target occlusion, and rapid target changes in dynamic and uncertain environments remains a challenge for multimodal target tracking systems.
[0007] For example, Chinese patent specification CN 117593620A discloses a multi-target detection method and device based on camera and lidar fusion, suitable for autonomous driving and robotics. By fusing lidar point cloud data with camera images, a deep learning model is used to classify and locate targets. While this solution solves the fusion of camera and radar data, the issue of perspective differences still exists. Due to the different operating principles and installation angles of cameras and lidars, there are significant discrepancies between the sensor data, resulting in significant deviations in the position of targets in the fields of view of different sensors.
[0008] Chinese patent specification CN 112731371A discloses an integrated target tracking system and method that integrates lidar and vision. While the proposed system tracks targets by fusing lidar and vision information, its fusion process relies primarily on spatial registration and target state prediction. This approach is prone to target positioning errors and tracking instability when sensor data is out of sync or experiences high dynamics.
[0009] In their paper "Fusion of LiDAR and Camera for 3D Object Detection in Autonomous Vehicles," published in IEEE Transactions on Intelligent Transportation Systems, Zhou, X., Sun, Z., et al. proposed a 3D object detection method based on the fusion of LiDAR and camera, aiming to address the object detection and localization challenges in autonomous vehicles. While this method has made progress in multi-sensor fusion, it still faces challenges such as viewpoint discrepancies, poor deduplication, and depth perception errors. In particular, duplicate detection and localization errors in areas where multiple sensors' fields of view overlap can hinder the execution of subsequent tasks.
[0010] Therefore, designing and implementing a target tracking system based on multimodal data fusion that can operate stably and accurately track multiple targets in a legged robot has become a key research topic. By effectively combining target detection, trajectory similarity calculation, multimodal data fusion, and optimization algorithms, the ability of legged robots to consistently and stably identify and track sensitive targets in complex and highly dynamic scenarios can be significantly improved, promoting their widespread application in fields such as automation, intelligent navigation, and the execution of complex tasks. Summary of the Invention
[0011] In response to the problems of low multi-modal target tracking accuracy, difficulty in data fusion, target occlusion and mismatching in existing legged robots in complex dynamic environments, this patent proposes a multi-target tracking method based on multi-modal data fusion for legged robots.
[0012] The technical solution of the present invention provides a multi-target tracking method based on multimodal data fusion for a legged robot, comprising the following steps:
[0013] S1. Data acquisition and preprocessing: Acquire data from multiple sensors and align the timestamps of the data in multiple sensor channels;
[0014] S2, data feature extraction: extract data features from the timestamp-aligned data of different sensor channels to generate high-dimensional feature representation;
[0015] S3, multi-channel feature fusion extraction: Based on the high-dimensional feature representation obtained in S2, the data features of the data from different sensor channels are fused to obtain a multimodal feature map, and the multi-task detection head processes the multimodal feature map to obtain the target detection result;
[0016] S4, Multi-target Tracking: The feature sharing module fuses the data features of different sensor channels and outputs them to the multi-target tracking module. The multi-target tracking module predicts the target's trajectory based on the target's state in the previous frame. Then, by calculating the target's motion information and spatial position, it correlates the matching degree between the target in the current frame and the tracked targets.
[0017] S5. Local trajectory generation: For each frame, the detected target is matched with the known target trajectory by calculating the similarity, and the successfully matched targets are associated to obtain a new local trajectory;
[0018] S6. Global trajectory generation: Calculate the similarity between the current trajectory and the existing trajectory based on the historical trajectory information of each target. If the similarity reaches the preset standard, they are merged into a global trajectory; otherwise, the trajectories remain independent.
[0019] Preferably, in the S2 data feature extraction step, the data features of the data from different sensor channels are fused to obtain a multimodal feature map, a convolutional neural network is used to extract image features obtained from the image, and a point cloud processing network is used to extract point cloud features containing spatial features and geometric information from the point cloud data.
[0020] Preferably, the S3 multi-channel feature fusion extraction step uses BEV pooling technology to map image features and point cloud features to a unified BEV space.
[0021] Preferably, the BEV pooling technology applies a coordinate attention mechanism to strengthen the correlation between features in the step of mapping image features and point cloud features into a unified BEV space.
[0022] Preferably, the S3 multi-channel feature fusion extraction step completes the accurate recognition of the target through the multi-task detection head and outputs the target's category, position, and direction information.
[0023] Preferably, the S4 multi-target tracking step maps the image features and point cloud features to a unified BEV space to obtain shared features through a feature sharing step. The fused shared features are passed to a subsequent target tracking module, which predicts the target's motion trajectory based on the target's state in the previous frame.
[0024] Preferably, the multi-target tracking step S4 further includes obtaining multi-scale features by performing multi-scale feature pyramid processing on the image features and the point cloud features in the target detection module;
[0025] The target tracking module reuses multi-scale features for target tracking.
[0026] Preferably, the target tracking module uses cascade matching to track the target, matching the detected target with the existing target trajectory to ensure the relevance of each target between consecutive frames and maintain the consistency of its trajectory.
[0027] Preferably, the IOU value between the targets matched at the current moment and the previous moment is also calculated, and the matching results are screened and filtered according to the IOU value. Only targets whose IOU value exceeds the set threshold will be retained as valid tracking targets.
[0028] Preferably, S5 local trajectory generation step: for each frame, the detected target is matched with the known target trajectory by calculating the similarity, and the successfully matched target is associated with a new local trajectory;
[0029] or
[0030] S6 global trajectory generation step: Calculate the similarity between each pair of trajectories. If the similarity is higher than the set similarity threshold, it is considered to be a continuous trajectory of the same target and merged into a global trajectory; otherwise, the trajectories are kept independent and continue to be tracked;
[0031] or
[0032] Before or after the timestamp alignment in the S1 data acquisition and preprocessing step, a preprocessing step is performed on the data of different sensor channels. The preprocessing step includes at least one of denoising and coordinate conversion operations on the different sensor channels.
[0033] The multi-target tracking method for a legged robot based on multimodal data fusion in this application extracts high-dimensional features from the data channels of different sensors by aligning their timestamps. Data fusion and target detection are then performed based on these extracted features, thereby achieving multi-target tracking. This data fusion process can improve the accuracy and robustness of subsequent target detection and tracking, enabling the continuous and stable identification and tracking of sensitive targets in complex and highly dynamic scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0035] Figure 2 The overall architecture of target detection and tracking of the present invention;
[0036] Figure 3 This is a diagram of the architecture of the multi-target tracking system of the present invention;
[0037] Figure 4 This is a lightweight deployment flow chart of the present invention. DETAILED DESCRIPTION
[0038] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. In this specification, the size ratios in the drawings do not represent the actual size ratios, but are only used to reflect the relative positional relationship and connection relationship between the various components. Components with the same name or the same number represent similar or identical structures and are only for illustrative purposes.
[0039] To address the challenges of low multimodal target tracking accuracy, data fusion difficulties, target occlusion, and mismatching in complex dynamic environments faced by existing legged robots, this patent proposes a multimodal data fusion-based multi-target tracking method and system for legged robots. This system integrates data from multiple sensors, including vision, lidar, and IMUs, and utilizes deep learning techniques for multimodal data fusion. This improves the quality of information acquired by the legged robot's sensing system, thereby enhancing the accuracy and robustness of subsequent target detection and tracking. To achieve consistent and stable identification and tracking of sensitive targets in complex, highly dynamic scenes, this patent proposes a multimodal collaborative multi-target detection and tracking algorithm based on a bird's-eye view, based on a legged robot system. By fusing information from multiple sensors, this algorithm overcomes the limitations of legged robots in environmental perception and enables effective target re-identification and association when targets are occluded or lost, ensuring accurate, real-time tracking of multiple targets. Furthermore, through optimized target detection, local trajectory generation, and similarity measurement modules, this patent effectively addresses interference and misidentification issues among multiple targets, improving the accuracy and robustness of target tracking.
[0040] A multi-target tracking method and system based on multimodal data fusion for a legged robot of the present invention comprises the following steps:
[0041] S1 Data Acquisition and Preprocessing. Acquire data from multiple sensors and align the timestamps of the data in multiple sensor channels. Generally speaking, multiple sensors may include data from lidar and RGB cameras, and the acquisition frequencies of the two may not be consistent. Timestamp alignment ensures that the processed data frames of different channels can be mapped to the same timestamp in a certain way. Before or after timestamp alignment, preprocessing can be performed on the data of different sensor channels, such as denoising and coordinate conversion, to ensure the standardization and validity of the data.
[0042] S2 Data Feature Extraction: Extract data features from the time-stamp aligned data of different sensor channels to generate high-dimensional feature representations.
[0043] S3 multi-channel feature fusion. Once the features of the multimodal data are extracted, the data features of the data from different sensor channels can be fused based on the high-dimensional feature representation to obtain a multimodal feature map. In the patent of this invention, the BEV (Bird's Eye View) pooling technology is first used to map the features of the RGB image and the LiDAR point cloud into a unified BEV space. At this time, the data of different modalities are converted to the same coordinate space to facilitate subsequent fusion. The model will detect the multimodal feature map through the multi-task detection head to complete the accurate recognition of the target, and output the target's category, position, direction and other information. This stage is mainly based on the detection and positioning of multiple targets based on the fused features. Simply put, the multi-task detection head processes the multimodal feature map to obtain the target detection result;
[0044] S4 Multi-target Tracking. The feature sharing module is used to fuse the data features of different sensor channels and output them to the multi-target tracking module. The multi-target tracking module first uses the multimodal feature map transmitted from the feature sharing module, combined with the state of the target in the previous frame (such as position, velocity, etc.), to predict the target's motion trajectory through a motion model (such as Kalman filtering). Then, by calculating the target's motion information (such as velocity, acceleration) and spatial position, the matching degree between the target in the current frame and the tracked target is accurately associated, thereby completing the continuous tracking of the target, generating a preliminary target trajectory and outputting it to the cascade matching.
[0045] S5 local trajectory generation.
[0046] For each frame, detected targets are matched against known target trajectories by calculating similarity. Successfully matched targets are then associated to form a new local trajectory. Trajectory updates can utilize a Kalman filter for prediction and correction, ensuring that each target's trajectory remains consistent across space. New targets are initialized, and missing targets are marked by setting a loss threshold (e.g., if a target fails to match across multiple consecutive frames).
[0047] S6. Global trajectory generation. Based on each target's historical trajectory information (such as position, velocity, acceleration, etc.), the similarity between the current trajectory and the existing trajectory is calculated. If the similarity meets the preset standard, they are merged into a global trajectory; otherwise, the trajectories remain independent and continue to be tracked.
[0048] Step 1: Multimodal data acquisition and preprocessing
[0049] Legged robots are usually equipped with multiple sensors, such as RGB cameras, LiDAR, IMU, etc., in order to obtain comprehensive perception information in complex environments. The purpose of this step is to obtain and process data from multiple sensors, laying the foundation for subsequent target detection, tracking and multimodal data fusion. It mainly includes collecting data from multiple sensors, performing time synchronization and calibration, preprocessing data (including denoising, size standardization, coordinate transformation, etc.), and performing data fusion and coordination to ensure that the data from all sensors can be efficiently integrated under a unified reference framework.
[0050] Specifically, multimodal data acquisition and preprocessing includes the following sub-steps.
[0051] Step 1.1 Multimodal Sensor Data Acquisition: Start the RGB camera and LiDAR sensor to begin real-time data acquisition. The RGB camera captures visible light image sequences, while the LiDAR captures 3D point cloud data. Point cloud images are typically composed of many 3D coordinates (x, y, z) and accurately reflect the scene geometry. They are particularly advantageous over RGB images in low-light or complex environments.
[0052] Step 1.2: Data preprocessing and timestamp alignment. First, perform denoising and grayscale conversion on the RGB image, and filter and remove invalid points on the radar point cloud. These processes remove redundant information, ensure sensor data quality, and provide effective input for subsequent feature extraction and data fusion. Next, align the timestamps of each sensor to ensure that data from different sensors can be compared and fused using the same time base. For example, if the RGB image acquisition frequency is 30 Hz and the radar's is 10 Hz, the radar point cloud data can be supplemented to the timestamp corresponding to the RGB image through interpolation or data multiplexing. Alternatively, alignment between the RGB camera and radar can be achieved by discarding intermediate RGB frames. Other acceptable alignment methods are also acceptable, as long as they can adjust the data from different channels to a format with the same time series.
[0053] Step 2: Multimodal feature extraction based on deep learning
[0054] This step is a precursor to data fusion, and its purpose is to extract useful features from the RGB image and LiDAR point cloud data collected in step 1 and generate high-dimensional feature representations.
[0055] The core of multimodal feature extraction is to design a suitable neural network architecture that can process RGB images and LiDAR point cloud data separately to extract discriminative features. Convolutional neural networks (CNNs) are used to extract features from images, such as object edges, texture, and color. Point cloud processing networks such as PointNet or PointNet++ are used to extract spatial features and geometric information from point cloud data.
[0056] Step 3: Multimodal collaborative BEV multi-target detection
[0057] Once the features of the multimodal data are extracted, the next step is to fuse these features. This patent first uses BEV (Bird's Eye View) pooling technology to map the features of the RGB image and LiDAR point cloud into a unified BEV space. At this point, the data from different modalities is converted to the same coordinate space, facilitating subsequent fusion.
[0058] Furthermore, considering the possible spatial correlation between features of different modalities, a coordinate attention mechanism is used to strengthen the correlation between features. This step enables the model to adaptively focus on the most important parts of the multimodal data. Finally, after BEV feature fusion, the model uses the multi-task detection head to accurately identify the target and output information such as the target's category, location, and direction. This stage mainly detects and locates multiple targets based on the fused features. The details are as follows:
[0059] Step 3.1: BEV pooling and alignment of multimodal data. After feature extraction, convolutional layers (e.g., 2D convolution) are used to process the RGB image features and map them to the BEV view space. A 3D convolution or projection algorithm is then performed on the LiDAR point cloud data to convert the 3D point cloud data into a 2D feature representation in the BEV space. Common methods include voxelization of the laser point cloud or projection of the point cloud.
[0060] The BEV view is a bird's-eye view suitable for multi-target detection tasks. Through BEV pooling, the spatial information of RGB images and LiDAR point clouds can be uniformly encoded in the same space. Then, geometric transformation is implemented based on camera calibration information or sensor installation position for feature alignment. Finally, the aligned RGB and LiDAR features are fused in the BEV space to form a multimodal joint representation.
[0061] Step 3.2 introduces the coordinate attention mechanism. In the BEV space after feature fusion, the coordinate attention mechanism weights features at different positions by introducing spatial coordinate information. Specifically, coordinate information (such as x and y coordinates) is used as input to the attention module to generate attention weights for each feature position. These weights are used to dynamically adjust the importance of features at different spatial positions during the feature fusion process. Based on the coordinate attention mechanism, features are weighted according to the attention weights of their spatial positions. High-weight areas correspond to more critical parts in the space, while low-weight areas are considered less important parts. In this way, the network can automatically focus on areas with higher semantic information, thereby enhancing the effect of target detection. Finally, the coordinate attention mechanism is combined with the original feature fusion module (such as multimodal feature fusion), and the weighted features are input into the multi-task detection head in step 3.3.
[0062] Step 3.3 Multi-task detection head design. In the multi-task detection head, all tasks share some input features. The multimodal feature map after BEV pooling and coordinate attention mechanism weighting will be fed into the multi-task detection head. At this stage, the network extracts global information through shared layers (such as convolutional layers or fully connected layers) to ensure that each task can obtain sufficient contextual information. These shared feature layers complete information extraction in the early stages of the network. Subsequently, different branches are assigned to different tasks for processing, each corresponding to a specific detection task:
[0063] Object classification branch: used to predict the category of the detected object.
[0064] Target position regression branch: used to predict the location of the target (usually the bounding box coordinates of the target).
[0065] Target direction regression branch: used to predict the direction or angle of the target.
[0066] The multi-task detection head operates in parallel through multiple output layers, completing object classification, localization, and orientation prediction within the same network. This ensures that all types of information are fully utilized during the detection process. Once the classification, position regression, and orientation regression outputs are generated by the multi-task detection head, the network further processes these outputs based on the task requirements. For each object, the classification result is combined with the bounding box position and orientation information to generate the final detection result.
[0067] Step 4: Multimodal collaborative 3D target tracking
[0068] Step 4.1 Feature Sharing. Since the RGB image and LiDAR features have been mapped to the unified BEV space in step 3.1, the tracking module now receives a pooled and aligned multimodal feature map as input, containing the combined information from the RGB image and LiDAR point cloud. To better fuse these different modal features and provide more effective input for the subsequent object tracking task, the feature sharing module further strengthens the integration of these two modal information.
[0069] Specifically, the feature sharing module fuses the feature maps of the RGB image and the LiDAR point cloud in the BEV space. These pooled and aligned feature maps can be further processed through convolution operations (such as 2D convolution), thereby achieving spatial weighting and fusion of features from different sensors. To achieve this goal, methods such as weighted summation, feature concatenation, or self-attention mechanisms can be used. These methods help highlight the important information of each modality, enhance their complementarity, and provide richer spatial information to the tracking module.
[0070] After this step is completed, the fused shared features are passed to the subsequent target tracking module. Through this feature sharing, the tracking module can simultaneously access data features from RGB and LiDAR, and fully utilize their spatial information and complementary features, thus providing multi-source data support for target tracking and avoiding information loss or asymmetry between modalities.
[0071] Step 4.2 Multi-scale Feature Fusion. In target tracking tasks, the target may change in scale, especially when the target moves quickly, is occluded, or appears different in size from different viewpoints. Therefore, capturing the changes in the target at different scales becomes particularly important. Multi-scale feature fusion helps the tracker maintain target consistency, especially when the target changes rapidly between consecutive frames or moves away from the sensor, thereby improving tracking accuracy and robustness.
[0072] In the object detection module, RGB images and LiDAR point cloud data are processed using a multi-scale feature pyramid, extracting multi-scale features rich in spatial information. Reusing these multi-scale features in the tracking module is crucial for consistently and accurately tracking the target. Because the tracking module's primary task is to correlate the target's position and motion across different frames, it eliminates the need for complex feature extraction and can directly reuse the multi-scale features extracted in the object detection module, reducing computational burden and improving efficiency.
[0073] By fusing these multi-scale features, the tracking module can handle variations in the target at different scales, making the tracking process more robust to challenges such as changes in target size and shape. This fusion helps enhance the tracking system's ability to perceive targets at different scales, distances, and motion states, improving the accuracy and robustness of target tracking.
[0074] Step 4.3 Multi-target Tracker. The tracker first uses the multimodal feature map passed from the feature sharing module, combined with the target's state in the previous frame (such as position and velocity), to predict the target's trajectory using a motion model (such as a Kalman filter). Then, by calculating the target's motion information (such as velocity and acceleration) and spatial position, it accurately correlates the matching degree between the target in the current frame and the tracked targets, thereby completing continuous target tracking and generating a preliminary target trajectory that is output to the cascade matching.
[0075] Step 4.4: Cascade Matching. During the multi-target tracker process in Step 4.3, cascade matching is used to enhance the accuracy and robustness of target tracking. This step matches detected targets with existing target trajectories to ensure the relevance of each target across consecutive frames and maintain the consistency of its trajectory. This multiple matching process effectively achieves stable target tracking, especially in situations with a large number of targets and frequent occlusions.
[0076] Step 4.5: Target Matching Screening. After the cascade matching is complete, the system also calculates the IOU (Intersection Over Union) value between the current and previously matched targets, and screens and filters the matching results based on the IOU value. Only targets whose IOU values exceed the set threshold are retained as valid tracking targets, resulting in the final target tracking result. This process ensures the accuracy of target matching, avoids false matches, and further improves the robustness and reliability of the system in complex scenarios.
[0077] The following are the specific steps for improving the cascade matching in the DeepSORT algorithm:
[0078] Calculate target similarity: First, calculate the similarity between the detected target and the existing trajectory. Usually it is measured by position similarity (Euclidean distance) and appearance similarity (through deep feature extraction). For target i and trajectory j, the similarity can be expressed as:
[0079] S ij =α·d(x i ,x j )+β·c(f i ,f j )
[0080] Among them, d(xi ,x j ) is the position distance, c(f i ,f j ) is the cosine similarity of the feature vectors, and α and β are weight coefficients that control the importance of position and appearance.
[0081] Preliminary matching (strict matching): Use the greedy algorithm or Hungarian algorithm to perform preliminary matching of targets based on the above similarity. For each pair of matches (detected target and existing track), if the similarity is greater than the set threshold S ij >τ, the match is considered successful.
[0082] Cascade matching (loose matching): If not enough matching pairs are found in the strict matching stage or there are missing targets, the threshold τ is relaxed in the next round of loose matching and the similarity between the target and the trajectory is recalculated. At this time, the tolerance for position and appearance features is increased, allowing some matches with larger errors to appear.
[0083] Fault-tolerant matching: In order to cope with short-term target occlusion or rapid changes, the system further introduces a fault-tolerant mechanism and expands the matching conditions, such as the position error d(x i ,x j ) to a larger value, or adopt a matching strategy based on speed or acceleration. Targets can be considered matched even if they are relatively far away.
[0084] Through this cascade matching process, the improved DeepSORT is able to track multiple targets more robustly in scenes where targets are occluded, lost, or changing rapidly.
[0085] Step 5: Local trajectory generation
[0086] Step 5.1 Local trajectory update: In the initial frame, an independent trajectory ID is generated for each target, and the target's position, velocity, acceleration and other state information are recorded. For each frame, the detected target is matched with the known target trajectory by calculating the similarity, and the successfully matched target is associated with a new local trajectory. If a target has appeared in the previous frames, it will be added to the corresponding trajectory to generate a new local trajectory. This local trajectory contains the position information of the target in the current frame, as well as its historical position and velocity information in the previous frame. Trajectory updates use Kalman filters for prediction and correction, so that the trajectory of each target remains consistent in space. New targets are initialized, and lost targets are marked by setting a loss threshold (such as failure to match for multiple consecutive frames).
[0087] Step 5.2 Output the final target trajectory: Finally, the system outputs the trajectory information of each target, including dynamic characteristics such as position and speed, for subsequent tracking tasks or decision-making systems. Represents the entire set of local trajectories:
[0088]
[0089] in represents the set of local trajectories observed for the i-th type of data.
[0090] Step 6: Global trajectory generation for multi-target tracking
[0091] Step 6.1: Trajectory similarity measurement
[0092] The purpose of trajectory similarity measurement is to calculate the similarity between each pair of trajectories to help determine whether the target is the same object and avoid trajectory splitting or merging. The specific implementation process is as follows:
[0093] First, the similarity between the current trajectory and the existing trajectory is calculated based on the historical trajectory information of each target (such as position, speed, acceleration, etc.). The trajectory similarity measurement module uses the output of step 5.2 as the As input, assume the trajectory and trajectory in and denote the local trajectories of trajectories i and j respectively, P i (t k ) and P j (t k ) represent the trajectory and At the time t, the Euclidean distance of the trajectory location information is calculated using the following formula to measure the location similarity:
[0094]
[0095] Then, construct the trajectory similarity matrix S, where each element of the matrix S ij Represents trajectory and The similarity score between them can be expressed as follows:
[0096]
[0097] Among them, σ is a hyperparameter that controls the speed of similarity decay. Smaller σ values make the similarity decrease rapidly when the distance is small, while larger σ values make the similarity more persistent. This method can convert the Euclidean distance into a similarity score in the range of [0,1]. Then, a clustering algorithm is applied to group the trajectories to identify which trajectories belong to the same target and which belong to different targets. Or by setting a similarity threshold α, if S ij >α, then it is considered and Belong to the same target. Trajectories with high similarity will be considered as continuous trajectories of the same target, while trajectories with low similarity may be different targets or interference targets.
[0098] Step 6.2: Based on the results of the trajectory similarity measurement in step 6.1, for local trajectories with high similarity, they represent the continuous movement of the same target in different frames, so they are merged into a global trajectory; for trajectories with low similarity, it means that they may represent different targets or the movement of the same target in different time frames, but for some reason (such as occlusion, fast movement), they cannot be successfully matched. These trajectories will remain independent and continue to be tracked. Their historical information will continue to be updated, and as new data is further confirmed to belong to new targets, if these independent trajectories match the new target detection results later, they may be merged again. Otherwise, they will continue to be tracked as independent trajectories.
[0099] The multi-target tracking algorithm based on multimodal data fusion of the present invention is deployed on a legged robot platform. Taking into account the real-time requirements of the system, the algorithm is hardware accelerated and software optimized. A lightweight model is designed in combination with the application scenario, and pruning, quantization, decomposition and distillation methods are used to reduce the model's computational complexity and memory usage; parallel computing is used to accelerate the calculation process, and mixed precision computing is used in the training and inference processes to reduce power consumption and latency. The lightweight deployment process of this solution is as follows: Figure 3 shown.
[0100] This patent proposes a multi-target tracking method and system based on multimodal data fusion for legged robots, aiming to solve a series of key technical problems in the collaborative fusion of multimodal data and multi-target tracking of legged robots in high-dynamic scenarios. Specifically, by fusing data from different types of sensors, accurate and stable target detection and tracking are achieved in high-dynamic scenarios. At the same time, combined with the multimodal pyramid feature module, an improved DeepSort algorithm is designed to improve the efficiency and accuracy of multi-target tracking. On this basis, taking into account the limitations of the embedded platform, the present invention also performs lightweight processing on the algorithm to achieve high-performance, low-latency deployment and operation.
[0101] With the help of the method and system of the present invention, legged robots can achieve efficient and accurate multimodal target detection and multi-target tracking in highly dynamic and complex scenarios. In particular, they can stably and continuously lock onto and track multiple targets in complex terrain and dynamic environments. In particular, the system can not only effectively fuse data from multiple sensors to improve the accuracy and robustness of target detection, but also achieve real-time and efficient multi-target tracking through an improved DeepSort algorithm. Using this system, legged robots can perform tasks such as multi-target detection and tracking in complex, highly dynamic, extreme, and dangerous situations, ensuring efficient perception and decision-making capabilities in dynamic environments, and realizing high-level functions such as autonomous navigation, obstacle avoidance, and task execution.
[0102] This technical solution constructs a multi-target BEV detection model based on multimodal fusion. Taking into account the sensor heterogeneity of visible light cameras and LiDAR sensors themselves, as well as the different data features between multi-source heterogeneous data, this method fully and complementary utilizes the feature information of multimodal data. The multimodal features of RGB images and LiDAR point cloud data are encoded into the same BEV space through BEV pooling to obtain features based on BEV unified representation. At the same time, considering the possible feature position relationship between multimodal data features, this method introduces a coordinate attention mechanism in the BEV feature fusion stage. Finally, the multi-task detection head completes the accurate recognition of multiple targets respectively, and outputs the category, position, direction and other information of each type of detection target. This method effectively solves the heterogeneity problem in multimodal data fusion and improves the accuracy and robustness of multi-target detection, especially the target recognition and tracking capabilities in complex and dynamic environments.
[0103] This technical solution targets the multi-target tracking task of legged robots and designs a DeepSORT legged robot tracking algorithm based on multimodal data fusion as a local trajectory generator. Building on the existing DeepSORT algorithm, it uses a deep neural network to extract features for target detection in RGB and LiDAR point clouds, and uses a feature pyramid to fuse multiple layers of features, thereby improving tracking accuracy and robustness. At the same time, this solution also provides a solution to the misidentification problem in multi-target tracking through a trajectory similarity measurement method. This method can effectively avoid misidentification and misassociation caused by factors such as changes in target appearance, rapid movement, or interfering targets. The system can determine which trajectories belong to the same target by quantifying the similarity between trajectories, thereby reducing the risk of mismatching and misidentification, and ensuring that in a multi-target environment, the legged robot can accurately identify and maintain tracking of each target.
[0104] This technical solution targets legged robots for use in complex environments, specifically considering the hardware limitations of embedded platforms and implementing a lightweight algorithm design. This optimization enables the system to operate stably in highly dynamic and complex scenarios, while also enabling low-latency, high-performance real-time processing on the embedded platform, enabling the legged robot to perform complex tasks such as autonomous navigation, obstacle avoidance, and multi-target tracking.
[0105] The above content only describes the preferred embodiments of the present invention and does not limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solution of the present invention by ordinary technicians in this field should fall within the scope of protection determined by the claims of the present invention.
Claims
1. A multi-target tracking method based on multimodal data fusion for a legged robot, characterized in that: The steps include: S1. Data acquisition and preprocessing: Acquire data from multiple sensors and align the timestamps of the data in multiple sensor channels; S2, data feature extraction: extract data features from the timestamp-aligned data of different sensor channels to generate high-dimensional feature representation; S3, multi-channel feature fusion extraction: Based on the high-dimensional feature representation obtained in S2, the data features of the data from different sensor channels are fused to obtain a multimodal feature map, and the multi-task detection head processes the multimodal feature map to obtain the target detection result; S4, Multi-target Tracking: The feature sharing module fuses data features from different sensor channels and outputs them to the multi-target tracking module. The multi-target tracking module predicts the target's trajectory based on its state in the previous frame. Then, by calculating the target's motion information and spatial position, it correlates the matching degree between the target in the current frame and the tracked targets. S5. Local trajectory generation: For each frame, the detected target is matched with the known target trajectory by calculating the similarity, and the successfully matched targets are associated to obtain a new local trajectory; S6. Global trajectory generation: Calculate the similarity between the current trajectory and the existing trajectory based on the historical trajectory information of each target. If the similarity reaches the preset standard, they are merged into a global trajectory; otherwise, the trajectories remain independent.
2. The multi-target tracking method based on multimodal data fusion for a legged robot according to claim 1, characterized in that: In the S2 data feature extraction step, the data features of data from different sensor channels are fused to obtain a multimodal feature map. A convolutional neural network is used to extract image features from the image, and a point cloud processing network is used to extract point cloud features containing spatial features and geometric information from the point cloud data.
3. The multi-target tracking method based on multimodal data fusion for a legged robot according to claim 2, characterized in that: The S3 multi-channel feature fusion extraction step uses the BEV pooling technique to map image features and point cloud features into a unified BEV space.
4. The multi-target tracking method based on multimodal data fusion for a legged robot according to claim 2, wherein: The BEV pooling technology maps image features and point cloud features into a unified BEV space and applies the coordinate attention mechanism to enhance the correlation between features.
5. The multi-target tracking method based on multimodal data fusion for a legged robot according to claim 2, characterized in that: The S3 multi-channel feature fusion extraction step uses a multi-task detection head to accurately identify the target and output the target's category, location, and direction information.
6. The multi-target tracking method based on multimodal data fusion for a legged robot according to claim 2, characterized in that: The S4 multi-target tracking step maps image features and point cloud features to a unified BEV space to obtain shared features through the feature sharing step. The fused shared features are passed to the subsequent target tracking module, which predicts the target's motion trajectory based on the target's state in the previous frame.
7. The multi-target tracking method based on multimodal data fusion for a legged robot according to claim 6, characterized in that: The S4 multi-target tracking step also includes obtaining multi-scale features by performing multi-scale feature pyramid processing on image features and point cloud features in the target detection module; The target tracking module reuses multi-scale features for target tracking.
8. The multi-target tracking method based on multimodal data fusion for a legged robot according to claim 6, characterized in that: The target tracking module uses cascade matching to track targets, matching the detected targets with existing target trajectories to ensure the relevance of each target between consecutive frames and maintain the consistency of its trajectory.
9. The multi-target tracking method based on multimodal data fusion for a legged robot according to claim 6, characterized in that: The IOU value between the current and previous matched targets is also calculated, and the matching results are screened and filtered according to the IOU value. Only targets whose IOU value exceeds the set threshold will be retained as valid tracking targets.
10. The multi-target tracking method based on multimodal data fusion for a legged robot according to claim 1, wherein: S5 local trajectory generation step: For each frame, the detected target is matched with the known target trajectory by calculating the similarity, and the successfully matched target is associated with a new local trajectory; or S6 global trajectory generation step: Calculate the similarity between each pair of trajectories. If the similarity is higher than the set similarity threshold, it is considered to be a continuous trajectory of the same target and merged into a global trajectory; otherwise, the trajectories are kept independent and continue to be tracked; or Before or after the timestamp alignment in the S1 data acquisition and preprocessing step, a preprocessing step is performed on the data of different sensor channels. The preprocessing step includes at least one of denoising and coordinate conversion operations on the different sensor channels.
Citation Information
Patent Citations
Laser radar and vision fusion integrated target tracking system and method
CN112731371A
Multi-target detection method and device based on camera and laser radar fusion
CN117593620A
Cited By
Multi-target continuous motion trail generation method and device for intersection scene
CN121259046A
Multi-target continuous motion trajectory generation method and device for intersection scene
CN121259046B
Method for determining motion parameters of humanoid robot and related equipment
CN121960563A