A 3D Occlusion Target Tracking Method Based on Multimodal Spatiotemporal Interaction

By employing a multimodal spatiotemporal interaction method, the problem of unstable detection boxes and trajectories in 3D occluded target tracking is solved, enabling accurate tracking of occluded targets and improving the robustness and accuracy of target recognition and tracking.

CN120747169BActive Publication Date: 2025-10-31NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511195162.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-10-31
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

In 3D moving target tracking, occlusion leads to low confidence in target detection boxes, unstable detection boxes and trajectories, making it difficult to accurately track occluded targets. Existing algorithms fail to effectively utilize the interaction features between point cloud and image data when fusing them, resulting in occluded targets being filtered out or inaccurate trajectory prediction.

Method used

A multimodal spatiotemporal interaction method is adopted. Point cloud and image features are extracted through a dual-branch structure, and three-dimensional spatial information and appearance texture information are fused to reconstruct the complete high-dimensional features of the occluded target. The detection box is processed by denoising convolution of reconstructed feature edges and feedforward neural network. The detection box and trajectory are matched by combining bipartite graph structure, adaptive channel graph convolution and Hungarian algorithm to achieve full-process tracking of the target.

Benefits of technology

It improves the accuracy and robustness of occluded target recognition, ensures the accuracy of detection boxes and the reliability of trajectories, reduces false detections and redundancy, and improves target tracking capabilities and matching efficiency in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747169B_ABST
    Figure CN120747169B_ABST
Patent Text Reader

Abstract

This application discloses a 3D occluded target tracking method based on multimodal spatiotemporal interaction, relating to the field of target tracking technology. The method includes acquiring and preprocessing point clouds and images to obtain global fusion features; processing with a region proposal network to obtain initial detection boxes and region of interest features; projecting non-empty voxel point clouds onto image features to reconstruct occluded target features; using convolution and neural network processing to obtain refined detection boxes; filtering valid detection boxes through distance calculation and validity judgment; calculating appearance association scores using bipartite graph and adaptive channel graph convolution; and matching detection boxes and trajectories based on the Hungarian algorithm to achieve full-process tracking. This application improves the accuracy and robustness of target recognition, addresses the occlusion problem, ensures the accuracy of detection boxes, improves association accuracy using bipartite graph and adaptive convolution, and combines geometric cost matrix matching nodes to achieve full-process tracking and error calibration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of target tracking technology, and in particular relates to a three-dimensional occlusion target tracking method based on multimodal spatiotemporal interaction. Background Technology

[0002] In the field of 3D moving target tracking technology, accurately identifying and establishing the correspondence between the same target across frames of data from a series of sensors is a core task. However, in practical applications, overlapping and occlusion between targets are extremely common, posing numerous challenges to 3D moving target tracking. Targets in occlusion states can be completely visible, partially occluded, or completely occluded. During sensor data acquisition, occluded targets cause a series of problems. For laser sensors, the laser beam from an occluded target will be reflected prematurely; for cameras, the light will be truncated. These phenomena directly lead to sparse point clouds in the occluded area, resulting in missing point cloud and image state of the target. Due to the incomplete target state, the detection difficulty of occluded targets increases significantly. Specifically, the detection box exhibits jitter, leading to unstable trajectories and low detection box confidence. Incomplete spatial states also reduce the distinguishability between targets, making the correlation between targets even more difficult.

[0003] To improve the reliability of detection boxes and ensure trajectory quality, most algorithms typically use thresholds to filter detection boxes with low confidence levels during the preprocessing module. However, occluded targets usually have low confidence levels, and this filtering method can lead to occluded targets being filtered out, thus preventing effective tracking. Most algorithms use a single motion model to predict the position, size, and heading angle of the bounding box. However, different state variables have different motion characteristics, and using the same motion model can lead to inaccurate trajectory predictions. For targets that reappear after complete occlusion, the position prediction deviation will accumulate and exceed the associated threshold, amplifying the inaccuracy and making the predicted box position deviation more pronounced, ultimately resulting in the inability to recover the trajectory after the occluded target reappears.

[0004] Point cloud data contains the 3D coordinates of a target's surface in the surrounding environment, possessing rich depth information and providing accurate localization and geometric features. Algorithms such as VoxelNet, SECOND, and PointPillars use LiDAR-based target detection algorithms to obtain accurate target locations. However, due to the inherent limitations of LiDAR sensors, point cloud data is relatively sparse, especially for occluded targets, where the point cloud is almost empty, making identification difficult. To compensate for the deficiencies of point cloud data, algorithms such as MVX-Net, MV3D, AVOD, RangeLVDet, EPNet++, and 3D-CVF incorporate image data, fusing point cloud and image data. However, the two types of data have fundamentally different representations, leading to numerous problems during the fusion process. Information conversion and representation in the point-level fusion process are complex and inefficient; detection result-level fusion requires two detectors to merge the tracking results of the image and point cloud to complete the final tracking, without considering the interaction between point cloud data and image data information, resulting in insufficient utilization of the potential features of point cloud and image fusion; feature-level fusion can achieve a balance between the two, but the occluded target point cloud state is incomplete, the usable information of the target after fusion is scarce, and the target is still difficult to identify. Summary of the Invention

[0005] The purpose of this application is to provide a 3D occluded target tracking method based on multimodal spatiotemporal interaction. Through steps such as multimodal data fusion, occluded target feature reconstruction, feature edge denoising and thinning, detection box filtering and legality judgment, legal bounding box screening and trajectory marking, bipartite graph structure and adaptive channel graph convolution, generalized intersection union and final correlation matrix, as well as Hungarian algorithm and trajectory prediction, this method solves the technical problem of difficult accurate tracking of occluded targets in 3D scenes.

[0006] To achieve the above objectives, embodiments of this application provide a method for tracking 3D occlusion targets based on multimodal spatiotemporal interaction, including:

[0007] The original point cloud and the original image are acquired at the same time. The original point cloud and the original image are preprocessed to obtain global fusion features.

[0008] The global fusion features are processed by a region proposal network to obtain initial detection boxes. The initial detection boxes and global fusion features are then processed by region of interest pooling to obtain global fusion features of the region of interest. The global fusion features of the region of interest are then separated to obtain global point cloud features and global image features.

[0009] Global point cloud features include non-empty voxel point cloud features. The non-empty voxel point cloud features are projected onto global image features to obtain the visible image features of the occluded target. The visible image features of the occluded target are reconstructed and decoded to obtain the complete high-dimensional features of the occluded target.

[0010] The complete high-dimensional features of the occluded target are smoothed by reconstructing feature edge denoising convolution to obtain the second fused feature; the second fused feature is then processed by a feedforward neural network to obtain a refined detection box.

[0011] The refined detection boxes are filtered to obtain the second refined detection boxes. The distance of the second refined detection boxes is calculated and the validity is judged to obtain the valid detection boxes. The valid detection boxes are grouped and density is calculated, and the valid bounding boxes are selected according to the density. The target is marked and the detection boxes are deleted according to the valid bounding boxes. The trajectory confidence is updated and the valid trajectory is marked.

[0012] A bipartite graph structure is adopted. The graph is constructed based on the global fusion feature, the second fusion feature, the detection box at time t and the trajectory at time t-1 to obtain the final features of the detection box and the final features of the trajectory. The appearance association score between the final features of the detection box and the final features of the trajectory is calculated by adaptive channel graph convolution.

[0013] The detection box and trajectory prediction box are calculated based on the generalized intersection joint method to obtain the geometric cost matrix. The appearance association score and geometric cost matrix are encapsulated to obtain the final association matrix.

[0014] The final correlation matrix is ​​matched using the Hungarian algorithm to obtain successfully matched detection boxes and trajectory prediction boxes, as well as unmatched detection boxes and trajectory prediction boxes.

[0015] Based on historical trajectory data and current detection box information, the center point position, size and heading angle of the target bounding box are processed by position motion model, size motion model and heading angle motion model respectively to predict the target position at the next moment of the trajectory and to track the target trajectory throughout the entire process.

[0016] The method described in the embodiments of this application may also have the following additional technical features:

[0017] Furthermore, the original point cloud and the original image are preprocessed to obtain global fusion features, including:

[0018] A dual-branch structure is used to extract the original point cloud and the original image separately, so as to obtain the three-dimensional spatial information of the original point cloud and the appearance texture information of the original image. The three-dimensional spatial information and appearance texture information are then fused to obtain a global fused feature that includes the three-dimensional spatial information and appearance texture information.

[0019] Furthermore, the refined detection boxes are filtered to obtain second refined detection boxes. Distance calculations and validity checks are then performed on these second refined detection boxes to obtain valid detection boxes, including:

[0020] Set a first preset threshold, filter out refined detection boxes with confidence scores less than the first preset threshold, obtain a second refined detection box, calculate the Euclidean distance between the second refined detection box and each trajectory prediction box, obtain the shortest Euclidean distance, and use the shortest Euclidean distance as the distance between the second refined detection box and the trajectory.

[0021] A second preset threshold is set. If the shortest Euclidean distance is less than the second preset threshold, the second refined detection box is considered a valid detection box. If the shortest Euclidean distance is not less than the second preset threshold, the second refined detection box is considered an invalid detection box.

[0022] Furthermore, the valid detection boxes are grouped and their density is calculated, and valid bounding boxes are selected based on the density, including:

[0023] Legitimate detection boxes are grouped according to category. Detection boxes of the same type of target are grouped together to obtain a candidate set of detection boxes for each category. For each group of candidate detection boxes, the geometric association cost between all detection boxes in the candidate set is calculated, and the maximum intersection of a detection box with other detection boxes is taken as the density of the detection box.

[0024] Set a third preset threshold, and consider detection boxes with a density less than the third preset threshold as valid bounding boxes; sort the remaining detection boxes in each group by confidence, and consider the detection box with the highest confidence as a valid bounding box, and use this detection box as the reference box for this round.

[0025] Furthermore, based on the valid bounding boxes, target marking and detection box deletion are performed, and trajectory confidence is updated and valid trajectory marking is performed, including:

[0026] Calculate the intersection between any two remaining detection boxes in each group, and consider any detection box whose intersection with any two detection boxes is greater than a third preset threshold as the same target as the reference box. Delete any detection boxes whose intersection with the reference box is greater than the third preset threshold until the candidate set of detection boxes in each group is empty.

[0027] A fourth preset threshold is set, and an initial confidence score is set for each trajectory. The initial confidence score is updated and calculated based on the confidence item of the detection box associated with the trajectory, the trajectory missing duration item, and the trajectory confidence time validity item to obtain the trajectory confidence score. Trajectories with a confidence score higher than the fourth preset threshold are marked as valid trajectories.

[0028] Furthermore, a bipartite graph structure is adopted. A graph is constructed based on the global fusion features, the second fusion features, the detection boxes at time t, and the trajectories at time t-1, to obtain the final features of the detection boxes and the final features of the trajectories, including:

[0029] Let the bipartite graph be ,in, This is a collection of bounding box nodes, with a size of [size missing]. ; The set of trajectory nodes, with a size of ; Let be the set of edges, representing the correlation score between the detection box and the trajectory;

[0030] Among them, the detection box node The feature consists of a first appearance feature and a first motion feature, the first appearance feature For detection box nodes The corresponding local fusion features, the first motion feature Obtained through two layers of MLP: ,in express The first moment One detection box;

[0031] trajectory nodes The features consist of a second appearance feature and a second motion feature, the second appearance feature ,in Indicates according to Time trajectory prediction Time trajectory prediction box This indicates context-aware dynamic enhancement of local features. Represents global fusion features; motion features Extracted using LSTM: ,in express The first moment A set of trajectories;

[0032] The first appearance feature and the first motion feature are concatenated and enhanced with multi-head attention (MHSA) and LSTM to obtain the final feature of the detection box; the second appearance feature and the second motion feature are concatenated and enhanced with multi-head attention (MHSA) and LSTM to obtain the final feature of the trajectory.

[0033] Furthermore, the appearance association score between the final features of the detection box and the final features of the trajectory is calculated through adaptive channel graph convolution, including:

[0034] By overlaying multiple ACGC node aggregation modules and performing multiple node information transfers and updates, the appearance correlation score between the detection box and the trajectory is calculated. For the first... The node aggregation module of layer ACGC takes the first node as input. Layer node features and learnable adjacency matrix The attention score for each edge is calculated using a self-attention mechanism:

[0035]

[0036] in Indicates splicing operation and Representing nodes respectively and The weights;

[0037] And add a learnable adjacency matrix ,

[0038] in, Represents a node and The correlation score between nodes; using attention scores for message passing and node updates:

[0039]

[0040] in Represents a node The set of neighboring nodes; This represents the feature set of node u at layer k-1 and other nodes;

[0041] pass and The function calculates the learnable adjacency matrix of this layer:

[0042]

[0043] in, The input for the next layer of node aggregation modules; the output of the final ACGC node aggregation module is the final set of appearance association scores between the bounding box and the trajectory:

[0044]

[0045] in Represents a node and The appearance correlation score between them.

[0046] Furthermore, the geometric cost matrix is ​​calculated based on the generalized intersection joint method for the detection box and the trajectory prediction box. The appearance association score and the geometric cost matrix are then encapsulated to obtain the final association matrix, which includes:

[0047] The geometric cost matrix is ​​obtained by calculating the detection box and trajectory prediction box through the three-dimensional generalized joint intersection or the bird's-eye view generalized joint intersection;

[0048] The appearance association score and geometric cost matrix are encapsulated to obtain the final association matrix, as shown in the formula below:

[0049]

[0050] in, and These represent the weights of the appearance association score and the set similarity score, respectively.

[0051] Furthermore, historical trajectory data includes Trajectory data at any given time, The target bounding box includes successfully matched detection boxes and trajectory prediction boxes, as well as unmatched detection boxes; the method also includes:

[0052] In the positional motion model, the target bounding box is defined with positioning jitter noise and added to the basic Kalman filter motion model to calibrate the errors caused by occlusion and target disappearance.

[0053] The 3D occlusion target tracking method based on multimodal spatiotemporal interaction provided in this application has the following advantages compared with the prior art:

[0054] This application embodiment extracts data from the original point cloud and the original image using a dual-branch structure, and fuses 3D spatial information and appearance texture information to obtain global fusion features. This fusion enables the target tracking system to simultaneously utilize the spatial information of the point cloud and the appearance information of the image, improving the accuracy and robustness of target recognition. The non-empty voxel point cloud features are projected onto the global image features to obtain the visible image features of the occluded target, and then reconstructed and decoded to obtain the complete high-dimensional features of the occluded target. This method effectively handles the occlusion problem and improves target tracking capabilities in complex scenes. Edge smoothing processing of the complete high-dimensional features is performed through edge denoising convolution of the reconstructed features to obtain the second fusion feature, further improving the accuracy and stability of the features.

[0055] This application embodiment reduces false detections by setting a confidence threshold to filter out refined detection boxes with low confidence. The Euclidean distance between the refined detection box and the trajectory prediction box is calculated, and the validity of the detection box is determined based on a preset threshold, ensuring the accuracy of the detection boxes.

[0056] This application's embodiments group legitimate detection boxes according to categories and calculate the geometric association cost and density between detection boxes to filter out legitimate bounding boxes, reducing redundant detection. The trajectory confidence is updated based on the confidence of the detection boxes associated with the trajectory, the trajectory missing duration, and the time validity of the trajectory confidence, ensuring the accuracy and reliability of the trajectory. A bipartite graph structure and adaptive channel graph convolution are used: a bipartite graph structure is used to construct the association graph between detection boxes and trajectories, and an appearance association score between detection boxes and trajectories is calculated through adaptive channel graph convolution, improving the accuracy of the association. The geometric cost matrix between detection boxes and trajectory prediction boxes is calculated based on the generalized intersection joint method, and combined with the appearance association score to obtain the final association matrix, providing a more comprehensive basis for node matching.

[0057] This application's embodiments use the Hungarian algorithm to perform node matching on the final correlation matrix, obtaining successfully matched detection boxes and trajectory prediction boxes, thus improving matching efficiency. Based on historical trajectory data and current detection box information, the center point position, size, and heading angle of the target bounding box are processed using positional motion models, size motion models, and heading angle motion models to predict the target position at the next moment of the trajectory, achieving full-process tracking of the target trajectory.

[0058] In this embodiment, positioning jitter noise is defined in the position motion model and added to the basic Kalman filter motion model to calibrate the errors caused by occlusion and target disappearance, thereby improving the accuracy of trajectory prediction. Attached Figure Description

[0059] Figure 1 A flowchart illustrating a three-dimensional occlusion target tracking method based on multimodal spatiotemporal interaction according to an embodiment of this application is shown. Detailed Implementation

[0060] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, it should be noted that, for ease of description, only the parts relevant to this application are shown in the accompanying drawings, not the entire structure. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.

[0061] The terms “comprising” and “having”, and any variations thereof, used in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.

[0062] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.

[0063] like Figure 1 As shown in the figure, this application provides a three-dimensional occlusion target tracking method based on multimodal spatiotemporal interaction, including the following steps:

[0064] Step 101: Obtain the original point cloud and the original image at the same time, and preprocess the original point cloud and the original image to obtain the global fusion features.

[0065] Step 101 involves acquiring the original point cloud and the original image at the same time, and preprocessing these data to obtain global fusion features. The original point cloud is usually acquired by devices such as LiDAR, containing point data in three-dimensional space, with each point having positional information (such as x, y, z coordinates). The original image is acquired by visual sensors such as cameras, containing two-dimensional image data, with appearance information such as color and texture.

[0066] This application employs a dual-branch structure to process the original point cloud and the original image separately. For the original point cloud, its 3D spatial information is extracted; for the original image, its appearance texture information is extracted. The extracted 3D spatial information and appearance texture information are fused to obtain a global fusion feature. This feature integrates information from different modalities, providing rich feature representations for subsequent target detection and tracking. Point cloud processing algorithms (such as PointNet, PointNet++, etc.) are used to extract features from the original point cloud to obtain its 3D spatial features. Convolutional neural networks (CNNs) are used to extract features from the original image to obtain its appearance texture features.

[0067] The 3D spatial features of the point cloud and the appearance and texture features of the image are fused. This can be achieved through a simple stitching operation or a more complex fusion strategy (such as attention mechanisms, weighted fusion, etc.). The fused features are called global fused features, which contain multimodal information from the point cloud and the image, providing a comprehensive feature representation for subsequent steps.

[0068] Specifically, in the process of feature extraction from the original point cloud, the input point cloud P is composed of discrete points in space. ,in Indicating the first point cloud A point, defined by three-dimensional coordinates Composition. To efficiently process point cloud data, a three-dimensional space S is defined, whose depth D, height H, and width W determine its size. Within space S, fixed-size three-dimensional regions are defined, called voxels, with dimensions d × h × w. Using these standard voxels, the point cloud P is divided into a voxelized mesh. The number of voxel meshes in the three coordinate directions (depth, height, and width) is calculated as follows:

[0069]

[0070] These quantities collectively constitute the dimensions of the voxelized mesh. Each voxel... It contains multiple points, and its quantity is expressed as The set of points within a voxel can be represented as For each voxel, a multilayer perceptron (MLP) is used to encode features of the points within the voxel. The MLP consists of a linear layer, a batch normalization (BN) layer, and a ReLU activation function. After feature encoding for each point, max pooling is used to aggregate the features of all points within the voxel into a unified voxel feature representation. To enhance the feature representation capability, the above accelerated feature extraction operation is repeated twice to form a deep network structure. The input of the second feature extraction layer is the output of the first feature extraction layer, and the output is the enhanced feature representation, resulting in the voxel feature embedding. The feature embeddings of all voxels together constitute the feature set. ,in It is the total number of voxels.

[0071] To enhance the spatial alignment between the point cloud and the image, a position embedding is added for each voxel. Considering the differences in the number and spatial distribution of points within a voxel, the centroid of the points in the voxel is chosen as the reference for the position embedding. centroid of individual elements The value is calculated as the average of the coordinates of all points within the voxel. The location embedding is performed by passing the centroid coordinates through a linear layer, a GeLU activation function, and another linear layer. Convert to vector representation The positional embeddings of all voxels constitute a set. .

[0072] To distinguish between point clouds and image modalities, a modal embedding is added to each point cloud voxel. The location embedding is added to the feature embedding, and the modal embedding is added through a concat operation to obtain the point cloud token, as shown in the formula below:

[0073]

[0074] in, Indicates positional embedding, Represents feature embedding, This represents point cloud modal embedding.

[0075] The tokens of all point cloud voxels are a set. The formula is shown below:

[0076]

[0077] Set of point cloud tokens The data is fed into the Point Cloud Encoder (PCE) to obtain the high-level latent feature set of the point cloud, as shown in the formula below:

[0078]

[0079] Structurally, PCE consists of multiple layers of Vision Transformer (ViT), which progressively learn richer contextual information and higher-level feature representations.

[0080] Specifically, in the process of feature extraction from the original image, for the image at the current time step of the input branch, the original image is first divided into several non-overlapping patches. Each patch represents a local region of the image, thus decomposing the image into a set of more easily processed units. Next, a convolutional neural network (CNN) is used to extract features from these image patches. Through convolution operations, CNNs can capture local features in the image, such as edges and textures. After processing by the convolutional layers, a feature map is obtained for each patch.

[0081] To convert these feature maps into vector form suitable for subsequent processing, the Flatten operation is used to unfold the feature maps into one-dimensional vectors. Then, to eliminate dimensional differences between different features, these vectors are normalized, mapping them to a D-dimensional feature space. The vectors obtained in this process are called patch feature embeddings, and the formula is shown below:

[0082]

[0083] in, This represents the mapping function from image patches to feature embeddings. Represents a single image block.

[0084] As well-organized structured data, the spatial location information of images is crucial for understanding their content. Therefore, a position embedding is assigned to each image patch, representing the patch's relative position within the image. The image center is chosen as a reference point, and the position information of each patch relative to the center is calculated and encoded as a position embedding. Furthermore, to distinguish between different modalities of data (such as RGB images, grayscale images, etc.), an image modality embedding is introduced. This embedding represents the image's modal information, helping the model maintain robustness when processing multimodal data. Adding the position embedding and modality embedding to the feature embedding yields the image token, as shown in the following formula:

[0085]

[0086] in, This represents the set of position embeddings for all image patches. Represents image modality embedding. This represents the set of image patch feature embeddings.

[0087] Finally, the obtained image tokens are fed into the Image Encoder (IE) to extract high-level latent features. The Image Encoder consists of multiple layers of Vision Transformer (ViT), a model based on a self-attention mechanism that excels at processing sequential data (here, image tokens are treated as sequence elements). Through the processing of multiple ViT layers, the Image Encoder can capture global features in the image and the correlation information between different patches, thereby generating high-level latent features, as shown in the following formula:

[0088]

[0089] in, Indicates an image encoder. This represents the set of input image tokens. This represents the high-level latent features of the output.

[0090] Furthermore, the high-level latent features of point clouds and high-level latent features of images The features are concatenated to form a joint latent feature vector, which is then fed into a point cloud image fusion encoder to learn the correlations and complementarities between features of different modalities. This encoder consists of multiple layers of Spatial Interaction Fusion (ViT), each designed to enhance the interaction of features across different dimensions and extract richer and more advanced multimodal feature representations. The ViT layers are connected using subsampling layers with a stride of 2, which helps reduce the size of the feature maps while preserving key information for more refined feature extraction in subsequent layers. The ViT layer itself consists of 2N feature interaction layers and cascaded segmentation-enhanced attention layers.

[0091] The feature interaction layer first uses depthwise convolutions for additional feature interactions, then maintains the feature size dimension invariant through a feedforward neural network layer. The depthwise convolutional layer consists of depthwise convolutions and batch normalization (BatchNorm) for initial feature extraction. The FFN layer consists of two depthwise convolutions, normalization, and a ReLU activation function for further feature extraction and more efficient communication between different feature channels. To mitigate the vanishing gradient problem during training, residual connections are used after both the depthwise convolutional and feedforward neural network layers.

[0092] The cascaded segmentation-enhanced attention layer is a special type of multi-head attention layer that segments the input features into h non-overlapping smaller features, with each attention head responsible for processing one of these smaller features. It learns projections onto features with richer information through Q, K, and V layers and applies a self-attention mechanism to jointly capture local and global relationships, further enhancing the feature representation. The heads interact in a cascaded manner to reduce redundant computation and progressively refine the feature representation. Finally, the outputs of the h attention heads are concatenated and projected back to the same dimension as the input. For the m-th layer of spatial interaction fusion (ViT), the computation process is as follows: First, the input features pass through a feature interaction layer to obtain an intermediate feature representation. Then, the intermediate feature representation is input into the cascaded segmentation-enhanced attention layer for multi-head attention computation, resulting in an enhanced feature representation. Finally, the enhanced feature representation passes through another feature interaction layer to obtain the output of the m-th layer of spatial interaction fusion (ViT). The formula is shown below:

[0093]

[0094] in, Represents the feature interaction layer. This indicates a cascaded segmentation enhanced attention layer. Indicates the first Layered spatial interaction integrates ViT's input. Indicates the first Layered spatial interaction and fusion of ViT output.

[0095] By stacking multiple spatial interactive fusion ViT layers, the point cloud image fusion encoder PCIFE can progressively extract richer and more advanced multimodal feature representations, obtaining global fusion features of point clouds and images. .

[0096] Step 102: The global fusion features are processed by the region proposal network to obtain the initial detection box. The initial detection box and the global fusion features are processed by region of interest pooling to obtain the global fusion features of the region of interest. The global fusion features of the region of interest are separated to obtain the global point cloud features and the global image features.

[0097] Step 102 mainly involves processing the global fusion features through a Region Proposal Network (RPN) to generate initial detection boxes, and further processing the initial detection boxes and global fusion features through Region of Interest Pooling (RoI Pooling) to obtain global fusion features of regions of interest. Subsequently, these global fusion features of regions of interest are separated to obtain global point cloud features and global image features, respectively.

[0098] RPN is used to generate candidate regions (i.e., initial detection boxes) that may contain targets on the global fused feature map. RPN traverses the global fused feature map using a sliding window, applies a set of anchor boxes to each location, and predicts whether these anchor boxes contain targets and their position adjustment parameters. RPN outputs a series of initial detection boxes, each containing location information (such as the coordinates of the bounding box) and a confidence score (indicating the probability that the box contains a target).

[0099] RoI Pooling extracts the feature regions corresponding to the initial bounding boxes from the global fused feature map; these are the regions of interest (ROIs) global fused features. For each initial bounding box, RoI Pooling maps it back onto the global fused feature map and performs pooling operations (such as max pooling or average pooling) on ​​that region to obtain a fixed-size feature map. RoIPooling outputs a series of ROI global fused features, each corresponding to an initial bounding box.

[0100] The global fusion features of the region of interest (ROI) are separated into global point cloud features and global image features for subsequent processing of point cloud and image information respectively. Point cloud-related information is extracted from the ROI global fusion features. This may involve identifying which features are associated with 3D spatial location, or extracting these features through specific network layers (such as point cloud feature extraction layers). Similarly, image-related information, such as appearance features like color and texture, is extracted from the ROI global fusion features. The separated global point cloud features and global image features will be used in subsequent point cloud processing and image processing steps, respectively.

[0101] Specifically, for global fusion features First, the initial detection box is obtained using RPN. Global fusion features of point cloud images of regions of interest are obtained using ROI Pooling. ,from The global point cloud features and global image features are separated from each other, as shown in the following formula:

[0102]

[0103] in, This represents the point cloud features of the region of interest after global fusion. This represents the image features of the region of interest after global fusion.

[0104] The reconstruction decoder module (RD) has a similar network structure to the encoder, employing a multi-layer MSF-ViT architecture to enhance the interaction of features across different dimensions and obtain semantic information at different scales. Within the region of interest (ROI) of an occluded target detection box, point cloud voxels contain numerous blank areas, while the ROI image features contain rich semantic information. However, these image features include not only the features of the occluded target but also the semantic information of other objects. To separate more refined ROI features, [the following text is missing from the original]... Voxels in the point cloud are projected onto the image to remove background interference. When projecting point cloud voxels onto the image, only non-empty voxels are projected. Image blocks corresponding to blank areas are identified as regions of non-interest, thereby reducing interference during the reconstruction process.

[0105] Step 103: The global point cloud features include non-empty voxel point cloud features. The non-empty voxel point cloud features are projected onto the global image features to obtain the visible image features of the occluded target. The visible image features of the occluded target are reconstructed and decoded to obtain the complete high-dimensional features of the occluded target.

[0106] Step 103 primarily processes the non-empty voxel point cloud features in the global point cloud features, and obtains the complete high-dimensional features of the occluded target through projection and reconstruction decoding. In step 101, the global fusion features were already obtained, which contain 3D spatial information extracted from the original point cloud. In this step, the global point cloud features are further separated from the global fusion features, with a particular focus on the non-empty voxel point cloud features in the point cloud data. Non-empty voxels refer to voxels that actually contain points in 3D space; these voxels carry the actual spatial location information of the target object.

[0107] The extracted non-empty voxel point cloud features are projected onto global image features. The purpose of this step is to map the point cloud information in three-dimensional space onto a two-dimensional image plane, thereby obtaining the visible portion of the occluded target in the image. The projection process may involve coordinate transformations, interpolation, and other operations to ensure that the point cloud features are accurately mapped onto the image.

[0108] By performing projection operations, the visible image features of the occluded target are obtained. These features reflect the unoccluded parts of the target object in the image and form the basis for subsequent reconstruction of the complete features of the occluded target.

[0109] The visible image features of the acquired occluded target are reconstructed and decoded. This reconstruction and decoding process may utilize deep learning models (such as autoencoders, generative adversarial networks, etc.) to learn the mapping relationship from the visible portion to the complete target. Through this process, complete high-dimensional features of the occluded target can be generated. These features not only contain information from the visible portion but also infer possible features of the occluded portion through the model's learning ability.

[0110] Complete high-dimensional features provide richer and more accurate information for subsequent target tracking. In occlusion situations, traditional target tracking methods may fail due to insufficient information. However, reconstructing the decoded complete high-dimensional features can, to some extent, compensate for the information loss caused by occlusion, improving the accuracy and robustness of target tracking.

[0111] Specifically, the non-empty voxel point cloud features of the region of interest in the i-th detection box are... Projecting onto the image, we obtain the visible image features of the occluded target. Next, the visible features of the occluded target are fed into the reconstruction decoder to predict the complete high-dimensional features of the occluded target. The formula is shown below:

[0112]

[0113] A complete high-dimensional feature consists of two parts, as shown in the following formula:

[0114]

[0115] One part of it is the completed point cloud features. One part is the image features after completion. .

[0116] Step 104: The complete high-dimensional features of the occluded target are smoothed by reconstructing feature edge denoising convolution to obtain the second fused feature; the second fused feature is processed by a feedforward neural network to obtain a refined detection box.

[0117] Step 104 involves performing edge smoothing on the complete high-dimensional features of the occluded target and obtaining a refined detection box through a feedforward neural network.

[0118] The complete high-dimensional features of the occluded target are processed using a reconstructed feature edge-denoising convolution. The purpose of this step is to smooth the edges of the features, reduce noise interference, and improve the stability and reliability of the features. The feature after edge smoothing is the second fused feature. Compared to the original complete high-dimensional feature, this feature is smoother at the edges, which is helpful for subsequent target detection and tracking.

[0119] The second fused feature is input into a feedforward neural network (FNN). A feedforward neural network is a type of multilayer perceptron (MLP) that performs nonlinear transformations and feature extraction on the input features. After processing the second fused feature, the feedforward neural network outputs refined detection boxes. These refined detection boxes are more accurate in position and size compared to the initial detection boxes, helping to improve the accuracy of target tracking.

[0120] Specifically, to improve the feature completion capability of the reconstruction decoder (RD), MBSS-GFM and RD are combined into a pre-trained reconstruction network. The pre-training process involves first acquiring point clouds using PCT and IT respectively. and images Both visible and invisible regions are uniformly masked, with the invisible regions remaining consistent to simulate a scene where the target is occluded. The remaining visible regions are then complementaryly masked. The visible point cloud and image tokens are fed into the encoder and decoder structures to train the encoder's feature representation capabilities and the decoder's learning of general features for similar objects. The reconstruction operation after the decoder output is the same as the MBSS-GFM pre-training process and will not be explained again.

[0121] The complete high-dimensional features are fed into the reconstructed feature edge denoising convolution NRConv to perform edge smoothing on the complete high-dimensional features, preventing inaccurate point cloud prediction at the target boundary from negatively impacting the accuracy of the detection box.

[0122] Specifically, NRConv employs a dual-branch structure to process the reconstructed point cloud features and image features. The point cloud branch uses a 3D geometric feature extraction layer to extract the 3D geometric semantic features of the point cloud space, as shown in the following formula:

[0123]

[0124] in, Represents a non-linear activation function. Represents sparse 3D convolution of submanifolds. This represents the 3D geometric semantic features of the point cloud of the region of interest of the i-th detection box.

[0125] The image branch uses a two-dimensional geometric feature extraction layer to extract two-dimensional geometric semantic features in the image space, as shown in the following formula:

[0126]

[0127] in, Represents two-dimensional convolution of a submanifold. This represents the two-dimensional geometric semantic features of the region of interest image of the i-th detection box.

[0128] Furthermore, the 3D geometric semantic features of the point cloud of the region of interest (ROI) of all detection boxes are concatenated with the 2D geometric semantic features of the ROI image, and the target edge noise is implicitly learned to obtain the second fused feature Y after the reconstruction feature edge smoothing process, as shown in the following formula:

[0129]

[0130] in, , This indicates the initial number of detection boxes. Finally, a refined set of detection boxes is obtained using a feedforward neural network (FFN). .

[0131] Step 105: Filter the refined detection boxes to obtain the second refined detection boxes; perform distance calculation and validity judgment on the second refined detection boxes to obtain valid detection boxes; group and calculate the density of the valid detection boxes, and filter out valid bounding boxes based on the density; mark the target and delete the detection boxes based on the valid bounding boxes, and update the trajectory confidence and mark the valid trajectory.

[0132] Step 105 involves further processing of the refined detection boxes, including filtering, distance calculation and validity judgment, grouping and density calculation, target labeling and detection box deletion, and trajectory confidence update and valid trajectory labeling.

[0133] First, a confidence threshold (first preset threshold) is set to filter out refined detection boxes with low confidence. Refined detection boxes with confidence below this threshold are removed, resulting in the remaining second-refined detection boxes. The Euclidean distance between each second-refined detection box and each trajectory prediction box is calculated to obtain the shortest Euclidean distance. A distance threshold (second preset threshold) is set; if the shortest Euclidean distance is less than this threshold, the second-refined detection box is considered a valid detection box; otherwise, it is considered an invalid detection box. Valid detection boxes are grouped according to the target category, grouping detection boxes of the same type of target into one group. For each group of detection boxes, the maximum intersection of each detection box with other detection boxes is calculated as the density of that detection box. A density threshold (third preset threshold) is set; detection boxes with a density less than this threshold are considered valid bounding boxes. Simultaneously, the remaining detection boxes within each group are sorted by confidence, and the detection box with the highest confidence is also considered a valid bounding box and used as the reference box for this round.

[0134] Calculate the intersection of the remaining detection boxes within each group. Detection boxes with an intersection greater than a third preset threshold are considered to be the same target as the reference box and are deleted. Repeat this process until the candidate set of detection boxes in each group is empty. Set an initial confidence score for each trajectory. Update the initial confidence score based on the confidence items of the detection boxes associated with the trajectory, the trajectory missing duration, and the trajectory confidence time validity item to obtain the trajectory confidence score. Set a trajectory confidence threshold (fourth preset threshold). Trajectories with a confidence score higher than this threshold are marked as valid trajectories.

[0135] Specifically, for the refined detection box obtained at the current time t Set the first preset threshold. If the filter confidence level is less than the first preset threshold The obvious false positive detection box was used to obtain the second refined detection box.

[0136] For the second refined detection box obtained after SF filtering, calculate its Euclidean distance with each trajectory prediction box. The Euclidean distance is used to measure the spatial proximity between the second refined detection box and the existing trajectory, and the shortest Euclidean distance is taken as the distance between the second refined detection box and the trajectory. If the Euclidean distance of the second refined detection box is less than a second preset threshold, then the second refined detection box is considered a valid detection box. Otherwise, it is considered an illegal detection frame that is far from the road. The formula is shown below:

[0137] in, and These represent the center point coordinates of the second refined detection box and the existing trajectory detection box, respectively.

[0138] Next, non-maximum suppression based on AGDIoU is used to remove overlapping detection boxes, and the valid detection boxes are sorted according to their categories. Grouping is performed to categorize detection boxes of the same type into groups, resulting in a candidate set of detection boxes for each category. For each candidate set of detection boxes, the geometric association cost between all detection boxes in the set is calculated.

[0139] The maximum intersection of a detection box with other detection boxes is used as the detection box density. Detection boxes with a density less than a third preset threshold are directly marked as valid and added to the valid bounding box set VD. The remaining detection boxes in each group are sorted, and the detection box with the highest confidence is marked as valid, added to the valid bounding box set, and used as the reference box for this round. Detection boxes with an intersection greater than the third preset threshold are considered to be the same target as the reference box and are deleted. This step is repeated from the remaining detection boxes until the candidate detection boxes are empty.

[0140] The second level of trajectory validity assessment (TVJ) assigns a confidence score to each trajectory, representing the probability that the trajectory is valid. A higher confidence score indicates a greater likelihood of the trajectory being valid; conversely, a lower score indicates a greater likelihood of the trajectory being a ghost trajectory. A confidence score is initialized for each trajectory. This is typically set to a small value to represent the uncertainty in the initial state. The legal confidence of the trajectory is dynamically adjusted based on factors such as the success rate of matching the trajectory with the detection box, the confidence of the matched detection box, and the trajectory interruption time. Trajectories with high confidence that are occluded are retained, while ghost trajectories are identified and deleted. This effectively balances the retention of low-confidence detection boxes of occluded targets with the avoidance of ghost trajectories, ensuring stable tracking even when the target is occluded.

[0141] Once the trajectory successfully matches the detection result in the new data frame, the trajectory status is updated, the confidence score of the trajectory is recalculated based on the association result, and the confidence score value of the trajectory is updated.

[0142] trajectory At any moment and time Each object was successfully associated with a specific bounding box, and the trajectory was updated at the specified time. At any moment Previously, at the moment and time Between, trajectory If it fails to successfully associate with any detection box at any time, then at that time... The formula for trajectory updates is as follows:

[0143]

[0144]

[0145] in, Representing the trajectory At any moment The confidence score, Representing the trajectory At any moment The confidence score, Indicates at time With trajectory The confidence level of the matched detection box. Indicates at time With trajectory The confidence level of the matched detection box. This indicates the time difference from the last trajectory update. This represents the duration of the missing trajectory ti after it disappears from time k.

[0146] The confidence score calculation process consists of three items: the confidence score of the detection boxes associated with the trajectory, the duration of trajectory missing data, and the time validity of the trajectory confidence score. Among them, the confidence score of the detection boxes associated with the trajectory and the time validity of the trajectory confidence score are positively correlated with the confidence score, while the duration of trajectory missing data is negatively correlated with the confidence score.

[0147] Confidence item of detection box associated with trajectory In the middle, the confidence of the detection box associated with the trajectory The larger the value, the higher the trajectory confidence. If the confidence of a detection box successfully associated with a ghost trajectory is high, the ghost trajectory is more likely to be mistakenly identified as a legitimate trajectory. Considering that the effectiveness of a trajectory gradually decreases as the duration of trajectory loss increases, a decay rate is introduced into the detection box confidence term for trajectory association. The decay rate increases with the duration of trajectory loss. The value decreases as the confidence level of the detection box increases, in order to neutralize the false positive impact caused by the increased confidence level of the detection box.

[0148] Trajectory missing duration item In the middle, the duration of missing trajectory The longer the interval between the successful matching of the trajectory and the detection results in the data frame, the more unstable the trajectory is, the greater the possibility that the trajectory is a ghost trajectory, and the greater the negative impact of the corresponding trajectory missing duration on the trajectory confidence. The smaller the value, the more frequently and often the trajectory matches the detection results in the data frame, the more frequently the trajectory is updated, the more accurate the trajectory prediction, the more stable the tracking result of this trajectory, and the higher the confidence score. If a trajectory is a valid trajectory at time m, but the target object corresponding to the trajectory is occluded between times (m+1) and (m+5), and the trajectory is not associated with the detection box during this period, the trajectory is missing for a relatively long time, and this trajectory will be incorrectly identified as a ghost trajectory. Considering that a valid trajectory after occlusion reconstruction has a high confidence score in the detection box after the target is reconstructed, an inverse proportional term to the detection box confidence score is introduced into the trajectory missing time term. This is used to neutralize the misjudgment caused by the missing duration of occlusion.

[0149] The trajectory confidence time validity term incorporates the trajectory confidence score from the last trajectory update into the current calculation, retaining the influence of historical trajectory confidence. Higher historical trajectory confidence indicates greater trajectory stability and... The higher the legitimacy.

[0150] When the confidence score of the trajectory Higher than the set fourth preset threshold Mark this trajectory as a valid trajectory, and add the prediction box to the valid trajectory. Proceed to the next step.

[0151] Step 106: Using a bipartite graph structure, a graph is constructed based on the global fusion feature, the second fusion feature, the detection box at time t, and the trajectory at time t-1 to obtain the final features of the detection box and the final features of the trajectory. The appearance association score between the final features of the detection box and the final features of the trajectory is calculated by adaptive channel graph convolution.

[0152] Step 106 involves using a bipartite graph structure, combining global fusion features, second fusion features, detection boxes at time t, and trajectories at time t-1 to construct a graph, and calculating the appearance association score between detection boxes and trajectories through adaptive channel graph convolution (ACGC).

[0153] Specifically, in this embodiment, the detection boxes and trajectories are represented as a graph structure to effectively capture potential dependencies between objects. To avoid high-dimensional computation and message redundancy, a bipartite graph structure is adopted. Let the bipartite graph be... ,in, This is a collection of bounding box nodes, with a size of [size missing]. ; The set of trajectory nodes, with a size of ; Let be the set of edges, representing the correlation score between the detection box and the trajectory;

[0154] Among them, the detection box node The feature consists of a first appearance feature and a first motion feature, the first appearance feature For detection box nodes The corresponding local fusion features, the first motion feature Obtained through two layers of MLP: ,in express The first moment One detection box.

[0155] trajectory nodes The features consist of a second appearance feature and a second motion feature, the second appearance feature ,in Indicates according to Time trajectory prediction Time trajectory prediction box This indicates context-aware dynamic enhancement of local features. Represents global fusion features; motion features Extracted using LSTM: ,in express The first moment A set of trajectories, It represents long and short-term memory.

[0156] The first appearance feature and the first motion feature are concatenated and enhanced with multi-head attention (MHSA) and LSTM to obtain the final feature of the detection box; the second appearance feature and the second motion feature are concatenated and enhanced with multi-head attention (MHSA) and LSTM to obtain the final feature of the trajectory.

[0157] After the graph construction is completed, multiple ACGC modules are overlaid. Through multiple node information transfers and updates, the appearance association score between the detection box and the trajectory is calculated. For the first... The node aggregation module of layer ACGC takes the first node as input. Layer node features and learnable adjacency matrix The attention score for each edge is calculated using a self-attention mechanism:

[0158] in Indicates splicing operation and Representing nodes respectively and The weights; This represents the attention mechanism module.

[0159] And add a learnable adjacency matrix ,

[0160] in, Represents a node and The correlation score between nodes; using attention scores for message passing and node updates:

[0161]

[0162] in Represents a node The set of neighboring nodes; Represents the feature set of node u at layer k-1 and other nodes; through and The function calculates the learnable adjacency matrix of this layer:

[0163]

[0164] in, Used as input for the next-level node aggregation module; This represents the weight of node m in the k-th layer. By stacking multiple ACGC layers, the node features of the same target become more similar, and the node features of different targets become more distinguishable. The output of the final ACGC node aggregation module is the final set of appearance association scores between the bounding box and the trajectory.

[0165]

[0166] in Represents a node and The appearance correlation score between them.

[0167] Step 107: Calculate the detection box and trajectory prediction box based on the generalized intersection joint method to obtain the geometric cost matrix. Encapsulate the appearance association score and the geometric cost matrix to obtain the final association matrix.

[0168] Step 107 involves calculating the detection box and trajectory prediction box based on the generalized intersection joint method to obtain the geometric cost matrix, and encapsulating the appearance association score and the geometric cost matrix to obtain the final association matrix.

[0169] Using three-dimensional generalized joint intersection Or the generalized joint intersection of bird's-eye view Calculate the area of ​​intersection and the area of ​​union for each pair of detection boxes and predicted trajectories.

[0170] According to the three-dimensional generalized joint intersection Or the generalized joint intersection of bird's-eye view The geometric cost is defined and calculated. The geometric costs of all bounding boxes and trajectory prediction boxes are organized into a matrix, resulting in the geometric cost matrix.

[0171] The appearance association score and geometric cost matrix are combined with specific weights. Typically, the appearance association score and geometric cost matrix are assigned different weights to reflect their importance in the final association determination. The resulting matrix, called the final association matrix, considers both the appearance and geometric similarity between the bounding box and the trajectory.

[0172] Specifically, the three-dimensional generalized joint intersection The generalized joint intersection of two bounding boxes in 3D space is given by the following formula:

[0173]

[0174] in, The ratio of geometric distance to total distance in three-dimensional space is expressed by the following formula:

[0175]

[0176] in, Indicates the geometric distance between bounding boxes. This represents the spatial distance between bounding boxes; the total distance is the sum of the two. This represents the ridge regression function, used for regularization during model optimization. It represents the ratio of spatial distance to total distance in three-dimensional space.

[0177] Geometric distance The formula is shown below:

[0178]

[0179] in, The intersection-union ratio (IU / U) of a 3D bounding box is expressed by the following formula:

[0180]

[0181] in, Represents 3D bounding box and The volume of the intersection portion; Represents 3D bounding box and The volume of the union of the subsets; Represents 3D bounding box and The minimum circumscribed cuboid volume.

[0182] spatial distance The formula is shown below:

[0183]

[0184] in, Represents bounding box and Difference in rotation angle; Represents 3D bounding box and The length of the body diagonal of the smallest circumscribed cuboid; Represents 3D bounding box and The distance from the center point.

[0185] Bird's-eye view generalized joint intersection The formula for the generalized joint intersection of two bounding boxes from a bird's-eye view is as follows:

[0186]

[0187] in, The ratio of geometric distance to total distance from a bird's-eye view is expressed by the following formula:

[0188]

[0189] in, Indicates the geometric distance between bounding boxes. This represents the spatial distance between bounding boxes; the total distance is the sum of the two. This represents the ridge regression function. This represents the ratio of spatial distance to total distance from a bird's-eye view perspective.

[0190] Geometric distance The formula is shown below:

[0191]

[0192] in, The intersection-union ratio (IU / U) of bounding boxes from a bird's-eye view is expressed by the following formula:

[0193]

[0194] in, Represents 3D bounding box and Area of ​​the intersection from a bird's-eye view; Represents 3D bounding box and Area of ​​the union portion from an aerial view; Represents 3D bounding box and The area of ​​the smallest circumscribed rectangle from a bird's-eye view.

[0195] spatial distance The formula is shown below:

[0196]

[0197] in, Represents bounding box and Difference in rotation angle; Represents 3D bounding box and The diagonal length of the smallest bounding rectangle from a bird's-eye view perspective; Represents 3D bounding box and Distance from the center point in a bird's-eye view.

[0198] Based on the target categories in the detection and prediction bounding boxes, a suitable boundary intersection method is selected to improve the correlation and distinction between targets and trajectories, and the geometric cost matrix is ​​calculated. .

[0199] Associating the bounding box with the appearance of the trajectory as a score and geometric similarity matrix The final association matrix formed by encapsulating them together The formula is shown below:

[0200]

[0201] in, This represents the geometric similarity matrix; since appearance association scores are defined as negative costs, subtraction is used to calculate feature association scores. and These represent the weights of the appearance association score and the set similarity score, respectively.

[0202] Step 108: Perform node matching on the final correlation matrix based on the Hungarian algorithm to obtain successfully matched detection boxes and trajectory prediction boxes, as well as unmatched detection boxes and unmatched trajectory prediction boxes.

[0203] Step 108 involves using the Hungarian algorithm to perform node matching on the final association matrix to determine successfully matched detection boxes and trajectory prediction boxes, as well as unmatched detection boxes and trajectory prediction boxes.

[0204] The final association matrix is ​​obtained by combining the appearance association score and the geometric cost matrix. The appearance association score reflects the visual similarity between the bounding box and the trajectory, while the geometric cost matrix reflects their spatial relationship. The encapsulation process typically involves weighted summation or other forms of combination of these scores to form a comprehensive association matrix.

[0205] The Hungarian algorithm is a classic algorithm for solving the maximum matching problem in bipartite graphs. In the context of object tracking, it can be used to find the optimal matching pair between a bounding box and a trajectory, maximizing the overall association score (or cost). The basic idea of ​​the algorithm is to continuously increase the number of matching pairs by finding augmenting paths until no more augmenting paths can be found.

[0206] The final association matrix is ​​viewed as an adjacency matrix of a bipartite graph, where the detection boxes and trajectory prediction boxes are the two parts of the bipartite graph, respectively. The Hungarian algorithm is used to perform maximum matching on this bipartite graph, finding successfully matched pairs of detection boxes and trajectory prediction boxes. The matching results will consist of two parts: one part contains successfully matched pairs of detection boxes and trajectory prediction boxes, and the other part contains unmatched pairs of detection boxes and trajectory prediction boxes.

[0207] For successfully matched detection boxes and trajectory prediction boxes, they can be considered to correspond to the same target and are further used for trajectory updating and prediction. Unmatched detection boxes may indicate the presence of a new target or that the target is temporarily occluded; unmatched trajectory prediction boxes may indicate that the target has left the scene or is temporarily undetectable. These unmatched results can serve as the basis for subsequent processing (such as target initialization, trajectory termination determination, etc.).

[0208] Step 109: Based on historical trajectory data and current detection box information, the center point position, size and heading angle of the target bounding box are processed by position motion model, size motion model and heading angle motion model respectively to predict the target position at the next moment of the trajectory and to track the target trajectory throughout the entire process.

[0209] Step 109 is based on historical trajectory data and current detection box information. It uses position motion model, size motion model and heading angle motion model to process the center point position, size and heading angle of the target bounding box in order to predict the target position at the next moment of the trajectory and to track the target trajectory throughout the entire process.

[0210] Historical trajectory data contains information such as the target's position, size, and heading angle at previous moments, which is used to train and calibrate the motion model. Current bounding box information represents the bounding box information obtained by the target detection algorithm at the current moment, including the center point position, size, and heading angle.

[0211] The positional motion model predicts the next moment's position of the target's center point. This model considers the target's velocity, acceleration, and potential external forces (such as wind and thrust). The dimensional motion model predicts the next moment's change in the target's dimensions. This model considers factors such as deformation, scaling, or rotation. The heading angle motion model predicts the next moment's change in the target's heading angle. This model considers the target's turning behavior, path planning, and the influence of the external environment.

[0212] Historical trajectory data and current bounding box information are input into the corresponding motion models. The position motion model predicts the center point position at the next moment based on the target's historical and current positions. The size motion model predicts the size change at the next moment based on the target's historical and current sizes. The heading angle motion model predicts the heading angle change at the next moment based on the target's historical and current heading angles.

[0213] Based on the predicted target position, size, and heading angle for the next moment, the target's bounding box information is updated. This updated bounding box information is then combined with historical trajectory data to form a complete trajectory tracking result. Throughout the tracking process, new detection box information can be continuously used to calibrate and update the motion model, improving prediction accuracy. In practical applications, prediction results may contain errors due to factors such as occlusion, target disappearance, or sensor noise. Algorithms such as basic Kalman filtering can be used to calibrate the prediction results, improving the stability and accuracy of trajectory tracking.

[0214] Specifically, by utilizing historical trajectory data and current detection box information, the center point position, size, and heading angle of the target bounding box are managed using the PoseMotion Model (PMM), SizeMotion Model (SMM), and YawMotion Model (YMM) respectively. This predicts the target position at the next moment of the trajectory and manages the target trajectory throughout its entirety, ensuring the stability and accuracy of the tracking system.

[0215] The basic Kalman filter motion model assumes that the state transition of the target from time t-1 to time t conforms to the state equation:

[0216]

[0217] in, This represents the true state at time t-1. Indicates the effect on The state transition model on, Represents the controller vector. Indication function The input-control model above, Let this represent process noise. Assume the process noise follows a multivariate normal distribution with a mean of 0 and a covariance matrix of... The formula is shown below:

[0218]

[0219] The true state at time t With observation status Between them, the observation equation is satisfied:

[0220]

[0221] in The observation model represents the mapping of the target from the real state space to the observation space. Let represent the observation noise, which follows a normal distribution with a mean of 0 and a covariance matrix of . The formula is shown below:

[0222]

[0223] The initial state and the noise at each time step are independent of each other, and the initial state is known.

[0224] The system state is estimated and the uncertainty of the system state is updated through two recursive steps: prediction and update. In the prediction step, the current state is predicted based on the previous state estimate using the system's state transition equation; and the covariance matrix of the state estimate, i.e., the uncertainty noise of the current state, is predicted based on the system's dynamic model and process noise. The formula for predicting the current state is shown below:

[0225]

[0226] in, This represents the predicted state value at time t. The predicted value has a large error compared to the estimated state value. The estimated value is obtained after correction using the observed value.

[0227] Current state prediction error covariance The formula is shown below:

[0228]

[0229] Among them, due to process noise Unknown, therefore this item is temporarily ignored; the actual state at time x-1. Unknown, estimated using time t-1 Substitute; matrix Represents the state transition model; For matrix Transpose of; Represents the controller vector; Indication function The input-control model above, This represents the covariance of the state estimation error at time t-1. The covariance matrix represents the process noise.

[0230] In the update step, the Kalman gain at time t is first calculated by comparing the observed values ​​with the predicted state, as shown in the following formula:

[0231]

[0232] in, This represents the covariance of the current state prediction error. Represents the observation model, This represents the observation noise covariance matrix.

[0233] Next, the observation residuals are calculated using the observation equation, as shown in the following formula:

[0234]

[0235] in Represents the observed value at time t. Let represent the predicted state value at time t. Then, the observed and predicted values ​​are weighted and fused according to the Kalman gain to update the state estimate at time t. The formula is shown below:

[0236]

[0237] Finally, update the state estimation error covariance matrix. The formula is shown below:

[0238]

[0239] in, The state estimate is an identity matrix. The updated state estimate takes into account the observations at the current time and is closer to the true state. At the same time, the state estimate error covariance matrix reflects the new uncertainties and measures the accuracy of the estimate.

[0240] Position Kalman Filter Motion Model (AMM): Considering the highly nonlinear motion characteristics of the target object's trajectory, using a constant-velocity motion model would result in significant trajectory prediction errors due to continuous missed detections or occlusions. Therefore, a constant acceleration-based motion model is used on top of the basic Kalman filter, with additional rigid structural constraints on the object introduced. The target's scene space is defined as a six-dimensional vector, as shown in the following formula:

[0241]

[0242] in, This represents the coordinates of the geometric center of the target object in three-dimensional space. Since the target's height coordinate (z) remains constant during ground travel, a z-direction coordinate is not included. and These represent the velocity and acceleration in the x-direction, respectively, which are consistent with the direction of the heading angle. These represent the velocity and acceleration in the y-direction, respectively.

[0243] Furthermore, when the target is occluded or leaves the sensor's field of view, the detection quality deteriorates, causing jitter in the detection box used to locate the target object. This leads to errors in the target's trajectory localization, resulting in trajectory drift. This deviation is caused by detection box localization jitter noise, which causes errors in the target object's motion process, ultimately affecting the state update and prediction of the Kalman filter.

[0244] Two types of detection box location trembling noise (LTN) are defined in the position motion model and incorporated into the basic Kalman filter motion model to calibrate errors caused by occlusion and target disappearance, improve the accuracy of the motion model, and maintain the stability of occluded target IDs. For the detection box location trembling noise, different sampling sensor parameters result in different specific parameters of the sampled data noise. In this embodiment, the parameters of this noise are calculated on the KITTI target tracking dataset. The formula is shown below:

[0245]

[0246]

[0247] Using the algebraic bias shown above for noise and The model was constructed, and the parameters of the noise were calculated on the KITTI target tracking dataset. The x and y coordinates represent the geometric center of the ground truth bounding box in the dataset. This represents the x and y coordinates of the geometric center of the bounding box in the target data obtained by the detection algorithm, where N is the number of categories in the dataset. This indicates the number of targets corresponding to each category.

[0248] The deviations in the x and y directions vary with distance, exhibiting a Gaussian distribution. The covariance matrix is ​​calculated using the formula shown below:

[0249]

[0250]

[0251] in, and These represent the covariance matrices along the x-axis and y-axis, respectively.

[0252] Define the matrix to obtain the detection noise covariance matrix on the KITTI dataset. The formula is shown below:

[0253]

[0254] in, This indicates the effect of noise on the detection frame.

[0255] This detection noise, or detector noise, affects the observation results and observation residuals. The observation residuals are updated using Kalman gain weighting for state estimation; therefore, the detector noise residuals are considered... The Kalman gain of the basic Kalman filter is added to solve the problem of jitter in the occlusion detection box.

[0256] The improved Kalman gain formula is shown below:

[0257] .

[0258] in, This indicates the improved Kalman gain; Let represent the covariance matrix.

[0259] Size Kalman Filter Motion Model (SMM): Theoretically, the size of the same target should remain constant. However, due to potential errors during perception, a separate constant-velocity Kalman filter motion model is set up for the size to ensure its stability and continuity. The target's height remains constant, and the size model only needs to predict the changes in the target object's length and width during motion. The target's size state vector is represented as a two-dimensional vector, as shown in the following formula:

[0260]

[0261] in, and These represent the length and width of the target object during its motion, respectively. The initialization, prediction, and update of the motion model in SMM are the same as those in the basic Kalman filter.

[0262] Kalman filter motion model for heading angle: For heading angle, the constant velocity Kalman filter motion model is also used, and the state variable formulas are as follows:

[0263]

[0264] The initialization, prediction, and update calculation process of the heading angle model state is the same as that of the size model, and will not be described again here.

[0265] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0266] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A three-dimensional occlusion target tracking method based on multimodal spatiotemporal interaction, characterized in that, The method includes: The original point cloud and the original image are acquired at the same time, and the original point cloud and the original image are preprocessed to obtain global fusion features; The global fusion features are processed by a region proposal network to obtain an initial detection box. The initial detection box and the global fusion features are then processed by region of interest pooling to obtain a global fusion feature of the region of interest. The global fusion feature of the region of interest is then separated to obtain global point cloud features and global image features. The global point cloud features include non-empty voxel point cloud features. The non-empty voxel point cloud features are projected onto the global image features to obtain the visible image features of the occluded target. The visible image features of the occluded target are reconstructed and decoded to obtain the complete high-dimensional features of the occluded target. The complete high-dimensional features of the occluded target are smoothed by reconstructing feature edge denoising convolution to obtain a second fused feature; the second fused feature is then processed by a feedforward neural network to obtain a refined detection box. The refined detection box is filtered to obtain a second refined detection box. Distance calculation and validity judgment are performed on the second refined detection box to obtain a valid detection box. The valid detection boxes are grouped and density is calculated, and valid bounding boxes are filtered out according to the density. Target marking and detection box deletion are performed according to the valid bounding boxes, and trajectory confidence is updated and valid trajectory is marked. A bipartite graph structure is adopted. A graph is constructed based on the global fusion feature, the second fusion feature, the detection box at time t and the trajectory at time t-1 to obtain the final features of the detection box and the final features of the trajectory. The appearance association score between the final features of the detection box and the final features of the trajectory is calculated by adaptive channel graph convolution. The detection box and trajectory prediction box are calculated based on the generalized intersection joint method to obtain the geometric cost matrix. The appearance association score and the geometric cost matrix are then encapsulated to obtain the final association matrix. The final correlation matrix is ​​matched using the Hungarian algorithm to obtain successfully matched detection boxes and trajectory prediction boxes, as well as unmatched detection boxes and unmatched trajectory prediction boxes. Based on historical trajectory data and current detection box information, the center point position, size and heading angle of the target bounding box are processed by position motion model, size motion model and heading angle motion model respectively to predict the target position at the next moment of the trajectory and to track the target trajectory throughout the entire process.

2. The three-dimensional occlusion target tracking method based on multimodal spatiotemporal interaction as described in claim 1, characterized in that, The preprocessing of the original point cloud and the original image to obtain global fusion features includes: A dual-branch structure is used to extract the original point cloud and the original image respectively, to obtain the three-dimensional spatial information of the original point cloud and the appearance texture information of the original image. Then, feature fusion is performed on the three-dimensional spatial information and the appearance texture information to obtain a global fused feature including the three-dimensional spatial information and the appearance texture information.

3. The three-dimensional occlusion target tracking method based on multimodal spatiotemporal interaction as described in claim 1, characterized in that, The process of filtering the refined detection box to obtain a second refined detection box, and then performing distance calculation and validity judgment on the second refined detection box to obtain a valid detection box includes: Set a first preset threshold, filter out refined detection boxes with confidence scores less than the first preset threshold, obtain a second refined detection box, calculate the Euclidean distance between the second refined detection box and each trajectory prediction box, obtain the shortest Euclidean distance, and use the shortest Euclidean distance as the distance between the second refined detection box and the trajectory. A second preset threshold is set. If the shortest Euclidean distance is less than the second preset threshold, the second refined detection box is considered a valid detection box. If the shortest Euclidean distance is not less than the second preset threshold, the second refined detection box is considered an invalid detection box.

4. The three-dimensional occlusion target tracking method based on multimodal spatiotemporal interaction as described in claim 1, characterized in that, The step of grouping and density calculation of the valid detection boxes, and filtering out valid bounding boxes based on the density, includes: The valid detection boxes are grouped according to their categories. Detection boxes of the same type of target are grouped together to obtain a candidate set of detection boxes for each category. For each group of candidate detection boxes, the geometric association cost between all detection boxes in the candidate set is calculated, and the maximum intersection of the detection box with other detection boxes is taken as the density of the detection box. Set a third preset threshold, and regard the detection boxes with a density less than the third preset threshold as valid bounding boxes; sort the remaining detection boxes in each group by confidence, regard the detection box with the highest confidence as a valid bounding box, and use the detection box as the reference box for this round.

5. The three-dimensional occlusion target tracking method based on multimodal spatiotemporal interaction as described in claim 4, characterized in that, The step of marking the target and deleting the detection box based on the legal bounding box, and updating the trajectory confidence and marking the legal trajectory includes: Calculate the intersection between every two remaining detection boxes in each group, and consider the detection boxes whose intersection between every two detection boxes is greater than the third preset threshold as the same target as the reference box, and delete the detection boxes whose intersection with the reference box is greater than the third preset threshold, until the candidate set of detection boxes in each group is empty; A fourth preset threshold is set, and an initial confidence score is set for each trajectory. The initial confidence score is updated and calculated based on the confidence item of the detection box associated with the trajectory, the trajectory missing duration item, and the trajectory confidence time validity item to obtain the trajectory confidence score. Trajectories with a trajectory confidence score higher than the fourth preset threshold are marked as valid trajectories.

6. The three-dimensional occlusion target tracking method based on multimodal spatiotemporal interaction as described in claim 1, characterized in that, The method employs a bipartite graph structure, constructing a graph based on the global fusion features, the second fusion features, the detection boxes at time t, and the trajectories at time t-1, to obtain the final features of the detection boxes and the final features of the trajectories, including: Let the bipartite graph be ,in, This is a collection of bounding box nodes, with a size of [size missing]. ; This is a set of trajectory nodes, with a size of [value missing]. ; Let be the set of edges, representing the correlation score between the detection box and the trajectory; Among them, the detection box node The feature consists of a first appearance feature and a first motion feature, the first appearance feature For the detection frame node The corresponding local fusion features, the first motion feature Obtained through two layers of MLP: ,in express The first moment One detection box; trajectory nodes The feature consists of a second appearance feature and a second motion feature, the second appearance feature ,in Indicates according to Time trajectory prediction Time trajectory prediction box This indicates context-aware dynamic enhancement of local features. Represents global fusion features; motion features Extracted using LSTM: ,in express The first moment A set of trajectories; The first appearance feature and the first motion feature are concatenated, and multi-head attention (MHSA) and LSTM are added for feature enhancement to obtain the final feature of the detection box; the second appearance feature and the second motion feature are concatenated, and multi-head attention (MHSA) and LSTM are added for feature enhancement to obtain the final trajectory feature.

7. The three-dimensional occlusion target tracking method based on multimodal spatiotemporal interaction as described in claim 6, characterized in that, The appearance association score between the final features of the detection box and the final features of the trajectory is calculated by adaptive channel graph convolution, including: By overlaying multiple ACGC node aggregation modules and performing multiple node information transfers and updates, the appearance correlation score between the detection box and the trajectory is calculated. For the first... The node aggregation module of layer ACGC takes the first node as input. Layer node features and learnable adjacency matrix The attention score for each edge is calculated using a self-attention mechanism: in Indicates splicing operation and Representing nodes respectively and The weights; And add a learnable adjacency matrix , in, Represents a node and The correlation score between nodes; using attention scores for message passing and node updates: in Represents a node The set of neighboring nodes; This represents the feature set of node u at layer k-1 and other nodes; pass and The function calculates the learnable adjacency matrix of this layer: in, Used as input for the next-level node aggregation module; This represents the weight of node m in the k-th layer; the output of the last layer ACGC node aggregation module is the final set of appearance association scores between the bounding box and the trajectory. in Represents a node and The appearance correlation score between them.

8. The three-dimensional occlusion target tracking method based on multimodal spatiotemporal interaction as described in claim 7, characterized in that, The method based on generalized intersection joint calculates the geometric cost matrix for the detection box and the trajectory prediction box. The appearance association score and the geometric cost matrix are then encapsulated to obtain the final association matrix, including: The geometric cost matrix is ​​obtained by calculating the detection box and the trajectory prediction box using the three-dimensional generalized joint intersection or the bird's-eye view generalized joint intersection; The appearance association score and the geometric cost matrix are encapsulated to obtain the final association matrix, as shown in the following formula: in, and These represent the weights of the appearance association score and the set similarity score, respectively. This represents the geometric similarity matrix.

9. The three-dimensional occlusion target tracking method based on multimodal spatiotemporal interaction as described in claim 1, characterized in that, The historical trajectory data includes Trajectory data at any given time, ; The target bounding box includes successfully matched detection boxes and trajectory prediction boxes, as well as unmatched detection boxes; the method further includes: In the positional motion model, the target bounding box is defined with positioning jitter noise and added to the basic Kalman filter motion model to calibrate the errors caused by occlusion and target disappearance.

Citation Information

Patent Citations

  • Multi-target tracking method, device and system and computer readable storage medium

    CN112883819A

  • Open vocabulary small target detection method based on cascade multistage refinement

    CN119360410A