Three-dimensional point cloud tracking method based on spatial-temporal feature fusion
Through the methods of spatial and temporal feature fusion and dynamic anchor point adjustment, the target tracking problem of three-dimensional point cloud data in complex scenarios is solved, achieving higher accuracy and robustness, especially stable tracking in target occlusion and fast motion.
Patent Information
- Application Number
- CN202510311461.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
The existing three-dimensional point cloud data lacks feature information in complex scenarios, resulting in imperfect data fusion, affecting the quality of target tracking, especially in poor performance under target occlusion, deformation and rapid movement.
The spatial and temporal feature fusion method is adopted to extract the features of multi-frame historical template point clouds and search area point clouds through the spatial and temporal attention mechanism, and introduce a dynamic weighting mechanism to weight the spatial features, temporal features and dynamic weighting mechanisms, and combine the candidate area generation method of dynamic anchor point adjustment to optimize the anchor point position and size to achieve accurate target tracking.
It improves the accuracy and robustness of three-dimensional point cloud target tracking, especially in complex scenarios, which can effectively deal with target occlusion and rapid movement, improving the stability and accuracy of tracking.
Smart Images

Figure CN120259363A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional point cloud tracking, and in particular to a three-dimensional point cloud target tracking method based on spatio-temporal feature fusion. Background Art
[0002] The statements in this section merely mention the background art related to the present invention and do not necessarily constitute prior art.
[0003] With the breakthrough progress of three-dimensional perception technology, three-dimensional point cloud data has become the core perception medium in fields such as autonomous driving, robot navigation, and augmented reality. Its unique geometric representation ability enables three-dimensional target tracking to get rid of the limitations of traditional two-dimensional images being sensitive to light and texture-dependent, and can directly analyze the physical pose of the target through the spatial coordinates of the point cloud, providing new possibilities for high-precision positioning in dynamic scenarios. However, the special nature of three-dimensional point cloud data and the diversity of complex scenes make it impossible to extract sufficient feature information, resulting in imperfect data fusion, loss of features, and affecting the tracking quality of three-dimensional point cloud target tracking. Summary of the Invention
[0004] To solve the deficiencies of the prior art, the present invention provides a three-dimensional point cloud target tracking method based on spatio-temporal feature fusion. By using a spatio-temporal attention mechanism to extract the features of multiple frames of historical template point clouds and the current search area point cloud, introducing a dynamic weight mechanism, and weighted fusing the spatial features, temporal features, and dynamic weight mechanism, enhanced features are generated. Finally, a candidate region generation method with dynamic anchor adjustment is used to dynamically adjust the anchor position through target motion prediction, optimize the anchor size in combination with the spatio-temporal attention mechanism, generate high-quality candidate boxes, and achieve accurate target tracking through feature extraction, target prediction, and confidence screening, and update the target prediction result to the historical template library. This method effectively addresses problems such as target occlusion, deformation, and fast movement, improving the accuracy and robustness of tracking.
[0005] In the first aspect, the present invention provides a three-dimensional point cloud target tracking method based on spatio-temporal feature fusion;
[0006] A three-dimensional point cloud target tracking method based on spatio-temporal feature fusion includes:
[0007] Obtain point cloud sequence data, create a historical template library, and define the search area point cloud. The features of multiple frames of historical template point clouds and the search area point cloud are extracted through a spatio-temporal attention mechanism;
[0008] Introduce a dynamic weight mechanism, weighted fuse the spatial attention, temporal attention, and dynamic weight mechanism, adaptively embed the fused spatio-temporal features of multiple frames of historical template point clouds into the search area point cloud, update the search area features, and enhance the expression ability of the target;
[0009] Based on the candidate region generation strategy, multiple candidate regions are extracted within the search region, and the three-dimensional bounding box, center point, and category of the target are predicted through the features of these regions, and the optimal result is screened through confidence scoring.
[0010] Further, the feature extraction operation combining spatio-temporal attention mechanism is performed on the multi-frame historical template point cloud and the search point cloud in parallel, specifically:
[0011] Input the multi-frame historical template point cloud and the search point cloud into a parallel feature extraction network for feature extraction;
[0012] Among them, the feature extraction network includes a spatial feature extraction module and a temporal feature extraction module, and the spatio-temporal attention mechanism is respectively embedded in both feature extraction modules.
[0013] Preferably, the spatio-temporal feature extraction module processes the multi-frame historical template point cloud and the search point cloud specifically as follows: Feature extraction is performed on the multi-frame historical template point cloud through the spatial feature extraction module and the temporal feature extraction module to obtain spatio-temporal features and input them into the spatio-temporal attention mechanism for weighting; Feature extraction is performed on the search point cloud through the spatial feature extraction module to obtain spatial features and input them into the spatio-temporal attention mechanism for weighting.
[0014] Further preferably, the spatio-temporal attention mechanism processes the spatio-temporal features specifically including: Passing the point cloud features through three different weight matrices to obtain the corresponding query vector, key vector, and value vector, and then calculating the similarity between the query and all keys, and normalizing it to the attention weight through an activation function. Using the attention weight to perform weighted summation on the values to obtain the weighted output of each query.
[0015] Preferably, the weight parameters are shared among the multi-frame historical template point cloud feature extraction networks. Further, the combination of the spatial feature, temporal feature, and dynamic weight mechanism is specifically: The feature enhanced by spatial attention After max pooling, it is combined with the temporal attention and the dynamic weight mechanism to generate an enhanced feature, and the spatio-temporal features of the multi-frame historical template point cloud are adaptively embedded into the search region point cloud features, thereby updating the search region features
[0016] Further, the target tracking prediction according to the candidate region generation method adjusted by dynamic anchors specifically includes: First, initialize the anchor and dynamically adjust the anchor point position through target motion prediction. Secondly, optimize the anchor size by combining the spatio-temporal attention mechanism to generate high-quality candidate boxes, and achieve accurate target tracking through feature extraction, target prediction, and confidence screening. Finally, predict the three-dimensional bounding box, center point, and category of the target.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0018] 1. The present invention introduces a dynamic weight update mechanism to optimize the spatio-temporal features by weighting according to the real-time performance of target tracking, enabling the multi-frame historical template point cloud and the search area point cloud to adaptively adjust the weights during the feature fusion process. This dynamic adjustment can effectively improve the accuracy and robustness of target tracking, especially in scenarios where the target moves rapidly or the environment changes significantly.
[0019] 2. The present invention embeds an attention mechanism in the spatial feature extraction module and the temporal feature extraction module through a parallel spatio-temporal attention mechanism, fully capturing the spatio-temporal relationship between single-frame point clouds and multi-frame point clouds, enhancing the spatio-temporal representation ability of the target, and being able to better understand the spatial structure and temporal evolution in the point cloud, thereby improving the tracking performance.
[0020] 3. The present invention introduces a dynamic anchor adjustment mechanism to adaptively optimize the anchor position according to the target motion prediction, enabling the candidate region to more accurately cover the target region. During the feature extraction and fusion process, the anchor size is optimized in combination with the spatio-temporal attention mechanism to adapt to the deformation of the target and environmental changes. This dynamic adjustment method effectively improves the accuracy and robustness of target tracking, especially showing more stability in complex scenarios where the target moves rapidly or is severely occluded. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only exemplary, and those of ordinary skill in the art can obtain other drawings according to the provided drawings without creative efforts.
[0022] Figure 1 It is a schematic flowchart of the three-dimensional point cloud target tracking method in the embodiment of the present invention;
[0023] Figure 2 It is a schematic flowchart of spatio-temporal feature extraction in the embodiment of the present invention;
[0024] Figure 3 It is a schematic flowchart of the spatio-temporal attention module in the embodiment of the present invention; DETAILED IMPLEMENTATION MANNER
[0025] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0026] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0028] Embodiment 1
[0029] The present invention discloses a three-dimensional point cloud object tracking method based on spatio-temporal feature fusion. Aiming at the problems of sparsity, disorder and lack of texture features of point cloud data, this method improves the accuracy and robustness of object tracking by combining spatio-temporal attention mechanisms.
[0030] First, obtain point cloud sequence data, create a historical template library, and extract the template point cloud of the historical template library T history as multi-frame historical template point clouds, and define the search area point cloud in the current frame. Then, construct a spatio-temporal attention mechanism module to extract features from the multi-frame historical template point clouds and the search area point cloud. On this basis, introduce a dynamic weight mechanism, combine the spatial features, temporal features and the dynamic weight mechanism, and adaptively embed the spatio-temporal features of the multi-frame historical template point clouds into the search area point cloud, so as to update the search area features. Finally, through the candidate region generation strategy, extract multiple candidate regions in the search area, use the features of these candidate regions to predict the three-dimensional bounding box, center point coordinates and object category of the target, and screen the optimal prediction results through confidence scoring, and finally realize the accurate tracking and prediction of the three-dimensional target object.
[0031] Next, combine Figures 1-3 A three-dimensional point cloud object tracking method based on spatio-temporal feature fusion disclosed in this embodiment is described in detail. The three-dimensional point cloud object tracking method based on spatio-temporal feature fusion includes the following steps:
[0032] S1. Obtain point cloud sequence data, and create a historical template library T history , and the historical template library is used to store the target point cloud information extracted from the historical frames.
[0033] S2. Extract the historical template library T historyThe template point cloud therein is used as a multi-frame historical template point cloud, and a search area point cloud is defined in the current frame.
[0034] S3. Extract spatio-temporal features from the multi-frame historical template point cloud and the search area point cloud through a spatio-temporal attention mechanism module; introduce a dynamic weight mechanism to dynamically adjust the weights of the historical frame features according to the motion pattern and context information of the target object.
[0035] Specifically, feature extraction is performed on the multi-frame historical template point cloud and the search area point cloud. The feature extraction network is PointNet++. Then, a spatio-temporal attention module is embedded on the basis of this feature extraction network, and the weight parameters are shared among the feature extraction networks of the multi-frame historical template point clouds. Among them, the spatio-temporal attention module is divided into a spatial attention module and a temporal attention module.
[0036] As an implementation manner, in combination with Figures 2-3 , step S3 is described in detail. The specific process includes:
[0037] S301. Input the multi-frame historical template point cloud and the search area point cloud into the PointNet++ network for feature extraction, and extract the multi-frame historical template point cloud F t-1 , F t-2 ..., F t-j and the point cloud feature F s of the search area point cloud.
[0038] S302. Pass the multi-frame historical template point cloud and the search area point cloud through the spatial attention module respectively. Specifically, the processing process of each frame is as follows. First, the corresponding query vector, key vector, and value vector are obtained through three different weight matrices. Then, the similarity between the query and all keys is calculated, and the attention weights are normalized through an activation function. The values are weighted and summed using the attention weights to obtain the weighted output of each query. The formula is as follows:
[0039]
[0040] Among them, is the spatial enhanced feature of the search point cloud, d k is the dimension of the feature, and Q, K, and V are the query, key, and value obtained through linear transformation respectively. Specifically:
[0041] Q = W q F s
[0042] K = W k F s
[0043] V = W v F s
[0044] Similarly, the spatial enhanced features of the historical frame point cloud The calculation is the same as above. W q ,W k ,W v is a learnable weight matrix, and three weight matrices are shared among historical frames.
[0045] S303. Input the spatial features of the multi-frame historical template point cloud extracted by the spatial attention module into the temporal attention module for temporal attention calculation. The temporal features of the search area point cloud are indirectly provided by the spatio-temporal features of the multi-frame historical template point cloud, and the calculation is not performed here.
[0046] Specifically, first use the self-attention mechanism to calculate the temporal relationship weights between frames. The formula is as follows:
[0047]
[0048] where, F t temporal is the temporal enhanced feature of the historical frame point cloud, d k is the dimension of the feature, and Q, K, V are obtained from the input features through linear transformation. Specifically:
[0049] Q = W q F s
[0050] K = W k ·{F t-1 , F t-2 ..., F t-j}
[0051] V = W v ·{F t-1 , F t-2 ..., F t-j}
[0052] Then, the dynamic weight mechanism dynamically adjusts the weights of the historical frame features according to the motion pattern and context information of the target object. The formula is as follows:
[0053] w = σ(F s W w )
[0054] where, σ is the Sigmoid function, ensuring w ∈ [0, 1]. W w is a learnable weight matrix.
[0055] S4. The features enhanced by spatial attention After max pooling, it is combined with the temporal attention and dynamic weight mechanism to generate enhanced features, and the spatio-temporal features of multiple-frame historical template point clouds are adaptively embedded into the search region point cloud features, thereby updating the search region features The specific formula is as follows:
[0056]
[0057]
[0058] S5. Finally, the fused depth features Through the candidate region generation strategy, multiple candidate regions are extracted within the search region. The features of these candidate regions are used to predict the 3D bounding box, center point coordinates, and target category of the target, and the optimal prediction result is selected through confidence scoring, ultimately achieving accurate tracking and prediction of the 3D target object.
[0059] S501. First, extract the position sequence {P t-1 , P t-2 ,..., P t-n} of the target from the multi-frame historical template point cloud, and use the Kalman filter to predict the target position of the current frame
[0060] P t = KalmanFilter(P t-1 , P t-2 ,.., P t-n )
[0061] S502. Based on the predicted target position dynamically adjust the anchor point A0 within the search region, move the initial position of the anchor point to the area where the target may appear, and reduce the number of invalid anchor points. The adjusted anchor point position A t is expressed as:
[0062] A t = A0 + ΔA
[0063] where ΔA is the offset calculated according to the target motion prediction
[0064] S503. Calculate the attention weight w s of each anchor point through the spatio-temporal attention mechanism STA(·), and dynamically adjust the anchor point size S i :
[0065] w s = STA(A t , F t temporal )
[0066] S i = S0·(1 + α·σ(w s - τ))
[0067] where S0 is the basic anchor size, τ is the preset attention weight threshold; α is the adjustment coefficient, controlling the variation range of the anchor size; is the Sigmoid activation function, making the adjustment smoothly transition.
[0068] S504. Generate candidate bounding box B using the adjusted anchor i , and filter out the low-weight candidate bounding boxes according to the attention weight w s , retaining the candidate regions with high confidence. The generation formula of candidate bounding box B i is:
[0069] B i = A t + S i
[0070] Extract the point cloud feature F for each candidate bounding box i , and regress the three-dimensional bounding box, center point coordinates and category of the target through the fully connected layer, and optimize the prediction result using the cross-entropy loss.
[0071] S505. Screen the optimal candidate bounding boxes through non-maximum suppression (NMS) and confidence score, output the three-dimensional position, size and category of the target, and complete the target tracking; and update the historical template library T according to the tracking result of the current frame history .
[0072] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not deviate from the spirit and scope of the present invention defined by the appended claims. It should be understood that the different dependent claims and the features described herein can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a single embodiment can be used in other described embodiments.
Claims
1. A three-dimensional point cloud tracking method based on spatio-temporal feature fusion, characterized in that The method includes the following steps: Step S1: Obtain point cloud sequence data and create a historical template library; Step S2: Use the target point clouds in the current frame and historical frames as multi-frame historical template point clouds, delimit a search area in subsequent frames, and extract the point clouds within the search area as search area point clouds; Step S3: Extract spatio-temporal features from the template point clouds and search area point clouds through a spatio-temporal attention mechanism module; Step S4: Adopt a method that combines spatial attention, temporal attention, and a dynamic weight mechanism to adaptively embed the spatio-temporal features of the template point clouds into the search area point clouds to complete the update of the search area features; Step S5: Use a candidate region generation method with dynamic anchor adjustment to extract multiple candidate regions within the search area, predict the three-dimensional bounding box, center point coordinates, and target category of the target using the features of the candidate regions, screen the optimal prediction results through confidence scoring, and update the target prediction results to the historical template library.
2. The three-dimensional point cloud object tracking method based on spatio-temporal feature fusion according to claim 1, wherein The feature extraction operation that combines the spatio-temporal attention mechanism is performed in parallel on the multi-frame historical template point clouds and search point clouds, specifically: Input multiple frames of historical template point clouds {P t-1 ,..., P t-k} and the search point cloud P s into a parallel feature extraction network for feature extraction; among them, a spatio-temporal attention mechanism is respectively embedded in each of the feature extraction networks; and the weight parameters are shared among the feature extraction networks of multiple frames of historical template point clouds.
3. The three-dimensional point cloud object tracking method based on spatio-temporal feature fusion according to claim 1, wherein The spatio-temporal feature extraction module processes the multi-frame historical template point cloud and the search point cloud as follows: The spatial feature extraction module extracts the spatial features F from the multi-frame historical template point cloud t , F t-1 , …, F t-k , and inputs them into the spatio-temporal attention mechanism for weighted output to obtain enhanced spatio-temporal features The spatial feature extraction module extracts features from the search point cloud P s to obtain the spatial feature F s , and inputs it into the spatio-temporal attention mechanism for weighted output to obtain enhanced spatial features 4. The three-dimensional point cloud object tracking method based on spatio-temporal feature fusion according to claim 1, wherein Introduce a dynamic weight mechanism, and perform weighted fusion of spatial attention, temporal attention, and the dynamic weight mechanism to generate enhanced features.
5. The 3D point cloud object tracking method based on spatio-temporal feature fusion according to claim 1, characterized in that The target position of the current frame is predicted first through Kalman filtering and according to the anchor point A0 is dynamically adjusted. Secondly, according to the weight w s the size S of the anchor point is dynamically adjusted i , and the candidate box B is generated by using the adjusted anchor point i . Then, through feature extraction, target prediction, and confidence screening of the candidate box B i , the three-dimensional bounding box, center point, and category of the target are finally predicted, and the target prediction result is updated to the historical template library
Citation Information
Cited By
Single target tracking method based on multi-scale feature fusion and channel attention mechanism
CN121095287A