A 3D single-object tracking method for spatio-temporal context information based on memory network

By constructing a three-dimensional single-object tracking method for spatiotemporal context information based on memory network, using external storage units to save the spatiotemporal features of historical frames, and combining with mask priors to enhance target feature representation, the problem of insufficient robustness in the existing methods is solved, and the accurate tracking effect in complex scenarios is achieved.

CN116596969BActive Publication Date: 2025-07-25ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310565602.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-19
Publication Date
2025-07-25
Estimated Expiration
2043-05-19

AI Technical Summary

Technical Problem

The existing three-dimensional single-object tracking method lacks robustness when dealing with target occlusion and long-term sequences, and cannot effectively utilize historical spatio-temporal information and geometric structures, resulting in a degradation of tracking performance.

Method used

The three-dimensional single-object tracking method of spatiotemporal context information based on memory network is adopted. By constructing a network model of the target tracking system, the spatiotemporal characteristics of historical frames are saved using external storage units, and the target feature representation is enhanced in combination with the mask prior to the target feature to achieve accurate tracking of the target.

Benefits of technology

Accurate and fast tracking of the goals in complex scenarios, improving success and precision scores on the KITTI dataset, showing excellent robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116596969B_ABST
    Figure CN116596969B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of 3D vision, and discloses a three-dimensional single-object tracking method for spatio-temporal context information based on a memory network, including the following steps: Step S1: Construct a target tracking system network model; Step S2: Set a memory set; Step S3: Obtain a query frame, and extract key-value encoding pairs of the query frame and memory frames; Step S4: Use a feature matching unit to perform matching calculations on the key-value encoding pairs of the query frame and the key-value encoding pairs of the memory frames in the external storage unit to obtain matching features, and the matching features are decoded to obtain the target tracking prediction result of the query frame; Step S5: Put the query frame into the memory set as a memory frame and continue tracking until the task ends; Step S6: Train the target tracking system network model to obtain a three-dimensional single-object tracking method for spatio-temporal context information based on a memory network. The present invention effectively solves the problem of effective encoding of historical spatio-temporal information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of 3D vision, and specifically relates to a three-dimensional single-object tracking method based on a memory network for spatio-temporal context information. Background Art

[0002] Three-dimensional single-object tracking (SOT) is a key task in 3D vision, which has made extensive contributions to various applications, such as autonomous driving, visual surveillance, and robot vision. In recent years, with the development of 3D information acquisition technology, point-cloud-based 3D SOT has attracted extensive attention. Compared with 2D trackers, point-cloud-based 3D trackers are not affected by changes in the natural environment (such as light and weather), and thus are more robust to changes in the surrounding environment. Improving the performance of 3D SOT is of great significance for 3D applications such as autonomous driving and robots.

[0003] Most existing 3D SOT methods follow the siamese network paradigm. The core idea of these methods is to extract features from previous and current frames through a backbone network with shared weights, and then perform feature matching or feature enhancement through a matcher. Although these methods have achieved satisfactory results, the siamese network cannot generate discriminative features for target missing or point-cloud self-occlusion scenarios.

[0004] Different from the above work, MMTrack introduces a motion-centered paradigm and proposes to predict the motion of the target between two consecutive frames. It segments the target points in two frames, then uses PointNet to predict the relative motion of the target, and converts the relative motion amount into a bounding box through a rigid body transformation. However, due to ignoring rich temporal context and geometric structure information, these methods cannot significantly reduce the challenges of 3D SOT.

[0005] In particular, the information contained in past frames in a video sequence is always ignored. Relying only on the previous frame lacks robustness and cannot handle complex situations such as target occlusion. Some methods combine the point clouds of the first frame and the previous frame as a template, but they may not work well on long-term sequences. In addition, P2B discusses the template generation method and attempts to combine all previous results as a template, but the tracking performance decreases instead. Therefore, how to effectively encode historical spatio-temporal information remains an open question.

[0006] In addition to spatio-temporal context features, the geometric information representation method of 3D objects is also worthy of in-depth exploration because it is a huge challenge to distinguish potential objects and backgrounds in sparse scenes. Some methods use shape information to handle object recognition. One specific method is to use a shape completion network to learn the dense geometric features of objects. Slightly different from this method, BAT proposed a BoxCloud representation method to utilize shape priors, which describe the distance between object points and box points (i.e., the corners and the center of the 3D BBox). However, these methods are not a reliable way to extract the geometric information of objects. Summary of the Invention

[0007] In view of the above problems, the present invention proposes a three-dimensional single-object tracking method based on a memory network for spatio-temporal context information, which realizes accurate and fast tracking of objects in many difficult actual scenarios.

[0008] To achieve the above object, the present invention provides a three-dimensional single-object tracking method based on a memory network for spatio-temporal context information, which is characterized by including the following steps:

[0009] Step S1: Construct a network model for the object tracking system, including a first feature extraction unit, a second feature extraction unit, a feature storage unit, and a feature matching unit;

[0010] Step S2: Set a memory set, obtain the first frame of point cloud and determine its tracking object, and then put the first frame of point cloud into the memory set as a memory frame;

[0011] Step S3: Obtain the next frame of point cloud as a query frame, extract the key-value encoding pair of the query frame through the first feature extraction unit, and extract the key-value encoding pair of the memory frame through the second feature extraction unit; the key-value encoding pair includes a key feature and a value feature; the key-value encoding pair of the memory frame is stored in an external storage unit; the key-value encoding pair of the query frame includes a query key feature and a query value feature; the key-value encoding pair of the memory frame includes a memory key feature and a memory value feature.

[0012] Step S4: Use the feature matching unit to perform matching calculations on the key-value encoding pair of the query frame and the key-value encoding pair of the memory frame in the external storage unit to obtain a matching feature, and the matching feature is decoded to obtain the target tracking prediction result of the query frame;

[0013] Step S5: Put the query frame into the memory set as a memory frame, and repeat steps S3 to S5 until the target tracking prediction result is obtained for the last query frame;

[0014] Step S6: Train the network model of the object tracking system, optimize the network parameters by reducing the network loss function until the network converges, and obtain a three-dimensional single-object tracking method based on a memory network for spatio-temporal context information.

[0015] Preferably, step S2 includes the following steps:

[0016] Step S21: Calculate the bounding box of the tracking target of the first-frame point cloud, and extract the target mask according to the bounding box of the tracking target of the first-frame point cloud. The target mask is the points within the bounding box.

[0017] Step S22: Put the first-frame point cloud together with its corresponding target mask into the memory set as the memory frame.

[0018] Step S5 includes the following steps:

[0019] Step S51: Obtain the predicted bounding box according to the target tracking prediction result of the query frame, and extract the target mask of the query frame according to the predicted bounding box.

[0020] Step S52: Put the query frame together with its corresponding target mask into the memory set as the memory frame, and initialize the memory set.

[0021] Preferably, the matching calculation step in step S4 includes:

[0022] Step S41: Calculate the similarity value between the memory key feature and the query key feature.

[0023] Step S42: Extract the memory value feature that best matches the query frame from the external storage unit as the matching value feature through the similarity value obtained in step S41, and calculate the similarity value between the query key feature and the matching value feature.

[0024] Step S43: Calculate the matching feature according to the matching value feature, the query value feature, and the similarity value obtained in step S42.

[0025] Preferably, the mathematical expression of the similarity value in step S41 is:

[0026]

[0027] S i,j is the similarity value between the i-th memory key feature and the j-th query key feature, K M is the memory key feature, is the i-th memory key feature, K Q is the query key feature, is the j-th query key feature.

[0028] Preferably, the mathematical expression of the matching feature in step S43 is:

[0029]

[0030]

[0031] is the matching feature of the j-th query frame, is the j-th query value feature, is the matching value feature, k is the subscript of the query key feature with the maximum similarity value, s kj is the similarity value between the query key feature and the matching value feature of the j-th query frame.

[0032] Preferably, step S3 includes the following steps:

[0033] Step S31: Input the query frame into the first feature extraction unit, first encode it through the query encoding unit, and then map the encoded output results through the first fully connected network and the second fully connected network to generate the query key feature and the query value feature respectively;

[0034] Step S32: Input the memory frame and its corresponding target mask into the second feature extraction unit. First, concatenate the memory frame and its corresponding target mask in the channel dimension, and then input them into the memory encoding unit for encoding to obtain the memory feature The mathematical expression is:

[0035]

[0036] is the memory feature, P l and M l respectively represent the l-th memory frame point cloud and the target mask in the memory set, and Concate represents the concatenation operation along the channel dimension;

[0037] Step S33: Encode the memory frame through the value encoding unit, and then concatenate the encoded output result with the memory feature in step S33 to obtain the concatenated feature Send the concatenated feature into the fusion unit for calculation to obtain the memory value feature;

[0038] Step S34: Map the result output by the value encoding unit in step S33 through the fully connected network to generate the memory key feature;

[0039] Step S35: Store the memory value feature and the memory key feature obtained in steps S33 and S34 into the external storage unit.

[0040] Preferably, the query encoding unit in step S31 is a point cloud feature encoder with an input channel of 3; the memory encoding unit in step S32 is a point cloud feature encoder with an input channel of 4; the value encoding unit in step S33 is a point cloud feature encoder with an input channel of 3, and the fusion unit is specifically a fusion unit with a channel of 64.

[0041] Preferably, in step S33, the fusion unit adopts the KNN algorithm in the feature dimension and extracts K points as neighbors in the feature dimension After that, Self-attention is used to extract the memory value feature, and the mathematical expression is:

[0042]

[0043] V l M is the value feature of the l-th memory frame, X p and Z p are respectively and 's position encodings.

[0044] Preferably, step S6 specifically includes the following steps:

[0045] Step S61: Use the server to obtain N frames of point clouds from the video;

[0046] Step S62: Set the memory set; obtain the first frame of point cloud and determine its tracking target, and then put the first frame of point cloud into the memory set as a memory frame;

[0047] Step S63: Obtain the next frame of point cloud as the query frame, extract the key-value encoding pair of the query frame through the first feature extraction unit, and extract the key-value encoding pair of the memory frame through the second feature extraction unit; the key-value encoding pair of the memory frame is stored in the external storage unit;

[0048] Step S64: Use the feature matching unit to perform matching calculation on the key-value encoding pair of the query frame and the key-value encoding pair of the memory frame in the external storage unit to obtain the matching feature, and the matching feature is decoded to obtain the target tracking prediction result of the query frame;

[0049] Step S65: Put the query frame into the memory set, and repeat steps S63 to S65 until the target tracking task prediction result is obtained for the last query frame, and calculate the tracking loss function L total , and its mathematical expression is:

[0050] L total = L center + L off + L z

[0051] L center represents the deviation between the ground truth and the prediction value under the bird's-eye view vision, L off represents the rotation angle regression deviation, L z represents the deviation on the z-axis;

[0052] Step S65: Optimize the tracking loss function using the server, and use the Adam optimizer to iteratively update the network parameters to reduce the tracking loss function until it converges to a local optimum. At this point, the training ends, and a three-dimensional single-object tracking method based on the memory network for spatio-temporal context information is obtained.

[0053] Preferably, step S61 specifically includes the following steps:

[0054] Step S611: Randomly extract 3 frames of point clouds at intervals from any video in multiple video datasets.

[0055] Step S612: Perform different affine transformations on the 3 frames of point clouds. The affine transformation includes translation and shear. Cut off the points 2m away from the target for each frame of point cloud, and randomly sample 1024 points from the cropped point cloud. As part of data augmentation, perform translation on each frame of point cloud.

[0056] Compared with the prior art, the beneficial effects of the present invention are:

[0057] A three-dimensional single-object tracking algorithm based on the memory network for spatio-temporal context information provided by the present invention, by utilizing spatio-temporal context information, using an external storage unit to save the spatio-temporal features of historical frames, and using mask priors to enhance the target feature representation, can accurately track the target in complex situations such as target disappearance or target occlusion. The success and precision scores on KITTI reach 66.5 and 83.4 respectively, and the success and precision scores on the nuScenes dataset reach 34.0 and 38.6 respectively, showing very good results. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is a schematic diagram of the algorithm framework of a three-dimensional single-object tracking method based on the memory network for spatio-temporal context information of the present invention;

[0059] Figure 2 It is a schematic diagram of the visualization result of a three-dimensional single-object tracking method based on the memory network for spatio-temporal context information of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0061] In view of the problems and deficiencies in the prior art, the present invention proposes a three-dimensional single-object tracking method based on a memory network, which mainly includes the implementation steps of four stages: external storage unit design, target feature matching unit design, model training, and model inference.

[0062] A three-dimensional single-object tracking method based on a memory network proposed by the present invention, as Figure 1 shown, includes the following steps:

[0063] Step S1: Construct a target tracking system network model, including a first feature extraction unit, a second feature extraction unit, a feature storage unit, and a feature matching unit;

[0064] Step S2: Set a memory set, obtain the first frame of point cloud and determine its tracking target, and then put the first frame of point cloud into the memory set as a memory frame;

[0065] Step S3: Obtain the next frame of point cloud as a query frame, extract the key-value encoding pair of the query frame through the first feature extraction unit, and extract the key-value encoding pair of the memory frame through the second feature extraction unit; the key-value encoding pair includes a key feature and a value feature; the key-value encoding pair of the memory frame is stored in the external storage unit; the key-value encoding pair (k Q , v Q ) of the query frame includes a query key feature k Q and a query value feature v Q , where the superscript Q refers to the query frame; the key-value encoding pair (k M , v M ) of the memory frame includes a memory key feature k M and a memory value feature v M , where the superscript M refers to the memory frame.

[0066] Step S4: Use the feature matching unit to perform matching calculations on the key-value encoding pair of the query frame and the key-value encoding pair of the memory frame in the external storage unit to obtain a matching feature, and the matching feature is decoded to obtain the target tracking prediction result of the query frame;

[0067] Step S5: Put the query frame into the memory set as a memory frame, and repeat Step S3 to Step S5 until the target tracking prediction result is obtained for the last query frame;

[0068] Step S6: Train the target tracking system network model, optimize the network parameters by reducing the network loss function until the network converges, and obtain a three-dimensional single-object tracking method based on a memory network.

[0069] A 3D single-object tracking algorithm based on a memory network provided by the present invention utilizes spatio-temporal context information, uses an external storage unit to save spatio-temporal features of historical frames, retains the information contained in past frames, has good robustness, and performs excellently in long-term object tracking tasks.

[0070] The following details the relevant steps.

[0071] Step S2 includes the following steps:

[0072] Step S21: Calculate the bounding box of the point cloud tracking target in the first frame, and extract the target mask according to the bounding box of the point cloud tracking target in the first frame. The target mask is the points within the bounding box.

[0073] Step S22: Put the first-frame point cloud as a memory frame together with its corresponding target mask into the memory set; Step S3 includes the following steps:

[0074] Step S31: Input the query frame into the first feature extraction unit. First, encode it through the query encoding unit, and then map the encoded output results through the first fully connected network and the second fully connected network respectively to generate the query key feature and the query value feature; The query encoding unit is a point cloud feature encoder with an input channel of 3; At the end of the point cloud feature encoder with an input channel of 3, there are two parallel fully connected networks for generating the query frame key-value encoding pair (k Q , v Q ); Figure 1 The key backbone network corresponding to the query frame shown is the query encoding unit of this embodiment;

[0075] Step S32: Input the memory frame and its corresponding target mask into the second feature extraction unit. First, concatenate the memory frame and its corresponding target mask in the channel dimension, and then input them into the memory encoding unit for encoding to obtain the memory feature Figure 1 The value backbone network shown is the memory encoding unit of this embodiment. The memory encoding unit is a point cloud feature encoder with an input channel of 4, The mathematical expression is:

[0076]

[0077] is the memory feature, P l and M l respectively represent the l-th frame memory frame point cloud and the target mask in the memory set, and Concate represents the concatenation operation along the channel dimension;

[0078] Step S33: Encode the memory frame through the value encoding unit, Figure 1The key backbone network of the corresponding memory frame shown is the value encoding unit of this embodiment, and then the result output after encoding is concatenated with the memory feature in step S33 to obtain the concatenated feature The concatenated feature is sent to the fusion unit for calculation to obtain the memory value feature; the value encoding unit is a point cloud feature encoder with 3 input channels; the fusion unit is specifically a fusion unit with 64 channels.

[0079] In step S33, the fusion unit uses the KNN algorithm in the feature dimension and extracts K points as neighbors in the feature dimension After that, Self-attention is used to extract the memory value feature, and the mathematical expression is:

[0080]

[0081] V l M is the value feature of the l-th memory frame, X p and Z p are respectively and the positional encodings of.

[0082] Step S34: Map the result output by the value encoding unit in step S33 through a fully connected network to generate a memory key feature;

[0083] Step S35: Store the memory value feature and the memory key feature obtained in steps S33 and S34 in an external storage unit.

[0084] The matching calculation steps in step S4 include:

[0085] Step S41: Calculate the similarity value between the memory key feature and the query key feature;

[0086] The mathematical expression of the similarity value in step S41 is:

[0087]

[0088] S i,j is the similarity value between the i-th memory key feature and the j-th query key feature, K M is the memory key feature, is the i-th memory key feature, K Q is the query key feature, is the j-th query key feature.

[0089] Step S42: Extract the memory value feature that best matches the query frame from the external storage unit through the similarity value obtained in step S41 as the matching value feature Calculate the similarity value s between the query key feature and the matching value feature kj ;

[0090] Step S43: According to the matching value feature Query value feature and the similarity value s obtained in step S42 kj Calculate the matching feature. Specifically, the matching value feature, the query value feature, and the similarity value are fed into a fully connected network to calculate the matching feature.

[0091] The mathematical expression of the matching feature in step S43 is:

[0092]

[0093]

[0094] is the matching feature of the j-th query frame, is the j-th query value feature, is the matching value feature, k is the subscript of the query key feature with the maximum similarity value, s kj is the similarity value between the query key feature and the matching value feature of the j-th query frame.

[0095] Step S5 includes the following steps:

[0096] Step S51: Obtain the predicted bounding box according to the target tracking prediction result of the query frame, and extract the target mask of the query frame according to the predicted bounding box;

[0097] Step S52: Put the query frame together with its corresponding target mask into the memory set as a memory frame, and initialize the memory set.

[0098] The present invention uses mask prior to enhance the target feature representation, and can accurately track the target in complex situations such as target disappearance or target occlusion.

[0099] Step S6 specifically includes the following steps:

[0100] Step S61: Use the server to obtain N frames of point clouds from the video;

[0101] Step S61 specifically includes the following steps:

[0102] Step S611: Randomly extract 3 frames of point clouds at intervals from any video of multiple video datasets;

[0103] Step S612: Perform different affine transformations on the three frames of point clouds. Affine transformations include translation and shearing. For each frame of point cloud, the points 2m away from the target are sheared off. 1024 points are randomly sampled from the cropped point cloud. As part of data enhancement, each frame of point cloud is translated.

[0104] Step S62: setting a memory set; obtaining the first frame point cloud and determining its tracking target, and then putting the first frame point cloud into the memory set as a memory frame;

[0105] Step S63: obtaining the next frame of point cloud as a query frame, extracting the key-value coding pair of the query frame through the first feature extraction unit, and extracting the key-value coding pair of the memory frame through the second feature extraction unit; the key-value coding pair of the memory frame is stored in the external storage unit;

[0106] Step S64: using a feature matching unit to match the key-value coding pair of the query frame with the key-value coding pair of the memory frame in the external storage unit to obtain matching features, and then decoding the matching features to obtain the target tracking prediction result of the query frame;

[0107] Step S65: put the query frame into the memory set, repeat steps S63 to S65 until the last query frame obtains the target tracking prediction result, and calculates the tracking loss function L of the target tracking system network model total , its mathematical expression is:

[0108] L total =L center +L off +L z

[0109] L center Indicates the deviation between the true value and the predicted value under the bird's-eye view vision, L off represents the rotation angle regression error, L z Indicates the deviation on the z-axis;

[0110] Step S65: Use the server to optimize the tracking loss function, and use the Adam optimizer to iteratively update the network parameters to reduce the tracking loss function until it converges to the local optimum. At this point, the training is completed, and a three-dimensional single target tracking method based on spatiotemporal context information of a memory network is obtained.

[0111] Using this embodiment to continuously track a given tracking target, it is found that the success and precision scores on the KITTI dataset reach 66.5 and 83.4 respectively, improving by 3.6 in terms of the success metric compared to the current state-of-the-art method. Success is defined as the accuracy of the three-dimensional bounding box output by the model, and precision measures the error between the center position of the three-dimensional bounding box output by the model and the center position of the actual true three-dimensional bounding box. Figure 2 It is a schematic diagram of the visualization result of the method of this embodiment; the first row and the second row are the results of successful and failed tracking under two different video sequences respectively, where T represents the subscript of the frame in a video sequence.

[0112] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not depart from the spirit and scope of the present invention as defined by the appended claims. It should be understood that different dependent claims and features in this document can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a single embodiment can be used in other embodiments.

Claims

1. A three-dimensional single-object tracking method based on a memory network for spatio-temporal context information, characterized in that Including the following steps: Step S1: Construct a target tracking system network model, including a first feature extraction unit, a second feature extraction unit, a feature storage unit, and a feature matching unit; Step S2: Set a memory set, obtain the first frame of point cloud and determine its tracking target, and then put the first frame of point cloud into the memory set as a memory frame; Step S3: Obtain the next frame of point cloud as a query frame, extract the key-value encoding pair of the query frame through the first feature extraction unit, and extract the key-value encoding pair of the memory frame through the second feature extraction unit; the key-value encoding pair includes a key feature and a value feature; the key-value encoding pair of the memory frame is stored in an external storage unit; Step S4: Use the feature matching unit to perform matching calculation on the key-value encoding pair of the query frame and the key-value encoding pair of the memory frame in the external storage unit to obtain a matching feature, and the matching feature is decoded to obtain the target tracking prediction result of the query frame; Step S5: Put the query frame into the memory set as a memory frame, and repeat steps S3 to S5 until the target tracking prediction result is obtained for the last query frame; Step S6: Train the target tracking system network model, optimize the network parameters by reducing the network loss function until the network converges, and obtain a three-dimensional single-object tracking method based on the memory network with spatio-temporal context information.

2. A three-dimensional single-object tracking method based on a memory network for spatio-temporal context information according to claim 1, characterized in that, The step S2 includes the following steps: Step S21: Calculate the bounding box of the tracking target of the first frame of point cloud, and extract the target mask according to the bounding box of the tracking target of the first frame of point cloud. The target mask is the points within the bounding box; Step S22: Put the first frame of point cloud into the memory set together with its corresponding target mask; The step S5 includes the following steps: Step S51: Obtain the predicted bounding box according to the target tracking prediction result of the query frame, and extract the target mask of the query frame according to the predicted bounding box; Step S52: Put the query frame into the memory set together with its corresponding target mask, and initialize the memory set.

3. A three-dimensional single-object tracking method based on a memory network for spatio-temporal context information according to claim 1, characterized in that, The matching calculation step in the step S4 includes: Step S41: Calculate the similarity value between the memory key feature and the query key feature; Step S42: Extract the memory value feature that best matches the query frame from the external storage unit through the similarity value obtained in the step S41 as the matching value feature, and calculate the similarity value between the query key feature and the matching value feature; Step S43: Calculate the matching feature according to the matching value feature, the query value feature, and the similarity value obtained in the step S42.

4. A three-dimensional single-object tracking method based on a memory network with spatio-temporal context information according to claim 3, characterized in that: The mathematical expression of the similarity value in the step S41 is: S i,j is the similarity value between the i-th memory key feature and the j-th query key feature, K M is the memory key feature, is the i-th memory key feature, K Q is the query key feature, is the j-th query key feature.

5. A three-dimensional single-object tracking method based on a memory network with spatio-temporal context information according to claim 4, characterized in that: The mathematical expression of the matching feature in the step S43 is: is the matching feature for the j-th query frame, is the j-th query value feature, is the matching value feature, s kj is the similarity value between the j-th query key feature and the matching value feature.

6. The three-dimensional single-object tracking method based on a memory network for spatio-temporal context information according to claim 2, wherein, The step S3 includes the following steps: Step S31: input the query frame into the first feature extraction unit, first encode it through the query encoding unit, and then map the encoded output result through the first fully connected network and the second fully connected network to generate query key features and query value features; Step S32: Input the memory frame and its corresponding target mask into the second feature extraction unit. First, concatenate the memory frame and its corresponding target mask in the channel dimension, and then input them into the memory encoding unit for encoding to obtain memory features The mathematical expression is: is the memory feature, P l and M l respectively represent the l-th memory frame point cloud in the memory set and its corresponding target mask, and Concate represents the concatenation operation along the channel dimension; Step S33: Encode the memory frame through the value encoding unit, and then splice the encoded output result with the memory feature in Step S33 to obtain the spliced feature Send the spliced feature to the fusion unit for calculation to obtain the memory value feature; Step S34: Generate memory key features by mapping the result output by the median encoding unit in step S33 through a fully connected network; Step S35: Store the memory value characteristics and memory key characteristics obtained in steps S33 and S34 into an external storage unit.

7. A three-dimensional single-object tracking method based on a memory network for spatio-temporal context information according to claim 6, characterized in that The query encoding unit in step S31 is a point cloud feature encoder with an input channel of 3; the memory encoding unit in step S32 is a point cloud feature encoder with an input channel of 4; the value encoding unit in step S33 is a point cloud feature encoder with an input channel of 3, and the fusion unit is specifically a fusion unit with 64 channels.

8. A three-dimensional single-object tracking method based on a memory network for spatio-temporal context information according to claim 7, characterized in that: In step S33, the fusion unit uses the KNN algorithm in the feature dimension and extracts K points as neighbors in the feature dimension After that, Self-attention is used to extract the memory value feature, and the mathematical expression is as follows: V l M is the value feature of the l-th memory frame, X p and Z p are respectively and position encodings of 9. A three-dimensional single-object tracking method for spatio-temporal context information based on a memory network according to claim 1, characterized in that The step S6 specifically comprises the following steps: Step S61: using the server to obtain N frames of point cloud from the video; Step S62: setting a memory set; obtaining a first frame of point cloud and determining its tracking target, and then putting the first frame of point cloud into the memory set as a memory frame; Step S63: obtaining the next frame of point cloud as a query frame, extracting the key-value coding pair of the query frame through the first feature extraction unit, and extracting the key-value coding pair of the memory frame through the second feature extraction unit; the key-value coding pair of the memory frame is stored in an external storage unit; Step S64: using a feature matching unit to match the key-value coding pair of the query frame with the key-value coding pair of the memory frame in the external storage unit to obtain matching features, and then decoding the matching features to obtain the target tracking prediction result of the query frame; Step S65: Put the query frame into the memory set, and repeat steps S63 to S65 until the target tracking prediction result is obtained for the last query frame, and calculate the tracking loss function L of the target tracking system network model total , and its mathematical expression is: L total = L center + L off + L z L center represents the deviation between the ground truth and the predicted value under the bird's-eye view, L off represents the regression deviation of the rotation angle, L z represents the deviation on the z-axis; Step S65: Use the server to optimize the tracking loss function, and use the Adam optimizer to iteratively update the network parameters to reduce the tracking loss function until it converges to the local optimum. At this point, the training is completed, and a three-dimensional single target tracking method based on spatiotemporal context information of a memory network is obtained.

10. A three-dimensional single-object tracking method based on a memory network for spatio-temporal context information according to claim 9, characterized in that The step S61 specifically includes the following steps: Step S611: randomly extracting 3 frames of point clouds from any video in the multiple video data sets at intervals; Step S612: Perform different affine transformations on the three frames of point clouds, including translation and shearing. For each frame of point cloud, the points 2m away from the target are sheared off, and 1024 points are randomly sampled from the cropped point cloud. As part of data enhancement, each frame of point cloud is translated.

Citation Information

Patent Citations

  • Spatio-temporal context target tracking method based on human brain memory mechanism

    CN107657627A

  • 4D target segmentation method based on point cloud space-time memory network

    CN115471651A