Class-independent target tracking method based on point cloud pre-training model
By building a category-uniform pre-training migration framework and combining geometry expert hybrid module and dynamic mask weighting module, the resource consumption and adaptability problems of category-specific learning in point cloud target tracking are solved, and high-precision and low resource consumption are achieved, which improves the generalization ability and robustness of the model.
Patent Information
- Application Number
- CN202510445830.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-29
AI Technical Summary
The existence of category-specific learning paradigms in point cloud target tracking makes computing resources consumed and difficult to adapt to new categories, and the existing methods combining 3D pre-trained models lack category uniformity, resulting in performance degradation.
Build a category-uniform pre-training migration framework, combining geometry expert hybrid module and dynamic mask weighting module, dynamically activate the geometry expert subnet and adaptively adjust the mask weight to improve tracking accuracy by lightening the adapter and freezing the pre-trained model parameters.
It realizes high-precision and low resource consumption category-independent target tracking, improves the generalization ability and unity of the model, and is robust to adapt to complex scenarios.
Smart Images

Figure CN120388215A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of point cloud object tracking, and particularly to a category-agnostic object tracking method based on a point cloud pre-trained model. Background Art
[0002] Large pre-trained models have developed rapidly, such as CLIP, LLaVA, and Llama, which have promoted the progress of core fields such as two-dimensional vision and natural language processing through their excellent representation learning capabilities. In the field of point clouds, some powerful large models have also emerged, including PointCLIP, Point-MAE, Point-BERT, CLIP2Point, and RECON, which have achieved excellent performance in tasks such as classification and generation.
[0003] The progress of large models has also promoted the development of downstream tasks. In addition to full fine-tuning, the PEFT technology has been widely adopted, enabling the model to be efficiently adapted with less resource consumption. In the field of point clouds, the research on PEFT is still relatively limited. Existing methods include IDAT, which combines DGCNN and instance-based prompt extraction to achieve geometric alignment; and DAPT, which combines dynamic adapters with prompt tuning through internal prompts to adaptively capture domain changes.
[0004] However, performing specific object tracking based on point clouds in a dynamic 3D scene is a challenging task, which requires the system to continuously and accurately locate the target relying only on sparse and irregular point cloud data. Different from RGB image tracking that utilizes rich texture and color information, the 3D lidar-based SOT faces unique difficulties: objects of different categories (such as cars, pedestrians) have huge differences in scale, motion patterns, and structural complexity. Most current methods adopt a category-specific learning paradigm, which can obtain high tracking accuracy, but this method is impractical for actual deployment because it requires huge computing resources and is difficult to adapt to newly emerging categories. In 3D single-object tracking of point clouds, the research on combining pre-trained models with PEFT is still rarely involved. As far as we know, the only relevant work is MemDisst, which is initialized through 3D pre-training. However, it lacks category-unified tracking, requires learning of the entire network, and relies on knowledge distilled from a 2D pre-trained tracker to ensure performance, rather than using the PEFT technology. This limits its efficiency and fails to fully utilize 3D pre-training knowledge. Directly attempting to use a unified model to handle all categories with existing methods will lead to a significant performance decline. Therefore, how to learn a representation that perceives geometric information and is not limited by categories without introducing manual biases has become an urgent problem to be solved. Summary of the Invention
[0005] To overcome the above deficiencies, the present invention aims to provide a class-agnostic object tracking method based on a point cloud pre-trained model, which makes full use of the combination of a 3D pre-trained model and the PEFT technology to solve the limitations of class-specific methods. By constructing a geometric expert mixture module and a dynamic mask weighting module, different geometric expert sub-networks are dynamically activated and the tracking accuracy is improved.
[0006] The present invention achieves the above object through the following scheme: A class-agnostic object tracking method based on a point cloud pre-trained model, comprising the following steps:
[0007] (1) Construct a pre-trained transfer framework with unified classes, including embedding a lightweight dual-path adapter into the Transformer layer of the pre-trained point cloud object tracking model;
[0008] (2) Construct the core module of the point cloud object tracking model, including:
[0009] (2.1) A geometric expert mixture module for dynamically activating specific geometric expert sub-networks to resolve conflicts between different geometric patterns;
[0010] (2.2) A dynamic mask weighting module for improving the tracking accuracy by adaptively learning the mask weights of the template frame and the search frame;
[0011] (3) Propose a training mechanism that, while freezing the parameters of the pre-trained point cloud object tracking model, only optimizes the newly added parameters; this training mechanism trains the point cloud object tracking model through a server, optimizes the objective function until the network converges, thereby obtaining locally optimal parameters, and finally generates a trained point cloud object tracking model;
[0012] (4) Input the point cloud data of the template frame and the search frame into the trained point cloud object tracking model, and accurately predict the target position through decoder decoding.
[0013] Preferably, the specific steps of the step (1) include:
[0014] (1.1) Given an original complete point cloud containing N0 points, apply farthest point sampling to select C points as the central points of the local region Then use K-nearest neighbors to select K nearest points relative to to form the corresponding input point cloud
[0015] (1.2) The pre-trained point cloud object tracking model uses a lightweight PointNet as the local region embedding layer and encodes it into an input embedding F0 ∈ R Nxd , where N and d represent the number of points and the feature dimension respectively;
[0016] (1.3) The encoder of the pre-trained point cloud object tracking model consists of X standard Transformer layers, which are used to encode the input feature F0;
[0017] (1.4) Add two adapters in each Transformer layer, which are parallel to the multi-head self-attention layer MHSA and the feed-forward network layer FFN respectively. For the i-th layer, the output result formula is as follows:
[0018]
[0019] where LN represents layer normalization and AD represents the adapter;
[0020] (1.5) The input point clouds of the template frame and the search frame respectively obtain their respective features through steps (1.1) and (1.2), are concatenated, and then feature extraction and matching are performed through a unified Transformer layer.
[0021] The adapter includes two paths: an adaptation path and a gating scoring path; the adaptation path includes a downsampling projection layer W dn ∈R d×r , a GeLU activation function, and an upsampling projection layer W up ∈R r×d , the gating scoring path includes a scoring weight matrix W s ∈R d×1 and a ReLU activation function. The formula for the processing process of the adapter AD is: AD(F i ) = ReLU(F i W s )⊙GeLU(F i W dn )W up ,
[0022] where d represents the number of channels of the input feature dimension, r is the dimension after downsampling, usually less than d, and ⊙ represents element-wise multiplication.
[0023] Preferably, the geometric expert mixture module in step (2.1) is added after the feed-forward network layer in each Transformer module.
[0024] Preferably, the geometric expert mixture module consists of a group of M geometric expert modules where L represents the total number of geometric experts, represents the m th -th expert in the j th -th layer.
[0025] Preferably, the geometric expert mixing module uses the Top-K gated routing mechanism as the router The router also includes an expert embedding matrix
[0026] Preferably, the output R j (Z j , K) of the geometric expert mixing module and the output E j (Z j ) of the router are expressed as:
[0027]
[0028] where represents the processing result of the m-th expert on the input data Z th in the j-th layer, th select the K values with the highest scores from the M expert scores j , Softmax is the activation function, means starting from m = 1 and accumulating until m = M.
[0029]
[0030] Preferably, the step (2.2) specifically includes the following steps:
[0030] (2.2.1) For the template frame, define a target-oriented mask where the target area is assigned 0.8 and the background area is assigned 0.2; for the search frame, define a uniform mask initialized to 0.5 representing the balance of foreground and background uncertainty;
[0031] (2.2.2) Introduce learnable weights β t and β s , respectively for the template and search regions, and adaptively adjust the mask by element-wise multiplication. The process of the entire dynamic mask weighting module is expressed as: where and represent the features of the template frame and the search frame respectively, and represent the initialized masks of the template frame and the search frame respectively, and β t and β s represent the learnable weights of the template frame and the search frame respectively, and ⊙ represents element-wise multiplication.
[0032] Preferably, in the step (3), all parameters of the pre-trained point cloud target tracking model for the original point cloud are frozen, and only the parameters of the adapter, geometric expert mixing module, dynamic mask weighting module, and decoder are learned and trained.
[0033] Preferably, the specific steps of step (4) include:
[0034] (4.1) Randomly select an initial frame from the trajectory, uniformly sample T frames from it in chronological order, stack the sampled point cloud frames on the batch dimension in chronological order to form the final input data, and the dimension of the input data is B×T×N×C, where C represents the number of feature channels and N represents the number of points in each frame of the point cloud;
[0035] (4.2) Use the trained encoder to extract features and perform feature interaction on the point cloud;
[0036] (4.3) Split the point cloud features obtained in step (4.2), input the search frame features into the decoder, and output the predicted bounding box of the target as the position information of the target.
[0037] The beneficial effects of the present invention are as follows: By applying the frozen large point cloud model to the target tracking task, controlling the number of learnable parameters to maintain the model efficiency, and designing a token propagation mechanism to enhance the transmission of time series information. Construct a geometric expert mixture module adapted to geometric characteristics to dynamically activate different geometric expert subnets and solve the conflicts between different geometric patterns; design a dynamic mask weighting module to improve the tracking accuracy by adaptively learning and adjusting the mask weights of the template frame and the search frame, effectively balancing the uncertainty of the foreground and the background. The method adopted by the present invention makes full use of the combination of the 3D pre-trained model and the PEFT technology, can accurately identify the target position in the point cloud data, solve the limitations of the category-specific method, and significantly improve the generalization ability and unity of the model, achieving state-of-the-art performance in multiple category tracking tasks, verifying its effectiveness in improving the target tracking accuracy and the robustness of processing complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic flow chart of the steps of the method of the present invention;
[0039] Figure 2 is a schematic diagram of the overall network framework of the method of the present invention;
[0040] Figure 3 is a schematic diagram of the structure of the adapter of the present invention;
[0041] Figure 4 is a schematic diagram of the structure of the geometric expert mixture module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0042] The present invention will be further described below in conjunction with specific implementation examples, but the protection scope of the present invention is not limited thereto:
[0043] Embodiment: The present invention proposes a class-agnostic object tracking method based on a point cloud pre-trained model. By introducing a geometric expert mixture module and a dynamic mask weighting module, the generalization ability and tracking accuracy of the model in complex geometric scenes are significantly improved. Among them, the geometric expert mixture module solves the conflict between different geometric patterns by dynamically activating specific geometric expert sub-networks; the dynamic mask weighting module effectively balances the uncertainty of the foreground and background by adaptively adjusting the mask weights of the template frame and the search frame. This method freezes the core parameters of the pre-trained model and only optimizes the newly added modules, reducing resource consumption while achieving high-precision tracking of point cloud objects, with good generalization ability and practical value.
[0044] As Figure 1 、 Figure 2 shown, the class-agnostic object tracking method based on a point cloud pre-trained model includes the following steps:
[0045] (1) Construct a pre-trained transfer framework with unified classes, including embedding a lightweight dual-path adapter into the Transformer layer of the pre-trained point cloud object tracking model.
[0046] In this step, by transferring the pre-trained point cloud object tracking model, while retaining the pre-trained point cloud object tracking model, a dual-path adapter is added at an appropriate position to obtain a point cloud object tracking model with performance superior to full training or full fine-tuning. The specific implementation process is as follows:
[0047] (1.1) Given an original complete point cloud containing N0 points, first apply farthest point sampling to select C points as the center points of the local area Subsequently, use K-nearest neighbors to select K nearest points relative to to form the corresponding input point cloud
[0048] (1.2) Then, the pre-trained point cloud object tracking model uses a lightweight PointNet as the local area embedding layer and encodes it into an input embedding F0 ∈ R N×d , where N and d represent the number of points and the feature dimension respectively.
[0049] (1.3) Subsequently, similar to ViT, the encoder of the pre-trained point cloud object tracking model consists of 12 layers of standard Transformers for encoding the input feature F0. Specifically, each layer of the Transformer mainly consists of a multi-head self-attention (MHSA) layer, a layer normalization (LN) layer, and a feed-forward network (FFN) layer. For the i-th layer, the formula is as follows:
[0050]
[0051] (1.4) To maintain consistency with the input of the pre-trained point cloud object tracking model, we adopt a unified modeling framework. The input point clouds of the template frame and the search frame respectively obtain their respective features through steps (1.1) and (1.2), and are concatenated, and then feature extraction and matching are performed through a unified Transformer module. In fact, the most direct method to transfer the pre-trained point cloud object tracking model is to fully fine-tune it. However, we found that this method may lead to sub-optimal performance and resource-intensive training because when the model covers the knowledge learned during pre-training, it may cause the degradation of its original capabilities. Therefore, we explored the PEFT method, which enables the model to adapt to new tasks through a small number of learnable parameters while retaining pre-trained knowledge by freezing its core parameters. Specifically, as Figure 3 shown, our adapter includes two paths: an adaptation path and a gating scoring path. The former contains a downsampling projection layer W dn ∈R d×r , a GeLU activation function, and an upsampling projection layer W up ∈R r×d . The gating scoring path contains a scoring weight matrix W s ∈R d×1 and a ReLU activation function, where d represents the number of channels of the input feature dimension, and r is the dimension after downsampling, usually less than d. This path aims to calculate a dynamic scaling factor for each token to mobilize the influence of the adaptation process in a data-driven manner. Then the outputs of these two paths are multiplied element-wise. Overall, for the input feature F i , the processing process of the adapter AD can be described as:
[0052] AD(F i ) = ReLU(F i W s ) ⊙ GeLU(F i W dn )W up , where ⊙ represents element-wise multiplication.
[0053] (1.5) This dual-path design ensures that the adapter module can effectively control the contribution of the adapted features. We add two adapters in each Transformer layer, in parallel with the multi-head self-attention layer and the feed-forward network layer respectively, as follows:
[0054]
[0055] (2) Build the core module of the point cloud object tracking model, including:
[0056] (2.1) Geometric Expert Mixing Module, which is used to dynamically activate specific geometric expert sub-networks to resolve conflicts between different geometric patterns. The geometric expert mixing module is added after the feed-forward network layer in each Transformer module. As Figure 4 shown, the geometric expert mixing module consists of a group of M geometric expert mixing modules where L represents a total of L geometric experts denotes the j th -th layer's m th -th expert, and its structure is similar to that of the FFN. The routing algorithm of the geometric expert mixing module determines which experts will process the input data. We adopt the Top-K gating routing mechanism as the router here to make decisions through a learnable gating network. The router includes an expert embedding matrix for converting features into scores. Specifically, the output of the geometric expert mixing module can be expressed as:
[0057]
[0058] where denotes the processing result of the input data Z th by the m th -th expert in the j j -th layer, selects the top K values with the highest scores from the M expert scores , Softmax is the activation function, denotes the accumulation starting from m = 1 and accumulating up to m = M.
[0059] (2.2) Dynamic Mask Weighting Module, which improves the tracking accuracy by adaptively learning the mask weights of the template frame and the search frame. The specific implementation steps include:
[0060] (2.2.1) To address the spatio-temporal variation problems of objects of different categories, we propose a dynamic mask weighting mechanism with learnable parameters. For the template frame, we define a target-oriented mask where the target region is assigned 0.8 and the background region is assigned 0.2; while for the search frame, we adopt a uniform mask initialized to 0.5 representing the uncertainty balance between the foreground and the background.
[0061] (2.2.2) Then, we introduce learnable weights β t and β s , for the template and search regions respectively, and adaptively adjust the masks by element-wise multiplication. The process of the entire dynamic mask weighting module can be expressed as:
[0062]
[0063] wherein and represent the features of the template frame and the search frame respectively, and represent the initialization masks of the template frame and the search frame respectively, and β t and β s represent the learnable weights of the template frame and the search frame respectively, and ⊙ represents element-wise multiplication.
[0064] (3) A training mechanism is proposed. While freezing the parameters of the pre-trained point cloud object tracking model, only the parameters of the adapter, the geometric expert hybrid module, the dynamic mask weighting module, and the decoder are optimized. This training mechanism trains the point cloud object tracking model through a server, optimizes the objective function until the network converges, thereby obtaining local optimal parameters, and finally generates a trained point cloud object tracking model. We train the model for 160 epochs on the KITTI dataset, and set the batch size to 32. The optimizer uses Adam, and the initial learning rate is 0.001.
[0065] (4) Input the point cloud data of the template frame and the search frame into the trained point cloud object tracking model, and accurately predict the target position through decoding by the decoder.
[0066] The specific implementation steps are as follows:
[0067] (4.1) First, perform frame sampling on the given point cloud trajectory data. For the frames in the trajectory, randomly select an initial frame, and uniformly sample T frames from the trajectory in chronological order. Stack the sampled point cloud frames in the batch dimension in chronological order, and finally form a set of input data with dimensions B, T, N, C, where C and N are the number of feature channels and the number of points in each frame of the point cloud respectively.
[0068] (4.2) Use the trained point cloud object tracking model to extract features and perform feature interaction on the point cloud.
[0069] (4.3) Input the search frame features obtained by splitting the point cloud features obtained in (4.2) into the object tracking head, and output the predicted bounding box of the target as the position information of the target.
[0070] This method is uniformly trained for all categories on the KITTI validation set, and the trained model is tested on the categories of Car, Pedestrian, Van, and Cyclist respectively. The Success and Precision of the four categories reach 73.4% / 85.2%, 59.6% / 85.6%, 70.0% / 82.8%, and 74.7% / 94.0% respectively, showing very good generalization and tracking effects.
[0071] The above are specific embodiments of the present invention and the technical principles applied. If changes are made according to the concept of the present invention and the functions and effects generated thereby do not exceed the spirit covered by the specification and the drawings, they shall still fall within the protection scope of the present invention.
Claims
1. A category-agnostic object tracking method based on a point cloud pre-trained model, characterized in that It includes the following steps: (1) Construct a pre-trained transfer framework with unified categories, including embedding lightweight dual-path adapters into the Transformer layer of the pre-trained point cloud object tracking model; (2) Construct the core module of the point cloud object tracking model, including: (2.1) Geometric expert mixture module, which is used to dynamically activate specific geometric expert sub-networks to resolve conflicts between different geometric patterns; (2.2) Dynamic mask weighting module, which improves the tracking accuracy by adaptively learning the mask weights of the template frame and the search frame; (3) Propose a training mechanism that only optimizes the newly added parameters while freezing the parameters of the pre-trained point cloud object tracking model; this training mechanism trains the point cloud object tracking model through the server, optimizes the objective function until the network converges, so as to obtain locally optimal parameters, and finally generates a trained point cloud object tracking model; (4) Input the point cloud data of the template frame and the search frame into the trained point cloud object tracking model, and accurately predict the target position through decoder decoding.
2. The category-agnostic object tracking method based on the point cloud pre-trained model according to claim 1, wherein The specific steps of step (1) include: (1.1) Given an original complete point cloud containing N0 points, apply farthest point sampling to select C points as the center points of the local regions Then use K-nearest neighbors to select the K nearest points relative to to form the corresponding input point cloud (1.2) The pre-trained point cloud object tracking model uses the lightweight PointNet as the local region embedding layer and encodes it into the input embedding F0 ∈ R N×d , where N and d represent the number of points and the feature dimension respectively; (1.3) The encoder of the pre-trained point cloud object tracking model consists of X standard Transformers, which are used to encode the input feature F0; (1.4) Add two adapters in each Transformer layer, which are parallel to the multi-head self-attention layer MHSA and the feed-forward network layer FFN respectively. For the i-th layer, its output result formula is as follows: Among them, LN represents layer normalization, and AD represents adapter; (1.5) The input point clouds of the template frame and the search frame respectively obtain their respective features through steps (1.1) and (1.2), and are concatenated, and then feature extraction and matching are performed through a unified Transformer layer.
3. The category - agnostic object tracking method based on the point - cloud pre - trained model according to any one of claims 1 - 2, characterized in that, The adapter includes two paths: an adaptation path and a gating score path; the adaptation path includes a downsampling projection layer $W$ dn ∈R d×r , a GeLU activation function, and an upsampling projection layer $W$ up ∈R r×d . The gating score path includes a scoring weight matrix $W$ s ∈R d×1 and a ReLU activation function. The formula for the processing process of the adapter AD is: AD(F i ) = ReLU(F i W s ) ⊙ GeLU(F i W dn )W up , Among them, d represents the number of channels of the input feature dimension, r is the dimension after downsampling, usually less than d, and ⊙ represents element-wise multiplication.
4. The category-agnostic object tracking method based on the point cloud pre-trained model according to claim 1, characterized in that The geometric expert mixture module in step (2.1) is added after the feed-forward network layer in each Transformer module.
5. The class-agnostic object tracking method based on the point cloud pre-trained model according to claim 4, wherein The geometric expert hybrid module consists of a set of M geometric expert modules where L represents that there are a total of L geometric experts, denotes the j th -th layer's m th -th expert.
6. The class-agnostic object tracking method based on the point cloud pre-trained model according to any one of claims 4, characterized in that The geometric expert mixing module uses the Top-K gating routing mechanism as the router The router also includes an expert embedding matrix 7. The category - agnostic object tracking method based on the point cloud pre - trained model according to any one of claims 4 - 6, characterized in that, Output R of the geometric expert mixing module j (Z j , K) and the output E of the router j (Z j ) are represented as: Among them represents the processing result of the m-th expert on the input data Z in the j-th layer th in the j-th layer th by the m-th expert j of the input data Z Select the K values with the highest scores from the scores of M experts. Softmax is the activation function and means starting from m = 1 and accumulating until m = M 8. The category-agnostic object tracking method based on the point cloud pre-trained model according to claim 1, wherein The specific steps of step (2.2) include the following steps: (2.2.1) For the template frame, define a target-oriented mask where the target area is assigned 0.8 and the background area is assigned 0.2; for the search frame, define a uniform mask initialized to 0.5 indicating a balance of uncertainty between foreground and background; (2.2.2) Introduce the learnable weight β t and β s , which are used for the template and the search region respectively, and adaptively adjust the mask by element-wise multiplication. The process of the entire dynamic mask weighting module is expressed as: where and represent the features of the template frame and the search frame respectively, and represent the initialization masks of the template frame and the search frame respectively, and β t and β s represent the learnable weights of the template frame and the search frame respectively, and ⊙ represents element-wise multiplication.
9. The category-agnostic object tracking method based on the point cloud pre-trained model according to claim 1, wherein, In step (3), all the parameters of the original point cloud pre-trained point cloud object tracking model are frozen, and only the parameters of the adapter, geometric expert mixture module, dynamic mask weighting module and decoder are learned and trained.
10. The category-agnostic object tracking method based on the point cloud pre-trained model according to claim 1, wherein, The specific steps of step (4) include: (4.1) Randomly select an initial frame from the trajectory, and uniformly sample T frames from it in chronological order. The sampled point cloud frames are stacked on the batch dimension in chronological order to form the final input data. The dimension of the input data is B×T×N×C, where C represents the number of feature channels and N represents the number of points in each frame of point cloud; (4.2) Use the trained encoder to perform feature extraction and feature interaction on the point cloud; (4.3) Split the point cloud features obtained in step (4.2), input the search frame features into the decoder, and output the predicted bounding box of the target as the position information of the target.