A Method and System for Parameter-Efficient Fine-Tuning of Three-Dimensional Pre-Trained Large Models

By block encoding of three-dimensional point cloud data and using point cloud prior library and geometry-aware adapter module to enhance prompt tokens, the problem of poor performance of efficient fine-tuning technology for the existing three-dimensional pre-trained model parameters is solved, and the effect of efficient fine-tuning in downstream tasks is achieved.

CN117252986BActive Publication Date: 2025-06-17SHANGHAI ARTIFICIAL INTELLIGENCE LABORATORY CO LTD

Patent Information

Application Number
CN202311258606.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-26
Publication Date
2025-06-17
Estimated Expiration
2043-09-26

AI Technical Summary

Technical Problem

The existing three-dimensional pre-trained model parameter efficient fine-tuning technology is designed only from the perspective of prompt adjustment, ignoring the unique knowledge in the three-dimensional field and the diversity of high-efficiency fine-tuning technology, resulting in the performance not being well improved when the number of parameters is relatively large.

Method used

A three-dimensional pre-training large model parameters is proposed. By blocking and encoding the three-dimensional point cloud data, a point cloud token sequence is formed, and the point cloud prior library and geometry perception adapter module are used to combine the parameterless attention mechanism and self-attention mechanism to enhance and adjust the learning prompt token and reduce the amount of learnable parameters.

Benefits of technology

Achieve better than full fine-tuning effects on various downstream tasks, significantly reducing the use of computing resources and improving the performance of the model in three-dimensional downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117252986B_ABST
    Figure CN117252986B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of three-dimensional pre-trained models, and particularly to a method and system for efficiently fine-tuning the parameters of a three-dimensional pre-trained large model. The method includes: partitioning and encoding three-dimensional point cloud data to form a point cloud token sequence; using the 3D features in the downstream task training dataset as prior knowledge to construct a point cloud prior library; and in the encoder module of the pre-trained model, adding learnable prompt tokens before the point cloud token sequence, adopting a parameter-free attention mechanism, and combining the prior knowledge in the point cloud prior library to enhance the learnable prompt tokens; clustering the enhanced prompt tokens through a geometric perception adapter, and performing local feature interaction through a self-attention mechanism to obtain adjusted tokens; inputting the adjusted tokens into the downstream task head to obtain a prediction output. The method of this application greatly reduces the number of learnable parameters while effectively improving the performance of the pre-trained model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of three-dimensional pre-trained models, and in particular to a method and system for efficiently fine-tuning parameters of a large three-dimensional pre-trained model. Background Art

[0002] The popularity of large 3D pre-trained models has overturned the traditional learning methods of downstream tasks in the 3D field. By performing unsupervised pre-training on large-scale 3D datasets and migrating to downstream tasks, considerable performance has been achieved.

[0003] The current mainstream method is to retrain all the parameters of the large model on the original basis when adapting to downstream tasks, which will result in expensive consumption of computing resources. In the field of two-dimensional image and language processing, some parameter-efficient fine-tuning techniques have been proposed, which can minimize the adaptation cost of downstream tasks and achieve good performance by fine-tuning some parameters.

[0004] The existing efficient fine-tuning techniques for parameters proposed for 3D pre-trained models are designed only from the perspective of prompt adjustment, ignoring the unique knowledge in the 3D field and the diversity of efficient fine-tuning techniques for parameters. When the number of parameters is relatively large, the performance is not well improved. Therefore, a proprietary efficient fine-tuning framework for large 3D pre-trained models is still under development. Summary of the invention

[0005] The embodiments of the present application provide a method and system for efficiently fine-tuning parameters of a three-dimensional pre-trained large model, which effectively improves the performance of the pre-trained model while greatly reducing the number of learnable parameters.

[0006] To solve the above technical problems, in the first aspect, an embodiment of the present application provides an efficient fine-tuning method for parameters of a three-dimensional pre-trained large model, which includes the following steps: first, the three-dimensional point cloud data is divided into blocks and encoded to form a point cloud token sequence; then, the 3D features in the downstream task training data set are used as prior knowledge to build a point cloud prior library; and the point cloud token sequence is used as the input of the pre-trained model. In the encoder module of the pre-trained model, before the learnable prompt token is added to the point cloud token sequence, a parameter-free attention mechanism is adopted, and the learnable prompt token is enhanced in combination with the prior knowledge in the point cloud prior library to obtain an enhanced prompt token; next, the enhanced prompt token is clustered through a geometric-aware adapter, and local feature interaction is performed through a self-attention mechanism to obtain an adjusted token; finally, the adjusted token is input into the downstream task head to obtain a predicted output.

[0007] In some exemplary embodiments, the three-dimensional point cloud data is segmented and encoded to form a point cloud token sequence, including the following steps: First, sample from the original point cloud to obtain a sub-point cloud; then, based on the spatial position information, divide the sub-point cloud into multiple point cloud blocks; finally, encode each point cloud block to obtain the representation of the point cloud block, and form a point cloud token sequence.

[0008] In some exemplary embodiments, a part of the points is sampled from the original point cloud by random sampling or farthest point sampling to obtain a sub-point cloud.

[0009] In some exemplary embodiments, each point cloud block includes a fixed number of points.

[0010] In some exemplary embodiments, the pre-trained model includes multiple encoder modules, and each encoder module includes a self-attention layer, a feed-forward network, and a geometric perception adapter connected in sequence; the self-attention layer and the feed-forward network are used to explore the global shape information and long-range dependencies in the point cloud, and enhance the ability of feature representation; the geometric perception adapter is used to complement the long-range dependencies of the self-attention layer, converge local geometric information, and capture fine-grained three-dimensional structures.

[0011] In some exemplary embodiments, the encoder module is composed of a transformer network structure.

[0012] In some exemplary embodiments, the enhanced prompt tokens are clustered by farthest point sampling and K-nearest neighbors.

[0013] In some exemplary embodiments, the downstream task training dataset is obtained by dividing the three-dimensional downstream dataset; the three-dimensional downstream dataset is a dataset for downstream three-dimensional scene tasks.

[0014] In a second aspect, the embodiments of the present application further provide a three-dimensional pre-trained large model parameter-efficient fine-tuning system, including a pre-trained model, which includes a three-dimensional token embedding module, a point cloud prior prompt module, a geometric perception adapter module, and a downstream task head connected in sequence; wherein, the three-dimensional token embedding module is used to segment and encode the three-dimensional point cloud data to form a point cloud token sequence; the point cloud prior prompt module is used to use the 3D features in the downstream task training dataset as prior knowledge to construct a point cloud prior library; and use the point cloud token sequence as the input of the pre-trained model. In the encoder module of the pre-trained model, add learnable prompt tokens before the point cloud token sequence, adopt a parameter-free attention mechanism, and combine the prior knowledge in the point cloud prior library to enhance the learnable prompt tokens to obtain enhanced prompt tokens; the geometric perception adapter module is used to cluster the enhanced prompt tokens, and after local feature interaction through the self-attention mechanism, obtain adjusted tokens; the downstream task head is used to obtain a prediction output according to the adjusted tokens.

[0015] In some exemplary embodiments, the above three-dimensional pre-trained large model parameter-efficient fine-tuning system further includes: a data processing module and a verification and application module; an output end of the data processing module is connected to an input end of the pre-trained model; an input end of the verification and application module is connected to an output end of the pre-trained model; the data processing module is configured to obtain a three-dimensional downstream data set and split the three-dimensional downstream data set to respectively obtain a downstream task training data set and a downstream task test data set; the verification and application module is configured to verify a prediction result output by the pre-trained model and apply the prediction result to a three-dimensional downstream task.

[0016] The technical solution provided by the embodiment of the present application has at least the following advantages:

[0017] The embodiment of the present application provides a three-dimensional pre-trained large model parameter-efficient fine-tuning method and system. The method includes the following steps: First, divide and encode three-dimensional point cloud data to form a point cloud token sequence; then, use 3D features in the downstream task training data set as prior knowledge to construct a point cloud prior library; and use the point cloud token sequence as the input of the pre-trained model. In the encoder module of the pre-trained model, before adding the learnable prompt tokens to the front of the point cloud token sequence, use a parameter-free attention mechanism and combine the prior knowledge in the point cloud prior library to enhance the learnable prompt tokens to obtain enhanced prompt tokens; next, cluster the enhanced prompt tokens through a geometric perception adapter, and perform local feature interaction through a self-attention mechanism to obtain adjusted tokens; finally, input the adjusted tokens into the downstream task head to obtain a prediction output.

[0018] The present application proposes a three-dimensional pre-trained large model parameter-efficient fine-tuning method and system, which uses extremely few learnable parameters to fine-tune the point cloud pre-trained model, so as to achieve better results than full fine-tuning on various downstream tasks. The present application achieves efficient fine-tuning by exploring how to effectively integrate downstream three-dimensional semantics into the pre-trained model. In view of the sparse and irregular characteristics of point clouds, the present application proposes a point cloud prior prompt module. Before each transformer block, the present application will add a set of learnable prompt tokens before the input point cloud features to inject downstream knowledge into the pre-trained model. In addition, the present application also proposes a geometric perception adapter module, which is inserted after the pre-trained self-attention layer and the feed-forward network, complements the long-range dependence of the pre-trained attention layer, aggregates local geometric information and captures fine-grained three-dimensional structures. This method freezes most of the pre-trained parameters, combines the specific knowledge of the three-dimensional domain and local feature interaction, and only fine-tunes the newly added modules and task heads on downstream tasks, proving to be more efficient and performant. Description of the Drawings

[0019] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Unless otherwise stated, the figures in the drawings do not constitute a scale limitation.

[0020] Figure 1 It is a schematic flow chart of a method for efficiently fine-tuning the parameters of a three-dimensional pre-trained large model provided by an embodiment of the present application;

[0021] Figure 2 It is a module structure diagram of a system for efficiently fine-tuning the parameters of a three-dimensional pre-trained large model provided by an embodiment of the present application;

[0022] Figure 3 It is an architecture flow chart of a method for efficiently fine-tuning the parameters of a three-dimensional pre-trained large model provided by an embodiment of the present application;

[0023] Figure 4 It is a schematic flow chart of the framework of a pre-trained model provided by an embodiment of the present application;

[0024] Figure 5 It is a schematic diagram of a geometric perception adapter module provided by an embodiment of the present application;

[0025] Figure 6 It is a schematic diagram of a point cloud prior prompt module provided by an embodiment of the present application;

[0026] Figure 7 It is a schematic diagram of the model quantitative experimental results provided by an embodiment of the present application;

[0027] Figure 8 It is a schematic diagram of the model qualitative experimental results provided by an embodiment of the present application. Detailed implementation manners

[0028] As can be seen from the background art, the existing parameter-efficient fine-tuning techniques proposed for three-dimensional pre-trained models are only designed from the perspective of prompt adjustment, ignoring the unique knowledge in the three-dimensional field and the diversity of parameter-efficient fine-tuning techniques. The performance has not been improved well when the number of parameters is relatively large.

[0029] In the field of two-dimensional image and language processing, some parameters are fine-tuned to minimize the adaptation cost of downstream tasks and achieve good performance. However, proprietary efficient fine-tuning methods and systems for large three-dimensional pre-trained models are still under development. In the field of computer vision, the purpose of large three-dimensional pre-trained models is to use a large number of three-dimensional objects without manual annotations for pre-training and to achieve knowledge transfer for downstream tasks through fine-tuning, thus overcoming the challenge of scarcity of three-dimensional data. This method can be widely used in fields such as autonomous driving and robot navigation. At present, pre-training methods in the three-dimensional field can be divided into two categories: methods based on contrastive learning and methods based on point cloud reconstruction.

[0030] (1) Pre-training methods based on contrastive learning: Due to the lack of large-scale high-quality annotated datasets and the lack of a unified backbone network architecture in the 3D field, the mainstream method is to train the target data from scratch in the 3D scene understanding task. This results in the need to redesign the network structure and train it for different tasks. Pre-training methods based on contrastive learning use data augmentation (rotation, cropping, flipping) and other methods to generate multiple different views for an instance data. The main network architecture uses a common deformer network and uses an attention mechanism to distinguish different views of a single instance from views of other instances. This method can effectively use the loss function to shorten the distance between different views of the same instance and increase the distance between different instances, effectively improving the general performance of the network architecture on multiple tasks. However, pre-training methods based on contrastive learning are likely to achieve overfitting in the absence of pre-training data and fail to bring about appropriate generalization performance.

[0031] (2) Pre-training methods based on point cloud reconstruction: Pre-training methods based on point cloud reconstruction are the mainstream direction of 3D pre-training. Similar to image and language features, partial structures of point cloud data can present local features, while the complete set of elements can constitute global features. Based on this, the pre-training method will first divide the input point cloud into irregular point cloud blocks, randomly mask the point cloud blocks at a high ratio to reduce data redundancy, and use an autoencoder to reconstruct explicit features (such as pixels) or implicit features (such as discrete tags) corresponding to the original masked content. This reconstruction task enables the autoencoder to learn high-level latent features from the unmasked data content. The autoencoder backbone adopts an asymmetric encoder-decoder structure, both of which are composed of a deformer network structure. The encoder processes the unmasked point cloud blocks, and then inputs the processed results and mask tags into a lightweight decoder with a simple prediction head to reconstruct the masked point cloud. Compared with pre-training methods based on contrastive learning, point cloud reconstruction methods are less dependent on data and can achieve better performance and generalization ability when the amount of data is small.

[0032] Currently, the mainstream method for adapting large models to downstream tasks is still full fine-tuning, which is very computationally intensive. Therefore, the parameter-efficient fine-tuning method has been proposed by scholars to address this challenge by freezing the trained weights and introducing new trainable modules. Parameter-efficient fine-tuning techniques include adapters, prompt tuning, low-rank adaptation (LoRA), bias tuning, and side tuning. Specifically, adapter tuning inserts additional bottleneck-shaped neural networks within the encoder layers of the pre-trained model to learn task-specific representations; prompt tuning promotes task adaptation by adding natural language prompts or learnable prompt tokens before the input; the LoRA technique uses low-rank factorization to learn the adaptation matrix in each block; bias tuning achieves performance comparable to full fine-tuning only by making the bias terms of the model learnable; side tuning only adjusts the lightweight module network parallel to the pre-trained network.

[0033] In the prior art, although parameter-efficient fine-tuning techniques have been proposed for 3D pre-trained models, they are only designed from the perspective of prompt tuning, ignoring the knowledge unique to the 3D domain and the diversity of parameter-efficient fine-tuning techniques. As a result, the performance has not been improved well when the number of parameters is relatively large. There are mainly two problems with the existing parameter-efficient fine-tuning techniques for 3D pre-trained models: First, they are all designed from a single perspective, such as starting from prompt tuning or adapters. Second, they ignore the knowledge unique to the 3D domain and the prior knowledge of the special geometric structure of 3D, which is very important for improving the performance of the model in 3D downstream tasks.

[0034] To solve the above technical problems, the present application provides a method for efficiently fine-tuning the parameters of a three-dimensional pre-trained large model, including the following steps: First, the three-dimensional point cloud data is segmented and encoded to form a point cloud token sequence; Then, using the 3D features in the downstream task training dataset as prior knowledge, a point cloud prior library is constructed; And using the point cloud token sequence as the input of the pre-trained model, in the encoder module of the pre-trained model, a learnable prompt token is added before the point cloud token sequence, a parameter-free attention mechanism is adopted, and the prior knowledge in the point cloud prior library is combined to enhance the learnable prompt token to obtain an enhanced prompt token; Next, the enhanced prompt token is clustered by a geometric perception adapter, and local feature interaction is performed through a self-attention mechanism to obtain an adjusted token; Finally, the adjusted token is input into the downstream task head to obtain a prediction output. Therefore, the parameter-efficient fine-tuning framework proposed in the present application starts from multiple technical designs at the same time, introduces pre-trained knowledge into downstream task training, pays attention to local structure feature interaction while retaining global interaction, further reduces the number of learnable parameters and greatly improves performance.

[0035] The following will elaborate on each embodiment of the present application with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present application, many technical details are presented to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.

[0036] See Figure 1 , the embodiments of the present application provide a method for efficiently fine-tuning the parameters of a three-dimensional pre-trained large model, including the following steps:

[0037] Step S1: Segment and encode the three-dimensional point cloud data to form a point cloud token sequence.

[0038] Step S2: Using the 3D features in the downstream task training dataset as prior knowledge, construct a point cloud prior library; and using the point cloud token sequence as the input of the pre-trained model, in the encoder module of the pre-trained model, add a learnable prompt token before the point cloud token sequence, adopt a parameter-free attention mechanism, and combine the prior knowledge in the point cloud prior library to enhance the learnable prompt token to obtain an enhanced prompt token.

[0039] Step S3: Cluster the enhanced prompt token through a geometric perception adapter, and perform local feature interaction through a self-attention mechanism to obtain an adjusted token.

[0040] Step S4: Input the adjusted token into the downstream task head to obtain a prediction output.

[0041] This application proposes a parameter-efficient fine-tuning method and system for 3D pre-trained models, exploring how to effectively integrate downstream 3D semantics into the pre-trained model to achieve efficient fine-tuning. Aiming at the sparse and irregular characteristics of point clouds, this application proposes a point cloud prior prompt module. Before each transformer block, this application adds a set of learnable prompt tokens before the input point cloud features to inject downstream knowledge into the pre-trained model. In addition, this application also proposes a geometric-aware adapter module, which is inserted after the pre-trained self-attention layer and feed-forward network, complements the long-range dependencies of the pre-trained attention layer, aggregates local geometric information, and captures fine-grained 3D structures.

[0042] In some embodiments, the 3D point cloud data is segmented and encoded to form a point cloud token sequence, including the following steps: First, sample from the original point cloud to obtain a sub-point cloud; then, based on the spatial position information, divide the sub-point cloud into multiple point cloud blocks; finally, encode each point cloud block to obtain the representation of the point cloud block and form a point cloud token sequence.

[0043] In some embodiments, a part of the points is sampled from the original point cloud by random sampling or farthest point sampling to obtain a sub-point cloud.

[0044] In some embodiments, each point cloud block includes a fixed number of points.

[0045] In some embodiments, the pre-trained model includes multiple encoder modules (also referred to as encoder blocks), and each encoder module includes a self-attention layer (also referred to as a self-attention layer), a feed-forward network, and a geometric-aware adapter connected in sequence; the self-attention layer and the feed-forward network are used to explore the global shape information and long-range dependencies in the point cloud and enhance the ability of feature representation; the geometric-aware adapter is used to complement the long-range dependencies of the self-attention layer, aggregate local geometric information, and capture fine-grained 3D structures.

[0046] In some embodiments, the encoder module is composed of a transformer network structure.

[0047] In some embodiments, the enhanced prompt tokens are clustered by farthest point sampling and K-nearest neighbors.

[0048] In some embodiments, the downstream task training dataset is obtained by dividing the 3D downstream dataset; the 3D downstream dataset is a dataset for downstream 3D scene tasks.

[0049] See Figure 2, the embodiment of the present application also provides a parameter-efficient fine-tuning system for a three-dimensional pre-trained large model, including a pre-trained model, which includes a three-dimensional token embedding module 101, a point cloud prior hint module 102, a geometric perception adapter module 103, and a downstream task head 104 connected in sequence; among them, the three-dimensional token embedding module 101 is used to block and encode three-dimensional point cloud data to form a point cloud token sequence; the point cloud prior hint module 102 is used to use the 3D features in the downstream task training dataset as prior knowledge to construct a point cloud prior library; and use the point cloud token sequence as the input of the pre-trained model. In the encoder module of the pre-trained model, before adding the learnable hint token to the point cloud token sequence, a parameter-free attention mechanism is adopted, and the prior knowledge in the point cloud prior library is combined to enhance the learnable hint token to obtain an enhanced hint token; the geometric perception adapter module 103 is used to cluster the enhanced hint token, and after local feature interaction through the self-attention mechanism, obtain an adjusted token; the downstream task head 104 is used to obtain a prediction output according to the adjusted token.

[0050] See Figure 3 , in some embodiments, the above-mentioned parameter-efficient fine-tuning system for a three-dimensional pre-trained large model further includes: a data processing module and a verification and application module; the output end of the data processing module is connected to the input end of the pre-trained model; the input end of the verification and application module is connected to the output end of the pre-trained model; the data processing module is used to obtain a three-dimensional downstream dataset and split the three-dimensional downstream dataset to obtain a downstream task training dataset and a downstream task test dataset respectively; the verification and application module is used to verify the prediction result output by the pre-trained model and apply the prediction result to a three-dimensional downstream task.

[0051] The parameter-efficient fine-tuning system for a three-dimensional pre-trained large model provided by the present application, as Figure 3 shown, the process of parameter-efficient fine-tuning generally includes three stages: data processing, model training, and verification and application.

[0052] In the data processing stage, the present application selects a dataset specifically for downstream three-dimensional scene tasks. The present application adopts the original training set and test set division of the downstream task dataset.

[0053] In the model training stage, the core part of the present application is proposed: a parameter-efficient fine-tuning framework for a three-dimensional pre-trained large model. The overall framework consists of four parts: a three-dimensional token embedding module, a point cloud prior hint module, a geometric perception adapter, and a downstream task head.

[0054] First, a detailed introduction to the 3D token embedding module is given. The 3D token embedding part is used to divide the 3D point cloud data into chunks and encode them into point cloud token embeddings as the input for the subsequent pre-trained model. Specifically, for the input point cloud, the present application first samples some points from the original point cloud to obtain a sub-point cloud. The sampling strategies include methods such as random sampling and farthest point sampling. Further, based on the spatial position information, the sampled sub-point cloud is further divided into smaller chunks, and each chunk contains a fixed number of points, which is called a point cloud chunk. Finally, the PointNet network is used to encode each point cloud chunk to obtain the representation of the point cloud chunk, that is, the token embedding of the 3D point cloud, as Figure 4 shown.

[0055] Then, a detailed introduction to the point cloud prior prompting module is given. In this part, the present application first uses the 3D features in the downstream task training dataset as prior knowledge to construct a point cloud prior bank. During the fine-tuning process, in each transformer block, as Figure 4 and Figure 6 shown, the present application adds a set of learnable prompt tokens to the front of the input point cloud token sequence. The present application uses a parameter-free attention mechanism to enhance the prompt tokens by combining the prior knowledge in the point cloud prior bank, so that the prompt tokens obtain 3D prior knowledge of a specific domain. This module endows the pre-trained model with a prior understanding of the 3D scene, thereby enhancing the point cloud representation learning ability.

[0056] Secondly, a detailed introduction to the geometric perception adapter module is given. First, the global shape information and long-range dependencies in the point cloud are explored through the pre-trained self-attention layer and feed-forward network to enhance the ability of feature representation. The design of the geometric perception adapter is inserted after the pre-training module, as Figure 5 shown. The input point cloud features are clustered by farthest point sampling and K-nearest neighbors, and the features within the class are locally interacted through the self-attention mechanism, and finally propagated to other point cloud features. Through this design, the local geometric information is effectively aggregated, and the fine-grained structure of the point cloud is learned.

[0057] Then, the adjusted tokens are input into the downstream task head. After passing through multiple encoder blocks with geometric perception adapters, the input is fed into the lightweight downstream task head to obtain the prediction output. During the entire fine-tuning process, only the point cloud prior prompting module, the geometric perception adapter, and the downstream task head can be fine-tuned, and the rest are frozen to maintain the pre-trained weights.

[0058] Finally, verification and application are the last stages of the technical solution process of this application. The feasibility and effectiveness of the proposed method will be verified through actual experiments. In the experimental verification stage, this application selects other 3D pre-training methods for quantitative and qualitative experiments. In the quantitative experiment part, the model of this application achieves the best results in most metrics; in the qualitative experiment, through the visualization of the results, this application discovers that the method of this application shows better attention to local geometric structures in some data.

[0059] This application proposes a parameter-efficient fine-tuning method and system specifically designed for 3D pre-trained models. Compared with full fine-tuning, this application has competitive performance and significantly reduces the use of computing resources. Pre-trained models are usually pre-trained on a large amount of data and then fully fine-tuned on specific tasks to adapt to specific applications or datasets, usually involving parameter updates across the entire model. This application adopts a more efficient method, only fine-tuning certain parts of the model, greatly reducing the computational burden. Secondly, this application designs a geometry-aware adapter for extracting fine-grained local geometric structures. At the same time, this application also designs a point prior prompt module equipped with a parameter-free attention mechanism to utilize domain-specific knowledge to promote fine-tuning performance in downstream tasks.

[0060] Compared with the prior art, the advantages of the method provided by this application are as follows: This application starts from multiple parameter-efficient fine-tuning techniques, introduces pre-trained knowledge into downstream task training, pays attention to feature interaction of local structures while retaining global interaction, further reduces the number of learnable parameters, and greatly improves performance. This technological advancement can further promote the development of fields and applications such as autonomous driving and embodied intelligence.

[0061] The method proposed in this application is used for quantitative experiments on a major dataset, such as Figure 7 shown. The results show that the model achieves better results than other current related technical methods. In addition, the method proposed in this application is used for qualitative experiments on the dataset, such as Figure 8 shown. Through the visualization of the results, this application discovers that the method of this application reaches the same level of results as other methods, and in some data, the method of this application shows better attention to local geometric structures, with red representing high attention.

[0062] With the above technical solutions, the embodiments of the present application provide a method and system for efficiently fine-tuning the parameters of a three-dimensional pre-trained large model. The method includes the following steps: First, the three-dimensional point cloud data is segmented and encoded to form a point cloud token sequence; then, using the 3D features in the downstream task training dataset as prior knowledge, a point cloud prior library is constructed; and with the point cloud token sequence as the input of the pre-trained model, in the encoder module of the pre-trained model, before adding the learnable prompt token to the front of the point cloud token sequence, a parameter-free attention mechanism is adopted, and the prior knowledge in the point cloud prior library is combined to enhance the learnable prompt token to obtain an enhanced prompt token; next, the enhanced prompt token is clustered through a geometric perception adapter, and after local feature interaction through the self-attention mechanism, an adjusted token is obtained; finally, the adjusted token is input into the downstream task head to obtain a prediction output.

[0063] The present application proposes a parameter-efficient fine-tuning framework for a three-dimensional pre-trained large model, which uses extremely few learnable parameters to fine-tune the point cloud pre-trained model to achieve better performance than full fine-tuning on various downstream tasks. The present application achieves efficient fine-tuning by exploring how to effectively integrate downstream three-dimensional semantics into the pre-trained model. In view of the sparse and irregular characteristics of the point cloud, the present application proposes a point cloud prior prompt module. Before each transformer block, the present application adds a set of learnable prompt tokens before the input point cloud features to inject downstream knowledge into the pre-trained model. In addition, the present application also proposes a geometric perception adapter module, which is inserted after the pre-trained self-attention layer and the feed-forward network, complements the long-range dependence of the pre-trained attention layer, aggregates local geometric information, and captures fine-grained three-dimensional structures. This method freezes most of the pre-trained parameters, combines the specific knowledge of the three-dimensional field and local feature interaction, and only fine-tunes the newly added modules and task heads on the downstream tasks, proving to be more efficient and performant.

[0064] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present application. Any person skilled in the art can make their own changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.

Claims

1. A method for parameter - efficient fine - tuning of a three - dimensional pre - trained large model, characterized in that, Including: Partitioning and encoding the three-dimensional point cloud data to form a point cloud token sequence; Using the 3D features in the downstream task training dataset as prior knowledge to construct a point cloud prior library; Taking the point cloud token sequence as the input of the pre-trained model. In the encoder module of the pre-trained model, adding learnable prompt tokens before the point cloud token sequence, adopting a parameter-free attention mechanism, and enhancing the learnable prompt tokens by combining the prior knowledge in the point cloud prior library to obtain enhanced prompt tokens; Clustering the enhanced prompt tokens through a geometric perception adapter, and performing local feature interaction through a self-attention mechanism to obtain adjusted tokens; the geometric perception adapter is used to cluster the input point cloud features through farthest point sampling and K-nearest neighbors, perform local interaction on the features within the class through a self-attention mechanism, and finally propagate to other point cloud features; Inputting the adjusted tokens into the downstream task head to obtain a prediction output.

2. The method for parameter - efficient fine - tuning of a three - dimensional pre - trained large model according to claim 1, characterized in that, The partitioning and encoding of the three-dimensional point cloud data to form a point cloud token sequence includes the following steps: Sampling from the original point cloud to obtain a sub-point cloud; Based on the spatial position information, dividing the sub-point cloud into multiple point cloud blocks; Encoding each point cloud block to obtain the representation of the point cloud block, forming a point cloud token sequence.

3. The method for parameter - efficient fine - tuning of a three - dimensional pre - trained large model according to claim 2, characterized in that, Each point cloud block includes a fixed number of points.

4. The method for parameter - efficient fine - tuning of a three - dimensional pre - trained large model according to claim 1, characterized in that, The pre-trained model includes multiple encoder modules, and each encoder module includes a self-attention layer, a feed-forward network, and a geometric perception adapter connected in sequence; The self-attention layer and the feed-forward network are used to explore the global shape information and long-range dependencies in the point cloud, enhancing the ability of feature representation; The geometric perception adapter is used to complement the long-range dependencies of the self-attention layer, converge local geometric information, and capture fine-grained three-dimensional structures.

5. The method for parameter - efficient fine - tuning of a three - dimensional pre - trained large model according to claim 1, characterized in that, The encoder module is composed of a transformer network structure.

6. The method for parameter - efficient fine - tuning of a three - dimensional pre - trained large model according to claim 1, characterized in that, The downstream task training dataset is obtained by dividing the three-dimensional downstream dataset; the three-dimensional downstream dataset is the dataset for the downstream three-dimensional scene task.

7. A system for parameter - efficient fine - tuning of a three - dimensional pre - trained large model, characterized in that, Including a pre-trained model, the pre-trained model includes a three-dimensional token embedding module, a point cloud prior prompt module, a geometric perception adapter module, and a downstream task head connected in sequence; where The three-dimensional token embedding module is used to partition and encode the three-dimensional point cloud data to form a point cloud token sequence; The point cloud prior prompt module is used to use the 3D features in the downstream task training dataset as prior knowledge to construct a point cloud prior library; and taking the point cloud token sequence as the input of the pre-trained model, adding learnable prompt tokens before the point cloud token sequence in the encoder module of the pre-trained model, adopting a parameter-free attention mechanism, and enhancing the learnable prompt tokens by combining the prior knowledge in the point cloud prior library to obtain enhanced prompt tokens; The geometric perception adapter module is used to cluster the enhanced prompt tokens, and after local feature interaction through the self-attention mechanism, obtain the adjusted tokens; the geometric perception adapter module is used to cluster the input point cloud features through farthest point sampling and K-nearest neighbors, and perform local interaction on the features within the class through the self-attention mechanism, and finally propagate them to other point cloud features; The downstream task head is used to obtain the prediction output according to the adjusted tokens.

8. The system for parameter - efficient fine - tuning of a three - dimensional pre - trained large model according to claim 7, characterized in that, It further includes a data processing module and a verification and application module; the output end of the data processing module is connected to the input end of the pre-trained model; the input end of the verification and application module is connected to the output end of the pre-trained model; The data processing module is used to obtain a three-dimensional downstream data set and split the three-dimensional downstream data set to obtain a downstream task training data set and a downstream task test data set respectively; The verification and application module is used to verify the prediction results output by the pre-trained model and apply the prediction results to three-dimensional downstream tasks.

Citation Information

Patent Citations

  • Three-dimensional point cloud data retrieval method and device based on category fusion and computer equipment

    CN114297237A

  • Image point cloud fusion three-dimensional target detection method based on cross attention mechanism

    CN115019043A

Cited By

  • Integrated point cloud denoising and analysis method based on point-level prompt

    CN121391650A