Spatial prior calibration and feature adaptation method and system for sparse perception architecture
Patent Information
- Application Number
- CN202611131434.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-29
AI Technical Summary
预训练阶段的位置嵌入与下游任务的分辨率不匹配时,直接插值虽然可以调整形状,但会导致密集的位置嵌入参数全部变为可训练状态,参数量从大幅增加至
(其中
为patch数量,
为通道维度),严重违背了参数高效微调的设计初衷
[0041] The beneficial effects of this application are as follows: The spatial prior calibration and feature adaptation method and system for sparse perception architectures in this application, by freezing the interpolated position embedding matrix into an untrainable state, completely preserves the spatial layout-sensitive representations (such as relative object positions and scale relationships) learned in the pre-training stage, avoiding the destruction of spatial priors caused by relearning in downstream tasks, and effectively reducing the risk of overfitting. It employs learnable 1×1 one-dimensional convolutions for calibration along the channel dimension, reducing the number of trainable parameters from those related to traditional resolution. Reduced to being only related to the channel dimension
This completely decouples the number of parameters from the input resolution. The same set of channel calibration parameters can be uniformly adapted to mixed configurations of multiple cameras, such as high-resolution front-view and low-resolution surround-view, without requiring separate storage or training of parameters for each resolution. The image features output after spatial prior protection and channel response calibration possess precise geometric structure, ensuring that the Deformable Attention mechanism in sparse sensing methods such as Sparse4D accurately locates and provides rich semantic features during four-dimensional keypoint sampling, thereby improving the overall performance of 3D object detection.
Smart Images

Figure CN122637386B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving perception technology, specifically to a spatial prior calibration and feature adaptation method and system for sparse perception architectures. Background Technology
[0002] In multi-view 3D object detection systems, the visual Transformer backbone network is typically pre-trained on large-scale 2D image datasets (such as ImageNet, with an input resolution of 224×224) to learn a general visual representation. However, in downstream autonomous driving 3D detection tasks, due to the need to meet the requirements of accurate detection of long-distance objects and recognition of small objects at close range, the input image resolution is usually much higher than that in the pre-training stage (common configurations are 320×800, 640×1600 or even higher), and the camera resolution configurations vary for different vehicle models and sensor solutions.
[0003] This resolution difference between the pre-training stage and the downstream task stage leads to the following technical problems:
[0004] (1) Mismatch in positional embeddings. Positional embeddings in the Vision Transformer (vit) encode the spatial location information of image patches and are crucial for the model to understand the spatial structure of images. When the positional embeddings in the pre-training stage do not match the resolution of the downstream task, direct interpolation, although it can adjust the shape, will cause all the dense positional embedding parameters to become trainable, increasing the number of parameters from... Significantly increased to (in For the number of patches, (For the channel dimension), which seriously violates the original design intention of efficient parameter fine-tuning.
[0005] (2) Destruction of spatial prior knowledge. During the pre-training process, the position embedding learns a sensitive representation of spatial layout (such as the relative position of objects, scale relationship, etc.). If the downstream task is made to completely relearn the position embedding, it will destroy the spatial prior knowledge accumulated in the pre-training process and reduce the transfer effect of the model.
[0006] (3) Difficulty in adapting to multiple resolutions. Autonomous vehicles often adopt a mixed resolution strategy (such as high resolution for the front-view camera and low resolution for the surround-view camera). Traditional methods are difficult to efficiently adapt to multiple different resolution inputs within a unified framework.
[0007] (4) Feature quality dependence of sparse sensing architecture. Sparse sensing methods such as Sparse4D perform sparse sampling on image feature maps through the DeformableAttention mechanism, and their detection performance is highly dependent on the image feature quality output by the backbone network. Summary of the Invention
[0008] In view of this, the purpose of this application is to provide a method and system for spatial prior calibration and feature adaptation for sparse perceptual architectures, so as to solve the problems in the background art.
[0009] To achieve the above objectives, this application adopts the following technical solution:
[0010] This application presents a spatial prior calibration and feature adaptation method for sparse perception architectures, used for feature adaptation in autonomous driving perception tasks, including:
[0011] The first position embedding matrix is extracted from the pre-trained visual transformer, and the image patch embedding vector from the autonomous driving perception backbone network is obtained. The visual transformer is pre-trained using visual training samples at a first resolution. The first position embedding matrix includes position embedding vectors for multiple positions. The image patch embedding vector from the autonomous driving perception backbone network comes from an environmental image at a second resolution, which is greater than the first resolution.
[0012] The first position embedding matrix is transformed into a second position embedding matrix aligned with the second resolution based on two-dimensional reshaping and bilinear interpolation.
[0013] The second position embedding matrix is frozen, and the second position embedding matrix after freezing parameters is calibrated to obtain a calibrated new position embedding matrix, wherein the calibrated new position embedding matrix includes new position embedding vectors for multiple positions.
[0014] The image patch embedding vector and the new position embedding vector are added and fused element by element to obtain the adaptation features, which are then input into the encoder of the visual transformer to obtain the adapted image features.
[0015] In one embodiment of this application, converting the first position embedding matrix into a second position embedding matrix aligned with the second resolution based on two-dimensional reshaping and bilinear interpolation includes:
[0016] Extract the classification token position embedding vector from the first position embedding matrix to obtain the intermediate position embedding matrix;
[0017] The intermediate position embedding matrix is converted into a two-dimensional tensor, and bilinear interpolation is performed on the two-dimensional tensor to obtain an interpolated position embedding matrix aligned with the second resolution.
[0018] The extracted classification token position embedding vector is re-attached to the front end of the interpolated position embedding matrix to obtain the second position embedding matrix.
[0019] In one embodiment of this application, bilinear interpolation is performed on the two-dimensional tensor to obtain an interpolated position embedding matrix aligned with a second resolution, including:
[0020] The two-dimensional tensor is stretched from a first resolution to a second resolution to obtain multiple stretched source grids and multiple interpolation points between the source grids;
[0021] For any interpolation point, take the four nearest vertices of the source mesh as the target vertex and obtain the position embedding vector of the target vertex;
[0022] The position embedding vectors of the four target vertices are weighted and fused to obtain the position embedding vector of the interpolation point.
[0023] A stretched two-dimensional tensor is constructed based on the position embedding vectors of multiple source grids and the position embedding vectors of multiple interpolation points. The stretched two-dimensional tensor is then flattened back into a sequence form to obtain an interpolated position embedding matrix aligned with the second resolution.
[0024] In one embodiment of this application, the second position embedding matrix after freezing parameters is calibrated to obtain a calibrated new position embedding matrix, including:
[0025] The second position embedding matrix is convolved along the channel dimension using a learnable 1×1 one-dimensional convolution module to obtain a calibrated new position embedding matrix, wherein the mathematical expression of the new position embedding matrix is:
[0026]
[0027] In the formula, Embed the new position in the matrix of the first The position, the The value of each channel, For location index, and All are channel indexes. For the number of channels, The second position after freezing parameters is embedded in the matrix of the first position. The position, the The value of each channel, This represents the learnable weights.
[0028] In one embodiment of this application, the mathematical expression of the adapted image features is:
[0029]
[0030] In the formula, For the first The adapted image features of each instance For feature scale index, The number of feature scales, 3D keypoint index for the target instance. The number of 3D keypoints in the target instance. For learnable weight parameters, For the first The first instance 3D key points This is a projection function used to project three-dimensional world coordinates onto two-dimensional image pixel coordinates. For the first Two-dimensional image feature maps at various feature scales. This indicates bilinear sampling.
[0031] In one embodiment of this application, the autonomous driving perception backbone network includes a feature pyramid, and the image patch embedding vector is a multi-scale feature vector.
[0032] In one embodiment of this application, the autonomous driving perception task adopts a multi-camera configuration, and the images captured by different cameras have different input resolutions; the parameter scale of the learnable 1×1 one-dimensional convolution module depends only on the channel dimension C and is independent of the number of positions N', and the same set of channel calibration parameters uniformly adapts to the position embedding of all cameras with different resolutions.
[0033] This application also provides a spatial prior calibration and feature adaptation system for sparse sensing architectures, including:
[0034] The acquisition module is used to extract a first position embedding matrix from a pre-trained visual transformer and acquire image patch embedding vectors from an autonomous driving perception backbone network. The visual transformer is pre-trained using visual training samples at a first resolution. The first position embedding matrix includes position embedding vectors for multiple positions. The image patch embedding vectors of the autonomous driving perception backbone network are obtained from an environmental image at a second resolution, which is greater than the first resolution.
[0035] A conversion module is used to convert the first position embedding matrix into a second position embedding matrix aligned with the second resolution based on two-dimensional reshaping and bilinear interpolation.
[0036] The calibration module is used to freeze the second position embedding matrix and convolve the second position embedding matrix along the channel dimension using a learnable 1×1 one-dimensional convolution module to obtain a calibrated new position embedding matrix, wherein the calibrated new position embedding matrix includes new position embedding vectors for multiple positions.
[0037] The feature adaptation module is used to perform element-wise addition and fusion of the image block embedding vector and the new position embedding vector to obtain the adaptation features, and input them into the encoder of the visual transformer to obtain the adapted image features.
[0038] This application also provides an electronic device, including: a processor and a memory;
[0039] The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory to cause the electronic device to perform the methods described above.
[0040] This application also provides a computer-readable storage medium having a computer program stored thereon, characterized in that: when the computer program is executed by a processor, it implements the method described above.
[0041] The beneficial effects of this application are as follows: The spatial prior calibration and feature adaptation method and system for sparse perception architectures in this application, by freezing the interpolated position embedding matrix into an untrainable state, completely preserves the spatial layout-sensitive representations (such as relative object positions and scale relationships) learned in the pre-training stage, avoiding the destruction of spatial priors caused by relearning in downstream tasks, and effectively reducing the risk of overfitting. It employs learnable 1×1 one-dimensional convolutions for calibration along the channel dimension, reducing the number of trainable parameters from those related to traditional resolution. Reduced to being only related to the channel dimension This completely decouples the number of parameters from the input resolution. The same set of channel calibration parameters can be uniformly adapted to mixed configurations of multiple cameras, such as high-resolution front-view and low-resolution surround-view, without requiring separate storage or training of parameters for each resolution. The image features output after spatial prior protection and channel response calibration possess precise geometric structure, ensuring that the Deformable Attention mechanism in sparse sensing methods such as Sparse4D accurately locates and provides rich semantic features during four-dimensional keypoint sampling, thereby improving the overall performance of 3D object detection. Attached Figure Description
[0042] The present application will be further described below with reference to the accompanying drawings and embodiments:
[0043] Figure 1 This is a flowchart of a spatial prior calibration and feature adaptation method for sparse perception architecture shown in one embodiment of this application;
[0044] Figure 2 This is a flowchart illustrating the implementation of a spatial prior calibration and feature adaptation method for a sparse perception architecture according to a specific embodiment of this application.
[0045] Figure 3This is a logical structure diagram of a spatial prior calibration and feature adaptation system for sparse perception architecture shown in one embodiment of this application. Detailed Implementation
[0046] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0047] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the drawings only show the layers related to this application and are not drawn according to the actual number, shape and size of the layers in the actual implementation. In the actual implementation, the form, number and proportion of each layer can be arbitrarily changed, and the layer layout may also be more complex.
[0048] Numerous details are explored in the following description to provide a more thorough explanation of embodiments of this application; however, it will be apparent to those skilled in the art that embodiments of this application may be practiced without these specific details.
[0049] While existing technologies have explored solutions to the problem of resolution adaptation in location embedding, they suffer from the following shortcomings:
[0050] (1) Traditional 2D interpolation schemes do not protect spatial prior knowledge: The 2D bilinear interpolation scheme proposed in the original ViT paper is a widely adopted method for position embedding resolution adaptation in the industry. This scheme reshapes the pre-trained position embedding into a two-dimensional tensor and adjusts it to the target size through bilinear interpolation. Although this method maintains the relative relationship of spatial positions, the interpolation operation smooths out the fine spatial pattern of the original position embedding. If the interpolated position embedding is set as a trainable parameter, the downstream task needs to relearn all parameters, which leads to the destruction of the spatial prior knowledge accumulated in the pre-training stage. In addition, when the fine-tuning resolution is much higher than the pre-training resolution, the interpolated position embedding will be over-smoothed, and the spatial discrimination will be significantly reduced, which directly affects the performance of perception methods such as Sparse4D that rely on precise spatial localization.
[0051] (2) Existing PEFT methods do not have a dedicated adaptation strategy for location embedding: Existing efficient parameter fine-tuning methods such as LoRA and Adapter mainly focus on low-rank adaptation of the Transformer weight matrix and do not have a dedicated adaptation strategy for location embedding. LoRA operates on the Attention and FFN weight matrix and does not directly handle location embedding resolution adaptation. Adapter inserts a lightweight bottleneck layer between Transformer layers and also does not involve location embedding resolution adaptation. ViTAR proposes a fuzzy location coding method, but it still needs to learn new parameters for each resolution and does not consider the decoupling of spatial layout and channel response. The SPC module of this invention decouples the location embedding role into two parts: spatial layout and channel response. It protects spatial prior knowledge by freezing the interpolated location embedding and learns the channel response adjustment after resolution change through a lightweight 1x1 Conv1D channel calibrator, thus realizing the organic combination of spatial prior protection and efficient parameter adaptation.
[0052] (3) Existing solutions do not consider co-optimization with sparse sensing architectures: Sparse sensing methods such as Sparse4D perform sparse sampling through the Deformable Attention mechanism, and the sampling quality is highly dependent on the fidelity of spatial structure information in the backbone network features. Existing location embedding adaptation schemes do not consider the special spatial accuracy requirements of sparse sensing architectures. After adaptation, features suffer from geometric distortion when the spatial resolution changes, leading to a shift in the Deformable Attention sampling position and affecting the accuracy of 3D target detection and localization. This invention, through co-design with Sparse4D, ensures that the features adapted by SPC support high-quality sparse sampling of four-dimensional key points, calibrating the channel response while protecting spatial priors, making the feature representation obtained by sparse sampling more accurate.
[0053] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a spatial prior calibration and feature adaptation method for sparse perceptual architectures. Its technical objectives include:
[0054] First, a Spatial Prior Calibration (SPC) module is designed to freeze the original spatial layout information during the process of embedding and adapting the pre-trained positions to the downstream task resolution. The module learns the channel response adjustment after the resolution change only through a lightweight channel calibrator, thereby protecting the spatial prior knowledge learned in the pre-training stage.
[0055] Second, minimize the number of parameters required for location embedding adaptation, reducing the traditional... Dense parameter learning transformed into only Channel calibration parameter learning reduces the number of parameters. times.
[0056] Third, through co-design with the Sparse4D sparse perception architecture, we ensure that the features adapted by SPC can support high-quality sparse sampling of four-dimensional key points, thereby improving the overall performance of multi-view three-dimensional object detection.
[0057] The SPC module in this application is used to address the mismatch between the pre-trained location embedding and the input resolution of the downstream task. To facilitate understanding of the technical essence of this invention, an analysis will first be conducted based on mathematical principles.
[0058] (1) Mathematical definition of position embedding.
[0059] Let the position embedding in the pre-training phase be... ,in Number of patches ( The image height and width, (for patch size) For channel dimensions. The input Transformer is the sum of the position embedding and the patch embedding:
[0060]
[0061] in Embedded in patch.
[0062] When the resolution of the downstream task becomes At that time, the number of patches became The shape of the position embedding is from Become It cannot be used directly.
[0063] (2) The parameter dilemma of traditional interpolation schemes.
[0064] Traditional solutions will Remodeling into a two-dimensional tensor (in , Adjusted to the target size using bilinear interpolation. Remodel back .
[0065] If If set as trainable parameters, then updates are required. One parameter. Using ViT-Base ( At 320×800 resolution ( , , For example, the trainable parameters are as follows: Much larger than during pre-training .
[0066] (3) SPC channel calibration mechanism.
[0067] The SPC module proposed in this invention decouples the role of position embedding into two parts:
[0068] Spatial layout role: embedded by frozen interpolation positions It is responsible for encoding the spatial relationships between tokens;
[0069] Channel response role: This is undertaken by a learnable channel calibrator, which adjusts the response distribution along the channel dimension.
[0070] Figure 1 This is a flowchart of a spatial prior calibration and feature adaptation method for sparse perception architecture in one embodiment of this application, as follows: Figure 1 As shown, the spatial prior calibration and feature adaptation method for sparse perceptual architectures in this application mainly includes the following steps:
[0071] S110, extract the first position embedding matrix from the pre-trained visual transformer and obtain the image patch embedding vector from the autonomous driving perception backbone network, wherein the visual transformer is pre-trained using visual training samples at a first resolution, the first position embedding matrix includes position embedding vectors for multiple positions, and the image patch embedding vector from the autonomous driving perception backbone network comes from an environmental image at a second resolution, the second resolution being greater than the first resolution;
[0072] In this application, the autonomous driving perception task adopts a multi-camera configuration, and the images captured by different cameras have different input resolutions; the parameter scale of the learnable 1×1 one-dimensional convolution module depends only on the channel dimension C and is independent of the number of positions N', and the same set of channel calibration parameters uniformly adapts to the position embedding of all cameras with different resolutions.
[0073] The first position embedding matrix comes from the pre-trained Visual Transformer (ViT) model itself, which is the parameter matrix learned by the model during the pre-training stage.
[0074] During the pre-training phase (e.g. on the ImageNet dataset with an input resolution of 224×224), ViT divides each image into fixed-size patches and assigns a learnable multidimensional vector to each patch to encode the patch’s position information in two-dimensional space (such as the row and column number).
[0075] These vectors, optimized through pre-training on large-scale data, gradually form a sensitive representation of spatial layout. For example, the model learns the spatial relative relationship between the "top left corner" and the "bottom right corner," as well as the approximate scale and positional distribution of objects in the image. These vectors are stored in the pre-trained model weights in matrix form, which is the first position embedding matrix.
[0076] This matrix comprises multiple location embedding vectors, each corresponding to a location in the image grid (e.g., 14×14=196 image patches, plus one classification token). It serves as the spatial prior knowledge source for subsequent adaptation operations.
[0077] The image patch embedding vector originates from the current input autonomous driving environment image itself and is a feature vector calculated in real time by the front-end processing module of the autonomous driving perception backbone network.
[0078] In downstream tasks of autonomous driving, the vehicle's onboard cameras capture high-resolution environmental images (e.g., 640×1600). These images are input to the autonomous driving perception backbone network (typically a ViT architecture or a variant thereof), first passing through a patch embedding layer:
[0079] Divide the high-resolution image into multiple image blocks;
[0080] Each image patch is mapped to a multidimensional vector through a linear mapping layer (usually a convolutional layer), which encodes the visual content information of the image patch (such as texture, edges, color, local structure of objects, etc.).
[0081] All these vectors are combined to form the set of image patch embedding vectors corresponding to the current image.
[0082] Furthermore, in this application, the autonomous driving perception backbone network includes a feature pyramid, and the image patch embedding vector is a multi-scale feature vector.
[0083] S120, the first position embedding matrix is transformed into a second position embedding matrix aligned with the second resolution based on two-dimensional reshaping and bilinear interpolation. The specific process includes:
[0084] S121, Extract the classification token position embedding vector from the first position embedding matrix to obtain the intermediate position embedding matrix;
[0085] The first row of the pre-trained position embedding matrix is the position embedding vector of the [CLS] token. This token was designed as a globally convergent feature during pre-training and does not have two-dimensional spatial coordinates (no height and width attributes). If it is forcibly included in a two-dimensional grid, it is impossible to determine which specific coordinate of the grid it should belong to, and the bilinear interpolation algorithm requires a strict H×W rectangular grid as input. Therefore, it is necessary to temporarily isolate it and perform interpolation only on the tile embeddings with clear spatial locations.
[0086] S122, the intermediate position embedding matrix is converted into a two-dimensional tensor, and bilinear interpolation is performed on the two-dimensional tensor to obtain an interpolated position embedding matrix aligned with the second resolution.
[0087] One-dimensional sequence form ( The tile positions are embedded and reshaped into a two-dimensional tensor according to the grid layout during pre-training. This process imbues the grid with image attributes. Then, bilinear interpolation is used to resample the grid points according to the target resolution. The core of this algorithm lies in the fact that the value of each point on the new grid is obtained by taking the inverse-weighted average of the values of the four nearest vertices in the original grid. This ensures that the stretched position embedding remains smooth and continuous in the spatial topology, without abrupt numerical changes. Specifically, this includes:
[0088] S1221, stretch the two-dimensional tensor from the first resolution to the second resolution to obtain multiple stretched source grids and multiple interpolation points between the source grids;
[0089] S1222, For any interpolation point, take the four nearest vertices of the source mesh as the target vertex and obtain the position embedding vector of the target vertex;
[0090] S1223, weighted fusion of the position embedding vectors of the four target vertices to obtain the position embedding vector of the interpolation point;
[0091] S1224, construct a stretched two-dimensional tensor based on the position embedding vectors of multiple source grids and the position embedding vectors of multiple interpolation points, and flatten the stretched two-dimensional tensor back into a sequence form to obtain an interpolated position embedding matrix aligned with the second resolution.
[0092] S123, the extracted classification token position embedding vector is re-attached to the front end of the interpolated position embedding matrix to obtain the second position embedding matrix.
[0093] The interpolation operation is performed only on the tile positions, with [CLS] isolated from the rest. After interpolation, the position embedding vector of [CLS] is placed back at the beginning of the sequence (index 0). This ensures that [CLS] still exists as the global origin (reference frame), and the sequence length of the entire position embedding matrix becomes N'+1N'+1, which is completely consistent with the structure of the current input sequence ([CLS] + tile sequence).
[0094] S130, freeze the second position embedding matrix and calibrate the second position embedding matrix after freezing parameters to obtain a calibrated new position embedding matrix, wherein the calibrated new position embedding matrix includes new position embedding vectors for multiple positions.
[0095] The calibration process includes:
[0096] The second position embedding matrix is convolved along the channel dimension using a learnable 1×1 one-dimensional convolution module to obtain a calibrated new position embedding matrix, wherein the mathematical expression of the new position embedding matrix is:
[0097]
[0098] In the formula, Embed the new position in the matrix of the first The position, the The value of each channel, For location index, and All are channel indexes. For the number of channels, The second position after freezing parameters is embedded in the matrix of the first position. The position, the The value of each channel, This represents the learnable weights.
[0099] S140, the image block embedding vector and the new position embedding vector are added and fused element by element to obtain the adaptation features, and then input into the encoder of the visual transformer to obtain the adapted image features.
[0100] The ViT backbone network outputs multi-scale feature maps through a feature pyramid network. The SPC module ensures that the location embeddings at each scale are adapted, enabling the Sparse4D Decoder to obtain image features of consistent quality when performing multi-scale four-dimensional keypoint sampling.
[0101] Sparse4D's Deformable Attention mechanism obtains the feature representation of the target instance by sparsely sampling on the image feature map:
[0102]
[0103] In the formula, For the first The adapted image features of each instance For feature scale index, The number of feature scales, 3D keypoint index for the target instance. The number of 3D keypoints in the target instance. For learnable weight parameters, For the first The first instance 3D key points This is a projection function used to project three-dimensional world coordinates onto two-dimensional image pixel coordinates. For the first Two-dimensional image feature maps at various feature scales. This indicates bilinear sampling.
[0104] The Conv1D channel calibrator in the SPC module is designed to be resolution-independent:
[0105]
[0106] Regardless of changes in input resolution, the interpolated position embedding size The channel calibrator will change accordingly, but it will always receive a fixed channel dimension. The input and output have the same dimensions. This characteristic allows for the adaptation of the same set of channel calibration parameters to position embedding at any resolution.
[0107] The comparison of the number of parameters in this application with those learned directly is shown in Table 1:
[0108] Table 1. Parameter Comparison
[0109]
[0110] As shown in Table 1, taking ViT-Base as an example, the number of parameters decreased from 192,000 to... (Note: Here) ,and SPC has slightly more parameters than directly learning position embeddings, but its key advantage lies in preserving spatial prior knowledge—directly learning position embeddings would destroy the pre-trained spatial prior, while SPC freezes the interpolated position embeddings. Only the channel calibration matrix is learned, thus preserving spatial priors. For ViT-Large ( ) and ViT-Huge ( ), at high resolution Much larger The advantages of SPC are even more obvious.
[0111] Figure 2 This is a flowchart illustrating the implementation of a spatial prior calibration and feature adaptation method for sparse sensing architectures according to a specific embodiment of this application. Figure 2 As shown, using ViT-B as the backbone network, StreamPETR as the detector, and the nuScenes dataset as the evaluation benchmark, the SPC module of this invention is implemented, and the process includes:
[0112] Step 1: Extract the position embedding P of the pre-trained ViT-B (N=197, C=768, corresponding to a 14x14 patch grid plus 1 class token).
[0113] Step 2: Remove the position embeddings of the class token and reshape the remaining 196 position embeddings into a 14x14x768 two-dimensional tensor.
[0114] Step 3: Use bilinear interpolation to adjust the position embedding to the target resolution (e.g., 320x800 corresponds to a 10x25 patch grid), and obtain the interpolated position embedding P_interp (250x768).
[0115] Step 4: Set P_interp to untrainable (frozen), and then cascade a 1x1 Conv1D channel calibrator (768x768 parameters) after it.
[0116] Step 5: Compare the performance differences between SPC and baseline schemes such as directly learned position embeddings (trainable) and fully frozen position embeddings.
[0117] The experimental results are shown in Table 2:
[0118] Table 2. Experimental Results - Performance Comparison of Embedding Strategies at Different Locations
[0119]
[0120] This invention has the following outstanding features and beneficial effects:
[0121] (1) Spatial Prior Protection: By freezing the interpolated position embeddings, the spatial layout prior knowledge learned in the pre-trained model is protected, avoiding the degradation of spatial perception caused by completely relearning the position embeddings. Ablation experiments show that the NDS is 53.2% when the position embeddings are completely frozen (without learning), while the NDS improves to 53.9% when SPC is used. This indicates that SPC effectively adapts to the downstream task resolution through channel calibration while protecting the spatial prior, achieving a 0.7% improvement in NDS. In contrast, although the NDS reaches 54.0% when all position embedding parameters are directly learned, it destroys the pre-trained spatial prior knowledge, increases the risk of overfitting, and the number of parameters is uncontrollable. SPC achieves the optimal balance between protecting the spatial prior and improving performance.
[0122] (2) High efficiency, controllability, and scalability of parameters: The SPC module transforms the number of trainable parameters for location embedding adaptation from N_prime×C to C×C channel calibration parameters. Taking ViT-Base (C=768) at a resolution of 320×800 (N_prime=250) as an example, the number of trainable parameters for directly learning location embedding is 250×768=192,000, while the number of SPC parameters is 768×768=589,824. The core advantage of SPC is that the parameter size is independent of the input resolution—when the resolution is increased to 640×1600 (N_prime=1000), the number of directly learned parameters surges to 1000×768=768,000, exceeding the 589,824 of SPC. For ViT-Large (C=1024) at a resolution of 640×1600, the number of parameters learned directly reaches 1,024,000, while SPC maintains 1,048,576. At a higher resolution of 800×1920 (N_prime=1500), the number of parameters learned directly reaches 1,536,000, while SPC still maintains 1,048,576, saving 31.8%. For ViT-Huge (C=1280), SPC has a parameter advantage at most resolutions commonly used in autonomous driving. More importantly, SPC protects the prior knowledge of the pre-training space by freezing the interpolation position embedding, a technical effect that direct learning schemes cannot achieve.
[0123] (3) Resolution Independence and Unified Adaptation for Multiple Cameras: The Conv1D channel calibrator in the SPC module operates only along the channel dimension, with a parameter dimension of C×C, independent of the input sequence length N_prime. Regardless of the input resolution (e.g., 320×800 corresponds to N_prime=250, 640×1600 corresponds to N_prime=1000, or 800×1920 corresponds to N_prime=1500), the same set of channel calibration parameters can be adapted. This makes SPC particularly suitable for autonomous driving multi-camera systems—for example, a forward-looking camera uses a high resolution (640×1600) to detect distant targets, while a surround-view camera uses a lower resolution (320×800) to cover near-field blind spots. SPC can uniformly adapt the position embeddings of all cameras without requiring separate training for each resolution. In multi-resolution mixed configurations, SPC can reduce the total number of position embedding adaptation parameters by 40%-70% compared to direct learning schemes (depending on the specific resolution combination), significantly reducing the model storage and training complexity of multi-camera systems.
[0124] (4) Deep Co-optimization with Sparse Perception Architecture: The SPC module ensures that the Deformable Attention of the Sparse4D Decoder can perform sparse sampling at the correct spatial location by protecting accurate spatial prior knowledge. Sparse4D relies on the Deformable Attention mechanism to perform sparse sampling on the image feature map to obtain the feature representation of the target instance. The sampling quality is highly dependent on the fidelity of the spatial structure information in the feature map. SPC maintains the geometric accuracy by freezing the interpolation position embedding, and at the same time adjusts the channel response distribution after the resolution change by channel calibration, so that the output features of the backbone network can meet both the spatial accuracy requirements and the channel adaptation requirements. Ablation experiment verification: On the StreamPETR detector, the NDS is 53.2% when using fully frozen position embedding, and the NDS is improved to 53.9% (+0.7%) when using SPC, which is close to the 54.0% of the direct learning scheme. However, SPC achieves this performance with lower risk and more controllable parameter amount. Furthermore, SPC is deeply integrated with the BEFT (BEVEfficient Fine-Tuning) framework. BEFT freezes most of the parameters of the ViT backbone during the training phase and achieves efficient parameter fine-tuning through the SLRT (Side Low-Rank Tuner) and SPC modules. SPC, as a position embedding adaptation component, is a key component of the BEFT technology system. The two work together to enable the 637M parameters of ViT-Huge to be trained on a 24G VRAM consumer GPU, achieving 61.4% NDS.
[0125] (5) Modular Design and Plug-and-Play Characteristics: The SPC module, as an independent component, can be plugged and played into any ViT-based perception system without relying on a specific detector architecture or backbone network size. The SPC only needs to be connected in series with a 1x1 Conv1D layer after location embedding interpolation, without modifying the internal structure of the Transformer or adjusting the training process. The SPC has been verified to be effective on ViT-B (88M parameters), ViT-L (307M parameters), and ViT-H (637M parameters), and is compatible with various multi-view 3D detectors such as StreamPETR, BEVFormer, and PETR. At the deployment level, the channel calibration parameters of the SPC can be stored in combination with the SLRT parameters of BEFT without adding extra storage overhead. The introduction of the SPC module has a negligible impact on inference speed (latency increase of less than 0.1ms), meeting the real-time requirements of autonomous driving systems. In addition, the SPC is independent of other components of the BEFT framework (such as Side Low-Rank Tuner and Bias Tuning), and can be flexibly combined and used according to actual needs, providing high flexibility for the deployment of perception systems under different resource constraints.
[0126] like Figure 3 As shown, this application also provides a spatial prior calibration and feature adaptation system for sparse sensing architectures, including:
[0127] The acquisition module is used to extract a first position embedding matrix from a pre-trained visual transformer and acquire image patch embedding vectors from an autonomous driving perception backbone network. The visual transformer is pre-trained using visual training samples at a first resolution. The first position embedding matrix includes position embedding vectors for multiple positions. The image patch embedding vectors of the autonomous driving perception backbone network are obtained from an environmental image at a second resolution, which is greater than the first resolution.
[0128] A conversion module is used to convert the first position embedding matrix into a second position embedding matrix aligned with the second resolution based on two-dimensional reshaping and bilinear interpolation.
[0129] The calibration module is used to freeze the second position embedding matrix and convolve the second position embedding matrix along the channel dimension using a learnable 1×1 one-dimensional convolution module to obtain a calibrated new position embedding matrix, wherein the calibrated new position embedding matrix includes new position embedding vectors for multiple positions.
[0130] The feature adaptation module is used to perform element-wise addition and fusion of the image block embedding vector and the new position embedding vector to obtain the adaptation features, and input them into the encoder of the visual transformer to obtain the adapted image features.
[0131] This application presents a spatial prior calibration and feature adaptation method and system for sparse perception architectures. By freezing the interpolated position embedding matrix into a non-trainable state, the spatial layout-sensitive representations (such as relative object positions and scale relationships) learned during the pre-training stage are fully preserved, avoiding the destruction of spatial priors caused by relearning in downstream tasks and effectively reducing the risk of overfitting. It employs learnable 1×1 one-dimensional convolutions along the channel dimension for calibration, reducing the number of trainable parameters from those dependent on traditional resolution. Reduced to being only related to the channel dimension This completely decouples the number of parameters from the input resolution. The same set of channel calibration parameters can be uniformly adapted to mixed configurations of multiple cameras, such as high-resolution front-view and low-resolution surround-view, without requiring separate storage or training of parameters for each resolution. The image features output after spatial prior protection and channel response calibration possess precise geometric structure, ensuring that the Deformable Attention mechanism in sparse sensing methods such as Sparse4D accurately locates and provides rich semantic features during four-dimensional keypoint sampling, thereby improving the overall performance of 3D object detection.
[0132] This embodiment also provides an electronic terminal, including: a processor and a memory;
[0133] The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory so that the terminal performs any of the methods in this embodiment.
[0134] As will be understood by those skilled in the art, the computer-readable storage medium described in this embodiment allows for the implementation of all or part of the steps in the above method embodiments by computer program-related hardware. The aforementioned computer program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0135] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication between them. The memory is used to store computer programs, the communication interface is used to perform communication, and the processor and the transceiver are used to run the computer programs, so that the electronic terminal performs the steps of the above method.
[0136] In this embodiment, the memory may include random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.
[0137] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0138] In the above embodiments, although the present application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. The embodiments of the present application are intended to cover all such substitutions, modifications, and variations falling within the broad scope of the appended claims.
[0139] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A spatial prior calibration and feature adaptation method for sparse perceptual architectures, characterized in that, Feature adaptation for autonomous driving perception tasks includes: The first position embedding matrix is extracted from the pre-trained visual transformer, and the image patch embedding vector from the autonomous driving perception backbone network is obtained. The visual transformer is pre-trained using visual training samples at a first resolution. The first position embedding matrix includes position embedding vectors for multiple positions. The image patch embedding vector from the autonomous driving perception backbone network comes from an environmental image at a second resolution, which is greater than the first resolution. The first position embedding matrix is transformed into a second position embedding matrix aligned with the second resolution based on two-dimensional reshaping and bilinear interpolation. The second position embedding matrix is frozen, and the second position embedding matrix after freezing parameters is calibrated to obtain a calibrated new position embedding matrix, wherein the calibrated new position embedding matrix includes new position embedding vectors for multiple positions. The image patch embedding vector and the new position embedding vector are added and fused element by element to obtain the adaptation features, which are then input into the encoder of the visual transformer to obtain the adapted image features.
2. The spatial prior calibration and feature adaptation method for sparse sensing architectures according to claim 1, characterized in that, The first position embedding matrix is transformed into a second position embedding matrix aligned with the second resolution based on two-dimensional reshaping and bilinear interpolation, including: Extract the classification token position embedding vector from the first position embedding matrix to obtain the intermediate position embedding matrix; The intermediate position embedding matrix is converted into a two-dimensional tensor, and bilinear interpolation is performed on the two-dimensional tensor to obtain an interpolated position embedding matrix aligned with the second resolution. The extracted classification token position embedding vector is re-attached to the front end of the interpolated position embedding matrix to obtain the second position embedding matrix.
3. The spatial prior calibration and feature adaptation method for sparse sensing architectures according to claim 2, characterized in that, Performing bilinear interpolation on the two-dimensional tensor to obtain an interpolated position embedding matrix aligned with the second resolution includes: The two-dimensional tensor is stretched from a first resolution to a second resolution to obtain multiple stretched source grids and multiple interpolation points between the source grids; For any interpolation point, take the four nearest vertices of the source mesh as the target vertex and obtain the position embedding vector of the target vertex; The position embedding vectors of the four target vertices are weighted and fused to obtain the position embedding vector of the interpolation point. A stretched two-dimensional tensor is constructed based on the position embedding vectors of multiple source grids and the position embedding vectors of multiple interpolation points. The stretched two-dimensional tensor is then flattened back into a sequence form to obtain an interpolated position embedding matrix aligned with the second resolution.
4. The spatial prior calibration and feature adaptation method for sparse sensing architectures according to claim 1, characterized in that, The second position embedding matrix after freezing parameters is calibrated to obtain a new calibrated position embedding matrix, including: The second position embedding matrix is convolved along the channel dimension using a learnable 1×1 one-dimensional convolution module to obtain a calibrated new position embedding matrix, wherein the mathematical expression of the new position embedding matrix is: In the formula, Embed the new position in the matrix of the first The position, the The value of each channel, For location index, and All are channel indexes. For the number of channels, The second position after freezing parameters is embedded in the matrix of the first position. The position, the The value of each channel, This represents the learnable weights.
5. The spatial prior calibration and feature adaptation method for sparse sensing architectures according to claim 1, characterized in that, The mathematical expression for the adapted image features is: In the formula, For the first The adapted image features of each instance For feature scale index, The number of feature scales, 3D keypoint index for the target instance. The number of 3D keypoints in the target instance. For learnable weight parameters, For the first The first instance 3D key points This is a projection function used to project three-dimensional world coordinates onto two-dimensional image pixel coordinates. For the first Two-dimensional image feature maps at various feature scales. This indicates bilinear sampling.
6. The spatial prior calibration and feature adaptation method for sparse sensing architectures according to claim 1, characterized in that, The autonomous driving perception backbone network includes a feature pyramid, and the image patch embedding vector is a multi-scale feature vector.
7. The spatial prior calibration and feature adaptation method for sparse sensing architectures according to claim 4, characterized in that, The autonomous driving perception task employs a multi-camera configuration, with images captured by different cameras having different input resolutions. The parameter size of the learnable 1×1 one-dimensional convolution module depends only on the channel dimension C and is independent of the number of positions N'. The same set of channel calibration parameters uniformly adapts to the position embedding of all cameras at different resolutions.
8. A spatial prior calibration and feature adaptation system for sparse sensing architectures, characterized in that, include: The acquisition module is used to extract a first position embedding matrix from a pre-trained visual transformer and acquire image patch embedding vectors from an autonomous driving perception backbone network. The visual transformer is pre-trained using visual training samples at a first resolution. The first position embedding matrix includes position embedding vectors for multiple positions. The image patch embedding vectors of the autonomous driving perception backbone network are obtained from an environmental image at a second resolution, which is greater than the first resolution. A conversion module is used to convert the first position embedding matrix into a second position embedding matrix aligned with the second resolution based on two-dimensional reshaping and bilinear interpolation. The calibration module is used to freeze the second position embedding matrix and convolve the second position embedding matrix along the channel dimension using a learnable 1×1 one-dimensional convolution module to obtain a calibrated new position embedding matrix, wherein the calibrated new position embedding matrix includes new position embedding vectors for multiple positions. The feature adaptation module is used to perform element-wise addition and fusion of the image block embedding vector and the new position embedding vector to obtain the adaptation features, and input them into the encoder of the visual transformer to obtain the adapted image features.
9. An electronic device, characterized in that, include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Fine-grained image recognition method and system based on shared visual backbone network
CN120689686A
Freezing visual basis model-based parameter efficient no-prompt remote sensing image semantic segmentation method and device
CN122336277A