A camera-based multi-task multi-view distillation method

CN118298395BActive Publication Date: 2026-09-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410380954.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-31
Publication Date
2026-09-25
Estimated Expiration
2044-03-31

AI Technical Summary

Technical Problem

Segment Anything Model(SAM)视觉分割基础模型,借助自然语言处理(Natural Language Processing,NLP)任务中的Prompt思路,通过给图像分割任务提供Prompt提示来完成任意目标的快速分割,使用提示工程来训练一个根据提示进行分割的预训练大模型,该模型具有在下游分割任务应用的潜力,并且可以与其他视觉任务组合形成新解决方案,尽管此基础模型拥有零样本泛化能力,由于缺乏适当的优化,它们无法理解自动驾驶场景中深度特征的特性,未能适应不同国家的不同驾驶场景,故仍未能良好应用于自动驾驶领域中

Benefits of technology

[0026]1、本发明通过对SAM大模型进行微调,在微调过程中添加重建损失与深度预测损失等少量科学参数,在保证实际部署成本和存储成本都很小的同时,使基础模型获得对图像的深度进行一定推理的能力,并获得深度信息,使本发明方法在多样化且从未出现过的环境中有效适应并表现良好的能力,从而缓解与不同领域相关的性能损失,提升模型的整体泛化能力;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118298395B_ABST
    Figure CN118298395B_ABST
Patent Text Reader

Abstract

The application discloses a camera-based multi-task multi-view distillation method, first, a benchmark model is established, a student model based on a camera is adopted, and a teacher model is Object-DGCNN; then, a depth loss and a reconstruction loss are calculated, and distillation of BEV features is completed; finally, a multi-task head is used for 3D target detection and semantic segmentation. The application realizes semantic segmentation and 3D detection tasks on an automatic driving data set, and widens the use scene of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autonomous driving technology, specifically relating to a multi-task, multi-view distillation method based on a camera. Background Technology

[0002] In recent years, autonomous driving technology has flourished, but the challenges vehicles face in complex and ever-changing road environments have become increasingly significant. The driving environment on roads is highly dynamic, characterized by constantly changing traffic conditions, diverse traffic signals, and varying road surface materials across different countries. The main participants in road traffic also differ, making it difficult for autonomous driving models to demonstrate good performance in scenarios outside their training domain. Foundation models, with their superior zero-shot and few-shot learning capabilities, have solved the problems of performance loss across different domains and the inability of models to generalize, achieving remarkable success and being widely applied in fields such as image classification, object detection, semantic and panoptic segmentation, video understanding, and few-shot learning. The Segment Anything Model (SAM) is a basic visual segmentation model that leverages the prompting approach from Natural Language Processing (NLP) tasks. It provides prompts to image segmentation tasks to achieve fast segmentation of arbitrary targets. It uses prompting engineering to train a large pre-trained model that segments based on prompts. This model has the potential to be applied to downstream segmentation tasks and can be combined with other visual tasks to form new solutions. Although this basic model has zero-shot generalization ability, due to a lack of proper optimization, it cannot understand the characteristics of deep features in autonomous driving scenarios and has failed to adapt to different driving scenarios in different countries. Therefore, it has not yet been well applied in the field of autonomous driving. Summary of the Invention

[0003] To overcome the shortcomings of existing technologies, this invention provides a camera-based multi-task, multi-view distillation method. First, a baseline model is established, employing a camera-based student model and an Object-DGCNN teacher model. Then, depth loss and reconstruction loss are calculated to distill BEV features. Finally, a multi-task head is used for 3D object detection and semantic segmentation. This invention implements semantic segmentation and 3D detection tasks on an autonomous driving dataset, broadening the application scenarios of the model.

[0004] The technical solution adopted by this invention to solve its technical problem is as follows:

[0005] Step 1: Establish a baseline model;

[0006] The student model is based on a camera and includes a SAM network with learnable parameters, an LSS module for converting camera view to bird's-eye view, and an independent decoder head for object detection and BEV segmentation; the teacher model is a radar point cloud processing network.

[0007] Step 2: Calculate the depth loss L depth and reconstruction losses L rec Complete the distillation of BEV characteristics;

[0008] Step 2-1: Introduce the depth prediction task, using discrete depth ground reality D obtained by back-projection from the lidar point cloud. GT The binary cross-entropy (BCE) loss is used as the depth loss function.

[0009] L depth =f BCE (D,D GT )

[0010] Where D represents the predicted depth value, obtained from image features via DepthNet; D GT This represents the true depth value obtained from the lidar point cloud;

[0011] The LSS method is used to transform image feature information into BEV features. First, the depth distribution of feature points after image plane downsampling is explicitly estimated for each camera image to obtain a view frustum with dimensions H×W×D×C containing image features. Here, H and W represent the spatial position of each image feature point in the vehicle coordinate system after transformation by combining camera intrinsic and extrinsic parameters, and C represents the semantic feature of each image feature point. The view frustums of all cameras are assigned to the BEV grid, and multiple view frustum points in each grid are summed and pooled to form a BEV feature map.

[0012] Step 2-2: Select the foreground region for normalized distillation and introduce a soft-supervised method to create a Gaussian distribution for each true center coordinate in the BEV space:

[0013]

[0014] Where, x i y i Represents the coordinates of each true center. They represent x respectively i ,y i The expected value of the mathematical value, Indicates the variance value;

[0015] Steps 2-3: Determine the BEV features generated by the student model and the teacher model; the student model uses the BEV feature map F generated by the BEVTransformer Encoder based on the large model SAM. 2D To align feature representations between teachers and students, the same BEV features F output by the Transformer Encoder of the teacher model are used. 3D Introducing reconstruction loss L rec Train feature extraction while training depth;

[0016]

[0017] Where H and W represent the width and height of the distillation feature map, respectively, and ||·||2 is the L2 norm; This represents the BEV feature vector of the 3D point cloud transformation in the teacher model. This represents the BEV feature vector derived from camera images in the student model;

[0018] Steps 2-4: The overall loss is expressed as:

[0019] L = L depth +αL rec

[0020] Where α represents the balancing loss term;

[0021] Step 3: Utilize multi-task heads to perform 3D object detection and semantic segmentation;

[0022] Preferably, the learnable parameters in the SAM network come from the VPT model, that is, learnable parameters for a specific task are introduced into the input space, and the entire pre-trained Transformer backbone is kept frozen during downstream training; the learnable parameters are added to the input sequence of each Transformer layer and learned together with the linear head.

[0023] Preferably, the teacher model is Object-DGCNN; the attention mechanism of DGCNN is replaced by a multi-scale attention module; the multi-scale attention module first projects 3D points onto the BEV plane, and then performs one-to-one supervision through Transformer-based label assignment; the teacher model is initialized from a pre-trained CenterPoint model and trained by fixing all parameters during knowledge distillation.

[0024] Preferably, the multi-task head is a 3D target detection head and a BEV segmentation head.

[0025] The beneficial effects of this invention are as follows:

[0026] 1. This invention fine-tunes the large SAM model by adding a small number of scientific parameters such as reconstruction loss and depth prediction loss during the fine-tuning process. While ensuring that the actual deployment cost and storage cost are very small, it enables the basic model to gain a certain ability to infer the depth of the image and obtain depth information. This allows the method of this invention to effectively adapt to and perform well in diverse and unprecedented environments, thereby alleviating the performance loss related to different domains and improving the overall generalization ability of the model.

[0027] 2. This invention makes full use of the rich geometric information of the LiDAR model. Through feature-dense reconstruction, sparse reconstruction of detected objects, and depth supervision, it transfers the advanced knowledge of the model in the BEV space, effectively enhancing the camera-based model's ability in geometric perception.

[0028] 3. After fine-tuning the SAM model, this invention enables semantic segmentation and 3D detection tasks on an autonomous driving dataset, thus broadening the application scenarios of the model. Attached Figure Description

[0029] Figure 1 This is a flowchart of the method of the present invention;

[0030] Figure 2 The following are examples of the embodiments of the present invention: (a) a radar point cloud visualization, (b) a camera visualization without learning, and (c) a BEV feature visualization after learning by the present invention.

[0031] Figure 3 This is a diagram showing the final detection results obtained in an embodiment of the present invention;

[0032] Figure 4 This is a BEV map segmentation result diagram according to an embodiment of the present invention;

[0033] Figure 5 This is a diagram showing the three-dimensional target detection results of an embodiment of the present invention. Detailed Implementation

[0034] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0035] To efficiently apply a base model to autonomous driving, enabling it to understand the characteristics of depth features in autonomous driving scenarios and meet the requirement that base models in the autonomous driving field can perform some inference on image depth and obtain depth information, this invention proposes a camera-based multi-task multi-view distillation method. This method, based on a cue engineering fine-tuning strategy, introduces a base model specifically designed to leverage its significant capabilities in zero-shot and few-shot learning. Building upon the large-scale image segmentation model Segment Anything Model (SAM), visual cue tuning (VPT) is used to fine-tune the SAM image encoder, cue encoder, and mask decoder on labeled autonomous driving datasets. This endows the autonomous driving system with the ability to effectively adapt and perform well in diverse scenarios outside the initially trained domain, thereby mitigating performance losses associated with different domains and improving the overall generalization ability of the large model. Specifically, during dense supervision of feature distillation, to provide the network with more opportunities for alignment information across different modalities, this invention delays supervision until after the transformer encoder encoding of the teacher model. When processing cross-modal features, view differences may lead to poor results, and it is considered that regional 3D features only provide meaningful information when points exist. This invention employs normalized distillation within the foreground region, introducing a soft-supervised method to create a Gaussian distribution for each ground truth center in the BEV space. For depth prediction tasks, this invention utilizes binary cross-entropy loss as the depth loss, derived from discrete ground reality derived from radar point clouds, enhancing convergence speed and proving crucial for fine-tuning the underlying model.

[0036] like Figures 1-5 As shown, the method steps of the present invention are as follows:

[0037] Step 1: Establish a baseline model;

[0038] This invention employs a camera-based student model, comprising a SAM network with learnable parameters, an LSS module for converting camera footage to bird's-eye view, and an independent decoder head for object detection and BEV segmentation; the teacher model is a radar point cloud processing network. The learnable parameters in the SAM network are derived from the Visual-Prompt Tuning (VPT) model, which introduces a small number of task-specific learnable parameters into the input space while keeping the entire pre-trained Transformer backbone frozen during downstream training. These additional learnable parameters are simply added to the input sequence of each Transformer layer and fine-tuned together with the linear head.

[0039] Step 2: Calculate the depth loss L depth and reconstruction losses L recComplete the distillation of BEV characteristics;

[0040] First, a depth prediction task is introduced, using discrete depth ground data (D) extracted from lidar point clouds. GT The binary cross-entropy (BCE) loss is used as the depth loss function.

[0041] L depth =f BCE (D,D GT )

[0042] Where D represents the predicted depth value, D GT This represents the true depth value obtained from the lidar point cloud;

[0043] The Lift-Splat-Shoot (LSS) method is used to transform image feature information into BEV features. First, the depth distribution of feature points after image plane downsampling is explicitly estimated for each camera image, resulting in a view frustum (point cloud) with dimensions H×W×D×C containing image features. Here, H and W represent the spatial position of each image feature point in the vehicle coordinate system obtained by combining camera intrinsic and extrinsic parameters, respectively, and C represents the semantic feature of each image feature point. The view frustums (point clouds) of all cameras are then assigned to the BEV grid, and sum-pooling is performed on multiple view frustum points in each grid to form the BEV feature map.

[0044] Determine the BEV features generated by the student model and the teacher model. The student model directly utilizes the BEV feature map F generated by the BEVTransformerEncoder. 2D To align feature representations between teachers and students, the same BEV features F output by the Transformer Encoder of the teacher model are used. 3D .

[0045] Then, the reconstruction loss L is introduced. rec Training feature extraction while training depth

[0046]

[0047] Where H and W represent the width and height of the distillation feature map, respectively, and ||·||2 is the L2 norm. This represents the BEV feature vector of the 3D point cloud transformation in the teacher model. This represents the BEV feature vector derived from camera images in the student model;

[0048] However, in the cross-modal feature setting, although view differences are eliminated by projecting the two features onto the BEV plane, domain differences still exist between different modalities. Considering that regional 3D features only provide meaningful information when points are present, and that foreground boundaries may also provide useful information, we choose to perform canonical distillation within the foreground regions and avoid hard supervision for each foreground region. Instead, we introduce a soft supervision method to create a Gaussian distribution for each ground truth center coordinate in the BEV space:

[0049]

[0050] Where, x i y i Represents the coordinates of each true center. They represent x respectively i ,y i The expected value of the mathematical value, Indicates the variance value;

[0051] Then the reconstruction loss L rec Can be integrated into

[0052]

[0053] The overall loss can be expressed as

[0054] L = L depth +αL rec

[0055] Step 3: Utilize multi-task heads to perform 3D object detection and semantic segmentation;

[0056] Multiple task-specific heads, such as 3D object detection heads and BEV segmentation heads, are applied to the fused BEV feature maps, making the method of this invention applicable to most 3D perception tasks.

[0057] Example:

[0058] Step 1: Establish a baseline model;

[0059] This invention uses a camera-based base model as the student model. This model comprises a SAM network with learnable parameters, an LSS module for camera-to-bird's-eye view (BEV) transformation, and independent decoder heads for object detection and BEV segmentation. The learnable parameters in the SAM network are derived from the Visual-Prompt Tuning (VPT) model. This method does not modify or fine-tune the pre-trained Transformer itself, but rather modifies the Transformer's input, introducing a small number of task-specific learnable parameters into the input space while keeping the entire pre-trained Transformer backbone frozen during downstream training. These additional learnable parameters are simply added to the input sequence of each Transformer layer and fine-tuned together with the linear head.

[0060] To maintain consistency with the student model, this invention selects Object-DGCNN as the teacher model. For simplification and generality, the attention mechanism of DGCNN is replaced with a standard multi-scale attention module. This module first projects 3D points onto the BEV plane, and then performs one-to-one supervision through Transformer-based label assignment. The model is initialized from a pre-trained CenterPoint model, and trained by fixing all parameters during knowledge distillation.

[0061] Step 2: Calculate the depth loss L depth and reconstruction losses L rec Complete the distillation of BEV characteristics;

[0062] BEVDistill employs a common knowledge distillation paradigm, using a 3D point cloud detector as the teacher model and an image detector as the student model. Unlike previous knowledge distillation methods, which typically maintain the same architecture for both the student and teacher models (except for the backbone structure), the SAMLSS method proposed in this invention addresses a more challenging scenario characterized by non-uniform representations.

[0063] Regarding BEV feature distillation, for intensive supervision, the initial step is to determine the BEV features generated by the two models. For the student model, this invention directly utilizes the BEV feature map F generated by BEVTransformerEncoder. 2D To align feature representations between teachers and students, this invention chooses to use the same BEV features F output by the Transformer Encoder of the teacher model. 3D .

[0064] Unlike LIGA-Stereo, which directly mimics the method of extracting BEV features from the 3D backbone, this invention postpones this supervision until after the Transformer Encoder. This delay provides the network with more opportunities to align information across different modalities. Previous research has achieved significant performance improvements by forcing student models to mimic the feature maps of teacher models.

[0065]

[0066] Here, H and W represent the width and height of the distilled feature map, respectively, and ||·||2 is the L2 norm. However, this strategy may not perform well in practice under cross-modal feature settings. Although view differences are eliminated by projecting the two features onto the BEV plane, domain differences still exist between different modalities. Even when capturing the same scene using LiDAR point clouds and camera images, the representations themselves may differ across modalities. For example, camera images contain pixels in both foreground and background regions, while LiDAR points only appear when there is light reflected from an object. Considering that regional 3D features only provide meaningful information when points are present, canonical distillation is chosen within the foreground region. Furthermore, noting that foreground boundaries may also provide useful information, a method of hard supervision for each foreground region is avoided, and instead, a soft supervision method similar to that proposed by Zhou et al. in 2019 is introduced. Specifically, this invention creates a Gaussian distribution for each ground truth center coordinate in the BEV space:

[0067]

[0068] Then the reconstruction loss L rec Can be integrated into

[0069]

[0070] Furthermore, this invention introduces a depth prediction task. Similar to previous studies, this invention uses discrete depth ground reality (DGT) extracted from LiDAR point clouds and employs binary cross-entropy (BCE) loss as the depth loss function to supervise the model inferring the depth of the image.

[0071] L depth =f BCE (D,D GT )

[0072] Therefore, the overall loss can be expressed as:

[0073] L = L depth +αL rec

[0074] The α-balanced loss term is included. It has been verified that depth loss can enhance convergence speed, which is crucial for fine-tuning the base model.

[0075] Step 3: Utilize multi-task heads to perform 3D object detection and semantic segmentation;

[0076] This invention applies multiple task-specific headers to the fused BEV feature map, making the method applicable to most 3D perception tasks. Two specific examples are shown here: 3D object detection and BEV map segmentation.

[0077] For object detection, a class-specific center heatmap head is used to predict the center position of all objects, while multiple regression heads are used to estimate the size, rotation and velocity of objects.

[0078] In terms of semantic segmentation, since different map categories may overlap, this problem is defined as multiple binary semantic segmentation tasks, one for each category. A CVT approach is employed, using standard focus loss to train the segmentation head.

Claims

1. A multi-task, multi-view distillation method based on a camera, characterized in that, Includes the following steps: Step 1: Establish a baseline model; The student model is based on a camera and includes a SAM network with learnable parameters, an LSS module for converting camera view to bird's-eye view, and an independent decoder head for object detection and BEV segmentation; the teacher model is a radar point cloud processing network. Step 2: Calculate the depth loss L depth and reconstruction losses L rec Complete the distillation of BEV characteristics; Step 2-1: Introduce the depth prediction task, using discrete depth ground reality D obtained by back-projection from the lidar point cloud. GT The binary cross-entropy (BCE) loss is used as the depth loss function. L depth =f BCE (D,D GT ) Where D represents the predicted depth value, obtained from image features via DepthNet; D GT This represents the true depth value obtained from the lidar point cloud; The LSS method is used to transform image feature information into BEV features. First, the depth distribution of feature points after image plane downsampling is explicitly estimated for each camera image to obtain a view frustum with dimensions H×W×D×C containing image features. Here, H and W represent the spatial position of each image feature point in the vehicle coordinate system after transformation by combining camera intrinsic and extrinsic parameters, and C represents the semantic feature of each image feature point. The view frustums of all cameras are assigned to the BEV grid, and multiple view frustum points in each grid are summed and pooled to form a BEV feature map. Step 2-2: Select the foreground region for normalized distillation and introduce a soft-supervised method to create a Gaussian distribution for each true center coordinate in the BEV space: Where, x i y i Represents the coordinates of each true center. They represent x respectively i ,y i The expected value of the mathematical value, Indicates the variance value; Steps 2-3: Determine the BEV features generated by the student model and the teacher model; the student model uses the BEV feature map F generated by the BEVTransformer Encoder based on the large model SAM. 2D To align feature representations between teachers and students, the same BEV features F output by the Transformer Encoder of the teacher model are used. 3D Introducing reconstruction loss L rec Feature extraction is trained simultaneously with depth training. Where H and W represent the width and height of the distillation feature map, respectively, and ||·|2 is the L2 norm; This represents the BEV feature vector of the 3D point cloud transformation in the teacher model. This represents the BEV feature vector derived from camera images in the student model; Steps 2-4: The overall loss is expressed as: L=L depth +αL red Where α represents the balancing loss term; Step 3: Utilize a multi-task head to perform 3D object detection and semantic segmentation.

2. The multi-task, multi-view distillation method based on a camera according to claim 1, characterized in that, The learnable parameters in the SAM network come from the VPT model, that is, learnable parameters for a specific task are introduced into the input space, and the entire pre-trained Transformer backbone is kept frozen during downstream training. The learnable parameters are added to the input sequence of each Transformer layer and learned together with the linear head.

3. The multi-task, multi-view distillation method based on a camera according to claim 1, characterized in that, The teacher model is Object-DGCNN; the attention mechanism of DGCNN is replaced by a multi-scale attention module; the multi-scale attention module first projects 3D points onto the BEV plane, and then performs one-to-one supervision through Transformer-based label assignment; the teacher model is initialized from a pre-trained CenterPoint model and trained by fixing all parameters during knowledge distillation.

4. The multi-task, multi-view distillation method based on a camera according to claim 1, characterized in that, The multi-task head consists of a 3D object detection head and a BEV segmentation head.

Citation Information

Patent Citations

  • Three-dimensional target detection model training and using method and device, medium and equipment

    CN115223117A

  • Method and device for training three-dimensional target detection model based on cross-modal knowledge distillation

    CN115690708A