Method and System for 3D Object Detection Using Multimodal Expert Knowledge

By building a vision-centric multimodal expert model and combining trajectory-based distillation and occupation reconstruction module, the problem of insufficient knowledge transfer caused by the gap in lidar and camera characteristics in 3D perception is solved, and the accuracy and performance of 3D object detection are improved.

CN117475425BActive Publication Date: 2025-06-24SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311202519.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-18
Publication Date
2025-06-24
Estimated Expiration
2043-09-18

AI Technical Summary

Technical Problem

In the autonomous driving perception task, the existing 3D perception methods are inadequately transferred knowledge and have low perception accuracy due to the domain gap between lidar and camera features.

Method used

A vision-centered multimodal expert model is constructed, and combined with a trajectory-based distillation module and an occupation reconstruction module, the knowledge of the expert model is transferred to a standard long-term visual detection model to alleviate the misalignment problem during the time fusion process.

Benefits of technology

The performance of vision-based 3D object detection model is improved, and the performance comparable to that of multimodal methods is achieved. At the same time, the model architecture is simplified and the dependence on the lidar backbone network is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117475425B_ABST
    Figure CN117475425B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer vision technology, and particularly to a method and system for 3D object detection using multi-modal expert knowledge. The method includes the following steps: First, construct an expert model; the expert model is a vision-centered multi-modal expert model; then, construct a trajectory-based distillation module and an occupancy reconstruction module; next, transfer the knowledge of the expert model to the apprentice model according to the trajectory-based distillation module and the occupancy reconstruction module; the apprentice model is a standard long-term vision detection model; finally, based on the apprentice model after knowledge transfer, perform 3D object detection. The method of this application proposes a framework for improving the camera-only apprentice model, including a multi-modal expert suitable for the apprentice and distillation supervision suitable for temporal fusion, so as to supervise static and dynamic objects to alleviate the misalignment problem in the long-term time fusion process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer vision technology, and in particular, to a method and system for 3D object detection using multi-modal expert knowledge. Background Art

[0002] Pure vision 3D perception methods, i.e., camera-based 3D perception methods, have received increasing attention in autonomous driving perception tasks. Although camera-only models have the advantages of low deployment cost and easy wide application, in terms of perception accuracy, they still lag behind the state-of-the-art models that utilize lidar sensors. Therefore, extraction methods have been adopted to transfer knowledge from powerful expert models to camera-only apprentice models, with the expectation of using the expertise of these stronger expert models to enhance the capabilities of camera-only models.

[0003] Existing 3D perception extraction methods usually adopt the best-performing expert models, such as point cloud-based models or multi-modal fusion models. However, the domain gap between lidar and camera features hinders knowledge transfer during the distillation process, resulting in limited improvements in practical applications. Summary of the Invention

[0004] The embodiments of the present application provide a method and system for 3D object detection using multi-modal expert knowledge, and propose a framework for improving a camera-only apprentice model, including a multi-modal expert suitable for the apprentice and distillation supervision suitable for temporal fusion, so as to supervise static and dynamic objects to mitigate the misalignment problem during the long-term time fusion process.

[0005] To solve the above technical problems, in a first aspect, the embodiments of the present application provide a method for 3D object detection using multi-modal expert knowledge, including: First, construct an expert model; the expert model is a vision-centered multi-modal expert model; Then, construct a trajectory-based distillation module and an occupancy reconstruction module; Next, transfer the knowledge of the expert model to an apprentice model according to the trajectory-based distillation module and the occupancy reconstruction module; the apprentice model is a standard long-term vision detection model; Based on the apprentice model after knowledge transfer, detect 3D objects.

[0006] In some exemplary embodiments, constructing the expert model includes the following steps: project lidar point cloud scans onto an image to obtain a temporal depth map; based on the image features, perform depth prediction on each pixel in the image to obtain the depth distribution of the image features; project the image features into the BEV space, and according to the depth distribution of the image features, obtain BEV features; the BEV features are bird's-eye view features; convert the BEV features of historical timestamps into the current BEV features; fuse all the BEV features with the temporal depth map to create a unified BEV representation for 3D object detection.

[0007] In some exemplary embodiments, constructing a trajectory-based distillation module includes the following steps: constructing a motion trajectory based on the transformed true object positions of all historical frames; obtaining normalized key sampling features by performing bilinear interpolation sampling on the sampled features and then performing normalization processing; the sampled features being the sampled features at the same points from the expert BEV features and the apprentice BEV features; calculating a trajectory-based distillation loss between the normalized key sampling features; using the motion trajectory as a query to perform trajectory-based distillation at the representative positions corresponding to the motion trajectory, enabling the expert to correct the motion misalignment problem in the apprentice.

[0008] In some exemplary embodiments, the calculation formula for the normalization processing is:

[0009]

[0010] Where, respectively represent the sampled features at the same point P E from the expert BEV feature F A and the apprentice BEV feature F i j′ ;

[0011] The calculation formula for the trajectory-based distillation loss is:

[0012]

[0013] Where, L TD represents the trajectory-based distillation loss; N represents the time interval between the current frame and the historical frames; K represents the number of objects to be measured.

[0014] In some exemplary embodiments, the occupancy reconstruction module establishes a grid occupancy state based on the depth information of the expert model and supervises the apprentice model based on the grid occupancy state.

[0015] In some exemplary embodiments, constructing an occupancy reconstruction module includes the following steps: predicting the depth of each image pixel based on a depth estimation module to obtain a depth map; back-projecting the depth map into a 3D point cloud and converting each image pixel into 3D coordinates; expanding the Gaussian distribution of each 3D coordinate into 3D space to obtain an accurate 3D object modeling; using the grid of the 3D object modeling as auxiliary supervision and using an intuitive L1 regularization loss to optimize the predicted grid occupancy state, thereby enhancing the depth prediction ability for static and dynamic objects.

[0016] In some exemplary embodiments, the conversion formula for converting each image pixel into 3D coordinates is:

[0017]

[0018] Among them, (u, v) represents the image pixel; D(u, v) represents the depth of the image pixel; c u , c v respectively represent the center points of the camera, and f u , f v respectively represent the horizontal and vertical focal lengths; the Gaussian distribution of each 3D coordinate is extended to the 3D space to obtain an accurate 3D object modeling; the mesh of the 3D object modeling is:

[0019]

[0020] Among them, (p x , p y , p z ) represents the center of the 3D object, and σ p represents the standard deviation of each object size.

[0021] An intuitive L1 regularization loss is used to optimize the predicted mesh occupancy state to obtain an occupancy reconstruction loss; its calculation formula is:

[0022]

[0023] Among them, L OR represents the occupancy reconstruction loss; G xyz represents the mesh of the 3D object modeling; G′ xyz represents the predicted mesh occupancy state.

[0024] In some exemplary embodiments, in the process of transferring the knowledge of the expert model to the apprentice model according to the trajectory-based distillation module and the occupancy reconstruction module, the semantic and geometric knowledge of the expert model to the apprentice model is promoted through the joint training loss; the joint training loss L Total is defined as:

[0025] L Total = L A + L TD + L OR (6)

[0026] Among them, L Total represents the joint training loss; L A represents the perceptual loss of the apprentice model; L TD represents the trajectory-based distillation loss; L OR represents the occupancy reconstruction loss.

[0027] Second aspect, the embodiments of the present application further provide a system for 3D object detection using multi-modal expert knowledge, including: a model construction module and a detection module connected to each other; wherein, the model construction module includes an expert model construction unit, a trajectory-based distillation module construction unit, and an occupancy reconstruction module construction unit; the detection module includes an apprentice model; the apprentice model is a standard long-term visual detection model; the expert model construction unit is used to construct a vision-centered multi-modal expert model; the trajectory-based distillation module construction unit and the occupancy reconstruction module construction unit are respectively used to construct a trajectory-based distillation module and an occupancy reconstruction module; the trajectory-based distillation module and the occupancy reconstruction module are used to transfer the knowledge of the expert model to the apprentice model; the detection module is used to detect 3D objects according to the apprentice model after knowledge transfer.

[0028] In some exemplary embodiments, the trajectory-based distillation module construction unit includes: a motion trajectory construction unit, a normalization processing unit, a calculation unit, and a distillation unit connected in sequence; the motion trajectory construction unit is used to construct a motion trajectory according to the transformed real object positions of all historical frames; the normalization processing unit is used to perform bilinear interpolation sampling on the sampled features and then perform normalization processing to obtain normalized key sampled features; the sampled features are the sampled features at the same points from the expert BEV features and the apprentice BEV features; the calculation unit is used to calculate the trajectory-based distillation loss between the normalized key sampled features; the distillation unit is used to use the motion trajectory as a query to perform trajectory-based distillation at the representative positions corresponding to the motion trajectory, so that the expert corrects the motion misalignment problem in the apprentice.

[0029] The technical solutions provided by the embodiments of the present application have at least the following advantages:

[0030] The embodiments of the present application provide a method and a system for 3D object detection using multi-modal expert knowledge. The method includes the following steps: First, construct an expert model; the expert model is a vision-centered multi-modal expert model; then, construct a trajectory-based distillation module and an occupancy reconstruction module; next, according to the trajectory-based distillation module and the occupancy reconstruction module, transfer the knowledge of the expert model to the apprentice model; the apprentice model is a standard long-term visual detection model; based on the apprentice model after knowledge transfer, detect 3D objects.

[0031] This application provides a method for 3D object detection using multi-modal expert knowledge. First, this application constructs a vision-centered multi-modal expert model that specifically encodes the image modality, thus eliminating the need to use a lidar backbone network. This application demonstrates for the first time that such an expert model can have comparable performance to other state-of-the-art multi-modal methods and is simpler. Due to its homogeneous characteristics and excellent performance, the vision-centered expert model has been proven to be very effective in extracting knowledge for vision-based models. This effect is very significant for various model sizes, from compact to more complex architectures. At the same time, this application also proposes trajectory-based knowledge extraction and occupancy reconstruction modules that supervise static and dynamic objects to mitigate misalignment problems during the long-term temporal fusion process. Combining the constructed expert model, this application improves the performance of vision-based models and achieves state-of-the-art results on the nuScenes validation and test leaderboards. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] One or more embodiments are illustrated by way of example in the accompanying drawings, which do not constitute a limitation on the embodiments unless otherwise stated, and the figures in the drawings do not constitute a scale limitation.

[0033] Figure 1A Schematic diagram of an existing 3D perception distillation process;

[0034] Figure 1B Schematic diagram of misalignment of moving objects during an existing 3D perception distillation process;

[0035] Figure 2 Schematic flowchart of a method for 3D object detection using multi-modal expert knowledge provided by an embodiment of this application;

[0036] Figure 3 Schematic architecture diagram of a method for 3D object detection using multi-modal expert knowledge provided by an embodiment of this application;

[0037] Figure 4 Structural diagram of a system for 3D object detection using multi-modal expert knowledge provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] As can be seen from the background art, in existing 3D perception distillation methods, due to the domain gap between lidar and camera features, which hinders knowledge transfer during the distillation process, the improvement in practical applications is limited.

[0039] In the prior art, in addition to using an expert model with the best performance, another expert model is a large-scale camera-only model. Although the domain gap between the camera-only expert model and the apprentice model is eliminated, due to the lack of accurate geometric information, the expert model performs poorly in terms of effectiveness. Similarly, it fails to bring satisfactory improvements to the apprentice model. Therefore, an ideal expert model should meet two basic requirements: achieving state-of-the-art performance and minimizing the domain gap.

[0040] In addition, the current distillation methods are insufficient in terms of compatibility with long-term temporal fusion in advanced camera-only 3D detectors. Long-term temporal modeling has shown significant potential to improve the accuracy of depth estimation and detection performance, but it introduces the problem of motion mismatch. Previous bird's-eye view distillation methods have adopted two different approaches: either distilling the entire bird's-eye view (BEV) space without sufficient attention to foreground objects, or only distilling the foreground object regions, thus ignoring the motion mismatch problem caused by long-term temporal fusion. As Figure 1B shown, this mismatch occurs when converting past scenes to the current scene coordinates based only on self-motion, assuming all objects are stationary. Figure 1B In, the green rectangles represent true positives, while the pink rectangles represent false positives. Mapping objects in the historical frame to the current timestamp results in incorrect positions in the current frame because it is assumed that the objects are stationary. x i represents different positions of the object at historical timestamps. However, in actual situations, dynamic objects cause mismatches, which interfere with the temporal fusion features. In the case of long-term temporal fusion, this is more challenging. Existing methods, such as Stream PETR, introduce Layer Norm for dynamic object modeling, but the effect of introducing velocity and time variables in the model is relatively small.

[0041] To solve the above technical problems, the present application provides a method and system for 3D object detection using multi-modal expert knowledge. The method includes the following steps: First, construct an expert model; the expert model is a vision-centered multi-modal expert model; then, construct a trajectory-based distillation module and an occupancy reconstruction module; next, transfer the knowledge of the expert model to the apprentice model according to the trajectory-based distillation module and the occupancy reconstruction module; the apprentice model is a standard long-term vision detection model; based on the apprentice model after knowledge transfer, detect 3D objects. By introducing a vision-based detector (VCD), the present application proposes a framework for improving the camera-only apprentice model, including a multi-modal expert suitable for the apprentice and distillation supervision suitable for temporal fusion. By constructing a vision-centered multi-modal expert model VCD-E (VCD-Expert), using the same structure as the camera-only apprentice to reduce feature differences, and using point cloud input as a depth prior to reconstruct the three-dimensional scene, performance comparable to other heterogeneous multi-modal experts is achieved.

[0042] The following will elaborate on each embodiment of the present application in conjunction with the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present application, many technical details are presented to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.

[0043] See Figure 2 , the embodiment of the present application provides a method for 3D object detection using multi-modal expert knowledge, including:

[0044] Step S1, construct an expert model; the expert model is a vision-centered multi-modal expert model.

[0045] Step S2, construct a trajectory-based distillation module and an occupancy reconstruction module.

[0046] Step S3, transfer the knowledge of the expert model to the apprentice model according to the trajectory-based distillation module and the occupancy reconstruction module; the apprentice model is a standard long-term vision detection model.

[0047] Step S4, detect 3D objects based on the apprentice model after knowledge transfer.

[0048] The present application addresses the technical problem in the prior art that "the domain gap between lidar and camera features hinders knowledge transfer during distillation, resulting in limited improvements in practical applications" by proposing a framework for improving the student model, including a distillation-friendly multi-modal expert model and temporal fusion-friendly distillation supervision. The existing 3D perception distillation process is as Figure 1A shown. As Figure 1AAs shown, the existing process requires the support of a camera and a point cloud backbone, while the method of this application does not require a point cloud backbone. Through the depth of point cloud projection, this application can directly convert image features into the BEV space and construct a vision-centered expert model. Figure 1B Shows the misalignment map of moving objects during the existing 3D perception distillation process; existing distillation methods either distill the entire BEV space or only distill the foreground object region, thus ignoring the motion mismatch problem caused by long-term temporal fusion. As Figure 1B Shown, this mismatch phenomenon occurs when converting past scenes to current scene coordinates based only on ego motion, assuming all objects are stationary. However, in reality, dynamic objects can cause mismatches, thus disturbing the temporally fused features. This is more challenging in the case of long-term temporal fusion.

[0049] Motivated by the success of unimodal distillation, an expert model suitable for the apprentice will mainly rely on camera features while achieving performance comparable to that of multimodal models. Therefore, this application constructs a vision-centered multimodal expert model that specifically encodes the image modality, thus eliminating the need to use a lidar backbone network. Due to its homogeneous characteristics and excellent performance, the vision-centered expert model has been proven to be very effective in extracting knowledge from vision-based models. This effect is very significant for various model sizes, from compact to more complex architectures. This application demonstrates for the first time that such an expert model can have performance comparable to other state-of-the-art multimodal methods and is more straightforward.

[0050] This application proposes trajectory-based knowledge extraction and occupancy reconstruction modules that supervise static and dynamic objects to mitigate the misalignment problem during the long-term temporal fusion process. Combining the constructed expert model, this application improves the performance of vision-based models and achieves state-of-the-art results on the nuScenes validation and test leaderboards.

[0051] The method for 3D object detection using multimodal expert knowledge provided by this application is described in detail below. The method of this application includes two main components: (1) a vision-centered expert model; (2) a trajectory-based distillation module and an occupancy reconstruction module. The flow architecture diagram of the method of this application is as Figure 3 shown.

[0052] The expert model and the apprentice model of this application adopt the same model architecture. The only difference is that the expert additionally utilizes the accurate depth map generated from the point cloud, while the apprentice model predicts the depth map from the image. Although the expert model of this application only uses the image backbone network to encode high-level scene information, it is comparable to the state-of-the-art multi-modal fusion methods that use several modality-specific backbone networks and complex interaction strategies. More importantly, this application eliminates the domain gap between the multi-modal expert and the camera-only apprentice model, which is considered one of the most challenging topics in the cross-modal distillation literature.

[0053] As Figure 3 shown, this application constructs a distillation framework between the expert network and the apprentice network. The features extracted by the image backbone network (Image Features) and the temporal depth map (Depth) projected from the lidar point cloud are fused to create a unified BEV representation for 3D object detection. Therefore, although this application adopts a cross-modal method for 3D object detection, the resulting representation is still consistent with the image modality features.

[0054] After obtaining the pre-trained vision-centered expert and the corresponding apprentice network, this application freezes the expert network and uses its intermediate features as the auxiliary supervision for the apprentice network. Since the current state-of-the-art vision-based detectors adopt long-term temporal modeling to achieve state-of-the-art performance, this application uses a standard long-term temporal vision detector based on the bird's-eye view depth map (BEV Depth) as the apprentice model. The expert model also utilizes long-term temporal modeling to ensure consistency and achieve higher performance.

[0055] In some embodiments, constructing the expert model in step S1 includes the following steps:

[0056] Step S101: Project the lidar point cloud scan onto the image to obtain the temporal depth map.

[0057] Step S102: Based on the image features, perform depth prediction for each pixel in the image to obtain the depth distribution of the image features.

[0058] Step S103: Project the image features into the BEV space and obtain the BEV features according to the depth distribution of the image features; the BEV features are bird's-eye view features.

[0059] Step S104: Convert the BEV features of the historical timestamp into the current BEV features.

[0060] Step S105: Fuse all the BEV features with the temporal depth map to create a unified BEV representation for 3D object detection.

[0061] The construction process of the expert model is described in detail below.

[0062] In this application, by integrating point cloud information as an accurate depth map input into a vision-based model, a vision-centered expert model is constructed. The vision-based detector serves as the main model, while the point cloud information supplements it by providing accurate depth information. This approach eliminates the need for complex training strategies or customized fusion modules, simplifying the fusion process.

[0063] For the expert model, this application projects multiple point cloud scans onto an image to obtain the corresponding depth map D. Since the depth map D generated from the point cloud cannot cover every pixel of the image, this application also predicts the depth distribution for each pixel based on image features. Then, this application projects the image features into the BEV space to obtain BEV features according to their depths. In addition, this application transforms the BEV features of the previous timestamp into the current BEV features to model long-term relationships. Here, N represents the time interval between the current frame and the historical frame. Then, the unified BEV features are combined to generate 3D object detection predictions. This expert model is trained using a multi-task loss function, considering 3D detection loss and depth estimation loss.

[0064] See Figure 3 , where the green area is the Expert model and the red area is the Apprentice model. The expert utilizes LiDAR (LiDAR point cloud) data to improve the accuracy of depth estimation before the view transformation in the BEV (Bird's Eye View) pipeline. The Apprentice model represents a standard long-term vision detection model. By establishing occupancy using the depth information of the expert model, this model serves as the supervision for the Apprentice model. The Motion Trajactory is constructed by wrapping the time series of GT (Ground Truth) queries into the current timestamp. Projecting the motion trajectory of each object into the BEV space can correct the misalignment of object motion. With the knowledge transferred from the expert, the Apprentice can perform better than before.

[0065] In the expert model part, first, image features are obtained through Depth Prediction; meanwhile, the expert obtains the temporal depth map (Depth) from LiDAR data before the view transformation in the BEV (Bird's Eye View) pipeline, and then fuses the image features (Image Features) extracted by the image backbone network and the temporal depth map (Depth) projected from the LiDAR point cloud to perform the View Transform to obtain the BEV features (BEVFeatures), that is, by fusing the image features extracted by the image backbone network and the temporal depth map projected from the LiDAR point cloud, a unified BEV representation is created for 3D object detection to improve the accuracy of depth estimation.

[0066] This application elucidates a methodology for overcoming the long-standing modeling limitations in multi-camera 3D object detection. This is achieved by introducing two innovative modules during the distillation process: the trajectory-based distillation module and the occupancy reconstruction module.

[0067] In some embodiments, constructing the trajectory-based distillation module in step S2 includes the following steps:

[0068] Step S201: Construct motion trajectories based on the transformed ground-truth object positions of all historical frames.

[0069] Step S202: Obtain the normalized key sampled features by performing bilinear interpolation sampling on the sampled features and then normalizing them; the sampled features are the sampled features at the same points from the expert BEV features and the apprentice BEV features.

[0070] Step S203: Calculate the trajectory-based distillation loss between the normalized key sampled features; use the motion trajectories as queries and perform trajectory-based distillation at the representative positions corresponding to the motion trajectories, enabling the expert to correct the motion misalignment problems in the apprentice.

[0071] See Figure 3 , for the fine-grained trajectory-based distillation module, this application aims to improve the detection of dynamic objects by focusing on the inconsistent parts of object motion. For the i-th historical frame containing K objects at timestamp t i , extract the j-th ground-truth object position P i j in the ego-coordinate system. Determine the actual ego-motion matrix M i between the current frame t0 and each historical frame t i . Apply the ego-motion transformation matrix M i to the ground-truth object position P i j to obtain the transformed position P i j′ in the current frame coordinate system:

[0072] P i j′ = M i P i j

[0073] This application pools the transformed ground-truth object positions P i j′ of all historical frames to construct motion trajectories. The trajectory of the j-th object can be represented as a sequence of object positions within the current frame.

[0074] Let respectively represent the same points P E on the expert BEV feature F A and the apprentice BEV feature F i j′ of the upsampled features. They are sampled by bilinear interpolation and then the normalized key sampling features are obtained through the calculation formula (1) of normalization processing.

[0075] The calculation formula of normalization processing is:

[0076]

[0077] where respectively represent the same points P E on the expert BEV feature F A and the apprentice BEV feature F i j′ of the upsampled features.

[0078] Then, the trajectory-based distillation loss L TD is calculated between the normalized key sampling features through formula (2).

[0079] The calculation formula of the trajectory-based distillation loss is:

[0080]

[0081] where L TD represents the trajectory-based distillation loss; N represents the time interval between the current frame and the historical frame; K represents the number of objects to be measured.

[0082] Finally, using the motion trajectory as a query, trajectory-based distillation is performed at these representative positions. This method enables the expert to correct the motion misalignment problem in the apprentice.

[0083] For the apprentice model (Apprentice) part and the occupancy reconstruction (Occupancy) module, as Figure 3 shown, before the view transform ((View Transform), the pictures obtained by the multi-view camera (Multi-view Camera) are first subjected to occupancy reconstruction (Reconstruction) to obtain the grid occupancy, which serves as the supervision of the apprentice model, thereby improving the model's ability to recognize the 3D geometric attributes of objects. The grid-based supervision signal can effectively guide the model to improve the prediction accuracy of the object depth. Different voxels within the occupancy structure aggregate the fused depth distributions from different perspective views, thereby enhancing the robustness to depth errors.

[0084] In some embodiments, the occupancy reconstruction module establishes a grid occupancy state based on the depth information of the expert model and supervises the apprentice model based on the grid occupancy state.

[0085] Specifically, constructing the occupancy reconstruction module includes the following steps: First, based on the depth estimation module, predict the depth of each image pixel to obtain a depth map; then, back-project the depth map into a 3D point cloud and convert each image pixel into 3D coordinates; next, expand the Gaussian distribution of each 3D coordinate into 3D space to obtain an accurate 3D object modeling; finally, use the grid of the 3D object modeling as auxiliary supervision and use an intuitive regularization loss to optimize the predicted grid occupancy state, thereby improving the depth prediction ability for static and dynamic objects.

[0086] The vision-centered expert model constructed in this application performs excellently in 3D object detection, thus achieving a more accurate 3D geometric representation of the object. This application uses a depth estimation module to predict the depth of each image pixel (u, v), denoted as D(u, v). Subsequently, the depth map D is back-projected into a 3D point cloud, and each image pixel (u, v) is converted into 3D coordinates (x, y, z) in the following way.

[0087] The conversion formula for converting each image pixel into 3D coordinates is:

[0088]

[0089] where (u, v) represents the image pixel; D(u, v) represents the depth of the image pixel; c u 、c v respectively represent the center points of the camera, and f u 、f v respectively represent the horizontal and vertical focal lengths;

[0090] The occupancy reconstruction module improves the model's ability to recognize the 3D geometric attributes of objects. The grid-based supervision signal can effectively guide the model to improve the prediction accuracy of the object depth. Different voxels within the occupancy structure aggregate the fused depth distributions from different perspective views, thereby enhancing the robustness to depth errors.

[0091] Inspired by CenterPoint, this application expands the Gaussian distribution applied to each target into 3D space to obtain an accurate 3D object modeling to achieve a more focused 3D object modeling.

[0092] Specifically, the grid of the 3D object modeling is:

[0093]

[0094] where (p x ,py , p z ) represents the center of the 3D object, and σ p represents the standard deviation of each object dimension.

[0095] The model uses these grids as auxiliary supervision and adopts an intuitive L1 regularization loss to optimize the predicted grid occupancy state G′ xyz , thereby enhancing the depth prediction ability for static and dynamic objects. Its calculation formula is:

[0096]

[0097] where, L OR represents the occupancy reconstruction loss; G xyz represents the grid for 3D object modeling; G′ xyz represents the predicted grid occupancy state.

[0098] In some embodiments, during the distillation process of transferring the knowledge of the expert model to the apprentice model according to the trajectory-based distillation module and the occupancy reconstruction module, through the joint training loss, the transfer of semantic and geometric knowledge of the expert model to the apprentice model is promoted; the joint training loss L Total is defined as:

[0099] L Total = L A + L TD + L OR (6)

[0100] where, L Total represents the joint training loss; L A represents the perceptual loss of the apprentice model; L TD represents the trajectory-based distillation loss; L OR represents the occupancy reconstruction loss.

[0101] See Figure 4, an embodiment of the present application also provides a system for 3D object detection using multi-modal expert knowledge, including: a model construction module 101 and a detection module 102 connected to each other; wherein, the model construction module 101 includes an expert model construction unit, a trajectory-based distillation module construction unit, and an occupancy reconstruction module construction unit; the detection module 102 includes an apprentice model; the apprentice model is a standard long-term visual detection model; the expert model construction unit is used to construct a vision-centered multi-modal expert model 1011; the trajectory-based distillation module construction unit and the occupancy reconstruction module construction unit are respectively used to construct a trajectory-based distillation module 1012 and an occupancy reconstruction module 1013; the trajectory-based distillation module 1012 and the occupancy reconstruction module 1013 are used to transfer the knowledge of the expert model to the apprentice model; the detection module 102 is used to detect 3D objects according to the apprentice model after knowledge transfer.

[0102] In some embodiments, the trajectory-based distillation module construction unit includes: a motion trajectory construction unit, a normalization processing unit, a calculation unit, and a distillation unit connected in sequence; the motion trajectory construction unit is used to construct a motion trajectory according to the transformed true object positions of all historical frames; the normalization processing unit is used to perform bilinear interpolation sampling on the sampled features and then perform normalization processing to obtain normalized key sampled features; the sampled features are the sampled features at the same points from the expert BEV features and the apprentice BEV features; the calculation unit is used to calculate the trajectory-based distillation loss between the normalized key sampled features; the distillation unit is used to use the motion trajectory as a query and perform trajectory-based distillation at the representative positions corresponding to the motion trajectory to correct the motion misalignment problem of the expert to the apprentice.

[0103] Compared with the prior art, the advantages of the present application are as follows: The present application constructs a vision-centered multi-modal expert model, which specifically encodes the image modality, thus eliminating the need to use a lidar backbone network. The present application proves for the first time that such an expert model can have comparable performance with other state-of-the-art multi-modal methods and is more simple.

[0104] Due to its homogeneous characteristics and excellent performance, the vision-centered expert model has been proven to be very effective in extracting knowledge from vision-based models. This effect is very significant for various model sizes, from compact to more complex architectures.

[0105] The present application proposes a trajectory-based knowledge extraction and occupancy reconstruction module, which supervises static and dynamic objects to alleviate the misalignment problem in the long-term time fusion process. Combined with the constructed expert model, the present application improves the performance of the vision-based model and achieves state-of-the-art results on the nuScenes validation and test leaderboards.

[0106] To verify the method for 3D object detection using multimodal expert knowledge provided in this application, experiments were conducted on a large-scale autonomous driving dataset (nuScenes). This dataset contains 700, 150, and 150 scenes, which are used for training, validation, and testing respectively. Each scene has a duration of approximately 20 seconds.

[0107] By comparing the method for 3D object detection using multimodal expert knowledge provided in this application with other methods, the results are shown in Tables 1, 2, and 3. VCD-A (VCD-Apprentice) achieved the best performance (State Of The Art, SOTA) on most key metrics and exceeded the baseline by 2 points in the nuScenes detection score (NDS). At the same time, VCD-E also achieved comparable performance to the SOTA that mainly relies on the point cloud backbone when only using the image backbone.

[0108] Table 1 Comparison results of the method in this application with other methods

[0109]

[0110] In Table 1, * represents the long-term baseline implemented by this application based on BEVDet 4D-Depth. VCD-A exceeded the previous SOTA by 2 points in terms of NDS and achieved SOTA under the same settings.

[0111] Table 2 Comparison results among camera methods on the nuScenes test set

[0112]

[0113] In Table 2, the method marked with * represents the long-term baseline method implemented by this application based on BEVDet 4D-Depth. represents the test-time augmentation technique adopted in the inference stage. VCD-A achieved SOTA on most key metrics and exceeded the baseline by 2 points in terms of NDS.

[0114] Table 3 Comparison results of multimodal methods on the nuScenes validation set

[0115]

[0116] VCD-E proposed in this application only uses the image backbone and at the same time achieves comparable performance to the state-of-the-art multimodal methods that mainly rely on the LiDAR backbone.

[0117] The experimental results of the ablation study are introduced below.

[0118] Table 4 presents an ablation study of the proposed distillation framework over different time lengths.

[0119]

[0120] The proposed VCD-E in this application only uses an image backbone and achieves performance comparable to the state-of-the-art multi-modal methods that mainly rely on LiDAR backbones. This design is applicable to time windows of different lengths, and the benefits increase with the increase in time length.

[0121] Table 5 Comparison of the improvement of apprentice performance by different experts

[0122]

[0123] In Table 5, CM represents cross-modal and UM represents unimodal. The performance improvement obtained by the apprentice is brought by different experts. As can be seen from Table 5, the success of unimodal distillation still exists.

[0124] Table 6 Training effects of the method in this application on different models

[0125]

[0126] In Table 6, all models are trained with the proposed VCD-E in this application as the expert. The method in this application significantly outperforms the previous SOTA methods.

[0127] Table 7 Comparison of the benefits of using different image backbone networks in multi-modal models

[0128]

[0129] From the benefit comparison data in Table 7, it can be seen that in multi-modal models, using a stronger backbone network can exhibit better performance.

[0130] For the above technical solutions, the embodiments of this application provide a method and system for 3D object detection using multi-modal expert knowledge. The method includes the following steps: First, construct an expert model; the expert model is a vision-centered multi-modal expert model; then, construct a trajectory-based distillation module and an occupancy reconstruction module; next, transfer the knowledge of the expert model to the apprentice model according to the trajectory-based distillation module and the occupancy reconstruction module; the apprentice model is a standard long-term vision detection model; based on the apprentice model after knowledge transfer, detect 3D objects.

[0131] The present application provides a method for 3D object detection using multi-modal expert knowledge. First, the present application constructs a vision-centric multi-modal expert model (VCD-E), which specifically encodes the image modality, thus eliminating the need to use a lidar backbone network. The present application demonstrates for the first time that such an expert model can have comparable performance to other state-of-the-art multi-modal methods and is simpler. Due to its homogeneous characteristics and excellent performance, the vision-centric expert model has been proven to be very effective in extracting knowledge into vision-based models. This effect is very significant for various model sizes, from compact to more complex architectures. At the same time, the present application also proposes trajectory-based knowledge extraction and occupancy reconstruction modules, which supervise static and dynamic objects to mitigate misalignment problems during the long-term temporal fusion process. Combining the constructed expert model, the present application improves the performance of vision-based models and achieves state-of-the-art results on the nuScenes validation and test leaderboards.

[0132] Those of ordinary skill in the art can understand that the above embodiments are specific examples for implementing the present application. In actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application. Any person skilled in the art can make their own changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.

Claims

1. A method for 3D object detection using multi-modal expert knowledge, characterized in that, It includes the following steps: Construct an expert model; the expert model is a vision-centered multi-modal expert model; Construct a trajectory-based distillation module and an occupancy reconstruction module. Constructing the occupancy reconstruction module includes the following steps: Based on a depth estimation module, predict the depth of each image pixel to obtain a depth map; back-project the depth map into a 3D point cloud and convert each image pixel into 3D coordinates; expand the Gaussian distribution of each 3D coordinate into 3D space to obtain an accurate 3D object modeling; use the mesh of the 3D object modeling as auxiliary supervision and use an intuitive L1 regularization loss to optimize the predicted mesh occupancy state, thereby improving the depth prediction ability for static and dynamic objects. Among them, use an intuitive L1 regularization loss to optimize the predicted mesh occupancy state to obtain an occupancy reconstruction loss. Its calculation formula is: Among them, L OR represents the occupancy reconstruction loss; G xyz represents the mesh for 3D object modeling; G' xyz represents the predicted mesh occupancy status; According to the trajectory-based distillation module and the occupancy reconstruction module, transfer the knowledge of the expert model to the apprentice model; the apprentice model is a standard long-term visual detection model; Based on the apprentice model after knowledge transfer, detect 3D objects.

2. The method for 3D object detection using multi-modal expert knowledge according to claim 1, wherein The construction of the expert model includes the following steps: Project the lidar point cloud scan onto the image to obtain a temporal depth map; Based on the image features, perform depth prediction on each pixel in the image to obtain the depth distribution of the image features; Project the image features into the BEV space and obtain BEV features according to the depth distribution of the image features; the BEV features are bird's-eye view features; Convert the BEV features of historical timestamps into the current BEV features; Fuse all the BEV features with the temporal depth map to create a unified BEV representation for 3D object detection.

3. The method for 3D object detection using multi-modal expert knowledge according to claim 1, characterized in that, The construction of the trajectory-based distillation module includes the following steps: Construct a motion trajectory based on the converted true object positions of all historical frames; Through bilinear interpolation sampling of the sampled features and then normalization processing, obtain normalized key sampled features; the sampled features are the sampled features at the same points from the expert BEV features and the apprentice BEV features; Calculate the trajectory-based distillation loss between the normalized key sampled features; Use the motion trajectory as a query and perform trajectory-based distillation at the representative positions corresponding to the motion trajectory to enable the expert to correct the motion misalignment problem in the apprentice.

4. The method for 3D object detection using multi-modal expert knowledge according to claim 3, characterized in that, The calculation formula for the normalization processing is: Among them, respectively represent the common points of the expert BEV feature F E and the apprentice BEV feature F A ; the upsampled features; The calculation formula for the trajectory-based distillation loss is: Among them, L TD represents the trajectory-based distillation loss; N represents the time interval between the current frame and the historical frame; K represents the number of objects to be measured.

5. The method for 3D object detection using multi-modal expert knowledge according to claim 1, characterized in that, The occupancy reconstruction module establishes a mesh occupancy state according to the depth information of the expert model and supervises the apprentice model based on the mesh occupancy state.

6. The method for 3D object detection using multi-modal expert knowledge according to claim 1, characterized in that, The conversion formula for converting each image pixel into 3D coordinates is: Among them, (u, v) represents an image pixel; D(u, v) represents the depth of the image pixel; c u , c v respectively represent the center points of the camera, and f u , f v respectively represent the horizontal and vertical focal lengths; Expand the Gaussian distribution of each 3D coordinate into 3D space to obtain an accurate 3D object modeling; the mesh of the 3D object modeling is: Among them, (p x , p y , p z ) represents the center of the 3D object, and σ p represents the standard deviation of each object dimension; 7. The method for 3D object detection using multi-modal expert knowledge according to claim 1, wherein During the process of transferring the knowledge of the expert model to the apprentice model according to the trajectory-based distillation module and the occupancy reconstruction module, through the joint training loss, promote the transfer of semantic and geometric knowledge of the expert model to the apprentice model; The combined training loss L Total is defined as: L Total = L A + L TD + L OR (6) Among them, L Total represents the joint training loss; L A represents the perception loss of the apprentice model; L TD represents the trajectory-based distillation loss; L OR represents the occupancy reconstruction loss.

8. A system for 3D object detection using multi-modal expert knowledge, characterized in that, It includes: A model construction module and a detection module connected to each other; among them, The model construction module includes an expert model construction unit, a trajectory-based distillation module construction unit, and an occupancy reconstruction module construction unit; the detection module includes an apprentice model; the apprentice model is a standard long-term visual detection model; The expert model construction unit is used to construct a vision-centered multi-modal expert model; The trajectory-based distillation module construction unit and the occupancy reconstruction module construction unit are respectively used to construct a trajectory-based distillation module and an occupancy reconstruction module. The steps of constructing the occupancy reconstruction module include: based on the depth estimation module, predicting the depth of each image pixel to obtain a depth map; back-projecting the depth map into a 3D point cloud and converting each image pixel into 3D coordinates; expanding the Gaussian distribution of each 3D coordinate into 3D space to obtain an accurate 3D object modeling; using the mesh of the 3D object modeling as auxiliary supervision and using an intuitive L1 regularization loss to optimize the predicted mesh occupancy state, so as to improve the depth prediction ability for static and dynamic objects. Among them, using the intuitive L1 regularization loss to optimize the predicted mesh occupancy state to obtain an occupancy reconstruction loss; its calculation formula is: Among them, L OR represents the occupancy reconstruction loss; G xyz represents the mesh for 3D object modeling; G' xyz represents the predicted mesh occupancy status; the trajectory-based distillation module and the occupancy reconstruction module are used to transfer the knowledge of the expert model to the apprentice model; The detection module is used to detect 3D objects according to the apprentice model after knowledge transfer.

9. The system for 3D object detection using multi-modal expert knowledge according to claim 8, wherein, The trajectory-based distillation module construction unit includes: a motion trajectory construction unit, a normalization processing unit, a calculation unit, and a distillation unit connected in sequence; The motion trajectory construction unit is used to construct a motion trajectory according to the converted real object positions of all historical frames; The normalization processing unit is used to perform bilinear interpolation sampling on the sampled features and then perform normalization processing to obtain normalized key sampled features; the sampled features are the sampled features at the same points from the expert BEV features and the apprentice BEV features; The calculation unit is used to calculate the trajectory-based distillation loss between the normalized key sampled features; The distillation unit is used to use the motion trajectory as a query to perform trajectory-based distillation at the representative positions corresponding to the motion trajectory, so that the expert corrects the motion misalignment problem in the apprentice.

Citation Information

Patent Citations

  • Multi-modal feature fusion-based incomplete shape symmetry prediction method and system

    CN115984364A

  • Around view camera model training method based on pre-training distillation

    CN116012797A