3D object detection method based on 3D point labeling multi-modal weak supervision learning
Patent Information
- Application Number
- CN202410163560.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-05
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-02-05
AI Technical Summary
[0005]这些背景技术存在的问题和局限性主要包括:单模式方法无法充分利用多模态数据的优势,而多模式方法虽然融合了不同模态的数据,但如何有效结合带有点先验的图像信息尚未得到深入探索
[0027]1.更高的性能提升:本发明(Point-DETR3D)在不同检测范围内相比于Point-DETR(P-DETR)展示出更高的性能提升。具体来说,Point-DETR3D在SPNDS和mAP指标上分别取得了最高130%和110%的提升,尤其在远距离目标检测中表现显著。
Smart Images

Figure CN117953205B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of 3D detection in autonomous driving, specifically to a multimodal weakly supervised learning 3D target detection method based on 3D point annotation. Background Technology
[0002] Single-modal and multi-modal 3D object detection methods: Existing 3D object detection methods are mainly divided into single-modal and multi-modal types. Single-modal methods use only point clouds (e.g., VoxelNet and PointPillar) or images as input. Multi-modal methods (such as TransFusion, BEVFusion, and DeepInteraction) improve 3D detection performance by fusing data from different modalities. While these methods are effective, how to utilize image information with point priors remains an underexplored area.
[0003] Semi-supervised 3D object detection: While supervised 3D object detection methods have achieved promising results, their reliance on large amounts of accurate annotations makes deployment in real-world scenarios difficult. Semi-supervised 3D detection offers a potential solution to alleviate the annotation cost problem, requiring only a small amount of labeled data. For example, the Mean Teacher-based SESS method and DetMatch employ a flexible framework for joint semi-supervised learning to generate more robust pseudo-labels.
[0004] Weakly supervised and semi-supervised object detection: Previous studies have typically used image-level annotations as weak supervision. However, the lack of location information severely impacts model performance. Recent research has proposed utilizing point annotations as weak supervision signals, such as WS3D and Point-Teaching. By using a small number of weakly annotated scenes and some precisely annotated object instances, competitive performance has been achieved.
[0005] The main problems and limitations of these background technologies include: single-modal methods cannot fully utilize the advantages of multimodal data, while multimodal methods, although fusing data from different modalities, have not yet explored how to effectively combine image information with point priors. In the semi-supervised and weakly supervised fields, existing methods still have room for improvement in handling the problem of missing location information, especially when using point annotations as weak supervision signals; improving the quality and accuracy of pseudo-labels remains a challenge. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides a multimodal weakly supervised learning 3D object detection method based on 3D point annotation. The main technical challenge of this invention is solving the fundamental task of 3D object detection in autonomous driving perception. Although current advanced 3D object detectors have made significant progress, they typically require training in multiple scenarios and precise 3D annotations to clearly define the object's position, size, and orientation. However, this manual annotation process is time-consuming and costly, especially when seven degrees of freedom (DoF) are needed to describe 3D objects. Therefore, this invention aims to reduce the reliance on large amounts of precise 3D annotations through innovative methods while maintaining or improving the performance of 3D object detection.
[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0008] A multimodal weakly supervised learning method for 3D object detection based on 3D point annotation is proposed. The method inputs radar point clouds and images containing 3D objects into a trained object detection model to obtain detection boxes and categories for the 3D objects. The training process of the object detection model includes the following steps:
[0009] Step 1: Collect multimodal data of 3D targets. Multimodal data includes radar point clouds and images. Point annotations are performed on the 3D targets in the radar point clouds of all multimodal data to obtain weakly annotated data. Boundary annotations are performed on the 3D targets in the radar point clouds of a subset of multimodal data to obtain fully annotated data. Point annotations include the spatial coordinates (x, y, z) of the annotated points of the 3D target, and the category c of the 3D target. Boundary annotations include the detection boxes surrounding the 3D target and the category c of the 3D target.
[0010] Step two, construct the teacher model, using fully labeled data and a corresponding number of weakly labeled data as the first training data to train the teacher model, specifically including:
[0011] The teacher model includes an image feature extractor, a first point cloud feature extractor, a 3D point encoder, an attention module, and a first detection head;
[0012] The images in the first training data are input into the image feature extractor to obtain image features; the radar point cloud in the first training data is input into the first point cloud feature extractor to obtain BEV features; each labeled point in the weakly labeled data is input into the 3D point encoder to obtain instance queries; the image features, BEV features, and instance queries are input into the attention module, and the RoI features of the image features and BEV features are fused through a cross-modal deformable attention mechanism. The final RoI features are input into the first detection head to obtain the pseudo detection boxes of the 3D targets predicted by the teacher model; the bounding boxes in the fully labeled data are used as ground truth to train the teacher model.
[0013] Step 3: Input the weakly labeled data other than the weakly labeled data of the first training data into the teacher model that has been trained to obtain the pseudo-detection boxes of the 3D targets in the weakly labeled data; combine the pseudo-detection boxes and the corresponding weakly labeled data to form pseudo-labeled data.
[0014] Step four: Construct a student model, which includes a second point cloud feature extractor and a second detection head; use fully labeled data and pseudo-labeled data as the second training data, and adopt a point-guided self-supervised learning method to perform standard comparison learning training on the student model; the completed student model is the target detection model.
[0015] Furthermore, the image feature extractor employs a ResNet50 network.
[0016] Furthermore, the first point cloud feature extractor and the second point cloud feature extractor adopt the Point Pillar model.
[0017] Further, in step two, the image features, BEV features, and instance queries are input into the attention module. The RoI features of the image features and BEV features are fused through a cross-modal deformable attention mechanism. The resulting final RoI features are then input into the first detection head to obtain the pseudo-detection box of the 3D target predicted by the teacher model. Specifically, this includes:
[0018] The input to the cross-modal deformable attention mechanism includes instance queries, image features, and BEV features. The instance query first projects the spatial coordinates of the points onto the image feature plane and the BEV feature plane respectively, generating a square kernel. Then, each position in the kernel is deformed through a learnable parameter to obtain the RoI features of the image features and the RoI features of the BEV features. Attention scores are calculated for the RoI features obtained at all positions to obtain the RoI features of the image features and the RoI features of the BEV features. The RoI features of the image features and the RoI features of the BEV features are fused across modal features through an instance-level self-attention mechanism to obtain the final RoI features. The final RoI features are then passed through the first detection head to obtain the pseudo-detection box of the 3D target predicted by the teacher model.
[0019] Furthermore, in step four, the fully labeled data and pseudo-labeled data are used as the second training data. A point-guided self-supervised learning method is employed to perform standard comparison training on the student model, specifically including:
[0020] Set up a first learning path and a second learning path; in the first learning path, the parameters of the student model are fixed; in the second learning path, the parameters of the student model are in a trainable state.
[0021] In the first learning path, the radar point cloud of the pseudo-labeled data is enhanced by the first amplitude and then input into the second point cloud feature extractor to obtain the first BEV feature.
[0022] In the second learning path, the radar point cloud of the pseudo-labeled data is enhanced by a second amplitude and then input into the second point cloud feature extractor to obtain the second BEV feature.
[0023] The first BEV feature and the second BEV feature are subjected to rotation and translation inverse transformation and unified in the original point cloud coordinate system;
[0024] The first BEV feature and the second BEV feature are reparameterized using a point-guided mask based on labeled points, and the reparameterized first BEV feature and second BEV feature are fed into the contrastive learning loss function to calculate the loss function, thereby achieving contrastive learning supervision.
[0025] Simultaneously, the first BEV feature and the second BEV feature, unified in the original point cloud coordinate system, are input into the second detection head to obtain the detection box and category of the 3D target predicted by the student model. The loss function is calculated with the ground truth to train the student model.
[0026] Compared with the prior art, the beneficial technical effects of the present invention are:
[0027] 1. Greater performance improvement: This invention (Point-DETR3D) demonstrates a greater performance improvement compared to Point-DETR (P-DETR) across different detection ranges. Specifically, Point-DETR3D achieves improvements of up to 130% and 110% in SPNDS and mAP metrics, respectively, with particularly significant performance in long-range target detection.
[0028] 2. Training process stability: Point-DETR3D employs a one-to-one label assignment strategy, which stabilizes the training process compared to the original Hungarian matching method. This approach simplifies the learning challenge, allowing the model to compete with methods using six decoders even when using only one decoder.
[0029] 3. Instance-level fusion effectiveness: By introducing a cross-modal fusion operation for point-centered deformable RoIs, Point-DETR3D effectively combines 3D priors and 2D data, improving instance-level fusion performance. This fusion strategy is particularly crucial for object detection in distant regions, significantly improving detection accuracy and robustness.
[0030] The aforementioned advantages demonstrate the significant advancements of Point-DETR3D in addressing key challenges in 3D object detection, particularly in weakly supervised and semi-supervised learning environments. It reduces reliance on precise 3D annotations while simultaneously improving the performance and stability of 3D object detection. Ultimately, using only 5% of fully annotated 3D bounding boxes, the student model in this invention achieves 90% of the performance of a fully supervised 3D detector. This indicates that Point-DETR3D of this invention can significantly reduce the workload of annotating 3D bounding boxes. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the target detection method in this invention;
[0032] Figure 2 This is a schematic diagram of the cross-attention flow of the deformable RoI in this invention;
[0033] Figure 3 This is a schematic diagram comparing the SPNDS, mAP, and performance improvement of the target detection method (Point-DETR3D) and Point-DETR (P-DETR) in different detection ranges in this invention;
[0034] Figure 4 This diagram illustrates the comparison of mAP between the target detection method (Point-DETR3D) in this invention and Point-DETR, as well as traditional fully supervised methods, on teacher and student models. (a) indicates that the teacher model has an advantage on the Point-DETR baseline; (b) indicates that the student model achieves performance comparable to the 100% fully supervised paradigm on the Center-Point baseline using only 10% of the fully labeled data.
[0035] Figure 5 This is a visual comparison diagram of the target detection method (Point-DETR3D) and Point-DETR (P-DETR) in this invention;
[0036] Figure 6 This is a visual comparison diagram of the target detection method (Point-DETR3D) and Point-DETR (P-DETR) in this invention. Detailed Implementation
[0037] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.
[0038] In this invention, 2D represents two-dimensional and 3D represents three-dimensional.
[0039] Box annotation refers to labeling 3D objects using boxes. Point annotation refers to labeling 3D objects using points instead of boxes.
[0040] ROI (region of interest) represents the region of interest.
[0041] The object detection method (Point-DETR3D) in this invention significantly improves the performance of 3D object detection models by effectively utilizing spatial point priors and image data. The main framework includes two key components: a point-centered teacher model to generate accurate 3D object pseudo-detection boxes, and a self-motivated student model utilizing point annotations. Extensive experiments demonstrate that this method significantly improves upon existing methods in scenarios with limited labeled data. Point-DETR3D shows great potential in reducing the annotation workload of 3D detection boxes.
[0042] First, this invention obtains a large number of 3D point annotations through manual annotation. Each 3D point annotation includes the spatial coordinates of an object and its category label. Each 3D point annotation is represented as:
[0043] (x, y, z, c);
[0044] Where (x, y, z) and c represent the position and category label of the labeled point in 3D space, respectively. It's important to note that the point labels are randomly sampled within the 3D box, following a normal distribution prior. This method can significantly alleviate the enormous burden of accurately annotating objects in 3D tasks.
[0045] Secondly, this invention trains a teacher model that generates 3D bounding boxes using 3D points, i.e., a point-centered teacher model. The technical implementation details of the main components of the teacher model are as follows:
[0046] 1. Explicit Location Query Initialization: While Point-DETR encodes coordinate and class information using point labels through a learnable projector, this approach offers only minor improvements when extended to the 3D domain. This limitation is likely due to the significant increase in the candidate space when transitioning from 2D to 3D. To address this issue, this invention employs an explicit location query initialization strategy instead of implicitly embedding prior information through a 3D point encoder. Specifically, this invention establishes an explicit binding between the spatial coordinates of the labeled points and each instance query. Thanks to the versatility of Object-DGCNN, this binding can be easily achieved using labeled points. To this end, n instance queries can be generated, each with its labeled points initialized as n ground truth (GT) point labels. Since there is a direct association between each instance query and the GT point labels, the originally complex bipartite matching can be replaced with a simple one-to-one matching based on this initialization, which significantly reduces the training instability observed in Point-DETR.
[0047] 2. Deformable RoI Cross-Modal Fusion: Despite obtaining information through 3D point priors, this invention empirically observes that the teacher model produces suboptimal results in distant regions. In fact, even skilled human experts struggle to determine the exact boundaries of target instances using only a few points. To compensate for this information deficiency, this invention introduces dense image data as a reference. Previous work has demonstrated significant performance improvements in 3D object detection by combining 2D and 3D data. However, most work focuses on general cross-modal fusion at the voxel or point level, while effective instance-level fusion with 3D priors has been rarely explored.
[0048] This invention proposes a novel point-centered deformable RoI cross-modal fusion operation (see...). Figure 2 It seamlessly aggregates instance-level features guided by 3D priors through point sampling of RoI regions and dynamically determines the reference of image regions, especially in distant regions. Specifically, given a labeled point... First, the radar coordinates of each camera are projected onto different image views using the projection matrices of each camera. The rotation and translation matrices from the radar coordinates to the camera coordinates of the j-th image view are respectively... and The intrinsic parameters of the j-th camera are represented as follows: The projection point of the radar onto the camera can be written as:
[0049]
[0050] Among them, the projection point Also known as the image coordinates of the labeled points. u represents the coordinates of a 3D labeled point in the image coordinate system. i vi These represent the x and y coordinates of the labeled point in the pixel coordinate system, respectively.
[0051] To obtain RoI reference points exist A square grid is generated around the area. Here, k represents the index of the RoI reference point in the RoI region, and K is the size of the square grid. Unlike previous methods, this invention does not simply use reference points to sample RoI features. Instead, this invention introduces a learnable offset Δ for each RoI reference point. ik It dynamically determines the aggregation of the most informative neighboring features (obtaining the coordinates of neighboring points through the RoI center point and offset, and then acquiring the features of that point), and is not limited to a fixed RoI region. The Deformable RoI Cross-Attention (DRoICA) process can be represented as:
[0052]
[0053] Where m represents the number of heads in the multi-head attention mechanism, and l represents the input feature level. Q i Represents an instance query, where F is the coordinate with the RoI centroid. Related multi-level image features; MSDeformAttn represents the multi-scale variability cross-attention mechanism; Let represent the k-th coordinate point in the i-th RoI within the l-th feature of the m-th head; Δmlik represents the offset of the k-th coordinate point in the i-th RoI within the l-th feature of the m-th heads. For BEV feature maps, this invention follows the same process to generate RoI meshes for each instance query in the 2D plane. Then, this invention applies deformable RoI cross-attention to obtain instance features F from camera and LiDAR features. l To preserve the spatial localization information from LiDAR RoI features and the texture information from image RoI features, the corresponding features of the two modalities interact at the instance level through the proposed cross-attention module.
[0054] Next, this invention uses the trained teacher model to generate 3D pseudo-detection boxes for all 3D point annotations. Finally, this invention uses the 3D pseudo-detection boxes to train a student model, namely a self-motivated student model. The technical implementation details of the student model are as follows:
[0055] In the original Point-DETR, the student model is trained only on a combination of fully labeled data and pseudo-labeled data generated by the teacher model. However, this approach ignores the potential benefits that the student model can gain from point labels during training. To maximize the use of prior information about points and enhance the final performance of the model, this invention proposes a simple and parameter-free self-supervised method called point center feature invariant learning. It focuses specifically on mitigating the impact of label noise and strengthening the robustness of the model representation. Drawing on previous contrastive learning methods, this invention adopts a standard contrastive learning training paradigm as the basic framework of this invention (see...). Figure 1 For a given input x, it is fed into two randomly augmented pipelines, corresponding to weak and strong augmentations respectively, producing two distinct inputs x1 and x2. x1 and x2 are processed by a feature extractor to derive BEV features. Subsequently, this invention reverses the BEV features according to the augmentation pipelines to generate corresponding paired features with consistency regularization. Previous CL work directly forced two models to have equal supervised mimicry feature maps:
[0056]
[0057] Where H and W represent the width and height of the image feature map, and ||·||² is the L2 norm. However, considering the characteristics of LiDAR points, there is a large amount of blank space in the BEV plane. Supervision on these background features does not contribute to model learning and may even degrade the final performance. Therefore, this invention utilizes point annotations as foreground region cues to guide the learning location.
[0058] Specifically, this invention is based on the annotation (x) at each point in the BEV space. i y i Generate a Gaussian distribution:
[0059]
[0060] Where σ i It is a constant (default value is 2). The x and y coordinates represent the Gaussian kernel. Since the feature map is class-agnostic, this invention treats all m... i,x,y Merge them into a single mask M. For different m at the same position i,x,y For overlapping regions, this invention simply takes their maximum values. After generating the mask M, this invention uses it to guide students in mimicking features for dense feature contrast learning:
[0061]
[0062] This point-guided reweighting strategy allows the model to focus on the foreground region from the teacher, while avoiding the useless imitation of 3D background noise in too many background regions.
[0063] In this embodiment, the specific implementation steps are as follows:
[0064] S1, Acquisition of 3D Point Annotations
[0065] First, 3D point annotation data is collected, with each annotation consisting of a labeled point and its category label. These point annotations are randomly sampled in 3D space, following a normal distribution prior.
[0066] S2, Construction of the Point-Centered Teacher Model
[0067] A point-centered teacher model is constructed, which uses an explicit location query initialization strategy. This means that each instance query is directly bound to its absolute position in 3D space. This approach helps improve the model's performance in distant object detection.
[0068] The teacher model introduces a point-centered deformable RoI cross-modal fusion operation, which aggregates instance-level features guided by 3D priors through point sampling of the RoI region and dynamically determines the reference of the image region.
[0069] The limited amount (10%, 5%, 2%) of fully labeled data and the corresponding amount (10%, 5%, 2%) of point-labeled data are fed into the teacher model as training data to train the teacher model. The remaining point-labeled data is then fed into the teacher model trained in the first stage to infer pseudo-labeled data.
[0070] The input data for the teacher model consists of images, radar point clouds, point annotations, and bounding box annotations. The images and radar point clouds are used as sensor input data, and image features and BEV features are obtained through the corresponding image feature extractor (ResNet50) and radar point cloud feature extractor (Point Pillar), respectively. Each point annotation data is encoded into an instance query through a 3D encoder. The fully annotated data is used as ground truth (GT) to supervise the training of the teacher model.
[0071] S3, Training of a Self-Motivated Student Model
[0072] A student model is trained using a small amount of fully labeled data and all pseudo-labeled data. The student model is the model of the conventional method, so all comparisons should mainly focus on the performance of the student model.
[0073] The student model is trained using a point-guided self-supervised learning approach, which involves using a standard contrastive learning training paradigm to generate different inputs through weak and strong data augmentation, and then processing them through a feature extractor to produce BEV features.
[0074] The student model's input consists of radar point clouds, point annotations, and bounding box annotations. The radar point clouds serve as sensor input data; point annotations are used to generate point guidance masks; and bounding box annotations serve as ground truth values to supervise model training. Figure 1 As shown, specifically in the student model, the model parameters are first duplicated. The parameters on the left are fixed, while the parameters on the right are set to a trainable state. The left side performs weak augmentation (small rotations and translations) on the input point cloud data, while the right side performs strong augmentation (large rotations and translations). Then, a radar point cloud feature extractor (Point Pillar) is used to obtain BEV features in two different coordinate systems. These are then unified to the original point cloud coordinate system through inverse rotation and translation transformations. Next, the two BEV features are reparameterized using point-guided masks, and the reparameterized BEV features are fed into a contrastive learning loss function for loss calculation and contrastive learning supervision. Simultaneously, similar to the normal model training process, the BEV features are fed into a detector to obtain the model's predicted 3D bounding boxes, which are then compared with the ground truth values to calculate the loss function and train the model.
[0075] Model training and evaluation:
[0076] All models were trained on a machine with eight NVIDIA A100 GPUs, with both teacher and student models undergoing 20 training epochs. The AdamW optimizer was used to optimize the overall network, with cosine annealing as the learning strategy, an initial learning rate of 1e-4, and a batch size of 2. The MMDetection3D toolkit was used as the codebase. SPNDS and mAP were used as evaluation metrics for model performance. To quantitatively assess the effectiveness of this invention, it was compared with methods such as Point-DETR on datasets such as nuScenes. In the quantitative analysis, this invention used only 5% of the fully labeled 3D detection boxes, and the student model of this invention achieved 90% of the performance of a fully supervised 3D detector, indicating that the method of this invention can significantly reduce the workload of 3D detection box annotation.
[0077] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.
[0078] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multimodal weakly supervised learning 3D object detection method based on 3D point annotation, characterized in that, The radar point cloud and image containing 3D targets are input into the trained target detection model to obtain the detection boxes and categories of the 3D targets. The training process of the target detection model includes the following steps: Step 1: Acquire multimodal data of 3D targets. Multimodal data includes radar point clouds and images. Point annotations are performed on the 3D targets in the radar point clouds of all multimodal data to obtain weakly annotated data. Boundary annotations are performed on the 3D targets in the radar point clouds of a subset of multimodal data to obtain fully annotated data. Point annotations include the spatial coordinates of the annotated points of the 3D targets. and the categories of 3D targets. The bounding box annotation includes the detection box surrounding the 3D target and the category of the 3D target. ; Step two, construct the teacher model, using fully labeled data and a corresponding number of weakly labeled data as the first training data to train the teacher model, specifically including: The teacher model includes an image feature extractor, a first point cloud feature extractor, a 3D point encoder, an attention module, and a first detection head; The images from the first training data are input into the image feature extractor to obtain image features; the radar point cloud from the first training data is input into the first point cloud feature extractor to obtain BEV features; each labeled point in the weakly labeled data is input into the 3D point encoder to obtain instance queries; the image features, BEV features, and instance queries are input into the attention module, and the RoI features of the image features and BEV features are fused through a cross-modal deformable attention mechanism. The resulting RoI features are then input into the first detection head to obtain the pseudo-detection boxes of the 3D targets predicted by the teacher model. The inputs to the cross-modal deformable attention mechanism include instance queries, image features, and BEV features; instance queries... First, the spatial coordinates of the points are projected onto the image feature plane and the BEV feature plane respectively, generating a square kernel. Then, each position in the kernel is deformed through a learnable parameter to obtain the RoI features of the image features and the RoI features of the BEV features. Attention scores are calculated on the RoI features obtained at all positions to obtain the RoI features of the image features and the RoI features of the BEV features. The RoI features of the image features and the RoI features of the BEV features are fused across modal features through an instance-level self-attention mechanism to obtain the final RoI features. The bounding box labels in the fully labeled data are used as ground truth to train the teacher model. Step 3: Input the weakly labeled data other than the weakly labeled data of the first training data into the teacher model that has been trained to obtain the pseudo-detection boxes of the 3D targets in the weakly labeled data; combine the pseudo-detection boxes and the corresponding weakly labeled data to form pseudo-labeled data. Step four: Construct a student model, which includes a second point cloud feature extractor and a second detection head; use fully labeled data and pseudo-labeled data as the second training data, and adopt a point-guided self-supervised learning method to perform standard comparison learning training on the student model; the completed student model is the target detection model.
2. The multimodal weakly supervised learning 3D target detection method based on 3D point annotation according to claim 1, characterized in that, The image feature extractor uses a ResNet50 network.
3. The multimodal weakly supervised learning 3D target detection method based on 3D point annotation according to claim 1, characterized in that, The first and second point cloud feature extractors use the Point Pillar model.
4. The multimodal weakly supervised learning 3D target detection method based on 3D point annotation according to claim 1, characterized in that, In step four, the fully labeled data and pseudo-labeled data are used as the second training data. A point-guided self-supervised learning method is employed to perform standard comparison training on the student model, specifically including: Set up a first learning path and a second learning path; in the first learning path, the parameters of the student model are fixed; in the second learning path, the parameters of the student model are in a trainable state. In the first learning path, the radar point cloud of the pseudo-labeled data is enhanced by the first amplitude and then input into the second point cloud feature extractor to obtain the first BEV feature. In the second learning path, the radar point cloud of the pseudo-labeled data is enhanced by the second amplitude and then input into the second point cloud feature extractor to obtain the second BEV feature; both the first amplitude enhancement and the second amplitude enhancement include rotation and translation transformation, and the rotation angle and translation transformation distance corresponding to the second amplitude are greater than the rotation angle and translation transformation distance corresponding to the first amplitude, respectively. The first BEV feature and the second BEV feature are subjected to rotation and translation inverse transformation and unified in the original point cloud coordinate system; The first BEV feature and the second BEV feature are reparameterized using a point-guided mask based on labeled points, and the reparameterized first BEV feature and second BEV feature are fed into the contrastive learning loss function to calculate the loss function, thereby achieving contrastive learning supervision. Simultaneously, the first BEV feature and the second BEV feature, unified in the original point cloud coordinate system, are input into the second detection head. The detection boxes and categories of the 3D targets predicted by the student model are obtained, and the loss function is calculated by comparing them with the ground truth values formed by the box annotations in the fully labeled data, and the student model is trained.