Three-dimensional bounding box processing method, device and equipment based on self-supervised learning

CN119723556BActive Publication Date: 2026-08-07ZHEJIANG WUWEN ZHIXING TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG WUWEN ZHIXING TECHNOLOGY CO LTD
Filing Date
2024-11-20
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]相关技术中,大多是采用后融合或交叉验证的方式实现3D检测框的自动标注,在不同的场景下后融合包含无穷无尽的场景特定规则,交叉验证无法真正明确到底应该相信点云检测结果还是图像检测结果,即便可以保证模型结果正确,不同模态的对齐问题也会引发标注结果的不一致,需要进行人工核对

Benefits of technology

[0020] The self-supervised learning-based 3D detection box processing method provided in this application embodiment acquires image iteration data and point cloud iteration data corresponding to the environmental scene of the intelligent vehicle at the current moment during the driving process of the intelligent vehicle. The image iteration data is used to describe the multi-view image data of the environmental scene collected by the intelligent vehicle at the current moment, and the point cloud iteration data is used to describe the laser point cloud data of the environmental scene collected by the intelligent vehicle. Based on the image iteration data, the first annotation information of the 3D detection box corresponding to the detection target is determined to obtain the 2D image contour coordinates of the detection target. The 2D image contour coordinates are used to describe the 2D boundary contour coordinates of the 3D detection box corresponding to the detection target determined by the image iteration data. Based on the point cloud iteration data, the second annotation information of the 3D detection box corresponding to the detection target is determined to obtain the 2D point cloud contour coordinates of the detection target. The 2D point cloud contour coordinates are used to describe the 2D boundary contour coordinates of the 3D detection box corresponding to the detection target determined by the point cloud iteration data. Self-supervised learning is performed on the 2D image contour coordinates and the 2D point cloud contour coordinates to obtain the target annotation information of the 3D detection box corresponding to the detection target. The target annotation information of the 3D detection box corresponding to the detection target includes: the center point coordinates, length, width, height, and orientation of the 3D detection box. Thus, by combining multi-view image data and laser point cloud data, and adopting a multimodal cross-iterative self-supervised approach, the two-dimensional boundary contour coordinates in each iteration are self-supervised to learn, so as to accurately obtain the target annotation information of the three-dimensional detection box. No manual verification is required throughout the process, which effectively improves the annotation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723556B_ABST
    Figure CN119723556B_ABST
Patent Text Reader

Abstract

The present disclosure provides a three-dimensional bounding box processing method, device and equipment based on self-supervised learning, comprising: obtaining image iteration data and point cloud iteration data corresponding to an environment scene of an intelligent vehicle at a current time during driving of the intelligent vehicle; determining first labeling information of a three-dimensional bounding box corresponding to a detection target based on the image iteration data to process two-dimensional image contour coordinates of the detection target; determining second labeling information of the three-dimensional bounding box corresponding to the detection target based on the point cloud iteration data to process two-dimensional point cloud contour coordinates of the detection target; and performing self-supervised learning on the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates to obtain target labeling information of the three-dimensional bounding box corresponding to the detection target. Thus, the target labeling information of the three-dimensional bounding box is accurately obtained, and no manual verification operation is required in the whole process, effectively improving the labeling efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure relate to the field of self-supervised learning technology, and more specifically, to a method, apparatus, and device for processing 3D detection boxes based on self-supervised learning. Background Technology

[0002] During vehicle operation, by identifying targets in the driving environment and annotating them with 3D bounding boxes, the specific location of the targets can be effectively determined, facilitating safe vehicle operation.

[0003] In related technologies, most of them use post-fusion or cross-validation to achieve automatic annotation of 3D detection boxes. Post-fusion contains an endless number of scene-specific rules in different scenarios, and cross-validation cannot truly determine whether to trust point cloud detection results or image detection results. Even if the model results can be guaranteed to be correct, the alignment problem of different modalities will cause inconsistencies in the annotation results, which require manual verification.

[0004] However, the existing method requires manual annotation and proofreading, which is costly and has a certain error rate, thus reducing annotation efficiency. Summary of the Invention

[0005] The embodiments described herein provide a method, apparatus, and device for processing 3D detection boxes based on self-supervised learning, which overcomes the aforementioned problems.

[0006] Firstly, based on the content of this disclosure, a method for processing 3D detection boxes based on self-supervised learning is provided, including:

[0007] During the operation of the intelligent vehicle, image iteration data and point cloud iteration data corresponding to the environmental scene of the intelligent vehicle at the current moment are acquired. The image iteration data is used to describe the multi-view image data of the environmental scene collected by the intelligent vehicle at the current moment, and the point cloud iteration data is used to describe the laser point cloud data of the environmental scene collected by the intelligent vehicle.

[0008] Based on the image iteration data, the first annotation information of the three-dimensional detection box corresponding to the detection target is determined, so as to obtain the two-dimensional image contour coordinates of the detection target. The two-dimensional image contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detection target determined by the image iteration data.

[0009] Based on the point cloud iterative data, the second annotation information of the three-dimensional detection box corresponding to the detection target is determined, so as to obtain the two-dimensional point cloud contour coordinates of the detection target. The two-dimensional point cloud contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detection target determined by the point cloud iterative data.

[0010] Self-supervised learning is performed on the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates to obtain the target annotation information of the three-dimensional detection box corresponding to the detected target.

[0011] The target annotation information of the three-dimensional detection box corresponding to the detection target includes: the center point coordinates, length, width, height and orientation of the three-dimensional detection box.

[0012] Secondly, according to the present disclosure, a 3D bounding box processing device based on self-supervised learning is provided, comprising:

[0013] The acquisition module is used to acquire image iteration data and point cloud iteration data of the intelligent vehicle corresponding to the environmental scene at the current moment during the driving process of the intelligent vehicle. The image iteration data is used to describe the multi-view image data of the environmental scene collected by the intelligent vehicle at the current moment, and the point cloud iteration data is used to describe the laser point cloud data of the environmental scene collected by the intelligent vehicle.

[0014] The first determining module is used to determine the first annotation information of the three-dimensional detection box corresponding to the detection target based on the image iteration data, so as to process and obtain the two-dimensional image contour coordinates of the detection target. The two-dimensional image contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detection target determined by the image iteration data.

[0015] The second determining module is used to determine the second annotation information of the three-dimensional detection box corresponding to the detection target based on the point cloud iterative data, so as to process and obtain the two-dimensional point cloud contour coordinates of the detection target. The two-dimensional point cloud contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detection target determined by the point cloud iterative data.

[0016] The learning module is used to perform self-supervised learning on the contour coordinates of the two-dimensional image and the contour coordinates of the two-dimensional point cloud to obtain the target annotation information of the three-dimensional detection box corresponding to the detected target.

[0017] The target annotation information of the three-dimensional detection box corresponding to the detection target includes: the center point coordinates, length, width, height and orientation of the three-dimensional detection box.

[0018] Thirdly, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the self-supervised learning-based 3D detection box processing method as described in any of the above embodiments.

[0019] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the three-dimensional detection box processing method based on self-supervised learning as described in any of the above embodiments.

[0020] The self-supervised learning-based 3D detection box processing method provided in this application embodiment acquires image iteration data and point cloud iteration data corresponding to the environmental scene of the intelligent vehicle at the current moment during the driving process of the intelligent vehicle. The image iteration data is used to describe the multi-view image data of the environmental scene collected by the intelligent vehicle at the current moment, and the point cloud iteration data is used to describe the laser point cloud data of the environmental scene collected by the intelligent vehicle. Based on the image iteration data, the first annotation information of the 3D detection box corresponding to the detection target is determined to obtain the 2D image contour coordinates of the detection target. The 2D image contour coordinates are used to describe the 2D boundary contour coordinates of the 3D detection box corresponding to the detection target determined by the image iteration data. Based on the point cloud iteration data, the second annotation information of the 3D detection box corresponding to the detection target is determined to obtain the 2D point cloud contour coordinates of the detection target. The 2D point cloud contour coordinates are used to describe the 2D boundary contour coordinates of the 3D detection box corresponding to the detection target determined by the point cloud iteration data. Self-supervised learning is performed on the 2D image contour coordinates and the 2D point cloud contour coordinates to obtain the target annotation information of the 3D detection box corresponding to the detection target. The target annotation information of the 3D detection box corresponding to the detection target includes: the center point coordinates, length, width, height, and orientation of the 3D detection box. Thus, by combining multi-view image data and laser point cloud data, and adopting a multimodal cross-iterative self-supervised approach, the two-dimensional boundary contour coordinates in each iteration are self-supervised to learn, so as to accurately obtain the target annotation information of the three-dimensional detection box. No manual verification is required throughout the process, which effectively improves the annotation efficiency.

[0021] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more obvious and understandable, specific implementation methods of this application are described below. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. It should be understood that the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure, wherein:

[0023] Figure 1 This is a flowchart illustrating a self-supervised learning-based 3D bounding box processing method provided in this disclosure.

[0024] Figure 2 This is a schematic diagram of a 3D detection box processing device based on self-supervised learning provided in this disclosure.

[0025] Figure 3 This is a schematic diagram of the structure of a computer device provided in this disclosure.

[0026] It should be noted that the elements in the attached diagram are schematic and not drawn to scale. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are also within the scope of protection of this disclosure.

[0028] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the specification and in the relevant art, and shall not be interpreted in an idealized or overly formal form unless otherwise explicitly defined herein. As used herein, the statement of “connecting” or “coupling” two or more parts together shall mean that these parts are directly joined together or joined through one or more intermediate components.

[0029] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of the phrase "embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0030] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists, A and B exist simultaneously, or B exists. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Terms such as "first" and "second" are only used to distinguish one component (or part of a component) from another component (or another part of a component).

[0031] In the description of this application, unless otherwise stated, "multiple" means two or more (including two), and similarly, "multiple groups" means two or more (including two groups).

[0032] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0033] Figure 1 This is a flowchart illustrating a self-supervised learning-based 3D bounding box processing method provided in an embodiment of this disclosure, as shown below. Figure 1 As shown, the specific process of the 3D bounding box processing method based on self-supervised learning includes:

[0034] S110. During the operation of the intelligent vehicle, acquire the image iteration data and point cloud iteration data of the intelligent vehicle corresponding to the environmental scene at the current moment.

[0035] Among them, image iteration data is used to describe multi-view image data of the environmental scene collected by the intelligent vehicle at the current moment, and point cloud iteration data is used to describe laser point cloud data of the same environmental scene collected by the intelligent vehicle at the current moment.

[0036] S120. Based on the image iteration data, determine the first annotation information of the three-dimensional detection box corresponding to the detection target, so as to process and obtain the two-dimensional image contour coordinates of the detection target.

[0037] The two-dimensional image contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detected target, determined by image iteration data. The two-dimensional image contour coordinates may include the position coordinates of all two-dimensional boundary contour points of the three-dimensional detection box corresponding to the detected target, determined by image iteration data.

[0038] It should be noted that the first annotation information is the ground truth data of the 3D detection box determined by the image iteration data at the current moment, which includes the center point coordinates, length, width, height and orientation of the 3D detection box.

[0039] In some embodiments, based on image iteration data, the first annotation information of the three-dimensional detection box corresponding to the detection target is determined, so as to process and obtain the two-dimensional image contour coordinates of the detection target, including:

[0040] Sparse visual processing is performed on the image iterative data to obtain the three-dimensional image vector representation corresponding to the image iterative data; the three-dimensional image vector representation corresponding to the image iterative data is decoded to obtain the first annotation information of the three-dimensional detection box corresponding to the detection target; boundary distance rendering is performed on the first annotation information of the three-dimensional detection box corresponding to the detection target to obtain the two-dimensional image contour coordinates of the detection target.

[0041] In the sparse visual processing of iterative image data, the iterative image data is input into a sparse model. The initial sparse queries undergo cross-attention calculation with multi-view image features (i.e., iterative image data) (image features serve as the key and value of the attention module, and sparse queries serve as the input queries to the attention module; the attention calculation uses QKV operations), outputting 3D query vectors (i.e., the 3D image vector representation corresponding to the iterative image data). A decoder is then used to decode the 3D image vector representation corresponding to the iterative image data to obtain 3dB boxes (i.e., the first annotation information).

[0042] Therefore, by processing the iterative image data, the ground truth value of the 3D detection box at the current moment is obtained, and the ground truth value of the 3D detection box is rendered into two dimensions to obtain the two-dimensional image contour coordinates.

[0043] In some embodiments, boundary distance rendering is performed on the first annotation information of the 3D detection box corresponding to the detection target to obtain the 2D image contour coordinates of the detection target, including:

[0044] Boundary distance prediction is performed on the first annotation information to obtain the image distance representation sequence between the detected target and the corresponding 3D detection box; 2D boundary rendering is performed on the image distance representation sequence between the detected target and the corresponding 3D detection box to obtain the 2D image contour coordinates of the detected target.

[0045] For example, the SDF (sign distance function) is used to calculate the function value of the first annotation information. A residual regression network is then used to regress the function value of the target's 3D true boundary and the difference between the target's 3D bounding box SDF (obtained by calculating the distance from each gridded point on the six boundary faces of the 3D bounding box to a reference point), resulting in an image distance representation sequence. This sequence is then rendered to 2D to obtain the 2D image contour coordinates of the detected target. Thus, the 2D image contour coordinates of the detected target are effectively obtained.

[0046] S130. Based on the point cloud iterative data, determine the second annotation information of the three-dimensional detection box corresponding to the detection target, so as to process and obtain the two-dimensional point cloud contour coordinates of the detection target.

[0047] The two-dimensional point cloud contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detected target, determined by the point cloud iteration data. The two-dimensional point cloud contour coordinates may include the position coordinates of all two-dimensional boundary contour points of the three-dimensional detection box corresponding to the detected target, determined by the point cloud iteration data.

[0048] It should be noted that the second annotation information is the ground truth data of the 3D detection box determined by the point cloud iteration data at the current moment, which includes the center point coordinates, length, width, height and orientation of the 3D detection box.

[0049] In some embodiments, based on point cloud iterative data, the second annotation information of the three-dimensional detection box corresponding to the detection target is determined to process and obtain the two-dimensional point cloud contour coordinates of the detection target, including:

[0050] Sparse visual processing is performed on the point cloud iterative data to obtain the 3D image vector representation corresponding to the point cloud iterative data; the 3D image vector representation corresponding to the point cloud iterative data is decoded to obtain the second annotation information of the 3D detection box corresponding to the detection target; boundary distance rendering is performed on the second annotation information of the 3D detection box corresponding to the detection target to obtain the 2D point cloud contour coordinates of the detection target.

[0051] In the sparse visual processing of point cloud iterative data, the point cloud iterative data is input into a sparse model. The initial sparse queries undergo cross-attention calculation with the point cloud features (i.e., the point cloud iterative data) (the point cloud features serve as the key and value of the attention module, and the sparse queries serve as the input queries to the attention module; the attention calculation uses QKV operations), outputting 3D query vectors (i.e., the 3D image vector representation corresponding to the point cloud iterative data). A decoder is then used to decode the 3D image vector representation corresponding to the point cloud iterative data to obtain 3dBBoxes (i.e., the second annotation information).

[0052] Therefore, by processing the point cloud iterative data, the ground truth value of the 3D detection box at the current moment is obtained, and the ground truth value of the 3D detection box is rendered into two dimensions to obtain the two-dimensional point cloud contour coordinates.

[0053] In some embodiments, boundary distance rendering is performed on the second annotation information of the 3D detection box corresponding to the detection target to obtain the 2D point cloud contour coordinates of the detection target, including:

[0054] Boundary distance prediction is performed on the second annotation information to obtain the point cloud distance representation sequence between the detected target and the corresponding 3D detection box; 2D boundary rendering is performed on the point cloud distance representation sequence between the detected target and the corresponding 3D detection box to obtain the 2D point cloud contour coordinates of the detected target.

[0055] For example, the sign distance function (SDF) is used to calculate the function value of the second annotation information. A residual regression network is then used to regress the function value of the target's 3D true boundary and the difference between the target's 3D bounding box SDF (obtained by calculating the distance between each meshed point on the six boundary faces of the 3D bounding box and a reference point), resulting in a point cloud distance representation sequence. This sequence is then rendered into 2D to obtain the 2D point cloud contour coordinates of the detected target. Thus, the 2D point cloud contour coordinates of the detected target are effectively derived.

[0056] S140. Perform self-supervised learning on the contour coordinates of the two-dimensional image and the contour coordinates of the two-dimensional point cloud to obtain the target annotation information of the three-dimensional detection box corresponding to the detected target.

[0057] The target annotation information of the 3D detection box corresponding to the detection target includes: the center point coordinates, length, width, height and orientation of the 3D detection box.

[0058] It should be noted that the target annotation information of the 3D detection box corresponding to the detected target is the most accurate ground truth data of the 3D detection box obtained by iterating through data from multiple time points. It may be the first or second annotation information determined at the current time point, as mentioned above, or it may be other annotation information determined at the next or subsequent time points.

[0059] In some embodiments, self-supervised learning is performed on the contour coordinates of the two-dimensional image and the contour coordinates of the two-dimensional point cloud to obtain the target annotation information of the three-dimensional detection box corresponding to the detected target, including:

[0060] Image instance segmentation is performed on the iterative image data to obtain a corresponding 2D instance mask. Similarly, point cloud instance segmentation is performed on the iterative point cloud data to obtain a corresponding 2D instance mask. Using the 2D instance mask from the iterative image data, the accuracy of the 2D image contour coordinates and the 2D point cloud contour coordinates is verified to obtain the ground truth value of the iterative image detection box for the 3D detection bounding box corresponding to the detected target. The accuracy of the iterative point cloud detection box is also verified using the 2D instance mask from the iterative point cloud data to obtain the ground truth value of the iterative point cloud detection box for the 3D detection bounding box corresponding to the detected target. Finally, the ground truth values ​​of the iterative image detection box and the iterative point cloud detection box are fused to obtain the target annotation information of the 3D detection bounding box corresponding to the detected target.

[0061] Specifically, a large-scale image instance segmentation model can be used to segment the image iterative data to obtain a two-dimensional instance mask corresponding to the image iterative data, and a large-scale point cloud instance segmentation model can be used to segment the point cloud iterative data to obtain a two-dimensional instance mask corresponding to the point cloud iterative data.

[0062] When using a two-dimensional instance mask corresponding to the image iteration data to perform accuracy verification on the two-dimensional image contour coordinates, if the coordinate difference between the object contour coordinates reflected by the two-dimensional instance mask corresponding to the image iteration data and the two-dimensional image contour coordinates is lower than a preset threshold, then the verification result of using the two-dimensional instance mask corresponding to the image iteration data to perform accuracy verification on the two-dimensional image contour coordinates is determined to be verification passed; if the coordinate difference between the object contour coordinates reflected by the two-dimensional instance mask corresponding to the image iteration data and the two-dimensional image contour coordinates is higher than or equal to the preset threshold, then the verification result of using the two-dimensional instance mask corresponding to the image iteration data to perform accuracy verification on the two-dimensional image contour coordinates is determined to be verification failed.

[0063] When using the 2D instance mask corresponding to the point cloud iteration data to perform accuracy verification on the 2D point cloud contour coordinates, if the coordinate difference between the object contour coordinates reflected by the 2D instance mask corresponding to the point cloud iteration data and the 2D point cloud contour coordinates is lower than a preset threshold, then the verification result of using the 2D instance mask corresponding to the point cloud iteration data to perform accuracy verification on the 2D point cloud contour coordinates is determined to be verification passed; if the coordinate difference between the object contour coordinates reflected by the 2D instance mask corresponding to the point cloud iteration data and the 2D point cloud contour coordinates is higher than or equal to the preset threshold, then the verification result of using the 2D instance mask corresponding to the point cloud iteration data to perform accuracy verification on the 2D point cloud contour coordinates is determined to be verification failed.

[0064] The method of fusing the ground truth values ​​of the image iterative detection bounding boxes and the point cloud iterative detection bounding boxes to obtain the target annotation information of the 3D detection bounding box corresponding to the detected target may include: using the same or different preset fusion ratios to fuse the ground truth values ​​of the image iterative detection bounding boxes and the point cloud iterative detection bounding boxes to obtain the target annotation information of the 3D detection bounding box corresponding to the detected target.

[0065] Therefore, by performing multimodal fusion on the obtained ground truth values ​​of the image iterative detection boxes and the point cloud iterative detection boxes, the target annotation information of the three-dimensional detection boxes corresponding to the detected target can be obtained effectively and accurately.

[0066] In some embodiments, a two-dimensional instance mask corresponding to the image iteration data is used to verify the accuracy of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates, respectively, in order to obtain the ground truth value of the image iteration detection box corresponding to the three-dimensional detection box of the detection target, including:

[0067] If the accuracy verification results for the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the image iteration data pass the verification, then the ground truth value of the image iteration detection box of the three-dimensional detection box corresponding to the detection target is determined based on the first annotation information and the second annotation information. If the first annotation information and the second annotation information are fused at different ratios (the fusion ratio of the first annotation information is higher than that of the second annotation information), the ground truth value of the image iteration detection box of the three-dimensional detection box corresponding to the detection target is obtained.

[0068] If the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the image iteration data are not passed, the iteration continues until the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates determined at the current time are passed. Then the ground truth value of the image iteration detection box corresponding to the three-dimensional detection box of the detection target is obtained.

[0069] For example, if the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the image iteration data are both unsuccessful, then the process continues to obtain the image iteration data and point cloud iteration data of the intelligent vehicle corresponding to the environmental scene at the current moment; based on the image iteration data, the first annotation information of the three-dimensional detection box corresponding to the detection target is determined, so as to obtain the two-dimensional image contour coordinates of the detection target; based on the point cloud iteration data, the second annotation information of the three-dimensional detection box corresponding to the detection target is determined, so as to obtain the two-dimensional point cloud contour coordinates of the detection target, until the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates corresponding to the two-dimensional instance mask corresponding to the image iteration data determined at the current moment are both unsuccessful, then the most accurate image iteration detection box ground value of the three-dimensional detection box is determined.

[0070] In some embodiments, a two-dimensional instance mask corresponding to the point cloud iterative data is used to perform accuracy checks on the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates, respectively, to obtain the ground truth value of the point cloud iterative detection box of the three-dimensional detection box corresponding to the detection target, including:

[0071] If the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the point cloud iterative detection data pass the verification, then the ground truth value of the point cloud iterative detection box of the three-dimensional detection box corresponding to the detection target is determined based on the first annotation information and the second annotation information. If the first annotation information and the second annotation information are fused at different ratios (the fusion ratio of the second annotation information is higher than that of the first annotation information), the ground truth value of the point cloud iterative detection box of the three-dimensional detection box corresponding to the detection target is obtained.

[0072] If the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the point cloud iteration data are not passed, the iteration continues until the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates determined at the current time are passed. Then the ground truth value of the point cloud iteration detection box of the three-dimensional detection box corresponding to the detection target is obtained.

[0073] For example, if the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the point cloud iteration data are both unsuccessful, then the process continues to obtain the image iteration data and point cloud iteration data of the intelligent vehicle corresponding to the environmental scene at the current moment; based on the image iteration data, the first annotation information of the three-dimensional detection box corresponding to the detection target is determined, so as to obtain the two-dimensional image contour coordinates of the detection target; based on the point cloud iteration data, the second annotation information of the three-dimensional detection box corresponding to the detection target is determined, so as to obtain the two-dimensional point cloud contour coordinates of the detection target, until the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates corresponding to the two-dimensional instance mask corresponding to the point cloud iteration data determined at the current moment are both unsuccessful, then the ground truth value of the most accurate point cloud iteration detection box is determined.

[0074] In this embodiment, during the operation of the intelligent vehicle, image iteration data and point cloud iteration data corresponding to the environmental scene at the current moment are acquired. The image iteration data describes the multi-view image data of the environmental scene collected by the intelligent vehicle at the current moment, and the point cloud iteration data describes the laser point cloud data of the environmental scene collected by the intelligent vehicle. Based on the image iteration data, the first annotation information of the three-dimensional detection box corresponding to the detection target is determined, and the two-dimensional image contour coordinates of the detection target are obtained. The two-dimensional image contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detection target determined by the image iteration data. Based on the point cloud iteration data, the second annotation information of the three-dimensional detection box corresponding to the detection target is determined, and the two-dimensional point cloud contour coordinates of the detection target are obtained. The two-dimensional point cloud contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detection target determined by the point cloud iteration data. Self-supervised learning is performed on the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates to obtain the target annotation information of the three-dimensional detection box corresponding to the detection target. The target annotation information of the three-dimensional detection box corresponding to the detection target includes: the center point coordinates, length, width, height, and orientation of the three-dimensional detection box. Thus, by combining multi-view image data and laser point cloud data, and adopting a multimodal cross-iterative self-supervised approach, the two-dimensional boundary contour coordinates in each iteration are self-supervised to learn, so as to accurately obtain the target annotation information of the three-dimensional detection box. No manual verification is required throughout the process, which effectively improves the annotation efficiency.

[0075] In summary, this embodiment implements a dual-path cross-validation paradigm for automatic 3D bounding box annotation. Through self-supervision, the bounding boxes generated from point clouds and images can complement and correct each other. A spatial transformation module aligns queries between the point cloud and image modules, avoiding gradient errors or jitter caused by camera and radar misalignment. The 3dsdf residual fills the gap between the 3D bounding boxes and the actual 3D boundaries of the objects, and projects it onto the image plane to establish consistency constraints with the 2D segmentation mask. 3D point cloud segmentation results and 2D image segmentation are obtained from a large model, and the results are cross-supervised after projection.

[0076] Figure 2 This is a schematic diagram of a 3D detection box processing device based on self-supervised learning provided in this embodiment. The 3D detection box processing device based on self-supervised learning may include: an acquisition module 210, a first determination module 220, a second determination module 230, and a learning module 240.

[0077] The acquisition module 210 is used to acquire image iteration data and point cloud iteration data of the intelligent vehicle corresponding to the environmental scene at the current moment during the driving process of the intelligent vehicle. The image iteration data is used to describe the multi-view image data of the environmental scene collected by the intelligent vehicle at the current moment, and the point cloud iteration data is used to describe the laser point cloud data of the environmental scene collected by the intelligent vehicle.

[0078] The first determining module 220 is used to determine the first annotation information of the three-dimensional detection box corresponding to the detection target based on the image iteration data, so as to process and obtain the two-dimensional image contour coordinates of the detection target. The two-dimensional image contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detection target determined by the image iteration data.

[0079] The second determining module 230 is used to determine the second annotation information of the three-dimensional detection box corresponding to the detection target based on the point cloud iterative data, so as to process and obtain the two-dimensional point cloud contour coordinates of the detection target. The two-dimensional point cloud contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detection target determined by the point cloud iterative data.

[0080] Learning module 240 is used to perform self-supervised learning on the contour coordinates of the two-dimensional image and the contour coordinates of the two-dimensional point cloud to obtain the target annotation information of the three-dimensional detection box corresponding to the detected target.

[0081] The target annotation information of the 3D detection box corresponding to the detection target includes: the center point coordinates, length, width, height and orientation of the 3D detection box.

[0082] In this embodiment, optionally, the first determining module 220 includes:

[0083] The first processing unit is used to perform sparse visual processing on the image iterative data to obtain the three-dimensional image vector representation corresponding to the image iterative data.

[0084] The first decoding unit is used to decode the three-dimensional image vector representation corresponding to the image iteration data to obtain the first annotation information of the three-dimensional detection box corresponding to the detection target.

[0085] The first rendering unit is used to render the boundary distance of the first annotation information of the three-dimensional detection box corresponding to the detection target, so as to obtain the two-dimensional image contour coordinates of the detection target.

[0086] In this embodiment, optionally, the first rendering unit is specifically used for:

[0087] Boundary distance prediction is performed on the first annotation information to obtain the image distance representation sequence between the detected target and the corresponding 3D detection box; 2D boundary rendering is performed on the image distance representation sequence between the detected target and the corresponding 3D detection box to obtain the 2D image contour coordinates of the detected target.

[0088] In this embodiment, optionally, the second determining module 230 includes:

[0089] The second processing unit is used to perform sparse visual processing on the point cloud iterative data to obtain the three-dimensional image vector representation corresponding to the point cloud iterative data.

[0090] The second decoding unit is used to decode the three-dimensional image vector representation corresponding to the point cloud iterative data to obtain the second annotation information of the three-dimensional detection box corresponding to the detection target.

[0091] The second rendering unit is used to render the boundary distance of the second annotation information of the three-dimensional detection box corresponding to the detection target, so as to obtain the two-dimensional point cloud contour coordinates of the detection target.

[0092] In this embodiment, optionally, the second rendering unit is specifically used for:

[0093] Boundary distance prediction is performed on the second annotation information to obtain the point cloud distance representation sequence between the detected target and the corresponding 3D detection box; 2D boundary rendering is performed on the point cloud distance representation sequence between the detected target and the corresponding 3D detection box to obtain the 2D point cloud contour coordinates of the detected target.

[0094] In this embodiment, optionally, the learning module 240 includes:

[0095] The segmentation unit is used to perform image instance segmentation on image iterative data to obtain a two-dimensional instance mask corresponding to the image iterative data, and to perform point cloud instance segmentation on point cloud iterative data to obtain a two-dimensional instance mask corresponding to the point cloud iterative data.

[0096] The verification unit is used to perform accuracy verification on the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates using the two-dimensional instance mask corresponding to the image iteration data, so as to obtain the true value of the image iteration detection box of the three-dimensional detection box corresponding to the detection target; and to perform accuracy verification on the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates using the two-dimensional instance mask corresponding to the point cloud iteration detection box of the three-dimensional detection box corresponding to the detection target.

[0097] The fusion unit is used to fuse the ground truth values ​​of the image iterative detection bounding boxes and the point cloud iterative detection bounding boxes to obtain the target annotation information of the 3D detection bounding box corresponding to the detected target.

[0098] In this embodiment, optionally, the verification unit is specifically used for:

[0099] If the accuracy verification results of the two-dimensional image contour coordinates and two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the image iteration data are passed, then the ground truth value of the image iteration detection box of the three-dimensional detection box corresponding to the detection target is determined based on the first annotation information and the second annotation information; if the accuracy verification results of the two-dimensional image contour coordinates and two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the image iteration data are not passed, the iteration continues until the accuracy verification results of the two-dimensional image contour coordinates and two-dimensional point cloud contour coordinates determined at the current time are passed, then the ground truth value of the image iteration detection box of the three-dimensional detection box corresponding to the detection target is obtained.

[0100] In this embodiment, optionally, the verification unit is specifically used for:

[0101] If the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the point cloud iteration data are passed, then the ground truth value of the point cloud iteration detection box of the three-dimensional detection box corresponding to the detection target is determined based on the first annotation information and the second annotation information; if the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the point cloud iteration data are failed, the iteration continues until the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates determined at the current time are passed, then the ground truth value of the point cloud iteration detection box of the three-dimensional detection box corresponding to the detection target is obtained.

[0102] The 3D detection box processing device based on self-supervised learning provided in this disclosure can execute the above-described method embodiments. Its specific implementation principle and technical effects can be found in the above-described method embodiments, and will not be repeated here.

[0103] This application also provides a computer device. Please refer to the following for details. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.

[0104] The computer device includes a memory 310 and a processor 320 that are communicatively connected to each other via a system bus. It should be noted that only a computer device with memory 310 and processor 320 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0105] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0106] The memory 310 includes at least one type of readable storage medium, including non-volatile memory or volatile memory, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. RAM may include static RAM or dynamic RAM. In some embodiments, the memory 310 may be an internal storage unit of a computer device, such as the hard disk or RAM of the computer device. In other embodiments, the memory 310 may also be an external storage device of the computer device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, or flash card equipped on the computer device. Of course, the memory 310 may include both internal storage units and external storage devices of the computer device. In this embodiment, the memory 310 is typically used to store the operating system and various application software installed on the computer device, such as the program code of the methods described above. Furthermore, the memory 310 may also be used to temporarily store various types of data that have been output or will be output.

[0107] Processor 320 is typically used to perform overall operations of a computer device. In this embodiment, memory 310 is used to store program code or instructions, including computer operation instructions, and processor 320 is used to execute the program code or instructions stored in memory 310 or process data, such as program code that runs the methods described above.

[0108] In this article, the bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus system can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0109] Another embodiment of this application also provides a computer-readable medium, which may be a computer-readable signal medium or a computer-readable medium. A processor in a computer reads computer-readable program code stored in the computer-readable medium, enabling the processor to execute the functional actions specified in each step or combination of steps in the above method; and to generate means for implementing the functional actions specified in each block or combination of blocks in the block diagram.

[0110] Computer-readable media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared memory or semiconductor systems, devices or apparatuses, or any suitable combination thereof, wherein the memory is used to store program code or instructions, the program code including computer operation instructions, and the processor is used to execute the program code or instructions of the above-described methods stored in the memory.

[0111] The definitions of memory and processor can be found in the description of the foregoing computer device embodiments, and will not be repeated here.

[0112] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0113] In the various embodiments of this application, the functional units or modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0114] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0115] In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" as described in this application does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims listing several means, several units of these means may be embodied by the same item of hardware. The use of "first," "second," and "third," etc., does not indicate any order and these words should be interpreted as names. Unless otherwise specified, the steps in the above embodiments should not be construed as limiting the order of execution.

[0116] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for processing 3D bounding boxes based on self-supervised learning, characterized in that, include: During the operation of the intelligent vehicle, image iteration data and point cloud iteration data corresponding to the environmental scene of the intelligent vehicle at the current moment are acquired. The image iteration data is used to describe the multi-view image data of the environmental scene collected by the intelligent vehicle at the current moment, and the point cloud iteration data is used to describe the laser point cloud data of the environmental scene collected by the intelligent vehicle. Based on the image iteration data, the first annotation information of the three-dimensional detection box corresponding to the detection target is determined, so as to obtain the two-dimensional image contour coordinates of the detection target. The two-dimensional image contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detection target determined by the image iteration data. Based on the point cloud iterative data, the second annotation information of the three-dimensional detection box corresponding to the detection target is determined, so as to obtain the two-dimensional point cloud contour coordinates of the detection target. The two-dimensional point cloud contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detection target determined by the point cloud iterative data. Self-supervised learning is performed on the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates to obtain the target annotation information of the three-dimensional detection box corresponding to the detected target. This process includes: performing image instance segmentation on the image iteration data to obtain a two-dimensional instance mask corresponding to the image iteration data; performing point cloud instance segmentation on the point cloud iteration data to obtain a two-dimensional instance mask corresponding to the point cloud iteration data; using the two-dimensional instance mask corresponding to the image iteration data, the accuracy of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates is verified to obtain the ground truth value of the image iteration detection box of the three-dimensional detection box corresponding to the detected target; using the two-dimensional instance mask corresponding to the point cloud iteration data, the accuracy of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates is verified to obtain the ground truth value of the point cloud iteration detection box of the three-dimensional detection box corresponding to the detected target; and fusing the ground truth value of the image iteration detection box and the ground truth value of the point cloud iteration detection box to obtain the target annotation information of the three-dimensional detection box corresponding to the detected target. The target annotation information of the three-dimensional detection box corresponding to the detection target includes: the center point coordinates, length, width, height and orientation of the three-dimensional detection box.

2. The method according to claim 1, characterized in that, The step of determining the first annotation information of the three-dimensional detection box corresponding to the detection target based on the image iteration data, and processing it to obtain the two-dimensional image contour coordinates of the detection target, includes: Sparse visual processing is performed on the image iterative data to obtain a three-dimensional image vector representation corresponding to the image iterative data; Decode the three-dimensional image vector representation corresponding to the image iteration data to obtain the first annotation information of the three-dimensional detection box corresponding to the detection target; The boundary distance is rendered on the first annotation information of the three-dimensional detection box corresponding to the detection target to obtain the two-dimensional image contour coordinates of the detection target.

3. The method according to claim 2, characterized in that, The step of rendering boundary distances on the first annotation information of the three-dimensional detection box corresponding to the detection target to obtain the two-dimensional image contour coordinates of the detection target includes: Boundary distance prediction is performed on the first annotation information to obtain the image distance representation sequence between the detected target and the corresponding three-dimensional detection box; Two-dimensional boundary rendering is performed on the image distance representation sequence between the detected target and the corresponding three-dimensional detection box to obtain the two-dimensional image contour coordinates of the detected target.

4. The method according to claim 1, characterized in that, The step of determining the second annotation information of the three-dimensional detection box corresponding to the detection target based on the point cloud iterative data, and processing it to obtain the two-dimensional point cloud contour coordinates of the detection target, includes: Sparse visual processing is performed on the point cloud iterative data to obtain the three-dimensional image vector representation corresponding to the point cloud iterative data. Decode the three-dimensional image vector representation corresponding to the point cloud iterative data to obtain the second annotation information of the three-dimensional detection box corresponding to the detection target; The boundary distance is rendered on the second annotation information of the three-dimensional detection box corresponding to the detection target to obtain the two-dimensional point cloud contour coordinates of the detection target.

5. The method according to claim 4, characterized in that, The step of rendering boundary distances on the second annotation information of the 3D detection box corresponding to the detection target to obtain the 2D point cloud contour coordinates of the detection target includes: Boundary distance prediction is performed on the second annotation information to obtain a point cloud distance representation sequence between the detected target and the corresponding 3D detection box; Two-dimensional boundary rendering is performed on the point cloud distance representation sequence between the detected target and the corresponding three-dimensional detection box to obtain the two-dimensional point cloud contour coordinates of the detected target.

6. The method according to claim 1, characterized in that, The step of using the two-dimensional instance mask corresponding to the image iteration data to perform accuracy verification on the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates respectively, in order to obtain the ground truth value of the image iteration detection box corresponding to the three-dimensional detection box of the detection target, includes: If the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the image iteration data are verified as passed, then the image iteration detection box ground value of the three-dimensional detection box corresponding to the detection target is determined based on the first annotation information and the second annotation information. If the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the image iteration data are both unsuccessful, the iteration continues until the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates determined at the current time are both successful. Then, the ground truth value of the image iteration detection box of the three-dimensional detection box corresponding to the detection target is obtained.

7. The method according to claim 1, characterized in that, The step of using the two-dimensional instance mask corresponding to the point cloud iterative data to perform accuracy verification on the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates respectively, in order to obtain the ground truth value of the point cloud iterative detection box of the three-dimensional detection box corresponding to the detection target, includes: If the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the point cloud iterative data are verified as passed, then the ground truth value of the point cloud iterative detection box of the three-dimensional detection box corresponding to the detection target is determined based on the first annotation information and the second annotation information. If the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the point cloud iterative data are both unsuccessful, the iteration continues until the accuracy verification results of the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates based on the two-dimensional instance mask corresponding to the point cloud iterative data determined at the current time are both successful. Then, the ground truth value of the point cloud iterative detection box of the three-dimensional detection box corresponding to the detection target is obtained.

8. A 3D bounding box processing device based on self-supervised learning, characterized in that, include: The acquisition module is used to acquire image iteration data and point cloud iteration data of the intelligent vehicle corresponding to the environmental scene at the current moment during the driving process of the intelligent vehicle. The image iteration data is used to describe the multi-view image data of the environmental scene collected by the intelligent vehicle at the current moment, and the point cloud iteration data is used to describe the laser point cloud data of the environmental scene collected by the intelligent vehicle. The first determining module is used to determine the first annotation information of the three-dimensional detection box corresponding to the detection target based on the image iteration data, so as to process and obtain the two-dimensional image contour coordinates of the detection target. The two-dimensional image contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detection target determined by the image iteration data. The second determining module is used to determine the second annotation information of the three-dimensional detection box corresponding to the detection target based on the point cloud iterative data, so as to process and obtain the two-dimensional point cloud contour coordinates of the detection target. The two-dimensional point cloud contour coordinates are used to describe the two-dimensional boundary contour coordinates of the three-dimensional detection box corresponding to the detection target determined by the point cloud iterative data. The learning module is used to perform self-supervised learning on the contour coordinates of the two-dimensional image and the contour coordinates of the two-dimensional point cloud to obtain the target annotation information of the three-dimensional detection box corresponding to the detected target. The learning module includes a segmentation unit, a verification unit, and a fusion unit. The segmentation unit is used to perform image instance segmentation on the image iterative data to obtain a two-dimensional instance mask corresponding to the image iterative data, and to perform point cloud instance segmentation on the point cloud iterative data to obtain a two-dimensional instance mask corresponding to the point cloud iterative data. The verification unit is used to use the two-dimensional instance mask corresponding to the image iterative data to perform accuracy verification on the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates, respectively, to obtain the ground truth value of the image iterative detection box of the three-dimensional detection box corresponding to the detection target; and to use the two-dimensional instance mask corresponding to the point cloud iterative data to perform accuracy verification on the two-dimensional image contour coordinates and the two-dimensional point cloud contour coordinates, respectively, to obtain the ground truth value of the point cloud iterative detection box of the three-dimensional detection box corresponding to the detection target. The fusion unit is used to fuse the ground truth value of the image iterative detection box and the ground truth value of the point cloud iterative detection box to obtain the target annotation information of the three-dimensional detection box corresponding to the detection target. The target annotation information of the three-dimensional detection box corresponding to the detection target includes: the center point coordinates, length, width, height and orientation of the three-dimensional detection box.

9. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the three-dimensional detection box processing method based on self-supervised learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Vehicle positioning method, device, equipment and medium

    CN117351079A