An Unsupervised 4D Automatic Annotation Method for Autonomous Driving Street Scenes Based on 3D Reconstruction

Through the unsupervised method of three-dimensional reconstruction and denoising processing, the high cost and domain differences of autonomous driving data annotation are solved, and efficient and accurate 4D annotation is achieved, and 3D object detection can be performed using only image information.

CN118967934BActive Publication Date: 2025-07-11SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Patent Information

Application Number
CN202411091554.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2025-07-11
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

The existing autonomous driving data labeling technology relies on true value labeling, resulting in waste of manpower, financial resources and time, and there are domain differences problems, so it is impossible to directly reason on unknown data.

Method used

Unsupervised method based on three-dimensional reconstruction is adopted, and a general pre-trained visual model is used to obtain pixel-by-pixel instance information and pseudo-point clouds, combined with implicit surface reconstruction and DBSCAN algorithm for denoising, realizing 3D object detection.

Benefits of technology

It realizes efficient and accurate 4D labeling without truth value labeling, avoids manual labeling costs and domain differences, and improves data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118967934B_ABST
    Figure CN118967934B_ABST
Patent Text Reader

Abstract

The present invention relates to an unsupervised 4D automatic annotation method for autonomous driving street scenes based on three-dimensional reconstruction, including: processing the information of the input frame-by-frame images to obtain instance information and pseudo-point clouds per pixel; performing three-dimensional reconstruction based on the instance information and pseudo-point clouds to obtain a reconstructed scene; sampling the reconstructed scene frame by frame to obtain instance point clouds per instance, and then performing denoising processing to output 3D annotation boxes for each frame, thereby obtaining a 4D annotation result for the entire sequence composed of frame-by-frame images. Compared with the prior art, the present invention can achieve 3D object detection only by using image information without ground truth annotation, essentially realizing unsupervised learning, directly eliminating the costs of manual annotation and data preprocessing, improving efficiency and saving a large amount of computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving data processing, and in particular to an unsupervised 4D automatic annotation method for autonomous driving street scenes based on three-dimensional reconstruction. Background Art

[0002] At present, the research and development of autonomous driving has not only stayed in the academic research stage, and the industrial implementation of technology has reflected many requirements. Among them, the most basic one is that for the collected real data, it is necessary to perform data annotation, including but not limited to the 4D boxes and semantic information of the objects in the scene (here the 4D box refers to the length, width, height coordinates and time dimension, that is, it is necessary to know the specific position of each object at different times). 4D annotation is a data annotation method that not only includes all the information of traditional 2D annotation but also introduces dimensions such as height and time to make the annotated data more rich and complete. In the field of autonomous driving, 4D annotation has a wide application prospect, but this work can only be completed manually or semi-automatically at present, which means that data annotators need to spend a lot of manpower, financial resources and time. Similar annotated data sets can refer to Waymo and nuScenes, etc.

[0003] The existing data annotation technical solutions are to use a data-driven training neural network method, based on a part of existing 2D / 3D / 4D annotation data, to train the model. After the training is completed, the model can complete the target function, that is, input the frame-by-frame images and point clouds to the model, and the model outputs the 3D information of the objects frame by frame. For example, works such as BEVfusion and Auto4D have well completed object detection in a fixed data set. Through the training of a certain data set, the model can detect the data in the current data set domain and the output effect is good, but it cannot directly reason on other data (non-training data domain); VSRD adopts an implicit reconstruction method to reconstruct a single-frame image. The inference process only needs a 2D image as input to output a 3D box. However, due to the characteristics of single-frame reconstruction, this solution still depends on true value data such as 2D boxes and the category identity information of objects, and cannot be completely unsupervised.

[0004] In summary, the existing technologies often rely on 3D true value annotation and 2D true value annotation for the target data to be detected, and cannot be completely unsupervised. In addition, most technologies have the problem of domain difference, that is, they cannot directly reason on unknown or unseen data by the model (the detection accuracy of forced reasoning is almost 0). Summary of the Invention

[0005] The objective of the present invention is to overcome the deficiencies of the above-mentioned existing technologies and provide an unsupervised 4D automatic annotation method for autonomous driving street scenes based on three-dimensional reconstruction, which can achieve 3D object detection only using image information without ground truth annotation.

[0006] The objective of the present invention can be achieved through the following technical solutions: An unsupervised 4D automatic annotation method for autonomous driving street scenes based on three-dimensional reconstruction, comprising the following steps:

[0007] S1. Perform information processing on the input frame-by-frame images to obtain per-pixel instance information and pseudo point clouds;

[0008] S2. According to the instance information and pseudo point clouds, perform three-dimensional reconstruction to obtain a reconstructed scene;

[0009] S3. Sample the reconstructed scene frame by frame to obtain per-instance point clouds, then perform denoising processing, and output 3D annotation boxes for each frame, thereby obtaining a 4D annotation result for the entire sequence composed of frame-by-frame images.

[0010] Furthermore, the step S1 specifically uses a general pre-trained visual model to perform information processing on the frame-by-frame images.

[0011] Furthermore, the general pre-trained visual model includes an image segmentation model and a depth estimation model. The image segmentation model is used to obtain the semantic information of each pixel point in the image, and then obtain instance information;

[0012] The depth estimation model is used to obtain the distance information of each pixel point in the image and convert it into the format of pseudo point clouds.

[0013] Furthermore, the working process of the image segmentation model is as follows:

[0014] For image data where k is the acquisition frames at different times, and j is different images at the same time;

[0015] Input into the image segmentation model to obtain per-pixel instance information The value in Figure 1 is the instance ID to which the current pixel belongs. To obtain instance information that is consistent across frames and views, project the image calculate the 2D box according to the instance, and project it into the internal hidden space of the image segmentation model according to the area of the 2D box respectively, to obtain the feature vectors corresponding to each box in each image m is the number of boxes in a single image of the current frame, and also the number of instances in the current image;

[0016] Subsequently, the Euclidean distance between the multi-view feature vectors is measured. To determine whether the current two bounding boxes belong to the same instance. If the distance is less than the threshold, they are considered to belong to the same instance, so as to unify the instances of all views in the current frame. The number of instances per frame is M, that is, the total number of instances under all views in the current frame.

[0017] Continue to measure the Euclidean distance between the feature vectors of different frames. For Update to achieve tracking on instances, so as to unify the instances of the entire scene.

[0018] Furthermore, the working process of the depth estimation model is as follows:

[0019] For the image data where k is the acquisition frame at different times and j is different images at the same time.

[0020] Input into the depth estimation model to obtain the per-pixel depth information, that is, the depth information, and then convert it into the pseudo point cloud format correspondingly.

[0021] Furthermore, the specific step S2 is to adopt the implicit surface reconstruction method streetsurf based on the signed distance field, and use the instance information to perform cross-frame full-scene instance reconstruction.

[0022] Furthermore, the specific process of the step S2 includes:

[0023] streetsurf encodes the scene information into the latent space vector to memorize and overfit the required information. The loss function used includes the per-pixel color regularization and the depth regularization is the color information of the current point predicted by the reconstruction network, is the depth information of the current point predicted by the reconstruction network;

[0024] In addition, for the instance information, an additional multi-layer perceptron model (MLP) is used to predict the instance information, and an additional instance regularization is the instance information predicted by the reconstruction network.

[0025] Furthermore, the step S3 includes the following steps:

[0026] S31. Sample the reconstructed scene frame by frame. By collecting the dense spatial point coordinates and querying the instance information corresponding to these spatial point coordinates, the per-frame instance point cloud data PC k ;

[0027] S32. Denoise the acquired instance point cloud to filter out outliers, and then calculate 3D bounding boxes frame by frame according to the instance category.

[0028] Further, the step S32 specifically uses the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm for denoising.

[0029] Further, the calculation formula for the 3D bounding box in the step S32 is:

[0030]

[0031] where R k is the 3D bounding box result corresponding to the k-th frame image, z represents the z-th bounding box, and M is the total number of bounding boxes.

[0032] Compared with the prior art, the present invention has the following advantages:

[0033] The present invention processes the input frame-by-frame images to obtain per-pixel instance information and pseudo point clouds for 3D reconstruction to obtain a reconstructed scene, then samples the reconstructed scene frame by frame to obtain per-instance point clouds, and then performs denoising processing to output frame-by-frame 3D bounding boxes, thereby obtaining 4D annotation results for the entire sequence composed of frame-by-frame images. Thus, unsupervised automatic annotation of images can be achieved, and only 2D temporal images and camera parameters are required to perform object box annotation on the scene.

[0034] The present invention uses a general pre-trained visual model to process the frame-by-frame images. Among them, the general pre-trained visual model includes an image segmentation model and a depth estimation model. The image segmentation model is used to obtain the semantic information of each pixel point in the image, and then instance information is obtained; the depth estimation model is used to obtain the distance information of each pixel point in the image and convert it into the format of pseudo point clouds. The present invention uses a pre-trained general visual model to process information, which not only avoids the need for the model for ground truth, but also avoids the data domain difference problem because these general visual models are not trained on a single data set.

[0035] The present invention further uses a tracking method for the acquired semantic information to perform similarity matching on the frame-by-frame data to achieve consistent instance information across frames, that is, the identity information of the same instance is the same at different times, which is beneficial to subsequent unifying the instances of the entire scene.

[0036] Considering that the pseudo point cloud is not dense enough, the present invention designs and adopts an implicit surface reconstruction method streetsurf based on the signed distance field, and uses instance information for cross-frame full-scene instance reconstruction. In addition, for the already reconstructed scene, frame-by-frame sampling is performed to obtain a denser three-dimensional semantic and instance point cloud, and the point cloud is denoised and outliers are filtered separately according to the instance category using the DBSCAN clustering algorithm, so as to accurately calculate the 3D bounding box. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic flowchart of the method of the present invention;

[0038] Figure 2 It is a schematic diagram of the application process of the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The present invention will be described in detail below with reference to the drawings and specific embodiments.

[0040] Embodiment

[0041] As Figure 1 shown, an unsupervised 4D automatic annotation method for autonomous driving street scenes based on three-dimensional reconstruction includes the following steps:

[0042] S1. Process the information of the input frame-by-frame images to obtain per-pixel instance information and pseudo point cloud;

[0043] S2. Perform three-dimensional reconstruction according to the instance information and pseudo point cloud to obtain a reconstructed scene;

[0044] S3. Perform frame-by-frame sampling on the reconstructed scene to obtain per-instance point cloud, and then perform denoising processing to output the 3D annotation bounding box for each frame, so as to obtain the 4D annotation result of the entire sequence composed of frame-by-frame images.

[0045] This embodiment applies the above scheme and is mainly divided into three parts:

[0046] First, process the information of the input frame-by-frame images, and use the general pre-trained visual models SAM (Segment Anything Model, an image segmentation model) and Depth-Anything (a depth estimation model) to obtain the semantic information and distance information of each pixel point in the image respectively. These general visual models are not trained on a single dataset, so there is no data domain difference problem. For the obtained semantic information, this scheme also uses a tracking method to perform similarity matching on the frame-by-frame data to achieve consistent instance information across frames, that is, the identity information of the same instance is the same at different times.

[0047] The second part is to perform three-dimensional reconstruction on the acquired information, and use the streetsurf reconstruction method to perform implicit surface reconstruction to reconstruct the surface information and semantic instance information of the street view.

[0048] The last part is to sample the reconstructed scene frame by frame to obtain a denser three-dimensional semantic and instance point cloud, and use the DBSCAN clustering algorithm to denoise and filter out outliers for each instance category of the point cloud respectively, so as to obtain accurate 3D bounding boxes.

[0049] Specifically, given a set of collected image data where k is the acquisition frame at different times, and j is different images at the same time, such as front, left, and right views.

[0050] As Figure 2 shown, in order to obtain the information required for reconstruction, is input into the segmenter S (i.e., the image segmentation model) to obtain per-pixel instance information The numbers in represent which instance the current pixel belongs to. Specifically, S is a detector pre-trained on general data, and only its weights need to be loaded to be directly used, without the problem of domain bias. Secondly, in order to obtain consistent instance information across frames and views Figure 1 For the images compute the 2D bounding boxes according to the instances, and project them into the internal latent space of S according to the regions of the 2D bounding boxes respectively to obtain the feature vectors corresponding to each box in each image m is the number of bounding boxes and also the number of instances in the current image. Then, measure the Euclidean distance between the multi-view feature vectors to determine whether the current two bounding boxes belong to the same instance. If the distance is less than the threshold, it is considered that the similarity is very high and they belong to the same instance, so as to unify the instances of all views in the current frame, and the number of instances per frame is M. Similarly, continue to measure the Euclidean distance between the feature vectors of different frames The updated can achieve tracking on instances, making the instances of the entire scene unified.

[0051] In addition to obtaining instance information, the images are also input into the pre-trained depth information detector D (i.e., the depth estimation model), and its usage method is similar to that of the detector S, and the pre-trained model can be directly used. Thus, per-pixel depth information, that is, pseudo point cloud, can be obtained for subsequent three-dimensional reconstruction.

[0052] Although there is already a pseudo point cloud and theoretically the calculation of the detection box can be directly carried out, considering that the pseudo point cloud is not dense enough, the accuracy of directly calculating the detection box is not high and the effect is not good. Therefore, the implicit surface reconstruction method streetsurf based on the signed distance field is adopted, and the instance information is used for cross-frame full-scene instance reconstruction. streetsurf encodes the scene information into an implicit space vector to memorize and overfit the required information, and the loss function used is mainly pixel-wise color regularization and depth regularization is the color information of the current point predicted by the reconstruction network, is the depth information predicted by the reconstruction network. In addition, for the instance information, an additional multi-layer perceptron model (MLP) is used to predict the instance information, and an additional instance regularization is the instance information predicted by the reconstruction network.

[0053] After training the reconstruction network (i.e., 3D reconstruction), dense spatial point coordinates are collected, and the instance information corresponding to these point coordinates is queried from the network to obtain the per-frame instance point cloud data PC k . Since the reconstruction network is not perfect and there are errors, there will be some outliers in these point clouds, which interfere with subsequent calculations. Therefore, the DBSCAN algorithm is adopted in this scheme to filter out these noise points. Finally, the 3D annotation box can be calculated frame by frame according to the instance:

[0054] To verify the effectiveness of this scheme, this embodiment is tested under the internationally common and recognized waymo dataset, and the results are shown in the following table:

[0055]

[0056] The results show that the unsupervised method proposed in this scheme has the advantages of high efficiency, low cost and accurate results.

[0057] In summary, for any data, this scheme can perform 3D object detection only by using image information without ground truth annotation (referring to 2D / 3D / 4D object detection boxes and corresponding semantic information). By inputting per-frame street view images and corresponding camera pose parameters, the 3D boxes and instance semantic information of each object in the street view image can be output. This scheme uses a pre-trained general vision model for information processing, which not only avoids the need for the model to the ground truth but also avoids the domain difference problem. This scheme realizes unsupervised automatic annotation of images. Only 2D temporal images and camera parameters are required to annotate the object boxes in the scene. Compared with the existing technology, this scheme completely does not require the annotation of 2D boxes and semantics, essentially realizes unsupervised, directly saves the cost of manual annotation and data preprocessing, improves the efficiency and saves a large amount of computing resources.

Claims

1. An unsupervised 4D automatic annotation method for autonomous driving street scenes based on 3D reconstruction, characterized in that, It includes the following steps: S1. Perform information processing on the input frame-by-frame images to obtain per-pixel instance information and pseudo point clouds; S2. Based on the instance information and pseudo point clouds, perform 3D reconstruction to obtain a reconstructed scene; S3. Sample the reconstructed scene frame by frame to obtain instance point clouds per frame, then perform denoising processing, and output 3D annotation boxes per frame, thereby obtaining 4D annotation results for the entire sequence composed of frame-by-frame images; Specifically, in step S1, a general pre-trained visual model is used to perform information processing on the frame-by-frame images. The general pre-trained visual model includes an image segmentation model, and the working process of the image segmentation model is as follows: For image data where k is the label of different moments and j is the label of different images at the same moment, which is the j-th image at the k-th moment; Input into the image segmentation model to obtain per-pixel instance information The value in is the instance ID to which the current pixel belongs. To obtain instance information that is consistent across frames and views, project the image Calculate 2D boxes according to instances, and project them into the internal latent space of the image segmentation model according to the regions of the 2D boxes respectively, to obtain the feature vectors corresponding to each box in each image m is the number of boxes in a single image of the current frame and also the number of instances in the current image; After that, the Euclidean distance between the multi-view feature vectors is measured For {p, q} ∈ m, to determine whether the current two bounding boxes belong to the same instance. If the distance is less than the threshold, they are considered to belong to the same instance, so as to unify the instances of all views in the current frame. The number of instances per frame is M, that is, the total number of instances under all views in the current frame; Continue to measure the Euclidean distance between the feature vectors of different frames For {p,q} ∈ M Update to achieve tracking on the instance, so that the instances of the entire scene are unified; Specifically, in step S2, the implicit surface reconstruction method streetsurf based on the signed distance field is adopted, and full-scene instance reconstruction across frames is performed using instance information. The specific process of step S2 includes: Streetsurf encodes scene information into a latent space vector to memorize and overfit the required information. The loss functions used include per-pixel color regularization and depth regularization is the j-th image at the k-th moment predicted by the reconstruction network, is the depth information of the j-th image at the k-th moment predicted by the reconstruction network, is the depth information of the j-th image at the k-th moment; In addition, for instance information, an additional multi-layer perceptron model is used to predict the instance information, and an additional instance regularization is the instance information predicted by the reconstruction network.

2. The unsupervised 4D automatic annotation method for autonomous driving street view based on 3D reconstruction according to claim 1, wherein The general pre-trained visual model also includes a depth estimation model; The depth estimation model is used to obtain the distance information of each pixel point in the image and convert it into the format of pseudo point clouds.

3. A method for unsupervised 4D automatic annotation of an autonomous driving street view based on three-dimensional reconstruction according to claim 2, characterized in that The working process of the depth estimation model is as follows: For image data where k is the label of different times and j is the label of different images at the same time which is the j-th image at the k-th time Input into the depth estimation model to obtain per-pixel depth information, i.e., depth information, and then convert it into the pseudo-point cloud format accordingly.

4. A method for unsupervised 4D automatic annotation of an autonomous driving street view based on three-dimensional reconstruction according to claim 1, characterized in that, Step S3 includes the following steps: S31. Sample each frame of the reconstruction scene. By collecting dense spatial point coordinates and querying the instance information corresponding to these spatial point coordinates, obtain the instance point cloud data PC for each frame k ; S32. Perform denoising processing on the obtained instance point clouds to filter out outliers, and then calculate 3D annotation boxes frame by frame according to the instance categories.

5. A method for unsupervised 4D automatic annotation of an autonomous driving street view based on three-dimensional reconstruction according to claim 4, characterized in that, Specifically, in step S32, the DBSCAN algorithm is adopted for denoising processing.

6. A method for unsupervised 4D automatic annotation of an autonomous driving street view based on three-dimensional reconstruction according to claim 5, characterized in that, The calculation formula for the 3D annotation box in step S32 is: where, R k is the 3D bounding box result corresponding to all images at the k-th moment, z represents the z-th bounding box, and M is the total number of bounding boxes, i.e., the total number of instances.

Citation Information

Patent Citations

  • Unsupervised monocular three-dimensional target detection method based on video sequence and pre-training instance segmentation

    CN116129318A

  • Systems and methods for generating and using visual datasets for training computer vision models

    US20220414928A1

Cited By

  • A method and system for verifying the consistency of velocity direction in 3D obstacle annotation data

    CN122574829A