Image processing method and device, equipment and storage medium
By obtaining the mask and point cloud in the target image and combining shape prior information for optimization processing, the problem of unstable marking results and poor accuracy in automatic marking of 3D objects is solved, and high-precision and stable 3D object marking is achieved.
Patent Information
- Application Number
- CN202311726862.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-17
AI Technical Summary
In the prior art, the labeling results are unstable and labeling accuracy are poor when automatically labeling 3D objects.
By obtaining the target mask in the target image, the target object point cloud in the target point cloud is obtained, and the initial pose and shape are determined based on the shape prior information, and then the target pose and shape are obtained through optimization processing.
Improve the accuracy and stability of 3D object labeling, ensuring data generalization and excellent accuracy.
Smart Images

Figure CN120163869A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and in particular to an image processing method, device, equipment and storage medium. Background Art
[0002] In modern robotics and autonomous driving technologies, existing algorithms require a large amount of labeled data to train deep learning models to understand 3D (3 Dimensions) scenes, especially dynamic objects such as vehicles and pedestrians. Therefore, there is an urgent need to improve the automation of 3D labeling. The current 3D automatic annotation methods can be divided into two categories according to whether true value bounding boxes are required: 1) Supervised automatic annotation methods: Such methods usually need to obtain manually annotated 3D bounding boxes, then train a 3D detection model based on these manually annotated 3D bounding boxes, and finally use the trained model to predict the 3D bounding boxes of the remaining data to achieve automatic annotation. However, due to the limitation of the amount of data during model training, using such methods to annotate 3D objects often has the problem of overfitting to a specific dataset, resulting in poor generalization performance when annotating data outside the data distribution. 2) Unsupervised automatic annotation methods: Compared with supervised annotation methods, although unsupervised automatic annotation methods do not require manually annotated 3D bounding boxes, due to the lack of effective supervision signals in unsupervised automatic annotation methods, the quality of the pseudo-labels generated by unsupervised automatic annotation methods is poor and the practical value is limited. Therefore, the above two types of methods have problems of unstable annotation results or poor annotation accuracy when automatically annotating 3D objects. Summary of the Invention
[0003] The purpose of the present application is to provide an image processing method, device, equipment and storage medium to solve the problems of unstable annotation results and poor annotation accuracy existing in the existing technical problems when annotating 3D objects.
[0004] In a first aspect, an embodiment of the present application provides an image processing method, including: obtaining a target mask corresponding to a target object in a target image; obtaining the point cloud of the target object in the target point cloud based on the target mask; determining the initial pose of the target object based on the point cloud of the target object, and determining the initial shape of the target object based on shape prior information; optimizing the initial pose and the initial shape of the target object based on the target mask and the point cloud of the target object to obtain the target pose and the target shape of the target object.
[0005] The image processing method provided by the embodiments of the present application determines the initial pose of the target object based on the point cloud of the target object and determines the initial shape of the target object based on the shape prior information. After obtaining the initial pose and shape, the initial pose and initial shape are optimized based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object. The initial shape obtained based on the shape prior information has a high accuracy, which is conducive to shortening the optimization time. At the same time, the target pose and target shape obtained by optimizing based on the initial shape and pose have good data generalization and excellent accuracy, and the accuracy of the target pose and target shape obtained by using the image processing method provided by the embodiments of the present invention has good stability.
[0006] In one implementation manner of the first aspect, obtaining the point cloud of the target object in the target point cloud based on the target mask includes: obtaining the target point cloud, where the target point cloud includes information of the target object and the environment where the target object is located; segmenting the target point cloud based on the target mask to obtain the point cloud of the target object.
[0007] Obtaining the point cloud of the object from the target point cloud based on the target mask can ensure the accuracy of the point cloud of the target object, so as to obtain a more accurate pose of the target object.
[0008] In one implementation manner of the first aspect, it further includes denoising the point cloud of the target object.
[0009] The point cloud of the target object obtained from the target point cloud through the target mask contains noise point clouds. Denoising the point cloud of the target object can further improve the accuracy of the pose of the target object.
[0010] In one implementation manner of the first aspect, denoising the point cloud of the target object includes: removing the point clouds whose distance from the center point of the point cloud of the target object is greater than the first threshold.
[0011] Regarding the point clouds whose distance from the center point of the point cloud of the target object is greater than the first threshold as noise points can effectively reduce the complexity of denoising.
[0012] In one implementation manner of the first aspect, obtaining the initial pose of the target object based on the point cloud of the target object includes: performing a mean process on the point cloud of the target object to obtain the initial pose of the target object.
[0013] Performing a mean process based on the point cloud of the target object is easy to implement, and the accuracy of the initial pose of the target object obtained is relatively high.
[0014] In one implementation manner of the first aspect, the shape prior information includes SDF values, and the shape prior information is obtained based on a preset dimensionality reduction algorithm.
[0015] Shape prior information can well represent the initial shape of the target object, facilitating the reduction of the complexity of subsequent optimization processing. To reduce the complexity of subsequent optimization processing, it is necessary to perform dimensionality reduction on the SDF values based on a preset dimensionality reduction algorithm, thereby obtaining shape prior information containing the SDF values.
[0016] In one implementation manner of the first aspect, the dimensionality reduction algorithm includes the PCA algorithm.
[0017] Performing dimensionality reduction using the PCA algorithm is conducive to implementation.
[0018] In one implementation manner of the first aspect, determining the initial shape of the target object based on the shape prior information includes: using the average shape of the shape prior information as the initial shape of the target object.
[0019] Using the average shape of the shape prior information as the initial shape of the target object can make the initial shape of the target object more compatible and have better applicability.
[0020] In one implementation manner of the first aspect, optimizing the initial pose and initial shape of the target object based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object includes: determining the initial pose and initial shape of the target object as the current pose and shape of the target object; based on the current pose and shape of the target object, obtaining the current mask of the target object; based on the current pose and shape of the target object, the current mask of the target object, the target mask, and the point cloud of the target object, obtaining the current loss information of the target object; when the current loss information of the target object meets the set conditions, determining the current pose and shape of the target object as the target pose and target shape of the target object.
[0021] The loss information can well verify the difference between the current shape and current pose and the actual shape and pose of the target object.
[0022] In one implementation manner of the first aspect, the loss information includes at least one of mask alignment loss information, point cloud alignment loss information, and height alignment loss information; wherein, the mask alignment loss information is used for the distance between the target mask and the current mask of the target object; the point cloud alignment loss information is used to represent the alignment degree between the object represented by the current pose and shape of the target object and the point cloud of the target object; the height alignment loss information is used to represent the distance between the height where the current pose and shape are located and the ground.
[0023] To improve the accuracy of the target pose and target shape, the loss information includes at least one of mask alignment loss information, point cloud alignment loss information, and height alignment loss information.
[0024] In an implementation of the first aspect, it further includes: when the current loss information of the target object does not meet the set conditions, updating the current pose and shape of the target object based on a preset update method, and repeating the steps of obtaining the current mask of the target object based on the current pose and shape of the target object, and the steps of obtaining the current loss information of the target object based on the current pose and shape of the target object, the current mask of the target object, the target mask, and the point cloud of the target object, until the current loss information of the target object meets the set conditions.
[0025] Updating the current pose and shape of the target object based on a preset update method can improve the efficiency of the optimization process.
[0026] In an implementation of the first aspect, obtaining the current mask of the target object based on the current pose and shape of the target object includes: obtaining the current mask of the target object based on a differentiable 3D rendering algorithm based on the SDF value, the current pose, and the shape of the target object.
[0027] The differentiable 3D rendering algorithm based on the SDF value can accurately obtain the current mask corresponding to the current shape and pose of the target object.
[0028] In an implementation of the first aspect, updating the current pose and shape of the target object based on a preset update method includes: updating the current pose and shape of the target object based on a preset gradient descent algorithm with a preset step size or a preset number of steps.
[0029] The gradient descent algorithm can improve the efficiency of the optimization process and shorten the time of the optimization process.
[0030] In an implementation of the first aspect, the set condition is that the loss information corresponding to the target object in the current pose and shape is convergent.
[0031] The fact that the loss information corresponding to the target object in the current pose and shape is convergent can improve the accuracy of the target shape and target pose of the target object.
[0032] In an implementation of the first aspect, it further includes: labeling the target object based on the target pose and the target shape.
[0033] After the image processing device obtains the target pose and the target shape, it can automatically label the target pose and the target shape corresponding to the target object in the target space.
[0034] Second aspect, an embodiment of the present application provides an image processing device, including: an acquisition unit and a processing unit. Among them, the acquisition unit is used to acquire the target mask of the target object in the target image. The acquisition unit is further used to obtain the point cloud of the target object in the target point cloud based on the target mask. The processing unit is used to determine the initial pose of the target object based on the point cloud of the target object, and determine the initial shape of the target object based on the shape prior information. The processing unit is further used to optimize the initial pose and the initial shape of the target object based on the target mask and the point cloud of the target object to obtain the target pose and the target shape of the target object.
[0035] In a possible implementation manner of the second aspect, the acquisition unit includes: a point cloud acquisition module and a separation module. Among them, the point cloud acquisition module is used to acquire the target point cloud. The separation module is used to segment the target point cloud based on the target mask to obtain the point cloud of the target object.
[0036] In a possible implementation manner of the second aspect, the acquisition unit further includes: a denoising module. The denoising module is used to perform denoising processing on the initial point cloud of the target object.
[0037] In a possible implementation manner of the second aspect, the denoising module is further used to remove the cloud points whose distance from the center point of the point cloud of the target object is greater than the first threshold. Thereby, the denoised point cloud of the target object can be obtained.
[0038] In a possible implementation manner of the second aspect, the processing unit is further used to perform mean processing on the point cloud of the target object to obtain the initial pose of the target object.
[0039] In a possible implementation manner of the second aspect, the processing unit is further used to use the average shape of the shape prior information as the initial shape of the target object.
[0040] In a possible implementation manner of the second aspect, the processing unit is further used to determine the initial pose and the initial shape of the target object as the current pose and shape of the target object, and obtain the current mask of the target object based on the current pose and shape of the target object, and obtain the current loss information of the target object based on the current pose and shape of the target object, the current mask of the target object, the target mask and the point cloud of the target object. When the current loss information of the target object meets the set conditions, the processing unit determines the current pose and shape of the target object as the target pose and shape of the target object.
[0041] In a possible implementation of the second aspect, the loss information includes at least one of mask alignment loss information, point cloud alignment loss information, and height alignment loss information. Among them, the mask alignment loss information is used for the distance between the target mask of the target object and the current mask. The point cloud alignment loss information is used to represent the distance between the current pose and shape of the target object and the point cloud of the target object. The height alignment loss information is used to represent the distance between the height at which the current pose and shape are located and the ground.
[0042] In a possible implementation of the second aspect, the processing unit is further configured to, when the current loss information of the target object does not meet the set conditions, update the current pose and shape of the target object based on a preset update method, and repeat the step of obtaining the current mask of the target object based on the current pose and shape of the target object, and the step of obtaining the current loss information of the target object based on the current pose and shape of the target object, the current mask of the target object, the target mask, and the point cloud of the target object, until the current loss information of the target object meets the set conditions.
[0043] In a possible implementation of the second aspect, the processing unit is further configured to obtain the current mask of the target object based on a differentiable 3D rendering algorithm of the SDF value, the current pose, and the shape of the target object.
[0044] In a possible implementation of the second aspect, the processing unit is further configured to update the current pose and shape of the target object based on a preset gradient descent algorithm with a preset step size or a preset number of steps.
[0045] In an embodiment of the present application, the image processing device further includes: an annotation unit. The annotation unit is configured to annotate the target object based on the target pose and the target shape.
[0046] In a third aspect, an embodiment of the present application provides an electronic device, including a memory for storing computer program instructions and a processor for executing the program instructions. When the computer program instructions are executed by the processor, the electronic device is enabled to implement the method of the first aspect or any one of the implementation manners in the first aspect.
[0047] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium includes a stored program. When the program is executed by at least one processor, the device where the computer-readable storage medium is located is controlled to implement the method of the first aspect or any one of the implementation manners in the first aspect.
[0048] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes executable instructions. When the executable instructions are executed by at least one processor, the computer is enabled to implement the method of the first aspect or any one of the implementation manners in the first aspect.
[0049] In a sixth aspect, an embodiment of the present application provides a vehicle equipped with a driving system or a navigation system, where the driving system or the navigation system includes image information, and the image information is obtained by the method of the first aspect or any one of the implementation manners in the first aspect.
[0050] The image processing method provided by the embodiment of the present application determines the initial pose of the target object based on the point cloud of the target object, and determines the initial shape of the target object based on the shape prior information. After obtaining the initial pose and shape, the initial pose and initial shape are optimized based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object. The initial shape obtained based on the shape prior information is beneficial to the optimization process, and the target pose and target shape obtained by optimizing based on the initial shape and pose have good data generalization and excellent accuracy, and the accuracy of the target pose and target shape obtained by using the image processing method provided by the embodiment of the present invention has good stability. Description of the Drawings
[0051] Figure 1 It is a flowchart of an image processing method provided by an embodiment of the present application.
[0052] Figure 2 It is a flowchart of obtaining the point cloud of the target object provided by an embodiment of the present application.
[0053] Figure 3 It is a flowchart of an optimization process provided by an embodiment of the present application.
[0054] Figure 4 It is a flowchart of an image processing method provided by an embodiment of the present application.
[0055] Figure 5 It is an effect diagram of 3D bounding box annotation using an image processing method provided by an embodiment of the present application.
[0056] Figure 6 It is an effect diagram of voxel occupancy grid annotation using an image processing method provided by an embodiment of the present application.
[0057] Figure 7 It is a comparison diagram of voxel occupancy grid annotation using an image processing method provided by an embodiment of the present application and the prior art method.
[0058] Figure 8 It is a schematic structural diagram of an image processing device provided by an embodiment of the present application.
[0059] Figure 9 It is a schematic structural diagram of an image processing device provided by an embodiment of the present application.
[0060] Figure 10 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific implementation manners
[0061] The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be construed as limiting the present invention.
[0062] The autonomous driving perception architecture is the eyes of autonomous driving and also the basic guarantee for the safety performance of autonomous driving. Existing 3D perception algorithms require a large amount of labeled data to train deep learning models to understand 3D scenes. However, the annotation of objects in 3D space is an extremely expensive and labor-intensive task.
[0063] In some technologies, the annotation method usually targets a specific data set and uses a rule-based or pre-trained model method. This often works well for a specific data set or data annotation pattern. When introducing a real data set or data that is quite different from the training data, the quality of the 3D bounding boxes generated by the algorithm will quickly decline, and even cause the model to fail.
[0064] In addition, existing annotation methods all focus on solving the problem of automatic annotation of 3D bounding boxes and do not support more fine-grained annotation, such as voxel occupancy grids. Although there are some algorithms that can also generate the three-dimensional shape of the target, due to the need for virtual data training and the fact that the data distribution has a certain gap with the real data, the shape quality is poor.
[0065] Please refer to Figure 1 , an embodiment of the present application provides an image processing method, which is applied to an image processing device and includes:
[0066] S101, the image processing device obtains a target mask corresponding to a target object in the target image.
[0067] In the embodiments of the present application, an image processing device obtains a target image through an image acquisition device, and the target image includes an image of a target object. In some implementation manners, the image acquisition device may be a camera, and the target image is an image captured by the camera. The target image may further include an image of the environment where the target object is located. The target object refers to an object to be processed or labeled. The target object may be a vehicle, a crowd, a building, or a building facility, etc. After obtaining the target image, the image processing device may perform segmentation processing on the target image to obtain an image of the target object, and thus may obtain a target mask corresponding to the target object according to the image of the target object. Of course, the image processing device may also directly obtain the target mask of the target object in the target image from other devices or other models. In some embodiments, the image processing device may obtain the target image, transmit the target image and the information of the target object to be labeled to other devices or other models, and then may obtain the target mask corresponding to the target object in the target image.
[0068] When the image processing device performs segmentation on the target image, it needs to first identify the target object in the target image before obtaining the target mask of the target object. At this time, in some embodiments, the user may label the target object in the target image. In this way, the image processing device may identify the target object in the target image according to the prompt information of the user in the target image, and then segment out the image of the target object in the target image, and obtain the target mask corresponding to the target object according to the image of the target object.
[0069] Alternatively, in some embodiments, the categories of the target objects to be labeled may also be preset in the image processing device. At this time, the image processing device may identify the target object in the target image based on the category of the target object, and then segment out the image of the target object, and obtain the target mask corresponding to the target object according to the image of the target object. As a possible implementation manner, the image processing device may use a preset algorithm to identify the target object in the target image based on the category of the target object. For example, a 2D perception algorithm or an image recognition algorithm, or other algorithms may be used, and the present application does not limit this. Among them, the category of the target object refers to the category to which the target object belongs, and the categories include: vehicles, people, animals, buildings, traffic signs, etc.
[0070] As a possible implementation, to reduce the complexity of obtaining the target mask corresponding to the target object, a pre-trained target network model is used to obtain the target mask corresponding to the target object. That is, the target network model is a pre-trained model for identifying the mask information of the target object. In some target network models, the image processing device needs to first mark the target object to be masked in the target image. At this time, the image processing device can first mark the target object in the obtained target image, and use the target image marked with the target object as the input of the target network model, and input it into the target network model. The target network model performs corresponding processing on the target object in the target image to obtain the target mask corresponding to the target object in the target image. In some target network models, for example, SAM (Segment Anything Model) can be used, and the target image and the prompt information of the target object can be input. That is to say, the image processing device can use the target image as the input of the target network model and input it into the target network model. The target network model processes the received target image, and during the processing of the target network model, the marking information of the target object needs to be obtained. At this time, the image processing device can input the prompt information of the target object into the target network model, or the user can input the prompt information of the target object into the target network model. This application does not limit this.
[0071] As a possible implementation, the prompt information of the target object includes points belonging to the target object or the bounding box of the target object.
[0072] As a possible implementation, the target image can be a 2D image, and the marking information of the target object includes points of the 2D image belonging to the target object or the 2D bounding box of the target object.
[0073] In some embodiments, the target network model can be a part of the image processing device, or other devices or equipment outside the image processing device. This application does not limit this.
[0074] S102. The image processing device obtains the point cloud of the target object in the target point cloud based on the target mask.
[0075] Generally, the target point cloud can be collected by a lidar. However, the target point cloud collected by the lidar contains not only the point cloud of the target object but also the point cloud of other objects. To improve the accuracy of obtaining the pose of the target object, the image processing device can obtain the point cloud of the target object based on the target mask.
[0076] Please refer to Figure 2 , in a possible implementation of this step, the specific steps for the image processing device to obtain the point cloud of the target object in the target point cloud based on the target mask include:
[0077] S201. Obtain the target point cloud, where the target point cloud includes information about the target object and the environment where the target object is located.
[0078] In the embodiments of the present application, the target point cloud can be collected by a lidar or other devices. The image processing device directly or indirectly obtains the target point cloud from the lidar or other devices. Among them, the image processing device indirectly obtaining the target point cloud means obtaining the target point cloud from other devices, and the other devices obtain the target point cloud from the lidar or the other devices. The other devices can be at least one of an RGB-D camera, an interferometric synthetic aperture radar, and a device capable of generating a point cloud through an image derivation method.
[0079] The target point cloud is used to represent the target object and the environmental information where the target object is located. The target point cloud includes a number of points, and each point includes the 3D information of the point, where the 3D information can be 3D coordinate information. The point cloud of the target object refers to the point cloud corresponding to the target object in the target point cloud collected by the lidar.
[0080] S202. The image processing device segments the target point cloud based on the target mask to obtain the point cloud of the target object.
[0081] Since the target mask is 2D mask information and the point cloud is 3D data, therefore, it is necessary to first convert the target mask into 3D mask information, and the image processing device uses the 3D mask information to segment the point cloud of the target object from the target point cloud. In some possible implementation manners, the image processing device can use the setting parameters of the corresponding sensors in the point cloud collection device (such as a lidar) and the target image collection device (such as a camera) to convert the target mask into 3D mask information.
[0082] In step 202, the image processing device uses the 2D mask information to estimate the position and shape of the target object in the 3D space. This operation is easy to implement, and the estimated position and shape results of the target object in the 3D space are relatively accurate, laying a good foundation for subsequent optimization.
[0083] The point cloud of the target object segmented from the target point cloud based on the target mask includes noise point clouds. In order to obtain a more accurate point cloud of the target object, in some implementation manners, after step S202, the image processing method further includes:
[0084] S203. The image processing device denoises the point cloud of the target object.
[0085] Since the point cloud of the target object includes noise point clouds, therefore, the image processing device needs to remove the noise point clouds in the point cloud of the target object to obtain the denoised point cloud of the target object.
[0086] In a possible implementation of step S203, the image processing device denoises the point cloud of the target object, including: the image processing device removes the cloud points whose distance from the center point of the point cloud of the target object is greater than the first threshold. Thereby, the point cloud of the target object after denoising can be obtained. The accuracy of the point cloud of the target object after denoising is higher.
[0087] Among them, the first threshold can be a preset threshold, which can be set by technicians according to actual needs. The center point of the point cloud of the target object can be the point cloud located on the median of the projection depth. Each point in the point cloud of the target object is a point in 3D space, and the information of this point includes 3D coordinate information. The 3D coordinate information can be the coordinate information in a preset point cloud coordinate system.
[0088] In a possible implementation, the image processing device denoises the point cloud of the target object, including: denoising the point cloud of the target object based on the voxel filtering method or the method based on the normal vector.
[0089] S103, the image processing device determines the initial pose of the target object based on the point cloud of the target object, and determines the initial shape of the target object based on the shape prior information.
[0090] The image processing device calculates based on the point cloud of the target object to obtain the initial pose of the target object.
[0091] In a possible implementation of step S103, obtaining the initial pose of the target object based on the point cloud of the target object includes:
[0092] Performing mean processing on the point cloud of the target object to obtain the initial pose of the target object.
[0093] The initial pose of the target object includes the initial position and the initial orientation of the target object. In this implementation, the average value of all points in the point cloud of the target object is calculated dimension by dimension to obtain the coordinate information of the center point of the target object, and the coordinate information of the center point of the target object is taken as the initial position of the target object. The initial orientation angle of the target object can be randomly specified as any value from 0 to 2π, and the initial position and the initial orientation angle constitute the initial pose of the object. For example, if the points in the point cloud of the target object correspond to points in 3D space, then the average value of all points in the point cloud of the target object in the first dimension (for example, the X dimension) is calculated to obtain The average value of all points in the point cloud of the target object in the second dimension (for example, the Y dimension) is calculated to obtain The average value of all points in the point cloud of the target object in the third dimension (for example, the Z dimension) is calculated to obtain Then the initial position of the target object is Combined with the randomly initialized orientation angle, the initial pose of the target object is obtained. In this step, the point cloud of the target object can be the point cloud of the target object before denoising or the point cloud of the target object after denoising. When the point cloud of the target object after denoising is used for mean processing, the accuracy of the initial pose of the target object obtained is higher.
[0094] The image processing device acquires the preset shape prior information corresponding to the target object, and determines the initial shape of the target object based on the shape prior information.
[0095] The shape prior information is the prior information of the target shape of the object. The shape prior information can be used to represent the prior shape of the object, and the prior shape is a prior fuzzy perception of the target shape of the object based on the training data set or past experience data. For example, the shape prior information of a car can be used to represent the prior shape of the car, and the prior shape of the car can be a fuzzy car shape, but the final shape of the car needs to be optimized based on this fuzzy perception. The image processing device can set the shape prior information according to the object category. The shape prior information corresponding to the same category of objects is the same. In some embodiments, the shape prior information may include SDF (Signed Distance Function) values. In a possible implementation, the shape prior information may include a function for representing the shape of the object, for example, the shape prior information includes a vector expression. In a possible implementation, the shape prior information may include an implicit expression vector of SDF values. The vector dimension of the implicit expression vector of SDF values can be a preset dimension. In some possible implementations, the image processing device can obtain the vector expression corresponding to the target object according to the category of the target object, and obtain the initial shape of the target object based on the preset initial value of the vector expression corresponding to the target object. Each vector dimension of the implicit expression vector of SDF values represents a shape parameter. Taking the implicit expression vector of SDF values with three dimensions as an example, one vector dimension represents length, another vector dimension represents volume, and the last vector dimension represents shape.
[0096] In a possible implementation, the shape prior information includes: SDF (Signed Distance Function) values, and the shape prior information is obtained based on a preset dimensionality reduction algorithm.
[0097] Since the SDF value is a numerical value with many dimensions, in order to improve the efficiency of obtaining the shape of the target object, it is necessary to perform dimensionality reduction processing on the SDF value based on the dimensionality reduction algorithm to obtain an implicit expression vector of SDF values with a preset dimension.
[0098] In a possible implementation, the shape prior information may include a vector expression with a preset dimension, for example, an implicit expression vector including SDF values with a preset dimension. In a possible implementation, the dimensionality reduction algorithm includes the PCA (Principal Component Analysis) algorithm. It should be noted that the dimensionality reduction algorithm may also be other dimensionality reduction algorithms (such as the reverse feature elimination algorithm, or the independent component analysis algorithm, etc.), which are not limited in this application.
[0099] In a possible implementation, the image processing device determines the initial shape of the target object based on the shape prior information, including: taking the average shape of the shape prior information as the initial shape of the target object.
[0100] In a possible implementation, the average shape of the shape prior information refers to: the shape represented by the average value of multiple shape prior information. The multiple shape prior information is the shape prior information corresponding to multiple samples. For example, for n samples, each sample corresponds to a shape prior information, and n samples correspond to n shape prior information. The average value of the n shape prior information is taken as the average shape of the shape prior information.
[0101] In a possible implementation, the image processing device determines the initial shape of the target object based on the shape prior information, including: the image processing device obtains the initial value of the shape prior information. Based on the initial value of the shape prior information, the initial shape of the target object is obtained.
[0102] To reduce the complexity of the image processing method, the initial value of the shape prior information corresponding to each category of objects can be set, and the initial shape of the object is obtained based on this initial value. In some embodiments, the initial value of the shape prior information may be a set value, for example, the default value after initializing the shape prior information. The image processing device obtaining the initial value of the shape prior information may be obtaining the initial value of the implicit expression vector of the SDF value. After obtaining the initial value of the shape prior information, the initial shape of the target object represented by the initial value of the shape prior information can be obtained. In some possible implementations, there is a preset correspondence between the shape prior information of the object and the shape of the object. Based on this preset correspondence and the current shape prior information of the object, the current shape of the target object can be obtained. In a possible implementation, the preset correspondence can be pre-set by a technician before the image device leaves the factory.
[0103] The image processing device can obtain the shape prior information corresponding to the target object according to the category of the target object. The image processing device can also obtain the shape prior information corresponding to the target object from other devices or other models according to the category of the target object.
[0104] In some possible implementation manners, when the image processing device segments the target image, it needs to first identify the category of the target object in the target image, and then, based on the category of the target object, obtain the shape prior information corresponding to the target object.
[0105] In some possible implementation manners, the categories of the target objects to be labeled can be preset in the image processing device. The image processing device obtains the category of the target object from the preset information, and based on the category of the target object, obtains the shape prior information corresponding to the target object.
[0106] In some possible implementation manners, the image processing device inputs the category of the target object to other devices or other models, and obtains the prior information of the target object and its corresponding shape fed back by the other devices or other models.
[0107] In a possible implementation manner, in order to reduce the complexity of obtaining the shape prior information corresponding to the target object, a pre-trained network model is used to obtain the shape prior information corresponding to the target object. In some embodiments, the target network model can be a shape model. The image processing device obtains the initial value of the shape prior information from the shape model based on the category of the target object. The initial value of the shape prior information can be the default value preset in the shape model. The shape model obtains the shape corresponding to the initial shape prior information based on the preset correspondence between the shape prior information and the shape, and this shape is the initial shape of the target object.
[0108] In a possible implementation manner, the acquisition process of the shape model is as follows: The 3D model acquisition device acquires several types of 3D models, and each type of 3D model includes several 3D model samples. The computing and processing device converts each acquired 3D model sample into the SDF expression. The computing and processing device performs dimensionality reduction processing on the SDF value expressions of several 3D model samples of each type one by one based on the PCA algorithm, that is, uses the PCA algorithm to reduce the dimension of each 3D model sample expressed by the SDF value to the target dimension. In some embodiments of the present application, the target dimension can be 5 dimensions. It should be noted that the target dimension being 5 dimensions is only an exemplary illustration and not a specific limitation, that is, the target dimension can be other values (such as 3 dimensions).
[0109] The initial value of the shape prior information preset in the shape model is the average value of several 3D model samples expressed based on SDF values of the corresponding category learned by using the PCA algorithm. At the same time, for the convenience of the driver's understanding, the shape model needs to present the shape prior information in the form of the shape of an object. In some embodiments of the present application, the initial values of the shape prior information of various categories, the correspondence between the shape prior information and the shape, and the shape of the object can be pre-stored in the shape model. After obtaining the shape prior information, the shape of the object is obtained through the pre-stored correspondence between the shape prior information and the shape. In other embodiments of the present application, the initial value of the shape prior information and the correspondence between the shape prior information and the shape can be pre-stored in the shape model. After obtaining the shape prior information, through the pre-stored correspondence between the shape prior information and the shape, the corresponding shape is calculated or drawn by itself. It should be noted that the data pre-stored in the above shape model is only an exemplary illustration and is not a specific limitation on the relevant parameters in the shape model. At the same time, the above shape model can be trained before leaving the factory and can be directly used after leaving the factory; the shape model can also be further optimized and trained during use, that is, the initial values of the shape priors corresponding to various categories in the shape model can be further optimized and updated regularly during use.
[0110] S104. The image processing device optimizes the initial pose and initial shape of the target object based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object.
[0111] To improve the accuracy of the shape and pose of the target object, after obtaining the initial shape of the target object, optimization processing is required to obtain the target pose and target shape of the target object. In some possible implementation manners, the image processing device takes the initial pose and initial shape of the target object as a starting point, uses an iterative algorithm to update the shape and pose of the target object, and uses a target optimization function for optimization verification. In one possible implementation manner, the target optimization function is used to represent the relationship between the current shape and current pose of the target object and the shape and pose in the ideal state.
[0112] Please refer to Figure 3 , in one possible implementation manner of step S104, the image processing device optimizes the initial pose and initial shape of the target object based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object, including:
[0113] S301. The image processing device determines the initial pose and initial shape of the target object as the current pose and shape of the target object.
[0114] S302. The image processing device obtains the current mask of the target object based on the current pose and shape of the target object.
[0115] To reduce the complexity of the optimization process, the image processing device performs dimensionality reduction based on the current shape and pose of the target object, and converts the 3D pose information and shape information into 2D mask information corresponding to the current state of the target object. Then, the current mask information and the target mask information are input into the optimization function for optimization verification. In some possible implementation manners, the dimensionality reduction process may be a projection process, and the projection process may be performed through a rendering algorithm.
[0116] In some possible implementation manners, the optimization function may be:
[0117]
[0118] where p * , s * respectively represent the target pose and shape of the target object, p and s respectively represent the current pose and shape of the target object, Y represents the current mask of the target object under the current pose and shape, and E is the loss information. The loss information E consists of three parts, namely mask alignment loss, point cloud alignment loss, and ground alignment loss.
[0119] In one possible implementation manner of step S302, based on the current pose and shape of the target object, obtaining the current mask of the target object includes: obtaining the current mask of the target object based on a differentiable 3D rendering algorithm based on the SDF value, the current pose and shape of the target object.
[0120] where the differentiable 3D rendering algorithm is:
[0121]
[0122] where P j represents the pixel point on the image plane, π(p j ) ∈ (0, 1) represents its rendering value, the set represents the sampling points on the ray passing through the camera optical center and the pixel p j , φ(x) represents the SDF value of the sampling point x, and ζ represents the parameter of the rendering algorithm, which is specified manually. If a ray passes through the object surface, it means that at least one of the sampling points on this ray has a positive SDF value, resulting in the rendering value of this pixel point approaching 1. Otherwise, it means that the SDF values of the sampling points on this ray are all negative, resulting in the rendering value of this pixel point approaching 0.
[0123] Substitute the shape prior information (such as the SDF value) corresponding to the shape of the target object into the above formula to obtain the current mask corresponding to the current shape of the object.
[0124] In S303, the image processing device obtains the current loss information of the target object based on the current pose and shape of the target object, the current mask of the target object, the target mask, and the point cloud of the target object.
[0125] In a possible implementation, the loss information includes at least one of mask alignment loss information, point cloud alignment loss information, and height alignment loss information. In some preferred embodiments, the loss information at least includes mask alignment loss information.
[0126] The mask alignment loss information is used to measure the distance between the target mask segmented by SAM and the current mask rendered by the current pose and shape. The mask alignment loss is defined as follows:
[0127]
[0128] Among them, Y proj represents the current mask obtained by the rendering algorithm, Y represents the target mask obtained by SAM, and O represents the occlusion mask of the target object, which can be calculated from the target mask obtained by SAM.
[0129] The point cloud alignment loss information is used to represent the alignment degree between the object represented by the current pose and shape of the target object and the point cloud of the target object. The point cloud alignment loss is defined as follows:
[0130]
[0131] Among them, represents the points belonging to the target obtained from the scene LiDAR point cloud (i.e., the target point cloud) by the target mask, that is, represents the points in the point cloud of the target object; represents the points where the light intersects the actual surface of the target object through these points; y represents the intersection point of the light and the surface of the target object at the current pose and shape; φ represents the SDF value function.
[0132] If only the mask alignment loss is optimized, it is often easy for the shape and pose of the object to fall into a local minimum, ultimately reducing the practicality of the method. Therefore, the method further proposes a point cloud alignment loss to improve the practicality of the method.
[0133] The height alignment loss information is used to represent the distance between the height where the current pose and shape are located and the ground. The definition of the height alignment loss information is as follows:
[0134]
[0135] Among them, represents the current 3D position of the target, represents that the current height of the object can be obtained from the object shape. The ground height function Obtained by the RANSAC (Random Sample Consensus) algorithm combined with the target point cloud.
[0136] In order to ensure that the target object should be located on the ground, the height alignment loss is proposed in the embodiments of the present application.
[0137] It should be understood that in the present application, distance is used to represent the difference value or degree of difference, rather than the length concept in physics. For example, the mask alignment loss information is used to measure the difference value or degree of difference between the target mask segmented by SAM and the current mask rendered by the current pose and shape; the height alignment loss information is used to represent the difference value or degree of difference between the height where the current pose and shape are located and the ground.
[0138] In a possible implementation manner of the present application, the definition of the loss information is as follows:
[0139] E = aE mask + bE pc + cE ground
[0140] Wherein, E is the loss information, a is the weight of the mask alignment loss information, b is the weight of the point cloud alignment loss information, and c is the weight of the height alignment loss information.
[0141] S304. When the current loss information of the target object meets the set conditions, the image processing device determines the current pose and shape of the target object as the target pose and target shape of the target object.
[0142] In order to obtain the target pose and target shape, the image processing device needs to set an optimization target. The optimization target can be reflected by the set conditions. In some embodiments, the optimization target can be reflected by a preset optimization function. Based on the preset conditions, the image processing device can know whether the current pose and current shape of the target object still need to be further optimized. If further optimization is required, the optimization steps are repeated. If no further optimization is required, that is, when the set conditions are met, the current pose and shape of the target object are determined as the target pose and target shape of the target object.
[0143] In a possible implementation manner of the present application, the set condition is: the loss information corresponding to the target object in the current pose and shape is convergent. That is, the value of the loss information corresponding to the target object in the current pose and shape is the smallest. The set condition can be reflected by the optimization function of the target object. Among them, the optimization function can be:
[0144]
[0145] Wherein, p * , s* respectively represent the target pose and shape of the target object, p and s respectively represent the current pose and shape of the target object, Y represents the current mask of the target object in the current pose and shape, and E is the loss information. The loss information E consists of three parts, namely the mask alignment loss, the point cloud alignment loss, and the ground alignment loss.
[0146] In a possible implementation manner of the present application, the set condition is: the loss information of the target object in the current pose and shape is the first loss information, the loss information of the target object in the previous pose and shape is the second loss information, and the difference between the first loss information and the second loss information is less than the set threshold. In a preferred implementation manner, the first loss information is less than the second loss information, and the value obtained by subtracting the first loss information from the second loss information is less than the set threshold.
[0147] In a possible implementation manner of step S104, step S104 further includes:
[0148] When the current loss information of the target object does not meet the set condition, update the current pose and shape of the target object based on a preset update method, and repeat the step of obtaining the current mask of the target object based on the current pose and shape of the target object, and the step of obtaining the current loss information of the target object based on the current pose and shape of the target object, the current mask of the target object, the target mask, and the point cloud of the target object until the current loss information of the target object meets the set condition.
[0149] In this implementation manner, updating the current pose and shape of the target object based on a preset update method includes: updating the current pose and shape of the target object based on a preset gradient descent algorithm with a preset step size or a preset number of steps.
[0150] In this implementation manner, update the SDF implicit expression vector corresponding to the position information, attitude, and shape of the target object based on a preset gradient descent algorithm with a set step size. The gradient descent algorithm has the advantages of wide applicability and strong scalability. Through the gradient descent algorithm, the target pose and shape of the target object can be obtained quickly and effectively, which is beneficial to improving the calculation efficiency.
[0151] In a possible implementation manner, the gradient descent algorithm can be implemented by an Adaptive Moment Estimation (Adam) optimizer. The gradient descent algorithm can be a batch gradient descent algorithm or a stochastic gradient descent algorithm, etc.
[0152] In a possible implementation manner of the present application, the current loss information of the target object becomes smaller as the current pose and shape of the target object are updated. When the loss information reaches the minimum value, the current pose and shape of the target object are the target pose and the target shape.
[0153] Please refer to Figure 4 , an embodiment of the present application provides an image processing method, including:
[0154] S401, the image processing device obtains a target mask corresponding to a target object in the target image.
[0155] For the specific implementation process of step S401, reference can be made to step S101, which will not be elaborated here.
[0156] S402, the image processing device obtains the point cloud of the target object in the target point cloud based on the target mask.
[0157] For the specific implementation process of step S402, reference can be made to step S102, which will not be elaborated here.
[0158] S403, the image processing device determines the initial pose of the target object based on the point cloud of the target object, and determines the initial shape of the target object based on the shape prior information.
[0159] For the specific implementation process of step S403, reference can be made to step S103, which will not be elaborated here.
[0160] S404, the image processing device optimizes the initial pose and initial shape of the target object based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object.
[0161] For the specific implementation process of step S404, reference can be made to step S104, which will not be elaborated here.
[0162] S405, the image processing device labels the target object based on the target pose and target shape.
[0163] After the image processing device obtains the target pose and target shape, it can automatically label the target pose and target shape corresponding to the target object in the target space. In some embodiments, the image processing device divides the target space into several grids. The position and pose of the target object can be marked from the corresponding grid in the target space through the target pose of the target object, and then the shape corresponding to the target object can be marked at the grid position and its surrounding grids by using the target shape, so that the target position can be marked in the target space. In a possible implementation manner, the image processing device can mark the position and pose of the target object in the target space through the target pose of the target object at the corresponding point in the target space (this point is the corresponding point of the center point of the target object in the target space), and then mark the points constituting the target shape around the corresponding point of the center point of the target object in the target space according to the target shape of the target object. In a possible implementation manner, the image processing device can also use the bounding box annotation algorithm and mark the bounding box corresponding to the target object in the target space based on the target shape and target pose of the target object.
[0164] In a possible implementation manner of this embodiment, the image processing device labels the target object in the 3D space based on the target pose and target shape. For example, the image processing device labels the 3D bounding box and category of the target object in the 3D space based on the target pose and target shape. The image processing device can also label the target object in the 3D space in the form of a voxel occupancy grid based on the target pose and target shape.
[0165] The image processing method provided by the embodiments of this application obtains the initial pose of the target object based on the point cloud of the target object and determines the initial shape of the target object based on the shape prior information. After obtaining the initial pose and shape, the initial pose and initial shape are optimized based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object. The initial shape obtained based on the shape prior information is beneficial to the optimization process, and the target pose and target shape obtained by optimizing based on the initial shape and pose have good data generalization and excellent accuracy, and the accuracy of the target pose and target shape obtained by using the image processing method provided by the embodiments of the present invention has good stability.
[0166] The method proposed by the present invention uses the object segmentation result and a small amount of vehicle CAD (Computer Aided Design) models as prior information, and generates accurate 3D poses and shapes through an iterative optimization method. Therefore, the technical solution provided by the embodiments of the present invention can provide an automatic or semi-automatic annotation solution for various data annotation requirements related to autonomous driving perception without the ground truth information of the target data domain.
[0167] To better understand the effect of the image processing method provided by the embodiments of the present application, the present application is verified by the publicly available KITTI dataset.
[0168] The KITTI dataset was jointly created by the Karlsruhe Institute of Technology in Germany and the Toyota Technological Institute in the United States. It is currently the largest computer vision algorithm evaluation dataset in the field of autonomous driving in the world. This dataset is used to evaluate the performance of computer vision technologies such as stereo images, optical flow, visual odometry, 3D object detection, and 3D tracking in a vehicle-mounted environment. KITTI contains real image data collected from scenarios such as urban areas, rural areas, and highways. Each image contains up to 15 vehicles and 30 pedestrians, as well as various degrees of occlusion and truncation. The entire dataset consists of 389 pairs of stereo images and optical flow maps, 39.2 km of visual odometry sequences, and images of more than 200k 3D annotated objects, sampled and synchronized at a frequency of 10 Hz.
[0169] Please refer to Figure 5 , during the verification process, for the automatic annotation of 3D bounding boxes, the present invention uses the average precision (AP BEV ) in the Bird's Eye View (BEV) perspective and the average precision (AP 3D ) of 3D bounding boxes as the measurement criteria. In the generalization comparison, the present invention uses the NuScenes detection score (NDS) and the mean average precision (mAP). Note that the larger the above metrics, the better.
[0170] Please refer to Figure 6 , the generalization and fine-grained annotation of the present invention are verified through the publicly available NuScenes dataset, which is a large-scale multi-modal dataset for 3D detection and map segmentation. The dataset is divided into 700 / 150 / 150 scenes for training / verification / testing. It contains data from multiple sensors, including six cameras, one lidar, and five radars. For camera input, each frame contains six views of the surrounding environment at a specific timestamp. The present invention adjusts the input views to a resolution of 256x704 and voxelizes the point clouds at 0.075m and 0.1m respectively for detection and segmentation.
[0171] By using the above publicly available datasets for verification, the following conclusions can be drawn:
[0172] Beneficial effect 1: The image processing method provided by the embodiments of the present application achieves optimal performance in the quality of generating pseudo-labels.
[0173] When directly comparing the pseudo-labels with the ground truth data, as can be seen from the upper part of Table 1, the test model of the present invention on the KITTI dataset achieved the best results in the automatic annotation task of 3D object detection.
[0174] VS3D (weakly supervised 3D object detection) in Table 1 is a weakly supervised 3D object detection model based on point clouds. VS3D proposes an unsupervised 3D object proposal module (UPM). UPM uses the geometric and density properties of the laser point cloud to find high-confidence regions that may contain the target object. In addition, to obtain a more accurate 3D bounding box, VS3D also designs a cross-modal transfer learning method: the detection network based on point clouds is regarded as a student and learns knowledge from an existing pre-trained off-the-shelf teacher image detection network. The 3D object proposals generated by UPM are projected onto the paired images and classified by the teacher network, and then the student network mimics the behavior of the teacher during training. However, the limitations of the VS3D framework are also obvious. First, in the VS3D framework, a pre-trained 2D object detection model is required. Therefore, when the data to be annotated is inconsistent with the pre-training, due to the domain difference, the performance of the object detection network will decline, ultimately leading to a decline in the performance of the 3D detection grid. Second, VS3D distills information from the teacher network of 2D object detection and lacks 3D supervision signals, resulting in poor performance of 3D detection and very limited practical applications.
[0175] The SDFLabel (signed distance fields Label) in Table 1 consists of a deep SDF network pre-trained on a virtual dataset and the RANSAC algorithm. This algorithm proposes a new differentiable rendering function based on SDF to train the deep SDF network to encode the object to be labeled. The input of the deep SDF network is the RGB (red green blue) patch of the object, and the output of the deep SDF network is the implicit expression vector of the SDF value of the object in the image. When performing labeling, the algorithm first obtains the 2D bounding box of the target object and the point cloud belonging to the object in the point cloud, and inputs the 2D bounding box into the deep SDF to obtain the SDF expression of the object in the box. Finally, the RANSAC algorithm is used, combined with the output of the deep SDF network, to estimate the 3D pose of the target object. Since this technical solution heavily relies on the quality of the pre-trained deep SDF network, but the training of the deep SDF is on a virtual dataset, which often has a large gap with the data in the real world, resulting in a low quality of the shape prediction of the target object on the real dataset, which directly affects the quality of the final pseudo-label. In addition, SDFLabel is not an end-to-end algorithm, and the result predicted by the deep SDF is fixed in the subsequent RANSAC optimization, which further reduces the accuracy of the 3D bounding box.
[0176] The FGR (Frustum Aware Geometric Reasoning) in Table 1 uses 2D bounding box hints and sparse point clouds to automatically label 3D bounding boxes. It consists of two components: coarse 3D point cloud segmentation and 3D bounding box estimation. In the coarse 3D segmentation stage, FGR proposes a context-aware adaptive region growing method, which uses context information and adaptively adjusts the threshold to roughly segment the target point cloud. In the 3D bounding box estimation stage, it designs a noise-resistant framework to estimate the bounding rectangle in the bird's-eye view and locate the key vertices. Then, a rule-based iterative method is used to estimate the 3D bounding box. The object bounding box used by FGR is not the common object detection box, but the complete object box with occlusion information. This kind of object bounding box usually contains certain object size information, and the current 2D object bounding box and detection algorithm cannot be used as the input of FGR, which limits the practical application of FGR. When FGR predicts the rotation information of the 3D bounding box, it cannot solve the problem of the bounding box flipping, so the true value information of the object rotation is required. In addition, when FGR estimates the 3D pose of the object, it uses a rule-based algorithm, which results in poor generalization performance of FGR.
[0177] LAD (Lifting 2d object locations to 3d and Discounting Lidar outliers) in Table 1 To overcome the problem of too much noise in the point cloud, a self-supervised method based on template matching is used. Specifically, the model first obtains the mask information of the object to be labeled through a Mask Region-based Convolutional Neural Network (Mask R-CNN), and obtains the noisy point cloud belonging to the object in the point cloud through the mask. The LAD framework consists of two parts: a point cloud segmentation network and a 3D pose regression network. LAD uses a fixed 3D model as a template to provide a supervision signal to train the point cloud segmentation and 3D pose regression networks. In addition, LAD can also use the temporal information of the video to refine the pseudo-labels. This scheme is trained using a self-supervised method based on template matching. It is crucial to select a suitable template 3D model, which directly determines the quality of the supervision signal and ultimately affects the quality of the 3D bounding box. In addition, the 3D model used as a template is fixed during training and cannot well match various different situations, which also makes LAD have poor generalization performance.
[0178] For the convenience of comparison, the image processing method provided in the embodiment of the present application is named SLF (Segment, Lift and Fit).
[0179] It can be seen from Table 1 that in the cases of medium difficulty and difficult difficulty, the image processing method (SLF) provided in the embodiment of the present application far outperforms the existing automatic annotation algorithms.
[0180] At the same time, the lower part also shows the performance of SLF in the automatic and semi-automatic annotation processes. Among them, SLF GTmask refers to the SLF automatic annotation algorithm with the ground truth mask as the input. SLF point represents the semi-automatic annotation process, that is, the annotator clicks on several points belonging to the target object. SLF box represents the fully automatic annotation process, that is, the 2D bounding box of the object is obtained by the detection algorithm; the "-" in the table indicates that there is no corresponding measurement result for this indicator. It can be found that SLF far exceeds the existing algorithms in both the automatic and semi-automatic annotation processes.
[0181] Table 1
[0182]
[0183] Beneficial effect 2: Through the image processing method provided by the embodiments of the present application, using pseudo-labels to train 3D detectors has excellent performance. Using the image processing method provided by the embodiments of the present application, pseudo-labels are used to train multiple different 3D detectors and evaluations are carried out. As shown in Table 2, among them, PointPillar is a point cloud object detection algorithm based on pillars; SECOND (Sparsely Embedded CONvolutional Detection) is a 3D detection algorithm based on sparse convolution; Part-A 2 (Part-Aware and Aggregation) is a 3D detection algorithm based on part awareness and aggregation; VoxelR-CNN is a region convolutional neural network based on voxel grids; GT (Ground truth) is the true value. It can be seen from Table 2 that: on the KITTI dataset, the evaluation indicators of all 5 3D detection algorithms trained with 3D pseudo-labels generated by the image processing method provided by the embodiments of the present application are significantly better than the unsupervised automatic annotation algorithm FGR, and can even be compared with the performance of detectors trained with KITTI true values.
[0184] Table 2
[0185]
[0186]
[0187] Beneficial effect 3: Using the image processing method provided by the embodiments of the present application for cross-dataset annotation can achieve excellent generalization performance.
[0188] The present invention compares the performance of the current best unsupervised annotation algorithm FGR and the supervised annotation algorithm MTrans (multimodal transformer) on cross-datasets in Table 3. It can be seen from the table that since FGR is a rule-based annotation algorithm, it has good results on KITTI, but it cannot obtain effective measurement results in data outside the data domain (nuScenes). And Mtrans, due to its training on the KITTI dataset, can obtain excellent results in KITTI data. However, as mentioned above, supervised annotation algorithms often overfit the annotation patterns of specific datasets and perform poorly in unseen data. SLF is an unsupervised annotation algorithm that does not require any 3D annotations in advance. It can be seen from the mAP index that the cross-dataset generalization performance of SLF is significantly better than that of Mtrans.
[0189] Table 3
[0190]
[0191] Technical effect 4: Conduct fine-grained 3D automatic annotation with excellent results.
[0192] Please refer to Figure 7 , the present invention conducts fine-grained annotation experiments in the nuScenes dataset. Specifically, it annotates the voxel occupancy grid of the target object. Since the ground truth information of the voxel occupancy grid is not provided in the nuScenes dataset, the present invention provides a qualitative comparison with existing annotation methods in Figure 10 . Among them, in Figure 7 , in the second and third columns are the ground truth information of the voxel occupancy grid provided by OFN (occupancy dataset for nuscenes) and OpenOccupancy respectively, and the fourth column is the ground truth information annotated by SLF. It can be clearly seen that compared with the existing annotation methods, the voxel occupancy truth annotated by SLF can provide more detailed and complete shape information.
[0193] Please refer to Figure 8 , an embodiment of the present application further provides an image processing device 100, including: an acquisition unit 101 and a processing unit 102. Among them, the acquisition unit 101 is used to acquire the target mask of the target object in the target image. The acquisition unit 101 is also used to obtain the point cloud of the target object in the target point cloud based on the target mask. The processing unit 102 is used to determine the initial pose of the target object based on the point cloud of the target object and determine the initial shape of the target object based on the shape prior information. The processing unit 102 is also used to optimize the initial pose and initial shape of the target object based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object.
[0194] In an embodiment of the present application, the acquisition unit 101 includes: a target mask acquisition module. The target mask acquisition module is used to acquire the target mask of the target object.
[0195] In an embodiment of the present application, the acquisition unit 101 includes: a point cloud acquisition module and a separation module. Among them, the point cloud acquisition module is used to acquire the target point cloud. The separation module is used to segment the target point cloud based on the target mask to obtain the point cloud of the target object.
[0196] In an embodiment of the present application, the acquisition unit 101 further includes: a denoising module. The denoising module is used to denoise the initial point cloud of the target object. Thus, the point cloud of the target object after denoising can be obtained.
[0197] Among them, the target point cloud includes information about the target object and the environment where the target object is located.
[0198] In an embodiment of the present application, the denoising module is further configured to remove cloud points whose distance from the center point of the point cloud of the target object is greater than a first threshold. Thereby, the denoised point cloud of the target object can be obtained.
[0199] In an embodiment of the present application, the processing unit 102 is further configured to perform a mean processing on the point cloud of the target object to obtain the initial pose of the target object.
[0200] In an embodiment of the present application, the processing unit 102 is further configured to obtain an initial value of the shape prior information and obtain the initial shape of the target based on the initial value of the shape prior information.
[0201] In an embodiment of the present application, the processing unit 102 is further configured to use the average shape of the shape prior information as the initial shape of the target object.
[0202] In an embodiment of the present application, the processing unit 102 is further configured to obtain the shape prior information corresponding to the target object based on a preset dimensionality reduction algorithm.
[0203] Wherein, the shape prior information includes SDF values. The dimensionality reduction algorithm includes the PCA algorithm.
[0204] In an embodiment of the present application, the processing unit 102 is further configured to determine the initial pose and initial shape of the target object as the current pose and shape of the target object, and obtain the current mask of the target object based on the current pose and shape of the target object, and obtain the current loss information of the target object based on the current pose and shape of the target object, the current mask of the target object, the target mask, and the point cloud of the target object. When the current loss information of the target object meets the set conditions, the processing unit 102 determines the current pose and shape of the target object as the target pose and target shape of the target object.
[0205] In one implementation, the loss information includes at least one of: mask alignment loss information, point cloud alignment loss information, and height alignment loss information. Among them, the mask alignment loss information is used for the distance between the target mask and the current mask of the target object. The point cloud alignment loss information is used to represent the distance between the current pose and shape of the target object and the point cloud of the target object. The height alignment loss information is used to represent the distance between the height where the current pose and shape are located and the ground.
[0206] In an embodiment of the present application, the processing unit 102 is further configured to, when the current loss information of the target object does not meet the set conditions, update the current pose and shape of the target object based on a preset update method, and repeat the steps of obtaining the current mask of the target object based on the current pose and shape of the target object, and the steps of obtaining the current loss information of the target object based on the current pose and shape of the target object, the current mask of the target object, the target mask, and the point cloud of the target object, until the current loss information of the target object meets the set conditions.
[0207] In an embodiment of the present application, the processing unit 102 is further configured to obtain the current mask of the target object based on a differentiable 3D rendering algorithm of the SDF value, the current pose, and the shape of the target object.
[0208] In an embodiment of the present application, the processing unit 102 is further configured to update the current pose and shape of the target object based on a preset gradient descent algorithm with a preset step size or a preset number of steps.
[0209] In a possible implementation manner of this embodiment, the set condition is that the loss information corresponding to the target object in the current pose and shape is convergent.
[0210] In a possible implementation manner of this embodiment, in a possible implementation manner of the present application, the set condition is that the loss information of the target object in the current pose and shape is the first loss information, the loss information of the target object in the previous pose and shape is the second loss information, and the difference between the first loss information and the second loss information is less than a set threshold. In a preferred implementation manner, the first loss information is less than the second loss information, and the value obtained by subtracting the first loss information from the second loss information is less than the set threshold.
[0211] Please refer to Figure 9 , in an embodiment of the present application, the image processing device 100 further includes: an annotation unit 103. The annotation unit 103 is configured to annotate the target object based on the target pose and the target shape.
[0212] An embodiment of the present application further provides an electronic device, including a memory for storing computer program instructions and a processor for executing the program instructions. When the computer program instructions are executed by the processor, the electronic device implements the method provided in any of the foregoing embodiments.
[0213] In an embodiment of the present application, the electronic device is a chip. In a possible implementation manner, the chip is applied to cloud processing to provide a driving basis for the driving system or navigation of a vehicle.
[0214] Please refer to Figure 10, in an embodiment of the present application, the electronic device 200 may include: a processor 201, a memory 202, and a communication unit 203. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the electronic device shown in the figure does not constitute a limitation on the embodiments of the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0215] Among them, the communication unit 203 is used to establish a communication channel so that the electronic device can communicate with other devices. Receive user data sent by other devices or send user data to other devices.
[0216] The processor 201 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing software programs, instructions, and / or modules stored in the memory 202, and by calling the data stored in the memory, it performs various functions of the electronic device and / or processes data. The processor may be composed of an integrated circuit (IC). For example, it may be composed of a single packaged IC, or may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 201 may only include a central processing unit (CPU). In the embodiments of the present invention, the CPU may be a single arithmetic core or may include multiple arithmetic cores.
[0217] The memory 202 is used to store the execution instructions of the processor 201. The memory 202 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0218] When the execution instructions in the memory 202 are executed by the processor 201, the electronic device 200 is enabled to execute Figure 1 some or all of the steps in the illustrated embodiment.
[0219] An embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program is executed by at least one processor, the device where the computer-readable storage medium is located is controlled to implement the method provided in any of the foregoing embodiments.
[0220] An embodiment of the present application further provides a vehicle. The vehicle is equipped with a driving system or a navigation system. The driving system or the navigation system includes image information, and the image information is obtained by the image processing method provided in any of the foregoing embodiments. Alternatively, the vehicle includes the above-mentioned image processing device or the above-mentioned electronic device.
[0221] An embodiment of the present application provides a computer program product. The computer program product contains executable instructions. When the executable instructions are executed on a computer, the computer is caused to execute the method of the foregoing embodiment.
[0222] The technical solution provided by the embodiment of the present invention can be applied to the following scenarios:
[0223] Application scenario 1: Automatic annotation of 3D object detection.
[0224] As an extremely important part of the visual perception system of autonomous vehicles, 3D object detection can provide traffic scene understanding by detecting common traffic participants such as pedestrians and vehicles. Existing automatic annotation technologies either require a small amount of 3D ground truth as starting data or predict 3D poses based on self-supervised algorithms. The former has poor cross-dataset generalization ability, and the latter has poor annotation quality. These drawbacks cannot meet the automatic annotation requirements of 3D data. The present invention uses object shape priors and 2D cues, and directly optimizes the 3D pose of the object using gradient descent, achieving better cross-data generalization and 3D bounding box quality compared to previous algorithms.
[0225] Application scenario 2: Automatic annotation of voxel occupancy grids.
[0226] The image processing technology of the present invention can generate accurate shape information and 3D poses of objects, and can be used for 3D voxel occupancy grid prediction tasks.
[0227] The structure, features and effects of the present invention have been described in detail based on the embodiments shown in the drawings. The above are only the preferred embodiments of the present invention, but the present invention is not limited to the implementation scope shown in the drawings. Any changes made according to the concept of the present invention, or modified into equivalent embodiments with equivalent changes, still within the spirit covered by the specification and the drawings, shall be within the protection scope of the present invention.
Claims
1. An image processing method, characterized in that, Including: Obtaining a target mask corresponding to a target object in a target image; Based on the target mask, obtaining the point cloud of the target object in the target point cloud; Based on the point cloud of the target object, determining the initial pose of the target object, and based on shape prior information, determining the initial shape of the target object; Based on the target mask and the point cloud of the target object, performing optimization processing on the initial pose and initial shape of the target object to obtain the target pose and target shape of the target object.
2. The image processing method according to claim 1, characterized in that, The obtaining the point cloud of the target object in the target point cloud based on the target mask includes: Obtaining the target point cloud, where the target point cloud includes information of the target object and the environment where the target object is located; Segmenting the target point cloud based on the target mask to obtain the point cloud of the target object.
3. The image processing method according to claim 2, characterized in that, It further includes denoising the point cloud of the target object.
4. The image processing method according to claim 3, characterized in that, The denoising the point cloud of the target object includes: Removing the point cloud whose distance from the center point of the point cloud of the target object is greater than a first threshold.
5. The image processing method according to any one of claims 1 to 4, characterized in that, The obtaining the initial pose of the target object based on the point cloud of the target object includes: Performing a mean process on the point cloud of the target object to obtain the initial pose of the target object.
6. The image processing method according to any one of claims 1 to 4, characterized in that, The shape prior information includes SDF values, and the shape prior information is obtained based on a preset dimensionality reduction algorithm.
7. The image processing method according to claim 6, characterized in that, The dimensionality reduction algorithm includes the PCA algorithm.
8. The image processing method according to any one of claims 1 to 4, characterized in that, The determining the initial shape of the target object based on the shape prior information includes: using the average shape of the shape prior information as the initial shape of the target object.
9. The image processing method according to any one of claims 1 to 4, characterized in that, The performing optimization processing on the initial pose and initial shape of the target object based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object includes: Determining the initial pose and initial shape of the target object as the current pose and shape of the target object; Based on the current pose and shape of the target object, obtaining the current mask of the target object; Based on the current pose and shape of the target object, the current mask of the target object, the target mask, and the point cloud of the target object, obtaining the current loss information of the target object; When the current loss information of the target object meets a set condition, determining the current pose and shape of the target object as the target pose and target shape of the target object.
10. The image processing method according to claim 9, characterized in that, The loss information includes at least one of mask alignment loss information, point cloud alignment loss information, and height alignment loss information; wherein, the mask alignment loss information is used for the distance between the target mask and the current mask of the target object; the point cloud alignment loss information is used to represent the alignment degree between the object represented by the current pose and shape of the target object and the point cloud of the target object; the height alignment loss information is used to represent the distance between the height where the current pose and shape are located and the ground.
11. The image processing method according to claim 9, characterized in that, It further includes: When the current loss information of the target object does not meet the set conditions, update the current pose and shape of the target object based on a preset update method, and repeat the steps of obtaining the current mask of the target object based on the current pose and shape of the target object, and the steps of obtaining the current loss information of the target object based on the current pose and shape of the target object, the current mask of the target object, the target mask, and the point cloud of the target object, until the current loss information of the target object meets the set conditions.
12. The image processing method according to any one of claims 9 to 11, characterized in that, Obtaining the current mask of the target object based on the current pose and shape of the target object includes: obtaining the current mask of the target object based on a differentiable 3D rendering algorithm based on SDF values, the current pose and shape of the target object.
13. The image processing method according to claim 11, characterized in that,The updating the current pose and shape of the target object based on a preset update method includes: Updating the current pose and shape of the target object based on a preset gradient descent algorithm with a preset step size or a preset number of steps.
14. The image processing method according to claim 9, wherein, The set condition is that the loss information corresponding to the target object in the current pose and shape is convergent.
15. The image processing method according to claim 1, wherein, It further includes: Labeling the target object based on the target pose and target shape.
16. An image processing apparatus, wherein, It includes: An acquisition unit for acquiring the target mask of the target object in the target image; The acquisition unit is further used to obtain the point cloud of the target object in the target point cloud based on the target mask; A processing unit for determining the initial pose of the target object based on the point cloud of the target object and determining the initial shape of the target object based on shape prior information; The processing unit is further used to optimize the initial pose and initial shape of the target object based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object.
17. The image processing apparatus according to claim 16, wherein, The acquisition unit includes a point cloud acquisition module and a separation module; The point cloud acquisition module is used to acquire the target point cloud, and the target point cloud includes information about the target object and the environment where the target object is located; The separation module is used to segment the target point cloud based on the target mask to obtain the point cloud of the target object.
18. The image processing apparatus according to claim 17, wherein, The acquisition unit further includes a denoising module; the denoising module is used to denoise the point cloud of the target object.
19. The image processing apparatus according to claim 18, wherein, The denoising module is used to remove the point cloud whose distance from the center point of the point cloud of the target object is greater than the first threshold.
20. The image processing apparatus according to any one of claims 16 to 19, wherein, The processing unit is used to perform mean processing on the point cloud of the target object to obtain the initial pose of the target object.
21. The image processing apparatus according to any one of claims 16 to 19, wherein, The processing unit is used to obtain the shape prior information based on a preset dimensionality reduction algorithm, and the shape prior information includes SDF values.
22. The image processing apparatus according to claim 21, wherein, The dimensionality reduction algorithm includes the PCA algorithm.
23. The image processing apparatus according to any one of claims 16 to 19, wherein, The processing unit is used to use the average shape of the shape prior information as the initial shape of the target object.
24. The image processing apparatus according to any one of claims 16 to 19, wherein, The processing unit is configured to determine the initial pose and initial shape of the target object as the current pose and shape of the target object, and based on the current pose and shape of the target object, obtain the current mask of the target object, and based on the current pose and shape of the target object, the current mask of the target object, the target mask, and the point cloud of the target object, obtain the current loss information of the target object; When the current loss information of the target object meets the set conditions, the processing unit determines the current pose and shape of the target object as the target pose and target shape of the target object.
25. The image processing apparatus according to claim 24, wherein, The loss information includes at least one of mask alignment loss information, point cloud alignment loss information, and height alignment loss information; wherein, the mask alignment loss information is used for the distance between the target mask and the current mask of the target object; the point cloud alignment loss information is used to represent the alignment degree between the object represented by the current pose and shape of the target object and the point cloud of the target object; the height alignment loss information is used to represent the distance between the height where the current pose and shape are located and the ground.
26. The image processing apparatus according to claim 24, wherein, The processing unit is further configured to, when the current loss information of the target object does not meet the set conditions, update the current pose and shape of the target object based on a preset update method, and repeat the step of obtaining the current mask of the target object based on the current pose and shape of the target object and the step of obtaining the current loss information of the target object based on the current pose and shape of the target object, the current mask of the target object, the target mask, and the point cloud of the target object, until the current loss information of the target object meets the set conditions.
27. The image processing apparatus according to any one of claims 24 to 26, wherein The processing unit is further configured to obtain the current mask of the target object based on a differentiable 3D rendering algorithm of the SDF value and the current pose and shape of the target object.
28. The image processing apparatus according to claim 26, wherein The processing unit is further configured to update the current pose and shape of the target object based on a preset gradient descent algorithm with a preset step size or preset number of steps.
29. The image processing apparatus according to claim 24, wherein The set condition is that the loss information corresponding to the target object in the current pose and shape is convergent.
30. The image processing apparatus according to claim 16, wherein It further includes: A labeling unit for labeling the target object based on the target pose and target shape.
31. An electronic device, wherein It includes a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to implement the method according to any one of claims 1-15.
32. A computer-readable storage medium, wherein The computer-readable storage medium includes a stored program, wherein when the program is executed by at least one processor, the device where the computer-readable storage medium is located is controlled to implement the method according to any one of claims 1-15.
Citation Information
Cited By
Image processing method, apparatus, device and storage medium
WO2025123872A1