Image processing method, apparatus, device and storage medium

By acquiring the target mask and point cloud in the target image, determining and optimizing the initial position and shape of the target object, the problem of unstable and poor accuracy of the automatic labeling of 3D objects in the prior art is solved, and the target position and shape with high accuracy and stability is achieved.

WO2025123872A1PCT designated stage expired Publication Date: 2025-06-19HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/121760
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2024-09-27
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The existing 3D object automatic labeling methods have problems with unstable labeling results and poor labeling accuracy, especially when labeling data outside the data distribution, the generalization performance is poor.

Method used

By obtaining the target mask in the target image, the target object point cloud in the target point cloud is obtained, the initial position and shape of the target object is determined based on the point cloud, and the shape prior information is used for optimization processing to obtain the target position and shape.

Benefits of technology

The accuracy and stability of the target position and target shape are improved, data generalization is enhanced, and the problems of unstable labeling results and poor accuracy are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024121760_19062025_PF_FP_ABST
    Figure CN2024121760_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are an image processing method, an apparatus, a device and a storage medium. The method comprises: acquiring a corresponding target mask of a target object in a target image; on the basis of the target mask, obtaining a point cloud of the target object from a target point cloud; on the basis of the point cloud of the target object, determining an initial pose of the target object, and, on the basis of shape prior information, determining an initial shape of the target object; and, on the basis of the target mask and the point cloud of the target object, optimizing the initial pose and the initial shape of the target object, so as to obtain a target pose and a target shape of the target object. The present application is used for solving the problems of unstable annotation results and low annotation accuracy during 3D object annotation in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method, device, equipment and storage medium

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 14, 2023, with application number 202311726862.8 and application name “An image processing method, device, equipment and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of image processing technology, and in particular to an image processing method, apparatus, device and storage medium. Background Art

[0003] In modern robotics and autonomous driving, existing algorithms require large amounts of labeled data to train deep learning models to understand 3D (three-dimensional) scenes, especially dynamic objects such as vehicles and pedestrians. Therefore, there is an urgent need to improve the automation of 3D labeling. Current 3D automatic labeling methods can be divided into two categories based on whether or not ground-truth bounding boxes are required: 1) Supervised automatic labeling methods: These methods typically require obtaining manually annotated 3D bounding boxes, then train a 3D detection model based on these manually annotated 3D bounding boxes. Finally, the trained model is used to predict 3D bounding boxes for the remaining data to achieve automatic labeling. However, due to the limited amount of data required for model training, these methods often suffer from overfitting to a specific dataset when labeling 3D objects, resulting in poor generalization performance when labeling data outside the data distribution. 2) Unsupervised automatic labeling methods: Compared to supervised labeling methods, although unsupervised automatic labeling methods do not require manually annotated 3D bounding boxes, they lack effective supervisory signals, resulting in poor pseudo-label quality and limited practical value. Therefore, the above two methods have the problem of unstable labeling results or poor labeling accuracy when automatically labeling 3D objects.

[0004] Summary of the Invention

[0005] The purpose of this application is to provide an image processing method, device, equipment and storage medium to solve the problems of unstable labeling results and poor labeling accuracy when labeling 3D objects in the existing technical problems.

[0006] In a first aspect, an embodiment of the present application provides an image processing method, comprising: obtaining a target mask corresponding to a target object in a target image; obtaining a point cloud of the target object in a target point cloud based on the target mask; determining an initial pose of the target object based on the point cloud of the target object, and determining an initial shape of the target object based on shape prior information; optimizing the initial pose and initial shape of the target object based on the target mask and the point cloud of the target object to obtain a target pose and target shape of the target object.

[0007] The image processing method provided in the embodiment of the present application determines the initial pose of the target object based on the point cloud of the target object, and determines the initial shape of the target object based on the shape prior information. After obtaining the initial pose and shape, the initial pose and initial shape are optimized based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object. The initial shape obtained based on the shape prior information has a high accuracy, which helps to shorten the optimization time. At the same time, the target pose and target shape obtained by optimizing the initial shape and pose have good data generalization and excellent accuracy, and the accuracy of the target pose and target shape obtained by the image processing method provided by the embodiment of the present invention has good stability.

[0008] In an implementation of the first aspect, obtaining a point cloud of a target object in a target point cloud based on a target mask includes: acquiring a target point cloud, the target point cloud including information about the target object and the environment in which the target object is located; and segmenting the target point cloud based on the target mask to obtain a point cloud of the target object.

[0009] Based on the target mask, the point cloud of the object is obtained from the target point cloud, which can ensure the accuracy of the point cloud of the target object and obtain the pose of the target object with higher accuracy.

[0010] In an implementation manner of the first aspect, the method further includes denoising the point cloud of the target object.

[0011] The point cloud of the target object obtained from the target point cloud by the target mask contains a noise point cloud. Denoising the point cloud of the target object can further improve the accuracy of the position and posture of the target object.

[0012] In an implementation of the first aspect, denoising the point cloud of the target object includes removing point clouds whose distance from a center point of the point cloud of the target object is greater than a first threshold.

[0013] Considering point clouds whose distance from the center point of the point cloud of the target object is greater than a first threshold as noise points can effectively reduce the complexity of denoising.

[0014] In an implementation of the first aspect, acquiring the initial pose of the target object based on the point cloud of the target object includes: performing mean processing on the point cloud of the target object to obtain the initial pose of the target object.

[0015] The mean processing based on the point cloud of the target object is easy to implement, and the accuracy of the initial position of the target object obtained is high.

[0016] In an implementation manner of the first aspect, the shape prior information includes an SDF value, and the shape prior information is obtained based on a preset dimensionality reduction algorithm.

[0017] Shape prior information can well represent the initial shape of the target object, making it easier to reduce the complexity of subsequent optimization processing. To reduce the complexity of subsequent optimization processing, it is necessary to reduce the dimension of the SDF value based on a preset dimensionality reduction algorithm to obtain shape prior information containing the SDF value.

[0018] In an implementation manner of the first aspect, the dimensionality reduction algorithm includes a PCA algorithm.

[0019] The PCA algorithm is used for dimensionality reduction, which is conducive to implementation.

[0020] In an implementation manner of the first aspect, determining the initial shape of the target object based on the shape prior information includes: taking an average shape of the shape prior information as the initial shape of the target object.

[0021] Using the average shape of the shape prior information as the initial shape of the target object can make the initial shape of the target object more compatible and more applicable.

[0022] In an implementation of the first aspect, the initial pose and initial shape of the target object are optimized based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object, including: determining the initial pose and initial shape of the target object as the current pose and shape of the target object; obtaining the current mask of the target object based on the current pose and shape of the target object; obtaining the current loss information of the target object based on the current pose and shape of the target object, the current mask of the target object, the target mask and the point cloud of the target object; when the current loss information of the target object meets the set conditions, determining the current pose and shape of the target object as the target pose and target shape of the target object.

[0023] The loss information can well verify the difference between the current shape and current posture and the actual shape and posture of the target object.

[0024] In an implementation of the first aspect, the loss information includes: at least one of: mask alignment loss information, point cloud alignment loss information, and height alignment loss information; wherein, the mask alignment loss information is used for the distance between the target mask and the current mask of the target object; the point cloud alignment loss information is used to indicate the degree of alignment between the object represented by the current posture and shape of the target object and the point cloud of the target object; the height alignment loss information is used to indicate the distance between the height of the current posture and shape and the ground.

[0025] In order to improve the accuracy of the target pose and target shape, the loss information includes at least one of mask alignment loss information, point cloud alignment loss information and height alignment loss information.

[0026] In an implementation method of the first aspect, it also includes: when the current loss information of the target object does not meet the set conditions, the current posture and shape of the target object are updated based on a preset update method, and the steps of obtaining the current mask of the target object based on the current posture and shape of the target object and the steps of obtaining the current loss information of the target object based on the current posture and shape of the target object, the current mask of the target object, the target mask and the point cloud of the target object are repeated until the current loss information of the target object meets the set conditions.

[0027] Updating the current position and shape of the target object based on a preset update method can improve the efficiency of the optimization process.

[0028] In an implementation of the first aspect, obtaining a current mask of the target object based on the current pose and shape of the target object includes: obtaining the current mask of the target object based on a differentiable 3D rendering algorithm based on an SDF value and the current pose and shape of the target object.

[0029] The differentiable 3D rendering algorithm based on SDF value can accurately obtain the current mask corresponding to the current shape and posture of the target object.

[0030] In an implementation of the first aspect, updating the current posture and shape of the target object based on a preset update method includes: updating the current posture and shape of the target object based on a preset gradient descent algorithm with a preset step size or a preset number of steps.

[0031] The gradient descent algorithm can improve the efficiency of the optimization process and shorten the time of the optimization process.

[0032] In an implementation of the first aspect, a condition is set as follows: the loss information corresponding to the target object in the current posture and shape is converged.

[0033] The loss information corresponding to the target object under the current posture and shape is convergent and can improve the accuracy of the target shape and target posture of the target object.

[0034] In an implementation of the first aspect, the method further includes: labeling the target object based on the target posture and target shape.

[0035] After the image processing device obtains the target posture and target shape, the target posture and target shape corresponding to the target object can be automatically marked in the target space.

[0036] In a second aspect, an embodiment of the present application provides an image processing device comprising: an acquisition unit and a processing unit. The acquisition unit is configured to acquire a target mask of a target object in a target image. The acquisition unit is further configured to obtain a point cloud of the target object in a target point cloud based on the target mask. The processing unit is configured to determine an initial pose of the target object based on the point cloud of the target object, and to determine an initial shape of the target object based on shape prior information. The processing unit is further configured to optimize the initial pose and initial shape of the target object based on the target mask and the point cloud of the target object to obtain a target pose and target shape of the target object.

[0037] In one possible implementation of the second aspect, the acquisition unit includes a point cloud acquisition module and a separation module. The point cloud acquisition module is configured to acquire a target point cloud. The separation module is configured to segment the target point cloud based on a target mask to obtain a point cloud of the target object.

[0038] In a possible implementation of the second aspect, the acquisition unit further includes a denoising module configured to perform denoising processing on the initial point cloud of the target object.

[0039] In a possible implementation of the second aspect, the denoising module is further configured to remove cloud points whose distance from the center point of the point cloud of the target object is greater than a first threshold, thereby obtaining a denoised point cloud of the target object.

[0040] In a possible implementation manner of the second aspect, the processing unit is further configured to perform mean processing on the point cloud of the target object to obtain an initial pose of the target object.

[0041] In a possible implementation manner of the second aspect, the processing unit is further configured to use an average shape of the shape prior information as an initial shape of the target object.

[0042] In one possible implementation of the second aspect, the processing unit is further configured to determine an initial pose and initial shape of the target object as the current pose and shape of the target object, obtain a current mask of the target object based on the current pose and shape of the target object, and obtain current loss information of the target object based on the current pose and shape of the target object, the current mask of the target object, the target mask, and a point cloud of the target object. When the current loss information of the target object meets set conditions, the processing unit determines the current pose and shape of the target object as the target pose and target shape of the target object.

[0043] In a possible implementation of the second aspect, the loss information includes at least one of mask alignment loss information, point cloud alignment loss information, and height alignment loss information. The mask alignment loss information represents the distance between a target mask and a current mask of the target object. The point cloud alignment loss information represents the distance between the current pose and shape of the target object and the point cloud of the target object. The height alignment loss information represents the distance between the height of the current pose and shape and the ground.

[0044] In a possible implementation of the second aspect, the processing unit is also used to update the current posture and shape of the target object based on a preset update method when the current loss information of the target object does not meet the set conditions, and repeat the steps of obtaining the current mask of the target object based on the current posture and shape of the target object and the steps of obtaining the current loss information of the target object based on the current posture and shape of the target object, the current mask of the target object, the target mask and the point cloud of the target object until the current loss information of the target object meets the set conditions.

[0045] In a possible implementation manner of the second aspect, the processing unit is further configured to obtain a current mask of the target object based on a differentiable 3D rendering algorithm of the SDF value and the current pose and shape of the target object.

[0046] In a possible implementation manner of the second aspect, the processing unit is further configured to update the current posture and shape of the target object based on a preset gradient descent algorithm with a preset step size or a preset number of steps.

[0047] In one embodiment of the present application, the image processing apparatus further includes: a labeling unit configured to label the target object based on the target posture and the target shape.

[0048] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein, when the computer program instructions are executed by the processor, the electronic device implements the method of the first aspect or any one of the implementation methods of the first aspect.

[0049] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a stored program, wherein when the program is executed by at least one processor, the device where the computer-readable storage medium is located is controlled to implement the first aspect or any one of the implementation methods of the first aspect.

[0050] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes executable instructions. When the executable instructions are executed by at least one processor, the computer implements the method of the first aspect or any one of the implementation methods of the first aspect.

[0051] In a sixth aspect, an embodiment of the present application provides a vehicle equipped with a driving system or a navigation system, the driving system or the navigation system including image information, and the image information is obtained by the method of the first aspect or any one of the implementation methods of the first aspect.

[0052] The image processing method provided in the embodiment of the present application determines the initial pose of the target object based on the point cloud of the target object, and determines the initial shape of the target object based on the shape prior information. After obtaining the initial pose and shape, the initial pose and initial shape are optimized based on the target mask and the point cloud of the target object to obtain the target pose and target shape of the target object. The initial shape obtained based on the shape prior information is conducive to optimization processing, and the target pose and target shape obtained by optimizing the initial shape and pose have good data generalization and excellent accuracy, and the accuracy of the target pose and target shape obtained by using the image processing method provided in the embodiment of the present invention has good stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] FIG1 is a flowchart of an image processing method provided in an embodiment of the present application.

[0054] FIG2 is a flowchart of obtaining a point cloud of a target object provided in an embodiment of the present application.

[0055] FIG3 is a flowchart of an optimization process provided in an embodiment of the present application.

[0056] FIG4 is a flowchart of an image processing method provided in an embodiment of the present application.

[0057] FIG5 is a diagram showing the effect of 3D bounding box annotation using an image processing method provided by an embodiment of the present application.

[0058] FIG6 is a diagram showing the effect of voxel occupancy grid annotation using an image processing method provided by an embodiment of the present application.

[0059] FIG7 is a comparison diagram of voxel occupancy grid annotation using an image processing method provided by an embodiment of the present application and a prior art method.

[0060] FIG8 is a schematic structural diagram of an image processing device provided in an embodiment of the present application.

[0061] FIG9 is a schematic structural diagram of an image processing device provided in an embodiment of the present application.

[0062] FIG10 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limiting the present invention.

[0064] The autonomous driving perception architecture is the eyes of autonomous driving and the fundamental guarantee for its safety. Existing 3D perception algorithms require large amounts of labeled data to train deep learning models to understand 3D scenes. However, labeling objects in 3D space is an extremely expensive and labor-intensive task.

[0065] In some technologies, annotation methods are often targeted at specific datasets, using rule-based or pre-trained models. This often works well for specific datasets or data annotation patterns. However, when real datasets or data that differs significantly from the training dataset are introduced, the quality of the 3D bounding boxes generated by the algorithm will rapidly degrade, and may even cause the model to fail.

[0066] Furthermore, existing annotation methods focus on automatically labeling 3D bounding boxes and lack support for finer-grained annotation, such as voxel occupancy grids. While some algorithms can generate the 3D shape of an object, these require training with virtual data, and the data distribution differs significantly from real data, resulting in poor shape quality.

[0067] Referring to FIG. 1 , an embodiment of the present application provides an image processing method, which is applied to an image processing device and includes:

[0068] S101: An image processing apparatus obtains a target mask corresponding to a target object in a target image.

[0069] In an embodiment of the present application, an image processing device acquires a target image through an image acquisition device, and the target image includes an image of a target object. In some implementations, the image acquisition device may be a camera, and the target image is an image captured by the camera. The target image may also include an image of the environment in which the target object is located. The target object refers to an object to be processed or labeled. The target object may be a vehicle, a group of people, a building, or a construction facility, etc. After acquiring the target image, the image processing device may segment the target image to obtain an image of the target object, thereby obtaining a target mask corresponding to the target object based on the image of the target object. Of course, the image processing device may also directly obtain the target mask of the target object in the target image from other devices or other models. In some embodiments, the image processing device may acquire the target image, transmit the target image and information about the target object to be labeled to other devices or other models, and then obtain the target mask corresponding to the target object in the target image.

[0070] When segmenting a target image, the image processing device must first identify the target object in the target image before obtaining a target mask for the target object. In some embodiments, the user can annotate the target object in the target image. This allows the image processing device to identify the target object based on the user's prompt information in the target image, segment the target object's image in the target image, and obtain a target mask corresponding to the target object based on the target object's image.

[0071] Alternatively, in some embodiments, the category of the target object that needs to be labeled can also be pre-set in the image processing device. At this time, the image processing device can identify the target object in the target image based on the category of the target object, and then segment the image of the target object, and obtain the target mask corresponding to the target object based on the image of the target object. As a possible implementation method, the image processing device can use a preset algorithm to identify the target object in the target image based on the category of the target object. For example, a 2D perception algorithm or an image recognition algorithm, or other algorithms can be used, and this application does not limit this. Among them, the category of the target object refers to the category to which the target object belongs, and the categories include: vehicles, people, animals, buildings, traffic signs, etc.

[0072] As a possible implementation method, in order to reduce the complexity of obtaining the target mask corresponding to the target object, a pre-trained target network model is used to obtain the target mask corresponding to the target object. That is, the target network model is a pre-trained model for identifying the mask information of the target object. In some target network models, the image processing device is required to first mark the target object for which the mask needs to be generated in the target image. In this case, the image processing device can first mark the target object in the acquired target image, and use the target image marked with the target object as the input of the target network model, and input it into the target network model. The target network model performs corresponding processing on the target object in the target image to obtain the target mask corresponding to the target object in the target image. In some target network models, for example, SAM (Segment Anything Model) can be used, and the target image and prompt information of the target object can be input. That is, the image processing device can use the target image as the input of the target network model and input it into the target network model. The target network model processes the received target image, and during the processing of the target network model, it is necessary to obtain the marking information of the target object. At this time, the image processing device can input prompt information of the target object into the target network model, or the user can input prompt information of the target object into the target network model. This application does not impose any restrictions on this.

[0073] As a possible implementation manner, the prompt information of the target object includes a point belonging to the target object or a bounding box of the target object.

[0074] As a possible implementation manner, the target image may be a 2D image, and the marking information of the target object includes points belonging to the 2D image of the target object or a 2D bounding box of the target object.

[0075] In some embodiments, the target network model may be a part of the image processing device, or may be another device or equipment other than the image processing device, which is not limited in this application.

[0076] S102: The image processing device obtains a point cloud of the target object in the target point cloud based on the target mask.

[0077] Typically, a target point cloud can be acquired using a LiDAR sensor. However, the target point cloud acquired by a LiDAR sensor contains not only the point cloud of the target object but also point clouds of other objects. To improve the accuracy of acquiring the target object's position and pose, the image processing device can acquire the target object's point cloud based on a target mask.

[0078] Referring to FIG. 2 , in a possible implementation of this step, the image processing device obtains a point cloud of a target object in a target point cloud based on a target mask, specifically steps including:

[0079] S201: Acquire a target point cloud, where the target point cloud includes information about a target object and the environment in which the target object is located.

[0080] In the embodiments of the present application, the target point cloud can be acquired by a laser radar or other device. The image processing device directly or indirectly acquires the target point cloud from the laser radar or other device. Indirect acquisition of the target point cloud by the image processing device refers to acquiring the target point cloud from another device, where the other device acquires the target point cloud from the laser radar or other device. The other device can be at least one of an RGB-D camera, an interferometric synthetic aperture radar, and a device capable of generating a point cloud using an image-derived method.

[0081] A target point cloud represents a target object and its surrounding environment. A target point cloud consists of several points, each of which contains 3D information, such as 3D coordinates. A target point cloud is defined as the point cloud corresponding to the target object within the target point cloud collected by the LiDAR.

[0082] S202: The image processing device segments the target point cloud based on the target mask to obtain a point cloud of the target object.

[0083] Because the target mask is 2D mask information, while the point cloud is 3D data, the target mask must first be converted into 3D mask information. The image processing device then uses this 3D mask information to segment the target object's point cloud from the target point cloud. In some possible implementations, the image processing device can convert the target mask into 3D mask information using the settings of the corresponding sensors in a point cloud acquisition device (e.g., a lidar) and a target image acquisition device (e.g., a camera).

[0084] In step 202, the image processing device uses the 2D mask information to estimate the position and shape of the target object in the 3D space. This operation is easy to implement, and the estimated position and shape of the target object in the 3D space are relatively accurate, laying a good foundation for subsequent optimization.

[0085] The point cloud of the target object segmented from the target point cloud based on the target mask includes a noise point cloud. In order to obtain a more accurate point cloud of the target object, in some implementations, after step S202, the image processing method further includes:

[0086] S203: The image processing device performs denoising on the point cloud of the target object.

[0087] Since the point cloud of the target object includes a noise point cloud, the image processing device needs to remove the noise point cloud from the point cloud of the target object to obtain a denoised point cloud of the target object.

[0088] In one possible implementation of step S203, the image processing apparatus denoising the point cloud of the target object includes removing cloud points whose distance from the center point of the point cloud of the target object is greater than a first threshold. This produces a denoised point cloud of the target object. The denoised point cloud of the target object has higher accuracy.

[0089] The first threshold may be a preset threshold value, which can be set by a technician based on actual needs. The center point of the point cloud of the target object may be a point cloud located at the median of the projected depth. Each point in the point cloud of the target object is a point in three-dimensional space, and the information of the point includes 3D coordinate information. The 3D coordinate information may be coordinate information in a preset point cloud coordinate system.

[0090] In a possible implementation, the image processing apparatus denoising the point cloud of the target object includes: denoising the point cloud of the target object based on a voxel filtering method or a normal vector-based method.

[0091] S103 : The image processing apparatus determines an initial position and posture of the target object based on the point cloud of the target object, and determines an initial shape of the target object based on the shape prior information.

[0092] The image processing device performs calculations based on the point cloud of the target object to obtain the initial posture of the target object.

[0093] In a possible implementation of step S103, obtaining the initial pose of the target object based on the point cloud of the target object includes:

[0094] The point cloud of the target object is averaged to obtain the initial pose of the target object.

[0095] The initial posture of the target object includes the initial position and initial posture of the target object. In this implementation, all points in the point cloud of the target object are averaged one by one according to the corresponding dimensions to obtain the coordinate information of the center point of the target object, and the coordinate information of the center point of the target object is taken as the initial position of the target object. The initial orientation angle of the target object can be randomly assigned any value from 0 to 2π, and the initial position and the initial orientation angle constitute the initial posture of the object. For example, if the points in the point cloud of the target object correspond to points in 3D space, then the average value of all points in the point cloud of the target object in the first dimension (for example, the X dimension) is processed to obtain The average value of all points in the point cloud of the target object in the second dimension (for example, the Y dimension) is obtained Perform the average value processing on all points in the point cloud of the target object in the third dimension (for example, the Z dimension) to obtain The initial position of the target object is Combined with the randomly initialized orientation angle, the initial pose of the target object is obtained. In this step, the point cloud of the target object can be either the pre-denoising point cloud or the post-denoising point cloud. When the denoising point cloud is used for averaging, the initial pose of the target object is more accurate.

[0096] The image processing device obtains preset shape prior information corresponding to the target object, and determines the initial shape of the target object based on the shape prior information.

[0097] Shape prior information is prior information on the target shape of an object. Shape prior information can be used to represent the prior shape of an object, and the prior shape is based on a prior fuzzy cognition of the target shape of the object based on a training data set or past experience data. For example, the shape prior information of a car can be used to represent the prior shape of a car, and the prior shape of a car can be a fuzzy car shape, but the final shape of the car needs to be optimized based on this fuzzy cognition. The image processing device can set the shape prior information according to the object category. Objects of the same category have the same shape prior information. In some embodiments, the shape prior information may include an SDF (Signed Distance Function) value. In one possible implementation, the shape prior information may include a function for representing the shape of an object, for example, the shape prior information includes a vector expression. In one possible implementation, the shape prior information may include an implicit expression vector of an SDF value. The vector dimension of the implicit expression vector of the SDF value may be a preset dimension. In some possible implementations, the image processing device can obtain a vector expression corresponding to the target object based on the category of the target object, and obtain the initial shape of the target object based on the preset initial value of the vector expression corresponding to the target object. Each vector dimension of the implicit expression vector of the SDF value represents a shape parameter. Taking the implicit expression vector of the three-dimensional SDF value as an example, one vector dimension represents length, another vector dimension represents volume, and the last vector dimension represents shape.

[0098] In a possible implementation, the shape prior information includes an SDF (Signed Distance Function) value, and the shape prior information is obtained based on a preset dimensionality reduction algorithm.

[0099] Since the SDF value is a number with many dimensions, in order to improve the efficiency of obtaining the shape of the target object, it is necessary to reduce the dimension of the SDF value based on the dimensionality reduction algorithm to obtain an implicit expression vector of the SDF value with preset dimensions.

[0100] In one possible implementation, the shape prior information may include a vector expression having a preset dimension, such as an implicit expression vector including an SDF value having a preset dimension. In one possible implementation, the dimensionality reduction algorithm includes a PCA (Principal Component Analysis) algorithm. It should be noted that the dimensionality reduction algorithm may also be other dimensionality reduction algorithms (such as an inverse feature elimination algorithm or an independent component analysis algorithm), which is not limited in this application.

[0101] In a possible implementation manner, the image processing apparatus determines the initial shape of the target object based on the shape prior information, including: taking an average shape of the shape prior information as the initial shape of the target object.

[0102] In one possible implementation, the average shape of the shape prior information refers to the shape represented by the average of multiple pieces of shape prior information. The multiple pieces of shape prior information are shape prior information corresponding to multiple samples. For example, if there are n samples, each of which corresponds to a piece of shape prior information, then n samples correspond to n pieces of shape prior information. The average of the n pieces of shape prior information is taken as the average shape prior information.

[0103] In a possible implementation, the image processing apparatus determines the initial shape of the target object based on the shape prior information, including: the image processing apparatus obtains an initial value of the shape prior information, and obtains the initial shape of the target object based on the initial value of the shape prior information.

[0104] In order to reduce the complexity of the image processing method, an initial value can be set for the shape prior information corresponding to each category of objects, and the initial shape of the object can be obtained based on the initial value. In some embodiments, the initial value of the shape prior information can be a set value, such as a default value after the shape prior information is initialized. The initial value of the shape prior information obtained by the image processing device can be the initial value of the implicit expression vector for obtaining the SDF value. After obtaining the initial value of the shape prior information, the initial shape of the target object represented by the initial value of the shape prior information can be obtained. In some possible implementations, there is a preset correspondence between the shape prior information of the object and the shape of the object, and the current shape of the target object can be obtained based on this preset correspondence and the current shape prior information of the object. In one possible implementation, the preset correspondence can be pre-set by a technician before the image device leaves the factory.

[0105] The image processing device can obtain the prior shape information corresponding to the target object according to the category of the target object. The image processing device can also obtain the prior shape information corresponding to the target object from other devices or other models according to the category of the target object.

[0106] In some possible implementations, when segmenting a target image, the image processing device needs to first identify the category of the target object in the target image, and then derive prior shape information corresponding to the target object based on the category of the target object.

[0107] In some possible implementations, the category of the target object to be labeled can be preset in the image processing device. The image processing device derives the category of the target object from the preset information and, based on the category of the target object, derives the prior information of the shape corresponding to the target object.

[0108] In some possible implementations, the image processing device inputs the category of the target object to other devices or other models, and obtains prior information of the target object and its corresponding shape from other devices or feedback from other devices.

[0109] In one possible implementation, to reduce the complexity of obtaining prior shape information corresponding to a target object, a pre-trained network model is used to obtain the prior shape information corresponding to the target object. In some embodiments, the target network model may be a shape model. The image processing device obtains an initial value for the prior shape information from the shape model based on the category of the target object. The initial value for the prior shape information may be a default value preset in the shape model. Based on the correspondence between the preset prior shape information and the shape, the shape model obtains a shape corresponding to the initial prior shape information; this shape is the initial shape of the target object.

[0110] In one possible implementation, the shape model acquisition process is as follows: a 3D model acquisition device acquires several types of 3D models, and each type of 3D model includes several 3D model samples. A computing and processing device is used to convert each of the acquired 3D model samples into an SDF expression. The computing and processing device performs dimensionality reduction processing on the expression of the SDF values ​​of several 3D model samples of each type based on the PCA algorithm, that is, the PCA algorithm is used to reduce the dimension of each 3D model sample expressed in SDF value to the target dimension. In certain embodiments of the present application, the target dimension may be 5 dimensions. It should be noted that the target dimension of 5 dimensions is only an exemplary description, not a specific limitation, that is, the target dimension may be other values ​​(for example, 3 dimensions).

[0111] The initial value of the shape prior information preset in the shape model is the average value of several 3D model samples expressed based on the SDF value of the corresponding category learned by the PCA algorithm. At the same time, in order to facilitate the driver's understanding, the shape model needs to present the shape prior information in the shape of the object. In certain embodiments of the present application, the shape model may pre-store the initial value of the shape prior information of each category, the correspondence between the shape prior information and the shape, and the shape of the object. After obtaining the shape prior information, the shape of the object is obtained through the correspondence between the pre-stored shape prior information and the shape. In other embodiments of the present application, the shape model may pre-store the initial value of the shape prior information and the correspondence between the shape prior information and the shape. After obtaining the shape prior information, the corresponding shape is automatically calculated or drawn through the correspondence between the pre-stored shape prior information and the shape. It should be noted that the data pre-stored in the above-mentioned shape model is only an exemplary description and is not a specific limitation on the relevant parameters in the shape model. At the same time, the above-mentioned shape model can be trained before leaving the factory and can be used directly after leaving the factory; the shape model can also be further optimized and trained during use, that is, the initial value of the shape prior corresponding to each category in the shape model can be further optimized and updated regularly during use.

[0112] S104 , the image processing device optimizes the initial posture and initial shape of the target object based on the target mask and the point cloud of the target object to obtain the target posture and target shape of the target object.

[0113] To improve the accuracy of the target object's shape and posture, optimization processing is required after obtaining the target object's initial shape to obtain the target posture and shape of the target object. In some possible implementations, the image processing device uses the target object's initial posture and initial shape as a starting point, updates the target object's shape and posture using an iterative algorithm, and performs optimization verification using a target optimization function. In one possible implementation, the target optimization function is used to represent the relationship between the target object's current shape and posture and its ideal shape and posture.

[0114] Referring to FIG. 3 , in a possible implementation of step S104 , the image processing apparatus optimizes the initial pose and initial shape of the target object based on the target mask and the point cloud of the target object, and obtains the target pose and target shape of the target object, including:

[0115] S301: The image processing apparatus determines the initial position and shape of the target object as the current position and shape of the target object.

[0116] S302: The image processing device obtains a current mask of the target object based on the current posture and shape of the target object.

[0117] To reduce the complexity of the optimization process, the image processing device performs dimensionality reduction based on the current shape and pose of the target object, converting the 3D pose and shape information into the 2D mask information corresponding to the current target object. The current mask information and the target mask information are then fed into the optimization function for optimization verification. In some possible implementations, the dimensionality reduction process can be a projection process, which can be performed using a rendering algorithm.

[0118] In some possible implementations, the optimization function may be:

[0119] Among them, p * ,s * where p and s represent the target pose and shape of the target object, respectively. p and s represent the current pose and shape of the target object, respectively. Y represents the current mask of the target object under the current pose and shape. E is the loss information. The loss information E consists of three parts: mask alignment loss, point cloud alignment loss, and ground alignment loss.

[0120] In a possible implementation of step S302, obtaining the current mask of the target object based on the current pose and shape of the target object includes: obtaining the current mask of the target object based on a differentiable 3D rendering algorithm based on an SDF value and the current pose and shape of the target object.

[0121] Among them, the differentiable 3D rendering algorithm is:

[0122] Among them, P j Represents the pixel point of the image plane, π(p j )∈(0,1) represents its rendering value, the set Indicates the time between the camera optical center and pixel p j The SDF value of the sampling point on the ray is φ(x), and ζ represents the manually specified parameters of the rendering algorithm. If a ray passes through the surface of an object, it means that at least one of the sampling points on this ray has a positive SDF value, resulting in a rendering value close to 1 for that pixel. Conversely, this means that the SDF values ​​of the sampling points on this ray are all negative, resulting in a rendering value close to 0 for that pixel.

[0123] Substitute the shape prior information (such as the SDF value) corresponding to the shape of the target object into the above formula to obtain the current mask corresponding to the shape of the current object.

[0124] S303: The image processing device obtains current loss information of the target object based on the current posture and shape of the target object, the current mask of the target object, the target mask and the point cloud of the target object.

[0125] In a possible implementation, the loss information includes at least one of mask alignment loss information, point cloud alignment loss information, and height alignment loss information. In some preferred embodiments, the loss information includes at least mask alignment loss information.

[0126] The mask alignment loss is used to measure the distance between the target mask segmented by SAM and the current mask rendered by the current pose and shape. The mask alignment loss is defined as follows:

[0127] Among them, Y proj Represents the current mask obtained by the rendering algorithm, Y represents the target mask obtained by SAM, and O represents the occlusion mask of the target object, which can be calculated by the target mask obtained by SAM.

[0128] Point cloud alignment loss information is used to indicate the degree of alignment between the object represented by the current pose and shape of the target object and the point cloud of the target object. Point cloud alignment loss is defined as follows:

[0129] in, represents the points belonging to the target obtained by the target mask in the scene LiDAR point cloud (i.e., target point cloud), i.e. A point in the point cloud representing the target object; Indicates the points where the light passes through these points and intersects the actual surface of the target object; y represents the intersection of the light and the surface of the target object in the current posture and shape; φ represents the SDF value function.

[0130] Optimizing the mask alignment loss alone can easily cause the object's shape and pose to fall into a local minimum, ultimately reducing the practicality of the method. Therefore, this method further proposes a point cloud alignment loss to improve the practicality of the method.

[0131] The height alignment loss information is used to represent the height of the current pose and shape relative to the ground. The definition of the height alignment loss information is as follows:

[0132] in, Indicates the current 3D position of the target, The current height of the object can be obtained from the shape of the object. Ground height function It is obtained by combining the RANSAC (Random Sample Consensus) algorithm with the target point cloud.

[0133] In order to ensure that the target object should be located on the ground, a height alignment loss is proposed in the embodiment of the present application.

[0134] It should be understood that in this application, distance is used to represent the difference value or degree of difference, rather than the physical concept of length. For example, mask alignment loss information is used to measure the difference value or degree of difference between the target mask segmented by SAM and the current mask rendered by the current pose and shape; height alignment loss information is used to represent the difference value or degree of difference between the height of the current pose and shape and the ground.

[0135] In a possible implementation of this application, the loss information is defined as follows:

[0136] E=aE mask +bE pc +cE ground

[0137] Among them, E is the loss information, a is the weight of the mask alignment loss information, b is the weight of the point cloud alignment loss information, and c is the weight of the height alignment loss information.

[0138] S304: When the current loss information of the target object meets the set conditions, the image processing device determines the current posture and shape of the target object as the target posture and target shape of the target object.

[0139] To obtain the target pose and shape, the image processing device must set an optimization target. This optimization target can be implemented through set conditions. In some embodiments, the optimization target can be implemented through a preset optimization function. Based on the preset conditions, the image processing device can determine whether the current pose and shape of the target object require further optimization. If so, the optimization steps are repeated. If no further optimization is required, that is, if the set conditions are met, the current pose and shape of the target object are determined as the target pose and shape of the target object.

[0140] In one possible implementation of the present application, the setting condition is that the loss information corresponding to the target object in the current position and shape is convergent. In other words, the value of the loss information corresponding to the target object in the current position and shape is minimized. The setting condition can be reflected by an optimization function of the target object. The optimization function can be:

[0141] Among them, p * ,s * where p and s represent the target pose and shape of the target object, respectively. p and s represent the current pose and shape of the target object, respectively. Y represents the current mask of the target object under the current pose and shape. E is the loss information. The loss information E consists of three parts: mask alignment loss, point cloud alignment loss, and ground alignment loss.

[0142] In one possible implementation of the present application, the conditions are as follows: the loss information of the target object in its current position and shape is the first loss information, the loss information of the target object in its previous position and shape is the second loss information, and the difference between the first and second loss information is less than a set threshold. In a preferred implementation, the first loss information is less than the second loss information, and the value of the second loss information minus the first loss information is less than the set threshold.

[0143] In a possible implementation of step S104, step S104 further includes:

[0144] When the current loss information of the target object does not meet the set conditions, the current posture and shape of the target object are updated based on a preset update method, and the steps of obtaining the current mask of the target object based on the current posture and shape of the target object and the step of obtaining the current loss information of the target object based on the current posture and shape of the target object, the current mask of the target object, the target mask and the point cloud of the target object are repeated until the current loss information of the target object meets the set conditions.

[0145] In this implementation, updating the current position and shape of the target object based on a preset updating method includes: updating the current position and shape of the target object based on a preset gradient descent algorithm with a preset step size or a preset number of steps.

[0146] In this implementation, a preset gradient descent algorithm is used to update the SDF implicit representation vector corresponding to the target object's position, pose, and shape with a set step size. The gradient descent algorithm has the advantages of wide applicability and strong scalability. It can quickly and effectively determine the target pose and shape of the target object, improving computational efficiency.

[0147] In one possible implementation, the gradient descent algorithm can be implemented by an adaptive moment estimation optimizer (Adam), and the gradient descent algorithm can be a batch gradient Xiangjiang algorithm or a stochastic gradient descent algorithm.

[0148] In one possible implementation of the present application, the current loss information of the target object decreases as the current posture and shape of the target object are updated until the loss information reaches a minimum value, and the current posture and shape of the target object are the target posture and target shape.

[0149] Referring to FIG4 , an embodiment of the present application provides an image processing method, including:

[0150] S401: The image processing apparatus obtains a target mask corresponding to a target object in a target image.

[0151] The specific implementation process of step S401 can refer to step S101 and will not be described in detail here.

[0152] S402: The image processing device obtains a point cloud of the target object in the target point cloud based on the target mask.

[0153] The specific implementation process of step S402 can refer to step S102 and will not be described in detail here.

[0154] S403: The image processing apparatus determines an initial pose of the target object based on the point cloud of the target object, and determines an initial shape of the target object based on the shape prior information.

[0155] The specific implementation process of step S403 can refer to step S103 and will not be described in detail here.

[0156] S404 : The image processing apparatus optimizes the initial posture and initial shape of the target object based on the target mask and the point cloud of the target object to obtain a target posture and target shape of the target object.

[0157] The specific implementation process of step S404 can refer to step S104 and will not be described in detail here.

[0158] S405: The image processing device labels the target object based on the target posture and target shape.

[0159] After the image processing device obtains the target posture and target shape, it can automatically mark the target posture and target shape corresponding to the target object in the target space. In some embodiments, the image processing device divides the target space into a plurality of grids, and the target posture of the target object can mark the position and posture of the target object from the corresponding grid in the target space, and then use the target shape to mark the shape corresponding to the target object at the grid position and its surrounding grids, so that the target position can be marked in the target space. In a possible implementation, the image processing device can mark the point corresponding to the position and posture of the target object in the target space (the point is the point corresponding to the center point of the target object in the target space) by the target posture of the target object, and then mark the points constituting the target shape around the point corresponding to the center point of the target object in the target space according to the target shape of the target object. In a possible implementation, the image processing device can also use a bounding box marking algorithm and mark the bounding box corresponding to the target object in the target space based on the target shape and target posture of the target object.

[0160] In one possible implementation of this embodiment, the image processing apparatus labels the target object in 3D space based on the target pose and shape. For example, the image processing apparatus labels the target object's 3D bounding box and category in 3D space based on the target pose and shape. The image processing apparatus may also label the target object in 3D space using a voxel-occupied grid based on the target pose and shape.

[0161] The image processing method provided in the embodiment of the present application obtains the initial posture of the target object based on the point cloud of the target object, and determines the initial shape of the target object based on the shape prior information. After obtaining the initial posture and shape, the initial posture and initial shape are optimized based on the target mask and the point cloud of the target object to obtain the target posture and target shape of the target object. The initial shape obtained based on the shape prior information is conducive to optimization processing, and the target posture and target shape obtained by optimizing the initial shape and posture have good data generalization and excellent accuracy, and the accuracy of the target posture and target shape obtained by the image processing method provided in the embodiment of the present invention has good stability.

[0162] The proposed method utilizes object segmentation results and a small number of vehicle CAD (Computer Aided Design) models as prior information to generate accurate 3D pose and shape through iterative optimization. Therefore, the technical solution provided by the embodiments of this invention can provide automatic or semi-automatic annotation solutions for various autonomous driving perception-related data annotation needs without requiring ground truth information in the target data domain.

[0163] In order to better understand the effect of the image processing method provided in the embodiments of the present application, the present application is verified by the public KITTI dataset.

[0164] The KITTI dataset, jointly developed by the Karlsruhe Institute of Technology and Toyota Research Institute America, is currently the world's largest evaluation dataset for computer vision algorithms in autonomous driving scenarios. It is used to evaluate the performance of computer vision technologies such as stereo, optical flow, visual odometry, 3D object detection, and 3D tracking in an in-vehicle environment. KITTI contains real-world image data collected from urban, rural, and highway scenes, with up to 15 vehicles and 30 pedestrians per image, and various degrees of occlusion and truncation. The dataset consists of 389 pairs of stereo and optical flow images, 39.2 km of visual odometry sequences, and over 200,000 images with 3D annotated objects, all sampled and synchronized at 10 Hz.

[0165] Please refer to FIG5 . In the verification process, for the automatic annotation of the 3D bounding box, the present invention uses the average precision (AP) of the bird's eye view (BEV) BEV ) and the average precision (AP) of the 3D bounding box 3D ) as the metric. We use the NuScenes Detection Score (NDS) and Mean Average Precision (mAP) for generalization comparisons. Note that for all of these metrics, larger values ​​are better.

[0166] Please refer to Figure 6. The generalization and fine-grained annotation of the present invention are verified by the public NuScenes dataset, which is a large-scale multimodal dataset for 3D detection and map segmentation. The dataset is divided into 700 / 150 / 150 scenes for training / validation / testing. It contains data from multiple sensors, including six cameras, one lidar and five radars. For camera input, each frame contains six views of the surrounding environment at a specific timestamp. The present invention adjusts the input view to 256x704 resolution and voxelizes the point cloud at 0.075m and 0.1m respectively for detection and segmentation.

[0167] By using the above-mentioned public dataset for verification, we can draw the following conclusions:

[0168] Beneficial effect 1: The image processing method provided by the embodiment of the present application achieves optimal performance in the quality of generating pseudo labels.

[0169] When directly using pseudo labels to compare with true data, it can be seen from the upper half of Table 1 that the model tested on the KITTI dataset achieves the best results in the automatic labeling task of 3D object detection.

[0170] VS3D (weakly supervised 3D object detection) in Table 1 is a weakly supervised 3D object detection model based on point clouds. VS3D proposes an unsupervised 3D object proposal module (UPM), which uses the geometric and density properties of laser point clouds to find high-confidence areas that may contain target objects. In addition, in order to obtain more accurate 3D bounding boxes, VS3D also designs a cross-modal transfer learning method: the point cloud-based detection network is regarded as a student and learns knowledge from an existing pre-trained off-the-shelf teacher image detection network. The 3D object proposals generated by UPM are projected onto paired images and classified by the teacher network, and then the student network imitates the teacher's behavior during training. However, the limitations of the VS3D framework are also obvious. First, the VS3D framework requires a pre-trained 2D object detection model. Therefore, when the unlabeled data is inconsistent with the pre-trained model, the performance of the object detection network will decline due to domain differences, ultimately leading to a decline in the performance of the 3D detection grid. Second, VS3D distills information from the teacher network of 2D object detection and lacks 3D supervision signals, resulting in poor 3D detection performance and very limited real-world applications.

[0171] The SDFLabel (signed distance fields labeling algorithm) in Table 1 consists of a deep SDF network pre-trained on a virtual dataset and the RANSAC algorithm. This algorithm proposes a novel SDF-based differentiable rendering function to train the deep SDF network to encode the objects to be labeled. The deep SDF network inputs RGB (red, green, and blue) image patches of the objects, and outputs a vector of implicit representations of the SDF values ​​of the objects in the image. During labeling, the algorithm first obtains the 2D bounding box of the target object and the points belonging to the object in the point cloud. The 2D bounding box is input into the deep SDF network to obtain the SDF representation of the object within the bounding box. Finally, the RANSAC algorithm is used to combine the output of the deep SDF network to estimate the 3D pose of the target object. Because this technical solution relies heavily on the quality of the pre-trained deep SDF network, which is often significantly different from real-world data, the shape prediction quality of the target object in the real dataset is poor, which directly affects the quality of the final pseudo-labeling. In addition, SDFLabel is not an end-to-end algorithm. The result of deep SDF prediction is fixed in the subsequent RANSAC optimization, which further reduces the accuracy of the 3D bounding box.

[0172] FGR (Frustum Aware Geometric Reasoning), listed in Table 1, automatically annotates 3D bounding boxes using 2D bounding box cues and sparse point clouds. It consists of two components: coarse 3D point cloud segmentation and 3D bounding box estimation. In the coarse 3D segmentation stage, FGR proposes a context-aware adaptive region growing method that leverages contextual information and adaptively adjusts thresholds to roughly segment the target point cloud. In the 3D bounding box estimation stage, a noise-resistant framework is designed to estimate bounding rectangles in a bird's-eye view image and locate key vertices. A rule-based iterative method is then used to estimate the 3D bounding box. The object bounding boxes used by FGR are not conventional object detection boxes, but rather complete object boxes with occlusion information. These object bounding boxes typically contain certain object size information. Current 2D object bounding boxes and detection algorithms cannot serve as input for FGR, limiting its practical application. When predicting the 3D bounding box's rotation, FGR cannot address the issue of bounding box flipping, and therefore requires ground-truth information about the object's rotation. In addition, FGR uses a rule-based algorithm when estimating the 3D pose of an object, which results in FGR having poor generalization performance.

[0173] To overcome the problem of excessive noise in point clouds, the LAD (Lifting 2D object locations to 3D and Discounting Lidar outliers) method in Table 1 uses a self-supervised approach based on template matching. Specifically, the model first obtains mask information of the object to be annotated through a Mask Region-based Convolutional Neural Network (Mask R-CNN). The mask is then used to obtain a noisy point cloud belonging to the object in the point cloud. The LAD framework consists of two parts: a point cloud segmentation network and a 3D pose regression network. LAD uses a fixed 3D model as a template, providing a supervisory signal to train the point cloud segmentation and 3D pose regression networks. In addition, LAD can also use the temporal information of the video to refine the pseudo-labels. This solution uses a self-supervised training method based on template matching. Selecting an appropriate template 3D model is crucial, as it directly determines the quality of the supervisory signal and ultimately affects the quality of the 3D bounding box. In addition, the 3D model used as a template is fixed during training and cannot match various situations well, which also makes LAD have poor generalization performance.

[0174] For the sake of comparison, the image processing method provided in the embodiment of the present application is named SLF (Segment, Lift and Fit).

[0175] It can be seen from Table 1 that in examples of medium difficulty and difficult difficulty, the image processing method (SLF) provided by the embodiment of the present application is significantly ahead of the existing automatic labeling algorithm.

[0176] The lower part also shows the performance of SLF in automatic and semi-automatic annotation processes. GTmask It refers to the SLF automatic labeling algorithm that takes the true value mask as input. point Indicates the semi-automatic labeling process, that is, the labeler points out several points belonging to the target object, SLF box Indicates the fully automatic annotation process, where the object's 2D bounding box is obtained by the detection algorithm. "-" in the table indicates that there is no corresponding measurement result for that metric. It can be seen that SLF significantly outperforms existing algorithms in both automatic and semi-automatic annotation processes.

[0177] Table 1

[0178] Beneficial Effect 2: Through the image processing method provided in the embodiment of the present application, the 3D detector trained using pseudo labels has excellent performance. Using the image processing method provided in the embodiment of the present application, multiple different 3D detectors were trained using pseudo labels and evaluated. As shown in Table 2, PointPillar is a point cloud target detection algorithm based on a cylinder; SECOND (Sparsely Embedded CONvolutional Detection) is a 3D detection algorithm based on sparse convolution; Part-A 2 Part-Aware and Aggregation (PAA) is a 3D detection algorithm based on part awareness and aggregation; VoxelR-CNN is a voxel-grid-based regional convolutional neural network; and GT (Ground Truth) is the true value. Table 2 shows that all five 3D detection algorithms trained on the KITTI dataset using the image processing method provided in this application to generate 3D pseudo-labels achieve significantly better evaluation metrics than the unsupervised automatic labeling algorithm FGR, and even rival the performance of detectors trained using KITTI ground truth.

[0179] Table 2

[0180] Beneficial effect 3: By using the image processing method provided in the embodiment of the present application to perform cross-dataset annotation, excellent generalization performance can be achieved.

[0181] In Table 3, the present invention compares the performance of the currently best unsupervised labeling algorithm FGR and the supervised labeling algorithm MTrans (multimodal transformer, multimodal self-attention neural network) on cross-datasets. As can be seen from the table, since FGR is a rule-based labeling algorithm, it works well on KITTI, but it cannot obtain effective measurement results in data outside the data domain (nuScenes). Mtrans, because it is trained on the KITTI dataset, can obtain excellent results in KITTI data, but as mentioned above, supervised labeling algorithms often overfit the labeling pattern of a specific dataset and perform poorly in unseen data. SLF is an unsupervised labeling algorithm that does not require any 3D labeling in advance. As can be seen from the mAP indicator, SLF's cross-dataset generalization performance is significantly better than Mtrans.

[0182] Table 3

[0183] Technical effect 4: Perform fine-grained 3D automatic labeling with excellent results.

[0184] Please refer to Figure 7. The present invention conducts a fine-grained annotation experiment in the nuScenes dataset, specifically annotating the voxel occupancy grid of the target object. Since the nuScenes dataset does not provide the true value information of the voxel occupancy grid, the present invention provides a qualitative comparison with the existing annotation method in Figure 10. Among them, in Figure 7, the second and third columns are the true value information of the voxel occupancy grid provided by OFN (occupancy dataset for nuscenes, an occupancy dataset based on nuscenes) and OpenOccupancy, respectively, and the fourth column is the true value information annotated by SLF. It can be clearly seen that compared with the existing annotation method, the voxel occupancy truth annotated by SLF can provide more detailed and complete shape information.

[0185] Referring to FIG8 , an embodiment of the present application further provides an image processing device 100, comprising: an acquisition unit 101 and a processing unit 102. The acquisition unit 101 is configured to acquire a target mask of a target object in a target image. The acquisition unit 101 is further configured to obtain a point cloud of the target object in a target point cloud based on the target mask. The processing unit 102 is configured to determine an initial pose of the target object based on the point cloud of the target object, and to determine an initial shape of the target object based on shape prior information. The processing unit 102 is further configured to optimize the initial pose and initial shape of the target object based on the target mask and the point cloud of the target object to obtain a target pose and target shape of the target object.

[0186] In an embodiment of the present application, the acquisition unit 101 includes: a target mask acquisition module. The target mask acquisition module is used to acquire the target mask of the target object.

[0187] In one embodiment of the present application, the acquisition unit 101 includes: a point cloud acquisition module and a separation module. The point cloud acquisition module is used to acquire a target point cloud. The separation module is used to segment the target point cloud based on a target mask to obtain a point cloud of the target object.

[0188] In one embodiment of the present application, the acquisition unit 101 further includes a denoising module. The denoising module is configured to denoise the initial point cloud of the target object, thereby obtaining a denoised point cloud of the target object.

[0189] The target point cloud includes information about the target object and the environment in which the target object is located.

[0190] In one embodiment of the present application, the denoising module is further configured to remove cloud points whose distance from the center point of the point cloud of the target object is greater than a first threshold, thereby obtaining a denoised point cloud of the target object.

[0191] In one embodiment of the present application, the processing unit 102 is further configured to perform mean processing on the point cloud of the target object to obtain an initial pose of the target object.

[0192] In an embodiment of the present application, the processing unit 102 is further configured to obtain an initial value of the shape prior information, and obtain an initial shape of the target based on the initial value of the shape prior information.

[0193] In an embodiment of the present application, the processing unit 102 is further configured to use the average shape of the shape prior information as the initial shape of the target object.

[0194] In one embodiment of the present application, the processing unit 102 is further configured to obtain shape prior information corresponding to the target object based on a preset dimensionality reduction algorithm.

[0195] The shape prior information includes an SDF value, and the dimensionality reduction algorithm includes a PCA algorithm.

[0196] In one embodiment of the present application, the processing unit 102 is further configured to determine the initial pose and initial shape of the target object as the current pose and shape of the target object, obtain a current mask of the target object based on the current pose and shape of the target object, and obtain current loss information of the target object based on the current pose and shape of the target object, the current mask of the target object, the target mask, and the point cloud of the target object. When the current loss information of the target object meets the set conditions, the processing unit 102 determines the current pose and shape of the target object as the target pose and target shape of the target object.

[0197] In one implementation, the loss information includes at least one of mask alignment loss, point cloud alignment loss, and height alignment loss. Mask alignment loss represents the distance between the target mask and the current mask of the target object. Point cloud alignment loss represents the distance between the current pose and shape of the target object and the point cloud of the target object. Height alignment loss represents the distance between the height of the current pose and shape and the ground.

[0198] In one embodiment of the present application, the processing unit 102 is also used to update the current posture and shape of the target object based on a preset update method when the current loss information of the target object does not meet the set conditions, and repeat the steps of obtaining the current mask of the target object based on the current posture and shape of the target object and the steps of obtaining the current loss information of the target object based on the current posture and shape of the target object, the current mask of the target object, the target mask and the point cloud of the target object until the current loss information of the target object meets the set conditions.

[0199] In one embodiment of the present application, the processing unit 102 is further configured to obtain a current mask of the target object based on a differentiable 3D rendering algorithm of the SDF value and the current pose and shape of the target object.

[0200] In one embodiment of the present application, the processing unit 102 is further configured to update the current position and shape of the target object based on a preset gradient descent algorithm with a preset step size or a preset number of steps.

[0201] In a possible implementation of this embodiment, a condition is set as follows: the loss information corresponding to the target object in the current posture and shape is converged.

[0202] In one possible implementation of this embodiment, in one possible implementation of the present application, a condition is set as follows: the loss information of the target object in the current posture and shape is the first loss information, the loss information of the target object in the previous posture and shape is the second loss information, and the difference between the first loss information and the second loss information is less than a set threshold. In a preferred implementation, the first loss information is less than the second loss information, and the value of the second loss information minus the first loss information is less than the set threshold.

[0203] 9 , in one embodiment of the present application, the image processing apparatus 100 further includes a labeling unit 103. The labeling unit 103 is configured to label the target object based on the target posture and target shape.

[0204] An embodiment of the present application also provides an electronic device, comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the electronic device implements the method provided in any of the aforementioned embodiments.

[0205] In one embodiment of the present application, the electronic device is a chip. In one possible implementation, the chip is used for cloud processing to provide driving basis for the vehicle's driving system or navigation.

[0206] Referring to Figure 10 , in one embodiment of the present application, an electronic device 200 may include: a processor 201, a memory 202, and a communication unit 203. These components communicate via one or more buses. Those skilled in the art will appreciate that the structure of the electronic device shown in the figure does not limit the embodiments of the present invention. The electronic device may have a bus structure or a star structure, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0207] The communication unit 203 is configured to establish a communication channel so that the electronic device can communicate with other devices, receive user data sent by other devices, or send user data to other devices.

[0208] The processor 201 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. It runs or executes software programs, instructions, and / or modules stored in the memory 202, and calls data stored in the memory to perform various functions of the electronic device and / or process data. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor 201 can only include a central processing unit (CPU). In an embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.

[0209] The memory 202 is used to store execution instructions of the processor 201. The memory 202 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0210] When the execution instructions in the memory 202 are executed by the processor 201 , the electronic device 200 is enabled to execute part or all of the steps in the embodiment shown in FIG. 1 .

[0211] An embodiment of the present application provides a computer-readable storage medium, which includes a stored program. When the program is executed by at least one processor, the device where the computer-readable storage medium is located is controlled to implement the method provided in any of the aforementioned embodiments.

[0212] The present application also provides a vehicle equipped with a driving system or a navigation system, wherein the driving system or the navigation system includes image information obtained by the image processing method provided in any of the aforementioned embodiments. Alternatively, the vehicle includes the aforementioned image processing device or the aforementioned electronic device.

[0213] An embodiment of the present application provides a computer program product, which includes executable instructions. When the executable instructions are executed on a computer, the computer executes the method of the aforementioned embodiment.

[0214] The technical solution provided by the embodiment of the present invention can be applied to the following scenarios:

[0215] Application scenario 1: Automatic labeling of 3D object detection.

[0216] 3D object detection is an extremely important part of the visual perception system of autonomous vehicles. It can provide traffic scene understanding by detecting common traffic participants such as pedestrians, vehicles and other targets. Existing automatic labeling technologies either require a small amount of 3D true values ​​as starting data, or predict 3D pose based on self-supervised algorithms. The former has poor cross-dataset generalization capabilities, and the latter has poor labeling quality. These shortcomings cannot meet the automatic labeling needs of 3D data. The present invention uses object shape priors and 2D hints, and uses gradient descent to directly optimize the 3D pose of the object, achieving better cross-data generalization and 3D bounding box quality than previous algorithms.

[0217] Application scenario 2: Automatic labeling of voxel occupation grids.

[0218] The image processing technology of the present invention can generate accurate shape information and 3D pose of an object, and can be used for 3D voxel occupancy grid prediction tasks.

[0219] The above describes in detail the structure, features and effects of the present invention based on the embodiments shown in the drawings. The above is only a preferred embodiment of the present invention, but the scope of implementation of the present invention is not limited to what is shown in the drawings. Any changes made in accordance with the concept of the present invention, or modifications to equivalent embodiments with equivalent changes, which do not exceed the spirit covered by the description and drawings, should be within the scope of protection of the present invention.

Claims

1. An image processing method, characterized in that: include: Obtain a target mask corresponding to the target object in the target image; Based on the target mask, obtaining a point cloud of the target object in the target point cloud; Determining an initial pose of the target object based on a point cloud of the target object, and determining an initial shape of the target object based on shape prior information; The initial position and shape of the target object are optimized based on the target mask and the point cloud of the target object to obtain the target position and shape of the target object.

2. The image processing method according to claim 1, characterized in that: The step of obtaining a point cloud of the target object in the target point cloud based on the target mask comprises: Acquire the target point cloud, wherein the target point cloud includes information about the target object and the environment in which the target object is located; The target point cloud is segmented based on the target mask to obtain a point cloud of the target object.

3. The image processing method according to claim 2, characterized in that: The method also includes denoising the point cloud of the target object.

4. The image processing method according to claim 3, characterized in that: The denoising of the point cloud of the target object comprises: The point clouds whose distance from the center point of the point cloud of the target object is greater than a first threshold are removed.

5. The image processing method according to any one of claims 1 to 4, characterized in that: The step of acquiring the initial pose of the target object based on the point cloud of the target object comprises: The point cloud of the target object is averaged to obtain an initial pose of the target object.

6. The image processing method according to any one of claims 1 to 4, characterized in that: The shape prior information includes an SDF value, and the shape prior information is obtained based on a preset dimensionality reduction algorithm.

7. The image processing method according to claim 6, characterized in that: The dimension reduction algorithm includes a PCA algorithm.

8. The image processing method according to any one of claims 1 to 4, characterized in that: Determining the initial shape of the target object based on shape prior information includes: taking an average shape of the shape prior information as the initial shape of the target object.

9. The image processing method according to any one of claims 1 to 4, characterized in that: The optimizing process of the initial posture and initial shape of the target object based on the target mask and the point cloud of the target object to obtain the target posture and target shape of the target object includes: Determining the initial position and shape of the target object as the current position and shape of the target object; Based on the current position and shape of the target object, obtain a current mask of the target object; Based on the current pose and shape of the target object, the current mask of the target object, the target mask and the point cloud of the target object, obtain current loss information of the target object; When the current loss information of the target object meets the set conditions, the current position and shape of the target object are determined as the target position and shape of the target object.

10. The image processing method according to claim 9, characterized in that: The loss information includes: at least one of mask alignment loss information, point cloud alignment loss information and height alignment loss information; wherein the mask alignment loss information is used for the distance between the target mask of the target object and the current mask; the point cloud alignment loss information is used to indicate the degree of alignment between the object represented by the current posture and shape of the target object and the point cloud of the target object; the height alignment loss information is used to indicate the distance between the height of the current posture and shape and the ground.

11. The image processing method according to claim 9, characterized in that: Also includes: When the current loss information of the target object does not meet the set conditions, the current position and shape of the target object are updated based on a preset update method, and the steps of obtaining the current mask of the target object based on the current position and shape of the target object and obtaining the current loss information of the target object based on the current position and shape of the target object, the current mask of the target object, the target mask and the point cloud of the target object are repeated until the current loss information of the target object meets the set conditions.

12. The image processing method according to any one of claims 9 to 11, characterized in that: Based on the current position and shape of the target object, obtaining the current mask of the target object includes: obtaining the current mask of the target object based on a differentiable 3D rendering algorithm of an SDF value and the current position and shape of the target object.

13. The image processing method according to claim 11, characterized in that: The updating of the current position and shape of the target object based on a preset updating method includes: The current position and shape of the target object are updated based on a preset gradient descent algorithm with a preset step size or a preset number of steps.

14. The image processing method according to claim 9, characterized in that: The setting condition is that the loss information corresponding to the target object in the current posture and shape is convergent.

15. The image processing method according to claim 1, characterized in that: Also includes: The target object is labeled based on the target posture and target shape.

16. An image processing device, characterized in that: include: An acquisition unit, used for acquiring a target mask of a target object in a target image; The acquisition unit is further used to obtain a point cloud of the target object in the target point cloud based on the target mask; A processing unit, configured to determine an initial position and posture of the target object based on a point cloud of the target object, and determine an initial shape of the target object based on shape prior information; The processing unit is further used to optimize the initial posture and initial shape of the target object based on the target mask and the point cloud of the target object to obtain the target posture and target shape of the target object.

17. The image processing device according to claim 16, characterized in that: The acquisition unit includes a point cloud acquisition module and a separation module; The point cloud acquisition module is used to acquire the target point cloud, and the target point cloud includes information about the target object and the environment in which the target object is located; The separation module is used to segment the target point cloud based on the target mask to obtain the point cloud of the target object.

18. The image processing device according to claim 17, characterized in that: The acquisition unit further includes a denoising module; the denoising module is used to denoise the point cloud of the target object.

19. The image processing device according to claim 18, characterized in that: The denoising module is used to remove point clouds whose distance from the center point of the point cloud of the target object is greater than a first threshold.

20. The image processing device according to any one of claims 16 to 19, characterized in that: The processing unit is used to perform mean processing on the point cloud of the target object to obtain an initial posture of the target object.

21. The image processing device according to any one of claims 16 to 19, characterized in that: The processing unit is used to obtain the shape prior information based on a preset dimensionality reduction algorithm, and the shape prior information includes an SDF value.

22. The image processing device according to claim 21, characterized in that: The dimension reduction algorithm includes a PCA algorithm.

23. The image processing device according to any one of claims 16 to 19, characterized in that: The processing unit is used for taking the average shape of the shape prior information as the initial shape of the target object.

24. The image processing device according to any one of claims 16 to 19, characterized in that: The processing unit is used to determine the initial posture and initial shape of the target object as the current posture and shape of the target object, and obtain the current mask of the target object based on the current posture and shape of the target object, and obtain the current loss information of the target object based on the current posture and shape of the target object, the current mask of the target object, the target mask and the point cloud of the target object; When the current loss information of the target object meets the set conditions, the processing unit determines the current position and shape of the target object as the target position and shape of the target object.

25. The image processing device according to claim 24, characterized in that: The loss information includes: at least one of mask alignment loss information, point cloud alignment loss information and height alignment loss information; wherein the mask alignment loss information is used for the distance between the target mask of the target object and the current mask; the point cloud alignment loss information is used to indicate the degree of alignment between the object represented by the current posture and shape of the target object and the point cloud of the target object; the height alignment loss information is used to indicate the distance between the height of the current posture and shape and the ground.

26. The image processing device according to claim 24, characterized in that: The processing unit is also used to update the current posture and shape of the target object based on a preset update method when the current loss information of the target object does not meet the set conditions, and repeat the steps of obtaining the current mask of the target object based on the current posture and shape of the target object and obtaining the current loss information of the target object based on the current posture and shape of the target object, the current mask of the target object, the target mask and the point cloud of the target object, until the current loss information of the target object meets the set conditions.

27. The image processing device according to any one of claims 24 to 26, characterized in that: The processing unit is further configured to obtain a current mask of the target object based on a differentiable 3D rendering algorithm of an SDF value and a current posture and shape of the target object.

28. The image processing device according to claim 26, characterized in that: The processing unit is also used to update the current posture and shape of the target object based on a preset gradient descent algorithm with a preset step size or a preset number of steps.

29. The image processing device according to claim 24, characterized in that: The setting condition is that the loss information corresponding to the target object in the current posture and shape is convergent.

30. The image processing device according to claim 16, characterized in that: Also includes: A labeling unit is used to label the target object based on the target posture and target shape.

31. An electronic device, characterized in that: The invention comprises a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the electronic device triggers the electronic device to implement the method according to any one of claims 1 to 15.

32. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed by at least one processor, the device where the computer-readable storage medium is located is controlled to implement the method described in any one of claims 1 to 15.

Citation Information

Patent Citations

  • Image processing method and device, equipment and storage medium

    CN120163869A

  • Object recognition and positioning method and device based on two-dimensional-three-dimensional fusion features

    CN112070838A

  • 6D pose labeling method and system and storage medium

    CN113034593A

  • Class level 6D attitude estimation method based on monocular RGB-D image

    CN114863573A

  • Object attitude estimation method and device, electronic equipment and storage medium

    CN115115700A