Pre-training Model Determination Method, Device, Equipment, and Storage Medium
By performing random masking and image semantic feature extraction on point cloud data, establishing the correspondence between point cloud and image, determining the reconstruction target and performing feature reconstruction, and generating a point cloud pre-training model, the problem of point cloud data processing difficulties and the poor application effect of Transformer in the point cloud field is solved, and the feature extraction and transfer learning ability of point cloud data is improved.
Patent Information
- Application Number
- CN202311768143.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-12-20
AI Technical Summary
The high-dimensional, sparse, disordered and heterogeneous characteristics of point cloud data make it very difficult to analyze and process it, and the application effect of Transformer in the point cloud field is not ideal, mainly due to the lack of large-scale annotation data sets and inductive bias on point clouds.
By obtaining multi-frame timing images and their corresponding original point cloud data, random masking operations and image semantic feature extraction are performed, the correspondence between point clouds and images is established, the point cloud reconstruction target is determined, and the image semantic features and geometric attribute features of the mask area are reconstructed to generate a point cloud pre-trained model.
The feature extraction ability of point cloud data and the transfer learning ability of model are improved, the semantic understanding and geometric alignment ability of the model for point cloud are enhanced, and the performance of point cloud processing tasks is improved.
Smart Images

Figure CN117745944B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as computer vision, deep learning, large models, etc., and particularly relates to a method, apparatus, device, and storage medium for determining a pre-trained model. Background Art
[0002] Currently, point cloud is a commonly used three-dimensional data representation form, which can be obtained from various sensors (such as lidar, depth camera, etc.) and is widely used in fields such as computer vision, robotics, and autonomous driving. However, point cloud data has characteristics such as high-dimensional, sparse, unordered, and heterogeneous, which bring great challenges to the analysis and processing of point clouds.
[0003] Transformer is a deep neural network structure based on the self-attention mechanism, which has achieved great success in the field of natural language processing and has also been gradually introduced into the field of point clouds, showing strong potential. However, due to the lack of large-scale labeled datasets in the field of point clouds and the lack of inductive bias of Transformer for point clouds, the direct application of Transformer on point clouds has unsatisfactory results. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, device, and storage medium for determining a pre-trained model.
[0005] According to the first aspect of the present disclosure, a method for determining a pre-trained model is provided, the method includes:
[0006] Obtain multiple frames of sequential images and multiple frames of original point cloud data corresponding to the multiple frames of sequential images;
[0007] Perform random masking operations on the multiple frames of the original point cloud data to obtain masked point cloud data;
[0008] Extract the image semantic features of the multiple frames of the sequential images to obtain a feature map;
[0009] Project the points in the masked point cloud data onto the feature map to obtain the image coordinates corresponding to the points;
[0010] Determine the point cloud reconstruction targets for the masked regions of the masked point cloud data according to the image semantic features corresponding to the image coordinates and the masked point cloud data, where the point cloud reconstruction targets include: semantic-level point cloud reconstruction targets, geometric-level point cloud reconstruction targets;
[0011] Reconstruct the image semantic features and geometric attribute features of the masked regions according to the point cloud reconstruction targets and the unmasked features of the unmasked regions on the masked point cloud data to generate a point cloud pre-trained model.
[0012] Further, performing random masking operation on multiple frames of the original point cloud data to obtain masked point cloud data includes:
[0013] Performing voxel feature encoding processing on the original point cloud data to obtain processed point cloud data;
[0014] Performing random masking operation on the processed point cloud data to obtain masked point cloud data.
[0015] Further, projecting the point cloud in the masked point cloud data onto the feature map to obtain the image coordinates corresponding to the point cloud includes:
[0016] Determining the correspondence between the point cloud in the masked point cloud data and the image semantic features based on a pre-determined internal and external parameter matrix;
[0017] Projecting the point cloud in the masked point cloud data onto the feature map according to the correspondence to obtain the corresponding image coordinates.
[0018] Further, determining the correspondence between the point cloud in the masked point cloud data and the image semantic features based on a pre-determined internal and external parameter matrix includes:
[0019] Calculating the average value of the three-dimensional point cloud coordinates within the voxel in the masked point cloud data to obtain the average value of the three-dimensional point cloud coordinates;
[0020] Determining the correspondence between the average value of the three-dimensional point cloud coordinates and the image semantic features based on the internal and external parameter matrix.
[0021] Further, determining the point cloud reconstruction target for the masked area of the masked point cloud data according to the image semantic features corresponding to the image coordinates includes:
[0022] Determining the position encoding information of the image semantic features corresponding to the image coordinates;
[0023] Determining the point cloud reconstruction target for the masked area of the masked point cloud data based on the position encoding information.
[0024] Further, the method further includes:
[0025] Using a semantic loss function to perform semantic alignment between the unmasked features in the unmasked area and the image semantic features.
[0026] Further, multiple frames of the temporal images are collected by an image sensor, and multiple frames of the original point cloud data are collected by a lidar, wherein the image sensor and the lidar have been pre-calibrated and timestamp aligned.
[0027] Further, the method further includes:
[0028] Adopt a point cloud feature extraction algorithm to extract features from the unmasked area in the masked point cloud data, and obtain the unmasked features of the unmasked area.
[0029] Further, according to the point cloud reconstruction target and the unmasked features of the unmasked area on the masked point cloud data, reconstruct the image semantic features and geometric attribute features of the masked area to obtain the point cloud pre-training model, including:
[0030] According to the point cloud reconstruction target and the unmasked features of the unmasked area on the masked point cloud data, reconstruct the image semantic features and geometric attribute features of the masked area to obtain the masked features of the masked area;
[0031] Generate the point cloud pre-training model according to the image semantic features, the masked features and the unmasked features.
[0032] According to the second aspect of the present disclosure, a pre-training model determination device is provided, and the device includes:
[0033] An acquisition unit for acquiring multiple frames of sequential images and multiple frames of original point cloud data corresponding to the multiple frames of sequential images;
[0034] A masking processing unit for performing a random masking operation on the multiple frames of the original point cloud data to obtain masked point cloud data;
[0035] An extraction unit for extracting the image semantic features of the multiple frames of the sequential images to obtain a feature map;
[0036] A projection processing unit for projecting the point cloud in the masked point cloud data onto the feature map to obtain the image coordinates corresponding to the point cloud;
[0037] A determination unit for determining the point cloud reconstruction target of the masked area of the masked point cloud data according to the image semantic features corresponding to the image coordinates;
[0038] A generation unit for reconstructing the image semantic features and geometric attribute features of the masked area according to the point cloud reconstruction target and the unmasked features of the unmasked area on the masked point cloud data, so as to generate a point cloud pre-training model.
[0039] Further, the masking processing unit includes:
[0040] A first processing sub-unit for performing voxel feature encoding processing on the original point cloud data to obtain processed point cloud data;
[0041] A second processing subunit, configured to perform a random masking operation on the processed point cloud data to obtain masked point cloud data.
[0042] Further, the projection processing unit includes:
[0043] A first determination subunit, configured to determine the correspondence between the point cloud in the masked point cloud data and the image semantic features based on a pre-determined internal and external parameter matrix;
[0044] A projection processing subunit, configured to project the point cloud in the masked point cloud data onto the feature map according to the correspondence to obtain corresponding image coordinates.
[0045] Further, the first determination subunit includes:
[0046] A calculation module, configured to calculate the average value of the three-dimensional point cloud coordinates within the voxel in the masked point cloud data to obtain the average value of the three-dimensional point cloud coordinates;
[0047] A determination module, configured to determine the correspondence between the average value of the three-dimensional point cloud coordinates and the image semantic features based on the internal and external parameter matrix.
[0048] Further, the determination unit includes:
[0049] A second determination subunit, configured to determine the position encoding information of the image semantic features corresponding to the image coordinates;
[0050] A third determination subunit, configured to determine the point cloud reconstruction target of the masked area of the masked point cloud data based on the position encoding information.
[0051] Further, the apparatus further includes:
[0052] An alignment processing unit, configured to perform semantic alignment between the unmasked features in the unmasked area and the image semantic features by using a semantic loss function.
[0053] Further, multiple frames of the temporal images are collected by an image sensor, and multiple frames of the original point cloud data are collected by a lidar, wherein the image sensor and the lidar have been pre-calibrated and timestamp-aligned.
[0054] Further, the apparatus further includes:
[0055] A feature extraction unit, configured to extract features from the unmasked area in the masked point cloud data by using a point cloud feature extraction algorithm to obtain the unmasked features in the unmasked area.
[0056] Further, the generation unit includes:
[0057] A reconstruction subunit, configured to reconstruct the image semantic features and geometric attribute features of the masked region according to the reconstruction target of the point cloud and the unmasked features of the unmasked region on the masked point cloud data, so as to obtain the masked features of the masked region;
[0058] A generation subunit, configured to generate the point cloud pre-training model according to the image semantic features, the masked features and the unmasked features.
[0059] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0060] At least one processor; and
[0061] A memory communicatively connected to the at least one processor; wherein,
[0062] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute any one of the methods.
[0063] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any one of the methods described above.
[0064] According to a fifth aspect of the present disclosure, there is provided a computer program product, including: a computer program, the computer program is stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to execute the method described in the first aspect.
[0065] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0067] Figure 1 is a flowchart of a method for determining a pre-training model according to an embodiment of the present disclosure;
[0068] Figure 2 is a schematic diagram of an implementation scenario of a method for determining a pre-training model that can implement the embodiment of the present disclosure;
[0069] Figure 3It is a flowchart of a method for determining a pre-trained model provided by an embodiment of the present disclosure;
[0070] Figure 4 It is a flowchart of a method for determining a pre-trained model provided by an embodiment of the present disclosure;
[0071] Figure 5 It is a flowchart of a method for determining a pre-trained model provided by an embodiment of the present disclosure;
[0072] Figure 6 It is a schematic framework diagram of a device for determining a pre-trained model provided by an embodiment of the present disclosure;
[0073] Figure 7 It is a schematic framework diagram of an optional device for determining a pre-trained model provided by an embodiment of the present disclosure;
[0074] Figure 8 It is a block diagram of an electronic device for implementing a method for determining a pre-trained model according to an embodiment of the present disclosure. Detailed implementation manners
[0075] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0076] First, the terms involved in this application are explained:
[0077] The CLIP (Contrastive Language-Image Pre-Training, hereinafter referred to as CLIP) model is a pre-trained neural network model released by OpenAI for matching images and texts, and can be said to be a classic in the field of multi-modal research in recent years.
[0078] Voxelization is to convert the geometric form representation of an object into a voxel representation form closest to the object, generating voxel data, which contains the surface information and internal attributes of the model.
[0079] Existing self-supervised pre-training schemes for point clouds based on mask reconstruction mainly start from the perspective of geometric attributes. However, the point cloud distribution is sparse and uneven, and the point distributions within different voxels may also be different. Therefore, the geometric relationships within the point cloud may also be unstable, and geometric reconstruction based on this may also be not robust.
[0080] Currently, the utilization of point cloud-image multimodal information mainly adopts methods such as contrast learning, joint reconstruction, or rendering. However, in the case of a small-scale dataset, it is difficult to ensure the generalization of the image features obtained by training the image branch. At the same time, there are also robustness issues with directly using color information.
[0081] To solve the above problems, the present disclosure provides a method, apparatus, device, and storage medium for determining a pre-trained model, which is applied to the field of artificial intelligence technology, specifically related to technical fields such as computer vision, deep learning, and large models, and can be applied to scenarios such as autonomous driving, so as to generate a point cloud pre-trained model through multiple frames of temporal images and their corresponding multiple frames of original point cloud data, which can provide a powerful feature extractor for downstream point cloud-related tasks and improve the transfer learning ability of the model.
[0082] Figure 1 is a flowchart of a method for determining a pre-trained model provided according to an embodiment of the present disclosure. As Figure 1 shown, a method for determining a pre-trained model provided by the present disclosure includes the following method steps:
[0083] S101, obtain multiple frames of temporal images and multiple frames of original point cloud data corresponding to the multiple frames of temporal images;
[0084] S102, perform random masking operations on the multiple frames of the above original point cloud data to obtain masked point cloud data;
[0085] S103, extract the image semantic features of the multiple frames of the above temporal images to obtain a feature map;
[0086] S104, project the points in the above masked point cloud data onto the above feature map to obtain the image coordinates corresponding to the above point cloud;
[0087] S105, determine the point cloud reconstruction targets for the masked areas of the above masked point cloud data according to the image semantic features corresponding to the above image coordinates and the above masked point cloud data, where the above point cloud reconstruction targets include: semantic-level point cloud reconstruction targets and geometric-level point cloud reconstruction targets;
[0088] S106, reconstruct the image semantic features and geometric attribute features of the above masked areas according to the above point cloud reconstruction targets and the unmasked features of the unmasked areas on the above masked point cloud data to generate a point cloud pre-trained model.
[0089] The pre-trained model determination method provided by the present disclosure example can specifically be a point cloud self-attention model pre-training method based on masked modeling. By performing random masking operations on the original point cloud data, the data augmentation ability of the point cloud can be enhanced, and the generalization performance of the model can be improved. Then, by extracting the image semantic features of the temporal images, the complementary information between the images and the point cloud can be utilized to improve the semantic understanding ability of the model. By projecting the masked point cloud data onto the feature map, the correspondence between the point cloud and the image can be established, and the geometric alignment ability of the model can be improved. After that, by determining the point cloud reconstruction target based on the image semantic features, the model can be guided to learn the relationships between different local structures, and the structure perception ability of the model can be improved. By reconstructing the masked area according to the point cloud reconstruction target and the unmasked features, the model can be trained to predict the masked point cloud structure from the visible local structure of the point cloud, and the reconstruction ability of the model can be improved. The finally generated point cloud pre-trained model can provide a powerful feature extractor for downstream point cloud-related tasks, and improve the transfer learning ability of the model.
[0090] Optionally, the method examples provided by the embodiments of the present disclosure can all be but are not limited to being used in urban traffic scenarios and autonomous driving scenarios to quickly, accurately, and stably detect surrounding vehicles, pedestrians, traffic obstacles, and other targets to ensure safe and efficient driving. This method can use multi-frame temporal images and point cloud data collected by in-vehicle cameras and lidar to generate a point cloud pre-trained model. However, it is not limited thereto, and it can also come from remote sensing images and point cloud data of satellites or drones. These data can provide the appearance and geometric information of the targets, as well as the motion trajectories and state changes of the targets (3D objects, etc.).
[0091] In an alternative example, the method steps of the label generation method based on detection boxes provided by the present disclosure example will be explained in more detail below, and some example application scenarios will be given.
[0092] As Figure 2 shown, the obtained multi-frame temporal images and the multi-frame original point cloud data corresponding to the multi-frame temporal images can come from in-vehicle image sensors (cameras) and lidar, or can also come from remote sensing images and point cloud data of satellites or drones.
[0093] In one example, a vehicle uses an image sensor to collect multiple frames of the above-mentioned temporal images and uses a lidar to collect multiple frames of the above-mentioned original point cloud data, where the above-mentioned image sensor and the above-mentioned lidar have been pre-calibrated and timestamp-aligned.
[0094] Optionally, the above multi-frame temporal images can be consecutive-frame RGB images, and the multi-frame original point cloud data corresponding to the multi-frame temporal images can be the point cloud data corresponding to the lidar, including point cloud coordinates, corresponding point reflection intensity information, and timestamp information. The RGB images can be acquired by a single camera or by multiple cameras.
[0095] By using an image sensor to acquire multi-frame temporal images, the high resolution and rich color information of the images can be utilized to improve the visual perception ability of the model; by using a lidar to acquire multi-frame original point cloud data, the high precision and sparsity of the point cloud can be utilized to improve the spatial perception ability of the model; through pre-calibration and timestamp alignment, accurate correspondence and synchronization between the image sensor and the lidar can be achieved, improving the cross-modal fusion ability of the model.
[0096] In the example of the present disclosure, a new multi-modal self-supervised pre-training method for image-point cloud that fuses semantic and spatial features is proposed for a large amount of unlabeled paired image-point cloud multi-modal data collected for the autonomous driving scenario. Using the paired image-point cloud data, combined with the characteristics of the CLIP model in achieving image-text semantic alignment, the point cloud semantic alignment is achieved by aligning the point cloud with the CLIP features of the image, so as to effectively mine the key multi-modal point pair information and effectively improve the quality of the model's self-supervised pre-training learning. At the same time, in order to ensure the original spatial geometric relationship characteristics of the point cloud data, a geometric attribute reconstruction target is further defined to be jointly trained with the semantic attribute alignment and reconstruction target, so as to extract more effective feature representations, provide better initial network parameters for downstream tasks, and improve the performance indicators of downstream tasks.
[0097] In the example of the present disclosure, a new multi-modal self-supervised pre-training scheme for image-point cloud that fuses semantic and spatial features is proposed for a large amount of unlabeled paired image-point cloud multi-modal data collected for the autonomous driving scenario. The purpose of this scheme is to utilize the correspondence between images and point clouds, extract semantic-rich image features through the CLIP model, and use them as the self-supervised pre-training targets for masked region semantic reconstruction and semantic alignment of point cloud features. At the same time, combined with the self-supervised signal of geometric attribute reconstruction, it is ensured that the point cloud features retain the characteristics of spatial geometric relationship description, and finally achieve better generalization. The point cloud model pre-trained using this method can obtain better performance gains after fine-tuning with a small amount of data in downstream tasks (such as: 3D object detection of point clouds, 3D object segmentation of point clouds, etc.). Therefore, this method will be widely applied to multiple application scenarios such as intelligent transportation and autonomous driving.
[0098] In an optional example, Figure 3 is a flowchart of an optional pre-training model determination method provided according to an embodiment of the present disclosure, as Figure 3As shown, the above-mentioned random masking operation on the multi-frame original point cloud data to obtain the masked point cloud data includes:
[0099] S301, performing voxel feature encoding processing on the original point cloud data to obtain the processed point cloud data;
[0100] S302, performing a random masking operation on the processed point cloud data to obtain the masked point cloud data.
[0101] Optionally, in the examples of the present disclosure, the calibrated and clock-synchronized image data and point cloud data obtained in the urban traffic scenario (intelligent transportation and autonomous driving) can be used as input. For the input discrete point cloud data, in the examples of the present disclosure, still as Figure 2 shown, first perform voxelization on it, convert it into the form of voxels or pillars and define it as the processed point cloud data ∈ R N×C×D×H×W , where the number of non-zero elements is N.
[0102] Considering that the network structure will perform downsampling, the present disclosure calculates the corresponding geometric attributes for each voxel after downsampling: the center point, whether it is non-zero, the surface normal, the normal vector, etc. These attributes will be used as the targets for geometric attribute reconstruction.
[0103] After that, the present disclosure performs a random masking operation on the processed point cloud data P with a masking ratio of 70%, and then obtains the processed point cloud data P m ∈R N×C×D×H×W , where the number of non-zero elements is N v , and the number of elements filtered out by the mask is N m =N - N v .
[0104] As an alternative example, Figure 4 is a flowchart of an alternative pre-training model determination method provided according to an embodiment of the present disclosure. As Figure 4 shown, the above-mentioned projection of the point cloud in the masked point cloud data onto the above-mentioned feature map to obtain the image coordinates corresponding to the point cloud includes:
[0105] S401, determining the correspondence between the points in the masked point cloud data and the above-mentioned image semantic features based on the pre-determined internal and external parameter matrices;
[0106] S402, projecting the point cloud in the masked point cloud data onto the above-mentioned feature map according to the above correspondence to obtain the corresponding image coordinates.
[0107] The present disclosure uses the image branch in the pre-trained CLIP model for feature extraction to obtain the feature map H img×W img ×C img , then, through the internal and external parameter matrices, project the points on the point cloud onto the image to obtain the corresponding image coordinates, and then obtain the corresponding image features through interpolation, which are used as the point cloud reconstruction target at the semantic level. Then, for the data P after mask processing m , the present disclosure uses a common 3D feature extraction network as a feature extractor to extract 3D features.
[0108] In the example of the present disclosure, by determining the correspondence between the point cloud in the masked point cloud data and the above-mentioned image semantic features based on the pre-determined internal and external parameter matrices, the point cloud in the masked point cloud data can be projected onto the feature map to obtain the corresponding image coordinates, which can enable the point cloud and the image to share the same coordinate space, thereby facilitating subsequent feature fusion and reconstruction and improving the feature representation ability of the model.
[0109] In an alternative example Figure 5 is a flowchart of a method for determining a pre-trained model provided according to an embodiment of the present disclosure, as Figure 5 shown, determining the correspondence between the point cloud in the masked point cloud data and the above-mentioned image semantic features based on the pre-determined internal and external parameter matrices includes:
[0110] S501, calculate the average value of the three-dimensional point cloud coordinates within the voxel in the masked point cloud data to obtain the average value of the three-dimensional point cloud coordinates.
[0111] S502, determine the correspondence between the average value of the three-dimensional point cloud coordinates and the above-mentioned image semantic features based on the internal and external parameter matrices.
[0112] In an example, for the original point cloud data, first perform voxelization or extract pillar encoding features to obtain the encoded point cloud features described as HxWxDxC. If it is a pillar feature, then D = 1, and the corresponding point cloud features are described as HxWxC. That is, N = HxWxD voxels or Pillars can be obtained, and the feature is C-dimensional. For the N features, perform a masking operation randomly according to a certain ratio to obtain N1 masked features and N2 unmasked features.
[0113] For the masked features, the present disclosure rearranges them into a structure of HxWxDxC or HxWxC, and directly uses a 3D convolutional neural network or a 2D convolutional neural network to extract features from the voxel or pillar to obtain the re-encoded 3D point cloud features. Note that here the present disclosure can also only perform subsequent operations on the N2 unmasked features.
[0114] An optional example is that the point cloud feature extraction network is mainly responsible for extracting point cloud features based on the input point cloud information. For the point cloud feature extractor, there can also be various different choices.
[0115] An optional embodiment Figure 6 is a flowchart of a method for determining a pre-trained model according to an embodiment of the present disclosure, as Figure 6 shown. Determining the point cloud reconstruction target of the masked area of the masked point cloud data based on the image semantic feature corresponding to the above image coordinate and the above masked point cloud data includes:
[0116] S601, determining the position encoding information of the image semantic feature corresponding to the above image coordinate;
[0117] S602, based on the above position encoding information, determining the semantic-level point cloud reconstruction target of the masked area of the masked point cloud data;
[0118] S603, determining the downsampling ratio corresponding to the network structure of the point cloud pre-trained model;
[0119] S604, determining the geometric-level point cloud reconstruction target of the masked point cloud data according to the point cloud within the minimum mask unit defined by the downsampling ratio.
[0120] By determining the position encoding information of the image semantic feature corresponding to the image coordinate, the spatial structure and semantic information of the image can be used to assist in the reconstruction of the point cloud, improving the cross-modal fusion ability of the model; based on the position encoding information, the semantic-level point cloud reconstruction target of the masked area of the masked point cloud data can be determined; in addition, considering that the point cloud feature extraction network will perform downsampling during feature extraction. Therefore, the present disclosure considers a block-level masking strategy when calculating the masked point cloud data. For example: if the downsampling ratio (K1, K2, K3) corresponding to the network structure of the point cloud pre-trained model, then the range of K1xK2xK3 is used as the minimum mask unit.
[0121] Determine the geometric-level point cloud reconstruction target of the masked point cloud data according to the point cloud within the minimum mask unit (such as a voxel block with a size of 4x4x4, etc.). According to the semantic-level point cloud reconstruction target and the geometric-level point cloud reconstruction target, the model can be trained to predict the masked point cloud structure from the visible local structure of the point cloud.
[0122] For the image data of point cloud pairing, the present disclosure uses the image branch of the pre-trained CLIP model as an encoder to extract image features. The size of the input image is 224x224, and the dimension of the output features is 14x14x512. The image features will serve as the target for point cloud feature alignment, that is, the target for mask reconstruction. For the semantic feature reconstruction target, first, the average coordinates need to be calculated from the 3D point coordinates within the voxel / pillar to obtain the average coordinates, then projected onto the image using the internal and external parameters, and finally, the reconstruction target corresponding to the final point cloud is obtained after interpolating the CLIP image features. Among them, the CLIP image features are pre-trained by image-text pairs and can effectively align semantic features. By aligning the image features, the point cloud features also achieve the goal of semantic alignment.
[0123] In the example of the present disclosure, the extracted unmasked features and position encoding information are used as the input to a decoder based on the Transformer structure to reconstruct the image semantic features and geometric attribute features in the masked area, and align the 3D features in the unmasked area with the image semantic features in the semantic space to implement the pre-training process. The obtained pre-training parameters will be used as the initialization parameters of the backbone network for downstream tasks, including 3D detection, segmentation, etc.
[0124] In the example of the present disclosure, by performing random masking operations on the original point cloud data, the occlusion and noise of the point cloud can be simulated, enhancing the robustness and generalization ability of the model; by performing random masking operations on the processed point cloud data, the model can learn from voxel features at different scales, improving the multi-scale perception ability of the model; by reconstructing the masked point cloud data, the model can be trained to predict the masked point cloud structure from the visible local structure of the point cloud, improving the reconstruction ability and self-attention ability of the model.
[0125] The present disclosure uses a 2D / 3D feature extraction network (point cloud feature extractor) to implement feature extraction. During feature extraction, downsampling operations are synchronously performed. Therefore, when performing geometric attribute extraction, the corresponding range of voxels or pillars is expanded according to the downsampling ratio. For example, if the downsampling is 4 times, the voxels within the previous 4x4x4 range need to be recombined into new voxels, and the point cloud therein is used to calculate the geometric attributes as the geometric attribute reconstruction target. For example, for the geometric attributes of the center point, the present disclosure averages all the points within the voxel to obtain the average center point.
[0126] As an optional embodiment, according to the above-mentioned point cloud reconstruction target and the unmasked features of the unmasked area on the masked point cloud data, the image semantic features and geometric attribute features in the masked area are reconstructed to obtain the above-mentioned point cloud pre-training model, including:
[0127] S701. Reconstruct the image semantic features and geometric attribute features of the masked region based on the above point cloud reconstruction target and the unmasked features of the unmasked region on the masked point cloud data to obtain the masked features of the masked region;
[0128] S702. Generate the above point cloud pre-training model according to the above image semantic features, the above masked features and the above unmasked features.
[0129] Optionally, the point cloud self-attention model pre-training method is a self-supervised learning method using unlabeled data, which can improve the generalization ability and transfer ability of the point cloud processing model. The application scenarios of this method mainly include:
[0130] 3D object recognition scenario: For the task of classifying or detecting 3D objects based on point cloud data, the point cloud self-attention model pre-training method can improve the model's semantic understanding and geometric alignment ability of 3D objects, thereby improving the accuracy of classification or detection.
[0131] 3D object segmentation scenario: For the task of semantic segmentation or instance segmentation of 3D objects based on point cloud data, the point cloud self-attention model pre-training method can improve the model's structural perception and reconstruction ability of 3D objects, thereby improving the accuracy and robustness of segmentation.
[0132] 3D object reconstruction scenario: For the task of fully or partially reconstructing 3D objects based on point cloud data, the point cloud self-attention model pre-training method can improve the model's multi-scale perception and generation ability of 3D objects, thereby improving the quality and naturalness of reconstruction.
[0133] By reconstructing the image semantic features and geometric attribute features of the masked region according to the point cloud reconstruction target and the unmasked features of the unmasked region on the masked point cloud data, the rich semantic information of the image and the local geometric relationship of the point cloud can be utilized to restore the masked point cloud structure and improve the model's representation ability; by obtaining the masked features of the masked region, they can be fused and aligned with the image semantic features and unmasked features, thereby realizing cross-modal feature representation and improving the model's semantic understanding ability and geometric consistency ability; by generating a point cloud pre-training model according to the image semantic features, masked features and unmasked features, and combining a large amount of unlabeled data, the generalization ability of the model can be further improved.
[0134] In the example of the present disclosure, a masked feature reconstruction decoder is used to calculate the cross-attention mechanism for decoding and reconstruction. Among them, the input query is the masked feature (masked Token), which can be obtained by initialization or by a 3D feature extraction network. The Key and Value are defined by the unmasked feature Z v That is: through the unmasked feature Zv to reconstruct the masked feature Z m The obtained feature is then adjusted in vector length through Linear Layer 1 to align with the dimension of the CLIP reconstructed feature to obtain Y m and Y v For geometric and semantic attributes, they can be completed by defining two independent decoders 1 and encoders 2, and the reconstructed masked features are respectively defined as Z m and Z ‘ m .
[0135] Adopting the examples of the present disclosure, by calculating the average value of the three-dimensional point cloud coordinates within the voxel in the masked point cloud data, the density of the point cloud can be reduced, the computational amount and memory consumption can be reduced, while the main structural information of the point cloud is retained; by based on the pre-determined internal and external parameter matrices, the geometric transformation between the known camera and lidar can be utilized to achieve precise alignment between the point cloud and the image, improving the geometric consistency ability of the model; by determining the correspondence between the average value of the three-dimensional point cloud coordinates and the image semantic features, the point cloud and the image can share the same feature space, thus facilitating subsequent feature fusion and reconstruction and improving the feature representation ability of the model.
[0136] As an alternative example, the above method further includes:
[0137] Adopting a semantic loss function to semantically align the unmasked features in the above unmasked region with the image semantic features.
[0138] In one example, for the reconstructed semantic masked feature Y m combined with the semantically unmasked feature Y v are both aligned with the image semantic features (T m , T v ), that is, semantic alignment is performed in the feature space. Two alignment losses are respectively defined as the reconstruction loss and the feature distillation loss, which are specifically described as follows:
[0139] ‖Y m - T m ‖2 + ‖Y v - T v ‖2
[0140] In one example, the geometric attribute loss function is obtained by the reconstructed masked geometric feature and the unmasked geometric feature. After being processed by an MLP layer, a predicted geometric center point with a dimension of 3 is obtained, and the linear layer 2 predicts the loss function L2-loss between the geometric center point and the true center point. The specific description is as follows:
[0141] ‖C pred - C target ‖2
[0142] Among them, C pred represents the property of the center point of the predicted voxel, and C target is represented as the ground truth label.
[0143] In the examples of the present disclosure, by adopting a semantic loss function, the model can pay more attention to the semantic information between the point cloud and the image during the training process, rather than just the appearance information, improving the semantic understanding ability of the model, improving the cross-domain semantic consistency of the model, and further improving the generalization ability of the model.
[0144] An optional example is that the above method further includes:
[0145] Adopt a point cloud feature extraction algorithm to extract features from the unmasked area in the above masked point cloud data to obtain the unmasked features of the unmasked area.
[0146] Optionally, the point cloud feature extraction algorithm refers to a class of algorithms for extracting useful information from point cloud data, usually including steps such as preprocessing of the point cloud, feature description, and feature matching.
[0147] By adopting a point cloud feature extraction algorithm, features can be extracted from the unmasked area in the masked point cloud data to obtain unmasked features, thereby enhancing the model's representation ability and discrimination ability for the point cloud; by extracting features from the unmasked area, the original information of the point cloud can be retained, while reducing the influence of noise and redundancy, improving the robustness and efficiency of the model; by obtaining unmasked features, they can be fused and aligned with the image semantic features, thereby realizing cross-modal feature representation and improving the model's semantic understanding ability and reconstruction ability.
[0148] The embodiments of the present disclosure propose a new self-supervised pre-training method for point cloud and image multi-modal based on the joint training of semantic alignment and geometric attributes; through the CLIP pre-training model that has achieved image-text semantic alignment, feature extraction is performed on the image branch to extract semantically aligned image features, and this is used as the target for reconstructing the point cloud features from a semantic perspective to drive the point cloud feature extraction to achieve the goal of semantic consistency. The feature correspondence relationship between the point cloud and the image is constructed by using the internal and external parameters. Among them, for multiple points within the pillar / voxel, the average method is used to calculate the center point, project it onto the image, and use the difference method to obtain the target image features. By extracting geometric attributes for each downsampled voxel, including: center point, surface attribute, occupancy status, etc., and combining semantic attribute reconstruction and semantic alignment as the learning objective of pre-training.
[0149] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0150] According to an embodiment of the present disclosure, Figure 7 is a schematic framework diagram of a pre-trained model determination device provided according to an embodiment of the present disclosure. As Figure 7 shown, the present disclosure also provides a pre-trained model determination device. The pre-trained model determination device 700 includes:
[0151] An acquisition unit 701, configured to acquire multiple frames of sequential images and multiple frames of original point cloud data corresponding to the multiple frames of sequential images;
[0152] A mask processing unit 702, configured to perform a random masking operation on the multiple frames of the original point cloud data to obtain masked point cloud data;
[0153] An extraction unit 703, configured to extract image semantic features of the multiple frames of the sequential images to obtain a feature map;
[0154] A projection processing unit 704, configured to project the points in the masked point cloud data onto the feature map to obtain image coordinates corresponding to the points;
[0155] A determination unit 705, configured to determine a point cloud reconstruction target for a masked area of the masked point cloud data according to the image semantic features corresponding to the image coordinates and the masked point cloud data, where the point cloud reconstruction target includes: a semantic-level point cloud reconstruction target and a geometric-level point cloud reconstruction target;
[0156] A generation unit 706, configured to reconstruct the image semantic features and geometric attribute features of the masked area according to the point cloud reconstruction target and unmasked features of an unmasked area on the masked point cloud data to generate a point cloud pre-trained model.
[0157] According to one or more examples of the present disclosure, the mask processing unit includes:
[0158] A first processing subunit, configured to perform voxel feature encoding processing on the original point cloud data to obtain processed point cloud data;
[0159] A second processing subunit, configured to perform a random masking operation on the processed point cloud data to obtain masked point cloud data.
[0160] According to one or more examples of the present disclosure, the projection processing unit includes:
[0161] A first determination subunit, configured to determine a correspondence between the points in the masked point cloud data and the image semantic features based on a pre-determined internal and external parameter matrix;
[0162] A projection processing subunit, configured to project the point cloud in the masked point cloud data onto the feature map according to the above correspondence relationship to obtain corresponding image coordinates.
[0163] According to one or more examples of the present disclosure, the above first determination subunit includes:
[0164] A calculation module, configured to calculate the average value of the three-dimensional point cloud coordinates within the voxel in the masked point cloud data to obtain the average value of the three-dimensional point cloud coordinates;
[0165] A determination module, configured to determine the correspondence relationship between the average value of the three-dimensional point cloud coordinates and the above image semantic features based on the above internal and external parameter matrices.
[0166] According to one or more examples of the present disclosure, the above determination unit includes:
[0167] A second determination subunit, configured to determine the position encoding information of the image semantic features corresponding to the above image coordinates;
[0168] A third determination subunit, configured to determine the semantic-level point cloud reconstruction target of the masked area of the masked point cloud data based on the above position encoding information;
[0169] A fourth determination subunit, configured to determine the downsampling ratio corresponding to the network structure of the point cloud pre-training model;
[0170] A fifth determination subunit, configured to determine the geometric-level point cloud reconstruction target of the masked point cloud data according to the point cloud within the minimum mask unit defined by the above downsampling ratio.
[0171] According to one or more examples of the present disclosure, the above device further includes:
[0172] An alignment processing unit, configured to perform semantic alignment on the unmasked features in the unmasked area and the image semantic features by using a semantic loss function.
[0173] According to one or more examples of the present disclosure, multiple frames of the above temporal images are acquired by an image sensor, and multiple frames of the above original point cloud data are acquired by a lidar, wherein the above image sensor and the above lidar have been pre-calibrated and timestamp-aligned.
[0174] According to one or more examples of the present disclosure, the above device further includes:
[0175] A feature extraction unit, configured to extract features from the unmasked area in the masked point cloud data by using a point cloud feature extraction algorithm to obtain the unmasked features of the unmasked area.
[0176] According to one or more examples of the present disclosure, the above generation unit includes:
[0177] A reconstruction subunit, configured to reconstruct the image semantic features and geometric attribute features of the masked region according to the above-mentioned point cloud reconstruction target and the unmasked features of the unmasked region on the above-mentioned masked point cloud data, so as to obtain the masked features of the masked region;
[0178] A generation subunit, configured to generate the above-mentioned point cloud pre-training model according to the above-mentioned image semantic features, the above-mentioned masked features and the above-mentioned unmasked features.
[0179] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0180] According to an embodiment of the present disclosure, the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of the above.
[0181] According to an embodiment of the present disclosure, the present disclosure provides a computer program product, which includes: a computer program, the computer program is stored in a readable storage medium, at least one processor of the electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program to cause the electronic device to execute the solution provided in any one of the above embodiments.
[0182] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0183] As Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 802 or computer programs loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0184] Multiple components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0185] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the pre-trained model determination method. For example, in some embodiments, the pre-trained model determination method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the pre-trained model determination method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the pre-trained model determination method in any other appropriate way (e.g., by means of firmware).
[0186] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0187] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0188] In the context of this disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0189] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0190] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0191] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.
[0192] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is made herein.
[0193] The above specific embodiments do not constitute a limitation to the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for determining a pre-trained model, the method comprising: Obtaining multiple frames of temporal images and multiple frames of original point cloud data corresponding to the multiple frames of temporal images; Performing a random masking operation on the multiple frames of the original point cloud data to obtain masked point cloud data; Extracting the image semantic features of the multiple frames of the temporal images to obtain a feature map; Calculating the average value of the three-dimensional point cloud coordinates within the voxels in the masked point cloud data to obtain the average value of the three-dimensional point cloud coordinates; Based on a pre-determined internal and external parameter matrix, determining the correspondence between the average value of the three-dimensional point cloud coordinates and the image semantic features; According to the correspondence, projecting the point cloud in the masked point cloud data onto the feature map to obtain the image coordinates corresponding to the point cloud; According to the image semantic features corresponding to the image coordinates and the masked point cloud data, determining the point cloud reconstruction target for the masked area of the masked point cloud data, wherein the point cloud reconstruction target includes: a semantic-level point cloud reconstruction target and a geometric-level point cloud reconstruction target; According to the point cloud reconstruction target and the unmasked features of the unmasked area on the masked point cloud data, reconstructing the image semantic features and geometric attribute features of the masked area to generate a point cloud pre-trained model, the point cloud pre-trained model being obtained through image-point cloud multi-modal self-supervised training by fusing image semantic features and spatial features, and the spatial features retaining the original spatial combination relationship characteristics of the point cloud data; Wherein, determining the geometric-level point cloud reconstruction target in the point cloud reconstruction target for the masked area of the masked point cloud data includes: Determining the downsampling ratio corresponding to the network structure of the point cloud pre-trained model; According to the point cloud within the minimum masking unit defined by the downsampling ratio, determining the geometric-level point cloud reconstruction target of the masked point cloud data, and the geometric level at least includes a center point, whether it is non-zero, and a surface normal or a normal vector.
2. The method according to claim 1, wherein The performing a random masking operation on the multiple frames of the original point cloud data to obtain masked point cloud data includes: Performing voxel feature encoding processing on the original point cloud data to obtain processed point cloud data; Performing a random masking operation on the processed point cloud data to obtain masked point cloud data.
3. The method according to claim 1, wherein Determining the semantic-level point cloud reconstruction target in the point cloud reconstruction target for the masked area of the masked point cloud data includes: Determining the position encoding information of the image semantic features corresponding to the image coordinates; Based on the position encoding information, determining the semantic-level point cloud reconstruction target for the masked area of the masked point cloud data.
4. The method according to any one of claims 1 to 3, wherein, The method further includes: Using a semantic loss function to semantically align the unmasked features of the unmasked area with the image semantic features.
5. According to the method according to any one of claims 1 to 3, wherein, The multiple frames of temporal images are collected by an image sensor, and the multiple frames of original point cloud data are collected by a lidar, wherein the image sensor and the lidar have been pre-calibrated and time-stamped aligned.
6. The method according to any one of claims 1 to 3, wherein The method further includes: Using a point cloud feature extraction algorithm, feature extraction is performed on the unmasked region in the masked point cloud data to obtain the unmasked features of the unmasked region.
7. The method according to any one of claims 1 to 3, wherein Based on the point cloud reconstruction target and the unmasked features of the unmasked region in the masked point cloud data, reconstructing the image semantic features and geometric attribute features of the masked region to obtain the point cloud pre-training model, including: Based on the point cloud reconstruction target and the unmasked features of the unmasked region in the masked point cloud data, reconstructing the image semantic features and geometric attribute features of the masked region to obtain the masked features of the masked region; Generating the point cloud pre-training model according to the image semantic features, the masked features, and the unmasked features.
8. A device for determining a pre-training model, the device includes: An acquisition unit, configured to acquire multiple frames of temporal images and multiple frames of original point cloud data corresponding to the multiple frames of temporal images; A masking processing unit, configured to perform a random masking operation on the multiple frames of the original point cloud data to obtain masked point cloud data; An extraction unit, configured to extract the image semantic features of the multiple frames of the temporal images to obtain a feature map; A projection processing unit, configured to project the point cloud in the masked point cloud data onto the feature map to obtain the image coordinates corresponding to the point cloud; A determination unit, configured to determine the point cloud reconstruction target of the masked region of the masked point cloud data according to the image semantic features corresponding to the image coordinates, where the point cloud reconstruction target includes: a semantic-level point cloud reconstruction target and a geometric-level point cloud reconstruction target; A generation unit, configured to reconstruct the image semantic features and geometric attribute features of the masked region according to the point cloud reconstruction target and the unmasked features of the unmasked region in the masked point cloud data to generate a point cloud pre-training model, where the point cloud pre-training model is obtained by multi-modal self-supervised training of image point clouds by fusing image semantic features and spatial features, and the spatial features retain the original spatial combination relationship characteristics of the point cloud data; The determination unit includes: A fourth determination subunit, configured to determine the downsampling ratio corresponding to the network structure of the point cloud pre-training model; A fifth determination subunit, configured to determine the geometric-level point cloud reconstruction target of the masked point cloud data according to the point cloud within the minimum masking unit defined by the above downsampling ratio, where the geometric level at least includes a center point, whether it is non-zero, and a surface normal or a normal vector; The projection processing unit is specifically configured to: Calculate the average value of the three-dimensional point cloud coordinates within the voxel in the masked point cloud data to obtain the average value of the three-dimensional point cloud coordinates; Based on the pre-determined internal and external parameter matrices, determine the correspondence between the average value of the three-dimensional point cloud coordinates and the image semantic features; According to the correspondence, project the point cloud in the masked point cloud data onto the feature map to obtain the image coordinates corresponding to the point cloud.
9. The device according to claim 8, wherein, The masking processing unit includes: A first processing subunit, configured to perform voxel feature encoding processing on the original point cloud data to obtain processed point cloud data; A second processing subunit, configured to perform a random masking operation on the processed point cloud data to obtain masked point cloud data.
10. The apparatus according to claim 8, wherein, The determining unit includes: A second determining subunit, configured to determine the position encoding information of the image semantic feature corresponding to the image coordinate; A third determining subunit, configured to determine the point cloud reconstruction target of the masked area of the masked point cloud data based on the position encoding information.
11. The device according to any one of claims 8 to 10, wherein, The apparatus further includes: An alignment processing unit, configured to perform semantic alignment between the unmasked feature of the unmasked area and the image semantic feature by using a semantic loss function.
12. The device according to any one of claims 8 to 10, wherein, The multiple frames of the temporal images are collected by an image sensor, and the multiple frames of the original point cloud data are collected by a lidar, wherein the image sensor and the lidar have been pre-calibrated and timestamp-aligned.
13. The device according to any one of claims 8 to 10, wherein The apparatus further includes: A feature extraction unit, configured to extract features from the unmasked area of the masked point cloud data by using a point cloud feature extraction algorithm to obtain the unmasked feature of the unmasked area.
14. The device according to any one of claims 8 to 10, wherein The generating unit includes: A reconstruction subunit, configured to reconstruct the image semantic feature and the geometric attribute feature of the masked area according to the point cloud reconstruction target and the unmasked feature of the unmasked area on the masked point cloud data to obtain the masked feature of the masked area; A generating subunit, configured to generate the point cloud pre-training model according to the image semantic feature, the masked feature, and the unmasked feature.
15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
17. A computer program product, comprising a computer program, which when executed by a processor, implements the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Model pre-training and training method, point cloud detection and segmentation method and device
CN117132850A