Model training method, point cloud encoding method, object processing method and device
By training a point cloud coding model using feature distribution differences with image features, the method addresses the limited feature expression of point cloud data, improving coding reliability and accuracy.
Patent Information
- Application Number
- CN202311110272.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-08-30
AI Technical Summary
The lack of texture information of point cloud data leads to insufficient feature expression capabilities of the encoding results, affecting the reliability of the encoding results.
By obtaining the difference in feature distribution between point cloud data and scene images, the point cloud encoding model is used for training, and the feature expression ability is improved in combination with the image segmentation model, and the training effect of the point cloud encoding model is enhanced.
The feature expression ability and reliability of point cloud encoding results are improved, and the accuracy of three-dimensional object detection and segmentation tasks are enhanced.
Smart Images

Figure CN117132964B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, specifically to technical fields such as computer vision, deep learning, and large models, and can be applied to scenarios such as autonomous driving and bird's-eye view perception. Specifically, it relates to a model training method, a point cloud encoding method, an object processing method, and an apparatus. Background Art
[0002] With the development of artificial intelligence technologies, three-dimensional object detection technologies and / or segmentation technologies have been continuously applied. For example, in scenarios such as autonomous driving and bird's-eye view perception, after collecting point cloud data through a lidar, the point cloud data can be encoded to obtain a point cloud encoding result, and based on this, a three-dimensional object detection task and / or segmentation task can be performed. Summary of the Invention
[0003] The present disclosure provides a model training method, a point cloud encoding method, an object processing method, and an apparatus.
[0004] According to one aspect of the present disclosure, a method for training a point cloud encoding model is provided, including:
[0005] Obtaining first point cloud data corresponding to a first training scenario;
[0006] Obtaining an image feature map obtained by processing the scene image of the first training scenario;
[0007] Encoding the first point cloud data through a point cloud encoding model to obtain a first point cloud feature map;
[0008] Training the point cloud encoding model based on the feature distribution difference between the first point cloud feature map and the image feature map to obtain a trained point cloud encoding model.
[0009] According to another aspect of the present disclosure, a method for training an object processing model is provided, including:
[0010] Obtaining second point cloud data corresponding to a second training scenario;
[0011] Encoding the second point cloud data through a target encoding model to obtain a second point cloud feature map; wherein, the target encoding model is a trained point cloud encoding model obtained through the point cloud encoding model training method;
[0012] Processing the second point cloud feature map through an object processing model to obtain a predicted processing result;
[0013] Training the object processing model based on the predicted processing result and the processing result label corresponding to the second point cloud data to obtain a trained object processing model.
[0014] According to another aspect of the present disclosure, there is provided a point cloud encoding method, including:
[0015] Obtaining a first point cloud to be encoded corresponding to a first target scene;
[0016] Encoding the first point cloud to be encoded through a target encoding model to obtain a first point cloud encoding result; wherein, the target encoding model is a trained point cloud encoding model obtained through a point cloud encoding model training method.
[0017] According to another aspect of the present disclosure, there is provided an object processing method, including:
[0018] Obtaining a second point cloud to be encoded corresponding to a second target scene;
[0019] Encoding the second point cloud to be encoded through a target encoding model to obtain a second point cloud encoding result; wherein, the target encoding model is a trained point cloud encoding model obtained through a point cloud encoding model training method;
[0020] Processing the second point cloud encoding result through a target processing model to obtain an object processing result; wherein, the target processing model is a trained object processing model obtained through an object processing model training method.
[0021] According to another aspect of the present disclosure, there is provided a point cloud encoding model training apparatus, including:
[0022] A first point cloud acquisition unit, configured to acquire first point cloud data corresponding to a first training scene;
[0023] A first image processing unit, configured to acquire an image feature map obtained by processing a scene image of the first training scene;
[0024] A first point cloud processing unit, configured to encode the first point cloud data through a point cloud encoding model to obtain a first point cloud feature map;
[0025] A first model training unit, configured to train the point cloud encoding model based on the feature distribution difference between the first point cloud feature map and the image feature map to obtain a trained point cloud encoding model.
[0026] According to another aspect of the present disclosure, there is provided an object processing model training apparatus, including:
[0027] A second point cloud acquisition unit, configured to acquire second point cloud data corresponding to a second training scene;
[0028] A second point cloud processing unit, configured to encode the second point cloud data through a target encoding model to obtain a second point cloud feature map; wherein, the target encoding model is a trained point cloud encoding model obtained through a point cloud encoding model training method;
[0029] A prediction processing result acquisition unit, configured to process the second point cloud feature map through an object processing model to obtain a prediction processing result;
[0030] A second model training unit, configured to train the object processing model based on the prediction processing result and a processing result label corresponding to the second point cloud data to obtain a trained object processing model.
[0031] According to another aspect of the present disclosure, there is provided a point cloud encoding device, including:
[0032] A first point cloud to be encoded acquisition unit, configured to acquire a first point cloud to be encoded corresponding to a first target scene;
[0033] A first point cloud to be encoded processing unit, configured to encode the first point cloud to be encoded through a target encoding model to obtain a first point cloud encoding result; wherein, the target encoding model is a trained point cloud encoding model obtained through the point cloud encoding model training method.
[0034] According to another aspect of the present disclosure, there is provided an object processing device, including:
[0035] A second point cloud to be encoded acquisition unit, configured to acquire a second point cloud to be encoded corresponding to a second target scene;
[0036] A second point cloud to be encoded processing unit, configured to encode the second point cloud to be encoded through a target encoding model to obtain a second point cloud encoding result; wherein, the target encoding model is a trained point cloud encoding model obtained through the method of any one of claims 1 to 8;
[0037] An object processing result acquisition unit, configured to process the second point cloud encoding result through a target processing model to obtain an object processing result; wherein, the target processing model is a trained object processing model obtained through the object processing model training method.
[0038] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0039] At least one processor;
[0040] A memory communicatively connected to the at least one processor;
[0041] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of any one of the embodiments of the present disclosure.
[0042] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are for causing the computer to execute the method according to any one of the embodiments of the present disclosure.
[0043] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, which implements the method according to any one of the embodiments of the present disclosure when executed by a processor.
[0044] Adopting the present disclosure can improve the reliability of the point cloud encoding result.
[0045] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0047] Figure 1 is a schematic flowchart of a method for training a point cloud encoding model provided by an embodiment of the present disclosure;
[0048] Figure 2 and Figure 3 is an auxiliary explanatory diagram of a method for training a point cloud encoding model provided by an embodiment of the present disclosure;
[0049] Figure 4 is a complete process auxiliary explanatory diagram of a method for training a point cloud encoding model provided by an embodiment of the present disclosure;
[0050] Figure 5 is a schematic diagram of a scenario of a method for training a point cloud encoding model provided by an embodiment of the present disclosure;
[0051] Figure 6 is a schematic flowchart of a method for training an object processing model provided by an embodiment of the present disclosure;
[0052] Figure 7 is an auxiliary explanatory diagram of a method for training an object processing model provided by an embodiment of the present disclosure;
[0053] Figure 8 is a schematic diagram of a scenario of a method for training an object processing model provided by an embodiment of the present disclosure;
[0054] Figure 9 is a schematic flowchart of a method for point cloud encoding provided by an embodiment of the present disclosure;
[0055] Figure 10An auxiliary explanatory diagram of a point cloud encoding method provided by an embodiment of the present disclosure;
[0056] Figure 11 A schematic diagram of a scenario of a point cloud encoding method provided by an embodiment of the present disclosure;
[0057] Figure 12 A schematic flow chart of an object processing method provided by an embodiment of the present disclosure;
[0058] Figure 13 An auxiliary explanatory diagram of an object processing method provided by an embodiment of the present disclosure;
[0059] Figure 14 A schematic diagram of a scenario of an object processing method provided by an embodiment of the present disclosure;
[0060] Figure 15 A schematic structural block diagram of a point cloud encoding model training device provided by an embodiment of the present disclosure;
[0061] Figure 16 A schematic structural block diagram of an object processing model training device provided by an embodiment of the present disclosure;
[0062] Figure 17 A schematic structural block diagram of a point cloud encoding device provided by an embodiment of the present disclosure;
[0063] Figure 18 A schematic structural block diagram of an object processing device provided by an embodiment of the present disclosure;
[0064] Figure 19 A schematic structural block diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0065] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0066] As described in the background art, in scenarios such as autonomous driving and Bird's Eye View (BEV) perception, after acquiring point cloud data through a lidar, the point cloud data can be encoded to obtain a point cloud encoding result, and based on this, a three-dimensional object detection task and / or segmentation task can be performed. However, through research by the inventors, it is found that since point cloud data usually only carries a small amount of feature information, for example, spatial position information and reflectivity information, and lacks rich texture information similar to that of images, therefore, when encoding point cloud data to obtain a point cloud encoding result, the feature information that can be relied on is less, and ultimately, it will affect the feature expression ability of the point cloud encoding result, that is, affect the reliability of the point cloud encoding result.
[0067] Based on the above research, embodiments of the present disclosure provide a method for training a point cloud encoding model, which can be applied to an electronic device. Hereinafter, with reference to Figure 1 the following flow schematic diagram, a method for training a point cloud encoding model provided by embodiments of the present disclosure will be described. It should be noted that although the logical order is shown in the flow schematic diagram, in some cases, the steps shown or described may also be executed in other orders.
[0068] Step S101, obtain first point cloud data corresponding to a first training scenario;
[0069] Step S102, obtain an image feature map obtained by processing a scene image of the first training scenario;
[0070] Step S103, encode the first point cloud data through a point cloud encoding model to obtain a first point cloud feature map;
[0071] Step S104, train the point cloud encoding model based on the feature distribution difference between the first point cloud feature map and the image feature map to obtain a trained point cloud encoding model.
[0072] Among them, the first training scenario can be any scenario including multiple three-dimensional objects, and the multiple three-dimensional objects can include objects such as buildings, road traffic facilities, motor vehicles, non-motor vehicles, and pedestrians.
[0073] Among them, the first point cloud data can be acquired through a lidar, which includes a plurality of scattered spatial points in three-dimensional space, and each spatial point has corresponding position information and reflectivity information; the scene image can be acquired through a camera, which belongs to a Red Green Blue (RGB) image and has rich texture information.
[0074] In the embodiments of the present disclosure, after obtaining a scene image, the scene image can be processed to obtain an image feature map corresponding to the scene image. In a specific example, processing the scene image aims to extract the features of each pixel point in the scene image, that is, high-dimensional feature expressions. At the same time, each pixel point in the scene image is assigned a pixel category to obtain an image feature map corresponding to the scene image. Based on this, it can be understood that in the embodiments of the present disclosure, each pixel point in the image feature map can carry its own feature and has a corresponding pixel category.
[0075] In addition, in the embodiments of the present disclosure, while obtaining the image feature map, the first point cloud data can be encoded by a point cloud encoding model to obtain a first point cloud feature map corresponding to the first point cloud data. In a specific example, encoding the first point cloud data aims to learn the spatial encoding of each spatial point in the first point cloud data, obtain the features of all spatial points, and then aggregate the features of all spatial points into a global point cloud feature as the first point cloud feature map corresponding to the first point cloud data. Based on this, in the embodiments of the present disclosure, the point cloud encoding model can be models such as PointNet, PointNet++, Second, or other models that can be used to encode point cloud data.
[0076] After obtaining the image feature map and the first point cloud feature map, the feature distribution difference between the first point cloud feature map and the image feature map can be obtained, and the point cloud encoding model can be trained based on this feature distribution difference, that is, the point cloud encoding model is guided to learn based on this feature distribution difference to obtain a trained point cloud encoding model. Among them, the feature distribution difference is used to characterize the difference in the feature distribution between the first point cloud feature map and the image feature map, and the purpose of training the cloud encoding model is to minimize the feature distribution difference.
[0077] Please combine Figure 2, by using the point cloud encoding model training method provided in the embodiments of the present disclosure, first point cloud data corresponding to a first training scenario can be obtained; image feature maps obtained by processing the scene image of the first training scenario can be obtained; first point cloud feature maps can be obtained by encoding the first point cloud data through the point cloud encoding model; the point cloud encoding model can be trained based on the feature distribution difference between the first point cloud feature maps and the image feature maps to obtain a trained point cloud encoding model. Among them, the scene image of the first training scenario has rich texture information. Therefore, the image feature maps obtained by processing the scene image have strong feature expression capabilities. Then, when encoding the first point cloud data through the point cloud encoding model to obtain first point cloud feature maps, obtaining the feature distribution difference between the first point cloud feature maps and the image feature maps, and training the point cloud encoding model based on this feature distribution difference (that is, guiding the point cloud encoding model to learn based on this feature distribution difference) to obtain a trained point cloud encoding model, in the application stage afterwards, when obtaining the to-be-encoded point cloud corresponding to the target scenario and encoding the to-be-encoded point cloud through the trained point cloud encoding model to obtain a point cloud encoding result, the defect that the to-be-encoded point cloud lacks rich texture information can be made up for, so as to improve the feature expression capability of the point cloud encoding result, that is, improve the reliability of the point cloud encoding result.
[0078] As described above, in the embodiments of the present disclosure, each pixel point in the image feature map can carry its own features and has a corresponding pixel category. Among them, the pixel category may not specify a specific semantic category, but is only used for the distinction of pixel categories.
[0079] Please combine Figure 3 , assume that there are 16 pixel points in the image feature map. Among them, pixel point A1, pixel point A2, and pixel point A3 belong to the same pixel category, specifically pixel category I, but pixel category I does not specify a specific semantic category; pixel point B1, pixel point B2, pixel point B3, and pixel point B4 belong to the same pixel category, specifically pixel category II, but pixel category II does not specify a specific semantic category; pixel point C1, pixel point C2, and pixel point C3 belong to the same pixel category, specifically pixel category III, but pixel category III does not specify a specific semantic category; pixel point D1, pixel point D2, and pixel point D3 belong to the same pixel category, specifically pixel category IV, but pixel category IV does not specify a specific semantic category; pixel point E1, pixel point E2, and pixel point E3 belong to the same pixel category, specifically pixel category V, but pixel category V does not specify a specific semantic category.
[0080] Based on this, it can be understood that in the embodiments of the present disclosure, the image feature map can actually be divided into multiple image feature regions, and all the pixel points in each image feature region belong to the same pixel category. Therefore, for each of the multiple image feature regions, the region category of the image feature region can also be defined by the pixel categories of all the pixel points in the image feature region.
[0081] Please refer to Figure 3 , the image feature map is actually divided into 5 image feature regions. Among them, the first image feature region 301 includes pixel points A1, A2, and A3. Therefore, the region category of the first image feature region 301 can be defined as region category I; the second image feature region 302 includes pixel points B1, B2, B3, and B4. Therefore, the region category of the second image feature region 302 can be defined as region category II; the third image feature region 303 includes pixel points C1, C2, and C3. Therefore, the region category of the third image feature region 303 can be defined as region category III; the fourth image feature region 304 includes pixel points D1, D2, and D3. Therefore, the region category of the fourth image feature region 304 can be defined as region category IV; the fifth image feature region 305 includes pixel points E1, E2, and E3. Therefore, the region category of the fifth image feature region 305 can be defined as region category V.
[0082] To achieve the above processing results, in some optional embodiments, "obtaining the image feature map obtained by processing the scene image of the first training scene" may include the following steps:
[0083] Performing visual segmentation on the scene image of the first training scene through an image segmentation model to obtain an image feature map; wherein, the image feature map includes multiple image feature regions.
[0084] Among them, the image segmentation model is pre-trained and has strong image segmentation capabilities.
[0085] In the embodiments of the present disclosure, the image segmentation model can be a model such as the "Segment Anything Model" (SAM), the "Segment Everything In Context" (SegGPT), the "Segment Everything Everywhere All At Once" (SEEM), etc.
[0086] After obtaining the scene image of the first training scene, the scene image can be directly input into the image segmentation model, and the output of the image segmentation model can be obtained as the image feature map corresponding to the scene image. Each pixel point in the image feature map can carry its own feature and has a corresponding pixel category. Therefore, the image feature map can be regarded as multiple image feature regions, and all pixel points in each image feature region belong to the same pixel category, that is, each image feature region corresponds to a region category.
[0087] Through the above steps, in the embodiments of the present disclosure, the scene image of the first training scene can be directly visually segmented by the image segmentation model to obtain the image feature map. Since the image segmentation model is pre-trained and has strong image segmentation ability, the segmentation accuracy of multiple image feature regions in the image feature map can be improved, and at the same time, the acquisition efficiency of the image feature map can be improved.
[0088] Based on the above processing results, in some alternative embodiments, "training the point cloud encoding model based on the feature distribution difference between the first point cloud feature map and the image feature map to obtain the trained point cloud encoding model" may include the following steps:
[0089] Obtain the correspondence between the first point cloud data and the scene image;
[0090] Segment the first point cloud feature map according to the correspondence to obtain multiple point cloud feature regions; wherein, the multiple point cloud feature regions correspond one-to-one to the multiple image feature regions;
[0091] Train the point cloud encoding model based on the feature distribution difference between the multiple point cloud feature regions and the multiple image feature regions to obtain the trained point cloud encoding model.
[0092] Among them, the correspondence between the first point cloud data and the scene image can be used to represent the pixel point corresponding to each spatial point in the first point cloud data in the scene image.
[0093] In the embodiments of the present disclosure, by segmenting the first point cloud feature map according to the correspondence, multiple point cloud feature regions can be obtained and the multiple point cloud feature regions correspond one-to-one to the multiple image feature regions. Therefore, the point cloud feature region and the image feature region with the correspondence belong to the same type of region pair, that is, they correspond to the same region category.
[0094] Please combine Figure 3, the first point cloud feature map can be segmented according to the corresponding relationship to obtain a first point cloud feature region 301', a second point cloud feature region 302', a third point cloud feature region 303', a fourth point cloud feature region 304', and a fifth point cloud feature region 305'. Among them, the first point cloud feature region 301' corresponds to the first image feature region 301 in the image feature map and its corresponding region category is I, the second point cloud feature region 302' corresponds to the second image feature region 302 in the image feature map and its corresponding region category is II, the third point cloud feature region 303' corresponds to the third image feature region 303 in the image feature map and its corresponding region category is III, the fourth point cloud feature region 304' corresponds to the fourth image feature region 304 in the image feature map and its corresponding region category is IV, and the fifth point cloud feature region 305' corresponds to the fifth image feature region 305 in the image feature map and its corresponding region category is V.
[0095] After obtaining multiple point cloud feature regions, the point cloud encoding model can be trained based on the feature distribution differences between the multiple point cloud feature regions and the multiple image feature regions to obtain a trained point cloud encoding model. In a specific example, for each point cloud feature region among the multiple point cloud feature regions, the feature distribution difference between it and each image feature region among the multiple image feature regions can be obtained respectively, and then the point cloud encoding model can be trained accordingly to obtain a trained point cloud encoding model.
[0096] Through the above steps, in the embodiments of the present disclosure, the first point cloud feature map can be segmented into multiple point cloud feature regions, and the multiple point cloud feature regions correspond one by one to the multiple image feature regions. Therefore, for each point cloud feature region among the multiple point cloud feature regions, the feature distribution difference between it and all image feature regions can be obtained respectively, and then the point cloud encoding model can be trained accordingly to obtain a trained point cloud encoding model. This is equivalent to reducing the calculation region granularity of the feature distribution difference while maintaining the overall calculation range of the original feature region difference, thereby improving the accuracy of the feature distribution difference and improving the training effect of the point cloud encoding model.
[0097] In some optional embodiments, "obtaining the corresponding relationship between the first point cloud data and the scene image" may include the following steps:
[0098] Obtain the external parameters between the lidar and the camera; where the lidar is the acquisition device for collecting the first point cloud data, and the camera is the acquisition device for collecting the scene image;
[0099] Obtain the internal parameters of the camera;
[0100] Based on the external parameters and the internal parameters, obtain the corresponding relationship between the first point cloud data and the scene image.
[0101] Among them, the external parameters are used to characterize the conversion relationship from the lidar coordinate system to the camera coordinate system. The lidar coordinate system is the coordinate system of the lidar, and the camera coordinate system is the coordinate system of the camera. In a specific example, constraints can be constructed by using the three-dimensional spatial points measured by the lidar and the three-dimensional coordinates of the calibration board measured by the camera, so as to realize the calibration of the external parameters. In another specific example, constraints can be constructed by using the three-dimensional spatial points measured by the lidar and the two-dimensional features (including point features, line segment features, etc.) of the corresponding images collected by the camera, so as to realize the calibration of the external parameters. The embodiments of the present disclosure will not elaborate on this.
[0102] Among them, the internal parameters are parameters related to the characteristics of the camera itself. For example, parameters such as the focal length and pixel size of the camera.
[0103] After obtaining the external parameters between the lidar and the camera and the internal parameters of the camera, first, based on the external parameters, the first point cloud data can be transformed from the lidar coordinate system to the camera coordinate system. Then, based on the internal parameters, the first point cloud data that has been transformed to the camera coordinate system can be projected onto the scene image, so as to determine the correspondence between the first point cloud data and the scene image, that is, to determine the pixel points corresponding to each spatial point in the first point cloud data in the scene image.
[0104] Through the above steps, in the embodiments of the present disclosure, the external parameters between the lidar and the camera and the internal parameters of the camera can be obtained, and directly based on the external parameters and the internal parameters, the correspondence between the first point cloud data and the scene image can be obtained. Since the external parameters and the internal parameters are fixed parameters of the lidar and the camera itself and have invariance, therefore, based on the external parameters and the internal parameters, obtaining the correspondence between the first point cloud data and the scene image can ensure the accuracy of the correspondence, thereby improving the segmentation accuracy of the first point cloud feature map.
[0105] In some optional embodiments, "training the point cloud encoding model based on the feature distribution differences between multiple point cloud feature regions and multiple image feature regions to obtain a trained point cloud encoding model" may include the following steps:
[0106] Each point cloud feature region and each image feature region are respectively used as the target region, and the feature distribution of the target region is calculated to obtain a plurality of first feature distributions and a plurality of second feature distributions. Among them, for each point cloud feature region, when the point cloud feature region is used as the target region, the obtained feature distribution is the first feature distribution, and for each image feature region, when the image feature region is used as the target region, the obtained feature distribution is the second feature distribution;
[0107] Calculate the loss between multiple first feature distributions and multiple second feature distributions as the feature distribution difference between multiple point cloud feature regions and multiple image feature regions;
[0108] Train the point cloud encoding model based on the feature distribution difference to obtain a trained point cloud encoding model.
[0109] Among them, the feature distribution of the target region is used to characterize the feature distribution on the target region. When the target region is a point cloud feature region, its feature distribution is defined as the first feature distribution, specifically used to characterize the feature distribution of spatial points on the point cloud feature region; when the target region is an image feature region, its feature distribution is defined as the second feature distribution, specifically used to characterize the feature distribution of pixel points on the image feature region.
[0110] Please combine Figure 3 , the image feature map includes multiple image feature regions, namely the first image feature region 301, the second image feature region 302, the third image feature region 303, the fourth image feature region 304, and the fifth image feature region 305; correspondingly, the first point cloud feature map is segmented into the first point cloud feature region 301', the second point cloud feature region 302', the third point cloud feature region 303', the fourth point cloud feature region 304', and the fifth point cloud feature region 305'.
[0111] After that, the first point cloud feature region 301', the second point cloud feature region 302', the third point cloud feature region 303', the fourth point cloud feature region 304', the fifth point cloud feature region 305', the first image feature region 301, the second image feature region 302, the third image feature region 303, the fourth image feature region 304, and the fifth image feature region 305 can be used as target regions respectively to calculate the feature distribution of the target regions. In this process, when the first point cloud feature region 301' is used as the target region, the obtained feature distribution is the first feature distribution, which can be specifically defined as the first feature distribution I, and so on. The first feature distribution II corresponding to the second point cloud feature region 302', the first feature distribution III corresponding to the third point cloud feature region 303', the first feature distribution IV corresponding to the fourth point cloud feature region 304', and the first feature distribution V corresponding to the fifth point cloud feature region 305' can be obtained; similarly, when the first image feature region 301 is used as the target region, the obtained feature distribution is the second feature distribution, which can be specifically defined as the second feature distribution I, and so on. The second feature distribution II corresponding to the second image feature region 302, the second feature distribution III corresponding to the third image feature region 303, the second feature distribution IV corresponding to the fourth image feature region 304, and the second feature distribution V corresponding to the fifth image feature region 305 can be obtained.
[0112] After obtaining multiple first feature distributions and multiple second feature distributions, the loss between the multiple first feature distributions and the multiple second feature distributions can be calculated as the feature distribution difference between the multiple point cloud feature regions and the multiple image feature regions, and the point cloud encoding model can be trained based on the feature distribution difference to obtain a trained point cloud encoding model. In a specific example, for each first feature distribution among the multiple first feature distributions, the loss between it and each second feature distribution among the multiple second feature distributions can be separately obtained to obtain the feature distribution difference between the first feature distribution and the multiple second feature distributions, and then the point cloud encoding model can be trained accordingly to obtain a trained point cloud encoding model.
[0113] Through the above steps, in the embodiments of the present disclosure, each point cloud feature region and each image feature region can be used as target regions, the feature distribution of the target regions can be calculated to obtain multiple first feature distributions and multiple second feature distributions, and the loss between the multiple first feature distributions and the multiple second feature distributions can be calculated as the feature distribution difference between the multiple point cloud feature regions and the multiple image feature regions, and then the point cloud encoding model can be trained based on the feature distribution difference to obtain a trained point cloud encoding model. That is to say, in the embodiments of the present disclosure, the first feature distribution of each point cloud feature region and the second feature distribution of each image feature region are calculated separately, with high accuracy. Therefore, when calculating the loss between the multiple first feature distributions and the multiple second feature distributions as the feature distribution difference between the multiple point cloud feature regions and the multiple image feature regions, the accuracy of the feature distribution difference can be further improved to improve the training effect of the point cloud encoding model.
[0114] In some alternative embodiments, "calculating the feature distribution of the target region" may include the following steps:
[0115] Perform pooling processing on the target region to obtain a region pooling result;
[0116] Calculate the similarity between the target region and the region pooling result as the feature distribution of the target region.
[0117] Among them, the pooling processing can be average pooling processing or max pooling processing, and the embodiments of the present disclosure do not make specific limitations on this. In addition, it can be understood that in the embodiments of the present disclosure, all processing on the target region can be understood as processing on the features of all points in the target region.
[0118] After obtaining the region pooling result of the target region, the similarity between the target region and the region pooling result can be calculated as the feature distribution of the target region. In a specific example, the cosine similarity between the target region and the region pooling result can be calculated as the feature distribution of the target region.
[0119] Taking the pooling process as the max-pooling process as an example, the step of "performing a pooling process on the target region to obtain a regional pooling result" can be characterized as:
[0120]
[0121] Among them, is used to represent the target region, and maxpool is used to represent performing a max-pooling process on the target region. is used to represent the regional pooling result corresponding to the target region.
[0122] Furthermore, in the embodiments of the present disclosure, the step of "calculating the similarity between the target region and the regional pooling result as the feature distribution of the target region" can be characterized as:
[0123]
[0124] Among them, is used to represent the target region. is used to represent the regional pooling result corresponding to the target region, and cos_sim is used to represent calculating the similarity between the target region and the regional pooling result. is used to represent the feature distribution of the target region.
[0125] Through the above steps, in the embodiments of the present disclosure, a pooling process can be performed on the target region to obtain a regional pooling result, and the similarity between the target region and the regional pooling result can be calculated as the feature distribution of the target region. In this process, the involved calculation logic is simple, which can improve the acquisition efficiency of the feature distribution of the target region. At the same time, when calculating the cosine similarity between the target region and the regional pooling result as the feature distribution of the target region, since the cosine similarity has good performance in the relevant data processing of high-dimensional feature expressions, the accuracy of the feature distribution of the target region can also be improved.
[0126] In some alternative embodiments, "performing a pooling process on the target region to obtain a regional pooling result" may include the following steps:
[0127] Performing a max-pooling process on the target region to obtain a regional pooling result.
[0128] That is, in the embodiments of the present disclosure, it is preferable to perform a max-pooling process on the target region to obtain a regional pooling result, and it is secondary to perform an average-pooling process on the target region to obtain a regional pooling result.
[0129] Through the above steps, in the embodiments of the present disclosure, max pooling processing can be performed on the target region to obtain a region pooling result. Since max pooling processing has less computational complexity compared to other pooling processing methods (e.g., average pooling processing), the acquisition efficiency of the feature distribution of the target region can be further improved.
[0130] In some alternative embodiments, "calculating the loss between a plurality of first feature distributions and a plurality of second feature distributions" may include the following steps:
[0131] Calculate the loss between a plurality of first feature distributions and a plurality of second feature distributions through a pre-constructed loss function; wherein, the construction principle of the loss function includes minimizing the loss between pairs of like features and maximizing the loss between pairs of unlike features. Pairs of like features include a first feature distribution and a second feature distribution with a corresponding relationship, and pairs of unlike features include a first feature distribution and a second feature distribution without a corresponding relationship.
[0132] Wherein, the loss function may be a cross-entropy loss function, that is, a Softmax loss function.
[0133] As described above, in the embodiments of the present disclosure, the construction principle of the loss function includes minimizing the loss between pairs of like features and maximizing the loss between pairs of unlike features. Pairs of like features include a first feature distribution and a second feature distribution with a corresponding relationship, and pairs of unlike features include a first feature distribution and a second feature distribution without a corresponding relationship. Among them, in the first feature distribution and the second feature distribution with a corresponding relationship, the point cloud feature region corresponding to the first feature distribution and the image feature region corresponding to the second feature distribution belong to a pair of like regions, that is, they correspond to the same region category; correspondingly, in the first feature distribution and the second feature distribution without a corresponding relationship, the point cloud feature region corresponding to the first feature distribution and the image feature region corresponding to the second feature distribution belong to a pair of unlike regions, that is, they correspond to different region categories.
[0134] Please refer to Figure 3 , the image feature map includes multiple image feature regions, namely a first image feature region 301, a second image feature region 302, a third image feature region 303, a fourth image feature region 304, and a fifth image feature region 305; correspondingly, the first point cloud feature map is segmented into a first point cloud feature region 301', a second point cloud feature region 302', a third point cloud feature region 303', a fourth point cloud feature region 304', and a fifth point cloud feature region 305'.
[0135] Among them, the first point cloud feature region 301' and the first image feature region 301 belong to the same type of region pair, and the first point cloud feature region 301' corresponds to the first feature distribution I, and the first image feature region 301 corresponds to the second feature distribution I; the second point cloud feature region 302' and the second image feature region 302 belong to the same type of region pair, and the second point cloud feature region 302' corresponds to the first feature distribution II, and the second image feature region 302 corresponds to the second feature distribution II; the third point cloud feature region 303' and the third image feature region 303 belong to the same type of region pair, and the third point cloud feature region 303' corresponds to the first feature distribution III, and the third image feature region 303 corresponds to the second feature distribution III; the fourth point cloud feature region 304' and the fourth image feature region 304 belong to the same type of region pair, and the fourth point cloud feature region 304' corresponds to the first feature distribution IV, and the fourth image feature region 304 corresponds to the second feature distribution IV; the fifth point cloud feature region 305' and the fifth image feature region 305 belong to the same type of region pair, and the fifth point cloud feature region 305' corresponds to the first feature distribution V, and the fifth image feature region 305 corresponds to the second feature distribution V.
[0136] Then, the first feature distribution I and the second feature distribution I belong to the same type of feature pair, the first feature distribution II and the second feature distribution II belong to the same type of feature pair, the first feature distribution III and the second feature distribution III belong to the same type of feature pair, the first feature distribution IV and the second feature distribution IV belong to the same type of feature pair, and the first feature distribution V and the second feature distribution V belong to the same type of feature pair. Other feature pairs other than these belong to different type of feature pairs. For example, the first feature distribution I and other second feature distributions other than the second feature distribution I belong to different type of feature pairs, and the second feature distribution I and other first feature distributions other than the first feature distribution I belong to different type of feature pairs.
[0137] Based on the above requirements, in the embodiments of the present disclosure, the step of "calculating the loss between multiple first feature distributions and multiple second feature distributions through a pre-constructed loss function" can be characterized as:
[0138]
[0139] where M is the total number of image feature regions in the image feature map, and is also equal to the total number of point cloud feature regions in the first point cloud feature map. is used to represent the second feature distribution corresponding to the i-th image feature region among the M image feature regions. is used to represent the first feature distribution corresponding to the i-th point cloud feature region among the M point cloud feature regions. is used to represent the first feature distribution corresponding to the j-th point cloud feature region among the M point cloud feature regions, Loss clUsed to characterize the loss between multiple (here, characterized as M) first feature distributions and multiple (here, characterized as M) second feature distributions.
[0140] Through the above steps, in the embodiments of the present disclosure, the loss between multiple first feature distributions and multiple second feature distributions can be calculated through a pre-constructed loss function. Since the construction principle of the loss function includes minimizing the loss between pairs of like features and maximizing the loss between pairs of unlike features, pairs of like features include corresponding first and second feature distributions, and pairs of unlike features include first and second feature distributions without a corresponding relationship. Therefore, during the training process of the point cloud encoding model, it can not only play a positive guiding role in the similarity of pairs of like features, but also play a reverse guiding role in the similarity of pairs of unlike features, so as to optimize the learning guiding effect on the point cloud encoding model, thereby further improving the training effect of the point cloud encoding model.
[0141] In addition, it should be noted that in the embodiments of the present disclosure, when calculating the loss between multiple first feature distributions and multiple second feature distributions as the feature distribution difference between multiple point cloud feature regions and multiple image feature regions, if the feature distribution difference satisfies the first convergence condition, the point cloud encoding model at this time is used as the trained point cloud encoding model; if the feature distribution difference does not satisfy the first convergence condition, the point cloud encoding model is trained based on the feature distribution difference (that is, the parameters of the point cloud encoding model are updated), and then enter the next round of training, that is, obtain a new feature distribution difference, until the new feature distribution difference satisfies the first convergence condition, and a trained point cloud encoding model is obtained. The new feature distribution difference that satisfies the first convergence condition can be defined as the target loss. Among them, the first convergence condition can be set according to actual application requirements, and the embodiments of the present disclosure do not make specific limitations on this.
[0142] Next, in combination with Figure 4 , the complete process of a point cloud encoding model training method provided by the embodiments of the present disclosure will be described.
[0143] (1) Obtain the scene image of the first training scene and the first point cloud data corresponding to the first training scene.
[0144] (2) Obtain the external parameters between the lidar and the camera; where the lidar is the acquisition device for collecting the first point cloud data, and the camera is the acquisition device for collecting the scene image; obtain the internal parameters of the camera; based on the external parameters and the internal parameters, obtain the correspondence between the first point cloud data and the scene image.
[0145] (3) Through the image segmentation model, perform visual segmentation on the scene image of the first training scene to obtain an image feature map; where the image feature map includes multiple image feature regions.
[0146] That is, after obtaining the scene image of the first training scene, the scene image can be directly input into the image segmentation model, and the output of the image segmentation model is obtained as the image feature map corresponding to the scene image. Each pixel point in the image feature map can carry its own feature and has a corresponding pixel category. Therefore, the image feature map can be regarded as multiple image feature regions, and all pixel points in each image feature region belong to the same pixel category. That is, each image feature region corresponds to a region category.
[0147] (4) Encode the first point cloud data through the point cloud encoding model to obtain the first point cloud feature map.
[0148] (5) Segment the first point cloud feature map according to the correspondence between the first point cloud data and the scene image to obtain multiple point cloud feature regions; among them, the multiple point cloud feature regions correspond one-to-one to the multiple image feature regions.
[0149] (6) Respectively use each point cloud feature region and each image feature region as the target region, and calculate the feature distribution of the target region to obtain multiple first feature distributions and multiple second feature distributions; among them, for each point cloud feature region, when the point cloud feature region is used as the target region, the obtained feature distribution is the first feature distribution, and for each image feature region, when the image feature region is used as the target region, the obtained feature distribution is the second feature distribution.
[0150] Among them, calculating the feature distribution of the target region includes: performing pooling processing on the target region to obtain the region pooling result; calculating the similarity between the target region and the region pooling result as the feature distribution of the target region.
[0151] (7) Through the pre-constructed loss function, calculate the loss between the multiple first feature distributions and the multiple second feature distributions as the feature distribution difference between the multiple point cloud feature regions and the multiple image feature regions; among them, the construction principle of the loss function includes minimizing the loss between similar feature pairs and maximizing the loss between dissimilar feature pairs. Similar feature pairs include the first feature distribution and the second feature distribution with a corresponding relationship, and dissimilar feature pairs include the first feature distribution and the second feature distribution without a corresponding relationship.
[0152] (8) Based on the feature distribution difference between the multiple point cloud feature regions and the multiple image feature regions, train the point cloud encoding model to obtain the trained point cloud encoding model.
[0153] Please refer to Figure 5 , which is a schematic diagram of the scene of a method for training a point cloud encoding model provided by an embodiment of the present disclosure.
[0154] As described above, the point cloud encoding model training method provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as servers, workstations, mainframe computers, desktop computers, laptop computers, or other suitable computers.
[0155] The electronic device can be used to:
[0156] Obtain first point cloud data corresponding to a first training scenario;
[0157] Obtain an image feature map obtained by processing a scene image of the first training scenario;
[0158] Encode the first point cloud data through a point cloud encoding model to obtain a first point cloud feature map;
[0159] Train the point cloud encoding model based on the feature distribution difference between the first point cloud feature map and the image feature map to obtain a trained point cloud encoding model.
[0160] Among them, the first training scenario can be any scenario including multiple three-dimensional objects, and the multiple three-dimensional objects can include objects such as buildings, road traffic facilities, motor vehicles, non-motor vehicles, and pedestrians.
[0161] Among them, the first point cloud data can be collected by a lidar, which includes a plurality of scattered spatial points in a three-dimensional space, and each spatial point has corresponding position information and reflectivity information; the scene image can be collected by a camera, which belongs to an RGB image and has rich texture information.
[0162] It should be noted that in the embodiments of the present disclosure, Figure 5 The shown scene schematic diagram is only illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 5 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.
[0163] The embodiments of the present disclosure provide an object processing model training method, which can be applied to an electronic device. Hereinafter, a method for training an object processing model provided by the embodiments of the present disclosure will be described in conjunction with Figure 6 the shown process schematic diagram. It should be noted that although the logical order is shown in the process schematic diagram, in some cases, the steps shown or described can also be executed in other orders.
[0164] Step S601, obtain second point cloud data corresponding to a second training scenario;
[0165] Step S602: Encode the second point cloud data through the target encoding model to obtain the second point cloud feature map. The target encoding model is a trained point cloud encoding model obtained through the point cloud encoding model training method.
[0166] Step S603: Process the second point cloud feature map through the object processing model to obtain the predicted processing result.
[0167] Step S604: Based on the predicted processing result and the processing result label corresponding to the second point cloud data, train the object processing model to obtain the trained object processing model.
[0168] Among them, the second training scenario can be any scenario including multiple three-dimensional objects, and the multiple three-dimensional objects can include objects such as buildings, road traffic facilities, motor vehicles, non-motor vehicles, and pedestrians. The second point cloud data can be collected by a lidar, which includes multiple spatial points scattered in three-dimensional space, and each spatial point has corresponding position information and reflectivity information.
[0169] In the embodiments of the present disclosure, after obtaining the second point cloud data corresponding to the second training scenario, the second point cloud data can be encoded through the target encoding model to obtain the second point cloud feature map corresponding to the second point cloud data. The target encoding model is a trained point cloud encoding model obtained through the point cloud encoding model training method. Therefore, in combination with the relevant descriptions in the embodiments of the point cloud encoding model training method, the target encoding model can be models such as PointNet, PointNet++, Second, or other models that can be used to encode point cloud data.
[0170] In addition, in the embodiments of the present disclosure, after obtaining the second point cloud feature map, the second point cloud feature map can be processed through the object processing model to obtain the predicted processing result for the second point cloud data. In a specific example, the object processing model is a three-dimensional object detection model, and the predicted processing result is the three-dimensional object detection result for the second point cloud data. In another specific example, the object processing model is a three-dimensional object segmentation model, and the predicted processing result is the three-dimensional object segmentation result for the second point cloud data. That is, in the embodiments of the present disclosure, the object processing model can be a three-dimensional object detection model or a three-dimensional object segmentation model to improve the applicable range of the object processing model training method. Among them, the three-dimensional object detection model can be models such as PointRCNN, Second, etc.; the three-dimensional object segmentation model can be the Mask3D model.
[0171] In addition, in the embodiments of the present disclosure, while obtaining the prediction processing result, a processing result label corresponding to the second point cloud data can be obtained, and based on the prediction processing result and the processing result label corresponding to the second point cloud data, the object processing model is trained to obtain a trained object processing model. Among them, when the object processing model is a three-dimensional object detection model, the processing result label corresponding to the second point cloud data can be an object detection label for the second point cloud data, which can include multiple object detection boxes; when the object processing model is a three-dimensional object segmentation model, the processing result label corresponding to the second point cloud data can be an object segmentation label for the second point cloud data, which can include the point category corresponding to each spatial point in the second point cloud data.
[0172] Please combine Figure 7 , and by using the object processing model training method provided in the embodiments of the present disclosure, the second point cloud data corresponding to the second training scenario can be obtained; the second point cloud data is encoded by the target encoding model to obtain a second point cloud feature map; the second point cloud feature map is processed by the object processing model to obtain a prediction processing result; the prediction processing result and the processing result label corresponding to the second point cloud data are used to train the object processing model to obtain a trained object processing model. Since the target encoding model is a trained point cloud encoding model obtained by the point cloud encoding model training method, the second point cloud feature map has strong feature expression ability, that is, has high reliability. Then, when the second point cloud feature map is processed by the object processing model to obtain a prediction processing result, and the prediction processing result and the processing result label corresponding to the second point cloud data are used to train the object processing model, the training effect of the object processing model can be improved, thereby improving the point cloud data processing ability of the trained object processing model.
[0173] In a specific example, when training the object processing model based on the prediction processing result and the processing result label corresponding to the second point cloud data, the task loss between the prediction processing result and the processing result label can be calculated. If the task loss meets the second convergence condition, the object processing model at this time is used as the trained object processing model; if the task loss does not meet the second convergence condition, the object processing model is trained based on the task loss (that is, the parameters of the object processing model are updated), and then enter the next round of training, that is, obtain a new task loss, until the new task loss meets the second convergence condition, and a trained object processing model is obtained. Among them, the second convergence condition can be set according to actual application requirements, and the embodiments of the present disclosure do not make specific limitations on this.
[0174] In another specific example, when training an object processing model based on the prediction processing result and the processing result label corresponding to the second point cloud data, after calculating the task loss between the prediction processing result and the processing result label, the sum of the task loss and the target loss described in the embodiment of the point cloud encoding model training method can also be obtained. This process can be represented as:
[0175] Loss = Loss cl ′ + Loss task
[0176] Where, Loss cl ′ is used to represent the target loss described in the embodiment of the point cloud encoding model training method, Loss task is used to represent the task loss between the prediction processing result and the processing result label, and Loss is used to represent the sum of the task loss and the target loss described in the embodiment of the point cloud encoding model training method.
[0177] If the sum of the losses meets the third convergence condition, the object processing model at this time is used as the trained object processing model; if the sum of the losses does not meet the third convergence condition, the object processing model is trained based on the sum of the losses (that is, the parameters of the object processing model are updated), and then enter the next round of training, that is, obtain a new sum of the losses, until the new sum of the losses meets the third convergence condition, and the trained object processing model is obtained. Among them, the third convergence condition can be set according to actual application requirements, and the embodiments of the present disclosure do not make specific limitations in this regard. In this example, at least partial compensation can be made to the target loss described in the embodiment of the point cloud encoding model training method through the task loss between the prediction processing result and the processing result label, so as to further improve the training effect of the object processing model, thereby improving the point cloud number processing ability of the trained object processing model.
[0178] In addition, it should be noted that in the embodiments of the present disclosure, when the object processing model is a three-dimensional object detection model, the task loss between the prediction processing result and the processing result label can be calculated through the cross-entropy loss function; when the object processing model is a three-dimensional object segmentation model, the task loss between the prediction processing result and the processing result label can be calculated through the cross-entropy loss function or the absolute value loss function (that is, the L1 loss function).
[0179] Please refer to Figure 8 , which is a schematic diagram of the scenario of a method for training an object processing model provided by the embodiments of the present disclosure.
[0180] As described above, the object processing model training method provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as servers, workstations, mainframe computers, desktop computers, laptop computers, or other suitable computers.
[0181] The electronic device can be used to:
[0182] Obtain second point cloud data corresponding to a second training scenario;
[0183] Encode the second point cloud data through a target encoding model to obtain a second point cloud feature map; wherein, the target encoding model is a trained point cloud encoding model obtained through a point cloud encoding model training method;
[0184] Process the second point cloud feature map through an object processing model to obtain a predicted processing result;
[0185] Train the object processing model based on the predicted processing result and the processing result label corresponding to the second point cloud data to obtain a trained object processing model.
[0186] Wherein, the second training scenario can be any scenario including multiple three-dimensional objects, and the multiple three-dimensional objects can include objects such as buildings, road traffic facilities, motor vehicles, non-motor vehicles, pedestrians, etc.; the second point cloud data can be collected by a lidar, and it includes a plurality of spatial points scattered in a three-dimensional space, and each spatial point has corresponding position information and reflectivity information.
[0187] It should be noted that in the embodiments of the present disclosure, Figure 8 The shown scenario schematic diagram is only illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 8 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.
[0188] The embodiments of the present disclosure provide a point cloud encoding method, which can be applied to an electronic device. Hereinafter, a point cloud encoding method provided by the embodiments of the present disclosure will be described in conjunction with Figure 9 the shown process schematic diagram. It should be noted that although the logical order is shown in the process schematic diagram, in some cases, the steps shown or described can also be executed in other orders.
[0189] Step S901, obtain a first point cloud to be encoded corresponding to a first target scenario;
[0190] Step S902, encode the first point cloud to be encoded through a target encoding model to obtain a first point cloud encoding result; wherein, the target encoding model is a trained point cloud encoding model obtained through a point cloud encoding model training method.
[0191] Among them, the first target scene can be any scene including multiple three-dimensional objects, and the multiple three-dimensional objects can include objects such as buildings, road traffic facilities, motor vehicles, non-motor vehicles, and pedestrians; the first point cloud to be encoded can be collected by a lidar, which includes a plurality of spatial points scattered in three-dimensional space, and each spatial point has corresponding position information and reflectivity information.
[0192] After obtaining the first point cloud to be encoded corresponding to the first target scene, the first point cloud to be encoded can be encoded by a target encoding model to obtain a first point cloud encoding result corresponding to the first point cloud to be encoded. Among them, the target encoding model is a trained point cloud encoding model obtained by a point cloud encoding model training method. Therefore, combining the relevant descriptions in the embodiments of the foregoing point cloud encoding model training method, the target encoding model can be models such as PointNet, PointNet++, Second, or other models that can be used to encode point cloud data.
[0193] Please combine Figure 10 , by using the point cloud encoding method provided in the embodiments of the present disclosure, the first point cloud to be encoded corresponding to the first target scene can be obtained; the target encoding model encodes the first point cloud to be encoded to obtain a first point cloud encoding result. Since the target encoding model is a trained point cloud encoding model obtained by a point cloud encoding model training method, the first point cloud encoding result has strong feature expression ability, that is, it has high reliability.
[0194] Please refer to Figure 11 , which is a schematic diagram of a scene of a point cloud encoding method provided in the embodiments of the present disclosure.
[0195] As mentioned above, the point cloud encoding method provided in the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as servers, workstations, mainframe computers, desktop computers, laptop computers, or other suitable computers. In addition, it should be noted that in the embodiments of the present disclosure, when the point cloud encoding method is applied to scenarios such as autonomous driving and BEV perception, the electronic device can also be an in-vehicle computer installed on an autonomous driving vehicle. The electronic device can be used for:
[0196] Obtain the first point cloud to be encoded corresponding to the first target scene;
[0197] Encode the first point cloud to be encoded by a target encoding model to obtain a first point cloud encoding result; among them, the target encoding model is a trained point cloud encoding model obtained by a point cloud encoding model training method.
[0198] Among them, the first target scenario can be any scenario including multiple three-dimensional objects, and the multiple three-dimensional objects can include objects such as buildings, road traffic facilities, motor vehicles, non-motor vehicles, and pedestrians; the first point cloud to be encoded can be collected by a lidar, which includes a plurality of spatial points scattered in a three-dimensional space, and each spatial point has corresponding position information and reflectivity information.
[0199] It should be noted that in the embodiments of the present disclosure, Figure 11 the shown scene schematic diagram is only illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 11 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.
[0200] The embodiments of the present disclosure provide an object processing method, which can be applied to an electronic device. Hereinafter, a method for processing an object provided by the embodiments of the present disclosure will be described in conjunction with Figure 12 the shown process schematic diagram. It should be noted that although the logical order is shown in the process schematic diagram, in some cases, the steps shown or described may also be executed in other orders.
[0201] Step S1201, obtain a second point cloud to be encoded corresponding to a second target scenario;
[0202] Step S1202, encode the second point cloud to be encoded through a target encoding model to obtain a second point cloud encoding result; wherein, the target encoding model is a trained point cloud encoding model obtained through a point cloud encoding model training method;
[0203] Step S1203, process the second point cloud encoding result through a target processing model to obtain an object processing result; wherein, the target processing model is a trained object processing model obtained through an object processing model training method.
[0204] Among them, the second target scenario can be any scenario including multiple three-dimensional objects, and the multiple three-dimensional objects can include objects such as buildings, road traffic facilities, motor vehicles, non-motor vehicles, and pedestrians; the second point cloud to be encoded can be collected by a lidar, which includes a plurality of spatial points scattered in a three-dimensional space, and each spatial point has corresponding position information and reflectivity information.
[0205] After obtaining the second point cloud to be encoded corresponding to the second target scene, the second point cloud to be encoded can be encoded by a target encoding model to obtain a second point cloud encoding result corresponding to the second point cloud to be encoded. Among them, the target encoding model is a trained point cloud encoding model obtained by a point cloud encoding model training method. Therefore, combining the relevant descriptions in the foregoing embodiments of the point cloud encoding model training method, the target encoding model can be models such as PointNet, PointNet++, Second, or other models that can be used to encode point cloud data.
[0206] After obtaining the second point cloud encoding result, the second point cloud encoding result can be processed by a target processing model to obtain an object processing result corresponding to the second point cloud to be encoded. Among them, the target processing model is a trained object processing model obtained by an object processing model training method. Combining the relevant descriptions in the foregoing embodiments of the object processing model training method, the object processing model can be a three-dimensional object detection model or a three-dimensional object segmentation model to improve the applicable range of the object processing model training method. Among them, the three-dimensional object detection model can be models such as PointRCNN, Second, etc.; the three-dimensional object segmentation model can be the Mask3D model.
[0207] Please combine Figure 13 , by using the object processing method provided in the embodiments of the present disclosure, the second point cloud to be encoded corresponding to the second target scene can be obtained; the target encoding model encodes the second point cloud to be encoded to obtain a second point cloud encoding result; the second point cloud encoding result is processed by the target processing model to obtain an object processing result. On the one hand, since the target encoding model is a trained point cloud encoding model obtained by a point cloud encoding model training method, the second point cloud encoding result has strong feature expression ability, that is, it has high reliability; on the other hand, since the target processing model is a trained object processing model obtained by an object processing model training method and has strong point cloud data processing ability, the reliability of the object processing result can be improved.
[0208] Please refer to Figure 14 , which is a schematic diagram of the scenario of an object processing method provided in the embodiments of the present disclosure.
[0209] As mentioned above, the object processing method provided in the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as servers, workstations, mainframe computers, desktop computers, laptop computers, or other suitable computers. In addition, it should be noted that in the embodiments of the present disclosure, when the point cloud encoding method is applied to scenarios such as autonomous driving and BEV perception, the electronic device can also be an in-vehicle computer installed on an autonomous driving vehicle.
[0210] An electronic device can be used for:
[0211] Obtaining a second point cloud to be encoded corresponding to a second target scenario;
[0212] Encoding the second point cloud to be encoded through a target encoding model to obtain a second point cloud encoding result; wherein, the target encoding model is a trained point cloud encoding model obtained through a point cloud encoding model training method;
[0213] Processing the second point cloud encoding result through a target processing model to obtain an object processing result; wherein, the target processing model is a trained object processing model obtained through an object processing model training method.
[0214] Wherein, the second target scenario can be any scenario including multiple three-dimensional objects, and the multiple three-dimensional objects can include objects such as buildings, road traffic facilities, motor vehicles, non-motor vehicles, pedestrians, etc.; the second point cloud to be encoded can be collected by a lidar, and it includes multiple spatial points scattered in a three-dimensional space, and each spatial point has corresponding position information and reflectivity information.
[0215] It should be noted that in the embodiments of the present disclosure, Figure 14 The shown scenario schematic diagram is only illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 14 The examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.
[0216] In order to better implement the point cloud encoding model training method, the embodiments of the present disclosure further provide a point cloud encoding model training device, which can be integrated in an electronic device. Hereinafter, with reference to Figure 15 The shown structural schematic diagram, a point cloud encoding model training device 1500 provided by the embodiments of the disclosure will be described.
[0217] The point cloud encoding model training device 1500 includes:
[0218] A first point cloud acquisition unit 1501, configured to acquire first point cloud data corresponding to a first training scenario;
[0219] A first image processing unit 1502, configured to acquire an image feature map obtained by processing the scene image of the first training scenario;
[0220] A first point cloud processing unit 1503, configured to encode the first point cloud data through a point cloud encoding model to obtain a first point cloud feature map;
[0221] A first model training unit 1504, configured to train the point cloud encoding model based on the feature distribution difference between the first point cloud feature map and the image feature map to obtain a trained point cloud encoding model.
[0222] In some alternative embodiments, the image feature map includes multiple image feature regions, and each image feature region corresponds to a region category; the first model training unit 1504 is configured to:
[0223] Obtain the correspondence between the first point cloud data and the scene image;
[0224] Segment the first point cloud feature map according to the correspondence to obtain multiple point cloud feature regions; wherein, the multiple point cloud feature regions correspond one-to-one to the multiple image feature regions;
[0225] Train the point cloud encoding model based on the feature distribution difference between the multiple point cloud feature regions and the multiple image feature regions to obtain a trained point cloud encoding model.
[0226] In some alternative embodiments, the first model training unit 1504 is configured to:
[0227] Respectively use each point cloud feature region and each image feature region as the target region, calculate the feature distribution of the target region to obtain multiple first feature distributions and multiple second feature distributions; wherein, for each point cloud feature region, when using the point cloud feature region as the target region, the obtained feature distribution is the first feature distribution, and for each image feature region, when using the image feature region as the target region, the obtained feature distribution is the second feature distribution;
[0228] Calculate the loss between the multiple first feature distributions and the multiple second feature distributions as the feature distribution difference between the multiple point cloud feature regions and the multiple image feature regions;
[0229] Train the point cloud encoding model based on the feature distribution difference to obtain a trained point cloud encoding model.
[0230] In some alternative embodiments, the first model training unit 1504 is configured to:
[0231] Perform pooling processing on the target region to obtain a region pooling result;
[0232] Calculate the similarity between the target region and the region pooling result as the feature distribution of the target region.
[0233] In some alternative embodiments, the first model training unit 1504 is configured to:
[0234] Perform max pooling processing on the target region to obtain a region pooling result.
[0235] In some alternative embodiments, the first model training unit 1504 is configured to:
[0236] Calculate the loss between multiple first feature distributions and multiple second feature distributions through a pre-constructed loss function; wherein, the construction principle of the loss function includes minimizing the loss between like feature pairs and maximizing the loss between unlike feature pairs. Like feature pairs include a first feature distribution and a second feature distribution with a corresponding relationship, and unlike feature pairs include a first feature distribution and a second feature distribution without a corresponding relationship.
[0237] In some alternative embodiments, the first model training unit 1504 is configured to:
[0238] Obtain the extrinsic parameters between the lidar and the camera; wherein, the lidar is the acquisition device for collecting the first point cloud data, and the camera is the acquisition device for collecting the scene image;
[0239] Obtain the intrinsic parameters of the camera;
[0240] Based on the extrinsic parameters and the intrinsic parameters, obtain the correspondence between the first point cloud data and the scene image.
[0241] In some alternative embodiments, the first image processing unit 1502:
[0242] Perform visual segmentation on the scene image of the first training scene through an image segmentation model to obtain an image feature map; wherein, the image feature map includes a plurality of image feature regions.
[0243] For the specific functions and examples of the units of the point cloud encoding model training device 1500 according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the foregoing embodiments of the point cloud encoding model training method, which will not be elaborated herein.
[0244] To better implement the point object processing model training method, the embodiments of the present disclosure further provide an object processing model training device, which may be integrated in an electronic device. Hereinafter, with reference to Figure 16 the structural schematic diagram shown, an object processing model training device 1600 provided by the embodiments of the present disclosure will be described.
[0245] The object processing model training device 1600 includes:
[0246] A second point cloud acquisition unit 1601, configured to acquire second point cloud data corresponding to a second training scene;
[0247] A second point cloud processing unit 1602, configured to encode the second point cloud data through a target encoding model to obtain a second point cloud feature map; wherein, the target encoding model is a trained point cloud encoding model obtained by the method of any one of claims 1 to 8;
[0248] A prediction processing result acquisition unit 1603, configured to process the second point cloud feature map through an object processing model to obtain a prediction processing result;
[0249] A second model training unit 1604, configured to train the object processing model based on the prediction processing result and a processing result label corresponding to the second point cloud data to obtain a trained object processing model.
[0250] In some alternative embodiments, the object processing model is a three-dimensional object detection model or a three-dimensional object segmentation model.
[0251] For the specific functions and examples of the units of the object processing model training apparatus 1600 according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the foregoing embodiments of the object processing model training method, which will not be elaborated herein.
[0252] To better implement the point cloud encoding method, the embodiments of the present disclosure further provide a point cloud encoding apparatus, which may be integrated in an electronic device. Hereinafter, with reference to Figure 17 the following structural schematic diagram, a point cloud encoding apparatus 1700 provided by the embodiments of the present disclosure will be described.
[0253] The point cloud encoding apparatus 1700 includes:
[0254] A first point cloud to be encoded acquisition unit 1701, configured to acquire a first point cloud to be encoded corresponding to a first target scene;
[0255] A first point cloud to be encoded processing unit 1702, configured to encode the first point cloud to be encoded through a target encoding model to obtain a first point cloud encoding result; wherein, the target encoding model is a trained point cloud encoding model obtained through the point cloud encoding model training method.
[0256] For the specific functions and examples of the units of the point cloud encoding apparatus 1700 according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the foregoing embodiments of the point cloud encoding method, which will not be elaborated herein.
[0257] To better implement the point object processing method, the embodiments of the present disclosure further provide an object processing apparatus, which may be integrated in an electronic device. Hereinafter, with reference to Figure 18 the following structural schematic diagram, an object processing apparatus 1800 provided by the embodiments of the present disclosure will be described.
[0258] The object processing apparatus 1800 includes:
[0259] A second point cloud to be encoded acquisition unit 1801, configured to acquire a second point cloud to be encoded corresponding to a second target scene;
[0260] The second point cloud to be encoded processing unit 1802 is configured to encode the second point cloud to be encoded through a target encoding model to obtain a second point cloud encoding result; wherein, the target encoding model is a trained point cloud encoding model obtained through a point cloud encoding model training method;
[0261] The object processing result acquisition unit 1803 is configured to process the second point cloud encoding result through a target processing model to obtain an object processing result; wherein, the target processing model is a trained object processing model obtained through an object processing model training method.
[0262] For the specific functions and examples of the units of the object processing device 1800 according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the foregoing embodiments of the object processing method, which will not be elaborated herein.
[0263] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0264] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0265] Figure 19 FIG. shows a schematic block diagram of an exemplary electronic device 1900 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0266] As Figure 19 shown, the device 1900 includes a computing unit 1901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1902 or a computer program loaded from a storage unit 1908 into a random access memory (RAM) 1903. In the RAM 1903, various programs and data required for the operation of the device 1900 can also be stored. The computing unit 1901, the ROM 1902, and the RAM 1903 are connected to each other through a bus 1904. An input / output (I / O) interface 1905 is also connected to the bus 1904.
[0267] Multiple components in device 1900 are connected to I / O interface 1905, including: an input unit 1906, such as a keyboard, a mouse, etc.; an output unit 1907, such as various types of displays, speakers, etc.; a storage unit 1908, such as a disk, an optical disc, etc.; and a communication unit 1909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1909 allows device 1900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0268] The computing unit 1901 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1901 include, but are not limited to, a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), various dedicated Artificial Intelligence (AI) computing chips, various computing units running machine learning model algorithms, a Digital Signal Processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1901 executes each of the methods and processes described above, such as at least one of the point cloud encoding model training method, the object processing model training method, the point cloud encoding method, and the object processing method. For example, in some embodiments, at least one of the point cloud encoding model training method, the object processing model training method, the point cloud encoding method, and the object processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1908. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1900 via the ROM 1902 and / or the communication unit 1909. When the computer program is loaded into the RAM 1903 and executed by the computing unit 1901, one or more steps of at least one of the point cloud encoding model training method, the object processing model training method, the point cloud encoding method, and the object processing method described above can be executed. Alternatively, in other embodiments, the computing unit 1901 can be configured to execute at least one of the point cloud encoding model training method, the object processing model training method, the point cloud encoding method, and the object processing method in any other suitable manner (such as by means of firmware).
[0269] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0270] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0271] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0272] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) display or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0273] The systems and techniques described herein can be implemented in a computing system that includes backend components (such as, for example, a data server), or a computing system that includes middleware components (such as, for example, an application server), or a computing system that includes frontend components (such as, for example, a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (such as, for example, a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0274] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0275] Embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute a point cloud encoding model training method.
[0276] Embodiments of the present disclosure also provide a computer program product, including a computer program which, when executed by a processor, implements a point cloud encoding model training method.
[0277] It should be understood that various forms of the processes shown above may be used, steps may be reordered, added, or deleted. For example, the steps recited in the present disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitations are imposed herein. In addition, in the present disclosure, relational terms such as "first", "second", "third", etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In addition, "a plurality" in the present disclosure can be understood as at least two.
[0278] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for training a point cloud encoding model, comprising: Obtaining first point cloud data corresponding to a first training scenario; Obtaining an image feature map obtained by processing the scene image of the first training scenario; Encoding the first point cloud data through a point cloud encoding model to obtain a first point cloud feature map; Training the point cloud encoding model based on the feature distribution difference between the first point cloud feature map and the image feature map to obtain a trained point cloud encoding model; The image feature map includes a plurality of image feature regions, and each of the image feature regions corresponds to a region category; the training the point cloud encoding model based on the feature distribution difference between the first point cloud feature map and the image feature map to obtain a trained point cloud encoding model includes: Obtaining the correspondence between the first point cloud data and the scene image; Segmenting the first point cloud feature map according to the correspondence to obtain a plurality of point cloud feature regions; wherein, the plurality of point cloud feature regions correspond one-to-one to the plurality of image feature regions; Training the point cloud encoding model based on the feature distribution difference between the plurality of point cloud feature regions and the plurality of image feature regions to obtain a trained point cloud encoding model; wherein, for each of the plurality of image feature regions, the feature distribution of the image feature region is used to characterize the feature distribution of the pixel points on the image feature region; for each of the plurality of point cloud feature regions, the feature distribution of the point cloud feature region is used to characterize the feature distribution of the spatial points on the point cloud feature region.
2. The method according to claim 1, wherein The training the point cloud encoding model based on the feature distribution difference between the plurality of point cloud feature regions and the plurality of image feature regions to obtain a trained point cloud encoding model includes: Respectively taking each of the point cloud feature regions and each of the image feature regions as a target region, and calculating the feature distribution of the target region to obtain a plurality of first feature distributions and a plurality of second feature distributions; wherein, for each of the point cloud feature regions, when the point cloud feature region is taken as the target region, the obtained feature distribution is the first feature distribution, and for each of the image feature regions, when the image feature region is taken as the target region, the obtained feature distribution is the second feature distribution; Calculating the loss between the plurality of first feature distributions and the plurality of second feature distributions as the feature distribution difference between the plurality of point cloud feature regions and the plurality of image feature regions; Training the point cloud encoding model based on the feature distribution difference to obtain a trained point cloud encoding model.
3. The method according to claim 2, wherein The calculating the feature distribution of the target region includes: Performing pooling processing on the target region to obtain a region pooling result; Calculating the similarity between the target region and the region pooling result as the feature distribution of the target region.
4. The method according to claim 3, wherein The performing pooling processing on the target region to obtain a region pooling result includes: Performing max pooling processing on the target region to obtain the region pooling result.
5. The method according to claim 2, wherein, Calculating the loss between the multiple first feature distributions and the multiple second feature distributions includes: Calculating the loss between the multiple first feature distributions and the multiple second feature distributions through a pre-constructed loss function; wherein, the construction principle of the loss function includes minimizing the loss between like feature pairs and maximizing the loss between unlike feature pairs, the like feature pairs include the first feature distribution and the second feature distribution with a corresponding relationship, and the unlike feature pairs include the first feature distribution and the second feature distribution without a corresponding relationship.
6. The method according to claim 1, wherein Obtaining the correspondence between the first point cloud data and the scene image includes: Obtaining the external parameters between the lidar and the camera; wherein, the lidar is the acquisition device for acquiring the first point cloud data, and the camera is the acquisition device for acquiring the scene image; Obtaining the internal parameters of the camera; Based on the external parameters and the internal parameters, obtaining the correspondence between the first point cloud data and the scene image.
7. The method according to any one of claims 1 to 6, wherein Obtaining the image feature map obtained by processing the scene image of the first training scene includes: Performing visual segmentation on the scene image of the first training scene through an image segmentation model to obtain the image feature map; wherein, the image feature map includes the multiple image feature regions.
8. A method for training an object processing model includes: Obtaining second point cloud data corresponding to a second training scene; Encoding the second point cloud data through a target encoding model to obtain a second point cloud feature map; wherein, the target encoding model is the trained point cloud encoding model obtained by the method according to any one of claims 1 to 7; Processing the second point cloud feature map through an object processing model to obtain a predicted processing result; Training the object processing model based on the predicted processing result and the processing result label corresponding to the second point cloud data to obtain a trained object processing model.
9. The method according to claim 8, wherein, The object processing model is a three-dimensional object detection model or a three-dimensional object segmentation model.
10. A point cloud encoding method includes: Obtaining a first point cloud to be encoded corresponding to a first target scene; Encoding the first point cloud to be encoded through a target encoding model to obtain a first point cloud encoding result; wherein, the target encoding model is the trained point cloud encoding model obtained by the method according to any one of claims 1 to 7.
11. An object processing method includes: Obtaining a second point cloud to be encoded corresponding to a second target scene; Encoding the second point cloud to be encoded through a target encoding model to obtain a second point cloud encoding result; wherein, the target encoding model is the trained point cloud encoding model obtained by the method according to any one of claims 1 to 7; Processing the second point cloud encoding result through a target processing model to obtain an object processing result; wherein, the target processing model is the trained object processing model obtained by the method according to claim 8 or 9.
12. A point cloud encoding model training device includes: The first point cloud acquisition unit is configured to acquire first point cloud data corresponding to a first training scenario; The first image processing unit is configured to acquire an image feature map obtained by processing a scene image of the first training scenario; The first point cloud processing unit is configured to encode the first point cloud data through a point cloud encoding model to obtain a first point cloud feature map; The first model training unit is configured to train the point cloud encoding model based on the feature distribution difference between the first point cloud feature map and the image feature map to obtain a trained point cloud encoding model; The image feature map includes a plurality of image feature regions, and each of the image feature regions corresponds to a region category; the first model training unit is configured to: Acquire the correspondence between the first point cloud data and the scene image; Segment the first point cloud feature map according to the correspondence to obtain a plurality of point cloud feature regions; wherein, the plurality of point cloud feature regions correspond to the plurality of image feature regions one by one; Train the point cloud encoding model based on the feature distribution difference between the plurality of point cloud feature regions and the plurality of image feature regions to obtain a trained point cloud encoding model; wherein, for each of the plurality of image feature regions, the feature distribution of the image feature region is used to characterize the feature distribution of the pixel points on the image feature region; for each of the plurality of point cloud feature regions, the feature distribution of the point cloud feature region is used to characterize the feature distribution of the spatial points on the point cloud feature region.
13. The training device according to claim 12, wherein, The first model training unit is configured to: Respectively use each of the point cloud feature regions and each of the image feature regions as a target region, calculate the feature distribution of the target region to obtain a plurality of first feature distributions and a plurality of second feature distributions; wherein, for each of the point cloud feature regions, when the point cloud feature region is used as the target region, the obtained feature distribution is the first feature distribution, and for each of the image feature regions, when the image feature region is used as the target region, the obtained feature distribution is the second feature distribution; Calculate the loss between the plurality of first feature distributions and the plurality of second feature distributions as the feature distribution difference between the plurality of point cloud feature regions and the plurality of image feature regions; Train the point cloud encoding model based on the feature distribution difference to obtain a trained point cloud encoding model.
14. The training device according to claim 13, wherein, The first model training unit is configured to: Perform pooling processing on the target region to obtain a region pooling result; Calculate the similarity between the target region and the region pooling result as the feature distribution of the target region.
15. The training device according to claim 14, wherein, The first model training unit is configured to: Perform max pooling processing on the target region to obtain the region pooling result.
16. The training device according to claim 13, wherein, The first model training unit is configured to: Calculate the loss between the multiple first feature distributions and the multiple second feature distributions through a pre-constructed loss function; wherein, the construction principle of the loss function includes minimizing the loss between similar feature pairs and maximizing the loss between dissimilar feature pairs. The similar feature pairs include a first feature distribution and a second feature distribution with a corresponding relationship, and the dissimilar feature pairs include a first feature distribution and a second feature distribution without a corresponding relationship.
17. The training device according to claim 12, wherein, The first model training unit is used for: Obtain the external parameters between the lidar and the camera; wherein, the lidar is the acquisition device for acquiring the first point cloud data, and the camera is the acquisition device for acquiring the scene image; Obtain the internal parameters of the camera; Based on the external parameters and the internal parameters, obtain the corresponding relationship between the first point cloud data and the scene image.
18. The training device according to any one of claims 12 to 17, wherein, The first image processing unit: Perform visual segmentation on the scene image of the first training scene through an image segmentation model to obtain the image feature map; wherein, the image feature map includes the multiple image feature regions.
19. An object processing model training device, comprising: A second point cloud acquisition unit for acquiring second point cloud data corresponding to a second training scene; A second point cloud processing unit for encoding the second point cloud data through a target encoding model to obtain a second point cloud feature map; wherein, the target encoding model is the trained point cloud encoding model obtained by the method according to any one of claims 1 to 7; A prediction processing result acquisition unit for processing the second point cloud feature map through an object processing model to obtain a prediction processing result; A second model training unit for training the object processing model based on the prediction processing result and the processing result label corresponding to the second point cloud data to obtain a trained object processing model.
20. The apparatus according to claim 19, wherein The object processing model is a three-dimensional object detection model or a three-dimensional object segmentation model.
21. A point cloud encoding device, comprising: A first point cloud to be encoded acquisition unit for acquiring a first point cloud to be encoded corresponding to a first target scene; A first point cloud to be encoded processing unit for encoding the first point cloud to be encoded through a target encoding model to obtain a first point cloud encoding result; wherein, the target encoding model is the trained point cloud encoding model obtained by the method according to any one of claims 1 to 7.
22. An object processing device, comprising: A second point cloud to be encoded acquisition unit for acquiring a second point cloud to be encoded corresponding to a second target scene; A second point cloud to be encoded processing unit for encoding the second point cloud to be encoded through a target encoding model to obtain a second point cloud encoding result; wherein, the target encoding model is the trained point cloud encoding model obtained by the method according to any one of claims 1 to 7; An object processing result acquisition unit for processing the second point cloud encoding result through a target processing model to obtain an object processing result; wherein, the target processing model is the trained object processing model obtained by the method according to claim 8 or 9.
23. An electronic device, comprising: at least one processor; a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 11.
24. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 11.
25. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Three-dimensional scene segmentation domain migration method and device based on multi-source heterogeneous data fusion
CN116246070A
Point cloud data processing method and related device
CN116468903A