A semantic segmentation method, device and electronic device based on three-dimensional point cloud
By rotating, sorting and feature extraction and combining the three-dimensional point cloud data captured by lidar, the irregularity and sparseness of point cloud data are solved, and a highly accurate semantic segmentation effect is achieved.
Patent Information
- Application Number
- CN202210530469.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-05-16
AI Technical Summary
How to effectively use the three-dimensional point cloud data captured by lidar for semantic segmentation, facing the irregularity, disorder and sparseness of point clouds.
By obtaining the three-dimensional point cloud data to be processed, rotating the preset angle, sorting it into a one-dimensional data queue, feature extraction and combination, and finally using a semantic segmentation network for analysis.
Semantic segmentation based on three-dimensional point cloud data is realized, the adjacent point information and receptive fields of each point in the point cloud are expanded, the loss of point cloud information is reduced, and the accuracy of semantic segmentation is improved.
Smart Images

Figure CN114863107B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent perception, and particularly to a semantic segmentation method, device, and electronic device based on three-dimensional point cloud. Background Art
[0002] LiDAR is widely used in the field of autonomous driving perception. The three-dimensional point cloud data captured by LiDAR provides important geometric information of complex environments for tasks such as semantic segmentation. Different from the conventional structured images in computer vision, point clouds are irregular, disordered, and sparse. How to use the three-dimensional point cloud data captured by LiDAR for semantic segmentation has become an urgent technical problem to be solved. Summary of the Invention
[0003] The purpose of the embodiments of the present invention is to provide a semantic segmentation method, device, and electronic device based on three-dimensional point cloud, which realizes semantic segmentation based on three-dimensional point cloud data. The specific technical solutions are as follows:
[0004] According to the first aspect of the embodiments of the present invention, a semantic segmentation method based on three-dimensional point cloud is provided. The method includes:
[0005] Obtain a group of three-dimensional point cloud data to be processed, where the group of three-dimensional point cloud data to be processed includes: three-dimensional coordinate information of each point;
[0006] Rotate the group of three-dimensional point cloud data to be processed along a preset direction at least once according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data;
[0007] For each group of three-dimensional point cloud data in the group of three-dimensional point cloud data to be processed and each group of rotated three-dimensional point cloud data, perform a first sorting on the points in the group of three-dimensional point cloud data according to the spatial positions of the points in the group of three-dimensional point cloud data to obtain a one-dimensional data queue of the group of three-dimensional point cloud data;
[0008] Extract features and combine features for each of the one-dimensional data queues to obtain combined target point cloud features;
[0009] Analyze the target point cloud features by using a semantic segmentation network to obtain a semantic segmentation result.
[0010] Optionally, the step of rotating the group of three-dimensional point cloud data to be processed along a preset direction at least once according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data includes:
[0011] Rotate the group of three-dimensional point cloud data to be processed around the z-axis direction to obtain at least one group of rotated three-dimensional point cloud data.
[0012] Optionally, performing a first sorting on each point in the three-dimensional point cloud data group according to the spatial positions of the points in the three-dimensional point cloud data group to obtain a one-dimensional data queue of the three-dimensional point cloud data group includes:
[0013] Performing a first sorting on each point in the three-dimensional point cloud data group according to the spatial positions of the points in the three-dimensional point cloud data group in the order from the Z-axis to the Y-axis and then to the X-axis, or from the Z-axis to the X-axis and then to the Y-axis, to obtain a one-dimensional data queue in which the points in the three-dimensional point cloud data group are arranged in sequence.
[0014] Optionally, the performing a first sorting on each point in the three-dimensional point cloud data group according to the spatial positions of the points in the three-dimensional point cloud data group in the order from the Z-axis to the Y-axis and then to the X-axis to obtain a one-dimensional data queue in which the points in the three-dimensional point cloud data group are arranged in sequence includes:
[0015] Calculating the score of each point in the three-dimensional point cloud data group by using the following space-filling curve formula:
[0016] socres i =k x ·round(x·r x )+k y ·round(y·r y )+k z ·round(z·r z )+k ρ ·ρ
[0017] where scores i is the score of point i; x, y, and z are the three-dimensional coordinates of point i respectively; round() is a rounding function, indicating the output of the integer closest to the value in the parentheses; r x , r y , r z are preset hyperparameters; k x , k y , k z , k ρ are preset fixed parameters, where k x >>k y >>k z >>k ρ , >> means much greater than;
[0018] Performing a first sorting on each point in the three-dimensional point cloud data group in ascending or descending order of the scores of the points in the three-dimensional point cloud data group to obtain a one-dimensional data queue of the three-dimensional point cloud data group.
[0019] Optionally, the feature extraction and feature combination of each of the one-dimensional data queues to obtain the combined target point cloud features include:
[0020] For each group of three-dimensional point cloud data groups, perform feature extraction on the one-dimensional data queue of the three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group;
[0021] Perform a second sorting on each of the point cloud features respectively to obtain the point cloud features after the second sorting;
[0022] Perform feature combination on the point cloud features after the second sorting of each of the three-dimensional point cloud data groups to obtain the combined target point cloud features.
[0023] Optionally, the performing feature extraction on the one-dimensional data queue of each three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group includes:
[0024] For each group of three-dimensional point cloud data groups, encode the spatial features and attribute features of each point in the one-dimensional data queue of the three-dimensional point cloud data group to obtain the pre-encoded features of the three-dimensional point cloud data group;
[0025] Perform downsampling and upsampling on the pre-encoded features of the three-dimensional point cloud data group to obtain the sampled features of the three-dimensional point cloud data group;
[0026] Merge the sampled features and pre-encoded features of the three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group.
[0027] Optionally, the performing feature extraction on the one-dimensional data queue of each three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group includes:
[0028] For each group of three-dimensional point cloud data groups, use the pre-encoding module in the feature extraction branch network corresponding to the three-dimensional point cloud data group to encode the spatial features and attribute features of each point in its one-dimensional data queue to obtain the pre-encoded features of the three-dimensional point cloud data group, where the feature extraction branch network corresponds to the three-dimensional point cloud data group one by one;
[0029] Use the sampling and encoding module in the feature extraction branch network corresponding to the three-dimensional point cloud data group to perform downsampling and upsampling on the pre-encoded features of the three-dimensional point cloud data group to obtain the sampled features of the three-dimensional point cloud data group;
[0030] Use the feature merging module in the feature extraction branch network corresponding to the three-dimensional point cloud data group to merge the sampled features and pre-encoded features of the three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group;
[0031] Separately performing a second sorting on each of the point cloud features to obtain the point cloud features after the second sorting, including:
[0032] For each group of three-dimensional point cloud data groups, using the sorting module in the feature extraction branch network corresponding to the three-dimensional point cloud data group to perform a second sorting on the point cloud features of the three-dimensional point cloud data group, so as to obtain the point cloud features after the second sorting of the three-dimensional point cloud data group;
[0033] Performing feature combination on the point cloud features after the second sorting of each of the three-dimensional point cloud data groups to obtain the combined target point cloud features, including:
[0034] Inputting the point cloud features after the second sorting of each of the three-dimensional point cloud data groups into a feature combination network for feature combination to obtain the combined target point cloud features.
[0035] Optionally, for any group of three-dimensional point cloud data groups, the pre-coded features of the three-dimensional point cloud data group include the coded features of each point in the one-dimensional data queue of the three-dimensional point cloud data group. For each one-dimensional data queue, the coded feature of the i-th point in the one-dimensional data queue includes the first spatial feature of the i-th point, the second spatial feature of the i-th point, and the attribute feature of the i-th point, where the first spatial feature of the i-th point represents the three-dimensional coordinates of the i-th point, the second spatial feature of the i-th point represents the three-dimensional coordinate difference between the i-th point and each of its adjacent points, and the adjacent points of the i-th point are the 2M points closest to the point in the one-dimensional data queue, and M is a positive integer.
[0036] According to the second aspect of the embodiments of the present invention, there is provided a semantic segmentation device based on three-dimensional point clouds, and the device includes:
[0037] An acquisition module, configured to acquire a group of three-dimensional point cloud data to be processed, where the three-dimensional point cloud data to be processed includes: three-dimensional coordinate information of each point;
[0038] A rotation module, configured to rotate the three-dimensional point cloud data to be processed along a preset direction by at least one time according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data groups;
[0039] A data queue determination module, configured to, for each group of three-dimensional point cloud data groups in the three-dimensional point cloud data group to be processed and each rotated three-dimensional point cloud data group, perform a first sorting on the points in the three-dimensional point cloud data group according to the spatial positions of the points in the three-dimensional point cloud data group to obtain a one-dimensional data queue of the three-dimensional point cloud data group;
[0040] A feature combination module, configured to perform feature extraction and feature combination on each of the one-dimensional data queues to obtain the combined target point cloud features;
[0041] A semantic segmentation module, configured to analyze the target point cloud features by using a semantic segmentation network to obtain a semantic segmentation result.
[0042] Optionally, the rotation module is specifically configured to rotate the three-dimensional point cloud data group to be processed around the z-axis direction to obtain at least one group of rotated three-dimensional point cloud data groups.
[0043] Optionally, the data queue determination module is specifically configured to perform a first sorting on the points in the three-dimensional point cloud data group according to the spatial positions of the points in the group, in the order from the Z-axis to the Y-axis and then to the X-axis, or from the Z-axis to the X-axis and then to the Y-axis, to obtain a one-dimensional data queue in which the points in the three-dimensional point cloud data group are arranged in sequence.
[0044] Optionally, the data queue determination module is specifically configured to:
[0045] Calculate the score of each point in the three-dimensional point cloud data group by using the following space filling curve formula:
[0046] socres i = k x ·round(x · r x ) + k y ·round(y · r y ) + k z ·round(z · r z ) + k ρ · ρ
[0047] where scores i is the score of point i; x, y, and z are the three-dimensional coordinates of point i respectively; round() is a rounding function, indicating the output of the integer closest to the value in the parentheses; r x , r y , r z are preset hyperparameters; k x , k y , k z , k ρ are preset fixed parameters, where k x >> k y >> k z >> k ρ , >> indicates much greater than;
[0048] Perform a first sorting on the points in the three-dimensional point cloud data group in ascending or descending order of the scores of the points in the group to obtain a one-dimensional data queue of the three-dimensional point cloud data group.
[0049] Optionally, the feature combination module includes: a feature combination network and a plurality of feature extraction branch networks, and the feature extraction branch networks correspond to the three-dimensional point cloud data groups one by one;
[0050] The feature extraction branch network is configured to extract features from the one-dimensional data queue of the corresponding three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group; perform a second sorting on the point cloud features corresponding to the three-dimensional point cloud data group to obtain the second sorted point cloud features of the three-dimensional point cloud data group;
[0051] The feature combination network is configured to perform feature combination on the second sorted point cloud features of each three-dimensional point cloud data group to obtain the combined target point cloud features.
[0052] Optionally, the feature extraction branch network includes a pre-coding module, a sampling and coding module, a feature merging module, and a sorting module;
[0053] Specifically, for the corresponding three-dimensional point cloud data group, the feature extraction branch network uses its own pre-coding module to encode the spatial features and attribute features of each point in the one-dimensional data queue of the three-dimensional point cloud data group to obtain the pre-coded features of the three-dimensional point cloud data group; uses its own sampling and coding module to perform downsampling and upsampling on the pre-coded features of the three-dimensional point cloud data group to obtain the sampling features of the three-dimensional point cloud data group; uses its own feature merging module to merge the sampling features and pre-coded features of the three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group; uses its own sorting module to perform a second sorting on the point cloud features corresponding to the three-dimensional point cloud data group to obtain the second sorted point cloud features of the three-dimensional point cloud data group.
[0054] Optionally, for any group of three-dimensional point cloud data groups, the pre-coded features of the three-dimensional point cloud data group include the encoded features of each point in the one-dimensional data queue of the three-dimensional point cloud data group. For each one-dimensional data queue, the encoded feature of the i-th point in the one-dimensional data queue includes the first spatial feature of the i-th point, the second spatial feature of the i-th point, and the attribute feature of the i-th point, where the first spatial feature of the i-th point represents the three-dimensional coordinates of the i-th point, the second spatial feature of the i-th point represents the three-dimensional coordinate difference between the i-th point and each neighboring point of the i-th point, and each neighboring point of the i-th point is the 2M points closest to the point in the one-dimensional data queue, and M is a positive integer.
[0055] According to the third aspect of the embodiments of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0056] The memory is used to store a computer program;
[0057] A processor, when executing a program stored in a memory, implements the method steps described in any one of the first aspects.
[0058] According to a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of the first aspects are implemented.
[0059] Beneficial effects of the embodiments of the present invention:
[0060] A semantic segmentation method based on 3D point cloud provided by the embodiments of the present invention obtains a 3D point cloud data group to be processed, where the 3D point cloud data group to be processed includes: 3D coordinate information of each point; rotating the 3D point cloud data group to be processed along a preset direction at least once according to a preset selection angle to obtain at least one group of rotated 3D point cloud data groups; for each group of 3D point cloud data groups in the 3D point cloud data group to be processed and each rotated 3D point cloud data group, performing a first sorting on the points in the 3D point cloud data group according to the spatial positions of the points in the 3D point cloud data group to obtain a one-dimensional data queue of the 3D point cloud data group; performing feature extraction and feature combination on each one-dimensional data queue to obtain a combined target point cloud feature; and analyzing the target point cloud feature by using a semantic segmentation network to obtain a semantic segmentation result. In the embodiments of the present invention, semantic segmentation based on 3D point cloud data is realized, and the 3D point cloud data group to be processed is rotated along a preset direction at least once according to a preset selection angle to obtain at least one group of rotated 3D point cloud data groups, so that the adjacent point information and receptive field of each point in the point cloud can be expanded. Sorting the points in each group of 3D point cloud data groups to obtain a one-dimensional data queue of the 3D point cloud data group; the sorted point cloud is regular and has clear adjacent point information. Using the one-dimensional data queues of multiple groups of 3D point cloud data groups to perform feature extraction and feature combination to obtain a combined target point cloud feature, and the target point cloud feature includes the features of 3D point cloud data groups at multiple angles, and the receptive field is expanded. Using the target point cloud feature for semantic segmentation can reduce the loss of point cloud information compared with only using one group of 3D point cloud data groups for semantic segmentation, and thus improve the accuracy of semantic segmentation.
[0061] Of course, when implementing any product or method of the present invention, it is not necessarily required to achieve all the above-mentioned advantages simultaneously. Description of the Drawings
[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0063] Figure 1a Flowchart of a semantic segmentation method based on 3D point cloud provided by an embodiment of the present invention;
[0064] Figure 1b Schematic diagram of converting a 3D point cloud data group into a one-dimensional data queue provided by an embodiment of the present invention;
[0065] Figure 2a Flowchart of obtaining target point cloud features based on the one-dimensional data queue of the 3D point cloud data group provided by an embodiment of the present invention;
[0066] Figure 2b Schematic diagram of the structure of a semantic segmentation model provided by an embodiment of the present invention;
[0067] Figure 3a Flowchart of obtaining point cloud features of a 3D point cloud data group based on a feature extraction branch network provided by an embodiment of the present invention;
[0068] Figure 3b Schematic diagram of the structure of a feature extraction branch network provided by an embodiment of the present invention;
[0069] Figure 4 Schematic diagram of the structure of a semantic segmentation device based on 3D point cloud provided by an embodiment of the present invention;
[0070] Figure 5 Provided by an embodiment of the present invention Figure 4 Schematic diagram of the structure of the semantic segmentation module in the illustrated embodiment;
[0071] Figure 6 Provided by an embodiment of the present invention Figure 5 Schematic diagram of the structure of the feature extraction branch network in the illustrated embodiment;
[0072] Figure 7 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0073] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art based on this application belong to the scope of protection of the present invention.
[0074] In the related art, in order to overcome the irregularity, disorder, and sparsity of point clouds, the irregular point clouds are converted into regular projection-based images or three-dimensional (3D) voxels. However, there is a problem of information loss in the 3D-2D projection-based method; while in the voxelization method, the voxelization calculation amount is large, small resolutions and large voxels are applied, and the features of specific points are ignored, resulting in relatively low semantic segmentation accuracy.
[0075] To solve at least one of the above problems and implement semantic segmentation based on three-dimensional point cloud data, embodiments of the present invention provide a semantic segmentation method, apparatus, and electronic device based on three-dimensional point clouds. Among them, the method includes:
[0076] Obtain a three-dimensional point cloud data group to be processed, where the three-dimensional point cloud data group to be processed includes: three-dimensional coordinate information of each point;
[0077] Rotate the three-dimensional point cloud data group to be processed along a preset direction at least once according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data groups;
[0078] For each group of three-dimensional point cloud data groups in the three-dimensional point cloud data group to be processed and each rotated three-dimensional point cloud data group, perform a first sorting on the points in the three-dimensional point cloud data group according to the spatial positions of the points in the three-dimensional point cloud data group to obtain a one-dimensional data queue of the three-dimensional point cloud data group;
[0079] Perform feature extraction and feature combination on each of the one-dimensional data queues to obtain a combined target point cloud feature;
[0080] Use a semantic segmentation network to analyze the target point cloud feature to obtain a semantic segmentation result.
[0081] In the embodiments of the present invention, semantic segmentation based on three-dimensional point cloud data is implemented, and the three-dimensional point cloud data group to be processed is rotated along a preset direction at least once according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data groups, which can expand the adjacent point information and receptive field of each point in the point cloud. Sort the points in each group of three-dimensional point cloud data groups to obtain a one-dimensional data queue of the three-dimensional point cloud data group; the sorted point cloud is regular and has clear adjacent point information. Use the one-dimensional data queues of multiple groups of three-dimensional point cloud data groups for feature extraction and feature combination to obtain a combined target point cloud feature. The target point cloud feature includes the features of three-dimensional point cloud data groups at multiple angles, and the receptive field is expanded. Using the target point cloud feature for semantic segmentation can reduce the loss of point cloud information compared to only using one group of three-dimensional point cloud data groups for semantic segmentation, and thus improve the accuracy of semantic segmentation.
[0082] The following is a detailed description. Figure 1a The figure is a flowchart of a semantic segmentation method based on 3D point cloud provided by an embodiment of the present invention, including the following steps:
[0083] Step S101: Obtain a 3D point cloud data group to be processed, where the 3D point cloud data group to be processed includes: 3D coordinate information of each point.
[0084] The semantic segmentation method based on 3D point cloud in the embodiment of the present invention can be implemented by an electronic device. In one example, the electronic device can be an in-vehicle electronic device, such as an in-vehicle intelligent radar or an in-vehicle intelligent algorithm device, etc.
[0085] Step S102: Rotate the 3D point cloud data group to be processed along a preset direction at least once according to a preset selection angle to obtain at least one group of rotated 3D point cloud data groups.
[0086] The preset direction can be custom-set according to the actual situation. For example, the X-axis direction, the Y-axis direction, the Z-axis direction, etc. In one example, the preset direction can be the Z-axis direction, where the Z-axis is the height axis of the object, that is, the axis perpendicular to the ground. The above step S102 may include: rotating the 3D point cloud data group to be processed around the z-axis direction according to a preset selection angle to obtain at least one group of rotated 3D point cloud data groups. In the actual scene, the object is always perpendicular to the ground along the height axis of the object. Therefore, it is considered that the probability that adjacent points in the Z-axis direction belong to the same object is greater than the probability that adjacent points in the X-axis direction and the Y-axis direction belong to the same object, that is, the features in the Z-axis direction are more important, and rotating around the z-axis does not change the coordinate of the point on the z-axis. The preset selection angle can be multiple angles. For example, the preset selection angle can be π / 4, π / 2, and 3π / 4. The 3D point cloud data group to be processed can be rotated around the z-axis by π / 4, π / 2, and 3π / 4 angles respectively to obtain three groups of rotated 3D point cloud data groups.
[0087] Step S103: For each group of 3D point cloud data groups in the 3D point cloud data group to be processed and each rotated 3D point cloud data group, perform a first sorting on the points in the 3D point cloud data group according to the spatial positions of the points in the 3D point cloud data group to obtain a one-dimensional data queue of the 3D point cloud data group;
[0088] In one example, a space-filling curve can be used to sort the points in the same data group. A space-filling curve (SFC) is a sorting method that maps high-dimensional data to a one-dimensional sequence. When mapping, it can preserve the characteristics of data points to a certain extent, but there will also be a problem of information loss. For example, SFCs such as MortonNet (a mapping algorithm from high dimensions to one dimension) have information loss during the process of mapping data from high dimensions to one dimension. Existing space-filling curves usually map two-dimensional data to one dimension, while the embodiments of the present invention are to process three-dimensional point cloud data groups. In order to extend the space-filling curve to three-dimensional autonomous driving scenarios, two assumptions are first introduced. The first assumption is that the adjacent points of each point should be as similar as possible to the category of this point. The second assumption is that except for ground points, the points in the object height z-axis direction are more likely to belong to the same object than the points in other directions. The first assumption can be understood as follows: Even if a certain point at the rear of a car is closer to a tree than to other positions such as the front of the car, its characteristics should be determined more by the points related to the car rather than the tree. The second assumption can be understood as that for objects perpendicular to the ground such as cars, trees, and buildings, the points in the object height z-axis direction are more likely to belong to the object than the points in other directions.
[0089] Therefore, the above step S103 may include: for each group of three-dimensional point cloud data groups, according to the spatial positions of the points in this group of three-dimensional point cloud data groups, in the order from the Z axis to the Y axis and then to the X axis, or from the Z axis to the X axis and then to the Y axis, perform a first sorting on the points in this group of three-dimensional point cloud data groups to obtain a one-dimensional data queue in which the points in this group of three-dimensional point cloud data groups are arranged in sequence. For each group of three-dimensional point cloud data groups, the conversion of the space-filling curve can be performed on the points in this group of three-dimensional point cloud data groups according to the priority order of the space-filling curve from the Z axis to the Y axis and then to the X axis to obtain a one-dimensional data queue in which the points in this group of three-dimensional point cloud data groups are arranged in sequence. Specifically, as Figure 1b shown, it is a schematic diagram of converting a three-dimensional point cloud data group into a one-dimensional data queue. Figure 1b In it, the points are first sorted along the z axis, then along the y axis, and finally along the x axis to obtain the one-dimensional data queue of this three-dimensional point cloud data group.
[0090] In this case, the following formula is proposed to calculate the score of each point in each group of three-dimensional point cloud data groups:
[0091] socres i =k x ·round(x·r x )+k y ·round(y·r y )+k z ·round(z·rz ) + k ρ ·ρ
[0092] Among them, scores i is the score of point i; x, y, and z are the three-dimensional coordinates of point i respectively; round() is a rounding function, indicating the integer closest to the value in the parentheses; r x , r y , r z are preset hyperparameters; k x , k y , k z , k ρ are preset fixed parameters, set by experience, where k x >> k y >> k z >> k ρ , >> means much greater than; Sort the points in the three-dimensional point cloud data group in ascending or descending order of the scores of each point in the group to obtain a one-dimensional data queue of the three-dimensional point cloud data group. k x >> k y >> k z >> k ρ is to ensure that each point is sorted along the z-axis first, then along the y-axis, and finally along the x-axis. Setting k ρ is to reduce the situation where point scores are the same. >> means much greater than. Taking k x >> k y as an example, for any point in the three-dimensional point cloud data group, it is satisfied that k x ·round(x·r x ) is greater than k y ·round(y·r y ).
[0093] In the above formula, the hyperparameters r x and r y determine the size of the body column. If the body column is small, each body column will only contain a small number of points, and then feature information will be lost in each body column. On the contrary, it contradicts the first assumption because a body column will include points of different objects. These parameters are determined according to experience during the implementation process.
[0094] In another example, for each group of three-dimensional point cloud data groups, the points in the group can also be transformed by a space-filling curve according to the priority order of the space-filling curve from the Z-axis to the X-axis and then to the Y-axis to obtain a one-dimensional data queue in which the points in the group are arranged in sequence. The processing method is similar to the above process.
[0095] Embodiments of the present invention directly apply the space filling curve method to the processing of large-scale three-dimensional point cloud data, which can reduce the problem of point cloud information loss caused by SFC, and do not use the K-Nearest Neighbor (KNN) algorithm or other complex pre- / post-processing methods. Therefore, the processing efficiency of semantic segmentation can be improved.
[0096] Step S104: Perform feature extraction and feature combination on each of the one-dimensional data queues to obtain the combined target point cloud features.
[0097] Perform feature extraction on each of the one-dimensional data queues respectively, and perform feature combination based on the extracted features to obtain the combined target point cloud features.
[0098] Step S105: Analyze the target point cloud features using a semantic segmentation network to obtain a semantic segmentation result.
[0099] The semantic segmentation network is pre-trained using sample point cloud features. In one example, the training process of the semantic segmentation network may include: obtaining a plurality of sample point cloud features and the semantic segmentation labels of each sample point cloud feature. The acquisition method of the sample point cloud features can refer to the acquisition method of the target point cloud features in the present invention, which will not be elaborated here. Select a sample point cloud feature and input it into the semantic segmentation network to obtain a predicted semantic segmentation result; calculate the loss of the network according to the semantic segmentation label of the currently selected sample point cloud feature and the current predicted semantic segmentation result, and adjust the parameters of the semantic segmentation network according to the loss; select sample point cloud features to continue training the semantic segmentation network until the loss of the semantic segmentation network converges to obtain a pre-trained semantic segmentation network. The specific structure of the semantic segmentation network can be custom-selected according to the actual situation. For example, the structure of the semantic segmentation network can adopt the classifier network structure in related technologies.
[0100] In an embodiment of the present invention, a three-dimensional point cloud data group to be processed is obtained, where the three-dimensional point cloud data group includes: three-dimensional coordinate information of each point; the three-dimensional point cloud data group to be processed is rotated at least once along a preset direction according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data groups; for each group of three-dimensional point cloud data groups, the points in the group are sorted according to the spatial positions of the points in the group to obtain a one-dimensional data queue of the three-dimensional point cloud data group; a pre-trained semantic segmentation model is used to analyze the one-dimensional data queues of each three-dimensional point cloud data group to obtain a semantic segmentation result. By rotating the three-dimensional point cloud data group to be processed at least once along a preset direction according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data groups, the adjacent point information and receptive field of each point in the point cloud can be expanded. The points in each group of three-dimensional point cloud data groups are sorted to obtain a one-dimensional data queue of the three-dimensional point cloud data group; an innovative sorting method is used instead of complex domain search methods such as KNN to obtain the adjacent points of each point. The sorted point cloud is regular and has clear adjacent point information. Using the one-dimensional data queues of multiple groups of three-dimensional point cloud data groups for semantic segmentation can reduce the loss of point cloud information compared to using only one group of three-dimensional point cloud data groups for semantic segmentation, thus improving the accuracy of semantic segmentation.
[0101] In a possible implementation manner, referring to Figure 2a , in step S104, feature extraction and feature combination are performed on the one-dimensional data queues of each three-dimensional point cloud data group to obtain combined target point cloud features, including the following steps:
[0102] Step S201, for each group of three-dimensional point cloud data groups, feature extraction is performed on the one-dimensional data queue of the three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group.
[0103] Step S202, the point cloud features are respectively subjected to a second sorting to obtain the point cloud features after the second sorting.
[0104] Step S203, the point cloud features after the second sorting of each three-dimensional point cloud data group are subjected to feature combination to obtain combined target point cloud features.
[0105] In a possible implementation manner, referring to Figure 3a , in step S201, for each group of three-dimensional point cloud data groups, feature extraction is performed on the one-dimensional data queue of the three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group, including:
[0106] Step S301: For each group of three-dimensional point cloud data, encode the points in the one-dimensional data queue of this group of three-dimensional point cloud data for spatial features and attribute features to obtain the pre-encoded features of this group of three-dimensional point cloud data;
[0107] Step S302: Perform downsampling and upsampling on the pre-encoded features of this group of three-dimensional point cloud data to obtain the sampled features of this group of three-dimensional point cloud data;
[0108] Step S303: Combine the sampled features and pre-encoded features of this group of three-dimensional point cloud data to obtain the point cloud features of this group of three-dimensional point cloud data.
[0109] The feature extraction, feature combination in step S104, and the semantic segmentation network in S105 can be combined to form a semantic segmentation model. The semantic segmentation model is pre-trained, and the specific type of the semantic segmentation model can be selected according to actual semantic segmentation requirements. In one example, the semantic segmentation model can adopt a structure of a pooling layer (feature extraction part) + a classifier (semantic segmentation network). In one example, the feature extraction part includes multiple feature extraction branch networks + a feature combination network. The process of training the semantic segmentation model includes: inputting the one-dimensional data queue of each sample three-dimensional point cloud data group into the semantic segmentation model to obtain a predicted semantic segmentation result, calculating the model loss based on the predicted semantic segmentation result and the true semantic segmentation result corresponding to each sample three-dimensional point cloud data group, and adjusting the parameters of the semantic segmentation model based on the model loss. Repeat the above training process until the preset training end condition is met to obtain the semantic segmentation model.
[0110] In a possible implementation manner, the semantic segmentation model includes a coordinate information rotation network, a one-dimensional mapping network, a feature combination network, a semantic segmentation network, and multiple feature extraction branch networks. Among them, the coordinate information rotation network is used to execute step S102, the one-dimensional mapping network is used to execute step S103, step S104 is implemented using the feature combination network and multiple feature extraction branch networks of the semantic segmentation model, and step S105 is implemented using the semantic segmentation network. The process of training the semantic segmentation model includes: inputting a group of sample three-dimensional point cloud data groups into the semantic segmentation model to obtain a predicted semantic segmentation result, calculating the model loss based on the predicted semantic segmentation result and the true semantic segmentation result of the sample three-dimensional point cloud data group, and adjusting the parameters of the semantic segmentation model based on the model loss. Repeat the above training process until the preset training end condition is met to obtain the semantic segmentation model.
[0111] The preset end condition can be customized according to the actual situation. For example, reaching the preset number of training times, or the loss of the semantic segmentation model converges, etc. In one example, when the result of model training reaches the preset accuracy, it is determined that the end condition is satisfied. For example, the preset accuracy can be 90%. Semantic segmentation in a three-dimensional scene is to segment three-dimensional point cloud data into three-dimensional regional blocks with certain semantic meanings, and identify the semantic categories of each regional block, realizing the semantic reasoning process from the bottom layer to the high layer. Finally, a three-dimensional segmentation image with point-by-point semantic annotation is obtained. For an image containing trees, roads, and pedestrians, the result of semantic segmentation is trees, roads, and pedestrians.
[0112] In an alternative embodiment, the semantic segmentation model includes a feature combination network, a semantic segmentation network, and multiple feature extraction branch networks. The feature extraction branch network includes a pre-coding module, a sampling coding module, a feature merging module, and a sorting module. The three-dimensional point cloud data groups correspond one-to-one with the feature extraction branch networks.
[0113] In the above step S104, feature extraction and feature combination are performed on the one-dimensional data queues of each of the three-dimensional point cloud data groups to obtain the combined target point cloud features, including the following steps:
[0114] Step 1, for each group of three-dimensional point cloud data groups (including both the three-dimensional point cloud data group to be processed and each rotated three-dimensional point cloud data group), use the feature extraction branch network corresponding to the three-dimensional point cloud data group to perform feature extraction on the one-dimensional data queue of the three-dimensional point cloud data group, obtain the point cloud features of the three-dimensional point cloud data group, and perform a second sorting on the point cloud features of the three-dimensional point cloud data group to obtain the second sorted point cloud features of the three-dimensional point cloud data group.
[0115] Among them, each feature extraction branch network corresponds to a group of three-dimensional point cloud data groups, and different feature extraction branch networks correspond to different three-dimensional point cloud data groups; the feature extraction branch network can be a one-dimensional convolutional neural network. For example, a sequence feature aggregation backbone (SFA). The point cloud features of the three-dimensional point cloud data group can include the coordinate information and attribute features of each point, where the attribute features can be color information, reflection intensity information, etc.
[0116] Before inputting the one-dimensional data queue of each group of three-dimensional point cloud data groups into the feature extraction branch network corresponding to the group of three-dimensional point cloud data groups, the three-dimensional point cloud data group is rotated by multiple preset angles around the z-axis, so that the order of the same point in each group of three-dimensional point cloud data groups is inconsistent. Therefore, a second sorting (reverse sorting) is required to keep the order of each point consistent in each preset angle channel, preparing for feature fusion in the next step.
[0117] Step 2: Input the second sorted point cloud features of each 3D point cloud data group into the feature combination network for feature combination to obtain the combined target point cloud features.
[0118] Among them, the feature combination network is used to combine the features of 3D point cloud data groups on multiple preset angle channels.
[0119] In one example, the structure of the semantic segmentation model in the embodiments of the present invention can be as Figure 2b shown. Through Figure 2b the model shown, the steps in Figure 2a can be implemented. The model includes a preprocessing module, which is used to rotate each 3D point cloud data group to be processed around the z-axis by angles of π / 4, π / 2, and 3π / 4 respectively, and perform a first sorting on each point in each group of 3D point cloud data groups to obtain a one-dimensional data queue for each group of 3D point cloud data groups. As shown in the figure, it includes four point cloud rotation modules (corresponding to the above rotation network) and four point cloud sorting modules (corresponding to the above one-dimensional mapping network). Input the one-dimensional data queues corresponding to the four angles into the sequence feature aggregation backbone SFA (corresponding to the above feature extraction branch network) respectively. The sequence feature aggregation backbone SFA outputs (N, 316), where N represents the number of point clouds, and 316 means that each point contains 316 features; input the point cloud features of each group of 3D point cloud data groups corresponding to the four angles into the reverse sorting module (corresponding to the above sorting network) for reverse sorting, and input the sorted point cloud features into the summation module (corresponding to the above feature combination network) for feature combination to obtain the combined target point cloud features; then decode the combined target point cloud features, which can be decoded by a multi-layer perceptron network with a kernel size of 1. Add semantic labels to each point in each group of 3D point cloud data groups after decoding to perform semantic segmentation and obtain the semantic segmentation result.
[0120] In the embodiments of the present invention, the sorted point cloud features of the 3D point cloud data groups are combined through the feature combination network, so that the combined point cloud features contain the point cloud features corresponding to different rotation angles. Furthermore, the adjacent point information and receptive field of each point are expanded, improving the accuracy of semantic segmentation. At the same time, since the feature extraction branch network, that is, the one-dimensional convolutional neural network, is used, the use of KNN or other complex pre- / post-processing methods is avoided, so the processing efficiency of semantic segmentation can be improved.
[0121] In an alternative embodiment, the feature extraction branch network includes a pre-coding module, a sampling and coding module, a feature merging module, and a sorting module. For each group of 3D point cloud data groups, performing feature extraction on the one-dimensional data queue of the 3D point cloud data group to obtain the point cloud features of the 3D point cloud data group includes:
[0122] Step A: For each group of three-dimensional point cloud data groups, use the pre-coding module in the feature extraction branch network corresponding to the three-dimensional point cloud data group to encode the points in its one-dimensional data queue for spatial features and attribute features, obtaining the pre-coded features of the three-dimensional point cloud data group. Here, the feature extraction branch network corresponds one-to-one with the three-dimensional point cloud data group.
[0123] Specifically, the encoding method can be explicit encoding. In a possible implementation, for any group of three-dimensional point cloud data groups, the pre-coded features of the three-dimensional point cloud data group include the encoded features of the points in the one-dimensional data queue of the three-dimensional point cloud data group. For each one-dimensional data queue, the encoded feature of the i-th point in the one-dimensional data queue includes the first spatial feature of the i-th point, the second spatial feature of the i-th point, and the attribute feature of the i-th point. Here, the first spatial feature of the i-th point represents the three-dimensional coordinates of the i-th point, the second spatial feature of the i-th point represents the difference in three-dimensional coordinates between the i-th point and its neighboring points, and the neighboring points of the i-th point are the 2M points closest to this point in the one-dimensional data queue, where M is a positive integer.
[0124] For each one-dimensional data queue, the feature of the i-th point in the one-dimensional data queue can be represented as x i =p i ⊕(p i -p i k )⊕p i other features , where p i is the three-dimensional coordinates of point i, p i k is the three-dimensional coordinates of the point whose distance from point i is k in the one-dimensional data queue, k ∈ (-M, -M + 1, …, -1, 1, …, M - 1, M), M is an integer greater than 1, p i other features is the attribute feature of point i. For example, the attribute feature can be the color, reflection intensity, etc. of point i. Since the relative position of each point with respect to other neighboring points remains unchanged, it remains unchanged, that is, the point cloud translation invariance property. ⊕ is the concatenation operation. That is, p i k is the (i - k)-th point in the one-dimensional data queue.
[0125] Step B: Use the sampling and encoding module in the feature extraction branch network corresponding to the three-dimensional point cloud data group to perform downsampling and upsampling processing on the pre-coded features of the three-dimensional point cloud data group, obtaining the sampled features of the three-dimensional point cloud data group.
[0126] Among them, downsampling processing can be understood as reducing the data dimension, and upsampling processing can be understood as increasing the data dimension. The sampling and encoding module includes multiple downsampling layers and multiple upsampling layers. In one example, the input of the first downsampling layer is the corresponding pre-encoded feature, the input of the (i + 1)-th downsampling layer is the output of the i-th downsampling layer, the output of the first downsampling layer is also connected to the input of the feature merging module, the output of the i-th downsampling layer is also connected to the input of the (i - 1)-th upsampling layer, and the outputs of each upsampling layer are all connected to the input of the feature merging module, where i is an integer greater than 1. After being processed by the sampling and encoding module, the pre-encoded feature can obtain sampling features of multiple different scales.
[0127] Step C: Use the feature merging module in the feature extraction branch network corresponding to the three-dimensional point cloud data group to merge the sampling features and pre-encoded features of the three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group.
[0128] In one example, Figure 3b is a schematic structural diagram of the feature extraction branch network provided by the embodiment of the present invention. Through Figure 3b the shown feature extraction branch network, each step in Figure 3a can be implemented. In an optional embodiment, the feature extraction branch network can be a one-dimensional convolutional neural network, and the one-dimensional convolutional neural network can be a sequence feature aggregation backbone. As Figure 3b shown, input the one-dimensional data queue of the three-dimensional point cloud data group into the pre-encoding module for encoding. In one example, the encoding can be neighborhood explicit encoding, and explicitly encode the position differences {pi - 4, pi - 3, pi - 2, pi - 1, pi + 1, pi + 2, pi + 3, pi + 4} between point i and its nearest 8 points. The pre-encoded feature output by the pre-encoding module is (N, 28), where N is the number of points in the point cloud, and 28 represents that each point contains 28 features. Perform 5 times of downsampling processing on the pre-encoded feature in sequence to obtain 5 downsampling features. The downsampling feature obtained by the first downsampling processing is (N, 32), and the downsampling features obtained by the 2nd, 3rd, 4th, and 5th downsampling processes are respectively and perform upsampling processing on the 2nd, 3rd, 4th, and 5th downsampling features. The 4 upsampling features obtained are all (N, 64); input the pre-encoded feature, the first downsampling feature, and the 4 upsampling features into the feature merging module for merging to obtain the point cloud features of the three-dimensional point cloud data group as (N, 316).
[0129] Separately performing a second sorting on each of the point cloud features to obtain the point cloud features after the second sorting, including: for each group of three-dimensional point cloud data groups, using the sorting module in the feature extraction branch network corresponding to the three-dimensional point cloud data group to perform a second sorting on the point cloud features of the three-dimensional point cloud data group, so as to obtain the point cloud features after the second sorting of the three-dimensional point cloud data group.
[0130] The sorting module can be a reverse sorting network. Before inputting the one-dimensional data queue of each group of three-dimensional point cloud data groups into the feature extraction branch network corresponding to the group of three-dimensional point cloud data groups, the three-dimensional point cloud data group is rotated by multiple preset angles around the z-axis, so that the sequence of the same point among the points in each group of three-dimensional point cloud data groups is inconsistent. Therefore, reverse sorting is required to keep the sequence of the points consistent on each preset angle channel and prepare for the feature fusion in the next step.
[0131] In the embodiment of the present invention, each point in the one-dimensional data queue of each group of three-dimensional point cloud data groups is encoded, downsampled, upsampled, and feature combined through the feature extraction branch network, so that the point cloud features of the three-dimensional point cloud data group output by the feature extraction branch network become rich. An innovative sorting method is used instead of complex domain search methods such as KNN to obtain the neighboring points of each point, thereby greatly improving the processing efficiency. At the same time, since the point cloud features are enriched, the accuracy of semantic segmentation can also be improved.
[0132] Figure 4 It is a schematic structural diagram of a semantic segmentation device based on three-dimensional point cloud provided by an embodiment of the present invention. Referring to Figure 4 , the device includes: an acquisition module 401, a rotation module 402, a data queue determination module 403, a feature combination module 404, and a semantic segmentation module 405, where
[0133] The acquisition module 401 is configured to acquire a group of three-dimensional point cloud data to be processed, where the group of three-dimensional point cloud data to be processed includes: three-dimensional coordinate information of each point.
[0134] The rotation module 402 is configured to rotate the group of three-dimensional point cloud data to be processed along a preset direction by at least one time according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data groups.
[0135] The data queue determination module 403 is configured to, for each group of three-dimensional point cloud data groups in the group of three-dimensional point cloud data to be processed and each group of rotated three-dimensional point cloud data groups, perform a first sorting on the points in the three-dimensional point cloud data group according to the spatial positions of the points in the three-dimensional point cloud data group to obtain the one-dimensional data queue of the three-dimensional point cloud data group.
[0136] The feature combination module 404 is configured to extract features and combine features from each of the one-dimensional data queues to obtain the combined target point cloud features.
[0137] The semantic segmentation module 405 is configured to analyze the target point cloud features using a semantic segmentation network to obtain a semantic segmentation result.
[0138] In an embodiment of the present invention, a three-dimensional point cloud data group to be processed is obtained through an acquisition module, and the three-dimensional point cloud data group to be processed is rotated at least once along a preset direction according to a preset selection angle through a rotation module to obtain at least one group of rotated three-dimensional point cloud data groups; for each group of three-dimensional point cloud data groups, the points in each group of three-dimensional point cloud data groups are sorted according to the spatial positions of the points in the group through a sorting module to obtain a one-dimensional data queue of the group of three-dimensional point cloud data groups; the one-dimensional data queues of each group of three-dimensional point cloud data groups are analyzed through a semantic segmentation module using a semantic segmentation model to obtain a semantic segmentation result. Among them, rotating the three-dimensional point cloud data group to be processed at least once along a preset direction according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data groups can expand the adjacent point information and receptive field of each point in the point cloud. Sorting the points in each group of three-dimensional point cloud data groups to obtain a one-dimensional data queue of the group of three-dimensional point cloud data groups; the sorted point cloud is regular and has clear adjacent point information. Using the one-dimensional data queues of multiple groups of three-dimensional point cloud data groups to perform feature extraction and feature combination to obtain the combined target point cloud features, the target point cloud features include the features of three-dimensional point cloud data groups at multiple angles, and the receptive field is expanded. Using the target point cloud features for semantic segmentation can reduce the loss of point cloud information compared with only using one group of three-dimensional point cloud data groups for semantic segmentation, thus improving the accuracy of semantic segmentation.
[0139] Optionally, the rotation module specifically rotates the three-dimensional point cloud data group to be processed at least once along a preset direction according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data groups.
[0140] Optionally, the data queue determination module is specifically configured to perform a first sorting on the points in the group of three-dimensional point cloud data groups according to the spatial positions of the points in the group of three-dimensional point cloud data groups in the order from the Z-axis to the Y-axis and then to the X-axis, or from the Z-axis to the X-axis and then to the Y-axis, to obtain a one-dimensional data queue in which the points in the group of three-dimensional point cloud data groups are arranged in sequence.
[0141] Optionally, the data queue determination module is specifically configured to:
[0142] Calculate the score of each point in the group of three-dimensional point cloud data groups using the following space filling curve formula:
[0143] scores i = k x · round(x · r x ) + k y · round(y · r y ) + k z · round(z · r z ) + k ρ · ρ
[0144] where scores i is the score of point i; x, y, and z are the three - dimensional coordinates of point i respectively; round() is the rounding function, representing the integer closest to the value in the parentheses; r x , r y , r z are preset hyperparameters; k x , k y , k z , k ρ are preset fixed parameters, where k x >> k y >> k z >> k ρ , >> means much greater than;
[0145] According to the ascending or descending order of the scores of each point in this group of three - dimensional point cloud data groups, the points in this group of three - dimensional point cloud data groups are sorted for the first time to obtain a one - dimensional data queue of this group of three - dimensional point cloud data groups.
[0146] Optionally, referring to Figure 5 , the feature combination module 404 includes: a feature combination network 502 and multiple feature extraction branch networks 501, and the feature extraction branch networks correspond to the three - dimensional point cloud data groups one by one;
[0147] The feature extraction branch network is used to extract features from the one - dimensional data queue of the corresponding three - dimensional point cloud data group to obtain the point cloud features of this three - dimensional point cloud data group; perform a second sorting on the point cloud features corresponding to this three - dimensional point cloud data group to obtain the second - sorted point cloud features of this three - dimensional point cloud data group;
[0148] The feature combination network is used to combine the second - sorted point cloud features of each three - dimensional point cloud data group to obtain the combined target point cloud features.
[0149] Optionally, referring to Figure 6 , the feature extraction branch network 501 includes a pre - coding module 601, a sampling and coding module 602, a feature merging module 603, and a sorting module 604;
[0150] The feature extraction branch network is specifically configured to, for the corresponding three-dimensional point cloud data group, use its own pre-coding module to encode the spatial features and attribute features of each point in the one-dimensional data queue of the three-dimensional point cloud data group, so as to obtain the pre-coded features of the three-dimensional point cloud data group; use its own sampling and encoding module to perform downsampling and upsampling on the pre-coded features of the three-dimensional point cloud data group to obtain the sampled features of the three-dimensional point cloud data group; use its own feature merging module to merge the sampled features and pre-coded features of the three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group; use its own sorting module to perform a second sorting on the point cloud features corresponding to the three-dimensional point cloud data group to obtain the second sorted point cloud features of the three-dimensional point cloud data group.
[0151] Optionally, for any group of three-dimensional point cloud data groups, the pre-coded features of the three-dimensional point cloud data group include the encoded features of each point in the one-dimensional data queue of the three-dimensional point cloud data group. For each one-dimensional data queue, the encoded features of the i-th point in the one-dimensional data queue include the first spatial feature of the i-th point, the second spatial feature of the i-th point, and the attribute feature of the i-th point. Among them, the first spatial feature of the i-th point represents the three-dimensional coordinates of the i-th point, the second spatial feature of the i-th point represents the three-dimensional coordinate differences between the i-th point and its neighboring points, and the neighboring points of the i-th point are the 2M points closest to the point in the one-dimensional data queue, where M is a positive integer.
[0152] Optionally, the sampling and encoding module includes a plurality of downsampling layers and a plurality of upsampling layers. Among them, the input of the first downsampling layer is the corresponding pre-coded features, the input of the (i + 1)-th downsampling layer is the output of the i-th downsampling layer, the output of the first downsampling layer is also connected to the input of the feature merging module, the output of the i-th downsampling layer is also connected to the input of the (i - 1)-th upsampling layer, and the outputs of each upsampling layer are all connected to the input of the feature merging module, where i is an integer greater than 1.
[0153] An embodiment of the present invention also provides an electronic device, as Figure 7 shown, including a processor 701, a communication interface 702, a memory 703, and a communication bus 704. Among them, the processor 701, the communication interface 702, and the memory 703 complete mutual communication through the communication bus 704.
[0154] The memory 703 is used to store a computer program.
[0155] When the processor 701 is used to execute the program stored on the memory 703, it implements the method steps of the above-mentioned semantic segmentation method based on three-dimensional point clouds.
[0156] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in representation, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.
[0157] The communication interface is used for communication between the above electronic device and other devices.
[0158] The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.
[0159] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0160] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above-mentioned semantic segmentation methods based on three-dimensional point clouds are implemented.
[0161] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which, when running on a computer, causes the computer to execute any of the semantic segmentation methods based on three-dimensional point clouds in the above embodiments.
[0162] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0163] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0164] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device, electronic device, and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the partial description of the method embodiments for the relevant parts.
[0165] The above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A semantic segmentation method based on 3D point cloud, characterized in that, The method includes: Obtaining a three-dimensional point cloud data group to be processed, where the three-dimensional point cloud data group to be processed includes: three-dimensional coordinate information of each point; Rotating the three-dimensional point cloud data group to be processed along a preset direction at least once according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data groups, so that the adjacent point information and receptive field of each point in the three-dimensional point cloud data group to be processed are expanded; For each three-dimensional point cloud data group in the three-dimensional point cloud data group to be processed and each rotated three-dimensional point cloud data group, performing a first sorting on the points in the three-dimensional point cloud data group according to the spatial positions of the points in the three-dimensional point cloud data group to obtain a one-dimensional data queue of the three-dimensional point cloud data group, so that the point clouds corresponding to the sorted one-dimensional data queues are regular and have clear adjacent point information; Performing feature extraction and feature combination on each of the one-dimensional data queues to obtain a combined target point cloud feature; Analyzing the target point cloud feature by using a semantic segmentation network to obtain a semantic segmentation result.
2. The method according to claim 1, wherein The step of rotating the three-dimensional point cloud data group to be processed along a preset direction at least once according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data groups includes: Rotating the three-dimensional point cloud data group to be processed around the z-axis direction to obtain at least one group of rotated three-dimensional point cloud data groups.
3. The method according to claim 1, wherein The step of performing a first sorting on the points in the three-dimensional point cloud data group according to the spatial positions of the points in the three-dimensional point cloud data group to obtain a one-dimensional data queue of the three-dimensional point cloud data group includes: According to the spatial positions of the points in the three-dimensional point cloud data group, performing a first sorting on the points in the three-dimensional point cloud data group in the order from the Z-axis to the Y-axis and then to the X-axis, or from the Z-axis to the X-axis and then to the Y-axis, to obtain a one-dimensional data queue in which the points in the three-dimensional point cloud data group are arranged in sequence.
4. The method according to claim 3, characterized in that, The step of performing a first sorting on the points in the three-dimensional point cloud data group according to the spatial positions of the points in the three-dimensional point cloud data group in the order from the Z-axis to the Y-axis and then to the X-axis to obtain a one-dimensional data queue in which the points in the three-dimensional point cloud data group are arranged in sequence includes: Calculating the score of each point in the three-dimensional point cloud data group by using the following space filling curve formula: socres i = k x · round(x · r x ) + k y · round(y · r y ) + k z · round(z · r z ) + k ρ · ρ Among them, scores i is the score of point i; x, y, and z are the three-dimensional coordinates of point i respectively; round() is a rounding function, indicating the integer closest to the value in the parentheses; r x , r y , r z are preset hyperparameters; k x , k y , k z , k ρ are preset fixed parameters, where k x >> k y >> k z >> k ρ , >> means much greater than; Performing a first sorting on the points in the three-dimensional point cloud data group in ascending or descending order of the scores of the points to obtain a one-dimensional data queue of the three-dimensional point cloud data group.
5. The method according to claim 1, characterized in that, The step of performing feature extraction and feature combination on each of the one-dimensional data queues to obtain a combined target point cloud feature includes: For each three-dimensional point cloud data group, performing feature extraction on the one-dimensional data queue of the three-dimensional point cloud data group to obtain a point cloud feature of the three-dimensional point cloud data group; Performing a second sorting on each of the point cloud features respectively to obtain each second-sorted point cloud feature; Performing feature combination on the second-sorted point cloud features of each of the three-dimensional point cloud data groups to obtain a combined target point cloud feature.
6. The method according to claim 5, wherein The step of performing feature extraction on the one-dimensional data queue of the three-dimensional point cloud data group for each three-dimensional point cloud data group to obtain a point cloud feature of the three-dimensional point cloud data group includes: For each group of three-dimensional point cloud data, encode the spatial features and attribute features of each point in the one-dimensional data queue of the three-dimensional point cloud data group to obtain the pre-encoded features of the three-dimensional point cloud data group; Perform downsampling and upsampling on the pre-encoded features of the three-dimensional point cloud data group to obtain the sampled features of the three-dimensional point cloud data group; Merge the sampled features and pre-encoded features of the three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group.
7. The method according to claim 5, wherein The feature extraction for each group of three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group includes: For each group of three-dimensional point cloud data, use the pre-encoding module in the corresponding feature extraction branch network of the three-dimensional point cloud data group to encode the spatial features and attribute features of each point in its one-dimensional data queue to obtain the pre-encoded features of the three-dimensional point cloud data group, where the feature extraction branch network corresponds to the three-dimensional point cloud data group one by one; Use the sampling and encoding module in the corresponding feature extraction branch network of the three-dimensional point cloud data group to perform downsampling and upsampling on the pre-encoded features of the three-dimensional point cloud data group to obtain the sampled features of the three-dimensional point cloud data group; Use the feature merging module in the corresponding feature extraction branch network of the three-dimensional point cloud data group to merge the sampled features and pre-encoded features of the three-dimensional point cloud data group to obtain the point cloud features of the three-dimensional point cloud data group.
8. The method according to claim 6 or 7, characterized in that, For any group of three-dimensional point cloud data, the pre-encoded features of the three-dimensional point cloud data group include the encoded features of each point in the one-dimensional data queue of the three-dimensional point cloud data group. For each one-dimensional data queue, the encoded feature of the i-th point in the one-dimensional data queue includes the first spatial feature of the i-th point, the second spatial feature of the i-th point, and the attribute feature of the i-th point, where the first spatial feature of the i-th point represents the three-dimensional coordinates of the i-th point, the second spatial feature of the i-th point represents the three-dimensional coordinate difference between the i-th point and its neighboring points, and the neighboring points of the i-th point are the 2M points closest to this point in the one-dimensional data queue, and M is a positive integer.
9. A semantic segmentation device based on three-dimensional point cloud, characterized in that, The device includes: An acquisition module for acquiring a three-dimensional point cloud data group to be processed, where the three-dimensional point cloud data group to be processed includes: three-dimensional coordinate information of each point; A rotation module for rotating the three-dimensional point cloud data group to be processed along a preset direction at least once according to a preset selection angle to obtain at least one group of rotated three-dimensional point cloud data groups, so as to expand the neighboring point information and receptive field of each point in the three-dimensional point cloud data group to be processed; A data queue determination module for, for each group of three-dimensional point cloud data in the three-dimensional point cloud data group to be processed and each rotated three-dimensional point cloud data group, perform a first sorting on the points in the three-dimensional point cloud data group according to the spatial positions of the points in the three-dimensional point cloud data group to obtain the one-dimensional data queue of the three-dimensional point cloud data group, so that the point clouds corresponding to the sorted one-dimensional data queues are regular and have clear neighboring point information; A feature combination module, configured to perform feature extraction and feature combination on each of the one-dimensional data queues to obtain combined target point cloud features; A semantic segmentation module, configured to analyze the target point cloud features by using a semantic segmentation network to obtain a semantic segmentation result.
10. The device according to claim 9, characterized in that, The feature combination module includes: A feature combination network and a plurality of feature extraction branch networks, where the feature extraction branch networks correspond one-to-one to the three-dimensional point cloud data groups; the feature extraction branch network includes a pre-coding module, a sampling and coding module, a feature merging module, and a sorting module; The feature extraction branch network is configured to, for the three-dimensional point cloud data group corresponding to itself, use its own pre-coding module to encode the points in the one-dimensional data queue of the three-dimensional point cloud data group to obtain pre-coded features of the three-dimensional point cloud data group; use its own sampling and coding module to perform downsampling and upsampling processing on the pre-coded features of the three-dimensional point cloud data group to obtain sampling features of the three-dimensional point cloud data group; use its own feature merging module to merge the sampling features and pre-coded features of the three-dimensional point cloud data group to obtain point cloud features of the three-dimensional point cloud data group; use its own sorting module to perform a second sorting on the point cloud features corresponding to the three-dimensional point cloud data group to obtain the second sorted point cloud features of the three-dimensional point cloud data group; The feature combination network is configured to perform feature combination on the second sorted point cloud features of each of the three-dimensional point cloud data groups to obtain combined target point cloud features.
Citation Information
Patent Citations
Method and device for extracting road edges
CN112016355A