Point cloud dynamic target segmentation method, electronic device, and storage medium
By using a multi-scale deep strip convolutional network and a motion-guided attention feature aggregation network, the problem of insufficient feature information in dynamic target segmentation of point clouds is solved, the segmentation accuracy is improved, and the environmental perception and decision-making capabilities of autonomous driving systems are enhanced.
Patent Information
- Application Number
- CN202410644624.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-23
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-05-23
AI Technical Summary
Existing point cloud dynamic target segmentation methods do not extract enough feature information, resulting in low segmentation accuracy.
We employ a multi-scale deep strip convolutional network and a motion-guided attention feature aggregation network. By acquiring the distance image and residual image of the current frame point cloud, we perform feature extraction and fusion. We then utilize a spatial attention mechanism to re-fuse the motion feature map, thereby improving segmentation accuracy.
It improves the accuracy of dynamic target segmentation in point clouds, enabling better identification of moving objects and enhancing the environmental perception and decision-making capabilities of autonomous driving systems.
Smart Images

Figure CN118608782B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of dynamic target segmentation, in particular to a point cloud dynamic target segmentation method, an electronic device and a storage medium. BACKGROUND
[0002] Point cloud dynamic target segmentation is to obtain point cloud data of the surrounding environment by laser radar scanning, and to segment the point cloud corresponding to the dynamic target, which is one of the important tasks in the field of computer vision. Point cloud dynamic target segmentation is beneficial to help the control system of an autonomous vehicle analyze the motion state of the objects around it and make timely control when necessary to ensure the safety of the autonomous driving process.
[0003] However, the existing point cloud dynamic target segmentation method extracts insufficient feature information of the point cloud, resulting in low accuracy of point cloud dynamic target segmentation. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a point cloud dynamic target segmentation method, an electronic device and a storage medium, which can improve the accuracy of point cloud dynamic target segmentation.
[0005] According to a first aspect of an embodiment of the present application, a point cloud dynamic target segmentation method is provided, comprising the following steps:
[0006] obtaining a distance image of a current frame point cloud and a residual image between the current frame point cloud and a historical frame point cloud;
[0007] inputting the distance image into a first multi-scale deep bar convolution network for feature extraction to obtain a semantic feature map; inputting the residual image into a second multi-scale deep bar convolution network for feature extraction to obtain a motion feature map;
[0008] inputting the semantic feature map and the motion feature map into a first fusion layer for feature fusion to obtain a first fusion feature map;
[0009] inputting the first fusion feature map and the motion feature map into a coding and decoding network for feature extraction to obtain a dynamic segmentation feature map;
[0010] inputting the motion feature map and the dynamic segmentation feature map into a motion-guided attention feature aggregation network for feature aggregation to obtain a dynamic segmentation result of the current frame point cloud.
[0011] According to a second aspect of an embodiment of the present application, a laser radar point cloud dynamic target segmentation method for autonomous driving is provided, comprising the following steps:
[0012] obtaining a plurality of frames of point cloud scanned by a laser radar on an autonomous vehicle on the surrounding environment;
[0013] The point cloud dynamic target segmentation method is used to segment each frame of point cloud to obtain a corresponding dynamic segmentation result.
[0014] According to a third aspect of the embodiments of the present application, an electronic device is provided, comprising a processor and a memory; the memory stores a computer program, the computer program is suitable for being loaded and executed by the processor to perform the steps of the method of the first aspect or the method of the second aspect.
[0015] According to a fourth aspect of the embodiments of the present application, a computer readable storage medium is provided, the computer readable storage medium stores a computer program, the computer program is executed by a processor to implement the steps of the method of the first aspect or the method of the second aspect.
[0016] The embodiments of the present application obtain the distance image of the current frame of point cloud and the residual image between the current frame of point cloud and the historical frame of point cloud; input the distance image into the first multi-scale deep bar convolution network to extract features and obtain a semantic feature map; input the residual image into the second multi-scale deep bar convolution network to extract features and obtain a motion feature map; input the semantic feature map and the motion feature map into the first fusion layer to fuse features and obtain a first fusion feature map; input the first fusion feature map and the motion feature map into the coding and decoding network to extract features and obtain a dynamic segmentation feature map; and input the motion feature map and the dynamic segmentation feature map into the motion guided attention feature aggregation network to aggregate features and obtain the dynamic segmentation result of the current frame of point cloud. The embodiments of the present application extract more abundant context information in the distance image and the residual image through the first multi-scale deep bar convolution network and the second multi-scale deep bar convolution network; and through the motion guided attention feature aggregation network, the motion features of the motion feature map are re-fused in the dynamic segmentation feature map in a spatial attention manner to further pay attention to some motion details, thereby improving the accuracy of point cloud dynamic target segmentation.
[0017] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application.
[0018] In order to better understand and implement, the present application is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 A flowchart of a point cloud dynamic target segmentation method provided by an embodiment of the present application is shown;
[0020] Figure 2 A flowchart of a laser radar point cloud dynamic target segmentation method for autonomous driving provided by an embodiment of the present application is shown;
[0021] Figure 3A structural block diagram of a point cloud dynamic target segmentation device provided by an embodiment of the present application is shown in FIG. 1.
[0022] Figure 4 A structural block diagram of a laser radar point cloud dynamic target segmentation device for automatic driving provided by an embodiment of the present application is shown in FIG. 1.
[0023] Figure 5 A structural schematic block diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 1. DETAILED DESCRIPTION
[0024] In order to make the objects, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application with reference to the accompanying drawings.
[0025] It should be clear that the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0026] The terms used in the embodiments of the present application are only for the purpose of describing the specific embodiments, and are not intended to limit the embodiments of the present application. The singular forms "a", "an" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein means and includes any or all possible combinations of one or more associated listed items.
[0027] The following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not necessarily describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0028] In addition, in the description of the present application, "multiple" means two or more, unless otherwise specified. "And / or" describes the association between the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally represents that the associated objects before and after are in an "or" relationship.
[0029] Please refer toFigure 1 Fig. 1 is a flowchart of a point cloud dynamic target segmentation method according to an embodiment of the present application. The point cloud dynamic target segmentation method according to an embodiment of the present application includes the following steps:
[0030] S10: Obtain a distance image of a current frame point cloud and a residual image between the current frame point cloud and a historical frame point cloud.
[0031] In the embodiment of the present application, the distance image of the current frame point cloud is a lightweight representation of data obtained by projecting the three-dimensional point cloud data of the current frame onto a two-dimensional space.
[0032] The residual image of the current frame point cloud refers to the residual image between the current frame point cloud and the historical frame point cloud. By using the provided pose information, the coordinates of the historical frame are converted into the coordinate system of the current frame, and the historical frame after coordinate conversion and the current frame are projected onto the corresponding distance image. The absolute difference of each pixel corresponding to the generated multiple distance images is calculated, and the residual image between the current frame point cloud and each historical frame point cloud is obtained.
[0033] S20: Input the distance image into a first multi-scale deep bar convolutional network for feature extraction to obtain a semantic feature map; input the residual image into a second multi-scale deep bar convolutional network for feature extraction to obtain a motion feature map.
[0034] Specifically, the distance image itself has a serious scale imbalance problem. Due to the scanning characteristics of the laser radar, the resolution of the vertical field of view is lower than that of the horizontal field of view. When the three-dimensional point cloud is projected into a two-dimensional image, a long strip-shaped distance image will be generated. In addition, it is worth noting that the traffic participants in the distance image are mostly cars and pedestrians, which are rectangular by themselves, so the cars and pedestrians in the distance image will inevitably be stretched and may have been distorted in space.
[0035] In the related art, an N*N convolution kernel is used to extract the semantic features of the distance image. Using an N*N convolution kernel, the feature extraction receptive field is a square receptive field, while the distance image is a long strip, and the objects in the distance image are also rectangular by themselves, which will lead to a decrease in feature extraction accuracy.
[0036] Therefore, compared with the existing square receptive field, the feature extraction receptive field of the multi-scale deep bar convolutional network is a multi-scale bar receptive field, which sufficiently expands the feature extraction receptive field and the multi-scale feature depth interaction, can solve the spatial distortion problem of the objects in the distance image, and improve the accuracy of feature extraction.
[0037] In the embodiment of the present application, the semantic feature map is used to represent the semantic information in the distance image, representing the information of the objective existing object. The motion feature map is used to represent the feature information in the residual image, representing the information of the moving object. The distance image is input into the first multi-scale deep bar convolution network for feature extraction, to obtain the semantic feature map. The residual image is input into the second multi-scale deep bar convolution network for feature extraction, to obtain the motion feature map. The network structures of the first multi-scale deep bar convolution network and the second multi-scale deep bar convolution network are the same.
[0038] S30: input the semantic feature map and the motion feature map into the first fusion layer for feature fusion, to obtain the first fusion feature map.
[0039] In the embodiment of the present application, the semantic feature map and the motion feature map are input into the first fusion layer for weighted summation, to obtain the first fusion feature map.
[0040] Specifically, the semantic feature map is multiplied by the first weight, to obtain the weighted semantic feature map. The motion feature map is multiplied by the second weight, to obtain the weighted motion feature map. The weighted semantic feature map and the weighted motion feature map are added, to obtain the first fusion feature map. The sum of the first weight and the second weight is 1. The second weight is greater than the first weight, to emphasize the importance of the motion feature map.
[0041] S40: input the first fusion feature map and the motion feature map into the coding and decoding network for feature extraction, to obtain the dynamic segmentation feature map.
[0042] The dynamic segmentation feature map is used to represent the rough dynamic segmentation result.
[0043] In the embodiment of the present application, the coding and decoding network is used to further extract the feature information of the first fusion feature map and the motion feature map. The first fusion feature map and the motion feature map are input into the coding and decoding network for feature extraction, to obtain the dynamic segmentation feature map.
[0044] S50: input the motion feature map and the dynamic segmentation feature map into the motion-guided attention feature aggregation network for feature aggregation, to obtain the dynamic segmentation result of the current frame point cloud.
[0045] The motion-guided attention feature aggregation network is used to fuse the motion information in the motion feature map in the dynamic segmentation feature map, to correct the rough dynamic segmentation result, to obtain the final dynamic segmentation result.
[0046] In the embodiment of the present application, the motion-guided attention feature aggregation network re-fuses the motion feature map in the dynamic segmentation feature map with the spatial attention mechanism, to further pay attention to some motion details, so that the dynamic segmentation result of the current frame point cloud is more accurate.
[0047] According to the embodiment of the present application, the distance image of the current frame point cloud and the residual image between the current frame point cloud and the historical frame point cloud are obtained; the distance image is input into the first multi-scale depth bar convolution network for feature extraction to obtain a semantic feature map; the residual image is input into the second multi-scale depth bar convolution network for feature extraction to obtain a motion feature map; the semantic feature map and the motion feature map are input into the first fusion layer for feature fusion to obtain a first fusion feature map; the first fusion feature map and the motion feature map are input into the coding and decoding network for feature extraction to obtain a dynamic segmentation feature map; and the motion feature map and the dynamic segmentation feature map are input into the motion-guided attention feature aggregation network for feature aggregation to obtain a dynamic segmentation result of the current frame point cloud. According to the first multi-scale depth bar convolution network and the second multi-scale depth bar convolution network, more abundant context information in the distance image and the residual image is extracted; and according to the motion-guided attention feature aggregation network, the motion features of the motion feature map are re-fused in the dynamic segmentation feature map in a spatial attention manner to further pay attention to some motion details, so that the accuracy of point cloud dynamic target segmentation is improved.
[0048] In an optional embodiment, step S10 includes steps S101-S105, which are specifically as follows:
[0049] S101: Obtain point cloud information of a current frame point cloud.
[0050] In the embodiment of the present application, the current frame point cloud is obtained by scanning through a laser radar sensor. The point cloud information of the current frame point cloud includes the three-dimensional coordinates (x, y, z) of each point cloud in the current frame point cloud, the distance r of each point cloud to the laser radar sensor, and the reflection intensity e of each point cloud to the laser.
[0051] S102: Map the point cloud information into two-dimensional coordinates of a distance image using spherical mapping to obtain a distance image of the current frame point cloud.
[0052] In the embodiment of the present application, the spherical mapping formula is as follows:
[0053]
[0054] wherein (u, v) represents the coordinates in the two-dimensional distance image, h and ω represent the height and width of the two-dimensional distance image respectively, and represents the mapping relationship between the three-dimensional coordinates (x, y, z) of each point cloud in the current frame point cloud and the two-dimensional coordinates (u, v) in the distance image. f represents the field of view of the laser radar sensor, f = f up +f down , f up represents the vertical field of view angle of the laser radar sensor, f down represents the horizontal field of view angle of the laser radar sensor.
[0055] The N points in the distance image can be expressed as follows with the three-dimensional coordinates (x, y, z) of the corresponding point cloud, the distance r of the point cloud to the lidar sensor, and the reflection intensity e of the point cloud to the laser as five channels of the distance image:
[0056]
[0057] S103: Obtain several historical frame point clouds before the current frame point cloud.
[0058] In the embodiments of the present application, several frame point clouds are obtained by scanning with a lidar sensor, and the several frame point clouds before the current frame point cloud are taken as several historical frame point clouds.
[0059] S104: Convert each historical frame point cloud to a coordinate system corresponding to the current frame point cloud, calculate the distance image of each historical frame point cloud, and obtain several historical distance images.
[0060] In the embodiments of the present application, for each historical frame point cloud converted in the coordinate system, a corresponding historical distance image is obtained by a spherical mapping method. Specifically, the several historical distance images are expressed as:
[0061]
[0062] wherein, represents the kth historical distance image of the K historical distance images.
[0063] S105: Calculate the residual error between each historical distance image and the distance image of the current frame point cloud, and obtain several residual error images.
[0064] In the embodiments of the present application, a residual error calculation formula is used to obtain the residual error image between the current frame point cloud and the historical frame point cloud.
[0065] wherein, the residual error calculation formula is as follows:
[0066]
[0067] wherein, represents the residual error image between the kth historical distance image and the distance image of the current frame point cloud, represents the distance image of the current frame point cloud.
[0068] By the spherical mapping formula and the residual error calculation formula, the distance image of the current frame point cloud and the residual error image between the current frame point cloud and the historical frame point cloud can be automatically and quickly obtained.
[0069] In an optional embodiment, the first multi-scale deep bar convolutional network comprises a first convolutional layer, a first activation layer, a first multi-scale deep bar convolutional layer, a second activation layer, a second convolutional layer and a batch normalization layer connected in sequence; the first multi-scale deep bar convolutional layer comprises a depth separable convolutional layer, a first pair of bar convolutional kernels, a second pair of bar convolutional kernels and a third pair of bar convolutional kernels; the step of inputting the distance image into the first multi-scale deep bar convolutional network to extract features and obtain the semantic feature map in step S20 comprises steps S201-S206, which are as follows:
[0070] S201: input the distance image into the first convolutional layer and the first activation layer to obtain a first convolutional feature map.
[0071] In the embodiment of the present application, the size of the convolutional kernel of the first convolutional layer is 1*1, and the first activation layer is a GELU activation function. The distance image is input into the first convolutional layer and the first activation layer for convolution and activation processing to obtain the first convolutional feature map.
[0072] S202: input the first convolutional feature map into the depth separable convolutional layer to obtain a second convolutional feature map.
[0073] In the embodiment of the present application, the size of the convolutional kernel of the depth separable convolutional layer is 5*5. The first convolutional feature map is input into the depth separable convolutional layer for convolution processing to obtain the second convolutional feature map.
[0074] S203: input the second convolutional feature map into the first pair of bar convolutional kernels, the second pair of bar convolutional kernels and the third pair of bar convolutional kernels simultaneously to obtain a third convolutional feature map, a fourth convolutional feature map and a fifth convolutional feature map.
[0075] In the embodiment of the present application, the first pair of bar convolutional kernels comprises two bar convolutional kernels with sizes of 1*7 and 7*1, the second pair of bar convolutional kernels comprises two bar convolutional kernels with sizes of 1*11 and 11*1, and the third pair of bar convolutional kernels comprises two bar convolutional kernels with sizes of 1*21 and 21*1.
[0076] The second convolutional feature map is input into the first pair of bar convolutional kernels for convolution processing to obtain the third convolutional feature map. The second convolutional feature map is input into the second pair of bar convolutional kernels for convolution processing to obtain the fourth convolutional feature map. The second convolutional feature map is input into the third pair of bar convolutional kernels for convolution processing to obtain the fifth convolutional feature map.
[0077] By using two consecutive bar convolutional kernels to simulate a large convolutional kernel and combining the depth separable convolution, the amount of training parameters is significantly reduced.
[0078] S204: add the second convolution feature map, the third convolution feature map, the fourth convolution feature map and the fifth convolution feature map to obtain a sixth convolution feature map.
[0079] In the embodiment of the application, elements at the same position in the second convolution feature map, the third convolution feature map, the fourth convolution feature map and the fifth convolution feature map are added to obtain the sixth convolution feature map.
[0080] S205: input the sixth convolution feature map into the second activation layer, the second convolution layer and the batch normalization layer to obtain a seventh convolution feature map.
[0081] In the embodiment of the application, the size of the convolution kernel of the second convolution layer is 1*1, and the second activation layer is a GELU activation function. The sixth convolution feature map is input into the second activation layer, the second convolution layer and the batch normalization layer for activation, convolution and batch normalization processing to obtain the seventh convolution feature map.
[0082] S206: add the seventh convolution feature map and the distance image to obtain a semantic feature map.
[0083] In the embodiment of the application, elements at the same position in the seventh convolution feature map and the distance image are added to obtain the semantic feature map.
[0084] By inputting the distance image into the first convolution layer, the first activation layer, the first multi-scale depth bar convolution layer, the second activation layer, the second convolution layer and the batch normalization layer connected in sequence, the semantic feature map can be automatically and quickly obtained.
[0085] In an optional embodiment, the coding and decoding network includes a first encoder, a second encoder, a third encoder and a decoder; the first encoder includes a first encoding layer, a first down-sampling layer and a second fusion layer connected in sequence and alternately; the second encoder includes a second encoding layer and a second down-sampling layer connected in sequence and alternately; the decoder includes an up-sampling layer, a splicing layer and a decoding layer connected in sequence and alternately; and step S40 includes steps S41-S47, which are specifically as follows:
[0086] S41: input the first fusion feature map into the first encoding layer to obtain a first encoding feature map; and input the first encoding feature map into the first down-sampling layer to obtain a first down-sampling feature map.
[0087] The coding and decoding network is a symmetrical network structure, and such symmetry helps to maintain the continuity and hierarchical relationship of features and allows better recovery of details of input features in the decoding process to ensure effective information flow between the encoder and the decoder.
[0088] In the embodiment of the present application, the first encoder comprises four first encoding layers, four first down-sampling layers and four second fusion layers connected alternately, the second encoder comprises four second encoding layers and four second down-sampling layers connected alternately, and the decoder comprises four up-sampling layers, four splicing layers and four decoding layers connected alternately.
[0089] Specifically, the first fusion feature map is input into the first encoding layer for feature encoding to obtain a first encoding feature map. The first encoding feature map is input into the first down-sampling layer for down-sampling processing to obtain a first down-sampling feature map.
[0090] S42: input the motion feature map into the second encoding layer to obtain a second encoding feature map; and input the second encoding feature map into the second down-sampling layer to obtain a second down-sampling feature map.
[0091] In the embodiment of the present application, the motion feature map is input into the second encoding layer for feature encoding to obtain a second encoding feature map. The second encoding feature map is input into the second down-sampling layer for down-sampling processing to obtain a second down-sampling feature map.
[0092] S43: input the first down-sampling feature map and the second down-sampling feature map into the second fusion layer to obtain a second fusion feature map.
[0093] In the embodiment of the present application, the first down-sampling feature map and the second down-sampling feature map are weighted and summed by the second fusion layer to obtain a second fusion feature map.
[0094] S44: input the second fusion feature map into a third encoder to obtain a third encoding feature map.
[0095] In the embodiment of the present application, the second fusion feature map is input into the third encoder for feature encoding to obtain a third encoding feature map.
[0096] S45: input the third encoding feature map into an up-sampling layer to obtain an up-sampling feature map.
[0097] In the embodiment of the present application, the third encoding feature map is input into the up-sampling layer for up-sampling processing to obtain an up-sampling feature map.
[0098] S46: input the up-sampling feature map, the first encoding feature map and the first encoding feature map into a splicing layer to obtain a first splicing feature map.
[0099] In the embodiment of the present application, the up-sampling feature map, the first encoding feature map and the first encoding feature map are spliced by the splicing layer to obtain a first splicing feature map.
[0100] S47: input the first splicing feature map into a decoding layer to obtain a dynamic segmentation feature map.
[0101] In the embodiment of the present application, the first spliced feature map is input into the decoding layer for decoding processing to obtain the dynamic segmentation feature map.
[0102] By inputting the first fusion feature map and the input motion feature map into the coding-decoding network, the dynamic segmentation feature map can be automatically and quickly obtained.
[0103] In an optional embodiment, the first encoding layer comprises a third convolutional layer, a third activation layer, a second multi-scale deep bar convolutional layer, a first convolutional module, a second convolutional module and a third convolutional module connected in sequence; and the step of inputting the first fusion feature map into the first encoding layer to obtain the first encoding feature map in step S41 comprises steps S411-S415, which are specifically as follows:
[0104] S411: input the first fusion feature map into the third convolutional layer, the third activation layer, the second multi-scale deep bar convolutional layer and the first convolutional module to obtain an eighth convolutional feature map.
[0105] In the embodiment of the present application, the size of the convolution kernel of the third convolutional layer is 1*1, the third activation layer is a GELU activation function, and the second multi-scale deep bar convolutional layer has the same structure as the first multi-scale deep bar convolutional layer described above, which will not be repeated here.
[0106] The size of the convolution kernel of the first convolutional module is 1*1.
[0107] The first fusion feature map is input into the third convolutional layer, the third activation layer, the second multi-scale deep bar convolutional layer and the first convolutional module connected in sequence for convolution processing to obtain an eighth convolutional feature map.
[0108] S412: input the eighth convolutional feature map into the second convolutional module to obtain a ninth convolutional feature map.
[0109] In the embodiment of the present application, the second convolutional module comprises a convolutional layer, an activation layer and a batch normalization layer connected in sequence. The size of the convolution kernel of the convolutional layer is 2*4, and the dilation coefficient is (2, 2). The activation layer is a GELU activation function.
[0110] The eighth convolutional feature map is input into the second convolutional module for convolution processing to obtain a ninth convolutional feature map.
[0111] S413: splice the eighth convolutional feature map and the ninth convolutional feature map to obtain a tenth convolutional feature map.
[0112] In the embodiment of the present application, the eighth convolutional feature map and the ninth convolutional feature map are spliced to retain all feature information of the eighth convolutional feature map and the ninth convolutional feature map to obtain a tenth convolutional feature map.
[0113] S414: input the tenth convolutional feature map into the third convolutional module to obtain an eleventh convolutional feature map.
[0114] In the embodiment of the present application, the tenth convolutional feature map is input into the third convolutional module for convolution processing to obtain the eleventh convolutional feature map.
[0115] S415: add the eleventh convolutional feature map and the first fusion feature map to obtain a first encoding feature map.
[0116] In the embodiment of the present application, the eleventh convolutional feature map and the elements at the same position in the first fusion feature map are added to obtain the first encoding feature map.
[0117] By inputting the first fusion feature map into the third convolutional layer, the third activation layer, the second multi-scale depth bar convolutional layer, the first convolutional module, the second convolutional module and the third convolutional module connected in sequence, the first encoding feature map can be automatically and quickly obtained.
[0118] In an optional embodiment, the motion-guided attention feature aggregation network comprises a first motion-guided attention feature aggregation module and a second motion-guided attention feature aggregation module; step S50 comprises steps S51-S55, which are as follows:
[0119] S51: splice the motion feature map and the dynamic segmentation feature map to obtain a second spliced feature map.
[0120] In the embodiment of the present application, the motion feature map and the dynamic segmentation feature map are spliced to retain all feature information of the motion feature map and the dynamic segmentation feature map, and the second spliced feature map is obtained.
[0121] S52: input the second spliced feature map into the first motion-guided attention feature aggregation module to obtain a first aggregated feature map.
[0122] In the embodiment of the present application, the second spliced feature map is input into the first motion-guided attention feature aggregation module to perform the first spatial attention mechanism guided fusion of the motion feature, and the first aggregated feature map is obtained.
[0123] S53: add the dynamic segmentation feature map and the first aggregated feature map to obtain a fourth feature map.
[0124] In the embodiment of the present application, the elements at the same position in the dynamic segmentation feature map and the first aggregated feature map are added to obtain the fourth feature map.
[0125] S54: splice the fourth feature map and the motion feature map to obtain a third spliced feature map.
[0126] In the embodiment of the present application, the fourth feature map is spliced with the motion feature map to retain all feature information of the fourth feature map and the motion feature map, and a third spliced feature map is obtained.
[0127] S55: inputting the third spliced feature map into a second motion-guided attention feature aggregation module to obtain a dynamic segmentation result of the current frame point cloud.
[0128] In the embodiment of the present application, the third spliced feature map is inputted into the second motion-guided attention feature aggregation module to perform a second fusion of the motion feature guided by the spatial attention mechanism, and a dynamic segmentation result of the current frame point cloud is obtained.
[0129] Through the first motion-guided attention feature aggregation module and the second motion-guided attention feature aggregation module, the dynamic segmentation feature map is continuously used twice by the spatial attention mechanism, and a dynamic segmentation result of the current frame point cloud can be automatically and quickly obtained.
[0130] In an optional embodiment, the first motion-guided attention feature aggregation module comprises a fourth convolutional layer, a second batch normalization layer, a third multi-scale depth bar convolutional layer and a fifth convolutional layer connected in sequence; and step S52 comprises steps S521-S522, which are specifically as follows:
[0131] S521: inputting the second spliced feature map into the fourth convolutional layer, the second batch normalization layer, the third multi-scale depth bar convolutional layer and the fifth convolutional layer to obtain an intermediate feature map.
[0132] In the embodiment of the present application, the convolution kernel size of the fourth convolutional layer and the fifth convolutional layer is 1*1, and the structure of the third multi-scale depth bar convolutional layer is the same as that of the first multi-scale depth bar convolutional layer, which will not be described herein.
[0133] The second spliced feature map is inputted into the fourth convolutional layer, the second batch normalization layer, the third multi-scale depth bar convolutional layer and the fifth convolutional layer for convolution processing to obtain an intermediate feature map.
[0134] S522: multiplying the intermediate feature map with the second spliced feature map to obtain a first aggregated feature map.
[0135] In the embodiment of the present application, the elements of the intermediate feature map are multiplied with the elements of the second spliced feature map to obtain the first aggregated feature map.
[0136] By inputting the second spliced feature map into the fourth convolutional layer, the second batch normalization layer, the third multi-scale depth bar convolutional layer and the fifth convolutional layer connected in sequence, the first aggregated feature map can be automatically and quickly obtained.
[0137] Please refer to Figure 2FIG. 1 is a flowchart of a method for segmenting dynamic targets in a laser radar point cloud for autonomous driving according to an embodiment of the present application. The method for segmenting dynamic targets in a laser radar point cloud for autonomous driving according to the embodiment of the present application includes the following steps:
[0138] S100: Obtain several frames of point clouds scanned by a laser radar on an autonomous vehicle on the surrounding environment.
[0139] In an autonomous driving scenario, moving objects such as pedestrians, cyclists, and moving vehicles often appear. On the one hand, moving objects can help an autonomous navigation system to predict the future state of the surrounding environment, avoid collisions, and plan tasks. On the other hand, moving objects can also interfere with the perception of the laser radar, resulting in poor performance of downstream tasks such as point cloud registration, simultaneous localization and mapping (SLAM), static map creation, and path planning. Therefore, moving object segmentation plays a crucial role in obtaining more accurate environmental perception and making more reliable decisions for the autonomous navigation system.
[0140] In the embodiment of the present application, the environment around the autonomous vehicle is scanned by the laser radar installed on the autonomous vehicle, and several frames of point clouds are obtained.
[0141] S200: Use the above-mentioned point cloud dynamic target segmentation method to segment the dynamic targets in each frame of point cloud, and obtain the corresponding dynamic segmentation result.
[0142] In the embodiment of the present application, the specific steps of the point cloud dynamic target segmentation method are described in steps S10-S50, which will not be repeated here. By using the point cloud dynamic target segmentation method to segment the dynamic targets in each frame of point cloud, the real-time dynamic segmentation result of each frame of point cloud can be obtained.
[0143] By obtaining several frames of point clouds scanned by a laser radar on an autonomous vehicle on the surrounding environment, using the above-mentioned point cloud dynamic target segmentation method to segment the dynamic targets in each frame of point cloud, and obtaining the corresponding dynamic segmentation result, the point cloud dynamic target can be segmented in real time in an autonomous driving scenario, and the accuracy of point cloud dynamic target segmentation is improved.
[0144] The following is an apparatus embodiment of the present application, which can be used to execute the content of the method in the embodiments of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the content of the method in the embodiments of the present application.
[0145] Please refer to Figure 3 FIG. 2 shows a structure diagram of a point cloud dynamic target segmentation device according to an embodiment of the present application. The point cloud dynamic target segmentation device 6 according to the embodiment of the present application includes:
[0146] The distance image acquisition module 61 is configured to acquire a distance image of the current frame point cloud and a residual image between the current frame point cloud and the historical frame point cloud.
[0147] The semantic feature map acquisition module 62 is configured to input the distance image into a first multi-scale depth bar convolution network to extract features and obtain a semantic feature map; and input the residual image into a second multi-scale depth bar convolution network to extract features and obtain a motion feature map.
[0148] The first fusion feature map acquisition module 63 is configured to input the semantic feature map and the motion feature map into a first fusion layer to fuse features and obtain a first fusion feature map.
[0149] The dynamic segmentation feature map acquisition module 64 is configured to input the first fusion feature map and the motion feature map into a coding-decoding network to extract features and obtain a dynamic segmentation feature map.
[0150] The dynamic segmentation result acquisition module 65 is configured to input the motion feature map and the dynamic segmentation feature map into a motion-guided attention feature aggregation network to aggregate features and obtain a dynamic segmentation result of the current frame point cloud.
[0151] By using the embodiments of the present application, the distance image of the current frame point cloud and the residual image between the current frame point cloud and the historical frame point cloud are acquired; the distance image is input into the first multi-scale depth bar convolution network to extract features and obtain the semantic feature map; the residual image is input into the second multi-scale depth bar convolution network to extract features and obtain the motion feature map; the semantic feature map and the motion feature map are input into the first fusion layer to fuse features and obtain the first fusion feature map; the first fusion feature map and the motion feature map are input into the coding-decoding network to extract features and obtain the dynamic segmentation feature map; and the motion feature map and the dynamic segmentation feature map are input into the motion-guided attention feature aggregation network to aggregate features and obtain the dynamic segmentation result of the current frame point cloud. The first multi-scale depth bar convolution network and the second multi-scale depth bar convolution network are used to extract more abundant context information in the distance image and the residual image; the motion-guided attention feature aggregation network is used to re-fuse the motion features of the motion feature map in the dynamic segmentation feature map in a manner of spatial attention to further pay attention to some motion details, thereby improving the accuracy of the point cloud dynamic target segmentation.
[0152] The following is an apparatus embodiment of the present application, which can be used to execute the content of the method in the embodiments of the present application. For details not disclosed in the apparatus embodiments of the present application, refer to the content of the method in the embodiments of the present application.
[0153] Please refer to Figure 4Fig. 7 shows a structural schematic diagram of a laser radar point cloud dynamic target segmentation device for autonomous driving provided by an embodiment of the present application. The laser radar point cloud dynamic target segmentation device 7 for autonomous driving provided by the embodiment of the present application comprises:
[0154] The point cloud acquisition module 71 is configured to acquire a plurality of frames of point clouds scanned by a laser radar on an autonomous driving vehicle on a surrounding environment.
[0155] The point cloud dynamic target segmentation module 72 is configured to perform point cloud dynamic target segmentation on each frame of point cloud by using the point cloud dynamic target segmentation method, and obtain a corresponding dynamic segmentation result.
[0156] By acquiring a plurality of frames of point clouds scanned by a laser radar on an autonomous driving vehicle on a surrounding environment, performing point cloud dynamic target segmentation on each frame of point cloud by using the point cloud dynamic target segmentation method, and obtaining a corresponding dynamic segmentation result, the embodiment of the present application can segment point cloud dynamic targets in real time in an autonomous driving scene, and improve the accuracy of point cloud dynamic target segmentation.
[0157] The following is an equipment embodiment of the present application, which can be used to execute the content of the method in the embodiment of the present application. For details not disclosed in the equipment embodiment of the present application, please refer to the content of the method in the embodiment of the present application.
[0158] Please refer to Figure 5 The present application also provides an electronic device 300, which can be specifically a computer, a mobile phone, a tablet computer, etc. In the exemplary embodiment of the present application, the electronic device 300 is a computer, which can include at least one processor 301, at least one memory 302, at least one display, at least one network interface 303, a user interface 304, and at least one communication bus 305.
[0159] The user interface 304 is mainly used to provide an interface for user input and obtain data input by the user. Optionally, the user interface can also include a standard wired interface and a wireless interface.
[0160] The network interface 303 can optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0161] The communication bus 305 is used to realize the connection and communication between the components.
[0162] The processor 301 can include one or more processing cores. The processor connects various parts in the entire electronic device through various interfaces and lines, executes various functions of the electronic device and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Alternatively, the processor can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor can be integrated with a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content to be displayed by the display layer; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor, but can be realized by a separate chip.
[0163] The memory 302 can include a random access memory (RAM) and a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store data involved in the above-mentioned various method embodiments, etc. The memory can also be at least one storage device located away from the above-mentioned processor. For example, Figure 3 The memory as a computer storage medium can include an operating system, a network communication module, a user interface module, and an operation application.
[0164] The processor can be used to call the application program of the point cloud dynamic target segmentation method stored in the memory, and specifically execute the method steps of the above-mentioned embodiments. The specific execution process can be referred to the specific description of the embodiments, which will not be repeated here.
[0165] It should also be noted that the terms "comprising", "comprises" or other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0166] The above embodiments are only used to illustrate the present application, but not to limit it. Instead of the above, various modifications and changes can be made to the application by those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall fall into the scope of the claims of the application.
Claims
1. A method for point cloud dynamic object segmentation, characterized in that, The method comprises the following steps: obtaining a distance image of a current frame point cloud and a residual image between the current frame point cloud and a historical frame point cloud; inputting the distance image into a first multi-scale deep bar convolution network for feature extraction to obtain a semantic feature map; inputting the residual image into a second multi-scale deep bar convolution network for feature extraction to obtain a motion feature map; inputting the semantic feature map and the motion feature map into a first fusion layer for feature fusion to obtain a first fusion feature map; inputting the first fusion feature map and the motion feature map into a coding and decoding network for feature extraction to obtain a dynamic segmentation feature map; inputting the motion feature map and the dynamic segmentation feature map into a motion-guided attention feature aggregation network for feature aggregation to obtain a dynamic segmentation result of the current frame point cloud; wherein the motion-guided attention feature aggregation network comprises a first motion-guided attention feature aggregation module and a second motion-guided attention feature aggregation module; the first motion-guided attention feature aggregation module comprises a fourth convolution layer, a second batch normalization layer, a third multi-scale deep bar convolution layer and a fifth convolution layer connected in sequence.
2. The point cloud dynamic target segmentation method according to claim 1, wherein: the first multi-scale deep bar convolution network comprises a first convolution layer, a first activation layer, a first multi-scale deep bar convolution layer, a second activation layer, a second convolution layer and a batch normalization layer connected in sequence; and the first multi-scale deep bar convolution layer comprises a depth separable convolution layer, a first bar convolution kernel pair, a second bar convolution kernel pair and a third bar convolution kernel pair; the step of inputting the distance image into the first multi-scale deep bar convolution network for feature extraction to obtain a semantic feature map comprises: inputting the distance image into the first convolution layer and the first activation layer to obtain a first convolution feature map; inputting the first convolution feature map into the depth separable convolution layer to obtain a second convolution feature map; simultaneously inputting the second convolution feature map into the first bar convolution kernel pair, the second bar convolution kernel pair and the third bar convolution kernel pair to obtain a third convolution feature map, a fourth convolution feature map and a fifth convolution feature map; adding the second convolution feature map, the third convolution feature map, the fourth convolution feature map and the fifth convolution feature map to obtain a sixth convolution feature map; inputting the sixth convolution feature map into the second activation layer, the second convolution layer and the batch normalization layer to obtain a seventh convolution feature map; adding the seventh convolution feature map and the distance image to obtain the semantic feature map.
3. The point cloud dynamic target segmentation method according to claim 1, wherein: the coding and decoding network comprises a first encoder, a second encoder, a third encoder and a decoder; the first encoder comprises a first encoding layer, a first down-sampling layer and a second fusion layer connected in sequence; the second encoder comprises a second encoding layer and a second down-sampling layer connected in sequence; and the decoder comprises an up-sampling layer, a splicing layer and a decoding layer connected in sequence. The step of inputting the first fusion feature map and the motion feature map into a coding and decoding network for feature extraction to obtain a dynamic segmentation feature map comprises: inputting the first fusion feature map into the first encoding layer to obtain a first encoding feature map; and inputting the first encoding feature map into the first down-sampling layer to obtain a first down-sampled feature map; inputting the motion feature map into the second encoding layer to obtain a second encoding feature map; and inputting the second encoding feature map into the second down-sampling layer to obtain a second down-sampled feature map; inputting the first down-sampled feature map and the second down-sampled feature map into the second fusion layer to obtain a second fusion feature map; inputting the second fusion feature map into the third encoder to obtain a third encoding feature map; inputting the third encoding feature map into the up-sampling layer to obtain an up-sampled feature map; inputting the up-sampled feature map, the first encoding feature map and the first encoding feature map into the splicing layer to obtain a first spliced feature map; inputting the first spliced feature map into the decoding layer to obtain the dynamic segmentation feature map.
4. The point cloud dynamic target segmentation method according to claim 3, wherein: the first encoding layer comprises a third convolutional layer, a third activation layer, a second multi-scale deep bar convolutional layer, a first convolutional module, a second convolutional module and a third convolutional module connected in sequence; the step of inputting the first fusion feature map into the first encoding layer to obtain a first encoding feature map comprises: inputting the first fusion feature map into the third convolutional layer, the third activation layer, the second multi-scale deep bar convolutional layer and the first convolutional module to obtain an eighth convolutional feature map; inputting the eighth convolutional feature map into the second convolutional module to obtain a ninth convolutional feature map; splicing the eighth convolutional feature map and the ninth convolutional feature map to obtain a tenth convolutional feature map; inputting the tenth convolutional feature map into the third convolutional module to obtain an eleventh convolutional feature map; adding the eleventh convolutional feature map and the first fusion feature map to obtain the first encoding feature map.
5. The point cloud dynamic target segmentation method according to claim 1, wherein: the step of inputting the motion feature map and the dynamic segmentation feature map into a motion-guided attention feature aggregation network for feature aggregation to obtain a dynamic segmentation result of the current frame point cloud comprises: splicing the motion feature map and the dynamic segmentation feature map to obtain a second spliced feature map; inputting the second spliced feature map into the first motion-guided attention feature aggregation module to obtain a first aggregated feature map; adding the dynamic segmentation feature map and the first aggregated feature map to obtain a fourth feature map; splicing the fourth feature map and the motion feature map to obtain a third spliced feature map; inputting the third spliced feature map into the second motion-guided attention feature aggregation module to obtain the dynamic segmentation result of the current frame point cloud.
6. The point cloud dynamic target segmentation method according to claim 5, wherein: The step of inputting the second spliced feature map into the first motion guided attention feature aggregation module to obtain a first aggregated feature map comprises: inputting the second spliced feature map into the fourth convolutional layer, the second batch normalization layer, the third multi-scale depth bar convolutional layer, and the fifth convolutional layer to obtain an intermediate feature map; multiplying the intermediate feature map and the second spliced feature map to obtain the first aggregated feature map.
7. The point cloud dynamic target segmentation method according to any one of claims 1 to 6, characterized in that: The step of obtaining the distance image of the current frame point cloud and the residual image between the current frame point cloud and the historical frame point cloud comprises: obtaining point cloud information of the current frame point cloud; using spherical mapping to map the point cloud information into two-dimensional coordinates of the distance image to obtain the distance image of the current frame point cloud; obtaining several historical frame point clouds before the current frame point cloud; converting each of the historical frame point clouds into a coordinate system corresponding to the current frame point cloud to calculate the distance image of each of the historical frame point clouds to obtain several historical distance images; calculating the residual between each of the historical distance images and the distance image of the current frame point cloud to obtain several residual images.
8. A method for dynamic target segmentation of LiDAR point clouds for autonomous driving, characterized in that, comprising the following steps: obtaining several frames of point clouds scanned by a laser radar on an autonomous vehicle on the surrounding environment; using the point cloud dynamic target segmentation method according to any one of claims 1 to 7 to perform point cloud dynamic target segmentation on each frame of point cloud to obtain a corresponding dynamic segmentation result.
9. An electronic device, comprising: comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1-9. The computer program is executed by the processor to implement the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Three-dimensional laser radar point cloud semantic segmentation method and device based on deep learning
CN116229057A
Point cloud dynamic target segmentation method based on regional-local self-attention
CN117237405A