A Point Cloud Automatic Annotation Method, Device and Medium Based on Neural Network

By performing relative coordinate system conversion, voxel coding and feature fusion on three-dimensional point cloud data, initial detection frames are generated and size adjusted, the problem of insufficient utilization of timing features in the prior art is solved, and detection accuracy and recall rate are improved.

CN119919929BActive Publication Date: 2025-06-13HANG ZHOU MINDFLOW TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510417137.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-13
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing automatic labeling method does not fully utilize the timing characteristics of three-dimensional point cloud data, resulting in unsatisfactory detection accuracy and recall, and the detection frame size is not accurate enough.

Method used

By obtaining the original point cloud sequence, dividing it into subsequences, the point cloud data of each frame is converted to the relative coordinate system, divided into voxel blocks for encoding, extracting spatial and timing features, performing feature fusion and convolution operations, generating an initial detection box, and adjusting the detection box size by tracking the trajectory.

Benefits of technology

Improve the accuracy and recall of point cloud detection results, obtain more accurate detection frame size, and reduce calculation complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919929B_ABST
    Figure CN119919929B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and medium for automatic annotation of point clouds based on a neural network, relating to the field of data processing. The method includes: obtaining an original point cloud sequence and dividing the point cloud sequence into a plurality of subsequences; converting each point in each frame of point cloud data from the world coordinate system to the relative coordinate system to obtain the relative coordinate system value of each point; evenly dividing each frame of point cloud data into a plurality of preset voxel blocks, and encoding the voxel blocks according to whether the voxel blocks contain valid points to obtain the valid voxels and empty voxels included in each frame of point cloud data; obtaining the spatial feature vector of each frame of point cloud data; using a neural network to extract the temporal feature vector of the subsequence; performing feature fusion to obtain a fused feature vector and generating an initial detection box; adjusting the size of the initial detection box according to the tracking trajectory to obtain a target detection box. The present invention can make full use of the spatial features and temporal features of the point clouds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and particularly to a method, device and medium for automatically annotating point clouds based on a neural network. Background Art

[0002] The annotation of three-dimensional point cloud data plays a crucial role in current deep learning systems, especially in the research and development processes of systems such as robots, embodied intelligence, and autonomous driving. At the same time, higher requirements are being put forward for data closed-loop and the efficiency and quality of data annotation on the data side. For three-dimensional point cloud data, the method of simply relying on annotators for data annotation generally has problems such as relatively difficult annotation efficiency, annotation fineness, and size estimation. In the case where the method of simply relying on annotators for data annotation can no longer meet the high-precision and high-efficiency annotation requirements proposed by high-precision algorithms, it has gradually become an important means to improve the efficiency of point cloud annotation and the quality of annotation results by using an automatic annotation method to annotate the data to be annotated and then having the annotators perform operations such as attribute supplementation, minor fine-tuning, filling in the gaps, and filtering errors.

[0003] However, some current mainstream automatic annotation methods do not make full use of the temporal features of point clouds. For example, some annotation methods use a single frame for detection output, and some annotation methods directly splice adjacent frames. These annotation methods will result in unsatisfactory detection accuracy and recall rate, and there are also situations where the size of the detection box is not precise enough.

[0004] Therefore, a method for data annotation of three-dimensional point cloud data needs to be provided. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method, device and medium for automatically annotating point clouds based on a neural network to solve the problem that the current automatic annotation method does not make full use of the temporal features of point clouds.

[0006] According to a first aspect, an embodiment of the present invention provides a method for automatically annotating point clouds based on a neural network, the method comprising:

[0007] Obtain an original point cloud sequence, and divide the point cloud sequence into a plurality of subsequences; each subsequence contains point cloud data of a plurality of consecutive frames;

[0008] According to the pose of each frame of point cloud data in the world coordinate system in the subsequence, convert each point in each frame of point cloud data from the world coordinate system to a relative coordinate system to obtain the relative coordinate value of each point; the origin of the relative coordinate system is the first frame of point cloud data of the corresponding subsequence, and the relative coordinate value includes the coordinate value and intensity value of each point in the relative coordinate system;

[0009] Divide each frame of the converted point cloud data into several preset voxel blocks uniformly, and encode the voxel blocks according to the valid points in the voxel blocks to obtain the valid voxels and empty voxels included in each frame of point cloud data; the voxel blocks containing valid points are valid voxels, and the voxel blocks without valid points are empty voxels;

[0010] Extract features from the valid voxels and convert the empty voxels into dense voxel features to obtain the spatial feature vectors of each frame of point cloud data, and obtain the spatial feature vectors of the subsequence according to the spatial feature vectors of each frame of point cloud data;

[0011] Use a neural network to extract the temporal feature vectors of the subsequence;

[0012] Fuse the spatial feature vectors and temporal feature vectors of the subsequence to obtain the fused feature vectors, perform convolution operations on the fused feature vectors to obtain the planar feature maps, and generate the initial detection boxes according to the planar feature maps;

[0013] Continuously track the initial detection boxes between each subsequence according to the spatial relationship of the initial detection boxes between each subsequence to obtain the tracking trajectories;

[0014] Adjust the sizes of the initial detection boxes according to the tracking trajectories to obtain the target detection boxes.

[0015] Combined with the first aspect, in the first embodiment of the first aspect, the method of converting each point in each frame of point cloud data from the world coordinate system to the relative coordinate system according to the pose of each frame of point cloud data in the subsequence in the world coordinate system specifically includes:

[0016] Determine the pose of each frame of point cloud data in the subsequence relative to the world coordinate system, and convert each point in each point cloud data in the subsequence into homogeneous coordinates;

[0017] Convert each point in the subsequence from the world coordinate system to the relative coordinate system according to the homogeneous coordinates corresponding to each point;

[0018] Convert the coordinate values in the homogeneous coordinates according to the point cloud data corresponding to each point and the pose of the first frame of point cloud data in the subsequence relative to the world coordinate system, and determine the coordinate values of each point in the relative coordinate system;

[0019] Normalize the intensity values in the homogeneous coordinates according to the maximum and minimum values of the preset point cloud intensity, and determine the intensity values of each point in the relative coordinate system.

[0020] Combined with the first aspect, in the second embodiment of the first aspect, the method of uniformly dividing each frame of converted point cloud data into a number of preset voxel blocks and encoding the voxel blocks according to the valid points in the voxel blocks to obtain the valid voxels and empty voxels included in each frame of point cloud data specifically includes:

[0021] Determine the valid area to be processed for each frame of point cloud data in each subsequence, uniformly divide the valid area into a number of preset voxel blocks, and determine the points under the point cloud data included in each voxel block; the preset voxel block is a cubic space with a size set as a cubic space, representing the size of the voxel block in the x-axis direction, representing the size of the voxel block in the y-axis direction, representing the size of the voxel block in the z-axis direction;

[0022] Use a preset encoder to perform voxel encoding on each voxel block to determine the valid voxels in the voxel block; the input data of the encoder is the relative coordinate system values of all the points in the voxel block.

[0023] Combined with the second embodiment of the first aspect, in the third embodiment of the first aspect, the method of using a neural network to extract the temporal feature vector of the subsequence specifically includes:

[0024] Uniformly divide the spatial feature vector corresponding to each frame of point cloud data into a number of sub-regions, and determine the voxel blocks included in each sub-region;

[0025] Concatenate the feature vectors of all the voxel blocks in the sub-region before and after to obtain the sub-space feature vector of the sub-region;

[0026] Pool the sub-space feature vectors of all the sub-regions of the subsequence, and perform temporal feature extraction on all the sub-space feature vectors included in the subsequence through a neural network to obtain the temporal feature vector.

[0027] Combined with the first aspect, in the fourth embodiment of the first aspect, the method of feature fusion is:

[0028] Obtain the point spatial feature vector of each point in the spatial feature vector and the point temporal feature vector of each point in the temporal feature vector, and concatenate the point spatial feature vector and the point temporal feature vector to obtain the point fusion vector;

[0029] Pool the point fusion vectors of each point in the subsequence to obtain the feature fusion vector.

[0030] Combined with the first aspect, in the fifth embodiment of the first aspect, converting the empty voxels into dense voxel features is achieved by assigning a value of 0 to all the empty voxels.

[0031] Combined with the first aspect, in the sixth implementation manner of the first aspect, adjusting the size of the initial detection box according to the tracking trajectory to obtain the target detection box specifically includes:

[0032] For each initial detection box in the tracking trajectory, crop the corresponding tracking point cloud at the corresponding position from the corresponding frame in the subsequence;

[0033] Convert the coordinates of the tracking point cloud to the detection box coordinate system with the initial detection box as the origin to obtain the converted point cloud;

[0034] Overlay the converted point clouds of all the initial detection boxes in the trajectory to form an adjusted point cloud;

[0035] Input the adjusted point cloud into the PointNet neural network for regression processing, and output the target size by the PointNet neural network;

[0036] Reassign the target size to all the initial detection boxes in the tracking trajectory to obtain the target detection box.

[0037] According to the second aspect, an embodiment of the present invention further provides a point cloud automatic annotation device based on a neural network. The device includes:

[0038] A sequence division module, configured to obtain the original point cloud sequence and divide the point cloud sequence into several subsequences; each subsequence contains point cloud data of several consecutive frames;

[0039] A coordinate conversion module, configured to convert each point in each frame of point cloud data from the world coordinate system to the relative coordinate system according to the pose of each frame of point cloud data in the world coordinate system, to obtain the relative coordinate value of each point; the origin of the relative coordinate system is the point cloud data of the first frame of the corresponding subsequence, and the relative coordinate value includes the coordinate value and intensity value of each point in the relative coordinate system;

[0040] A voxel encoding module, configured to uniformly divide each frame of converted point cloud data into several preset voxel blocks, and encode the voxel blocks according to the valid points in the voxel blocks to obtain the valid voxels and empty voxels included in each frame of point cloud data; the voxel blocks containing valid points are valid voxels, and the voxel blocks not containing valid points are empty voxels;

[0041] A spatial feature module, configured to extract features from the valid voxels and convert the empty voxels into dense voxel features to obtain the spatial feature vector of each frame of point cloud data, and obtain the spatial feature vector of the subsequence according to the spatial feature vector of each frame of point cloud data;

[0042] A temporal feature module, configured to extract the temporal feature vector of the subsequence by using a neural network;

[0043] A feature fusion module is used to perform feature fusion on the spatial feature vector and the temporal feature vector of the subsequence to obtain a fused feature vector, perform a convolution operation on the fused feature vector to obtain a planar feature map, and generate an initial detection box according to the planar feature map;

[0044] A trajectory tracking module is used to continuously track the initial detection boxes between subsequences according to the spatial relationship of the initial detection boxes between subsequences to obtain tracking trajectories;

[0045] A size optimization module is used to adjust the size of the initial detection box according to the tracking trajectory to obtain a target detection box.

[0046] According to a third aspect, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the neural network-based point cloud automatic annotation method as described in any one of the above are implemented.

[0047] According to a fourth aspect, an embodiment of the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the neural network-based point cloud automatic annotation method as described in any one of the above are implemented.

[0048] The neural network-based point cloud automatic annotation method, device, and medium of the present invention can make full use of the spatial and temporal features of the point cloud. The accuracy and recall rate of the detection results generated based on the fusion of these two features are significantly improved. Using the tracking results to optimize the size of the original point cloud data can obtain a more accurate detection box size. Since the detected object has the continuity of local features in time, when extracting temporal features, the method of dividing the subspace first and then extracting is adopted, rather than calculating and extracting in the entire space, which can greatly reduce the computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings. The drawings are schematic and should not be construed as limiting the present invention in any way. In the drawings:

[0050] Figure 1 A flowchart showing the neural network-based point cloud automatic annotation method provided by the present invention is shown;

[0051] Figure 2 A schematic diagram of a subsequence after voxel encoding in the neural network-based point cloud automatic annotation method provided by the present invention is shown;

[0052] Figure 3Shows a schematic diagram of extracting temporal features from subsequences in the automatic point cloud annotation method based on neural network provided by the present invention;

[0053] Figure 4 Shows a schematic diagram of dividing sub - partitions and performing tiling mapping in the automatic point cloud annotation method based on neural network provided by the present invention;

[0054] Figure 5 Shows a schematic diagram of the subspace feature vectors of subsequences in the automatic point cloud annotation method based on neural network provided by the present invention;

[0055] Figure 6 Shows a schematic diagram of performing temporal feature extraction in the automatic point cloud annotation method based on neural network provided by the present invention;

[0056] Figure 7 Shows a schematic structural diagram of the automatic point cloud annotation method based on neural network provided by the present invention;

[0057] Figure 8 Is a schematic hardware structure diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0058] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] With the continuous acceleration of the iteration cycle in the development of AI products represented by robots and autonomous driving, there are increasingly high requirements for data closed-loop and the efficiency and quality of data annotation on the data side. Among them, the annotation of 3D point cloud data plays a crucial role in current deep learning systems, especially in the R & D processes of systems such as robots, embodied intelligence, and autonomous driving. For example, in robot and autonomous driving systems, lidar is an important sensor often used to achieve functions such as environmental perception, dynamic and static obstacle recognition, map matching, positioning, and mapping. The realization of these functions highly depends on various high-precision and high-performance point cloud algorithms, and training these algorithms requires a large amount of labeled lidar point cloud data; in the pure vision autonomous driving solution based on Bird’s Eye View (BEV) + Transformer, true values in 3D space are also needed to train the algorithm, and 3D point cloud data is exactly an excellent carrier for providing 3D space expression. Therefore, high-efficiency and high-quality point cloud data annotation has become an important link to improve the iteration speed of algorithms and products, optimize the training effect of algorithms, and increase the upper limit of algorithm capabilities.

[0060] For 3D point cloud data, the current method of simply relying on annotators for data annotation generally has problems such as relatively difficult annotation efficiency, annotation fineness, and size estimation. This is because the subjectivity of humans is relatively large, and it is difficult to achieve consistent operations for the annotation of the same object in the time dimension. There are also problems such as jitter and deformation in the trajectory annotation of the object to be annotated.

[0061] Therefore, in the case where the method of simply relying on annotators for data annotation can no longer meet the high-precision and high-efficiency annotation requirements put forward by high-precision algorithms, annotating the data to be annotated through an automatic annotation method and then having annotators perform operations such as attribute supplementation, minor fine-tuning, filling in the blanks, and filtering errors has gradually become an important means to improve the efficiency of point cloud annotation and the quality of annotation results.

[0062] However, some current mainstream automatic annotation methods do not make full use of the temporal features of point clouds. For example, some annotation methods use single-frame detection output, and some annotation methods directly splice adjacent frames. These annotation methods will all result in less than ideal detection accuracy and recall rate, and there are also situations where the size of the detection box is not precise enough.

[0063] In summary, a method for data annotation of 3D point cloud data needs to be provided.

[0064] To solve the above problems, in this embodiment, a method for automatically annotating point clouds based on a neural network is provided, aiming to better annotate three-dimensional point cloud data. The method for automatically annotating point clouds based on a neural network in the embodiment of the present invention can be used in an electronic device, and the electronic device includes but is not limited to a computer, a mobile terminal, etc., Figure 1 is a schematic flowchart of the method for automatically annotating point clouds based on a neural network according to an embodiment of the present invention, as Figure 1 shown, the method may include the following steps:

[0065] S10. Obtain an original point cloud sequence, and divide the point cloud sequence into several subsequences, where each subsequence contains point cloud data of several consecutive frames.

[0066] First, obtain the original continuous point cloud sequence , which are the original point cloud data collected at N consecutive times from 1 to N respectively. Then, divide the original continuous point cloud sequence into several subsequences, and each subsequence contains at most point cloud data of consecutive M frames, that is, point cloud data of M frames at consecutive times.

[0067] In this embodiment, the subsequences are divided in the following way: sequentially extract consecutive M frames of data from the original continuous point cloud sequence according to the time order to obtain a subsequence. In the last extraction, if there are not enough M frames of point cloud data, the actual remaining number of frames is used for extraction, so as to obtain M subsequences.

[0068] For any one of the subsequences , where is the th frame of the th subsequence, it can be seen that in total has frames of point cloud data, is the th frame of point cloud data of the th subsequence.

[0069] S20. According to the pose of each frame of point cloud data in the world coordinate system, convert each point in each frame of point cloud data from the world coordinate system to the relative coordinate system to obtain the relative coordinate system value of each point, where the origin of the relative coordinate system is the first frame of point cloud data of the corresponding subsequence, and the relative coordinate system value includes the coordinate value and intensity value of each point in the relative coordinate system.

[0070] In this embodiment, each frame of point cloud data in the obtained original continuous point cloud sequence corresponds to the pose of this frame of point cloud data relative to the world coordinate system , and the pose is a 4×4 coordinate transformation matrix. Then for the subsequence obtained by partitioning , in this step, it is necessary to uniformly transfer the point cloud coordinate system in the subsequence to the relative coordinate system with the origin set to , and perform intensity normalization operations subsequently.

[0071] Specifically, assume that is the th frame of point cloud data in the th subsequence, and its pose in the world coordinate system is . For any point in the point cloud data, it is necessary to convert this point to the relative coordinate value in the relative coordinate system with the origin set to (the first frame of point cloud data in the th subsequence), where represents the coordinate of point in the x-axis direction in the relative coordinate system, represents the coordinate of point in the y-axis direction in the relative coordinate system, represents the coordinate of point in the z-axis direction in the relative coordinate system, represents the intensity value (scale factor) of point in the relative coordinate system.

[0072] In this embodiment, step S20 specifically includes:

[0073] S21. Determine the pose of each frame of point cloud data in the subsequence relative to the world coordinate system, and convert each point in each point cloud data in the subsequence to homogeneous coordinates.

[0074] First, convert the point to homogeneous coordinates , where represents the coordinate of point in the x-axis direction in the world coordinate system, represents the coordinate of point in the y-axis direction in the world coordinate system, represents the coordinate of point in the z-axis direction in the world coordinate system, represents the intensity value (scale factor) of point in the world coordinate system.

[0075] S22. According to the homogeneous coordinates corresponding to each point, convert each point in the subsequence from the world coordinate system to the relative coordinate system.

[0076] S23. According to the point cloud data corresponding to each point and the pose of the first frame of point cloud data of the subsequence relative to the world coordinate system, perform coordinate transformation on the coordinate values in homogeneous coordinates to determine the coordinate values of each point in the relative coordinate system.

[0077] Next, perform coordinate transformation on homogeneous coordinates to transform the point from the world coordinate system to the relative coordinate system with the origin set to . After coordinate transformation, the homogeneous coordinate of the point is . At this time, the coordinate value of takes the x, y, and z coordinate values of .

[0078] S24. According to the maximum and minimum values of the preset point cloud intensity, perform normalization processing on the intensity values in homogeneous coordinates to determine the intensity values of each point in the relative coordinate system. The relative coordinate values include the coordinate values and intensity values of each point in the relative coordinate system.

[0079] Finally, normalize the intensity value of the point . Assume that the maximum value of the preset point cloud intensity is , and the preset minimum value is . Scale linearly to make the intensity value between 0 and 1, that is: , so as to finally obtain the new coordinate values and intensity values of each point .

[0080] S30. Perform voxel division on the subsequence after coordinate transformation processing, evenly divide each frame of point cloud data into several preset voxel blocks, and perform coding processing on the voxel blocks according to whether they contain valid points to obtain the valid voxels and empty voxels included in each frame of point cloud data, where the voxel blocks containing valid points are valid voxels, and the voxel blocks containing valid points are empty voxels.

[0081] In this embodiment, step S30 specifically includes:

[0082] S31. Determine the effective area to be processed for each frame of point cloud data in each subsequence, evenly divide the effective area into several preset voxel blocks, and determine the points under the point cloud data included in each voxel block. Among them, the preset voxel block is a cubic space with the size set to , represents the size of the voxel block in the x-axis direction, represents the size of the voxel block in the y-axis direction, represents the size of the voxel block in the z-axis direction.

[0083] S32. Use a preset encoder to perform voxel encoding on each voxel block to determine the valid voxels in the voxel block. The input data of the encoder is the relative coordinate values of all the points within the voxel block.

[0084] Suppose the effective region RS to be processed by a certain subsequence is a cubic space, where represents the size of the effective region RS in the x-axis direction, represents the size of the effective region RS in the y-axis direction, represents the size of the effective region RS in the z-axis direction. The sizes of a single preset voxel block in the three directions of the x, y, and z axes are respectively , and . Then, the space represented by the effective region RS can be evenly divided into several voxel blocks. After even division, the numbers of voxel blocks in the three directions of the x, y, and z axes are respectively , and , and = / , = / , = / , totaling preset voxel blocks. Due to the sparse characteristics of the point cloud, only some of these voxel blocks contain valid points, which are valid voxels, and the remaining voxel blocks are empty, which are empty voxels.

[0085] After completing the spatial voxel division, use a preset encoder to perform voxel encoding on each valid voxel. The input data of the encoder is the coordinate values and intensity values of all the points within a voxel block in the relative coordinate system, that is, the relative coordinate values of the points. When the voxel block is a valid voxel, the output data of the encoder is a vector with a length of .

[0086] In this embodiment, the encoder can adopt static encoding in voxelnet or dynamic encoding in the DynamicVoxel network structure. As Figure 2 shown, after completing the voxel encoding, the subsequence is transformed into voxels in space and the encoding of the voxels. Figure 2 In , the gray part is the empty voxel, and the green part is the valid voxel. Each valid voxel contains a vector with a length of

[0087] S40. Extract features from the valid voxels in each frame of point cloud data and convert the empty voxels in each frame of point cloud data into dense voxel features to obtain the spatial feature vector of each frame of point cloud data, and obtain the spatial feature vector of the subsequence according to the spatial feature vector of each frame of point cloud data.

[0088] In this embodiment, a typical structure in the neural network, 3D sparse convolution, is used to extract features from the previously generated voxel encoding. Each convolution operation only operates and outputs on the valid voxels, further improving the efficiency of point cloud data annotation. While continuously performing convolutions, Stride convolution is used to gradually reduce the spatial size of the voxel feature map. Finally, the spatial size of the obtained voxel feature map is 、 and , where represents the size of the voxel feature map in the x-axis direction, represents the size of the voxel feature map in the y-axis direction, represents the size of the voxel feature map in the z-axis direction.

[0089] Then, by assigning a value of 0 to all empty voxels, the sparse voxels are transformed into dense voxel features. Finally, the length of each voxel feature vector is , that is, through this step, each frame of point cloud data in the subsequence obtains a four-dimensional tensor with a size of , denoted as the spatial feature vector , and this tensor is used to represent the spatial voxel features of this frame of point cloud data.

[0090] S50. According to the spatial feature vector corresponding to each frame of point cloud data in the subsequence and using a neural network, extract the temporal feature vector of the subsequence.

[0091] In this embodiment, step S50 specifically includes:

[0092] S51. Uniformly divide the spatial feature vector corresponding to each frame of point cloud data into several sub-partitions and determine the voxel blocks included in each sub-partition.

[0093] S52. Concatenate the feature vectors of all voxel blocks in the sub-partition before and after to obtain the subspace feature vector of the sub-partition.

[0094] S53. Pool the subspace feature vectors of all sub-partitions of the subsequence, and perform temporal feature extraction on all subspace feature vectors included in the subsequence through a neural network to obtain the temporal feature vector.

[0095] To better extract temporal features and minimize the computational amount as much as possible, this step first performs subspace partitioning on the spatial feature vector. Specifically, it is uniformly divided into Sub - partitions, each sub - partition has dimensions in the x, y, and z - axis directions of , and . As Figure 3 shown, Figure 3 each color in = / / / .

[0096] For each sub - partition, the feature vectors of all voxel blocks in the sub - partition are concatenated front - to - back to obtain a feature vector of length = . Thus, for each frame of point - cloud data, sub - space feature vectors of length can be obtained. As Figure 4 shown, Figure 4 a sub - space on the left is mapped to Figure 4 a voxel on the right, Figure 4 the length of the feature vector contained in each voxel on the left is , Figure 4 the length of the feature vector contained in each voxel on the right is .

[0097] A subsequence contains M frames of point - cloud data. Thus, a subsequence can generate sub - space feature vectors, with the vector length of . As Figure 5 shown, where is the time dimension, representing M frames of point - cloud data at different times in a subsequence, while is the space dimension, representing that each frame has sub - spaces, with the feature vector length of . Thus, it can be regarded as a three - dimensional tensor of

[0098] As Figure 6 shown, by transposing the three - dimensional tensor in the time dimension and the space dimension, it is regarded as time - series sequences of size M. Each M - dimensional time - series sequence is subjected to time - series feature extraction through the Transformer or Recurrent Neural Network (RNN) in the neural network to obtain a feature vector (with a length of ). Finally, feature vectors of length The temporal feature vector, denoted as the temporal feature vector .

[0099] S60. Feature fusion is performed on the spatial feature vector and the temporal feature vector of the subsequence to obtain a fused feature vector. Several convolution operations are performed on the fused feature vector until a planar feature map with a size of 1 in the z-axis direction is formed, and an initial detection box is generated based on the planar feature map.

[0100] In this embodiment, the temporal feature vector and the spatial feature vector are feature concatenated to obtain a fused feature vector . It can be understood that the size of the fused feature vector is the same as that of the spatial feature vector .

[0101] Preferably, the method of feature concatenation is as follows: for any point in the fused feature vector , first obtain the point spatial feature vector (the length of the feature vector is ) at the same coordinate in the spatial feature vector according to the coordinates of this point, and then divide and tile-map the coordinates according to the sub-partition shown in Figure 4 to obtain the point temporal feature vector (the length of the feature vector is ) of this coordinate in the temporal feature vector . Concatenate these two point vectors to obtain a new feature vector with a length of = + . Finally, the size of the fused feature is .

[0102] After obtaining the fused feature vector of each frame of point cloud data, several 3D convolution operations are performed. Among them, the 3D convolution has a convolution stride of stide = 2 in the z-axis direction until the size in the z-axis direction becomes 1, forming a planar feature map. Based on this planar feature map, an initial detection box is generated. The generation method of the initial detection box is to use several regression neural networks to sequentially output parameters such as the center heat map of the detection box, the classification category, the size of the target box, and the orientation. In this way, a corresponding detection box is generated for each subsequence.

[0103] The 3D convolution operation is actually a convolution operation that will be performed in the three directions of the x, y, and z axes, and the convolution stride stide = 2 is set in the z-axis direction. In this way, each time after a convolution, the size in the z-axis direction will be reduced by half until the size in the z-axis direction becomes 1, which becomes a planar feature map.

[0104] S70. Continuously track the initial detection frames among subsequences according to the spatial relationship between the initial detection frames among subsequences to obtain tracking trajectories.

[0105] Use the multi-object detection method to obtain the spatial relationship between the initial detection frames among subsequences. Finally, a set of tracking trajectories can be obtained. Each trajectory consists of a series of detection frames in different frames, representing the positions of the same target object in different frames.

[0106] S80. Adjust the size of the initial detection frame according to the tracking trajectory to obtain the target detection frame.

[0107] In this embodiment, step S80 specifically includes:

[0108] Crop the corresponding tracking point cloud at the corresponding position from the corresponding frame in the subsequence for each initial detection frame in the tracking trajectory, and then convert the coordinates of these tracking point clouds to the detection frame coordinate system with the initial detection frame as the origin to obtain the converted point cloud. After that, stack the converted point clouds of all initial detection frames in the tracking trajectory to form a new adjusted point cloud. Next, input the adjusted point cloud into the PointNet neural network of the point cloud to finally regress the more refined size of the detection frame, that is, the target size. Assign the generated target size to all initial detection frames in this trajectory to obtain a more refined size result, that is, obtain the target detection frame.

[0109] The detection frames in the generated trajectory are converted into the format required by the annotation platform or algorithm according to the requirements of the annotation platform and saved. Subsequently, the annotation platform can directly read this data to display the annotation result, and algorithm researchers can directly read the format for training or testing.

[0110] The automatic point cloud annotation method based on neural network of the present invention can make full use of the spatial and temporal features of the point cloud. The accuracy and recall rate of the detection results generated based on the fusion of these two features are significantly improved. Using the tracking results to optimize the size of the original point cloud data can obtain a more accurate detection frame size. Since the detected object has the continuity of local features in time, when extracting temporal features, the method of dividing the subspace first and then extracting is adopted, rather than calculating and extracting in the whole space, which can greatly reduce the computational complexity.

[0111] Next, the device provided by the embodiment of the present invention will be described. The device described below can be correspondingly referred to the method described above.

[0112] Please refer to Figure 7 , Figure 7 which shows the structural schematic diagram of the automatic point cloud annotation device based on neural network of the embodiment of the present invention. The device may include:

[0113] The sequence division module 10 is used to obtain the original point cloud sequence and divide the point cloud sequence into several subsequences, where each subsequence contains the point cloud data of several consecutive frames.

[0114] First, obtain the original continuous point cloud sequence , which are the original point cloud data collected at N consecutive moments from 1 to N respectively. Then, divide the original continuous point cloud sequence into several subsequences, and each subsequence contains at most the point cloud data of continuous M frames, that is, the point cloud data of M frames at consecutive moments.

[0115] In this embodiment, the division method of the subsequence is to sequentially extract continuous M-frame data from the original continuous point cloud sequence to obtain a subsequence. In the last extraction, if there are not enough M-frame point cloud data, the actual remaining number of frames is used for extraction, so as to obtain M subsequences.

[0116] For any one of the subsequences , where is the th frame of the subsequence. It can be seen that there are frames of point cloud data in total, is the th frame of point cloud data of the subsequence.

[0117] The coordinate transformation module 20 is used to convert each point in each frame of point cloud data from the world coordinate system to the relative coordinate system according to the pose of each frame of point cloud data in the world coordinate system, so as to obtain the relative coordinate value of each point. The origin of the relative coordinate system is the first frame of point cloud data of the corresponding subsequence, and the relative coordinate value includes the coordinate value and intensity value of each point in the relative coordinate system.

[0118] In this embodiment, each frame of point cloud data in the obtained original continuous point cloud sequence corresponds to the pose of this frame of point cloud data relative to the world coordinate system , and the pose is a 4×4 coordinate transformation matrix. Then for the divided subsequence , in this step, it is necessary to uniformly transfer the point cloud coordinate system in of the subsequence to the relative coordinate system with the origin set to , and perform intensity normalization operation.

[0119] Specifically, assume that is the th frame in the Frame point cloud data, whose pose in the world coordinate system is , for any point in the point cloud , it is necessary to convert this point based on the pose to the relative coordinate value in the relative coordinate system with the origin set to (the first frame of point cloud data in the th subsequence), where , represents the coordinate of point in the x-axis direction in the relative coordinate system, represents the coordinate of point in the y-axis direction in the relative coordinate system, represents the coordinate of point in the z-axis direction in the relative coordinate system, represents the intensity value (scale factor) of point in the relative coordinate system.

[0120] The voxel encoding module 30 is used to uniformly divide each frame of point cloud data into several preset voxel blocks, and encode the voxel blocks according to whether the voxel blocks contain valid points, so as to obtain the valid voxels and empty voxels included in each frame of point cloud data, where the voxel blocks containing valid points are valid voxels, and the voxel blocks not containing valid points are empty voxels.

[0121] The spatial feature module 40 is used to extract features from the valid voxels in each frame of point cloud data and convert the empty voxels into dense voxel features, so as to obtain the spatial feature vector of each frame of point cloud data, and obtain the spatial feature vector of the subsequence according to the spatial feature vector of each frame of point cloud data.

[0122] In this embodiment, a typical structure in the neural network, 3D sparse convolution, is used to extract features from the previously generated voxel encoding. Each convolution operation only operates and outputs on the valid voxels, which further improves the efficiency of point cloud data annotation. While continuously performing convolution, Stride convolution is used to gradually reduce the spatial size of the voxel feature map. Finally, the spatial size of the obtained voxel feature map is , and , where represents the size of the voxel feature map in the x-axis direction, represents the size of the voxel feature map in the y-axis direction, represents the size of the voxel feature map in the z-axis direction.

[0123] Then, the sparse voxels are converted into dense voxel features by assigning 0 values to all empty voxels. Finally, the length of each voxel feature vector is , that is, a four-dimensional tensor with a size of is obtained from each frame of point cloud data in this subsequence of steps, denoted as the spatial feature vector , and this tensor is used to represent the spatial voxel features of this frame of point cloud data.

[0124] The temporal feature module 50 is used to extract the temporal feature vector of the subsequence according to the spatial feature vector corresponding to each frame of point cloud data in the subsequence and by using a neural network.

[0125] The feature fusion module 60 is used to perform feature fusion on the spatial feature vector and the temporal feature vector of the subsequence to obtain a fused feature vector, perform convolutional processing on the fused feature vector in the z-axis direction until a planar feature map with a size of 1 in the z-axis direction is formed, and generate an initial detection box according to the planar feature map.

[0126] In this embodiment, the temporal feature vector and the spatial feature vector are concatenated to obtain a fused feature vector . It can be understood that the size of the fused feature vector is the same as that of the spatial feature vector .

[0127] Preferably, the method of feature concatenation is as follows: for any point in the fused feature vector , first obtain the point spatial feature vector (the length of the feature vector is ) at the same coordinate in the spatial feature vector according to the coordinates of this point, and then divide and tile-map the coordinates according to the sub-partition shown in Figure 4 to obtain the point temporal feature vector (the length of the feature vector is ) of this coordinate in the temporal feature vector . Concatenate these two point vectors to obtain a new feature vector with a length of = + . Finally, the size of the fused feature is .

[0128] After obtaining the fused feature vector of each frame of point cloud data, several 3D convolution operations are performed. Among them, the 3D convolution has a convolution stride of stide = 2 in the z-axis direction until the size in the z-axis direction becomes 1, forming a planar feature map. According to this planar feature map, an initial detection box is generated. The generation method of the initial detection box is to use several regression neural networks to sequentially output parameters such as the center heat map of the detection box, the classification category, the size of the target box, and the orientation. In this way, a corresponding detection box is generated for each subsequence.

[0129] The 3D convolution operation is actually a convolution operation that is performed in three directions: the x, y, and z axes. The convolution stride is set to 2 in the z-axis direction. In this way, every time a convolution is performed, the size in the z-axis direction is reduced by half until the size in the z-axis direction becomes 1, and it becomes a planar feature map.

[0130] The trajectory tracking module 70 is used to continuously track the initial detection frames between subsequences based on the spatial relationship between the initial detection frames of subsequences, and obtain tracking trajectories.

[0131] Using the method of multi-object detection to obtain the spatial relationship between the initial detection frames of subsequences, a set of tracking trajectories can be finally obtained. Each trajectory is composed of a series of detection frames in different frames, representing the positions of the same target object in different frames.

[0132] The size optimization module 80 is used to adjust the size of the initial detection frame according to the tracking trajectory to obtain the target detection frame.

[0133] The detection frames in the generated trajectories are converted into the format required by the annotation platform or algorithm according to the requirements of the annotation platform and saved. Subsequently, the annotation platform can directly read this data to display the annotation results, and algorithm researchers can directly read the format for training or testing.

[0134] The neural network-based point cloud automatic annotation device of the present invention can make full use of the spatial and temporal features of the point cloud. The accuracy and recall rate of the detection results generated based on the fusion of these two features are both significantly improved. Using the tracking results to optimize the size of the original point cloud data can obtain more accurate detection frame sizes. Since the detected objects have the continuity of local features in time, when extracting temporal features, the method of dividing the subspace first and then extracting is adopted, rather than calculating and extracting in the entire space, which can greatly reduce the computational complexity.

[0135] Figure 8 Illustrates a schematic diagram of the physical structure of an electronic device, as Figure 8 shown. The electronic device may include: a processor 810 (processor), a communication interface 820 (Communications Interface), a memory 830 (memory), and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 can call the logical commands in the memory 830 to execute the neural network-based point cloud automatic annotation method, and this method includes:

[0136] Obtain the original point cloud sequence and divide the point cloud sequence into several subsequences; each subsequence contains the point cloud data of several consecutive frames;

[0137] According to the pose of each frame of point cloud data in the world coordinate system, convert each point in each frame of point cloud data from the world coordinate system to the relative coordinate system to obtain the relative coordinate system value of each point; the origin of the relative coordinate system is the first frame of point cloud data of the corresponding subsequence, and the relative coordinate system value includes the coordinate value and intensity value of each point in the relative coordinate system;

[0138] Evenly divide the converted point cloud data of each frame into several preset voxel blocks, and encode the voxel blocks according to the valid points in the voxel blocks to obtain the valid voxels and empty voxels included in the point cloud data of each frame; the voxel block containing valid points is a valid voxel, and the voxel block not containing valid points is an empty voxel;

[0139] Extract features from the valid voxels and convert the empty voxels into dense voxel features to obtain the spatial feature vector of each frame of point cloud data, and obtain the spatial feature vector of the subsequence according to the spatial feature vector of each frame of point cloud data;

[0140] Use a neural network to extract the temporal feature vector of the subsequence;

[0141] Fuse the spatial feature vector and temporal feature vector of the subsequence to obtain a fused feature vector, perform a convolution operation on the fused feature vector to obtain a planar feature map, and generate an initial detection box according to the planar feature map;

[0142] According to the spatial relationship of the initial detection boxes between subsequences, continuously track the initial detection boxes between each subsequence to obtain a tracking trajectory;

[0143] Adjust the size of the initial detection box according to the tracking trajectory to obtain the target detection box.

[0144] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memor), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0145] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the neural network-based point cloud automatic annotation method provided by the above-mentioned various methods. The method includes:

[0146] Obtain the original point cloud sequence and divide the point cloud sequence into several subsequences; each subsequence contains point cloud data of several consecutive frames;

[0147] According to the pose of each frame of point cloud data in the world coordinate system in the subsequence, convert each point in each frame of point cloud data from the world coordinate system to the relative coordinate system to obtain the relative coordinate system value of each point; the origin of the relative coordinate system is the first frame of point cloud data of the corresponding subsequence, and the relative coordinate system value includes the coordinate value and intensity value of each point in the relative coordinate system.

[0148] Evenly divide the converted point cloud data of each frame into several preset voxel blocks, and encode the voxel blocks according to the valid points in the voxel blocks to obtain the valid voxels and empty voxels included in each frame of point cloud data; the voxel blocks containing valid points are valid voxels, and the voxel blocks not containing valid points are empty voxels;

[0149] Extract features from the valid voxels and convert the empty voxels into dense voxel features to obtain the spatial feature vector of each frame of point cloud data, and obtain the spatial feature vector of the subsequence according to the spatial feature vector of each frame of point cloud data;

[0150] Use a neural network to extract the temporal feature vector of the subsequence;

[0151] Fuse the spatial feature vectors and temporal feature vectors of the subsequences to obtain fused feature vectors, perform a convolution operation on the fused feature vectors to obtain a planar feature map, and generate initial detection boxes based on the planar feature map;

[0152] Continuously track the initial detection boxes between subsequences according to the spatial relationship of the initial detection boxes between subsequences to obtain tracking trajectories;

[0153] Adjust the size of the initial detection boxes according to the tracking trajectories to obtain target detection boxes.

[0154] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it is configured to execute the above-provided method for automatically annotating point clouds based on a neural network, and the method includes:

[0155] Obtain an original point cloud sequence and divide the point cloud sequence into several subsequences; each subsequence contains point cloud data of several consecutive frames;

[0156] According to the pose of each frame of point cloud data in the world coordinate system in the subsequence, convert each point in each frame of point cloud data from the world coordinate system to the relative coordinate system to obtain the relative coordinate values of each point; the origin of the relative coordinate system is the first frame of point cloud data of the corresponding subsequence, and the relative coordinate values include the coordinate values and intensity values of each point in the relative coordinate system;

[0157] Evenly divide the converted point cloud data of each frame into several preset voxel blocks, and perform encoding processing on the voxel blocks according to the valid points in the voxel blocks to obtain the valid voxels and empty voxels included in each frame of point cloud data; the voxel blocks containing valid points are valid voxels, and the voxel blocks not containing valid points are empty voxels;

[0158] Extract features from the valid voxels and convert the empty voxels into dense voxel features to obtain the spatial feature vectors of each frame of point cloud data, and obtain the spatial feature vectors of the subsequences according to the spatial feature vectors of each frame of point cloud data;

[0159] Use a neural network to extract the temporal feature vectors of the subsequences;

[0160] Fuse the spatial feature vectors and temporal feature vectors of the subsequences to obtain fused feature vectors, perform a convolution operation on the fused feature vectors to obtain a planar feature map, and generate initial detection boxes based on the planar feature map;

[0161] Continuously track the initial detection boxes between subsequences according to the spatial relationship of the initial detection boxes between subsequences to obtain tracking trajectories;

[0162] Adjust the size of the initial detection box according to the tracking trajectory to obtain the target detection box.

[0163] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A point cloud automatic annotation method based on neural network, characterized in that: The method comprises: Obtain the original point cloud sequence, and divide the point cloud sequence into several subsequences; each subsequence contains several consecutive frames of point cloud data; According to the pose of each frame of point cloud data in the subsequence in the world coordinate system, each point in each frame of point cloud data is converted from the world coordinate system to the relative coordinate system to obtain the relative coordinate system value of each point; the origin of the relative coordinate system is the first frame of point cloud data in the corresponding subsequence, and the relative coordinate system value includes the coordinate value and intensity value of each point in the relative coordinate system; Each frame of point cloud data after conversion is evenly divided into a number of preset voxel blocks, and the voxel blocks are encoded according to the valid points in the voxel blocks to obtain the valid voxels and empty voxels contained in each frame of point cloud data; the voxel blocks containing valid points are valid voxels, and the voxel blocks without valid points are empty voxels; Extract features from valid voxels and convert empty voxels into dense voxel features to obtain the spatial feature vector of each frame of point cloud data, and obtain the spatial feature vector of the subsequence based on the spatial feature vector of each frame of point cloud data; Extract the temporal feature vector of the subsequence using neural network; Perform feature fusion on the spatial feature vector and the temporal feature vector of the subsequence to obtain a fused feature vector, perform convolution operation on the fused feature vector to obtain a plane feature map, and generate an initial detection frame based on the plane feature map; According to the spatial relationship between the initial detection frames of the subsequences, the initial detection frames between the subsequences are continuously tracked to obtain the tracking trajectory; The size of the initial detection frame is adjusted according to the tracking trajectory to obtain the target detection frame.

2. The method for automatic point cloud annotation based on neural network according to claim 1, characterized in that: According to the pose of each frame of point cloud data in the subsequence in the world coordinate system, each point in each frame of point cloud data is converted from the world coordinate system to the relative coordinate system to obtain the relative coordinate system value of each point, specifically including: Determine the pose of each frame of point cloud data in the subsequence relative to the world coordinate system, and convert each point under each point cloud data in the subsequence into homogeneous coordinates; According to the homogeneous coordinates corresponding to each point, each point in the subsequence is converted from the world coordinate system to the relative coordinate system; According to the point cloud data corresponding to each point and the position of the first frame of the subsequence point cloud data relative to the world coordinate system, the coordinate values ​​in the aligned secondary coordinates are converted to determine the coordinate value of each point in the relative coordinate system; According to the preset maximum and minimum values ​​of the point cloud intensity, the intensity values ​​in the secondary coordinates are normalized to determine the intensity value of each point in the relative coordinate system.

3. The method for automatic point cloud annotation based on neural network according to claim 1, characterized in that: The converted point cloud data of each frame is evenly divided into a number of preset voxel blocks, and the voxel blocks are encoded according to the valid points in the voxel blocks to obtain the valid voxels and empty voxels contained in each frame of point cloud data, specifically including: Determine the effective area to be processed for each frame of point cloud data in each subsequence, and evenly divide the effective area into a number of preset voxel blocks, and determine the points under the point cloud data contained in each voxel block; the preset voxel block is set to a size of Cube space, represents the size of the voxel block in the x-axis direction, represents the size of the voxel block in the y-axis direction, Indicates the size of the voxel block in the z-axis direction; Each voxel block is voxel-encoded using a preset encoder to determine valid voxels in the voxel block; the input data of the encoder is the relative coordinate system values ​​of all points in the voxel block.

4. The method for automatic point cloud annotation based on neural network according to claim 3, characterized in that: The extracting of the time series feature vector of the subsequence by using a neural network specifically includes: The spatial feature vector corresponding to each frame of point cloud data is evenly divided into several sub-partitions, and the voxel blocks contained in each sub-partition are determined; Concatenate the feature vectors of all voxel blocks in the subpartition forward and backward to obtain the subspace feature vector of the subpartition; The subspace feature vectors of all subpartitions of the subsequence are collected, and the time series features of all subspace feature vectors contained in the subsequence are extracted through a neural network to obtain a time series feature vector.

5. The method for automatic point cloud annotation based on neural network according to claim 1, characterized in that: The feature fusion method is as follows: Obtain the point spatial feature vector of each point in the spatial feature vector and the point temporal feature vector of each point in the temporal feature vector, and concatenate the point spatial feature vector and the point temporal feature vector to obtain a point fusion vector; The point fusion vectors of each point in the subsequence are collected to obtain the feature fusion vector.

6. The method for automatic point cloud annotation based on neural network according to claim 1, characterized in that: Converting empty voxels to dense voxel features is achieved by assigning 0 values ​​to all empty voxels.

7. The method for automatic point cloud annotation based on neural network according to claim 1, characterized in that: The step of adjusting the size of the initial detection frame according to the tracking trajectory to obtain the target detection frame specifically includes: Using each initial detection frame in the tracking trajectory, a tracking point cloud at a corresponding position is cut out from the corresponding frame in the subsequence; Convert the coordinates of the tracking point cloud to the detection frame coordinate system with the initial detection frame as the origin to obtain the converted point cloud; Superimpose the transformed point clouds of all initial detection boxes in the trajectory to form an adjusted point cloud; Input the adjusted point cloud into the point cloud PointNet neural network for regression processing, and the target size is obtained by the output of the point cloud PointNet neural network; Reassign the target size to all initial detection boxes in the tracking trajectory to obtain the target detection box.

8. A point cloud automatic annotation device based on a neural network, characterized in that: The device comprises: The sequence division module is used to obtain the original point cloud sequence and divide the point cloud sequence into several subsequences; each subsequence contains several consecutive frames of point cloud data; The coordinate conversion module is used to convert each point in each frame of point cloud data from the world coordinate system to the relative coordinate system according to the position of each frame of point cloud data in the subsequence in the world coordinate system, and obtain the relative coordinate system value of each point; the origin of the relative coordinate system is the first frame of point cloud data in the corresponding subsequence, and the relative coordinate system value includes the coordinate value and intensity value of each point in the relative coordinate system; The voxel encoding module is used to evenly divide each frame of point cloud data after conversion into a number of preset voxel blocks, and encode the voxel blocks according to the valid points in the voxel blocks to obtain the valid voxels and empty voxels contained in each frame of point cloud data; the voxel blocks containing valid points are valid voxels, and the voxel blocks not containing valid points are empty voxels; The spatial feature module is used to extract features from valid voxels and convert empty voxels into dense voxel features to obtain the spatial feature vector of each frame of point cloud data, and obtain the spatial feature vector of the subsequence according to the spatial feature vector of each frame of point cloud data; A time series feature module is used to extract the time series feature vector of the subsequence using a neural network; A feature fusion module is used to fuse the spatial feature vectors and the temporal feature vectors of the subsequences to obtain a fused feature vector, perform a convolution operation on the fused feature vector to obtain a plane feature map, and generate an initial detection frame based on the plane feature map; A trajectory tracking module is used to continuously track the initial detection frames between the subsequences according to the spatial relationship between the initial detection frames between the subsequences to obtain a tracking trajectory; The size optimization module is used to adjust the size of the initial detection frame according to the tracking trajectory to obtain the target detection frame.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the point cloud automatic labeling method based on a neural network as described in any one of claims 1 to 7 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for automatic point cloud annotation based on a neural network as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • 3D target detection method based on convolutional long short-term memory network

    CN116758534A

  • Multi-mode automatic driving three-dimensional target tracking method based on space-time fusion

    CN119274167A