A laser point cloud data interval labeling method and system for intelligent driving

By using the interval annotation method to manually annotate key frames in intelligent driving, and then using detection algorithms and feature fusion to generate high-precision annotation boxes, the problem of low efficiency and low accuracy of traditional annotation methods is solved, and efficient and accurate laser point cloud data annotation is achieved.

CN121640251BActive Publication Date: 2026-07-28YIQING AUTOMOTIVE SAFETY SYST (SUZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YIQING AUTOMOTIVE SAFETY SYST (SUZHOU) CO LTD
Filing Date
2025-12-25
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing technologies for labeling laser point cloud data in intelligent driving are inefficient and inaccurate. Traditional manual labeling methods are labor-intensive and inconsistent, while automated pre-labeling methods frequently result in false detections and missed detections in complex scenarios, making it difficult to meet high-precision requirements.

Method used

An interval annotation method is adopted, which involves manually annotating key frames and generating interpolation boxes and detection boxes for non-key frames using a 3D target detection algorithm. These are then fused with geometric features and temporal motion features and input into a two-stage fine-tuning network for optimization to generate high-precision annotation boxes.

Benefits of technology

It significantly reduces the cost of manual annotation, improves annotation efficiency and accuracy, maintains high recall and low false detection rate in complex scenarios, and provides stable, high-quality annotated data for training intelligent driving models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640251B_ABST
    Figure CN121640251B_ABST
Patent Text Reader

Abstract

The application discloses a kind of laser point cloud data interval marking method and system for intelligent driving, three-dimensional target frame of unmarked frame is inferred by artificial marking information based on small amount of key frame, realize the efficient automatic marking of long time sequence point cloud data.The method first selects interval frame as key frame, generates detection frame on the basis of multiple frame point cloud superposition, and forms interpolation frame in combination with the time sequence interpolation result of key frame true value under world coordinate system;Subsequently, the detection frame and the interpolation frame form a candidate region, the coarse frame and its point cloud are encoded and optimized by two-stage fine-tuning network, multi-dimensional information is extracted using proxy point features, geometric relationship features and time sequence motion features, thereby outputting accurate three-dimensional true value frame.The application can significantly reduce the construction cost of large-scale point cloud data set for intelligent driving while maintaining the accuracy comparable to manual annotation, and is suitable for training and evaluation of various automatic driving perception tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent driving and three-dimensional environment perception technology, and in particular relates to a laser point cloud data interval annotation method and system for intelligent driving. Background Technology

[0002] With the rapid evolution of intelligent driving technology, environmental perception models are increasingly reliant on high-precision training data. This is especially true for tasks such as multi-target detection, tracking, and behavior understanding in 3D point cloud scenarios, which place unprecedented demands on the completeness, accuracy, and consistency of data annotation. In typical autonomous driving data acquisition processes, LiDAR can continuously acquire a large amount of temporal point cloud information. However, if this data is to be used for model training, it must undergo high-precision frame-by-frame annotation. Due to the characteristics of point cloud data, such as high sparsity, high dimensionality, large variations in object scale, and severe occlusion, traditional manual annotation methods are not only extremely labor-intensive but also struggle to maintain stable consistency among different annotators, leading to fluctuations in annotation accuracy and consequently affecting the generalization performance of downstream models.

[0003] To improve annotation efficiency, the industry has gradually introduced automated or semi-automated pre-annotation methods. These methods generate rough bounding boxes using detection models, which are then manually corrected, thus reducing manual workload. However, the performance of these methods is usually limited by the accuracy of the detection model itself. When dealing with long-term point cloud data, the pre-annotated bounding boxes are easily affected by factors such as occlusion, sparse scanning, and changes in the LiDAR viewpoint, leading to frequent false positives and false negatives, and the burden of manual correction remains heavy. In addition, automated pre-annotation models are often limited by the distribution of training data, making it difficult to maintain stable performance in complex scenes (such as distant targets, targets with weak reflectivity, or heavily occluded areas), resulting in overall annotation quality that fails to meet the high-precision requirements of intelligent driving systems.

[0004] Temporal point clouds contain a large amount of continuity and motion patterns, which can provide potential advantages for improving annotation efficiency. However, traditional annotation methods rely almost entirely on independent frame-by-frame processing, failing to fully utilize the spatiotemporal consistency of objects across frames. Although some studies have attempted to use linear interpolation or trajectory tracking methods to assist annotation, these methods often fail to accurately infer the position and size of objects in unannotated frames when faced with complex speed changes, turning behaviors, occlusion disappearance and reappearance in real driving scenarios, making it difficult to directly use the interpolation results as reliable annotations.

[0005] For real-world autonomous driving data, the time span is long and the frame rate is high, leading to an exponential increase in the cost of complete frame-by-frame annotation. Existing technologies based on independent detection or simple interpolation struggle to achieve an effective balance between cost and accuracy, failing to significantly reduce manual input or consistently provide high-quality annotation results that can replace manual annotation. Against this backdrop, there is an urgent need for a novel annotation technology that can fully leverage cross-frame correlations in long-term time-series point cloud data, improve automated inference capabilities, significantly reduce manual workload, and still guarantee stable, high-precision output. This would enable intelligent driving perception systems to acquire high-quality annotated data suitable for model training at a lower cost. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides a method and system for laser point cloud data interval annotation for intelligent driving. Specifically, the technical solution provided by this invention is as follows: A method for interval annotation of laser point cloud data for intelligent driving includes the following steps: S1. Select several point cloud frames from the laser point cloud sequence to be labeled as key frames according to the set frame interval for manual labeling, obtain the three-dimensional bounding box parameters of each target in each key frame, and the unselected point cloud frames are regarded as non-key frames. S2. For the 3D bounding box parameters of the same target in adjacent key frames, generate the 3D interpolation box of the same target in each non-key frame between the adjacent key frames, and obtain the 3D interpolation box parameters. S3. Use a 3D target detection algorithm to perform target detection on each non-key frame, generate 3D detection boxes for each target in the non-key frame, and obtain the parameters of the 3D detection boxes. S4. The 3D interpolation boxes and 3D detection boxes in the same non-key frame are fused and filtered to remove redundant boxes and obtain the 3D rough boxes of each target in the non-key frame. S5. For each 3D rough bounding box in each non-key frame, geometric feature encoding is performed based on the spatial positional relationship between each point cloud in the rough bounding box and the corner and center points of the rough bounding box to obtain geometric features that reflect the shape of the target; at the same time, temporal motion features that can characterize the target's motion trend and cross-frame correlation are extracted. S6. After fusing geometric and motion features, input them into the trained two-stage fine-tuning network to regress the center position, size, and orientation of each rough box in non-keyframes, and obtain the optimized 3D box, which is then used as the annotation box.

[0007] Furthermore, in step S2, for the same target i Assume that the 3D bounding box parameters in two adjacent keyframes are as follows:

[0008]

[0009] in, k 1 and k 2 as the target i The time index of two adjacent keyframes. and The target i In keyframe k 1 and k The center point coordinates of the 3D bounding box in Figure 2 and The target i In keyframe k 1 and k 2. The length, width, height, and orientation angle of the 3D bounding box; Then for those in k 1 and k Non-keyframes between 2 t The target i The three-dimensional interpolation frame parameters are obtained by linear interpolation based on time. ,in: Let be the coordinates of the center point of the 3D interpolation frame, and , ,

[0010] Let the length, width, and height of this 3D interpolation frame be , and , , , Let be the orientation angle of the 3D interpolation frame, and .

[0011] Furthermore, in step S3, before performing target detection on each non-critical frame using the 3D target detection algorithm, the insufficient information in a single frame caused by the sparse point cloud of the lidar is alleviated by superimposing multiple point cloud frames. That is, for a non-critical frame that needs to be detected, the frame and several neighboring point cloud frames before and after it are selected and superimposed, and the superimposed point cloud frames are input into the 3D target detection algorithm to obtain the 3D detection box of each target.

[0012] Preferably, the three-dimensional target detection algorithm in step S3 uses the CenterPoint network.

[0013] Further, in step S4, assuming that the 3D interpolation boxes and 3D detection boxes of each target in a non-key frame constitute an interpolation box set and a detection box set respectively, the fusion filtering is as follows: For any detection box in the detection box set, if there is an interpolation box in the interpolation box set with an IoU greater than a set threshold, then the two are considered as bounding boxes of the same target and the detection box is retained, or the box obtained by weighted averaging of the two is taken as the replacement box and retained; otherwise, the detection box is directly retained. For any interpolation box in the interpolation box set, if there is a detection box in the detection box set with an IoU greater than a set threshold, then the two are considered as bounding boxes of the same target and the detection box is retained, or the box obtained by weighted averaging of the two is taken as the replacement box and retained; otherwise, the interpolation box is directly retained. Finally, all the retained bounding boxes are used as the 3D rough bounding boxes of each target in the non-keyframe.

[0014] Furthermore, in step S5: The rough box is represented by 9 points consisting of its center point and 8 corner points. For each target point cloud in the rough box, the difference between its coordinates and each of the frame points in each dimension is calculated, and the coordinates of the target point cloud itself are attached to form the geometric feature vector of the target point cloud. The geometric feature vectors of all target point clouds in the rough box together constitute the geometric features of the rough box. Alternatively, the set abstraction method in PoinNet++ can be used to aggregate the geometric features onto surrogate points. With each surrogate point as the center, points in the neighborhood are searched to form local point groups for grouping. Then, local features are extracted by a multilayer perceptron and aggregated using max pooling to obtain the final geometric features. The surrogate points are selected by dividing the rough box into several equal parts along its coordinate axes, with each division point corresponding to a surrogate point.

[0015] Further, in step S5, the temporal motion feature is as follows: for the nine points of the rough box, calculate the difference of each dimension coordinate between each proxy point of the rough box for the same target in the previous frame or multiple previous frames, and attach the corresponding point cloud frame time code to jointly constitute the temporal motion feature; or the temporal motion feature is further input into the multilayer perceptron to output the final temporal motion feature.

[0016] Furthermore, in step S6, the geometric features and motion features are directly concatenated into vectors or their corresponding dimensions are added to obtain the fused geometric temporal motion features, which are then input into the trained two-stage fine-tuning network. The two-stage fine-tuning network includes an input module, an attention enhancement module, and a regression output module. The input module is used to input the geometric temporal motion features corresponding to each rough box. The attention enhancement module uses the Transformer attention mechanism to interactively fuse the geometric temporal motion features. By learning the importance distribution between different feature dimensions, it adaptively adjusts geometric or motion cues that are highly correlated with the box optimization, thereby improving the regression stability in complex scenes. The regression output module includes a regression head composed of a multilayer perceptron, which is used to regress the center position, size, and orientation of the rough box based on the output of the attention enhancement module.

[0017] Preferably, the regression output module includes multiple regression heads, which are used to predict the center position offset, size correction, and orientation angle correction of the rough box, respectively. The prediction output results are superimposed on the corresponding parameters of the rough box in the form of residuals, thereby obtaining the refined three-dimensional bounding box as the annotation box of the corresponding target.

[0018] A laser point cloud data interval annotation system based on the above method, the system includes the following modules: The data input module is used to receive and organize the laser point cloud frame sequence arranged in chronological order, store the point cloud frames in a structured manner, and provide a unified time index and access interface. The keyframe management and manual annotation module is used to select multiple keyframes from the point cloud frame sequence according to a preset strategy, and to annotate the targets in the keyframes with three-dimensional bounding boxes through an external annotation platform. The interpolation inference module is used to perform interpolation inference on non-key frames based on the geometric parameter changes of the same target in adjacent key frames, and generate the corresponding set of 3D interpolation boxes. The temporal detection module is used to select the point cloud of each non-key frame and several neighboring frames before and after it for fusion processing, and input the fused point cloud frame into the three-dimensional target detection algorithm to output the three-dimensional detection box set of the non-key frame. The candidate ROI fusion and point cloud extraction module is used to fuse the interpolation box and the detection box to construct a set of candidate coarse boxes for non-key frames, and extract the corresponding point cloud in each coarse box. The feature encoding module is used to calculate geometric features based on the point cloud data and bounding box feature points within the rough box, further extract temporal motion features by combining cross-frame motion information, and fuse and encode the geometric features and temporal features. The two-stage bounding box optimization module receives the fused encoded features and inputs them into a bounding box optimization network that includes an attention mechanism and a residual regression structure. It performs fine regression on the center position, size parameters, and orientation angle of each rough bounding box, and outputs accurate 3D bounding box results for automatic annotation of interval frames.

[0019] This invention introduces an interval annotation inference mechanism into long-term laser point cloud data, enabling unlabeled frames to obtain stable and reliable 3D ground truth boxes based on prior information from keyframes. This significantly reduces the reliance on frame-by-frame manual annotation, achieving a balance between cost and accuracy. Compared to traditional methods that rely on independent detection models or simple interpolation, this invention fully utilizes the spatiotemporal continuity between multiple point cloud frames, effectively fusing the complementary advantages of detection boxes and interpolation boxes. This allows the inferred candidate boxes to maintain high accuracy and completeness even in complex scenes. By introducing surrogate point feature encoding, geometric relationship encoding, and motion feature extraction based on temporal variations, this invention can stably recover the target position, shape, and orientation under various occlusion, sparse point cloud, or long-distance weak echo conditions, enabling unlabeled frames to achieve annotation quality close to that of manual annotation.

[0020] The interval annotation framework of this invention can generate results comparable to or even better than manual frame-by-frame annotation with only a small number of keyframes manually annotated, thus significantly improving data annotation efficiency and greatly alleviating the reliance of intelligent driving systems on costly manual annotation. Under different sampling intervals, this invention maintains high recall and low false detection rates, significantly improving the efficiency of constructing large-scale point cloud datasets and providing more stable, reliable, and consistent training samples for intelligent driving perception models. Furthermore, because this invention employs a unified two-stage optimization mechanism, the model possesses stronger generalization capabilities in spatial geometry, motion consistency, and cross-frame feature relationships, making it widely applicable to dynamic scenes, complex urban roads, and challenging conditions involving mixed multi-object types. Attached Figure Description

[0021] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0022] Figure 1 This is a schematic diagram of the laser point cloud data interval annotation method framework provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the target detection process of the CenterPoint network provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the geometric feature encoding provided in this embodiment of the invention, which is a proxy point. Figure 4 This is a schematic diagram illustrating the construction of temporal motion features provided in an embodiment of the present invention; Figure 5 This is a rendering of the conventional automated annotation provided in an embodiment of the present invention; Figure 6 This is a diagram showing the annotation effect of the interval annotation method provided in the embodiment of the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative effort are all within the scope of protection of the present invention.

[0024] Example 1 This embodiment provides a method for interval annotation of laser point cloud data for intelligent driving, such as... Figure 1 As shown, the method mainly includes the following steps: Step 1: Keyframe Selection and Manual Annotation A LiDAR system installed on the vehicle continuously scans the driving scene to acquire a 3D point cloud sequence arranged in chronological order. To achieve efficient annotation of long-term LiDAR point cloud data, this embodiment first performs a keyframe selection operation within the entire point cloud sequence.

[0025] Suppose the entire point cloud sequence has T Frames are composed of, denoted as ,in Indicates the first t Frame point cloud data. Keyframes are selected according to a preset interval strategy, such as selecting one frame every 10, 20, or 30 frames as a keyframe, and the remaining frames are recorded as non-keyframes. A fixed keyframe interval parameter is set. Then the keyframe index set can be represented as: .

[0026] For each keyframe in the keyframe set Precise annotation is performed manually using 3D point cloud annotation platforms (such as LabelCloud, OpenPCDet, and other visualization annotation tools). The annotation content includes, but is not limited to, the 3D bounding boxes of targets such as vehicles, pedestrians, and non-motorized vehicles, resulting in ground truth boxes corresponding to keyframes. Each 3D bounding box is defined by its center point coordinates (…). x , y , z ), size parameters (length) l ,Width w ,high h ) and orientation angle θ Description, therefore, the first k The first frame i Each annotation box can be represented as .

[0027] Step 2: Generation of interpolation boxes based on keyframe ground truth values Based on the changes in the 3D position, size, and orientation of the same target in adjacent keyframes, a temporal trajectory of the target is constructed along the time axis. For non-keyframes between two adjacent keyframes, interpolation is used based on frame time information to calculate parameters such as the target center point coordinates and orientation, and the size parameters are smoothed, thereby generating corresponding interpolation boxes on unlabeled non-keyframes.

[0028] Specifically, for the same goal (e.g., the first...) i (One target), whose truth box parameters can be extracted in two adjacent keyframes, that is:

[0029]

[0030] in, k 1 and k 2 as the target i The time index of two adjacent keyframes. Then, for those in... k 1 and k Non-keyframes between 2 t The target i The parameters of its 3D bounding box can be adjusted using linear interpolation over time. Calculations are performed.

[0031] Center point coordinate interpolation: , ,

[0032] Smoothing of bounding box size parameters: , ,

[0033] Similarly, orientation angle interpolation:

[0034] Step 3: One-stage point cloud detection with multiple frames superimposed For non-critical frames requiring annotation, the point clouds of this frame and several neighboring frames are overlaid in the temporal dimension. The resulting point cloud is then input into the CenterPoint network for target detection, outputting the 3D bounding boxes of each target as coarse candidate annotation results for the non-critical frame. Considering the sparsity, occlusion, and weak long-distance reflection of LiDAR point clouds, this step uses multi-frame point cloud overlay to recover the complete geometric structure of the target even when single-frame information is insufficient.

[0035] Among them, the CenterPoint network is an end-to-end 3D object detection network based on the concept of center points, such as... Figure 2 As shown, the network mainly includes voxelization feature extraction, BEV feature generation, and a center-point-based detection head module, which can efficiently locate the center of an object in three-dimensional space and regress its size, orientation, and other attributes. CenterPoint was first proposed by TianweiYin et al. in the paper "Center-based 3D Object Detection and Tracking" in 2020, and has been widely deployed on public platforms as an open-source project. Its implementation has been integrated into various open-source 3D perception frameworks (such as OpenPCDet, MMDetection3D, etc.), and is a mature detection model that can be directly called and reproduced by those skilled in the art.

[0036] Step 4: Constructing the candidate ROI set and extracting point clouds After completing keyframe interpolation and the first-stage detection, this step fuses the 3D bounding box results from the two sources (interpolated boxes and detection boxes) and extracts a subset of the point cloud inside each coarse box from the corresponding frame point cloud.

[0037] 1. Fusion processing of interpolation boxes and detection boxes Let the current non-keyframe be the th frame. t Frame, construct the following two initial box sets: The set of 3D bounding boxes generated by interpolation: , The set of 3D bounding boxes output by the first-stage detection: .

[0038] The two sets of boxes above are spatially joined and merged. The IoU (Intersection over Union) criterion is used to remove redundant boxes with high overlap. For any detection box If there is an interpolation box in the interpolation box set whose IoU is greater than a set threshold (e.g., 0.6), If both are considered as the same target, the one with higher quality (e.g., higher detection box confidence) is retained first, or the box obtained by weighted averaging of their geometric centers is used as the replacement box. If the detection box does not have an interpolation box with an IoU greater than a set threshold, or the interpolation box does not have a detection box with an IoU greater than a set threshold, then the detection box or interpolation box is retained. Finally, all the retained fusion results are used as the coarse boxes of the non-keyframe.

[0039] 2. Extraction of point cloud inside rough frame Let the original point cloud of the current non-keyframe be... , , Let be the three-dimensional coordinates of the point. For each rough bounding box... Define its point cloud within the bounding box as . judge Does it belong to The method is as follows: First calculate the points. Regarding rough outline Relative to its center point in length, width and height directions The differences in each dimension, namely:

[0040] like Then the point is considered Fall into the rough frame Inside, among them, , l , w , h These represent the orientation angle and length, width, and height of the rough frame, respectively.

[0041] Step 5: Rough Box Feature Encoding and Box Optimization For each rough bounding box in each non-keyframe, geometric features are first encoded based on the spatial relationship between the point cloud within the rough bounding box and the corner and center points of the box, yielding geometric features reflecting the target's shape and internal structure. Then, by combining positional changes and temporal encoding under a unified coordinate system at different times, temporal motion features that characterize the target's motion trend and cross-frame correlation are extracted. The geometric and motion features are fused and input into a two-stage fine-tuning network. At the output, a fine regression is performed on the center position, size, and orientation of each rough bounding box in the non-keyframe, achieving fine optimization of the rough bounding box.

[0042] 1. Geometric Feature Coding In this embodiment, the rough bounding box is represented by its center point and 8 corner points, totaling 9 points. For each target point cloud within the rough bounding box, the coordinate differences between it and each point of the rough bounding box in each dimension are calculated, and the target point cloud's own coordinates are appended to form a (1, 10, 3) dimensional vector. If the rough bounding box contains a total of N If there are 1 target point cloud, then a total of ( ) will be formed N The geometric eigenvectors of (,10,3).

[0043] Considering that the number of target point clouds within the rough bounding box is usually large, the geometric feature vectors obtained directly using the above method increase by up to ten times compared to the original point cloud data, resulting in low efficiency in processing such a large amount of data. Therefore, this embodiment further employs the Set Abstraction (SA) method from PoinNet++ to aggregate the aforementioned geometric features onto surrogate points, such as... Figure 3As shown, each proxy point is used as the center, and points in its neighborhood are grouped to form local point clusters. Then, a multilayer perceptron extracts local features and aggregates them using max pooling, thus aggregating the geometric features onto each proxy point to obtain the final geometric features. Regarding proxy point selection, this embodiment performs a step-by-step selection along the XYZ coordinate axes for each rough bounding box. n p Divide the box into equal parts, with each division point corresponding to a surrogate point; that is, each rough bounding box can be formed. n p × n p × n p =Np agent points.

[0044] It should be noted that PointNet++ is a point cloud processing method proposed in the 2017 paper "PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space," and has been widely used in 3D object recognition, segmentation, and detection tasks. Since the structure and training methods of PointNet++ are fully disclosed in literature and open-source frameworks, and those skilled in the art can directly call and implement them based on publicly available information, this invention will not elaborate further here, and this will not affect the feasibility of the invention.

[0045] Using the above method, this embodiment effectively completed the rough bounding box and point cloud encoding, successfully aggregated the point cloud features and rough bounding box features onto the surrogate points, and obtained the final geometric features.

[0046] 2. Temporal motion feature fusion coding The geometric feature encoding described above is limited to a single moment corresponding to a non-keyframe, while combining it with temporal information is also very important and effective. For example... Figure 4 As shown, in this embodiment, for the first... t For the nine points (center and eight corner points) of the rough bounding box in the non-key frame, calculate the coordinate differences of each dimension between each surrogate point of the rough bounding box for the same target in the previous frame (or multiple previous frames), and add additional temporal (frame order) encoding. Then, for Np surrogate points, a (Np, 10, 3) dimensional vector is formed. This vector can be directly used as a temporal motion feature, or it can be input into a multilayer perceptron for further extraction and output as a temporal motion feature.

[0047] Finally, the temporal motion features corresponding to each rough box in the non-keyframes are fused with the aforementioned geometric features to obtain fused geometric-temporal motion features, which are then input into the two-stage fine-tuning network to optimize each rough box. During fusion, feature vectors can be directly concatenated or their corresponding dimensions can be added.

[0048] 3. Two-stage fine-tuning network The two-stage fine-tuning network is used to refine the rough bounding boxes obtained by fusing one-stage detection and interpolation inference in non-key frames. Its core objective is to fully utilize the geometric structure information of the point cloud within the rough bounding boxes and the temporal variation information across frames to perform high-precision regression correction on the center position, spatial size, and orientation parameters of the rough bounding boxes, thereby generating 3D bounding boxes that can be directly used as ground truth annotations. This network does not re-perform global object detection; instead, it focuses on the geometric and temporal motion features corresponding to each rough bounding box in the non-key frames, making it a typical candidate box-level refinement network.

[0049] In terms of network structure, the two-stage fine-tuning network mainly includes an input module, an attention enhancement module, and a regression output module. The input module is used to input the geometric temporal motion features corresponding to each coarse box contained in non-keyframes. The attention enhancement module uses the Transformer attention mechanism to interactively fuse the geometric temporal motion features. By learning the importance distribution between different feature dimensions, it adaptively adjusts geometric or motion cues that are highly correlated with box optimization, thereby improving regression stability in complex scenes. The introduction of the attention mechanism enables the network to dynamically adjust the feature contribution ratio under conditions such as sparse point clouds and severe occlusion, avoiding error amplification caused by a single feature dominating.

[0050] In the regression output stage, the attention-enhanced features are input into a regression head composed of a multilayer perceptron (MLP). This regression head predicts the center position offset, size correction, and orientation angle correction of the rough bounding box, respectively. The output results are superimposed on the original rough bounding box parameters in the form of residuals to obtain the refined 3D bounding box. Through this residual regression method, the network's learning objective changes from directly predicting absolute values ​​to predicting correction values, which helps improve convergence speed and numerical stability.

[0051] During the training phase, the two-stage fine-tuning of the network uses keyframes or manually calibrated data as supervision signals. By matching rough bounding boxes with their corresponding ground truth boxes, the differences in center position, size, and orientation are calculated, and a regression loss function is constructed to train the network end-to-end. The network parameters can be trained using conventional deep learning optimization methods, and after training, the parameters are fixed for the inference phase.

[0052] During training, random perturbations to the rough bounding boxes can be incorporated to enable the network to learn correction capabilities under different bias conditions, thereby improving generalization performance. Specifically, for each non-keyframe rough bounding box in the training data, it is first randomly decided whether to apply a perturbation to the box based on a set perturbation probability. For the selected rough bounding box to be perturbed, its center coordinates are calculated. Dimensions and yaw angle Adding a small perturbation term to the top, the new bounding box after perturbation can be represented as: , , , ,

[0053] in, Indicates the interval [ a , b Uniform distribution on ] a x , a y , a z , , For the preset disturbance amplitude control factor (e.g.) , , This ensures that the disturbance is representative but does not completely disrupt the target's location.

[0054] To prevent the generated perturbation boxes from deviating too far from the real target or causing training failure, this embodiment also introduces a dual-screening constraint mechanism. First, the 3D IoU value between the perturbation box and its corresponding ground truth box is calculated. If the IoU is less than a set threshold (e.g., 0.25), the perturbation is considered too strong and may mislead the model's learning, so it is discarded. Second, the number of point clouds contained within the perturbation box is counted. If this number is less than a preset lower limit (e.g., 5-10 points), it indicates that the box fails to cover the effective object point cloud, and this perturbation is also discarded. Finally, the coarse boxes that are successfully perturbed under the above constraints will be used as augmented samples in training, allowing the network to fully engage with data samples of various bias types during the training phase, thereby possessing stronger error robustness and correction capabilities during the inference phase.

[0055] The above summarizes the main content of the interval annotation method of this invention. To verify the effectiveness of this method, it was compared with traditional automated annotation methods that rely on independent detection models. Figure 5 and Figure 6 The images show the annotation results of traditional automated annotation methods and the method of this invention for the same scene. By comparison, it can be seen that traditional methods have serious omissions in complex scenes, while the present invention can make full use of the spatiotemporal continuity between multiple frames of point clouds, so that the optimized and inferred annotation boxes can still maintain high accuracy and completeness in complex scenes.

[0056] The core of this invention lies in combining the advantages of detection boxes and interpolation boxes to maintain high accuracy and recall under challenging scenarios, significantly reducing annotation costs. Experiments show that with annotation intervals of 10 frames, the recall rate (IOU 0.7) reaches 0.986, representing improvements of 31% and 15% respectively compared to detection boxes and interpolation boxes, and virtually eliminating false negatives and missed detections, requiring almost no manual adjustment of the detection results. Even with manual annotation of one frame every 30 frames, the recall rate of IOU 0.7 still reaches 0.87. Compared to the accuracy of fully manual annotation (IOU 0.7 recall rate 0.864), the interval annotation performance of this invention is comparable, and is of great significance for reducing annotation costs in intelligent driving systems.

[0057] Example 2 Based on the above method, this embodiment provides a laser point cloud data interval annotation system for intelligent driving. This system is designed for long-term temporal 3D point cloud annotation tasks and possesses a complete processing chain including keyframe selection and fine-tuning, cross-frame interpolation inference, multi-frame target detection, interpolation detection fusion, candidate bounding box point cloud extraction, feature encoding and fusion, and rough box optimization. It can significantly improve the annotation efficiency and consistency of point cloud data, and is particularly suitable for high-quality construction of large-scale autonomous driving scenario datasets. Specifically, this system mainly includes a data input module, a keyframe management and manual annotation module, an interpolation inference module, a temporal detection module, a candidate ROI fusion and point cloud extraction module, a feature encoding module, and a two-stage bounding box optimization module. Multiple functional modules are connected in series through a unified data interface protocol, supporting both offline batch processing and online collaborative operation modes.

[0058] First, the data input module serves as the system's initial interface, receiving time-series point cloud data from external sources. This module supports standardized 3D point cloud data sequence input (such as .bin, .pcd, .npy, etc.), and can parse and organize it into a frame-structured format for storage, forming a list of point cloud frames arranged in chronological order. Each point cloud frame contains a timestamp and a set of 3D spatial points, with each point containing at least 3D coordinates (x, y, z), and optionally additional attributes such as reflection intensity and radar channel ID. This module also provides a point cloud loading caching mechanism, time index management, and segmentation control functions, facilitating subsequent modules to operate on data within the target time period and improving the overall data processing throughput and access efficiency of the system.

[0059] Next, the keyframe management and manual annotation module performs inter-frame analysis on the input point cloud sequence, selecting several representative frames as keyframes based on the set keyframe selection strategy (such as fixed interval, velocity change rate threshold, curvature trigger, etc.). The system can be configured with different sampling parameters, supporting strategies such as selecting one frame every 10, 20, or 30 frames, and generates a keyframe index set. After obtaining the keyframe list, the system calls the integrated 3D annotation platform interface (compatible with tools such as LabelCloud and OpenPCDet GUI) to push the keyframe data to the manual annotation interface for manual drawing and attribute labeling of 3D bounding boxes. The annotation box of each target includes parameters such as its spatial center point coordinates, length, width, height, and orientation angle, and is bound to a category label. The system automatically performs format verification, normalization, and mapping of the annotation data to a unified coordinate system as the basis for subsequent inference calculations.

[0060] After receiving the labeled keyframe bounding box data, the interpolation inference module analyzes and reconstructs the cross-frame trajectory of each target on the time axis. This module uses methods such as linear interpolation, spline interpolation, or Bezier interpolation to estimate the position, size, and orientation of targets in non-keyframes, generating a set of interpolation boxes. During the interpolation process, the system considers the temporal continuity and local velocity smoothness of the targets. Acceleration constraints and attitude change constraints can be used to improve the physical rationality and structural consistency of the interpolation, thereby ensuring that the interpolation boxes have strong reference value and preliminary geometric accuracy. The interpolation results will serve as important prior information input into the system backbone for subsequent detection and optimization.

[0061] In parallel, the temporal detection module is responsible for performing multi-frame fusion 3D target detection for each non-critical frame. This module selects several neighboring frames (e.g., two frames before and two before, for a total of five frames) centered on the target non-critical frame, and fuses their corresponding point clouds along the time axis to form an enhanced frame point cloud input. The fusion method employs a coordinate alignment strategy based on LiDAR extrinsic parameters, rotating and translating the point clouds of different frames to superimpose them onto the target frame reference frame, retaining timestamps or frame numbers as additional features. Then, the system inputs the enhanced frame point cloud into a first-stage 3D target detection network to extract potential target bounding boxes in the scene, outputting preliminary detection results. This module is compatible with different detection networks such as CentPoint, PointPillars, SECOND, and PV-RCNN, and can adapt to different types of LiDAR equipment and point cloud densities by configuring and switching the network backend.

[0062] The candidate ROI fusion and point cloud extraction module receives interpolated bounding boxes and detection boxes generated by the interpolation inference module and the temporal detection module, respectively. It performs IoU overlap determination and redundant box removal operations, fusing the two result sets to construct a unified coarse candidate box set. For each coarse box, the system performs a spatial envelope query in the original point cloud of the corresponding frame, extracting a subset of the point cloud inside the box. This extraction process combines the box's length, width, height, and orientation angle information, using coordinate transformation from the point to the box center to determine whether a point is within the bounding box. It supports GPU-accelerated batch spatial filtering algorithms, achieving highly efficient point cloud cropping and cleaning.

[0063] The feature encoding module is responsible for constructing structural input features for each rough bounding box and its internal point cloud. The system first transforms the bounding box into its nine feature points (i.e., the center point and eight corner points). Then, it calculates the spatial coordinate difference between each point within the box and these feature points, and appends its own point cloud coordinates to form the original geometric feature tensor. To reduce dimensionality redundancy and improve computational efficiency, the system uses the Set Abstraction mechanism in PointNet++ to encode the high-dimensional point set into a local representation of a surrogate point set, and utilizes max pooling and MLP to construct spatially invariant geometric features. Furthermore, the system extracts temporal motion features by utilizing the temporal displacement, angular offset, and size change trends of the rough bounding box relative to adjacent frame bounding boxes. Finally, it merges the geometric and temporal features through feature fusion operations to form the final candidate bounding box input representation.

[0064] Finally, the two-stage bounding box optimization module utilizes the aforementioned feature inputs to optimize the structure of each coarse bounding box through a fine regression network. This module introduces a Transformer-based attention mechanism network to dynamically learn the importance weights of each feature dimension in different scenarios, effectively suppressing interference from invalid features. In the regression stage, the system employs a residual regression strategy to output corrections for position, size, and orientation, which are then superimposed on the original coarse bounding box parameters to generate the final refined bounding box result. During the training stage, the system uses manually annotated keyframes as supervision targets, constructs multi-task loss functions such as IoU loss and L1 regression loss to jointly optimize network parameters, and combines perturbation enhancement and dual screening mechanisms to enhance the model's robustness and generalization ability.

[0065] The above system can execute the laser point cloud data interval annotation method for intelligent driving described in Embodiment 1, and has the corresponding functional modules and beneficial effects of the method. For technical details not described in detail in this embodiment, please refer to the laser point cloud data interval annotation method provided in Embodiment 1 of the present invention.

[0066] This laser point cloud data interval annotation system, relying only on manual annotation of a small number of keyframes, achieves high-precision automated annotation of unannotated frames in complex environments through mechanisms such as interpolation prior, detection inference, multi-frame enhancement, and geometric-temporal joint modeling. It boasts extremely high annotation efficiency and application value. This system is suitable for the data construction phase in the intelligent driving data closed loop and can be widely deployed in autonomous driving perception training platforms, urban road simulation systems, and multi-sensor fusion annotation pipelines.

[0067] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; under the concept of the present invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the present invention as described above, which are not provided in detail for the sake of brevity; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for interval labeling of laser point cloud data for intelligent driving, characterized in that, Including the following steps: S1. Select several point cloud frames from the laser point cloud sequence to be labeled as key frames according to the set frame interval for manual labeling, obtain the three-dimensional bounding box parameters of each target in each key frame, and the unselected point cloud frames are regarded as non-key frames. S2. For the 3D bounding box parameters of the same target in adjacent key frames, generate the 3D interpolation box of the same target in each non-key frame between the adjacent key frames, and obtain the 3D interpolation box parameters. S3. Use a 3D target detection algorithm to perform target detection on each non-key frame, generate 3D detection boxes for each target in the non-key frame, and obtain the parameters of the 3D detection boxes. S4. The 3D interpolation boxes and 3D detection boxes in the same non-key frame are fused and filtered to remove redundant boxes and obtain the 3D rough boxes of each target in the non-key frame. S5. For each 3D rough box in each non-key frame, geometric feature encoding is performed based on the spatial positional relationship between each point cloud in the rough box and the corner and center points of the rough box to obtain geometric features that reflect the shape of the target. Simultaneously, extract temporal motion features that can characterize the target's motion trend and cross-frame correlation; S6. After fusing geometric and motion features, input them into the trained two-stage fine-tuning network to regress the center position, size, and orientation of each rough box in non-keyframes, and obtain the optimized 3D box, which is then used as the annotation box.

2. The laser point cloud data interval annotation method as described in claim 1, characterized in that, In step S2, for the same target i , suppose its three-dimensional bounding box parameters in two adjacent key frames are respectively: wherein, k 1 and k 2 are target i time indices of two adjacent key frames, and are target i center point coordinates of three-dimensional bounding boxes in key frames k 1 and k 2, and are target i length, width, height, and orientation angle of three-dimensional bounding boxes in key frames k 1 and k 2, respectively. Then for those in k 1 and k Non-keyframes between 2 t The target i The three-dimensional interpolation frame parameters are obtained by linear interpolation based on time. ,in: Let be the coordinates of the center point of the 3D interpolation frame, and , , Let the length, width, and height of this 3D interpolation frame be , and , , , Let be the orientation angle of the 3D interpolation frame, and .

3. The laser point cloud data interval annotation method as described in claim 1, characterized in that, In step S3, before using the 3D target detection algorithm to perform target detection on each non-critical frame, the insufficient information of a single frame caused by the sparse point cloud of the lidar is alleviated by superimposing multiple point cloud frames. That is, for a non-critical frame that needs to be detected, the frame and several neighboring point cloud frames before and after it are selected and superimposed, and the superimposed point cloud frames are input into the 3D target detection algorithm to obtain the 3D detection box of each target.

4. The laser point cloud data interval annotation method as described in claim 1, characterized in that, The three-dimensional target detection algorithm described in step S3 uses the CenterPoint network.

5. The laser point cloud data interval annotation method as described in claim 1, characterized in that, In step S4, assuming that the 3D interpolation boxes and 3D detection boxes of each target in a non-key frame constitute an interpolation box set and a detection box set respectively, the fusion filtering is as follows: For any detection box in the detection box set, if there is an interpolation box in the interpolation box set with an IoU greater than a set threshold, then the two are considered as bounding boxes of the same target and the detection box is retained, or the box obtained by weighted averaging of the two is taken as the replacement box and retained; otherwise, the detection box is directly retained. For any interpolation box in the interpolation box set, if there is a detection box in the detection box set with an IoU greater than a set threshold, then the two are considered as bounding boxes of the same target and the detection box is retained, or the box obtained by weighted averaging of the two is taken as the replacement box and retained; otherwise, the interpolation box is directly retained. Finally, all the retained bounding boxes are used as the 3D rough bounding boxes of each target in the non-keyframe.

6. The laser point cloud data interval annotation method as described in claim 1, characterized in that, In step S5: The rough box is represented by 9 points consisting of its center point and 8 corner points. For each target point cloud in the rough box, the difference between its coordinates and each of the frame points in each dimension is calculated, and the coordinates of the target point cloud itself are attached to form the geometric feature vector of the target point cloud. The geometric feature vectors of all target point clouds in the rough box together constitute the geometric features of the rough box. Alternatively, the set abstraction method in PoinNet++ can be used to aggregate the geometric features onto surrogate points. With each surrogate point as the center, points in the neighborhood are searched to form local point groups for grouping. Then, local features are extracted by a multilayer perceptron and aggregated using max pooling to obtain the final geometric features. The surrogate points are selected by dividing the rough box into several equal parts along its coordinate axes, with each division point corresponding to a surrogate point.

7. The laser point cloud data interval annotation method as described in claim 6, characterized in that, In step S5, the temporal motion feature is as follows: for the nine points of the rough box, calculate the difference of each dimension coordinate between each point and each proxy point of the rough box for the same target in the previous frame or multiple previous frames, and attach the corresponding point cloud frame time code to jointly constitute the temporal motion feature; or the temporal motion feature is further input into the multilayer perceptron to output the final temporal motion feature.

8. The laser point cloud data interval annotation method as described in claim 1, characterized in that, In step S6, the geometric features and motion features are directly concatenated into vectors or their corresponding dimensions are added to obtain the fused geometric temporal motion features, which are then input into the trained two-stage fine-tuning network. The two-stage fine-tuning network includes an input module, an attention enhancement module, and a regression output module, wherein the input module is used to input the geometric temporal motion features corresponding to each rough box; The attention enhancement module uses the Transformer attention mechanism to interactively fuse geometric temporal motion features. By learning the importance distribution between different feature dimensions, it adaptively adjusts geometric or motion cues that are highly correlated with the frame optimization, thereby improving regression stability in complex scenes. The regression output module includes a regression head composed of a multilayer perceptron, used to regress the center position, size, and orientation of the rough box based on the output of the attention enhancement module.

9. The laser point cloud data interval annotation method as described in claim 8, characterized in that, The regression output module includes multiple regression heads, which are used to predict the center position offset, size correction, and orientation angle correction of the rough bounding box, respectively. The prediction output results are superimposed on the corresponding parameters of the rough bounding box in the form of residuals, thereby obtaining the refined three-dimensional bounding box as the annotation box of the corresponding target.

10. A laser point cloud data interval annotation system based on the method of any one of claims 1 to 9, characterized in that, Includes the following modules: The data input module is used to receive and organize the laser point cloud frame sequence arranged in chronological order, store the point cloud frames in a structured manner, and provide a unified time index and access interface. The keyframe management and manual annotation module is used to select multiple keyframes from the point cloud frame sequence according to a preset strategy, and to annotate the targets in the keyframes with three-dimensional bounding boxes through an external annotation platform. The interpolation inference module is used to perform interpolation inference on non-key frames based on the geometric parameter changes of the same target in adjacent key frames, and generate the corresponding set of 3D interpolation boxes. The temporal detection module is used to select the point cloud of each non-key frame and several neighboring frames before and after it for fusion processing, and input the fused point cloud frame into the three-dimensional target detection algorithm to output the three-dimensional detection box set of the non-key frame. The candidate ROI fusion and point cloud extraction module is used to fuse the interpolation box and the detection box to construct a set of candidate coarse boxes for non-key frames, and extract the corresponding point cloud in each coarse box. The feature encoding module is used to calculate geometric features based on the point cloud data and bounding box feature points within the rough box, further extract temporal motion features by combining cross-frame motion information, and fuse and encode the geometric features and temporal features. The two-stage bounding box optimization module receives the fused encoded features and inputs them into a bounding box optimization network that includes an attention mechanism and a residual regression structure. It performs fine regression on the center position, size parameters, and orientation angle of each rough bounding box, and outputs accurate 3D bounding box results for automatic annotation of interval frames.