Point cloud target self-labeling method and device, equipment and medium
The point cloud target box is determined through the inter-frame transfer algorithm and model inference algorithm, and the cross-validation algorithm is filtered, and the time-consuming and labor-intensive problem of manual labeling is solved, efficient and low-cost point cloud target self-labeling is achieved, and the labeling quality is improved.
Patent Information
- Application Number
- CN202510052468.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, manual labeling of point cloud data is time-consuming and labor-intensive and costly, which limits the application and development of point cloud target recognition in autonomous driving.
The inter-frame transfer algorithm and model inference algorithm are used to determine the point cloud delivery box and the point cloud inference box, and the two are matched and filtered through the cross-validation algorithm to obtain the final point cloud target box.
The efficient and low-cost self-marking of point cloud goals have been achieved, the quality of labeling has been improved, and the demand for manual labeling has been reduced.
Smart Images

Figure CN119964110A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of point cloud data annotation, and in particular to a point cloud target self-annotation method, device, equipment and medium. Background Art
[0002] In the field of autonomous driving, the ability to accurately perceive the physical world is crucial. Point cloud data has become an indispensable source of information for autonomous driving because it can provide accurate spatial structure information. However, in order for data-driven autonomous driving algorithms to effectively understand and process point cloud data, a large amount of accurate 4D (3D space plus time series) true value annotations are required for supervised learning. The existing manual annotation of this type is time-consuming, labor-intensive and extremely costly, which limits the application and development of point cloud object recognition. Summary of the invention
[0003] The purpose of this application is to provide a point cloud target self-labeling method, device, equipment and medium, which can realize self-labeling of point cloud targets efficiently and at low cost.
[0004] To achieve the above objectives, this application provides the following solutions:
[0005] In a first aspect, the present application provides a point cloud target self-labeling method, comprising:
[0006] Get the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame;
[0007] According to the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame, an inter-frame transfer algorithm is used to determine the point cloud transfer frame of the current frame; the point cloud transfer frame is a set of self-annotated point cloud target frames;
[0008] According to the point cloud data of the current frame, a model inference algorithm is used to determine the point cloud inference frame of the current frame; the point cloud inference frame is another set of self-annotated point cloud target frames;
[0009] A cross-validation algorithm is used to match and screen the point cloud transfer frame of the current frame and the point cloud inference frame of the current frame to obtain a point cloud proposal frame of the current frame; the point cloud proposal frame is the final point cloud target frame.
[0010] Optionally, according to the point cloud data of the current frame and the point cloud data and the point cloud target frame of the previous frame, an inter-frame transfer algorithm is used to determine the point cloud transfer frame of the current frame, specifically including:
[0011] The point cloud target in the point cloud data of the previous frame is aligned with the point cloud target in the point cloud data of the current frame, the spatial transformation relationship between the point cloud targets of the two frames is determined, and the spatial transformation relationship between the point cloud targets of the two frames is used as the spatial transformation relationship of the point cloud target annotation boxes of the two frames. The point cloud target box of the previous frame tracks the point cloud target box of the current frame to obtain the point cloud transfer box of the current frame.
[0012] Optionally, the point cloud target in the point cloud data of the previous frame is registered with the point cloud target in the point cloud data of the current frame, the spatial transformation relationship between the point cloud targets of the two frames is determined, and the spatial transformation relationship between the point cloud targets of the two frames is used as the spatial transformation relationship of the point cloud target annotation frames of the two frames, and the point cloud target frame of the previous frame tracks the point cloud target frame of the current frame to obtain the point cloud transfer frame of the current frame, specifically including:
[0013] Determine the previous frame data, current frame data, previous frame transformation parameters and current frame transformation parameters according to the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame; the previous frame data includes: the entire point cloud of the previous frame, the local coordinate system of the ego vehicle of the previous frame and the point cloud target frame of the previous frame; the current frame data includes: the entire point cloud of the current frame and the local coordinate system of the ego vehicle of the current frame; the previous frame transformation parameters include: the rotation matrix and translation vector of the local coordinate system of the ego vehicle of the previous frame transformed to the global coordinate system; the current frame transformation parameters include: the rotation matrix and translation vector of the local coordinate system of the ego vehicle of the current frame transformed to the global coordinate system;
[0014] Determine all point clouds within the point cloud target frame of the previous frame in the local coordinate system of the vehicle in the previous frame according to the previous frame data to obtain a first point cloud;
[0015] Convert the first point cloud to a global coordinate system according to the previous frame transformation parameters to obtain a second point cloud;
[0016] The point cloud target frame of the previous frame is converted to the global coordinate system, and the converted target frame is determined as the point cloud target frame initialized in the current frame;
[0017] Determine all point clouds within the point cloud target frame initialized in the vehicle local coordinate system of the current frame according to the current frame data and the point cloud target frame initialized in the current frame, and obtain a third point cloud;
[0018] Converting the third point cloud to a global coordinate system according to the current frame transformation parameters to obtain a fourth point cloud;
[0019] Registering the second point cloud with the fourth point cloud, and calculating a spatial transformation relationship between the fourth point cloud and the second point cloud;
[0020] Calculate the inverse transformation relationship according to the spatial transformation relationship between the fourth point cloud and the second point cloud, and determine the inverse transformation relationship as the spatial transformation relationship between the point cloud objects of the two frames;
[0021] The point cloud target frame initialized in the current frame is transformed according to the spatial transformation relationship between the point cloud targets of the two frames to obtain the point cloud transfer frame of the current frame.
[0022] Optionally, a cross-validation algorithm is used to match and filter the point cloud transfer box of the current frame and the point cloud inference box of the current frame to obtain a point cloud proposal box of the current frame, specifically including:
[0023] Match the point cloud transfer frame of the current frame with the point cloud inference frame of the current frame to determine the intersection-union ratio;
[0024] Determine the point cloud inference frame whose intersection-over-union ratio is greater than or equal to the first set value as a matched frame, and determine the point cloud inference frame whose intersection-over-union ratio is less than the first set value as an unmatched frame;
[0025] Determine a matching box whose confidence is less than or equal to a second set value as a death box, and determine a matching box whose confidence is greater than the second set value as a cross box;
[0026] Determine an unmatched frame whose confidence is less than or equal to a third set value as an eliminated frame, and determine an unmatched frame whose confidence is greater than the third set value as a new frame;
[0027] The cross frame and the new frame are merged to obtain the point cloud proposal frame of the current frame.
[0028] Optionally, according to the point cloud data of the current frame, a model inference algorithm is used to determine the point cloud inference frame of the current frame, specifically including:
[0029] According to the point cloud data of the current frame, a deep learning point cloud target detection model is used to determine the point cloud inference box of the current frame; the deep learning point cloud target detection model is a Det3D target detection framework, a CenterPoint model or a dynamic sparse voxel transformer.
[0030] Optionally, the point cloud transfer frame includes: the center point coordinates of the frame, the rotation matrix of the frame, the size of the frame and the category of the frame; the point cloud inference frame includes: the center point coordinates of the frame, the rotation matrix of the frame, the size of the frame, the category of the frame and the confidence of the frame.
[0031] Optionally, the first setting value is 0.7; the second setting value is 0.5; and the third setting value is 0.9.
[0032] In a second aspect, the present application provides a point cloud target self-labeling device, comprising:
[0033] A data acquisition module is used to acquire the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame;
[0034] An inter-frame transfer module is used to determine the point cloud transfer frame of the current frame using an inter-frame transfer algorithm based on the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame; the point cloud transfer frame is a set of self-annotated point cloud target frames;
[0035] A model inference module is used to determine a point cloud inference frame of the current frame using a model inference algorithm according to the point cloud data of the current frame; the point cloud inference frame is another set of self-annotated point cloud target frames;
[0036] The cross-validation module is used to match and filter the point cloud transfer frame of the current frame and the point cloud inference frame of the current frame using a cross-validation algorithm to obtain a point cloud proposal frame of the current frame; the point cloud proposal frame is the final point cloud target frame.
[0037] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the above-described point cloud target self-labeling methods.
[0038] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described point cloud target self-labeling methods.
[0039] According to the specific embodiments provided in this application, this application has the following technical effects:
[0040] The present application provides a method, apparatus, device and medium for self-labeling of point cloud targets. For the point cloud data of the current frame, an inter-frame transfer algorithm and a model inference algorithm are used to determine two groups of self-labeled point cloud target frames, and a cross-validation algorithm is used to match and screen the two groups of self-labeled point cloud target frames to obtain the final point cloud target frame of the current frame. This method can not only realize self-labeling of point cloud targets efficiently and at low cost, but also has high labeling quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1 This is an application environment diagram of a point cloud target self-labeling method in one embodiment of the present application;
[0043] Figure 2 A schematic diagram of a process of a point cloud target self-labeling method provided in one embodiment of the present application;
[0044] Figure 3 An overall block diagram of a point cloud target self-labeling method provided in one embodiment of the present application;
[0045] Figure 4 A schematic diagram of the process of an inter-frame transfer algorithm provided in an embodiment of the present application;
[0046] Figure 5 A schematic diagram of an F1 point cloud target frame and an F2 point cloud target frame provided in an embodiment of the present application;
[0047] Figure 6 A schematic diagram of the overall structure of a point cloud frame cross-validation algorithm provided in an embodiment of the present application;
[0048] Figure 7 A schematic diagram of functional modules of a point cloud object self-labeling device provided by another embodiment of the present application;
[0049] Figure 8 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0051] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0052] In response to the problem that manual labeling is time-consuming, labor-intensive and extremely costly, with the advancement of deep learning, model-based point cloud self-labeling methods have emerged, which automatically generate labels through model reasoning, which can significantly reduce the need for manual labeling. However, this self-labeling method may produce mislabeling in the absence of external verification. Therefore, developing a point cloud 4D true value self-labeling method that can cross-validate from multiple information sources to ensure the quality of the labeling has important research value and broad application prospects.
[0053] The point cloud target self-labeling method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown, the terminal 102 communicates with the server 104 through a network. The data storage system can store data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the point cloud data of the current frame and the point cloud data and the point cloud target frame of the previous frame to the server 104. After the server 104 receives the point cloud data of the current frame and the point cloud data of the previous frame, the server 104 uses an inter-frame transfer algorithm to determine the point cloud transfer frame of the current frame based on the point cloud data of the current frame and the point cloud data of the previous frame; the point cloud transfer frame is a group of self-annotated point cloud target frames; based on the point cloud data of the current frame, a model inference algorithm is used to determine the point cloud inference frame of the current frame; the point cloud inference frame is another group of self-annotated point cloud target frames; a cross-validation algorithm is used to match and screen the point cloud transfer frame of the current frame and the point cloud inference frame of the current frame to obtain a point cloud proposal frame of the current frame; the point cloud proposal frame is the final point cloud target frame.
[0054] The server 104 may feed back the obtained point cloud proposal frame of the current frame to the terminal 102. In addition, in some embodiments, the point cloud target self-labeling method may also be implemented by the server 104 or the terminal 102 alone, such as the terminal 102 may directly process the point cloud data of the current frame and the point cloud data and the point cloud target frame of the previous frame, or the server 104 may obtain the point cloud data of the current frame and the point cloud data and the point cloud target frame of the previous frame from the data storage system, and process the point cloud data of the current frame and the point cloud data and the point cloud target frame of the previous frame.
[0055] The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.
[0056] In an exemplary embodiment, Figure 2 As shown, a point cloud target self-labeling method is provided, which is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or can be executed by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used as an example to illustrate the method, which includes the following steps 201 to 204. Among them:
[0057] Step 201, obtaining the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame.
[0058] Step 202, based on the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame, an inter-frame transfer algorithm is used to determine the point cloud transfer frame of the current frame; the point cloud transfer frame is a set of self-labeled point cloud target frames.
[0059] Step 203, based on the point cloud data of the current frame, a model inference algorithm is used to determine the point cloud inference frame of the current frame; the point cloud inference frame is another set of self-labeled point cloud target frames.
[0060] In step 204, a cross-validation algorithm is used to match and filter the point cloud transfer frame of the current frame and the point cloud inference frame of the current frame to obtain a point cloud proposal frame of the current frame; the point cloud proposal frame is the final point cloud target frame.
[0061] The implementation of the above steps 201 to 204 can not only realize the self-labeling of the point cloud target efficiently and at low cost, but also achieve high labeling quality.
[0062] The embodiments of the present application involve three algorithms as a whole: a point cloud frame inter-frame transfer algorithm, a point cloud frame model inference algorithm, and a point cloud frame cross-validation algorithm. Point cloud frame inter-frame transfer algorithm: According to the target frame position annotated in the previous frame, the target frame position (point cloud transfer frame) of the next frame is found through point cloud registration, coordinate transformation and other processes; Point cloud frame model inference method: Input the point cloud, perform inference based on the point cloud target detection deep learning model, and obtain the point cloud target prediction frame (point cloud inference frame) of the frame; Point cloud target frame cross-validation algorithm: According to the point cloud transfer frame and point cloud inference frame obtained by the above two algorithms, cross-validate and determine a set of point cloud proposal frames as the automatic annotation result of the point cloud 4D annotation frame of this frame. The overall block diagram of the point cloud target self-annotation method is shown below. Figure 3 shown.
[0063] In another exemplary embodiment of the present application, the inter-frame transfer algorithm in step 202 is introduced in detail.
[0064] Point cloud target frame inter-frame transfer algorithm: Given the point cloud target, point cloud target frame of the previous frame, and point cloud target of the next frame, after target point cloud registration, the spatial transformation relationship between the point cloud targets of the two frames is found, that is, the spatial transformation relationship between the point cloud target annotation frames of the two frames, so that the point cloud target frame of the next frame is tracked by the point cloud target frame of the previous frame, and the point cloud frame is transferred between consecutive frames, such as Figure 4 As shown in , F1 and F2 are the point clouds of the 1st and 2nd frames respectively. Figure 5As shown, the point cloud in the F1 point cloud target frame is registered to obtain the coordinate transformation relationship between the two frame targets, that is, the coordinate transformation relationship between the two frame target frames, and then the F2 point cloud target frame is obtained by transforming the F1 point cloud target frame. The F1 point cloud target frame is as shown in Figure 5 As shown in part (a) of Figure 5 As shown in part (b) of .
[0065] (1) Point cloud target frame inter-frame transfer algorithm object definition.
[0066] f1 and f2 are two adjacent frames, f1 is a reference frame (reference), and f2 is a source frame (source).
[0067] Related objects in frame f1: P r is the entire point cloud of f1; C r is the local coordinate system of the vehicle f1; B r is a reference target frame of f1; B r =(c r , r r , d r , class), where: c r is the coordinate of the center point of the box, c r =[x r ,y r , z r ],r r is the rotation matrix of the box, indicating the orientation angle of the box, d r is the size of the frame, d r =[l r , w r ,h r ], class is the annotation box category, such as "Car", "Pedestrian", "Bus", etc. In the point cloud box transfer algorithm, the previous and next two frames B r c r , r r The properties will change, r , the class attribute remains unchanged; is the reference frame set within f1; O r For f1 in C r Target frame B in coordinate system r All point clouds within.
[0068] Related objects in frame f2: P s is the entire point cloud of f2; C s is the local coordinate system of f2; B s is a reference target frame of f2, B s =(c s , rs , d s , class), where each attribute is the same as B r , is the reference frame set in f2, and is composed of Obtained through the proposed point cloud frame tracking algorithm; s For f2 in C s Target frame B in coordinate system s All point clouds within.
[0069] (2) Definition of transformation relationship.
[0070] There are three coordinate systems involved in the algorithm: C r , C s , C w , respectively, are the local coordinate system and global coordinate system of f1 and f2, let R rw , T rw It is C r Transform to C w The rotation matrix and translation vector, that is, C r →C w :[R rw , T rw ], similarly: C s →C w :[R sw , T sw There are two transformation methods:
[0071] The first transformation method: coordinate system transformation of point cloud.
[0072] Point p r For example, C r The point below r Convert to c w Point p w The process is:
[0073] p w =R rw *p r +T rw .
[0074] Point Cloud O r For example, C r Point cloud O r Convert to C w Point cloud below The process is:
[0075]
[0076] The second transformation method: coordinate system transformation of the annotation box.
[0077] Frame Br =(c r , r r , d r ) as an example:
[0078] c w =R rw *c r +T rw ;
[0079] r w =R rw *r r ;
[0080] d w =d r ;
[0081] Then C w B r ,Right now
[0082] (3) Definition of calculation method.
[0083] First calculation method: point cloud target O r , O s Cut.
[0084] Initialize O r is an empty set, in C r In the coordinate system, traverse P r All points in r , judge p r Is it in B r If so, point p r Add O r ;O s Same reason.
[0085] The second calculation method: point cloud registration.
[0086] Input: Reference point cloud target O in the same coordinate system r , source point cloud target O s .
[0087] Output: Source point cloud O s To the reference point cloud O r The transformation relationship between them is denoted as: s →O r : [R, T], including the rotation matrix R and the translation vector T.
[0088] The point cloud registration algorithm uses the reference point cloud as the benchmark, transforms the source point cloud position to register the source point cloud to the reference point cloud, and finally returns the transformation relationship [R, T] between the source point cloud and the reference point cloud, that is, the source point cloud O sAfter the transformation relationship [R, T] is transformed to the reference point cloud O r , the transformation relationship is recorded as: O s →O r :[R,T].
[0089] The specific algorithm of point cloud registration can use the existing mature point cloud registration algorithm, such as ICP, Go-ICP, 4PCS, FGR and other methods.
[0090] The third calculation method: find the inverse transformation relationship.
[0091] Input: O s to r The position transformation relationship, that is, O s →O r : [R, T]; Output: O r to s The position transformation relationship, that is, O r →O s :[R T , -R T T], where R T is the transposed matrix of R.
[0092] (4) The process of the point cloud target frame inter-frame transfer algorithm is as follows.
[0093] Input: f1: P r , C r , B r ,
[0094] f2: P s , C s ,
[0095] C r →C w :[R rw , T rw ],
[0096] C s →C w :[R sw , T sw ].
[0097] Output: B s .
[0098] Specific:
[0099] 1.P r , C r , B r The first calculation method is After the first transformation, we get
[0100] 2. After the second transformation, we get initialization P s , C s , B s The first calculation method is After the first transformation, we get
[0101] 3. The position transformation relationship between point cloud targets is obtained through the second calculation method [R, T].
[0102] 4. [R, T] After the third calculation method, the inverse transformation relationship [R T , -R T T], R T , -R T T] is The transformation relationship is also The transformation relationship is also transformation relationship.
[0103] 5. With [R T , -R T T] is the transformation relationship, and after the second transformation method, we get Output
[0104] Based on the above inter-frame transfer algorithm, step 202 specifically includes:
[0105] The point cloud target in the point cloud data of the previous frame is registered with the point cloud target in the point cloud data of the current frame, the spatial transformation relationship between the point cloud targets of the two frames is determined, and the spatial transformation relationship between the point cloud targets of the two frames is used as the spatial transformation relationship of the point cloud target annotation boxes of the two frames. The point cloud target box of the previous frame tracks the point cloud target box of the current frame to obtain the point cloud transfer box of the current frame. Specifically:
[0106] (1) Determine the previous frame data, current frame data, previous frame transformation parameters and current frame transformation parameters according to the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame; the previous frame data includes: the entire point cloud of the previous frame, the local coordinate system of the ego vehicle of the previous frame and the point cloud target frame of the previous frame; the current frame data includes: the entire point cloud of the current frame and the local coordinate system of the ego vehicle of the current frame; the previous frame transformation parameters include: the rotation matrix and translation vector of the local coordinate system of the ego vehicle of the previous frame transformed to the global coordinate system; the current frame transformation parameters include: the rotation matrix and translation vector of the local coordinate system of the ego vehicle of the current frame transformed to the global coordinate system.
[0107] (2) Determine all point clouds within the point cloud target frame of the previous frame in the local coordinate system of the vehicle in the previous frame according to the previous frame data to obtain a first point cloud. This step is implemented using the first calculation method.
[0108] (3) According to the transformation parameters of the previous frame, the first point cloud is transformed into a global coordinate system to obtain a second point cloud. This step is implemented using the first transformation method.
[0109] (4) The point cloud target frame of the previous frame is converted to the global coordinate system, and the converted target frame is determined as the point cloud target frame initialized in the current frame. This step is implemented using the second transformation method.
[0110] (5) Determine all point clouds within the point cloud target frame initialized in the vehicle local coordinate system of the current frame according to the current frame data and the point cloud target frame initialized in the current frame to obtain a third point cloud. This step is implemented using the first calculation method.
[0111] (6) According to the current frame transformation parameters, the third point cloud is transformed into a global coordinate system to obtain a fourth point cloud. This step is implemented using the first transformation method.
[0112] (7) Aligning the second point cloud with the fourth point cloud, and calculating the spatial transformation relationship between the fourth point cloud and the second point cloud. This step is implemented using the second calculation method.
[0113] (8) Calculate the inverse transformation relationship based on the spatial transformation relationship between the fourth point cloud and the second point cloud, and determine the inverse transformation relationship as the spatial transformation relationship between the two frame point cloud objects. This step is implemented using the third calculation method.
[0114] (9) The point cloud target frame initialized in the current frame is transformed according to the spatial transformation relationship between the point cloud targets of the two frames to obtain the point cloud transfer frame of the current frame. This step is implemented using the second transformation method. The point cloud transfer frame includes: the coordinates of the center point of the frame, the rotation matrix of the frame, the size of the frame, and the category of the frame.
[0115] In another exemplary embodiment of the present application, step 203 specifically includes: according to the point cloud data of the current frame, using a deep learning point cloud target detection model to determine the point cloud reasoning box of the current frame; the deep learning point cloud target detection model is a Det3D target detection framework, a CenterPoint model or a dynamic sparse voxel transformer (DSVT). The input of this step is a single frame or continuous frame point cloud, and the output is a 3D target reasoning box of the traffic participant in the current frame point cloud.
[0116] Target reasoning box definition: The perception result obtained by the point cloud target detection model is the target reasoning box, B iis a target reasoning box of f1, defined as: B i =(c i , r i , d i , class, score), class is the type of box, which needs to be cross-validated by category later, and the calculation range is narrowed by category restriction. At the same time, the category is also a necessary attribute of the labeled box, and score is the confidence, which is output by the model. i The higher the score, the more confident the model is about B. i The correctness of is the inference box set of f1. It can be seen that the point cloud inference box includes: the center point coordinates of the box, the rotation matrix of the box, the size of the box, the category of the box and the confidence of the box.
[0117] In another exemplary embodiment of the present application, step 204 specifically includes:
[0118] (1) Match the point cloud transfer box of the current frame with the point cloud inference box of the current frame to determine the intersection-union ratio.
[0119] (2) The point cloud inference frame whose intersection-over-union ratio is greater than or equal to a first set value is determined as a matching frame, and the point cloud inference frame whose intersection-over-union ratio is less than the first set value is determined as an unmatched frame. The first set value may be 0.7.
[0120] (3) A matching box whose confidence is less than or equal to a second setting value is determined as a dead box, and a matching box whose confidence is greater than the second setting value is determined as a cross box. The second setting value may be 0.5.
[0121] (4) Determine the unmatched frames whose confidence is less than or equal to a third set value as eliminated frames, and determine the unmatched frames whose confidence is greater than the third set value as new frames. The third set value may be 0.9.
[0122] (5) Merge the cross frame and the new frame to obtain the point cloud proposal frame of the current frame.
[0123] It should be noted that the values of the first setting value, the second setting value and the third setting value are only used as examples and are not intended to be limiting. In practical applications, their values are determined according to actual needs.
[0124] Combine the following Figure 6 , further introduces the point cloud box cross-validation algorithm in practical applications.
[0125] like Figure 6As shown in the figure, two sets of point cloud annotation boxes are obtained through different self-annotation methods, and the two sets are matched and screened to determine the final self-annotation results. Cross-validation of different data sources is used to improve the reliability of the self-annotation results. Among them, matching: the reasoning boxes are divided into matching boxes and unmatched boxes according to the intersection and union ratio; screening: matching boxes and unmatched boxes are retained or eliminated according to the confidence.
[0126] (1) Point cloud box matching.
[0127] Input: A set of reasoning boxes of the same category Transfer box collection set up There are m boxes in There are n boxes in .
[0128] in,
[0129] Output: and The matching results.
[0130] Matching process:
[0131] and Construct the IoU (Intersection over Union) matrix, which is the ratio of the intersection area to the union area of the two annotation boxes, referred to as the intersection over union ratio.
[0132] (2) Point cloud frame screening.
[0133] According to the IoU matrix, determine and The matching relationship between elements between the two boxes is realized by maximizing the sum of the IoU values corresponding to all matching boxes through the Hungarian matching algorithm. and Best match for the frame.
[0134] According to the matching results, Divide into matching sets Does not match the set That is, IoU ≥ 0.7 enter Otherwise enter
[0135] right Filter by:
[0136] The boxes in the figure represent boxes detected by the target detection model and considered to exist by the transfer algorithm from the previous frame to the next frame. This is a set of point cloud boxes that have been double-verified. The credibility is high, but because the point cloud target has a process of generation, existence and exit in the scene, it is necessary to use the score (confidence) to filter again
[0137] The following takes the second set value as 0.5 and the third set value as 0.9 as an example to illustrate the screening process.
[0138] score≥0.5: a point cloud box that has been cross-validated and has high confidence is classified as a cross box; score<0.5: a point cloud box that has been cross-validated but has low confidence, which may be a point cloud target that is about to exit the scene. The point cloud is occluded and sparse, and the annotation accuracy is not high. It should be excluded and classified as a dead box.
[0139] right Filter by:
[0140] The box in the figure indicates that the target detection model detects that the target exists, but the transfer algorithm does not match the box passed down from the previous frame.
[0141] score≥0.9: There is no cross-matching, but the model confidence is very high. It may be the frame that enters the scene in the first frame and is classified as a new frame. score<0.9: There is no cross-matching and the confidence is not high enough. It is classified as an eliminated frame.
[0142] Final target proposal frame determination:
[0143] The cross-validated cross box and the new box with high confidence are merged into the target proposal box. The target proposal box is the final self-annotation result of the frame point cloud. At the same time, the target proposal box is also the initial box for the next frame point cloud to execute the transfer algorithm.
[0144] The present application also provides an autonomous driving application scenario, which applies the above-mentioned point cloud target self-labeling method. Specifically: obtain the point cloud data of the current frame in the autonomous driving application scenario and the point cloud data and point cloud target frame of the previous frame; according to the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame, use the inter-frame transfer algorithm to determine the point cloud transfer frame of the current frame; according to the point cloud data of the current frame, use the model inference algorithm to determine the point cloud inference frame of the current frame; use the cross-validation algorithm to match and filter the point cloud transfer frame of the current frame and the point cloud inference frame of the current frame to obtain the point cloud proposal frame of the current frame; the point cloud proposal frame is the final point cloud target frame. Among them, the targets in the autonomous driving application scenario are traffic participants, such as vehicles, pedestrians, etc. The present application realizes the self-labeling of traffic participants in this application scenario, and has never achieved accurate perception of the physical world.
[0145] At present, there are several related methods for automatic annotation of point cloud targets: one is to use multiple different types of point cloud deep learning models to perform target detection separately, and then fuse the detection results of multiple models to obtain the final detection result. The present application obtains two sets of self-annotated target boxes through a rule-based inter-frame transfer algorithm and a model-based reasoning algorithm, and then uses the proposed cross-validation algorithm for screening and verification. It not only uses different methods to test and screen the point cloud annotation results, but also can effectively reduce the overfitting risk of the model-based method, and enhance the generalization ability for unknown data. The second is to use point cloud and image modalities for joint detection and annotation. In terms of data collection, additional data space calibration and time synchronization are required, and additional training of image modality deep learning models is required. However, the present application only needs point cloud data for multi-information source cross-validation. It can be seen that the present application not only has more accurate annotation results, but also has strong generalization ability and is simple and easy to implement.
[0146] Based on the same inventive concept, the embodiment of the present application also provides a point cloud target self-labeling device for implementing the point cloud target self-labeling method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more point cloud target self-labeling device embodiments provided below can refer to the limitations of the point cloud target self-labeling method above, and will not be repeated here.
[0147] In an exemplary embodiment, Figure 7 As shown, a point cloud target self-labeling device is provided, comprising:
[0148] The data acquisition module 701 is used to acquire the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame.
[0149] The inter-frame transfer module 702 is used to determine the point cloud transfer frame of the current frame using the inter-frame transfer algorithm based on the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame; the point cloud transfer frame is a set of self-labeled point cloud target frames.
[0150] The model inference module 703 is used to determine the point cloud inference frame of the current frame using a model inference algorithm based on the point cloud data of the current frame; the point cloud inference frame is another set of self-labeled point cloud target frames.
[0151] The cross-validation module 704 is used to match and filter the point cloud transfer frame of the current frame and the point cloud inference frame of the current frame using a cross-validation algorithm to obtain a point cloud proposal frame of the current frame; the point cloud proposal frame is the final point cloud target frame.
[0152] In this embodiment, for the original point cloud modal data, two sets of self-annotated target boxes are obtained through a rule-based inter-frame transfer algorithm and a model-based reasoning algorithm, and then the proposed cross-validation algorithm is used for screening and verification to ensure the annotation quality.
[0153] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a point cloud target self-labeling method is implemented.
[0154] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0155] In an exemplary embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0156] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0157] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0158] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0159] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, etc., but is not limited thereto.
[0160] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0161] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A point cloud target self-labeling method, characterized in that: The point cloud target self-labeling method comprises: Get the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame; According to the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame, an inter-frame transfer algorithm is used to determine the point cloud transfer frame of the current frame; the point cloud transfer frame is a set of self-annotated point cloud target frames; According to the point cloud data of the current frame, a model inference algorithm is used to determine the point cloud inference frame of the current frame; the point cloud inference frame is another set of self-annotated point cloud target frames; A cross-validation algorithm is used to match and screen the point cloud transfer frame of the current frame and the point cloud inference frame of the current frame to obtain a point cloud proposal frame of the current frame; the point cloud proposal frame is the final point cloud target frame.
2. The point cloud target self-labeling method according to claim 1, characterized in that: According to the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame, the point cloud transfer frame of the current frame is determined by using the inter-frame transfer algorithm, which specifically includes: The point cloud target in the point cloud data of the previous frame is aligned with the point cloud target in the point cloud data of the current frame, the spatial transformation relationship between the point cloud targets of the two frames is determined, and the spatial transformation relationship between the point cloud targets of the two frames is used as the spatial transformation relationship of the point cloud target annotation boxes of the two frames. The point cloud target box of the previous frame tracks the point cloud target box of the current frame to obtain the point cloud transfer box of the current frame.
3. The point cloud target self-labeling method according to claim 2, characterized in that: The point cloud target in the point cloud data of the previous frame is registered with the point cloud target in the point cloud data of the current frame, the spatial transformation relationship between the point cloud targets of the two frames is determined, and the spatial transformation relationship between the point cloud targets of the two frames is used as the spatial transformation relationship of the point cloud target annotation frames of the two frames. The point cloud target frame of the previous frame tracks the point cloud target frame of the current frame to obtain the point cloud transfer frame of the current frame, which specifically includes: Determine the previous frame data, current frame data, previous frame transformation parameters and current frame transformation parameters according to the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame; the previous frame data includes: the entire point cloud of the previous frame, the local coordinate system of the ego vehicle of the previous frame and the point cloud target frame of the previous frame; the current frame data includes: the entire point cloud of the current frame and the local coordinate system of the ego vehicle of the current frame; the previous frame transformation parameters include: the rotation matrix and translation vector of the local coordinate system of the ego vehicle of the previous frame transformed to the global coordinate system; the current frame transformation parameters include: the rotation matrix and translation vector of the local coordinate system of the ego vehicle of the current frame transformed to the global coordinate system; Determine all point clouds within the point cloud target frame of the previous frame in the local coordinate system of the vehicle in the previous frame according to the previous frame data to obtain a first point cloud; Convert the first point cloud to a global coordinate system according to the previous frame transformation parameters to obtain a second point cloud; The point cloud target frame of the previous frame is converted to the global coordinate system, and the converted target frame is determined as the point cloud target frame initialized in the current frame; Determine all point clouds within the point cloud target frame initialized in the vehicle local coordinate system of the current frame according to the current frame data and the point cloud target frame initialized in the current frame, and obtain a third point cloud; Converting the third point cloud to a global coordinate system according to the current frame transformation parameters to obtain a fourth point cloud; Registering the second point cloud with the fourth point cloud, and calculating a spatial transformation relationship between the fourth point cloud and the second point cloud; Calculate the inverse transformation relationship according to the spatial transformation relationship between the fourth point cloud and the second point cloud, and determine the inverse transformation relationship as the spatial transformation relationship between the point cloud objects of the two frames; The point cloud target frame initialized in the current frame is transformed according to the spatial transformation relationship between the point cloud targets of the two frames to obtain the point cloud transfer frame of the current frame.
4. The point cloud target self-labeling method according to claim 1, characterized in that: The cross-validation algorithm is used to match and filter the point cloud transfer frame of the current frame and the point cloud inference frame of the current frame to obtain the point cloud proposal frame of the current frame, which includes: Match the point cloud transfer frame of the current frame with the point cloud inference frame of the current frame to determine the intersection-union ratio; Determine the point cloud inference frame whose intersection-over-union ratio is greater than or equal to the first set value as a matched frame, and determine the point cloud inference frame whose intersection-over-union ratio is less than the first set value as an unmatched frame; Determine a matching box whose confidence is less than or equal to a second set value as a death box, and determine a matching box whose confidence is greater than the second set value as a cross box; Determine an unmatched frame whose confidence is less than or equal to a third set value as an eliminated frame, and determine an unmatched frame whose confidence is greater than the third set value as a new frame; The cross frame and the new frame are merged to obtain the point cloud proposal frame of the current frame.
5. The point cloud target self-labeling method according to claim 1, characterized in that: According to the point cloud data of the current frame, the model inference algorithm is used to determine the point cloud inference frame of the current frame, including: According to the point cloud data of the current frame, a deep learning point cloud target detection model is used to determine the point cloud inference box of the current frame; the deep learning point cloud target detection model is a Det3D target detection framework, a CenterPoint model or a dynamic sparse voxel transformer.
6. The point cloud target self-labeling method according to claim 1, characterized in that: The point cloud transfer frame includes: the center point coordinates of the frame, the rotation matrix of the frame, the size of the frame and the category of the frame; the point cloud inference frame includes: the center point coordinates of the frame, the rotation matrix of the frame, the size of the frame, the category of the frame and the confidence of the frame.
7. The point cloud target self-labeling method according to claim 4, characterized in that: The first setting value is 0.7; the second setting value is 0.5; and the third setting value is 0.
9.
8. A point cloud target self-labeling device, characterized in that: The point cloud target self-labeling device comprises: A data acquisition module is used to acquire the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame; An inter-frame transfer module is used to determine the point cloud transfer frame of the current frame using an inter-frame transfer algorithm based on the point cloud data of the current frame and the point cloud data and point cloud target frame of the previous frame; the point cloud transfer frame is a set of self-annotated point cloud target frames; A model inference module is used to determine a point cloud inference frame of the current frame using a model inference algorithm according to the point cloud data of the current frame; the point cloud inference frame is another set of self-annotated point cloud target frames; The cross-validation module is used to match and filter the point cloud transfer frame of the current frame and the point cloud inference frame of the current frame using a cross-validation algorithm to obtain a point cloud proposal frame of the current frame; the point cloud proposal frame is the final point cloud target frame.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the point cloud target self-labeling method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the point cloud target self-labeling method described in any one of claims 1 to 7 is implemented.