Dynamic target object labeling method and device, equipment and storage medium
By acquiring multi-frame point cloud data and utilizing automatic annotation models and user operations, the problem of high cost and low efficiency in dynamic object annotation is solved, and fast and easy dynamic target object annotation is achieved, which is suitable for the training of intelligent annotation platforms and dynamic perception models.
Patent Information
- Application Number
- CN202510812442.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
The existing technology for labeling dynamic objects is costly and inefficient, especially the high cost and low efficiency caused by manual frame-by-frame labeling.
By acquiring multi-frame point cloud data with a time-series relationship, using the automatic annotation model to output the first annotation information, determining the first point cloud set from the multi-frame point cloud based on the dynamic target and displaying it on the annotation interface, and performing automatic annotation operations in response to user operations, rapid annotation of dynamic targets can be achieved.
It achieves fast and easy labeling of dynamic targets, reduces labeling costs, and improves labeling efficiency. It is suitable for the training of intelligent labeling platforms and dynamic perception models.
Smart Images

Figure CN120708223A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a method, device, equipment, and storage medium for labeling dynamic targets. Background Art
[0002] Model training is an indispensable part of model application, and the model training process usually requires a large amount of labeled data. Therefore, how to label a large amount of sample data is a very important link.
[0003] During model training, a large amount of sample data must be accurately labeled to obtain the true value of the data. This true value is then used for model training, thereby improving the accuracy of model prediction or perception. Labeling a large amount of sample data, including dynamic objects, is a common practice. Manual labeling of dynamic objects in sample data typically increases manual labeling costs, resulting in relatively low labeling efficiency. Summary of the Invention
[0004] In order to solve the above technical problems, the present disclosure provides a method and apparatus for labeling dynamic targets, a device, and a storage medium to solve the problems of high labeling cost and low labeling efficiency for labeling dynamic objects.
[0005] A first aspect of the present disclosure provides a method for labeling a dynamic target, comprising: obtaining a plurality of frame point clouds having a temporal relationship, and first labeling information for the dynamic target in at least one frame point cloud in the plurality of frame point clouds; based on the dynamic target, determining a first point cloud set containing the dynamic target from the plurality of frame point clouds and displaying the result in a labeling interface; and in response to a labeling operation in the labeling interface, performing an automatic labeling operation on at least one frame point cloud in the first point cloud set based on the first labeling information to obtain a labeling result for the dynamic target in the first point cloud set.
[0006] A second aspect of the present disclosure provides a device for labeling dynamic targets, including: a data acquisition module for acquiring multiple frames of point clouds with a time-series relationship, and first labeling information for the dynamic target in at least one frame of the multiple frames of point clouds; a determination and display module for determining, based on the dynamic target, a first point cloud set containing the dynamic target from the multiple frames of point clouds and displaying it in a labeling interface; an automatic labeling module for responding to a labeling operation in the labeling interface and performing an automatic labeling operation on at least one frame of point cloud in the first point cloud set based on the first labeling information to obtain a labeling result of the dynamic target in the first point cloud set.
[0007] A third aspect of the present disclosure provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it is used to implement the dynamic target object labeling method provided in the first aspect.
[0008] The fourth aspect of the present disclosure provides an electronic device comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the dynamic target object labeling method provided in the first aspect above.
[0009] A fifth aspect of the present disclosure provides a computer program product. When instructions in the computer program product are executed by a processor, the dynamic target object labeling method provided in the first aspect is executed.
[0010] Based on the method for labeling dynamic targets provided by the present disclosure, by obtaining multiple frame point clouds with a time-sequential relationship, and first labeling information for the dynamic target from at least one frame point cloud in the multiple frame point clouds; based on the dynamic target, a first point cloud set containing the dynamic target is determined from the multiple frame point clouds and displayed in the labeling interface; in response to the labeling operation in the labeling interface, an automatic labeling operation is performed on at least one frame point cloud in the first point cloud set based on the first labeling information, to obtain the labeling result of the dynamic target in the first point cloud set. It can be seen that the present disclosure utilizes the first labeling information for the dynamic target in the multiple frame time-sequential point clouds to automatically label the dynamic target in other frames of the first point cloud set containing the dynamic target, so as to complete the one-time labeling of the dynamic target in the multiple frame point clouds, thereby simply and quickly labeling the dynamic target in the first point cloud set, avoiding the problem of high labeling cost and low labeling efficiency caused by manually labeling the dynamic target one by one in all frames. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1A A flowchart of a marking method provided by an exemplary embodiment of the present disclosure.
[0012] Figure 1B A schematic diagram of a multi-frame point cloud acquisition method provided by an exemplary embodiment of the present disclosure.
[0013] Figure 2 A schematic diagram of a annotation interface provided by an exemplary embodiment of the present disclosure.
[0014] Figure 3 A flowchart of a marking method provided by another exemplary embodiment of the present disclosure.
[0015] Figure 4 A flowchart of a marking method provided in accordance with another exemplary embodiment of the present disclosure is provided.
[0016] Figure 5A A flowchart of a marking method provided in accordance with another exemplary embodiment of the present disclosure is provided.
[0017] Figure 5BA schematic diagram illustrating the principle of an automatic annotation function provided by an exemplary embodiment of the present disclosure.
[0018] Figure 5C A schematic diagram of an automatically labeled scene provided by an exemplary embodiment of the present disclosure.
[0019] Figure 6 A flowchart of a marking method provided in accordance with another exemplary embodiment of the present disclosure is provided.
[0020] Figure 7 A schematic structural diagram of a labeling device provided by an exemplary embodiment of the present disclosure.
[0021] Figure 8 A schematic structural diagram of a labeling device provided by another exemplary embodiment of the present disclosure.
[0022] Figure 9 This is a schematic structural diagram of a labeling device provided by yet another exemplary embodiment of the present disclosure.
[0023] Figure 10 The present invention provides a structural diagram of an electronic device according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0024] To explain the present disclosure, example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. It should be understood that the present disclosure is not limited to the example embodiments.
[0025] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.
[0026] Application Overview
[0027] Typically, related technologies use frame slicing to achieve dynamic object detection and tracking. However, this approach requires manual labeling of dynamic objects frame by frame, resulting in low labeling efficiency and a lack of quality verification.
[0028] In order to solve the above problems, an embodiment of the present disclosure provides a method for labeling dynamic targets, which obtains multi-frame point clouds with a time-sequential relationship and first labeling information for the dynamic target in at least one frame of the multi-frame point cloud; based on the dynamic target, a first point cloud set containing the dynamic target is determined from the multi-frame point cloud and displayed in the labeling interface; in response to the labeling operation in the labeling interface, an automatic labeling operation is performed on at least one frame of the point cloud in the first point cloud set based on the first labeling information to obtain the labeling result of the dynamic target in the first point cloud set. It can be seen that the present disclosure utilizes the first labeling information for the dynamic target in the multi-frame time-sequential point cloud to automatically label the dynamic target in other frames of the first point cloud set containing the dynamic target, so as to complete the one-time labeling of the dynamic target in the multi-frame point cloud, thereby simply and quickly labeling the dynamic target in the first point cloud set, avoiding the problem of high labeling cost and low labeling efficiency caused by manual labeling of the dynamic target in all frames one by one.
[0029] Furthermore, the dynamic object labeling method provided in the embodiments of the present disclosure can be applied to an intelligent labeling platform having a visual labeling interface. In the embodiments of the present disclosure, a multi-frame point cloud having a temporal relationship and first labeling information of at least one dynamic object in the multi-frame point cloud can be imported into the intelligent labeling platform, so that the intelligent labeling platform can classify the at least one dynamic object based on the first labeling information and automatically label the at least one dynamic object based on the user's labeling operation on the labeling interface, thereby obtaining a labeling result for the at least one dynamic object in the multi-frame point cloud.
[0030] In this way, the above-mentioned dynamic target object labeling method can be used to obtain sample data with labeling results, and the sample data with labeling results can be used as training samples to train the dynamic perception model until the preset training completion conditions are met, thereby obtaining a trained perception model that can be used for downstream dynamic perception tasks.
[0031] Exemplary Methods
[0032] Figure 1A This is a flow chart of a method for marking dynamic objects provided by an embodiment of the present disclosure. This embodiment can be applied to any electronic device such as a local terminal device, a cloud server, etc., and can also be applied to terminal devices and cloud servers in a distributed manner, such as Figure 1A As shown, the method includes the following steps S101 to S103.
[0033] Step S101: Acquire multiple frames of point clouds having a time sequence relationship, and first annotation information of at least one frame of the multiple frames of point clouds for a dynamic target.
[0034] For example, the multi-frame point cloud with a time sequence relationship in the embodiment of the present disclosure refers to point cloud data collected at different time points. These data are continuous in time and can reflect the dynamic changes of objects in the collection environment. Figure 1B As shown, the vehicle 10 is provided with a plurality of cameras 11 in different directions and a three-dimensional scanning device 12 (for example, a lidar). During the driving process of the vehicle 10, the cameras 11 in different directions on the vehicle 10 can respectively capture images of the surrounding environment at a first preset frequency to obtain image sequences at multiple different perspectives; wherein, the image frames in each image sequence are sorted according to the image acquisition time or the driving trajectory of the vehicle 10. In this way, the image sequences at multiple different perspectives can be processed by multi-view reconstruction to obtain multi-frame point clouds with a temporal relationship. During the driving process of the vehicle 10, the three-dimensional scanning device 12 can also capture point cloud data of the surrounding environment at a second preset frequency to obtain multi-frame point clouds with a temporal relationship. wherein, the first preset frequency and the second preset frequency may be the same or different. It should be noted that, Figure 1B This is only an example, and the embodiments of the present disclosure Figure 1B There is no limitation on the number and position of the multiple cameras 11 and the three-dimensional scanning device 12. In actual use, the multiple cameras 11 include but are not limited to: a vehicle's front-view wide-angle camera, a vehicle's front-view narrow-angle camera, and a vehicle's side / rear / surround / peripheral cameras.
[0035] In some examples, multi-frame point cloud data with a time-series relationship is usually collected by sensors at different time points, and each frame of point cloud is used to record the environmental information at the corresponding collection moment. These point cloud data are combined in chronological order to construct dynamic information such as the motion trajectory and orientation change of dynamic objects. It should be noted that the specific method of obtaining multi-frame point clouds with a time-series relationship in the embodiments of the present disclosure is not limited. Those skilled in the art can choose an appropriate method to obtain multi-frame time-series point clouds according to actual conditions. In some examples, the dynamic targets in the embodiments of the present disclosure include but are not limited to: pedestrians, vehicles, bicycles, motorcycles, animals, etc.
[0036] In some examples, multiple frames of point clouds can be input into a trained automatic annotation model, and the output of the automatic annotation model is the first annotation information for the dynamic target object in at least one frame of the multiple frame point cloud. The first annotation information includes, but is not limited to: a 3D (3D) annotation box of the dynamic target object, an ID (Identification) of the dynamic target object, and the attributes of the dynamic target object. The automatic annotation model includes, for example, a target detection and tracking model, an instance segmentation model, and the like. The attributes of the dynamic target object may include the category information of the dynamic target object, such as whether the dynamic target object is a vehicle or a pedestrian, or whether the dynamic target object is a large vehicle or a small vehicle when the dynamic target object is a vehicle. Of course, the attributes of the dynamic target object may also include the size information of the dynamic target object, such as the size of the vehicle when the dynamic target object is a vehicle. It should be noted that the specific implementation method for obtaining the first annotation information of the dynamic target object in the embodiment of the present disclosure is not limited. Those skilled in the art can use the automatic annotation model to obtain the first annotation information of the dynamic target object or use other methods to obtain the first standard information of the dynamic target object.
[0037] For example, when the dynamic target is another vehicle A, the sensors on the vehicle can be used to collect data on the surrounding environment at different time points to obtain a multi-frame point cloud with a time sequence relationship. The multi-frame point cloud with a time sequence relationship includes point cloud frames 1 to 50, that is, 50 frames of point cloud data. Point cloud frames 1 to 50 can be input into a trained automatic annotation model so that the automatic annotation model outputs the annotation information of point cloud frames 1 to 50. However, due to structural defects in the automatic annotation model or low training accuracy of the automatic annotation model, the automatic annotation results are inaccurate or missing. For example, the automatic annotation model only outputs the first annotation information of the other vehicle A in point cloud frames 1 to 25, and / or the first annotation information of the other vehicle A in point cloud frames 40 to 50, and omits or lacks the annotation results of other point cloud frames.
[0038] Step S102: Based on the dynamic target, a first point cloud set including the dynamic target is determined from the multi-frame point cloud and displayed in the annotation interface.
[0039] For example, if the dynamic target in the ego vehicle's surroundings is another vehicle A, the ego vehicle's sensors can be used to collect data about the surroundings at different acquisition times to obtain a multi-frame point cloud with a time-series relationship. This multi-frame point cloud with a time-series relationship includes point cloud frames 1 through 100, i.e., 100 frames of point cloud data. Among them, the other vehicle A entered the ego vehicle's surroundings at the acquisition time corresponding to point cloud frame 1 and exited the surroundings at the acquisition time corresponding to point cloud frame 51. Furthermore, if only point cloud frames 1 through 50 in the multi-frame point cloud contain the other vehicle A, the first point cloud set determined is the point cloud set including point cloud frames 1 through 50. Furthermore, point cloud frames 1 through 50 containing the other vehicle A can be displayed in the annotation interface.
[0040] In some examples, the first point cloud set containing the dynamic target can be visualized so that it can be displayed on a visual interactive interface (i.e., the annotation interface in step S102 above). For example, the first point cloud set and the determined first annotation information can be visualized using the visualization module in the PCL (Point Cloud Library) or Cloud Compare (point cloud registration) visualization software. For example, the first point cloud set including point cloud frames 1 to 50 can be used as the tracking set for the other vehicle A. Furthermore, point cloud frames 1 to 50 can be displayed in the annotation interface. At the same time, if the above-mentioned point cloud frames 1 to 100 are input into the automatic annotation model, the first annotation information of the other car A in the point cloud frames 1 to 25 and the point cloud frames 40 to 50 is obtained, then the first annotation information of the other car A can be displayed in the point cloud frames 1 to 25 and the point cloud frames 40 to 50 of the annotation interface at the same time (such as displaying the 3D detection frame of the other car A and the tracking ID of the other car A).
[0041] Figure 2 Schematic diagram of the visualization interface of the first point cloud set and the first annotation information in the embodiment of the present disclosure, such as Figure 2 As shown, the tracking ID of the dynamic target 21a is 51, and the first point cloud set 22a containing the dynamic target 21a is displayed on the right side of the visualization interface. The first point cloud set 22a includes point cloud frames 2 to 51, that is, 50 frames of point cloud data. Among them, the first annotation information of the dynamic target 21a is contained in point cloud frames 2, 8-9, 12-22, 24-25, and 29-51. Figure 2 As shown in FIG, the frame numbers corresponding to the multiple point cloud frames with the first annotation information of the dynamic target 21a are highlighted. Similarly, since the first annotation information of the dynamic target 21 is not present in the point cloud frames 3-7, 10-11, 23, and 26-28 (because the automatic annotation model may have false detection or missed detection), Figure 2As shown in FIG, the frame number corresponding to the multi-frame point cloud in which the first annotation information of the dynamic target object 21a does not exist can be low-lit. Figure 2 As shown, in response to the user's click operation on the frame number corresponding to the point cloud frame 8 in the annotation interface, the first annotation information of the dynamic target object 21a in the point cloud frame 8 can be displayed on the left side of the interface. The first annotation information includes the annotation box of the dynamic target object 21a and the tracking ID 51 of the dynamic target object 21a.
[0042] In actual applications, there may be multiple dynamic targets in a multi-frame point cloud with a time sequence relationship. In the embodiment of the present disclosure, the first point cloud set corresponding to each dynamic target can be determined and displayed in the annotation interface. Figure 2 The right side of the annotation interface shows that in addition to displaying the first point cloud set 22a of the dynamic target 21a (tracking ID is 51), the first point cloud sets (tracking sets) of other dynamic targets can also be displayed. That is, the first point cloud set 23a of the dynamic target with tracking ID 82 and the first point cloud set 24a of the dynamic target with tracking ID 84 can also be displayed. Figure 2 The diagram shows a first point cloud set 23a of a dynamic target with tracking ID 82, which includes 19 frames of point clouds; and a first point cloud set 24a of a dynamic target with tracking ID 84, which includes 28 frames of point clouds.
[0043] Step S103 : In response to the annotation operation in the annotation interface, an automatic annotation operation is performed on at least one frame of point cloud in the first point cloud set based on the first annotation information to obtain an annotation result of the dynamic target object in the first point cloud set.
[0044] For example, the above-mentioned visual interactive interface (i.e., the annotation interface of step S102) can be used not only to display the determined first point cloud set and the first annotation information, but also to automatically annotate point cloud frames in the first point cloud set that do not contain the first annotation information in response to a user's annotation operation (e.g., a user clicking an automatic annotation function button on the visual interactive interface). The visual interactive interface can display data via a display (e.g., a liquid crystal display or plasma display) and can also receive user operation commands via a keyboard, mouse, touch screen, etc.
[0045] For example, Figure 2As shown, in response to the user clicking the automatic annotation function button in the annotation interface, the point cloud frames 3-7, point cloud frames 10-11, point cloud frames 23 and point cloud frames 26-28 in the first point cloud set 22a can be automatically annotated to mark out the 3D annotation box and tracking ID 51 of the dynamic target object 21a in the point cloud frames 3-7, point cloud frames 10-11, point cloud frames 23 and point cloud frames 26-28.
[0046] That is to say, in the embodiment of the present disclosure, for a multi-frame point cloud with a time-series relationship, it can be input into the automatic annotation model to obtain the first annotation information (i.e., pre-flash data) and then imported into the annotation tool. Assuming that there are 3D annotation boxes and tracking IDs of dynamic targets such as other vehicles and pedestrians in the pre-flash data, the point cloud frames corresponding to the dynamic targets with the same tracking ID can be quickly classified into the corresponding tracking set in the annotation tool; wherein, multiple dynamic targets correspond to their own tracking sets (i.e., the first point cloud set). Furthermore, you can browse through the tracking set to see which point cloud frames are included in a certain classified tracking ID, thereby realizing the one-click automatic annotation function; at the same time, you can also manually adjust the existing first annotation information on the annotation interface. For example, the first annotation information can be adjusted under three views (top view, front view, side view), so that the 3D box of the first annotation information can be displayed from each perspective to facilitate user adjustment.
[0047] The method for labeling a dynamic target provided by an embodiment of the present disclosure obtains a multi-frame point cloud having a time-sequential relationship, and first labeling information for the dynamic target from at least one frame of the multi-frame point cloud; based on the dynamic target, a first point cloud set containing the dynamic target is determined from the multi-frame point cloud and displayed in the labeling interface; in response to the labeling operation in the labeling interface, an automatic labeling operation is performed on at least one frame of the point cloud in the first point cloud set based on the first labeling information, to obtain the labeling result of the dynamic target in the first point cloud set. It can be seen that the present disclosure utilizes the first labeling information for the dynamic target in the multi-frame time-sequential point cloud to automatically label the dynamic target in other frames of the first point cloud set containing the dynamic target, so as to complete the one-time labeling of the dynamic target in the multi-frame point cloud, thereby simply and quickly labeling the dynamic target in the first point cloud set, avoiding the problem of high labeling cost and low labeling efficiency caused by manually labeling the dynamic target one by one in all frames.
[0048] like Figure 3 As shown in the above Figure 1A Based on the illustrated embodiment, step S102 may include the following steps S1021 to S1023.
[0049] Step S1021: Based on the first annotation information of the dynamic target object, determine a second point cloud set containing the first annotation information from the multi-frame point cloud.
[0050] For example, after processing multiple point cloud frames with a temporal relationship, first annotation information for a dynamic object can be obtained for at least one point cloud frame in the multiple point cloud frames. Furthermore, a point cloud frame containing the first annotation information can be determined from the multiple point cloud frames, and the multiple point cloud frames containing the first annotation information can be determined as a second point cloud set.
[0051] For example, point cloud frames 1 to 25 including the first annotation information and point cloud frames 40 to 50 including the first annotation information among the point cloud frames 1 to 100 may be determined as the second point cloud set.
[0052] Step S1022: Based on the second point cloud set, determine a first frame point cloud and a second frame point cloud from the second point cloud set; wherein the first frame point cloud is the frame point cloud in the second point cloud set that includes the first annotation information for the first time, and the second frame point cloud is the frame point cloud in the second point cloud set that includes the first annotation information for the last time.
[0053] Exemplarily, in an embodiment of the present disclosure, based on the acquisition time of each frame point cloud in the second point cloud set, the frame point cloud that first includes the first annotation information and the frame point cloud that last includes the first annotation information can be determined from the second point cloud set. Among them, the point cloud frame that first includes the first annotation information refers to the point cloud frame in which a dynamic target object is detected for the first time when target detection is performed on multiple frame point clouds with a temporal relationship, and the point cloud frame that last includes the first annotation information refers to the point cloud frame in which the dynamic target object is detected for the last time when target detection is performed on multiple frame point clouds with a temporal relationship. For example, if the second point cloud set includes point cloud frames 1 to 25, and point cloud frames 40 to 50, then the point cloud frame that first includes the first annotation information is point cloud frame 1, and the point cloud frame that last includes the first annotation information is point cloud frame 50. Thus, point cloud frame 1 can be determined as the first frame point cloud, and point cloud frame 50 can be determined as the second frame point cloud.
[0054] It should be noted that the first point cloud frame in the disclosed embodiment is the first point cloud frame in the second point cloud set to include the first annotation information, not the first point cloud frame acquired. Similarly, the second point cloud frame is the last point cloud frame in the second point cloud set to include the first annotation information, and is also unrelated to the acquisition time.
[0055] Step S1023: Based on the first frame point cloud and the second frame point cloud, a first point cloud set is determined from the multiple frame point clouds and displayed in the annotation interface.
[0056] For example, since the first point cloud frame is the first point cloud frame in the second point cloud set to include the first annotation information, and the second point cloud frame is the last point cloud frame in the second point cloud set to include the first annotation information, the first point cloud frame was acquired earlier than the second point cloud frame, and the multiple point cloud frames with a temporal relationship are point cloud data acquired continuously in a temporal sequence. Furthermore, the first point cloud frame, the second point cloud frame, and the point cloud frames acquired continuously between the first and second point cloud frames can be identified as the first point cloud set, and this first point cloud set can be displayed in the annotation interface.
[0057] For example, if the first frame of point cloud is point cloud frame 1 and the second frame of point cloud is point cloud frame 50, point cloud frame 1, point cloud frame 2, point cloud frame 3 ... point cloud frame 50 can be used as the first point cloud set and displayed in the annotation interface.
[0058] The method for labeling dynamic targets provided by the embodiment of the present disclosure determines, based on the first labeling information of the dynamic target, a second point cloud set containing the first labeling information from multiple frame point clouds; based on the second point cloud set, determines the first frame point cloud and the second frame point cloud from the second point cloud set; wherein, the first frame point cloud is the point cloud in the second point cloud set that includes the first labeling information for the first time, and the second frame point cloud is the point cloud in the second point cloud set that includes the first labeling information for the last time; based on the first frame point cloud and the second frame point cloud, determines the first point cloud set from the multiple frame point clouds and displays it in the labeling interface; in this way, the frame point cloud that includes the first labeling information for the first time and the frame point cloud that includes the first labeling information for the last time in the pre-flash data can be used to quickly construct a tracking set for the dynamic target, thereby realizing a one-click automatic labeling function.
[0059] In some embodiments, based on the second point cloud set, the first frame point cloud and the second frame point cloud are determined from the second point cloud set, including: sorting each frame point cloud in the second point cloud set to obtain the sorting results corresponding to each frame point cloud; based on the sorting results corresponding to each frame point cloud, the first frame point cloud and the second frame point cloud are determined from the second point cloud set.
[0060] For example, since multiple point clouds with a temporal relationship are point cloud data collected continuously in a time series, the point clouds in the second point cloud set can be sorted from earliest to latest according to their acquisition time to obtain sorting results for each point cloud. Furthermore, based on the sorting results for each point cloud, the point cloud at the top of the sorting results can be used as the first point cloud, and the point cloud at the bottom of the sorting results can be used as the second point cloud.
[0061] like Figure 4 As shown in the above Figure 1A Based on the illustrated embodiment, step S103 may include the following steps S1031 to S1033.
[0062] Step S1031 : In response to the annotation operation in the annotation interface, a third frame point cloud and a fourth frame point cloud are determined from the first point cloud set.
[0063] Exemplarily, the third frame point cloud determined from the first point cloud set is not adjacent to the fourth frame point cloud, and there is a frame point cloud between the third frame point cloud and the fourth frame point cloud that does not include the first annotation information.
[0064] For example, Figure 2 As shown, the third frame point cloud can be point cloud frame 2, and the fourth frame point cloud can be point cloud frame 8. The third frame point cloud can be point cloud frame 9, and the fourth frame point cloud can be point cloud frame 12. The third frame point cloud can also be point cloud frame 2, and the fourth frame point cloud can also be point cloud frame 51. In other words, there must be a point cloud frame that does not contain the first annotation information between the third frame point cloud and the fourth frame point cloud. There can be a point cloud frame containing the first annotation information between the third frame point cloud and the fourth frame point cloud, and there may also be no point cloud frame containing the first annotation information between the third frame point cloud and the fourth frame point cloud. The embodiments of the present disclosure do not impose any restrictions on this.
[0065] Step S1032: Determine the first pose information of the dynamic target object in the third frame point cloud and the second pose information of the dynamic target object in the fourth frame point cloud based on the first annotation information.
[0066] Exemplarily, the pose information of a dynamic object includes at least the position and orientation of the dynamic object. Therefore, the first pose information is the position and orientation information of the dynamic object in the third frame point cloud, and the second pose information is the position and orientation information of the dynamic object in the fourth frame point cloud.
[0067] Step S1033: Based on the first pose information and the second pose information, perform an automatic labeling operation on at least one frame of point cloud in the first point cloud set to obtain a labeling result of the dynamic target in the first point cloud set.
[0068] Exemplarily, based on the first pose information of the dynamic target object in the third frame point cloud and the second pose information of the dynamic target object in the fourth frame point cloud, the motion trajectory of the dynamic target object between the third frame point cloud and the fourth frame point cloud can be determined, and then the first pose information and / or the second pose information can be combined to determine the pose information corresponding to each acquisition moment on the motion trajectory. In this way, an automatic labeling operation can be performed on at least one frame of point cloud in the first point cloud set to obtain the labeling result of the dynamic target object in the first point cloud set.
[0069] The method for labeling dynamic targets provided by the embodiment of the present disclosure determines a third frame point cloud and a fourth frame point cloud from a first point cloud set by responding to a labeling operation in a labeling interface; determines the first pose information of the dynamic target in the third frame point cloud and the second pose information of the dynamic target in the fourth frame point cloud based on the first labeling information; performs an automatic labeling operation on at least one frame point cloud in the first point cloud set based on the first pose information and the second pose information to obtain a labeling result of the dynamic target in the first point cloud set; in this way, the pose information of the dynamic target in the point cloud frame that does not contain the first labeling information can be determined based on the pose information of the dynamic target in the third frame point cloud and the fourth frame point cloud, thereby realizing the function of automatically labeling the point cloud frame that does not contain the first labeling information, and further realizing the labeling of dynamic targets conveniently, quickly and accurately.
[0070] like Figure 5A As shown in the above Figure 4 Based on the illustrated embodiment, step S1033 may include the following steps S11 to S14.
[0071] Step S11: Based on the third frame point cloud and the fourth frame point cloud, determine at least one frame target point cloud from the first point cloud set.
[0072] For example, if there is a frame point cloud that does not contain the first annotation information between the third frame point cloud and the fourth frame point cloud in the embodiment of the present disclosure, then the at least one frame target point cloud in step S11 may be a frame point cloud that is located between the third frame point cloud and the fourth frame point cloud and does not contain the first annotation information. Figure 2 As shown, the third frame point cloud is point cloud frame 2, and the fourth frame point cloud is point cloud frame 51. Then, the at least one frame target point cloud in step S11 can be at least one frame point cloud among the 11 frames point cloud frames 3-7, point cloud frames 10-11, point cloud frame 23, and point cloud frames 26-28.
[0073] Step S12: Determine third pose information of the dynamic target object in at least one frame of target point cloud based on the first pose information and the second pose information.
[0074] For example, the pose information of the frame point cloud that is located between the third frame point cloud and the fourth frame point cloud and does not contain the first annotation information can be determined based on the first pose information of the dynamic target object in the third frame point cloud and the second pose information of the dynamic target object in the fourth frame point cloud. Figure 2 As shown, if the third point cloud is point cloud frame 2 and the fourth point cloud is point cloud frame 51, the pose information of the dynamic target 21a in point cloud frame 2 and the pose information of the dynamic target 21a in point cloud frame 51 can be used to determine the pose information of the dynamic target 21a contained in at least one point cloud frame among point cloud frames 3-7, point cloud frames 10-11, point cloud frame 23, and point cloud frames 26-28. In other words, Figure 2 At least one frame of target point cloud may be point cloud frames 3-7, point cloud frames 10-11, point cloud frame 23, and point cloud frames 26-28.
[0075] In some examples, such as Figure 2 As shown, the pose information of the dynamic target object 21a in point cloud frame 2 and the pose information of the dynamic target object 21a in point cloud frame 8 can be determined, thereby automatically annotating point cloud frames 3-7. Similarly, the pose information of the dynamic target object 21a in point cloud frame 9 and the pose information of the dynamic target object 21a in point cloud frame 12 can be determined, thereby automatically annotating point cloud frames 10-11. In other words, the embodiment of the present disclosure does not limit the selection method of the third frame point cloud and the fourth frame point cloud. Those skilled in the art can select the appropriate third frame point cloud and fourth frame point cloud according to actual conditions to perform automatic annotation operations on the first point cloud set.
[0076] Step S13: Determine second annotation information of the dynamic target object in the at least one frame of target point cloud based on the third pose information of the dynamic target object in the at least one frame of target point cloud.
[0077] For example, in the embodiment of the present disclosure, after determining the third pose information of the dynamic target in at least one frame of target point cloud (such as the position, orientation, etc. of the dynamic target in at least one frame of target point cloud), the second annotation information of the dynamic target in at least one frame of target point cloud can be determined based on the position, orientation, etc. in the third pose information, such as the 3D detection frame of the dynamic target in at least one frame of target point cloud, the tracking ID of the dynamic target in at least one frame of target point cloud, etc. Figure 2 As shown, the second annotation information of the dynamic target object 21a in point cloud frames 3-7, point cloud frames 10-11, point cloud frames 23, and point cloud frames 26-28 can be determined respectively based on the posture information of the dynamic target object 21a in point cloud frames 3-7, point cloud frames 10-11, point cloud frames 23, and point cloud frames 26-28.
[0078] In some examples, the pose information of the dynamic target 21a in point cloud frame 2 and the pose information of the dynamic target 21a in point cloud frame 51 can also be used to determine the pose information of the dynamic target 21a contained in any of the point cloud frames 3-50, and then determine the automatic annotation results corresponding to each pose information. If both pre-brush data (i.e., first annotation information) and automatic annotation results (i.e., second annotation information) exist in the frame point cloud, the automatic annotation results can be used to replace the pre-brush data to complete the correction of the pre-brush data.
[0079] Step S14: Obtain a labeling result of the dynamic target object in the first point cloud set based on the first labeling information and the second labeling information.
[0080] Exemplarily, the first annotation information of the dynamic target in the embodiment of the present disclosure is the annotation information of the dynamic target obtained after pre-processing the data of the multi-frame point cloud with a time sequence relationship (such as using the cloud-based automatic annotation model to process the multi-frame point cloud with a time sequence relationship). The second annotation information is the annotation information obtained after the multi-frame point cloud and the pre-flush data (first annotation information) are imported into the annotation tool, and then, in response to the annotation operation in the annotation interface, an automatic annotation operation is performed on at least one frame of the point cloud in the first point cloud set based on the first annotation information. Furthermore, based on the pre-flush data and the automatic annotation data (second annotation information), the annotation results of the dynamic target in the first point cloud set containing the dynamic target can be obtained.
[0081] For example, Figure 2 As shown, first annotation information for dynamic object 21a is present in point cloud frames 2, 8-9, 12-22, 24-25, and 29-51, while second annotation information for dynamic object 21a is present in point cloud frames 3-7, 10-11, 23, and 26-28. Based on these first and second annotation information, the annotation result for dynamic object 21a in the first point cloud set 22a can be obtained.
[0082] The method for labeling dynamic targets provided by the embodiment of the present disclosure determines at least one frame of target point cloud from the first point cloud set based on the third frame point cloud and the fourth frame point cloud; determines the third pose information of the dynamic target in the at least one frame target point cloud based on the first pose information and the second pose information; determines the second labeling information of the dynamic target in the at least one frame target point cloud based on the third pose information of the dynamic target in the at least one frame target point cloud; obtains the labeling result of the dynamic target in the first point cloud set based on the first labeling information and the second labeling information; in this way, the labeling of dynamic targets in multiple frame point clouds with a time series relationship can be achieved based on pre-flash data and automatic labeling of dynamic targets that do not contain pre-flash data in the labeling tool.
[0083] In some embodiments, based on the first pose information and the second pose information, the third pose information of the dynamic target object in at least one frame of target point cloud is determined, including: based on the first pose information and the second pose information, determining the spatial distance and angle information between the dynamic target object in the third frame point cloud and the dynamic target object in the fourth frame point cloud; based on the spatial distance and angle information, determining the predicted trajectory of the dynamic target object in at least one frame of target point cloud; based on the first pose information, the second pose information, and the predicted trajectory of the dynamic target object in at least one frame of target point cloud, determining the third pose information of the dynamic target object in at least one frame of target point cloud.
[0084] For example, Figure 5BFIG. 1 is a schematic diagram showing the principle of the automatic annotation function in an embodiment of the present disclosure. Figure 5B As shown, first, the position A of the dynamic target object 51 in the third frame point cloud (i.e. Figure 5B Similarly, the position B of the dynamic target object 51 in the fourth frame point cloud (i.e., Figure 5B At this time, the spatial distance between the dynamic target object 51 in the third frame point cloud and the dynamic target object 51 in the fourth frame point cloud, that is, the line segment AB, can be determined, and then the perpendicular bisector 52 of the line segment AB can be obtained. Secondly, based on the orientation of the dynamic target object 51 in the first pose information and the orientation of the dynamic target object 51 in the second pose information, the angle information between the dynamic target object 51 in the third frame point cloud and the dynamic target object 51 in the fourth frame point cloud can be determined, and then based on the angle information, a point D can be determined on the perpendicular bisector 52 as the center of the circle corresponding to the predicted trajectory of the dynamic target object 51 in at least one frame target point cloud. Among them, the distance between the center point D and point A (or point B) is the radius of the circle, and then the predicted trajectory S of the dynamic target object 51 between the third frame point cloud and the fourth frame point cloud can be determined based on the center point D and the radius of the circle. Finally, after determining the predicted trajectory S, the third pose information of the dynamic target object 51 in each frame between point A and point B can be obtained by using the equal-division interpolation method, and the second annotation information of the dynamic target object 51 in each frame between point A and point B can be determined based on the third pose information, so that at least one frame of target point cloud between the third frame point cloud and the fourth frame point cloud can be automatically annotated.
[0085] Of course, if Figure 5B As shown, based on the orientation of the dynamic target object 51 at position A (i.e., the orientation information in the first pose information), a perpendicular line 53 corresponding to the orientation can be obtained, and the intersection point between the perpendicular midline 52 and the perpendicular line 53 can be determined as the center of the circle D. Similarly, based on the orientation of the dynamic target object 51 at position B (i.e., the orientation information in the second pose information), a perpendicular line 54 corresponding to the orientation can be obtained, and the intersection point between the perpendicular midline 52 and the perpendicular line 54 can be determined as the center of the circle D.
[0086] It should be noted that Figure 5B The automatic annotation principle in the text is only an example, and those skilled in the art can use Figure 5B The dynamic target objects can be automatically labeled according to the automatic labeling principle in the embodiment, or the dynamic target objects can be automatically labeled based on the first pose information of the dynamic target objects in the third frame point cloud and the second pose information of the dynamic target objects in the fourth frame point cloud according to other methods. The embodiments of the present disclosure do not limit this.
[0087] Figure 5CThis is a schematic diagram of a scenario of the automatic annotation function in an embodiment of the present disclosure, such as Figure 5C As shown, schematic diagram 55 is a schematic diagram of the automatic labeling function of the dynamic target 51 in a straight-ahead scenario. In this scenario, automatic labeling can be performed using linear interpolation, and the expected labeling result is shown in the dotted box in schematic diagram 55. Schematic diagram 56 is a schematic diagram of the automatic labeling function of the dynamic target 51 in a turning scenario. In this scenario, the difference in orientation of the dynamic target 51 in the third frame point cloud and the fourth frame point cloud is less than 90 degrees. Automatic labeling can be performed using arc interpolation of the starting direction. The expected labeling result is shown in the dotted box in schematic diagram 56. Schematic diagram 57 is a schematic diagram of the automatic labeling function of the dynamic target 51 in another turning scenario. In this scenario, the difference in orientation of the dynamic target 51 in the third frame point cloud and the fourth frame point cloud is equal to 90 degrees. Automatic labeling can be performed using arc interpolation of the starting direction. The expected labeling result is shown in the dotted box in schematic diagram 57. Diagram 58 illustrates the automatic labeling of a dynamic object 51 in another turning scenario. In this scenario, the difference in orientation between the third and fourth frame point clouds is greater than 90 degrees and less than 180 degrees. Automatic labeling can be performed using arc interpolation of the starting direction, with the expected labeling result shown in the dashed box in Diagram 58. Diagram 59 illustrates the automatic labeling of a dynamic object 51 in an unusual scenario: the orientations of the dynamic object 51 in the third and fourth frame point clouds are completely opposite or opposite. After calculating the included angle, automatic labeling can be performed using small arc interpolation and angle averaging. The expected labeling result is shown in the dashed box in Diagram 59. While the position of this labeling result is correct, the rotation angle may be incorrect, requiring confirmation of the same direction before labeling. Diagram 60 illustrates the automatic labeling of a dynamic object 51 in a receding scenario. In this scenario, labeling is normal and the labeling result is as expected, with the expected labeling result shown in the dashed box in Diagram 60.
[0088] It should be noted that if Figure 5C As shown, the dynamic target 51 marked as “start” in schematic diagrams 55 to 60 is the dynamic target in the third frame point cloud, and the dynamic target 51 marked as “end” is the dynamic target in the fourth frame point cloud.
[0089] like Figure 6 As shown in the above Figure 1A On the basis of the embodiment shown, the following steps S104 to S106 may also be included.
[0090] Step S104: Acquire at least one image sequence corresponding to the multi-frame point cloud.
[0091] For example, a multi-frame point cloud with a temporal relationship can be obtained by processing image sequences from multiple different perspectives using a multi-view reconstruction method. In this case, the at least one image sequence corresponding to the multi-frame point cloud in step S104 can be an image sequence captured by a camera on a vehicle. A multi-frame point cloud with a temporal relationship can also be obtained by collecting point cloud data using a three-dimensional scanning device. In this case, the multi-frame point cloud with a temporal relationship obtained by scanning can be processed using projection mapping to obtain the at least one image sequence corresponding to the multi-frame point cloud in step S104. Of course, the specific implementation method for obtaining the at least one image sequence corresponding to the multi-frame point cloud is not limited in the embodiments of the present disclosure, and those skilled in the art can make settings based on actual usage scenarios.
[0092] Step S105 : Projecting the annotation result of the dynamic target object to at least one image sequence to obtain the annotation result of the dynamic target object in the at least one image sequence.
[0093] For example, if at least one image sequence contains an image frame containing a dynamic object, the annotation results (e.g., 3D annotation boxes) of the dynamic object in the first point cloud set can be projected onto the image sequence to obtain the annotation results of the dynamic object in the image sequence, so that the user can check the accuracy of the annotation results. Furthermore, if the annotation results do not meet the required accuracy, the annotation results of the dynamic object in the first point cloud set can be projected onto the three-view image for adjustment.
[0094] Of course, in addition to projecting the annotation results onto the image sequence for verification, they can also be quickly played back within the annotation tool for quality inspection (because each frame of the point cloud in the first point cloud set is a continuous frame point cloud with a time sequence relationship). This allows for high-quality annotation ground truth values for dynamic objects. For example, the annotation results can be played back for each dynamic object, which can more comprehensively ensure the quality of the annotation data.
[0095] Step S106: verify the annotation box and identification information in the annotation result to determine the annotation accuracy of the annotation result.
[0096] For example, the electronic device may perform a quality check on the annotation results (e.g., 2D annotation boxes and tracking IDs) of the dynamic objects in at least one image sequence. For example, the segmentation results may be matched with the 2D annotation boxes to verify the annotation results of the dynamic objects in the image sequence.
[0097] Of course, the quality of the annotation box and the tracking ID in the annotation result can also be manually checked, and the embodiment of the present disclosure does not limit this.
[0098] The method for labeling dynamic targets provided by the embodiments of the present disclosure obtains at least one image sequence corresponding to a multi-frame point cloud; projects the labeling results of the dynamic targets to the at least one image sequence to obtain the labeling results of the dynamic targets in the at least one image sequence; verifies the labeling box in the labeling results and the identification information in the labeling results to determine the labeling accuracy of the labeling results; in this way, the labeling results of the dynamic targets in the multi-frame point cloud can be projected in real time to the image space to check the correctness of the labeling results, thereby obtaining high-quality labeling truth values.
[0099] Exemplary devices
[0100] Figure 7 A dynamic target marking device provided by the embodiment of the present disclosure is as follows Figure 7 As shown, the labeling device 700 includes a data acquisition module 701 , a determination and display module 702 , and an automatic labeling module 703 .
[0101] The data acquisition module 701 is configured to acquire multiple frames of point clouds having a temporal relationship, and first annotation information of at least one frame of the multiple frames of point clouds for a dynamic target object;
[0102] A determination and display module 702 is configured to determine, based on the dynamic target, a first point cloud set containing the dynamic target from the multi-frame point cloud and display the first point cloud set in the annotation interface;
[0103] The automatic annotation module 703 is used to respond to the annotation operation in the annotation interface and perform an automatic annotation operation on at least one frame of point cloud in the first point cloud set based on the first annotation information to obtain the annotation results of the dynamic target objects in the first point cloud set.
[0104] In some embodiments, as Figure 8 As shown, the determination and display module 702 includes a set determination unit 7021 , a first point cloud determination unit 7022 and a determination and display unit 7023 .
[0105] The set determining unit 7021 is configured to determine, based on the first annotation information of the dynamic target object, a second point cloud set including the first annotation information from the multi-frame point cloud;
[0106] The first point cloud determining unit 7022 is configured to determine, based on the second point cloud set, a first frame point cloud and a second frame point cloud from the second point cloud set; wherein the first frame point cloud is the frame point cloud in the second point cloud set that first includes the first annotation information, and the second frame point cloud is the frame point cloud in the second point cloud set that last includes the first annotation information;
[0107] The determination display unit 7023 is used to determine a first point cloud set from multiple frames of point clouds based on the first frame point cloud and the second frame point cloud, and display the first point cloud set in the annotation interface.
[0108] In some embodiments, the point cloud determination unit 7022 is specifically used to sort each frame point cloud in the second point cloud set to obtain a sorting result corresponding to each frame point cloud; based on the sorting result corresponding to each frame point cloud, a first frame point cloud and a second frame point cloud are determined from the second point cloud set; wherein the first frame point cloud is the frame point cloud in the second point cloud set that includes the first annotation information for the first time, and the second frame point cloud is the frame point cloud in the second point cloud set that includes the first annotation information for the last time.
[0109] In some embodiments, as Figure 9 As shown, the automatic labeling module 703 includes a second point cloud determination unit 7031 , a vehicle posture determination unit 7032 and an automatic labeling unit 7033 .
[0110] The second point cloud determining unit 7031 is configured to determine a third frame point cloud and a fourth frame point cloud from the first point cloud set in response to a marking operation in the marking interface;
[0111] The vehicle pose determination unit 7032 is configured to determine, based on the first annotation information, the first pose information of the dynamic target object in the third frame point cloud and the second pose information of the dynamic target object in the fourth frame point cloud;
[0112] The automatic labeling unit 7033 is used to perform an automatic labeling operation on at least one frame of point cloud in the first point cloud set based on the first pose information and the second pose information to obtain a labeling result of the dynamic target object in the first point cloud set.
[0113] In some embodiments, the automatic labeling unit 7033 is specifically used to determine at least one frame of target point cloud from the first point cloud set based on the third frame point cloud and the fourth frame point cloud; determine the third pose information of the dynamic target object in the at least one frame target point cloud based on the first pose information and the second pose information; determine the second labeling information of the dynamic target object in the at least one frame target point cloud based on the third pose information of the dynamic target object in the at least one frame target point cloud; and obtain the labeling result of the dynamic target object in the first point cloud set based on the first labeling information and the second labeling information.
[0114] In some embodiments, the automatic labeling unit 7033 is specifically used to determine at least one frame of target point cloud from the first point cloud set based on the third frame point cloud and the fourth frame point cloud; determine the spatial distance and angle information between the dynamic target in the third frame point cloud and the dynamic target in the fourth frame point cloud based on the first pose information and the second pose information; determine the predicted trajectory of the dynamic target in the at least one frame target point cloud based on the spatial distance and the angle information; determine the third pose information of the dynamic target in the at least one frame target point cloud based on the first pose information, the second pose information, and the predicted trajectory of the dynamic target in the at least one frame target point cloud; determine the second labeling information of the dynamic target in the at least one frame target point cloud based on the third pose information of the dynamic target in the at least one frame target point cloud; and obtain the labeling result of the dynamic target in the first point cloud set based on the first labeling information and the second labeling information.
[0115] In some embodiments, the labeling device 700 further includes:
[0116] An image acquisition module, configured to acquire at least one image sequence corresponding to a multi-frame point cloud;
[0117] A result projection module, configured to project the annotation results of the dynamic target object onto at least one image sequence to obtain the annotation results of the dynamic target object in the at least one image sequence;
[0118] The quality verification module is used to verify the annotation box and identification information in the annotation result to determine the annotation accuracy of the annotation result.
[0119] The beneficial technical effects corresponding to the exemplary embodiment of the dynamic target labeling device 700 can be found in the corresponding beneficial technical effects of the exemplary method section above, and will not be repeated here.
[0120] Exemplary electronic devices
[0121] Figure 10 This is a structural diagram of an electronic device 100 provided in an embodiment of the present disclosure, including at least one processor 101 and a memory 102.
[0122] The processor 101 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 100 to perform desired functions.
[0123] The memory 102 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 101 may execute the one or more computer program instructions to implement the above-mentioned methods for labeling dynamic targets and / or other desired functions of the various embodiments of the present disclosure.
[0124] In one example, the electronic device 100 may further include an input device 103 and an output device 104 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0125] The input device 103 may also include, for example, a keyboard, a mouse, and the like.
[0126] The output device 104 can output various information to the outside, and may include, for example, a display, a speaker, a printer, a communication network and its connected remote output devices, etc.
[0127] Of course, to simplify, Figure 10 Only some of the components related to the present disclosure in the electronic device 100 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 100 may further include any other appropriate components according to specific application scenarios.
[0128] Exemplary computer program products and computer-readable storage media
[0129] In addition to the above-mentioned methods and devices, embodiments of the present disclosure may also provide a computer program product, including computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the dynamic target labeling method of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.
[0130] The computer program product may be written in any combination of one or more programming languages to implement the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0131] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the steps in the dynamic target labeling method of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.
[0132] Computer readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium is, for example, but not limited to, a system, device or component comprising electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0133] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be considered as essential to each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0134] Those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.
Claims
1. A method for labeling a dynamic target, comprising: Acquire multiple frames of point clouds having a temporal relationship, and first annotation information of at least one frame of the multiple frames of point clouds for a dynamic target object; Based on the dynamic target, determining a first point cloud set including the dynamic target from the multi-frame point cloud and displaying the first point cloud set in the annotation interface; In response to the annotation operation in the annotation interface, an automatic annotation operation is performed on at least one frame of point cloud in the first point cloud set based on the first annotation information to obtain an annotation result of the dynamic target object in the first point cloud set.
2. The method according to claim 1, wherein The determining, based on the dynamic target, a first point cloud set including the dynamic target from the multiple frame point clouds and displaying the first point cloud set in the annotation interface includes: Based on the first annotation information of the dynamic target object, determining a second point cloud set including the first annotation information from the multi-frame point cloud; Based on the second point cloud set, determining a first frame point cloud and a second frame point cloud from the second point cloud set; wherein the first frame point cloud is a frame point cloud in the second point cloud set that includes the first annotation information for the first time, and the second frame point cloud is a frame point cloud in the second point cloud set that includes the first annotation information for the last time; Based on the first frame point cloud and the second frame point cloud, a first point cloud set is determined from the multiple frames of point cloud and displayed in the annotation interface.
3. The method according to claim 2, wherein: The step of determining a first frame point cloud and a second frame point cloud from the second point cloud set based on the second point cloud set includes: Sorting each frame of point cloud in the second point cloud set to obtain a sorting result corresponding to each frame of point cloud; Based on the sorting results corresponding to the frame point clouds, the first frame point cloud and the second frame point cloud are determined from the second point cloud set.
4. The method according to claim 1, wherein In response to the annotation operation in the annotation interface, performing an automatic annotation operation on at least one frame of point cloud in the first point cloud set based on the first annotation information to obtain an annotation result of the dynamic target in the first point cloud set includes: In response to the annotation operation in the annotation interface, determining a third frame point cloud and a fourth frame point cloud from the first point cloud set; Determine, based on the first annotation information, first pose information of the dynamic target object in the third frame point cloud and second pose information of the dynamic target object in the fourth frame point cloud; Based on the first pose information and the second pose information, an automatic labeling operation is performed on at least one frame of point cloud in the first point cloud set to obtain a labeling result of the dynamic target in the first point cloud set.
5. The method according to claim 4, wherein The step of performing an automatic labeling operation on at least one frame of point cloud in the first point cloud set based on the first pose information and the second pose information to obtain a labeling result of the dynamic target in the first point cloud set includes: Based on the third frame point cloud and the fourth frame point cloud, determining at least one frame target point cloud from the first point cloud set; Determining third pose information of the dynamic target object in the at least one frame of target point cloud based on the first pose information and the second pose information; Determining second labeling information of the dynamic target object in the at least one frame of target point cloud based on third pose information of the dynamic target object in the at least one frame of target point cloud; Based on the first annotation information and the second annotation information, a annotation result of the dynamic target object in the first point cloud set is obtained.
6. The method according to claim 5, wherein: The determining, based on the first pose information and the second pose information, third pose information of the dynamic target object in the at least one frame of target point cloud includes: Determining, based on the first pose information and the second pose information, a spatial distance and angle information between the dynamic target object in the third frame point cloud and the dynamic target object in the fourth frame point cloud; Determining a predicted trajectory of the dynamic target object in the at least one frame of target point cloud based on the spatial distance and the angle information; Based on the first pose information, the second pose information, and the predicted trajectory of the dynamic target object in the at least one frame of target point cloud, third pose information of the dynamic target object in the at least one frame of target point cloud is determined.
7. The method according to any one of claims 1 to 6, further comprising: Acquire at least one image sequence corresponding to the multi-frame point cloud; Projecting the annotation result of the dynamic target object to the at least one image sequence to obtain the annotation result of the dynamic target object in the at least one image sequence; The annotation box in the annotation result and the identification information in the annotation result are verified to determine the annotation accuracy of the annotation result.
8. A dynamic target marking device, comprising: A data acquisition module is configured to acquire multiple frames of point clouds having a temporal relationship, and first annotation information of at least one frame of the multiple frames of point clouds for a dynamic target object; a determination and display module, configured to determine, based on the dynamic target, a first point cloud set including the dynamic target from the multi-frame point cloud and display the first point cloud set in a labeling interface; An automatic labeling module is used to respond to the labeling operation in the labeling interface and perform an automatic labeling operation on at least one frame of point cloud in the first point cloud set based on the first labeling information to obtain the labeling result of the dynamic target object in the first point cloud set.
9. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, is used to implement the dynamic target labeling method according to any one of claims 1 to 7.
10. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the dynamic target marking method described in any one of claims 1 to 7.