Target marker labeling method and device, equipment and storage medium
By constructing three-dimensional point cloud data and user interface interaction, semi-automated annotation of traffic markers is realized, solving the problems of high cost and low accuracy in existing methods, and improving the labeling efficiency and accuracy.
Patent Information
- Application Number
- CN202510639233.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-12
AI Technical Summary
The existing traffic marker annotation methods are mainly carried out on the basis of images, and lack of timing information, resulting in high labeling costs and long-tail fusion problems and factor misalignment.
By constructing three-dimensional point cloud data of the target scene, determining the predicted location of the target marker, and combining interactions on the user interface, a semi-automated human-computer interaction method is realized for annotation, including checksum user adjustment of the predicted location, and the final annotation result is generated.
While ensuring the quality of labeling, it significantly improves labeling efficiency, reduces labeling costs and improves labeling accuracy.
Smart Images

Figure CN120472049A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data annotation, and in particular to a method, apparatus, device, and storage medium for labeling target markers. Background Art
[0002] Model training is essential for achieving intelligent driving, and the model training process typically requires a large amount of labeled data. Therefore, data labeling is a very important step. As we all know, the more accurate the data labeling, the better the algorithm model training effect.
[0003] In intelligent driving perception, the quality and accuracy of traffic sign perception (such as traffic lights and signs) are crucial for the safety of advanced assisted driving, requiring a large number of accurately labeled samples for model training. However, most current traffic sign annotation methods in the industry are based on images, lacking temporal information. This not only results in high annotation costs, but also leads to long-tail fusion problems and feature misalignment when converting multiple image domain annotation results to the vector space domain. Summary of the Invention
[0004] At present, the annotation methods for traffic signs in intelligent driving scenarios are mostly based on images. Since a large number of annotation samples are required, the annotation cost is very high.
[0005] In order to solve the above technical problems, the present disclosure provides a target marker labeling method and device, equipment, and storage medium, which can solve the problem of high labeling cost of existing traffic sign labeling methods.
[0006] A first aspect of the present disclosure provides a method for labeling a target marker, comprising: determining time series data corresponding to a target scene; constructing three-dimensional point cloud data of the target scene based on the time series data, and determining a predicted position of the target marker in the three-dimensional point cloud data; obtaining a first labeling result of the target marker in the three-dimensional point cloud data; determining target information corresponding to the target marker based on the predicted position and the first labeling result, and displaying the target information on a user interface; and determining a second labeling result of the target marker in the three-dimensional point cloud data in response to detecting a labeling operation on the target information on the user interface.
[0007] According to a second aspect of the present disclosure, a device for marking a target marker is provided, comprising: a time series data determination module for determining time series data corresponding to a target scene; a predicted position determination module for constructing three-dimensional point cloud data of the target scene based on the time series data, and determining a predicted position of the target marker in the three-dimensional point cloud data; a pre-flash result determination module for obtaining a first marking result of the target marker in the three-dimensional point cloud data; a marking display module for determining target information corresponding to the target marker based on the predicted position and the first marking result, and displaying the target information on a user interface; and a marking result determination module for determining a second marking result of the target marker in the three-dimensional point cloud data in response to detecting a marking operation on the target information on the user interface.
[0008] A third aspect of the present disclosure provides a computer-readable storage medium storing a computer program for executing the target marker labeling method provided in the first aspect.
[0009] The fourth aspect of the present disclosure provides an electronic device comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the target marker labeling method provided in the first aspect above.
[0010] A fifth aspect of the present disclosure provides a computer program product. When the instructions in the computer program product are executed by a processor, the target marker labeling method provided in the first aspect is executed.
[0011] Based on the target marker labeling method provided by the present invention, the time series data corresponding to the target scene is determined; based on the time series data, three-dimensional point cloud data of the target scene is constructed, and the predicted position of the target marker in the three-dimensional point cloud data is determined; a first labeling result of the target marker in the three-dimensional point cloud data is obtained; based on the predicted position and the first labeling result, the target information corresponding to the target marker is determined, and the target information is displayed on a user interface; in response to detecting the labeling operation of the target information on the user interface, a second labeling result of the target marker in the three-dimensional point cloud data is determined; in this way, the three-dimensional point cloud data constructed by the time series data and the pre-brush labeling result of the target marker (the target information corresponding to the target marker) can be displayed on the user interface so that manual labeling adjustment operations can be performed, thereby completing the rapid labeling of the target marker based on a semi-automatic human-computer interaction method, while ensuring the quality of the target marker labeling and improving the labeling efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1A schematic diagram of an application scenario provided for an exemplary embodiment of the present disclosure.
[0013] Figure 2A A flowchart of a marking method provided by an exemplary embodiment of the present disclosure.
[0014] Figure 2B A schematic diagram of a user interface provided by an exemplary embodiment of the present disclosure.
[0015] Figure 3 A flowchart of a marking method provided by another exemplary embodiment of the present disclosure.
[0016] Figure 4A A flowchart of a marking method provided in accordance with another exemplary embodiment of the present disclosure is provided.
[0017] Figure 4B A schematic diagram of three views is provided for yet another exemplary embodiment of the present disclosure.
[0018] Figure 4C A schematic diagram of a user interface provided by another exemplary embodiment of the present disclosure.
[0019] Figure 5A A flowchart of a marking method provided in accordance with another exemplary embodiment of the present disclosure is provided.
[0020] Figure 5B A schematic diagram of a user interface provided by yet another exemplary embodiment of the present disclosure.
[0021] Figure 5C A schematic diagram of an image projection result provided by an exemplary embodiment of the present disclosure.
[0022] Figure 6 A flowchart of a marking method provided in accordance with another exemplary embodiment of the present disclosure is provided.
[0023] Figure 7 A flowchart of a marking method provided in accordance with another exemplary embodiment of the present disclosure is provided.
[0024] Figure 8 A schematic structural diagram of a labeling device provided by an exemplary embodiment of the present disclosure.
[0025] Figure 9 A schematic structural diagram of a labeling device provided by another exemplary embodiment of the present disclosure.
[0026] Figure 10 The present invention provides a structural diagram of an electronic device according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0027] To explain the present disclosure, example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. It should be understood that the present disclosure is not limited to the example embodiments.
[0028] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.
[0029] Application Overview
[0030] The target marker labeling method provided by the embodiment of the present disclosure can be applied to the scenario of true value labeling during model training, and any other feasible scenario. For example, it can be applied to the training samples required for perception model training: when training a perception model for detecting target markers (such as traffic signs, traffic lights, etc.), the target markers can be labeled by first labeling at least one image sequence to be labeled that is collected by an image sensor (also called: visual sensor) in different orientations through the embodiment of the present disclosure, and the labeling result can be used as a training sample to train the perception model until the preset training completion conditions are met, thereby obtaining a trained perception model that can be used for downstream perception tasks.
[0031] Figure 1 This is a schematic diagram of an application scenario provided by an exemplary embodiment of the present disclosure. Figure 1As shown, a vehicle 10 is equipped with multiple image sensors 11 (e.g., cameras) and a 3D scanning device 12 (e.g., a lidar) in different orientations. The vehicle 10 passes through a target scene 13 at a preset speed. While the vehicle 10 is traveling, each image sensor 11 on the vehicle 10 captures images of the target scene 13 at a first preset frequency, generating a sequence of images to be labeled. Therefore, the multiple image sensors 11 generate multiple image sequences to be labeled. Each image sequence to be labeled includes multiple frames of images to be labeled, and the multiple frames in each image sequence to be labeled are sorted according to the image acquisition time or the driving trajectory of the vehicle 10. While the vehicle 10 is traveling, the 3D scanning device 12 can also capture point cloud data of the target scene 13 at a second preset frequency, generating a point cloud dataset. The first and second preset frequencies can be the same or different. Based on the multiple image sequences to be labeled and / or the point cloud dataset, the target scene 13 is reconstructed in three dimensions to generate three-dimensional point cloud data of the target scene 13. Then, the target markers are annotated on the 3D point cloud data using the target marker annotation method provided in the embodiments of the present disclosure, obtaining annotation results for the target markers (e.g., traffic signs, traffic lights, buildings, etc.) in the 3D point cloud. Furthermore, the target marker annotation results in the 3D point cloud data can be projected onto the multiple image sequences to be annotated, obtaining multiple annotated image sequences with the target marker annotation results.
[0032] It should be noted that Figure 1 This is only an example, and the embodiments of the present disclosure Figure 1 There is no limitation on the number and location of the multiple image sensors 11 and the 3D scanning device 12. In actual use, the image sensor 11 can be the image sensor corresponding to the vehicle's forward wide-angle camera, the image sensor 11 can be the image sensor corresponding to the vehicle's forward narrow-angle camera, the image sensor 11 can also be the image sensor corresponding to the side view camera, etc.
[0033] That is, an embodiment of the present disclosure provides a method for labeling a target marker, by determining time series data corresponding to a target scene; constructing three-dimensional point cloud data of the target scene based on the time series data, and determining a predicted position of the target marker in the three-dimensional point cloud data; obtaining a first labeling result of the target marker in the three-dimensional point cloud data; determining target information corresponding to the target marker based on the predicted position and the first labeling result, and displaying the target information on a user interface; in response to detecting a labeling operation on the target information on the user interface, determining a second labeling result of the target marker in the three-dimensional point cloud data; in this way, the three-dimensional point cloud data constructed by the time series data and the pre-brush labeling result of the target marker (target information corresponding to the target marker) can be displayed on the user interface so that manual labeling adjustment operations can be performed, thereby completing rapid labeling of the target marker based on a semi-automatic human-computer interaction method, while ensuring the quality of target marker labeling and improving labeling efficiency.
[0034] Exemplary Methods
[0035] Figure 2A This is a flow chart of a method for marking target markers provided by an embodiment of the present disclosure. This embodiment can be applied to any electronic device such as a local terminal device, a cloud server, etc., and can also be applied to terminal devices and cloud servers in a distributed manner, such as Figure 2A As shown, the method includes the following steps S201-S205.
[0036] Step S201: Determine the time series data corresponding to the target scene.
[0037] For example, the target scene may include a scene of a certain section of a highway or a scene of a certain intersection on a city road. The present disclosure does not limit the actual situation of the target scene. Factors affecting the target scene include, but are not limited to, driving location, road type, weather conditions, traffic conditions, etc.
[0038] In some examples, such as Figure 1 As shown, the vehicle 10 passes through the target scene 13 at a preset speed, and each image sensor 11 on the vehicle 10 respectively captures images of the target scene 13 at a first preset frequency, obtaining a plurality of image sequences to be labeled, which are the time series data in step S201; wherein, the frequency of image acquisition by each image sensor 11 can be the same or different, and the embodiment of the present disclosure does not limit this. Similarly, if Figure 1As shown, vehicle 10 passes through target scene 13 at a preset speed. A 3D scanning device 12 on vehicle 10 scans target scene 13 at a second preset frequency, obtaining a plurality of point cloud data sets having a time-series relationship. This plurality of point cloud data sets is the time-series data in step S201. In other words, the time-series data corresponding to the target scene in the disclosed embodiments includes, but is not limited to, an image sequence acquired at the first preset frequency and / or a point cloud data set acquired at the second preset frequency in the target scene.
[0039] Step S202: construct three-dimensional point cloud data of the target scene based on the time series data, and determine the predicted position of the target marker in the three-dimensional point cloud data.
[0040] Exemplarily, the time series data corresponding to the target scene can be a plurality of image sequences to be annotated, and / or a plurality of point cloud data with a time series relationship. On this basis, the target scene can be reconstructed in three dimensions using a plurality of point cloud data with a time series relationship to obtain three-dimensional point cloud data of the target scene. Among them, the three-dimensional scanning equipment includes but is not limited to Lidar (Light Laser Detection and Ranging), structured light sensor, TOF (Time of Flight) camera, etc. Alternatively, the target scene can be reconstructed in three dimensions using at least one image sequence to be annotated to obtain three-dimensional point cloud data of the target scene. Of course, it is also possible to simultaneously use a plurality of point cloud data with a time series relationship and at least one image sequence to be annotated for three-dimensional reconstruction to improve the accuracy of the three-dimensional reconstruction result. The specific implementation method of constructing the three-dimensional point cloud data of the target scene is not limited in the embodiment of the present disclosure.
[0041] In some examples, there will be one or more target markers (such as traffic lights, traffic signs, etc.) in the target scene of the embodiments of the present disclosure. For example, multiple traffic lights at a certain intersection in a city road; another example, signs on a certain section of a highway, etc. Furthermore, the predicted position of the target marker in the three-dimensional point cloud data can be determined based on the three-dimensional point cloud data of the target scene. The predicted position is used to indicate that there is a target marker at that location in the three-dimensional point cloud data. The predicted position can be expressed in the form of a prediction box, so the predicted position can include the three-dimensional coordinate information of each point in the prediction box in the three-dimensional point cloud data. It should be noted that this is only an example, and the specific expression of the predicted position of the target marker in the three-dimensional point cloud data in the embodiments of the present disclosure is not limited.
[0042] Step S203: Obtain a first annotation result of the target marker in the three-dimensional point cloud data.
[0043] For example, the first annotation result of the target marker in the 3D point cloud data can be the first 3D (three-dimensional) annotation box of the target marker in the 3D point cloud data. For example, a large cloud-based automatic annotation model can be pre-trained, and the input of the cloud-based automatic annotation model is the 3D point cloud data, and the output is the first 3D annotation box of the target marker in the 3D point cloud data. In this way, the pre-brush result (first 3D annotation box) of the target marker can be automatically obtained, thereby eliminating the need for manual annotation and saving annotation costs and time.
[0044] It should be noted that the embodiments of the present disclosure do not impose any restrictions on the model structure of the cloud-based automatic annotation large model, the training samples of the cloud-based automatic annotation large model, the training iteration termination conditions of the cloud-based automatic annotation large model, and the detection accuracy of the cloud-based automatic annotation large model. In actual use, those skilled in the art can set them according to actual conditions. In the embodiments of the present disclosure, the cloud-based automatic annotation large model is used to process the three-dimensional point cloud data to obtain the corresponding pre-flush results, which refer to the initial annotation results of the target markers in the three-dimensional point cloud data. Of course, other methods can also be used to obtain the pre-flush results of the target markers in the three-dimensional point cloud data, and the embodiments of the present disclosure do not impose any restrictions on this.
[0045] Step S204: Based on the predicted position and the first annotation result, determine the target information corresponding to the target marker, and display the target information on the user interface.
[0046] Exemplarily, the first annotation result can be obtained by a large model for automatic annotation in the cloud. Those skilled in the art know that under normal circumstances, the model will not reach 100% accuracy due to various reasons (such as insufficient or too complex training data, certain defects in the structure of the model, etc.). Therefore, the output of the large model for automatic annotation in the cloud may also have certain errors. For example, there is a target marker at position A in the three-dimensional point cloud data, but the first 3D annotation box output by the large model for automatic annotation in the cloud is located at position B, and position A and position B are different positions in the three-dimensional point cloud data. For another example, there are multiple target markers in the three-dimensional point cloud data, but the large model for automatic annotation in the cloud only outputs the first 3D annotation box corresponding to one target marker, and the other target markers are not detected. Therefore, in the embodiment of the present disclosure, the predicted position is used to verify the first annotation result to obtain a corresponding verification result; thereby, the target information corresponding to the target marker is determined based on the verification result, and the target information is displayed on the user interface. Among them, the target information is related to the predicted position, or the target information is related to the first annotation result.
[0047] In some examples, the three-dimensional point cloud data corresponding to the target scene and the determined target information can be visualized so that the three-dimensional point cloud data of the target scene and the determined target information can be displayed on a visual interactive interface. For example, the three-dimensional point cloud data of the target scene and the determined target information can be visualized through the visualization module in PCL (Point Cloud Library) and the Cloud Compare (point cloud registration) visualization software. In addition, the above-mentioned visual interactive interface (i.e., the user interface of step S204) is used to display the three-dimensional point cloud data and the determined target information of the target scene, and can also be used to edit and reorganize the target information displayed in the three-dimensional point cloud data according to the user's operation instructions (such as marking instructions, adjustment instructions, etc.). Among them, the visual interactive interface can display data through a display (for example, a liquid crystal display, a plasma display), and the visual interactive interface can also receive user operation instructions through a keyboard, a mouse, a touch screen, etc.
[0048] In some examples, Figure 2B FIG. 1 is an exemplary schematic diagram of target information in three-dimensional point cloud data according to an embodiment of the present disclosure. Figure 2B As shown, the target marker 21 is a traffic light, and the target information 22 in the three-dimensional point cloud data can be displayed on the visual interactive interface 23. Figure 2B The target information 22 in the image is the first 3D annotation box of the target marker 21. For example, the first 3D annotation box may be a annotation box for surrounding a traffic light. Figure 2B There are multiple target markers and target information corresponding to each target marker, but Figure 2B Only one target marker and its corresponding target information are shown for exemplary explanation.
[0049] Step S205: In response to detecting a labeling operation on the target information on the user interface, determine a second labeling result of the target marker in the three-dimensional point cloud data.
[0050] For example, after the 3D point cloud data of the target scene and the determined target information are displayed on the visual interactive interface, the user can use the visual interactive interface to annotate the target information in the 3D point cloud data, thereby determining a second annotation result for the target landmark in the 3D point cloud data. The second annotation result is obtained by the user annotating the target information, and therefore, in general, the second annotation result is more accurate than the first annotation result.
[0051] The target marker labeling method provided by the embodiment of the present disclosure determines the time series data corresponding to the target scene; constructs the three-dimensional point cloud data of the target scene based on the time series data, and determines the predicted position of the target marker in the three-dimensional point cloud data; obtains a first labeling result of the target marker in the three-dimensional point cloud data; determines the target information corresponding to the target marker based on the predicted position and the first labeling result, and displays the target information on a user interface; in response to detecting the labeling operation of the target information on the user interface, determines the second labeling result of the target marker in the three-dimensional point cloud data; in this way, the three-dimensional point cloud data constructed by the time series data and the pre-brush labeling result of the target marker (the target information corresponding to the target marker) can be displayed on the user interface so that the labeling adjustment operation can be manually performed, thereby completing the rapid labeling of the target marker based on a semi-automatic human-computer interaction method, while ensuring the quality of the target marker labeling and improving the labeling efficiency.
[0052] like Figure 3 As shown in the above Figure 2A Based on the illustrated embodiment, step S204 may include the following steps S2041 and S2042.
[0053] Step S2041: Match the predicted box of the predicted position with the marked box of the first marking result to obtain a matching result.
[0054] Exemplarily, the prediction box of the predicted position can be matched (compared) with the annotation box of the first annotation result (such as the first 3D annotation box) to obtain a matching result; wherein the matching result includes that the position of the annotation box in the three-dimensional point cloud data is the same as the position of the prediction box in the three-dimensional point cloud data, or the position of the annotation box in the three-dimensional point cloud data is different from the position of the prediction box in the three-dimensional point cloud data. For example, if the matching result is that there is a annotation box at the position of the prediction box, then the matching result indicates that the first annotation result is available. For another example, if the matching result is that there is no annotation box at the position of the prediction box, then the matching result indicates that the first annotation result is unavailable.
[0055] Step S2042: Based on the matching result, determine the target information corresponding to the target marker, and display the target information on the user interface.
[0056] For example, in embodiments of the present disclosure, target information corresponding to a target marker can be determined based on different situations included in the matching results. For example, if the matching result shows that the position of the annotation box in the 3D point cloud data is the same as the position of the prediction box in the 3D point cloud data, then the target information is associated with the first annotation result. For another example, if the position of the annotation box in the 3D point cloud data is different from the position of the prediction box in the 3D point cloud data, then the target information is associated with the predicted position.
[0057] The target marker labeling method provided by the embodiment of the present disclosure obtains a matching result by matching the prediction box of the predicted position with the labeling box of the first labeling result; based on the matching result, the target information corresponding to the target marker is determined; in this way, the automatically labeled first labeling result can be verified, and the manually labeled labeling content can be determined based on the verification result, thereby improving the accuracy of the labeling.
[0058] In some embodiments, based on the matching result, the target information corresponding to the target marker is determined, including: if the matching result shows that the position of the prediction box is the same as that of the annotation box, the annotation box is determined as the target information corresponding to the target marker; or, if the matching result shows that the position of the prediction box is different from that of the annotation box, the prediction box is determined as the target information corresponding to the target marker.
[0059] For example, if the matching result indicates that the position of the annotation box in the 3D point cloud data is the same as the position of the prediction box in the 3D point cloud data, the target information is the annotation box of the first annotation result. For another example, if the position of the annotation box in the 3D point cloud data is different from the position of the prediction box in the 3D point cloud data, the target information is the prediction box at the predicted position. Furthermore, if the target information is the annotation box of the first annotation result, the 3D point cloud data and the annotation box of the target scene are displayed in the user interface; if the target information is the prediction box at the predicted position, the 3D point cloud data of the target scene and the prediction box annotation box are displayed in the user interface. For example, if the matching result indicates that the first 3D annotation box exists at the location of the prediction box, the 3D point cloud data and the first 3D annotation box are displayed in the visual interactive interface. For another example, if the matching result indicates that the first 3D annotation box does not exist at the location of the prediction box, the 3D point cloud data and the prediction box are displayed in the visual interactive interface. In this way, the corresponding target information can be displayed based on the automatic pre-flush result or the prediction result for manual annotation, conveniently, quickly, and accurately achieving the annotation of target landmarks.
[0060] like Figure 4A As shown in the above Figure 2A Based on the illustrated embodiment, step S205 may include the following steps S2051a to S2053a.
[0061] Step S2051a: When the target information is the annotation frame of the first annotation result, in response to detecting an adjustment operation on the first sub-annotation frame in the annotation frame on the user interface, determining the adjusted first sub-annotation frame.
[0062] For example, if the target information is a label frame of a first labeling result, the three-dimensional point cloud data and the label frame are displayed on a visual interactive interface (user interface), allowing the user to use the visual interactive interface to adjust the first sub-label frame of the label frame in the three-dimensional point cloud data to obtain an adjusted first sub-label frame. The label frame of the first labeling result includes a first sub-label frame and a second sub-label frame, and the first sub-label frame and the second sub-label frame are respectively used to annotate different contents in the positioning target marker.
[0063] In some examples, when the time series data includes at least one image sequence to be annotated, at least one frame of the image to be annotated in the at least one image sequence to be annotated can be determined. Furthermore, the annotation box (including the first sub-annotation box and the second sub-annotation box) of the first annotation result can be projected to at least one frame of the image to be annotated, so that the user can check the annotation accuracy of the first annotation result. If the inspection result shows that the geometric information and attribute information of the first sub-annotation box and the second sub-annotation box are both annotated correctly, then in response to detecting a first click operation on the user interface, the first annotation result is determined as the second annotation result. Among them, the geometric information includes but is not limited to: position, size (length, width, height), etc. The attribute information includes but is not limited to: orientation, shape, type, etc. The first click operation is used to trigger the event of determining the first annotation result as the second annotation result.
[0064] In some examples, if the verification result indicates that the geometric information and / or attribute information of the first sub-label box is incorrect, then in response to detecting an adjustment operation on the first sub-label box on the user interface, an adjusted first sub-label box is determined.
[0065] For example, the target landmark may include a traffic light, and the traffic light includes a light box and a light bulb. Therefore, the first sub-annotation box is a light box annotation box, and the second sub-annotation box is a light bulb annotation box. Furthermore, if the light box type marked in the light box annotation box is correct, in response to detecting an adjustment operation on the orientation, position, and size of the light box annotation box on the user interface, an adjusted light box annotation box is obtained. If the light box type marked in the light box annotation box is incorrect, in response to detecting an adjustment operation on the light box type in the light box annotation box on the user interface, a light box annotation box with an adjusted type is obtained; further, in response to detecting an adjustment operation on the orientation, position, and size of the light box annotation box on the user interface, an adjusted light box annotation box is obtained.
[0066] In some examples, the first sub-annotation box can be adjusted in three views (top view, front view, and side view), so that the first sub-annotation box can be displayed from various angles to facilitate user adjustment.
[0067] Step S2052a: Automatically generate a second sub-labeling frame based on the adjusted first sub-labeling frame.
[0068] For example, if the inspection results indicate that the geometric information and / or attribute information of the first sub-annotation box is incorrectly labeled, then in response to detecting an adjustment operation on the first sub-annotation box on the user interface, an adjusted first sub-annotation box is determined. Typically, the first sub-annotation box and the second sub-annotation box are associated with each other, for example, the second sub-annotation box is located within the first sub-annotation box. Therefore, if the geometric information and / or attribute information of the first sub-annotation box is incorrectly labeled, the second sub-annotation box may also be incorrectly labeled. Furthermore, in response to detecting a delete operation on the second sub-annotation box on the user interface, a second sub-annotation box may be automatically generated based on the adjusted first sub-annotation box. Of course, the second sub-annotation box may also be automatically generated based on the adjusted first sub-annotation box in response to detecting a second click operation on the user interface. The second click operation is used to trigger the event of automatically generating the second sub-annotation box. It should be noted that in the embodiments of the present disclosure, the triggering conditions for automatically generating the second sub-annotation box based on the adjusted first sub-annotation box are not limited, and the above two are merely exemplary.
[0069] In some examples, the target marker may be a traffic light, which includes a light box and a light bulb. Therefore, the first sub-annotation frame is the light box annotation frame, the second sub-annotation frame is the light bulb annotation frame, and the light bulb annotation frame is located within the light box annotation frame. Furthermore, after the light box annotation frame is adjusted, the light bulb annotation frame may be automatically generated based on the adjusted light box annotation frame. For example, a light bulb 3D annotation frame may be automatically generated based on the adjusted light box 3D annotation frame in accordance with the national standard for traffic lights. For example, a light bulb 3D annotation frame with the same orientation may be generated by shrinking the light box by a certain proportion and then dividing it equally. In this way, accurate and consistent true value data of parameters such as position and orientation can be obtained, which can improve annotation efficiency while ensuring annotation quality.
[0070] In some examples, after the second sub-annotation frame is automatically generated, it can also be fine-tuned through three views (top view, front view, side view), and the fine-tuned second sub-annotation frame can be projected into the image domain to check the annotation quality. Figure 4B As shown, Picture 41 is a light bulb in the 3D point cloud data under the top view, Picture 42 is a light bulb in the 3D point cloud data under the side view, and Picture 43 is a light bulb in the 3D point cloud data under the front view; the boxes in the pictures are used to indicate the position of the light bulb in the 3D point cloud data, and the arrows are used to indicate the direction of the light bulb. Then, the position and size of the light bulb can be fine-tuned in the above three views, and the annotation quality can be checked by projecting it into the image domain during the fine-tuning process, and finally the light bulb annotation box after fine-tuning is obtained. Figure 4C As shown, the 3D marking frame 44 is the final light bulb marking frame, and the 3D marking frame 45 is the adjusted light box marking frame.
[0071] Step S2053a: Determine a second annotation result of the target marker in the three-dimensional point cloud data based on the second sub-annotation frame and the adjusted first sub-annotation frame.
[0072] For example, the adjusted first sub-annotation frame and the automatically generated second sub-annotation frame can be used to determine the second annotation structure of the target marker in the three-dimensional point cloud data. Figure 4C As shown, the light bulb 3D annotation frame 44 and the light box 3D annotation frame 45 together constitute a second 3D annotation frame of the traffic light.
[0073] The target marker marking method provided by the embodiment of the present disclosure determines the adjusted first sub-marking frame in response to detecting an adjustment operation of the first sub-marking frame in the marking frame on the user interface when the target information is the marking frame of the first marking result; automatically generates a second sub-marking frame based on the adjusted first sub-marking frame; and determines the second marking result of the target marker in the three-dimensional point cloud data based on the second sub-marking frame and the adjusted first sub-marking frame; in this way, the second sub-marking frame can be automatically generated based on the adjusted first sub-marking frame in the automatic pre-flush result, thereby conveniently, quickly and accurately realizing the marking of the target marker.
[0074] like Figure 5A As shown in the above Figure 2A Based on the illustrated embodiment, step S205 may include the following steps S2051b to S2052b.
[0075] S2051b. When the target information is a prediction box of a predicted position, in response to detecting a selection operation on the user interface, determine the position information of the prediction box and the type of the target marker corresponding to the selection operation.
[0076] Exemplarily, if the target information is a prediction box of a predicted position, the three-dimensional point cloud data and the prediction box are displayed on a visual interactive interface (user interface), so that the user can manually mark the target marker using the position information of the prediction box on the visual interactive interface to obtain a corresponding second marking result.
[0077] In some examples, the selection operation in step S2051b is used to select the type of the target marker. For example, the type of the target marker includes but is not limited to a traffic light, a traffic sign, etc. In the embodiment of the present disclosure, the type of the target marker may further include the types of different contents in the target marker. For example, when the target marker is a traffic light, the traffic light includes a light box and a light bulb. Furthermore, the type of the light box may also be included. Figure 5B As shown, the operator can select whether the predicted box is marked with a traffic light or a signboard, and the light box type of the traffic light in the annotation directory in the user interface 51. Figure 5BAs can be seen in the figure, the types of traffic light boxes include but are not limited to three longitudinal signal lights, three transverse signal lights, two longitudinal signal lights, two transverse signal lights, single signal lights and other lights.
[0078] S2052b. Based on the position information of the prediction box and the type of the target marker, automatically generate a second annotation result of the target marker in the three-dimensional point cloud data.
[0079] For example, based on the position information of the prediction box in the 3D point cloud data and the type of the selected target marker, a second annotation result of the target marker in the 3D point cloud data can be automatically generated at the location indicated by the position information. Figure 5B As shown, when the operator selects the target marker as a traffic light and the light box type as a horizontal double signal light in the user interface 51, a light box 3D annotation frame corresponding to the horizontal double signal light can be automatically generated at the location indicated by the position information in response to the operator's selection operation. When the operator determines that the light box 3D annotation frame is accurate, the light box can be retracted by a certain proportion and then divided into two equal parts in response to the operator's second click operation to automatically generate a light bulb 3D annotation frame with the same orientation as the light box 3D annotation frame. That is to say, in the embodiment of the present disclosure, by selecting the light box type and pasting the light box 3D annotation frame of the corresponding type to the position where the traffic light is predicted in the three-dimensional point cloud data, a light bulb 3D annotation frame with the same orientation can be automatically generated, and then the operator can adjust the orientation, position and size of the light box 3D annotation frame and / or the light bulb 3D annotation frame through the three-view drawing, and project it to the image domain in real time during the adjustment process to check the annotation quality.
[0080] In some examples, when the time series data includes at least one image sequence to be annotated, at least one frame of the image sequence to be annotated can be determined. Furthermore, the automatically generated second annotation result can be projected onto the at least one frame of the image to be annotated, so that the user can check the annotation accuracy of the automatically generated second annotation result. If the inspection result shows that the annotation accuracy of the automatically generated second annotation result does not meet the requirements, the automatically generated second annotation result can be manually adjusted on the user interface until the projection meets the requirements. Figure 5CAs shown, Picture 52 is the projection result of the light bulb 3D annotation box in the second annotation result in one image to be annotated, and Picture 53 is the projection result of the light bulb 3D annotation box in the second annotation result in another image to be annotated. Moreover, it can be seen from Pictures 52 and 53 that the annotation quality of the light bulb 3D annotation box in the three-dimensional point cloud data projected to the light bulb 2D (2Dimensional, two-dimensional) annotation box 54 in the image to be annotated meets the requirements (such as the light bulb 2D annotation box 54 and the light bulb of the traffic light in the image to be annotated meet the requirements), and thus it can be verified that the annotation accuracy of the light bulb 3D annotation box in the three-dimensional point cloud data meets the requirements. It should be noted that Pictures 52 and 53 include multiple projected 2D annotation boxes, but Figure 5C Only some 2D annotation boxes are marked as examples.
[0081] The target marker labeling method provided by the embodiment of the present disclosure determines the position information of the prediction box and the type of the target marker corresponding to the selection operation in response to detecting a selection operation on the user interface when the target information is a prediction box of the predicted position; based on the position information of the prediction box and the type of the target marker, a second labeling result of the target marker in the three-dimensional point cloud data is automatically generated; in this way, the second labeling result can be automatically generated based on the predicted position of the target marker and the user's selection operation, thereby conveniently, quickly and accurately realizing the labeling of the target marker.
[0082] like Figure 6 As shown in the above Figure 2A Based on the illustrated embodiment, step S202 may include the following steps S2021 to S2024.
[0083] Step S2021: construct three-dimensional point cloud data of the target scene based on the time series data, and perform semantic segmentation processing on the three-dimensional point cloud data to obtain semantic information of the three-dimensional point cloud data.
[0084] For example, semantic segmentation can be performed on the 3D point cloud data of the target scene to obtain semantic information for each point in the 3D point cloud data. For example, a trained semantic segmentation model can be used to process the 3D point cloud data to produce an image of the same size as the input point cloud, where each pixel corresponds to the semantic label of a point in the input point cloud. Common semantic segmentation algorithms include PointNet, PointNet++, and Pillar-based methods.
[0085] It should be noted that the embodiments of the present disclosure do not limit the specific implementation of semantic segmentation, nor do they limit the model structure of the semantic segmentation model.
[0086] In some examples, the semantic information of three-dimensional point cloud data refers to the content category of each point cloud in the three-dimensional point cloud. For example, the semantic information of point cloud A in the three-dimensional point cloud data is a traffic light, the semantic information of point cloud B is a tree, the semantic information of point cloud C is a building, and the semantic information of point cloud D is a pedestrian, etc.
[0087] It is understandable that the specific implementation method of constructing the three-dimensional point cloud data of the target scene based on the time series data can be found in the description of the above step S202, which will not be repeated here.
[0088] Step S2022: Grid the three-dimensional point cloud data to obtain multiple point cloud blocks.
[0089] For example, the three-dimensional point cloud data may be gridded to obtain a plurality of point cloud blocks, wherein each of the plurality of point cloud blocks represents a small area of space and contains a portion of the point cloud.
[0090] It should be noted that the embodiments of the present disclosure do not limit the granularity of the gridding process, that is, the size of the point cloud block, and those skilled in the art can set it according to actual use requirements and the size of the three-dimensional point cloud data.
[0091] Step S2023: Based on the semantic information of the three-dimensional point cloud data, determine at least one target point cloud block from the multiple point cloud blocks.
[0092] For example, if the target landmark in the disclosed embodiment is a traffic light, then based on the semantic information of the three-dimensional point cloud data, a point cloud block containing semantic information about the traffic light can be identified from multiple point cloud data blocks as the target point cloud block. For example, when the target landmark is a traffic light, the semantic information of each point cloud can be compared with the preset semantic information (traffic light) to obtain a semantic comparison result for each point cloud. Furthermore, based on the semantic comparison result of each point cloud and the point cloud block to which the point cloud belongs, the target point cloud block is determined.
[0093] Step S2024: cluster at least one target point cloud block to obtain a predicted position of the target marker in the three-dimensional point cloud data.
[0094] For example, after determining at least one target point cloud block from a plurality of point cloud blocks, clustering processing may be performed on the at least one target point cloud block to obtain a predicted position of the target marker in the three-dimensional point cloud data.
[0095] In some examples, clustering at least one target point cloud block in the disclosed embodiments involves dividing the at least one target point cloud block into a number of groups or clusters based on similarity. After the division is complete, each group or cluster represents a target marker, and the predicted location of the at least one target marker in the 3D point cloud data can be obtained.
[0096] The target marker labeling method provided by the embodiment of the present disclosure performs semantic segmentation processing on three-dimensional point cloud data to obtain semantic information of the three-dimensional point cloud data; performs grid processing on the three-dimensional point cloud data to obtain multiple point cloud blocks; based on the semantic information of the three-dimensional point cloud data, determines at least one target point cloud block from the multiple point cloud blocks; performs clustering processing on the at least one target point cloud block to obtain a predicted position of the target marker in the three-dimensional point cloud data; in this way, the position of the target marker in the three-dimensional point cloud can be obtained by segmenting and clustering the three-dimensional point cloud data, so that the position can be used to verify the pre-brush result or use the position as an anchor point to draw labeling information.
[0097] In some embodiments, based on the semantic information of three-dimensional point cloud data, at least one target point cloud block is determined from multiple point cloud blocks, including: determining the point cloud included in each point cloud block in the multiple point cloud blocks, and the semantic information of the point cloud included in each point cloud block; matching the semantic information of the point cloud included in each point cloud block with the target marker to obtain the semantic matching results of each point cloud block; and determining the point cloud block with a successful semantic matching result as the target point cloud block.
[0098] For example, the semantic information of each point cloud in the point cloud block can be matched with the target landmark to obtain a semantic matching result for each point cloud. The point cloud block containing the point cloud with a successful semantic match is then determined as the target point cloud block. In this way, the target point cloud block can be determined based on the semantic information, and the predicted position of the target landmark can be determined.
[0099] like Figure 7 As shown in the above Figure 2A Based on the embodiment shown, the following steps S206 and S207 are also included.
[0100] Step S206 : When the time series data includes at least one image sequence to be labeled, determine a plurality of frames of images to be labeled in the at least one image sequence to be labeled.
[0101] For example, Figure 1 As shown, a vehicle 10 passes through a target scene 13 at a preset speed. Each image sensor 11 on the vehicle 10 captures images of the target scene 13 at a preset frequency, thereby obtaining multiple image sequences to be labeled. Furthermore, multiple frames of images to be labeled can be determined for each of the multiple image sequences to be labeled.
[0102] Step S207 : Project the second annotation result of the target marker in the three-dimensional point cloud data to each frame of the image to be annotated, and obtain the annotation information of the target marker in each frame of the image to be annotated.
[0103] Exemplarily, the second annotation result (such as a 3D annotation box) of the target marker in the three-dimensional point cloud data can be projected onto each frame of the image to be annotated in the multiple frames of the image to be annotated to obtain the annotation information (such as a 2D annotation box) of the target marker in each frame of the image to be annotated. For example, based on the position information of the target marker in the three-dimensional point cloud data, the corresponding pixel of the target marker in the image to be annotated can be determined by coordinate system conversion, and then the second annotation result of the target marker in the three-dimensional point cloud can be projected onto the corresponding pixel in the image to be annotated to complete the annotation of the target marker on the image to be annotated.
[0104] The target marker labeling method provided by the embodiment of the present disclosure determines multiple frames of images to be labeled in at least one sequence of images to be labeled; projects the second labeling result of the target marker in the three-dimensional point cloud data to each frame of the image to be labeled, and obtains the labeling information of the target marker in each frame of the image to be labeled; in this way, it is possible to implement only one labeling of the target marker in the three-dimensional point cloud data constructed from the time series data by projection, and obtain the labeling results of the target marker in multiple images to be labeled in the target scene, thereby effectively improving the image labeling efficiency.
[0105] Exemplary devices
[0106] Figure 8 A target marker marking device provided in an embodiment of the present disclosure, such as Figure 8 As shown, the annotation device 800 includes a time series data determination module 801 , a predicted position determination module 802 , a pre-flash result determination module 803 , an annotation display module 804 and an annotation result determination module 805 .
[0107] A time series data determination module 801 is used to determine the time series data corresponding to the target scene;
[0108] The predicted position determination module 802 is used to construct three-dimensional point cloud data of the target scene based on the time series data and determine the predicted position of the target landmark in the three-dimensional point cloud data;
[0109] The pre-brush result determination module 803 is used to obtain a first annotation result of the target marker in the three-dimensional point cloud data;
[0110] The annotation display module 804 is used to determine target information corresponding to the target marker based on the predicted position and the first annotation result, and display the target information on the user interface;
[0111] The annotation result determination module 805 is configured to determine a second annotation result of the target marker in the three-dimensional point cloud data in response to detecting an annotation operation on the target information on the user interface.
[0112] In some embodiments, as Figure 9 As shown, the annotation display module 804 includes a matching unit 8041 , a determination unit 8042 and a display unit 8043 .
[0113] a matching unit 8041, configured to match the predicted box of the predicted position with the labeled box of the first labeled result to obtain a matching result;
[0114] a determination unit 8042, configured to determine target information corresponding to the target marker based on the matching result;
[0115] The display unit 8043 is used to display the target information on the user interface.
[0116] In some embodiments, the determination unit 8042 is specifically used to determine the labeled box as the target information corresponding to the target marker if the matching result shows that the predicted box and the labeled box are in the same position; or, if the matching result shows that the predicted box and the labeled box are in different positions, determine the predicted box as the target information corresponding to the target marker.
[0117] In some embodiments, when the target information is the annotation box of the first annotation result, the annotation result determination module 805 is specifically used to determine the adjusted first sub-annotation box in response to detecting an adjustment operation on the first sub-annotation box in the annotation box on the user interface; based on the adjusted first sub-annotation box, automatically generate a second sub-annotation box; based on the second sub-annotation box and the adjusted first sub-annotation box, determine the second annotation result of the target marker in the three-dimensional point cloud data.
[0118] In some embodiments, when the target information is a prediction box of a predicted position, the annotation result determination module 805 is specifically used to determine the position information of the prediction box and the type of the target marker corresponding to the selection operation in response to detecting a selection operation on the user interface; based on the position information of the prediction box and the type of the target marker, automatically generate a second annotation result of the target marker in the three-dimensional point cloud data.
[0119] In some embodiments, the predicted position determination module 802 is used to construct three-dimensional point cloud data of the target scene based on time series data; perform semantic segmentation processing on the three-dimensional point cloud data to obtain semantic information of the three-dimensional point cloud data; perform gridding processing on the three-dimensional point cloud data to obtain multiple point cloud blocks; based on the semantic information of the three-dimensional point cloud data, determine at least one target point cloud block from the multiple point cloud blocks; perform clustering processing on the at least one target point cloud block to obtain the predicted position of the target marker in the three-dimensional point cloud data.
[0120] In some embodiments, the predicted position determination module 802 is specifically used to construct three-dimensional point cloud data of the target scene based on time series data; perform semantic segmentation processing on the three-dimensional point cloud data to obtain semantic information of the three-dimensional point cloud data; perform gridding processing on the three-dimensional point cloud data to obtain multiple point cloud blocks; determine the point cloud included in each point cloud block in the multiple point cloud blocks, and the semantic information of the point cloud included in each point cloud block; match the semantic information of the point cloud included in each point cloud block with the target marker to obtain the semantic matching result of each point cloud block; determine the point cloud block with a successful semantic matching result as the target point cloud block; perform clustering processing on at least one target point cloud block to obtain the predicted position of the target marker in the three-dimensional point cloud data.
[0121] In some embodiments, as Figure 9 As shown, when the time series data includes at least one image sequence to be annotated, the annotation device 800 further includes a frame image determination module 806 and a projection module 807 .
[0122] The frame image determination module 806 is configured to determine a plurality of frames of images to be labeled in at least one sequence of images to be labeled;
[0123] The projection module 807 is used to project the second annotation result of the target marker in the three-dimensional point cloud data to each frame of the image to be annotated, so as to obtain the annotation information of the target marker in each frame of the image to be annotated.
[0124] The beneficial technical effects corresponding to the exemplary embodiment of the target marker labeling device 800 can be found in the corresponding beneficial technical effects of the exemplary method section above, and will not be repeated here.
[0125] Exemplary electronic devices
[0126] Figure 10 This is a structural diagram of an electronic device 100 provided in an embodiment of the present disclosure, including at least one processor 101 and a memory 102.
[0127] The processor 101 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 100 to perform desired functions.
[0128] The memory 102 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 101 may execute one or more computer program instructions to implement the target marker labeling method and / or other desired functions of the various embodiments of the present disclosure described above.
[0129] In one example, the electronic device 100 may further include an input device 103 and an output device 104 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0130] The input device 103 may also include, for example, a keyboard, a mouse, and the like.
[0131] The output device 104 can output various information to the outside, and may include, for example, a display, a speaker, a printer, a communication network and its connected remote output devices, etc.
[0132] Of course, to simplify, Figure 10 Only some of the components related to the present disclosure in the electronic device 100 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 100 may further include any other appropriate components according to specific application scenarios.
[0133] Exemplary computer program products and computer-readable storage media
[0134] In addition to the above-mentioned methods and devices, embodiments of the present disclosure may also provide a computer program product, comprising computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the target marker labeling method of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.
[0135] The computer program product may be written in any combination of one or more programming languages to implement the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0136] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the steps in the target marker labeling method of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.
[0137] Computer readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium is, for example, but not limited to, a system, device or component comprising electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0138] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be considered as essential to each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0139] Those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.
Claims
1. A method for labeling a target marker, comprising: Determine the time series data corresponding to the target scenario; Constructing three-dimensional point cloud data of the target scene based on the time series data, and determining a predicted position of a target marker in the three-dimensional point cloud data; Obtaining a first annotation result of the target marker in the three-dimensional point cloud data; Determining target information corresponding to the target marker based on the predicted position and the first annotation result, and displaying the target information on a user interface; In response to detecting a labeling operation on the target information on the user interface, a second labeling result of the target marker in the three-dimensional point cloud data is determined.
2. The method according to claim 1, wherein The determining, based on the predicted position and the first annotation result, target information corresponding to the target marker includes: Matching the predicted box of the predicted position with the marked box of the first marked result to obtain a matching result; Based on the matching result, target information corresponding to the target marker is determined.
3. The method according to claim 2, wherein: Determining target information corresponding to the target marker based on the matching result includes: If the matching result indicates that the predicted box and the marked box are located at the same position, the marked box is determined as the target information corresponding to the target marker; or If the matching result indicates that the positions of the predicted box and the marked box are different, the predicted box is determined as the target information corresponding to the target marker.
4. The method according to claim 3, wherein: When the target information is a labeling frame of the first labeling result, determining a second labeling result of the target marker in the three-dimensional point cloud data in response to detecting a labeling operation on the target information on the user interface includes: In response to detecting an adjustment operation on the first sub-label frame in the label frame on the user interface, determining the adjusted first sub-label frame; Automatically generating a second sub-annotation frame based on the adjusted first sub-annotation frame; Based on the second sub-annotation frame and the adjusted first sub-annotation frame, a second annotation result of the target marker in the three-dimensional point cloud data is determined.
5. The method according to claim 3, wherein When the target information is a prediction frame of the predicted position, in response to detecting a labeling operation on the target information on the user interface, determining a second labeling result of the target marker in the three-dimensional point cloud data includes: In response to detecting a selection operation on the user interface, determining position information of the prediction box and a type of the target marker corresponding to the selection operation; Based on the position information of the prediction box and the type of the target marker, a second annotation result of the target marker in the three-dimensional point cloud data is automatically generated.
6. The method according to any one of claims 1 to 5, wherein Determining the predicted position of the target marker in the three-dimensional point cloud data includes: Performing semantic segmentation processing on the three-dimensional point cloud data to obtain semantic information of the three-dimensional point cloud data; Performing grid processing on the three-dimensional point cloud data to obtain a plurality of point cloud blocks; Determining at least one target point cloud block from the plurality of point cloud blocks based on semantic information of the three-dimensional point cloud data; Clustering is performed on the at least one target point cloud block to obtain a predicted position of a target marker in the three-dimensional point cloud data.
7. The method according to claim 6, wherein: The determining, based on the semantic information of the three-dimensional point cloud data, at least one target point cloud block from the plurality of point cloud blocks comprises: Determining a point cloud included in each of the plurality of point cloud blocks, and semantic information of the point cloud included in each of the point cloud blocks; Matching the semantic information of the point cloud included in each point cloud block with the target marker to obtain a semantic matching result of each point cloud block; The point cloud block with a successful semantic matching result is determined as the target point cloud block.
8. The method according to any one of claims 1 to 5, wherein the time series data comprises at least one image sequence to be labeled, and the method further comprises: Determining a plurality of frames of images to be labeled in the at least one sequence of images to be labeled; The second annotation result of the target marker in the three-dimensional point cloud data is projected onto each frame of the image to be annotated to obtain annotation information of the target marker in each frame of the image to be annotated.
9. A target marker marking device, comprising: A time series data determination module is used to determine the time series data corresponding to the target scene; A predicted position determination module, configured to construct three-dimensional point cloud data of the target scene based on the time series data, and determine a predicted position of a target marker in the three-dimensional point cloud data; a pre-brush result determination module, configured to obtain a first annotation result of the target marker in the three-dimensional point cloud data; a marking display module, configured to determine target information corresponding to the target marker based on the predicted position and the first marking result, and display the target information on a user interface; The labeling result determination module is used to determine a second labeling result of the target marker in the three-dimensional point cloud data in response to detecting a labeling operation on the target information on the user interface.
10. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the target marker labeling method according to any one of claims 1 to 8.
11. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the target marker labeling method described in any one of claims 1-8.