Labeling method and device of ground elements, equipment and storage medium
By constructing three-dimensional point cloud data to generate bird's-eye view and elevation map, combining user interface and elevation map, the problem of high cost and difficulty of marking ground elements in the existing technology is solved, and efficient and accurate labeling of ground elements in intelligent driving scenarios is achieved.
Patent Information
- Application Number
- CN202510637386.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, the ground element annotation method has the problem of high labeling cost and difficulty in 2D images and 3D point cloud spaces. Especially in intelligent driving scenarios, the existing methods cannot efficiently and accurately obtain the labeling truth value of the ground element.
By constructing three-dimensional point cloud data of the target scene, a bird's-eye view and an elevation map are generated, and a bird's-eye view is displayed on the user interface, responding to the user's annotation operation of ground elements, and determining the annotation results of ground elements in combination with the elevation map, the semi-automated human-computer interaction labeling process is realized.
High-quality and fast ground element labeling is achieved, reducing the cost and difficulty of labeling, improving the labeling efficiency and accuracy, and able to obtain high-quality ground element labeling truth values in intelligent driving scenarios.
Smart Images

Figure CN120472048A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of intelligent driving, and in particular to a method, apparatus, device, and storage medium for labeling ground elements. Background Art
[0002] Model training is essential for achieving intelligent driving, and the model training process typically requires a large amount of labeled data. Therefore, data labeling is a very important step. As we all know, the more accurate the data labeling, the better the algorithm model training effect.
[0003] Ground elements are frequently used in intelligent driving. There are two main methods for labeling ground elements in existing technologies. One method involves directly labeling ground elements on multiple acquired 2D (two-dimensional) images, which is inaccurate and expensive. The other method involves labeling ground elements in 3D (three-dimensional) point clouds, which is more difficult and time-consuming. Summary of the Invention
[0004] Currently, most methods for annotating ground elements in intelligent driving scenarios are performed in 2D images or 3D point clouds. This is very costly due to the large number of sample images required for 2D images. Furthermore, the quality of 3D point cloud data is generally poor and prone to noise, which increases the difficulty of annotation.
[0005] In order to solve the above technical problems, the present disclosure provides a ground element labeling method and device, equipment, and storage medium, which can solve the problems of high labeling cost and difficulty caused by existing ground element labeling methods.
[0006] A first aspect of the present disclosure provides a method for labeling ground elements, comprising: determining time series data corresponding to a target scene, and constructing three-dimensional point cloud data of the target scene based on the time series data; determining a bird's-eye view and an elevation map corresponding to the three-dimensional point cloud data based on the three-dimensional point cloud data, and displaying the bird's-eye view on a user interface; in response to detecting a labeling operation for a ground element in the bird's-eye view, determining first labeling information of the ground element; and determining a labeling result of the ground element based on the first labeling information and the elevation map.
[0007] The second aspect of the present disclosure provides a ground element labeling device, comprising: a scene construction module, for determining time series data corresponding to a target scene, and constructing three-dimensional point cloud data of the target scene based on the time series data; a data determination module, for determining a bird's-eye view and an elevation map corresponding to the three-dimensional point cloud data based on the three-dimensional point cloud data, and displaying the bird's-eye view on a user interface; an information determination module, for determining first labeling information of the ground element in response to detecting a labeling operation for the ground element in the bird's-eye view; and a result determination module, for determining the labeling result of the ground element based on the first labeling information and the elevation map.
[0008] A third aspect of the present disclosure provides a computer-readable storage medium storing a computer program for executing the ground element labeling method provided in the first aspect.
[0009] The fourth aspect of the present disclosure provides an electronic device, which includes: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the ground element labeling method provided in the first aspect above.
[0010] A fifth aspect of the present disclosure provides a computer program product. When instructions in the computer program product are executed by a processor, the method for labeling ground elements provided in the first aspect is executed.
[0011] Based on the ground element labeling method provided by the present invention, by determining the time series data corresponding to the target scene, and constructing the three-dimensional point cloud data of the target scene based on the time series data; based on the three-dimensional point cloud data, determining the bird's-eye view and elevation map corresponding to the three-dimensional point cloud data, and displaying the bird's-eye view on the user interface; in response to detecting the labeling operation for the ground elements in the bird's-eye view, determining the first labeling information of the ground elements; based on the first labeling information and the elevation map, determining the labeling result of the ground elements; in this way, the bird's-eye view obtained through the three-dimensional point cloud data can be displayed on the user interface so that the labeling operation can be manually performed to obtain the 2D labeling information of the ground elements, and then the elevation value of the 2D labeling information is obtained from the elevation map obtained through the three-dimensional point cloud data to obtain the complete 3D labeling information of the ground elements, so that high-quality ground element labeling truth values can be obtained conveniently and quickly based on semi-automatic human-computer interaction and combined with the elevation map. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A schematic diagram of an application scenario provided by an exemplary embodiment of the present disclosure.
[0013] Figure 2A flowchart of a method for labeling ground elements provided by an exemplary embodiment of the present disclosure.
[0014] Figure 3A A schematic flow chart of a method for labeling ground elements provided in accordance with another exemplary embodiment of the present disclosure.
[0015] Figure 3B A schematic diagram of a user interface provided by an exemplary embodiment of the present disclosure.
[0016] Figure 4 A schematic flow chart of a method for labeling ground elements provided in yet another exemplary embodiment of the present disclosure.
[0017] Figure 5A A schematic flow chart of a method for labeling ground elements provided in yet another exemplary embodiment of the present disclosure.
[0018] Figure 5B A schematic diagram of a user interface provided by another exemplary embodiment of the present disclosure.
[0019] Figure 6 A schematic flow chart of a method for labeling ground elements provided in yet another exemplary embodiment of the present disclosure.
[0020] Figure 7 A schematic structural diagram of a ground element labeling device provided by an exemplary embodiment of the present disclosure.
[0021] Figure 8 A schematic structural diagram of a ground element labeling device provided by another exemplary embodiment of the present disclosure.
[0022] Figure 9 A schematic structural diagram of a ground element labeling device provided by yet another exemplary embodiment of the present disclosure.
[0023] Figure 10 The present invention provides a structural diagram of an electronic device according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0024] To explain the present disclosure, example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. It should be understood that the present disclosure is not limited to the example embodiments.
[0025] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.
[0026] Application Overview
[0027] The ground element labeling method provided by the embodiment of the present disclosure can be applied to the front end of the model training scenario, and any other feasible scenario. For example, it can be applied to generate the training samples required for the perception model training: when training the perception model for detecting ground elements (such as lane lines, ground signs, curbs, etc.), the labeling results of the ground elements in the sample image can be obtained through the embodiment of the present disclosure, and the sample image with the labeling results can be used as a training sample to train the perception model until the preset training completion conditions are met, thereby obtaining a trained perception model that can be used for downstream perception tasks.
[0028] Figure 1 This is a schematic diagram of an application scenario provided by an exemplary embodiment of the present disclosure. Figure 1 As shown, a vehicle 10 is equipped with multiple image sensors 11 (e.g., cameras) and a 3D scanning device 12 (e.g., a lidar) in different orientations. Vehicle 10 travels through a target scene 13 at a preset speed. The target scene 13 includes a specific intersection or road section. While vehicle 10 is traveling, each image sensor 11 on vehicle 10 captures images of the target scene 13 at a first preset frequency, generating a sequence of images to be labeled. Consequently, the multiple image sensors 11 generate multiple image sequences to be labeled. Each image sequence to be labeled includes multiple frames of images to be labeled, and the multiple frames in each sequence are ordered by image acquisition time or the vehicle 10's travel trajectory. While vehicle 10 is traveling, the 3D scanning device 12 can also capture point cloud data of the target scene 13 at a second preset frequency, generating a point cloud dataset. The first and second preset frequencies can be the same or different. Based on the multiple image sequences to be labeled and / or the point cloud dataset, 3D scene reconstruction is performed on the target scene 13 to generate 3D point cloud data of the target scene 13. Furthermore, the three-dimensional point cloud data is processed by the ground element labeling method provided in the embodiment of the present disclosure to obtain the corresponding bird's-eye view and elevation map, and the bird's-eye view and elevation map are used to obtain the labeling results of the ground elements (such as lane lines, right turn arrows, curbs, etc.).
[0029] It should be noted that Figure 1 This is only an example, and the embodiments of the present disclosure Figure 1 There is no limitation on the number and location of the multiple image sensors 11 and the 3D scanning device 12. In actual use, the image sensor 11 can be the image sensor corresponding to the vehicle's forward wide-angle camera, the image sensor 11 can be the image sensor corresponding to the vehicle's forward narrow-angle camera, the image sensor 11 can also be the image sensor corresponding to the side view camera, etc.
[0030] That is, the embodiment of the present disclosure provides a method for labeling ground elements, by determining the time series data corresponding to the target scene, and constructing the three-dimensional point cloud data of the target scene based on the time series data; based on the three-dimensional point cloud data, determining the bird's-eye view and elevation map corresponding to the three-dimensional point cloud data, and displaying the bird's-eye view on the user interface; in response to detecting the labeling operation for the ground elements in the bird's-eye view, determining the first labeling information of the ground elements; based on the first labeling information and the elevation map, determining the labeling result of the ground elements; in this way, the bird's-eye view obtained through the three-dimensional point cloud data can be displayed on the user interface so that the labeling operation can be manually performed to obtain the 2D labeling information of the ground elements, and then the elevation value of the 2D labeling information is obtained from the elevation map obtained through the three-dimensional point cloud data to obtain the complete 3D labeling information of the ground elements, so that high-quality ground element labeling truth values can be obtained conveniently and quickly based on semi-automatic human-computer interaction and combined with the elevation map.
[0031] Exemplary Methods
[0032] Figure 2 This is a flow chart of a method for labeling ground elements provided by an embodiment of the present disclosure. This embodiment can be applied to any electronic device such as a local terminal device, a cloud server, etc., and can also be applied to terminal devices and cloud servers in a distributed manner, such as Figure 2 As shown, the method includes the following steps S201-S204.
[0033] Step S201: Determine the time series data corresponding to the target scene, and construct three-dimensional point cloud data of the target scene based on the time series data.
[0034] For example, the target scene may include a scene of a certain section of a highway or a scene of a certain intersection on a city road. The present disclosure does not limit the actual situation of the target scene. Factors affecting the target scene include, but are not limited to, driving location, road type, weather conditions, traffic conditions, etc.
[0035] In some examples, such as Figure 1 As shown, the vehicle 10 passes through the target scene 13 at a preset speed, and each image sensor 11 on the vehicle 10 respectively captures images of the target scene 13 at a first preset frequency to obtain a plurality of image sequences to be labeled, which are the time series data corresponding to the target scene in step S201; wherein, the frequency of image acquisition by each image sensor 11 can be the same or different, and the embodiment of the present disclosure does not impose any limitation on this. Similarly, if Figure 1As shown, vehicle 10 passes through target scene 13 at a preset speed. A 3D scanning device 12 on vehicle 10 scans target scene 13 at a second preset frequency, obtaining a plurality of point cloud data having a time-series relationship. This plurality of point cloud data is the time-series data corresponding to the target scene in step S201. In other words, the time-series data corresponding to the target scene in the disclosed embodiments includes, but is not limited to, an image sequence acquired at the first preset frequency and / or a point cloud dataset acquired at the second preset frequency in the target scene.
[0036] Exemplarily, the time series data corresponding to the target scene can be multiple image sequences to be annotated, and / or multiple point cloud data with a time series relationship. On this basis, the multiple point cloud data can be used to perform three-dimensional scene reconstruction on the target scene to obtain three-dimensional point cloud data of the target scene. Among them, the three-dimensional scanning equipment includes but is not limited to Lidar (Light Laser Detection and Ranging), structured light sensor, TOF (Time of Flight, flight ranging) camera, etc. Alternatively, the target scene can be reconstructed in three dimensions using multiple image sequences to be annotated to obtain three-dimensional point cloud data of the target scene. Of course, some auxiliary information may be needed in the three-dimensional scene reconstruction process, such as information measured by a vehicle wheel speedometer, GPS positioning information, etc. In some cases, multiple point cloud data with a time series relationship and at least one image sequence to be annotated can also be used simultaneously for three-dimensional reconstruction to improve the accuracy of the three-dimensional reconstruction result. The specific implementation method of constructing the three-dimensional point cloud data of the target scene in the embodiment of the present disclosure is not limited.
[0037] Step S202: Based on the three-dimensional point cloud data, determine the bird's-eye view and elevation map corresponding to the three-dimensional point cloud data, and display the bird's-eye view on the user interface.
[0038] Exemplarily, in the embodiments of the present disclosure, the BEV (Bird's Eye View) map and elevation map corresponding to the target scene can be determined based on the three-dimensional point cloud data of the target scene. For example, the three-dimensional point cloud data of the target scene can be flattened on the ground with a BEV perspective, and then the flattened result can be rasterized to obtain the BEV map corresponding to the three-dimensional point cloud data. At the same time, after performing projection, de-occlusion, and denoising on the three-dimensional point cloud data, the elevation map (Height map) corresponding to the three-dimensional point cloud data can be output; wherein, the BEV map and the elevation map constitute the reference information marked in the embodiments of the present disclosure. Of course, the bird's-eye view and elevation map corresponding to the three-dimensional point cloud data can also be determined by other algorithms, and the embodiments of the present disclosure do not limit this.
[0039] In some examples, since both the BEV map and the elevation map are obtained through three-dimensional point cloud data (i.e., the BEV map and the elevation map are of the same origin), the sizes of the BEV map and the elevation map are the same. Moreover, the pixels at the same position in the BEV map and the elevation map represent the same 3D position point in the three-dimensional point cloud data. Furthermore, the height information of the pixel coordinates (X, Y) can be determined by the correspondence between the pixel coordinates (X, Y) in the BEV map and the elevation map. For example, the elevation value of point (0, 0) in the elevation map can be used as the height information of point (0, 0) in the BEV map.
[0040] In some examples, since the BEV image is an image obtained by looking down at the ground from a high viewpoint, it can be easily used to perform 2D annotation of ground elements after visualization, thus obtaining 2D annotation information of ground elements. At the same time, since the elevation map can show the elevation distribution of each point on the ground, the height information of this 2D annotation information can be obtained from the elevation map, forming a complete 3D annotation information. In this way, after performing 2D annotation of ground elements on the BEV image and combining it with the elevation map, complete 3D annotation information of ground elements can be obtained.
[0041] In some examples, the BEV map (bird's-eye view) corresponding to the target scene can be visualized so that the BEV map of the target scene can be displayed on the visual interaction interface. For example, the BEV map of the target scene can be visualized through the visualization module in PCL (Point Cloud Library) and the Cloud Compare (point cloud registration) visualization software. In addition, the above-mentioned visual interaction interface (i.e., the user interface of step S202) is used to display the BEV map of the target scene. It can also be used to edit and reorganize the annotation information of the ground elements displayed in the BEV map according to the user's operating instructions (such as annotation instructions, adjustment instructions, etc.), so as to obtain the final annotation results of the ground elements. Among them, the visual interaction interface can display data through a display (for example, a liquid crystal display, a plasma display), and the visual interaction interface can also receive user operating instructions through a keyboard, a mouse, a touch screen, etc.
[0042] Step S203: In response to detecting a labeling operation on a ground element in the bird's-eye view, determine first labeling information of the ground element.
[0043] Exemplarily, the labeling operation includes an adjustment operation. For example, a bird's-eye view image can be input into a trained large model, and the output of the large model is the labeling information of the ground elements in the bird's-eye view image; then, the BEV image and the pre-flashed labeling information in the BEV image can be displayed on the user interface. At this time, the operator can manually adjust the pre-flashed labeling information on the user interface to make the labeling information have higher labeling accuracy. Furthermore, the electronic device can determine the first labeling information of the ground element in response to detecting the adjustment operation for the labeling information of the ground element; wherein, the first labeling information is the labeling information obtained after the operator adjusts the preset labeling information.
[0044] In some examples, the labeling operation includes a fully manual labeling operation. For example, an operator can manually label a ground element in the bird's-eye view through a user interface, thereby obtaining first labeling information for the ground element in the bird's-eye view. For example, lane lines in a BEV image can be manually labeled using lines; another example is that arrows in a BEV image can be manually labeled using 2D boxes. Furthermore, in response to detecting the fully manual labeling operation for the ground element, the electronic device can determine the first labeling information for the ground element.
[0045] In some examples, the labeling operation includes a semi-automatic manual labeling operation. For example, the operator can perform semi-automatic manual labeling on the ground elements in the bird's-eye view through the user interface, thereby obtaining the first labeling information of the ground elements in the bird's-eye view. For example, if there is a lane line in the bird's-eye view, the operator can click on the lane line; then, in response to detecting the click operation, the electronic device performs segmentation processing on the lane line according to the intensity information of the pixel points in the bird's-eye view to obtain a segmentation result, and automatically extracts the lane line based on the segmentation result for labeling. For another example, if there is a ground indicator arrow in the bird's-eye view, the operator can click on the ground indicator arrow; then, in response to detecting the click operation, the electronic device can perform edge detection or intensity clustering processing on the ground indicator arrow to obtain a processing result; then, the minimum enclosing rectangle of the ground indicator arrow is determined based on the processing result, and the minimum enclosing rectangle is marked in the bird's-eye view.
[0046] It should be noted that the specific type of annotation operation in step S203 is not limited in the present embodiment. It is sufficient that the first annotation information of the ground element can be determined according to the annotation operation. The first annotation information is 2D annotation information in the BEV map; for example, the first annotation information can be vectorized data such as points, lines, or surfaces.
[0047] Step S204: Determine the labeling results of the ground elements based on the first labeling information and the elevation map.
[0048] For example, the first annotation information is 2D annotation information in the BEV map, so the 3D annotation information corresponding to the 2D annotation information can be obtained in combination with the elevation map. Then, the 3D annotation information is determined as the annotation result of the ground element.
[0049] In some examples, the first annotation information is 2D annotation information in the BEV map, so the 3D annotation information corresponding to the 2D annotation information can be obtained by combining the elevation map. Figure 1 The labeling information of the ground elements in each image frame is obtained from the image frames of the plurality of image sequences to be labeled, and the labeling information of the ground elements in each image frame is determined as the labeling result of the ground elements.
[0050] The method for labeling ground elements provided by the embodiments of the present disclosure determines the time series data corresponding to the target scene and constructs the three-dimensional point cloud data of the target scene based on the time series data; determines the bird's-eye view and elevation map corresponding to the three-dimensional point cloud data based on the three-dimensional point cloud data, and displays the bird's-eye view on the user interface; determines the first labeling information of the ground element in response to detecting the labeling operation for the ground element in the bird's-eye view; determines the labeling result of the ground element based on the first labeling information and the elevation map; in this way, the bird's-eye view obtained through the three-dimensional point cloud data can be displayed on the user interface so that the labeling operation can be manually performed to obtain the 2D labeling information of the ground element, and then the elevation value of the 2D labeling information is obtained from the elevation map obtained through the three-dimensional point cloud data to obtain the complete 3D labeling information of the ground element, so that high-quality ground element labeling truth values can be obtained conveniently and quickly based on semi-automatic human-computer interaction and combined with the elevation map.
[0051] like Figure 3A As shown in the above Figure 2 Based on the illustrated embodiment, step S203 may include the following steps S2031 to S2033.
[0052] Step S2031: Process the three-dimensional point cloud data or the bird's-eye view to obtain vectorized data corresponding to ground elements.
[0053] Exemplarily, the three-dimensional point cloud data or the bird's-eye view can be input into a trained large model, and the output of the large model is the vectorized data corresponding to the ground elements in the three-dimensional point cloud data or the vectorized data corresponding to the ground elements in the bird's-eye view. For example, the large model can be a cloud-based automatic annotation large model: the three-dimensional point cloud data of the target scene or the bird's-eye view of the target scene is input into the cloud-based automatic annotation large model, and the vectorized annotation results of the ground elements in the target scene can be obtained. The vectorized data of the ground elements refers to the data generated by vectorizing the ground elements; the vectorized data includes but is not limited to points, lines, and surfaces. For example, a lane line can be represented by a line or a series of points. For another example, a ground indicator arrow can be represented by a box composed of lines.
[0054] It should be noted that the specific implementation method for obtaining ground element vectorized data from 3D point cloud data or bird's-eye view images is not limited in the disclosed embodiments. Those skilled in the art may obtain ground element vectorized data using a large model or other methods. Furthermore, the disclosed embodiments do not limit the model structure corresponding to the large model, the training samples used, the set iteration termination conditions, and the like.
[0055] Step S2032: Display the vectorized data corresponding to the ground elements in the bird's-eye view.
[0056] For example, vectorized data corresponding to ground elements may be displayed in a bird's-eye view, and both the bird's-eye view and the vectorized data may be displayed on a user interface. Figure 3B A user interface diagram is provided for an exemplary embodiment of the present disclosure, as shown in FIG. Figure 3B As shown, the BEV map 31 corresponding to the target scene is displayed on the user interface. The BEV map 31 contains lane lines 32 and a plurality of lane line vectorized data 33 corresponding to the lane lines 32. The lane line vectorized data 33 is a lane line mark composed of lines and points. It should be noted that Figure 3B There are multiple lane line vectorized data in , but only one is marked for example.
[0057] Step S2033: In response to detecting a labeling operation on the vectorized data, first labeling information of the ground element is determined.
[0058] For example, those skilled in the art understand that models typically do not achieve 100% accuracy due to various reasons (such as insufficient or overly complex training data, structural flaws in the model, etc.). Therefore, the large model used to determine the vectorized data of ground elements may have certain detection errors. For example, a ground element exists at position A in the 3D point cloud data, but the vectorized data output by the large model is at position B. Position A and position B are different locations in the 3D point cloud data. Another example is that multiple ground elements exist in a bird's-eye view, but the large model only outputs vectorized data for one ground element, leaving the others undetected. Furthermore, the vectorized data obtained due to various reasons is not necessarily accurate. Therefore, the operator can perform annotation operations on the vectorized data on the user interface, such as adjustment, deletion, full manual annotation, semi-automatic manual annotation, etc. Furthermore, in response to detecting the annotation operation on the vectorized data, the electronic device can determine first annotation information for the ground element, which is the 2D annotation information of the ground element in the bird's-eye view.
[0059] In some examples, the annotation operation in step S2031 includes at least an adjustment operation. For example, if the vectorized data does not meet the conditions, the operator can manually adjust the vectorized data on the user interface. Furthermore, the electronic device can determine the first annotation information for the ground element in response to detecting the adjustment operation on the vectorized data on the user interface.
[0060] The method for labeling ground elements provided by the embodiments of the present disclosure obtains vector data corresponding to the ground elements by processing three-dimensional point cloud data or a bird's-eye view; the vector data corresponding to the ground elements are displayed in the bird's-eye view; and in response to detecting a labeling operation on the vector data, first labeling information of the ground elements is determined; in this way, the labeling of ground elements can be implemented based on pre-flush data (i.e., vectorized data) from the perspective of the bird's-eye view, thereby improving the labeling efficiency while ensuring the labeling quality.
[0061] like Figure 4 As shown in the above Figure 3A Based on the illustrated embodiment, step S2033 may include the following steps S21a to S23a.
[0062] Step S21a: In response to detecting a deletion operation on the vectorized data, the vectorized data is deleted in the bird's-eye view.
[0063] For example, if the annotation accuracy of the vectorized data is low, indicating that the vectorized data is completely unusable, the operator can delete the vectorized data on the user interface. In response to detecting the deletion operation on the vectorized data, the electronic device can delete the vectorized data in the bird's-eye view.
[0064] Step S22a: In response to detecting a click operation in the bird's-eye view, determining a first target point corresponding to the click operation.
[0065] For example, when no vectorized data exists in the bird's-eye view, the operator can click on the bird's-eye view in the user interface; the click location corresponds to a ground element. For example, the operator can click on a lane line or a ground indicator arrow in the bird's-eye view. The first target point corresponding to the click is the click location in the bird's-eye view.
[0066] Step S23a: Determine first annotation information of the ground element based on the first target point corresponding to the click operation.
[0067] For example, based on the first target point corresponding to the click operation, the corresponding ground element (i.e., the ground element where the first target point is located) can be determined from the bird's-eye view image, and then the ground element can be automatically labeled to obtain the first labeling information of the ground element; in this way, semi-automatic manual labeling of the ground element can be achieved. For example, the operator clicks on a lane line in the bird's-eye view image, and the electronic device can determine the click position (first target point) corresponding to the click operation, and perform element segmentation on the pixel intensity information around the click position to automatically extract the lane line and label it, to obtain the first labeling information of the lane line. Pixel intensity information refers to the brightness value of the pixel, which is usually represented by a value between 0 and 255. For another example, the operator clicks on an indicator arrow in the bird's-eye view image, and the electronic device can determine the click position (first target point) corresponding to the click operation, and perform edge extraction on the indicator arrow based on the click position to obtain an edge extraction result; finally, the minimum circumscribed rectangle of the indicator arrow is determined based on the edge extraction result and labeled to obtain the first labeling information of the indicator arrow.
[0068] The method for labeling ground elements provided by the embodiment of the present disclosure deletes vector data in a bird's-eye view in response to detecting a deletion operation on vector data; determines a first target point corresponding to the click operation in response to detecting a click operation in the bird's-eye view; and determines first labeling information of the ground element based on the first target point corresponding to the click operation; in this way, the labeling information of the ground elements in the bird's-eye view can be determined by using a semi-automatic interactive extraction method, thereby conveniently, quickly and accurately obtaining 2D labeling information of the ground elements in the bird's-eye view.
[0069] In some embodiments, if there is no vectorized data in the bird's-eye view, the first annotation information of the ground element can also be determined through the operations in steps S22a to S23a. The specific implementation of steps S22a and S23a can be found in the above description and will not be repeated here.
[0070] In some embodiments, based on the first target point corresponding to the click operation, the first annotation information of the ground element is determined, including: based on the first target point corresponding to the click operation, determining the ground element to be labeled in the bird's-eye view; performing image segmentation processing on the ground element to be labeled to obtain the image segmentation result corresponding to the ground element to be labeled; based on the image segmentation result, automatically generating the first annotation information of the ground element to be labeled in the bird's-eye view.
[0071] For example, based on the click location (first target point) corresponding to the click operation, the ground element at the click location in the bird's-eye view can be determined, and the ground element at the click location can be determined as the ground element to be labeled. Furthermore, image segmentation processing can be performed on the ground element to be labeled to obtain a segmentation result for the ground element to be labeled in the bird's-eye view. The segmentation result of the ground element to be labeled in the bird's-eye view includes, but is not limited to, a segmentation mask for the ground element and semantic information for the ground element. The segmentation mask for the ground element is used to separate the ground element from the bird's-eye view, and the semantic information for the ground element is used to indicate the content of the ground element, thereby providing a prerequisite for subsequent automatic labeling. For example, pixel intensity information around the click location can be segmented to obtain a segmentation result for the ground element to be labeled. For another example, edge extraction can be performed on the ground element to be labeled to obtain a segmentation result for the ground element to be labeled. Furthermore, after obtaining the segmentation result for the ground element, first labeling information for the ground element to be labeled can be automatically generated in the bird's-eye view based on the segmentation result. For example, points or lines for labeling lane lines can be automatically generated in the bird's-eye view based on the segmentation result for lane lines. For example, based on the segmentation results of the ground indicator arrows, a 2D annotation box can be automatically generated in the bird's-eye view to annotate the ground indicator arrows. In this way, the annotation information of the ground elements can be extracted semi-automatically and interactively.
[0072] like Figure 5A As shown in the above Figure 2 Based on the illustrated embodiment, step S204 may include the following steps S2041 to S2043.
[0073] Step S2041: Determine the height information of the ground element based on the first annotation information of the ground element and the elevation map.
[0074] Exemplarily, the first annotation information of the ground element is the 2D annotation information in the bird's-eye view. Since the bird's-eye view and the elevation map are of the same origin (that is, the bird's-eye view and the elevation map are both obtained through the three-dimensional point cloud data of the same target scene), the bird's-eye view not only has the same size as the elevation map, but the position points on the bird's-eye view and the elevation map are also one-to-one corresponding. For example, the (0,1) coordinate point in the bird's-eye view and the (0,1) coordinate point in the elevation map represent the same 3D position point in the three-dimensional point cloud data. Therefore, the (X,Y) coordinates of each annotation point in the first annotation information can be determined first, and then the height information of each annotation point can be determined in the elevation map based on the (X,Y) coordinates of each annotation point. That is, the height information of each annotation point in the bird's-eye view can be determined by the correspondence between the coordinate points (X,Y) in the bird's-eye view and the elevation map, and then the height information of the ground element can be determined. In other words, the height information of each annotation point in the first annotation information can be obtained by the elevation value of the corresponding (X,Y) coordinate of each annotation point in the elevation map. Furthermore, the height information of each marked point in the first marked information constitutes the height information of the ground element. For example, if the coordinate position of a marked point in the first marked information in the bird's-eye view is (1, 2), then the height information of the marked point is the elevation value corresponding to the coordinate point (1, 2) in the elevation map.
[0075] In some examples, since the BEV image is an image obtained by looking down at the ground from a high viewpoint, it can be easily used to perform 2D annotation of ground elements after visualization, thus obtaining 2D annotation information of ground elements. At the same time, since the elevation map can show the elevation distribution of each point on the ground, the height information of this 2D annotation information can be obtained from the elevation map, forming a complete 3D annotation information. In this way, after performing 2D annotation of ground elements on the BEV image and combining it with the elevation map, complete 3D annotation information of ground elements can be obtained.
[0076] Step S2042: Determine second annotation information of the ground element based on the first annotation information and the height information of the ground element.
[0077] For example, the first annotation information for the ground element is 2D annotation information in the bird's-eye view. The height information of each point in the first annotation information can be determined through step S2041. Furthermore, 3D annotation information for the ground element can be obtained using the 2D annotation information and the height information of each point. This 3D annotation information is the second annotation information for the ground element.
[0078] Step S2043: Determine the labeling result of the ground element based on the second labeling information and the time series data.
[0079] For example, Figure 1As shown, during the driving process of the vehicle 10, each image sensor 11 on the vehicle 10 respectively captures images of the target scene 13 at a first preset frequency to obtain a sequence of images to be labeled. Therefore, multiple image sensors 11 correspond to multiple sequences of images to be labeled; wherein each sequence of images to be labeled includes multiple frames of images to be labeled, and the multiple frames of images to be labeled in each sequence of images to be labeled are sorted according to the image acquisition time or the driving trajectory of the vehicle 10. In addition, since the second annotation information is 3D annotation information, the 3D annotation information can be projected onto each frame of the images to be labeled in the multiple sequences of images to be labeled to obtain the annotation results of the ground elements in each frame of the images to be labeled. For example, the 3D annotation information of the ground element can be used to determine the pixel corresponding to the ground element in the image to be labeled by means of coordinate system conversion, and then the 3D annotation information of the ground element can be projected onto the pixel corresponding to it in the image to be labeled to complete the annotation of the ground element on the image to be labeled.
[0080] In some examples, such as Figure 5B As shown, the second annotation information of the lane line can be projected onto an image to be annotated 51 to obtain an annotation result 52 of the lane line on the image to be annotated 51. It can be seen that the annotation result 52 of the lane line on the image to be annotated 51 is composed of points and lines. It should be noted that Figure 5B There are other lane lines and other lane line corresponding annotation results in , but Figure 5B Only the annotation result corresponding to one lane line is used as an example.
[0081] The method for labeling ground elements provided by the embodiments of the present disclosure determines the height information of the ground elements based on the first labeling information of the ground elements and the elevation map; determines the second labeling information of the ground elements based on the first labeling information and the height information of the ground elements; and determines the labeling results of the ground elements based on the second labeling information and time series data. In this way, by only labeling the ground elements in the bird's-eye view, it is possible to label multiple images to be labeled corresponding to the target scene at one time, thereby effectively improving the image labeling efficiency.
[0082] In some embodiments, when the time series data includes at least one image sequence, the labeling results of the ground elements are determined based on the second labeling information and the time series data, including: determining at least one image frame including the ground elements based on the at least one image sequence; and projecting the second labeling information to the at least one image frame to obtain the labeling results of the ground elements.
[0083] For example, Figure 1As shown, vehicle 10 passes through target scene 13 at a preset speed. Each image sensor 11 on vehicle 10 captures images of target scene 13 at a first preset frequency, generating multiple image sequences to be labeled. Each image sequence to be labeled includes multiple frames of images to be labeled. Furthermore, at least one image frame can be identified from the multiple frames of images to be labeled, and the second annotation information can be projected onto the at least one image frame to obtain the annotation results for the ground elements.
[0084] In some examples, the 3D annotation information (second annotation information) of ground elements can be projected onto the image frame containing the ground elements, thereby achieving the effect of obtaining multiple ground truth data through a single annotation. For example, the 3D annotation information of a lane line can be projected onto the image frames corresponding to multiple track passes to obtain the ground truth data for each image frame corresponding to the track pass.
[0085] In some embodiments, the second annotation information is projected onto at least one image frame to obtain the annotation results of the ground elements, including: determining a target image frame from at least one image frame, and projecting the second annotation information onto the target image frame to determine the annotation accuracy of the second annotation information; in response to detecting an adjustment operation on the second annotation information on the user interface, determining the adjusted second annotation information; and projecting the adjusted second annotation information onto at least one image frame to obtain the annotation results of the ground elements.
[0086] For example, Figure 1 As shown, a vehicle 10 passes through a target scene 13 at a preset speed. Each image sensor 11 on the vehicle 10 captures images of the target scene 13 at a first preset frequency, generating multiple image sequences to be annotated. Each image sequence to be annotated includes multiple frames of images to be annotated. Furthermore, the target image frame can be any one of the multiple frames to be annotated, or it can be a frame of the multiple frames to be annotated that meets specific conditions. This is not a limitation in the present embodiment. In the present embodiment, the target image frame serves to verify the accuracy of the second annotation information. Furthermore, the second annotation information can be projected onto the target image frame to determine its accuracy. If the accuracy of the second annotation information meets the requirements, the second annotation information can be projected onto at least one image frame to obtain the annotation results for the ground elements. If the accuracy of the second annotation information does not meet the requirements, the electronic device, in response to detecting an adjustment operation on the user interface for the second annotation information, can project the adjusted second annotation information onto at least one image frame to obtain the annotation results for the ground elements. In this way, the accuracy of the second annotation information can be verified by projection, and the second annotation information can be converted into a final annotation result by projection. In the projection process, the camera parameters of the image sensor 11 need to be used.
[0087] like Figure 6 As shown in the above Figure 5A Based on the illustrated embodiment, step S2041 may include the following steps S21b to S23b.
[0088] Step S21b: Determine at least one second target point in the first annotation information, and the coordinate position of each second target point in the bird's-eye view.
[0089] For example, in the disclosed embodiments, the second target point may be a specific target point in the first annotation information, and the first annotation information can be obtained by using at least one second target point. For example, if the first annotation information is a line segment, the at least one second target point in the first annotation information may be a node in the line segment. For another example, if the first annotation information is a 2D annotation box, the at least one second target point in the first annotation information may be the four vertices of the annotation box.
[0090] In some examples, after determining at least one second target point, it is also necessary to determine the XY coordinate position of each second target point in the bird's-eye view.
[0091] Step S22b: Determine the elevation value corresponding to each second target point based on the coordinate position of each second target point in the bird's-eye view and the elevation map.
[0092] For example, since the bird's-eye view and the elevation map have the same size and each position point in the image corresponds one to one, for example, the (0,1) coordinate point in the bird's-eye view and the (0,1) coordinate point in the elevation map represent the same 3D position point in the three-dimensional point cloud data. Therefore, after determining the XY coordinate position of each second target point in the bird's-eye view, the electronic device can determine the elevation value of the corresponding position from the elevation map based on the XY coordinate position of each second target point in the bird's-eye view; in this way, the complete 3D geometric information of the ground element can be constructed, and then the 3D geometric information can be projected into the image domain to check the accuracy of the annotation result in real time, thereby fully ensuring the annotation quality of the data.
[0093] Step S23b: If the elevation value meets the preset condition, determine the height information of the ground element based on the elevation value corresponding to each second target point.
[0094] For example, since the elevation values of each point in the elevation map may deviate, for example, there may be a location with an elevation value of -128. In this case, the elevation value of this location is invalid and cannot reflect the actual elevation information, that is, the elevation value does not meet the preset conditions. Therefore, it is necessary to determine the height information of the ground element based on the elevation value corresponding to each second target point when the elevation value meets the preset conditions. Generally, the elevation value corresponding to each second target point constitutes the height information of the ground element.
[0095] In some examples, when annotating points, lines, or surfaces in the user interface (annotation tool), each location point (XY) will obtain the corresponding height information from the elevation map and determine the validity of the height information. If the height information is invalid, the drawing cannot be completed and the XY location point must be re-determined. If the height information is valid, the annotation is valid and the 3D annotation information of the ground element is obtained.
[0096] In some examples, the ground element annotation method in the embodiment of the present disclosure, based on data visualization, realizes the editing or new annotation of pre-brush elements (vectorized data) through the editing capabilities of points / lines / surfaces. During the drawing process, invalid elevation information will be filtered in combination with the elevation map. During the drawing process and after the drawing is completed, the 3D annotation information in the visualization interface will be projected onto the images corresponding to all trajectories in real time. The operator can check in real time whether the annotation results meet the requirements and make necessary adjustments. From the above, it can be seen that only one annotation of the ground elements on the BEV map is required to achieve multi-pass data annotation of the trajectory image.
[0097] The method for labeling ground elements provided by the embodiments of the present disclosure determines at least one second target point and the coordinate position of each second target point in the bird's-eye view in the first labeling information; determines the elevation value corresponding to each second target point based on the coordinate position of each second target point in the bird's-eye view and the elevation map; if the elevation value meets the preset conditions, determines the height information of the ground element based on the elevation value corresponding to each second target point; in this way, the height information of the ground element can be determined through at least one second target point to obtain the complete 3D geometric information of the ground element.
[0098] Exemplary devices
[0099] Figure 7 A ground element marking device provided by an embodiment of the present disclosure, such as Figure 7 As shown, the labeling device 700 includes a scene construction module 701 , a data determination module 702 , an information determination module 703 and a result determination module 704 .
[0100] A scene construction module 701 is used to determine the time series data corresponding to the target scene and construct the three-dimensional point cloud data of the target scene based on the time series data;
[0101] A data determination module 702 is configured to determine a bird's-eye view and an elevation map corresponding to the three-dimensional point cloud data based on the three-dimensional point cloud data, and to display the bird's-eye view on a user interface;
[0102] An information determining module 703 is configured to determine first annotation information of a ground element in response to detecting a annotation operation on the ground element in the bird's-eye view;
[0103] The result determination module 704 is configured to determine the labeling result of the ground element based on the first labeling information and the elevation map.
[0104] In some embodiments, as Figure 8 As shown, the information determination module 703 includes a data pre-flash unit 7031 , a data display unit 7032 and an information determination unit 7033 .
[0105] The pre-flush unit 7031 is used to process the 3D point cloud data or the bird's-eye view image to obtain vectorized data corresponding to the ground elements;
[0106] The data display unit 7032 is used to display the vectorized data corresponding to the ground elements in the bird's-eye view;
[0107] The information determining unit 7033 is configured to determine first labeling information of a ground element in response to detecting a labeling operation on the vectorized data.
[0108] In some embodiments, the information determination unit 7033 is used to delete the vectorized data in the bird's-eye view in response to detecting a deletion operation on the vectorized data; determine the first target point corresponding to the click operation in response to detecting a click operation in the bird's-eye view; and determine the first annotation information of the ground element based on the first target point corresponding to the click operation.
[0109] In some embodiments, the information determination unit 7033 is specifically used to delete the vectorized data in the bird's-eye view in response to detecting a deletion operation on the vectorized data; determine the first target point corresponding to the click operation in response to detecting a click operation in the bird's-eye view; determine the ground elements to be labeled in the bird's-eye view based on the first target point corresponding to the click operation; perform image segmentation processing on the ground elements to be labeled to obtain image segmentation results corresponding to the ground elements to be labeled; and automatically generate first labeling information of the ground elements to be labeled in the bird's-eye view based on the image segmentation results.
[0110] In some embodiments, as Figure 9 As shown, the result determination module 704 includes a height determination unit 7041 , a three-dimensional information determination unit 7042 and a labeling result determination unit 7043 .
[0111] The height determination unit 7041 is configured to determine the height information of the ground element based on the first annotation information of the ground element and the elevation map;
[0112] a three-dimensional information determining unit 7042, configured to determine second annotation information of the ground element based on the first annotation information and the height information of the ground element;
[0113] The labeling result determining unit 7043 is configured to determine the labeling result of the ground element based on the second labeling information and the time series data.
[0114] In some embodiments, when the time series data includes at least one image sequence, the annotation result determination unit 7043 is used to determine at least one image frame including ground elements based on the at least one image sequence; and project the second annotation information to the at least one image frame to obtain the annotation result of the ground elements.
[0115] In some embodiments, the annotation result determination unit 7043 is specifically used to determine at least one image frame including ground elements based on at least one image sequence; determine a target image frame from the at least one image frame, and project the second annotation information to the target image frame to determine the annotation accuracy of the second annotation information; in response to detecting an adjustment operation on the second annotation information on the user interface, determine the adjusted second annotation information; project the adjusted second annotation information to at least one image frame to obtain the annotation result of the ground element.
[0116] In some embodiments, the height determination unit 7041 is specifically used to determine at least one second target point in the first annotation information, and the coordinate position of each second target point in the bird's-eye view; based on the coordinate position of each second target point in the bird's-eye view and the elevation map, determine the elevation value corresponding to each second target point; if the elevation value meets the preset conditions, determine the height information of the ground element based on the elevation value corresponding to each second target point.
[0117] The beneficial technical effects corresponding to the exemplary embodiment of the ground element labeling device 700 can be found in the corresponding beneficial technical effects of the exemplary method section above, and will not be repeated here.
[0118] Exemplary electronic devices
[0119] Figure 10 This is a structural diagram of an electronic device 100 provided in an embodiment of the present disclosure, including at least one processor 101 and a memory 102.
[0120] The processor 101 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 100 to perform desired functions.
[0121] The memory 102 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 101 may execute the one or more computer program instructions to implement the ground element labeling method and / or other desired functions of the various embodiments of the present disclosure described above.
[0122] In one example, the electronic device 100 may further include an input device 103 and an output device 104 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0123] The input device 103 may also include, for example, a keyboard, a mouse, and the like.
[0124] The output device 104 can output various information to the outside, and may include, for example, a display, a speaker, a printer, a communication network and its connected remote output devices, etc.
[0125] Of course, to simplify, Figure 10 Only some of the components related to the present disclosure in the electronic device 100 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 100 may further include any other appropriate components according to specific application scenarios.
[0126] Exemplary computer program products and computer-readable storage media
[0127] In addition to the above-mentioned methods and devices, embodiments of the present disclosure may also provide a computer program product, including computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the ground element labeling method of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.
[0128] The computer program product may be written in any combination of one or more programming languages to implement the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0129] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the steps in the ground element labeling method of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.
[0130] Computer readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium is, for example, but not limited to, a system, device or component comprising electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0131] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be considered as essential to each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0132] Those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.
Claims
1. A method for labeling ground elements, comprising: Determining time series data corresponding to a target scene, and constructing three-dimensional point cloud data of the target scene based on the time series data; Determining, based on the three-dimensional point cloud data, a bird's-eye view and an elevation map corresponding to the three-dimensional point cloud data, and displaying the bird's-eye view on a user interface; In response to detecting a labeling operation on a ground element in the bird's-eye view, determining first labeling information of the ground element; Based on the first annotation information and the elevation map, a annotation result of the ground element is determined.
2. The method according to claim 1, wherein In response to detecting a labeling operation on a ground element in the bird's-eye view, determining first labeling information of the ground element includes: Processing the three-dimensional point cloud data or the bird's-eye view to obtain vectorized data corresponding to the ground elements; Displaying the vectorized data corresponding to the ground elements in the bird's-eye view; In response to detecting a labeling operation on the vectorized data, first labeling information of the ground element is determined.
3. The method according to claim 2, wherein: In response to detecting a labeling operation on the vectorized data, determining first labeling information of the ground element includes: In response to detecting a deletion operation on the vectorized data, deleting the vectorized data in the bird's-eye view; In response to detecting a click operation in the bird's-eye view, determining a first target point corresponding to the click operation; Based on the first target point corresponding to the click operation, first annotation information of the ground element is determined.
4. The method according to claim 3, wherein: The determining, based on the first target point corresponding to the click operation, first annotation information of the ground element includes: Determining, based on the first target point corresponding to the click operation, a ground element to be marked in the bird's-eye view; Performing image segmentation processing on the ground element to be labeled to obtain an image segmentation result corresponding to the ground element to be labeled; Based on the image segmentation result, first labeling information of the ground element to be labeled is automatically generated in the bird's-eye view.
5. The method according to claim 1, wherein The determining, based on the first annotation information and the elevation map, the annotation result of the ground element includes: Determining height information of the ground element based on the first annotation information of the ground element and the elevation map; Determining second annotation information of the ground element based on the first annotation information and the height information of the ground element; Based on the second labeling information and the time series data, a labeling result of the ground element is determined.
6. The method according to claim 5, wherein: When the time series data includes at least one image sequence, determining the labeling result of the ground element based on the second labeling information and the time series data includes: determining, based on the at least one image sequence, at least one image frame including the ground element; The second annotation information is projected onto the at least one image frame to obtain an annotation result of the ground element.
7. The method according to claim 6, wherein: The projecting the second annotation information onto the at least one image frame to obtain the annotation result of the ground element includes: Determining a target image frame from the at least one image frame, and projecting the second annotation information onto the target image frame to determine the annotation accuracy of the second annotation information; In response to detecting an adjustment operation on the second annotation information on the user interface, determining the adjusted second annotation information; The adjusted second annotation information is projected onto the at least one image frame to obtain an annotation result of the ground element.
8. The method according to claim 5, wherein The determining, based on the first annotation information of the ground element and the elevation map, the height information of the ground element includes: Determining at least one second target point in the first annotation information, and a coordinate position of each second target point in the bird's-eye view; Determining the elevation value corresponding to each second target point based on the coordinate position of each second target point in the bird's-eye view and the elevation map; If the elevation value meets a preset condition, the height information of the ground element is determined based on the elevation value corresponding to each second target point.
9. A ground element marking device, comprising: A scene construction module, configured to determine time series data corresponding to a target scene and construct three-dimensional point cloud data of the target scene based on the time series data; a data determination module, configured to determine, based on the three-dimensional point cloud data, a bird's-eye view and an elevation map corresponding to the three-dimensional point cloud data, and display the bird's-eye view on a user interface; An information determining module, configured to determine first annotation information of the ground element in response to detecting a annotation operation on the ground element in the bird's-eye view; A result determination module is used to determine the labeling result of the ground element based on the first labeling information and the elevation map.
10. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the ground element labeling method according to any one of claims 1 to 8.
11. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the ground element labeling method described in any one of claims 1 to 8.
Citation Information
Cited By
Ground identification detection method and device and medium
CN121582890A
A ground marking detection method, device and medium
CN121582890B