Static-element labeling method and apparatus, and electronic device and storage medium

By using three-dimensional point cloud data to determine the position of the acquisition device and aligning the image data, efficient and accurate labeling of static elements is achieved, solving the problems of complex and low accuracy of road surface reconstruction in the prior art.

WO2025129577A1PCT designated stage expired Publication Date: 2025-06-26YINWANG INTELLIGENT TECHNOLOGIES CO LTD

Patent Information

Application Number
PCT/CN2023/140726
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In the prior art, the pavement reconstruction method has complex processes and low accuracy, making it difficult to effectively identify static elements on the pavement.

Method used

The position of the acquisition device is determined through the three-dimensional point cloud data collected by the acquisition device, and the image data collected multiple times is aligned based on the position, so as to realize the labeling of static elements.

Benefits of technology

This reduces the complexity of image data alignment, improves the efficiency and accuracy of static feature labeling, and solves the problems of complex and low accuracy of road surface reconstruction processes in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023140726_26062025_PF_FP_ABST
    Figure CN2023140726_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a static-element labeling method and apparatus, and an electronic device and a storage medium, which solve the problem in the prior art of the process of a pavement reconstruction method being complex, and the accuracy of same being relatively low. The method comprises: on the basis of three-dimensional point cloud data, which is collected by a collection device, of a target region, determining a specified pose of the collection device for collecting the three-dimensional point cloud data and image data; and on the basis of the pose and the image data collected by the collection device, obtaining a static-element labeling result in which a road geometry of the target region and the specific positions and types of static elements are labeled. The alignment of image data collected at different times is realized by means of reusing a pose acquired on the basis of three-dimensional point cloud data, thereby reducing the complexity of aligning the image data collected at different times, and improving the efficiency of labeling static elements. Moreover, the pose acquired on the basis of the three-dimensional point cloud data makes a pose result more accurate, such that an alignment result is also more accurate, thereby improving the accuracy of labeled static elements.
Need to check novelty before this filing date? Find Prior Art

Description

Static element annotation method, device, electronic device and storage medium Technical Field

[0001] The present application relates to the field of data processing, and more particularly to a static element annotation method, device, electronic device, and storage medium. Background Art

[0002] Autonomous driving technology often requires the assistance of visual information. During the autonomous driving process, the vehicle needs to use the video captured by the vehicle's cameras and, through perception algorithms, identify static elements on the road. These static elements can then be used to make autonomous driving decisions. Static elements include lane markings, road signs, and zebra crossings. Traditional autonomous driving perception algorithms often use images from a perspective view (PV). PV images often make it difficult to determine the distance between objects, and large, close-range targets in multiple images are often not effectively identified. Therefore, visual perception using a bird's-eye view (BEV) is gradually becoming the mainstream visual perception framework for next-generation autonomous driving.

[0003] In the existing technology, static elements on the road are generally identified through artificial intelligence models. In order for vehicles to be able to identify static elements through artificial intelligence models, the artificial intelligence model needs to be trained first through training data marked with static elements.

[0004] A current mainstream method for acquiring training data labeled with static features involves using a vehicle to capture road surface image data and then reconstructing the road surface in 3D. This results in highly accurate 3D spatial information of the road surface with semantic categories. This 3D spatial information is then used to label static features. The semantic categories define the static features on the road surface. However, existing road reconstruction methods still suffer from complex processes and low accuracy.

[0005] Summary of the Invention

[0006] The embodiments of the present application provide a static feature annotation method, device, electronic device and storage medium, which solve the problems of complex processes and low accuracy of road reconstruction methods in the prior art.

[0007] According to the first aspect of the embodiment of the present application, a method for labeling static elements in three-dimensional modeling is provided, in which the posture of the acquisition device in collecting three-dimensional point cloud data and image data is determined by acquiring three-dimensional point cloud data of the target area, and the static element labeling results marked with the road geometry of the target area and the specific position and type of each static element are obtained through the posture and the image data collected by the acquisition device.

[0008] In this application, based on the three-dimensional point cloud data of the target area collected, the posture of the collection device when collecting the three-dimensional point cloud data is determined, and based on the posture, the image data of the target area collected multiple times by the collection device are aligned, so that the geometric shape of the road surface and the various static elements on the road surface are determined through the aligned image data. The alignment of the image data collected multiple times is achieved through the posture obtained by the three-dimensional point cloud data. Compared with the prior art that only uses image data to align the image data collected multiple times, the method of this application reduces the complexity of aligning the image data collected multiple times and improves the efficiency of annotating static elements. In addition, the posture result obtained by the three-dimensional point cloud data is more accurate, which makes the alignment result more accurate and improves the accuracy of the annotated static elements.

[0009] In an optional embodiment, the density of the three-dimensional point cloud data is less than a preset threshold. In the case where the density of the three-dimensional point cloud data is less than the preset threshold, the solution of the present application can still use the three-dimensional point cloud data with lower density to align the posture, and can still ensure the accuracy and efficiency of the aligned posture. It can be seen that in the case where the density of the three-dimensional point cloud data is less than the preset threshold in the present application, the present application can use a lower-cost sensor to collect the three-dimensional point cloud data, solving the problem of high cost in the prior art.

[0010] In an optional embodiment, the present application also generates first three-dimensional spatial information of the target area including semantic category data based on the three-dimensional point cloud data. The process of generating the static feature annotation result includes generating second three-dimensional spatial information of the target area including semantic category data based on the pose image data and the three-dimensional point cloud data, and fusing the first three-dimensional spatial information and the second three-dimensional spatial information to obtain the target three-dimensional spatial information used to generate the static feature annotation result. This fully utilizes the advantages of multi-source data and makes the generated target three-dimensional spatial information more accurate.

[0011] In an optional embodiment, the process of generating the static feature annotation result includes: generating second three-dimensional spatial information of the target area for generating the static feature annotation result based on the posture, image data and three-dimensional point cloud data.

[0012] In an optional embodiment, the image data may be pre-processed to obtain a semantic recognition result indicating the type of static elements included in the image data. The semantic recognition result may assist in the generation of three-dimensional spatial information in subsequent steps.

[0013] In an optional embodiment, the process of generating the second three-dimensional spatial information includes: dividing the target area into multiple sub-areas; generating three-dimensional spatial information of the target area based on the three-dimensional modeling model of each sub-area in the multiple sub-areas, and using the three-dimensional spatial information of the target area as the second three-dimensional spatial information; wherein the three-dimensional modeling model takes the first coordinate value and the second coordinate value of the grid point included in the sub-area as input, and takes the third coordinate value and semantic category data of the grid point as output, and supervises the output of the three-dimensional modeling model through the semantic recognition results of the three-dimensional point cloud data and the image data to train the three-dimensional modeling model, and the semantic recognition results are used to indicate the types of static elements included in the image data. By realizing three-dimensional reconstruction through an artificial intelligence model, the efficiency and quality of three-dimensional reconstruction can be improved. In this way, when the artificial intelligence model is limited and the reconstruction area is large, the partition-based reconstruction method can improve the accuracy of the generated second three-dimensional spatial information.

[0014] In an optional embodiment, the method further includes: determining multiple paths for data acquisition by the acquisition device based on the pose; for each of the multiple sub-areas, using an area within a preset range of each of the multiple paths as a reconstruction area corresponding to the sub-area; and training a 3D model for each of the multiple sub-areas based on the reconstruction area corresponding to the sub-area. This ensures the reliability of the training data and improves the accuracy of the 3D spatial information.

[0015] In an alternative embodiment, when training a 3D model for a sub-region, the 3D model is iteratively trained based on multiple iteration regions selected from the reconstructed region. This way, during each iteration, only the data for one iteration region needs to be loaded, rather than the entire data for the entire reconstructed region / sub-region, ensuring training efficiency.

[0016] In an optional embodiment, the process of generating the first three-dimensional spatial information of the target area includes: generating the first three-dimensional spatial information based on semantic recognition results of the three-dimensional point cloud data and the image data. The semantic recognition results of the image data can ensure that the first three-dimensional spatial information carries accurate semantic category data.

[0017] In one optional embodiment, the process of fusing the first and second 3D spatial information to generate target 3D spatial information includes: matching and aligning the first and second 3D spatial information, and superimposing the aligned first and second 3D spatial information to obtain the target 3D spatial information. By fusing the first and second 3D spatial information by matching and aligning first and second 3D spatial information and then superimposing them, the fused result, i.e., the target 3D spatial information, can be more accurate.

[0018] In an optional embodiment, the process of matching and aligning the first and second three-dimensional spatial information includes: converting the first and second three-dimensional spatial information into bird's-eye-view images, respectively; and matching and aligning the bird's-eye-view images corresponding to the first and second three-dimensional spatial information. Compared to matching three-dimensional spatial information by converting the three-dimensional spatial information into bird's-eye-view images and then aligning the bird's-eye-view image sets, matching and alignment is more efficient. Moreover, the offset between the first and second three-dimensional spatial information is primarily in the horizontal direction. This allows for faster matching while ensuring matching accuracy, thereby improving the efficiency of generating target three-dimensional spatial information.

[0019] In an optional implementation, training data for the target area can be generated based on the static feature annotation results of the target area. This training data is used to train an artificial intelligence model for identifying static road features. This method can quickly obtain accurate training data and reduce the cost of acquiring training data.

[0020] In an optional embodiment, a high-precision map of the target area can be generated based on the static element annotation results of the target area, and the target area includes the road area. Obtaining a high-precision map by the above method can reduce the cost of obtaining a high-precision map and improve the speed of obtaining a high-precision map.

[0021] According to the second aspect of the embodiment of the present application, an electronic device is provided, including a memory and a processor, the memory is used to store a computer program, and the processor is used to execute the computer program to implement the method of the first aspect of the embodiment of the present application or any implementation method of the first aspect.

[0022] According to the third aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by at least one processor, the computer implements the method of the first aspect of the embodiments of the present application or any implementation method of the first aspect.

[0023] According to the fourth aspect of the embodiment of the present application, a static feature labeling device is provided, including: a determination module for determining the posture of the acquisition device when collecting each frame of three-dimensional point cloud data in the three-dimensional point cloud data based on the three-dimensional point cloud data obtained by the acquisition device through multiple collections of the target area; a generation module for generating a static feature labeling result of the target area based on at least the posture and the image data obtained by the acquisition device through multiple collections of the target area, the static feature labeling result being used to indicate the road geometry of the target area and the position and type of the static feature.

[0024] It should be understood that the technical solutions provided in the above-mentioned second, third and fourth aspects and their technical features can all correspond to the methods provided in the first aspect and its optional implementation methods, so the beneficial effects that can be achieved are similar and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] FIG1 is a simplified schematic diagram of a system architecture provided in an embodiment of the present application;

[0026] FIG2 is a simplified schematic diagram of a collection device provided in an embodiment of the present application;

[0027] FIG3 is a simplified schematic diagram of a server provided in an embodiment of the present application;

[0028] FIG4 is a flow chart of a static element annotation method provided in an embodiment of the present application;

[0029] FIG5 is a flow chart of a method for generating training data provided in an embodiment of the present application;

[0030] FIG6 is a flowchart of a static element annotation method provided in an embodiment of the present application;

[0031] FIG7 is a schematic diagram of a projection provided by an embodiment of the present application;

[0032] FIG8 is a schematic diagram of a method for fusing first three-dimensional spatial information and second three-dimensional spatial information provided by an embodiment of the present application;

[0033] FIG9 is a flowchart of a method for generating three-dimensional space information provided in an embodiment of the present application;

[0034] FIG10 is a schematic diagram of a reconstruction area provided in an embodiment of the present application;

[0035] FIG11 is a schematic diagram of a model training method provided in an embodiment of the present application;

[0036] FIG12 is a block diagram of a static feature labeling device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0037] In autonomous driving technology, vehicles are required to automatically decide routes and handle accidents. In order to achieve these functions, vehicles need the assistance of visual information. Vehicles need to use images or videos taken by cameras on the vehicle and use perception algorithms to identify static elements on the road. For example, vehicles need to identify lane lines and signs on the road to ensure that the vehicle can drive in the appropriate lane when making autonomous driving decisions.

[0038] In traditional autonomous driving technology, the images or videos recognized by perception algorithms are often PV images captured directly by cameras. PV images tend to appear larger when near objects are larger than when far away. However, as two-dimensional images, PV images are not easy to determine the distance between objects. Furthermore, lane markings cover a large area, so the perception results from a single image often cannot contain all the required perception information. Lane markings in different images, and lane markings recognized in PV images, are often not usable by other processing processes.

[0039] To address the aforementioned issues with PV images, existing technologies are gradually beginning to use BEVs for visual perception. In existing technologies, artificial intelligence models are generally used to input PV images and output recognition results of static road features of the BEV centered on the vehicle itself. To train the aforementioned artificial intelligence models, supervised training is required. To achieve supervised training, ground truth (GT) is required to supervise the training of the model. Generally, the BEV's road static feature annotation results can be used as GT to supervise the training of the aforementioned models. The training data generated later in this article is also called GT.

[0040] The mainstream method in existing technologies is to obtain high-precision road geometry and semantic categories through three-dimensional road reconstruction, thereby realizing the automatic labeling of static road elements under BEV. Among them, three-dimensional reconstruction is a method of restoring the three-dimensional spatial information of a scene based on sensory information such as laser and vision. In addition, in order to complete data labeling, the above-mentioned three-dimensional reconstruction is also a semantic three-dimensional reconstruction method, that is, based on the three-dimensional reconstruction, a semantic category is assigned to each part of the scene. For example, semantic three-dimensional reconstruction can label each part as a lane line, a sign, or an ordinary road surface.

[0041] Road surface 3D reconstruction is typically performed using data collected by a collection device on the road surface. The mainstream method in the prior art is to complete large-scale 3D reconstruction of the road surface by collecting multiple passes of data. Compared to collecting only single-pass data, collecting multiple passes can solve the problem that a single pass cannot complete a complete 3D reconstruction of the road surface. For example, at an intersection, a single pass of data may only be able to reconstruct a portion of the road surface in 3D, but not the entire road surface. With multi-pass data, during data collection, different collection devices or the same device may pass over the same road area multiple times. For example, at an intersection, one pass captures a left turn, another a right turn, and another a straight ahead. Each of these passes covers the same intersection area, but the motion trajectories of the different passes are different. This makes large-scale 3D reconstruction more convenient. Furthermore, the 3D reconstruction results completed using multiple passes can be stored. If a new collection device collects new data corresponding to the reconstruction result, the existing multi-pass reconstruction results can be directly used to annotate the newly collected results.

[0042] It can be seen that the mainstream method in the existing technology is to complete the three-dimensional reconstruction of the road surface through multiple trips of data, and rely on the data of the three-dimensional reconstruction of the road surface to realize automatic labeling, so as to automatically provide training data for the above-mentioned artificial intelligence model used to identify static elements of the road surface under BEV.

[0043] Next, we will introduce a conventional method for 3D road surface reconstruction. This method performs 3D reconstruction based solely on visual information. Multiple passes of road surface image data are acquired through the camera of a capture device. Three-dimensional reconstruction is performed for each pass of road surface image data, generating 3D spatial information of the road surface. Based on the position of the capture device determined from the road surface image data, the corresponding 3D spatial information of the multiple passes is aligned, resulting in a large-scale 3D spatial information of the road surface. This 3D spatial information can also be considered a 3D model. This 3D spatial information can then be further vectorized to obtain the desired vectorized road surface static feature annotation results.

[0044] This method suffers from the following issues: Because it's impossible to accurately align multiple passes of image or video data using visual information alone, it first performs a 3D reconstruction of the road surface based on a single pass of image or video data, then performs alignment and fusion based on the 3D spatial information corresponding to multiple passes. However, this approach complicates the process of aligning the 3D spatial information corresponding to multiple passes. Furthermore, the reliability of the results of 3D road surface reconstruction using only visual data is low. Furthermore, because this two-stage process involves errors in the single-pass reconstruction process, errors in the subsequent multi-pass alignment can affect the quality of the final reconstruction.

[0045] It can be seen that the method shown in the prior art has the problems of complex process and low accuracy of road surface three-dimensional reconstruction.

[0046] The above-mentioned solutions in the prior art are mainly due to the relatively complex process of alignment, resulting in low implementation efficiency. To solve this problem, this application uses the three-dimensional point cloud data collected by the acquisition device to align the posture of the acquisition device when collecting the three-dimensional point cloud data. The aligned posture can then be used to realize the subsequent three-dimensional reconstruction process of the image data. This alignment method also makes the alignment result more accurate. The error in the alignment result will not affect the subsequent reconstruction work. This solves the problem of the prior art solution of first rebuilding and then aligning, where the error caused by reconstruction will affect the alignment.

[0047] This application proposes a static feature labeling method, which determines the posture of the acquisition device when collecting three-dimensional point cloud data and image data through the three-dimensional point cloud data of the target area collected by the acquisition device, and obtains the static feature labeling results that are marked with the road geometry of the target area and the specific position and type of each static feature through the posture and the image data collected by the acquisition device.

[0048] In this application, based on the three-dimensional point cloud data of the target area collected, the posture of the collection device when collecting the three-dimensional point cloud data is determined, and based on the posture, the image data of the target area collected multiple times by the collection device are aligned, so that the geometric shape of the road surface and the various static elements on the road surface are determined through the aligned image data. The alignment of the image data collected multiple times is achieved through the posture obtained by the three-dimensional point cloud data. Compared with the prior art that only uses image data to align the image data collected multiple times, the method of this application reduces the complexity of aligning the image data collected multiple times and improves the efficiency of annotating static elements. In addition, the posture result obtained by the three-dimensional point cloud data is more accurate, which makes the alignment result more accurate and improves the accuracy of the annotated static elements.

[0049] The implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0050] FIG1 is a simplified schematic diagram of a system architecture to which embodiments of the present application can be applied. As shown in FIG1 , the system architecture may include: at least one acquisition device 11 and a server 12 .

[0051] The acquisition device 11 is a removable device used to collect 3D point cloud data and image data of the target area. The acquisition device 11 can be a vehicle, a robot, an aircraft, or the like, and this application does not limit the specific form of the acquisition device 11. In the case of 3D road surface reconstruction, as shown in Figure 1, the acquisition device 11 can be a vehicle.

[0052] Figure 2 is a simplified schematic diagram of an acquisition device according to an embodiment of the present application. As shown in Figure 2, the acquisition device 11 is equipped with a first sensor 111 and a second sensor 112. The first sensor 111 is used to acquire 3D point cloud data, and the second sensor 112 is used to acquire image data.

[0053] The first sensor 111 can be any 3D scanner for collecting 3D point cloud data, such as a laser radar, millimeter-wave radar, structured light sensor, ultrasonic sensor, etc. In the present application, when the density of the 3D point cloud data is less than a preset threshold, the first sensor 111 can be a low-cost sensor that collects 3D point cloud data with a density less than the preset threshold, such as a low-line-count laser radar, millimeter-wave radar, etc.

[0054] The second sensor 112 may be a sensor for collecting image or video data, such as a camera.

[0055] The first sensor 111 and the second sensor 112 can be deployed separately on the acquisition device 11, or the two sensors can be integrated into the same sensor device, which can realize the acquisition of images and three-dimensional point cloud data.

[0056] The acquisition device 11 can collect data while moving. For example, the acquisition device 11 can be in motion all the time, and the first sensor 111 and the second sensor 112 on the acquisition device 11 can collect data at preset time intervals. Then, during the movement, the acquisition device 11 can obtain multiple frames of three-dimensional point cloud data, as well as multiple images or multiple video frames. The multiple images or multiple video frames can be collectively regarded as image data.

[0057] In addition to the first sensor 111 and the second sensor 112, the acquisition device 11 may also be equipped with other sensors, such as a third sensor for obtaining the position and posture of the acquisition device 11 during its travel. The position and posture acquired by the third sensor will serve as the initial posture. In subsequent processes, the initial posture can be corrected using the three-dimensional point cloud data to obtain the precise posture of the acquisition device 11 when acquiring each frame of three-dimensional point cloud data.

[0058] Exemplarily, the third sensor may include a sensor for collecting position data and a sensor for collecting attitude data. The sensor for collecting position data may specifically be a positioning sensor such as the Beidou navigation satellite system (BDS) and the global positioning system (GPS), and the sensor for collecting attitude data may be a sensor such as an inertial measurement unit (IMU).

[0059] The data acquired in this application is multiple passes, for example, data on multiple paths can be acquired through multiple passes of data collection. To collect data on multiple paths at a low cost, a single collection device 11 can collect data on multiple paths separately. To improve collection efficiency, multiple collection devices 11 can also collect data on multiple paths simultaneously. The number of collection devices 11 can be the same as the number of paths, or it can be less than the number of paths. This application does not specifically limit the number of collection devices 11.

[0060] The collection device 11 can also communicate with the server 12 in a wired or wireless manner to transmit the data collected by the collection device 11 to the server 12, and the server 12 obtains the static feature annotation results.

[0061] Optionally, the system used in the embodiment of the present application may also only include the acquisition device 11, and the acquisition device 11 performs data acquisition and annotation to obtain static element annotation results.

[0062] Server 12 is primarily used for annotating static elements. Figure 3 is a simplified schematic diagram of a server according to an embodiment of the present application. As shown in Figure 3, server 12 may include a processor 121 and a memory 122. Memory 122 is used to store computer programs, and processor 121 is used to execute computer programs to implement the methods described herein.

[0063] The processor 121 is the control center of the server 12. For example, the processor 121 can be any one or a combination of multiple types of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), and an application-specific integrated circuit (ASIC).

[0064] The memory 122 may include a hard disk, memory / video memory (also known as graphics card memory), and cache. The hard disk is the primary storage medium for the server 12. The server can use the hard disk to store the 3D point cloud data and image data acquired from the acquisition device 11, as well as the computer program used to implement the method of the present application. The cache is primarily used to improve the read and write performance of the system.

[0065] It should be noted that the system architecture shown in Figures 2 and 3 does not represent a limitation on the present application. The collection device 11 and server 12 in the present application may not only include the hardware shown in Figures 2 and 3, but also include other hardware for implementing the operation of the collection device 11 and the server 12. For example, when the collection device 11 is a collection vehicle, the collection device 11 may also include an engine and a processor, etc., and the server 12 may also include a communication module for communication, etc.

[0066] After briefly describing the architecture of the system to which the method of the present application is applied, an exemplary embodiment of the present application will be described in conjunction with FIG4 . As shown in FIG4 , the present application illustrates a static feature annotation method, including the following steps:

[0067] Step 401 : Based on the three-dimensional point cloud data acquired by the acquisition device through multiple acquisitions of the target area, the posture of the acquisition device when acquiring each frame of the three-dimensional point cloud data is determined.

[0068] The method of the present application determines a relatively accurate posture by collecting three-dimensional point cloud data collected by the acquisition device, and then uses the accurate posture to perform data alignment in subsequent steps.

[0069] The method of the present application can be executed by an acquisition device or by an electronic device with computing capabilities such as a server. The present application does not limit the execution entity of the three-dimensional modeling method.

[0070] The target area is the area where static feature annotation results need to be generated. As for the specific form of the target area, in some embodiments, the collection device can be a collection vehicle, and the target area can be a road area. When the target area is a road area, the static feature annotation method shown in this application is used to annotate various static features on the road, such as marking the positions of lane lines, signs, etc. on the road. In addition, the target area can also be any area that requires static feature annotation. For example, the target area can be a village, a town, etc. The static feature annotation method shown in this application is used to annotate the static features of a village or town, such as annotating the types of buildings in the village or town.

[0071] For the description of the acquisition device, please refer to the description of the acquisition device 11 in Figure 2 above, which will not be repeated here. 3D point cloud data refers to a set of vectors in a three-dimensional coordinate system. 3D point cloud data is recorded in the form of points, each point contains three-dimensional coordinates, and can carry other information about the attributes of the point, such as color, reflectivity, intensity, etc. In this application, the three-dimensional spatial information mentioned later may also include a three-dimensional point cloud, which includes a number of points, each point contains a three-dimensional coordinate, and each point also carries semantic category data.

[0072] 3D point cloud data consists of several frames of 3D point cloud data. As mentioned above, 3D point cloud data is actually collected at preset intervals, and the acquisition device has a corresponding position and posture when collecting each frame of 3D point cloud data. The position of the acquisition device refers to the position and attitude of the acquisition device. The position is used to determine the position of the acquisition device in the 3D world when collecting data, and the attitude is used to indicate the pitch angle, yaw angle, and roll angle of the acquisition device when collecting data. In addition, the data indicating attitude may be different for different acquisition devices. For example, for a collection vehicle, the attitude can be represented by the heading (also understood as the yaw angle).

[0073] Exemplarily, step 401 can be implemented by a simultaneous localization and mapping (SLAM) method in the prior art, thereby obtaining an aligned posture when the acquisition device collects three-dimensional point cloud data of each frame of the target area multiple times. The specific implementation method is not described in detail here.

[0074] Step 402 : generating static feature annotation results for the target area based on at least the posture and image data of the target area acquired by the acquisition device multiple times.

[0075] The static feature annotation results are used to indicate the road geometry of the target area and the location and type of static features. Both the 3D point cloud data and image data are obtained by collecting data from the target area multiple times.

[0076] After obtaining the pose of the acquisition device when capturing each frame of 3D point cloud data in step 401, this pose can be used to align the image data, thereby resolving the complex alignment process in existing technologies. Furthermore, the accurate pose will not introduce errors into the subsequent annotation of static features, thus ensuring the accuracy of the annotated static features.

[0077] As mentioned above, in addition to collecting three-dimensional point cloud data, the acquisition device can also collect image data. It is easy to understand that the three-dimensional point cloud data and image data can be collected by the acquisition device simultaneously and in parallel. Then the posture of the acquisition device when collecting each frame of three-dimensional point cloud data is also the posture of the acquisition device when collecting each frame of image data. In addition, in order to make the determined posture of the acquisition device when collecting each frame of image data more accurate, the first sensor and the second sensor with the same sampling frequency can be selected, or the sampling frequency of the first sensor can be selected to be much greater than the sampling frequency of the second sensor, so as to ensure that the posture of the acquisition device when collecting each frame of image data can be directly determined based on the posture of the acquisition device when collecting each frame of three-dimensional point cloud data. As mentioned above, the first sensor is a sensor for collecting three-dimensional point cloud data, and the second sensor is a sensor for collecting image data.

[0078] Furthermore, both image data and 3D point cloud data are multi-pass data, meaning they are acquired by the acquisition device over multiple passes. For each pass, the acquisition device can select multiple different paths for acquisition, or it can select similar or identical paths for acquisition. Regarding the relationships between different paths, there can be overlap, but no two paths are exactly the same. Thus, by sampling multiple paths with the acquisition device, relatively complete data about the target area can be obtained, thereby generating 3D spatial information about the target area.

[0079] It should be noted that multiple acquisitions by an acquisition device may refer to multiple acquisitions by one acquisition device, or multiple acquisition devices performing multiple acquisitions, where each acquisition device performs at least one acquisition. The number of acquisition devices is not limited in this application.

[0080] Compared with labeling static road features with single-pass data, labeling static road features with multi-pass data can improve efficiency and make the labeling results more complete.

[0081] Static feature annotation results describe the visual characteristics of the target area. Road geometry, also known as the geometric shape of the road, describes the road's shape, including the presence of forks and turns. Road geometry can be represented through bird's-eye view images or three-dimensional spatial information. Static features can be static elements on the road, such as zebra crossings, lane markings, and left-turn signs. Static feature annotation results indicate the road geometry and the location and type of static features, describing the specific visual characteristics of the road surface in the target area.

[0082] As for the specific form of the static feature annotation results, the static feature annotation results can be an image of the road surface from a bird's-eye view, or it can be three-dimensional spatial information with semantic category data. This three-dimensional spatial information is referred to as the target three-dimensional spatial information below. The semantic category data is used to annotate the type of static feature corresponding to each point in the three-dimensional spatial information.

[0083] In the case where the static feature annotation result is an image of the road surface from a bird's-eye view, step 402 can be to align the image data collected by the acquisition device according to the posture, and then convert the image data collected by the acquisition device into an image of the road surface from a bird's-eye view according to the internal and external parameters of the camera.

[0084] Next, the static element annotation method in this application will be explained by taking the method for generating target three-dimensional spatial information as an example.

[0085] The target three-dimensional spatial information can also be understood as the three-dimensional point cloud data of the target area. Each point in the three-dimensional point cloud data carries semantic category data, and the semantic category data is used to indicate the semantic category of the point. From another perspective, the semantic category data can also be used to indicate the type of static element corresponding to the grid point in the three-dimensional spatial information. The grid point refers to the point in the three-dimensional spatial information, and the static element is the static object in the three-dimensional spatial information. For example, in the scene of three-dimensional reconstruction of the road surface, the semantic data of the corresponding three-dimensional spatial information can be used to indicate whether each point is an ordinary road surface, a lane line, a sign, etc.

[0086] In addition, the “target” of the target three-dimensional spatial information is used to distinguish it from other three-dimensional spatial information, and does not limit the form of the three-dimensional spatial information.

[0087] Regarding the specific implementation of step 402, step 402 can be implemented as steps 603 and 604 of the following embodiment. That is, firstly, firstly, the first 3D spatial information is generated by the 3D point cloud data, and second 3D spatial information is generated by the image data, and then the two 3D spatial information are fused to generate the target 3D spatial information.

[0088] Alternatively, target 3D spatial information can be generated directly from image data. Specifically, step 402 includes generating second 3D spatial information of the target area based on the pose, image data, and 3D point cloud data. This second 3D spatial information is used to generate static feature annotation results. For details on how this is generated, see the description of step 603 below. Generating static feature annotation results from the second 3D spatial information can generate static feature annotation results in the form of an image, or it can be used as a static feature annotation result.

[0089] Furthermore, in the two aforementioned implementations of step 402, the method for generating the second three-dimensional spatial information using image data may not employ the implementation described in step 603 below, but may employ other implementations. For example, a method similar to that employed in the prior art may be employed, whereby the three-dimensional spatial information corresponding to each single-pass image data is first generated using the single-pass image data, and then the position determined in step 401 is reused to determine the relative relationship between the multiple three-dimensional spatial information to generate target three-dimensional spatial information for the target area.

[0090] For another example, the posture determined in step 401 may be used to align the image data collected by multiple paths, that is, in a large coordinate system, the posture of the acquisition device when each image is collected in the data of multiple paths of the target area collected by the acquisition device is determined, and then the target three-dimensional spatial information of the target area is generated based on the aligned image data.

[0091] The above example is only an example of step 402 and does not limit the present application.

[0092] Furthermore, after obtaining the target 3D spatial information in step 402, the training data can be automatically labeled based on the target 3D spatial information. Compared to the prior art that only uses visual information to perform 3D road reconstruction, the method presented in this application has improved accuracy and efficiency.

[0093] The target area in the present application can be various areas. In the case where the target area in the present application is a road area, the three-dimensional spatial information obtained by the method of the present application can have many applications, such as automatic annotation of training data, generation of high-precision maps, etc. Next, a specific embodiment will be used to illustrate the application scenario of automatically annotating training data through the three-dimensional modeling method of the present application. In some examples, the target area can not only be a road area, but the three-dimensional modeling method of the present application can also be executed alone, that is, only steps 501 to 503 in the following method are executed. The following examples do not represent limitations on the present application. The various steps in Figure 5 below can be executed by the server 12 shown in Figure 1, or by the acquisition device 11 shown in Figure 1. The present application does not limit the execution subject.

[0094] As shown in FIG5 , the method may specifically include the following steps:

[0095] Step 501: Acquire data collected by sensors.

[0096] Specifically, as shown in FIG2 , the collection device 11 may be equipped with a plurality of sensors, and the sensors on the collection device 11 may be used to collect road surface data.

[0097] In this embodiment, the first sensor may be a low-line-count laser radar, the second sensor may be a camera, and the third sensor may include an IMU and a GPS. The low-line-count laser radar may be used to collect laser three-dimensional point cloud (hereinafter referred to as a three-dimensional point cloud) data, and the camera may be used to collect image data, which may be composed of several video frames. In addition, in order to determine the posture of the camera when the image data was captured, the camera's internal and external parameters may also be obtained. The IMU and GPS may be used to obtain the position and posture of the acquisition device when collecting data, respectively.

[0098] Step 502: pre-process the data collected by the sensor.

[0099] Pre-processing refers to pre-processing the data collected by the sensor in step 501. Pre-processing can mainly include: first, pre-processing the 3D point cloud data; second, semantic segmentation and pre-labeling of the image data; and third, calibrating the camera's internal and external parameters.

[0100] Preprocessing 3D point cloud data primarily involves removing dynamic target point clouds and non-road area point clouds. Specific removal methods can be found in prior art methods and will not be detailed in this application. This application primarily generates static feature annotations for the road surface. Preprocessing the 3D point cloud data to remove unnecessary content improves the accuracy of the pose determined using the 3D point cloud data in subsequent steps, as well as the processing efficiency of subsequent steps.

[0101] Semantic segmentation and pre-labeling are performed on the image data. This involves labeling the semantic categories of various road surface components in the PV image captured by the camera. For example, this can be used to label which parts of the image are signs and which are normal road surfaces. Semantic segmentation and pre-labeling methods can be found in the prior art and will not be further described in this application. By performing semantic segmentation and pre-labeling on the image data, semantic recognition results for the image data can be obtained. The purpose of the semantic recognition results for the image data is described in detail below.

[0102] In other words, the static element labeling method of the present application further includes: preprocessing the image data to obtain a semantic recognition result of the image data, and the semantic recognition result is used to indicate the type of static elements included in the image data.

[0103] Calibrating the camera's internal and external parameters can more accurately determine the camera's posture relative to the acquisition device, thereby more accurately determining the projection relationship between the image data captured by the camera and the three-dimensional world, thereby making the semantic category data determined in subsequent steps more accurate.

[0104] Step 503: Complete the three-dimensional modeling through the data collected multiple times.

[0105] That is, 3D modeling of the road surface is performed based on the pre-processing results of step 502. The implementation of this step can be found in the description of Figure 4 above and in the description of Figure 6 below, and will not be repeated here. By performing 3D modeling in step 503, 3D spatial information of the target area can be obtained.

[0106] Step 504: perform vectorization processing on the three-dimensional space information.

[0107] Specifically, the three-dimensional spatial information obtained in step 503 is vectorized to obtain a vectorized annotation result covering the road surface of the target area. This means that each static feature of the road surface is annotated using vectors. Conventional maps are generally vector maps. Therefore, after obtaining the three-dimensional spatial information of the target area, the three-dimensional spatial information can be converted into a vector map in step 504 to facilitate automated annotation of the map in subsequent steps.

[0108] In addition, the processing result of step 504 can also be used as a high-precision map of the target area. That is, when the target area is a road area, a high-precision map can be generated by executing steps 501 to 504.

[0109] In other words, this application also generates a high-precision map of the target area based on the static element annotation results, and the target area includes the road surface area.

[0110] Step 505: back-projection annotation.

[0111] That is, the vectorized high-precision map generated in step 504 is back-projected and annotated. By relocating the data of a single pass, the vectorized high-precision map of the target area is back-projected onto the data of the single pass, and the static elements of the single pass data are automatically annotated.

[0112] Therefore, in one application scenario of this application, it can be used to annotate training data. Since training data is ultimately required, and training data is used to train an artificial intelligence model to identify static road features, it is necessary to project the vectorized high-precision map onto the single-pass data to generate training data annotated with static features.

[0113] In other words, the present application also generates training data for the target area based on the static element labeling results. The training data is a bird's-eye view image including semantic category data. The training data is used to train an artificial intelligence model for identifying static elements on the road surface.

[0114] The specific implementation of step 504 and step 505 can refer to the methods in the prior art, and this application will not go into details.

[0115] After explaining the main process of generating training data, the above step 503 will be explained in detail. As shown in FIG6 , in this embodiment, the three-dimensional modeling method includes the following steps:

[0116] Step 601: Generate first three-dimensional spatial information of a target area based at least on three-dimensional point cloud data.

[0117] The first three-dimensional spatial information includes semantic category data.

[0118] Step 602 : Based on the three-dimensional point cloud data of the target area collected by the collection device, determine the posture of the collection device when collecting each frame of the three-dimensional point cloud data.

[0119] Step 603: Generate second three-dimensional spatial information of the target area based on the posture, image data and three-dimensional point cloud data.

[0120] The second three-dimensional spatial information includes semantic category data.

[0121] Step 604: Fuse the first three-dimensional spatial information and the second three-dimensional spatial information to generate target three-dimensional spatial information.

[0122] The target 3D spatial information is used to generate static feature annotation results. For example, the target 3D spatial information can be converted into an image-based static feature annotation result, or the target 3D spatial information with semantic category data can be directly used as the static feature annotation result.

[0123] Step 603 and step 604 may together constitute step 402 shown in FIG. 4 .

[0124] In this method, two 3D spatial information sets are generated for the target area based on the 3D point cloud data and the image data, respectively. These two sets of information are then fused and used as the final output of the target 3D spatial information. This fully utilizes the advantages of multiple data sources, making the generated target 3D spatial information more accurate.

[0125] Furthermore, in an optional embodiment, the density of the three-dimensional point cloud data may be less than a preset threshold. The preset threshold is used to distinguish between sparse three-dimensional point cloud data and dense three-dimensional point cloud data. Three-dimensional point cloud data with a density less than the preset threshold is also sparse three-dimensional point cloud data. In an optional embodiment, the preset threshold can be determined based on the density of the three-dimensional point clouds that can be collected by high-line-count lidar and low-line-count lidar. The sensor that collects three-dimensional point cloud data with a density less than the preset threshold can be specifically described above for the first sensor 111 and will not be further elaborated here.

[0126] Selecting 3D point cloud data with a density below a preset threshold can address another existing technical solution. Specifically, another existing approach for 3D road reconstruction utilizes traditional high-precision map production methods. It's easy to understand that traditional high-precision map production methods also generate 3D spatial information about the road surface, so this method can be reused to annotate static road features. Specifically, this method uses data collected by a high-beam laser radar and a precise positioning module to reconstruct 3D spatial information.

[0127] This solution has the following problems: it requires the acquisition equipment to be equipped with a high-line-count lidar and a precise positioning module. The high cost of the high-line-count lidar and the precise positioning module makes it impossible to achieve large-scale three-dimensional road reconstruction at a low cost.

[0128] In the solution of the present application, the three-dimensional point cloud data collected by the laser radar or other sensors is mainly used to align the posture. It is easy to understand that the density of the three-dimensional point cloud data of the present application can be large or small. In the case that the density of the three-dimensional point cloud data is less than the preset threshold, the solution of the present application can still use the three-dimensional point cloud data with lower density to align the posture, and still ensure the accuracy and efficiency of the aligned posture. It can be seen that in the present application, when the density of the three-dimensional point cloud data is less than the preset threshold, the present application can use a lower-cost sensor to collect the three-dimensional point cloud data, while ensuring the accuracy of the alignment, and solves the problem of high cost in the existing technology. And the selection of the above-mentioned three-dimensional point cloud data combines the advantages of three-dimensional point cloud data and image data, uses the characteristics of three-dimensional point cloud data to provide accurate self-vehicle posture, reuses the posture and uses image data with dense semantics for three-dimensional modeling, thereby obtaining accurate static element annotation results.

[0129] In this way, the acquisition device can be equipped with a low-cost first sensor for collecting 3D point cloud data. For example, the first sensor can be a low-line-count LiDAR, and the acquisition device can be a mass-produced vehicle equipped with a low-line-count LiDAR. The sparse 3D point cloud data collected by the first sensor is used to align the poses of data collected from multiple paths, and first 3D spatial information is generated based on the sparse 3D point cloud data. After obtaining the pose determined by the 3D point cloud data, this pose is reused for visual 3D reconstruction to generate second 3D spatial information.

[0130] Correspondingly, in an optional embodiment, the density of the first three-dimensional spatial information obtained by three-dimensional point cloud data having a density less than a preset threshold can be less than the point cloud density of the second three-dimensional spatial information, that is, the point cloud density of the second three-dimensional spatial information can be greater than the point cloud density of the first three-dimensional spatial information. In an embodiment of the present application, the first three-dimensional spatial information is used to align image data collected multiple times, that is, to determine the positional relationship between image data collected multiple times. Therefore, the density of the first three-dimensional spatial information can be low, and in subsequent steps, the sparse part of the first three-dimensional information can be supplemented by the dense second three-dimensional spatial information.

[0131] In this implementation, accurate 3D modeling and pose are acquired from 3D point cloud data, and image data alignment is achieved using this pose. The image data provides dense semantics, supplementing the missing details of the sparse 3D point cloud data. The fusion of the two 3D spatial information fully utilizes the potential of both 3D point cloud data and image data, achieving efficient, high-quality 3D road reconstruction.

[0132] After briefly summarizing the method of FIG. 6 , the specific implementation of each step in FIG. 6 will be described next.

[0133] The semantic category data of the first three-dimensional spatial information and the second three-dimensional spatial information are the same as the semantic category data of the target three-dimensional spatial information. In other words, the semantic category information is an attribute of the three-dimensional spatial information, and is used to indicate the semantic category of the points in the three-dimensional spatial information, that is, to indicate the type of static element corresponding to the grid point in the three-dimensional spatial information.

[0134] Regarding the specific implementation of steps 601 and 602, positioning and 3D modeling can be performed based on the initial pose using existing SLAM-like methods. Specifically, the position and pose of the acquisition device at the time of data acquisition are obtained using the third sensor of the acquisition device. This position and pose is referred to as the initial pose. The initial pose is then corrected based on the relationship between each frame of 3D point cloud data to determine the pose of the acquisition device at the time of each frame of 3D point cloud data. Multiple frames of 3D point cloud data are then matched and aligned to obtain 3D spatial information.

[0135] In addition, after obtaining the three-dimensional spatial information through the SLAM method, it is also necessary to add semantic category data to the three-dimensional spatial information to generate the first three-dimensional spatial information.

[0136] The method of adding semantic category data will be described here using a specific example, which does not limit the present application. Specifically, it is possible to: determine the semantic recognition results of the image data; and generate first three-dimensional spatial information based on the semantic recognition results of the three-dimensional point cloud data and the image data.

[0137] That is, the projection relationship between the three-dimensional spatial information / three-dimensional point cloud data and the image data can be determined based on the camera's internal and external parameters, and the points in the three-dimensional spatial information can be projected into the image data. For example, point A in the three-dimensional spatial information is projected into the image data as point B in the image data. Based on the semantic recognition results of the image data, the semantic category data corresponding to point B can be determined. As shown in Figure 7, point B is a lane line, and the semantic category data corresponding to point B is used as the semantic category data of point A, that is, the semantic category data corresponding to point A is considered to be a lane line. In this way, the first three-dimensional spatial information including the semantic category data of each point can be obtained.

[0138] Regarding the specific implementation of step 604, the first 3D spatial information and the second 3D spatial information can be directly superimposed. Furthermore, since the process of generating the second 3D spatial information based on the image data depends on parameters such as camera intrinsic and extrinsic parameters, if these parameters are inaccurate, or if errors occur during the process of generating the second 3D spatial information based on the image data, the coordinates of the first 3D spatial information and the second 3D spatial information may not be completely aligned. Therefore, step 604 may also involve first matching and aligning the first 3D spatial information and the second 3D spatial information, and then superimposing the aligned first and second 3D spatial information.

[0139] In other words, step 604 may include: matching and aligning the first 3D spatial information and the second 3D spatial information, and based on the alignment result, superimposing the first 3D spatial information and the second 3D spatial information to obtain target 3D spatial information.

[0140] By fusing the first three-dimensional spatial information and the second three-dimensional spatial information in a manner of first matching and then superimposing, the fusion result, ie, the target three-dimensional spatial information, can be made more accurate.

[0141] As for the method of matching the first three-dimensional spatial information and the second three-dimensional spatial information, the matching method will be described below through a specific example. It should be noted that the following example does not limit the present application.

[0142] The first and second 3D spatial information can be converted into bird's-eye-view images, and then the two images can be aligned to determine the offset between the two images. In other words, the alignment process includes: converting the first and second 3D spatial information into bird's-eye-view images, respectively; and aligning the bird's-eye-view images corresponding to the first and second 3D spatial information.

[0143] As shown in Figure 8, the first three-dimensional spatial information can be converted into the image shown in (a) in Figure 8, and the second three-dimensional spatial information can be converted into the image shown in (b) in Figure 8. It can also be seen from Figure 8 that the first three-dimensional spatial information is sparser than the second three-dimensional spatial information. Then, by matching and superimposing the two, the target three-dimensional spatial information as shown in (c) in Figure 8 can be obtained. It should be noted that in Figure 8, for the convenience of display, the target three-dimensional spatial information is also displayed in the form of a BEV image.

[0144] Compared with matching three-dimensional spatial information, matching two-dimensional images is more efficient, and the offset between the first three-dimensional spatial information and the second three-dimensional spatial information is mainly in the horizontal direction. This can complete the matching faster while ensuring matching accuracy, thereby improving the efficiency of generating target three-dimensional spatial information.

[0145] As for specific matching methods, center-masked template matching can be used. Specifically, a template with the center portion removed can be selected. The two images are then matched based on this template, with the center portion removed. The remaining portions are then matched, an offset is determined, and the center portions of the two images are then superimposed based on this offset. This matching method ensures matching accuracy. Furthermore, the first and second 3D spatial information can be converted into multiple images, and images at different locations can be matched separately, further improving matching accuracy.

[0146] In addition, in some cases, the point clouds of the superimposed first three-dimensional spatial information and the second three-dimensional spatial information at different positions may be different. For example, in the first three-dimensional spatial information, there is a point A with coordinate values ​​of (x, y, z1), and the semantic category data corresponding to point A is ordinary road surface. The second three-dimensional spatial information corrected based on the offset has a point B with (x, y, z1), and the semantic category data corresponding to point B is a lane line. It can be seen that the heights and semantic category data of points A and B are different. In the case of the above problem, the point can be marked, and the specific height and semantic category of the point can be determined manually, or the result of the point can be corrected by filtering. This application does not limit the specific processing method for the above situation. It can be understood that the method used in this application has high accuracy and a low probability of generating the above error, that is, there are fewer points with the above problem. Even if manual error correction is performed, the efficiency of generating three-dimensional spatial information will not be reduced too much.

[0147] Regarding the specific implementation of step 603, in addition to the example of step 402 above, the second three-dimensional space information can also be generated in the following manner. It should be noted that the following manner is only an example and does not limit the present application.

[0148] As shown in FIG9 , the process of generating the second three-dimensional space information includes the following steps:

[0149] Step 901: Divide the target area into multiple sub-areas, and divide each of the multiple sub-areas into multiple grid points.

[0150] Step 902: Based on the posture, determine multiple paths for the acquisition device to collect data.

[0151] Step 903 : For each of the multiple sub-regions, an area within a preset range of each of the multiple paths is used as a reconstruction area corresponding to the sub-region.

[0152] Step 904 : Generate three-dimensional spatial information of the target area based on the three-dimensional modeling model of each sub-area in the multiple sub-areas, and use the three-dimensional spatial information of the target area as second three-dimensional spatial information.

[0153] The 3D modeling model uses the first and second coordinate values ​​of the grid points included in the sub-region as input and outputs the third coordinate values ​​of the grid points and semantic category data. The output of the 3D modeling model is supervised by the semantic recognition results of the 3D point cloud data and image data to train the 3D modeling model. The semantic recognition results are used to indicate the types of static elements included in the image data. The 3D modeling model of each of the multiple sub-regions is trained based on the reconstructed area corresponding to the sub-region.

[0154] In other words, first determine a sub-area covered by a three-dimensional model. Within this sub-area, align multiple paths through posture to determine the relationship between the multiple paths. Based on the multiple paths, determine the modeling range where the data is reliable. Within this modeling range, supervise the height and semantic category data respectively through the semantic recognition results of the three-dimensional point cloud data and image data to train a three-dimensional modeling model that can output the parameters of each point in the three-dimensional spatial information. The three-dimensional modeling model also summarizes the road surface characteristics of the sub-area. Different three-dimensional modeling models describe different road surface ranges. When the model converges, the three-dimensional spatial information of the sub-area is generated through the three-dimensional modeling model, and the three-dimensional reconstruction of the road surface is completed through the image data. If the model does not converge, re-execute step 904. By combining the three-dimensional spatial information of multiple sub-areas, the three-dimensional spatial information of the target area can be obtained.

[0155] It is easy to understand that although the above process does not complete the three-dimensional reconstruction of the road surface only through image data, image data plays a greater role in the reconstruction process than three-dimensional point cloud data and posture. Therefore, the process shown in Figure 9 can be called visual three-dimensional reconstruction.

[0156] After a general description of the above process, the specific implementation of the steps shown in FIG9 will be described next.

[0157] In step 901, the target area is divided into multiple sub-areas. This can be done by determining the size of a sub-area and dividing it accordingly. Alternatively, for multiple collected data passes and multiple intersections, an area within a certain range from each intersection can be selected as a sub-area. For example, a range of 400 meters by 400 meters from the intersection can be selected as a sub-area. Based on the determined pose, the image data and 3D point cloud data located within the sub-area can be considered training data for that sub-area.

[0158] After determining the sub-region, the sub-region can be divided into grids in the horizontal direction, for example, the grid can be divided into 1 cm * 1 cm in size, and each grid point is used as the point in the second three-dimensional spatial information for which height and semantic category data need to be determined.

[0159] In subsequent steps, the 3D model is trained based on each of the multiple sub-regions. The 3D model includes the 3D model trained for each of the multiple sub-regions. In other words, a 3D model is trained for each sub-region, and the 3D model for different sub-regions may be different.

[0160] The reason for dividing the image into multiple sub-regions in step 901 is due to the limitations of the AI ​​model's expressive capabilities. By training a separate 3D model (i.e., an AI model) for each sub-region, the image can be reconstructed using different 3D models to represent the characteristics of different locations within the reconstructed region. This improves the accuracy of the generated second 3D spatial information even when the AI ​​model is limited and the reconstruction region is large.

[0161] Of course, in other embodiments, step 901 may be omitted, that is, a three-dimensional modeling model that can describe the height and semantic category data characteristics of the target area is directly trained for the target area.

[0162] The multiple paths in step 902 are the multiple paths that the collection device travels when collecting data multiple times. The multiple paths can be the same or different.

[0163] For step 902, the precise position of the acquisition device is determined in step 602, so multiple paths collected by the acquisition device can be determined based on the precise position of the acquisition device. The paths determined in step 901 are used to determine the sub-areas in the subsequent steps.

[0164] For step 903, the area within the preset range of the distance trajectory is used as the sub-area corresponding to the target area. It can be understood that the target area and the reconstruction area can be the same or different. The target area refers to the area that can be collected by the acquisition device, but the reconstruction area is the area where the road surface is three-dimensionally reconstructed.

[0165] Regarding the specific implementation of step 903, for each path, the area within the circumscribed rectangle of the trajectory can be determined as the reconstruction area. For example, as shown in Figure 10, line 1001 represents a trajectory, and dashed box 1002 represents the corresponding reconstruction area. Alternatively, based on the pose data, a 20m x 20m range from each location can be selected as the 3D modeling area, and the 3D modeling areas corresponding to multiple points in the sub-area can be combined to form the reconstruction area.

[0166] The reason why step 903 divides the reconstruction area is that although both the image data and the three-dimensional point cloud data can capture a wider area, the image of a position far away from the acquisition device trajectory may be unreliable, so the reconstruction area needs to be divided to ensure the accuracy of the generated three-dimensional spatial information.

[0167] After the reconstruction area is determined in step 903, in the subsequent training process, the training of the three-dimensional modeling model of each sub-area among the multiple sub-areas is performed based on the reconstruction area corresponding to the sub-area.

[0168] In other embodiments, step 903 may not be performed. For example, all data of the sub-region may be used as training data for step 904 .

[0169] After explaining the generation process, the training process of the 3D modeling model will be explained next.

[0170] The specific training process may include: selecting multiple grid points in the sub-area, and looping through the following steps until convergence: for each grid point in the multiple grid points, taking the first coordinate value and the second coordinate value of the grid point as input, and obtaining the third coordinate value and semantic category data of the grid point through a three-dimensional modeling model; determining the true value of the third coordinate value of the grid point based on the three-dimensional point cloud data, and supervising the third coordinate value according to the first loss function and the true value of the third coordinate value; projecting the grid point into the image data according to the first coordinate value, the second coordinate value and the third coordinate value of the grid point, determining the true value of the semantic category data of the grid point according to the semantic recognition result of the image data, and supervising the semantic category data of the grid point according to the second loss function and the true value of the semantic category data.

[0171] In addition, in some cases, considering the limited computing power, multiple iteration areas can be selected from each reconstruction area, and multiple training can be performed through multiple iteration areas. After each iteration area has iterated itself several times, other iteration areas can be reselected for training. In this way, in each iteration process, only the data of one iteration area needs to be loaded, and there is no need to load all the data of the entire reconstruction area / sub-area, which can ensure training efficiency.

[0172] In other words, taking the example of selecting multiple iteration regions from each reconstruction region, when training the 3D model of the sub-region, the 3D model is iteratively trained based on the multiple iteration regions selected from the reconstruction region.

[0173] It should be noted that if there are iteration regions, then the selection of multiple grid points in the subregions mentioned above can be done by selecting several grid points from each iteration region. After the grid points selected in one iteration region are iterated multiple times, grid points from other iteration regions can be selected and the above training steps can be repeated.

[0174] Of course, in other embodiments, when the computing power is relatively high, it is also possible not to select an iteration region, but to directly select a number of grid points from each sub-region / reconstruction region to train the three-dimensional model.

[0175] In this method, several grid points are first selected and their two coordinate values ​​are used as input. Based on this input, the 3D model outputs the corresponding height value (also known as the third coordinate value) and semantic category data. The height value output is then supervised by the 3D point cloud data, and the semantic category data is supervised by the semantic recognition results of the image data. These steps are repeated until convergence occurs. If convergence does not occur, multiple grid points are selected in the subregion and training is repeated.

[0176] Specifically, as shown in FIG11 , step 904 includes the following steps:

[0177] 1. As shown in Figure 11, first select multiple grid points from the grid of the iteration area from a bird's-eye view. The large box corresponding to 1 in Figure 11 is a sub-area, and the small box is an iteration area. Normalize the first and second coordinate values ​​of the multiple grid points and input them into the 3D modeling model, which can be a multilayer perceptron (MLP) network. The grid size and number of selected grid points in Figure 11 do not represent the actual size and number, but are for reference only. The (x, y) in Figure 11 are the first and second coordinate values ​​of each grid point input.

[0178] The first coordinate value and the second coordinate value can be two horizontal coordinate values, such as an x-coordinate value and a y-coordinate value. Generally, in a three-dimensional world, a horizontal grid point generally corresponds to a unique height value. Using two horizontal coordinate values ​​as input can improve training efficiency.

[0179] 2. Obtain the output height value (third coordinate value) and semantic category data through the MLP network.

[0180] Next, the process of outputting height values ​​and semantic category data will be explained through a specific MLP example. It should be noted that the following example does not limit the present application.

[0181] As shown in Figure 11, the MLP network in this embodiment is a multi-head MLP structure, the input is the normalized valid grid point position, and the output is the height and semantic category of the corresponding grid point. Among them, 2.1 is an MLP structure, which contains a total of 8 layers, of which the first layer upgrades the 2-dimensional position feature to a 256-dimensional feature, and maps the original two-dimensional position to a high dimension through a high-frequency function, so that the feature carries more high-frequency information, so that the model can ensure that the output is more accurate based on these high-frequency information. Then, 7 layers of full connection are connected and the 256-dimensional feature dimension is maintained unchanged. At the same time, a linear correction unit (rectified linear unit, relu) layer is connected after each layer of full connection, and a skip connection structure is designed, that is, the input feature is directly spliced ​​with the fully connected output of the 4th layer as the input of the fully connected 5th layer. It should also be noted that the number of layers shown in Figure 11 is only for reference and does not represent the actual number of layers.

[0182] 2.2 in Figure 11 is a fully connected layer that reduces the feature dimension from 256 to 1, obtaining a height estimate for a specific grid point. 2.3 in Figure 11 uses positional encoding (PE) based on the normalized grid point positions using sin and cos functions, increasing the 2-dimensional features to 42 dimensions. 2.4 in Figure 11 is an MLP structure consisting of 17 layers. The first layer converts the 298-dimensional features obtained by the positional feature extractor (256-dimensional) and the 42-dimensional features obtained by the positional encoding into 256 dimensions. This is followed by 15 fully connected layers that maintain the 256-dimensional features, and finally by a fully connected layer that converts the 256-dimensional features into semantic category dimensions. Each fully connected layer is followed by a Relu layer, and a skip connection structure is designed. Specifically, the input features are directly concatenated with the fully connected output of the 4th layer as the input to the 5th fully connected layer; the output of the 4th layer is concatenated with the output of the 8th layer as the input to the 9th layer; and the output of the 8th layer is concatenated with the output of the 12th layer as the input to the 13th layer. Thus, specific semantic category data is obtained.

[0183] 3. Grid point height supervision. If corresponding 3D point cloud data exists for a grid point, the 3D point cloud data is projected into the coordinate system corresponding to the grid point, combined with the pose data, to determine the true height value corresponding to the grid point. The regressed grid point height is supervised using the mean squared error (MSE) loss function.

[0184] 4. Semantic category data supervision: Based on the first and second coordinate values ​​of the grid point and the height value output by the MLP network, the grid point is projected onto the image data. Based on the semantic recognition results of the image data, the true value of the semantic recognition result corresponding to the grid point is determined. The semantic recognition results are supervised using the cross entropy (CE) loss function.

[0185] In addition, when the density of three-dimensional point cloud data is less than the preset threshold, the selected grid points may not have corresponding true height values. In this case, supervision can be performed only through the semantic category true value. Since the height value output by the MLP network is also used in the semantic category supervision process, the CE function can not only supervise the semantic recognition results, but also supervise the height value to a certain extent, so that the model can be trained more accurately.

[0186] After obtaining the 3D model of the reconstructed area through the above steps, step 904 can be used to input each grid point of the reconstructed area into the 3D model of the reconstructed area to obtain the height value and semantic category data corresponding to each grid point in the reconstructed area. The 3D spatial information corresponding to the multiple reconstructed areas is combined to generate the second 3D spatial information.

[0187] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of specific implementation of the functions. It is understandable that, in order to realize the above functions, the server or acquisition device includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0188] In the embodiment of the present application, the server or acquisition device can be divided into functional modules according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.

[0189] In the case of dividing each functional module into corresponding functional modules, as shown in FIG12 , the present application further provides a static element annotation device, which includes:

[0190] A determination module 1201 is configured to determine the pose of each frame of three-dimensional point cloud data acquired by the acquisition device based on the three-dimensional point cloud data acquired by the acquisition device multiple times acquiring the target area;

[0191] The generation module 1202 is used to generate static feature annotation results of the target area based on at least the posture and image data of the target area obtained by the acquisition device multiple times. The static feature annotation results are used to indicate the road geometry of the target area and the position and type of the static features.

[0192] In an optional implementation, the density of the three-dimensional point cloud data is less than a preset threshold.

[0193] In an optional embodiment, generation module 1202 is further configured to generate first three-dimensional spatial information of the target area based at least on the three-dimensional point cloud data; the first three-dimensional spatial information includes semantic category data. When generating static feature annotation results, generation module 1202 is specifically configured to generate second three-dimensional spatial information of the target area based on the pose, image data, and three-dimensional point cloud data; the second three-dimensional spatial information includes semantic category data. The first three-dimensional spatial information and the second three-dimensional spatial information are then fused to generate target three-dimensional spatial information, which is used to generate the static feature annotation results.

[0194] In an optional implementation, the generation module 1202 is specifically configured to generate second three-dimensional spatial information of the target area based on the posture, image data, and three-dimensional point cloud data, and the second three-dimensional spatial information is used to generate static feature annotation results.

[0195] In an optional embodiment, the device further includes an acquisition module 1203 (not shown in the figure) for preprocessing the image data to obtain a semantic recognition result of the image data, where the semantic recognition result is used to indicate the type of static elements included in the image data.

[0196] In an optional embodiment, the generation module 1202 is specifically used to divide the target area into multiple sub-areas, and divide each sub-area in the multiple sub-areas into multiple grid points when generating the second three-dimensional spatial information; based on the three-dimensional modeling model of each sub-area in the multiple sub-areas, generate the three-dimensional spatial information of the target area, and use the three-dimensional spatial information of the target area as the second three-dimensional spatial information; wherein, the three-dimensional modeling model takes the first coordinate value and the second coordinate value of the grid point included in the sub-area as input, and takes the third coordinate value and semantic category data of the grid point as output, and supervises the output of the three-dimensional modeling model through the semantic recognition results of the three-dimensional point cloud data and the image data to train the three-dimensional modeling model.

[0197] In an optional embodiment, the determination module 1201 is also used to determine multiple paths when the acquisition device collects data based on the posture; for each sub-region in the multiple sub-regions, the area within a preset range of each path in the multiple paths in the sub-region is used as the reconstruction area corresponding to the sub-region; the training of the three-dimensional modeling model of each sub-region in the multiple sub-regions is performed based on the reconstruction area corresponding to the sub-region.

[0198] In an optional implementation, when performing the three-dimensional modeling model training for the sub-region, the three-dimensional modeling model is iteratively trained based on a plurality of iterative regions selected from the reconstruction region.

[0199] In an optional implementation, when generating the first three-dimensional spatial information, the generating module 1202 is specifically configured to generate the first three-dimensional spatial information based on the semantic recognition results of the three-dimensional point cloud data and the image data.

[0200] In an optional embodiment, the generation module 1202 is specifically used to convert the first three-dimensional spatial information and the second three-dimensional spatial information into images from a bird's-eye view, respectively, when generating target three-dimensional spatial information; match and align the images from the bird's-eye view corresponding to the first three-dimensional spatial information and the second three-dimensional spatial information, respectively; and based on the alignment result, superimpose the first three-dimensional spatial information and the second three-dimensional spatial information to obtain the target three-dimensional spatial information.

[0201] In an optional embodiment, the generation module 1202 is also used to generate training data for the target area based on the static element labeling results. The training data is a bird's-eye view image that includes semantic category data. The training data is used to train an artificial intelligence model for identifying static elements on the road surface.

[0202] In an optional implementation, the generation module 1202 is further configured to generate a high-precision map of the target area based on the static feature annotation results, where the target area includes a road surface area.

[0203] It should be noted that all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.

[0204] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0205] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical function division. In actual implementation, other division methods may be used. For example, multiple modules may be combined or integrated into another device, or some features may be ignored or not performed.

[0206] The present application also provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to implement the above-mentioned method. The electronic device can be a server 12 as shown in FIG1 of the present application. In addition, the electronic device can also be a collection device 11 as shown in FIG1. ​​The present application does not limit the specific form of the electronic device. Any device that can execute the above-mentioned static feature annotation method of the present application can be used as the electronic device in the present application.

[0207] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by at least one processor, the computer implements the above method.

[0208] The present application also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.

[0209] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A static element annotation method, characterized in that, The method includes: Based on the three-dimensional point cloud data obtained by the acquisition device collecting the target area multiple times, determining the pose of the acquisition device when collecting each frame of the three-dimensional point cloud data in the three-dimensional point cloud data; Generating a static element annotation result of the target area based at least on the pose and the image data obtained by the acquisition device collecting the target area multiple times, where the static element annotation result is used to indicate the road geometry of the target area and the position and type of static elements.

2. The method according to claim 1, wherein The density of the three-dimensional point cloud data is less than a preset threshold.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Generating first three-dimensional spatial information of the target area based at least on the three-dimensional point cloud data; the first three-dimensional spatial information includes semantic category data; The generating the static element annotation result of the target area based at least on the pose and the image data of the target area collected by the acquisition device includes: Generating second three-dimensional spatial information of the target area based on the pose, the image data, and the three-dimensional point cloud data; the second three-dimensional spatial information includes semantic category data; Fusing the first three-dimensional spatial information and the second three-dimensional spatial information to generate target three-dimensional spatial information, where the target three-dimensional spatial information is used to generate the static element annotation result.

4. The method according to claim 1 or 2, characterized in that, The generating the static element annotation result of the target area based at least on the pose and the image data of the target area collected by the acquisition device includes: Generating second three-dimensional spatial information of the target area based on the pose, the image data, and the three-dimensional point cloud data, where the second three-dimensional spatial information is used to generate the static element annotation result.

5. The method according to claim 3 or 4, characterized in that, The method further includes: Preprocessing the image data to obtain a semantic recognition result of the image data, where the semantic recognition result is used to indicate the type of static elements included in the image data.

6. The method according to claim 5, characterized in that The generating the second three-dimensional spatial information of the target area based on the pose, the image data, and the three-dimensional point cloud data includes: Dividing the target area into multiple sub-areas, and dividing each sub-area in the multiple sub-areas into multiple grid points; Generating second three-dimensional spatial information of the target area based on the three-dimensional modeling model of each sub-area in the multiple sub-areas; Wherein, the three-dimensional modeling model takes the first coordinate value and the second coordinate value of the grid points included in the sub-area as inputs, and takes the third coordinate value of the grid points and semantic category data as outputs, and supervises the output of the three-dimensional modeling model through the semantic recognition result of the three-dimensional point cloud data and the image data to train the three-dimensional modeling model.

7. The method according to claim 6, characterized in that, The method further includes: Based on the pose, determining multiple paths when the acquisition device collects data; For each sub-area in the multiple sub-areas, taking the area within a preset range of each path in the multiple paths within the sub-area as the reconstruction area corresponding to the sub-area; The training of the three-dimensional modeling model of each sub-area in the multiple sub-areas is performed based on the reconstruction area corresponding to the sub-area.

8. The method according to claim 7, wherein When performing the iterative training of the 3D modeling model for the sub-region, the 3D modeling model is iteratively trained based on multiple iterative regions selected from the reconstructed region.

9. The method according to claim 5, wherein The generating of the first 3D spatial information of the target region based at least on the 3D point cloud data includes: Generating the first 3D spatial information based on the 3D point cloud data and the semantic recognition result of the image data.

10. The method according to claim 3, characterized in that, The fusing of the first 3D spatial information and the second 3D spatial information to generate the target 3D spatial information includes: Converting the first 3D spatial information and the second 3D spatial information into images from a bird's-eye view respectively; Matching and aligning the images from a bird's-eye view corresponding to the first 3D spatial information and the second 3D spatial information respectively; Based on the alignment result, superimposing the first 3D spatial information and the second 3D spatial information to obtain the target 3D spatial information.

11. The method according to any one of claims 1 to 10, characterized in that, The method further includes: Generating training data for the target region based on the static element annotation result, where the training data is an image from a bird's-eye view including semantic category data, and the training data is used to train an artificial intelligence model for recognizing static elements on the road surface.

12. The method according to any one of claims 1-10, characterized in that, The method further includes: Generating a high-precision map for the target region based on the static element annotation result, where the target region includes a road surface region.

13. A static element labeling device, characterized in that, The device includes: A determination module, configured to determine the pose of the acquisition device when acquiring each frame of 3D point cloud data in the 3D point cloud data acquired by the acquisition device for the target region multiple times based on the 3D point cloud data acquired by the acquisition device; A generation module, configured to generate a static element annotation result for the target region based at least on the pose and the image data acquired by the acquisition device for the target region multiple times, where the static element annotation result is used to indicate the road geometry of the target region and the position and type of static elements.

14. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory is used to store a computer program, and the processor is used to execute the computer program to implement the static element annotation method according to any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by at least one processor, the computer is enabled to implement the static element annotation method according to any one of claims 1-12.

Citation Information

Patent Citations

  • High-precision map construction method and device, electronic equipment and storage medium

    CN113034566A

  • Method and device for constructing semantic map

    CN114440856A

  • Bird-eye view generation method and device of driving scene, equipment and storage medium

    CN114898313A

  • High-precision map generation method and device, electronic equipment and medium

    CN115841552A

  • Method and system using lidar and camera to enhance depth information about image feature point

    WO2021025364A1

Cited By

  • Multi-dimensional map generation method and device based on multi-sensor tight coupling SLAM

    CN121113032A

  • Information alignment labeling method and system for vector data and panoramic data, terminal and medium

    CN121527173A