Road surface element reconstruction method, apparatus, device, and medium
By acquiring sensor data to determine the feature information and pose of road surface elements, a mesh reconstruction method was used to solve the problem of difficult point cloud data annotation, achieving structured storage and efficient annotation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING HORIZON INFORMATION TECH CO LTD
- Filing Date
- 2022-06-24
- Publication Date
- 2026-05-19
AI Technical Summary
In existing technologies, point cloud data is stored as individual points as basic units when reconstructing road surface elements, which leads to difficulties in annotation and a large amount of data processing.
By acquiring sensor data, the characteristic information of the target road segment and the pose of the mobile device are determined. The road surface elements are described using a grid reconstruction method. The grid is used to represent the data geometry, forming a structured storage, and the grid is directly labeled.
It reduces the amount of data processing, improves the convenience and efficiency of annotation, and is suitable for the accurate reconstruction of road surface elements.
Smart Images

Figure CN115063767B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of autonomous driving technology, and in particular to a method and system for reconstructing road elements, electronic devices, and storage media. Background Technology
[0002] In the field of autonomous driving, map building is an important factor in improving vehicle driving safety. An important step in map building is to extract semantic or instance information from road elements perceived around the vehicle (such as lane lines, zebra crossings, etc.), and then to perform 3D modeling of the road elements around the vehicle, providing a foundation for vehicle localization and trajectory planning.
[0003] Therefore, how to accurately extract road surface elements is a technical issue that deserves attention. Summary of the Invention
[0004] To address the aforementioned technical issues, the solutions in the relevant technologies have at least the following drawbacks: when reconstructing road surface elements, point cloud data is stored only as individual points as basic units, making annotation difficult.
[0005] To address the aforementioned technical problems, the present disclosure provides a solution. Embodiments of this disclosure provide a method and system for reconstructing road surface elements, an electronic device, and a storage medium.
[0006] According to one aspect of the present disclosure, a method for reconstructing road surface elements is provided, comprising: acquiring sensor data collected by a mobile device for a target road segment; determining feature information of the target road segment and the pose of the mobile device on the target road segment based on the sensor data; and performing mesh reconstruction of the road surface elements of the target road segment based on the feature information and the pose to obtain a mesh reconstruction result representing the road surface elements of the target road segment.
[0007] According to another aspect of the present disclosure, a road surface element reconstruction device is provided, comprising: a data acquisition module configured to acquire sensor data collected by a mobile device for a target road segment; a data processing module configured to determine feature information of the target road segment and the pose of the mobile device on the target road segment based on the sensor data; and a mesh reconstruction module configured to perform mesh reconstruction of the road surface elements of the target road segment based on the feature information and the pose, thereby obtaining a mesh reconstruction result representing the road surface elements of the target road segment.
[0008] According to another aspect of the present disclosure, a computer-readable storage medium is provided, the storage medium storing a computer program for performing the road element reconstruction method of any of the above embodiments of the present disclosure.
[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the road element reconstruction method of any of the above embodiments of the present disclosure.
[0010] Based on the road surface element reconstruction method, apparatus, electronic device, and storage medium provided in the above embodiments of this disclosure, by determining the feature information of the target road segment and the pose of the mobile device on the target road segment, and then performing mesh reconstruction on the road surface elements of the target road segment based on the feature information and pose, a mesh reconstruction result of the road surface elements of the target road segment can be obtained. Mesh reconstruction refers to representing the road surface elements as a mesh, and then optimizing the mesh-represented road surface elements using an objective function as a constraint to obtain the mesh reconstruction result. The mesh is generally based on a data geometry structure (e.g., it can be a triangular data format, described as: M = (V, E, F), where V represents a vertex, E represents an edge, and F represents a face, as can be referred to...). Figure 4 Using points as the basic unit, road surface elements are described, and this description method can form structured data for storage. In contrast, traditional technical solutions only store individual points in point cloud data as basic units, which does not meet the conditions for structured storage. In addition, when annotating based on the grid reconstruction results, the grid can be used as the basic unit (for example, the grid is represented as a triangular piece data format) for direct annotation, without the need to annotate each point individually as in traditional point cloud data annotation. Therefore, the data processing volume of this disclosure is relatively reduced, which is beneficial to improving the convenience of annotation.
[0011] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0012] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to offer a further understanding of the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0013] Figure 1 This is a schematic flowchart of a road surface element reconstruction method provided in an exemplary embodiment of this disclosure;
[0014] Figure 2 yes Figure 1 A flowchart illustrating an exemplary implementation of the road surface element reconstruction method provided in this embodiment;
[0015] Figure 3 yes Figure 1A flowchart illustrating another exemplary implementation of the road surface element reconstruction method provided in this embodiment;
[0016] Figure 4 It is to utilize Figure 1 A partial exemplary structural diagram of the grid reconstruction result of road elements obtained by the road element reconstruction method provided in the embodiment;
[0017] Figure 5 yes Figure 1 A flowchart illustrating another exemplary implementation of the road surface element reconstruction method provided in this embodiment;
[0018] Figure 6 yes Figure 1 A flowchart illustrating another exemplary implementation of the road surface element reconstruction method provided in this embodiment;
[0019] Figure 7 This is a schematic diagram of the structure of a road surface element reconstruction device provided in an exemplary embodiment of the present disclosure;
[0020] Figure 8 yes Figure 6 A schematic diagram of an exemplary embodiment of the road element reconstruction device provided in the embodiments;
[0021] Figure 9 yes Figure 6 A schematic diagram of another exemplary embodiment of the road surface element reconstruction device provided in the embodiment;
[0022] Figure 10 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation
[0023] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0024] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0025] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.
[0026] It should also be understood that in the embodiments disclosed herein, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0027] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.
[0028] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.
[0029] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.
[0030] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0031] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0032] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, they should be considered part of the specification.
[0033] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0034] The embodiments disclosed herein can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate together with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems, etc.
[0035] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are executed by remote processing devices linked through communication networks. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.
[0036] Application Overview
[0037] In the process of realizing this disclosure, the inventors discovered that most of the related technologies are based on the traditional Simultaneous Localization and Mapping (SLAM) method to reconstruct the ground, thereby obtaining the vehicle's pose and sparse / dense point clouds of ground elements.
[0038] The aforementioned technologies have at least the following problems: the sparse / dense point cloud data of the obtained ground elements are stored as individual points as basic units, which cannot be stored in a structured manner. When annotating the data, each point in the point cloud data needs to be annotated one by one, resulting in a large amount of data processing and making annotation difficult.
[0039] Exemplary methods
[0040] Figure 1 This is a schematic flowchart of a road surface element reconstruction method provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices, such as... Figure 1 As shown, the road surface element reconstruction method includes the following steps:
[0041] S1. Acquire sensor data collected by the mobile device for the target road segment.
[0042] This disclosure does not limit the type of mobile device. For example, the mobile device can be a vehicle or an aircraft (e.g., a quadcopter drone). When the mobile device is an aircraft, the aircraft can assist the vehicle in collecting sensor data of the target road segment.
[0043] In addition, this disclosure does not limit the type of the target road segment, which can be any road segment that needs to be reconstructed; for example, it can include road segments where urban roads, factory and mining roads, test roads, competition roads, and automobile test roads are located, and can also include, but are not limited to, road segments where traffic accidents occur frequently, road segments frequently used by users, or road segments that have been reopened after construction.
[0044] This disclosure does not limit the types of elements included in the target road segment, which can generally include road surface elements and non-road surface elements; among them, road surface elements may include, but are not limited to, lane lines, zebra crossings, curbs, etc.; non-road surface elements may include, for example, the sky, trees, buildings, vehicles, pedestrians, etc.
[0045] The device or module performing step S1 can communicate with the sensor to acquire sensor data. This disclosure does not limit the communication method; for example, wired or wireless communication can be used. Wired communication can utilize a data cable, for example; wireless communication can include, but is not limited to, using Wi-Fi, NFC (Near Field Communication), Bluetooth, etc.
[0046] S2. Based on sensor data, determine the feature information of the target road segment and the pose of the mobile device on the target road segment.
[0047] Feature information refers to the spatial or image information of the target road segment, such as semantic segmentation results and road surface element point clouds, which will be described in detail below. Additionally, pose refers to the position and orientation of the mobile device in the world coordinate system; the orientation describes the direction in which the mobile device is facing.
[0048] S3. Based on feature information and pose, perform grid reconstruction on the pavement elements of the target road segment to obtain the grid reconstruction results representing the pavement elements of the target road segment.
[0049] It should be noted that mesh reconstruction refers to representing road surface elements as a mesh, using an objective function as a constraint, and optimizing the mesh-represented road surface elements to obtain the mesh reconstruction result. The mesh representation is a data geometry format, such as a triangular data format (data format description: M = (V, E, F), where V represents vertices, E represents edges, and F represents faces, see reference...). Figure 4 ) is used as the basic unit to describe road surface elements.
[0050] The road surface element reconstruction method provided in this disclosure can reconstruct the road surface elements of a target road segment based on feature information and pose, and obtain the grid reconstruction result representing the road surface elements of the target road segment. Since the road surface elements are represented by a grid in the grid reconstruction result, the grid representation is a data geometry format, thus forming structured data storage. Compared with the traditional technical solution where point cloud data is stored using each point as a basic unit, this disclosure has the conditions for structured storage, which facilitates data annotation.
[0051] In addition, when annotating based on the grid reconstruction results, the grid as the basic unit can be directly annotated, unlike the annotation method for point cloud data in traditional technology, which requires annotating each point one by one. Therefore, the amount of data processing is relatively reduced, which is conducive to improving annotation efficiency.
[0052] In the above Figure 1 Based on the embodiments, as an optional implementation, the mobile device includes various types of sensors, specifically including at least one camera, lidar, IMU (Inertial Measurement Unit), and GNSS (Global Navigation Satellite System), etc. Accordingly, the sensor data includes inertial measurement data collected by the IMU, positioning data determined by the GNSS, and at least one of the following: road images of the target road segment collected by each camera in at least one camera based on its corresponding viewpoint, and point cloud data collected by the radar for the target road segment.
[0053] The placement of at least one camera can vary depending on the type of mobile device. For example, if the mobile device is a vehicle, at least one camera can be a vehicle-mounted camera. If there are multiple vehicle-mounted cameras, they can be positioned at the front, rear, or side of the vehicle, where they can capture the target road segment. If the mobile device is an aircraft, at least one camera can be a camera mounted on the aircraft. The aircraft-mounted camera can be positioned at the side and / or lower part of the aircraft, where it is advantageous to capture the target road segment.
[0054] This disclosure does not limit the number of at least one camera, and can be adapted to the load capacity of the mobile device. For example, if the mobile device is a vehicle, there can be multiple cameras; if the mobile device is an aircraft, there can be one camera.
[0055] Furthermore, this disclosure does not limit the type of at least one camera. For example, the at least one camera may be a plurality of surround-view cameras.
[0056] Understandably, taking a mobile device as an example, a vehicle can have multiple surround-view cameras distributed at different locations within the vehicle. Each camera has a field of view, and the combined field of view captures the environment surrounding the vehicle. As the vehicle moves along the target road segment, for each moving position (i.e., a pose of the vehicle), the multiple cameras can capture a set of road images from the corresponding field of view. This set of road images corresponds to a section of the target road segment, and the edge regions between multiple sets of road images can partially overlap (to ensure the continuity of the road images). Thus, as the vehicle moves along the target road segment, the multiple cameras can capture road images of a continuous area of the target road segment.
[0057] Based on sensor data, the characteristic information of the target road segment and the pose of the mobile device in the target road segment can be determined, thus providing conditions for subsequent steps to perform grid reconstruction of the road surface elements of the target road segment.
[0058] In the above Figure 1 Based on the embodiments, as another optional implementation method, refer to Figure 2 Step S2, "Based on sensor data, determine the feature information of the target road segment and the pose of the mobile device in the target road segment," may include the following steps:
[0059] S21. Using a semantic segmentation model, perform semantic segmentation on each road image to obtain the semantic segmentation results.
[0060] It should be noted that the above-mentioned road images are acquired by each of the at least one camera based on its corresponding viewpoint for the target road segment. This disclosure does not limit the type of semantic segmentation model. For example, the semantic segmentation model can be an image semantic segmentation model based on Convolutional Neural Networks (CNN), specifically including but not limited to image semantic segmentation models such as FCN, SegNet, U-Net, PSPNet, and DeepLab.
[0061] The semantic segmentation results include classification information representing the categories of each element in the road image; the element categories can be found in [reference needed]. Figure 1 The embodiments are described in detail here and will not be repeated.
[0062] In addition, semantic segmentation models require pre-training. The pre-training process mainly includes the following steps: 1) Constructing a training dataset. This includes multiple pairs of road images. Each pair includes a road sample image and its annotation result; the annotation result is obtained by manually labeling the element categories in the road sample image. 2) Constructing an initialization model. For example, initializing the parameters of the input layer, hidden layer, and output layer of the neural network in the model. 3) Executing model training. For example, inputting the road sample images from the training dataset into the initialization model to obtain the corresponding output predicted classification results of the road sample images, then calculating the difference between the predicted classification results and the annotation results of the road sample images, and then optimizing the parameters of the initialization model based on the difference. After iterative optimization to meet the preset conditions, the trained semantic segmentation model can be obtained.
[0063] S22. Based on sensor data, establish the point cloud of road surface elements of the target road segment using Simultaneous Localization and Mapping (SLAM) and determine the pose of the mobile device.
[0064] Here, step S22 can be implemented in different ways depending on the data selected from the sensor data:
[0065] For example, in an optional example, the data selected from the sensor data includes: case (1) selecting road images, inertial measurement data and positioning data of the mobile device collected from the perspective of each camera in at least one camera for the target road segment; case (2) selecting radar point cloud data of the target road segment, inertial measurement data and positioning data of the mobile device from the sensor data; and case (3) selecting road images, radar point cloud data of the target road segment, inertial measurement data and positioning data of the mobile device collected from the perspective of each camera in at least one camera for the target road segment.
[0066] For case (1), visual SLAM can be used to establish the point cloud of road elements in the target road segment and determine the pose of the mobile device; for case (2), laser SLAM can be used to establish the point cloud of road elements in the target road segment and determine the pose of the mobile device; for case (3), multimodal fusion SLAM can be used to establish the point cloud of road elements in the target road segment and determine the pose of the mobile device. Since the above visual SLAM, laser SLAM and multimodal fusion SLAM methods can all be implemented using any applicable method in the relevant technologies, the specific implementation process of each method will not be described here to save space.
[0067] In addition, feature information may include semantic segmentation results and point clouds of road surface elements.
[0068] As described above, through the implementation of step S2, semantic segmentation results, road surface element point clouds, and mobile device poses can be obtained, thereby providing data support for subsequent steps to perform grid reconstruction of road surface elements of the target road segment.
[0069] In the above Figure 1 Based on the embodiments and implementation methods, as another optional implementation method, refer to... Figure 3 Step S3, "Based on feature information and pose, perform mesh reconstruction of the road surface elements of the target road segment," may include the following steps:
[0070] S31. Based on the semantic segmentation results, non-road elements in each road image are filtered to obtain each road surface image and each road surface label image corresponding to each road image. In this step, each road image collected by each camera in at least one camera from its corresponding viewpoint for the target road segment is filtered for non-road elements. The resulting road surface image includes road surface elements and their pixel values. The road surface label image is an image obtained by classifying the road surface elements in the road surface image.
[0071] The "filtering" process can be achieved by using a masking algorithm to occlude non-road elements in each frame of the road image and labeling them as background, thus obtaining a road image that only includes road surface elements. Furthermore, the road surface elements in the road surface image are categorized to obtain corresponding road surface label images.
[0072] For information on pavement elements and non-pavement elements, please refer to [link / reference]. Figure 1 The descriptions in the embodiments will not be repeated here.
[0073] S32. Create an initial ground grid. The initial ground grid is used to represent the pavement elements of the target road segment.
[0074] It should be noted that the ground grid (i.e., Mesh) can be built based on a spatial rectangular coordinate system. The data format of the Mesh can be described as: M = (V, E, F), where V represents vertices, E represents edges, and F represents faces.
[0075] Reference Figure 4This is an exemplary local structural diagram of the mesh reconstruction result of road surface elements, wherein the local region includes three faces (i.e., triangles) F1, F2, and F3; each face includes three vertices, for example, F1 includes vertices V1, V2, and V4, F2 includes vertices V2, V3, and V4, and F3 includes vertices V1, V3, and V4; the lines connecting the vertices form edges, for example, the line connecting vertices V1 and V2 forms edge E1, the line connecting vertices V2 and V3 forms edge E2, and the line connecting vertices V1 and V3 forms edge E3.
[0076] Furthermore, this disclosure does not limit the method of creating the initial grid.
[0077] For example, during the initialization phase, the height, pixel value, and category probability value of each point in the ground grid can be set to zero; the category probability value is the probability value corresponding to the category label, which is used to characterize the probability that a road element or a point in the initial ground grid belongs to a certain category.
[0078] For example, firstly, based on a preset bias, the coordinates of each point in the road element point cloud are transformed to the coordinate system of the initialized ground grid; secondly, the correspondence between each point in the road element point cloud and each point in the initialized ground grid is determined; finally, based on the correspondence, the height value of each point in the road element point cloud is used as the height value of the corresponding point in the initialized ground grid. The preset bias is determined based on the relationship between the coordinate system of the initialized ground grid and the pose coordinates of the mobile device.
[0079] As mentioned above, the mesh reconstruction results of the road surface elements of the target road segment contain topological and neighborhood information between each vertex. Compared with the pure point cloud data in the existing technology, it is conducive to the structured storage of data, and is more suitable for rendering and easier to label.
[0080] The reason why the above mesh reconstruction results are easy to annotate is that the number of triangles in the mesh reconstruction results is much smaller than the number of points in the pure point cloud, thus reducing the workload of annotation.
[0081] S33. Based on the pose, initialize the ground grid, each road surface image, and each road surface label image, perform grid reconstruction on the road surface elements of the target road segment.
[0082] In step S33, "grid reconstruction of the pavement elements of the target road segment" can be performed in any available manner. For example, as an optional example, refer to... Figure 5 Step S33 may include the following steps:
[0083] S331. Align the pose with the origin of the initial ground grid coordinate system, and determine the corrected extrinsic parameters of each camera in at least one camera relative to the initial ground grid coordinate system.
[0084] This disclosure does not limit the implementation method of step S331.
[0085] For example, this can be implemented as follows: First, based on the pose coordinates, determine the movement trajectory of the mobile device on the target road segment and the center point of the movement trajectory; second, based on the center point, normalize the movement trajectory towards the origin of the pose coordinate system to obtain the normalized movement trajectory; wherein, the normalized movement trajectory is distributed around the origin of the pose coordinate system; third, move the normalized movement trajectory according to a preset offset to align it with the origin of the initial ground grid coordinate system; wherein, the preset offset is determined based on the relationship between the initial ground grid coordinate system and the pose coordinate system.
[0086] For ease of understanding, assume that the movement trajectory is a straight line from point (100, 100) to point (500, 500) in the pose coordinate system. Then the normalization process can be described as follows: First, determine the center point of the movement trajectory as (300, 300), and subtract 300 from the x-coordinate and y-coordinate of each point in the entire movement trajectory, so that the normalized movement trajectory is distributed around the origin of the pose coordinate system.
[0087] In addition, the rotation matrix and / or translation matrix between the coordinate system of the initialized ground grid and the coordinate system of the pose can be pre-calculated as a preset offset.
[0088] "Alignment" means that a preset offset can be superimposed on the coordinates of the pose to transform the pose from the coordinate system of the pose to the coordinate system of the initialized ground grid.
[0089] In addition, the corrected extrinsic parameters can be obtained by superimposing a preset bias amount on the camera's calibration extrinsic parameters.
[0090] It is understandable that by aligning the pose of the mobile device with the origin of the initial ground grid coordinate system and determining the corrected extrinsic parameters of each camera relative to the initial ground grid coordinate system, it is beneficial to unify subsequent calculations under the initial ground grid coordinate system and ensure the validity of the results.
[0091] S332. Using the corrected extrinsic parameters of each camera, the initialized ground grid is projected onto the road surface images from the perspective of each camera, thus obtaining the corresponding projected grid images and projected grid label images.
[0092] It should be explained that the modified extrinsic parameters can be expressed as the transformation relationship between the camera coordinate system and the coordinate system of the initialized ground grid, such as the rotation matrix and / or translation matrix between the two.
[0093] Based on this, step S332 can be implemented as follows: the coordinates of each point in the initialized ground grid are rotated and / or translated according to the corrected extrinsic parameters of each camera to obtain the corresponding projected grid images and projected grid label images from the perspective of each camera.
[0094] It should be noted that here, projection refers to field-of-view transformation. The projected grid image is the image from the perspective of the camera corresponding to the initialized ground grid, used for comparison with the road surface image from the perspective of the corresponding camera. Furthermore, by categorizing the elements in the projected grid image (e.g., road surface elements), a corresponding projected grid label image can be obtained.
[0095] S333: Based on the method of minimizing the difference between each projected grid image and each road surface image from the perspective of each camera, and the method of minimizing the difference between each projected grid label image and each road surface label image, the initial ground grid is optimized.
[0096] Here, step S333 can be implemented in any available manner. For example, as an optional example, refer to Figure 6 Step S333 may include the following steps:
[0097] S3331. Based on each projected grid image and the corresponding road surface image, construct a photometric loss function, and based on each projected grid label image and the corresponding road surface label image, construct a cross-entropy loss function.
[0098] It should be noted that the photometric loss function is used to estimate the pixel value difference between a selected region in the projected grid image and the corresponding region in the road surface image. The cross-entropy loss function is used to estimate the difference in class probability values between a selected region in the projected grid label image and the corresponding region in the road surface label image. The class probability value is used to characterize the probability that a point in the road surface element or the initial ground grid belongs to a certain class.
[0099] S3332. Randomly adjust the pose and / or initialize the height, pixel value, and category probability value of each point in the ground grid. Based on the adjusted projected grid images and each road surface image, estimate the change in the photometric loss function value. Based on the adjusted projected grid label images and each road surface label image, estimate the change in the cross-entropy loss function value until the photometric loss function and the cross-entropy loss function converge.
[0100] Optionally, step S3332 can be implemented in the following way:
[0101] First, following the direction that converges the photometric loss function and the cross-entropy loss function, iteratively adjust the height, pixel value, and class probability value of each point in the initial ground grid, or iteratively adjust the pose and the height, pixel value, and class probability value of each point in the initial ground grid, to determine the estimated values of the photometric loss function and the cross-entropy loss function for each iteration. Second, calculate the first difference between the estimated value of the photometric loss function in the current iteration and the estimated value of the photometric loss function in the previous iteration, and the second difference between the estimated value of the cross-entropy loss function in the current iteration and the estimated value of the cross-entropy loss function in the previous iteration. If both the first and second differences are less than the corresponding preset error thresholds, then it is determined that the photometric loss function and the cross-entropy loss function have converged in the current iteration.
[0102] It should be noted that the corresponding iteration range can be determined based on the range of adjusted parameters. For example, if the adjusted parameters are the height, pixel value, and class probability value of each point in the initial ground grid, then the iteration range is steps S332 to S333. As another example, if the adjusted parameters are the height, pixel value, class probability value, and pose of each point in the initial ground grid, then the iteration range is steps S331 to S333. This iteration range includes the purpose of step S331: since the pose of the mobile device determined by Simultaneous Localization and Mapping (SLAM) may have deviations, after each pose adjustment iteration, it is necessary to start the iteration from step S331, "that is, align the pose with the origin of the coordinate system of the initial ground grid," to improve the accuracy of road element reconstruction.
[0103] In addition, starting from the second iteration, the initialized ground grid becomes the ground grid for the intermediate process of the corresponding iteration. Correspondingly, the height value, pixel value, class probability value, and pose of each point change from the initial value to the intermediate value.
[0104] Optionally, after each iteration, before performing the step of calculating the first difference and the second difference, the gradients of the photometric loss function and the cross-entropy loss function can be calculated first; under the condition of backpropagation of the gradients of the photometric loss function and the cross-entropy loss function, the height value, pixel value, class probability value and / or pose of each point in the initialized ground grid are updated.
[0105] Here, it needs to be explained that gradient backpropagation refers to the gradient descent of the photometric loss function and the cross-entropy loss function, which means that the rate of change of the function decreases, indicating that the function gradually approaches convergence.
[0106] It is understandable that iteratively optimizing the height, pixel value, and category probability value of each point in the initial ground grid can gradually make the initial ground grid approximate the real road surface image.
[0107] In addition, since the pose of the mobile device determined by Simultaneous Localization and Mapping (SLAM) may be inaccurate, the pose can be corrected by iteratively optimizing the pose. Thus, during the mesh reconstruction process, by aligning the pose with the origin of the coordinate system of the initialized ground mesh, the accuracy of subsequent steps in reconstructing road elements can be improved.
[0108] S3333: In response to the convergence of the photometric loss function and the cross-entropy loss function, the height value, pixel value, class probability value and / or pose of each point in the initial ground grid determined when the photometric loss function and the cross-entropy loss function converge, are used as the optimal parameters for grid reconstruction of the road surface elements of the target road segment.
[0109] The road surface element reconstruction method of this disclosure describes road surface elements using a mesh. Utilizing a loss function as a constraint, the method gradually optimizes the mesh representation of road surface elements by adjusting the height, pixel value, class probability value, and / or pose of each point in the initial ground mesh, making it progressively closer to the real road surface image, thereby obtaining the mesh reconstruction result of the road surface elements for the target road segment. As described above, since the mesh reconstruction result uses triangular pieces as basic units to represent road surface elements, it can form structured data storage and is easy to annotate. Furthermore, the pose can be optimized during the mesh reconstruction process, thereby improving the accuracy of road surface element reconstruction.
[0110] Any of the road surface element reconstruction methods provided in this disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, any of the road surface element reconstruction methods provided in this disclosure can be executed by a processor, such as a processor executing any of the image processing methods mentioned in this disclosure by calling corresponding instructions stored in memory. Further details will not be elaborated below.
[0111] Exemplary device
[0112] It should be understood that the road element reconstruction method described in the foregoing embodiments herein can also be similarly applied to the following road element reconstruction apparatus for similar extensions; for the sake of simplicity, it is not described in detail.
[0113] Figure 7 This is a schematic diagram of the structure of a road element reconstruction device provided in an exemplary embodiment of this disclosure. Figure 7As shown, the road surface element reconstruction device may include: a data acquisition module 610, configured to acquire sensor data collected by sensors on a mobile device for a target road segment; a data processing module 620, configured to determine the feature information of the target road segment and the pose of the mobile device on the target road segment based on the sensor data; and a mesh reconstruction module 630, configured to perform mesh reconstruction on the road surface elements of the target road segment based on the feature information and pose, to obtain the mesh reconstruction result representing the road surface elements of the target road segment.
[0114] In an optional example, the sensor data includes at least one of road images of the target road segment acquired from the respective viewpoints of at least one camera in the field of view and radar point cloud data acquired by radar of the target road segment, as well as inertial measurement data and positioning data of the mobile device.
[0115] In an optional example, refer to Figure 8 The data processing module 620 includes: a semantic segmentation submodule 6210, configured to: use a semantic segmentation model to perform semantic segmentation on road images collected from the target road segment at the corresponding viewpoints of at least one camera, and obtain semantic segmentation results; and a localization and mapping submodule 6220, configured to: establish a road surface element point cloud of the target road segment based on sensor data using a real-time localization and map building method, and determine the pose of the mobile device; wherein, the feature information includes the semantic segmentation results and the road surface element point cloud.
[0116] In an optional example, refer to Figure 9 The mesh reconstruction module 630 includes: a filtering submodule 6310, configured to: filter non-road elements in each road image based on semantic segmentation results to obtain road surface images and surface label images corresponding to each road image; wherein, the road surface image includes road surface elements and pixel values of the road surface elements; the road surface label image is an image obtained by classifying the road surface elements in the road surface image; a ground mesh creation submodule 6320, configured to: create an initial ground mesh, which is used to represent the road surface elements of the target road segment; and a mesh reconstruction execution submodule 6330, configured to: perform mesh reconstruction on the road surface elements of the target road segment based on pose, the initial ground mesh, each road surface image, and each road surface label image.
[0117] In an optional example, the mesh reconstruction execution submodule 6330 is further configured to: align the pose with the origin of the coordinate system of the initialized ground mesh, determine the corrected extrinsic parameters of each camera relative to the coordinate system of the initialized ground mesh; use the corrected extrinsic parameters of each camera to project the initialized ground mesh onto the road surface image from the viewpoint of each corresponding camera, thereby obtaining each projected mesh image and the projected mesh label image; and optimize the initialized ground mesh based on the method of minimizing the difference between each projected mesh image and the corresponding road surface image, and the method of minimizing the difference between each projected mesh label image and the corresponding road surface label image.
[0118] In an optional example, the mesh reconstruction execution submodule 6330 is further configured to: construct a photometric loss function based on each projected mesh image and its corresponding road surface image, and construct a cross-entropy loss function based on each projected mesh label image and its corresponding road surface label image; wherein, the photometric loss function is used to estimate the pixel value difference between a selected region in the projected mesh image and the corresponding region in the road surface image; the cross-entropy loss function is used to estimate the class probability value difference between a selected region in the projected mesh label image and the corresponding region in the road surface label image; the class probability value is used to characterize the probability that a road surface element or a point in the initialized ground mesh belongs to a certain class; and randomly adjust the pose. And / or initialize the height, pixel value, and class probability value of each point in the ground grid. Based on the adjusted projected grid image and road surface image, estimate the change in the photometric loss function value, and based on the adjusted projected grid label image and the corresponding road surface label image, estimate the change in the cross-entropy loss function value, until the photometric loss function and cross-entropy loss function converge. In response to the convergence of the photometric loss function and cross-entropy loss function, use the height, pixel value, class probability value, and / or pose of each point in the initialized ground grid determined when the photometric loss function and cross-entropy loss function converge as the optimal parameters for grid reconstruction of the road surface elements of the target road segment.
[0119] In an optional example, the mesh reconstruction execution submodule 6330 is further configured to: determine the movement trajectory of the mobile device on the target road segment and the center point of the movement trajectory based on the pose coordinates; normalize the movement trajectory to the origin of the pose coordinate system based on the center point to obtain a normalized movement trajectory; wherein the normalized movement trajectory is distributed around the origin of the pose coordinate system; and move the normalized movement trajectory to align with the origin of the initial ground mesh coordinate system according to a preset offset; wherein the preset offset is determined based on the relationship between the initial ground mesh coordinate system and the pose coordinate system.
[0120] In an optional example, the mesh reconstruction execution submodule 6330 is further configured to: initialize the coordinates of each point in the ground mesh, rotate and / or translate them according to the corrected extrinsic parameters of each camera, and obtain each projected mesh image and the corresponding projected mesh label image.
[0121] In an optional example, the mesh reconstruction execution submodule 6330 is further configured to: iteratively adjust the height, pixel value, and class probability value of each point in the initialized ground mesh, or iteratively adjust the pose and the height, pixel value, and class probability value of each point in the initialized ground mesh, in the direction that makes the photometric loss function and the cross-entropy loss function converge; determine the estimated values of the photometric loss function and the cross-entropy loss function corresponding to each iteration; calculate the first difference between the estimated value of the photometric loss function in the current iteration and the estimated value of the photometric loss function in the previous iteration, and the second difference between the estimated value of the cross-entropy loss function in the current iteration and the estimated value of the cross-entropy loss function in the previous iteration; if both the first difference and the second difference are less than the corresponding preset error thresholds, then determine that the photometric loss function and the cross-entropy loss function have converged in the current iteration; and / or, after each iteration, calculate the gradient of the photometric loss function and the cross-entropy loss function; and, under the condition of backpropagation of the gradient of the photometric loss function and the cross-entropy loss function, update the height, pixel value, class probability value, and / or pose of each point in the initialized ground mesh.
[0122] In an optional example, the ground grid creation submodule 6320 is configured to: transform the coordinates of each point in the road element point cloud to the coordinate system of the initialized ground grid according to a preset offset; determine the correspondence between each point in the road element point cloud and each point in the initialized ground grid; and use the height value of each point in the road element point cloud as the height value of the corresponding point in the initialized ground grid according to the correspondence.
[0123] The road surface element reconstruction apparatus of this disclosure describes road surface elements using a mesh. Utilizing a loss function as a constraint, it gradually optimizes the mesh representation of road surface elements by adjusting the height, pixel value, class probability value, and / or pose of each point in the initial ground mesh, making it progressively closer to the real road surface image, thereby obtaining the mesh reconstruction result of the road surface elements for the target road segment. As described above, since the mesh reconstruction result uses triangular pieces as basic units to represent road surface elements, it can form structured data that is easy to annotate. Furthermore, the pose can be optimized during the mesh reconstruction process, thereby improving the accuracy of road surface element reconstruction.
[0124] Exemplary electronic devices
[0125] Below, for reference Figure 10This describes an electronic device according to embodiments of the present disclosure. The electronic device may be either or both of a first device and a second device, or a standalone device independent of them, which may communicate with the first device and the second device to receive acquired input signals from them.
[0126] Figure 10 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0127] like Figure 10 As shown, the electronic device includes one or more processors 101 and memory 102.
[0128] The processor 101 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0129] The memory 102 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 101 may execute the program instructions to implement the road element reconstruction methods of the various embodiments of this disclosure described above and / or other desired functions. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.
[0130] In one example, the electronic device may also include an input device 103 and an output device 104, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0131] For example, when the electronic device is a first device or a second device, the input device 103 can be the microphone or microphone array described above, used to capture the input signal from the sound source. When the electronic device is a standalone device, the input device 103 can be a communication network connector, used to receive the acquired input signals from the first device 100 and the second device 200.
[0132] In addition, the input device 103 may also include, for example, a keyboard, a mouse, etc.
[0133] The output device 104 can output various information to the outside, including determined distance information, direction information, etc. The output device 104 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0134] Of course, for the sake of simplicity, Figure 10 Only some of the components of the electronic device relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.
[0135] Exemplary computer program products and computer-readable storage media
[0136] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps in the road element reconstruction methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.
[0137] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0138] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the road element reconstruction methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section above.
[0139] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0140] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0141] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0142] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0143] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the method is for illustrative purposes only, and the steps of the method of this disclosure are not limited to the order specifically described above, unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the method according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the method according to this disclosure.
[0144] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0145] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0146] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A method for reconstructing road surface elements, wherein, include: Acquire sensor data collected by a mobile device for a target road segment, the sensor data including road images; Based on the sensor data, the feature information of the target road segment and the pose of the mobile device in the target road segment are determined. The feature information includes the semantic segmentation result obtained by semantic segmentation of the road image. Based on the feature information and the pose, the road surface elements of the target road segment are reconstructed by mesh to obtain the mesh reconstruction result representing the road surface elements of the target road segment. The mesh reconstruction refers to representing the road surface elements by mesh, using an objective function as a constraint, optimizing the mesh-represented road surface elements to obtain the mesh reconstruction result. The mesh representation is in a data geometry format. The step of reconstructing the road surface elements of the target road segment based on the feature information and the pose includes: Based on the semantic segmentation results, non-road elements in each road image are filtered to obtain a road surface image and a road surface label image corresponding to each road image; wherein, the road surface image includes the road surface elements and the pixel values of the road surface elements; the road surface label image is an image obtained by classifying the road surface elements in the road surface image; Create an initial ground grid, which is used to characterize the pavement elements of the target road segment; Based on the pose, the initialized ground grid, each of the road surface images, and each of the road surface label images, the road surface elements of the target road segment are reconstructed using a grid.
2. The road surface element reconstruction method according to claim 1, wherein, The sensor data includes road images of the target road segment collected by each of the cameras in at least one of the cameras of the mobile device from their respective corresponding viewpoints, point cloud data collected by the radar of the mobile device for the target road segment, and inertial measurement data and positioning data of the mobile device.
3. The road surface element reconstruction method according to claim 2, wherein, The feature information includes the semantic segmentation result and the road surface element point cloud. The step of determining the feature information of the target road segment and the pose of the mobile device on the target road segment based on the sensor data includes: Using a semantic segmentation model, semantic segmentation is performed on each of the road images to obtain the semantic segmentation results; Based on the sensor data, a point cloud of road surface elements for the target road segment is established using real-time positioning and mapping, and the pose of the mobile device is determined.
4. The road surface element reconstruction method according to claim 3, wherein, The process of reconstructing the road surface elements of the target road segment based on the pose, the initialized ground grid, each of the road surface images, and each of the road surface label images includes: Align the pose with the origin of the coordinate system of the initialized ground grid, and determine the corrected extrinsic parameters of each of the cameras relative to the coordinate system of the initialized ground grid; Using the corrected extrinsic parameters of each camera, the initialized ground grid is projected onto the road surface image at the corresponding viewpoint of each camera, thereby obtaining each projected grid image and each projected grid label image. The initial ground grid is optimized based on the method of minimizing the difference between each of the projected grid images and the corresponding road surface images, and the method of minimizing the difference between each of the projected grid label images and the corresponding road surface label images.
5. The road surface element reconstruction method according to claim 4, wherein, The optimization of the initialized ground grid based on the method of minimizing the difference between each of the projected grid images and the corresponding road surface images, and the method of minimizing the difference between each of the projected grid label images and the corresponding road surface label images, includes: Based on each of the projected grid images and the corresponding road surface images, a photometric loss function is constructed, and based on each of the projected grid label images and the corresponding road surface label images, a cross-entropy loss function is constructed. The photometric loss function is used to estimate the pixel value difference between a selected region in the projected grid image and the corresponding region in the road surface image. The cross-entropy loss function is used to estimate the class probability value difference between a selected region in the projected grid label image and the corresponding region in the road surface label image. The class probability value represents the probability that a road surface element or a point in the initialized ground grid belongs to a certain class. Randomly adjust the pose and / or the height, pixel value, and category probability value of each point in the initialized ground grid. Based on each adjusted projected grid image and each road surface image, estimate the change in the photometric loss function value. Also, based on each adjusted projected grid label image and each road surface label image, estimate the change in the cross-entropy loss function value until the photometric loss function and the cross-entropy loss function converge. In response to the convergence of the photometric loss function and the cross-entropy loss function, the height value, pixel value, class probability value and / or pose of each point in the initial ground grid determined when the photometric loss function and the cross-entropy loss function converge are used as the optimal parameters for grid reconstruction of the road surface elements of the target road segment.
6. The road surface element reconstruction method according to claim 4, wherein, Aligning the pose with the origin of the coordinate system of the initialized ground grid includes: Based on the coordinates of the pose, determine the movement trajectory of the mobile device on the target road segment and the center point of the movement trajectory; Based on the center point, the movement trajectory is normalized to the origin of the coordinate system of the pose to obtain the normalized movement trajectory; wherein, the normalized movement trajectory is distributed around the origin of the coordinate system of the pose. The normalized trajectory is moved to align with the origin of the coordinate system of the initialized ground grid according to a preset offset; wherein the preset offset is determined based on the relationship between the coordinate system of the initialized ground grid and the coordinate system of the pose.
7. The road surface element reconstruction method according to claim 4, wherein, The step of projecting the initialized ground grid onto the road surface image from the viewpoint of each camera using the corrected extrinsic parameters of each camera, thereby obtaining each projected grid image and each projected grid label image, includes: The coordinates of each point in the initialized ground grid are rotated and / or translated according to the corrected extrinsic parameters of each camera to obtain the corresponding projected grid images and projected grid label images.
8. The road surface element reconstruction method according to claim 5, wherein, The random adjustment of the pose and / or the height, pixel value, and category probability value of each point in the initialized ground grid, estimating the change in the photometric loss function value based on the adjusted projected grid images and the corresponding road surface images, and estimating the change in the cross-entropy loss function value based on the adjusted projected grid label images and the corresponding road surface label images, until the photometric loss function and the cross-entropy loss function converge, includes: Following the direction that makes the photometric loss function and the cross-entropy loss function converge, iteratively adjust the height value, pixel value, and class probability value of each point in the initial ground grid, or iteratively adjust the pose and the height value, pixel value, and class probability value of each point in the initial ground grid, and determine the estimated value of the photometric loss function and the estimated value of the cross-entropy loss function corresponding to each iteration; Calculate the first difference between the estimated value of the photometric loss function in the current iteration and the estimated value of the photometric loss function in the previous iteration, and the second difference between the estimated value of the cross-entropy loss function in the current iteration and the estimated value of the cross-entropy loss function in the previous iteration. If both the first difference and the second difference are less than the corresponding preset error threshold, then it is determined that the photometric loss function and the cross-entropy loss function have converged in the current iteration. And / or, After each iteration, the gradients of the photometric loss function and the cross-entropy loss function are calculated; Under the gradient backpropagation of the photometric loss function and the cross-entropy loss function, update the height value, pixel value, class probability value and / or pose of each point in the initialized ground grid.
9. The road surface element reconstruction method according to claim 6, wherein, The creation of the initial ground grid also includes: Based on the preset offset, the coordinates of each point in the road surface element point cloud are transformed to the coordinate system of the initialized ground grid; Determine the correspondence between each point in the road surface element point cloud and each point in the initialized ground grid; Based on the correspondence, the height value of each point in the road surface element point cloud is used as the height value of the corresponding point in the initialized ground grid.
10. A road surface element reconstruction device, wherein, include: The data acquisition module is configured to acquire sensor data collected by the mobile device for the target road segment, the sensor data including road images; The data processing module is configured to: determine the feature information of the target road segment and the pose of the mobile device in the target road segment based on the sensor data, wherein the feature information includes the semantic segmentation result obtained by semantic segmentation of the road image; The mesh reconstruction module is configured to: perform mesh reconstruction on the road surface elements of the target road segment based on the feature information and the pose, and obtain a mesh reconstruction result representing the road surface elements of the target road segment. The mesh reconstruction refers to representing the road surface elements with a mesh, optimizing the mesh-represented road surface elements with an objective function as a constraint, and obtaining the mesh reconstruction result. The mesh representation is in a data geometry format. The grid reconstruction module includes: The filtering submodule is configured to filter non-road elements in each of the road images based on the semantic segmentation results, to obtain a road surface image and a road surface label image corresponding to each road image; wherein, the road surface image includes the road surface elements and the pixel values of the road surface elements; the road surface label image is an image obtained by classifying the road surface elements in the road surface image; The ground grid creation submodule is configured to create an initial ground grid, which is used to characterize the pavement elements of the target road segment; The mesh reconstruction execution submodule is configured to perform mesh reconstruction on the road surface elements of the target road segment based on the pose, the initialized ground mesh, each of the road surface images, and each of the road surface label images.
11. A computer-readable storage medium storing a computer program for performing the road element reconstruction method according to any one of claims 1-9.
12. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the road element reconstruction method according to any one of claims 1-9.