Semantic understanding-based automatic driving high-precision map construction method and device
By generating an initial point cloud map from multi-source perception data and pose data, and performing semantic segmentation and feature recognition, the problem of insufficient accuracy and reliability in the construction of high-precision maps for autonomous driving is solved. This achieves accurate parsing and consistency of high-precision maps, thereby improving the safety and reliability of autonomous driving systems.
Patent Information
- Application Number
- CN202511894604.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-02-27
AI Technical Summary
Existing high-precision maps for autonomous driving suffer from insufficient accuracy and reliability, particularly in areas such as lane boundary misalignment, inconsistent semantic annotations, and incomplete topological relationships.
The system acquires multi-source sensing data and pose data through a data acquisition unit, generates an initial point cloud map using LiDAR, image acquisition devices, and positioning devices, performs semantic segmentation and feature recognition, constructs a high-precision map with semantic information, and optimizes the map data through semantic relationships to achieve incremental updates and directional optimization.
It improves the accuracy and reliability of high-precision maps, ensures accurate parsing and consistency of map data, and enhances the safety and reliability of autonomous driving systems.
Smart Images

Figure CN121577016A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of high-precision maps, in particular to an automatic driving high-precision map construction method and device based on semantic understanding. BACKGROUND
[0002] With the rapid development of automatic driving technology, vehicles need to obtain high-precision environmental information and reliable positioning data during driving to realize safe and stable automatic driving functions. Traditional electronic navigation maps usually only have meter-level precision and only provide basic information such as road center lines and road types, which cannot meet the needs of automatic driving systems for centimeter-level positioning accuracy, lane-level geometric structures and semantic information. Therefore, high-precision maps have become important basic data for automatic driving systems.
[0003] In the existing automatic driving high-precision map construction process, multiple collection trajectories need to be spliced, cross-period data needs to be fused, and collection results need to be uniformly modeled. However, due to differences in processing procedures, problems such as lane boundary misplacement, inconsistent semantic labeling, and incomplete topological relationships often occur, and manual verification and labeling in the existing process are also prone to introduce human errors, further reducing the accuracy and stability of the map.
[0004] The existing automatic driving high-precision map construction has the technical problem of insufficient accuracy and reliability. SUMMARY
[0005] The purpose of the present application is to provide an automatic driving high-precision map construction method and device based on semantic understanding, to solve the technical problem of insufficient accuracy and reliability in the existing automatic driving high-precision map construction.
[0006] In view of the above problems, the present application provides an automatic driving high-precision map construction method and device based on semantic understanding.
[0007] The first aspect of the present application provides an automatic driving high-precision map construction method based on semantic understanding, which comprises: collecting multi-source perception data and pose data along a target road by a data collection unit, the data collection unit comprising at least one image collection device, a laser radar collection device and a positioning device; based on the pose data and the multi-source perception data, generating an initial point cloud map of the target road by simultaneous localization; performing semantic segmentation and feature recognition on the initial point cloud map, vectorizing static road features including at least lane lines, traffic signs and lane topological relationships, and constructing a high-precision map with semantic information.
[0008] Optionally, the crowdsourcing perception data of the vehicle terminal is received, the crowdsourcing perception data is compared and identified with the established high-precision map layer to identify element changes of the target road, the high-precision map layer is incrementally updated based on the element changes, and the high-precision map is constructed.
[0009] Optionally, high-density three-dimensional point cloud data of the road environment is acquired through the laser radar acquisition device; image recognition features including feature points and object boundaries are extracted by processing continuous image data acquired by the image acquisition device; and positioning data used for analyzing the spatial position and motion posture of the data acquisition unit is acquired through the positioning device.
[0010] Optionally, laser radar point cloud data in the multi-source perception data is denoised and motion distortion corrected to obtain time-series point cloud frames; image data in the multi-source perception data is subjected to semantic feature extraction and instance segmentation to obtain image feature points containing semantic labels; the continuous motion posture of the data acquisition unit is estimated in real time by fusing the pose data, the corrected time-series point cloud frames and the image feature points; the continuous motion posture, laser radar point cloud matching constraints, image feature re-projection constraints and absolute position constraints of the positioning data are used to construct a common solving optimization problem, and an optimized posture trajectory globally consistent is obtained by optimization search; and based on the optimized posture trajectory, each time-series point cloud frame is spliced to generate a globally consistent initial point cloud map of the target road.
[0011] Optionally, based on the initial point cloud map, multi-view image perception data or fused point cloud data are unified to a bird's eye view perspective through perspective conversion to generate bird's eye view features; according to the bird's eye view features, lane lines and traffic sign elements are detected and contour-extracted as independent instances through an instance segmentation network; the extracted instance contours are fitted as parameterized vector elements, a connection topological relationship between lane lines is established, and a vectorized high-precision map layer with semantic labels is output.
[0012] Optionally, the bird's eye view features are subjected to coordinate transformation and feature projection with a driving perspective direction as a reference axis to generate a driving bird's eye view feature map; according to the driving bird's eye view feature map, the driving direction historical trajectory of the data acquisition unit or road direction prior information are combined to perform directional correlation optimization processing on the semantics and geometric attributes of lane lines, arrows and traffic sign elements, and a structured layer feature rich in directional semantics and optimized in the driving direction perspective is output, which is used for constructing or updating the high-precision map.
[0013] Optionally, a semantic element association relationship is established, a structured scene graph is constructed based on the semantic association relationship, and traffic control understanding of an associated scene is performed based on the structured scene graph to generate traffic management understanding information, which includes a target lane associated with a traffic sign, a traffic flow direction associated with a lane line, and an allowed driving direction indicated by a road arrow, and the traffic management understanding information is labeled in the high-definition map for semantic labeling.
[0014] In a second aspect of the present application, an automatic driving high-definition map construction device based on semantic understanding is provided, which includes: a data acquisition module configured to acquire multi-source perception data and pose data along a target road by a data acquisition unit, the data acquisition unit including at least an image acquisition device, a laser radar acquisition device, and a positioning device; an initial point cloud map generation module configured to generate an initial point cloud map of the target road by simultaneous localization based on the pose data and the multi-source perception data; and a high-definition map construction module configured to perform semantic segmentation and element recognition on the initial point cloud map, vectorize static road elements including at least lane lines, traffic signs, and lane topological relationships, and construct a high-definition map with semantic information.
[0015] The one or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0016] The method provided in the embodiments of the present application acquires multi-source perception data and pose data along a target road by a data acquisition unit, the data acquisition unit including at least an image acquisition device, a laser radar acquisition device, and a positioning device; generates an initial point cloud map of the target road by simultaneous localization based on the pose data and the multi-source perception data; performs semantic segmentation and element recognition on the initial point cloud map, vectorizes static road elements including at least lane lines, traffic signs, and lane topological relationships, and constructs a high-definition map with semantic information. The technical effect of improving the accuracy and reliability of high-definition map construction by accurately analyzing semantic association is achieved.
[0017] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the embodiments of the present application can be implemented in accordance with the content of the description, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application are described below. It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only exemplary, and other drawings can be obtained by those skilled in the art without creative effort on the basis of the provided drawings.
[0019] Figure 1 The flowchart of the automatic driving high-precision map construction method based on semantic understanding provided by the application.
[0020] Figure 2 The structural diagram of the automatic driving high-precision map construction device based on semantic understanding provided by the application.
[0021] Legend: data acquisition module 11, initial point cloud map generation module 12, high-precision map construction module 13. DETAILED DESCRIPTION
[0022] The application provides an automatic driving high-precision map construction method and device based on semantic understanding, which is used to solve the technical problems of insufficient accuracy and reliability in the prior art of automatic driving high-precision map construction. The technical effect of improving the accuracy and reliability of high-precision map construction is achieved by accurately analyzing semantic association.
[0023] The technical solutions in the application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are only part of the embodiments of the application, not all embodiments of the application. It should be understood that the application is not limited by the example embodiments described herein. Based on the embodiments of the application, all other embodiments obtained by those skilled in the art without creative effort fall within the scope of protection of the application. In addition, it should be noted that only parts related to the application are shown in the drawings, not all.
[0024] Embodiment one, as shown in the application, an automatic driving high-precision map construction method based on semantic understanding is provided, which comprises: Figure 1
[0025] The data acquisition unit collects multi-source perception data and pose data along the target road, and the data acquisition unit at least includes an image acquisition device, a laser radar acquisition device and a positioning device.
[0026] Further, the data acquisition unit collects multi-source perception data and pose data along the target road, including: acquiring high-density three-dimensional point cloud data of the road environment through the laser radar acquisition device; processing the continuous image data obtained by the image acquisition device to extract image recognition features containing feature points and object boundaries; collecting positioning data for analyzing the spatial position and motion attitude of the data acquisition unit through the positioning device.
[0027] Specifically, the data acquisition unit is configured and installed on a mobile measurement vehicle. The data acquisition unit collects data along the target road, and at least includes an image acquisition device, a laser radar acquisition device, and a positioning device. The image acquisition device, such as a camera, is installed at appropriate positions in front, back, left and right of the vehicle and inside the vehicle to obtain omnidirectional visual information, such as road conditions in front, vehicle conditions behind, pedestrian dynamics on the side, etc. The laser radar acquisition device is installed on the top of the vehicle to expand the scanning range and comprehensively perceive the three-dimensional environment around the vehicle. The positioning device is installed at a relatively stable position inside the vehicle to ensure that the position and attitude information of the vehicle can be accurately obtained. Multi-source perception data and pose data are collected by the data acquisition unit, the multi-source perception data includes high-density three-dimensional point cloud data and image recognition features, and the pose data includes position and motion attitude data.
[0028] The laser radar acquisition device uses laser ranging principle to emit a large number of laser beams to the target road environment, accurately calculates the distance to the surrounding objects by measuring the time difference from emission to reflection of the laser beam, and further obtains laser radar point cloud data of the road environment, i.e. high-density three-dimensional point cloud data. The high-density three-dimensional point cloud data refers to a data set composed of multiple three-dimensional coordinate information points, which can comprehensively and truly depict the geometric shape and spatial structure of the target road environment, including the ups and downs of the road surface, the outline of the surrounding buildings, the shape of the trees, etc., providing an accurate geometric basis for constructing a high-precision map.
[0029] At the same time, the image acquisition device continuously captures continuous image data of the target road, and uses image processing algorithms, such as feature extraction algorithms and edge detection algorithms, to analyze and process the collected continuous image data, and extracts image recognition features containing feature points and object boundaries from the road images. The feature points are points in the image with unique properties that can be accurately matched between different images, such as corner points, spots, etc. The object boundaries clearly define the contour range of each object in the image. By extracting image recognition features, the visual information of the road environment can be reflected, such as the color and shape of traffic signs, the texture and direction of lane lines, etc., providing rich visual details and semantic information for map construction.
[0030] The positioning device collects positioning data for analyzing the spatial position and motion posture of the data collection unit in real time through a global positioning system (GPS) and an inertial measurement unit (IMU). The spatial position specifies the specific location of the data collection unit in the earth coordinate system, and the motion posture reflects the direction, inclination angle and other information during automatic driving, ensuring that different data correspond accurately in a unified coordinate system, thereby ensuring the accuracy and consistency of the high-precision map constructed.
[0031] Through the series of operations of acquiring high-density three-dimensional point cloud data by the laser radar acquisition device, extracting image recognition features by the image acquisition device, and collecting positioning data by the positioning device, the multi-source perception data and pose data of the target road are comprehensively and accurately collected, providing comprehensive data support for the construction of the high-precision map for autonomous driving.
[0032] Based on the pose data and the multi-source perception data, an initial point cloud map of the target road is generated through simultaneous localization.
[0033] Further, based on the pose data and the multi-source perception data, an initial point cloud map of the target road is generated through simultaneous localization, including: denoising and motion distortion correction of the laser radar point cloud data in the multi-source perception data to obtain a time series point cloud frame; performing semantic feature extraction and instance segmentation on the image data in the multi-source perception data to obtain image feature points containing semantic labels; fusing the pose data, the corrected time series point cloud frame and the image feature points to estimate the continuous motion pose of the data collection unit in real time; constructing a common solving optimization problem by using the continuous motion pose, laser radar point cloud matching constraints, image feature re-projection constraints and absolute position constraints of the positioning data, and performing optimization search to obtain a globally consistent optimized pose trajectory; and based on the optimized pose trajectory, splicing each time series point cloud frame to generate a globally consistent initial point cloud map of the target road.
[0034] Specifically, due to various factors, such as the reflection of suspended particles in the environment, electronic noise of the device itself, and the like, the laser radar will be disturbed during data collection, resulting in some discrete noise points in the point cloud data that do not conform to the actual scene. At the same time, when the laser radar moves with the data collection unit, such as the vehicle carrying it, when scanning different parts of the same object, the motion will cause motion distortion of the point cloud data, so that the point cloud cannot accurately reflect the true shape and position of the object. Using a filtering algorithm based on statistical characteristics, such as a statistical outlier removal algorithm, for each point in the laser radar point cloud data, the number of neighborhood points or distance distribution within a certain radius range around the point is counted. According to a predetermined threshold, points with too few neighborhood points or significantly abnormal distance distribution are determined as noise points and removed. For example, set the radius parameter to 0.5 meters and the neighborhood point number threshold to 5. For each point, count the number of neighborhood points within the radius range. If the number is lower than the threshold, the point is considered to be an isolated noise point and is deleted. Or use a density-based filtering algorithm, such as a voxel grid-based filtering algorithm, to divide the point cloud data into multiple voxel grids, calculate the density of points in each voxel grid, and according to a set density threshold, points in voxel grids with density lower than the threshold are considered as noise points and removed.
[0035] The complete scanning process of the laser radar is divided into multiple time intervals, and the point cloud data collected in each time interval is called a sub-frame. According to the pose data, the motion change of the data collection unit between the start and end time of each sub-frame is calculated, including translation and rotation. For each point in each sub-frame, according to the time difference of its collection time relative to the start time of the sub-frame, the motion state of the data collection unit at that time is calculated using an interpolation method, and the point is motion compensated. The process of motion compensation is to convert the coordinates of the point from the local coordinate system, i.e. the laser radar coordinate system, to the global coordinate system, such as the world coordinate system, while considering the motion state of the data collection unit at that time. By converting, the point cloud data collected at different times is unified under the same coordinate system, eliminating the influence of motion distortion, thereby obtaining a time-series point cloud frame. The time-series point cloud frame can more truly and accurately reflect the three-dimensional structure of the road environment, improving the quality and accuracy of the point cloud data.
[0036] The deep learning-based convolutional neural network (CNN) extracts semantic features from image data in multi-source perception data. A pre-trained model, such as ResNet, VGG, etc., is selected as the base network. The pre-trained model has been fully trained on a large-scale image dataset and has learned rich image feature representation capabilities. To better adapt to the semantic feature extraction task of the target road scene, the pre-trained model can be fine-tuned by collecting a large amount of target road scene image data and labeling it to label various semantic elements in the image, such as traffic signs, lane lines, pedestrians, vehicles, and other category information. The pre-trained model is retrained using the labeled data to adjust the parameters of the pre-trained model so that it can more accurately extract semantic features in the target road scene.
[0037] The image data is input into the pre-trained model, and the low-level features and high-level semantic features of the image are gradually extracted through the convolutional layer and the pooling layer, forming the semantic features of the image data. Among them, the low-level features are edges, textures, etc., and the high-level semantic features are the shape and category of the object, etc.
[0038] The image data is also instance segmented, which means that different individuals belonging to the same semantic category in the image are distinguished and each individual is assigned an independent instance label. The instance segmentation method, such as Mask R-CNN, generates multiple candidate regions that may contain objects in the image through the region proposal network (RPN). For each candidate region, the RoIAlign layer is used to map its features to a fixed-size feature map, ensuring the spatial correspondence of the features and avoiding quantization errors. The feature map is input into the classification branch, the bounding box regression branch, and the mask prediction branch, respectively. The classification branch is used to predict the category of the object in the candidate region, the bounding box regression branch is used to adjust the position and size of the candidate region to make it more accurately surround the object, and the mask prediction branch is used to predict the mask of the object, i.e., the accurate pixel-level distribution of the object in the image. Through segmentation, Mask R-CNN simultaneously realizes the tasks of object detection and instance segmentation, generating an accurate mask for each detected object instance. Then, a feature point detection algorithm such as SIFT, SURF, or ORB is used to detect feature points on the instance-segmented image based on the extracted semantic features. The feature points are assigned corresponding semantic labels based on the object instance mask corresponding to the pixel region where the feature points are located, and image feature points containing semantic labels are obtained. For example, if the extracted feature point is located within the mask region of an object instance identified as a vehicle, the semantic label of the feature is a vehicle.
[0039] The pose data, the corrected time-series point cloud frames and the image feature points are fused to estimate the continuous motion pose of the data collection unit in real time. The pose data provides the spatial position and attitude information of the data collection unit at different time, the time-series point cloud frames reflect the three-dimensional geometric changes of the road environment, and the image feature points reflect the semantic information of the environment. By fusing the multi-source data and using a simultaneous localization and mapping (SLAM) algorithm, such as a SLAM algorithm based on graph optimization, the continuous motion pose of the data collection unit in the data collection process, i.e., the position and attitude of the data collection unit at each time, is estimated in real time and accurately.
[0040] The relative matching constraints between the laser radar point clouds, the re-projection error constraints of the image feature points in the three-dimensional space and the absolute position constraints provided by the positioning device are jointly introduced into the same optimization framework to construct a nonlinear optimization problem for joint solution. Among them, the laser radar point cloud matching constraints are established by performing feature extraction and point-plane / point-point constraints on the point clouds of adjacent or overlapping frames to form error terms reflecting spatial geometric consistency, the image feature re-projection constraints are obtained by projecting the three-dimensional point cloud corresponding to the feature points or instance contours obtained after semantic segmentation to the image plane and calculating the re-projection error between the three-dimensional point cloud and the actual image feature position to improve the visual consistency of the pose estimation, and the absolute position constraints of the positioning data take the coordinates provided by the positioning device as the reference datum, limit the deviation of the optimization variables from the absolute position range, and enhance the spatial globality and scale stability of the overall map.
[0041] By integrating the above multi-source constraint terms, a nonlinear least squares objective function with pose parameters as optimization variables is constructed, and an optimization search algorithm is used for iterative search so that the multi-source constraints are minimized in the unified optimization space. For example, the LM nonlinear least squares iterative optimization method is adopted, the continuous motion pose is taken as the optimization variable, the laser radar point cloud matching error, the image feature re-projection error and the absolute position error of the positioning data are uniformly constructed into a residual function, and the Jacobian matrix of each residual with respect to the pose variable is calculated by linearization. In each iteration process, the LM algorithm solves the incremental update according to the current pose estimation, adopts the Gauss-Newton update strategy when the overall residual is reduced after the update, otherwise automatically adjusts the damping factor and degenerates into the gradient descent mode to ensure that the optimal solution is gradually approached when the multi-source constraints are excessively conflicted or locally unstable. By repeatedly repeating the above iteration process, the overall objective function is finally minimized in the unified optimization space under global consistency, and a stable and convergent optimized pose trajectory is obtained, which maintains geometric consistency, semantic consistency and absolute coordinate consistency in the global range, and provides a stable and reliable spatial reference for subsequent time-series point cloud stitching and high-precision map construction.
[0042] The optimized pose trajectory provides accurate spatial transformation relationship for each time point cloud frame, and the point cloud frames collected at different times are spliced according to the spatial transformation relationship, that is, the local point cloud data is integrated into a complete and globally consistent initial point cloud map, which not only contains accurate three-dimensional geometric information of the road environment, but also integrates rich semantic information.
[0043] By using the pose data and multi-source perception data, simultaneous localization is realized and a globally consistent initial point cloud map of the target road is generated, which provides accurate and reliable data basis for automatic driving high-precision map construction based on semantic understanding.
[0044] The initial point cloud map is subjected to semantic segmentation and element recognition, and the static road elements including at least lane lines, traffic signs and lane topological relationships are vectorized to construct a high-precision map with semantic information.
[0045] Further, the initial point cloud map is subjected to semantic segmentation and element recognition, and the static road elements including at least lane lines, traffic signs and lane topological relationships are vectorized to construct a high-precision map with semantic information, including: based on the initial point cloud map, the multi-view image perception data or the fused point cloud data are unified to the bird's eye view perspective through perspective conversion to generate bird's eye view features; according to the bird's eye view features, the lane line and traffic sign elements are detected and contour extracted as independent instances through an instance segmentation network; the extracted instance contour is fitted as a parameterized vector element, the connection topological relationship between lane lines is established, and a vectorized high-precision map layer with semantic labels is output.
[0046] Specifically, based on the initial point cloud map, the multi-view image perception data or the fused point cloud data are subjected to perspective normalization processing, and through steps such as geometric projection, camera extrinsic calibration and point cloud height normalization, three-dimensional scene features are uniformly converted to the bird's eye view perspective to obtain continuous and position-consistent bird's eye view features. The bird's eye view perspective refers to a unified observation perspective in which the original three-dimensional scene is projected onto the ground plane in a vertical downward overhead manner through spatial transformation of multi-source perception data, and the bird's eye view feature refers to a structured feature representation of fused geometric information and semantic information expressed in a two-dimensional ground plane as a coordinate system under the bird's eye view perspective.
[0047] The instance segmentation network can use Mask R-CNN to perform element-level identification on the bird's eye view features. The bird's eye view features are input into the ResNet-FPN backbone network to extract a multi-scale feature pyramid. The region proposal network (RPN) generates candidate regions containing potential elements such as lane lines and traffic signs on the feature map. The classification branch distinguishes the categories of each candidate region, and the bounding box regression branch accurately locates the spatial position of each candidate region in the bird's eye plane. The Mask branch performs pixel-level segmentation on each candidate region to generate an instance mask for the corresponding element. During the training phase, the instance segmentation network uses instance annotations of elements such as lane lines, traffic signs, and road arrows as supervision signals to enable it to distinguish different instances of the same element. Through connected component analysis and contour tracking algorithms such as pixel gradient-based boundary extraction or polygon approximation-based contour reconstruction on the instance mask, the continuous contours of elements such as lane lines and traffic signs are obtained, achieving detection and contour extraction of instances such as lane lines and traffic sign elements.
[0048] The lane line and traffic sign instance contours extracted by the instance segmentation network are preprocessed to check the integrity of the contours. For contours with missing or broken points, interpolation algorithms such as linear interpolation, spline interpolation, etc. are used to repair them to ensure the continuity of the contours. Gaussian filtering or median filtering methods are used to smooth the contours to remove minor fluctuations caused by noise or segmentation errors. According to the shape characteristics of different road elements, appropriate parametric models are selected for fitting. For lane lines, common models include straight line models and polynomial curve models. The straight line model is suitable for straight lane lines, and the slope and intercept parameters of the straight line are obtained by least squares fitting. The polynomial curve model is used to fit curved lane lines, such as least squares multi-segment polyline fitting, Bezier curve fitting, or cubic spline curve fitting. For traffic signs, geometric models such as circles and rectangles are commonly used for fitting, e.g., the least squares method is used to fit the center coordinates and radius of a circular sign.
[0049] The preprocessed contour point coordinates are input into the selected parametric model, and optimization algorithms such as gradient descent and Newton's method are used to solve the model parameters to minimize the error between the fitted model and the original contour points. During the fitting process, an error threshold can be set, and when the fitting error is less than the threshold, the fitting result is considered to meet the requirements. For the fitted lane line vector elements, key features such as the starting point, ending point coordinates, curvature, and direction are extracted. Based on the characteristics of the lane lines, the spatial positional relationship between the lanes is analyzed, including parallel, intersection, branching, and merging. For example, by calculating the distance and direction change between two lane lines, it is determined whether they are parallel. If the distance between two lane lines at a certain point is less than a certain threshold and the direction changes significantly, it is considered that the two lane lines intersect, branch, or merge at that point.
[0050] The topological relationship between the lane lines obtained by analysis is stored in a structured manner, for example, using a graph data structure, the lane lines are taken as nodes in the graph, the topological relationship is taken as an edge, the attributes of the edge record the type of the topological relationship and the related geometric information, the type of the topological relationship is parallel, intersection, etc., and the geometric information is intersection coordinates, etc. And semantic labels are added to the fitted vector elements and the established topological relationship. For lane lines, corresponding semantic labels are given according to the lane line type and function, wherein the lane line type is a dashed line, a solid line, a double yellow line, etc., and the lane line function is a driving lane line, a parking line, etc. For traffic signs, corresponding semantic information is added according to the meaning of the traffic signs, such as speed limit signs, no-passing signs, etc. The semantic labels can be stored through pre-defined coding rules.
[0051] The vectorized lane lines, traffic signs and other elements, and the corresponding topological relationship and semantic labels are integrated into a unified map layer, for example, using geographic information system (GIS) software or a special map construction tool, data integration is performed according to a certain data format, such as shapefile, GeoJSON, etc., to form a vectorized, semantic-labeled high-precision map layer. The high-precision map layer can accurately describe the geometric shape, topological structure and semantic information of the road, and provide basic data support for automatic driving, intelligent transportation and other applications.
[0052] Further, the high-precision map with semantic information is constructed, including: receiving crowdsourcing perception data of a vehicle terminal, identifying element changes of the target road by comparing the crowdsourcing perception data with the established high-precision map layer; based on the element changes, incrementally updating the high-precision map layer to construct the high-precision map.
[0053] Specifically, crowdsourcing perception data collected by a plurality of vehicle terminals during driving is received, the data includes but is not limited to vehicle-mounted camera images, radar point clouds, positioning information and motion state parameters of the vehicle itself. After pre-processing and coordinate alignment, the crowdsourcing perception data is projected into the established high-precision map coordinate system, and using semantic segmentation, instance extraction or element matching algorithm, the latest perception results of the road elements such as the type and direction of the lane lines, the position state of the traffic signs, and the road arrows are identified from the crowdsourcing data.
[0054] The received crowdsourcing perception data is compared with the established high-precision map layer. During the comparison process, a multi-level and multi-dimensional method is adopted. For example, comparison is made from the spatial position dimension. Through coordinate conversion and matching algorithms, the position information of the road element in the crowdsourcing perception data is accurately aligned with the position of the corresponding element in the high-precision map layer, and it is checked whether there is a position deviation. Comparison is made from the semantic information dimension. The semantic labels of the road elements extracted in the crowdsourcing perception data, such as lane line types and traffic sign meanings, are analyzed to determine whether they are consistent with the semantic labels of the corresponding elements in the high-precision map layer. At the same time, image recognition and point cloud matching technologies are used to compare the shape, texture and other characteristics of the road elements, and to detect changes in the target road elements.
[0055] The element change type of the target road is determined through geometric shape difference, position offset, semantic category change and other criteria, including addition, disappearance, shape update or position adjustment. For example, if it is found through comparison that the position of a certain lane line deviates significantly between the high-precision map layer and the crowdsourcing perception data, and the deviation exceeds a preset threshold, it is determined that the position of the lane line has changed. If a new traffic sign is detected in the crowdsourcing perception data, but there is no corresponding sign at this position in the high-precision map layer, it is identified as a new traffic sign.
[0056] For the identified element changes, incremental update operations are performed in the map layer according to the change type, and the corresponding vector elements and their semantic attributes are dynamically replaced or adjusted, ensuring that the high-precision map is continuously updated without the need for overall reconstruction, thereby outputting an updated high-precision map with real-time and semantic integrity, providing more reliable, timely and accurate map services for autonomous vehicles, and greatly improving the safety and reliability of autonomous driving.
[0057] Further, constructing a high-precision map with semantic information also includes: taking the driving perspective direction as the reference axis, performing coordinate transformation and feature projection on the aerial view features to generate a driving aerial view feature map; according to the driving aerial view feature map, combining the driving direction historical trajectory of the data acquisition unit or the road direction prior information, performing direction correlation optimization processing on the semantic and geometric attributes of lane lines, arrows and traffic sign elements, and outputting a structured layer feature optimized in the driving direction perspective and rich in directional semantics, which is used to construct or update the high-precision map.
[0058] Specifically, the driving perspective direction refers to the vehicle forward direction or the road center line direction. The driving perspective direction is taken as a reference axis to perform coordinate transformation on the generated bird's eye view features. The bird's eye view features are converted from a geographic coordinate system to a driving coordinate system with the driving direction as the longitudinal reference through a rotation matrix and a translation vector, and the multi-source fused bird's eye view features are re-projected and arranged in the driving coordinate system, so as to generate a driving bird's eye view feature map conforming to the driving perspective of the vehicle. The driving bird's eye view feature map can more intuitively reflect the relative position and shape of road elements in the driving process.
[0059] The driving direction history track of the data acquisition unit is acquired. The driving direction history track records the driving direction and position information of the vehicle at different times. At the same time, road direction prior information such as the road direction determined in advance in road planning and design can also be acquired. The driving bird's eye view feature map is fused and analyzed with the driving direction information. The semantic attributes and geometric shapes of the lane line, road arrow, traffic sign and other elements are analyzed for direction correlation, including judging the front and back pointing of the elements, the left and right attribution relationship, the consistency with the road direction and the direction offset in the driving coordinate system, and based on the direction correlation information, the geometric parameters and semantic classification of the elements are optimized, so that they have clear direction attributes under the driving reference axis. For example, for the lane line element, it is judged whether it is the lane line of the current driving direction according to the driving direction, and then its semantic label is optimized to clearly mark it as the main lane line, the overtaking lane line, etc. At the same time, the geometric attributes of the lane line are corrected in combination with the driving direction, such as adjusting the curvature and direction of the lane line, so that it is more in line with the visual feeling in actual driving. For the arrow element, it is determined whether its indicating direction is correct according to the driving direction. If there is a deviation, it is corrected, and its semantic information is optimized, such as clearly marking it as a straight arrow, a left turn arrow, etc. For the traffic sign element, it is judged whether it is within the visible range of the current driving route according to the driving direction. Its semantic label is optimized, such as marking it as a speed limit sign ahead, a no-passing sign, etc., and its geometric attributes are adjusted, such as size and position, so that it is more accurate and reasonable under the driving perspective.
[0060] Finally, the optimized structured layer features rich in directional semantics are output for high-precision map construction or layer updating. The structured layer features are stored in a structured data form and contain the accurate geometric information and semantic categories of the lane line, arrow, traffic sign and other elements.
[0061] By taking the driving perspective as the reference to perform coordinate transformation and feature projection to generate a driving bird's eye view feature map, and combining the driving direction information to optimize the direction association of the semantic and geometric attributes of lane lines, arrows, traffic signs and other elements, a structured layer feature rich in directional semantics can be output, so that the constructed or updated high-precision map is more in line with the actual driving scene, providing more accurate and intuitive navigation and decision-making basis for autonomous vehicles, effectively improving the perception and planning capabilities of the autonomous driving system in complex road environments, and enhancing the safety and reliability of driving.
[0062] Further, the method further comprises: establishing an association relationship between semantic elements, constructing a structured scene graph based on the semantic association relationship; based on the structured scene graph, understanding the traffic control of the associated scene, generating traffic management understanding information, the traffic management understanding information including: the target lane associated with the control of the traffic sign, the traffic flow direction associated with the constraint of the lane line, and the allowed driving direction indicated by the road arrow, labeling the traffic management understanding information to the high-precision map for semantic labeling.
[0063] Specifically, based on the spatial position relationship, direction attribute and semantic category between the extracted and vectorized lane lines, traffic signs, road arrows and other elements, through steps such as spatial proximity analysis, direction consistency determination and semantic rule matching, the semantic association relationship between the elements is constructed, including the effective range relationship of the sign to the lane, the direction consistency relationship of the arrow and the corresponding lane, the topological and driving constraint relationship between the lane lines, etc. For example, the shortest distance and the visual projection relationship between the traffic sign and each lane center line or lane boundary line are calculated, and when the traffic sign is located within the effective range of a certain lane and the direction is consistent with the driving direction of the vehicle, the sign and the target lane are associated with the control relationship. For road arrows, by analyzing the angle between the arrow vector direction and the lane center line direction, the position overlap degree and the lane area where the arrow is located, the corresponding lane and its allowed driving direction indicated by the arrow are determined, and the pointing association is established. For lane lines, the topological association and driving constraint association are constructed according to the front and rear connection order, left and right adjacent relationship and split-flow structure characteristics of the vector curve.
[0064] According to all elements and their associated relationships, a structured scene graph is constructed, different road elements are represented by nodes, and the functional or semantic relationship between elements is represented by edges, so as to realize the structured expression of the road scene. Based on the structured scene graph, the traffic control understanding of the associated scene is performed: the control lane applied by the traffic sign is determined according to the associated edge between the sign node and the lane node, the traffic flow direction and the driving changeable range are analyzed according to the direction attribute and the topological relationship of the lane line node, and the specific driving direction allowed by the road is inferred according to the graphic semantics represented by the arrow node. The traffic management understanding information such as the traffic sign control information, the lane direction constraint, and the driving direction indicated by the arrow is corresponded to the corresponding map element, and is labeled into the high-precision map layer in the form of text labeling and icon labeling, so that the high-precision map not only contains the geometric information of the road, but also has complete traffic management semantic information. For example, in the high-precision map, the target lane number controlled by the traffic sign is labeled beside the traffic sign, the traffic flow direction constrained by the lane line is labeled beside the lane line, and the driving direction allowed by the road arrow is labeled beside the road arrow.
[0065] By establishing the associated relationship between the semantic elements and constructing the structured scene graph, the traffic control situation of the associated scene can be associated, the traffic management understanding information can be accurately generated and labeled into the high-precision map, the high-precision map has more rich semantic information, the more comprehensive and accurate traffic rules and driving guidance can be provided for the autonomous vehicle, and the decision-making ability and safety of the autonomous driving in the complex traffic environment can be effectively improved.
[0066] In the second embodiment, based on the same inventive concept as the semantic understanding based high-precision map construction method in the foregoing embodiments, as shown in Figure 2 The application provides a semantic understanding based high-precision map construction device for autonomous driving, wherein the semantic understanding based high-precision map construction device for autonomous driving comprises:
[0067] A data acquisition module 11 is configured to acquire multi-source perception data and pose data along a target road by a data acquisition unit, wherein the data acquisition unit at least comprises an image acquisition device, a laser radar acquisition device, and a positioning device; an initial point cloud map generation module 12 is configured to generate an initial point cloud map of the target road by simultaneous localization based on the pose data and the multi-source perception data; and a high-precision map construction module 13 is configured to perform semantic segmentation and element recognition on the initial point cloud map, vectorize static road elements containing at least lane lines, traffic signs, and lane topological relationships, and construct a high-precision map with semantic information.
[0068] Further, the high-precision map construction module 13 is further configured to: receive crowdsourcing perception data of the vehicle terminal, identify element changes of the target road by comparing the crowdsourcing perception data with the established high-precision map layer, and perform incremental update on the high-precision map layer based on the element changes to construct the high-precision map.
[0069] Further, the data acquisition module 11 is further configured to: acquire high-density three-dimensional point cloud data of a road environment through the laser radar acquisition device, extract image recognition features containing feature points and object boundaries by processing continuous image data acquired by the image acquisition device, and acquire positioning data for analyzing spatial positions and motion postures of the data acquisition unit through the positioning device.
[0070] Further, the initial point cloud map generation module 12 is further configured to: perform denoising and motion distortion correction on laser radar point cloud data in the multi-source perception data to obtain time-series point cloud frames, perform semantic feature extraction and instance segmentation on image data in the multi-source perception data to obtain image feature points containing semantic labels, fuse the pose data, the corrected time-series point cloud frames and the image feature points to estimate continuous motion poses of the data acquisition unit in real time, construct a common solving optimization problem by using the continuous motion poses, laser radar point cloud matching constraints, image feature re-projection constraints and absolute position constraints of the positioning data, and perform optimization search to obtain a globally consistent optimized pose trajectory, and splice each time-series point cloud frame based on the optimized pose trajectory to generate a globally consistent initial point cloud map of the target road.
[0071] Further, the high-precision map construction module 13 is further configured to: based on the initial point cloud map, unify multi-view image perception data or fused point cloud data to a bird's eye view perspective through perspective conversion to generate bird's eye view features, detect and contour extract lane lines and traffic sign elements as independent instances through an instance segmentation network according to the bird's eye view features, fit the extracted instance contours into parameterized vector elements, establish connection topological relations between lane lines, and output a vectorized high-precision map layer with semantic labels.
[0072] Further, the high-precision map construction module 13 is further configured to: take a driving perspective direction as a reference axis, perform coordinate transformation and feature projection on the bird's eye view features to generate a driving bird's eye view feature map, and perform direction correlation optimization processing on semantic and geometric attributes of lane lines, arrows and traffic sign elements according to the driving bird's eye view feature map in combination with a driving direction historical trajectory of the data acquisition unit or road direction prior information to output a structured layer feature rich in directional semantics and optimized in a driving direction perspective, which is used for constructing or updating the high-precision map.
[0073] Further, the high-definition map construction module 13 is further configured to: establish a semantic element association relationship, construct a structured scene graph based on the semantic association relationship, and perform traffic control understanding of an associated scene based on the structured scene graph to generate traffic management understanding information, wherein the traffic management understanding information includes: a target lane associated with a traffic sign, a traffic flow direction associated with a lane line, and an allowed driving direction indicated by a road arrow, and the traffic management understanding information is labeled to the high-definition map for semantic labeling.
[0074] Each of the embodiments in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The semantic understanding based automatic driving high-definition map construction method and specific examples in the first embodiment are also applicable to the semantic understanding based automatic driving high-definition map construction device of the present embodiment. Through the foregoing detailed description of the semantic understanding based automatic driving high-definition map construction method, those skilled in the art can clearly understand the semantic understanding based automatic driving high-definition map construction device of the present embodiment. Therefore, for the sake of brevity of the specification, the semantic understanding based automatic driving high-definition map construction device of the present embodiment will not be described in detail.
[0075] The above description of disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
[0076] Obviously, for those skilled in the art, without departing from the principles of the present application, the present application can be improved and modified in several ways, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A method for constructing high-precision maps for autonomous driving based on semantic understanding, characterized in that, include: The data acquisition unit collects multi-source sensing data and pose data along the target road through a data acquisition unit, which includes at least an image acquisition device, a lidar acquisition device, and a positioning device. Based on the pose data and the multi-source perception data, an initial point cloud map of the target road is generated through synchronous positioning. The initial point cloud map is semantically segmented and feature identified, and static road features, including at least lane lines, traffic signs and lane topology, are vectorized to construct a high-precision map with semantic information.
2. The method for constructing high-precision maps for autonomous driving based on semantic understanding according to claim 1, characterized in that, Constructing high-precision maps with semantic information includes: Receive crowdsourced perception data from vehicle terminals, and compare the crowdsourced perception data with the established high-precision map layer to identify changes in the elements of the target road; Based on the changes in the aforementioned elements, the high-precision map layer is incrementally updated to construct the high-precision map.
3. The method for constructing high-precision maps for autonomous driving based on semantic understanding according to claim 1, characterized in that, Multi-source sensing data and pose data are collected along the target road using a data acquisition unit, including: The laser radar acquisition device is used to acquire high-density three-dimensional point cloud data of the road environment; By processing the continuous image data acquired by the image acquisition device, image recognition features containing feature points and object boundaries are extracted. The positioning device collects positioning data for analyzing the spatial position and motion posture of the data acquisition unit.
4. The method for constructing high-precision maps for autonomous driving based on semantic understanding according to claim 3, characterized in that, Based on the pose data and the multi-source sensing data, an initial point cloud map of the target road is generated through synchronous positioning, including: Denoising and motion distortion correction are performed on the lidar point cloud data in the multi-source sensing data to obtain a temporal point cloud frame; Semantic feature extraction and instance segmentation are performed on the image data in the multi-source perception data to obtain image feature points containing semantic labels; By fusing the pose data, the corrected temporal point cloud frames, and image feature points, the continuous motion pose of the data acquisition unit is estimated in real time. The continuous motion pose, lidar point cloud matching constraints, image feature reprojection constraints, and absolute position constraints of the localization data are used to construct a common optimization problem, and an optimization search is performed to obtain a globally consistent optimized pose trajectory. Based on the optimized pose trajectory, the point cloud frames of each time series are stitched together to generate a globally consistent initial point cloud map of the target road.
5. The method for constructing high-precision maps for autonomous driving based on semantic understanding according to claim 4, characterized in that, The initial point cloud map is semantically segmented and feature identified, and static road features, including at least lane lines, traffic signs, and lane topology, are vectorized to construct a high-precision map with semantic information, including: Based on the initial point cloud map, multi-view image perception data or fused point cloud data are unified to the bird's-eye view through viewpoint transformation to generate bird's-eye view features. Based on the bird's-eye view features, lane lines and traffic sign elements are detected and their contours extracted as independent instances using an instance segmentation network. The extracted instance contours are fitted into parameterized vector elements to establish the connection topology between lane lines, and a vectorized, semantically labeled high-precision map layer is output.
6. The method for constructing high-precision maps for autonomous driving based on semantic understanding according to claim 5, characterized in that, Constructing high-precision maps with semantic information also includes: Using the driving perspective direction as a reference axis, coordinate transformation and feature projection are performed on the bird's-eye view features to generate a driving bird's-eye view visual feature map. Based on the driving bird's-eye view visual feature map, and combined with the driving direction historical trajectory or road direction prior information of the data acquisition unit, the semantic and geometric attributes of lane lines, arrows, and traffic signs are optimized for directional correlation. The resulting structured layer features, optimized for driving direction perspective and rich in directional semantics, are used to construct or update the high-precision map.
7. The method for constructing high-precision maps for autonomous driving based on semantic understanding according to claim 6, characterized in that, Also includes: Establish the relationships between semantic elements, and construct a structured scene graph based on these relationships; Based on the structured scene graph, traffic control is understood in the associated scene to generate traffic management understanding information. The traffic management understanding information includes: the target lanes controlled by traffic signs, the traffic flow directions constrained by lane lines, and the permitted driving directions indicated by road arrows. The traffic management understanding information is then annotated in the high-precision map for semantic annotation.
8. A high-precision map construction device for autonomous driving based on semantic understanding, characterized in that, The steps for implementing the semantic understanding-based high-precision map construction method for autonomous driving according to any one of claims 1 to 7 include: The data acquisition module is used to collect multi-source sensing data and pose data along the target road through the data acquisition unit. The data acquisition unit includes at least an image acquisition device, a lidar acquisition device, and a positioning device. The initial point cloud map generation module is used to generate an initial point cloud map of the target road based on the pose data and the multi-source perception data through synchronous positioning. The high-precision map construction module is used to perform semantic segmentation and feature recognition on the initial point cloud map, vectorize static road elements that include at least lane lines, traffic signs and lane topology, and construct a high-precision map with semantic information.
Citation Information
Cited By
Map data updating method and device, equipment and storage medium
CN122045323A
Map data updating method and device, equipment and storage medium
CN122045323B