Semantic map construction method, device and equipment and computer program product
Through multi-sensor data fusion and two-dimensional grid semantic map construction methods, combined with image and laser point cloud data, the problems of large calculation volume, low accuracy and low storage efficiency in semantic map construction are solved, and efficient and accurate semantic map construction is achieved, which is suitable for autonomous driving systems.
Patent Information
- Application Number
- CN202510581511.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-12
AI Technical Summary
The existing semantic map construction methods are difficult to meet the real-time and high-precision requirements of autonomous driving due to the problems of large computing volume, long processing time, low accuracy and low storage efficiency.
Multi-sensor data fusion method is adopted, combined with image data and laser point cloud data for semantic segmentation, and through two-dimensional grid semantic map construction and pose map optimization, the semantic segmentation accuracy and map construction efficiency are improved, and visual semantic information and laser point cloud information are integrated.
It improves the construction accuracy and efficiency of semantic maps, reduces storage requirements, enhances the robustness and real-timeness of the map, and is suitable for autonomous driving systems.
Smart Images

Figure CN120467360A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of semantic map construction, and in particular to a semantic map construction method, apparatus and device, and computer program product. Background Art
[0002] With the rapid development of autonomous driving technology, semantic maps serve as a crucial foundation for autonomous vehicles to perceive the environment, plan paths, and make decisions. The quality of their construction directly impacts the overall performance of autonomous driving systems. Semantic maps not only contain road geometry information but also incorporate rich semantic information, such as lane markings, traffic signs, pedestrians, and vehicles. This provides autonomous vehicles with a more comprehensive and accurate understanding of the environment.
[0003] However, the current construction of semantic maps faces many challenges:
[0004] 1) Construction efficiency: Traditional semantic map construction methods typically perform semantic segmentation directly on point cloud data. This process is computationally intensive and time-consuming, making it difficult to meet the real-time requirements of autonomous driving applications. This efficiency issue is particularly prominent when constructing large-scale maps or dynamically updating them.
[0005] 2) Map accuracy: The accuracy of semantic maps is directly related to the safety and comfort of autonomous vehicles. However, due to factors such as sensor noise, environmental changes, and occlusion, it is often difficult to obtain high-precision semantic maps by performing semantic segmentation directly on point cloud data.
[0006] 3) Map size and storage efficiency issues: As the amount of information contained in semantic maps continues to increase, the amount of map data has expanded dramatically, bringing tremendous pressure to storage and transmission. Summary of the Invention
[0007] The embodiments of the present application provide a semantic map construction method, apparatus and device, and computer program product to improve the accuracy and efficiency of semantic map construction and improve map storage efficiency.
[0008] The embodiments of this application adopt the following technical solutions:
[0009] In a first aspect, an embodiment of the present application provides a method for constructing a semantic map, the method comprising:
[0010] Acquiring multi-sensor data, wherein the multi-sensor data includes image data, laser point cloud data, inertial navigation data, and GPS data;
[0011] Performing semantic segmentation on road surface element information in the image data to obtain an image semantic segmentation result;
[0012] Performing data association on the laser point cloud data and the image semantic segmentation result to construct a two-dimensional grid semantic map of the current frame;
[0013] Performing pose graph optimization based on the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the GPS data to obtain a pose graph optimization result;
[0014] A global two-dimensional grid semantic map is constructed according to the pose graph optimization result.
[0015] Optionally, the performing data association on the laser point cloud data and the image semantic segmentation result to construct a two-dimensional grid semantic map of the current frame includes:
[0016] Determine the transformation relationship between laser point cloud data and image data;
[0017] Projecting the laser point cloud data into the image data based on a transformation relationship between the laser point cloud data and the image data;
[0018] Determining a correspondence between the laser point cloud data and the image semantic segmentation result according to a projection result of the laser point cloud data, wherein the correspondence includes image semantic information corresponding to each pixel point in the image data and laser point cloud intensity information;
[0019] According to the correspondence between the laser point cloud data and the image semantic segmentation result, the image semantic segmentation result is subjected to two-dimensional gridding processing to obtain a two-dimensional grid semantic map of the current frame.
[0020] Optionally, performing pose graph optimization according to the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the GPS data to obtain a pose graph optimization result includes:
[0021] Determine the initial position of the current frame according to the GPS data;
[0022] Constructing a pose constraint for the current frame based on the two-dimensional grid semantic map of the current frame, the inertial navigation data, and a local two-dimensional grid semantic map corresponding to the current frame;
[0023] According to the initial pose of the current frame and the pose constraints of the current frame, the optimized pose of the current frame is solved using the preset pose optimization algorithm;
[0024] The pose graph is optimized according to the optimized pose of the current frame and the GPS data to obtain a pose graph optimization result.
[0025] Optionally, the pose constraint includes a first pose constraint, and constructing the pose constraint of the current frame according to the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the local two-dimensional grid semantic map corresponding to the current frame includes:
[0026] Calculating a pose change between the two-dimensional grid semantic map of the current frame and the local two-dimensional grid semantic map;
[0027] The first pose constraint is constructed according to a pose change between the two-dimensional grid semantic map of the current frame and the local two-dimensional grid semantic map.
[0028] Optionally, the pose constraint includes a second pose constraint, and constructing the pose constraint of the current frame according to the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the local two-dimensional grid semantic map corresponding to the current frame includes:
[0029] Based on the inertial navigation data, calculating the pose change of the local two-dimensional grid semantic map to the current frame;
[0030] The second pose constraint is constructed according to the pose change from the local two-dimensional grid semantic map to the current frame.
[0031] Optionally, performing pose graph optimization according to the optimized pose of the current frame and the GPS data to obtain a pose graph optimization result includes:
[0032] Determining whether the current frame meets the key frame condition according to the optimized posture of the current frame;
[0033] When the current frame meets the key frame condition, the pose graph is optimized according to the optimized pose of the current frame and the GPS data to obtain a pose graph optimization result.
[0034] Optionally, constructing a global two-dimensional grid semantic map according to the pose graph optimization result includes:
[0035] Update the local two-dimensional grid semantic map corresponding to the current frame according to the pose graph optimization result;
[0036] A global two-dimensional grid semantic map is constructed based on the updated local two-dimensional grid semantic map.
[0037] In a second aspect, an embodiment of the present application further provides a semantic map construction device, the semantic map construction device comprising:
[0038] An acquisition unit, configured to acquire multi-sensor data, wherein the multi-sensor data includes image data, laser point cloud data, inertial navigation data, and GPS data;
[0039] A semantic segmentation unit, configured to perform semantic segmentation on the road surface element information in the image data to obtain an image semantic segmentation result;
[0040] A first construction unit is configured to perform data association on the laser point cloud data and the image semantic segmentation result to construct a two-dimensional grid semantic map of the current frame;
[0041] an optimization unit, configured to optimize a pose graph according to the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the GPS data, to obtain a pose graph optimization result;
[0042] The second construction unit is used to construct a global two-dimensional grid semantic map according to the pose graph optimization result.
[0043] In a third aspect, an embodiment of the present application further provides a device, including:
[0044] A processor; and a memory arranged to store computer-executable instructions, which, when executed, cause the processor to perform any of the aforementioned semantic map construction methods.
[0045] In a fourth aspect, an embodiment of the present application further provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements any of the aforementioned semantic map construction methods.
[0046] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: The semantic map construction method of the embodiments of the present application first obtains multi-sensor data, which includes image data, laser point cloud data, inertial navigation data, and GPS data; then semantically segments the road element information in the image data to obtain image semantic segmentation results; then, data association is performed on the laser point cloud data and the image semantic segmentation results to construct a two-dimensional grid semantic map of the current frame; then, pose graph optimization is performed based on the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the GPS data to obtain pose graph optimization results; finally, a global two-dimensional grid semantic map is constructed based on the pose graph optimization results. The semantic map construction method of the embodiments of the present application combines image data and laser point cloud data to construct a semantic map, thereby improving the accuracy of semantic segmentation and the efficiency of semantic map construction. By constructing a two-dimensional grid semantic map, the storage efficiency of the semantic map is improved, and each grid in the semantic map simultaneously stores visual semantic information and laser point cloud information, thereby improving the accuracy of semantic map construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0048] Figure 1 A flowchart of a method for constructing a semantic map according to an embodiment of the present application is shown;
[0049] Figure 2 This is a schematic diagram of the structure of a semantic map construction device in an embodiment of the present application;
[0050] Figure 3 This is a structural diagram of a device in an embodiment of the present application. DETAILED DESCRIPTION
[0051] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0052] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0053] The present application provides a method for constructing a semantic map. As shown in the figure, a flow chart of a method for constructing a semantic map in the present application is provided. The method for constructing a semantic map includes at least the following steps S110 to S150:
[0054] Step S110 , acquiring multi-sensor data, wherein the multi-sensor data includes image data, laser point cloud data, inertial navigation data, and GPS data.
[0055] When building a semantic map, data must first be acquired from a variety of sensors installed on vehicles, mobile devices, or fixed roadside equipment. For example, on-board cameras can capture image data, which can capture detailed visual information about the environment surrounding the axle. On-board lidar can be used to obtain laser point cloud data to accurately measure the distance and shape of surrounding objects. The vehicle's inertial navigation system can be used to obtain inertial navigation data, which is used to record the device's motion state and posture changes. Meanwhile, the GPS module can be used to obtain the vehicle's GPS data to provide global position information.
[0056] Because different sensors may have different sampling frequencies and data output times, image data, laser point cloud data, and GPS data can be synchronized in time and space based on high-frequency inertial navigation data (INS data typically has a higher sampling frequency and can provide more continuous and accurate time information). For example, using the INS data's timestamp as a reference, interpolation or sampling adjustments can be performed on image data and laser point cloud data to align them in time with the INS data. For GPS data, data fusion or time calibration can be performed based on its update frequency and the time relationship with the INS data to ensure the temporal and spatial consistency of all sensor data, providing an accurate foundation for subsequent data processing and analysis.
[0057] Step S120 , performing semantic segmentation on the road surface element information in the image data to obtain an image semantic segmentation result.
[0058] The synchronized image data is fed into a trained semantic segmentation model. The model then classifies each pixel in the image, determining its category of road surface elements, such as lane markings, arrows, and stop signs. This yields image semantic segmentation results, represented as pixel-level labels that clearly indicate the location (x, y) and category of each road surface element in the image.
[0059] The semantic segmentation model can be derived from a pre-trained convolutional neural network (CNN) model, such as U-Net or DeepLab. These models are trained on large amounts of annotated image data and can learn the features and semantic information of different objects in the image. Of course, the specific semantic segmentation model used is flexibly selected by those skilled in the art based on practical needs and is not specifically limited here.
[0060] Step S130 , performing data association on the laser point cloud data and the image semantic segmentation result to construct a two-dimensional grid semantic map of the current frame.
[0061] After obtaining the image semantic segmentation information, a correspondence between the laser point cloud data and the image semantic segmentation results can be established based on the internal and external parameters of the laser and camera. Based on this correspondence, a 2D grid semantic map of the current frame is constructed. Each grid in the 2D grid semantic map contains both the image semantic segmentation information and the laser point cloud information, thereby improving the accuracy of semantic map construction.
[0062] The above two-dimensional grid can be divided according to the size and accuracy requirements of the actual application scenario. For example, the map is divided into grid cells of a certain size, and each grid cell records information such as the location, category, and laser intensity of the road elements in the area.
[0063] Step S140 , performing pose graph optimization based on the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the GPS data to obtain a pose graph optimization result.
[0064] Both inertial navigation data and GPS data can provide constraint information for pose optimization to a certain extent. Therefore, the pose graph can be optimized by combining the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the GPS data. The purpose of pose graph optimization is to improve the accuracy of pose estimation. For example, the pose graph can be optimized using nonlinear optimization algorithms such as least squares. By minimizing the error of all edges in the pose graph (i.e., the difference between the actual measured value and the predicted value), the position and posture of the nodes are adjusted to make the entire pose graph more consistent and accurate. The optimized pose graph can eliminate noise and errors in the sensor data, improve the accuracy and stability of the device pose estimation, and obtain the pose graph optimization result for the current frame.
[0065] Step S150: constructing a global two-dimensional grid semantic map based on the pose graph optimization result.
[0066] By continuously optimizing the pose graph, we can obtain the optimized poses of multiple frames. By combining the image semantic information of each frame and the corresponding laser point cloud information, we can establish a global two-dimensional grid semantic map.
[0067] The semantic map construction method of the embodiment of the present application combines image data and laser point cloud data to construct a semantic map, thereby improving the semantic segmentation accuracy and the efficiency of semantic map construction. By constructing a two-dimensional grid semantic map, the storage efficiency of the semantic map is improved, and each grid in the semantic map simultaneously stores visual semantic information and laser point cloud information, thereby improving the accuracy of semantic map construction.
[0068] In some embodiments of the present application, the data association of the laser point cloud data and the image semantic segmentation result to construct a two-dimensional grid semantic map of the current frame includes: determining a transformation relationship between the laser point cloud data and the image data; projecting the laser point cloud data into the image data based on the transformation relationship between the laser point cloud data and the image data; determining a correspondence between the laser point cloud data and the image semantic segmentation result according to the projection result of the laser point cloud data, the correspondence including image semantic information and laser point cloud intensity information corresponding to each pixel in the image data; and performing two-dimensional gridding processing on the image semantic segmentation result according to the correspondence between the laser point cloud data and the image semantic segmentation result to obtain a two-dimensional grid semantic map of the current frame.
[0069] When associating laser point cloud data with image semantic segmentation results, the transformation relationship between the laser point cloud data and the image semantic segmentation results can be determined first. Here, the external parameters of the laser radar (the transformation relationship from the laser radar to the vehicle body) and the external parameters of the camera (the transformation relationship from the camera to the vehicle body) as well as the transformation relationship from the camera to the image coordinate system can be considered as pre-calibrated parameters. Then, by combining the external parameters of the laser radar, the external parameters of the camera and the transformation relationship from the camera to the image coordinate system, the transformation relationship between the laser point cloud data and the image semantic segmentation results can be determined.
[0070] After determining the transformation relationship between the laser point cloud data and the image semantic segmentation results, the laser point cloud data can be projected into the image, so that the image pixel points corresponding to each point in the point cloud data can be obtained, thereby establishing a correspondence between the semantic segmentation information of the image pixel points and the intensity information of the laser point cloud.
[0071] Based on the correspondence between the semantic segmentation information of the above-mentioned image pixels and the intensity information of the laser point cloud, this information is further two-dimensionally gridded in the image dimension. For example, the image can be divided into a two-dimensional grid of a fixed size, thus obtaining a two-dimensional grid semantic map of the current frame.
[0072] Each pixel in the grid stores the following two pieces of information, which serve as the basis for constructing a global semantic map:
[0073] 1) Image semantic information: such as category labels of road signs such as “lane line”, “arrow”, and “stop line”.
[0074] 2) Laser point cloud intensity information: such as reflectivity, distance, etc., used to enhance the robustness of semantic features.
[0075] By combining the high-resolution features of image semantic segmentation and information such as the intensity of laser point clouds, the error of a single sensor is reduced and the accuracy of semantic map construction is improved.
[0076] In some embodiments of the present application, the pose graph optimization is performed based on the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the GPS data to obtain the pose graph optimization result, including: determining the initial pose of the current frame based on the GPS data; constructing the pose constraints of the current frame based on the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the local two-dimensional grid semantic map corresponding to the current frame; solving the optimized pose of the current frame using a preset pose optimization algorithm based on the initial pose of the current frame and the pose constraints of the current frame; performing pose graph optimization based on the optimized pose of the current frame and the GPS data to obtain the pose graph optimization result.
[0077] The initial pose of the current frame can be determined based on GPS data. GPS provides the device's approximate location on the Earth's surface. Using specific algorithms and coordinate transformations, GPS data is converted into an estimate of the initial pose of the current frame in a specific coordinate system. For example, GPS may provide information such as the device's latitude, longitude, and altitude, which, after processing, yields an initial three-dimensional pose (position and attitude). However, because GPS positioning is susceptible to environmental influences, the pose determined based on GPS data may contain errors, and therefore is only used here as an initial pose.
[0078] The current frame's 2D grid semantic map contains information about the road environment, such as the location and category of road surface elements in the current frame's area. The local 2D grid semantic map for the current frame refers to the 2D grid semantic map within a specific area of the current frame and is constructed based on data from previous moments.
[0079] Specifically, during the movement of vehicles and other equipment, sensors (such as lidar, cameras, etc.) will continuously collect raw data of the surrounding environment, use semantic segmentation algorithms to perform semantic analysis on the processed data, and identify semantic information such as the location and category of different road elements. Based on the collected time series data, the environmental data with semantic information is gradually constructed into a two-dimensional grid semantic map. At the same time, as the equipment continues to move, new data is continuously collected, and the local two-dimensional grid semantic map will continue to be updated to reflect the latest environmental information within a certain area currently covered. This "certain area range" can be set according to actual application scenarios and needs, such as a circular area with a certain radius centered on the current location of the device, or a rectangular area of a fixed size.
[0080] Therefore, a local 2D grid semantic map is constructed and continuously updated based on data from previous moments. It contains semantic information about the area a device passes through during its movement within a certain timeframe, reflecting the overall environmental characteristics of that area over that period. The data range is relatively large, encompassing the area of the current frame and a certain range around it. The 2D grid semantic map for the current frame, on the other hand, is generated for the moment corresponding to the current frame and only contains semantic information such as the location and category of road elements in the current frame's area. The data range is limited to the field of view or perception range of the current frame, making it smaller than a local 2D grid semantic map.
[0081] By matching the 2D grid semantic map of the current frame with the corresponding local 2D grid semantic map, the constraints on the pose change can be obtained.
[0082] Inertial navigation data also provides information such as the acceleration and angular velocity of the current frame. By integrating this data and performing other processing, the relative motion of the device between adjacent frames can be estimated, thereby also obtaining constraints on posture changes. For example, inertial navigation data can provide the distance and direction a device has moved in a short period of time. This information can be used to construct posture constraints between adjacent frames.
[0083] Based on the pose constraints described above, a pose optimization algorithm, such as the least squares method, can be used to optimize the initial pose of the current frame, thereby obtaining the optimized pose of the current frame. The goal of the pose optimization algorithm is to minimize the error in the pose constraints so that the optimized pose can better meet all constraints. For example, through iterative calculations, the pose parameters of the current frame are continuously adjusted to minimize the error between the observation value calculated based on the pose and the actual observation value.
[0084] After obtaining the optimized pose for the current frame, the pose graph is further optimized based on the optimized pose for the current frame and GPS data. A pose graph is a graph structure that represents the relationship between the poses of a device at different times, where nodes represent poses and edges represent constraints between poses. The optimized pose for the current frame is added as a new node to the pose graph, and the entire pose graph structure is further optimized based on the GPS data. For example, by taking into account the global position information provided by GPS data, the node positions in the pose graph are globally adjusted, making the pose graph more accurate and consistent.
[0085] By comprehensively utilizing GPS data, two-dimensional grid semantic maps, and inertial navigation data, obtaining pose information from multiple angles and constructing pose constraints can more comprehensively consider the movement and positional relationships of the device in the environment. Compared to using a single data source alone, this multi-source data fusion method can significantly improve the accuracy of pose estimation. In practical applications, various data sources may be subject to different interferences and errors. By constructing pose constraints and performing pose graph optimization, these errors and interferences can be effectively handled, enhancing the robustness of pose estimation. In addition, the pose graph optimization process takes into account the pose relationships between different frames and the global position information provided by GPS data, ensuring the global consistency of the entire pose estimation result.
[0086] In some embodiments of the present application, the posture constraint includes a first posture constraint, and constructing the posture constraint of the current frame based on the two-dimensional grid semantic map of the current frame, the inertial navigation data and the local two-dimensional grid semantic map corresponding to the current frame includes: calculating the posture change between the two-dimensional grid semantic map of the current frame and the local two-dimensional grid semantic map; constructing the first posture constraint based on the posture change between the two-dimensional grid semantic map of the current frame and the local two-dimensional grid semantic map.
[0087] A pose constraint constructed in this application can be derived from a pose transformation constraint between semantic map information. Specifically, based on the two-dimensional grid semantic map of the current frame and the local two-dimensional grid semantic map, particle filtering technology or histogram filtering technology can be used to calculate the pose change between the two-dimensional grid semantic map of the current frame and the local two-dimensional grid semantic map, and then the first pose constraint can be constructed based on the pose change between the two-dimensional grid semantic map of the current frame and the local two-dimensional grid semantic map.
[0088] When using particle filtering to calculate pose changes, a set of particles can be generated based on the initial pose estimate. Each particle represents a possible pose state, including position and attitude. The number and initial distribution of particles can be adjusted based on the specific application scenario. For each particle, the particle weight is updated based on the degree of match between the current frame's 2D semantic map and the local semantic map. The degree of match can be measured by calculating the similarity between the two semantic maps, for example, using metrics such as cross-entropy and mutual information. The higher the similarity, the greater the particle weight. Resampling is then performed based on the particle weights, retaining particles with larger weights and eliminating those with smaller weights. The resampling process increases the probability of sampling the correct pose state, improving the accuracy of the pose estimate. Finally, based on the resampled particle set, the weighted average pose of the particles is calculated as the pose change estimate between the current frame's 2D semantic map and the local semantic map, i.e., the first pose constraint.
[0089] When calculating pose changes using histogram filtering, the vehicle's pose state space is first discretized into a set of grids. Each grid represents a possible pose state. The probability of each grid is updated based on the inertial navigation data. The inertial navigation data provides a prior probability distribution for pose changes. Using Bayes' theorem, the prior probability is combined with the matching probability between the current frame's 2D semantic map and the local semantic map to obtain the posterior probability of each grid. Similar to particle filtering, the matching probability of each grid is determined by calculating the similarity between the two semantic maps. Finally, the pose state corresponding to the grid with the largest posterior probability is selected as the pose change estimate between the current frame's 2D semantic map and the local semantic map, i.e., the first pose constraint.
[0090] Particle filtering and histogram filtering techniques are adaptable to dynamic environments. When dynamic obstacles appear in the environment, the semantic map can be updated promptly, and the filtering algorithm can dynamically adjust the pose estimate to avoid the impact of dynamic obstacles. Furthermore, by using semantic map information as an independent constraint source, combined with particle filtering or histogram filtering techniques, the initial pose estimate can be continuously corrected, improving the robustness of the pose estimate.
[0091] In some embodiments of the present application, the posture constraint includes a second posture constraint, and the posture constraint of the current frame is constructed based on the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the local two-dimensional grid semantic map corresponding to the current frame, including: calculating the posture change from the local two-dimensional grid semantic map to the current frame based on the inertial navigation data; and constructing the second posture constraint based on the posture change from the local two-dimensional grid semantic map to the current frame.
[0092] Another posture constraint constructed in the present application can be derived from inertial navigation data, and the track-reckoning technology is used to perform integral calculations based on the linear acceleration and angular velocity information in the inertial navigation data. First, the linear acceleration is integrated twice to obtain the displacement change of the vehicle, and the angular velocity is integrated once to obtain the rotation angle change of the vehicle. Then, these changes are added to the posture of the previous moment to obtain the posture estimate of the vehicle at the current moment relative to the local two-dimensional grid semantic map. The posture change from the local two-dimensional grid semantic map obtained by track-reckoning to the current frame is expressed as a constraint condition. For example, a transformation matrix can be used to represent the posture transformation relationship from the local two-dimensional grid semantic map coordinate system to the current frame coordinate system. The transformation matrix contains a translation vector and a rotation matrix.
[0093] The constructed second pose constraint is integrated into the pose estimation system together with other pose constraints such as the first pose constraint in the aforementioned embodiment for subsequent pose optimization and fusion.
[0094] Inertial navigation data has a high-frequency output. Dead reckoning technology can calculate the device's position and attitude changes in real time, providing high-frequency position and attitude updates. In a short period of time, dead reckoning technology can provide relatively high pose estimation accuracy. Because inertial navigation data directly reflects the vehicle's motion state and has low cumulative error over a short period of time, dead reckoning results based on inertial navigation data can accurately describe the vehicle's motion within a local area, thereby providing accurate position and attitude constraints.
[0095] In some embodiments of the present application, the pose graph optimization is performed according to the optimized pose of the current frame and the GPS data to obtain the pose graph optimization result, which includes: determining whether the current frame meets the key frame condition according to the optimized pose of the current frame; when the current frame meets the key frame condition, the pose graph optimization is performed according to the optimized pose of the current frame and the GPS data to obtain the pose graph optimization result.
[0096] Before optimizing the pose graph using the current frame's optimized pose, you can first determine whether the current frame meets the keyframe conditions. In semantic map construction, keyframes are selected frames that can significantly reflect environmental characteristics or changes in the target's motion state. By determining keyframe conditions, the amount of data required for processing is reduced, and the speed of map construction is improved. The selection and optimization of keyframes helps cope with complex environments and improves the robustness and accuracy of map construction.
[0097] The conditions for selecting key frames in the embodiment of the present application may include, for example, distance constraints and angle constraints. Among them, the distance constraint refers to setting a distance threshold and calculating the position distance between the current frame and the previous key frame. By comparing the optimized posture (including position information) of the current frame with the position information of the previous key frame, the Euclidean distance between the two is calculated. If the distance is greater than the set threshold, the current frame meets the distance constraint.
[0098] An angle constraint also sets an angle threshold and calculates the change in pose angle between the current frame and the previous keyframe. The pose can be represented using quaternions or rotation matrices, and the angle difference between the two is calculated. If the change in pose angle between the current frame and the previous keyframe is greater than the set angle threshold, the current frame satisfies the angle constraint.
[0099] After obtaining the optimized pose of the current frame, the distance constraint and the angle constraint are checked at the same time. When the current frame satisfies both constraints, it can be considered as a key frame.
[0100] When the current frame is determined to be a keyframe, its optimized pose is added as a node in the pose graph. At the same time, GPS data is used to provide global position information constraints. GPS data provides the device's absolute position in the Earth's coordinate system, which serves as a global reference for the pose graph.
[0101] The pose graph is optimized using pose graph optimization algorithms such as least squares and g2o. These algorithms adjust the position and attitude of pose nodes by minimizing the errors between pose nodes and the global constraint error, making the entire pose graph more accurate and consistent. For example, during the optimization process, the relative motion constraints between pose nodes (such as the motion information between adjacent frames provided by odometer data) and the global position constraints provided by GPS data are considered. The parameters of the pose nodes are continuously adjusted through iterative optimization algorithms to minimize the total error of the pose graph.
[0102] After iterative calculations of the pose graph optimization algorithm, the optimized pose graph is finally obtained. This result contains the optimized pose information of all keyframes, as well as their relative relationships and global position constraints. This optimized pose information can be used for subsequent tasks such as map construction and path planning.
[0103] By setting keyframe conditions to select keyframes, we can ensure that they contain sufficient information and are representative. Keyframes contain important environmental features and motion information. Using these keyframes for pose graph optimization can more accurately estimate the device's pose. Furthermore, keyframes contain rich environmental information. By performing pose graph optimization on keyframes, we can build more accurate and complete large-scale maps.
[0104] In addition, GPS data provides global position information constraints, so that pose graph optimization not only relies on local relative pose information, but also considers global position consistency, which helps to reduce the cumulative error of pose estimation and improve the overall accuracy and robustness of pose estimation.
[0105] In some embodiments of the present application, constructing a global two-dimensional grid semantic map based on the pose graph optimization results includes: updating the local two-dimensional grid semantic map corresponding to the current frame based on the pose graph optimization results; and constructing a global two-dimensional grid semantic map based on the updated local two-dimensional grid semantic map.
[0106] After the pose graph optimization process is complete, the pose graph optimization result is obtained, which contains the optimized pose information of all keyframes, as well as their relative relationships and global position constraints. The pose graph optimization result is parsed to extract the optimized pose data corresponding to each keyframe, including position (x, y, z coordinates) and attitude (represented by quaternions or Euler angles).
[0107] The pose information in the local 2D grid semantic map is updated using the optimized pose data corresponding to each keyframe in the local pose graph optimization results. Based on the optimized keyframe poses, the updated local 2D grid semantic map is spliced with the previous 2D grid semantic map. During the splicing process, the overlapping areas between the maps need to be considered to ensure map continuity and consistency. Because the local 2D grid semantic map integrates image semantic information and laser point cloud intensity information, the splicing of local semantic maps can produce an information-rich and high-precision global 2D grid semantic map.
[0108] By locally updating the two-dimensional grid semantic map, the efficiency of map construction is greatly improved, and by integrating image semantic information and laser point cloud intensity information, the accuracy of semantic map construction is improved.
[0109] The embodiment of the present application also provides a semantic map construction device 200, such as Figure 2As shown, a schematic diagram of the structure of a semantic map construction device in an embodiment of the present application is provided. The semantic map construction device 200 includes at least: an acquisition unit 210, a semantic segmentation unit 220, a first construction unit 230, an optimization unit 240, and a second construction unit 250, wherein:
[0110] An acquisition unit 210 is configured to acquire multi-sensor data, wherein the multi-sensor data includes image data, laser point cloud data, inertial navigation data, and GPS data;
[0111] A semantic segmentation unit 220 is configured to perform semantic segmentation on the road surface element information in the image data to obtain an image semantic segmentation result;
[0112] A first construction unit 230 is configured to perform data association on the laser point cloud data and the image semantic segmentation result to construct a two-dimensional grid semantic map of the current frame;
[0113] An optimization unit 240 is configured to perform pose graph optimization based on the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the GPS data to obtain a pose graph optimization result;
[0114] The second construction unit 250 is used to construct a global two-dimensional grid semantic map according to the pose graph optimization result.
[0115] In some embodiments of the present application, the first construction unit is specifically used to: determine the transformation relationship between the laser point cloud data and the image data; project the laser point cloud data into the image data based on the transformation relationship between the laser point cloud data and the image data; determine the correspondence between the laser point cloud data and the image semantic segmentation result according to the projection result of the laser point cloud data, the correspondence including the image semantic information and laser point cloud intensity information corresponding to each pixel in the image data; perform two-dimensional gridding on the image semantic segmentation result according to the correspondence between the laser point cloud data and the image semantic segmentation result to obtain a two-dimensional grid semantic map of the current frame.
[0116] In some embodiments of the present application, the optimization unit is specifically used to: determine the initial pose of the current frame based on the GPS data; construct the pose constraints of the current frame based on the two-dimensional grid semantic map of the current frame, the inertial navigation data and the local two-dimensional grid semantic map corresponding to the current frame; solve the optimized pose of the current frame using a preset pose optimization algorithm based on the initial pose of the current frame and the pose constraints of the current frame; perform pose graph optimization based on the optimized pose of the current frame and the GPS data to obtain a pose graph optimization result.
[0117] In some embodiments of the present application, the posture constraint includes a first posture constraint, and the optimization unit is specifically used to: calculate the posture change between the two-dimensional grid semantic map of the current frame and the local two-dimensional grid semantic map; and construct the first posture constraint based on the posture change between the two-dimensional grid semantic map of the current frame and the local two-dimensional grid semantic map.
[0118] In some embodiments of the present application, the posture constraint includes a second posture constraint, and the optimization unit is specifically used to: calculate the posture change from the local two-dimensional grid semantic map to the current frame based on the inertial navigation data; and construct the second posture constraint according to the posture change from the local two-dimensional grid semantic map to the current frame.
[0119] In some embodiments of the present application, the optimization unit is specifically used to: determine whether the current frame meets the key frame condition based on the optimized posture of the current frame; when the current frame meets the key frame condition, perform posture graph optimization based on the optimized posture of the current frame and the GPS data to obtain a posture graph optimization result.
[0120] In some embodiments of the present application, the second construction unit is specifically used to: update the local two-dimensional grid semantic map corresponding to the current frame according to the pose graph optimization result; and construct a global two-dimensional grid semantic map based on the updated local two-dimensional grid semantic map.
[0121] It can be understood that the above-mentioned semantic map construction device can implement each step of the semantic map construction method provided in the above-mentioned embodiment. The relevant explanations about the semantic map construction method are applicable to the semantic map construction device and will not be repeated here.
[0122] Figure 3 This is a schematic diagram of the structure of a device in the embodiment of the present application. Figure 3 As shown, the device includes one or more processors (or processing units), may further include one or more memories coupled to the processors, and may further include a communication module coupled to the processors.
[0123] The communication module can be used to communicate with other devices or apparatuses, such as sending or receiving data and / or signals. The communication module can include at least one communication module for communication. The communication module can include any interface necessary for communicating with other devices. Exemplarily, the communication module can be a transceiver, circuit, bus, module, or other type of communication module.
[0124] The processor may include, but is not limited to, at least one of the following: a general-purpose computer, a special-purpose computer, a microcontroller, a digital signal processor (DSP), or one or more of a controller-based multi-core controller architecture. A device may have multiple processors, such as application-specific integrated circuit chips, which are time-slave to a clock synchronized with a main processor.
[0125] The memory may include one or more non-volatile memories and one or more volatile memories. Examples of non-volatile memories include, but are not limited to, at least one of the following: read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, hard disk, compact disc (CD), digital video disc (DVD), or other magnetic storage and / or optical storage. Examples of volatile memories include, but are not limited to, at least one of the following: random access memory (RAM), or other volatile memories that do not persist during a power outage.
[0126] A computer program includes computer-executable instructions that are executed by an associated processor. The program may be stored in ROM. The processor may perform any suitable actions and processes by loading the program into RAM.
[0127] The possible implementation of the present application can be realized by means of a program, so that the communication device can perform any process discussed in the above embodiments. The possible implementation of the present application can also be realized by hardware or by a combination of software and hardware.
[0128] In some embodiments, the program may be tangibly contained in a computer-readable storage medium that may be included in the device (such as in a memory) or other storage device accessible by the device. The program may be loaded from the computer-readable storage medium into RAM for execution. The computer-readable storage medium may include any type of tangible non-volatile memory, such as ROM, EPROM, flash memory, hard disk, CD, DVD, etc.
[0129] The present application also provides a computer-readable storage medium having computer instructions or program codes stored thereon, which, when executed by a processor, causes the processor to perform the methods and functions described in any of the above embodiments. A computer-readable medium may be any tangible medium containing or storing a program for or related to an instruction execution system, apparatus, or device. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. Computer-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination thereof. The computer-readable storage medium may be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more available media integrated therein. More detailed examples of computer-readable storage media include electrical connections with one or more wires, magnetic media (e.g., magnetic disks, floppy disks, hard disks, tapes, magnetic storage devices), optical media (e.g., optical storage devices, DVDs), semiconductor media (e.g., solid-state drives), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), or any suitable combination thereof.
[0130] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The embodiments of the present application also provide at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes one or more computer-executable instructions, such as instructions included in a program module, which are executed in a device on a real or virtual processor of the target to perform the processes, methods and functions involved in any of the above embodiments. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method.
[0131] The present application also provides a computer program product, including a computer program or instructions, which, when run on a computer, causes the computer to perform the processes, methods, and functions in the above-described embodiments. Typically, a program module includes routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of the program modules can be combined or divided between program modules as needed. The machine executable instructions for the program modules can be executed in local or distributed devices. In distributed devices, the program modules can be located in local and remote storage media.
[0132] In general, various embodiments of the present application can be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software, which can be executed by a controller, microprocessor, or other computing device. Although various aspects of the embodiments of the present disclosure are shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that the blocks, devices, systems, techniques, or methods described herein can be implemented as, by way of non-limiting example, hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0133] It should be noted that although the embodiments of the present application are described above in conjunction with the accompanying drawings, the above embodiments are not independent of each other, and they can also be combined to obtain other embodiments. The methods, situations, categories, and divisions of the embodiments in the embodiments of the present application are only for the convenience of description and should not constitute special limitations. The features of the various methods, categories, situations, and embodiments can be combined with each other when they are logical. The various embodiments of the present application can be combined arbitrarily to achieve different technical effects. The embodiments of the present application no longer list various combinations.
[0134] In addition, although the operations of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that these operations must be performed in this particular order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can change the order of execution. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step, and / or one step can be decomposed into multiple steps. It should also be noted that the features and functions of two or more devices according to the present disclosure can be embodied in one device. Conversely, the features and functions of a device described above can be further divided into being embodied by multiple devices.
[0135] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0136] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A semantic map construction method, characterized in that: The semantic map construction method comprises: Acquiring multi-sensor data, wherein the multi-sensor data includes image data, laser point cloud data, inertial navigation data, and GPS data; Performing semantic segmentation on road surface element information in the image data to obtain an image semantic segmentation result; Performing data association on the laser point cloud data and the image semantic segmentation result to construct a two-dimensional grid semantic map of the current frame; Performing pose graph optimization based on the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the GPS data to obtain a pose graph optimization result; A global two-dimensional grid semantic map is constructed according to the pose graph optimization result.
2. The semantic map construction method according to claim 1, characterized in that: The step of associating the laser point cloud data with the image semantic segmentation result to construct a two-dimensional grid semantic map of the current frame includes: Determine the transformation relationship between laser point cloud data and image data; Projecting the laser point cloud data into the image data based on a transformation relationship between the laser point cloud data and the image data; Determining a correspondence between the laser point cloud data and the image semantic segmentation result according to a projection result of the laser point cloud data, wherein the correspondence includes image semantic information corresponding to each pixel point in the image data and laser point cloud intensity information; According to the correspondence between the laser point cloud data and the image semantic segmentation result, the image semantic segmentation result is subjected to two-dimensional gridding processing to obtain a two-dimensional grid semantic map of the current frame.
3. The semantic map construction method according to claim 1, characterized in that: The performing pose graph optimization according to the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the GPS data to obtain a pose graph optimization result includes: Determine the initial position of the current frame according to the GPS data; Constructing a pose constraint for the current frame based on the two-dimensional grid semantic map of the current frame, the inertial navigation data, and a local two-dimensional grid semantic map corresponding to the current frame; According to the initial pose of the current frame and the pose constraints of the current frame, the optimized pose of the current frame is solved using the preset pose optimization algorithm; The pose graph is optimized according to the optimized pose of the current frame and the GPS data to obtain a pose graph optimization result.
4. The semantic map construction method according to claim 3, characterized in that: The pose constraint includes a first pose constraint, and constructing the pose constraint of the current frame according to the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the local two-dimensional grid semantic map corresponding to the current frame includes: Calculating a pose change between the two-dimensional grid semantic map of the current frame and the local two-dimensional grid semantic map; The first pose constraint is constructed according to a pose change between the two-dimensional grid semantic map of the current frame and the local two-dimensional grid semantic map.
5. The semantic map construction method according to claim 3, characterized in that: The pose constraint includes a second pose constraint, and constructing the pose constraint of the current frame according to the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the local two-dimensional grid semantic map corresponding to the current frame includes: Based on the inertial navigation data, calculating the pose change of the local two-dimensional grid semantic map to the current frame; The second pose constraint is constructed according to the pose change from the local two-dimensional grid semantic map to the current frame.
6. The method for constructing a semantic map according to claim 3, wherein: The pose graph optimization is performed according to the optimized pose of the current frame and the GPS data to obtain the pose graph optimization result, including: Determining whether the current frame meets the key frame condition according to the optimized posture of the current frame; When the current frame meets the key frame condition, the pose graph is optimized according to the optimized pose of the current frame and the GPS data to obtain a pose graph optimization result.
7. The semantic map construction method according to claim 1, characterized in that: The constructing of a global two-dimensional grid semantic map according to the pose graph optimization result includes: Update the local two-dimensional grid semantic map corresponding to the current frame according to the pose graph optimization result; A global two-dimensional grid semantic map is constructed based on the updated local two-dimensional grid semantic map.
8. A semantic map construction device, characterized in that: The semantic map construction device comprises: An acquisition unit, configured to acquire multi-sensor data, wherein the multi-sensor data includes image data, laser point cloud data, inertial navigation data, and GPS data; A semantic segmentation unit, configured to perform semantic segmentation on the road surface element information in the image data to obtain an image semantic segmentation result; A first construction unit is configured to perform data association on the laser point cloud data and the image semantic segmentation result to construct a two-dimensional grid semantic map of the current frame; an optimization unit, configured to optimize a pose graph according to the two-dimensional grid semantic map of the current frame, the inertial navigation data, and the GPS data, to obtain a pose graph optimization result; The second construction unit is used to construct a global two-dimensional grid semantic map according to the pose graph optimization result.
9. A device comprising: processor; and a memory arranged to store computer-executable instructions, which, when executed, cause the processor to perform the semantic map construction method according to any one of claims 1 to 7.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the semantic map construction method according to any one of claims 1 to 7 is implemented.