Semantic map construction method and device, electronic equipment and storage medium

By collecting and fusing vehicle positioning information, image panoramic segmentation results, and radar point clouds, and combining high-precision map data to correct errors, a custom semantic map is constructed, which solves the positioning deficiencies of visual SLAM and laser SLAM and achieves high-precision autonomous driving positioning.

CN115493602BActive Publication Date: 2026-04-17ZHIDAO NETWORK TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHIDAO NETWORK TECH (BEIJING) CO LTD
Filing Date
2022-09-28
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing autonomous driving technologies, visual SLAM is affected by dynamic objects and changes in lighting, while laser SLAM is costly and its localization performance degrades in areas with few features, failing to provide high-precision longitudinal position information. Vision-based high-precision map matching and localization elements are too few to provide centimeter-level longitudinal position information over a long period of time.

Method used

By collecting real-time vehicle positioning information, lane line information from high-precision maps, panoramic image segmentation results, and radar point cloud perception fusion, and combining high-precision map data as prior information, a semantic map is constructed. IPM transformation is used to correct camera intrinsic and extrinsic parameter errors, and combined with radar extrinsic parameter adjustment, the LiDAR extrinsic parameters are adaptively adjusted to construct a customized semantic map.

Benefits of technology

It reduces localization errors, provides more accurate semantic maps, ensures high-precision localization in high-speed and feature-sparse areas, and enhances the reliability and accuracy of vertical localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115493602B_ABST
    Figure CN115493602B_ABST
Patent Text Reader

Abstract

The application discloses a semantic map construction method, a vehicle positioning method, a device, an electronic device, and a storage medium. The semantic map construction method comprises the following steps: collecting self-vehicle positioning information, high-precision map lane line information, image panoramic segmentation results, and real-time perception fusion results between visual images and radar point clouds in real time according to a preset collection route; aligning the radar point cloud image panoramic segmentation results and the perception fusion results in time stamps, and associating roadside elements in the image panoramic segmentation results with the real-time perception fusion results; projecting the target elements in the image panoramic segmentation results by combining the self-vehicle positioning information, the high-precision map lane line information, and the association results of the roadside elements in the image panoramic segmentation results and the real-time perception fusion results, and constructing a semantic map. The self-defined semantic map obtained by the semantic map construction method of the application realizes accurate vehicle high-precision positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, and in particular to a semantic map construction method, device, electronic device, and storage medium. Background Technology

[0002] With the development of various vehicle sensors and corresponding positioning technologies, the positioning scheme for autonomous vehicles has been transformed from traditional integrated navigation positioning to multi-sensor fusion positioning, including: laser SLAM, visual SLAM, and vision-based high-precision map matching positioning.

[0003] Typically, multi-sensor fusion positioning frameworks integrate positioning information from multiple positioning technologies, including GNSS / RTK, as observations, along with high-frequency IMU predictions, to output high-frequency absolute positioning information. In this process, the various observations complement and correct each other, mitigating the instability and uncertainty inherent in observations from a single sensor.

[0004] Among related technologies, visual SLAM's feature tracking is affected by dynamic objects, changes in lighting, and excessive vehicle speed, and is only used in conjunction with other positioning technologies in low-speed scenarios such as closed parks and parking lots. Laser SLAM technology is relatively mature, but the cost of lasers is too high to be widely deployed. Furthermore, in areas with sparse or similar features, such as highways and tunnels, laser positioning performance degrades, failing to provide high-precision relocation information.

[0005] Furthermore, while vision-based high-precision map matching and positioning matches traffic elements such as lane lines, arrows, stop lines, or signs identified by semantic recognition models with the high-precision map and calculates the vehicle's relative position on the high-precision map using the PNP algorithm, it still cannot provide centimeter-level vertical position information for extended periods due to the limited number of traffic elements available for longitudinal positioning in the high-precision map. Summary of the Invention

[0006] This application provides a semantic map construction method, a vehicle positioning method, a device and electronic device, and a storage medium to establish a customizable semantic map and achieve high-precision positioning information based on the vertical positioning elements provided by the semantic map.

[0007] The embodiments of this application adopt the following technical solutions:

[0008] In a first aspect, embodiments of this application provide a semantic map construction method, wherein the method includes:

[0009] According to the preset acquisition route, the system collects in real time vehicle positioning information, high-precision map lane line information, image panoramic segmentation results, and real-time perception fusion results between visual images and radar point clouds. The vehicle positioning information includes positioning information with a preset positioning accuracy obtained by post-processing of the positioning post-processing ground truth device or by the combined navigation module. The high-precision map lane line information is extracted from high-precision map data pre-configured by the preset acquisition route. The image panoramic segmentation results include semantic segmentation results of target elements in each frame of the image. The target elements include at least one of the following: target road surface elements and target roadside elements. The real-time perception fusion results between visual images and radar point clouds include the target object and the relative positional relationship between the vehicle and the target object. The target object is the perception fusion result of the camera and radar for the same object.

[0010] The image panoramic segmentation result is timestamped and aligned with the real-time perception fusion result between the visual image and the radar point cloud. The roadside elements in the image panoramic segmentation result are associated with the real-time perception fusion result.

[0011] Using the vehicle positioning information and the lane line information of the high-precision map, combined with the association results of the roadside elements in the panoramic image segmentation results and the real-time perception fusion results, the target elements in the panoramic image segmentation results are projected to construct a semantic map.

[0012] In some embodiments, the roadside elements in the panoramic image segmentation result further include target roadside elements added according to preset requirements. The projection of the target elements in the panoramic image segmentation result using the vehicle positioning information and the high-precision map lane line information, combined with the association result between the roadside elements in the panoramic image segmentation result and the real-time perception fusion result, includes:

[0013] The top view obtained by projecting the target road surface element after IPM transformation is corrected using the lane line information of the high-precision map.

[0014] And / or,

[0015] The target roadside element is projected into the high-precision map data using preset semantic elements in the high-precision map data.

[0016] In some embodiments, correcting the top view obtained after IPM transformation of the target road surface element using the lane line information from the high-precision map includes:

[0017] The parameters of the IPM transform are corrected by using at least one of the following as prior information: the parallel relationship of lane lines, the distance between lane lines, and the relationship between the lane lines and the heading angle of the vehicle in the lane line information of the high-precision map.

[0018] In some embodiments, projecting the target roadside element onto the high-precision map using preset semantic elements in the high-precision map data includes:

[0019] Using some preset semantic elements in the high-precision map data as prior information, and combining the distance information between the vehicle and the target object sensed by the radar, the target roadside element is corrected, and the roadside element is projected onto the corresponding position in the high-precision map.

[0020] In some embodiments, after constructing the semantic map, the method further includes:

[0021] The semantic map is post-processed according to preset requirements;

[0022] If the coverage area of ​​the semantic map is determined to meet the preset positioning conditions, then the established semantic map is used;

[0023] If the coverage area of ​​the semantic map does not meet the preset positioning conditions, the semantic map is thinned or some semantic elements in the map are outlined or vectorized.

[0024] In some embodiments, the panoramic image segmentation result includes a variety of semantic elements, which include at least one of the following: lane lines, road arrows, stop lines, signs, poles, traffic lights, and custom objects.

[0025] Secondly, embodiments of this application also provide a vehicle localization method based on semantic map construction, wherein the semantic map obtained by the semantic map construction method is used, and the vehicle localization method includes:

[0026] During vehicle operation, vehicle localization is performed by matching the results with the semantic map point cloud, which includes preset semantic elements and custom object elements.

[0027] Thirdly, embodiments of this application also provide a semantic map construction apparatus, wherein the apparatus includes:

[0028] The real-time acquisition module is used to acquire vehicle positioning information, high-precision map lane line information, image panoramic segmentation results, and real-time perception fusion results between visual images and radar point clouds in real time according to a preset acquisition route. The image panoramic segmentation results include semantic segmentation results of target elements in each frame of the image. The target elements include at least one of the following: target road surface elements and target roadside elements.

[0029] The post-processing module is used to timestamp-align the panoramic image segmentation result with the real-time perception fusion result between the visual image and the radar point cloud, and to associate the roadside elements in the panoramic image segmentation result with the real-time perception fusion result.

[0030] The mapping module is used to project the target elements in the panoramic image segmentation result using the vehicle positioning information and the lane line information of the high-precision map, combined with the association result of the roadside elements in the panoramic image segmentation result and the real-time perception fusion result, to construct a semantic map.

[0031] Fourthly, embodiments of this application also provide an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, which, when executed, cause the processor to perform the above-described method.

[0032] Fifthly, embodiments of this application also provide a computer-readable storage medium that stores one or more programs, which, when executed by an electronic device including multiple applications, cause the electronic device to perform the above-described method.

[0033] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects:

[0034] In constructing the semantic map, the system first collects vehicle positioning information, high-precision map lane line information, panoramic image segmentation results, and real-time perception fusion results between visual images and radar point clouds according to a preset acquisition route. Then, the panoramic image segmentation results and radar point cloud perception fusion results are timestamped, and roadside elements in the panoramic image segmentation results are associated with the real-time perception fusion results. Finally, using the vehicle positioning information and the high-precision map lane line information, combined with the association results between the roadside elements in the panoramic image segmentation results and the real-time perception fusion results, the target elements in the panoramic image segmentation results are projected to construct the semantic map. By using high-precision map data as prior information for semantic map construction, and combining the association results between the roadside elements in the panoramic image segmentation results and the real-time perception fusion results, the positioning error caused by camera intrinsic and extrinsic parameter errors and the distance error caused by lidar extrinsic parameter errors are reduced, resulting in a more accurate semantic map. Attached Figure Description

[0035] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0036] Figure 1 This is a schematic diagram of a semantic map construction method in an embodiment of this application;

[0037] Figure 2 This is a schematic diagram of a semantic map construction device according to an embodiment of this application;

[0038] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0040] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0041] This application provides a semantic map construction method, such as... Figure 1 The diagram illustrates a semantic map construction method in an embodiment of this application. The method includes at least the following steps S110 to S130:

[0042] Step S110: According to the preset acquisition route, the vehicle positioning information, high-precision map lane line information, image panoramic segmentation results, and real-time perception fusion results between visual images and radar point clouds are acquired in real time. The vehicle positioning information includes positioning information with a preset positioning accuracy obtained by post-processing of the positioning post-processing ground truth device or by the combined navigation module. The high-precision map lane line information is extracted from high-precision map data pre-configured by the preset acquisition route. The image panoramic segmentation results include semantic segmentation results of target elements in each frame of the image. The target elements include at least one of the following: target road surface elements and target roadside elements. The real-time perception fusion results between visual images and radar point clouds include the target object and the relative positional relationship between the vehicle and the target object. The target object is the perception fusion result of the camera and radar for the same object.

[0043] To create a semantic map, you typically need to use a data collection vehicle to collect the necessary information along a set route, then perform post-processing, and finally complete the map creation.

[0044] Therefore, the first step is to set up a data acquisition vehicle to collect real-time information along the designated acquisition route.

[0045] The data acquisition vehicle is equipped with several hardware components, including a positioning and post-processing ground truth device, at least one calibrated monocular camera (with intrinsic and extrinsic parameters), and at least one calibrated lidar (with intrinsic and extrinsic parameters).

[0046] Accurate positioning information can be obtained through the aforementioned post-positioning truth-checking device.

[0047] Using a calibrated monocular camera (with both intrinsic and extrinsic parameters) and a calibrated lidar (with both intrinsic and extrinsic parameters), the distance between the vehicle and the target object can be determined based on the perception fusion results of the two. It is understandable that perception fusion requires pre-labeling of the objects to be fused and training of the fusion model.

[0048] In addition, several software components are required, including a high-precision map of the data collection route, an image-radar perception fusion algorithm, and an image panoramic segmentation algorithm.

[0049] By configuring a high-precision map of the data collection route, the high-precision map data corresponding to the data collection route can be directly obtained, including but not limited to lane lines, stop lines, and other information that can be used for precise positioning.

[0050] An image-radar perception fusion algorithm is configured to perform object perception fusion.

[0051] By configuring an image panoramic segmentation algorithm, semantic segmentation processing is performed on each frame of image captured by the camera to obtain semantic segmentation results.

[0052] Based on the above preparation process, the vehicle's positioning information, high-precision map lane line information, panoramic image segmentation results, and real-time perception fusion results between visual images and radar point clouds can be collected in real time according to the preset collection route.

[0053] Specifically, the vehicle positioning information includes positioning information with a preset positioning accuracy obtained through post-processing by a ground truth device or through a combined navigation module. Considering scenarios such as tunnels and underpasses (if present in the data collection route), the positioning accuracy may not reach centimeter levels, requiring post-processing by a ground truth device to meet the accuracy requirements. However, in ordinary open environments, the location provided by the combined navigation device can be used directly.

[0054] The high-precision map lane line information is extracted from pre-configured high-precision map data along the preset data collection route. In other words, high-precision map lane line information can be extracted from pre-configured high-precision map data. For example, by combining the vehicle's current location information, high-precision map lane line information can be extracted from the pre-configured high-precision map data.

[0055] The panoramic image segmentation result includes the semantic segmentation result of the target elements in each frame of the image. The target elements include at least one of the following: target road surface elements and target roadside elements. That is, there is a corresponding semantic segmentation result for each target element in each frame of the image, and the target elements include target road surface elements and target roadside elements.

[0056] The real-time perception fusion result between the visual image and the radar point cloud includes the target object and the relative positional relationship between the vehicle and the target object. The target object is the perception fusion result of the camera and radar targeting the same object. The real-time perception fusion result can be obtained according to the image-radar perception fusion algorithm. Since the real-time perception fusion result includes the target object and the relative positional relationship between the vehicle and the target object, the target object is the perception fusion result of the camera and radar targeting the same object. For example, the perception fusion result of the vehicle relative to a traffic light ahead, or the perception fusion result of the vehicle relative to a pole on the side of the road.

[0057] Step S120: The image panoramic segmentation result is timestamped and the real-time perception fusion result between the visual image and the radar point cloud is timestamped and the roadside elements in the image panoramic segmentation result are associated with the real-time perception fusion result.

[0058] After acquiring the above data, it is also necessary to align the timestamps of the image panoramic segmentation results with the radar point cloud perception fusion results. This timestamp synchronization ensures that the visual image and the radar point cloud are perception fusion results for the target object at the same time.

[0059] Furthermore, it is necessary to associate the roadside elements in the panoramic image segmentation result with the real-time perception fusion result. The panoramic image segmentation result includes semantic segmentation results of various elements. Here, it is mainly necessary to associate the roadside elements in the panoramic image segmentation result with the acquired real-time perception fusion result. That is, to ensure that the visual image and the radar point cloud are perception fusion results for the same target object at the same time.

[0060] After aligning the timestamps of the perceived data and associating the segmented roadside elements with the perceived fusion results, we can begin to select the semantic elements needed for mapping each frame of panoramic segmentation results.

[0061] Step S130: Using the vehicle positioning information and the lane line information of the high-precision map, combined with the association result of the roadside elements in the panoramic image segmentation result and the real-time perception fusion result, the target elements in the panoramic image segmentation result are projected to construct a semantic map.

[0062] Specifically, during map construction, the vehicle positioning information (centimeter-level positioning) and the lane line information of the high-precision map (centimeter-level positioning) are used as prior conditions. Then, the roadside elements in the panoramic image segmentation result are combined with the correlation results of the real-time perception fusion result to project the relevant target elements in the panoramic image segmentation result. This process is the process of constructing a semantic map.

[0063] It is important to note that the process of building a semantic map described above is for building a local semantic map, not a global semantic map. After the semantic map is built, it is mainly used for the localization of autonomous vehicles.

[0064] Furthermore, since the lane lines in the semantic map established in the above manner have clear dashed outlines and relatively many road surface elements, both lateral and longitudinal positioning can maintain long-term effectiveness and accuracy when using vehicle odometer and Yawrate for prediction.

[0065] Using the above method, a custom localization semantic map is established based on the prior information from a high-precision map and various semantic elements derived from panoramic image segmentation. The vehicle (autonomous driving vehicle) is then localized based on this custom semantic map.

[0066] In one embodiment of this application, the roadside elements in the panoramic image segmentation result further include target roadside elements added according to preset requirements. The projection of the target elements in the panoramic image segmentation result by combining the vehicle positioning information and the lane line information of the high-precision map with the correlation result of the roadside elements in the panoramic image segmentation result and the real-time perception fusion result includes: correcting the top view obtained by projecting the target road surface element after IPM transformation using the lane line information of the high-precision map to reduce the error caused by the intrinsic and extrinsic parameters of the camera; and / or, projecting the target roadside elements into the high-precision map data using preset semantic elements in the high-precision map data to adaptively adjust the extrinsic parameters of the radar.

[0067] In practical implementation, feature tracking in visual SLAM is affected by dynamic objects, changes in lighting, and excessive vehicle speed. While laser SLAM technology is relatively mature, its positioning effectiveness degrades in areas with sparse or similar features, such as highways and tunnels, failing to provide high-precision relocation information. Here, an IPM transformation is used to obtain a top-view of the road surface elements, and the top-view is corrected using lane line data from a high-precision map. This correction primarily aims to reduce errors caused by camera intrinsic and extrinsic parameters. This compensates for the shortcomings of visual SLAM feature tracking, which is affected by dynamic objects, changes in lighting, and excessive vehicle speed. Specifically, the top-view projected from the target road surface elements after IPM transformation is corrected using the lane line information from the high-precision map to reduce errors caused by the camera's intrinsic and extrinsic parameters.

[0068] Furthermore, by combining the distance information provided by the LiDAR, roadside elements can also be projected onto their relative positions on the map, thereby adaptively adjusting the LiDAR's extrinsic parameters. This compensates for the shortcomings of related technologies such as LiDAR SLAM, where the laser positioning effect degrades in areas with sparse or similar features, such as highways and tunnels, failing to provide high-precision relocation information. Specifically, by using preset semantic elements in the high-precision map data, the target roadside element is projected onto the high-precision map data to adaptively adjust the radar's extrinsic parameters.

[0069] In some embodiments, the panoramic image segmentation result includes multiple semantic elements, which include at least one of the following: lane lines, road arrows, stop lines, signs, poles, traffic lights, and custom objects. It is understood that target roadside elements here include, but are not limited to, signs, poles, and traffic lights. Target road surface elements include, but are not limited to, lane lines, road arrows, and stop lines. Furthermore, custom-added target roadside elements include, but are not limited to, flower beds, buildings, bus stop signs, gas stations, etc.

[0070] As can be seen from the above, roadside semantic elements can be customized for mapping and positioning according to user needs, so that vertical positioning has more and more reliable information sources, and the resulting semantic map is a customized semantic map.

[0071] In one embodiment of this application, the step of correcting the top view obtained after the target road surface element has undergone IPM transformation using the high-precision map lane line information includes: using the lane line parallelism, the distance between lane lines, and the relationship between the lane lines and the heading angle of the vehicle in the high-precision map lane line information as prior information to correct the parameters of the IPM transformation.

[0072] In practice, correcting the top view obtained after the target road surface elements undergo IPM transformation is mainly to reduce errors caused by the camera's (monocular) intrinsic and extrinsic parameters. Therefore, using the lane line parallelism, distance between lane lines, and the relationship between lane lines and the vehicle's heading angle from the high-precision map lane line information as prior information, the parameters of the IPM transformation are corrected to obtain a more accurate top view.

[0073] As can be seen from the above, using high-precision map data as prior information for semantic map construction can reduce projection errors caused by camera intrinsic and extrinsic parameters, and at the same time reduce global errors caused by global optimization.

[0074] In one embodiment of this application, the step of projecting the target roadside element into the high-precision map using preset semantic elements in the high-precision map data includes: using some preset semantic elements in the high-precision map data as prior information, combining the distance information between the vehicle and the target object sensed by the radar to correct the target roadside element, and projecting the roadside element into the corresponding position in the high-precision map.

[0075] In practice, by projecting the target roadside element onto the high-precision map data using preset semantic elements, the extrinsic parameters of the LiDAR can be adaptively adjusted. It can be understood that this projection process can also use some semantic elements existing in the high-precision map as prior information to adaptively adjust the extrinsic parameters of the (LiDAR).

[0076] As can be seen from the above, using LiDAR combined with high-precision maps to map roadside elements can prevent semantic element distance errors caused by pure IPM transformation.

[0077] In one embodiment of this application, after constructing the semantic map, the method further includes: post-processing the semantic map according to preset requirements; if it is determined that the coverage area of ​​the semantic map meets the preset positioning conditions, then the established semantic map is used directly; if it is determined that the coverage area of ​​the semantic map does not meet the preset positioning conditions, then the semantic map is thinned or some semantic elements in the map are outlined or vectorized.

[0078] In practice, after the semantic map (containing point cloud information) is established, post-processing can be performed according to actual needs. If the coverage area of ​​the semantic map meets the preset positioning conditions, the established semantic map can be used directly, for example, if the coverage area is small and belongs to a local area.

[0079] Conversely, if the coverage area of ​​the semantic map does not meet the preset positioning conditions, the semantic map is thinned or some semantic elements in the map are outlined or vectorized. For example, if the map coverage area is large and the volume of the semantic point cloud map is too large, the point cloud map can be thinned or some semantic elements can be outlined or vectorized.

[0080] This application also provides a vehicle localization method based on semantic map construction, wherein the semantic map is obtained using the semantic map construction method described above, and the vehicle localization method includes:

[0081] During vehicle operation, vehicle localization is performed by matching the results with the semantic map point cloud, which includes preset semantic elements and custom objects.

[0082] Specifically, based on the semantic map (partial), the vehicle can be located by point cloud matching during vehicle operation. The semantic map includes preset semantic elements and custom objects.

[0083] It is understandable that preset semantic elements can include lane lines, road arrows, stop lines, signs, poles, traffic lights, etc., while custom objects can include flower beds, buildings, bus stops, gas stations, etc. on the road.

[0084] Vehicle localization is achieved through point cloud matching using the semantic map described above. Since the semantic map provides more accurate localization and reduces errors, it can also provide high-precision localization information.

[0085] When constructing a custom semantic map, the data collection vehicle extracts high-precision maps of the corresponding road segments based on its own location. According to the vehicle's positioning information and the lane line information of the high-precision map, the semantic elements after semantic segmentation are projected to generate a custom (multi-semantic information) semantic point cloud map. The vehicle is then positioned based on this point cloud map, which improves the accuracy of positioning.

[0086] This application embodiment also provides a semantic map construction apparatus 200, such as Figure 2 As shown, a schematic diagram of the semantic map construction device in this embodiment is provided. The semantic map construction device 200 includes at least: a real-time acquisition module 210, a post-processing module 220, and a mapping module 230, wherein:

[0087] In one embodiment of this application, the real-time acquisition module 210 is specifically used to: acquire in real time vehicle positioning information, high-precision map lane line information, image panoramic segmentation results, and real-time perception fusion results between visual images and radar point clouds according to a preset acquisition route, wherein the image panoramic segmentation results include semantic segmentation results of target elements in each frame of the image, and the target elements include at least one of the following: target road surface elements and target roadside elements.

[0088] To create a semantic map, you typically need to use a data collection vehicle to collect the necessary information along a set route, then perform post-processing, and finally complete the map creation.

[0089] Therefore, the first step is to set up a data acquisition vehicle to collect real-time information along the designated acquisition route.

[0090] The data acquisition vehicle is equipped with several hardware components, including a positioning and post-processing ground truth device, at least one calibrated monocular camera (with intrinsic and extrinsic parameters), and at least one calibrated lidar (with intrinsic and extrinsic parameters).

[0091] Accurate positioning information can be obtained through the aforementioned post-positioning truth-checking device.

[0092] Using a calibrated monocular camera (with both intrinsic and extrinsic parameters) and a calibrated lidar (with both intrinsic and extrinsic parameters), the distance between the vehicle and the target object can be determined based on the perception fusion results of the two. It is understandable that perception fusion requires pre-labeling of the objects to be fused and training of the fusion model.

[0093] In addition, several software components are required, including a high-precision map of the data collection route, an image-radar perception fusion algorithm, and an image panoramic segmentation algorithm.

[0094] By configuring a high-precision map of the data collection route, the high-precision map data corresponding to the data collection route can be directly obtained, including but not limited to lane lines, stop lines, and other information that can be used for precise positioning.

[0095] An image-radar perception fusion algorithm is configured to perform object perception fusion.

[0096] By configuring an image panoramic segmentation algorithm, semantic segmentation processing is performed on each frame of image captured by the camera to obtain semantic segmentation results.

[0097] Based on the above preparation process, the vehicle's positioning information, high-precision map lane line information, panoramic image segmentation results, and real-time perception fusion results between visual images and radar point clouds can be collected in real time according to the preset collection route.

[0098] Specifically, the vehicle positioning information includes positioning information with a preset positioning accuracy obtained through post-processing by a ground truth device or through a combined navigation module. Considering scenarios such as tunnels and underpasses (if present in the data collection route), the positioning accuracy may not reach centimeter levels, requiring post-processing by a ground truth device to meet the accuracy requirements. However, in ordinary open environments, the location provided by the combined navigation device can be used directly.

[0099] The high-precision map lane line information is extracted from pre-configured high-precision map data along the preset data collection route. In other words, high-precision map lane line information can be extracted from pre-configured high-precision map data. For example, by combining the vehicle's current location information, high-precision map lane line information can be extracted from the pre-configured high-precision map data.

[0100] The panoramic image segmentation result includes the semantic segmentation result of the target elements in each frame of the image. The target elements include at least one of the following: target road surface elements and target roadside elements. That is, there is a corresponding semantic segmentation result for each target element in each frame of the image, and the target elements include target road surface elements and target roadside elements.

[0101] The real-time perception fusion result between the visual image and the radar point cloud includes the target object and the relative positional relationship between the vehicle and the target object. The target object is the perception fusion result of the camera and radar targeting the same object. The real-time perception fusion result can be obtained according to the image-radar perception fusion algorithm. Since the real-time perception fusion result includes the target object and the relative positional relationship between the vehicle and the target object, the target object is the perception fusion result of the camera and radar targeting the same object. For example, the perception fusion result of the vehicle relative to a traffic light ahead, or the perception fusion result of the vehicle relative to a pole on the side of the road.

[0102] In one embodiment of this application, the post-processing module 220 is specifically used to timestamp-align the panoramic image segmentation result with the real-time perception fusion result between the visual image and the radar point cloud, and to associate the roadside elements in the panoramic image segmentation result with the real-time perception fusion result.

[0103] After acquiring the above data, it is also necessary to align the timestamps of the image panoramic segmentation results with the radar point cloud perception fusion results. This timestamp synchronization ensures that the visual image and the radar point cloud are perception fusion results for the target object at the same time.

[0104] Furthermore, it is necessary to associate the roadside elements in the panoramic image segmentation result with the real-time perception fusion result. The panoramic image segmentation result includes semantic segmentation results of various elements. Here, it is mainly necessary to associate the roadside elements in the panoramic image segmentation result with the acquired real-time perception fusion result. That is, to ensure that the visual image and the radar point cloud are perception fusion results for the same target object at the same time.

[0105] After aligning the timestamps of the perceived data and associating the segmented roadside elements with the perceived fusion results, we can begin to select the semantic elements needed for mapping each frame of panoramic segmentation results.

[0106] In one embodiment of this application, the mapping module 230 is specifically used to: project the target element in the panoramic image segmentation result using the vehicle positioning information and the lane line information of the high-precision map, combined with the association result of the roadside elements in the panoramic image segmentation result and the real-time perception fusion result, to construct a semantic map.

[0107] Specifically, during map construction, the vehicle positioning information (centimeter-level positioning) and the lane line information of the high-precision map (centimeter-level positioning) are used as prior conditions. Then, the roadside elements in the panoramic image segmentation result are combined with the correlation results of the real-time perception fusion result to project the relevant target elements in the panoramic image segmentation result. This process is the process of constructing a semantic map.

[0108] It is important to note that the process of building a semantic map described above is for building a local semantic map, not a global semantic map. After the semantic map is built, it is mainly used for the localization of autonomous vehicles.

[0109] Furthermore, since the lane lines in the semantic map established in the above manner have clear dashed outlines and relatively many road surface elements, both lateral and longitudinal positioning can maintain long-term effectiveness and accuracy when using vehicle odometer and Yawrate for prediction.

[0110] Using the aforementioned device, and with prior information from a high-precision map, a custom positioning semantic map is established using various semantic elements derived from panoramic image segmentation. The vehicle (autonomous driving vehicle) is then positioned based on this custom semantic map.

[0111] It is understood that the semantic map construction device described above can implement each step of the semantic map construction method provided in the foregoing embodiments. The relevant explanations of the semantic map construction method are applicable to the semantic map construction device and will not be repeated here.

[0112] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Please refer to it. Figure 3 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations.

[0113] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0114] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.

[0115] The processor reads the corresponding computer program from non-volatile memory into main memory and then runs it, forming a semantic map construction device at the logical level. The processor executes the program stored in memory and specifically performs the following operations:

[0116] According to the preset acquisition route, the vehicle positioning information, high-precision map lane line information, image panoramic segmentation results, and real-time perception fusion results between visual images and radar point clouds are acquired in real time. The image panoramic segmentation results include the semantic segmentation results of target elements in each frame of the image. The target elements include at least one of the following: target road surface elements and target roadside elements.

[0117] The image panoramic segmentation result is timestamped and aligned with the real-time perception fusion result between the visual image and the radar point cloud. The roadside elements in the image panoramic segmentation result are associated with the real-time perception fusion result.

[0118] By combining the vehicle positioning information and the lane line information of the high-precision map, and the correlation results between the roadside elements in the panoramic image segmentation result and the real-time perception fusion result, the target elements in the panoramic image segmentation result are projected to construct a semantic map. This is as described in this application. Figure 1 The semantic map construction apparatus disclosed in the illustrated embodiments can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0119] The electronic device can also perform Figure 1 The method executed by the semantic map building device, and the implementation of the semantic map building device in Figure 1 The functions of the embodiments shown are not described in detail here.

[0120] This application also proposes a computer-readable storage medium that stores one or more programs, the programs including instructions that, when executed by an electronic device including multiple applications, enable the electronic device to perform... Figure 1 The semantic map construction apparatus in the illustrated embodiment executes a method, specifically for performing:

[0121] According to the preset acquisition route, the vehicle positioning information, high-precision map lane line information, image panoramic segmentation results, and real-time perception fusion results between visual images and radar point clouds are acquired in real time. The image panoramic segmentation results include the semantic segmentation results of target elements in each frame of the image. The target elements include at least one of the following: target road surface elements and target roadside elements.

[0122] The image panoramic segmentation result is timestamped and aligned with the real-time perception fusion result between the visual image and the radar point cloud. The roadside elements in the image panoramic segmentation result are associated with the real-time perception fusion result.

[0123] Using the vehicle positioning information and the lane line information of the high-precision map, combined with the association results of the roadside elements in the panoramic image segmentation results and the real-time perception fusion results, the target elements in the panoramic image segmentation results are projected to construct a semantic map.

[0124] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0125] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0126] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0127] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0128] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0129] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0130] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0131] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0132] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0133] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for constructing a custom semantic map, the method comprising: According to the preset acquisition route, the vehicle positioning information, high-precision map lane line information, image panoramic segmentation results, and real-time perception fusion results between visual images and radar point clouds are acquired in real time. The image panoramic segmentation results include the semantic segmentation results of target elements in each frame of the image. The target elements include at least one of the following: target road surface elements and target roadside elements. The image panoramic segmentation result is timestamped and aligned with the real-time perception fusion result between the visual image and the radar point cloud. The roadside elements in the image panoramic segmentation result are associated with the real-time perception fusion result. Using the vehicle positioning information and the lane line information of the high-precision map, combined with the association results of the roadside elements in the panoramic image segmentation results and the real-time perception fusion results, the target elements in the panoramic image segmentation results are projected to construct a local semantic map. The roadside elements in the panoramic image segmentation result also include target roadside elements that are added custom-defined according to preset requirements. The step of projecting the target element in the panoramic image segmentation result using the vehicle positioning information and the lane line information of the high-precision map, combined with the correlation result of the roadside elements in the panoramic image segmentation result and the real-time perception fusion result, includes: Using the lane line information from the high-precision map, the top view obtained by projecting the target road surface element after IPM transformation is corrected. At least one of the following in the lane line information from the high-precision map—the parallel relationship of lane lines, the distance between lane lines, and the relationship between lane lines and the heading angle of the vehicle—is used as prior information to correct the parameters of the IPM transformation. The target roadside element is projected onto the high-precision map data using preset semantic elements in the high-precision map data. Using some preset semantic elements in the high-precision map data as prior information, the target roadside element is corrected by combining the distance information between the vehicle and the target object sensed by the radar. The roadside element is then projected onto the corresponding position in the high-precision map. The preset semantic elements include lane lines, road arrows, stop lines, signs, poles, and traffic lights.

2. The method as described in claim 1, wherein, After constructing the semantic map, the following is also included: The semantic map is post-processed according to preset requirements; If the coverage area of ​​the semantic map is determined to meet the preset positioning conditions, then the established semantic map is used; If the coverage area of ​​the semantic map does not meet the preset positioning conditions, the semantic map is thinned or some semantic elements in the map are outlined or vectorized.

3. The method of claim 1, wherein, The panoramic image segmentation result includes a variety of semantic elements, which include at least one of the following: lane lines, road arrows, stop lines, signs, poles, traffic lights, and custom objects.

4. A vehicle positioning method based on a semantic map construction, wherein, The semantic map obtained by the semantic map construction method as described in any one of claims 1 to 3, wherein the vehicle localization method includes: During vehicle operation, the vehicle is located by matching the results with the semantic map point cloud. The semantic map includes preset semantic elements and custom object elements. The preset semantic elements include lane lines, road arrows, stop lines, signs, poles, and traffic lights.

5. A self-defining semantic map building apparatus wherein, The device includes: The real-time acquisition module is used to acquire vehicle positioning information, high-precision map lane line information, image panoramic segmentation results, and real-time perception fusion results between visual images and radar point clouds in real time according to a preset acquisition route. The image panoramic segmentation results include semantic segmentation results of target elements in each frame of the image. The target elements include at least one of the following: target road surface elements and target roadside elements. The post-processing module is used to timestamp-align the panoramic image segmentation result with the real-time perception fusion result between the visual image and the radar point cloud, and to associate the roadside elements in the panoramic image segmentation result with the real-time perception fusion result. The mapping module is used to project the target elements in the panoramic image segmentation result using the vehicle positioning information and the lane line information of the high-precision map, combined with the association result of the roadside elements in the panoramic image segmentation result and the real-time perception fusion result, to construct a local semantic map. The roadside elements in the panoramic image segmentation result also include target roadside elements that are added custom-defined according to preset requirements. The step of projecting the target element in the panoramic image segmentation result using the vehicle positioning information and the lane line information of the high-precision map, combined with the correlation result of the roadside elements in the panoramic image segmentation result and the real-time perception fusion result, includes: Using the lane line information from the high-precision map, the top view obtained by projecting the target road surface element after IPM transformation is corrected. At least one of the following is used as prior information: the parallel relationship of lane lines, the distance between lane lines, and the relationship between the lane lines and the heading angle of the vehicle. The parameters of the IPM transformation are corrected. Using preset semantic elements in the high-precision map data, the target roadside element is projected into the high-precision map data. Using some preset semantic elements in the high-precision map data as prior information, combined with the distance information between the vehicle and the target object sensed by the radar, the target roadside element is corrected, and the roadside element is projected into the corresponding position in the high-precision map. The preset semantic elements include lane lines, road arrows, stop lines, signs, poles, and traffic lights.

6. An electronic device, comprising: processor; as well as A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method of any one of claims 1 to 3.

7. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of applications, cause the electronic device to perform the method of any one of claims 1 to 3.

Citation Information

Patent Citations

  • Map data updating method and device, electronic equipment and storage medium

    CN111797187A

  • Navigation information display method and device, lane line tracking method and device and storage medium

    CN113705305A