Semantic map construction methods, devices and electronic equipment, and storage media
By combining multi-dimensional optimization models of image, laser, inertial navigation and satellite positioning data, a high-precision semantic map is constructed, which solves the problems of insufficient construction efficiency and accuracy in existing technologies and improves the positioning capabilities of autonomous vehicles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHIDAO NETWORK TECH (BEIJING) CO LTD
- Filing Date
- 2023-03-17
- Publication Date
- 2026-05-26
AI Technical Summary
Existing semantic map building technologies are insufficient in terms of building efficiency, map accuracy, and storage efficiency, making it difficult to meet the high-precision positioning requirements of autonomous vehicles.
By combining image data, laser data, inertial navigation data, and satellite positioning data, a multi-dimensional nonlinear optimization model is constructed to optimize vehicle pose. Furthermore, three-dimensional semantic information is constructed by combining image and laser data, and a post-processing strategy is adopted to optimize the semantic map.
It improves the accuracy and efficiency of semantic map construction, reduces the amount of semantic map data, and enhances the positioning accuracy and safety of autonomous vehicles.
Smart Images

Figure CN116310174B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of map construction technology, and in particular to a semantic map construction method, apparatus, electronic device, and storage medium. Background Technology
[0002] In the field of autonomous driving, in addition to relying on their onboard sensors for pathfinding and obstacle avoidance, autonomous vehicles also need a sufficiently accurate map to determine their location and plan their routes.
[0003] Semantic maps have wide applications in driver assistance and autonomous driving, and are of great significance for ensuring the positioning accuracy and driving safety of autonomous vehicles.
[0004] However, semantic maps currently built based on technologies such as laser SLAM face challenges related to construction efficiency, map accuracy, map size, and storage efficiency. Summary of the Invention
[0005] This application provides a semantic map construction method, apparatus, electronic device, and storage medium to improve the accuracy and efficiency of semantic map construction.
[0006] The embodiments of this application adopt the following technical solutions:
[0007] In a first aspect, embodiments of this application provide a semantic map construction method, wherein the method includes:
[0008] Obtain the source data for the semantic map, which includes image data, laser data, inertial navigation data, and satellite positioning data;
[0009] The vehicle pose corresponding to the image data is determined based on the image data and the inertial navigation data, and the vehicle pose corresponding to the laser data is determined based on the laser data and the inertial navigation data.
[0010] A first nonlinear optimization model is constructed based on the vehicle pose corresponding to the image data, the vehicle pose corresponding to the laser data, and the satellite positioning data, and the optimized vehicle pose is determined based on the first nonlinear optimization model.
[0011] Determine the three-dimensional semantic information corresponding to the image data based on the image data and the laser data;
[0012] A semantic map is constructed based on the optimized vehicle pose and the corresponding 3D semantic information of the image data.
[0013] Optionally, determining the vehicle pose corresponding to the image data based on the image data and the inertial navigation data includes:
[0014] The image data is semantically segmented using a preset semantic segmentation model to obtain semantic segmentation results, and the image data is feature extracted using a preset image feature extraction algorithm to obtain image feature extraction results.
[0015] A second nonlinear optimization model is constructed based on the semantic segmentation results, the image feature extraction results, and the inertial navigation data.
[0016] The vehicle pose corresponding to the image data is determined based on the second nonlinear optimization model.
[0017] Optionally, the semantic segmentation result is the semantic segmentation result of the current frame, the image feature extraction result is the image feature extraction result of the current frame, and the step of constructing a second nonlinear optimization model based on the semantic segmentation result, the image feature extraction result, and the inertial navigation data includes:
[0018] Obtain the semantic segmentation result of the previous frame, and construct semantic constraints based on the semantic segmentation result of the current frame and the semantic segmentation result of the previous frame;
[0019] Obtain the image feature extraction result of the previous frame, and construct image feature constraints based on the image feature extraction result of the current frame and the image feature extraction result of the previous frame;
[0020] Obtain the inertial navigation data between the image data of the current frame and the image data of the previous frame, and construct the inertial navigation pre-integration constraints corresponding to the image data based on the inertial navigation data between the image data of the current frame and the image data of the previous frame;
[0021] The second nonlinear optimization model is constructed based on the semantic constraints, the image feature constraints, and the inertial navigation pre-integration constraints corresponding to the image data.
[0022] Optionally, determining the vehicle pose corresponding to the laser data based on the laser data and the inertial navigation data includes:
[0023] The laser data is subjected to feature extraction using a preset point cloud feature extraction algorithm to obtain point cloud feature extraction results;
[0024] A third nonlinear optimization model is constructed based on the point cloud feature extraction results and the inertial navigation data.
[0025] The vehicle pose corresponding to the laser data is determined based on the third nonlinear optimization model.
[0026] Optionally, the point cloud feature extraction result is the point cloud feature extraction result of the current frame, and the step of constructing a third nonlinear optimization model based on the point cloud feature extraction result and the inertial navigation data includes:
[0027] Obtain the point cloud feature extraction result of the previous frame, and construct point cloud feature constraints based on the point cloud feature extraction result of the current frame and the point cloud feature extraction result of the previous frame.
[0028] Obtain the inertial navigation data between the laser data of the current frame and the laser data of the previous frame, and construct the inertial navigation pre-integration constraint corresponding to the laser data based on the inertial navigation data between the laser data of the current frame and the laser data of the previous frame;
[0029] The third nonlinear optimization model is constructed based on the point cloud feature constraints and the inertial navigation pre-integration constraints corresponding to the laser data.
[0030] Optionally, constructing the first nonlinear optimization model based on the vehicle pose corresponding to the image data, the vehicle pose corresponding to the laser data, and the satellite positioning data includes:
[0031] Construct relative change constraints between image data based on the vehicle pose corresponding to the image data;
[0032] Based on the vehicle pose corresponding to the laser data, construct the relative change constraints between the laser data;
[0033] Construct relative variation constraints between satellite positioning data based on the satellite positioning data;
[0034] The first nonlinear optimization model is constructed based on the relative variation constraints between the image data, the relative variation constraints between the laser data, and the relative variation constraints between the satellite positioning data.
[0035] Optionally, after constructing a semantic map based on the optimized vehicle pose and the 3D semantic information corresponding to the image data, the method further includes:
[0036] The semantic map is post-processed using a preset post-processing strategy to obtain a post-processed semantic map. The preset post-processing strategy includes performing voxel filtering on the semantic map and / or determining the semantic map that is outside the vehicle's current field of view and extracting semantic information contours from the semantic map that is outside the vehicle's current field of view.
[0037] Secondly, embodiments of this application also provide a semantic map construction apparatus, wherein the apparatus includes:
[0038] The acquisition unit is used to acquire the source data of the semantic map, which includes image data, laser data, inertial navigation data and satellite positioning data;
[0039] The first determining unit is configured to determine the vehicle pose corresponding to the image data based on the image data and the inertial navigation data, and to determine the vehicle pose corresponding to the laser data based on the laser data and the inertial navigation data.
[0040] The second determining unit is used to construct a first nonlinear optimization model based on the vehicle pose corresponding to the image data, the vehicle pose corresponding to the laser data, and the satellite positioning data, and to determine the optimized vehicle pose based on the first nonlinear optimization model.
[0041] The third determining unit is used to determine the three-dimensional semantic information corresponding to the image data based on the image data and the laser data;
[0042] The construction unit is used to construct a semantic map based on the optimized vehicle pose and the three-dimensional semantic information corresponding to the image data.
[0043] Thirdly, embodiments of this application also provide an electronic device, including:
[0044] Processor; and
[0045] A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform any of the methods described above.
[0046] Fourthly, embodiments of this application also provide a computer-readable storage medium that stores one or more programs, which, when executed by an electronic device including multiple applications, cause the electronic device to perform any of the methods described above.
[0047] The at least one technical solution adopted in this application embodiment can achieve the following beneficial effects: The semantic map construction method of this application embodiment first acquires the source data of the semantic map, which includes image data, laser data, inertial navigation data, and satellite positioning data; then, it determines the vehicle pose corresponding to the image data based on the image data and inertial navigation data, and determines the vehicle pose corresponding to the laser data based on the laser data and inertial navigation data; then, it constructs a first nonlinear optimization model based on the vehicle pose corresponding to the image data, the vehicle pose corresponding to the laser data, and the satellite positioning data, and determines the optimized vehicle pose based on the first nonlinear optimization model; then, it determines the three-dimensional semantic information corresponding to the image data based on the image data and laser data; finally, it constructs a semantic map based on the optimized vehicle pose and the three-dimensional semantic information corresponding to the image data. The semantic map construction method of this application embodiment combines data from multiple sensors to optimize the vehicle pose, improving the construction accuracy of the semantic map, and combines image data and laser data to construct the semantic map, improving the construction efficiency of the semantic map. Attached Figure Description
[0048] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0049] Figure 1 This is a flowchart illustrating a semantic map construction method according to an embodiment of this application;
[0050] Figure 2 This is a schematic diagram of the structure of a semantic map construction device according to an embodiment of this application;
[0051] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0053] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0054] This application provides a semantic map construction method, such as... Figure 1 The diagram illustrates a flowchart of a semantic map construction method according to an embodiment of this application. The method includes at least the following steps S110 to S150:
[0055] Step S110: Obtain the source data of the semantic map, which includes image data, laser data, inertial navigation data and satellite positioning data.
[0056] In constructing a semantic map, this embodiment requires first acquiring source map data for the semantic map. This data can include image data captured by a camera, laser data collected by a LiDAR, inertial navigation data output by an IMU, and satellite positioning data output by GPS. Since different sensors have different data output frequencies, the image data, laser data, inertial navigation data, and satellite positioning data can be spatiotemporally synchronized to ensure the accuracy of subsequent processing. Specifically, data from the sensor with the lowest output frequency can be used as a reference.
[0057] Step S120: Determine the vehicle pose corresponding to the image data based on the image data and the inertial navigation data, and determine the vehicle pose corresponding to the laser data based on the laser data and the inertial navigation data.
[0058] Image data and laser data can provide the information needed in semantic maps from different dimensions. Therefore, in this embodiment, the vehicle pose of the current frame can be determined from the dimensions of image data and laser data respectively. That is, a nonlinear optimization model of vehicle pose can be constructed based on image data and inertial navigation data. The vehicle pose calculated based on image data can be obtained by solving the model. A nonlinear optimization model of vehicle pose can be constructed based on laser data and inertial navigation data. The vehicle pose calculated based on laser data can be obtained by solving the model. Finally, the vehicle poses calculated from these two dimensions are used as the basis for subsequent joint optimization.
[0059] Step S130: Construct a first nonlinear optimization model based on the vehicle pose corresponding to the image data, the vehicle pose corresponding to the laser data, and the satellite positioning data, and determine the optimized vehicle pose based on the first nonlinear optimization model.
[0060] Based on the vehicle pose corresponding to the image data and the vehicle pose corresponding to the laser data obtained in the aforementioned steps, a joint nonlinear optimization model is further constructed by combining satellite positioning data. Existing nonlinear least squares solution algorithms such as Gauss-Newton (GN) and Levenberg-Marquardt (LM) are used to solve the nonlinear optimization model, thereby obtaining the final optimized vehicle pose. That is, the optimized vehicle pose is obtained by jointly optimizing sensor positioning data from multiple dimensions such as images, laser point clouds, inertial navigation, and satellite positioning. Therefore, the calculation accuracy of vehicle pose is improved, thereby improving the construction accuracy of semantic maps.
[0061] Step S140: Determine the three-dimensional semantic information corresponding to the image data based on the image data and the laser data.
[0062] Vehicle localization based on semantic maps is absolute localization. Therefore, the constructed semantic map needs to provide three-dimensional semantic information. In this embodiment, the laser point cloud can be projected onto the image based on the pre-calibrated extrinsic parameters between the camera and the LiDAR to obtain the projection result of the laser point cloud on the image. Then, it is matched with the two-dimensional semantic information extracted from the image to obtain the 3D coordinates corresponding to all two-dimensional semantic information in the image.
[0063] Since laser point clouds are sparser than image data, not all two-dimensional semantic pixels extracted from an image can be matched with corresponding three-dimensional information. Therefore, this application embodiment can perform equation fitting based on the two-dimensional semantic pixels that can be matched with corresponding three-dimensional information, and then solve for the three-dimensional information corresponding to more other two-dimensional semantic pixels based on the fitted equation. In other words, it is equivalent to assigning three-dimensional information to all two-dimensional semantic pixels in the image based on the projection result of laser point cloud data in the image, rather than directly assigning two-dimensional semantic information from the image to laser point cloud data, thereby improving the richness of three-dimensional semantic information and the accuracy of the semantic map.
[0064] Step S150: Construct a semantic map based on the optimized vehicle pose and the three-dimensional semantic information corresponding to the image data.
[0065] Based on the vehicle pose, the relative transformation relationship between the semantic map information of two adjacent frames can be determined. Therefore, based on the optimized vehicle pose, the three-dimensional semantic information corresponding to the image data can be stitched together to construct a semantic map, which can be continuously updated as subsequent map source data is collected.
[0066] The semantic map construction method of this application combines data from multiple sensors to optimize vehicle pose, thereby improving the accuracy of semantic map construction. By combining image data and laser data to construct the semantic map, the construction efficiency of the semantic map is improved.
[0067] In some embodiments of this application, determining the vehicle pose corresponding to the image data based on the image data and the inertial navigation data includes: performing semantic segmentation on the image data using a preset semantic segmentation model to obtain a semantic segmentation result; and performing feature extraction on the image data using a preset image feature extraction algorithm to obtain an image feature extraction result; constructing a second nonlinear optimization model based on the semantic segmentation result, the image feature extraction result, and the inertial navigation data; and determining the vehicle pose corresponding to the image data based on the second nonlinear optimization model.
[0068] When calculating the vehicle pose corresponding to the image data based on image data and inertial navigation data, the semantic information and visual features in the image data can be segmented and extracted first. That is, on the one hand, the existing semantic segmentation model can be used to perform semantic segmentation on the image data to obtain semantic segmentation results such as lane lines, arrows, stop lines, and 3D markings in the image. On the other hand, the existing image feature extraction algorithm can be used to extract features from the image data to obtain image feature extraction results such as corner features in the image.
[0069] The aforementioned semantic segmentation model can, for example, employ the LaneAF model. LaneAF is primarily based on semantic segmentation combined with clustering post-processing techniques. First, semantic segmentation is used to binary classify pixels, determining whether they belong to lane lines or the background. Then, lane line pixels are clustered to generate different lane line instances. Image feature extraction algorithms are mainly used to extract corner features from images; therefore, existing corner feature extraction algorithms such as SIFT, SURF, and FAST can be used. Of course, those skilled in the art can flexibly choose the specific method for semantic segmentation and feature extraction based on existing technologies, and no specific limitations are imposed here.
[0070] Based on the semantic segmentation results, image feature extraction results, and inertial navigation data, a second nonlinear optimization model can be constructed. Then, VIO (Visual-Inertial Odometry) is used for optimization. Specifically, the existing nonlinear least squares solution algorithm can be used to solve the second nonlinear optimization model to obtain the optimized vehicle pose, which is the vehicle pose calculated based on image data.
[0071] In some embodiments of this application, the semantic segmentation result is the semantic segmentation result of the current frame, the image feature extraction result is the image feature extraction result of the current frame, and the step of constructing a second nonlinear optimization model based on the semantic segmentation result, the image feature extraction result, and the inertial navigation data includes: obtaining the semantic segmentation result of the previous frame and constructing semantic constraints based on the semantic segmentation result of the current frame and the semantic segmentation result of the previous frame; obtaining the image feature extraction result of the previous frame and constructing image feature constraints based on the image feature extraction result of the current frame and the image feature extraction result of the previous frame; obtaining the inertial navigation data between the image data of the current frame and the image data of the previous frame and constructing inertial navigation pre-integration constraints corresponding to the image data based on the inertial navigation data between the image data of the current frame and the image data of the previous frame; and constructing the second nonlinear optimization model based on the semantic constraints, the image feature constraints, and the inertial navigation pre-integration constraints corresponding to the image data.
[0072] In this embodiment of the application, when constructing the second nonlinear optimization model, the terms to be optimized and the constraints of the second nonlinear optimization model can be determined first. The terms to be optimized are the vehicle pose, and the constraints can specifically include:
[0073] 1) Semantic constraints: When constructing semantic constraints, it is necessary to obtain the semantic segmentation result of the previous frame. By matching the semantic segmentation result of the current frame with the semantic segmentation result of the previous frame, the relative semantic transformation relationship between adjacent frames is calculated and used as semantic constraints.
[0074] 2) Image feature constraints: When constructing image feature constraints, it is necessary to obtain the image feature extraction results of the previous frame. By matching the image feature extraction results of the current frame with the image feature extraction results of the previous frame, the image feature transformation relationship between adjacent frames is calculated and used as the image feature constraints.
[0075] 3) Inertial navigation pre-integration constraints: Since the output frequency of inertial navigation data is different from that of image data, and the output frequency of inertial navigation data is usually higher, when calculating vehicle pose based on image data, the output frequency of image data can be used as a reference to obtain all inertial navigation data between corresponding moments of two adjacent image frames for pre-integration calculation. For example, if there are 30 IMU measurements (M_0~M_29) between two image frames, then starting from the moment of the previous image frame, the IMU measurements are continuously integrated: M_0*delta_t+M_1*delta_t+...+M_29*delta_t, and finally the relative pose change between the two frames is obtained, which serves as the inertial navigation pre-integration constraint corresponding to the image data.
[0076] In some embodiments of this application, determining the vehicle pose corresponding to the laser data based on the laser data and the inertial navigation data includes: extracting features from the laser data using a preset point cloud feature extraction algorithm to obtain point cloud feature extraction results; constructing a third nonlinear optimization model based on the point cloud feature extraction results and the inertial navigation data; and determining the vehicle pose corresponding to the laser data based on the third nonlinear optimization model.
[0077] When calculating the vehicle pose corresponding to the laser data based on laser data and inertial navigation data, existing point cloud feature extraction algorithms such as the LeGO-LOAM algorithm can be used to extract features from the laser data first. Based on the point cloud feature extraction results and inertial navigation data, a third nonlinear optimization model can be constructed. Similarly, existing nonlinear least squares solution algorithms can be used to solve the third nonlinear optimization model to obtain the optimized vehicle pose, which is the vehicle pose calculated based on the laser data.
[0078] A third nonlinear optimization model is constructed based on the point cloud feature extraction results and inertial navigation data. Then, LIO (Lidar-Inertial Odometry) is used for optimization. Similarly, the existing nonlinear least squares solution algorithm can be used to solve the second nonlinear optimization model to obtain the optimized vehicle pose, which is the vehicle pose calculated based on laser data.
[0079] In some embodiments of this application, the point cloud feature extraction result is the point cloud feature extraction result of the current frame. The step of constructing a third nonlinear optimization model based on the point cloud feature extraction result and the inertial navigation data includes: obtaining the point cloud feature extraction result of the previous frame, and constructing point cloud feature constraints based on the point cloud feature extraction result of the current frame and the point cloud feature extraction result of the previous frame; obtaining inertial navigation data between the laser data of the current frame and the laser data of the previous frame, and constructing inertial navigation pre-integration constraints corresponding to the laser data based on the inertial navigation data between the laser data of the current frame and the laser data of the previous frame; and constructing the third nonlinear optimization model based on the point cloud feature constraints and the inertial navigation pre-integration constraints corresponding to the laser data.
[0080] In this embodiment of the application, when constructing the second nonlinear optimization model, the terms to be optimized and the constraints of the second nonlinear optimization model can be determined first. The terms to be optimized are the vehicle pose, and the constraints can specifically include:
[0081] 1) Point cloud feature constraints: When constructing point cloud feature constraints, it is necessary to obtain the point cloud feature extraction results of the previous frame. By matching the point cloud feature extraction results of the current frame with the point cloud feature extraction results of the previous frame, the point cloud feature transformation relationship between adjacent frames is calculated and used as the point cloud feature constraints.
[0082] 2) Inertial navigation pre-integration constraint: Since the output frequency of inertial navigation data is different from that of laser data, the output frequency of inertial navigation data is usually higher. Therefore, when calculating the vehicle pose based on laser data, the output frequency of laser data can be used as a reference to obtain all inertial navigation data between the corresponding times of two adjacent laser data frames for pre-integration calculation. That is, starting from the time of the previous laser frame, the IMU measurement value is continuously integrated to finally obtain the relative pose change between the two frames, which serves as the inertial navigation pre-integration constraint corresponding to the laser data.
[0083] In some embodiments of this application, the step of constructing a first nonlinear optimization model based on the vehicle pose corresponding to the image data, the vehicle pose corresponding to the laser data, and the satellite positioning data includes: constructing relative change constraints between image data based on the vehicle pose corresponding to the image data; constructing relative change constraints between laser data based on the vehicle pose corresponding to the laser data; constructing relative change constraints between satellite positioning data based on the satellite positioning data; and constructing the first nonlinear optimization model based on the relative change constraints between image data, the relative change constraints between laser data, and the relative change constraints between satellite positioning data.
[0084] In this embodiment of the application, when constructing the first nonlinear optimization model, the terms to be optimized and the constraints of the first nonlinear optimization model can also be determined first. The terms to be optimized are the vehicle pose, and the constraints can specifically include:
[0085] 1) Relative change constraints between image data: By matching the vehicle pose corresponding to the image data of the current frame with the optimized vehicle pose of the previous frame, the relative pose transformation relationship of the vehicle pose corresponding to the image data of the current frame relative to the previous frame can be calculated, which serves as the relative change constraint between image data.
[0086] 2) Relative change constraints between laser data: By matching the vehicle pose corresponding to the laser data of the current frame with the optimized vehicle pose of the previous frame, the relative pose transformation relationship of the vehicle pose corresponding to the laser data of the current frame relative to the previous frame can be calculated, which serves as the relative change constraint between laser data.
[0087] 3) Relative change constraints between satellite positioning data: By matching the vehicle pose corresponding to the satellite positioning data of the current frame with the optimized vehicle pose of the previous frame, the relative pose transformation relationship of the vehicle pose corresponding to the satellite positioning data of the current frame relative to the previous frame can be calculated, which serves as the relative change constraint between satellite positioning data.
[0088] By jointly optimizing vehicle pose based on the constraint information of the above-mentioned dimensions, the information that cameras, LiDAR and GPS can provide under different road scenarios is fully considered, thereby improving the accuracy and efficiency of semantic map construction.
[0089] In some embodiments of this application, determining the three-dimensional semantic information corresponding to the image data based on the image data and the laser data includes: projecting the laser data into the image data according to the calibration relationship between the camera and the lidar to obtain the projection result of the laser data in the image data; matching the projection result of the laser data in the image data with the two-dimensional semantic information in the image data to obtain the 3D coordinates of the first two-dimensional semantic information in the image data; performing equation fitting on the 3D coordinates corresponding to the first two-dimensional semantic information in the image data using a preset fitting algorithm; and determining the 3D coordinates of the second two-dimensional semantic information in the image data based on the equation fitting result.
[0090] Vehicle localization based on semantic maps is absolute localization. Therefore, the constructed semantic map needs to provide three-dimensional semantic information. In this embodiment, the laser point cloud can be projected onto the image based on the pre-calibrated extrinsic parameters between the camera and the LiDAR to obtain the projection result of the laser point cloud on the image. The projection result of the laser point cloud on the image is matched with the two-dimensional semantic information extracted from the image to obtain the 3D coordinates corresponding to all two-dimensional semantic information in the image.
[0091] Since laser point clouds are sparser than image data, not all two-dimensional semantic pixels extracted from an image can be matched with corresponding three-dimensional information. Therefore, this application embodiment can perform equation fitting based on the two-dimensional semantic pixels that can be matched with corresponding three-dimensional information, and then solve for the three-dimensional information corresponding to more other two-dimensional semantic pixels based on the fitted equation. In other words, it is equivalent to assigning three-dimensional information to all two-dimensional semantic pixels in the image based on the projection result of laser point cloud data in the image, rather than directly assigning two-dimensional semantic information from the image to laser point cloud data, thereby improving the richness of three-dimensional semantic information and the accuracy of the semantic map.
[0092] In some embodiments of this application, after constructing a semantic map based on the optimized vehicle pose and the three-dimensional semantic information corresponding to the image data, the method further includes: post-processing the semantic map using a preset post-processing strategy to obtain a post-processed semantic map. The preset post-processing strategy includes performing voxel filtering on the semantic map and / or determining the semantic map beyond the vehicle's current field of view and extracting semantic information contours from the semantic map beyond the vehicle's current field of view.
[0093] As the information contained in the semantic map increases during the map building process, the data volume of the semantic map becomes larger and larger, resulting in low storage efficiency. When positioning based on the semantic map, the map loading efficiency also decreases and the processing time increases, thus affecting the real-time positioning performance.
[0094] Based on this, the embodiments of this application can perform some post-processing operations on the semantic map constructed in the aforementioned embodiments. For example, Voxel voxel filtering can be used to process the semantic map to reduce the amount of data in the local semantic map. Semantic information contours can also be extracted from the local semantic map corresponding to the area that is beyond the field of view, that is, the area behind the vehicle that has been driven past, thereby further reducing the amount of data in the semantic map, improving map storage efficiency, and improving the efficiency of simultaneous mapping of multiple vehicles.
[0095] This application also provides a semantic map construction apparatus 200, such as... Figure 2 As shown, a schematic diagram of a semantic map construction device according to an embodiment of this application is provided. The device 200 includes: an acquisition unit 210, a first determination unit 220, a second determination unit 230, a third determination unit 240, and a construction unit 250, wherein:
[0096] The acquisition unit 210 is used to acquire the map source data of the semantic map, wherein the map source data includes image data, laser data, inertial navigation data and satellite positioning data;
[0097] The first determining unit 220 is configured to determine the vehicle pose corresponding to the image data based on the image data and the inertial navigation data, and to determine the vehicle pose corresponding to the laser data based on the laser data and the inertial navigation data.
[0098] The second determining unit 230 is used to construct a first nonlinear optimization model based on the vehicle pose corresponding to the image data, the vehicle pose corresponding to the laser data, and the satellite positioning data, and to determine the optimized vehicle pose based on the first nonlinear optimization model.
[0099] The third determining unit 240 is used to determine the three-dimensional semantic information corresponding to the image data based on the image data and the laser data;
[0100] The construction unit 250 is used to construct a semantic map based on the optimized vehicle pose and the three-dimensional semantic information corresponding to the image data.
[0101] In some embodiments of this application, the first determining unit 220 is specifically used for: performing semantic segmentation on the image data using a preset semantic segmentation model to obtain a semantic segmentation result, and performing feature extraction on the image data using a preset image feature extraction algorithm to obtain an image feature extraction result; constructing a second nonlinear optimization model based on the semantic segmentation result, the image feature extraction result, and the inertial navigation data; and determining the vehicle pose corresponding to the image data based on the second nonlinear optimization model.
[0102] In some embodiments of this application, the semantic segmentation result is the semantic segmentation result of the current frame, and the image feature extraction result is the image feature extraction result of the current frame. The first determining unit 220 is specifically used for: obtaining the semantic segmentation result of the previous frame, and constructing semantic constraints based on the semantic segmentation result of the current frame and the semantic segmentation result of the previous frame; obtaining the image feature extraction result of the previous frame, and constructing image feature constraints based on the image feature extraction result of the current frame and the image feature extraction result of the previous frame; obtaining inertial navigation data between the image data of the current frame and the image data of the previous frame, and constructing inertial navigation pre-integration constraints corresponding to the image data based on the inertial navigation data between the image data of the current frame and the image data of the previous frame; and constructing the second nonlinear optimization model based on the semantic constraints, the image feature constraints, and the inertial navigation pre-integration constraints corresponding to the image data.
[0103] In some embodiments of this application, the first determining unit 220 is specifically used to: extract features from the laser data using a preset point cloud feature extraction algorithm to obtain point cloud feature extraction results; construct a third nonlinear optimization model based on the point cloud feature extraction results and the inertial navigation data; and determine the vehicle pose corresponding to the laser data based on the third nonlinear optimization model.
[0104] In some embodiments of this application, the point cloud feature extraction result is the point cloud feature extraction result of the current frame. The first determining unit 220 is specifically used to: obtain the point cloud feature extraction result of the previous frame, and construct point cloud feature constraints based on the point cloud feature extraction result of the current frame and the point cloud feature extraction result of the previous frame; obtain inertial navigation data between the laser data of the current frame and the laser data of the previous frame, and construct inertial navigation pre-integration constraints corresponding to the laser data based on the inertial navigation data between the laser data of the current frame and the laser data of the previous frame; and construct the third nonlinear optimization model based on the point cloud feature constraints and the inertial navigation pre-integration constraints corresponding to the laser data.
[0105] In some embodiments of this application, the second determining unit 230 is specifically used to: construct relative change constraints between image data based on the vehicle pose corresponding to the image data; construct relative change constraints between laser data based on the vehicle pose corresponding to the laser data; construct relative change constraints between satellite positioning data based on the satellite positioning data; and construct the first nonlinear optimization model based on the relative change constraints between image data, the relative change constraints between laser data, and the relative change constraints between satellite positioning data.
[0106] In some embodiments of this application, the apparatus further includes: a post-processing unit, configured to, after constructing a semantic map based on the optimized vehicle pose and the three-dimensional semantic information corresponding to the image data, perform post-processing on the semantic map using a preset post-processing strategy to obtain a post-processed semantic map, wherein the preset post-processing strategy includes performing voxel filtering on the semantic map, and / or, determining the semantic map beyond the vehicle's current field of view and extracting semantic information contours from the semantic map beyond the vehicle's current field of view.
[0107] It is understood that the semantic map construction device described above can implement each step of the semantic map construction method provided in the foregoing embodiments. The relevant explanations of the semantic map construction method are applicable to the semantic map construction device and will not be repeated here.
[0108] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Please refer to it. Figure 3 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations.
[0109] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0110] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.
[0111] The processor reads the corresponding computer program from non-volatile memory into main memory and then runs it, forming a semantic map construction device at the logical level. The processor executes the program stored in memory and specifically performs the following operations:
[0112] Obtain the source data for the semantic map, which includes image data, laser data, inertial navigation data, and satellite positioning data;
[0113] The vehicle pose corresponding to the image data is determined based on the image data and the inertial navigation data, and the vehicle pose corresponding to the laser data is determined based on the laser data and the inertial navigation data.
[0114] A first nonlinear optimization model is constructed based on the vehicle pose corresponding to the image data, the vehicle pose corresponding to the laser data, and the satellite positioning data, and the optimized vehicle pose is determined based on the first nonlinear optimization model.
[0115] Determine the three-dimensional semantic information corresponding to the image data based on the image data and the laser data;
[0116] A semantic map is constructed based on the optimized vehicle pose and the corresponding 3D semantic information of the image data.
[0117] The above is as stated in this application. Figure 1The semantic map construction apparatus disclosed in the illustrated embodiments can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0118] The electronic device can also perform Figure 1 The method executed by the semantic map building device, and the implementation of the semantic map building device in Figure 1 The functions of the embodiments shown are not described in detail here.
[0119] This application also proposes a computer-readable storage medium that stores one or more programs, the programs including instructions that, when executed by an electronic device including multiple applications, enable the electronic device to perform... Figure 1 The semantic map construction apparatus in the illustrated embodiment executes a method, specifically for performing:
[0120] Obtain the source data for the semantic map, which includes image data, laser data, inertial navigation data, and satellite positioning data;
[0121] The vehicle pose corresponding to the image data is determined based on the image data and the inertial navigation data, and the vehicle pose corresponding to the laser data is determined based on the laser data and the inertial navigation data.
[0122] A first nonlinear optimization model is constructed based on the vehicle pose corresponding to the image data, the vehicle pose corresponding to the laser data, and the satellite positioning data, and the optimized vehicle pose is determined based on the first nonlinear optimization model.
[0123] Determine the three-dimensional semantic information corresponding to the image data based on the image data and the laser data;
[0124] A semantic map is constructed based on the optimized vehicle pose and the corresponding 3D semantic information of the image data.
[0125] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0126] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0127] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0128] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0129] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0130] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0131] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0132] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0133] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0134] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A semantic map construction method, wherein, The method includes: Obtain the source data for the semantic map, which includes image data, laser data, inertial navigation data, and satellite positioning data; The vehicle pose corresponding to the image data is determined based on the image data and the inertial navigation data, and the vehicle pose corresponding to the laser data is determined based on the laser data and the inertial navigation data. A first nonlinear optimization model is constructed based on the vehicle pose corresponding to the image data, the vehicle pose corresponding to the laser data, and the satellite positioning data, and the optimized vehicle pose is determined based on the first nonlinear optimization model. Determine the three-dimensional semantic information corresponding to the image data based on the image data and the laser data; A semantic map is constructed based on the optimized vehicle pose and the three-dimensional semantic information corresponding to the image data. Determining the vehicle pose corresponding to the image data based on the image data and the inertial navigation data includes: The image data is semantically segmented using a preset semantic segmentation model to obtain semantic segmentation results, and the image data is feature extracted using a preset image feature extraction algorithm to obtain image feature extraction results. A second nonlinear optimization model is constructed based on the semantic segmentation results, the image feature extraction results, and the inertial navigation data. The vehicle pose corresponding to the image data is determined based on the second nonlinear optimization model.
2. The method as described in claim 1, wherein, The semantic segmentation result is the semantic segmentation result of the current frame, and the image feature extraction result is the image feature extraction result of the current frame. The step of constructing a second nonlinear optimization model based on the semantic segmentation result, the image feature extraction result, and the inertial navigation data includes: Obtain the semantic segmentation result of the previous frame, and construct semantic constraints based on the semantic segmentation result of the current frame and the semantic segmentation result of the previous frame; Obtain the image feature extraction result of the previous frame, and construct image feature constraints based on the image feature extraction result of the current frame and the image feature extraction result of the previous frame; Obtain the inertial navigation data between the image data of the current frame and the image data of the previous frame, and construct the inertial navigation pre-integration constraints corresponding to the image data based on the inertial navigation data between the image data of the current frame and the image data of the previous frame; The second nonlinear optimization model is constructed based on the semantic constraints, the image feature constraints, and the inertial navigation pre-integration constraints corresponding to the image data.
3. The method as described in claim 1, wherein, Determining the vehicle pose corresponding to the laser data based on the laser data and the inertial navigation data includes: The laser data is subjected to feature extraction using a preset point cloud feature extraction algorithm to obtain point cloud feature extraction results; A third nonlinear optimization model is constructed based on the point cloud feature extraction results and the inertial navigation data. The vehicle pose corresponding to the laser data is determined based on the third nonlinear optimization model.
4. The method as described in claim 3, wherein, The point cloud feature extraction result is the point cloud feature extraction result of the current frame. The step of constructing a third nonlinear optimization model based on the point cloud feature extraction result and the inertial navigation data includes: Obtain the point cloud feature extraction result of the previous frame, and construct point cloud feature constraints based on the point cloud feature extraction result of the current frame and the point cloud feature extraction result of the previous frame. Obtain the inertial navigation data between the laser data of the current frame and the laser data of the previous frame, and construct the inertial navigation pre-integration constraint corresponding to the laser data based on the inertial navigation data between the laser data of the current frame and the laser data of the previous frame; The third nonlinear optimization model is constructed based on the point cloud feature constraints and the inertial navigation pre-integration constraints corresponding to the laser data.
5. The method as described in claim 1, wherein, The step of constructing a first nonlinear optimization model based on the vehicle pose corresponding to the image data, the vehicle pose corresponding to the laser data, and the satellite positioning data includes: Construct relative change constraints between image data based on the vehicle pose corresponding to the image data; Based on the vehicle pose corresponding to the laser data, construct the relative change constraints between the laser data; Construct relative variation constraints between satellite positioning data based on the satellite positioning data; The first nonlinear optimization model is constructed based on the relative variation constraints between the image data, the relative variation constraints between the laser data, and the relative variation constraints between the satellite positioning data.
6. The method of claim 1, wherein, After constructing a semantic map based on the optimized vehicle pose and the corresponding 3D semantic information of the image data, the method further includes: The semantic map is post-processed using a preset post-processing strategy to obtain a post-processed semantic map. The preset post-processing strategy includes voxel filtering of the semantic map and / or determining the semantic map that is outside the vehicle's current field of view and extracting semantic information contours from the semantic map that is outside the vehicle's current field of view.
7. A semantic map construction apparatus, wherein, The device includes: The acquisition unit is used to acquire the map source data of the semantic map, which includes image data, laser data, inertial navigation data and satellite positioning data; The first determining unit is configured to determine the vehicle pose corresponding to the image data based on the image data and the inertial navigation data, and to determine the vehicle pose corresponding to the laser data based on the laser data and the inertial navigation data. The second determining unit is used to construct a first nonlinear optimization model based on the vehicle pose corresponding to the image data, the vehicle pose corresponding to the laser data, and the satellite positioning data, and to determine the optimized vehicle pose based on the first nonlinear optimization model. The third determining unit is used to determine the three-dimensional semantic information corresponding to the image data based on the image data and the laser data; A construction unit is used to construct a semantic map based on the optimized vehicle pose and the three-dimensional semantic information corresponding to the image data. The first determining unit is specifically used for: The image data is semantically segmented using a preset semantic segmentation model to obtain semantic segmentation results, and the image data is feature extracted using a preset image feature extraction algorithm to obtain image feature extraction results. A second nonlinear optimization model is constructed based on the semantic segmentation results, the image feature extraction results, and the inertial navigation data. The vehicle pose corresponding to the image data is determined based on the second nonlinear optimization model.
8. An electronic device, comprising: processor; as well as A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method of any one of claims 1 to 6.
9. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of applications, cause the electronic device to perform the method of any one of claims 1 to 6.