Map construction method and device, equipment, medium and product
The parking lot stable feature point cloud was extracted by surround viewing camera and semantic segmentation model, and the raster map was constructed in combination with the Cartographer algorithm, which solved the problem of visual SLAM being sensitive to scene changes and high cost of laser SLAM, and achieved low-cost and efficient SLAM map construction.
Patent Information
- Application Number
- CN202510527966.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-01
AI Technical Summary
In the existing SLAM technology, in scenarios where GPS signals are weak, such as indoor parking lots, visual SLAM is sensitive to scene changes, while laser SLAM is expensive and difficult to popularize in ordinary vehicles.
The surround view camera is used to collect surround view images, extract stable feature point clouds through semantic segmentation model, and build a raster map with the Cartographer algorithm to simulate lidar point cloud data.
Without relying on lidar, efficient and low-cost parking map construction is achieved, which alleviates the impact of scene changes on positioning and provides stable navigation information.
Smart Images

Figure CN120403602A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of map data, and particularly to a map construction method, device, equipment, medium and product. Background Art
[0002] Simultaneous Localization and Mapping (SLAM) technology is a technology for constructing a map of an unknown area and providing a relative pose. This technology is generally applicable to scenarios such as indoor parking lots where the Global Positioning System (GPS) cannot provide effective positioning information. There are two main applications of SLAM technology, namely visual SLAM and lidar SLAM. Among them, it is difficult for the technology based on visual SLAM to solve the problem of scene change, and the scene during map construction will affect the SLAM positioning effect due to changes in vehicles and light. The robustness of lidar SLAM is better than that of visual SLAM, and it is less sensitive to changes in vehicles and light to a certain extent. However, the disadvantage of lidar sensors is that they are expensive and are currently only installed on a small number of high-end models. Summary of the Invention
[0003] Embodiments of the present invention provide a map construction method, device, equipment, medium and product, which realizes map construction using a set algorithm without installing a lidar, and alleviates the problem of scene change in SLAM to a certain extent.
[0004] In a first aspect, an embodiment of the present invention provides a map construction method, which includes:
[0005] Establish a bird's-eye view of a parking lot according to at least one direction of panoramic images collected during driving, where the panoramic images are collected by a panoramic camera installed on a vehicle;
[0006] Determine a target point cloud representing the stable features of the parking lot according to the bird's-eye view and a pre-constructed semantic segmentation model;
[0007] Obtain a grid map of the parking lot according to the target point cloud representing the stable features of the parking lot and a set simultaneous localization and mapping algorithm.
[0008] In a second aspect, an embodiment of the present invention provides a map construction device, which includes:
[0009] A bird's-eye view construction module, configured to establish a bird's-eye view of a parking lot according to at least one direction of panoramic images collected during driving, where the panoramic images are collected by a panoramic camera installed on a vehicle;
[0010] A point cloud determination module, configured to determine a target point cloud representing stable features of the parking lot according to the bird's-eye view and a pre-built semantic segmentation model;
[0011] A map construction module, configured to obtain a grid map of the parking lot according to the target point cloud representing stable features of the parking lot and a set simultaneous localization and mapping algorithm.
[0012] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the map construction method according to any embodiment of the present invention.
[0016] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores computer instructions for enabling a processor to implement the map construction method according to any embodiment of the present invention when executed.
[0017] In a fifth aspect, an embodiment of the present invention further provides a computer program product, which includes a computer program that implements the map construction method according to any embodiment of the present invention when executed by a processor.
[0018] An embodiment of the present invention provides a map construction method, device, equipment, medium and product. The method includes: establishing a bird's-eye view of a parking lot according to at least one panoramic image collected during driving, and the panoramic image is collected by a panoramic camera installed on a vehicle; determining a target point cloud representing stable features of the parking lot according to the bird's-eye view and a pre-built semantic segmentation model; obtaining a grid map of the parking lot according to the target point cloud representing stable features of the parking lot and a set simultaneous localization and mapping algorithm. In the above technical solution, the panoramic image is projected into the bird's-eye view space, and the pre-built semantic segmentation model is used to extract stable features such as lane lines and parking space lines, and convert them into the point cloud of a lidar. Then, based on the set algorithm, map construction is performed on the point cloud. It is realized that point cloud data can be obtained even without a lidar, so that the set simultaneous localization and mapping algorithm can also be used to perform map construction on the point cloud data, and to a certain extent, the problem of scene change in SLAM mapping is alleviated.
[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 It is a schematic flowchart of a map construction method provided in Embodiment 1 of the present invention;
[0022] Figure 2 It is a schematic diagram of a panoramic image of a parking lot during the execution of a map construction method provided in Embodiment 1 of the present invention;
[0023] Figure 3 It is a schematic diagram of an aerial view of a parking lot during the execution of a map construction method provided in Embodiment 1 of the present invention;
[0024] Figure 4 It is a schematic diagram of the rendering of a map constructed for the underground parking lot of a large shopping mall provided in Embodiment 1 of the present invention;
[0025] Figure 5 It is a schematic flowchart of another map construction method provided in Embodiment 2 of the present invention;
[0026] Figure 6 It is an aerial view of a parking lot and the corresponding semantic segmentation map during the execution of a map construction method provided in Embodiment 2 of the present invention;
[0027] Figure 7 It is a comparison diagram of the effects of map construction using the prior art and using the technical solution of the present invention provided in Embodiment 2 of the present invention;
[0028] Figure 8 It is a flow example diagram of a map construction method in a certain application scenario provided in Embodiment 2 of the present invention;
[0029] Figure 9 It is a schematic structural diagram of a map construction device provided in Embodiment 3 of the present invention;
[0030] Figure 10 It is a schematic structural diagram of an electronic device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0032] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0033] It should be clear that Cartographer is an open-source set of SLAM codes, which is widely used for mapping, positioning and navigation in the mobile robot industry. SLAM technology is a technology that constructs a map of an unknown area and provides relative poses in real time. This technology is commonly applied to scenarios such as indoor parking lots where GPS cannot provide effective positioning information. For example, the navigation and positioning functions of many indoor cleaning robots and food delivery robots are developed based on Cartographer codes. Cartographer codes use lidar as the main sensor input, that is, the point cloud data of lidar is used as the input. However, vehicle-mounted lidar is expensive and is generally installed on vehicles with a price of more than 300,000 yuan. The robustness of laser SLAM is better than that of visual SLAM, and it is not so sensitive to changes in vehicles and light to a certain extent. However, the disadvantage of laser sensors is that they are expensive and are currently only installed on a small number of high-end models. On the other hand, surround-view cameras are inexpensive and have become standard equipment for the vast majority of vehicles. However, existing vision-based SLAM technologies are difficult to solve the problem of scene changes, and the scenes during mapping will affect the positioning effect of SLAM due to changes in vehicles and light. Therefore, a map construction method is needed to solve the above problems.
[0034] Embodiment 1
[0035] Figure 1FIG. 0 is a schematic flowchart of a map construction method provided in Embodiment 1 of the present invention. This method is applicable to the situation of constructing a map. This method can be executed by a map construction device, which can be implemented in the form of hardware and / or software and is generally integrated in an electronic device.
[0036] As Figure 1 shown, a map construction method provided in Embodiment 1 specifically may include the following steps:
[0037] S101. Establish an aerial view of a parking lot according to the omnidirectional images collected in at least one direction during driving.
[0038] In this embodiment, the omnidirectional images are collected by omnidirectional cameras installed on the vehicle. During driving, the vehicle-mounted cameras are used to collect the surrounding environment images, which are recorded as omnidirectional images. In order to obtain more comprehensive images, omnidirectional cameras can be arranged in different directions of the vehicle so that the omnidirectional images collected by each omnidirectional camera can completely cover the surrounding environment. Preferably, the omnidirectional cameras are respectively installed on the front side, rear side, left side and right side of the vehicle. Exemplarily, Figure 2 FIG. 13 is a schematic diagram of the omnidirectional images of a parking lot during the execution of a map construction method provided in Embodiment 1 of the present invention. As Figure 2 shown, the parking lot is imaged respectively based on the front side, rear side, left side and right side of the vehicle to obtain omnidirectional images in four directions. In this embodiment, taking advantage of the fact that omnidirectional cameras are usually installed on the vehicle, the omnidirectional cameras are used to simulate the input of lidar, which is equivalent to transforming the omnidirectional images collected by the omnidirectional cameras into point cloud data of lidar.
[0039] Continuing with the above description, after obtaining the omnidirectional images in each direction during driving, the inverse projection transformation can be performed on each omnidirectional image, and the omnidirectional images can be spliced into the aerial view space to establish an aerial view of the parking lot. Exemplarily, the steps of splicing each omnidirectional image to obtain an aerial view of the parking lot can be described as follows: when the coordinates of each pixel point on the known image, as well as the internal parameters and external parameters of the camera are known, the two-dimensional pixel points can be reversely projected into the three-dimensional space through the following steps. Specifically: through the inverse of the internal parameter matrix of the camera, the coordinates of the two-dimensional pixel points are converted into the coordinates in the camera coordinate system, and then the external parameter matrix is used to convert the coordinates in the camera coordinate system into the points in the world coordinate system.
[0040] Then, project the three-dimensional coordinates of each pixel in the world coordinate system onto the two-dimensional camera plane to obtain an aerial view of the parking lot. It should be noted that this method assumes that the objects in the scene are all on the ground. If there are some objects higher than the ground, they will be projected onto the ground, which may cause errors. However, considering that the features used for map construction, such as guiding arrows and lane lines, are all on the ground, the existence of some objects higher than the ground will not affect these stable features such as guiding arrows and lane lines. Exemplarily, Figure 3 is the aerial view of the parking lot during the execution of a map construction method provided in the first embodiment of the present invention, as Figure 3 shown. By processing the panoramic image of the parking lot shown in Figure 2 , a schematic diagram of the aerial view of the parking lot shown in Figure 3 can be obtained.
[0041] S102. According to the aerial view and the pre-constructed semantic segmentation model, determine the target point cloud representing the stable features of the parking lot.
[0042] In this embodiment, the semantic segmentation model is specifically understood as a model for performing semantic segmentation on images, and the semantic segmentation model is pre-constructed and trained. Exemplarily, the semantic segmentation model used is the BiSeg network. BiSeg is a method for simultaneously performing instance segmentation and semantic segmentation based on a fully convolutional network. BiSeg predicts instance segmentation as the posterior in Bayesian inference, where semantic segmentation is used as the prior. It extends the concept of the position-sensitive score map used in recent methods to the fusion of multiple score maps at different scales and partition patterns, and uses it as a robust likelihood end-to-end solution for instance segmentation inference. Since both Bayesian inference and map fusion are performed on each pixel, BiSeg is a fully convolutional end-to-end solution, inheriting all the advantages of the fully convolutional network, such as being able to process input images of any size and being able to efficiently perform feature extraction and segmentation prediction. In this embodiment, the BiSeg network is selected as the initial network model, and the initial network model is trained based on the training sample set to obtain a trained semantic segmentation model. The input data of the semantic segmentation model is an image, and the output data is the semantic segmentation result.
[0043] In this embodiment, after obtaining the aerial view of the parking lot, the semantic segmentation model can be used to perform semantic segmentation on the aerial view to extract stable features such as lane lines and parking space lines. Then, using the assumption of pixel projection on the plane, measure the distance from the pixel to the center of the vehicle body and convert it into the point cloud of the lidar. The lidar point cloud is obtained without using a laser sensor, and the point cloud representing the stable features is recorded as the target point cloud.
[0044] Specifically, the bird's-eye view is input into the semantic segmentation model, which performs semantic segmentation on the bird's-eye view and assigns an attribute value to each pixel point in the bird's-eye view. For example, the pixel point may be background, drivable area, no-go area, curb, speed bump, arrow, lane line, parking space line, or other lines, etc. Pixel points in different regions correspond to different attribute values. For example, background, drivable area, no-go area, curb, speed bump, arrow, lane line, parking space line, and other lines, etc. correspond to different attribute values respectively. The attribute value corresponding to each pixel point in the bird's-eye view is obtained, and the pixel points representing the stable features of the parking lot are extracted as target pixel points for subsequent map construction. For example, what can represent the stable features of the parking lot can be guiding arrows, lane lines, or other lines, etc. According to the above description, for each pixel point extracted representing the stable features, using the assumption that the pixel point projects onto a plane, the distance from the pixel point to the center of the vehicle body is measured and converted into the point cloud of the lidar to obtain the target point cloud representing the stable features of the parking lot.
[0045] S103. According to the target point cloud representing the stable features of the parking lot and the set simultaneous localization and mapping algorithm, obtain the grid map of the parking lot.
[0046] Among them, the set simultaneous localization and mapping algorithm is the Cartographer algorithm. The Cartographer code is an open-source library for simultaneous localization and mapping. It is mainly used for the creation of 2D maps, especially in the case of using lidar sensors. In this embodiment, after converting the surround-view image into the target point cloud representing the stable features of the parking lot, the open-source Cartographer code can be used to perform map construction with the determined target point cloud to obtain the map of the parking lot. In addition, the Cartographer code can use a grid map to represent the environment. Therefore, based on the Cartographer algorithm and the target point cloud, the grid map of the parking lot can be constructed. The constructed grid map contains passable information, such as labels like navigation arrows, lane lines, etc., for path planning.
[0047] It can be understood that using the map construction method provided in this embodiment to establish a map avoids the high cost brought by lidar. At the same time, compared with the pure vision solution, it is more robust to scene changes. Because in scenarios like underground parking lots, lane lines generally do not change significantly, which enables the map to remain available for a long time. In addition, the grid map itself is also a very efficient representation method. Exemplarily, Figure 4 is a schematic diagram of the effect diagram of the map constructed for the underground parking lot of a large shopping mall provided in Embodiment 1 of the present invention, as Figure 4As shown, it is the rendering of the map constructed for the underground parking lot of a large shopping mall. In the figure, the green trajectory represents the trajectory generated by constructing the map using the existing technology, and the red represents the trajectory generated by constructing the map using the technical solution of the present invention. The storage space occupied by the entire map is less than 100 MB, saving the occupied resources.
[0048] In the above technical solution, the surround-view images are projected into the bird's-eye view space, and using the pre-constructed semantic segmentation model, stable features such as lane lines and parking space lines are extracted and converted into the point cloud of the lidar. Then, based on the set algorithm, map construction is performed on the point cloud. It realizes obtaining point cloud data even without carrying a lidar, and thus can also use the set simultaneous localization and mapping algorithm to perform map construction on the point cloud data, and to a certain extent alleviates the problem of scene change in SLAM.
[0049] Embodiment 2
[0050] [[ID= It is a schematic flowchart of another map construction method provided by Embodiment 2 of the present invention. This embodiment is a further optimization of the above embodiment. In this embodiment, the limitation of "establishing a bird's-eye view of the parking lot according to the surround-view images collected in at least one direction during driving" is further optimized, and the limitation of "determining the target point cloud representing the stable features of the parking lot according to the bird's-eye view and the pre-constructed semantic segmentation model" is optimized, and the limitation of "obtaining the grid map of the parking lot according to the target point cloud representing the stable features of the parking lot and the set simultaneous localization and mapping algorithm" is optimized.
[0051] As shown, Embodiment 2 provides a map construction method, which specifically includes the following steps:
[0052] S201. For the calibration parameters of each surround-view camera, establish an inverse projection matrix representing the relationship between the image coordinate system and the world coordinate system.
[0053] It should be clear that the inverse projection transformation is an operation opposite to the projection transformation, which restores the information in the two-dimensional image to the three-dimensional space. The following is the general process of generating a bird's-eye view using the inverse projection transformation. First, it is necessary to calibrate the camera to determine the internal and external parameters of the camera. The internal parameters of the camera represent the transformation of the image coordinate system relative to the camera coordinate system, and the internal parameters include the focal length, principal point coordinates, etc. The external parameters of the camera represent the transformation of the camera coordinate system relative to the world coordinate system, and the external parameters include the rotation and translation parameters of the camera. These parameters can be obtained through experiments using known calibration objects (such as checkerboards), or approximate values can be obtained according to the camera's specification manual. The purpose of camera calibration is to establish the relationship between image coordinates and world coordinates and clarify the inverse projection matrix used. Exemplarily, the KB8 model can be adopted. The KB8 model is a camera model used to process the image distortion of fisheye lenses and is used for projection and back-projection operations. In this embodiment, it is necessary to establish an inverse projection matrix that characterizes the relationship between the two-dimensional image coordinate system and the three-dimensional world coordinate system for the calibration parameters of each surround-view camera.
[0054] S202. Multiply each pixel point in the surround-view image collected by the surround-view camera during driving by the inverse projection matrix to obtain the three-dimensional coordinates of each pixel point in the world coordinate system.
[0055] In this embodiment, for each pixel point in the surround-view image corresponding to the surround-view camera, its three-dimensional coordinates in the world coordinate system can be obtained by multiplying by the inverse projection matrix, that is, its coordinates in the three-dimensional space are obtained. The coordinates obtained here are usually uncertain because depth information is lost during the projection process. Generally, it is necessary to determine this value according to the prior knowledge of the scene or other constraint conditions. Therefore, it is necessary to assume that all pixels are on the ground here. Since the subsequent semantic segmentation model only detects ground signs such as lane lines, parking space lines, and guiding arrows, and lane lines do indeed only exist on the ground, the assumption that pixel points are all on the plane becomes reasonable by only detecting ground signs such as lane lines, parking space lines, and guiding arrows.
[0056] S203. Project the three-dimensional coordinates of each pixel point in the world coordinate system onto the camera plane to obtain a bird's-eye view of the parking lot.
[0057] In this embodiment, after obtaining the coordinates of each pixel point in the three-dimensional space, project all points onto a plane parallel to the camera plane. This plane can be regarded as the plane of the bird's-eye view. Then project the three-dimensional points onto this plane, and thus a bird's-eye view of the parking lot is obtained.
[0058] S204. Determine at least one target pixel point that characterizes the stable features of the parking lot according to the bird's-eye view and the pre-constructed semantic segmentation model.
[0059] Specifically, the bird's-eye view is input into the semantic segmentation model, which performs semantic segmentation on the bird's-eye view and assigns an attribute value to each pixel in the bird's-eye view. For example, a pixel may belong to the background, drivable area, no-entry area, road edge, speed bump, arrow, lane line, parking space line, or other lines, etc. In this embodiment, the semantic segmentation model can perform semantic segmentation on each pixel in the bird's-eye view and assign a value to each pixel. It can be understood that pixels in different regions correspond to different attribute values. For example, the background, drivable area, no-entry area, road edge, speed bump, arrow, lane line, parking space line, and other lines, etc. respectively correspond to different attribute values.
[0060] Continuing with the above description, after performing semantic segmentation on the bird's-eye view using the semantic segmentation model, the attribute value corresponding to each pixel in the bird's-eye view can be obtained, and the pixels representing the stable features of the parking lot are extracted as target pixels for subsequent map construction. For example, what can represent the stable features of the parking lot can be guiding arrows, lane lines, parking space lines, or other lines, etc.
[0061] As a specific implementation, the step of determining at least one target pixel representing the stable features of the parking lot according to the bird's-eye view and the pre-constructed semantic segmentation model can be optimized, including:
[0062] a1) Input the bird's-eye view into the semantic segmentation model to obtain the attribute values corresponding to each pixel in the bird's-eye view.
[0063] Specifically, when the bird's-eye view is input into the semantic segmentation model, the semantic segmentation model assigns an attribute value to each pixel in the bird's-eye view, and different attribute values represent different meanings. Preferably, each attribute value corresponds to a regional attribute respectively, and the regional attributes include the background, drivable area, no-parking area, road edge, speed bump, guiding arrow, lane line, parking space line, or other lines respectively. Exemplarily, assuming the attribute values are represented by numbers, 0 represents the background, 1 represents the drivable area, 2 represents the no-parking area, 3 represents the road edge, 4 represents the speed bump, 5 represents the guiding arrow, 6 represents the lane line, 7 represents the parking space line, and 8 represents other lines. When the bird's-eye view is input into the semantic segmentation model, the semantic segmentation model will segment out what the attribute value of each pixel is, and further it can be determined which area the pixel belongs to. It should be noted that what semantics the semantic segmentation model can segment the bird's-eye view into is related to the training of the semantic segmentation model according to the actual situation.
[0064] b1) Extract the pixels with the target attribute value as the target pixels representing the stable features of the parking lot.
[0065] In this embodiment, the target attribute value can be specifically understood as the value of the area attribute related to the construction of the parking lot map trajectory. Preferably, the target area attributes corresponding to the target attribute values include guiding arrows, lane lines or other lines. The guiding arrows, lane lines or other lines are the signs that vehicles need to refer to when driving, and these signs are relatively stable features on the ground. These area attributes are recorded as target area attributes. It is necessary to screen out the pixel points belonging to the target area attributes. This step is used to extract the pixel points with the attribute value being the target attribute value based on the attribute values of each pixel point, as the pixel points representing the stable features of the parking lot, denoted as target pixel points. Continuing to describe according to the above example, the attribute value 5 corresponding to the guiding arrow, the attribute value 6 corresponding to the lane line, the attribute value 7 corresponding to the parking space line, and the attribute value 8 corresponding to other lines are used as the target attribute values, and the pixel points with the attribute value being the target attribute value are used as the target pixel points representing the stable features of the parking lot.
[0066] Exemplarily, FIG. is a bird's-eye view of a parking lot and the corresponding semantic segmentation map during the execution of a map construction method provided in Embodiment 2 of the present invention. As shown, by performing semantic segmentation processing on the bird's-eye view of the parking lot, the attribute corresponding to each pixel point can be obtained, corresponding to the semantic segmentation map. It should be noted that, for the convenience of visualization, the value of each pixel is magnified by 35 times.
[0067] The above technical solution concretizes the use of a semantic segmentation model to segment pixel points representing stable features from a bird's-eye view, and then converts these pixel points into point clouds.
[0068] S205. Convert each target pixel point into a target point cloud representing the stable features of the parking lot.
[0069] In this embodiment, for each target pixel point, using the assumption that the pixel point is projected on a plane, the distance from the pixel point to the center of the vehicle body is measured and converted into a point cloud of a lidar, denoted as the target point cloud representing the stable features of the parking lot.
[0070] As a specific implementation manner, the step of converting each target pixel point into each target point cloud representing the stable features of the parking lot can be optimized, including:
[0071] a2) For each target pixel point, calculate the distance from the target pixel point to the center of the vehicle body.
[0072] In this embodiment, for each target pixel, the distance from the target pixel to the center of the vehicle body is calculated. The number of pixels from the target pixel to the center of the vehicle body can be determined, and the size of each pixel is also known. Therefore, by multiplying the pixel size by the number of pixels from the target pixel to the center of the vehicle body, the distance from the target pixel to the center of the vehicle body on the X and Y axes can be obtained. These two distances can be considered the X-axis coordinate and Y-axis coordinate of the point cloud corresponding to the target pixel. For example, assuming that the length and width of each pixel are 3 mm, respectively, and a target pixel is 300 pixels away from the center of the vehicle body on the X axis and 500 pixels away on the Y axis, the X-axis distance of the point cloud corresponding to the target pixel is: 3 mm * 300 = 90 cm, and the Y-axis distance of the point cloud corresponding to the target pixel is: 3 mm * 500 = 150 cm.
[0073] b2) Determine a target point cloud representing stable features of the parking lot based on the distance between the target pixel point and the center of the vehicle body.
[0074] Continuing with the above description, the distance from the target pixel point to the center of the vehicle body is calculated here to simulate the principle of collecting point clouds by a laser radar installed on the vehicle body. After determining the distance from the target pixel point to the center of the vehicle body, the distance from the target pixel point to the center of the vehicle body can be used as the X-axis coordinate and Y-axis coordinate of the target pixel point. Considering that the pixel points of stable features such as lane lines and parking space lines are located on the ground, zero is directly used as the Z-axis coordinate of the target pixel point. Based on this, the coordinates of the target pixel point on the X-axis, Y-axis and Z-axis are obtained, and the three-dimensional coordinates are used as the target point cloud corresponding to the target pixel point. Each target pixel point can calculate its corresponding target point cloud, and the target pixel points representing the stable features of the parking lot are converted into target point clouds, thereby obtaining a target point cloud representing the stable features of the parking lot.
[0075] The above technical solution specifies the steps for converting pixels representing stable features into point clouds, realizing the conversion of surround view images into lidar point clouds. Even without a lidar, point clouds of stable features such as parking spaces and lane lines can be indirectly obtained, providing basic data for subsequent mapping using the Cartographer code.
[0076] S206 , performing voxel filtering on the target point cloud to obtain a filtered target point cloud.
[0077] Voxel filtering is a point cloud downsampling method used to reduce the size of point cloud data while preserving its geometric properties. This step uses voxel filtering to thin out the segmented target point cloud, resulting in a sparse target point cloud. This thinned target point cloud ensures the real-time performance of the Cartographer code.
[0078] S207. Input the filtered target point cloud into a set simultaneous localization and mapping algorithm to obtain a grid map of the parking lot.
[0079] Specifically, inputting the filtered target point cloud into a set simultaneous localization and mapping algorithm can ensure the real-time nature of map building and obtain a grid map of the parking lot.
[0080] Exemplarily, to more clearly illustrate the effect of the map building method provided in the embodiments of the present invention, an example of map building in an actual scenario is used for illustration. FIG. is a comparison diagram of the effects of map building using the prior art and using the technical solution of the present invention in the second embodiment of the present invention. As shown, the green trajectory represents the trajectory generated by map building using the prior art, that is, the trajectory generated by the odometer. The red represents the trajectory generated by map building using the technical solution of the present invention. It can be seen that the data set has circled the parking lot twice in total. The two green trajectories in the figure do not completely overlap, with a large error. While the two red trajectories basically overlap, and there is no obvious ghosting in the lane lines and parking space lines, which proves the correctness of map building and the error is small.
[0081] The above technical solution specifies the steps of stitching panoramic images into a bird's-eye view image, the steps of performing semantic segmentation on the bird's-eye view image to obtain the target point cloud, and the steps of performing map building based on the target point cloud and the simultaneous localization and mapping algorithm. Project the panoramic image into the bird's-eye view space to obtain a bird's-eye view image of the parking lot, and then use the semantic segmentation model to extract stable features such as lane lines and parking space lines from the bird's-eye view image. Finally, use the assumption of pixel projection on a plane to convert pixel points into the point cloud of a lidar. By using inverse projection transformation and semantic segmentation technology, the problem of scene changes in visual SLAM is solved, and the lidar point cloud is obtained without using a laser sensor. Thus, the set simultaneous localization and mapping algorithm can also be used to perform map building on the point cloud data.
[0082] To more clearly illustrate the map building method provided in the embodiments of the present invention, an example of an actual application scenario of map building is used for illustration. Exemplarily, FIG. is a flow example diagram of the map building method in an application scenario provided in the second embodiment of the present invention. As shown, the execution steps of the map building method can specifically include:
[0083] S1. For the calibration parameters of each panoramic camera, establish an inverse projection matrix representing the relationship between the image coordinate system and the world coordinate system.
[0084] S2. Multiply each pixel point in the panoramic image collected by the panoramic camera during driving by the inverse projection matrix to obtain the three-dimensional coordinates of each pixel point in the world coordinate system.
[0085] S3. Project the three-dimensional coordinates of each pixel point in the world coordinate system onto the camera plane to obtain an aerial view of the parking lot.
[0086] Among them, the surround view image is acquired by the surround view camera installed on the vehicle.
[0087] S4. Input the aerial view into the semantic segmentation model to obtain the attribute values corresponding to each pixel point in the aerial view.
[0088] S5. Extract the pixel points with the target attribute value as the target pixel points representing the stable features of the parking lot.
[0089] S6. For each target pixel point, calculate the distance from the target pixel point to the center of the vehicle body.
[0090] S7. Determine the target point cloud representing the stable features of the parking lot according to the distance from the target pixel point to the center of the vehicle body.
[0091] S8. Perform voxel filtering on the target point cloud to obtain the filtered target point cloud.
[0092] S9. Input the filtered target point cloud into the set simultaneous localization and mapping algorithm to obtain the grid map of the parking lot.
[0093] Embodiment III
[0094] As shown in the schematic structural diagram of a map construction device provided in Embodiment III of the present invention, this device can be applied to the situation of map construction. This map construction device can be implemented in the form of hardware and / or software and is generally integrated in an electronic device. As shown, this device includes: an aerial view construction module 31, a point cloud determination module 32, and a map construction module 33, where
[0095] The aerial view construction module 31 is configured to establish an aerial view of the parking lot according to the surround view images collected in at least one direction during the driving process. The surround view images are acquired by the surround view camera installed on the vehicle;
[0096] The point cloud determination module 32 is configured to determine the target point cloud representing the stable features of the parking lot according to the aerial view and the pre-constructed semantic segmentation model;
[0097] The map construction module 33 is configured to obtain the grid map of the parking lot according to the target point cloud representing the stable features of the parking lot and the set simultaneous localization and mapping algorithm.
[0098] In the above technical solution, the surround view images are projected into the bird's-eye view space, and the pre-built semantic segmentation model is used to extract stable features such as lane lines and parking space lines, and convert them into the point cloud of the lidar. Then, based on the set algorithm, map construction is performed on the point cloud. It realizes obtaining point cloud data even without carrying a lidar, so that the set simultaneous localization and mapping algorithm can also be used to perform map construction on the point cloud data, and to a certain extent alleviates the problem of scene changes in SLAM.
[0099] Optionally, the point cloud determination module 32 may include:
[0100] A pixel point determination unit, configured to determine at least one target pixel point representing the stable features of the parking lot according to the bird's-eye view and the pre-built semantic segmentation model;
[0101] A point cloud conversion unit, configured to convert each target pixel point into a target point cloud representing the stable features of the parking lot.
[0102] Optionally, the pixel point determination unit is specifically configured to:
[0103] Input the bird's-eye view into the semantic segmentation model to obtain the attribute values corresponding to each pixel point in the bird's-eye view;
[0104] Extract the pixel points with the target attribute value as the target pixel points representing the stable features of the parking lot.
[0105] Optionally, each attribute value corresponds to a regional attribute respectively, and the regional attributes include background, drivable area, no-parking area, curb, speed bump, guiding arrow, lane line, parking space line or other lines respectively. The target regional attribute corresponding to the target attribute value includes guiding arrow, lane line, parking space line or other lines.
[0106] Optionally, the point cloud conversion unit is specifically configured to:
[0107] For each target pixel point, calculate the distance from the target pixel point to the center of the vehicle body;
[0108] According to the distance from the target pixel point to the center of the vehicle body, determine the target point cloud representing the stable features of the parking lot.
[0109] Optionally, the bird's-eye view construction module 31 is specifically configured to:
[0110] For the calibration parameters of each surround view camera, establish an inverse projection matrix representing the relationship between the image coordinate system and the world coordinate system;
[0111] Multiply each pixel point in the surround view image collected by the surround view camera during driving by the inverse projection matrix to obtain the three-dimensional coordinates of each pixel point in the world coordinate system;
[0112] Project the three-dimensional coordinates of each pixel point in the world coordinate system onto the camera plane to obtain an aerial view of the parking lot.
[0113] Optionally, the map construction module 33 is specifically configured to:
[0114] Perform voxel filtering on the target point cloud to obtain a filtered target point cloud;
[0115] Input the filtered target point cloud into a set simultaneous localization and mapping algorithm to obtain a grid map of the parking lot.
[0116] Optionally, the surround cameras are respectively installed on the front side, rear side, left side, and right side of the vehicle.
[0117] The map construction device provided by the embodiments of the present invention can execute the map construction method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0118] Embodiment IV
[0119] It is a schematic structural diagram of an electronic device provided by Embodiment IV of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0120] As shown, the electronic device 40 includes at least one processor 41, and a memory communicatively connected to at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc. Among them, the memory stores a computer program executable by at least one processor. The processor 41 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other through a bus 44. The input / output (I / O) interface 45 is also connected to the bus 44.
[0121] Multiple components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disc, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0122] The processor 41 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the map construction method.
[0123] In some embodiments, the map construction method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the map construction method described above can be executed. Alternatively, in other embodiments, the processor 41 can be configured to execute the map construction method by any other suitable means (e.g., by means of firmware).
[0124] The various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0125] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0126] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0127] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).
[0128] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0129] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0130] An embodiment of the present invention also provides a computer program product, including a computer program, which when executed by a processor, implements the map construction method provided in any embodiment of the present invention.
[0131] In the process of implementing the computer program product, computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user computer, partially on the user computer, executed as an independent software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).
[0132] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0133] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for map construction, characterized in that, including: establishing an aerial view of a parking lot according to panoramic images in at least one direction collected during driving, where the panoramic images are collected by panoramic cameras installed on a vehicle; determining a target point cloud representing stable features of the parking lot according to the aerial view and a pre-constructed semantic segmentation model; obtaining a grid map of the parking lot according to the target point cloud representing stable features of the parking lot and a set simultaneous localization and mapping algorithm.
2. The method according to claim 1, characterized in that, The determining a target point cloud representing stable features of the parking lot according to the aerial view and a pre-constructed semantic segmentation model includes: determining at least one target pixel point representing stable features of the parking lot according to the aerial view and a pre-constructed semantic segmentation model; converting each of the target pixel points into a target point cloud representing stable features of the parking lot.
3. The method according to claim 2, wherein The determining at least one target pixel point representing stable features of the parking lot according to the aerial view and a pre-constructed semantic segmentation model includes: inputting the aerial view into the semantic segmentation model to obtain attribute values corresponding to each pixel point in the aerial view; extracting pixel points with the attribute values being target attribute values as target pixel points representing stable features of the parking lot.
4. The method according to claim 3, wherein Each of the attribute values corresponds to a regional attribute respectively, and the regional attributes respectively include background, drivable area, no-parking area, curb, speed bump, guiding arrow, lane line, parking space line or other lines, and the target regional attribute corresponding to the target attribute value includes guiding arrow, lane line, parking space line or other lines.
5. The method according to claim 2, wherein The converting each of the target pixel points into each target point cloud representing stable features of the parking lot includes: for each target pixel point, calculating the distance from the target pixel point to the center of the vehicle body; determining a target point cloud representing stable features of the parking lot according to the distance from the target pixel point to the center of the vehicle body.
6. The method according to claim 1, characterized in that The establishing an aerial view of a parking lot according to panoramic images in at least one direction collected during driving includes: establishing an inverse projection matrix representing the relationship between the image coordinate system and the world coordinate system for the calibration parameters of each of the panoramic cameras; multiplying each pixel point in the panoramic images collected by the panoramic cameras during driving by the inverse projection matrix to obtain three-dimensional coordinates of each of the pixel points in the world coordinate system; projecting the three-dimensional coordinates of each of the pixel points in the world coordinate system onto the camera plane to obtain the aerial view of the parking lot.
7. The method according to claim 1, wherein The obtaining a grid map of the parking lot according to the target point cloud representing stable features of the parking lot and a set simultaneous localization and mapping algorithm includes: performing voxel filtering on the target point cloud to obtain a filtered target point cloud; inputting the filtered target point cloud into the set simultaneous localization and mapping algorithm to obtain the grid map of the parking lot.
8. The method according to claim 1, characterized in that, The panoramic cameras are respectively installed on the front side, rear side, left side and right side of the vehicle.
9. A map construction device, characterized in that, including: an aerial view construction module for establishing an aerial view of a parking lot according to panoramic images in at least one direction collected during driving, where the panoramic images are collected by panoramic cameras installed on a vehicle; A point cloud determination module, configured to determine a target point cloud representing the stable features of the parking lot according to the bird's-eye view and a pre-built semantic segmentation model; A map construction module, configured to obtain a grid map of the parking lot according to the target point cloud representing the stable features of the parking lot and a set simultaneous localization and mapping algorithm; 10. An electronic device, characterized in that, Including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the map construction method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to implement the map construction method according to any one of claims 1-8 when executed.
12. A computer program product, characterized in that, The computer program product includes a computer program which, when executed by a processor, implements the map construction method according to any one of claims 1-8.
Citation Information
Patent Citations
Map construction method and device, electronic device and computer readable storage medium
CN113865580A
Parking lot semantic map construction method and device and electronic equipment
CN116052127A
Global semantic map building and updating method
CN118776572A