Map construction method, device, equipment and storage medium
By combining the image data of vehicle-mounted cameras and aerial drones, a three-dimensional map is constructed, which solves the problem of insufficient map accuracy in the existing technology, and realizes high coverage and high precision three-dimensional map generation.
Patent Information
- Application Number
- CN202111210541.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-10-18
AI Technical Summary
The maps generated in the prior art are limited in accuracy and are difficult to meet user needs, especially in the construction of high-precision maps in non-road areas.
A three-dimensional map is constructed by combining vehicle-mounted cameras and aerial drones. First, the vehicle-mounted camera and the global satellite inertial navigation system are used to generate the initial feature map, and then the high-precision three-dimensional map is constructed through aerial photography drone data update and optimization, combining depth processing and grid model.
A three-dimensional map construction with high coverage and high precision is achieved, improving the completeness and accuracy of the map, especially in reconstruction accuracy in non-road areas.
Smart Images

Figure CN113920263B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of visual positioning technology, and relate to, but are not limited to, a method, device, equipment and storage medium for map construction. Background Art
[0002] With the continuous development of computer and communication technologies, maps have provided great assistance for people's travel. In the related art, methods such as conventional acquisition and manual processing are used to generate maps. The maps created in this way have limited accuracy and cannot well meet the needs of users. Summary of the Invention
[0003] The embodiments of the present application provide a technical solution for map construction.
[0004] The technical solution of the embodiments of the present application is implemented as follows:
[0005] The embodiments of the present application provide a method for map construction, the method includes:
[0006] Obtain first image data of the area to be drawn and navigation data corresponding to the first image data;
[0007] Based on the first image data and the navigation data, determine a first feature map;
[0008] Based on second image data of the area to be drawn, update the first feature map to obtain a second feature map; wherein, the acquisition methods of the first image data and the second image data are different;
[0009] Based on the second image data, the second feature map and the first image data, construct a three-dimensional map of the area to be drawn.
[0010] In some embodiments, the determining a first feature map based on the first image data and the navigation data includes: extracting features from the first image data to obtain image feature points and description information of the image feature points; based on the description information of the image feature points, match different images in the first image data to obtain an image relationship library representing the matching relationship between the different images; based on the image relationship library and the navigation data, determine the first feature map. In this way, by combining the feature point coordinates in the image relationship library with the navigation data, it is possible to simply and quickly construct a first feature map representing the spatial positions of the image feature points in the image data.
[0011] In some embodiments, determining the first feature map based on the image relationship library and the navigation data includes: determining the first pose of the acquisition device of the first image data based on the navigation data and a preset first acquisition parameter; wherein the preset first acquisition parameter includes: the external parameter for acquiring the first image data and the external parameter for acquiring the navigation data; determining the first feature map based on the image feature points in the relationship database, the first pose, and a second acquisition parameter; wherein the second acquisition parameter includes the internal parameter for acquiring the first image data. In this way, by combining the feature matching relationship with the initial image pose provided by the navigation data and based on the triangulation principle, the feature map of the image data can be quickly constructed.
[0012] In some embodiments, determining the first feature map based on the image feature points in the relationship database, the first pose, and the second acquisition parameter includes: triangulating the first true position of the image feature points in the image to obtain the three-dimensional coordinates of the image feature points in the world coordinate system based on the first pose and the second acquisition parameter; constructing the first feature map based on the three-dimensional coordinates.
[0013] In some embodiments, constructing the first feature map based on the three-dimensional coordinates includes: constructing an initial feature map representing the spatial positions of the image feature points based on the three-dimensional coordinates; determining the first predicted position of the spatial positions projected onto the first image data based on the conversion parameter from the navigation data to the camera coordinate system and the second acquisition parameter; determining the first difference between the first true position and the first predicted position; adjusting the spatial positions of the feature points in the initial feature map based on the first difference to obtain the first feature map. In this way, the spatial positions of the feature points and the image pose in the initial feature map are optimized by the first difference, making the obtained first feature map more accurate.
[0014] In some embodiments, updating the first feature map based on the second image data of the area to be drawn to obtain a second feature map includes: updating an image relationship library based on the matching relationship between the second image data and the first image data to obtain an updated image relationship library; determining a to-be-registered image in the updated image relationship library; wherein there are target feature points in the to-be-registered image corresponding to the three-dimensional points in the first feature map; determining the image pose of the to-be-registered image based on the target feature points; registering the image pose into the first feature map to obtain a registered feature map; adjusting the registered feature map based on other feature points in the to-be-registered image and the second image data to obtain the second feature map; wherein the other feature points are feature points other than the target feature points in the to-be-registered image. In this way, the first feature map is updated using second image data from different sources to obtain a second feature map with higher coverage.
[0015] In some embodiments, the registering the image pose into the first feature map to obtain a registered feature map further includes: determining the number of target feature points included in each to-be-registered image; determining the registration order of each to-be-registered image based on the number; registering the image poses of each to-be-registered image into the first feature map based on the registration order to obtain the registered feature map. In this way, by determining the registration order of the to-be-registered images according to the number of target features, the accuracy of the registered image poses can be improved.
[0016] In some embodiments, the adjusting the registered feature map based on other feature points in the to-be-registered image and the second image data to obtain the second feature map includes: sampling the other feature points to obtain sampled feature points; triangulating the sampled feature points to determine the three-dimensional coordinates of the sampled feature points in the world coordinate system; adjusting the registered feature map based on the three-dimensional coordinates of the sampled feature points and the second image data to obtain the second feature map. In this way, by uniformly sampling other feature points in the sampling area, the over-concentration of image features can be reduced and the complexity of global optimization can be lowered.
[0017] In some embodiments, adjusting the registered feature map based on the three-dimensional coordinates of the sampled feature points and the second image data to obtain the second feature map includes: determining a second predicted position where the three-dimensional coordinates of the sampled feature points are projected onto the first image data and a third predicted position where the three-dimensional coordinates of the sampled feature points are projected onto the second image data based on the conversion parameters from the navigation data to the camera coordinate system; determining a second difference between the second predicted position and the second true position of the sampled feature point in the corresponding image, and a third difference between the third predicted position and the second true position; and adjusting the spatial positions of the feature points in the registered feature map based on the second difference and the third difference to obtain the second feature map. In this way, the spatial positions of the feature points in the first feature map are optimized through the second difference and the third difference, making the second feature map more complete.
[0018] In some embodiments, constructing a three-dimensional map of the area to be drawn based on the second image data, the second feature map, and the first image data includes: performing depth processing on the second image data, the second feature map, and the first image data to generate point cloud data representing the first image data and the second image data; constructing an initial mesh model representing the connection relationships between the point cloud data based on the point cloud data; and determining the three-dimensional map based on the initial mesh model. In this way, the high-precision characteristics of in-vehicle data can be retained while combining the high coverage rate of the aerial perspective, making the obtained three-dimensional map more accurate.
[0019] In some embodiments, performing depth processing on the second image data, the second feature map, and the first image data to generate point cloud data representing the first image data and the second image data includes: performing depth estimation on the second image data, the second feature map, and the first image data to obtain depth maps; and fusing the depth maps to obtain the point cloud data. In this way, the obtained point cloud data is more abundant.
[0020] In some embodiments, constructing an initial mesh model representing the connection relationships between the point cloud data based on the point cloud data includes: determining the visibility of each point in the point cloud data in the second feature map and the reprojection error of each point; determining target points in the point cloud data whose visibility and reprojection error meet preset conditions; using the target points as vertices and connecting the vertices to obtain the initial mesh model. In this way, by constructing tetrahedrons from the points in the point cloud data according to the visibility and reprojection error of the points, the obtained initial mesh model can fully reflect the connection relationships between the target points.
[0021] In some embodiments, determining the three-dimensional map based on the initial grid model includes: in the initial grid model, determining the acquisition sources of the point cloud data corresponding to multiple vertices of each face; based on the acquisition sources, determining the weight representing the penetration of each face by the line of sight; and determining the three-dimensional map based on the faces with weights less than a preset threshold. Therefore, by analyzing the weight of a face in a tetrahedron being penetrated by the line of sight and selecting the faces with smaller weights to construct the three-dimensional map, the constructed three-dimensional map is smoother.
[0022] In some embodiments, the acquisition device for the first image data includes an in-vehicle camera, and / or, the acquisition device for the second image data is an aerial photography device.
[0023] An embodiment of the present application provides a map construction device, and the device includes:
[0024] A first acquisition module, configured to acquire first image data of a to-be-drawn area and navigation data corresponding to the first image data;
[0025] A first determination module, configured to determine a first feature map based on the first image data and the navigation data;
[0026] A first update module, configured to update the first feature map based on second image data of the to-be-drawn area to obtain a second feature map; wherein, the acquisition methods of the first image data and the second image data are different;
[0027] A first construction module, configured to construct a three-dimensional map of the to-be-drawn area based on the second image data, the second feature map, and the first image data.
[0028] Correspondingly, an embodiment of the present application provides a computer storage medium, on which computer-executable instructions are stored. After being executed, the computer-executable instructions can implement the above-mentioned method steps.
[0029] An embodiment of the present application provides an electronic device, which includes a memory and a processor. When the processor runs the computer-executable instructions stored on the memory, the steps of the above method can be implemented.
[0030] The embodiments of the present application provide a map construction method, device, equipment and storage medium. For a to-be-drawn area that needs to draw a 3D map, first, by using the navigation data collected for the to-be-drawn area as the image pose of the first image data, a first feature map representing the image feature space pose of the first image data can be quickly constructed; then, by optimizing and updating the first feature map based on the second image data with different acquisition methods, the information in the second feature map can be made more abundant; finally, by combining two types of image data with different acquisition methods with the second feature map, a 3D map with higher coverage and higher accuracy can be constructed. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic flowchart of the implementation process of the map construction method provided by the embodiments of the present application;
[0032] Figure 2A It is another schematic flowchart of the implementation process of the map creation method provided by the embodiments of the present application;
[0033] Figure 2B It is yet another schematic flowchart of the implementation process of the map construction method provided by the embodiments of the present application;
[0034] Figure 3 It is a schematic flowchart of the implementation process of the map construction method provided by the embodiments of the present application;
[0035] Figure 4 It is a schematic flowchart of the implementation process of initially constructing a feature map provided by the embodiments of the present application;
[0036] Figure 5 It is a schematic flowchart of the implementation process of updating a feature map provided by the embodiments of the present application;
[0037] Figure 6 It is a schematic flowchart of the implementation process of building a grid model provided by the embodiments of the present application;
[0038] Figure 7 It is a schematic diagram of the structural composition of the map construction device provided by the embodiments of the present application;
[0039] Figure 8 It is a schematic diagram of the composition structure of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the invention in detail with reference to the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0041] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0042] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0044] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are subject to the following explanations.
[0045] 1) Computer vision: refers to machine vision that uses cameras and computers to replace the human eye to identify, track, and measure targets, etc., and further performs graphic processing to make the computer process into images that are more suitable for human eye observation or transmission to instrument detection.
[0046] 2) Structure-from-Motion (SFM): Since images are disordered, it is necessary to associate images with overlapping relationships. The output of the structure-from-motion part is a set of associated images verified geometrically, and the image projection points corresponding to the map points.
[0047] 3) Feature map: represents the environment using relevant geometric features (such as points, lines, planes), and is mostly used in Simultaneous Localization and Mapping (SLAM) and SFM. In the embodiments of this application, point features are used to represent, including information such as the spatial position, color, feature description, and visible images of the points.
[0048] 4) Triangular mesh: is a type of polygon mesh, and polygon mesh is also known as "Mesh", which is a data structure used in computer graphics to build models for various irregular objects. In essence, it uses a large number of small triangular patches to represent continuous objects in reality.
[0049] The following describes an exemplary application of the map construction device provided in the embodiments of the present application. Among them, the device provided in the embodiments of the present application can be implemented as various types of electronic devices such as a laptop computer, a tablet computer, a desktop computer, and a mobile device (for example, a personal digital assistant, a dedicated messaging device, a portable gaming device) with an image acquisition function.
[0050] Next, an exemplary application will be described when the map construction device is implemented as an electronic device.
[0051] Figure 1 It is a schematic flow chart of the implementation of the map construction method provided in the embodiments of the present application. As Figure 1 shown, it will be described in conjunction with the steps shown in Figure 1 as follows:
[0052] Step S101, in the area to be drawn, obtain the first image data and the navigation data corresponding to the first image data.
[0053] In some embodiments, the image acquisition device of the first image data can be any device or equipment with an image acquisition function; for example, an in-vehicle camera, an in-vehicle camera, or a roadside camera, etc. Obtaining the first image data of the area to be drawn can be through real-time acquisition by the image acquisition device, or can also be the image data sent by other devices received. The first image data can be an outdoor image collected in the area to be drawn, can be a simple image whose picture content includes an outdoor scene, or can also be a complex image whose picture content includes an outdoor scene. The first image data can include multiple images, and the navigation data is the data of the navigation device when each image is collected. The navigation data and the first image data can be synchronously collected. For example, an in-vehicle camera and an in-vehicle navigation device are used to synchronously collect the first image data and the navigation data for the area to be drawn. Or, the navigation data and the first image data can also be asynchronously collected. For example, after collecting the first image data for the area to be drawn, the navigation device is used to collect the navigation data of this area; in this way, the pixel coordinates of each point in the navigation data can be used as the pose of this point in the first image data.
[0054] Step S102, based on the first image data and the navigation data, determine the first feature map.
[0055] In some embodiments, the first feature map represents the spatial positions of the feature points of the first image data. By using the navigation data of the navigation device when collecting the first image data as the initial image pose of the first image data; for multiple images in the first image data, determining the matching feature points in the images; and building an image relationship library based on the matching feature points between the images. By combining the image pose of the first image data obtained by the navigation device and the external parameters of the acquisition device of the first image data, the pose of the acquisition device is determined; and based on this, triangulating the feature points in the first image data to obtain the spatial positions of the feature points in the world coordinate system, thus initially establishing a complete first feature map representing the spatial positions of the feature points in the first image data.
[0056] Step S103: Update the first feature map based on the second image data of the area to be drawn to obtain a second feature map.
[0057] In some embodiments, the acquisition methods of the first image data and the second image data are different, including: the sources of the first image data and the second image data are different, the acquisition devices of the second image data and the first image data are different, or the acquisition angles of the first image data and the second image data are different. For example, the acquisition device of the first image data is an in-vehicle camera, and the acquisition device of the second image data is an aerial photography drone. In some possible implementation manners, the first image data may be data collected on the ground, and the second image data may be data collected in the air. By matching the first image data and the second image data with different sources, and based on the matched first image data and second image data, the first feature map is updated by using the method of incremental structure from motion to obtain a second feature map.
[0058] Step S104: Construct a three-dimensional map of the area to be drawn based on the second image data, the second feature map, and the first image data.
[0059] In some embodiments, first, the first image data and the second image data are converted into point cloud data; then, by analyzing the visibility and reprojection error of each point of the point cloud data, it is determined whether the point can be used as a vertex to construct a tetrahedron; finally, by integrating the characteristics of the first image data and the second image data, the constructed tetrahedron is optimized to obtain a more planar three-dimensional map.
[0060] In the embodiment of the present application, first, by using the collected navigation data as the image pose of the first image data, a first feature map representing the image feature space pose of the first image data can be quickly constructed; then, by using second image data with different acquisition methods to optimize and update the first feature map, the feature points included in the second feature map can be made more abundant; finally, by combining two image data with different acquisition methods with the second feature map, a three-dimensional map with higher coverage and higher accuracy can be constructed.
[0061] In some embodiments, to quickly construct the first feature map, step S102 above can be implemented through Figure 2A the steps shown in Figure 2A which is another schematic diagram of the implementation process of the map creation method provided by the embodiment of the present application. In combination with Figure 1 and 2A the following description is given:
[0062] Step S201: Extract features from the first image data to obtain image feature points and description information of the image feature points.
[0063] In some embodiments, for each image in the first image data, feature points of Scale Invariant Feature Transform (SIFT) are extracted; and description information of the feature points is obtained. The description information includes the position of the feature point in the image (for example, the two-dimensional coordinates in the image), the scale invariant of the feature point, the rotation invariant, etc. Among the image feature points extracted from the first image data, the position of the image feature point in the image is included, which can be represented in the form of two-dimensional coordinates.
[0064] Step S202: Based on the description information of the image feature points, match different images in the first image data to obtain an image relationship library representing the matching relationship between different images.
[0065] In some embodiments, by analyzing whether the description information of the image feature points of different images is consistent, it can be determined whether the image feature points in the different images are the same feature points. Furthermore, based on the number of feature points with consistent description information, the similarity between the different images can be determined. For example, for any two images (Image A and Image B) in the first image data, by analyzing the number of image feature points with consistent description information in Image A and Image B, the similarity between Image A and Image B can be determined. If there are more image feature points with consistent description information between Image A and Image B, it is determined that the similarity between Image A and Image B is higher and the tightness is higher. Based on the similarity between the different images, an image relationship library is created. The similarity between the image feature points of different images is used as the association between the image feature points. If there are more image feature points with consistent description information between two images, it can be considered that the connection between the two images is stronger. Based on such an association relationship and the images in the first image data, an image relationship library capable of representing the similarity between different images is built. For example, first, extract the SIFT features of the images in the input first image data. Then, establish the association between the feature points of different images. Finally, construct a scene graph as the image relationship library using the bag-of-words model.
[0066] Step S203: Based on the image relationship library and the navigation data, determine the first feature map.
[0067] In some embodiments, by first using the navigation data as the initial image pose of the first image data, and combining the external parameters of the navigation device that collects the navigation data and the external parameters of the acquisition device that collects the first image data, the pose of the acquisition device is determined. Then, the two-dimensional coordinates of the image feature points in the image relationship library (i.e., feature observations), the pose of the acquisition device, and the internal parameters of the acquisition device are used as the input for triangulation processing, so that the spatial position of each feature point in the world coordinate system can be determined. Furthermore, a feature map representing the spatial position of the image feature points in the first image data, that is, the first feature map, is preliminarily constructed.
[0068] In the embodiments of the present application, by matching the description information of the image feature points in different images, the similarity between different images can be obtained. Thus, according to the similarity between different images, a bag-of-words model is used to construct an image relationship library. Furthermore, by combining the feature point coordinates in the image relationship library with the navigation data, a first feature map representing the spatial position of the image feature points in the image data can be simply and quickly constructed.
[0069] In some embodiments, by combining the navigation data with the parameters of the set acquisition device and determining the spatial position of the image feature points through triangulation, that is, the above step 204 can be implemented through the following steps:
[0070] Step S141: Determine the first pose of the first acquisition device of the first image data based on the navigation data and the preset first acquisition parameters.
[0071] In some embodiments, the preset first acquisition parameters include: the external parameters for acquiring the first image data and the external parameters for acquiring the navigation data. For example, if the device for acquiring the first image data is an in-vehicle camera and the device for acquiring the navigation data is a global satellite inertial navigation system, then determine the external parameters of the in-vehicle camera and the external parameters of the global satellite inertial navigation system. Taking the navigation data as the initial image pose of the first image data, and combining the external parameters of the in-vehicle camera and the external parameters of the global satellite inertial navigation system, the pose of the in-vehicle camera, i.e., the first pose C, can be solved. i 。
[0072] Step S142: Determine the first feature map based on the image feature points, the first pose, and the second acquisition parameters in the relational database.
[0073] In some embodiments, the second acquisition parameters include the internal parameters for acquiring the first image data. The second acquisition parameter is the internal parameter of the device for acquiring the first image data. For example, if the device is an in-vehicle camera, then the internal parameter can be the camera internal parameter matrix. Based on the camera internal parameter matrix, the two-dimensional coordinates of the image feature points, and the pose of the camera as the first pose, triangulate the image feature points to determine the three-dimensional coordinates of the image feature points in the world coordinate system, i.e., the spatial position of the image feature points. In this way, through the feature matching relationship and the initial image pose provided by the navigation data, based on the triangulation principle, the feature map of the image data can be quickly constructed.
[0074] In some possible implementation manners, to improve the accuracy of the finally obtained first feature map, optimize and adjust the initial feature map through the predicted coordinates of the image feature points and the coordinates of the image feature points in the image to obtain the first feature map. That is, the above step S142 can be implemented through the following steps:
[0075] The first step: Triangulate the first true position of the image feature points in the image to which they belong based on the first pose and the second acquisition parameters to obtain the three-dimensional coordinates of the image feature points in the world coordinate system.
[0076] In some embodiments, use the first true position of the image feature points in the image to which they belong, the first pose of the acquisition device of the first image data, and the internal parameter matrix of the acquisition device as the input for triangulating the image feature points to obtain the three-dimensional coordinates of the image feature points in the world coordinate system, thereby realizing the triangulation of each image feature point.
[0077] Step 2: Based on the three-dimensional coordinates, construct the first feature map.
[0078] In some embodiments, after obtaining the three-dimensional coordinates of each image feature point in the image, construct a feature map of the image feature point to represent the spatial position of the image feature point with the feature map; the feature map can be directly used as the first feature map for subsequent map updates; it can also be the first feature map obtained by optimizing the feature point coordinates and image pose of the feature map. In this way, by triangulating the true position of the image feature point, the triangulated feature point is obtained, and the first feature map can be quickly constructed based on the three-dimensional coordinates of the triangulated feature point.
[0079] In some possible implementation manners, after triangulating the image feature points, initially construct an initial feature map, and obtain the first feature map by optimizing and adjusting the initial feature map, so that the accuracy of the first feature map is more accurate. That is, the above Step 2 can be implemented through the following process:
[0080] First, based on the three-dimensional coordinates, construct an initial feature map representing the spatial position of the image feature point.
[0081] Secondly, based on the conversion parameters of the navigation data to the camera coordinate system and the second acquisition parameter, determine the first predicted position where the spatial position is projected onto the first image data.
[0082] In some embodiments, the conversion parameters of the navigation data to the camera coordinate system include: the pose of the navigation device of the navigation data (including the direction R i and position t i ), the rotation matrix and translation vector from the navigation device of the navigation data to the camera coordinate system; for example, when the navigation device is a global satellite inertial navigation system, the conversion parameters include: the direction R i and position t i , the rotation matrix and translation vector from the satellite inertial navigation system to the camera coordinate system (T R and T t ). The second acquisition parameter is the internal parameter matrix K of the acquisition device for acquiring the first image data, and the three-dimensional coordinate X j of the j-th image feature point in the initial feature map. Based on this, determine the first predicted position where the image feature point is projected from the spatial position to the image as K(T R R i X j + T R t i + T t ).
[0083] Next, determine a first difference between the first true position and the first predicted position.
[0084] In some embodiments, the first true position x of an image feature point in the relational database in the corresponding image ij , by determining the difference between the first predicted position and the first true position, the rationality of the spatial position of the feature point in the initial feature map can be estimated; based on this difference, the image feature points with unreasonable spatial positions in the initial feature map are optimized and adjusted, so as to obtain a first feature map with higher accuracy. For example, for each feature point in the initial feature map, determine the first predicted position where the feature point is projected back to the image from the spatial position, that is, the position predicted in the first image data; based on the first true position of the feature point in the image, estimate the prediction error of the first predicted position.
[0085] Finally, based on the first difference, adjust the spatial positions of the feature points in the initial feature map to obtain the first feature map.
[0086] In some embodiments, based on this first difference, the three-dimensional coordinates of the image feature points in the initial feature map and the image pose can be adjusted to complete the construction of the feature map based on the vehicle-mounted acquisition data. In this way, by determining the first predicted position where the feature point in the initial feature map is projected back to the image from the spatial position and the first true position of the feature point in the image, the spatial position and the image pose of the feature points in the initial feature map can be optimized, so that the accuracy of the optimized first feature map is higher.
[0087] In some embodiments, to improve the high coverage of the constructed first feature map, the first feature map is updated using second image data from different sources to obtain a second feature map with higher coverage; that is, the above step S103 can be implemented through the steps as Figure 2B shown, Figure 2B is another schematic implementation flowchart of the map construction method provided by the embodiments of the present application. In combination with Figure 1 and 2B shown steps, the following description is made:
[0088] Step S221, update the image relationship database based on the matching relationship between the second image data and the first image data to obtain an updated image relationship database.
[0089] In some embodiments, the method of updating the image relationship library using the second image data is substantially the same as the implementation manners of step S201 and step S202. Taking the second image data as aerial image data and the first image data as ground image data collected by a vehicle-mounted camera as an example, by performing feature extraction on the second image data, according to the description information of the feature points extracted and the description information of the feature points in the first image data, the similarity between the images in the first image data and the second image data is determined, that is, the matching relationship between the images is obtained. Based on this matching relationship, the image relationship library established according to the matching relationship between the images in the first image data is updated, so that the similarity degree between the images in the first image data can be reflected in the updated image relationship library, and the similarity degree between the first image data and the second image data can also be reflected.
[0090] Step S222, determine the image to be registered in the updated image relationship library.
[0091] In some embodiments, there are target feature points corresponding to the three-dimensional points in the first feature map in the image to be registered. In the updated image relationship library, for each image, determine whether the feature points in the image can find corresponding three-dimensional points in the first feature map. If the feature points in the image can find corresponding three-dimensional points in the first feature map, determine that the image is the image to be registered.
[0092] Step S223, determine the image pose of the image to be registered based on the target feature points.
[0093] In some embodiments, for these target feature points that can find corresponding three-dimensional points in the first feature map, a random sampling method is used to determine the image pose of the image to which the target feature point belongs, that is, the image pose of the image to be registered.
[0094] Step S224, register the image pose into the first feature map to obtain a registered feature map.
[0095] In some embodiments, the image pose determined according to the target feature points is registered into the first feature map so that the spatial position of the picture content of the image to be registered to which the target feature point belongs can be presented in the registered feature map. In this way, by combining the first image data and the second image data from different sources, the advantages of vehicle-mounted camera and drone data can be effectively utilized to improve the coverage and integrity of mapping.
[0096] In some possible implementation manners, by counting the number of target feature points, the registration order of the images to be registered is determined, and thus the image poses of the images to be registered are sequentially registered into the first feature map according to this registration order, which can be implemented through the following steps:
[0097] Step 1: Determine the number of target feature points included in each image to be registered.
[0098] For each frame of the image to be registered, count the number of target feature points in the image to be registered. The more target feature points in the image to be registered, the higher the similarity between the image to be registered and the area to be drawn, that is, the higher the possibility that the image to be registered is an image collected for the area to be drawn, and it also indicates a higher overlap degree between the image to be registered and the three-dimensional points in the first feature map. Then, the image to be registered can be preferentially registered.
[0099] Step 2: Based on the quantity, determine the registration order of each image to be registered.
[0100] Determine the registration order of the images to be registered in descending order of the quantity. Arrange the image to be registered with the largest number of target features first and preferentially register the image pose.
[0101] Step 3: Based on the registration order, register the image pose of each image to be registered into the first feature map to obtain the registered feature map.
[0102] For example, for the image to be registered with the largest number of target features, first register the image pose into the first feature map, and then register the image pose of the image to be registered with the second largest number of target features into the first feature map. In this way, by determining the registration order of the images to be registered according to the number of target features, the accuracy of the registered image pose can be improved.
[0103] Step S225: Based on the other feature points in the image to be registered and the second image data, adjust the registered feature map to obtain the second feature map.
[0104] In some embodiments, the other feature points are the feature points other than the target feature points in the image to be registered.
[0105] In some possible implementation manners, to further optimize the registered feature map, by upsampling the other feature points and performing triangulation processing on one sampled feature point, and optimizing and adjusting the registered feature map according to the result of the triangulation processing, the above step S225 further includes the following steps:
[0106] Step 1: Sample the other feature points to obtain sampled feature points.
[0107] Perform uniform sampling on the other feature points in the image to be registered, and retain a small number of feature points in each sampling area. For example, retain one sampled feature point in each sampling area. In this way, it is possible to reduce the over-concentration of image features and reduce the complexity of global optimization.
[0108] In the second step, triangulate the sampled feature points to determine the three-dimensional coordinates of the sampled feature points in the world coordinate system.
[0109] Triangulate the sampled feature points by using the pose of the acquisition device corresponding to the sampled feature point, the internal parameter matrix of the acquisition device, and the coordinates of the sampled feature point in the image, so as to determine the spatial pose of the sampled feature point, that is, the three-dimensional coordinates of the sampled feature point in the world coordinate system.
[0110] In the third step, adjust the registered feature map based on the three-dimensional coordinates of the sampled feature points and the second image data to obtain the second feature map.
[0111] By combining the spatial position of the sampled feature point and the pose of the acquisition device of the second image data, etc., the difference between the actual position and the estimated spatial position of the sampled feature point can be further predicted, and then the position of the three-dimensional points in the registered feature map can be optimized based on this difference to obtain the second feature map. In some possible implementation manners, it can be implemented through the following steps:
[0112] In the first step, based on the conversion parameters from the navigation data to the camera coordinate system, determine the second predicted position where the three-dimensional coordinates of the sampled feature point are projected onto the first image data and the third predicted position where the three-dimensional coordinates of the sampled feature point are projected onto the second image data.
[0113] The method for determining the second predicted position is similar to the method for determining the first predicted position. Determine the rotation matrix and translation vector (T R and T t ) from the navigation data to the camera coordinate system, the direction R i and position t i of the navigation device, the internal parameter matrix K C of the acquisition device for acquiring the first image data (for example, an in-vehicle camera), and the three-dimensional coordinates X j of the sampled feature point; based on this, determine that the second predicted position where the sampled feature point is projected from the spatial position to the first image is K C (T R R i X j +T R t i +T t ).
[0114] Based on the internal parameter matrix K A of the acquisition device for acquiring the second image data (for example, an aerial drone), the direction R i and position t i of the navigation device, and the three-dimensional coordinates X j, determine the third predicted position K in the second image where the sampling feature point is projected from the spatial position A (R k X j +t k ).
[0115] In the second step, determine the second difference between the second predicted position and the second true position of the sampling feature point in the image to which it belongs, and the third difference between the third predicted position and the second true position.
[0116] For each sampling feature point obtained in each sampling area, determine the second predicted position where the sampling feature point is projected back from the spatial position to the first image data, that is, the position predicted in the first image data; based on the second true position of the sampling feature point in the image to which it belongs, estimate the prediction error of the second predicted position. Similarly, for this sampling feature point, estimate the prediction error at its third predicted position, that is, the third difference.
[0117] In the third step, based on the second difference and the third difference, adjust the spatial positions of the feature points in the registered feature map to obtain the second feature map.
[0118] Based on this second difference, the three-dimensional coordinates corresponding to the image feature points in the first image data in the registered feature map, and the image pose of the first image data can be adjusted. Based on this third difference, the three-dimensional coordinates corresponding to the image feature points in the second image data in the registered feature map, and the image pose of the second image data can be adjusted. In this way, by uniformly sampling in each sampling area to obtain a sampling feature point, and through the difference between the predicted position of the sampling feature point and its actual position in the image, the spatial positions of the feature points in the first feature map are optimized, making the integrity of the second feature map higher.
[0119] In some embodiments, based on the second feature map, by performing depth recovery, point cloud generation, mesh construction, etc. on the image data, a high-precision three-dimensional map is finally created. That is, step S105 above can be implemented through the following steps:
[0120] Step S151, perform depth processing on the second image data, the second feature map, and the first image data to generate point cloud data representing the first image data and the second image data.
[0121] In some embodiments, the image pose of the second image is included in the second image data, and the image pose of the first image is included in the first image data; the first type of three-dimensional feature points corresponding to the image feature points in the first image data in the second feature map, and the second type of three-dimensional feature points corresponding to the image feature points in the second image data in the second feature map are included in the second feature map. The first image pose and the first type of three-dimensional feature points are taken as a group for depth estimation and depth map fusion to obtain the point cloud data representing the first image data; the second image pose and the second type of three-dimensional feature points are taken as a group for depth estimation and depth map fusion to obtain the point cloud data representing the second image data.
[0122] In some possible implementation manners, the point cloud data can be generated through the following steps:
[0123] First step, perform depth estimation on the second image data, the second feature map, and the first image data to obtain a depth map.
[0124] Combine the second image, the second image pose, and the second type of three-dimensional feature points in the second image data for depth estimation to obtain the depth map of the second image; combine the first image, the first image pose, and the first type of three-dimensional feature points in the first image data for depth estimation to obtain the depth map of the first image. The depth map of the first image and the depth map of the second image are used as the depth map obtained in this step.
[0125] Second step, fuse the depth maps to obtain the point cloud data.
[0126] Perform depth fusion on the depth map of the first image and the depth map of the second image respectively to generate the point cloud data corresponding to the first image and the point cloud data corresponding to the second image; perform point cloud integration on these two types of point cloud data to obtain the final point cloud data. In this way, through depth estimation and depth fusion of two types of images, the obtained point cloud data is richer.
[0127] Step S152, based on the point cloud data, construct an initial mesh model representing the connection relationship between the point cloud data.
[0128] In some embodiments, the connection relationship between the point cloud data is whether the points are connected in the gateway model. According to the point cloud data integrated based on the two types of image data, and the visibility and reprojection error of the points, determine whether the points can be used as vertices to construct tetrahedrons, so that multiple tetrahedrons can be constructed based on multiple vertices to serve as the initial mesh model.
[0129] In some possible implementation manners, the primary gateway model can be built through the following steps:
[0130] First step, determine the visibility of each point in the point cloud data in the second feature map and the reprojection error of each point.
[0131] The visibility of each point in the point cloud data in the second feature map, that is, whether there is a corresponding 3D feature point for this point in the second feature map. If there is a corresponding 3D feature point for this point in the second feature map, it means this point is visible; if there is no corresponding 3D feature point for this point in the second feature map, it means this point is invisible. In some embodiments, regarding the reprojection error, through feature matching pairs, it can be known that the observed values A and B are a set of feature matching pairs, and they are projections of the same spatial point C. A belongs to one map and B belongs to another map. B` is the projection point on the coordinate system to which B belongs after converting A to the first global coordinate in the coordinate system to which B belongs. There is a certain distance between the projection B` of A and the observed value B, and this is the reprojection error. Among them, the reprojection error of each point in the point cloud data is the projections A and B of this point in two types of images. For example, A is in the first image data and B is in the second image data; the reprojection error of this point is the distance between the projection point on the coordinate system to which B belongs after converting A to the first global coordinate in the coordinate system to which B belongs and the observed value B.
[0132] Second step, in the point cloud data, determine the target points whose visibility and reprojection error meet the preset conditions.
[0133] In the point cloud data, determine that the visibility of a point is that this point can find a corresponding 3D feature point in the second feature map, and the target points with a reprojection error less than a certain threshold, and multiple target points can be obtained.
[0134] Third step, use the target points as vertices and connect the vertices to obtain the initial mesh model.
[0135] Use these multiple target points as vertices. Connecting any vertices in space can form a tetrahedron, obtaining multiple tetrahedrons, and use this tetrahedron as the initial mesh model. In this way, by constructing tetrahedrons from the points in the point cloud data according to the visibility and reprojection error of the points, the obtained initial mesh model can fully reflect the connection relationship between the target points.
[0136] Step S153, based on the initial mesh model, determine the 3D map.
[0137] In some embodiments, since the initial mesh model includes multiple tetrahedrons, according to the confidence that each face in the tetrahedron is the surface of the real object, among the faces of these multiple tetrahedrons, an optimal set of faces is selected, and this optimal set of faces is connected to form a 3D map. In this way, through depth recovery, point cloud generation, and mesh construction, a high-precision dense point cloud and mesh model are finally produced, which can not only retain the high-precision characteristics of vehicle-mounted data but also combine the advantage of high coverage rate of the aerial photography perspective, making the obtained 3D map more accurate.
[0138] In some possible implementation manners, the construction of the 3D map can be achieved through the following steps:
[0139] First step, in the initial mesh model, determine the acquisition sources of the point cloud data corresponding to the multiple vertices of each face.
[0140] Among the multiple faces of the multiple tetrahedrons included in the initial mesh model, determine the sources of the three vertices of each face. Taking this face as a triangular patch, analyze the acquisition status of the three vertices of this face. Whether the three vertices come from different vehicle-mounted cameras, or whether all three vertices come from the same vehicle-mounted camera or navigation device (such as an aerial photography drone).
[0141] Second step, based on the acquisition sources, determine the weight representing the penetration of each face by the line of sight.
[0142] If the three vertices come from different vehicle-mounted cameras, it indicates that the three vertices of this triangular patch are less likely to belong to the same object, that is, the triangular patch is more likely to be the surface of the real object; then the weight of the penetration of this patch by the line of sight is greater; where the line of sight is a line emitted from the acquisition device to the patch.
[0143] Third step, based on the faces with weights less than the preset threshold, determine the 3D map.
[0144] For each patch, a weight representing the penetration of this face by the line of sight is determined, and a set of faces with the smallest sum of weights is selected; these faces are connected to obtain the 3D map. In this way, by analyzing the weight of the penetration of a face in the tetrahedron by the line of sight and selecting the faces with smaller weights to construct the 3D map, the constructed 3D map is smoother.
[0145] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described. Taking the construction of a high-precision map based on a global satellite inertial navigation system, a camera, and an aerial photography drone as an example, it will be described.
[0146] With the continuous development of computer and communication technologies, maps have provided great assistance for people's travel. The information they provide includes, but is not limited to, road information, building information, traffic information, etc. Among them, high-precision maps have attracted much attention due to their centimeter-level reconstruction accuracy and rich hierarchical information. High-precision maps are mostly used in the field of autonomous driving. In related technologies, maps are mostly generated by conventional collection plus manual post-processing methods. For example, both high-precision professional equipment collection and crowdsourcing collection methods require manual alignment and other operations after data collection. Moreover, the collection of high-precision map data is mostly carried out by collection vehicles within the range of paved roads, and the reconstruction accuracy of non-road areas is limited, making it difficult to meet the requirements for constructing high-precision maps of the entire area. The integrity and hierarchical structure of the overall reconstruction are not yet perfect.
[0147] Based on this, in the embodiments of the present application, combining the advantages of comprehensive coverage and convenient collection of aerial drones with ground collection, a method for constructing a high-precision map with full-area coverage of non-homogeneous data from the air and the ground is proposed. In the embodiments of the present application, a method and a collection system for automatically completing the construction of a high-precision map can be achieved by using an on-vehicle global satellite inertial navigation system, a camera, and an aerial drone.
[0148] Figure 3 FIG. is a schematic flowchart of the implementation process of the map construction method provided by the embodiments of the present application. From Figure 3 it can be seen that this method can be implemented through the following steps:
[0149] Step S301, obtain the pose data of the on-vehicle global satellite inertial navigation system, and obtain the on-vehicle camera image data.
[0150] Calibrate and synchronize the pose data of the on-vehicle global satellite inertial navigation system and the on-vehicle camera image data. The pose data of the on-vehicle global satellite inertial navigation system can be used as the navigation data in the above embodiments.
[0151] Step S302, construct a feature map of the on-vehicle collected data based on the pose data and the camera image.
[0152] In some embodiments, the feature map of the on-vehicle collected data can be used as the first feature map in the above embodiments. By combining the feature matching relationship between the images collected by the on-vehicle camera with the pose provided by the global satellite inertial navigation system, based on the triangulation principle, a feature map of the ground data is constructed. It mainly includes the following three processes: feature extraction and matching; triangulation of feature points; adjustment and optimization. As Figure 4 shown, Figure 4 FIG. is a schematic flowchart of the implementation process of the preliminary construction of the feature map provided by the embodiments of the present application. In combination with Figure 4 the steps shown below for the following description:
[0153] Step S401: Extract and match features from the in-vehicle camera image data 41.
[0154] In some embodiments, the matching quantity of Scale Invariant Feature Transform (SIFT) feature points is used to measure the tightness of image association. First, extract the SIFT features of the input image; then, establish the association of feature points between images; finally, construct a Scene Graph using the bag-of-words model.
[0155] Step S402: Triangulate the initial feature map based on the pose data 42 of the global satellite inertial navigation system and the external parameters of the camera.
[0156] Combine the pose data obtained by the global satellite inertial navigation system with the external parameters of the in-vehicle global satellite inertial navigation system and the camera to solve the pose (R i and t i ) of camera C i and triangulate the feature points based on this. Initially construct the initial feature map of the in-vehicle data.
[0157] Step S403: Optimize and adjust the triangulated initial feature map.
[0158] The process of adjusting the feature point coordinates and image poses of the initial feature map is shown in formula (1):
[0159]
[0160] where n is the number of images, m is the number of map points, K is the camera intrinsic matrix, T R and T t are the rotation matrix and translation vector from the satellite inertial navigation system to the camera coordinate system respectively, X is the three-dimensional coordinate of the map point, x is the feature observation,
[0161] Thus, optimize and update the feature point coordinates and image poses of the feature map to achieve the construction link of the feature map based on the in-vehicle acquisition data.
[0162] Step S303: Update the feature map based on the UAV aerial photography data 31.
[0163] In some embodiments, extract the feature points of the images obtained by UAV aerial photography and match them with the in-vehicle camera image data to update the scene graph and the feature map. This process mainly includes two parts: feature extraction and matching of aerial photography images, and feature map update.
[0164] Part 1: Similar to step S401, extract the features of the aerial images obtained by the drone's aerial photography, match the image features with the features of the on-vehicle camera images, and update the original scene graph.
[0165] Combine the new scene graph with the existing feature map, and use the method of incremental structure from motion to update the feature map, which can be achieved through Figure 5 the steps shown as follows:
[0166] Step S501: Obtain the first feature map created based on the on-vehicle collected images and the updated scene graph.
[0167] Step S502: Select the image of the frame to be registered in the updated scene graph.
[0168] The selection of the frame to be registered is mainly based on the number of visible map points. The number of visible map points reflects the similarity degree of the observed area. When most of the feature points in the newly registered image can find corresponding 3D points in the current feature map, it indicates a high degree of observation overlap and can be preferentially registered.
[0169] Step S503: Triangulate other feature points in the image to be registered.
[0170] For the image to be registered, for the feature points with map observations, use pnp+ransac to solve the image pose and register it to the feature map. Then, uniformly sample the feature points without map observations, and only retain one feature point in each sampling area. Solve the spatial position of the feature points through triangulation. On the basis of ensuring the complete coverage of the feature map, avoid the over-concentration of image features and reduce the complexity of global optimization.
[0171] Step S504: Based on the triangulated feature points, perform local / global optimization adjustment on the first feature map to obtain the second feature map.
[0172] The process of optimizing and adjusting the 3D feature points and image poses in the first feature map is shown in formula (2):
[0173]
[0174] where O is the number of aerial images, K A is the internal parameter of the aerial image, R k and t k are the rotation matrix and translation vector corresponding to the image, v kj and v ij both represent the visibility of feature point j in the corresponding image, and the meanings of the remaining parameters are the same as those in formula (1).
[0175] Step S304: Generate a mesh model based on the high-precision dense point cloud of sensor fusion.
[0176] Based on the existing feature maps, through depth recovery, point cloud generation, and mesh construction, a high-precision dense point cloud and mesh model are finally produced. Based on multi-view geometry, the vehicle-mounted acquisition data and UAV aerial photography data are jointly processed, retaining the characteristics of high precision of the vehicle-mounted data and the advantage of high aerial photography perspective coverage rate. The implementation process is as Figure 6 shown. For the vehicle-mounted camera image data 61, the second feature map, and the image pose 62, and the UAV aerial photography image data 63, the depth estimation and depth map fusion are independently completed, as well as the two processes of point cloud merging to generate a dense point cloud and jointly constructing a mesh model.
[0177] Step S601: Perform depth estimation on the vehicle-mounted camera image data 61, the image pose of this data, and the three-dimensional feature points corresponding to this data in the second feature map to obtain a depth map.
[0178] Step S602: Perform depth map fusion on the depth map obtained in this step S601 to obtain the point cloud of the vehicle-mounted camera image data 61.
[0179] Step S603: Perform depth estimation on the UAV aerial photography image data 63, the image pose of this data, and the three-dimensional feature points corresponding to this image in the second feature map to obtain a depth map.
[0180] Step S604: Perform depth map fusion on the depth map obtained in step S603 to obtain the point cloud of the UAV aerial photography image 63.
[0181] Using the depth estimation method from the image and the corresponding pose, independently estimate the depth maps of the images collected by the vehicle and the UAV. The obtained depth maps are fused into point clouds to generate the point cloud of the vehicle-mounted acquisition data and the point cloud of the UAV aerial photography data. Integrate the point clouds obtained in step S602 and step S604 to obtain point cloud data.
[0182] Step S605: Integrate the point cloud obtained by the vehicle-mounted camera acquisition with the point cloud obtained by the UAV aerial photography to obtain point cloud data.
[0183] This point cloud data is a dense point cloud.
[0184] Step S606: Generate a mesh model based on this point cloud data.
[0185] Step S607: Optimize and adjust the mesh model according to the set weight to obtain a three-dimensional map.
[0186] The point cloud insertion process integrates the point cloud collected by the vehicle-mounted camera and the point cloud obtained from the UAV aerial photography. On this basis, it is judged whether to use this point as a vertex according to two indicators: the visibility of the point and the reprojection error, and a tetrahedron is constructed. When using the Graph-Cut algorithm to extract triangular patches from the tetrahedron set, combined with the characteristics of UAV aerial photography and vehicle-mounted camera data, a weight term is added, as shown in formula (3):
[0187]
[0188] where b << r, making the finally generated 3D map smoother.
[0189] In the embodiments of the present application, first, the pose obtained by the global satellite inertial navigation system is used as the initial value, and combined with the final adjustment and optimization, a feature map is quickly constructed. In this way, the iterative registration and optimization processes are simplified, and the inertial navigation information is fully utilized to improve the reconstruction efficiency. Then, combining the advantages of different cameras, making full use of the advantages of high coverage of aerial images and sufficient observation of vehicle-mounted data, a high-precision and high-coverage dense point cloud and grid model are constructed; in this way, the coverage and integrity of mapping can be improved, and a more complete 3D model is finally generated.
[0190] The embodiments of the present application provide a map construction device, Figure 7 which is a schematic structural composition diagram of the map construction device provided by the embodiments of the present application, as Figure 7 shown. The map construction device 700 includes:
[0191] A first acquisition module 701, configured to acquire first image data of an area to be drawn and navigation data corresponding to the first image data;
[0192] A first determination module 702, configured to determine a first feature map based on the first image data and the navigation data;
[0193] A first update module 703, configured to update the first feature map based on second image data of the area to be drawn to obtain a second feature map; wherein, the acquisition methods of the first image data and the second image data are different;
[0194] A first construction module 704, configured to construct a 3D map of the area to be drawn based on the second image data, the second feature map, and the first image data.
[0195] In some embodiments, the first determination module 702 includes:
[0196] A first extraction sub-module, configured to extract features from the first image data to obtain image feature points and description information of the image feature points;
[0197] A first matching submodule, configured to match different images in the first image data based on the description information of the image feature points, and obtain an image relationship library representing the matching relationship between the different images;
[0198] The first determination submodule is used to determine the first feature map based on the image relationship library and the navigation data.
[0199] In some embodiments, the first determining submodule includes:
[0200] A first determination unit is configured to determine a first pose of a device for acquiring the first image data based on the navigation data and a preset first acquisition parameter; wherein the preset first acquisition parameter includes: an extrinsic parameter for acquiring the first image data and an extrinsic parameter for acquiring the navigation data;
[0201] A second determination unit is configured to determine the first feature map based on the image feature points in the relational database, the first posture, and second acquisition parameters; wherein the second acquisition parameters include intrinsic parameters for acquiring the first image data.
[0202] In some embodiments, the second determining unit includes:
[0203] A first determining subunit is used to triangulate a first true value position of the image feature point in the corresponding image based on the first pose and the second acquisition parameter to obtain a three-dimensional coordinate of the image feature point in a world coordinate system;
[0204] The first construction subunit is used to construct the first feature map based on the three-dimensional coordinates.
[0205] In some embodiments, the first construction subunit is further used to: construct an initial feature map representing the spatial position of the image feature point based on the three-dimensional coordinates; determine a first predicted position of the spatial position projected to the first image data based on the conversion parameters of the navigation data to the camera coordinate system and the second acquisition parameters; determine a first difference between the first true value position and the first predicted position; and adjust the spatial position of the feature point in the initial feature map based on the first difference to obtain the first feature map.
[0206] In some embodiments, the first updating module 703 includes:
[0207] A first updating submodule, configured to update an image relationship library based on a matching relationship between the second image data and the first image data to obtain an updated image relationship library;
[0208] A second determination sub-module, configured to determine a to-be-registered image in the updated image relationship library; wherein, there are target feature points in the to-be-registered image corresponding to the three-dimensional points in the first feature map;
[0209] A third determination sub-module, configured to determine the image pose of the to-be-registered image based on the target feature points;
[0210] A first registration sub-module, configured to register the image pose into the first feature map to obtain a registered feature map;
[0211] A first adjustment sub-module, configured to adjust the registered feature map based on other feature points in the to-be-registered image and the second image data to obtain the second feature map; wherein, the other feature points are feature points other than the target feature points in the to-be-registered image.
[0212] In some embodiments, the first registration sub-module further includes:
[0213] A third determination unit, configured to determine the number of target feature points included in each to-be-registered image;
[0214] A fourth determination unit, configured to determine the registration order of each to-be-registered image based on the number;
[0215] A first registration unit, configured to register the image pose of each to-be-registered image into the first feature map based on the registration order to obtain the registered feature map.
[0216] In some embodiments, the first adjustment sub-module includes:
[0217] A first sampling unit, configured to sample the other feature points to obtain sampled feature points;
[0218] A fifth determination unit, configured to triangulate the sampled feature points to determine the three-dimensional coordinates of the sampled feature points in the world coordinate system;
[0219] A first adjustment unit, configured to adjust the registered feature map based on the three-dimensional coordinates of the sampled feature points and the second image data to obtain the second feature map.
[0220] In some embodiments, the first adjustment unit includes:
[0221] A second determination sub-unit, configured to determine a second predicted position of the three-dimensional coordinates of the sampled feature points projected onto the first image data and a third predicted position of the three-dimensional coordinates of the sampled feature points projected onto the second image data based on the conversion parameters from the navigation data to the camera coordinate system;
[0222] A third determination subunit, configured to determine a second difference between the second predicted position and a second true position of the sampling feature point in the image to which it belongs, and a third difference between the third predicted position and the second true position;
[0223] A first adjustment subunit, configured to adjust the spatial positions of the feature points in the registered feature map based on the second difference and the third difference, to obtain the second feature map.
[0224] In some embodiments, the first construction module 704 includes:
[0225] A first processing sub-module, configured to perform depth processing on the second image data, the second feature map, and the first image data, to generate point cloud data representing the first image data and the second image data;
[0226] A first construction sub-module, configured to construct an initial mesh model representing the connection relationship between the point cloud data based on the point cloud data;
[0227] A second determination sub-module, configured to determine the three-dimensional map based on the initial mesh model.
[0228] In some embodiments, the first processing sub-module includes:
[0229] A first estimation unit, configured to perform depth estimation on the second image data, the second feature map, and the first image data, to obtain a depth map;
[0230] A first fusion unit, configured to fuse the depth map, to obtain the point cloud data.
[0231] In some embodiments, the first construction sub-module includes:
[0232] A sixth determination unit, configured to determine the visibility of each point in the point cloud data in the second feature map and the reprojection error of each point;
[0233] A seventh determination unit, configured to determine target points in the point cloud data whose visibility and reprojection error meet preset conditions;
[0234] A first connection unit, configured to use the target points as vertices, and connect the vertices, to obtain the initial mesh model.
[0235] In some embodiments, the second determination sub-module includes:
[0236] An eighth determination unit, configured to determine the acquisition sources of the point cloud data corresponding to the multiple vertices of each face in the initial mesh model;
[0237] A ninth determination unit, configured to determine a weight representing that each surface is penetrated by a line of sight based on the acquisition source;
[0238] A tenth determination unit, configured to determine the three-dimensional map based on the surfaces with weights less than a preset threshold.
[0239] In some embodiments, the acquisition device of the first image data includes an in-vehicle camera, and / or, the acquisition device of the second image data is an aerial photography device.
[0240] It should be noted that the description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0241] It should be noted that in the embodiments of the present application, if the above map construction method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a terminal, a server, etc.) to execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0242] Correspondingly, the embodiments of the present application further provide a computer program product, where the computer program product includes computer-executable instructions, and after the computer-executable instructions are executed, the steps in the map construction method provided by the embodiments of the present application can be implemented.
[0243] Correspondingly, the embodiments of the present application further provide a computer storage medium, where computer-executable instructions are stored on the computer storage medium, and when the computer-executable instructions are executed by a processor, the steps of the map construction method provided by the above embodiments are implemented.
[0244] Correspondingly, the embodiments of the present application provide an electronic device, Figure 8 is a schematic structural diagram of the electronic device provided by the embodiments of the present application, as Figure 8As shown, the electronic device 800 includes: a processor 801, at least one communication bus, a communication interface 802, at least one external communication interface, and a memory 803. Among them, the communication interface 802 is configured to implement connection communication between these components. Among them, the communication interface 802 may include a display screen, and the external communication interface may include a standard wired interface and a wireless interface. The processor 801 is configured to execute an image processing program in the memory to implement the steps of the map construction method provided in the above embodiments.
[0245] The descriptions of the above embodiments of the map construction device, electronic device, and storage medium are similar to the descriptions of the above method embodiments, and have similar technical descriptions and beneficial effects to the corresponding method embodiments. Due to space limitations, reference may be made to the records of the above method embodiments, so they will not be repeated here. For the technical details not disclosed in the embodiments of the map construction device, electronic device, and storage medium of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0246] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. It should be understood that in various embodiments of the present application, the order numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments. It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0247] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0248] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0249] In addition, each functional unit in the embodiments of this application can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units. Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical discs and other various media that can store program codes.
[0250] Alternatively, if the above integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as removable storage devices, ROM, magnetic disks, or optical discs that can store program codes. As described above, only the specific implementation manners of the present application are provided, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for map construction, characterized in that, The method includes: Obtaining first image data of the area to be mapped and navigation data corresponding to the first image data; Determining a first feature map based on the first image data and the navigation data; Updating an image relationship library based on a matching relationship between second image data of the area to be mapped and the first image data to obtain an updated image relationship library; wherein, the first image data and the second image data are acquired in different ways; Determining an image to be registered in the updated image relationship library; wherein, there are target feature points in the image to be registered corresponding to three-dimensional points in the first feature map; Determining the image pose of the image to be registered based on the target feature points; registering the image pose into the first feature map to obtain a registered feature map; adjusting the registered feature map based on other feature points in the image to be registered and the second image data to obtain a second feature map; wherein, the other feature points are feature points other than the target feature points in the image to be registered; Constructing a three-dimensional map of the area to be mapped based on the second image data, the second feature map, and the first image data.
2. The method according to claim 1, characterized in that, The determining the first feature map based on the first image data and the navigation data includes: Performing feature extraction on the first image data to obtain image feature points and description information of the image feature points; Matching different images in the first image data based on the description information of the image feature points to obtain an image relationship library representing the matching relationship between the different images; Determining the first feature map based on the image relationship library and the navigation data.
3. The method according to claim 2, characterized in that, The determining the first feature map based on the image relationship library and the navigation data includes: Determining a first pose of the acquisition device of the first image data based on the navigation data and preset first acquisition parameters; wherein, the preset first acquisition parameters include: external parameters for acquiring the first image data and external parameters for acquiring the navigation data; Determining the first feature map based on the image feature points in the relationship database, the first pose, and second acquisition parameters; wherein, the second acquisition parameters include internal parameters for acquiring the first image data.
4. The method according to claim 3, wherein The determining the first feature map based on the image feature points in the relationship database, the first pose, and second acquisition parameters includes: Triangulating a first true position of the image feature points in the corresponding image based on the first pose and the second acquisition parameters to obtain three-dimensional coordinates of the image feature points in the world coordinate system; Constructing the first feature map based on the three-dimensional coordinates.
5. The method according to claim 4, characterized in that The constructing the first feature map based on the three-dimensional coordinates includes: Constructing an initial feature map representing the spatial positions of the image feature points based on the three-dimensional coordinates; Determining a first predicted position where the spatial position is projected onto the first image data based on conversion parameters from the navigation data to the camera coordinate system and the second acquisition parameters; Determining a first difference between the first true position and the first predicted position; Adjust the spatial positions of the feature points in the initial feature map based on the first difference to obtain the first feature map.
6. The method according to claim 1, wherein The step of registering the image pose into the first feature map to obtain the registered feature map further includes: Determine the number of target feature points included in each image to be registered; Based on the number, determine the registration order of each image to be registered; Based on the registration order, register the image poses of each image to be registered into the first feature map to obtain the registered feature map.
7. The method according to claim 1, characterized in that, The step of adjusting the registered feature map based on the other feature points in the image to be registered and the second image data to obtain the second feature map includes: Sample the other feature points to obtain sampled feature points; Triangulate the sampled feature points to determine the three-dimensional coordinates of the sampled feature points in the world coordinate system; Adjust the registered feature map based on the three-dimensional coordinates of the sampled feature points and the second image data to obtain the second feature map.
8. The method according to claim 7, wherein The step of adjusting the registered feature map based on the three-dimensional coordinates of the sampled feature points and the second image data to obtain the second feature map includes: Based on the conversion parameters from the navigation data to the camera coordinate system, determine the second predicted position of the three-dimensional coordinates of the sampled feature points projected onto the first image data and the third predicted position of the three-dimensional coordinates of the sampled feature points projected onto the second image data; Determine the second difference between the second predicted position and the second true position of the sampled feature points in the corresponding image, and the third difference between the third predicted position and the second true position; Adjust the spatial positions of the feature points in the registered feature map based on the second difference and the third difference to obtain the second feature map.
9. The method according to any one of claims 1 to 8, characterized in that, The step of constructing a three-dimensional map of the area to be drawn based on the second image data, the second feature map, and the first image data includes: Perform depth processing on the second image data, the second feature map, and the first image data to generate point cloud data representing the first image data and the second image data; Based on the point cloud data, construct an initial mesh model representing the connection relationships between the point cloud data; Based on the initial mesh model, determine the three-dimensional map.
10. The method according to claim 9, wherein The step of performing depth processing on the second image data, the second feature map, and the first image data to generate point cloud data representing the first image data and the second image data includes: Perform depth estimation on the second image data, the second feature map, and the first image data to obtain depth maps; Fuse the depth maps to obtain the point cloud data.
11. The method according to claim 9, wherein The step of constructing an initial mesh model representing the connection relationships between the point cloud data based on the point cloud data includes: Determine the visibility of each point in the point cloud data in the second feature map and the reprojection error of each point; In the point cloud data, determine the target points whose visibility and reprojection error meet the preset conditions; Using the target point as a vertex and connecting the vertices to obtain the initial mesh model.
12. The method according to claim 9, wherein Determining the three-dimensional map based on the initial mesh model includes: In the initial mesh model, determining the acquisition sources of the point cloud data corresponding to multiple vertices of each face; Based on the acquisition sources, determining the weights representing the penetration of each face by the line of sight; Based on the faces with weights less than a preset threshold, determining the three-dimensional map.
13. The method according to any one of claims 1 to 8, characterized in that, The acquisition device of the first image data includes an in-vehicle camera, and / or the acquisition device of the second image data is an aerial photography device.
14. A map construction device, characterized in that, The device includes: A first acquisition module, configured to acquire first image data of a to-be-drawn area and navigation data corresponding to the first image data; A first determination module, configured to determine a first feature map based on the first image data and the navigation data; A first update module, configured to update an image relationship library based on a matching relationship between second image data of the to-be-drawn area and the first image data to obtain an updated image relationship library; wherein, the acquisition methods of the first image data and the second image data are different; in the updated image relationship library, determining a to-be-registered image; wherein, there are target feature points corresponding to three-dimensional points in the first feature map in the to-be-registered image; based on the target feature points, determining the image pose of the to-be-registered image; registering the image pose into the first feature map to obtain a registered feature map; based on other feature points in the to-be-registered image and the second image data, adjusting the registered feature map to obtain a second feature map; wherein, the other feature points are feature points other than the target feature points in the to-be-registered image; A first construction module, configured to construct a three-dimensional map of the to-be-drawn area based on the second image data, the second feature map, and the first image data.
15. A computer storage medium, characterized in that, Computer-executable instructions are stored on the computer storage medium, and after being executed, can implement the method steps of any one of claims 1 to 13.
16. An electronic device, characterized in that, The electronic device includes a memory and a processor. When the computer-executable instructions stored on the memory are run by the processor, the method steps of any one of claims 1 to 13 can be implemented.
Citation Information
Patent Citations
Map construction method and device, computer readable storage medium and electronic equipment
CN112270709A
Method and device for realizing virtual-real fusion, electronic equipment and storage medium
CN113409473A