A three-dimensional map construction method and related device

By performing pose estimation and matching information optimization on 3D maps, high-quality images are selected, solving the problem of poor quality caused by data errors in traditional 3D mapping and improving the accuracy and reliability of 3D maps.

CN116051767BActive Publication Date: 2026-08-04SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
Filing Date
2023-01-03
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In traditional 3D mapping methods, the presence of numerous errors or inaccuracies in the mapping data leads to poor pose optimization, thus affecting the quality of the final 3D map.

Method used

By optimizing the pose estimation and pose matching information of the first 3D map, the pose and evaluation information of the optimized image are obtained, the target image is selected and its pose is determined, abnormal images are removed, and the image quality is improved.

Benefits of technology

It improves the quality of 3D maps, reduces the impact of abnormal images on subsequent mapping operations, and provides a better data foundation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116051767B_ABST
    Figure CN116051767B_ABST
Patent Text Reader

Abstract

This application discloses a method for constructing a 3D map. In this method, the pose of the first 3D map is optimized based on pose estimation information and / or pose matching information corresponding to the first 3D map. This optimizes the pose of each of the multiple first images corresponding to the optimized first 3D map, and provides first evaluation information for the optimized first 3D map. The first evaluation information is used to evaluate the optimization quality of the first optimized pose of each first image. Based on the first optimized pose and the first evaluation information, a target image is determined from the multiple first images, and a target 3D map is obtained based on the target pose corresponding to the target image. Thus, based on the first optimized pose and the first evaluation information, a target image with a relatively accurate pose can be determined from the multiple first images, thereby reducing the impact of abnormal images on subsequent mapping operations and improving the quality of the final obtained target 3D map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of 3D mapping technology, specifically to a 3D map construction method and related equipment. Background Technology

[0002] With the continuous development of computer vision technology, 3D mapping methods are now widely used in fields such as robotics, drones, virtual reality, and augmented reality.

[0003] Traditional 3D mapping methods typically involve acquiring mapping data from sensors such as cameras, inertial measurement units, odometers, and lidar, and then building a 3D map based on all this data. However, if the mapping data contains numerous errors or large inconsistencies, the pose optimization will be poor, resulting in a low-quality 3D map. Summary of the Invention

[0004] This application provides a method for constructing a 3D map, which can obtain a high-quality 3D map. This application also provides corresponding apparatus, devices, computer-readable storage media, and computer program products.

[0005] A first aspect of this application provides a method for constructing a three-dimensional map, the method comprising: optimizing the pose of a first three-dimensional map based on pose estimation information and / or pose matching information corresponding to the first three-dimensional map, to obtain a first optimized pose for each of a plurality of first images corresponding to the optimized first three-dimensional map, and first evaluation information corresponding to the optimized first three-dimensional map, the first evaluation information being used to evaluate the optimization quality of the first optimized pose of each first image; determining a target image from the plurality of first images based on the first optimized pose and the first evaluation information, and determining a target pose corresponding to each target image; and obtaining a target three-dimensional map based on the target pose corresponding to the target image.

[0006] In the first aspect, the pose of the first 3D map can be optimized based on pose estimation information and / or pose matching information to obtain an optimized first 3D map, and first evaluation information corresponding to the optimized first 3D map can be obtained. Optimizing the first 3D map based on pose estimation information and / or pose matching information ensures the accuracy of the first optimized pose, providing a better data foundation for obtaining a high-quality target 3D map. At this time, the first evaluation information can reflect the degree of matching between the first optimized pose corresponding to each first image and the optimized first 3D map, or in other words, it can reflect the possibility of anomalies in the first optimized pose corresponding to a given first image. Thus, based on the first optimized pose and the first evaluation information, a target image with a relatively accurate pose can be determined from multiple first images for obtaining the target 3D map, thereby reducing the impact of abnormal images on subsequent mapping operations and improving the quality of the final target 3D map.

[0007] In one possible implementation of the first aspect, the first evaluation information includes the weights corresponding to each first image in the optimized first three-dimensional map.

[0008] In this possible implementation, the weight corresponding to each first image can reflect the degree to which the first optimized pose corresponding to the first image matches the optimized first 3D map. Furthermore, the weight corresponding to each first image can reflect its potential for use in subsequent mapping operations (e.g., constructing a target 3D map).

[0009] In one possible implementation of the first aspect, the first evaluation information includes feature point matching information corresponding to the optimized first three-dimensional map, and the feature point matching information includes matching information corresponding to at least one feature point in a plurality of first images.

[0010] In this possible implementation, the matching information corresponding to at least one feature point in multiple first images may include matching information between feature points and / or matching information between feature points and spatial points in the optimized first 3D map.

[0011] In one possible implementation of the first aspect, the feature point matching information includes a first value and / or a second value; the first value is the number or proportion of feature points among multiple feature points contained in multiple first images that have matching spatial points among multiple spatial points in the optimized first 3D map; the second value is the number of matching point pairs contained in at least one set of first image pairs, wherein the two images in each set of first image pairs are contained in multiple first images, and the number of matching point pairs in each set of first image pairs is the number or proportion of matching feature point pairs between the two images of the corresponding first image pair.

[0012] In this possible implementation, the first value can be determined based on the reprojection error between the optimized first 3D map and each first image. The first value can be considered as the number or proportion of feature points in the optimized first 3D map that satisfy 3D-2D matching. The second value can reflect the 2D-2D matching result corresponding to the optimized first 3D map.

[0013] In one possible implementation of the first aspect, determining a target image from multiple first images based on a first optimized pose and first evaluation information, and determining a target pose corresponding to each target image, includes: if the first evaluation information does not meet a first preset condition, determining multiple second images from multiple first images based on the weights corresponding to each first image in the optimized first 3D map, and determining the weights corresponding to each second image; performing pose optimization based on the weights corresponding to each second image and the first optimized pose corresponding to each second image to obtain an updated first optimized pose corresponding to each second image and second evaluation information about each updated first optimized pose, until the second evaluation information meets a second preset condition; and determining a target image from multiple second images based on the second evaluation information that meets the second preset condition and the updated first optimized pose corresponding to the second evaluation information that meets the second preset condition, and determining a target pose corresponding to each target image, wherein the second evaluation information includes the updated weights corresponding to each second image.

[0014] In this possible implementation, when the first evaluation information does not meet the first preset condition, it can be considered that there is a high probability of images with abnormal poses among the multiple first images, which may affect the quality of the optimized first 3D map. Therefore, pose optimization needs to be continued. In this possible implementation, the method can identify images with possible abnormal poses among the multiple first images through one or more pose optimizations based on the weights corresponding to the poses (e.g., further reducing low weights during pose optimization). This allows for the determination of the target image with a more accurate pose and the acquisition of a more accurate target pose, thus providing a better data foundation for subsequent mapping operations and reducing the impact of images with abnormal poses on subsequent mapping operations.

[0015] In one possible implementation of the first aspect, determining a target image from a plurality of first images based on a first optimized pose and first evaluation information, and determining a target pose corresponding to each target image, includes: determining a target image from a plurality of first images when the first evaluation information satisfies a first preset condition, and determining the first optimized pose corresponding to each target image as the target pose of the corresponding target image.

[0016] In this possible implementation, when the first evaluation information meets the first preset condition, it can be considered that most or even all of the multiple first images match the optimized first 3D map well, and the probability of images with abnormal poses appearing in the multiple first images is low. At this time, it can be considered that the first optimized poses corresponding to each first image have reached the expected optimization state. Therefore, the target image can be determined from the multiple first images, and the first optimized pose corresponding to each target image can be determined as the more accurate target pose of the corresponding target image.

[0017] In one possible implementation of the first aspect, the method further includes: outputting indication information based on first evaluation information, the indication information being used to indicate the acquisition method of the mapping data; acquiring target mapping data obtained based on the indication information; and obtaining a target 3D map based on the target pose corresponding to the target image, including: obtaining the target 3D map based on the target mapping data and the target pose corresponding to the target image.

[0018] In this possible implementation, the first evaluation information generated during the 3D map construction process can be used to identify potential problems in the corresponding mapping data, thereby guiding the corresponding data acquisition process, promptly correcting anomalies in the current mapping data, ensuring the quality of the mapping data, providing a better data foundation for obtaining high-quality 3D maps, and avoiding the final mapping failure due to problems in the acquisition process.

[0019] In one possible implementation of the first aspect, the pose estimation information includes global pose information and / or relative pose information; the global pose information includes pose estimation information of multiple images corresponding to the first three-dimensional map in the coordinate system corresponding to the first three-dimensional map; the relative pose information includes the relative pose information between two images in at least one set of second image pairs, and the images in at least one set of second image pairs are included in the multiple images corresponding to the first three-dimensional map.

[0020] In one possible implementation of the first aspect, the pose matching information includes matching information between the first three-dimensional map and at least one image among a plurality of images corresponding to the first three-dimensional map, and / or, in at least one set of second image pairs, each set of second image pairs contains relative pose matching information between two images, and the images in at least one set of second image pairs are included in the plurality of images corresponding to the first three-dimensional map.

[0021] A second aspect of this application provides a three-dimensional map construction apparatus, which has the function of implementing the method described in the first aspect or any possible implementation of the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described function, such as an optimization module, a determination module, a processing module, an output module, and an acquisition module.

[0022] A third aspect of this application provides an electronic device including at least one processor, a memory, and computer-executable instructions stored in the memory and executable on the processor. When the computer-executable instructions are executed by the processor, the processor performs a method as described in the first aspect or any possible implementation thereof.

[0023] The fourth aspect of this application provides a computer-readable storage medium storing one or more computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the processor performs a method as described in the first aspect or any possible implementation thereof.

[0024] The fifth aspect of this application provides a computer program product storing one or more computer execution instructions, the computer program product including computer execution instructions, which, when executed by a processor, enable the processor to perform a method as described in the first aspect or any possible implementation thereof.

[0025] A sixth aspect of this application provides a chip system including a processor for supporting electronic devices in implementing the functions described in the first aspect or any possible implementation thereof. In one possible design, the chip system may further include a memory for storing necessary program instructions and data. This chip system may be composed of chips or may include chips and other discrete devices.

[0026] The technical effects of the second to sixth aspects or any of their possible implementations can be found in the first aspect or the technical effects of its related possible implementations, and will not be repeated here. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of an embodiment of the three-dimensional map construction method provided in this application;

[0028] Figure 2a This is an exemplary schematic diagram of obtaining a first three-dimensional map provided in an embodiment of this application;

[0029] Figure 2b This is another exemplary schematic diagram of obtaining a first three-dimensional map provided in the embodiments of this application;

[0030] Figure 2c This is yet another exemplary schematic diagram of obtaining a first three-dimensional map provided in the embodiments of this application;

[0031] Figure 2d This is yet another exemplary schematic diagram illustrating the acquisition of a first three-dimensional map provided in the embodiments of this application;

[0032] Figure 3 This is a schematic diagram of another embodiment of the three-dimensional map construction method provided in this application;

[0033] Figure 4 This is a schematic diagram of an embodiment of the three-dimensional map building apparatus provided in this application;

[0034] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0035] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0036] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of singular or plural items. The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion, such that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not expressly listed or inherent to those processes, methods, products, or apparatus.

[0037] With the continuous development of computer vision technology, 3D mapping methods are now widely used in fields such as robotics, drones, virtual reality, and augmented reality.

[0038] Traditional 3D mapping methods typically involve acquiring mapping data from sensors such as cameras, inertial measurement units, odometry, and lidar, then building a map based on all this data and optimizing all poses globally to obtain the final 3D map. However, if the mapping data contains numerous errors or large inconsistencies, the pose optimization will generally be poor, resulting in a low-quality 3D map.

[0039] To address the aforementioned issues, this application provides a method for constructing a three-dimensional map, which can yield a high-quality three-dimensional map.

[0040] The 3D map construction method of this application embodiment can be applied to various scenarios. For example, it can be applied to augmented reality (AR) and / or virtual reality (VR) related application scenarios, such as AR navigation, AR games, and other application scenarios based on AR and / or VR interaction. The method can also be applied to robots, for example, to implement mapping tasks for robots such as inspection robots, drones, and automated guided vehicles (AGVs). Furthermore, the method can also be used for mapping tasks in autonomous driving scenarios to achieve high-precision positioning and tasks such as AR-head-up display (AR-HUD).

[0041] The type of electronic device that performs the three-dimensional map construction method of the embodiments of this application is not limited herein.

[0042] For example, the electronic device can be a single server, a server cluster, an edge device, a terminal device, etc., or it can be a virtual machine (VM) or a container.

[0043] For example, if the electronic device is a terminal device, the type of the terminal device can be a robot, mobile phone, tablet, computer with wireless transceiver function, virtual reality (VR) terminal, augmented reality (AR) terminal, terminal in industrial control, terminal in self-driving, terminal in remote medical care, terminal in smart grid, terminal in transportation safety, terminal in smart city, terminal in smart home, terminal in Internet of Things (IoT), etc.

[0044] like Figure 1 As shown, based on this electronic device, a three-dimensional map construction method according to an embodiment of this application may include steps 101-104.

[0045] Step 101: Obtain the first 3D map.

[0046] In this embodiment of the application, the first three-dimensional map can be generated in various ways.

[0047] For example, mapping data can be acquired to construct a first 3D map based on the mapping data.

[0048] The mapping data may include multiple images captured by a camera. These multiple images may be two-dimensional images or depth images.

[0049] In addition, in some examples, the mapping data may include, but is not limited to, data collected by one or more of the following sensors:

[0050] Inertial measurement unit (IMU), global navigation satellite system (GNSS), wheel odometry, lidar.

[0051] After obtaining the mapping data, a first 3D map can be constructed based on the mapping data.

[0052] In this embodiment, the specific method of constructing the first three-dimensional map is not limited. For example, the first three-dimensional map can be constructed based on existing and future three-dimensional mapping technologies such as simultaneous localization and mapping (SLAM) technology, according to the mapping data.

[0053] For example, in one instance, feature points and descriptors can be extracted from each of the multiple images in the mapping data, and feature matching between the images can be performed based on the feature points and descriptors of each image. Then, a first 3D map can be constructed based on the feature matching results between the images, and the pose of each image in the coordinate system corresponding to the first 3D map can be estimated. The constructed first 3D map can be a point cloud map, which can specifically include the position information and normal vector information of multiple spatial points in the corresponding scene.

[0054] In this embodiment of the application, the entity performing the first three-dimensional map generation operation can be of various types, and the method of obtaining the first three-dimensional map is not limited here.

[0055] In one example, the first three-dimensional map may be generated by an electronic device that performs the three-dimensional map construction method of the embodiments of this application, and the electronic device may be a terminal device.

[0056] For example, such as Figure 2a As shown, the terminal device can collect mapping data about a specified scene through its built-in camera and one or more sensing devices such as IMU, GNSS, wheel odometer, and lidar. It can also construct a first three-dimensional map based on the collected mapping data using algorithms such as SLAM, and store the information of the first three-dimensional map after it is generated.

[0057] Or, such as Figure 2b As shown, the terminal device can also communicate with one or more other devices to receive mapping data collected by cameras and one or more sensing devices such as IMU, GNSS, wheel odometer, and lidar on the one or more other devices. Based on the collected mapping data, it constructs a first three-dimensional map using algorithms such as SLAM, and stores the information of the first three-dimensional map after it is generated.

[0058] In another example, the electronic device that performs the three-dimensional map construction method of this application embodiment can be a server or server cluster or edge device located in the cloud.

[0059] Taking electronic devices as cloud servers as an example, such as Figure 2cAs shown, the terminal device can collect mapping data through cameras and one or more sensing devices such as IMU, GNSS, wheeled odometers, and LiDAR, and upload the collected mapping data to a cloud server. The cloud server can then construct a first 3D map based on the received mapping data using algorithms such as SLAM.

[0060] In another example, an electronic device executing the three-dimensional map construction method of this application embodiment can obtain the sub-maps corresponding to each of the multiple terminal devices, and then obtain a first three-dimensional map based on each sub-map.

[0061] For example, in such Figure 2d In the example shown, terminal device 1 collects mapping data 1 using its built-in camera and one or more sensing devices such as IMU, GNSS, wheeled odometer, and lidar, and constructs a sub-location based on the mapping data 1. Figure 1 Then, terminal device 1 can transmit the sub-ground Figure 1 The data is uploaded to a cloud server. Terminal device 2 collects mapping data 2 using its built-in camera and one or more sensing devices such as IMU, GNSS, wheel odometer, and LiDAR, and constructs a sub-map 2 based on the mapping data 2; then, terminal device 2 can upload the sub-map 2 to the cloud server.

[0062] The cloud server receives the sub-location Figure 1 After sub-map 2, the sub-map can be... Figure 1 Align with sub-map 2 to obtain a first 3D map, and perform pose estimation on the image corresponding to the first 3D map.

[0063] Of course, the above examples are merely illustrative descriptions of obtaining a first 3D map, and are not intended to limit the scope of this application. In other scenarios, electronic devices may also obtain a first 3D map based on other information interaction methods and data processing methods.

[0064] Step 102: Based on the pose estimation information and / or pose matching information corresponding to the first three-dimensional map, perform pose optimization on the first three-dimensional map to obtain the first optimized pose of each of the multiple first images corresponding to the optimized first three-dimensional map, and the first evaluation information corresponding to the optimized first three-dimensional map.

[0065] The first evaluation information is used to evaluate the optimization quality of the first optimized pose for each first image.

[0066] In this embodiment of the application, pose estimation information and / or pose matching information corresponding to the first three-dimensional map can be obtained, so as to optimize the pose of the first three-dimensional map according to the pose estimation information and / or pose matching information corresponding to the first three-dimensional map.

[0067] The pose estimation information may include pose information of multiple images corresponding to the first 3D map; and the pose estimation information may include both global pose information in the coordinate system corresponding to the first 3D map and relative pose information between the images corresponding to the first 3D map.

[0068] In some embodiments, the pose estimation information includes global pose information and / or relative pose information; the global pose information includes pose estimation information of multiple images corresponding to the first three-dimensional map in the coordinate system corresponding to the first three-dimensional map; the relative pose information includes the relative pose information between two images in each of at least one set of second image pairs, and the images in the at least one set of second image pairs are included in the multiple images corresponding to the first three-dimensional map.

[0069] Global pose refers to the position and orientation of an object in a reference coordinate system. In the embodiments of this application, the reference coordinate system can be the coordinate system corresponding to a first three-dimensional map, and the coordinate system corresponding to the first three-dimensional map can be a three-dimensional coordinate system.

[0070] The relative pose information may include information on the relative pose between two images in at least one set of second image pairs, wherein the images in at least one set of second image pairs are included in multiple images corresponding to the first three-dimensional map.

[0071] In this embodiment, the multiple images corresponding to the first 3D map can be images obtained by re-estimating the pose of the original image used to generate the first 3D map when obtaining the first 3D map. These multiple images can be included in the original image used to generate the first 3D map, and the global pose corresponding to these multiple images is determined based on the first 3D map.

[0072] In this embodiment of the application, each group of second image pairs includes two images. The method for determining each group of second image pairs is not limited here.

[0073] In one example, the multiple original images used to generate the first 3D map may include an image sequence acquired by a camera, where the original images in the image sequence are arranged sequentially based on the order of their acquisition time. After obtaining the first 3D map, the multiple images corresponding to the first 3D map may also contain an image sequence. In this case, any two adjacent frames in the image sequence can be considered as a second image pair.

[0074] As can be seen in this example, in each second image pair, there is at least one image sequence formed by multiple images corresponding to the first three-dimensional map. The relative pose information of two adjacent frames includes the information of the relative pose between the two adjacent frames.

[0075] Pose matching information can indicate the pose matching status between images corresponding to the first 3D map.

[0076] In some embodiments, the pose matching information includes matching information between at least one image in a plurality of images corresponding to the first 3D map, and / or, in at least one set of second image pairs, each set of second image pairs contains relative pose matching information between two images, and the images in at least one set of second image pairs are included in the plurality of images corresponding to the first 3D map.

[0077] Since the image plane is a two-dimensional plane and the first three-dimensional map is a three-dimensional point cloud map, the matching information between the first three-dimensional map and at least one of the multiple images can also be called three-dimensional (3D) - two-dimensional (2D) matching information.

[0078] The 3D-2D matching information may include the matching results between spatial points in the first 3D map and at least one image among multiple images. For example, it may specifically include the reprojection error between the first 3D map and at least one image among multiple images. The reprojection error between the first 3D map and any image refers to the difference between the projection of at least one spatial point in the 3D point cloud of the first 3D map onto that image and the corresponding feature point on that image.

[0079] It can be seen that the reprojection error between the first 3D map and at least one of the multiple images can reflect the matching situation between the corresponding image and the first 3D map.

[0080] In some examples, the 3D-2D matching between each image and the first 3D map can be calculated based on the global pose of each image in multiple images, so as to comprehensively evaluate the 3D-2D matching of the 3D point cloud of the first 3D map.

[0081] The relative pose matching result between two images in any set of second image pairs can reflect the matching of feature points between the corresponding two images. For example, the relative pose matching result between two images in any set of second image pairs may include the number and / or proportion of matching feature points between the two images in any set of second image pairs.

[0082] Since the image plane is a two-dimensional plane, the relative pose matching result between two images in at least one second image pair can reflect the 2D-2D matching result corresponding to the first three-dimensional map.

[0083] In this embodiment of the application, there are multiple ways to generate the pose estimation information and / or pose matching information corresponding to the first three-dimensional map, and no limitation is made here.

[0084] For example, global pose estimation information can be obtained based on the pose estimation information of multiple images in the first three-dimensional map in the coordinate system corresponding to the first three-dimensional map.

[0085] Furthermore, the 3D-2D matching information between the first 3D map and at least one of the multiple images can be calculated using the Ransac-PnP pose estimation algorithm.

[0086] Furthermore, in some examples, any two adjacent frames in an image sequence captured by a camera can be considered as an image pair. In this case, based on information collected by sensors such as IMUs and wheel odometers, the relative pose information between any two adjacent frames can be obtained using methods such as visual-inertial odometry (VIO), visual odometry (VO), or inertial odometry (IO). Based on this relative pose information, the number and / or proportion of matching feature points between the two adjacent frames can be determined to obtain the relative pose matching result between them.

[0087] In this embodiment of the application, the first evaluation information is used to evaluate the optimization quality of the first optimized pose of each first image.

[0088] In this embodiment of the application, the optimization quality of the first optimized pose of each first image can reflect the degree to which the first optimized pose corresponding to the first image matches the optimized first 3D map, or in other words, it can reflect the possibility that the first optimized pose corresponding to the first image is abnormal.

[0089] In this way, the likelihood of each first image being used for subsequent mapping operations (such as constructing a target 3D map) can be determined through the first evaluation information, avoiding abnormal images from affecting subsequent mapping operations and improving the quality of the final target 3D map.

[0090] The specific content of the first evaluation information can vary. For example, the priority, weight, and / or score corresponding to the first image can be obtained based on the degree to which the first optimized pose corresponding to the first image matches the optimized first 3D map, and used as the corresponding first evaluation information.

[0091] In one embodiment, the first evaluation information includes the weights corresponding to each first image in the optimized first three-dimensional map, and / or the feature point matching information corresponding to the optimized first three-dimensional map, wherein the feature point matching information includes the matching information corresponding to at least one feature point in the plurality of first images.

[0092] In this embodiment, the weight corresponding to each first image can reflect the degree to which the first optimized pose corresponding to the first image matches the optimized first 3D map. Furthermore, the weight corresponding to each first image can reflect the likelihood of it being used for subsequent mapping operations (e.g., for constructing a target 3D map).

[0093] For example, the matching information corresponding to at least one feature point in a plurality of first images may include matching information between feature points and / or matching information between feature points and spatial points in the optimized first 3D map.

[0094] In some embodiments, the feature point matching information includes a first value and / or a second value;

[0095] The first value is the number or proportion of feature points that match spatial points in multiple spatial points of the optimized first 3D map among the multiple feature points contained in the multiple first images.

[0096] The second value is the number of matching point pairs contained in at least one set of first image pairs, where the two images in each set of first image pairs are contained in multiple first images, and the number of matching point pairs in each set of first image pairs is the number or proportion of matching feature point pairs between the two images of the corresponding first image pair.

[0097] In the embodiments of this application, the numerical value can be a quantity or a proportion.

[0098] For example, the first value can be determined based on the reprojection error between the optimized first 3D map and each first image. By calculating the reprojection error between the optimized first 3D map and each first image, the error between the projection of a spatial point in the optimized first 3D map onto the plane of the corresponding first image and the corresponding feature point of the corresponding first image can be determined. If the error is less than a specified error threshold, it can be determined that the spatial point matches the corresponding feature point. In this way, the number or proportion of feature points among the multiple feature points contained in multiple first images that have matching spatial points among the multiple spatial points in the optimized first 3D map can be determined, thereby obtaining the first value.

[0099] The multiple spatial points in the optimized first 3D map are points in 3D space, while the feature points in the first image are usually points on a 2D image plane. Therefore, the first value can be considered as the number or proportion of feature points that satisfy 3D-2D matching corresponding to the optimized first 3D map.

[0100] The second value can be the number or proportion of matching point pairs contained in at least one set of first image pairs. The two images in each set of first image pairs are contained in multiple first images. Furthermore, the number or proportion of matching point pairs in each set of first image pairs is the number or proportion of matching feature point pairs between the two images of the corresponding first image pair.

[0101] For example, the plurality of first images may include an image sequence, and the two images in any pair of first images may be two adjacent images in the image sequence of the plurality of first images. Furthermore, information collected by sensors such as IMUs and wheel odometers can be used to obtain the number or proportion of matching feature points between the two images in any pair of first images through methods such as VIO, VO, or IO.

[0102] Since the image plane is a two-dimensional plane, the second value can reflect the 2D-2D matching result corresponding to the optimized first three-dimensional map.

[0103] In this embodiment of the application, there are various ways to optimize the pose of the first three-dimensional map based on the pose estimation information and / or pose matching information corresponding to the first three-dimensional map.

[0104] For example, the pose of the first three-dimensional map that makes the parameters in the first evaluation information optimal can be calculated based on the expectation-maximization algorithm, so as to obtain the optimized first three-dimensional map and the first evaluation information corresponding to the optimized first three-dimensional map.

[0105] The following example illustrates an exemplary method for optimizing the pose of a first 3D map based on pose estimation information and / or pose matching information corresponding to the first 3D map.

[0106] In this example, the first 3D map can be optimized based on the following optimization function:

[0107]

[0108] in, x refers to the value that minimizes the subsequent expression. k w k The value of is: k represents the time corresponding to the k-th frame in the image sequence, and k-1 represents the time corresponding to the (k-1)-th frame, which is also the time of the previous frame corresponding to the k-th frame; X k For the first optimized pose, w k P represents the weights corresponding to the first optimized pose. k Let ||P| represent the global pose estimation result corresponding to the k-th frame image. k -X k || represents the global pose error; F k This represents the feature point detection result of the k-th frame image, where λ1 is a preset coefficient, and λ1π(X) k F k ) represents the global 3D-2D reprojection error; Z k-1,k The relative pose estimation result between the k-th frame and the (k-1)-th frame is represented by ||h(X). k-1 X k )-Z k-1,k || represents the relative pose error; Q k-1,k This represents the feature point matching result between the k-th frame image and the (k-1)-th frame image, where λ2 is a preset coefficient, and λ 2p (X k-1 X k Q k-1,k ) represents the matching error between the k-th frame image and the (k-1)-th frame image.

[0109] During the optimization process, the expectation-maximization algorithm can be used for optimization.

[0110] The Expectation-Maximization (EM) algorithm is an iterative algorithm for parameter estimation. Each iteration consists of two steps: expectation and maximization. After defining the optimization function, it mainly consists of the following two steps: 1. Adjusting the model based on the parameters (E step); 2. Adjusting the parameters based on the model (M step). The E step and the M step are performed alternately until the optimal solution is obtained.

[0111] In this example, step E is optimized with the following function:

[0112]

[0113] Among them, U 2 (·) refers to the regularization operation.

[0114] The M-step is optimized based on the following function:

[0115]

[0116] in, To optimize w after the t-th iteration k , Let x be the result of the (t+1)th iteration. k .

[0117] Based on the above optimization process, the pose of each frame in multiple images can be optimized. After optimization, the corresponding first optimized pose and the weight corresponding to each first optimized pose can be obtained. At the same time, the first value and the second value can also be obtained, that is, the number or proportion of feature points that satisfy 3D-2D matching and the number or proportion of feature points that satisfy 2D-2D matching can be obtained.

[0118] As can be seen, in this example, the first 3D map can be optimized based on one or more of the following: global pose, relative pose, global 3D-2D matching information, and relative matching information. This ensures the accuracy of the optimization and provides a good data foundation for obtaining a higher quality target 3D map in the future.

[0119] Step 103: Based on the first optimized pose and the first evaluation information, determine the target image from multiple first images, and determine the target pose corresponding to each target image.

[0120] In this embodiment, the first evaluation information can reflect the degree to which each first image matches the optimized first 3D map, or in other words, it can reflect the possibility that the first optimized pose corresponding to the first image is abnormal. Therefore, the first evaluation information can be used to determine the possibility of each first image being used for subsequent mapping operations (e.g., for constructing a target 3D map). Thus, a target image is determined based on multiple first images, and the target pose corresponding to each target image is determined based on the first optimized pose corresponding to each first image. At this point, after filtering and processing multiple first images based on the first evaluation information, the target poses corresponding to each target image are usually globally accurate. Therefore, subsequent mapping operations can be performed based on relatively accurate poses, avoiding the influence of images with abnormal poses on subsequent mapping operations, thereby ensuring the quality of the final obtained 3D map.

[0121] In some embodiments, step 103 above includes:

[0122] If the first evaluation information meets the first preset condition, a target image is determined from multiple first images, and the first optimized pose corresponding to each target image is determined as the target pose of the corresponding target image.

[0123] In this embodiment of the application, the first preset condition can indicate that each first image matches the optimized first three-dimensional map to a high degree.

[0124] For example, the first preset condition may include at least one of the following conditions:

[0125] In each first image, the proportion of weights greater than a preset weight threshold is greater than a preset proportion; the first value is greater than the first value threshold; and the second value is greater than the second value threshold.

[0126] Furthermore, in one example, the first 3D map has been optimized multiple times, and after each optimization, the weights corresponding to the respective images can be obtained. In this case, the first preset condition may include: the weights corresponding to each first image converge to the desired state.

[0127] In this embodiment of the application, when the first evaluation information meets the first preset condition, it can be considered that most or even all of the images in the multiple first images have a good matching degree with the optimized first three-dimensional map, and the probability of images with abnormal poses appearing in the multiple first images is low. At this time, it can be considered that the first optimized pose corresponding to each first image has reached the expected optimization state. Therefore, the target image can be determined from the multiple first images, and the first optimized pose corresponding to each target image can be determined as the more accurate target pose of the corresponding target image.

[0128] In this case, each first image can be determined as the target image, or the first image whose weight is greater than a preset weight threshold can be determined as the target image.

[0129] In some embodiments, the first evaluation information includes the weights corresponding to each first image in the optimized first three-dimensional map;

[0130] Step 103 above includes:

[0131] If the first evaluation information does not meet the first preset condition, multiple second images are determined from the multiple first images according to the weights corresponding to each first image in the optimized first three-dimensional map, and the weights corresponding to each second image are determined.

[0132] Based on the weights corresponding to each second image and the first optimized pose corresponding to each second image, pose optimization is performed to obtain the updated first optimized pose corresponding to each second image and the second evaluation information about each updated first optimized pose, until the second evaluation information meets the second preset condition. Based on the second evaluation information that meets the second preset condition and the updated first optimized pose corresponding to the second evaluation information that meets the second preset condition, a target image is determined from multiple second images, and the target pose corresponding to each target image is determined. The second evaluation information includes the updated weights corresponding to each second image.

[0133] When the first evaluation information does not meet the first preset condition, it can be considered that there is a high probability that an image with abnormal pose appears among the multiple first images, which may affect the quality of the optimized first 3D map. Therefore, pose optimization needs to be continued.

[0134] Specifically, in this embodiment of the application, multiple second images can be determined from multiple first images based on the weights corresponding to each first image in the optimized first three-dimensional map, and the weights corresponding to each second image can be determined.

[0135] For example, the first image whose weight is greater than a specified weight threshold is identified as the second image. In this case, the pose of each second image is the corresponding first optimized pose.

[0136] Alternatively, each first image can be designated as a second image. If the weight of a first image is less than a specified weight threshold, then after the first image is designated as a second image, the weight of the second image can be made less than the weight of the first image. In other words, the weight of the first image can be reduced before it is designated as a second image.

[0137] Then, pose optimization can be performed based on the weights corresponding to each second image and the first optimized pose corresponding to each second image. The method for pose optimization can refer to the method described in the example above. For instance, the pose optimization can be performed on each second image using the expectation-maximization algorithm. After obtaining the corresponding optimization results, the first optimized pose corresponding to each second image can be updated based on the optimization results, and second evaluation information about each updated first optimized pose can be obtained. Based on this second evaluation information, the weights corresponding to each second image can then be updated.

[0138] After obtaining the updated first optimized pose, the updated weights, and the corresponding second evaluation information for each second image, it can be determined whether the second evaluation information satisfies the second preset condition.

[0139] If the current second evaluation information meets the second preset condition, it can be considered that the updated first optimized pose corresponding to each second image has been optimized to the desired state. Then, the target image can be determined from multiple second images, and the target pose corresponding to each target image can be determined.

[0140] If the current second evaluation information does not meet the second preset condition, further optimization is required. At this point, the first optimized pose corresponding to each second image and the weights corresponding to each second image can be updated (for example, if the weight of a certain second image in the current second evaluation information is less than a specified weight threshold, the updated weight of that second image is further reduced). Then, the steps of optimizing the pose based on the weights of each second image and the first optimized pose corresponding to each second image, as well as subsequent steps, are re-executed until the second evaluation information obtained after a pose optimization meets the second preset condition.

[0141] The specific content of the second assessment information may be similar to that of the first assessment information, and the second preset condition may be similar to that of the first preset condition. For details, please refer to the relevant descriptions of the first assessment information and the first preset condition.

[0142] In the method of this application embodiment, based on the weights corresponding to the pose (e.g., further reducing low weights during pose optimization), through one or more pose optimizations, images with possible pose abnormalities in multiple first images can be identified, a target image with a more accurate pose can be determined, and a more accurate target pose of the target image can be obtained. This provides a better data foundation for subsequent composition operations and reduces the impact of images with abnormal poses on subsequent composition operations.

[0143] Step 104: Obtain a 3D map of the target based on the target pose corresponding to the target image.

[0144] In this embodiment of the application, the target pose corresponding to each target image can be considered as a relatively accurate pose. In this way, a map can be constructed based on the target pose corresponding to the target image to obtain a target 3D map.

[0145] The specific method for mapping based on the target pose corresponding to the target image is not limited here. In one example, a 3D map of the target can be obtained by mapping based on the target pose corresponding to the target image through methods such as triangulation.

[0146] As can be seen, in this embodiment, the first 3D map can be optimized based on pose estimation information and / or pose matching information to obtain an optimized first 3D map, and the first evaluation information corresponding to the optimized first 3D map can be obtained. At this time, the first evaluation information can reflect the degree to which the first optimized pose corresponding to each first image matches the optimized first 3D map, or in other words, it can reflect the possibility that the first optimized pose corresponding to the corresponding first image is abnormal. Thus, based on the first optimized pose and the first evaluation information, a target image with a relatively accurate pose is determined from multiple first images to obtain the target 3D map, thereby reducing the impact of abnormal images on subsequent mapping operations and improving the quality of the finally obtained target 3D map.

[0147] In addition, such as Figure 3 As shown, in some embodiments, the method further includes steps 301-302.

[0148] Step 301: Output instruction information based on the first evaluation information.

[0149] The instruction information is used to indicate the method of collecting mapping data.

[0150] Step 302: Obtain the target mapping data based on the indication information.

[0151] Step 104 above includes step 1041.

[0152] Step 1041: Obtain a 3D map of the target based on the target mapping data and the target pose corresponding to the target image.

[0153] In this embodiment of the application, based on the weights corresponding to each first image in the first evaluation information and / or the feature point matching information corresponding to the optimized first three-dimensional map, potential problems with the mapping data used to generate the first three-dimensional map can be determined.

[0154] For example, based on this first evaluation information, it can be determined that the pose weight of the image acquired in a certain area of ​​the corresponding scene is low, indicating that the pose of the image acquired in that area may be abnormal. At this time, an instruction can be generated to instruct the adjustment of the mapping data acquisition method and to re-acquire data in that area in a timely manner to provide accurate mapping data for that area.

[0155] Alternatively, based on this first evaluation information, it can be determined that the number of images collected in a certain area of ​​the corresponding scene is relatively small, and the number of matched feature points is also small. In this case, an instruction can be generated to instruct the mapping data collection method to be adjusted, and more data to be collected in that area to promptly compensate for the lack of mapping data in that area.

[0156] The specific data format of the instruction information is not limited here. For example, the instruction information may include one or more of the following: text, tables, images, or even voice. The specific method of outputting the instruction information is also not limited here. For example, the instruction information may be displayed on the screen of an electronic device, or it may be sent to a target device (such as a designated client device) through an electronic device.

[0157] After obtaining the target mapping data, pose optimization and mapping can be performed based on the target mapping data and the target pose corresponding to the target image to obtain a 3D map of the target.

[0158] In this embodiment of the application, the first evaluation information generated during the 3D map construction process can be used to identify potential problems in the corresponding mapping data, thereby guiding the corresponding data acquisition process, promptly correcting anomalies in the current mapping data, ensuring the quality of the mapping data, providing a better data foundation for obtaining a high-quality 3D map, and avoiding the final mapping failure due to problems in the acquisition process.

[0159] The above embodiments of this application have described the three-dimensional map construction method from multiple aspects. The three-dimensional map construction device of this application will be described below with reference to the accompanying drawings.

[0160] like Figure 4 As shown, this application embodiment provides a three-dimensional map building device 40, which can be applied to the electronic devices in the above embodiments.

[0161] The device 40 includes:

[0162] The optimization module 401 is used to optimize the pose of the first three-dimensional map according to the pose estimation information and / or pose matching information corresponding to the first three-dimensional map, so as to obtain the first optimized pose of each first image in the plurality of first images corresponding to the optimized first three-dimensional map, and the first evaluation information corresponding to the optimized first three-dimensional map, wherein the first evaluation information is used to evaluate the optimization quality of the first optimized pose of each first image.

[0163] The determination module 402 is used to determine a target image from multiple first images based on the first optimized pose and the first evaluation information, and to determine the target pose corresponding to each target image;

[0164] The processing module 403 is used to obtain a three-dimensional map of the target based on the target pose corresponding to the target image.

[0165] Optionally, the first evaluation information includes the weights corresponding to each first image in the optimized first 3D map.

[0166] Optionally, the first evaluation information includes feature point matching information corresponding to the optimized first 3D map, and the feature point matching information includes matching information corresponding to at least one feature point in a plurality of first images.

[0167] Optionally, the feature point matching information includes a first value and / or a second value;

[0168] The first value is the number or proportion of feature points that match spatial points in multiple spatial points of the optimized first 3D map among the multiple feature points contained in the multiple first images.

[0169] The second value is the number of matching point pairs contained in at least one set of first image pairs, where the two images in each set of first image pairs are contained in multiple first images, and the number of matching point pairs in each set of first image pairs is the number or proportion of matching feature point pairs between the two images of the corresponding first image pair.

[0170] Optionally, the determining module 402 is used for:

[0171] If the first evaluation information does not meet the first preset condition, multiple second images are determined from the multiple first images according to the weights corresponding to each first image in the optimized first three-dimensional map, and the weights corresponding to each second image are determined.

[0172] Based on the weights corresponding to each second image and the first optimized pose corresponding to each second image, pose optimization is performed to obtain the updated first optimized pose corresponding to each second image and the second evaluation information about each updated first optimized pose, until the second evaluation information meets the second preset condition. Based on the second evaluation information that meets the second preset condition and the updated first optimized pose corresponding to the second evaluation information that meets the second preset condition, a target image is determined from multiple second images, and the target pose corresponding to each target image is determined. The second evaluation information includes the updated weights corresponding to each second image.

[0173] Optionally, the determining module 402 is used for:

[0174] If the first evaluation information meets the first preset condition, a target image is determined from multiple first images, and the first optimized pose corresponding to each target image is determined as the target pose of the corresponding target image.

[0175] Optionally, the device 40 further includes an output module 404 and an acquisition module 405;

[0176] Output module 404 is used to: output indication information based on the first evaluation information, the indication information being used to indicate the method of collecting mapping data;

[0177] The acquisition module 405 is used to: acquire target mapping data based on indication information;

[0178] The processing module 403 is used to: obtain a three-dimensional map of the target based on the target mapping data and the target pose corresponding to the target image.

[0179] Optionally, the pose estimation information includes global pose information and / or relative pose information;

[0180] Global pose information includes pose estimation information of multiple images corresponding to the first 3D map in the coordinate system corresponding to the first 3D map;

[0181] The relative pose information includes information on the relative pose between two images in at least one set of second image pairs, and the images in at least one set of second image pairs are included in multiple images corresponding to the first three-dimensional map.

[0182] Optionally, the pose matching information includes matching information between the first 3D map and at least one image among a plurality of images corresponding to the first 3D map, and / or, in at least one set of second image pairs, each set of second image pairs contains relative pose matching information between two images, and the images in at least one set of second image pairs are included in the plurality of images corresponding to the first 3D map.

[0183] Figure 5 The diagram shown is a possible logical structure schematic of an electronic device 50 provided in an embodiment of this application. This electronic device 50 is used to implement the functions of the electronic device involved in any of the above embodiments. The electronic device 50 includes: a memory 501, a processor 502, a communication interface 503, and a bus 504. The memory 501, processor 502, and communication interface 503 are interconnected via the bus 504.

[0184] The memory 501 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 501 may store a program, and when the program stored in the memory 501 is executed by the processor 502, the processor 502 and the communication interface 503 are used to perform one or more steps of the above-described three-dimensional map construction method embodiment.

[0185] The processor 502 can be a central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), graphics processing unit (GPU), digital signal processing (DSP), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any combination thereof, to execute relevant programs to achieve the functions required by the optimization module, determination module, processing module, output module, and acquisition module in the 3D map building device of the above embodiments, or to execute one or more steps of the method embodiments of this application. The steps of the method disclosed in the embodiments of this application can be executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 501. The processor 502 reads the information in memory 501 and, in conjunction with its hardware, executes one or more steps of the 3D map building method embodiments described above.

[0186] The communication interface 503 uses transceiver devices, such as, but not limited to, transceivers, to enable communication between the electronic device 50 and other devices or communication networks.

[0187] Bus 504 enables the transmission of information between various components of electronic device 50 (e.g., memory 501, processor 502, and communication interface 503). Bus 504 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 5 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0188] In another embodiment of this application, a computer-readable storage medium is also provided, which stores computer-executable instructions. When the processor of the device executes the computer-executable instructions, the device performs the aforementioned... Figure 5 The steps performed by the processor in the process.

[0189] In another embodiment of this application, a computer program product is also provided, which includes computer-executable instructions stored in a computer-readable storage medium; when the processor of the device executes the computer-executable instructions, the device performs the above-described... Figure 5 The steps performed by the processor in the process.

[0190] In another embodiment of this application, a chip system is also provided, the chip system including a processor for implementing the above. Figure 5 The steps performed by the processor. In one possible design, the chip system may also include memory for storing necessary program instructions and data. The chip system can be composed of chips or may include chips and other discrete devices.

[0191] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.

[0192] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0193] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0194] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0195] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0196] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of this application, essentially, or the parts that contribute to the prior art, or parts of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0197] The above are merely specific implementation methods of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto.

Claims

1. A method for constructing a three-dimensional map, characterized in that, include: Based on the pose estimation information and / or pose matching information corresponding to the first three-dimensional map, the pose of the first three-dimensional map is optimized to obtain the first optimized pose of each of the multiple first images corresponding to the optimized first three-dimensional map, and the first evaluation information corresponding to the optimized first three-dimensional map, wherein the first evaluation information is used to evaluate the optimization quality of the first optimized pose of each first image. Based on the first optimized pose and the first evaluation information, a target image is determined from the plurality of first images, and a target pose corresponding to each target image is determined. The first evaluation information includes the weight corresponding to each first image in the optimized first three-dimensional map. The weight corresponding to each first image is used to characterize the degree to which the first optimized pose corresponding to the first image matches the optimized first three-dimensional map. A 3D map of the target is obtained based on the target pose corresponding to the target image.

2. The method according to claim 1, characterized in that, The first evaluation information includes feature point matching information corresponding to the optimized first 3D map, and the feature point matching information includes matching information corresponding to at least one feature point in the plurality of first images.

3. The method according to claim 2, characterized in that, The feature point matching information includes a first value and / or a second value; The first value is the number or proportion of feature points that match spatial points among the multiple feature points contained in the multiple first images in the multiple spatial points of the optimized first three-dimensional map; The second value is the number of matching point pairs contained in at least one set of first image pairs, where the two images in each set of first image pairs are contained in the plurality of first images, and the number of matching point pairs in each set of first image pairs is the number or proportion of matching feature point pairs between the two images of the corresponding first image pair.

4. The method according to claim 1, characterized in that, The step of determining a target image from the plurality of first images based on the first optimized pose and the first evaluation information, and determining the target pose corresponding to each target image, includes: If the first evaluation information does not meet the first preset condition, a plurality of second images are determined from the plurality of first images according to the weight corresponding to each first image in the optimized first three-dimensional map, and the weight corresponding to each second image is determined. Based on the weights corresponding to each second image and the first optimized pose corresponding to each second image, pose optimization is performed to obtain the updated first optimized pose corresponding to each second image and second evaluation information about each updated first optimized pose, until the second evaluation information satisfies a second preset condition. Based on the second evaluation information that satisfies the second preset condition and the updated first optimized pose corresponding to the second evaluation information that satisfies the second preset condition, a target image is determined from the plurality of second images, and a target pose corresponding to each target image is determined. The second evaluation information includes the updated weights corresponding to each second image.

5. The method according to claim 1, characterized in that, The step of determining a target image from the plurality of first images based on the first optimized pose and the first evaluation information, and determining the target pose corresponding to each target image, includes: If the first evaluation information meets the first preset condition, a target image is determined from the plurality of first images, and the first optimized pose corresponding to each target image is determined as the target pose of the corresponding target image.

6. The method according to any one of claims 1-5, characterized in that, Also includes: Based on the first evaluation information, an indication information is output, which is used to indicate the method of collecting mapping data; Obtain target mapping data based on the indicated information; The step of obtaining a target 3D map based on the target pose corresponding to the target image includes: A 3D map of the target is obtained based on the target mapping data and the target pose corresponding to the target image.

7. The method according to any one of claims 1-5, characterized in that, The pose estimation information includes global pose information and / or relative pose information; The global pose information includes pose estimation information of multiple images corresponding to the first three-dimensional map in the coordinate system corresponding to the first three-dimensional map; The relative pose information includes information on the relative pose between two images in at least one set of second image pairs, and the images in the at least one set of second image pairs are included in multiple images corresponding to the first three-dimensional map.

8. The method according to any one of claims 1-5, characterized in that, The pose matching information includes matching information between the first 3D map and at least one image among a plurality of images corresponding to the first 3D map, and / or, in at least one set of second image pairs, the relative pose matching information between two images contained in each set of second image pairs, wherein the images in the at least one set of second image pairs are contained in a plurality of images corresponding to the first 3D map.

9. A three-dimensional map building device, characterized in that, include: An optimization module is used to optimize the pose of the first three-dimensional map based on the pose estimation information and / or pose matching information corresponding to the first three-dimensional map, so as to obtain the first optimized pose of each of the multiple first images corresponding to the optimized first three-dimensional map, and the first evaluation information corresponding to the optimized first three-dimensional map, wherein the first evaluation information is used to evaluate the optimization quality of the first optimized pose of each first image. The determination module is used to determine a target image from the plurality of first images based on the first optimized pose and the first evaluation information, and to determine the target pose corresponding to each target image. The first evaluation information includes the weight corresponding to each first image in the optimized first three-dimensional map. The weight corresponding to each first image is used to characterize the degree to which the first optimized pose corresponding to the first image matches the optimized first three-dimensional map. The processing module is used to obtain a three-dimensional map of the target based on the target pose corresponding to the target image.

10. The apparatus according to claim 9, characterized in that, The first evaluation information includes feature point matching information corresponding to the optimized first 3D map, and the feature point matching information includes matching information corresponding to at least one feature point in the plurality of first images.

11. The apparatus according to claim 10, characterized in that, The feature point matching information includes a first value and / or a second value; The first value is the number or proportion of feature points that match spatial points among the multiple feature points contained in the multiple first images in the multiple spatial points of the optimized first three-dimensional map; The second value is the number of matching point pairs contained in at least one set of first image pairs, where the two images in each set of first image pairs are contained in the plurality of first images, and the number of matching point pairs in each set of first image pairs is the number or proportion of matching feature point pairs between the two images of the corresponding first image pair.

12. The apparatus according to claim 9, characterized in that, The determining module is used for: If the first evaluation information does not meet the first preset condition, a plurality of second images are determined from the plurality of first images according to the weight corresponding to each first image in the optimized first three-dimensional map, and the weight corresponding to each second image is determined. Based on the weights corresponding to each second image and the first optimized pose corresponding to each second image, pose optimization is performed to obtain the updated first optimized pose corresponding to each second image and second evaluation information about each updated first optimized pose, until the second evaluation information satisfies a second preset condition. Based on the second evaluation information that satisfies the second preset condition and the updated first optimized pose corresponding to the second evaluation information that satisfies the second preset condition, a target image is determined from the plurality of second images, and a target pose corresponding to each target image is determined. The second evaluation information includes the updated weights corresponding to each second image.

13. The apparatus according to claim 9, characterized in that, The determining module is used for: If the first evaluation information meets the first preset condition, a target image is determined from the plurality of first images, and the first optimized pose corresponding to each target image is determined as the target pose of the corresponding target image.

14. The apparatus according to any one of claims 9-13, characterized in that, The device also includes an output module and an acquisition module; The output module is used to: output indication information based on the first evaluation information, wherein the indication information is used to indicate the method of collecting mapping data; The acquisition module is used to: acquire target mapping data based on the indication information; The processing module is used to: obtain a target 3D map based on the target mapping data and the target pose corresponding to the target image.

15. The apparatus according to any one of claims 9-13, characterized in that, The pose estimation information includes global pose information and / or relative pose information; The global pose information includes pose estimation information of multiple images corresponding to the first three-dimensional map in the coordinate system corresponding to the first three-dimensional map; The relative pose information includes information on the relative pose between two images in at least one set of second image pairs, and the images in the at least one set of second image pairs are included in multiple images corresponding to the first three-dimensional map.

16. The apparatus according to any one of claims 9-13, characterized in that, The pose matching information includes matching information between the first 3D map and at least one image among a plurality of images corresponding to the first 3D map, and / or, in at least one set of second image pairs, the relative pose matching information between two images contained in each set of second image pairs, wherein the images in the at least one set of second image pairs are contained in a plurality of images corresponding to the first 3D map.

17. An electronic device, characterized in that, The electronic device includes at least one processor, a memory, and instructions stored in the memory and executable by the at least one processor, wherein the at least one processor executes the instructions to implement the steps of the method according to any one of claims 1-8.

18. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-8.