Fusion positioning method based on multiple types of maps and electronic device

By combining offline panoramic maps and volumetric model maps into a fusion positioning method, the problem of low image positioning accuracy of terminal devices is solved, achieving high-precision and stable 6-DOF positioning, adapting to various environments and data states, and improving map quality.

CN116363196BActive Publication Date: 2026-05-15HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2021-12-22
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing positioning technologies struggle to acquire the 6-DOF position and attitude information of images captured by terminal devices and are easily affected by environmental interference, resulting in low positioning accuracy.

Method used

A multi-map fusion positioning method is adopted, which combines offline panoramic maps and volumetric model maps. By fusing multiple positioning methods, the 6-DOF pose of the image is obtained, and accurate positioning is achieved in different environments using different types of map data.

Benefits of technology

It achieves high-precision and stable positioning in different environments, can gradually improve map quality, adapt to various data states, and provide a smooth positioning experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116363196B_ABST
    Figure CN116363196B_ABST
Patent Text Reader

Abstract

The application provides a multi-type map-based fusion positioning method and an electronic device. The method comprises: acquiring a target image to be positioned; acquiring a target position at which the target image is captured; determining a target map quality of a map located at the target position in a map database composed of multiple different types of map data; determining a target positioning mode from a plurality of different positioning modes preset according to the target map quality, the target positioning mode being composed of basic positioning modes corresponding to each map data included in the map database; and positioning the target image by using the target positioning mode to obtain 6DOF of the target image. Thus, when the image is positioned, a suitable positioning scheme can be selected from multiple positioning schemes according to different map qualities corresponding to positions of the image, a positioning effect with higher precision and stability is realized, map quality is upgraded without perception in time and space, and user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to a fusion positioning method and electronic device based on multiple types of maps. Background Technology

[0002] Indoor and outdoor positioning technology has always been a fundamental service in terminal business. Accurate positioning can achieve excellent results in fields such as maps, navigation, and virtual-real integration.

[0003] Traditional positioning services primarily rely on satellite signals such as the Global Positioning System (GPS) and BeiDou, as well as signals from communication base stations, Wi-Fi, and Bluetooth. However, the positioning accuracy obtained through these technologies is relatively low. Furthermore, these technologies are easily affected by the environment and, in most cases, only provide positional information, not attitude information. While attitude sensors in terminal devices (such as gyroscopes and magnetometers) are typically inexpensive but of lower performance, their errors are often significant in environments with magnetic field interference; for example, magnetometer errors often exceed 30 degrees. Therefore, current positioning solutions struggle to obtain the 6 degrees of freedom (6DOF) position and attitude information corresponding to the images captured by the terminal. Thus, obtaining the 6DOF information corresponding to the images captured by the terminal is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] This application provides a multi-type map fusion positioning method, a map upgrade / update method, an electronic device, a computer storage medium, and a computer program product, which can accurately acquire the 6DOF corresponding to the images collected by the electronic device.

[0005] Firstly, this application provides a fusion positioning method based on multiple map types. The method includes: acquiring a target image to be positioned; acquiring the target location of the captured target image; determining the target map quality of the map at the target location in a map database, wherein the map database consists of multiple different types of map data, each type of map data corresponding to a basic positioning method; determining a target positioning method from multiple pre-set different positioning methods based on the target map quality, wherein the target positioning method consists of the basic positioning methods corresponding to each map data contained in the map database; positioning the target image using the target positioning method to obtain the 6-DOF pose of the target image; and outputting the 6DOF of the target image. Therefore, when positioning an image, using maps composed of different types of map data allows for the selection of a suitable positioning scheme from multiple positioning options based on the different map qualities corresponding to the image location, achieving higher accuracy and more stable positioning results, realizing a seamless upgrade transition of map quality in time and space, and improving the user experience.

[0006] In one possible implementation, the map database consists of two different types of map data. One type of map data corresponds to a first basic positioning method, and the other type of map data corresponds to a second basic positioning method. The target image is located using a target positioning method, specifically including: locating the target image using the first basic positioning method to obtain a first 6DOF of the target image; extracting 2D features and descriptors of the target image using the second basic positioning method; determining 3D features in the map database that match the 2D features of the target image based on the first 6DOF, the map data in the map database, and the 2D features and descriptors of the target image; and determining the 6DOF of the target image based on the 2D features of the target image, the 3D features that match the 2D features of the target image, and historical image data. The historical image data includes historical data of images acquired before the target image was obtained, including the 6DOF, 2D features, and 3D features of the images.

[0007] In one possible implementation, the map database consists of two different types of map data. One type of map data corresponds to a first basic positioning method, and the other type of map data corresponds to a second basic positioning method. The target image is located using the target positioning method, specifically including: locating the target image using the first basic positioning method to obtain a first 6DOF of the target image; extracting features from the target image using the second basic positioning method to obtain global and local features of the target image; using the global features of the target image and the first 6DOF, searching the map data in the map database to obtain images similar to the target image; processing the local features of the target image and the images similar to the target image to obtain the number of inliers or the projection error between the target image and the images similar to the target image; when the proportion of inliers is greater than a preset proportion, or the projection error is within a preset range, the 6DOF of the target image is determined to be the first 6DOF.

[0008] In one possible implementation, the map database consists of two different types of map data. One type of map data corresponds to a first basic positioning method, and the other type of map data corresponds to a second basic positioning method. The target image is located using the target positioning method, specifically including: locating the target image using the first basic positioning method to obtain a first 6DOF of the target image; rendering the image from the first 6DOF perspective using model data from the map database to obtain a first textureless feature map corresponding to the target image; processing the target image using the second basic positioning method according to a preset processing method to obtain a second textureless map corresponding to the target image. The preset processing method includes at least one of semantic segmentation, instance segmentation, and depth estimation. When the target parameters between the first and second textureless maps are within a preset range, the 6DOF of the target image is determined to be the first 6DOF. For example, the target parameters may include a loss function between the first and second textureless maps, which may include one or more of semantic intersection-union ratio, instance intersection-union ratio, contour distance, or depth error.

[0009] In one possible implementation, the map database consists of two different types of map data. One type of map data corresponds to a first basic positioning method, and the other type of map data corresponds to a second basic positioning method. The target image is located using the target positioning method, specifically including: processing the target image using the first basic positioning method to obtain point cloud data of the target image; processing the target image using the second basic positioning method to obtain the image with the highest similarity to the target image in the map database; determining the data in the map database that matches the point cloud data; and using a first algorithm to process the image with the highest similarity to the target image and the data that matches the point cloud data to obtain the 6DOF of the target image.

[0010] In one possible implementation, the map database consists of two different types of map data. One type of map data corresponds to a first basic positioning method, and the other type of map data corresponds to a second basic positioning method. The target image is located using a target positioning method, specifically including: locating the target image using the first basic positioning method to obtain a first 6DOF of the target image; processing the target image using the second basic positioning method, combined with the position 3DOF and orientation angle from the first 6DOF, to obtain an image similar to the target image in the map database; and processing the image similar to the target image using a first algorithm, combined with the pose 3DOF from the first 6DOF, to obtain the 6DOF of the target image.

[0011] In one possible implementation, the method further includes: determining that the number of images contained in the image data reaches a preset number; locating each image in the image data to obtain pose information of at least a portion of the images in the image data; determining the correspondence between 2D features and 3D features of each image in the image data from a map database based on the pose information of at least a portion of the images in the image data; matching at least a portion of the images in the map database based on the pose information of at least a portion of the images in the image data to obtain images with a co-view relationship with each of the at least a portion of the images; for images in the image data for which pose information has not been obtained, searching the map database to obtain images similar to the images for which pose information has not been obtained, and performing 2D feature matching between the obtained images and the images for which pose information has not been obtained to obtain images with a co-view relationship with the images for which pose information has not been obtained; obtaining the pose information of the images for which pose information has not been obtained through the pose information of the images for which pose information has not been obtained; and optimizing the pose and 3D features of the images for which pose information has been obtained. In this way, when the number of images obtained reaches the preset number, the map can be updated / upgraded, so that the map quality of each area in the map will gradually improve as the collected data and / or crowdsourced data accumulate.

[0012] In one possible implementation, the method further includes: reading point cloud data of at least a portion of the map in the map database as the target point cloud, and using the 3D features corresponding to the images in the constructed image data as the source point cloud for point cloud registration; determining the coordinates of the 3D features corresponding to the remaining 2D features in the images with existing pose information in the image data, and storing the images with existing pose information and the target parameters of the images with existing pose information in the image data into the map database, wherein the target parameters include at least one of global features, pose information, 2D features, and 3D features.

[0013] Secondly, this application provides a map upgrade / update method, which further includes: determining that the number of images contained in the image data reaches a preset number; locating each image in the image data to obtain pose information of at least a portion of the images in the image data; determining the correspondence between 2D features and 3D features of each image in the image data from a map database based on the pose information of at least a portion of the images in the image data; matching at least a portion of the images in the map database based on the pose information of at least a portion of the images in the image data to obtain images with a co-view relationship with each image in the at least a portion of the images; for images in the image data for which pose information has not been obtained, searching in the map database to obtain images similar to images for which pose information has not been obtained, and performing 2D feature matching between the obtained images and images for which pose information has not been obtained to obtain images with a co-view relationship with images for which pose information has not been obtained; obtaining the pose information of images for which pose information has not been obtained through the pose information of images with a co-view relationship with images for which pose information has not been obtained; and optimizing the pose and 3D features of images for which pose information has been obtained. In this way, when the number of images obtained reaches the preset number, the map can be updated / upgraded, so that the map quality of each area in the map will gradually improve as the collected data and / or crowdsourced data accumulate.

[0014] In one possible implementation, the method further includes: reading point cloud data of at least a portion of the map in the map database as the target point cloud, and using the 3D features corresponding to the images in the constructed image data as the source point cloud for point cloud registration; determining the coordinates of the 3D features corresponding to the remaining 2D features in the images with existing pose information in the image data, and storing the images with existing pose information and the target parameters of the images with existing pose information in the image data into the map database, wherein the target parameters include at least one of global features, pose information, 2D features, and 3D features.

[0015] Thirdly, this application provides an electronic device, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is used to execute the multi-type map fusion positioning method provided in the first or second aspect.

[0016] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when run on an electronic device, causes the electronic device to execute the multi-type map-based fusion positioning method provided in the first or second aspect.

[0017] Fifthly, this application provides a computer program product, characterized in that, when the computer program product is run on an electronic device, it causes the electronic device to execute the multi-type map fusion positioning method provided in the first or second aspect.

[0018] It is understood that the beneficial effects of the third to sixth aspects mentioned above can be found in the relevant descriptions in the first or second aspects mentioned above, and will not be repeated here. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the system framework of a visual positioning method provided in an embodiment of this application;

[0020] Figure 2 This is a schematic diagram illustrating a process for updating and / or upgrading a map, as provided in an embodiment of this application.

[0021] Figure 3 This is a schematic diagram of a visual positioning method provided in an embodiment of this application;

[0022] Figure 4 This is a schematic diagram of a visual positioning method provided in an embodiment of this application;

[0023] Figure 5 This is a schematic diagram of a visual positioning method provided in an embodiment of this application;

[0024] Figure 6 This is a schematic diagram of a visual positioning method provided in an embodiment of this application;

[0025] Figure 7 This is a schematic diagram of a visual positioning method provided in an embodiment of this application;

[0026] Figure 8 This is a schematic diagram of a visual positioning method provided in an embodiment of this application;

[0027] Figure 9 This is a flowchart illustrating a multi-type map fusion positioning method provided in an embodiment of this application. Detailed Implementation

[0028] In this article, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The symbol " / " in this article indicates that the related objects are in an "or" relationship; for example, A / B means A or B.

[0029] The terms "first" and "second," etc., used in the specification and claims herein are used to distinguish different objects, not to describe a specific order of objects. For example, "first response message" and "second response message," etc., are used to distinguish different response messages, not to describe a specific order of response messages.

[0030] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0031] In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, for example, multiple processing units means two or more processing units, multiple elements means two or more elements, etc.

[0032] In some embodiments, positioning can be achieved using GPS / BeiDou satellite signals, communication base station signals, etc. This positioning method is based on the signal strength of different satellites and base stations at the current location; positioning can be achieved when the number of satellites / base stations is greater than four. Currently, most map and navigation software uses this method for positioning. Taking the most widely used GPS positioning as an example, its principle is to receive electromagnetic wave signals emitted by satellites through devices such as mobile phones, and calculate the current location of the device using the known positions of multiple satellites and signal propagation time. Because the electromagnetic wave signals emitted by satellites / base stations are affected by atmospheric ionosphere interference, the distance calculated based on time is not the actual distance; therefore, signals from at least four satellites are needed to obtain relatively accurate location information. Furthermore, since positioning via satellite or base stations fundamentally relies on calculating the propagation distance of electromagnetic wave signals, when in a city with many tall buildings, electromagnetic wave signals are easily interfered with by building surfaces. This makes the calculated propagation distance not the actual distance to the satellite, severely affecting positioning accuracy. Additionally, this positioning method alone cannot determine the device's attitude.

[0033] In some embodiments, a visual positioning method based on offline panoramic maps can also be used to obtain the 6DOF corresponding to the image acquired by the terminal. For example, the 6DOF corresponding to the image acquired by the terminal can be obtained through global visual positioning system (GVPS). For GVPS, image information of the area is first collected by devices such as satellites, aerial drones, and ground acquisition vehicles to build a map database; then, the single image acquired by the terminal is combined with the information given by the GPS and sensors when the image was acquired by the terminal to perform matching and positioning within a certain range of the map database; finally, geometric relationships are used to obtain the precise position information and attitude angle information corresponding to the image, thereby realizing the 6DOF positioning service.

[0034] For example, based on an offline panoramic map, panoramic photos can be collected at a certain density within the service area, and the camera extrinsic parameters of the panoramic photos, i.e., the absolute position and orientation of the camera, can be recorded. (For ease of subsequent processing, the panoramic map may be divided into multiple images for storage; the panoramic image mentioned below can refer to a single panoramic photo, a group of multiple photos containing 360-degree panoramic information, or even data containing abstract image features.) After taking a photo containing coarse positional information (usually GPS-provided location information), a global feature search is performed on the panoramic map within a certain range to obtain images with similar content. Through local feature point extraction and matching, the relative pose change between the current photo and the photos in the database is obtained. Then, based on the absolute position and orientation of the images stored in the database, the precise pose of the current photo is calculated. Thus, the 6DOF of the photo can be obtained.

[0035] While offline panoramic map-based visual positioning can obtain the precise pose of images acquired by the terminal, it has the following drawbacks: 1. High acquisition cost of panoramic maps: Because the stored images need precise absolute pose, specialized equipment and professional personnel are required; furthermore, acquisition in public areas (such as squares and roads) involves complex processes. 2. Large data volume and high operating costs: This solution requires storing image information or abstract feature point information within the images, resulting in a large data volume and high server operating costs. 3. Feature point-based matching is easily affected by environmental changes, such as seasonal changes.

[0036] In some embodiments, a level-of-detail (LoD) model-based visual positioning system can be used to obtain the 6DOF (6 Degrees of Rendering) corresponding to the images captured by the terminal. For example, a city-level building model can be used, and images and semantic maps from a virtual perspective can be constructed through rendering to achieve relatively accurate visual positioning and orientation, thereby realizing 6DOF positioning services. This scheme mainly uses satellite and aerial imagery to construct a city-level 3D model. However, due to the high viewing angle and low actual resolution of the images, it is impossible to construct good textures for buildings and roads (especially the side textures of buildings). However, the overall geometric accuracy of the 3D model is relatively high, and it can distinguish and identify individual buildings. Finally, based on the building-level-of-detail (LoD) model, semantic images and instance images are rendered at certain intervals from a ground perspective. By comparing these with the semantic maps and instance images extracted from the images captured by the terminal, accurate positioning is achieved.

[0037] While block-based visual positioning can obtain the precise pose of images acquired by the terminal, it cannot achieve the positioning accuracy of offline panoramic map-based visual positioning due to limitations in map precision. Furthermore, city models are limited by satellite and aerial imagery, resulting in asymmetry between the information from mobile phone images. This means that the solution cannot achieve positioning in certain scenarios, such as indoor scenes, outdoor scenes with severe obstruction from ceilings or trees.

[0038] To improve the accuracy of visual positioning, this application provides a visual positioning scheme. This scheme integrates the advantages of visual positioning methods based on offline panoramic maps and those based on volumetric models, achieving a fusion positioning solution based on multiple map types. This scheme can adapt to various data types and quality levels, achieving smooth spatial and temporal transitions under different data conditions.

[0039] The smooth spatial transition refers to the fact that it is difficult to achieve large-scale coverage of urban scenes using visual positioning methods based on offline panoramic maps, while visual positioning methods based on volumetric models cannot solve the experience problems of special areas. Therefore, when a large area of ​​the city is covered by the LoD model (i.e., when using visual positioning methods based on volumetric models), a small number of key areas and special areas are covered by offline panoramic maps (i.e., when using visual positioning methods based on offline panoramic maps). The boundary between the two coverage areas is covered by this solution, thereby achieving a smooth transition effect.

[0040] Smooth transition in time refers to the boundary of the coverage area of ​​an offline panoramic map (i.e., when using a visual positioning method based on an offline panoramic map), or the coverage area constructed from low-precision, low-quality map data, which often leads to a poor positioning experience. In this case, this solution can gradually improve the positioning experience through a large amount of crowdsourced data or data recorded by personnel using simple data collection devices. During the data accumulation process, this solution can employ different positioning methods for different data states to achieve a smooth transition in positioning performance. Crowdsourced data can be understood as a large amount of image data provided by users. For example, images uploaded by users through their terminals can be understood as user-provided images.

[0041] The data sources for this solution can include two main categories and multiple acquisition methods:

[0042] a) Mass model map data. This data can be acquired through various methods such as satellite, aerial photography, and ground-based data collection, and there are already mature technologies available.

[0043] b) Image-based 3D feature map data (such as 3D point features, 3D line features, etc.). This data can be obtained through professional laser equipment, panoramic equipment, street view images, textured city model renderings, etc., or it can be constructed based on block model maps using crowdsourced data or data collected by simple acquisition devices.

[0044] This solution has the following main advantages:

[0045] a) Using multiple data sources as map data enables rapid and low-cost coverage.

[0046] b) By accumulating data, map quality and positioning experience can be gradually improved.

[0047] c) A good experience can be found in the transition areas between different data types.

[0048] For example, Figure 1 A system framework for a visual localization method is shown. For example... Figure 1 As shown, the system framework mainly includes: map building service, map data management service and location service.

[0049] The map building service is primarily used to produce map data, and different building methods are adopted depending on the input data. When building a map, raw map data uploaded by the administrator can be used, such as satellite imagery, aerial imagery, and other raw data, as well as data collected using methods such as vehicles, carts, and backpacks.

[0050] For example, there are two main ways to construct map data: a) Obtaining map data by rendering a 3D model. This approach can use raw data such as satellite imagery and aerial imagery to construct a city-level 3D model. The model's level of detail can be, but is not limited to, LOD2. Typically, due to image resolution limitations, this model cannot include texture information, or can only include low-quality texture information. However, the model itself has good geometric accuracy, so it can include semantic, instance, and depth information. In addition, to improve the efficiency of single visual positioning, a certain density of virtual viewpoints can be selected on the ground to render semantic maps, instance maps, and depth maps, which are then stored in a database. b) Using panoramic, laser, and other equipment to collect ground data, and then recovering 3D element information through adjustment and other methods to obtain map data. This approach can use vehicles, carts, backpacks, etc., to collect panoramic images, laser, and GPS data. Then, feature matching can be performed between images, or point cloud data can be fused to construct a 3D scene. Next, by combining city control points, the coordinates of the 3D data in actual space can be obtained. Finally, for the 2D features (such as 2D feature points, 2D feature lines, etc.) in the panoramic image (or panoramic slice image), the coordinates of the corresponding 3D features (such as 3D feature points, 3D feature lines, etc.) in the constructed scene are obtained and stored in the database.

[0051] Map data management services primarily involve comprehensive evaluation and management of map data, recording the status of map data for each region, such as 3D features, rendered images / semantic maps, building models, etc. When certain conditions are met (e.g., the map quality of a region meets certain requirements), the map data for that region can be used to request the map building service to upgrade and rebuild the map data for that region. In addition, administrators can also directly upload maps built using various offline methods to the map database, or upload other collected map data to the map building service, which will then archive the completed map data to the map database. For example, image data uploaded by users via their terminals can be stored in the map data management service, such as in temporary data. When the amount of data in the temporary data reaches a certain level, this data can be used to upgrade and / or update the constructed map.

[0052] The positioning service is primarily user-oriented. It can receive user-uploaded images, GPS data, data detected by inertial measurement units (IMUs) (such as pose information), and camera intrinsic and extrinsic parameters. Additionally, it can read map data for positioning, and after positioning is complete, return the corresponding 6DOF image to the user-uploaded image. Simultaneously, it stores the image (or its abstract features) and other information in the temporary data of the map database of the map data management service.

[0053] exist Figure 1 Within the system framework shown, crowdsourced data can be used, and existing map construction schemes can be used to upgrade and / or update the constructed maps, thereby improving map accuracy. This allows for high-precision map construction results using a low-cost data collection method based on existing map data.

[0054] In addition, Figure 1 Within the system framework shown, the map quality of various regions in the constructed map can also be evaluated from multiple dimensions. For example, from image quality, image acquisition density, and other factors. Figure 3 Multiple indicators, such as D-element accuracy and the number of 3D elements, are used to evaluate map data for a certain area. Then, through quantitative evaluation of different parameters, different positioning strategies can be automatically selected to achieve the best positioning effect under different map qualities.

[0055] In addition, Figure 1 Within the system framework shown, when locating user-uploaded images, different positioning methods can be used depending on the map quality of the area corresponding to the uploaded image. This allows for positioning of the image across different map levels, resulting in a better positioning experience and more cost-effective coverage. Based on this solution, more flexible deployment is possible. For example, the LoDVPS positioning service can be initially implemented across the entire system, followed by gradual upgrades and updates to address positioning issues in specific scenarios and build high-precision 3D maps.

[0056] Next, based on the content described above, we will introduce map upgrades and / or updates, map quality assessment, and different positioning methods.

[0057] (1) Map upgrades and / or updates

[0058] When upgrading and / or updating maps, the primary approach is to rely on crowdsourced data. This involves using low-cost data collection methods combined with existing data to upgrade and update the map data. Specifically, for example... Figure 2 As shown, upgrading and / or updating a map may include the following steps:

[0059] S201. Visually locate each image in the image data using an existing positioning scheme to obtain pose information for at least some of the images.

[0060] Specifically, but not limited to, visual localization can be performed on individual images in the image data using existing localization schemes to obtain pose information for at least some images. For example, the pose information of an image may include its 6DOF (6 digits of direction of view). In one example, at this stage, due to low map data quality or the availability of only textureless model data, localization results may be obtained for only some images in the image data, while localization of the remaining images may fail.

[0061] As one possible implementation, it can also be obtained from the data acquired when locating user-uploaded images.

[0062] In some embodiments, S201 can be executed every time, or it can be executed selectively according to the actual situation; no limitation is made here. When the pose information of each image has been obtained in advance, the pose information of each image can be read directly.

[0063] S202. Construct the correspondence between 2D features and 3D features of each image in the image data.

[0064] Specifically, for images successfully located in S201, the 3D features corresponding to the 2D features of these images can be directly read from the database, thereby obtaining the initial coordinates of the 3D features corresponding to the 2D features of these images. The database can store images, global features, pose information, 2D features, 3D features, and map-related data (such as coordinate information at different locations).

[0065] For images that fail to be localized in S201, the coordinates of the 3D features corresponding to the 2D features of these images on the model can be solved by using ray intersection.

[0066] This establishes the correspondence between 2D and 3D features of each image in the image data. For example, 2D features may include 2D feature points, 2D feature lines, etc.; 3D features may include 3D feature points, 3D feature lines, etc.

[0067] S203. Match the descriptors of 2D features between images in the image data to construct 2D-2D matching between images with co-viewing relationships.

[0068] Specifically, for the image successfully located in S201, image pairs with a co-view relationship can be filtered using the pose results. For example, when filtering using pose results, two images with overlapping areas can be considered as one image pair.

[0069] For images that failed to be located in S201, image retrieval can be used to obtain image pairs, which are then verified using 2D-2D matching. In one example, since the 6DOF of an image that failed to be located cannot be determined, image retrieval can be used to obtain image pairs. For instance, for any image that failed to be located, the similarity between that image and each successfully located image can be obtained, and then that image can be paired with the successfully located image that has the highest similarity. Finally, after determining the image pairs, the descriptors of the 2D features of the two images in the pair can be used for verification to ensure that the two images have a co-view relationship. For example, verification can be performed using the similarity between the descriptors of the 2D features of the two images, such as using the cosine similarity of the descriptors of the two 2D features. When the similarity is higher than a certain value, the verification is considered successful.

[0070] S204. Optimize the pose and initial coordinates of the 3D features of the image with acquired pose information to minimize the reprojection error.

[0071] Specifically, existing camera models and camera intrinsic parameters can be used to calculate the coordinates of the 3D features of the image with pose information (i.e., the successfully localized image) projected onto the image. The projection error is then calculated as the loss in the optimization problem to optimize the pose and initial coordinates of the 3D features of the image with pose information. For example, but not limited to, the pose and initial coordinates of the 3D features of the image with pose information obtained through bundle adjustment (BA) can be processed for optimization.

[0072] S205. For images that have not obtained pose information but have a common viewing relationship, obtain the initial pose of the images that have not obtained pose information but have a common viewing relationship through 2D-3D matching.

[0073] Specifically, for images with shared view relationships but without pose information, the initial pose of these images can be obtained through 2D-3D matching. For example, for images with shared view relationships but without pose information, the initial pose of the images with shared view relationships and pose information can be obtained by combining them with images that have shared view relationships and pose information with the images without pose information, using the PnP (perspective-n-point) algorithm.

[0074] In some embodiments, the process stops when the number of registered images reaches a certain value, or when all remaining images cannot be registered. Image registration can be understood as obtaining the initial pose of an image. Failure to register an image can be understood as the inability to obtain the initial pose of the image. For example, if an image has no co-view relationship with any other image, it can be determined that the image cannot be registered.

[0075] S206. Repeat S204 and S205 until all remaining images cannot be registered. At this point, the 3D information of the scene has been constructed from the images. For example, the 3D information may include the coordinates corresponding to the 2D features of the images, the coordinates corresponding to the 3D features, etc.

[0076] S207. Read at least some of the existing model or point cloud data in the map as the target point cloud, and use the constructed 3D features as the source point cloud for point cloud registration.

[0077] Specifically, the location of each image in the image data can be indexed from the database to determine the location of the region to which each image belongs. This allows the model or point cloud data corresponding to the region identified in the database to be determined, thus obtaining the target point cloud. Additionally, the database can also identify the constructed 3D features within the region to which each image belongs, and these 3D features can be used as the source point cloud. After obtaining the target and source point clouds, a point cloud registration algorithm can be used to process them for point cloud registration. This corrects minor errors in the pose and 3D features of each image in the image data, further improving the alignment accuracy between the constructed map and the real-world map.

[0078] S208. Determine the coordinates of the 3D features corresponding to the remaining 2D features in the image with existing pose information, and store the image, global features, pose information, 2D features, and 3D features in the database.

[0079] Specifically, for any image with existing pose information, the coordinates of the 3D features corresponding to each 2D feature in the image can be determined by triangulation or ray intersection of the image's pose (i.e., the pose corrected in S207). Afterwards, the image, global features, pose information, 2D features, and 3D features can be stored in the database.

[0080] It is understandable that the coordinates of the 3D features corresponding to the 2D features in the image determined in the steps prior to S208 may only be the coordinates of the 3D features corresponding to a portion of the 2D features in the image. That is to say, there is another part of the 2D features in the image that does not have corresponding 3D feature coordinates. Therefore, the coordinates of the 3D features corresponding to the other part of the 2D features in the image can be determined through S208, and then the coordinates of the 3D features corresponding to all the 2D features in the image can be obtained.

[0081] This enables the construction of high-precision maps. As collected data and / or crowdsourced data accumulate, the map quality of each area will gradually improve.

[0082] (2) Map quality assessment

[0083] The success rate and accuracy of visual positioning are closely related to map quality. Map data collected by high-precision equipment can make positioning more reliable, but images collected through rendering, street view, or handheld devices cannot guarantee the quality of data collection and map accuracy. Therefore, the map quality of a certain area can be described from multiple perspectives to enable the rational use of different positioning strategies. Since there are already relevant standards and definitions for the quality of textureless models (LoDx Models), the map quality described in this solution can refer to the map data quality of maps with 3D features.

[0084] In this scheme, the following dimensions (pose accuracy, 3D feature accuracy, 3D feature density and distribution) can be used to describe the map quality of a certain area:

[0085] a) Pose accuracy

[0086] Pose accuracy refers to the precision of the 6DOF data corresponding to the image. Since there is no true pose value, this indicator is evaluated through the following two aspects. First, high-precision equipment is used in the early stage to collect data with true pose values, and the pose accuracy level of different data is determined by evaluating the pose of different map data. Second, during the mapping process, the number of co-view relationships, the number of matched inliers, the average reprojection error, and the feature depth can indirectly reflect the stability and accuracy of the pose.

[0087] Based on the two methods described above, the pose accuracy of map data can be divided into five levels. Data acquired using professional laser equipment has the highest accuracy after processing, at level 1. Images with only initial poses have the lowest pose accuracy, at level 5. After performing local bundle adjustment, the pose accuracy is improved to level 4. After global alignment, the pose accuracy is improved to level 3.

[0088] b) 3D Feature Accuracy

[0089] Similar to pose accuracy, 3D feature accuracy can also be graded using the two methods described above, which will not be repeated here.

[0090] c) 3D Feature Density and Distribution

[0091] The density and distribution of 3D features are mainly related to the number of feature matches during localization and / or the accuracy of the calculated pose. Scenes with rich 3D features have more usable information during localization, resulting in more robust localization. The distribution of 3D features is equally important; the more uniform the distribution, the stronger the constraint on the pose during localization, and the more accurate and robust the pose calculation results.

[0092] The density and distribution of 3D features can be defined and calculated using the following formula. Specifically, the 3D space of a city can be divided based on a voxel model, and based on the city model, voxels on the city surface can be labeled as voxels of interest, which can form a set I. Next, the number F of different types of features in each voxel of interest is counted. i Let i = 1, 2, 3, ..., s, where s can be the number of feature types. Then, the 3D feature density and distribution can be obtained using the following formula:

[0093] ε=ave(P(z))

[0094] θ = std(P(z))

[0095] Where ε is the 3D feature density, θ is the 3D distribution, and z∈I, λ is the weight of the number of features corresponding to a single interest voxel, and s is the number of feature types.

[0096] After obtaining the 3D feature density and distribution in the map, the map quality level of each region can be classified according to the 3D feature density and distribution in each region.

[0097] d) Number of images

[0098] Image quantity refers to the number of images per unit area within the road network (i.e., the area accessible to users). Once the image quantity is obtained, the classification level of a specific area can be set based on the number of images within that area.

[0099] e) Image coverage

[0100] Image coverage can be defined in the following ways:

[0101] First, the road network area within a certain range can be divided into N non-overlapping planar regions. Each small region is further divided into M non-overlapping orientation regions based on its orientation, for a total of N*M small regions. Then, based on the pose of the image, the image is marked in the corresponding small region, and the number of images in each small region is recorded to obtain the mapping P between the region and the number of images in that region. Finally, the image coverage can be obtained using the following formula:

[0102]

[0103] Where γ is the image coverage, ρ is the image coverage density, and γ s For calibration parameters, such as γ s It can be set to 65.26%.

[0104] After obtaining the image coverage and / or image coverage density in the map, the map quality level of each region can be classified according to the image coverage and / or image coverage density of each region.

[0105] f) Image quality

[0106] Image quality can refer to image resolution, sharpness, etc., which can be set according to the shooting device.

[0107] After obtaining the image quality in the map, the map quality level for each region can be classified according to the image quality of each region. In some embodiments, when determining the map quality level by image resolution, the level can be determined by the average angular resolution of the image. For example, a high-definition (1440*1080) image taken by a regular mobile phone has an angular resolution of approximately 0.05° / pixel, which can be defined as level 3; a consumer-grade panoramic camera has an angular resolution of approximately 0.07° / pixel, which can be defined as level 4; a professional panoramic acquisition camera has an angular resolution of up to 0.03° / pixel, which can be defined as level 1, and so on.

[0108] After obtaining map quality data through one or more of the above methods, it can be stored as follows: First, divide the geographical area into blocks of a certain size, such as 50m. Next, use the map data identifier for each block to indicate whether the area has map data; for areas outside the service area, this identifier is "no". Then, the map data within the service area contains the aforementioned indicators; the specific values ​​of these indicators are for reference only and are not limited. For different floors that may exist within the area, multiple sets of data need to be stored. Finally, the overall data can be stored using a two-dimensional table; considering the continuity of geographical features, a quadtree can be used to reduce the data's footprint.

[0109] (3) Positioning method

[0110] In this solution, different positioning methods can be designed for different map qualities to achieve visual positioning across the entire scene. Furthermore, using feature information (such as 2D features and 3D features) contained in the map database of this solution can also improve the positioning effect to some extent. Several positioning methods are introduced below.

[0111] a) Visual localization based on textureless models (i.e., block models) (LOD-VPS)

[0112] Texture-free models primarily utilize semantic information for localization. Specifically, such as... Figure 3 As shown, this localization process mainly involves acquiring the image uploaded by the user (i.e., the localization image shown in the figure), and then performing semantic segmentation on the image using a pre-trained neural network model to obtain point cloud data. The point cloud data is then corrected based on a pre-defined algorithm. Next, the corrected point cloud data is registered using an iterative closest point (ICP) algorithm. The registered point cloud data is then used to retrieve matching data from semantic data stored in a database, and the pose information corresponding to this matching data is used as the 6DOF of the image corresponding to the point cloud data.

[0113] In some embodiments, semantic segmentation can be extended to instance segmentation and depth estimation, and the loss corresponding to instance and depth features can be fused during retrieval and ICP to further improve the applicability and accuracy of the localization algorithm.

[0114] b) Image feature-based visual positioning system (GVPS)

[0115] Image-based visual localization primarily utilizes 3D features for localization. Specifically, for example... Figure 4 As shown, this positioning process mainly involves obtaining the image uploaded by the user (i.e., the positioning image shown in the figure), and then extracting global and local features of the image using, but not limited to, a pre-trained neural network model. Next, the global features are used to perform feature retrieval in the map data of the database to find images similar to the original image. Then, local features are used to perform feature matching from the retrieved images to obtain one or more images with the highest similarity. Finally, the PNP / BA algorithm is used to process the original image and the images with the highest similarity to obtain the 6DOF of the image.

[0116] In some embodiments, line features of the image can be incorporated during feature matching to improve the positioning accuracy and success rate in indoor weak texture scenes.

[0117] c) Local adjustment optimization positioning

[0118] Local adjustment optimization positioning methods, such as Figure 5As shown, after obtaining the image uploaded by the user that needs to be located (i.e., the location image shown in the figure), the 6DOF (pose data) of the image can be obtained first through the LOD-VPS-based method mentioned in "Solution a"); and the 2D features and their descriptors of the image can be extracted through a pre-trained neural network model. Then, the 2D features of the image, the determined 6DOF, and the data of the map model in the database are processed by ray intersection to obtain the coordinates of the coarse 3D features corresponding to the 2D features of the image. After obtaining the coordinates of the 2D features and the coarse 3D features corresponding to the 2D features of the image, the data of the area where the image is located in the database can be combined, and the coordinates of the 2D features and the coarse 3D features corresponding to the 2D features of the image can be processed by the PNP / BA algorithm to obtain the optimized 6DOF, 2D features, and 3D features. The optimized pose can be used as the output result, while the 2D and 3D features are stored in the database for subsequent positioning.

[0119] In some embodiments, this scheme can be applied to areas in the database with low image density, poor pose accuracy, and poor 3D point accuracy. The positioning effect of this method is no worse than the LOD-VPS-based method mentioned in "Scheme a") above, and the positioning effect will gradually improve as the number of images in the area gradually increases.

[0120] d) Cross-validation localization

[0121] When using cross-validation for localization, after obtaining the image uploaded by the user (i.e., the localization image shown in the figure), a primary localization method can be selected from "Solution a) based on the map quality of the area corresponding to the image, using either the LOD-VPS-based method mentioned in "Solution a") or the GVPS-based method mentioned in "Solution b" as the primary localization method. Then, the localization result obtained from the primary method is validated using another localization method. If the validation passes, the localization result is output; otherwise, a localization failure result is output. Thus, this cross-validation method greatly improves the robustness of the algorithm.

[0122] When using "the LOD-VPS-based approach mentioned in "Solution a)" as the primary solution, such as Figure 6As shown in (A), the 6DOF of the image can be obtained using the above-described "Scheme a)", and the global and local features of the image can be extracted using the method described in the above-described "Scheme b)". Then, using the obtained global features and combined with the obtained 6DOF, feature retrieval and matching are performed in the map data in the database to obtain images similar to the image. Afterwards, a pre-set algorithm can be used to process the local features of the image and the obtained similar images to obtain the number of inliers and / or projection error between the image and the obtained similar images. Finally, the obtained 6DOF of the image can be verified using the obtained number of inliers and / or projection error. Specifically, if the projection error is within a preset range and / or the proportion of inliers is greater than a preset proportion, the obtained 6DOF is determined to be accurate, and the 6DOF can be output; otherwise, the obtained 6DOF is determined to be inaccurate, and a positioning failure message can be output.

[0123] When using the GVPS-based approach mentioned in "Solution b)" as the primary solution, such as Figure 6 As shown in (B), the 6DOF of the image can be obtained using the above-mentioned "Scheme a)". Next, according to the camera parameters corresponding to the image, the semantic map, instance map, and depth map (hereinafter collectively referred to as textureless feature maps) from the database can be rendered using model data to obtain the textureless feature maps from the 6DOF perspective. Simultaneously, the textureless feature maps corresponding to the acquired image can also be obtained through semantic segmentation, instance segmentation, and / or depth estimation. Finally, the accuracy of the obtained 6DOF of the image is determined by comparing these two sets of textureless feature maps.

[0124] One possible implementation is to determine the accuracy of the 6DOF of the obtained image using a loss function between these two sets of textureless feature maps. For example, the loss function for these two sets of textureless feature maps can include one or more of the following: semantic intersection-over-union (SUI), instance intersection-over-union (IUI), contour distance, and depth error. In one example, if the SUI and / or IUI of the two sets are greater than a preset value, the obtained 6DOF is considered accurate, and the 6DOF can be output; otherwise, the obtained 6DOF is considered inaccurate, and a localization failure message can be output. Similarly, if the contour distance and / or depth error between the two sets are less than a preset value, the obtained 6DOF is considered accurate, and the 6DOF can be output; otherwise, the obtained 6DOF is considered inaccurate, and a localization failure message can be output.

[0125] e) Multi-type Loss Fusion Positioning

[0126] Multi-type loss fusion positioning mainly combines the LOD-VPS-based approach mentioned in "Solution a") and the GVPS-based approach mentioned in "Solution b") for positioning. Specifically, such as... Figure 7 As shown, after obtaining the user-uploaded image requiring localization (i.e., the localization image shown in the figure), the point cloud data corresponding to the image can be obtained through the LOD-VPS-based method mentioned in "Scheme a" above. Simultaneously, images similar to the image and the image with the highest similarity to the image can be obtained through the GVPS-based method mentioned in "Scheme b" above. Furthermore, the feature matching correspondence between the image and its similar images can also be obtained. Finally, the obtained point cloud data is registered using the ICP algorithm, and the registered point cloud data is used to retrieve semantic data stored in the database to obtain data matching the point cloud data. Then, the bundle adjustment algorithm is used to process the obtained data and the image with the highest similarity to the image to obtain the 6DOF corresponding to the image. In one example, the initial alignment pose between the image and point cloud data can be obtained through search and ICP algorithms. Since point cloud data has semantic, instance, and depth information, it can provide semantic loss, instance loss, and depth loss in the image. By obtaining the correspondence of feature points through feature matching, the reprojection error loss of 3D features (such as 3D point features, 3D line features, etc.) on the image can be calculated. Finally, the various losses are fused together through configurable weights and jointly optimized using an optimization solver to obtain the 6DOF corresponding to the image.

[0127] The front end of the algorithm can be executed in parallel: 1. Extracting textureless feature maps of the image, such as semantic maps, instance maps, depth maps, etc.; 2. Extracting global feature vectors of the image and completing database feature retrieval; 3. Extracting local feature points / lines of the image and completing matching.

[0128] The algorithm's backend can integrate the LOSS of LODVPS and GVPS schemes, including semantic IOU, contour LOSS, feature point line reprojection error, etc., for bundle adjustment.

[0129] In some embodiments, this scheme can be applied to situations with high map quality. Different fusion loss weights can be set according to the map quality to improve the positioning accuracy and stability of the algorithm.

[0130] f) Rotation and translation separation estimation positioning

[0131] Rotation-translation separation estimation and localization also combines the LOD-VPS-based method mentioned in "Scheme a") and the GVPS-based method mentioned in "Scheme b") for localization. Specifically, as... Figure 8 As shown, after obtaining the image uploaded by the user that needs to be located (i.e., the location image shown in the figure), the 6DOF corresponding to the image can be obtained through the LOD-VPS-based method mentioned in "Solution a" above. Then, the global features obtained through the GVPS-based method mentioned in "Solution b" above can be used to perform feature retrieval in the database to find images similar to the image. During the retrieval, the position 3DOF and orientation angle in the obtained 6DOF of the image can be combined to reduce the retrieval range and improve efficiency. At the same time, the retrieval results can be filtered to improve the retrieval accuracy. Finally, the image with the highest similarity to the given image obtained through local feature matching in "Scheme b) based on GVPS" is processed using the PNP / BA algorithm to obtain the 6DOF of the image. During this processing, the 3DOF of the attitude (i.e., pitch, yaw, and roll) from the previously obtained 6DOF can be combined, thus fusing the initial attitude information. This decouples translation and rotation estimation, and utilizes the strong constraint of the model contour on the angles to achieve attitude constraint. For example, during the fusion process, the angle can be calculated first using feature points, and then the translation can be calculated while keeping the angle fixed.

[0132] In some embodiments, the process of this rotation-translation separation estimation and localization method can be as follows: First, feature extraction is performed, such as extracting textureless feature maps, global features, and local features. Next, an initial pose is obtained based on the LODVPS scheme, and the pose is decomposed into three-axis position and three angles: roll, pitch, and heading. Then, global features are used, along with the position and orientation information from the initial pose, to retrieve map data from the database. Afterward, local features are used to perform feature matching on the retrieved data. Finally, the results of feature matching are used, along with the initial pose information (i.e., 3DOF pose), to decouple translation and rotation estimation, utilizing the strong constraint of the model contour on the angles; where the angles can be calculated first by combining feature points, and then the translation can be calculated while keeping the angles fixed. Finally, the calculated angles and translation information are merged into 6DOF information for output. It is understandable that adding filtering of the initial position (xyz) and orientation angles during feature retrieval (e.g., filtering by the distance and angle difference between the corresponding pose of the image already labeled in the database and the current initial result) can effectively filter out some erroneous retrieval results. Furthermore, the model map data provides strong angular constraints on the localization results. Therefore, in the final localization process, compared to the traditional approach of simultaneously estimating rotation and translation, the rotation component can be estimated first. This fully utilizes the strong constraints of existing information to improve the angular localization effect. After the angle calculation is completed, the translation component is calculated using the matching information of 2D-3D features. With known constraints, optimizing only the translation component yields more accurate and robust location results.

[0133] In some embodiments, this scheme can be applied to scenarios with high map quality. By fusing point and line features and model structure features, it can improve positioning stability and accuracy to a certain extent, and solve positioning problems in indoor repetitive texture and weak texture scenarios, as well as outdoor ultra-long-range scene scenarios.

[0134] In some embodiments, the various computational methods described above can be implemented, but are not limited to, through a pre-trained neural network model.

[0135] Next, based on the content described above, a fusion positioning method based on multiple map types provided in this application will be introduced. It is understood that this method is proposed based on the content described above, and some or all of its content can be found in the relevant descriptions above.

[0136] Please see Figure 9 , Figure 9This is a flowchart illustrating a multi-type map fusion positioning method provided in an embodiment of this application. It is understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. For ease of explanation, the following description uses server execution as an example. It is understood that the server can be replaced with other devices, and the replaced solution is still within the protection scope of this application. Figure 9 As shown, this multi-type map-based fusion positioning method may include:

[0137] S901. Obtain the target image to be located.

[0138] Specifically, when an electronic device captures a target image, it can upload the target image to a server, thus allowing the server to obtain the target image with location information.

[0139] S902, Obtain the target location of the captured target image.

[0140] Specifically, the data that electronic devices can upload to the server can also include the target location of the currently captured target image, so that the server can obtain the target location when the target image was captured.

[0141] S903. Determine the target map quality of the map located at the target location in the map database. The map in the map database consists of various types of map data, and each type of map data corresponds to a basic positioning method.

[0142] Specifically, after the server obtains the target location, it can determine the target map quality of the map located at that location in the map database. The map quality in the map database can be determined in advance, for example, through the methods described in the "Map Quality Assessment" section above. For instance, the map in the map database can consist of various types of map data, each corresponding to a basic positioning method. For example, the map in the map database can consist of volumetric model map data and image-based 3D feature map data (such as 3D point features, 3D line features, etc.). The positioning orientation corresponding to the volumetric model map data can be the "Image Feature-Based Visual Positioning System (GVPS)" method described above, while the positioning method for the image-based 3D feature map data can be the "Texture-Free Model-Based Visual Positioning (LOD-VPS)" method described above.

[0143] S904. Based on the quality of the target map, determine the target positioning method from multiple pre-set different positioning methods, wherein the target positioning method consists of the basic positioning methods corresponding to each map data contained in the map database.

[0144] Specifically, after obtaining the target map quality, a target positioning method can be determined from multiple pre-defined positioning methods based on the target map quality. This target positioning method is composed of the basic positioning methods corresponding to various map data contained in the map database. It can be understood that each of the pre-defined positioning methods is composed of the basic positioning methods corresponding to various map data contained in the map database. This composition of basic positioning methods can be understood as being formed through all or part of the steps in each of the multiple basic positioning methods.

[0145] In some embodiments, map quality evaluation can be divided into six dimensions: pose accuracy, 3D feature accuracy, 3D feature density and distribution, number of images, image coverage, and image quality. Each dimension can be divided into five levels, with levels 1 to 5 representing high, relatively high, medium, relatively low, and low, respectively (this is just an example to indicate the relative levels of the indicators).

[0146] Ideally, with professional equipment used for data collection, the map quality in all dimensions is at the highest level (Level 1), allowing for direct use of the GVPS positioning solution. When only model data is available, the LODVPS positioning solution can be used.

[0147] When a region is covered by LODVPS service and has a certain amount of user crowdsourced data, all indicators in the map data are at level 5. The "local adjustment optimization positioning" described above can be used for joint positioning. At the same time, the existing map data (pose, 3D points) can be optimized to improve accuracy.

[0148] As user data gradually increases and the number of images and pose accuracy reach level 4, the positioning scheme based on LODVPS in the "cross-validation positioning" described above can be adopted to verify the results using existing images.

[0149] As the number and coverage of images gradually increase, reaching Level 3, map construction can proceed. After construction, the map's pose accuracy and 3D feature accuracy can be improved to Level 3. At this point, Cocoa adopts the GVPS-based localization scheme from the "cross-validation localization" described above. Under these map quality conditions, GVPS achieves higher accuracy than LODVPS.

[0150] Once the map is more complete, with coverage reaching level 2, 3D feature density and distribution reaching level 3, and pose and 3D point accuracy reaching level 3, the positioning scheme described above in the "multi-class loss fusion positioning" method can be used for positioning. At this point, the map density is high, and the accuracy has reached a certain standard. After fusing the retrieved 3D features into the positioning optimization objective function, it can provide better constraints and improve positioning accuracy and stability.

[0151] When the 3D feature distribution is improved to level 2, or when data obtained through other methods has high accuracy but suffers from defects in coverage density or image quality, such as street view data with a coverage density of only level 3, or images rendered from textured models constructed from aerial imagery with low quality, reaching only level 4-5, the positioning scheme described above in the "rotation-translation separation estimation positioning" method can be used for localization. This scheme, based on the GVPS scheme, adds constraints on attitude from the model map data to improve the angular accuracy of localization; it also adds prior knowledge of image retrieval from the model data to improve the problems of weak and repetitive textures.

[0152] S905. Using the target localization method, the target image is located to obtain the 6-DOF pose of the target image.

[0153] Specifically, once the target localization method is determined, the target image can be located using this method, thereby obtaining the 6-DOF pose of the target image.

[0154] S906, Output the target image in 6DOF format.

[0155] Specifically, after obtaining the 6DPF of the target image, the server can output the 6DOF of the target image, for example, by sending the 6DOF of the target image to an electronic device for display on the electronic device.

[0156] Therefore, when locating images, using maps composed of different types of map data allows for the selection of a suitable positioning scheme from multiple options based on the different map qualities corresponding to the image's location. This achieves more accurate and stable positioning results, enables seamless upgrades in map quality over time and space, and enhances the user experience.

[0157] In some embodiments, the map database consists of two different types of map data, one type of map data corresponding to a first basic positioning method, and the other type of map data corresponding to a second basic positioning method. For example, the first basic positioning method can be the "LOD-VPS" scheme described above, and the second basic positioning method can be the "GVPS" scheme described above.

[0158] The target image localization method can be implemented as follows: First, the target image is located using a first basic localization method to obtain the first 6DOF (6 Degrees of Fiber). Then, the 2D features and descriptors of the target image are extracted using a second basic localization method. Next, based on the first 6DOF, map data in the map database, and the 2D features and descriptors of the target image, 3D features matching the 2D features of the target image in the map database are determined. Finally, based on the 2D features of the target image, the matching 3D features, and historical image data, the 6DOF of the target image is determined. The historical image data includes historical data of images acquired before the target image was obtained, including the 6DOF, 2D features, and 3D features of the images. This localization method can be understood as the "local adjustment optimization localization" scheme described above. Figure 5 The schemes shown are detailed in the above description and will not be repeated here.

[0159] In some embodiments, the map database consists of two different types of map data, one type of map data corresponding to a first basic positioning method, and the other type of map data corresponding to a second basic positioning method. For example, the first basic positioning method can be the "LOD-VPS" scheme described above, and the second basic positioning method can be the "GVPS" scheme described above.

[0160] The target image localization method can be implemented as follows: First, the target image is located using a first basic localization method to obtain the first 6DOF (6 Dimensions of Freedom). Then, feature extraction is performed on the target image using a second basic localization method to obtain global and local features. Next, using the global features and the first 6DOF, a search is conducted in the map database to obtain images similar to the target image. Then, the local features of the target image and the similar images are processed to obtain the number of inliers or the projection error between the target image and the similar images. Finally, when the proportion of inliers is greater than a preset proportion, or the projection error is within a preset range, the 6DOF of the target image is determined as the first 6DOF. This localization method can be understood as the localization method with "LOD-VPS" as the main scheme in the "cross-validation localization" described above. Figure 6 The scheme shown in (A) is described above and will not be repeated here.

[0161] In some embodiments, the map in the map database consists of two different types of map data, one type of map data corresponding to a first basic positioning method, and the other type of map data corresponding to a second basic positioning method. For example, the first basic positioning method can be the "GVPS" scheme described above, and the second basic positioning method can be the "LOD-VPS" scheme described above.

[0162] The target image localization method can be implemented as follows: First, the target image is localized using a first basic localization method to obtain a first 6DOF (6th Dimension of Field of View). Next, the image is rendered using model data from a map database at the first 6DOF perspective to obtain a first textureless feature map corresponding to the target image. Then, the target image is processed using a second basic localization method according to a preset processing method to obtain a second textureless map corresponding to the target image. The preset processing method includes at least one of semantic segmentation, instance segmentation, and depth estimation. Finally, when the target parameters between the first and second textureless maps are within a preset range, the 6DOF of the target image is determined to be the first 6DOF. For example, the target parameters may include a loss function between the first and second textureless maps, which may include one or more of semantic intersection-union ratio (CIU), instance intersection-union ratio (CIU), contour distance, or depth error. This localization method can be understood as the localization method with "GVPS" as the main scheme in the "cross-validation localization" described above. Figure 6 The scheme shown in (B) is described above and will not be repeated here.

[0163] In some embodiments, the map database consists of two different types of map data, one type of map data corresponding to a first basic positioning method, and the other type of map data corresponding to a second basic positioning method. For example, the first basic positioning method can be the "LOD-VPS" scheme described above, and the second basic positioning method can be the "GVPS" scheme described above.

[0164] The target image localization method can be implemented as follows: First, the target image is processed using a first basic localization method to obtain point cloud data. Then, the target image is processed using a second basic localization method to obtain the image with the highest similarity to the target image in the map database. Next, the data in the map database that matches the point cloud data is determined. Finally, the first algorithm is used to process the image with the highest similarity to the target image and the data that matches the point cloud data to obtain the 6DOF of the target image. This localization method can be understood as the "multi-class loss fusion localization" scheme described above, i.e. Figure 7 The schemes shown are detailed in the above description and will not be repeated here.

[0165] In some embodiments, the map database consists of two different types of map data, one type of map data corresponding to a first basic positioning method, and the other type of map data corresponding to a second basic positioning method. For example, the first basic positioning method can be the "LOD-VPS" scheme described above, and the second basic positioning method can be the "GVPS" scheme described above.

[0166] The target localization method for locating the target image can be as follows: First, the target image is located using a first basic localization method to obtain the first 6DOF of the target image. Then, using a second basic localization method, combined with the position 3DOF and orientation angle from the first 6DOF, the target image is processed to obtain an image similar to the target image in the map database. Finally, using a first algorithm, combined with the pose 3DOF from the first 6DOF, the image similar to the target image is processed to obtain the 6DOF of the target image. This localization method can be understood as the "rotation-translation separation estimation localization" scheme described above, i.e. Figure 8 The schemes shown are detailed in the above description and will not be repeated here. For example, the first algorithm can be the PNP / BA algorithm.

[0167] In some embodiments, when the number of target images contained in the image data reaches a preset number, the map can be upgraded / updated. Specifically, when upgrading / updating the map, it can be first determined that the number of images contained in the image data has reached the preset number. Then, each image in the image data is located to obtain the pose information of at least a portion of the images in the image data. Next, based on the pose information of at least a portion of the images in the image data, the correspondence between the 2D features and 3D features of each image in the image data is determined from the map database. Next, based on the pose information of at least a portion of the images in the image data, the at least a portion of the images are matched in the map database to obtain images that have a co-view relationship with each of the at least a portion of the images. Next, for images in the image data for which pose information has not been obtained, a search is performed in the map database to obtain images similar to those for which pose information has not been obtained, and 2D feature matching is performed between the obtained images and the images for which pose information has not been obtained to obtain images that have a co-view relationship with those images. Next, for images without pose information, the pose information of the images without pose information is obtained by using the pose information of images with co-view relationships with the images without pose information. Finally, the pose and 3D features of the images with pose information are optimized, thus completing the map upgrade / update. The map upgrade / update process can be the "map upgrade and / or update" method described above, i.e. Figure 2 The solutions described herein are detailed above and will not be repeated here.

[0168] Furthermore, after optimizing the pose and 3D features of the image with acquired pose information, point cloud data from at least a portion of the map in the map database can be read as target point clouds. The 3D features corresponding to the images in the constructed image data can then be used as source point clouds for point cloud registration. Additionally, the coordinates of the 3D features corresponding to the remaining 2D features in the image with existing pose information are determined, and the images with existing pose information and their target parameters are stored in the map database. The target parameters include at least one of the following: global features, pose information, 2D features, and 3D features.

[0169] It is understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. Furthermore, in some possible implementations, each step in the above embodiments may be selectively executed according to actual circumstances; it may be partially or fully executed, without limitation here. Moreover, all or part of any feature of any embodiment of this application can be freely and arbitrarily combined without contradiction. The combined technical solutions are also within the scope of this application.

[0170] Based on the methods described in the above embodiments, this application also provides an electronic device. The electronic device may include: at least one memory for storing a program; and at least one processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor executes the methods described in the above embodiments.

[0171] It should be understood that each step of the above method embodiments can be completed by hardware logic circuits or software instructions in a processor.

[0172] It is understood that the processor in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.

[0173] The method steps in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0174] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)). It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application.

Claims

1. A fusion positioning method based on multiple map types, characterized in that, The method includes: Acquire the target image to be located; Obtain the target location when capturing the target image; The target map quality of the map located at the target location is determined in the map database. The map in the map database consists of various types of map data, and each type of map data corresponds to a basic positioning method. Based on the target map quality, a target positioning method is determined from a plurality of pre-set different positioning methods, wherein the target positioning method includes a fusion positioning strategy, which is generated by dynamically selecting and fusing at least two basic positioning methods based on the target map quality; Using the target localization method, the target image is localized to obtain the 6-DOF pose of the target image; Output the 6DOF of the target image.

2. The method according to claim 1, characterized in that, The maps in the map database consist of two different types of map data. One type of map data corresponds to a first basic positioning method, and the other type of map data corresponds to a second basic positioning method. The step of locating the target image using the target localization method specifically includes: The target image is located using the first basic positioning method to obtain the first 6DOF of the target image; The 2D features and descriptors of the target image are extracted using the second basic localization method; Based on the first 6DOF, the map data in the map database, the 2D features and descriptors of the target image, determine the 3D features in the map database that match the 2D features of the target image; The 6DOF of the target image is determined based on the 2D features of the target image, the 3D features that match the 2D features of the target image, and historical image data. The historical image data includes historical data of images acquired before the target image was acquired, and the historical data includes the 6DOF, 2D features, and 3D features of the images.

3. The method according to claim 1, characterized in that, The maps in the map database consist of two different types of map data. One type of map data corresponds to a first basic positioning method, and the other type of map data corresponds to a second basic positioning method. The step of locating the target image using the target localization method specifically includes: The target image is located using a first basic positioning method to obtain the first 6DOF of the target image; The target image is subjected to feature extraction using the second basic positioning method to obtain the global and local features of the target image; Using the global features of the target image and the first 6DOF, a search is performed in the map data of the map database to obtain an image similar to the target image; The local features of the target image and images similar to the target image are processed to obtain the number of interior points or projection error between the target image and images similar to the target image; When the proportion of the number of interior points is greater than a preset proportion, or when the projection error is within a preset range, the 6DOF of the target image is determined to be the first 6DOF.

4. The method according to claim 1, characterized in that, The maps in the map database consist of two different types of map data. One type of map data corresponds to a first basic positioning method, and the other type of map data corresponds to a second basic positioning method. The step of locating the target image using the target localization method specifically includes: The target image is located using the first basic positioning method to obtain the first 6DOF of the target image; The image under the first 6DOF viewpoint is rendered using the model data in the map database to obtain the first textureless image corresponding to the target image; The target image is processed by the second basic localization method according to a preset processing method to obtain a second textureless image corresponding to the target image. The preset processing method includes at least one of semantic segmentation, instance segmentation and depth estimation. When the target parameters between the first textureless image and the second textureless image are within a preset range, the 6DOF of the target image is determined to be the first 6DOF.

5. The method according to claim 1, characterized in that, The maps in the map database consist of two different types of map data. One type of map data corresponds to a first basic positioning method, and the other type of map data corresponds to a second basic positioning method. The step of locating the target image using the target localization method specifically includes: The target image is processed using a first basic positioning method to obtain point cloud data of the target image; The target image is processed using a second basic positioning method to obtain the image with the highest similarity to the target image in the map database; Identify data in the map database that matches the point cloud data; The first algorithm is used to process the image with the highest similarity to the target image and the data that matches the point cloud data to obtain the 6DOF of the target image.

6. The method according to claim 1, characterized in that, The maps in the map database consist of two different types of map data. One type of map data corresponds to a first basic positioning method, and the other type of map data corresponds to a second basic positioning method. The step of locating the target image using the target localization method specifically includes: The target image is located using a first basic positioning method to obtain the first 6DOF of the target image; By using the second basic positioning method, and combining the position 3DOF and orientation angle in the first 6DOF, the target image is processed to obtain an image similar to the target image in the map database; Using the first algorithm and combining it with the pose 3DOF in the first 6DOF, an image similar to the target image is processed to obtain the 6DOF of the target image.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Determine that the number of images contained in the image data reaches a preset number; The images in the image data are located to obtain the pose information of at least some of the images in the image data; Based on the pose information of at least some images in the image data, determine the correspondence between the 2D features and 3D features of each image in the image data from the map database; Based on the pose information of at least some images in the image data, the at least some images are matched in the map database to obtain images that have a co-view relationship with each of the at least some images; For images in the image data for which pose information has not been obtained, a search is performed in the map database to obtain images similar to the images for which pose information has not been obtained, and 2D feature matching is performed on the obtained images and the images for which pose information has not been obtained to obtain images with co-view relationship with the images for which pose information has not been obtained. For images for which pose information has not been obtained, the pose information of the images for which pose information has not been obtained is obtained by using the pose information of images that have a co-view relationship with the images for which pose information has not been obtained. The pose and 3D features of the image with acquired pose information are optimized.

8. The method according to claim 7, characterized in that, The method further includes: Read the existing point cloud data of at least a portion of the map in the map database as the target point cloud, and use the 3D features corresponding to the image in the constructed image data as the source point cloud for point cloud registration. The coordinates of the 3D features corresponding to the remaining 2D features in the image with existing pose information in the image data are determined, and the image with existing pose information in the image data and the target parameters of the image with existing pose information are stored in the map database. The target parameters include at least one of global features, pose information, 2D features, and 3D features.

9. A map upgrade / update method, characterized in that, The method further includes: Determine that the number of images contained in the image data reaches a preset number; The images in the image data are located to obtain the pose information of at least some of the images in the image data; Based on the pose information of at least some images in the image data, determine the correspondence between the 2D features and 3D features of each image in the image data from the map database; Based on the pose information of at least some images in the image data, the at least some images are matched in the map database to obtain images that have a co-view relationship with each of the at least some images; For images in the image data for which pose information has not been obtained, a search is performed in the map database to obtain images similar to the images for which pose information has not been obtained, and 2D feature matching is performed on the obtained images and the images for which pose information has not been obtained to obtain images with co-view relationship with the images for which pose information has not been obtained. For images for which pose information has not been obtained, the pose information of the images for which pose information has not been obtained is obtained by using the pose information of images that have a co-view relationship with the images for which pose information has not been obtained. The pose and 3D features of the image with acquired pose information are optimized.

10. The method according to claim 9, characterized in that, The method further includes: Read the existing point cloud data of at least a portion of the map in the map database as the target point cloud, and use the 3D features corresponding to the image in the constructed image data as the source point cloud for point cloud registration. The coordinates of the 3D features corresponding to the remaining 2D features in the image with existing pose information in the image data are determined, and the image with existing pose information in the image data and the target parameters of the image with existing pose information are stored in the map database. The target parameters include at least one of global features, pose information, 2D features, and 3D features.

11. An electronic device, characterized in that, include: At least one memory for storing programs; At least one processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method as described in any one of claims 1-10.

12. A computer-readable storage medium storing a computer program that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1-10.

13. A computer program product, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-10.