Visual positioning method, device, storage medium and program product

By identifying the preset target in the first image during visual positioning and matching it in the target image set, the problem of low image matching efficiency caused by the preset target not occupying the main area is solved, and faster scene recognition and image retrieval are achieved.

CN117710702BActive Publication Date: 2025-09-05HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310955226.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-31
Publication Date
2025-09-05
Estimated Expiration
2043-07-31

AI Technical Summary

Technical Problem

In visual positioning, when a preset target in a preset scene does not occupy the main area of ​​the first image, the existing technology leads to low image matching efficiency, affecting the processing speed.

Method used

By identifying whether the first image includes a preset target and matching images with a similarity greater than or equal to a first threshold in the target image set in the database, the point cloud data and spatial pose are determined in combination with the feature association relationship to improve the scene recognition accuracy and image retrieval efficiency.

Benefits of technology

When the preset target does not occupy the main area, the accuracy of scene recognition and image retrieval efficiency are improved, thereby speeding up the processing speed of visual positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117710702B_ABST
    Figure CN117710702B_ABST
Patent Text Reader

Abstract

The present application provides a visual positioning method, apparatus, storage medium, and program product. In this method, a preset target corresponding to a preset scene in a first image is identified to identify the scene corresponding to the first image. When the preset target does not occupy the main area of ​​the first image, the accuracy of identifying the scene corresponding to the first image is improved compared to identifying the scene corresponding to the first image using a scene recognition method. Furthermore, if the first image includes the preset target, an image search is performed in a target image set in a database to match a second image whose similarity to the first image is greater than or equal to a first threshold, rather than performing an image search in the entire database or images related to the scene. This improves the efficiency of image retrieval and thus increases the processing speed of visual positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of visual positioning technology, and in particular to a visual positioning method, device, storage medium and program product. Background Art

[0002] A visual positioning system (VPS) is a technology that locates electronic devices based on images they capture. For example, a device captures an image and matches it with images in a database to obtain a matching image. The captured image's position is then determined based on the position of the matched image. The current location of the electronic device can then be determined based on the captured image's position, achieving the goal of locating the electronic device.

[0003] Currently, in the process of matching a captured image with an image in a database, the efficiency of image retrieval can be improved by identifying the scene corresponding to the captured image and performing image matching in images related to the scene in the database.

[0004] However, if the captured image includes feature points of a preset target in a preset scene, but the preset target in the preset scene does not occupy the main area of ​​the first image, the scene corresponding to the captured image identified at this time is not accurate, and image matching is required in the entire database, resulting in low image retrieval efficiency, thereby affecting the processing speed of visual positioning. Summary of the Invention

[0005] The present application provides a visual positioning method, device, storage medium and program product, which improve the accuracy of identifying the scene corresponding to the first image and can improve the image retrieval efficiency, thereby increasing the processing speed of the visual positioning method.

[0006] To achieve the above objectives, this application adopts the following technical solutions:

[0007] In a first aspect, a visual positioning method is provided, including: determining whether a preset target is included in a first image, the preset target corresponds to a preset scene, and the first image is captured by an electronic device; if the preset target is included in the first image, matching a second image for the first image in a target image set in a database, any image in the target image set includes the preset target, and the similarity between the second image and the first image is greater than or equal to a first threshold; determining point cloud data corresponding to the first image based on an association relationship between features in the first image and features in the second image; and determining the spatial posture of the electronic device when the first image was captured based on the point cloud data corresponding to the first image.

[0008] Based on the above technical solution, in the embodiment of the present application, by identifying the preset target corresponding to the preset scene in the first image, the scene corresponding to the first image is identified. When the preset target does not occupy the main area of ​​the first image, the accuracy of identifying the scene corresponding to the first image is improved compared to using a scene recognition method to identify the scene corresponding to the first image (when using a scene recognition method to identify the scene corresponding to the first image, the identified scene is greatly affected by the feature points of the main area in the first image, that is, it is greatly affected by other feature points in the first image that are unrelated to the feature points of the preset target in the preset scene, and the identified scene corresponding to the first image is not accurate). If the first image includes the preset target, an image search is performed in the target image set in the database to match a second image whose similarity with the first image is greater than or equal to a first threshold, rather than performing an image search in the entire database or images related to the scene, thereby improving the efficiency of image retrieval and thus improving the processing speed of visual positioning.

[0009] In a possible implementation of the first aspect, after determining whether the preset target is included in the first image, it also includes: if the first image does not include the preset target, identifying whether the scene corresponding to the first image is the preset scene; if the scene corresponding to the first image is the preset scene, matching a third image for the first image in the scene image set in the database, wherein the third image corresponds to the same scene as the first image, the images in the scene image set are obtained by pre-shooting the area where the preset scene is located, and the scene image set includes multiple image subsets, and the multiple image subsets include the target image set and correspond one-to-one to multiple preset targets; determining the point cloud data corresponding to the first image based on the correlation between the features in the first image and the features in the third image; and determining the spatial posture of the electronic device when shooting the first image based on the point cloud data corresponding to the first image.

[0010] Based on the above technical solution, in the embodiment of the present application, it is pre-determined whether the first image includes a preset target. If the first image does not include the preset target, the classification model is used to identify the scene corresponding to the first image. This can make up for the shortcoming that the scene corresponding to the first image identified by the classification model is not accurate, thereby improving the accuracy of identifying the scene corresponding to the first image.

[0011] In a possible implementation of the first aspect, after identifying whether the scene corresponding to the first image is the preset scene, it also includes: if the scene corresponding to the first image is not the preset scene, matching the third image for the first image among all images in the database.

[0012] Based on the above technical solution, in the embodiment of the present application, image retrieval is performed in all images in the database. The scope of image retrieval is large, and all third images whose similarity with the first image is greater than or equal to the first threshold can be fully retrieved.

[0013] Alternatively, after identifying whether the scene corresponding to the first image is the preset scene, the method further includes: if the scene corresponding to the first image is not the preset scene, matching the third image for the first image in other images in the database except the scene image set.

[0014] Based on the above technical solution, in this embodiment, it has been determined that the scene corresponding to the first image is not the preset scene. Therefore, there is no need to repeatedly perform image retrieval in the scene image set corresponding to the preset scene. Directly performing image retrieval in other images in the database except the scene image set can further improve the image retrieval efficiency.

[0015] In a possible implementation of the first aspect, the second image is matched for the first image in a target image set in a database, any image in the target image set includes the preset target, and the similarity between the second image and the first image is greater than or equal to a first threshold, including: obtaining a global feature descriptor of the first image; calculating the similarity between the first image and each image in the target image set based on the global feature descriptor of the first image and the global feature descriptor of each image in the target image set in the database; and determining the image in the target image set whose similarity is greater than or equal to the first threshold as the second image.

[0016] In a possible implementation of the first aspect, the similarity between the first image and each image in the target image set is calculated based on the global feature descriptor of the first image and the global feature descriptor of each image in the target image set of the database, including: determining the distance between the global feature descriptor of the first image and the global feature descriptor of each image in the target image set of the database; determining the similarity based on the distance, wherein the smaller the distance, the greater the similarity.

[0017] In a possible implementation of the first aspect, determining whether the first image includes a preset target includes: using a pre-trained target detection model to determine whether the first image includes the preset target.

[0018] In a possible implementation of the first aspect, identifying whether the scene corresponding to the first image is the preset scene includes: using a pre-trained classification model to identify whether the scene corresponding to the first image is the preset scene.

[0019] In a second aspect, a visual positioning device is provided, including: a processing module, used to determine whether a preset target is included in a first image, the preset target corresponds to a preset scene, and the first image is captured by an electronic device; if the preset target is included in the first image, a second image is matched for the first image in a target image set in a database of a storage module, any image in the target image set includes the preset target, and the similarity between the second image and the first image is greater than or equal to a first threshold; it is also used to determine the point cloud data corresponding to the first image based on the correlation relationship between the features in the first image and the features in the second image; and based on the point cloud data corresponding to the first image, determine the spatial posture of the electronic device when the first image was captured.

[0020] Based on the above technical solution, in the embodiment of the present application, by identifying the preset target corresponding to the preset scene in the first image, the scene corresponding to the first image is identified. When the preset target does not occupy the main area of ​​the first image, the accuracy of identifying the scene corresponding to the first image is improved compared to using a scene recognition method to identify the scene corresponding to the first image (when using a scene recognition method to identify the scene corresponding to the first image, the identified scene is greatly affected by the feature points of the main area in the first image, that is, it is greatly affected by other feature points in the first image that are unrelated to the feature points of the preset target in the preset scene, and the identified scene corresponding to the first image is not accurate). If the first image includes the preset target, an image search is performed in the target image set in the database to match a second image whose similarity with the first image is greater than or equal to a first threshold, rather than performing an image search in the entire database or images related to the scene, thereby improving the efficiency of image retrieval and thus improving the processing speed of visual positioning.

[0021] In a possible implementation of the second aspect, after determining whether the first image includes a preset target, the processing module is further used to identify whether the scene corresponding to the first image is the preset scene if the first image does not include the preset target; if the scene corresponding to the first image is the preset scene, a third image is matched for the first image in the scene image set in the database of the storage module, wherein the third image corresponds to the same scene as the first image, the images in the scene image set are obtained by pre-shooting the area where the preset scene is located, and the scene image set includes multiple image subsets, and the multiple image subsets include the target image set and correspond one-to-one to multiple preset targets; it is also used to determine the point cloud data corresponding to the first image based on the correlation between the features in the first image and the features in the third image, and determine the spatial posture of the electronic device when the first image was shot based on the point cloud data corresponding to the first image.

[0022] Based on the above technical solution, in the embodiment of the present application, it is pre-determined whether the first image includes a preset target. If the first image does not include the preset target, the classification model is used to identify the scene corresponding to the first image. This can make up for the shortcoming that the scene corresponding to the first image identified by the classification model is not accurate, thereby improving the accuracy of identifying the scene corresponding to the first image.

[0023] In a possible implementation of the second aspect, after the processing module identifies whether the scene corresponding to the first image is the preset scene, if the scene corresponding to the first image is not the preset scene, the processing module is further used to match the third image for the first image among all images in the database of the storage module.

[0024] Based on the above technical solution, in the embodiment of the present application, image retrieval is performed in all images in the database. The scope of image retrieval is large, and all third images whose similarity with the first image is greater than or equal to the first threshold can be fully retrieved.

[0025] After identifying whether the scene corresponding to the first image is the preset scene, the processing module is further used to match the third image for the first image among other images other than the scene image set in the database of the storage module if the scene corresponding to the first image is not the preset scene.

[0026] Based on the above technical solution, in this embodiment, it has been determined that the scene corresponding to the first image is not the preset scene. Therefore, there is no need to repeatedly perform image retrieval in the scene image set corresponding to the preset scene. Directly performing image retrieval in other images in the database except the scene image set can further improve the image retrieval efficiency.

[0027] In a possible implementation of the second aspect, the processing module is specifically used to obtain a global feature descriptor of the first image; calculate the similarity between the first image and each image in the target image set based on the global feature descriptor of the first image and the global feature descriptor of each image in the target image set of the database; and determine the image in the target image set whose similarity is greater than or equal to a first threshold as the second image.

[0028] In a possible implementation of the second aspect, the processing module is specifically used to determine the distance between the global feature descriptor of the first image and the global feature descriptor of each image in the target image set of the database; determine the similarity based on the distance, wherein the smaller the distance, the greater the similarity.

[0029] In a possible implementation of the second aspect, the processing module is specifically used to use a pre-trained target detection model to determine whether the first image includes the preset target.

[0030] In a possible implementation manner of the second aspect, the processing module is specifically configured to use a pre-trained classification model to identify whether the scene corresponding to the first image is the preset scene.

[0031] In a third aspect, a visual positioning device is provided, comprising a memory and a processor, wherein the memory is used to store instructions. When the instructions are executed by the processor, the visual positioning device executes the visual positioning method in the first aspect or any possible implementation of the first aspect.

[0032] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed, the visual positioning method in the first aspect or any possible implementation of the first aspect is implemented.

[0033] In a fifth aspect, a computer program product is provided, comprising: a computer program code, which, when executed on a computer, enables the computer to execute the visual positioning method in the first aspect or any possible implementation of the first aspect.

[0034] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is an exemplary flow chart of a visual positioning method provided in an embodiment of the present application;

[0036] Figure 2 It is a floor plan diagram of a certain floor of a shopping mall;

[0037] Figure 3 This is a schematic diagram of a scenario in which the visual positioning method provided in an embodiment of the present application is applied to AR navigation;

[0038] Figure 4 yes Figure 3 a schematic diagram of the electronic device in the scenario schematic diagram;

[0039] Figure 5 Schematic diagram of the structure of a visual positioning device provided in an embodiment of the present application;

[0040] Figure 6 This is a structural diagram of another visual positioning device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0041] The technical solution in this application will be described below with reference to the accompanying drawings.

[0042] In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a way to describe the association relationship of associated objects, indicating that three relationships may be included, for example, A and / or B can mean: including A alone, including A and B at the same time, and including B alone.

[0043] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.

[0044] With the maturity of augmented reality (AR) technology, services based on AR technology have appeared in all aspects of people's lives and have demonstrated outstanding superiority. For example, AR is used to implement navigation services. Compared with traditional navigation, AR navigation can provide users with more intuitive and accurate navigation services. It is understandable that the accuracy and timeliness of AR navigation depend on the accuracy and timeliness of visual positioning.

[0045] Visual positioning is a method for determining position based on images, implemented through a visual positioning system (VPS). VPS can be used in AR navigation. When a user or the system needs to know their own position, the user simply raises their phone to capture an image of the scene. The VPS then generates the camera pose at the time the image was taken, which is the user's own position.

[0046] The specific process of visual positioning is as follows: the electronic device captures an image, and by matching the image with the image in the database, an image matching the captured image is obtained. The shooting position of the captured image is determined according to the position of the matched image. The current position of the electronic device can be determined according to the shooting position of the captured image to achieve the purpose of positioning the electronic device.

[0047] Currently, in the process of matching a captured image with an image in a database, the efficiency of image retrieval can be improved by identifying the scene corresponding to the captured image and performing image matching in images related to the scene in the database.

[0048] However, if the captured image includes feature points of a preset target in a preset scene, but the preset target in the preset scene does not occupy the main area of ​​the first image, the scene corresponding to the captured image identified at this time is not accurate, and image matching is required in the entire database, resulting in low image retrieval efficiency, thereby affecting the processing speed of visual positioning.

[0049] In view of this, an embodiment of the present application provides a visual positioning method, including: determining whether a preset target is included in a first image, the preset target corresponds to a preset scene, and the first image is captured by an electronic device; if the preset target is included in the first image, matching a second image for the first image in a target image set in a database, any image in the target image set includes the preset target, and the similarity between the second image and the first image is greater than or equal to a first threshold; determining point cloud data corresponding to the first image based on the correlation relationship between the features in the first image and the features in the second image; and determining the spatial posture of the electronic device when the first image was captured based on the point cloud data corresponding to the first image.

[0050] Based on the above technical solution, in the embodiment of the present application, by identifying the preset target corresponding to the preset scene in the first image, the scene corresponding to the first image is identified. When the preset target does not occupy the main area of ​​the first image, the accuracy of identifying the scene corresponding to the first image is improved compared to using a scene recognition method to identify the scene corresponding to the first image (when using a scene recognition method to identify the scene corresponding to the first image, the identified scene is greatly affected by the feature points of the main area in the first image, that is, it is greatly affected by other feature points in the first image that are unrelated to the feature points of the preset target in the preset scene, and the identified scene corresponding to the first image is not accurate). If the first image includes the preset target, an image search is performed in the target image set in the database to match a second image whose similarity with the first image is greater than or equal to a first threshold, rather than performing an image search in the entire database or images related to the scene, thereby improving the efficiency of image retrieval and thus improving the processing speed of visual positioning.

[0051] The following describes in detail a visual positioning method provided by an embodiment of the present application with reference to the accompanying drawings. Figure 1 An exemplary flowchart of the visual positioning method.

[0052] Step S101: Determine whether a first image includes a preset target, where the preset target corresponds to a preset scene. If it is determined that the first image includes the preset target, execute step S102; if it is determined that the first image does not include the preset target, execute step S105.

[0053] Specifically, the preset scene can be a scene that appears frequently in a large environment. For example, assuming the large environment is a shopping mall, garage, or park, the preset scene can be an entrance or exit or a bathroom. Scenes such as entrances or exits or bathrooms appear frequently in large environments such as shopping malls, garages, and parks.

[0054] A preset target is a target that is representative of a preset scene and is different from other scenes. In one example, assuming the environment is a garage or a shopping mall, the preset scene can be an entrance or exit, and the preset targets corresponding to the preset scene can be: an elevator, an escalator, a zebra crossing, etc. It can be understood that the preset scene corresponding to an image of any preset target such as an elevator, an escalator, a zebra crossing, etc. is an entrance or exit. In another example, assuming the environment is a park, the preset scene can be a bathroom, and the preset targets corresponding to the preset scene can be: a road sign, a road sign, a wall bathroom sign, a washbasin, etc. It can be understood that the preset scene corresponding to an image of any preset target such as a road sign, a road sign, a wall bathroom sign, a washbasin, etc. is a bathroom.

[0055] In the embodiment of the present application, the first image is captured by an electronic device, and whether the first image corresponds to a preset scene can be determined by determining whether the first image captured by the electronic device includes a preset target. The preset target in the first image can be one or multiple. If there are multiple preset targets in the first image, these multiple preset targets may correspond to the same preset scene or multiple preset scenes.

[0056] By identifying a preset target corresponding to a preset scene in a first image and identifying the scene corresponding to the first image, when the preset target does not occupy the main body area of ​​the first image, the accuracy of identifying the scene corresponding to the first image is improved compared to using a scene recognition method to identify the scene corresponding to the first image (when using a scene recognition method to identify the scene corresponding to the first image, the identified scene is greatly affected by the feature points of the main body area in the first image, that is, it is greatly affected by other feature points in the first image that are unrelated to the feature points of the preset target in the preset scene, and the identified scene corresponding to the first image is not accurate). If the first image includes the preset target, an image search is performed in the target image set in the database to match a second image whose similarity with the first image is greater than or equal to a first threshold, rather than performing an image search in the entire database or images related to the scene, thereby improving the efficiency of image retrieval and thus improving the processing speed of visual positioning.

[0057] In a possible implementation, determining whether the first image includes a preset target includes: using a pre-trained target detection model to determine whether the first image includes the preset target.

[0058] Specifically, training images marked with preset targets can be input into the target detection model for training until the target detection model can accurately identify the preset targets in any image. At this point, the target detection model training is complete. The trained target detection model can be directly used to identify whether the first image contains the preset target and the category of the preset target contained in the first image. For example, if the first image contains the preset target of an elevator, the pre-trained target detection model can not only determine that the first image contains the preset target, but also determine that the preset target contained in the first image is an elevator.

[0059] Realistically, the target detection model can be implemented using a region convolutional neural network (R-CNN) or an object detection algorithm (You Only Look Once, YOLO).

[0060] Step S102: matching a second image for a first image in a target image set in a database, wherein any image in the target image set includes a preset target, and a similarity between the second image and the first image is greater than or equal to a first threshold.

[0061] In the embodiments of the present application, a database of macro-environments is pre-established. Images in the database are obtained by photographing the area containing the macro-environment. The images in the database include at least one scene image set, each corresponding to a preset scene. The images in the scene image set are obtained by photographing the area containing the preset scene. The scene image set includes multiple image subsets, each of which is a target image set. A target image set corresponds to a preset target, and any image in the target image set includes the preset target.

[0062] For example, assuming the macro environment is a shopping mall, the preset scenes are entrances and exits and restrooms. The preset targets corresponding to the preset scene "entrances and exits" are: elevators, escalators, zebra crossings, etc., and the preset targets corresponding to the preset scene "restroom" are: road signs, road signs, wall restroom signs, washbasins, etc.

[0063] The images in the pre-established shopping mall database were obtained by photographing the mall environment. The database includes the following content, as shown in Table 1: a first scene image set corresponding to the preset "entrance / exit" scene, a second scene image set corresponding to the preset "restroom" scene, and other images not corresponding to preset scenes. The images in the first scene image set were obtained by photographing the area where the preset "entrance / exit" scene is located, while the images in the second scene image set were obtained by photographing the area where the preset "restroom" scene is located.

[0064] The first scene image set corresponding to the preset scene of "entrance and exit" includes multiple image subsets, as well as other images that do not correspond to preset targets but correspond to the preset scene of "entrance and exit". The multiple image subsets include a first target image set, a second target image set, a third target image set, etc., among which the first target image set corresponds to the preset target of "elevator", the second target image set corresponds to the preset target of "escalator", and the third target image set corresponds to the preset target of "zebra crossing".

[0065] The second scene image set corresponding to the preset scene of "bathroom" also includes multiple image subsets, as well as other images that do not correspond to the preset targets but correspond to the preset scene of "bathroom". The multiple image subsets include a fourth target image set, a fifth target image set, a sixth target image set, a seventh target image set, etc., among which the fourth target image set corresponds to the preset target of "road sign", the fifth target image set corresponds to the preset target of "road sign", the sixth target image subset corresponds to the preset target of "wall bathroom sign", and the seventh target image set corresponds to the preset target of "washbasin".

[0066] Table 1

[0067]

[0068] Specifically, if the first image includes a preset target, image retrieval is performed in the target image set in the database to match a second image whose similarity with the first image is greater than or equal to a first threshold, rather than performing image retrieval in the entire database or images related to the scene. This improves the efficiency of image retrieval and thus improves the processing speed of visual positioning.

[0069] In practice, if the first image includes multiple preset targets, multiple target image sets corresponding to the multiple preset targets are determined in the database. Image retrieval is performed in each of these multiple target image sets. A second image having a similarity with the first image greater than or equal to a first threshold is matched from each target image set. The second images matched from each target image set may include multiple images. Steps S103 and S104 are then executed for the second images matched from each target image set.

[0070] In one possible implementation, a second image is matched for a first image in a target image set in a database, any image in the target image set includes a preset target, and the similarity between the second image and the first image is greater than or equal to a first threshold, including: obtaining a global feature descriptor of the first image; calculating the similarity between the first image and each image in the target image set based on the global feature descriptor of the first image and the global feature descriptor of each image in the target image set in the database; and determining an image in the target image set whose similarity is greater than or equal to the first threshold as the second image.

[0071] Specifically, matching a first image with a second image in a database can be achieved using feature matching. Feature matching can be divided into global feature matching and local feature matching. Global features refer to the overall attributes of an image. Common global features include color, texture, and shape features, such as intensity histograms. Local features are features extracted from local regions of an image, including edges, corners, lines, curves, and regions with special attributes. Common local features fall into two categories: corner and region. Compared to global features, using local features for feature matching offers higher image matching accuracy, higher matching accuracy, and stronger resistance to interference (such as flipping, re-photographing, color shifts, and background interference). This generally meets normal target matching requirements. However, when searching for images in databases containing tens or even hundreds of millions of images, the time and space overhead of using local features becomes unacceptable. Therefore, matching a first image with a second image in a database is generally performed using global features. Global features are represented by global feature descriptors, such as the Histogram of Oriented Gradient (HOG) and the Deformable Part Model (DPM).

[0072] In the embodiment of the present application, a global feature descriptor of the first image is obtained, and based on the global feature descriptor of the first image and the global feature descriptors of each image in the target image set in the database, a similarity between the first image and each image in the target image set in the database is calculated, thereby determining an image in the target image set whose similarity is greater than or equal to a first threshold as the second image. The first threshold can be set according to actual circumstances.

[0073] Realistically, the global feature descriptor of the first image and the global feature descriptor of each image in the target image set of the database can be obtained through a deep neural network, for example, a Network Vector of Locally Aggregated Descriptors (NetVLAD).

[0074] Realistically, each image in the database stores a global feature descriptor corresponding to the image.

[0075] In one possible implementation, the similarity between the first image and each image in the target image set is calculated based on the global feature descriptor of the first image and the global feature descriptors of each image in the target image set of the database, including: determining the distance between the global feature descriptor of the first image and the global feature descriptors of each image in the target image set of the database; and determining the similarity based on the distance, wherein the smaller the distance, the greater the similarity.

[0076] Specifically, the similarity can be determined based on the distance between the global feature descriptor of the first image and the global feature descriptors of each image in the target image set. The smaller the distance, the greater the similarity. Specifically, a linear relationship can be set between distance and similarity. For example, it can be assumed that when the distance is 0, the similarity is 100%; when the distance is 100, the similarity is 0%. In this way, when the distance is known, the similarity can be calculated based on this linear relationship.

[0077] It is feasible to sort the images in the target image set according to the distance from small to large, that is, the similarity from large to small, and select the first N images to be determined as the second image; or set a first threshold and determine the images in the target image set with a similarity greater than or equal to the first threshold as the second image.

[0078] Step S103: determining point cloud data corresponding to the first image based on the association relationship between the features in the first image and the features in the second image.

[0079] Specifically, after matching a second image in the database whose similarity to the first image is greater than or equal to a first threshold, local features of the first image are obtained, and feature matching is performed on the local features of the first image and the local features of the second image to determine the association relationship between the local features of the first image and the local features of the second image. Feature matching here refers to matching local features in the first image with the same local features in the second image. Therefore, the point cloud data of the local features in the second image can be used as the point cloud data of the same local features in the first image.

[0080] Realistically, each image in the database stores a plurality of local features corresponding to the image and point cloud data corresponding to each local feature.

[0081] Step S104: determining the spatial posture of the electronic device when taking the first image based on the point cloud data corresponding to the first image.

[0082] Specifically, after determining the point cloud data for each local feature in the first image, the three-dimensional point cloud data corresponding to each two-dimensional local feature in the first image can be obtained. Based on the three-dimensional point cloud data corresponding to each two-dimensional local feature in the first image, the spatial pose of the electronic device when the first image was captured can be calculated. This spatial pose consists of six degrees of freedom, also known as the 6-DOF camera pose. The six degrees of freedom are composed of the camera's coordinates (x, y, z) relative to the 3D world, as well as the angular deflection, pitch, and roll around the three coordinate axes.

[0083] Realistically, methods for calculating the spatial pose based on the three-dimensional point cloud data corresponding to each two-dimensional local feature in the first image include but are not limited to: Random Sample Consensus (RANSAC), perspective-n-point (PNP), direct linear transformation method, etc.

[0084] Step S105: Identify whether the scene corresponding to the first image is a preset scene. If the scene corresponding to the first image is a preset scene, execute step S106; if the scene corresponding to the first image is not a preset scene, execute step S107.

[0085] Specifically, if the first image does not include a preset target, the scene corresponding to the first image is identified as a preset scene to determine whether the scene the electronic device was in when the image was taken was a preset scene. For example, if the electronic device captures the first image in a shopping mall environment and the scene corresponding to the first image is identified as an "entrance" or "restroom," the scene the electronic device was in when the image was taken was an "entrance" or "restroom."

[0086] Realistically, identifying whether the scene corresponding to the first image is a preset scene includes: using a pre-trained classification model to identify whether the scene corresponding to the first image is the preset scene.

[0087] In this embodiment, training images pre-labeled with preset scenes can be input into the classification model for training until the classification model can accurately identify the preset scenes in any image. At this point, the classification model training is complete. The trained classification model can be directly used to identify whether the scene corresponding to the first image is a preset scene, and to which preset scene in the larger environment the preset scene in the first image belongs. For example, the pre-trained classification model can not only determine that the scene corresponding to the first image is a preset scene, but also determine that the preset scene corresponding to the first image is a "bathroom" or "entrance / exit."

[0088] Optionally, the classification model may adopt a visual geometry group model (Visual Geometry Group, VGG), a support vector machine model (SVM), a logistic regression model (LR), a decision tree model, a random forest model or a gradient boosting tree model, etc.

[0089] It is worth noting that the classification model determines the scene corresponding to the first image based on the various feature points in the first image. If the first image includes feature points of a preset target in the preset scene, but the preset target in the preset scene does not occupy the main area of ​​the first image, the classification model will be greatly affected by the feature points in the main area of ​​the first image, which can easily lead to inaccurate identification of the scene corresponding to the first image. However, in the present application, whether the first image includes the preset target is determined in advance. If the preset target is not included in the first image, the classification model is used to identify the scene corresponding to the first image. This can compensate for the disadvantage of the classification model identifying the scene corresponding to the first image being inaccurate, thereby improving the accuracy of identifying the scene corresponding to the first image.

[0090] Step S106: Matching a third image for the first image in the scene image set in the database.

[0091] Among them, the third image and the first image correspond to the same scene, the images in the scene image set are obtained by pre-shooting the area where the preset scene is located, the scene image set includes multiple image subsets, the multiple image subsets include a target image set and correspond one-to-one to multiple preset targets.

[0092] Specifically, if the scene corresponding to the first image is identified as a preset scene, image retrieval is performed within a set of scene images in the database, rather than searching the entire database, thereby improving image retrieval efficiency. After the image retrieval, a third image can be matched whose similarity to the first image is greater than or equal to a first threshold. The matched third image corresponds to the same preset scene as the first image, and the third image may include multiple images.

[0093] It is worth noting that the scene image set includes not only multiple image subsets corresponding one-to-one to multiple target image sets, but also other images that do not correspond to preset targets but correspond to preset scenes. These images do not belong to any target image set.

[0094] It should be noted that if the scene corresponding to the first image is identified as a preset scene, image retrieval can be performed among all images in the scene image set in the database, so as to find a third image whose similarity with the first image is greater than or equal to the first threshold as comprehensively as possible; or, image retrieval can be directly performed among other images in the scene image set that do not correspond to the preset target but correspond to the preset scene, so as to further improve the retrieval efficiency.

[0095] It can be seen that when the image captured by the electronic device includes a "pre-set target", the visual positioning method in the embodiment of the present application has higher image retrieval efficiency and faster processing speed than when the image captured by the electronic device does not include a "pre-set target".

[0096] Step S107: Matching a third image to the first image in all images in the database, or matching a third image to the first image in other images in the database except the scene image set.

[0097] In one possible implementation, if it is identified that the scene corresponding to the first image is not a preset scene, an image search may be performed among all images in the database to match a third image whose similarity with the first image is greater than or equal to a first threshold. The matched third image corresponds to the same scene as the first image. The scenes corresponding to the first and third images mentioned here are other scenes in the larger environment that are different from the preset scenes. It should be noted that all images in the database include images in all scene image sets, and also include other images that do not correspond to the preset scenes. In this embodiment, image search is performed among all images in the database. The scope of image search is relatively large, and all third images whose similarity with the first image is greater than or equal to the first threshold can be comprehensively retrieved.

[0098] In another possible implementation, if it is identified that the scene corresponding to the first image is not the preset scene, image retrieval may be performed in other images in the database other than the scene image set, that is, image retrieval may be performed in other images that do not correspond to the preset scene, and a third image whose similarity with the first image is greater than or equal to a first threshold is matched. The matched third image corresponds to the same scene as the first image. The scenes corresponding to the first and third images mentioned here are other scenes in the large environment that are different from the preset scenes. In this embodiment, it has been determined that the scene corresponding to the first image is not the preset scene. Therefore, there is no need to repeatedly perform image retrieval in the scene image set corresponding to the preset scene. Directly performing image retrieval in other images in the database other than the scene image set can further improve image retrieval efficiency.

[0099] Step S108: determining point cloud data corresponding to the first image based on the association relationship between the features in the first image and the features in the third image.

[0100] Step S109: Determine the spatial posture of the electronic device when taking the first image based on the point cloud data corresponding to the first image.

[0101] The above step S108 is substantially the same as the above step S103, and the above step S109 is substantially the same as the above step S104. To avoid repetition, they will not be described again.

[0102] It is worth noting that the visual positioning method in the embodiment of the present application can be executed by the electronic device itself, such as a mobile phone, a mobile computer, a robot, a car, etc. The visual positioning method in the embodiment of the present application can also be executed by a cloud server. After the electronic device captures the first image, the first image is sent to the cloud server. The cloud server receives the first image from the electronic device, and then the cloud server executes the following steps: Figure 1 The steps shown determine the spatial posture of the electronic device when taking the first image, and then the cloud server returns the spatial posture of the electronic device when taking the first image to the electronic device.

[0103] In one example, in the visual positioning of AR navigation for a shopping mall, it is assumed that there are about 30,000 images in the database, about 10,000 images in a scene image set, and about 3,000 images in a target image set in the scene image set. Directly searching in the database based on the first image taken by the electronic device, matching images whose similarity with the first image is greater than or equal to the first threshold, the retrieval takes about 1 second. However, through the visual positioning method in the embodiment of the present application, if the first image includes a preset target, the scope of image retrieval can be directly narrowed from the relevant images of the entire shopping mall to the images including the preset target, that is, directly narrowing the scope of image retrieval from more than 30,000 images in the database to more than 3,000 images in the target image set corresponding to the preset target, compared to image retrieval in all images of the entire data (about 30,000 images) or image retrieval in the images of the scene image set (about 10,000 images), the image retrieval efficiency can be further improved.

[0104] For example, Figure 2The figure shows a floor plan of a shopping mall floor. It shows that this floor includes multiple shops, such as 1 to 25G, the Shoe Garden, the Beauty and Makeup Source, and the Women's Bags area, as well as escalators, elevators, and entrances and exits. Suppose an electronic device captures an image in front of shop 5B. The image includes shop 5B and the escalator in front of it, but the escalator is not in the main area of ​​the image. In this case, if scene recognition methods are used to identify the scene corresponding to the image, since the escalator is not in the main area of ​​the image, the scene recognition method can only identify the scene in the image as related to the "shop." In this case, the search scope is limited to the area where the shop is located. However, since there are many shops in the mall, there are many images related to the "shop." Therefore, searching within images related to the "shop" will be inefficient.

[0105] Using the visual positioning method in the embodiments of this application, when constructing a database for the shopping mall, we can pre-set a preset scene, "Entrance and Exit," and treat the objects "escalator," "elevator," and "entrance and exit" in the image as preset objects within this preset scene. The constructed database includes a scene image set corresponding to "entrance and exit," which includes a target image set corresponding to "escalator," a target image set corresponding to "elevator," and a target image set corresponding to "entrance and exit."

[0106] Similarly, let's take the example of an electronic device taking an image in front of shop 5B, and the image includes shop 5B and the escalator in front of shop 5B, but the escalator is not in the main area of ​​the image. After obtaining the captured image, it is determined that the image includes the preset target "escalator". At this time, the search scope can be narrowed down to the area where the escalator is located. The number of "escalators" in the mall is less than the number of "shops". Therefore, the number of images in the target image set corresponding to "escalators" is much smaller than the number of images related to "shops". Therefore, performing image retrieval in the target image set corresponding to "escalators" in the database greatly improves the image retrieval efficiency compared to performing image retrieval in images related to "shops" using scene recognition methods.

[0107] The following combination Figure 3 and Figure 4 The application of the visual positioning method in the embodiment of the present application in AR navigation is described.

[0108] like Figure 3 As shown, the user holds the electronic device 300 and performs AR navigation outdoors. When the user opens the map navigation and enters the AR navigation interface, see Figure 4The display interface of the electronic device 300 is divided into two parts by an arc 301. The upper part of the electronic device 300 displays the current environment being photographed, and the lower part of the electronic device 300 displays a navigation map. The position of the black short arrow 302 in the navigation map is the current position of the user in the map, and the direction indicated by the direction indication arrow 303 is the navigation direction.

[0109] The application of the visual positioning method in AR navigation is illustrated by taking the cloud server as an example. When the user enters the AR navigation interface, the electronic device 300 captures the image of the environment (with the image of the environment). Figure 3 The upper half of the image of the electronic device 300 is the same as the image of the upper half of the electronic device 300 in the image), and the image captured by the electronic device 300 is sent to the cloud server. The cloud server determines the spatial posture of the electronic device 300 in the pre-set database based on the captured image, that is, the spatial posture of the user itself. The cloud server returns the spatial posture to the electronic device 300. After receiving the spatial posture, the electronic device 300 adjusts the spatial posture according to the spatial posture. Figure 4 The position of the black short arrow 302 in the navigation map is adjusted, that is, the current position of the user in the navigation map is adjusted to achieve accurate positioning of the user's position.

[0110] This method can obtain the spatial position of the electronic device only through the environmental information around it. It is a low-cost positioning method and is particularly suitable for indoor scenarios such as large venues where the global positioning system (GPS) signal is unstable.

[0111] It should be understood that the above examples are intended to help those skilled in the art understand the embodiments of the present application, and are not intended to limit the embodiments of the present application to the specific numerical values ​​or specific scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or variations based on the above examples, and such modifications or variations also fall within the scope of the embodiments of the present application.

[0112] Figure 5 A visual positioning device 400 provided in an embodiment of the present application includes a processing module 410 and a storage module 420 .

[0113] The processing module 410 is used to determine whether a preset target is included in a first image, where the preset target corresponds to a preset scene, and the first image is captured by an electronic device; if the preset target is included in the first image, a second image is matched for the first image in a target image set in a database of the storage module 420, any image in the target image set includes the preset target, and a similarity between the second image and the first image is greater than or equal to a first threshold.

[0114] The processing module 410 is also used to determine the point cloud data corresponding to the first image based on the correlation between the features in the first image and the features in the second image; and to determine the spatial posture of the electronic device when taking the first image based on the point cloud data corresponding to the first image.

[0115] In the embodiment of the present application, the processing module 410 identifies a preset target corresponding to a preset scene in the first image and identifies the scene corresponding to the first image. When the preset target does not occupy the main area of ​​the first image, the accuracy of identifying the scene corresponding to the first image is improved compared to using a scene recognition method to identify the scene corresponding to the first image (when using a scene recognition method to identify the scene corresponding to the first image, the identified scene is greatly affected by the feature points of the main area in the first image, that is, it is greatly affected by other feature points in the first image that are unrelated to the feature points of the preset target in the preset scene, and the identified scene corresponding to the first image is not accurate). If the first image includes the preset target, the processing module 410 performs an image search in the target image set in the database to match a second image whose similarity with the first image is greater than or equal to a first threshold, rather than performing an image search in the entire database or images related to the scene, thereby improving the efficiency of image retrieval and thus improving the processing speed of visual positioning.

[0116] In practice, the processing module 410 is specifically configured to determine whether the first image includes the preset target using a pre-trained target detection model. The target detection model may be implemented using a Regions with CNN features (R-CNN) or an object detection algorithm (You Only Look Once, YOLO).

[0117] Optionally, in some embodiments, after determining whether the first image includes a preset target, the processing module 410 is further used to identify whether the scene corresponding to the first image is the preset scene if the first image does not include the preset target; if the scene corresponding to the first image is the preset scene, a third image is matched for the first image in the scene image set in the database of the storage module 420, wherein the third image corresponds to the same scene as the first image, the images in the scene image set are obtained by pre-shooting the area where the preset scene is located, and the scene image set includes multiple image subsets, and the multiple image subsets include the target image set and correspond one-to-one to multiple preset targets; it is also used to determine the point cloud data corresponding to the first image based on the correlation between the features in the first image and the features in the third image, and determine the spatial posture of the electronic device when the first image was shot based on the point cloud data corresponding to the first image.

[0118] Based on the above technical solution, in the embodiment of the present application, it is pre-determined whether the first image includes a preset target. If the first image does not include the preset target, the classification model is used to identify the scene corresponding to the first image. This can make up for the shortcoming that the scene corresponding to the first image identified by the classification model is not accurate, thereby improving the accuracy of identifying the scene corresponding to the first image.

[0119] In practice, the processing module 410 is specifically configured to use a pre-trained classification model to identify whether the scene corresponding to the first image is the preset scene. Optionally, the classification model may be a Visual Geometry Group (VGG) model, a Support Vector Machine (SVM) model, a Logistic Regression (LR) model, a Decision Tree model, a Random Forest model, or a Gradient Boosting Tree model.

[0120] Realistically, the processing module 410 is specifically used to obtain the global feature descriptor of the first image; calculate the similarity between the first image and each image in the target image set based on the global feature descriptor of the first image and the global feature descriptor of each image in the target image set of the database; and determine the image in the target image set whose similarity is greater than or equal to a first threshold as the second image.

[0121] Specifically, matching a first image with a second image in a database can be achieved using feature matching. Feature matching can be divided into global feature matching and local feature matching. Global features refer to the overall attributes of an image. Common global features include color, texture, and shape features, such as intensity histograms. Local features are features extracted from local regions of an image, including edges, corners, lines, curves, and regions with special attributes. Common local features fall into two categories: corner and region. Compared to global features, using local features for feature matching offers higher image matching accuracy, higher matching accuracy, and stronger resistance to interference (such as flipping, re-photographing, color shifts, and background interference). This generally meets normal target matching requirements. However, when searching for images in databases containing tens or even hundreds of millions of images, the time and space overhead of using local features becomes unacceptable. Therefore, matching a first image with a second image in a database is generally performed using global features. Global features are represented by global feature descriptors, such as the Histogram of Oriented Gradient (HOG) and the Deformable Part Model (DPM).

[0122] In the embodiment of the present application, a global feature descriptor of the first image is obtained, and based on the global feature descriptor of the first image and the global feature descriptors of each image in the target image set in the database, a similarity between the first image and each image in the target image set in the database is calculated, thereby determining an image in the target image set whose similarity is greater than or equal to a first threshold as the second image. The first threshold can be set according to actual circumstances.

[0123] Realistically, the global feature descriptor of the first image and the global feature descriptor of each image in the target image set of the database can be obtained through a deep neural network, for example, a Network Vector of Locally Aggregated Descriptors (NetVLAD).

[0124] In one example, the processing module 410 is specifically used to determine the distance between the global feature descriptor of the first image and the global feature descriptors of each image in the target image set of the database; determine the similarity based on the distance, wherein the smaller the distance, the greater the similarity.

[0125] Specifically, the similarity can be determined based on the distance between the global feature descriptor of the first image and the global feature descriptors of each image in the target image set. The smaller the distance, the greater the similarity. Specifically, a linear relationship can be set between distance and similarity. For example, it can be assumed that when the distance is 0, the similarity is 100%; when the distance is 100, the similarity is 0%. In this way, when the distance is known, the similarity can be calculated based on this linear relationship.

[0126] It is feasible to sort the images in the target image set according to the distance from small to large, that is, the similarity from large to small, and select the first N images to be determined as the second image; or set a first threshold and determine the images in the target image set with a similarity greater than or equal to the first threshold as the second image.

[0127] In one implementation, after identifying whether the scene corresponding to the first image is the preset scene, the processing module 410 is further configured to match the third image for the first image among all images in the database of the storage module 420 if the scene corresponding to the first image is not the preset scene.

[0128] Based on the above technical solution, in the embodiment of the present application, image retrieval is performed in all images in the database. The scope of image retrieval is large, and all third images whose similarity with the first image is greater than or equal to the first threshold can be fully retrieved.

[0129] In another implementation, after identifying whether the scene corresponding to the first image is the preset scene, the processing module 410 is further used to match the third image for the first image among other images other than the scene image set in the database of the storage module 420 if the scene corresponding to the first image is not the preset scene.

[0130] Based on the above technical solution, in this embodiment, it has been determined that the scene corresponding to the first image is not the preset scene. Therefore, there is no need to repeatedly perform image retrieval in the scene image set corresponding to the preset scene. Directly performing image retrieval in other images in the database except the scene image set can further improve the image retrieval efficiency.

[0131] The visual positioning device 400 of the embodiment of the present application may correspond to the visual positioning method described in the embodiment of the present application, and the above and other operations and / or functions of each unit in the visual positioning device 400 are respectively to achieve Figure 1 For the sake of brevity, the corresponding process of the method in will not be repeated here.

[0132] Figure 6 This is a schematic structural block diagram of a visual positioning device 500 provided in an embodiment of the present application. The visual positioning device 500 includes: a processor 510, a memory 520, a communication interface 530, and a bus 540.

[0133] It should be understood that Figure 6 The visual positioning device 500 shown can be an electronic device, such as a mobile phone, a laptop computer, a robot, etc., or can be a cloud server.

[0134] It should be understood that Figure 6 The processor 510 in the visual positioning device 500 shown may correspond to Figure 5 The processing module 410 in the visual positioning device 400; the memory 520 in the visual positioning device 500 may correspond to the storage module 420 in the visual positioning device 400.

[0135] The processor 510 may be connected to a memory 520. The memory 520 may be used to store the program code and data. Therefore, the memory 520 may be a storage unit within the processor 510, an external storage unit independent of the processor 510, or a component including both a storage unit within the processor 510 and an external storage unit independent of the processor 510.

[0136] Optionally, the visual positioning device 500 may further include a bus 540. The memory 520 and the communication interface 530 may be connected to the processor 510 via the bus 540. The bus 540 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus 540 may be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 5 The fact that only one line is used does not mean that there is only one bus or one type of bus.

[0137] It should be understood that in the embodiment of the present application, the processor 510 can adopt a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Alternatively, the processor 510 uses one or more integrated circuits to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0138] The memory 520 may include a read-only memory and a random access memory, and provides instructions and data to the processor 510. A portion of the processor 510 may also include a non-volatile random access memory. For example, the processor 510 may also store information about the device type.

[0139] When the visual positioning device is running, the processor 510 executes the computer-executable instructions in the memory 520 to utilize the hardware resources in the visual positioning device to perform the operation steps of the above-mentioned visual positioning method.

[0140] It should be understood that the visual positioning device 500 according to the embodiment of the present application may correspond to the visual positioning device 400 in the embodiment of the present application, and may correspond to the device that performs the operation according to the embodiment of the present application. Figure 1 The corresponding subjects in the method shown, and the above and other operations and / or functions of each module in the visual positioning device 500 are respectively to achieve Figure 1 For the sake of brevity, the corresponding process of the method in will not be repeated here.

[0141] The present application also provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed, the visual positioning method provided in the embodiment of the present application is implemented.

[0142] The present application also provides a computer program product, which includes: computer program code, which, when executed on a computer, enables the computer to execute the visual positioning method provided in an embodiment of the present application.

[0143] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid state drive (SSD).

[0144] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of this application.

[0145] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0146] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0147] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.

[0148] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may be physically included separately, or two or more units may be integrated into one unit.

[0149] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a memory (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0150] The above description is merely a specific implementation of the embodiments of the present application, but the scope of protection of the embodiments of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the embodiments of the present application should be included in the scope of protection of the embodiments of the present application. Therefore, the scope of protection of the embodiments of the present application should be based on the scope of protection of the claims.

Claims

1. A visual positioning method, characterized in that: include: A database of the macro environment is pre-established, wherein the images in the database are obtained by photographing the area where the macro environment is located. The images in the database include at least one scene image set, wherein each scene image set corresponds to a preset scene, wherein the preset scene is a scene in the macro environment, and the images in the scene image set are obtained by photographing the area where the preset scene is located. The scene image set includes multiple target image sets, wherein each target image set corresponds to a preset target, wherein the preset target is a target that is representative of the preset scene and distinguishable from other scenes, and any image in the target image set includes the corresponding preset target. Determining whether a first image includes a preset target, the preset target corresponding to a preset scene, and the first image is captured by an electronic device in the large environment; If the first image includes the preset target, matching a second image for the first image in the target image set corresponding to the preset target, wherein a similarity between the second image and the first image is greater than or equal to a first threshold; determining point cloud data corresponding to the first image based on an association relationship between features in the first image and features in the second image; If the first image does not include the preset target, identifying whether the scene corresponding to the first image is the preset scene; If the scene corresponding to the first image is the preset scene, matching a third image for the first image in the scene image set corresponding to the preset scene, wherein the third image and the first image correspond to the same preset scene; If the scene corresponding to the first image is not the preset scene, matching a third image for the first image from images other than the scene image set in the database, wherein the similarity between the third image and the first image is greater than the first threshold, and the third image and the first image correspond to scenes other than the preset scene; determining point cloud data corresponding to the first image based on an association relationship between features in the first image and features in the third image; The spatial posture of the electronic device when taking the first image is determined based on the point cloud data corresponding to the first image.

2. The visual positioning method according to claim 1, characterized in that: The matching of the first image with the second image in the target image set corresponding to the preset target includes: Obtaining a global feature descriptor of the first image; Calculating the similarity between the first image and each image in the target image set based on the global feature descriptor of the first image and the global feature descriptor of each image in the target image set corresponding to the preset target; An image in the target image set whose similarity is greater than or equal to a first threshold is determined as the second image.

3. The visual positioning method according to claim 2, characterized in that: The calculating, based on the global feature descriptor of the first image and the global feature descriptor of each image in the target image set corresponding to the preset target, the similarity between the first image and each image in the target image set includes: Determining a distance between a global feature descriptor of the first image and a global feature descriptor of each image in the target image set corresponding to the preset target; The similarity is determined according to the distance, wherein the smaller the distance is, the greater the similarity is.

4. The visual positioning method according to any one of claims 1 to 3, characterized in that: The determining whether the first image includes the preset target includes: A pre-trained target detection model is used to determine whether the first image includes the preset target.

5. The visual positioning method according to any one of claims 1 to 3, characterized in that: The identifying whether the scene corresponding to the first image is the preset scene includes: A pre-trained classification model is used to identify whether the scene corresponding to the first image is the preset scene.

6. A visual positioning device, characterized in that: include: A storage module, wherein the storage module stores a pre-established database of the macro environment, wherein the images in the database are obtained by photographing the area where the macro environment is located, and wherein the images in the database include at least one scene image set, wherein each scene image set corresponds to a preset scene, wherein the preset scene is a scene in the macro environment, and wherein the images in the scene image set are obtained by photographing the area where the preset scene is located; wherein the scene image set includes a plurality of target image sets, wherein each target image set corresponds to a preset target, wherein the preset target is a target that is representative of the preset scene and is distinguishable from other scenes, and wherein any image in the target image set includes the corresponding preset target; A processing module, configured to determine whether a first image includes a preset target, the preset target corresponding to a preset scene, and the first image is captured by an electronic device in the large environment; If the first image includes the preset target, matching a second image for the first image in the target image set corresponding to the preset target, wherein a similarity between the second image and the first image is greater than or equal to a first threshold; further configured to determine point cloud data corresponding to the first image based on an association relationship between features in the first image and features in the second image; The processing module is further configured to, after determining whether the first image includes the preset target, identify whether the scene corresponding to the first image is the preset scene if the first image does not include the preset target; further configured to, if the scene corresponding to the first image is the preset scene, match a third image for the first image in the scene image set corresponding to the preset scene, wherein the third image and the first image correspond to the same preset scene; further configured to, if the scene corresponding to the first image is not the preset scene, match a third image for the first image from other images in the database of the storage module, excluding the set of scene images, wherein the similarity between the third image and the first image is greater than the first threshold, and the third image and the first image correspond to other scenes different from the preset scene; further configured to determine point cloud data corresponding to the first image based on an association relationship between features in the first image and features in the third image; It is also used to determine the spatial posture of the electronic device when taking the first image based on the point cloud data corresponding to the first image.

7. The visual positioning device according to claim 6, characterized in that: The processing module is specifically used to obtain the global feature descriptor of the first image; calculate the similarity between the first image and each image in the target image set corresponding to the preset target based on the global feature descriptor of the first image and the global feature descriptor of each image in the target image set; and determine the image in the target image set whose similarity is greater than or equal to a first threshold as the second image.

8. The visual positioning device according to claim 7, characterized in that: The processing module is specifically used to determine the distance between the global feature descriptor of the first image and the global feature descriptor of each image in the target image set corresponding to the preset target; determine the similarity based on the distance, wherein the smaller the distance, the greater the similarity.

9. The visual positioning device according to any one of claims 6 to 8, characterized in that: The processing module is specifically configured to use a pre-trained target detection model to determine whether the first image includes the preset target.

10. The visual positioning device according to any one of claims 6 to 8, characterized in that: The processing module is specifically configured to use a pre-trained classification model to identify whether the scene corresponding to the first image is the preset scene.

11. A visual positioning device, characterized in that: The visual positioning device includes a memory and a processor, wherein the memory is used to store instructions. When the instructions are executed by the processor, the visual positioning device executes the visual positioning method according to any one of claims 1 to 5.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed, the visual positioning method according to any one of claims 1 to 5 is implemented.

13. A computer program product, characterized in that The computer program product comprises: a computer program code, and when the computer program code is run on a computer, the computer is caused to execute the visual positioning method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Visual positioning method and system and computer readable storage medium

    CN111046125A

  • Indoor positioning method and device of intelligent mobile terminal and electronic equipment

    CN111199564A

  • Cognitive navigation method and system for structured scene expression

    CN111369688A