Method and apparatus for positioning

The matching algorithm in the image database determines the pose of the image under local and world coordinate systems, which solves the time-consuming and labor-intensive problem of building a 3D point cloud feature library in the prior art, and achieves low-cost and high-efficiency AR positioning.

CN114565663BActive Publication Date: 2025-07-08HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011271315.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-13
Publication Date
2025-07-08
Estimated Expiration
2040-11-13

AI Technical Summary

Technical Problem

The 3D point cloud feature library for building an environment in existing AR technology requires relying on professionals and equipment, which is time-consuming and labor-intensive and expensive, making it difficult to achieve high-efficiency and low-cost positioning.

Method used

By acquiring the local and global feature information of the image, the matching algorithm in the image database determines the position of the image under the local and world coordinate systems, and realizes the positioning of the image without the need to build a global 3D point cloud feature library.

Benefits of technology

It reduces the cost and time of image database construction, realizes high-efficiency and low-cost positioning, and is suitable for terminal devices such as mobile phones to build and position image databases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114565663B_ABST
    Figure CN114565663B_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus for positioning. In the embodiments of the present application, at least one second image matching the first image captured by the first terminal is determined in an image database, as well as the poses of the first image and the at least one second image in the local coordinate system of the first terminal. Then, according to the pose of the first image in the first local coordinate system, the poses of the at least one second image in the first local coordinate system, and the poses of the at least one second image in the world coordinate system, the pose of the first image in the world coordinate system is output, thereby realizing the positioning of the first image. Therefore, the embodiments of the present application can obtain the relationship between the local coordinate system and the world coordinate system based on the poses of a part of the images in the local coordinate system and the world coordinate system, so as to position the first image in the world coordinate system without collecting 3D point clouds for each image. The present application can be applied to fields such as virtual reality (VR) and augmented reality (AR).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of augmented reality (AR) technology, and more specifically, to a positioning method and device. Background Art

[0002] With the gradual maturity of the fifth generation (5G) communication technology and the development of mobile phone camera hardware and computing power, intelligent application services based on visual AR technology are becoming more and more abundant. AR technology is a technology that cleverly integrates virtual information with the real world. It widely uses a variety of technical means such as multimedia, three-dimensional modeling, real-time tracking and registration, intelligent interaction, and sensing. It simulates computer-generated text, images, three-dimensional models, music, videos and other virtual information and applies them to the real world. The two types of information complement each other, thereby achieving "enhancement" of the real world.

[0003] In order to achieve wide coverage of AR applications, it is necessary to build a map feature library for posture acquisition. For example, first, in the offline database construction stage, professional acquisition equipment such as road measurement vehicles, drones, measuring laser scanners, optical cameras, etc. can be used to collect three-dimensional (3D) point cloud (x, y, z) information of the surrounding environment to build a map feature library. Then, in the online positioning stage, you can take a picture of the environment, extract the two-dimensional (2D) feature points of the picture, and match them with the map feature library built in the offline database construction stage. Retrieve a series of 2D-3D matching point pairs, and then use a posture solution algorithm, such as PnP (point-n-point), to obtain the terminal's posture.

[0004] However, the above solutions rely on the construction of a map feature library containing point cloud information that is complete and highly consistent with the environment. However, the construction of a map feature library containing point cloud information requires the use of professionals and acquisition equipment, which is time-consuming, labor-intensive and costly. Therefore, a high-efficiency and low-cost positioning solution is urgently needed. Summary of the invention

[0005] The present application provides a positioning method and device, which can obtain the relationship between the local coordinate system and the world coordinate system based on the pose of a part of the image in the local coordinate system and the pose in the world coordinate system, so as to position the image to be positioned in the world coordinate system without collecting 3D point clouds for each image.

[0006] In a first aspect, a positioning method is provided, which can be applied to a terminal or the cloud. In this method, first, a first image captured by a first terminal can be obtained. According to the first information of the first image and the first information of the images in the image database, at least one second image matching the first image can be determined in the image database, where the first information is used to indicate the global features of the image, and the image database includes the first information of multiple frames of images, the second information of multiple frames of images, and the poses of multiple frames of images in the world coordinate system, where the second information is used to indicate the local features of the image. Then, according to the second information of the first image and the second information of at least one second image, the poses of the first image and at least one second image in the first local coordinate system of the first terminal can be determined. After that, according to the pose of the first image in the first local coordinate system, the poses of at least one second image in the first local coordinate system, and the poses of at least one second image in the world coordinate system, the pose of the first image in the world coordinate system can be output.

[0007] Therefore, in the embodiment of the present application, by determining at least one second image matching the first image captured by the first terminal in the image database, and the poses of the first image and at least one second image in the local coordinate system of the first terminal, and according to the pose of the first image in the first local coordinate system, the poses of at least one second image in the first local coordinate system, and the poses of at least one second image in the world coordinate system, the pose of the first image in the world coordinate system is output, thereby realizing the positioning of the first image. Therefore, the embodiment of the present application can obtain the relationship between the local coordinate system and the world coordinate system based on the poses of a part of the images in the local coordinate system and in the world coordinate system, so as to position the image to be positioned (such as the first image) in the world coordinate system without collecting 3D point clouds for each image.

[0008] In the embodiment of the present application, the pose of the first image in the world coordinate system is the pose of the first terminal in the world coordinate system when the first image is captured, and can also be referred to as the pose of the camera.

[0009] In some embodiments, the pose of the above at least one second image in the world coordinate system and the second information can be obtained from the image database, and the present application does not make any limitations thereto.

[0010] As an example, the first terminal in the embodiment of the present application can be a mobile device such as a mobile phone or an autonomous driving vehicle, and the present application does not make any limitations thereto.

[0011] Compared with the solution of constructing a 3D point cloud feature library of the environment (including global 3D point cloud features) in the existing solution and calculating the pose based on feature point matching, the embodiment of the present application does not utilize the 3D point cloud features of the image, but uses the pose of the image for positioning. Since constructing a 3D point cloud feature library of the environment requires relying on professionals and acquisition equipment, which is time-consuming, laborious and costly, and the present application does not need to obtain 3D point cloud features. Therefore, on the one hand, the present application can use low-cost devices, such as cameras with low resolution and small field of view (such as mobile phone cameras), to construct an image database without relying on professionals and professional equipment, and can reduce the data volume of the image database, thereby reducing the cost of constructing the image database; on the other hand, the present application can shorten the time for constructing the database, which helps to construct the image database with high efficiency. Therefore, the positioning solution based on the pose of the image and the image database in the present application can help to perform high-efficiency and low-cost positioning.

[0012] In some embodiments, the above at least one second image may form an image set, and this image set may be referred to as a similar image set of the first image.

[0013] In some embodiments, the above first image and at least one second image may form an image set, and this image set may be referred to as a local image set. Among them, the above first local coordinate system, that is, the camera coordinate system of the first terminal, may be a relative coordinate system constructed with any frame of image in the local image set as the origin of the first local coordinate system.

[0014] As a possible implementation manner, the mapping relationship (or also referred to as the conversion relationship) between the first local coordinate system and the world coordinate system may be determined according to the pose of the above at least one second image in the first local coordinate system and the pose of the at least one second image in the world coordinate system. Then, according to this mapping relationship, the pose of the first image in the first local coordinate system may be converted to determine the pose of the first image in the world coordinate system, and this pose may be output.

[0015] As an example, the mapping relationship between the world coordinate system and the first local coordinate system may be described by a rotation matrix or a translation vector, and the present application does not limit this.

[0016] In the embodiment of the present application, the world coordinate system may also be referred to as an absolute coordinate system, a global coordinate system, etc. As an example, the world coordinate system may be used as a reference coordinate system in the environment to describe the position of a camera (which may also be a terminal including this camera), or may also be used to describe the position of any object in the environment.

[0017] As a possible implementation, the image database may include multiple mapping relationships, where one mapping relationship may indicate the correspondence between the identifier of an image, the first information of the image, the second information of the image, and the pose of the image in the world coordinate system. This application does not limit this.

[0018] In combination with the first aspect, in some implementations of the first aspect, the method may further include establishing the above-mentioned image database.

[0019] In combination with the first aspect, in some implementations of the first aspect, the video captured by the second terminal may be obtained, and the poses of multiple frames of images in the video in the second local coordinate system of the second terminal may be obtained. Then, based on the third information and the poses of multiple frames of images in the video in the second local coordinate system, the poses of multiple frames of images in the video in the world coordinate system may be determined. Wherein, the third information is used to indicate the positions of at least some images in the video in the map, and the map is associated with the world coordinate system. After that, the first information and the second information of multiple frames of images in the video may be obtained. In this way, the establishment of the image database can be achieved.

[0020] Therefore, the embodiments of this application can determine the positions of multiple frames of images in the video in the world coordinate system by obtaining the poses of multiple frames of images in the video captured by the second terminal in the second local coordinate system of the second terminal, and combining the relevant positions of multiple frames of images in the video in the map, and obtain the first information and the second information of these images to establish the image database. Therefore, when constructing the image database in this application, only the pose of the image in the world coordinate system, as well as the first information and the second information of the image, need to be obtained, without the need to obtain the 3D point cloud features of the environment (such as global 3D point cloud features).

[0021] It should be noted that in the embodiments of this application, the above-mentioned first terminal and second terminal may be the same terminal device or different terminal devices, which is not limited.

[0022] In some embodiments, after obtaining the pose of the first image in the world coordinate system, the first information of the first image, the second information of the first image, and the pose of the first image in the world coordinate system may be added to the image database, so as to update the image database.

[0023] As an example, the third information may be used to indicate the position of the first frame image in the video on the map and / or the position of the last frame image on the map. Exemplarily, the map may be a map obtained through the global positioning system (GPS) or a map obtained through the global navigation satellite system (GNSS). This application does not make any limitation thereto.

[0024] As a possible implementation manner, a video may be collected based on the camera of the second terminal, and the pose of each frame image in the video in the second local coordinate system of the second terminal may be obtained through a simultaneous localization and mapping (SLAM) algorithm. Further, the pose of each image in the video in the world coordinate system may be calculated by combining the position of the first frame image in the video on the map and the position of the last frame image in the video on the map. Here, the second local coordinate system, that is, the camera coordinate system of the second terminal, may be a relative coordinate system constructed with any frame image in the video as the origin of the second local coordinate system.

[0025] Exemplarily, the video may be collected according to a planned collection route. For example, the user may hold the terminal, start from the starting point of the route, continuously take pictures of the environment along the route until reaching the end point of the route to obtain the video corresponding to the route. For some or all of the video images in the video, the SLAM algorithm may be used to obtain the pose of the image in the local coordinate system of the camera. Meanwhile, the terminal may also obtain the position of the starting point of the route on the map and / or the position of the end point of the route on the map to be used for calculating the pose of the image in the video in the world coordinate system.

[0026] Combined with the first aspect, in some implementation manners of the first aspect, the first position of the first image may further be obtained, and at least one image may be determined in the image database according to the first position of the first image. The distance between the position of each image in the at least one image and the first position of the first image is less than a first threshold, and the first position is from the GPS module or the wireless fidelity (WiFi) module of the first terminal.

[0027] Among them, as a specific implementation manner of matching the first information of the first image with the first information of the images in the image database to determine at least one second image that matches the first image in the image database, the first information of the first image may be matched with the first information of the at least one image, and at least one second image that matches the first image may be obtained from the at least one image.

[0028] It should be noted that the above first position is a "coarse positioning" position with relatively low accuracy, and the error can be, for example, 3 to 10 meters.

[0029] Therefore, in the embodiment of the present application, at least one image is determined in the image database according to the first position of the first image, and then the first information of the first image is matched with the first information of the at least one image, and at least one second image matching the first image is obtained from the at least one image. In this way, it is not necessary to match the first information of the first image with the first information of all images in the image database, thereby reducing the computing amount of the terminal, reducing the matching time consumption, and being beneficial to improving the efficiency of obtaining the second image.

[0030] As an example, the at least one image obtained according to the first position in the image data may be an image set, and this image set may be referred to as an image candidate set.

[0031] As a possible implementation manner for determining at least one second image according to the first information of the first image and the first information of the images in the image database or the image candidate set, the similarity between the first information of the first image and the first information of the images in the image database or the image candidate set may be calculated, and then at least one second image is determined according to the similarity. The present application does not limit this.

[0032] As a possible implementation manner, after obtaining the similarity between the images in the image database or the image candidate set and the first image, the calculated multiple similarities may be sorted to obtain the top m images with the highest similarity as the second images, where m is an integer greater than 1. As an example, m may be 20. Or, in some other implementation manners, the images with a similarity greater than a preset threshold to the first image may be used as the second images. The present application does not limit this.

[0033] Here, the process of calculating the similarity between the first information of the first image and the first information of the images in the image database or the image candidate set may be referred to as matching the first image with the images in the image database or the image candidate set. The present application does not limit this.

[0034] In combination with the first aspect, in some implementations of the first aspect, as an implementation of determining at least one second image that matches the first image in the image database based on the first information of the first image and the first information of the images in the image database, a plurality of third images that match the first image can be determined in the image database based on the first information of the first image and the first information of the images in the image database, and then images in the plurality of third images that have a distance greater than a second threshold from the cluster center of the plurality of third images are deleted to obtain the at least one second image. The cluster center is determined based on the positions of the plurality of third images in the world coordinate system.

[0035] Exemplarily, an image that has a distance greater than the second threshold from the cluster center of the plurality of third images is an outlier image among the plurality of third images.

[0036] Here, the set of images composed of the plurality of third images can also be referred to as a set of similar images, and the set of similar images includes the set of similar images composed of the at least one second image described above. Additionally, when there are no outlier images among the plurality of third images, the operation of deleting images may not be performed. In this case, the third image is the second image.

[0037] Since images at different positions may have very similar textures, errors may occur during the similarity calculation process, which may lead to outlier images in the set of similar images. Therefore, by removing these outlier images in this application, errors in the images in the set of similar images can be eliminated, which helps to obtain a more accurate set of similar images. Here, obtaining a more accurate set of similar images can help make the pose of the subsequent obtained second image in the first local coordinate system more accurate, which in turn helps to improve the accuracy of the pose of the first image in the world coordinate system.

[0038] In combination with the first aspect, in some implementations of the first aspect, as an implementation of determining at least one second image that matches the first image in the image database based on the first information of the first image and the first information of the images in the image database, a plurality of fourth images that match the first image can be determined in the image database based on the first information of the first image and the first information of the images in the image database, and then images in the plurality of fourth images with an angle less than a third threshold and / or a distance less than a fourth threshold are deleted to obtain the at least one second image. The angle is the difference in the poses of at least two images in the world coordinate system, and the distance is the difference in the positions of at least two images in the world coordinate system. As an example, the angle can be the difference in the pitch, yaw, or roll angles of the poses of at least two images in the world coordinate system.

[0039] Exemplarily, among the multiple fourth images, the images with an angle less than the third threshold and / or a distance less than the fourth threshold are redundant images among the multiple second images.

[0040] Here, the image set composed of the multiple fourth images can also be referred to as a set of similar images, and the set of similar images includes the set of similar images composed of at least one of the above-mentioned second images. Additionally, when there are no redundant images among the multiple fourth images, the operation of deleting images may not be performed. At this time, the fourth images are the second images.

[0041] Since the images with an angle less than the preset threshold and / or a distance less than the preset threshold have a high degree of overlap in spatial distribution, there are redundant images in the set of similar images. Therefore, in the embodiments of the present application, by deleting the images with an angle less than the preset value and / or a distance less than the preset value in the set of similar images, the overlap degree of the remaining second images in the set of similar images can be made moderate, and the surrounding space can be evenly covered, which further helps to obtain a more refined set of similar images. Here, obtaining a more refined set of similar images can help reduce the computational amount of the poses of the subsequent obtained second images in the first local coordinate system, and further help improve the efficiency of obtaining the pose of the first image in the world coordinate system.

[0042] In some embodiments, after the outlier images are deleted from the multiple second images, redundant images may be further deleted, and the present application does not limit this.

[0043] In a second aspect, an embodiment of the present application provides a positioning device for performing the method in the first aspect or any possible implementation manner of the first aspect. Specifically, the device includes modules for performing the method in the first aspect or any possible implementation manner of the first aspect. The device may include an acquisition unit, a processing unit, and an output unit.

[0044] The acquisition unit is configured to acquire a first image captured by a first terminal.

[0045] The processing unit is configured to determine at least one second image that matches the first image in the image database according to the first information of the first image and the first information of the images in the image database, where the first information is used to indicate the global features of the image, and the image database includes the first information of multiple frames of images, the second information of the multiple frames of images, and the poses of the multiple frames of images in the world coordinate system, where the second information is used to indicate the local features of the image.

[0046] The processing unit is further configured to determine the poses of the first image and the at least one second image in the first local coordinate system of the first terminal according to the second information of the first image and the second information of the at least one second image.

[0047] An output unit, configured to output the pose of the first image in the world coordinate system according to the pose of the first image in the first local coordinate system, the pose of the at least one second image in the first local coordinate system, and the pose of the at least one second image in the world coordinate system.

[0048] In combination with the second aspect, in some implementation manners of the second aspect, it further includes a building unit, configured to build the image database.

[0049] In combination with the second aspect, in some implementation manners of the second aspect, the building unit is specifically configured to obtain a video captured by a second terminal, obtain the poses of multiple frames of images in the video in a second local coordinate system of the second terminal, and determine the poses of the multiple frames of images in the video in the world coordinate system according to third information and the poses of the multiple frames of images in the video in the second local coordinate system. Wherein, the third information is used to indicate the positions of at least some of the images in the video in a map, and the map is associated with the world coordinate system.

[0050] The building unit is specifically further configured to obtain the first information and the second information of the multiple frames of images in the video.

[0051] In combination with the second aspect, in some implementation manners of the second aspect, the obtaining unit is further configured to obtain a first position of the first image, where the first position comes from a GPS module or a WiFi module of the first terminal.

[0052] Wherein, the obtaining unit can receive data sent by the GPS module or the WiFi module, such as the above-mentioned first position.

[0053] Optionally, the obtaining unit can also send a request message to the GPS module or the WiFi module, and the request message is used to request the GPS module or the WiFi module to send the first position of the first image collected by it to the obtaining unit. In response to the request message, the GPS module or the WiFi module can send the first position to the obtaining unit.

[0054] The processing unit is further configured to determine at least one image in the image database according to the first position of the first image, and the distance between the position of each image in the at least one image and the first position of the first image is less than a first threshold.

[0055] The processing unit is further configured to match the first information of the first image with the first information of the at least one image, and obtain at least one second image that matches the first image from the at least one image.

[0056] In combination with the second aspect, in some implementations of the second aspect, the processing unit is specifically configured to: determine, in the image database, a plurality of third images that match the first image according to the first information of the first image and the first information of the images in the image database, and delete the images in the plurality of third images whose distance from the cluster center of the plurality of third images is greater than a second threshold, so as to obtain the at least one second image, where the cluster center is determined according to the positions of the plurality of third images in the world coordinate system.

[0057] In combination with the second aspect, in some implementations of the second aspect, the processing unit is specifically configured to: determine, in the image database, a plurality of fourth images that match the first image according to the first information of the first image and the first information of the images in the image database, and delete the images in the plurality of fourth images whose angle is less than a third threshold and / or whose distance is less than a fourth threshold, so as to obtain the at least one second image, where the angle is the difference in the poses of at least two images in the world coordinate system, and the distance is the difference in the positions of at least two images in the world coordinate system.

[0058] In a third aspect, an embodiment of the present application provides a positioning device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method in the first aspect or any possible implementation manner of the first aspect as described above.

[0059] In a fourth aspect, an embodiment of the present application provides a computer-readable medium for storing a computer program, where the computer program includes instructions for executing the method in the first aspect or any possible implementation manner of the first aspect.

[0060] In a fifth aspect, an embodiment of the present application further provides a computer program product including instructions, when the computer program product runs on a computer, the computer is caused to execute the method in the first aspect or any possible implementation manner of the first aspect.

[0061] It should be understood that for the beneficial effects obtained from the descriptions of the second to fifth aspects and the corresponding implementation manners of the present application, reference may be made to the beneficial effects obtained from the first aspect and the corresponding implementation manners of the present application, which will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a schematic block diagram of a positioning device provided by an embodiment of the present application;

[0063] Figure 2 is a schematic block diagram of a system architecture of a solution applied to an embodiment of the present application;

[0064] Figure 3 It is a schematic flowchart of a positioning method provided by an embodiment of the present application;

[0065] Figure 4 It is an example of an indoor acquisition route in an embodiment of the present application;

[0066] Figure 5 It is an example of a video extracted from a video in an embodiment of the present application;

[0067] Figure 6 It is an example of a visualization result of an image database provided by an embodiment of the present application;

[0068] Figure 7 It is an example of retrieving a second image in an embodiment of the present application;

[0069] Figure 8 It is a specific example of performing region clustering on multiple frames of second images;

[0070] Figure 9 It is an example of further deleting redundant images after deleting outlier images in multiple second images;

[0071] Figure 10 It is a specific example of performing pose solution in an embodiment of the present application;

[0072] Figure 11 It is a specific example of a real-time positioning result in an embodiment of the present application;

[0073] Figure 12 It is a schematic flowchart of another positioning method provided by an embodiment of the present application;

[0074] Figure 13 It is a schematic block diagram of another positioning device in an embodiment of the present application;

[0075] Figure 14 It is a structural schematic diagram of another positioning device in an embodiment of the present application. Detailed implementation manners

[0076] Next, the technical solutions in the present application will be described with reference to the accompanying drawings.

[0077] First, relevant terms involved in the present application will be briefly introduced.

[0078] Pose: Information used to indicate the position and orientation of a certain image or a certain content in the image. Among them, the content can be an object, a person, a building, an animal, etc., and this application does not make any restrictions. As an example, the pose can be 6 degrees of freedom (6DOF). The position can be represented by coordinates (X, Y, Z) in the Euclidean space, and the orientation is represented by the pitch angle, yaw angle, and roll angle of the rotation coordinates.

[0079] Simultaneous Localization and Mapping (SLAM): Start moving from an unknown position in an unknown environment, perform self-localization according to position estimation and the map during the movement, and at the same time build an incremental map based on the self-localization to achieve autonomous positioning and navigation. Through SLAM tracking and positioning, the pose of the camera in the relative space (i.e., the relative coordinate system) can be obtained.

[0080] Figure 1 It is a schematic block diagram of a positioning device 100 provided by an embodiment of this application. The device 100 can be applied to terminals, such as mobile phones, wearable devices, virtual reality (VR) devices, AR devices, in-vehicle intelligent terminals, etc. The device 100 can also be applied to the cloud, such as a server, and this application embodiment does not make any limitations in this regard. As an example, the device 100 can be used for visual positioning.

[0081] As Figure 1 shown, the device 100 includes an image retrieval module 110 and a pose calculation module 120.

[0082] Among them, the image retrieval module 110 is used to obtain a first image captured by a first terminal (which can also be called an image to be located, an image that needs to be located, etc.), and determine at least one second image that matches the first image in the image database according to the first information of the first image and the first information of the images in the image database.

[0083] Among them, the first information is used to indicate the global feature of the image, and can also be called global feature information. The image database includes the first information of multiple frames of images, the second information of multiple frames of images, and the poses of multiple frames of images in the world coordinate system. Among them, the second information is used to indicate the local feature of the image, and can also be called local feature information.

[0084] In some embodiments, the above at least one second image can form an image set, and this image set can be called a similar image set of the first image.

[0085] In the embodiments of the present application, the image database may also be referred to as an image pose library, an image pose feature library, a pose feature library, etc., and no limitation is imposed thereon.

[0086] The pose calculation module 120 may determine the poses of the first image and at least one second image in the first local coordinate system of the first terminal according to the second information of the first image and the second information of at least one second image. Thereafter, according to the pose of the first image in the first local coordinate system, the poses of at least one second image in the first local coordinate system, and the poses of at least one second image in the world coordinate system, the pose of the first image in the world coordinate system may be output.

[0087] As a possible implementation, the mapping relationship (or also referred to as the conversion relationship) between the first local coordinate system and the world coordinate system may be determined according to the poses of the at least one second image in the first local coordinate system and the poses of the at least one second image in the world coordinate system. Then, according to this mapping relationship, the pose of the first image in the first local coordinate system may be converted to determine the pose of the first image in the world coordinate system, and this pose may be output.

[0088] As an example, the mapping relationship between the world coordinate system and the first local coordinate system may be described by a rotation matrix or a translation vector, and no limitation is imposed thereon in the present application.

[0089] In the embodiments of the present application, the world coordinate system may also be referred to as an absolute coordinate system. As an example, the world coordinate system may be used as a reference coordinate system in the environment to describe the position of a camera (which may also be a terminal including the camera), or may also be used to describe the position of any object in the environment.

[0090] As a possible implementation, the image database may include multiple mapping relationships, and one of the mapping relationships may indicate the correspondence between the identifier of an image, the first information of the image, the second information of the image, and the pose of the image in the world coordinate system, and no limitation is imposed thereon in the present application.

[0091] As an example, the pose calculation module 120 may obtain the poses of the at least one second image in the world coordinate system and the second information from the image database, and no limitation is imposed thereon in the present application.

[0092] In the embodiments of the present application, the pose of the first image in the world coordinate system is the pose of the first terminal in the world coordinate system when the first image is captured, and may also be referred to as the pose of the camera.

[0093] In some embodiments, the above-mentioned first image and at least one second image can form an image set, which can be referred to as a local image set. Among them, the above-mentioned first local coordinate system, that is, the camera coordinate system of the first terminal, can be a relative coordinate system constructed with the origin of the first local coordinate system being any frame of image in the local image set. As an example, SLAM can be used to construct the first local coordinate system, and the present application does not limit this.

[0094] Therefore, in the embodiments of the present application, by determining at least one second image that matches the first image captured by the first terminal in the image database, and the poses of the first image and the at least one second image in the local coordinate system of the first terminal, and according to the pose of the first image in the first local coordinate system, the poses of the at least one second image in the first local coordinate system, and the poses of the at least one second image in the world coordinate system, the pose of the first image in the world coordinate system is output, so as to realize the positioning of the first image. Therefore, the embodiments of the present application can obtain the relationship between the local coordinate system and the world coordinate system based on the poses of a part of the images in the local coordinate system and in the world coordinate system, so as to position the image to be positioned (such as the first image) in the world coordinate system, without collecting 3D point clouds for each image.

[0095] Compared with the existing solution of constructing a 3D point cloud feature library of the environment (which includes global 3D point cloud features) and calculating the pose based on feature point matching, the embodiments of the present application do not utilize the 3D point cloud features of the images, but utilize the poses of the images for positioning. Since constructing a 3D point cloud feature library of the environment requires relying on professionals and acquisition equipment, which is time-consuming, laborious, and costly, and the present application does not need to obtain 3D point cloud features. Therefore, on the one hand, the present application can use low-cost devices, such as cameras with low resolution and small field of view (such as mobile phone cameras), to construct the image database, without relying on professionals and professional equipment, and can reduce the data volume of the image database, thereby reducing the cost of constructing the image database; on the other hand, the present application can shorten the time for constructing the database, which helps to construct the image database with high efficiency. Therefore, the positioning solution based on the poses of the images and the image database in the present application can help to achieve high-efficiency and low-cost positioning.

[0096] Figure 2 It is a schematic block diagram of a system architecture 200 to which the solution of the embodiments of the present application is applied. As Figure 2As shown in the figure, the system architecture 200 may include a hardware abstraction layer data interface 210, a pose acquisition module 220, an application service 230, and a data processing module 240. Among them, the pose acquisition module 220 may include an image retrieval module 210 and a pose calculation module 220. The image retrieval module 210 may further include a feature extraction unit 2211 and an image retrieval unit 2212, and the pose calculation module 220 may further include a local pose calculation unit 2221 and a coordinate conversion unit 2222. The data processing module 240 may include an image database construction module 241 and a video 242.

[0097] It should be understood that Figure 2 The figure shows modules or units of a system architecture applicable to the embodiments of the present application, but these modules or units are only examples. The embodiments of the present application may also include other parts or Figure 2 deformations of each part therein, or may not necessarily include Figure 2 all the modules or units therein.

[0098] In Figure 2 the present application, the pose acquisition module 220 may be an example of the above-mentioned device 100, the image retrieval module 210 may be an example of the above-mentioned image retrieval module 110, and the pose calculation module 222 may be an example of the pose calculation module 120. The present application does not limit this.

[0099] In some embodiments, the pose acquisition module 220 and the image database construction module 241 may exist in the form of a binary software package. In addition, the pose acquisition module 220 may be deployed in the framework layer of the terminal's operating system and provide positioning information for service applications in the application layer through an interface, such as an application programming interface (API).

[0100] It should be noted that the embodiments of the present application may be implemented based on hardware such as a global positioning system (GPS) / magnetometer, gyroscope, wireless fidelity (WiFi) chip, and camera chip built into the terminal. The hardware drivers or data reading and writing modules corresponding to these hardware or chips may interact with data and control through the hardware abstraction layer data interface 210 according to the standard system interface and the upper-layer positioning software service program.

[0101] In some embodiments, the positioning solution provided by the embodiments of the present application may implement pose calculation through the framework of the offline library construction - online positioning stage.

[0102] In the offline database construction stage, the image database construction module 241 can construct an image database, for example, based on the video 242, construct an image database. As an example, the image database construction module 241 can obtain the pose of each image frame in the video 242 in the local coordinate system of the camera through the SLAM algorithm, and fuse the position of the video 242 in the map to obtain a sequence of image frames with poses in the world coordinate system. Further, the image database construction module 241 can also extract the global features (also referred to as global feature information, first information) and local features (also referred to as local feature information, second information) of the image frames in the video 242. Here, the local coordinate system can be a local coordinate system constructed with any one image in the video 242 as the origin of the second local coordinate system.

[0103] Among them, the above video 242 can be referred to as a data source for constructing an image database, for example, it can be a mobile phone video.

[0104] As an example, after the image database construction module 241 establishes an image database, it can output an image feature library. Optionally, the image database can be stored in a memory, and the present application does not limit this.

[0105] Therefore, when constructing an image database in the embodiments of the present application, only the pose of the image in the world coordinate system, as well as the global feature information and local feature information of the image, need to be obtained, and there is no need to obtain 3D point cloud features (such as global 3D point cloud features) that are highly unified with the environment. Therefore, on the one hand, the present application can use low-cost devices, such as cameras with low resolution and small field of view (such as mobile phone cameras), to construct an image database, without relying on professionals and professional equipment, and can reduce the data volume of the image database, thereby reducing the cost of constructing the image database; on the other hand, the present application can shorten the time for constructing the database, which helps to construct the image database with high efficiency. Therefore, the positioning solution based on the pose of the image and the image database in the present application can help to perform high-efficiency and low-cost positioning.

[0106] In the online positioning stage, an image can be obtained in real time, and the pose of the image can be obtained based on the image.

[0107] As an example, the hardware abstraction layer data interface 210 can be used to obtain the image data collected by the camera, such as obtaining the image to be positioned, the first image, etc. Optionally, the hardware abstraction layer data interface 210 can also be used to obtain sensor signals (such as magnetometer, gyroscope parameters, etc.), WiFi chip parameters, which are not limited. As an example, the hardware abstraction layer data interface 210 can obtain the above information from the standard API extracted from the hardware abstraction layer of the terminal's operating system, and the present application does not limit this.

[0108] After obtaining the first image, the Hardware Abstraction Layer Data Interface 210 can send the first image to the Pose Acquisition Module 220. After the Pose Acquisition Module 220 obtains the first image, the Feature Extraction Unit 2211 can extract the global feature information of the first image, and the Image Retrieval Unit 2212 can determine at least one second image that matches the first image in the image database according to the global feature information of the first image and the global feature information of the images in the image database. As an example, the set of second images can be referred to as a similar image set.

[0109] The Local Pose Solution Unit 2221 can determine the poses of the first image and at least one second image in the local coordinate system according to the local feature information of the first image and the local feature information of at least one second image. Here, the local coordinate system can be, for example, the first local coordinate system in the above text. For specific details, please refer to the description of the first local coordinate system and will not be elaborated here. After that, the Coordinate Transformation Unit 2222 can determine the mapping relationship between the local coordinate system and the world coordinate system according to the pose of at least one second image in the world coordinate system and the pose of at least one second image in the local coordinate system, and then, according to this mapping relationship, transform the pose of the first image in the local coordinate system to obtain the pose of the first image in the world coordinate system.

[0110] In some alternative embodiments, after obtaining the pose of the first image, the Pose Acquisition Module 220 can send the pose to the Application Service 230. The Application Service 230 can use this pose to provide a pose positioning service for the user.

[0111] In some embodiments, the Application Service 230 can also initiate a request for a pose positioning service to the Pose Acquisition Module 220. In response to this request, the Pose Acquisition Module 220 can obtain an image from the Hardware Abstraction Layer Data Interface 210 and perform positioning on the image to obtain the pose of the image in the world coordinate system.

[0112] As an example, the Application Service 230 can be various location based services (LBS), AR / VR application services, etc., such as including but not limited to dedicated positioning applications, various e-commerce shopping applications, various social communication applications, various car-hailing applications, online to offline (O2O) door-to-door service applications, museum self-guided tour applications, family anti-lost applications, emergency rescue service applications, audio-visual entertainment applications, game applications, etc. that require precise positioning. Among them, a typical application scenario is, for example, AR navigation, that is, navigation icons are superimposed on the real world captured by the camera of the terminal to provide more indicative navigation information.

[0113] In Figure 2In the system architecture 200 shown, the pose acquisition module 220 can be implemented on a client (such as a terminal) to support various application services on the client. As an example, in response to a positioning service request initiated by the application service 230, which is a positioning software as a system service, the pose acquisition module 220 starts running and uses the positioning method provided in the embodiments of the present application to obtain the image pose in real time on a smart phone.

[0114] As a possible implementation manner, the system architecture 200 can support the client-server mode. As an example, the data processing module 140 can be stored on the server side (such as a server), and the video data acquired by the client (such as a terminal) can be uploaded to the server side through a network connection. After the construction of the image database is completed on the server side, the image database or part of the data in the image database can be downloaded to the client through a network connection.

[0115] As another possible implementation manner, the system architecture 200 can support the client mode, that is, the positioning process can be directly performed offline. At this time, the data processing module 140 can be stored on the client to form a pure offline client mode.

[0116] Figure 3 The schematic flowchart of a positioning method 300 provided in the embodiments of the present application is shown. As an example, the method 300 is described by taking the method 300 including an offline database construction stage and an online positioning stage as an example. Next, the method 300 will be described in combination with Figure 2 the system architecture 200 in.

[0117] It should be understood that Figure 3 the steps or operations of the positioning method are shown, but these steps or operations are only examples, and the embodiments of the present application can also perform other operations or Figure 3 variations of each operation in. In addition, Figure 3 each step in can be executed in a different order from that Figure 3 presented in, and it is possible that not all the operations in Figure 3 need to be executed.

[0118] As Figure 3 shown, the method 300 can include step 301 to step 310. Among them, the offline composition stage can include step 301 to step 304, and the online positioning stage can include step 305 to step 310.

[0119] Step 301, video acquisition. Here, the acquired video can be used as the input of the offline database construction stage.

[0120] As an example, video acquisition can be performed indoors. As a possible implementation, according to the indoor floor plan, the acquisition route can be planned. For example, multiple acquisition routes can be planned, and these acquisition routes can be located on one floor or multiple floors, without limitation.

[0121] Figure 4 An example of an indoor acquisition route is shown. Exemplarily, the user can hold a terminal (such as a mobile phone) and start from the starting point of the route, continuously shoot the environment along the route until reaching the end point of the route to end the shooting.

[0122] While acquiring the video, the terminal can also obtain the position of the route on the map. For example, when starting to shoot the video, the user can manually activate the positioning module in the terminal to obtain the position of the starting point of the route on the map. When reaching the end point of the route, the user can manually activate the positioning module to obtain the position of the end point of the route on the map. Or, when starting to shoot the video, the positioning module in the terminal can automatically obtain the position of the starting point of the route on the map. When reaching the end point of the route, the positioning module can automatically obtain the position of the end point of the route on the map. Here, the position of the starting point of the route on the map can be referred to as the position of the first frame image in the acquired video on the map, and the position of the end point of the route on the map can be referred to as the position of the last frame image in the acquired video on the map.

[0123] In the embodiments of the present application, the position information of the starting point and / or the end point on the map can be used for auxiliary calculation, for example, for calculating the pose of the image in the video in the world coordinate system. Exemplarily, the positioning module can be at least one of a GPS / magnetometer, a gyroscope, a WiFi chip, etc.

[0124] It should be noted that here an indoor scenario is taken as an example for illustration, and this solution is equally applicable to an outdoor scenario. That is, in step 301, video acquisition can also be performed outdoors. For example, an outdoor acquisition route can be planned, and the present application makes no limitation thereto.

[0125] Step 302, frame extraction of video images.

[0126] As an example, during the process of using a terminal to shoot and obtain video images, the terminal can also automatically perform frame extraction processing on the video images to obtain multiple frames of images. Figure 5 An example of multiple frames of images extracted from a video is shown.

[0127] In some embodiments, when the client-server mode is adopted, the terminal can upload the acquired video to the server. When the client mode is adopted, the terminal can store the video and process the video.

[0128] Step 303, constructing an image database.

[0129] As an example, step 303 can be executed by the image database construction module 241 on the data processing module 240 described above. Refer to Figure 2 and step 303 can further include step 3031 and step 3032. Figure 3

[0130] Step 3031, SLAM path tracking.

[0131] As an example, for each frame of video image, the SLAM algorithm can be used to obtain the pose of the image in the local coordinate system of the camera (i.e., the terminal). Here, the local coordinate system is the relative coordinate system of the camera during the construction of the image database.

[0132] In some other embodiments, the AR engine (AR Engine) algorithm can also be used to obtain the pose of each frame of image in the local coordinate system of the camera, and the present application does not limit this.

[0133] After obtaining the pose of each frame of image in the local coordinate system of the camera, image coordinate registration can be performed in combination with the position of the image in the map, and the pose of the image in the camera coordinate system can be converted to the world coordinate system. For example, the position of the starting point of the video acquisition route in the map (i.e., the position of the first frame of image in the map in the video), and / or the position of the end point of the video acquisition route in the map (i.e., the position of the last frame of image in the map in the video) can be used for image coordinate registration.

[0134] Figure 6 shows an example of the visualization result of the image database. As Figure 6 shown, the acquisition route can be selected on the left side of the interface. When a certain acquisition route is selected, the acquisition route can be presented on the right side of the interface. For example, the position of the acquisition route in the map can be presented. As an example, Figure 6 shows 3 acquisition routes collected on the 4th floor of Building A. Among them, as an example, the name of acquisition route 1 can be 0a1c55c4-9303-46a5-b785-3ad77d80d809_20200415-090323, the name of acquisition route 2 can be 0a1c55c4-9303-46a5-b785-3ad77d80d809_20200415-091659, and the name of acquisition route 3 can be 0a1c55c4-9303-46a5-b785-3ad77d80d809_20200415-093745.

[0135] Optionally, the folder where the image data corresponding to the acquisition route is located can also be opened through the visualization display interface. For example, it can be through Figure 6 ​Click the "Open Folder" button in the upper right corner to open the folder where the image data corresponding to the selected acquisition route is located.

[0136] Exemplarily, after selecting the trajectory tree - Building A - 4th Floor - Trajectory - 0ale55c4 - 9303 - 46a5 - b785 - 3ad77d80d809_202004150 - 090323 trajectory and choosing to open the folder, Table 1 below can be displayed. Among them, pos_x, pos_y, pos_z represent position information; quat_x, quat_y, quat_z, quat_w represent the direction represented by quaternions, which can be converted into attitude information, such as yaw, roll, pitch.

[0137] Table 1

[0138] FileName pos_x pos_y pos_z quat_x quat_y quat_z quat_w 000001 1.8536 0.0189 76.6551 0.4322 -0.5769 -0.5329 0.4430 000002 2.8263 0.0043 76.6297 0.4190 -0.5702 -0.5573 0.4341 000003 3.3308 0.0002 76.5465 0.4543 -0.5904 -0.53556 0.3975 …

[0139] Optionally, in the visualization result of the image database, the image frames, global feature information, local feature information, etc. obtained on each acquisition route can also be displayed. This application does not limit this.

[0140] Step 3032, global / local feature extraction.

[0141] Specifically, global feature information and local feature information can be extracted from each frame of the image collected by the terminal. Among them, the global feature information is image retrieval feature information, such as NetVLAD, BoVW, etc., which is not limited. The local feature information is, for example, scale - invariant feature transform (SIFT), ORB, SUPERPOINT, D2Net, etc., which is not limited.

[0142] Step 304, output the image database.

[0143] As an example, the image database can include the pose, global features, and local features of each frame of the image in the world coordinate system in the video collected along the planned path. As an example, the images in the image database can be represented as (R, T, F l , F g ), where, R represents position information, T represents angle information, F l represents global features, F g represents local features.

[0144] Step 305, image acquisition. Here, the collected image is the first image, which can also be called the image to be located or others. This application does not limit this.

[0145] As an example, the user can use a terminal to capture an image. Accordingly, the hardware abstraction layer data interface 210 in the smart phone can obtain the image data collected by the camera and send the image data to the pose acquisition module 220.

[0146] It should be noted that the terminal in step 305 and the terminal in step 301 can be the same terminal or different terminals, and this application does not make any limitation in this regard.

[0147] Optionally, in step 306, the initial position is obtained.

[0148] As an example, when the image is captured in step 305, the hardware abstraction layer data interface 210 can also obtain the positioning information collected by the WiFi chip or the GPS module and send the positioning information to the pose acquisition module 220. This positioning information can be referred to as the initial position of the first image.

[0149] It should be noted that the position provided by the WiFi chip or the GPS module is a "rough positioning" position with relatively low accuracy, and the error can be, for example, 3 to 10 meters. This position can be used as the initial position of the first image.

[0150] In addition, the hardware abstraction layer data interface 210 can also send the information of the gyroscope and the magnetometer to the pose acquisition module 220, so that the pose acquisition module 220 can jointly determine the initial position of the first image according to the positioning information collected by the WiFi chip or the GPS module and the acquisition information of the gyroscope and the magnetometer.

[0151] Optionally, in step 307, the floor is identified. As an example, when the initial position of the first image obtained in step 306 includes height information, the floor corresponding to the first image can be identified based on this initial position.

[0152] Step 308, image retrieval.

[0153] As an example, Figure 2 the image retrieval module 221 in can determine at least one second image matching the first image in the image database based on the global feature information of the first image obtained in step 305 and the global feature information of the images in the image database. In the embodiments of the present application, determining at least one second image matching the first image in the image data can also be described as retrieving at least one second image matching the first image in the image database.

[0154] As an example, the set of the at least one second image can be referred to as a similar image set.

[0155] In some possible implementations, the image retrieval module 221 may perform image retrieval based on the first image and the initial position obtained in step 306.

[0156] In some possible implementations, the image retrieval module 221 may perform image retrieval based on the first image, the initial position obtained in step 306, and the floor obtained in step 307.

[0157] Continue to refer to Figure 3 , as an example, step 308 may include the following steps 3081 to 3083. As an example, step 3081 may be executed by the feature extraction unit 2211, and steps 3082 and 3083 may be executed by the image retrieval unit 2212.

[0158] Step 3081, global feature extraction.

[0159] As an example, a deep learning algorithm may be used to extract the global feature information of the first image. For example, a NetVLAD layer may be added after the convolutional neural network framework to implement the extraction of the global descriptor (an example of global feature information). Among them, the global feature of the image is the overall attribute of the image, and may include, for example, color features, texture features, and shape features, etc., without limitation.

[0160] Step 3082, preliminary image retrieval.

[0161] As a possible implementation, at least one second image matching the first image may be retrieved in the image database according to the global feature information obtained in step 3081 to obtain a set of similar images.

[0162] As another possible implementation, at least one image may be determined in the image database according to the initial position obtained in step 306, where the distance between the position of each image in the at least one image and the initial position of the first image is less than a preset threshold. Then, at least one second image matching the first image is retrieved from the at least one image to obtain a set of similar images.

[0163] As a specific example, a region (for example, it can be called a buffer) may be constructed according to the initial position of the first image. For example, a circular buffer may be constructed with the initial position of the first image as the center and a radius of 5 meters (or 10 meters, or 30 meters, without limitation). Then, at least one second image matching the first image is retrieved from the images in the image database located in the buffer to obtain a set of similar images. As an example, the set of images located in the buffer may be called an image candidate set.

[0164] Therefore, in the embodiment of the present application, at least one image is determined in the image database according to the first position of the first image, and then the first information of the first image is matched with the first information of the at least one image, and at least one second image matching the first image is obtained from the at least one image. In this way, it is not necessary to match the first information of the first image with the first information of all images in the image database, thereby reducing the computing amount of the terminal, reducing the matching time, and being beneficial to improving the efficiency of obtaining the second image.

[0165] As a possible implementation manner, the similarity between the global feature information of the first image and the global feature information of the images in the image database or the image candidate set can be calculated, and the second image can be determined according to the similarity.

[0166] Exemplarily, the similarity can be calculated by the following formula (1):

[0167]

[0168] where similarity(X, Y) represents the similarity between X and Y, X represents the global feature information of the first image, Y represents the global feature information of the image to be matched in the feature library, x i represents each component in the feature encoding of X, y i represents each component in the feature encoding of Y, and n represents the total number of components in a feature encoding, and n is a positive integer.

[0169] After obtaining the similarities between each frame of image and the first image, the calculated multiple similarities can be sorted to obtain the top m images with the highest similarities as the second images, where m is an integer greater than 1. As an example, m can be 20. Or, in some other implementation manners, the images with similarities greater than a preset threshold to the first image can be used as the second images.

[0170] Figure 7 Shows an example of retrieving the second image. As Figure 7 shown, "×" represents the global feature information of the image of the cluster center in the one cluster region (which can be represented as C k VLAD ), and the first image is included in the cluster region. All images in the cluster region can be retrieved to obtain the second image. Among them, "★" represents the first image (which can be represented as The global feature information of "□" represents the global feature information of the first image. The global feature information of "○" represents the global feature information of the image closest to the first image in the clustering region (which can be called the closest image). The global feature information of "●" represents the global feature information of the image that is the second closest to the first image in the buffer (which can be called the second closest image). When the ratio of the distance between the closest image and the first image to the distance between the second closest image and the first image is greater than the threshold, it indicates that the closest image matches the first image. Optionally, according to this method, 20 images that can match the first image can be sequentially determined in the buffer as the second images.

[0171] Optionally, in step 3083, pose-aware image refinement.

[0172] As an example, when the number of the obtained second images is multiple, the multiple second images can be further refined. For example, the multiple second images can be refined based on the poses of the multiple second images in the world coordinate system.

[0173] As a possible implementation, outlier images in the second images (i.e., the similar image set) can be deleted based on the poses of the second images in the world coordinate system. Here, the outlier images in the second images refer to the images whose distances from the clustering center of the multiple second images are greater than a preset threshold, where the clustering center is determined according to the positions of the multiple second images in the world coordinate system.

[0174] Exemplarily, the positions of each second image in the world coordinate system can be obtained from the image database, and according to these positions, the clustering center of the multiple second images can be determined. The images whose distances from the clustering center are greater than the preset threshold among the multiple second images are determined as outlier images.

[0175] It should be noted that the process of determining the images whose distances from the clustering center are greater than the second threshold (i.e., determining outlier images) is the process of judging whether there are images whose distances from the clustering center are greater than the second threshold among the multiple second images (i.e., whether there are outlier images). As a possible judgment result, there may be no outlier images among the multiple second images. As another possible judgment result, there may be at least one outlier image among the multiple second images.

[0176] In some embodiments, when there are no outlier images among the multiple second images, the operation of deleting images may not be performed.

[0177] Since the images at different positions may have very similar textures, errors may occur during the similarity calculation, which may lead to outlier images in the similar image set. Therefore, in this application, by removing these outlier images, the errors in the images in the similar image set can be eliminated, which helps to obtain a more accurate similar image set. Here, obtaining a more accurate similar image set can help to make the pose of the subsequent obtained second image in the first local coordinate system more accurate, which in turn helps to improve the accuracy of the pose of the first image in the world coordinate system.

[0178] Figure 8 A specific example of region clustering for multiple frames of second images is shown. Among them, through clustering calculation of multiple frames of second images in the similar image set, it can be obtained that most of the second images are clustered in regions 1 and 2, and only a small part of the second images are scattered in regions 3 and 4. At this time, the second images corresponding to the clustering points in regions 2 and 3 can be deleted, and the second images corresponding to the clustering points in regions 1 and 2 can be used as the second images in the updated similar image set.

[0179] As another possible implementation, redundant images in the second images (i.e., the similar image set) can be deleted based on the pose of the second images in the world coordinate system. As an example, the pose of each second image can be obtained from the image database, and then according to this pose, the images with an angle less than a preset angle threshold and / or a distance less than a preset distance threshold are determined among the multiple frames of second images as the redundant images in the multiple second images. Wherein, the angle is the difference in the pose of at least two images in the world coordinate system, and the distance is the difference in the position of at least two images in the world coordinate system.

[0180] As an example, the angle can be the difference in the pitch, yaw or roll angle of the pose of at least two images in the world coordinate system.

[0181] It should be noted that the process of determining the images with an angle less than a preset angle threshold and / or a distance less than a preset distance threshold (i.e., determining redundant images) is the process of judging whether there are images with an angle less than a preset angle threshold and / or a distance less than a preset distance threshold in the multiple second images (i.e., whether there are redundant images). As a possible judgment result, there may be no redundant images in the multiple second images. As another possible judgment result, there may be at least one frame of redundant image in the multiple second images.

[0182] In some embodiments, when there are no redundant images in the multiple second images, the operation of deleting images may not be performed.

[0183] Since images with an angle less than a preset threshold and / or a distance less than a preset threshold have a high degree of overlap in spatial distribution, it can be recognized that there are redundant images in the similar image set due to the high degree of overlap in spatial distribution of images with an angle less than a preset threshold and / or a distance less than a preset threshold. Therefore, in the embodiments of the present application, by deleting images with an angle less than a preset value and / or a distance less than a preset value in the similar image set, the overlap of the second images remaining in the similar image set can be made moderate, and the surrounding space can be evenly covered, which further helps to obtain a more refined similar image set. Here, obtaining a more refined similar image set can help reduce the computational amount of the poses of the second images obtained subsequently in the first local coordinate system, and further help improve the efficiency of obtaining the pose of the first image in the world coordinate system.

[0184] In some embodiments, after removing the outlier images from the multiple second images, redundant images can be further removed. Figure 9 An example of further removing redundant images after removing outlier images from multiple second images is shown. Among them, the positions where the circles are located can represent the positions of the second images, and the arrows can represent the poses of the second images. As a specific example, the angle threshold can be set to ±15°, and the distance threshold can be set to 0.5 m. Correspondingly, when the difference (i.e., the angle) between the poses of any two or more of at least two frames of second images is less than 15° and / or the position distance is less than 0.5 m, these two or more frames of second images can be considered redundant images (for example, it can correspond to Figure 9 the images corresponding to the positions of the white circles in). Correspondingly, through angle and position filtering, the redundant images in the second images will be removed. At this time, the angles between the second images remaining in the multiple second images are greater than or equal to 15° and / or the position distances are greater than or equal to 0.5 m, which can evenly cover the surrounding space and have a moderate degree of overlap.

[0185] As a specific example, after steps 3081 and 3082, the 20 frames of second images in the original similar image set can be reduced to 8 frames of second images. Exemplarily, these 8 images can be respectively represented as (R1, T1), (R2, T2), (R3, T3), (R4, T4), (R5, T5), (R6, T6), (R7, T7), (R8, T8). At this time, the first image can be represented as (R q , T q ).

[0186] Continue to refer to Figure 3 , after the image retrieval module 221 obtains the similar image set, the above-mentioned first image and at least one second image can be sent to the pose calculation module 222.

[0187] Step 309, camera pose solution.

[0188] Exemplarily, after the pose calculation module 222 obtains the first image and at least one second image, camera pose solution can be performed. Refer to Figure 3 , step 309 may further include steps 3091 and 3092.

[0189] Step 3091, local pose calculation.

[0190] As an example, the local pose calculation unit 2221 in the image retrieval module 221 may perform local pose calculation based on the first image and at least one second image.

[0191] Specifically, the local pose calculation unit 2221 may determine the poses of the first image and at least one second image in the local coordinate system of the first terminal according to the local feature information of the first image and the local feature information of at least one second image. Exemplarily, the local pose calculation unit 2221 may use the first image and at least one second image as a local image set, and obtain the relative poses of each image in the local image set in the local coordinate system according to the local feature information of each image in the local image set. Here, the local coordinate system may be a relative coordinate system of the camera constructed with any frame image in the local image set as the origin of the local coordinate system. Among them, the local feature of the image is a local expression of the image feature, which can reflect the local characteristics of the image.

[0192] As a possible implementation manner, the local pose calculation unit 2221 may match the images in the local image set according to the local feature information in the local image set, obtain the poses of the images in the local image set in the local coordinate system, and the point cloud information of each image mapped into the three-dimensional space. Here, the poses of the images in the local image set in the local coordinate system and the point cloud information mapped into the three-dimensional space may be referred to as the local structure information of the local image set. In some embodiments, the process of obtaining the local structure information of the local image set may be referred to as the recovery process of the local structure information of each image in the local image set.

[0193] In the embodiments of the present application, by recovering the local structure information of the images in the local image set, on the one hand, the construction of the global 3D point cloud feature library can be avoided, saving the library construction cost, and on the other hand, the error caused by aligning a large number of images using the global feature library can be avoided, so that more accurate poses of the images in the local image set can be obtained.

[0194] Step 3092, coordinate transformation.

[0195] Exemplarily, the coordinate transformation unit 2222 may obtain the pose of the second image in the world coordinate system in the similar image set, and determine the mapping relationship (i.e., transformation relationship) between the local coordinate system and the world coordinate system according to the pose of the second image in the local coordinate system obtained in step 3091 above. Then, according to this mapping relationship, the pose of the first image in the local coordinate system is coordinate-transformed to obtain the pose of the first image in the world coordinate system.

[0196] As a possible implementation, the coordinate transformation unit 2222 may obtain the pose of the second image in the world coordinate system from the image database output in step 304.

[0197] Figure 10 Fig. shows a specific example of pose calculation. Among them, in Fig. (a), it represents the local image set input during pose calculation, which includes n frames of second images in the similar image set, denoted as (R1, T1)…(R n , T n ), and the first image, denoted as (R q , T q ). As an example, in Figure 10 , n is taken as 8 for description, but the present application is not limited thereto.

[0198] Meanwhile, during pose calculation, it is also necessary to input the pose and local feature information of the n frames of second images in the similar image set in the world coordinate system. As an example, the pose and local feature information of the n frames of second images in the world coordinate system can be obtained from the image database. Fig. (e) shows an example of the image trajectory distribution of 8 frames of second images in the world coordinate system. Among them, the 8 boxes constituting the image trajectory respectively correspond to the poses of the 8 frames of second images.

[0199] As an example, for the input local image set, the SFM algorithm can be used to estimate the poses of the images in the local image set in the local coordinate system.

[0200] First, the image feature points of the first image in the local image set can be extracted to obtain the local feature information of the first image. Then, the local feature information of the first image and the local feature information of the second image are matched. Figure 9 Fig. (b) in

[0201] Then, incremental camera parameter solution can be performed on the feature points matched between the images in the local image set. As an example, Figure 9Figure (c) shows an example of the incremental camera parameters corresponding to three frames of images in the local image set. Among them, the upper image in Figure (c) contains 7 feature points. The left image at the lower part of Figure (c) is a frame of image in the local image set, containing 4 feature points (for example, the 4 left feature points among the above 7 feature points). The incremental parameter of this left image is P0 = K[I|0], where K represents the camera memory, I represents the initial direction (0, 0, 0, 1), and 0 represents the initial position (0, 0, 0). That is to say, a local coordinate system of this local image set can be constructed with this left image as the origin. The middle image at the lower part of Figure (c) is a frame of image in the local image set, containing these 4 feature points (for example, the 4 front feature points among the above 7 feature points). The incremental parameter of this middle image is P1 = K[R1|t1], where R1 represents the position of this middle image, and t1 represents the angle of this middle image. The right image at the lower part of Figure (c) is a frame of image in the local image set, containing these 4 feature points (for example, the 4 right feature points among the above 7 feature points). The incremental parameter of this right image is P i = K[R i |t i , where R i represents the position of this middle image, and t i represents the angle of this middle image.

[0202] After that, according to the obtained incremental camera parameters, the poses and 3D point clouds of each image in the local image set in the local coordinate system can be obtained. Figure (d) shows a specific example of the image trajectory distribution and 3D point cloud of 9 frames of images in the local coordinate system. Among them, the 9 boxes constituting the image trajectory respectively correspond to the poses of these 9 frames of images. Among them, the black box represents the pose of the first image in the local coordinate system, and the white boxes represent the poses of 8 frames of second images in the local coordinate system.

[0203] After that, the poses (or image trajectory distributions) of each image in the local image set in the local coordinate system, as well as the poses (or image trajectory distributions) of the second images in the similar image set in the world coordinate system, can be input into the coordinate conversion unit 2222, and the coordinate conversion unit 2222 performs coordinate conversion. Figure (f) shows a specific example of the coordinate conversion. Among them, according to the poses of the above 8 frames of second images in the world coordinate system and the poses of these 8 frames of second images in the local coordinate system, the mapping relationship between the local coordinate system and the world coordinate system can be obtained. As an example, this mapping relationship can be a similarity transformation matrix (R i , t i , α i ) from the local coordinate system to the world coordinate system. Further, by multiplying the pose of the first image in the local coordinate system by this similarity transformation matrix (R i, t i , α i ), the pose of the first image in the world coordinate system can be obtained.

[0204] 310, output the positioning result.

[0205] Specifically, after obtaining the pose of the first image in the world coordinate system, the pose can be output, that is, the positioning result of the first image is output, thereby completing the positioning of the first image.

[0206] Figure 11 Shows a specific example of the real-time positioning result. Among them, the real scene in the figure is the location where the terminal is currently located. The small circle in the lower map represents the visualization result of the current location in the two-dimensional map, and the arrow represents the direction information.

[0207] Therefore, in the embodiment of the present application, by determining at least one second image that matches the first image captured by the first terminal in the image database, and the poses of the first image and the at least one second image in the local coordinate system of the first terminal, and according to the pose of the first image in the first local coordinate system, the poses of the at least one second image in the first local coordinate system, and the poses of the at least one second image in the world coordinate system, the pose of the first image in the world coordinate system is output, realizing the positioning of the first image. Therefore, the embodiment of the present application can obtain the relationship between the local coordinate system and the world coordinate system based on the poses of a part of the images in the local coordinate system and in the world coordinate system, so as to locate the image to be positioned (such as the first image) in the world coordinate system, without collecting 3D point clouds for each image.

[0208] Based on the above positioning scheme, the embodiment of the present application can make full use of existing terminal devices, such as user-level receiver devices (such as mobile phones) or navigation-level receiver devices (such as vehicles), to collect images, extract image features, and perform robust image retrieval and matching to achieve positioning (such as 6DoF positioning). During the image retrieval process, the embodiment of the present application can, on the basis of image retrieval according to global feature information, further consider the spatial distribution of the retrieved images, jointly construct a double-threshold image refinement algorithm with poses to delete outlier images and redundant images (which can also be called secondary retrieval) to obtain a refined similar image set. For the local image set, the embodiment of the present application can execute the SFM algorithm, recover the relative poses of the images in the local image set in the local coordinate system based on local structure information, and combine the poses of the images in the world coordinate system in the pose feature library to convert the relative pose of the image to be positioned into the world coordinate system to solve the camera pose.

[0209] The embodiments of the present application can solve the problem that it is difficult to construct a point cloud database for LBS / AR / VR application services and large-scale promotion cannot be achieved. Specifically, the embodiments of the present application propose a solution for indoor and outdoor positioning that does not rely on a 3D point cloud feature library. Compared with the positioning solution that requires the construction of a point cloud database, the embodiments of the present application can have the advantages of high positioning accuracy and low cost, breaking the high-cost barrier of traditional global 3D point cloud feature library construction. In addition, the embodiments of the present application do not need to align each image to construct a global 3D point cloud database, but only need to obtain the pose of the image, and extract the global and local features of the image to construct an image pose library. In the online positioning stage of the embodiments of the present application, first, the image to be positioned is retrieved in the database, then the local structure information of the retrieved image and the image to be positioned is restored, and the relative pose of the image to be positioned and the retrieved image in the relative coordinate system is obtained. Finally, the pose feature library is combined to realize the conversion of the pose from the relative coordinate system to the world coordinate system, and the solution of the camera pose is realized.

[0210] According to the above embodiments, an example of the positioning accuracy of the present application in Area A (such as Nanjing) and Area B (such as Xi'an) is shown in Table 2 below. Among them, the error of position positioning is at the centimeter level, and the proportion of the angle positioning accuracy within 3° is 48.1% to 67.7%, and the proportion of the angle positioning accuracy within 10% is more than 90%.

[0211] Table 2

[0212] <1° <3° <10° X(m) Y(m) Z(m) Location A 10% 48.1% 91.9% 0.17 0.01 0.13 Location B 7.5% 67.7% 93.7% -0.01 0.004 -0.019

[0213] The reason for achieving the above technical effects is that the image pose library (a model), the image retrieval algorithm based on pose perception, and the camera pose solution algorithm based on local structure information (two algorithms) proposed in the present application have a solid theoretical basis and practical basis. The theoretical basis is that the embodiments of the present application can search the evenly distributed image data set in the image database based on the global feature information and position of the image to be positioned, provide stable input data for SMF to restore local structure information, ensure that as many images as possible participate in the reconstruction, and at the same time do not generate redundant information. Combining the global prior information provided by the image pose library, the transformation relationship (i.e., the mapping relationship) from the local coordinate system to the global coordinate system can be obtained, and accurate pose calculation can be realized. The practical basis is that the terminal (such as a mobile phone), as an easily accessible image acquisition device, combined with the easily operable acquisition method of collecting videos, can greatly improve the efficiency of data acquisition, simplify the data acquisition process and labor costs.

[0214] Figure 12FIG. 0 shows a schematic flowchart of a positioning method 1200 provided by an embodiment of the present application. Among them, the method 1200 can be applied to a terminal or the cloud. As an example, the method 1200 can be executed by Figure 1 the positioning device 100 therein, or by Figure 2 the system 200 therein. As shown in Figure 12 , the method 1200 includes step 1210 to step 1240.

[0215] 1210, obtain a first image captured by a first terminal.

[0216] 1220, determine at least one second image matching the first image in the image database according to the first information of the first image and the first information of the images in the image database, where the first information is used to indicate the global features of the image, and the image database includes the first information of multiple frames of images, the second information of the multiple frames of images, and the poses of the multiple frames of images in the world coordinate system, where the second information is used to indicate the local features of the image.

[0217] 1230, determine the poses of the first image and the at least one second image in the first local coordinate system of the first terminal according to the second information of the first image and the second information of the at least one second image.

[0218] 1240, output the pose of the first image in the world coordinate system according to the pose of the first image in the first local coordinate system, the poses of the at least one second image in the first local coordinate system, and the poses of the at least one second image in the world coordinate system.

[0219] Therefore, in the embodiment of the present application, by determining at least one second image matching the first image captured by the first terminal in the image database, and the poses of the first image and the at least one second image in the local coordinate system of the first terminal, and according to the pose of the first image in the first local coordinate system, the poses of the at least one second image in the first local coordinate system, and the poses of the at least one second image in the world coordinate system, output the pose of the first image in the world coordinate system, so as to realize the positioning of the first image. Therefore, the embodiment of the present application can obtain the relationship between the local coordinate system and the world coordinate system based on the poses of a part of the images in the local coordinate system and the poses in the world coordinate system, so as to locate the image to be positioned (such as the first image) in the world coordinate system, without collecting 3D point clouds for each image.

[0220] Compared with the solution of constructing a 3D point cloud feature library of the environment (including global 3D point cloud features) in the existing solution and calculating the pose based on feature point matching, the embodiment of the present application does not utilize the 3D point cloud features of the image, but uses the pose of the image for positioning. Since constructing a 3D point cloud feature library of the environment requires relying on professionals and acquisition equipment, which is time-consuming, laborious and costly, and the present application does not need to obtain 3D point cloud features. Therefore, on the one hand, the present application can use low-cost equipment, such as a camera with low resolution and small field of view (such as a mobile phone camera) to construct an image database, without relying on professionals and professional equipment, and can reduce the data volume of the image database, thereby reducing the cost of constructing the image database; on the other hand, the present application can shorten the time for constructing the database, which helps to construct the image database with high efficiency. Therefore, the positioning solution based on the pose of the image and the image database in the present application can help to perform high-efficiency and low-cost positioning.

[0221] In some possible implementation manners, method 1200 further includes establishing an image database.

[0222] In some possible implementation manners, it is possible to obtain a video captured by a second terminal, and obtain the poses of multiple frames of images in the video in a second local coordinate system of the second terminal. Then, according to the third information and the poses of multiple frames of images in the video in the second local coordinate system, the poses of multiple frames of images in the video in the world coordinate system can be determined. Wherein, the third information is used to indicate the positions of at least some images in the video in the map, and the map is associated with the world coordinate system. After that, the first information and the second information of multiple frames of images in the video can be obtained. In this way, establishing the image database can be achieved.

[0223] In some possible implementation manners, method 1200 may further include obtaining a first position of the first image, and determining at least one image in the image database according to the first position of the first image. Wherein, the distance between the position of each image in the at least one image and the first position of the first image is less than a first threshold, and the first position is from the GPS module or WIFI module of the first terminal.

[0224] Wherein, as a specific implementation manner of matching the first information of the first image with the first information of the images in the image database to determine at least one second image that matches the first image in the image database, the first information of the first image can be matched with the first information of the at least one image, and at least one second image that matches the first image can be obtained from the at least one image.

[0225] In some possible implementations, multiple third images matching the first image may be determined in the image database according to the first information of the first image and the first information of the images in the image database. Then, images among the multiple third images that have a distance greater than a second threshold from the cluster center of the multiple third images are deleted to obtain the at least one second image. The cluster center is determined according to the positions of the multiple third images in the world coordinate system.

[0226] Here, an image that has a distance greater than the second threshold from the cluster center of the multiple third images may be referred to as an outlier image. That is to say, the third images may include outlier images and the at least one second image. In some possible descriptions, the third images may be described as multiple second images in the image database that match the first image and for which outlier images have not been deleted.

[0227] In some possible implementations, multiple fourth images matching the first image may be determined in the image database according to the first information of the first image and the first information of the images in the image database. Then, images among the multiple fourth images that have an angle less than a third threshold and / or a distance less than a fourth threshold are deleted to obtain the at least one second image. The angle is the difference in the poses of at least two images in the world coordinate system, and the distance is the difference in the positions of at least two images in the world coordinate system. As an example, the angle may be the difference in the pitch, yaw, or roll angles of the poses of at least two images in the world coordinate system.

[0228] Here, an image that has an angle less than the third threshold and / or a distance less than the fourth threshold may be referred to as a redundant image. That is to say, the fourth images may include redundant images and the at least one second image. In some possible descriptions, the fourth images may be described as multiple second images in the image database that match the first image and for which redundant images have not been deleted.

[0229] Specifically, Figure 12 All relevant content of each step involved in the positioning method 1200 shown may be referred to the relevant functions of each module in the above Figure 1 or Figure 2 in, or the description of the positioning method 300 shown in Figure 3 is not elaborated herein.

[0230] As described above in connection with Figures 1 to 12 the positioning method provided in the embodiments of the present application has been described in detail. Next, the positioning device in the embodiments of the present application will be introduced in connection with Figure 13 and Figure 14 It should be understood that Figure 13 and Figure 14The positioning device in can execute each step in the positioning method in the embodiments of the present application. To avoid repetition, the following will appropriately omit the repeated descriptions when introducing the Figure 13 and Figure 14 positioning devices in.

[0231] Figure 13 FIG. is a schematic block diagram of the positioning device 1300 in the embodiments of the present application. The device 1300 includes an acquisition unit 1310, a processing unit 1320, and an output unit 1330.

[0232] Specifically, when the positioning device 1300 executes the positioning method, the acquisition unit 1310 is configured to acquire a first image captured by a first terminal.

[0233] The processing unit 1320 is configured to determine at least one second image matching the first image in the image database according to the first information of the first image and the first information of the images in the image database, where the first information is used to indicate the global features of the image, and the image database includes the first information of multiple frames of images, the second information of the multiple frames of images, and the poses of the multiple frames of images in the world coordinate system, where the second information is used to indicate the local features of the image.

[0234] The processing unit 1320 is further configured to determine the poses of the first image and the at least one second image in the first local coordinate system of the first terminal according to the second information of the first image and the second information of the at least one second image.

[0235] The output unit 1330 is configured to output the pose of the first image in the world coordinate system according to the pose of the first image in the first local coordinate system, the poses of the at least one second image in the first local coordinate system, and the poses of the at least one second image in the world coordinate system.

[0236] In some possible implementation manners, the device 1300 further includes a building unit configured to build the image database.

[0237] In some possible implementation manners, the building unit is specifically configured to:

[0238] Acquire a video captured by a second terminal; acquire the poses of multiple frames of images in the video in the second local coordinate system of the second terminal; determine the poses of the multiple frames of images in the video in the world coordinate system according to the third information and the poses of the multiple frames of images in the video in the second local coordinate system, where the third information is used to indicate the positions of at least some images in the video in the map, and the map is associated with the world coordinate system; acquire the first information and the second information of the multiple frames of images in the video.

[0239] In some possible implementations, the obtaining unit 1310 is further configured to obtain a first position of the first image, where the first position is from the GPS module or the WiFi module of the first terminal.

[0240] Wherein, the obtaining unit can receive data sent by the GPS module or the WiFi module, such as the above-mentioned first position.

[0241] Optionally, the obtaining unit may also send a request message to the GPS module or the WiFi module, and the request message is used to request the GPS module or the WiFi module to send the first position of the first image collected by it to the obtaining unit. In response to the request message, the GPS module or the WiFi module may send the first position to the obtaining unit.

[0242] The processing unit 1320 is further configured to determine at least one image in the image database according to the first position of the first image, and the distance between the position of each image in the at least one image and the first position of the first image is less than a first threshold.

[0243] The processing unit 1320 is further configured to match the first information of the first image with the first information of the at least one image, and obtain at least one second image that matches the first image from the at least one image.

[0244] In some possible implementations, the processing unit 1320 is specifically configured to:

[0245] Determine a plurality of third images that match the first image in the image database according to the first information of the first image and the first information of the images in the image database; delete the images in the plurality of third images whose distance from the cluster center of the plurality of third images is greater than a second threshold, so as to obtain the at least one second image, where the cluster center is determined according to the positions of the plurality of third images in the world coordinate system.

[0246] In some possible implementations, the processing unit 1320 is specifically configured to:

[0247] Determine a plurality of fourth images that match the first image in the image database according to the first information of the first image and the first information of the images in the image database; delete the images in the plurality of fourth images whose angle is less than a third threshold and / or whose distance is less than a fourth threshold, so as to obtain the at least one second image. Wherein, the angle is the difference in the poses of at least two images in the world coordinate system, and the distance is the difference in the positions of the poses of at least two images in the world coordinate system.

[0248] Specifically,Figure 13 All relevant content of each unit related to the positioning device 1300 shown (such as implementation examples or technical effects) can be referred to in the above text Figure 1 or Figure 2 the relevant functions of each module in Figure 3 or the relevant description of the positioning method 300 shown in

[0249] Figure 14 is a schematic structural diagram of the positioning device 1400 according to an embodiment of the present application. As Figure 14 shown, the device 1400 includes a communication module 1410, a sensor 1420, a user input module 1430, an output module 1440, a processor 1450, an audio-video input module 1460, a memory 1470, and a power supply 1480.

[0250] The communication module 1410 may include at least one module capable of enabling the device 1400 to communicate with other devices (such as other computer systems or mobile terminals). For example, the communication module 1410 may include a wired network interface, a broadcast receiving module, a mobile communication module, a wireless Internet module, a local area communication module, and a location (or positioning) information module, etc., one or more of them. There are various implementations of these various modules in the prior art, and the present application will not describe them one by one.

[0251] The sensor 1420 can sense the current state of the device 1400, such as position, whether there is contact with the user, direction, and acceleration / deceleration, etc. Exemplarily, the sensor 1420 can send the sensed current state of the device 1400 to the GPS module or the WiFi module.

[0252] The user input module 1430 is used to receive input digital information, character information, or contact touch operations / non-contact gestures, and receive signal inputs related to user settings and function control of the device, etc. The user input module 1430 includes a touch panel and / or other input devices.

[0253] The output module 1440 includes a display panel for displaying information input by the user, information provided to the user, or various menu interfaces of the system, etc. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED), etc. In some other embodiments, the touch panel can cover the display panel to form a touch display screen. In addition, the output module 1440 may further include an audio output module, an alarm, and a tactile module, etc.

[0254] Exemplarily, the output module 1440 may implement the related functions of the output unit 1330 in the device 1300. For example, it may be used to output the pose of the first image in the world coordinate system. As a specific example, the visualization result of the current position of the terminal in the map may be displayed on the display panel of the output module 1440.

[0255] The audio-video input module 1460 is used to input an audio signal or a video signal. The audio-video input module 1460 may include a camera and a microphone. Exemplarily, the camera may be used to capture the first image, and the present application does not limit this.

[0256] The power supply 1480 may receive external power and internal power under the control of the processor 1450, and provide the power required for the operation of each component of the system.

[0257] The processor 1450 may indicate one or more processors. For example, the processor 1450 may include one or more central processing units, or include a central processing unit and a graphics processing unit, or include an application processor and a coprocessor (such as a micro control unit). When the processor 1450 includes multiple processors, these multiple processors may be integrated on the same chip or may be independent chips respectively. A processor may include one or more physical cores, where the physical core is the smallest processing module.

[0258] Exemplarily, the processor 1450 may implement the functions of the processing unit 1320 in the device 1300. Optionally, the processor 1450 may also implement the functions of the establishment unit in the device 1300, and the present application does not limit this.

[0259] Exemplarily, the processor 1450 may also be used to implement the functions of the acquisition unit 1310 in the device 1300, such as obtaining the first image captured by the terminal from the imaging unit (such as a camera), and / or obtaining the first position of the first image from the GPS module or the WiFi module, etc., and the present application does not limit this.

[0260] The memory 1470 stores computer programs, which include an operating system program 1472, an application program 1471, etc. Typical operating systems such as Windows developed by Microsoft Corporation and MacOS developed by Apple Inc. for desktop or laptop systems, and Android developed by Google Inc. for mobile terminal systems based on Android systems, etc. The methods provided in the foregoing embodiments may be implemented in software, and may be considered as the specific implementation of the application program 1471 and / or the operating system program 1472.

[0261] Memory 1470 can be one or more of the following types: flash memory, hard disk type memory, micro multimedia card type memory, card type memory (such as SD or XD memory), random access memory (RAM), static random access memory (SRAM), read only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable ROM (PROM), magnetic memory, magnetic disk or optical disk. In some other embodiments, memory 1470 can also be a network storage device on the Internet, and the system can perform operations such as updating or reading on the memory 1470 on the Internet.

[0262] Processor 1450 is used to read the computer program in memory 1470, and then execute the method defined by the computer program. For example, processor 1450 reads the operating system program 1472 to run the operating system in the system and implement various functions of the operating system, or reads one or more application programs 1471 to run applications on the system.

[0263] Memory 1470 also stores other data 1473 in addition to the computer program, such as the image database involved in this application.

[0264] Figure 14 The connection relationship between the modules in [the device] is only an example. The method provided in any embodiment of this application can also be applied to a positioning device with other connection methods, such as all modules are connected through a bus.

[0265] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0266] The embodiments of the present application also provide a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a computer, the steps of the positioning method in any of the above embodiments are implemented.

[0267] The embodiments of the present application also provide a computer program product, and when the computer program product is executed by a computer, the steps of the positioning method in any of the above embodiments are implemented.

[0268] The various embodiments in the present application can be used independently or in combination, and this is not limited here.

[0269] In addition, various aspects or features of the present application can be implemented as a method, an apparatus, or an article of manufacture using standard programming and / or engineering techniques. The term "article of manufacture" as used in this application encompasses a computer program accessible from any computer-readable device, carrier, or medium. For example, computer-readable media can include, but are not limited to: magnetic storage devices (such as hard disks, floppy disks, or magnetic tapes, etc.), optical discs (such as compact discs (CDs), digital versatile discs (DVDs), etc.), smart cards, and flash memory devices (such as erasable programmable read-only memories (EPROMs), cards, sticks, or key drives, etc.). Additionally, the various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable media" can include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.

[0270] It should be noted that in each of the embodiments provided in this application, there is no time limit relationship between the steps, and each step can be regarded as a solution, or can be combined with one or more other steps to form a solution. This application does not make any limitations in this regard.

[0271] The various embodiments in this application can be used independently or jointly. For example, any one or more steps in different embodiments can be combined to form an embodiment alone, and no limitations are made here.

[0272] It should be understood that in the embodiments shown above, the first and the second are only for facilitating the distinction of different objects and should not constitute any limitation to this application.

[0273] It should also be understood that in the embodiments of this application, the magnitude of the sequence numbers of the above processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of this application.

[0274] It should also be understood that "and / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one" means one or more; "at least one of A and B" is similar to "A and / or B" and describes the association relationship of associated objects, indicating that there can be three relationships. For example, at least one of A and B can represent: A exists alone, A and B exist simultaneously, and B exists alone.

[0275] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0276] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0277] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0278] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0279] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0280] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0281] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A positioning method, characterized in that, Including: Obtain a first image captured by a first terminal; Determine at least one second image matching the first image in the image database according to the first information of the first image and the first information of the images in the image database, where the first information is used to indicate the global features of the image, and the image database includes the first information of multiple frames of images, the second information of the multiple frames of images, and the poses of the multiple frames of images in the world coordinate system, where the second information is used to indicate the local features of the image; Determine the poses of the first image and the at least one second image in a first local coordinate system of the first terminal according to the second information of the first image and the second information of the at least one second image; Output the pose of the first image in the world coordinate system according to the pose of the first image in the first local coordinate system, the poses of the at least one second image in the first local coordinate system, and the poses of the at least one second image in the world coordinate system.

2. The method according to claim 1, wherein Also including: Establish the image database.

3. The method according to claim 2, wherein The establishing of the image database includes: Obtain a video captured by a second terminal; Obtain the poses of multiple frames of images in the video in a second local coordinate system of the second terminal; Determine the poses of the multiple frames of images in the video in the world coordinate system according to third information and the poses of the multiple frames of images in the video in the second local coordinate system, where the third information is used to indicate the positions of at least some of the images in the video in a map, and the map is associated with the world coordinate system; Obtain the first information and the second information of the multiple frames of images in the video.

4. The method according to claim 1, wherein The method further includes: Obtain a first position of the first image, where the first position comes from a global positioning system (GPS) module or a Wi-Fi module of the first terminal; Determine at least one image in the image database according to the first position of the first image, and the distance between the position of each image in the at least one image and the first position of the first image is less than a first threshold; Wherein, the determining of at least one second image matching the first image in the image database according to the first information of the first image and the first information of the images in the image database includes: Match the first information of the first image with the first information of at least one of the images, and obtain at least one second image matching the first image among the at least one image.

5. The method according to any one of claims 1-4, characterized in that, The determining of at least one second image matching the first image in the image database according to the first information of the first image and the first information of the images in the image database includes: Determine multiple third images matching the first image in the image database according to the first information of the first image and the first information of the images in the image database; Delete the images in the multiple third images that have a distance greater than a second threshold from the cluster center of the multiple third images to obtain the at least one second image, where the cluster center is determined according to the positions of the multiple third images in the world coordinate system.

6. The method according to any one of claims 1-4, characterized in that The determining, in the image database, of at least one second image that matches the first image according to the first information of the first image and the first information of the images in the image database includes: Determine a plurality of fourth images that match the first image in the image database according to the first information of the first image and the first information of the images in the image database; Delete the images in the plurality of fourth images that have an angle less than a third threshold and / or a distance less than a fourth threshold to obtain the at least one second image, where the angle is the difference in the poses of at least two images in the world coordinate system, and the distance is the difference in the positions of the poses of at least two images in the world coordinate system.

7. A positioning device, characterized in that, Includes: An acquisition unit configured to acquire a first image captured by a first terminal; A processing unit configured to determine, in the image database, at least one second image that matches the first image according to the first information of the first image and the first information of the images in the image database, where the first information is used to indicate the global features of the image, and the image database includes the first information of multiple frames of images, the second information of the multiple frames of images, and the poses of the multiple frames of images in the world coordinate system, and the second information is used to indicate the local features of the image; The processing unit is further configured to determine the poses of the first image and the at least one second image in a first local coordinate system of the first terminal according to the second information of the first image and the second information of the at least one second image; An output unit configured to output the pose of the first image in the world coordinate system according to the pose of the first image in the first local coordinate system, the poses of the at least one second images in the first local coordinate system, and the poses of the at least one second images in the world coordinate system.

8. The device according to claim 7, characterized in that Further includes a building unit configured to build the image database.

9. The device according to claim 8, characterized in that, The building unit is specifically configured to: Acquire a video captured by a second terminal; Acquire the poses of multiple frames of images in the video in a second local coordinate system of the second terminal; Determine the poses of the multiple frames of images in the video in the world coordinate system according to third information and the poses of the multiple frames of images in the video in the second local coordinate system, where the third information is used to indicate the positions of at least some of the images in the video in a map, and the map is associated with the world coordinate system; Acquire the first information and the second information of the multiple frames of images in the video.

10. The apparatus according to claim 7, wherein The acquisition unit is further configured to acquire a first position of the first image, where the first position comes from a global positioning system (GPS) module or a wireless fidelity (WiFi) module of the first terminal; The processing unit is further configured to determine at least one image in the image database according to the first position of the first image, wherein the distance between the position of each image in the at least one image and the first position of the first image is less than a first threshold; The processing unit is further configured to match the first information of the first image with the first information of the at least one image, and obtain at least one second image that matches the first image from the at least one image.

11. The device according to any one of claims 7 to 10, characterized in that, Specifically, the processing unit is configured to: Determine a plurality of third images that match the first image in the image database according to the first information of the first image and the first information of the images in the image database; Delete the images in the plurality of third images whose distance from the cluster center of the plurality of third images is greater than a second threshold, so as to obtain the at least one second image, wherein the cluster center is determined according to the positions of the plurality of third images in the world coordinate system.

12. The device according to any one of claims 7 to 10, characterized in that Specifically, the processing unit is configured to: Determine a plurality of fourth images that match the first image in the image database according to the first information of the first image and the first information of the images in the image database; Delete the images in the plurality of fourth images whose angle is less than a third threshold and / or whose distance is less than a fourth threshold, so as to obtain the at least one second image, wherein the angle is the difference in the poses of at least two images in the world coordinate system, and the distance is the difference in the positions of at least two images in the world coordinate system.

13. A terminal device, characterized in that, Comprising: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method for calculating pose of vehicle at curve

    CN107704821A