Multi-feature visual positioning method

By utilizing BIM models and photo feature extraction technology on mobile devices, combined with tightly coupled features for visual positioning, the problem of poor positioning accuracy in urban areas and indoors has been solved, achieving a high-precision, low-cost positioning solution.

CN115527044BActive Publication Date: 2026-04-07THE HONG KONG POLYTECHNIC UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-08
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing positioning technologies have poor positioning accuracy in urban areas and indoor environments, require high-end hardware, consume a lot of power, have a small user base, and have low reliability.

Method used

By using BIM models and photos taken by mobile devices, feature points, objects, materials, and line features are extracted. Visual positioning is then performed by combining tightly coupled features, and the user's location is determined by calculating similarity through a central server.

Benefits of technology

It improves indoor and outdoor positioning accuracy, reduces hardware investment, lowers power consumption, and enhances positioning reliability and coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115527044B_ABST
    Figure CN115527044B_ABST
Patent Text Reader

Abstract

Provided is a multi-feature visual positioning method, comprising: receiving a raw photo taken by a user and photographing parameters associated therewith and a positioning request, wherein the photographing parameters comprise a preliminary position of a user photographing device and attitude data; determining a plurality of candidate poses within a certain range with the preliminary position as the center; generating an image for each candidate pose according to the attitude data using a BIM model; calculating the similarity of the generated image of each candidate pose and the raw photo; and determining the position in the candidate pose where the image with the maximum similarity is located as the user position. The method improves the visual positioning accuracy in urban areas, including outdoors and indoors, without increasing hardware devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual positioning, and more particularly to a method for visual positioning using multiple features. Background Technology

[0002] In recent years, various methods for positioning both indoors and outdoors have emerged. Global Navigation Satellite Systems (GNSS) can typically achieve positioning accuracy of less than 1 meter. However, in urban areas, GNSS positioning signals are obstructed by buildings, trees, and other obstructions, leading to decreased accuracy. This is especially true in densely populated cities like Beijing and Hong Kong, where tall buildings block or reflect signals, severely degrading positioning accuracy.

[0003] In the field of indoor positioning, several technologies have been developed that utilize radio frequency (RF) signals for positioning, including Wi-Fi, Bluetooth, Bluetooth Low Energy (BLE), and Ultra Wideband (UWB) technologies. However, these methods have relatively high requirements for hardware and other infrastructure, requiring investment in a significant amount of hardware signal equipment or modification of existing equipment. Furthermore, these devices need to scan signals at high frequencies (e.g., transmitting a positioning signal every 1 to 4 seconds). This method suffers from drawbacks such as high power consumption and low positioning accuracy; for example, positioning using Wi-Fi has an error range of 5 to 15 meters.

[0004] In addition to the aforementioned drawbacks such as poor positioning accuracy and high cost of basic equipment, existing positioning technologies also suffer from limitations such as a small user base, low reliability, and overlapping coverage areas of positioning devices. Summary of the Invention

[0005] To address the aforementioned shortcomings of existing technologies, the inventors of this invention propose a multi-feature visual positioning method based on BIM models and photographs taken by mobile devices.

[0006] This invention extracts four types of features from images taken by users on-site: edge, object, material, and line. These features are then compared with the corresponding features stored in the BIM database, thereby improving the user's visual positioning accuracy both indoors and outdoors.

[0007] According to one aspect of the present invention, a central server (or processing device) connected to a network receives an original image taken by a user, along with associated photographing parameters and a positioning request. The photographing parameters include the pose of the user's photographing device, i.e., a preliminary position and pose data. A plurality of candidate positions are determined within a certain range centered on the preliminary position. An image is generated for each candidate position using a BIM model based on the pose data, and this image is similar to the received original image taken by the user. Then, the central server calculates the similarity between the generated image and the original image for each candidate position. The candidate position containing the image with the highest similarity is determined as the user's position. Finally, the determined user position is sent to the user.

[0008] The steps for calculating the similarity between the generated image and the original image at each candidate location include: calculating the similarity between multiple features of the generated image at each candidate location and the corresponding multiple features in the original image.

[0009] In another implementation, tightly coupled features are used for positioning. Attached Figure Description

[0010] The various embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which include:

[0011] Figure 1a This is an example of a photograph taken with a mobile phone used for positioning in one embodiment of the present invention;

[0012] Figure 1b It is extraction Figure 1a Line feature diagram formed after the line features in the image;

[0013] Figure 1c Is Figure 1a Based on the display of feature points, a diagram is shown.

[0014] Figure 1d This is an example of a candidate image generated using a BIM system based on the camera-related parameters of a user's mobile phone in one embodiment of the present invention;

[0015] Figure 2 This is a schematic diagram illustrating the angular coordinates defining the orientation of the shooting equipment in three-dimensional space;

[0016] Figure 3 yes Figure 1a The example shown is a partial top-view of the city where the user needs to be located, generated by BIM. In this view, multiple dots distributed in a grid represent multiple candidate locations.

[0017] Figure 4This is a flowchart of a visual positioning method according to an embodiment of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below through embodiments and in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0019] Smartphones all have camera functions and the ability to upload photos to the internet. Other portable devices, such as digital cameras and iPads, can also take photos and upload them online. In this invention, these photos taken on-site and uploaded to the internet are used to accurately locate the photographer (user).

[0020] Figure 1a An example image (original image) taken by a user on a city street using a mobile phone is shown. In the following description, a photo taken with a mobile phone is used as an example, but those skilled in the art will know that the image can be any suitable image (such as a video, a frame from a video, etc.) taken by a mobile phone or any other suitable electronic device (such as a tablet computer). In this invention, such a digital photo is linked to the mobile phone's photo-related parameters being sent to a central server (or similar data processing device) connected to the Internet to accurately locate the user's position when taking the photo. The photo-related parameters of the mobile phone include the focal length f of the mobile phone camera and the phone's pose data, including attitude and coarse position. Here, the coarse position (or preliminary position, including longitude, latitude, and altitude) can be obtained by existing positioning methods, such as GNSS. As mentioned above, in urban areas, existing positioning methods have a large error in obtaining the preliminary position. This invention can further refine the positioning based on the coarse position. Furthermore, the mobile phone's attitude can be determined by three attitude angles in a three-dimensional coordinate system: pitch angle (ψ), yaw angle (θ), and roll angle. Figure 2 A schematic diagram showing the pitch, yaw, and roll angles is provided. The phone's attitude, determined by these three-dimensional coordinates, determines the angle at which the phone's camera takes the picture. The inertial measurement unit (IMU) in the phone can measure the phone's pitch (ψ), yaw (θ), and roll angles in real time. In this invention, the photos (original images) taken by the user are transmitted to a central server (or similar processing device) via a network, along with the aforementioned photo-related parameters of the mobile phone.

[0021] Next, the visual positioning method of the present invention determines the user's location based on the received camera-related parameters from the mobile phone and utilizes visual positioning technology. One embodiment of the present invention employs Building Information Modeling (BIM) as the visual positioning technology; however, those skilled in the art will know that any other suitable visual positioning technology can also be used. Specifically, one embodiment of the present invention utilizes Building Information Modeling (BIM) to provide multiple corresponding candidate images, and then compares these candidate images with the image actually taken by the user to find the corresponding candidate image that best approximates the user's actual photo. Depending on actual needs, candidate images can be provided through real-time calculation or by retrieving images stored in a database.

[0022] BIM technology is a data-driven tool applied to engineering design, construction, and management. By integrating the data and information models of buildings, it enables sharing and transmission throughout the entire lifecycle of project planning, operation, and maintenance, allowing engineering technicians to correctly understand and efficiently respond to various building information.

[0023] The core of BIM is to create a virtual 3D model of the building project and use digital technology to provide this model with a complete and accurate database of building project information. This database not only contains geometric information, professional attributes, and status information describing building components, but also status information of non-component objects (such as spatial and movement behaviors).

[0024] For example, Figure 1d This is an example of an image generated by BIM, a view taken from a street location looking upwards. After determining the latitude, longitude, altitude, and viewing angle of the observation point, BIM can generate a corresponding visualization image. This visualization image corresponds to an image taken with a mobile phone (or other camera equipment) at the same latitude, longitude, altitude, and viewing angle in the actual scene. Figure 3 It is generated by BIM in and Figure 1d A top view of a location close to the position.

[0025] BIM technology can generate not only outdoor scene images but also indoor scene images. Furthermore, the BIM database is dynamic, constantly being updated, enriched, and expanded during application.

[0026] The central server (or similar processing device) for positioning processing in this invention is connected to the aforementioned BIM system, or the central server (or similar processing device) of this invention is equipped with the aforementioned BIM system and is capable of calling programs in the BIM system and data in the database to generate data for positioning, for example, generating a candidate image (or candidate image, such as...) for each candidate pose. Figure 1d As shown in the image, each candidate image is similar to the original image.

[0027] Specifically, as mentioned above, the central server receives photos sent by the user via mobile phone (e.g., Figure 1a The image shown includes a photograph and associated photograph-related parameters, including the phone's approximate location and the phone's attitude data measured by its IMU. Next, the central server sets several candidate locations within a radius L, centered on the phone's approximate location. As mentioned above, the phone's approximate location can be obtained using existing positioning methods (such as GNSS). Figure 3 This is a top-down view of the city streets near the user's approximate location, generated by BIM, where grid-like dots represent candidate locations. Specifically, candidate locations are determined by a central server within a circle centered on the user's approximate location and with radius L. Wherein, because... Figure 1a The example photo shows an outdoor street scene. The central server determines the street and building locations based on the BIM database and sets the candidate locations on the street, not inside the buildings. Each candidate location may have different latitude and longitude but the same altitude as the approximate location received from the user's mobile phone. In another embodiment, the latitude and longitude of each candidate location are set by the central server, and then the altitude corresponding to each candidate location's latitude and longitude is read from the BIM database.

[0028] The radius L can be set manually. The value of L can be 20 meters, 50 meters, 100 meters, or any other distance.

[0029] Next, the central server, using each candidate location as the user's assumed location and the received mobile phone posture data as the shooting angle, generates a photo (or image) using the BIM model, such as... Figure 1d The image shown.

[0030] The central server is Figure 3 Each candidate location shown generates an image as follows Figure 1dThe images shown are similar, but distinct, because they simulate images taken from similar locations and angles. Then, the similarity between each generated image and the original image (i.e., the image received from the user) is compared one by one. This involves calculating the similarity between each generated image and the original image, and determining the candidate location with the highest similarity as the user's location.

[0031] In calculating the similarity between each generated image and the original image, existing calculation methods can be used to obtain the similarity between the two images, or the improved calculation method in the preferred embodiment of the present invention can be used to calculate the similarity between the generated image and the original image, so as to improve the calculation speed and accuracy.

[0032] In a preferred embodiment, the central server... Figure 1a The image shown was taken with a mobile phone and analyzed to extract line features and remove other irrelevant elements, such as... Figure 1b As shown. In a further preferred embodiment, extraction Figure 1b The feature points (edges) in the line feature map include the intersection of two lines and the turning point of a line. Figure 1c exist Figure 1a The image shows the feature points, with each small box representing a feature point. Line features and feature points are used to compare with the corresponding image generated by BIM, thus more quickly determining the similarity between the actual captured image and the BIM-generated image.

[0033] In yet another preferred embodiment, the central server employs an image semantic analysis method to... Figure 1a The image shown is captured by a mobile phone and interpreted, analyzed, or segmented to identify each category of objects. Each category is assigned a color, with different categories receiving different colors, thus generating a segmented image. In the generated segmented image, each object's pixels have the same color, but different objects have different colors. For example, in the image above, tables are one category, buildings are another, and shelves are yet another. In the segmented image, different categories of objects are assigned different colors. Similarly, images generated by BIM can also be segmented to obtain segmented images. Using segmented images facilitates comparison between BIM-generated images and the original images captured by the mobile phone.

[0034] In addition, existing artificial intelligence deep learning methods can be used for object classification. For example, see the deep learning methods described at the following website:

[0035] https: / / paperswithcode.com / sota / semantic-segmentation-on-nyu-depth-v2

[0036] The following describes a scheme for calculating the similarity between two images in an improved visual positioning method according to a preferred embodiment of the present invention. Figure 4 A flowchart of a visual positioning method according to an embodiment of the present invention is shown.

[0037] In step S110, the central server receives the original image captured by the user on the spot, along with the associated photo-related parameters, and also receives the user's location request. The user's mobile phone only needs to have wireless internet access to send the original image, associated photo-related parameters, and location request to the central server, without necessarily requiring a communication access point like Wi-Fi or Bluetooth.

[0038] In step S120, the central server sets a circle of radius L centered on the approximate location, and sets several candidate locations within the circle. The number of candidate locations can be set manually as needed; for example, the horizontal distance between adjacent candidate locations can be between 0.1 meters and 10 meters, or any other distance.

[0039] In step S130, for each candidate location, and based on the received pose data, corresponding image data (such as...) is generated using the BIM model. Figure 1d (As shown).

[0040] In step S140, for each candidate location, the similarity between the image data generated for each candidate location and the original image received in step S110 is calculated. The specific method for calculating the similarity will be described in detail below.

[0041] In step S150, the similarity between the images of all candidate locations and the original image is compared, and the candidate location with the highest similarity is determined as the user's accurate location.

[0042] In step S160, the obtained user location information is sent to the customer, for example, to the customer's mobile phone.

[0043] The step S140 above, which calculates the similarity between the generated image data and the original photo, involves a significant amount of computation. The method for calculating this similarity will be further described below.

[0044] Refer again Figure 1aThe photograph shown is a street scene taken at a certain upward angle. It includes several buildings, doors, sky, tables, chairs, shelves, and a transparent ceiling with a mesh pattern. Some buildings have glass facades, while others have exterior walls made of ordinary building materials. Each building has a specific outline curve. This invention uses image semantic analysis methods to interpret, analyze, or segment this photograph, extracting various feature elements. These feature elements include: material, object, line, and edge. Material can include, for example, glass, brick, cement, etc.; object features include, for example, buildings, doors, sky, shelves, chairs, tables, etc.; line features include the outlines of buildings, doors, sky, chairs, and tables; and edge features include points characterizing the features of objects in the photograph, such as... Figure 1c The intersections and turning points of the lines shown.

[0045] Then, the central server generates a corresponding image for each candidate location using the BIM database, and extracts the corresponding features of each element from each image using image semantic analysis methods, including material, object, line, and edge (this process is similar to the processing of the original image described above). When calculating the similarity between the candidate image generated for each candidate location and the original image, this invention preferably calculates the similarity between each feature of the candidate image and each corresponding feature of the original image.

[0046] In one embodiment of the present invention, the total similarity between the generated candidate image and the original image is calculated using the following formula.

[0047] prob(x j ) = prob material (x j )·prob object (x j )·prob line (x j )·prob edge (x j (1)

[0048] in,

[0049]

[0050] Where, x j This represents the candidate pose of the user's mobile phone, i.e., its position and attitude. The subscript j indicates the number of each candidate pose, lat represents latitude, lon represents longitude, alt represents altitude, ψ represents pitch angle, and θ represents yaw angle. This indicates the roll angle; prob represents the probability, i.e., similarity; material represents the feature material; object represents the feature object; line represents the feature line; and edge represents the feature point. Additionally, all x... j Form a set X, that is, x j It is an element of set X, x j ∈X.

[0051] In formula (1), prob(x) j ) is the candidate pose x j Total similarity, i.e., for candidate pose x j The final total similarity between the generated image and the original image. This total similarity is the product of the similarities of four individual elements. Before calculating the total similarity, in one embodiment of the invention, the similarity of four individual elements in the generated image is first calculated, namely, the similarity between the material features in the generated image and the material in the original image, i.e., prob... material (x j The similarity between objects in the generated image and objects in the original image, i.e., the probability density function (prob). object (x j The similarity between the lines in the generated image and the lines in the original image, i.e., prob... line (x j The similarity between feature points in the generated image and feature points in the original image, i.e., prob edge (x j The total similarity is the product of the similarities of these four individual elements.

[0052] In another embodiment, the total similarity of only two or three of the four feature elements can be calculated.

[0053] The candidate pose with the highest similarity among all candidate poses is determined using the following formula:

[0054]

[0055] Where arg max is the function for finding the maximum value; The candidate pose with the highest similarity.

[0056] The following describes the calculation of the four types of single-element similarity. The calculation of each single-element similarity uses the score of each single element. Therefore, the calculation of the score of each single element will be introduced first.

[0057] Material characteristic element score

[0058] In one embodiment, the Jaccard metric is used to compare two segmented images of a material, calculating the similarity for each material class:

[0059]

[0060] Where sim is the similarity function. It is the original image (i.e., the image taken by a mobile phone or other digital camera) and the candidate pose x j The similarity index of the generated image (candidate image) within a certain material class; Img cam_seg The segmented image representing the original image. Indicates the candidate pose x j The segmented image of the generated candidate image, where the superscript 3DM_seg indicates the segmented image of the image generated by BIM.

[0061] Calculate the above similarity index for each material class, and then sum the similarity indices for each class using a weighted method to obtain the total score for the material class:

[0062]

[0063] Indicates the candidate pose x j The generated candidate image is compared with the score of a certain material class in the original image. The above formula (5) is a weighted average based on the number of pixels occupied by a certain material class in the candidate image, where N total This represents the total number of pixels in the candidate image. This indicates the number of pixels in an image representing a particular type of material.

[0064] Finally, summing the results for each material class calculated using formula (5) yields a candidate pose x. j Material characteristic score:

[0065]

[0066] score material (x j ) represents the candidate pose x j The overall material characteristic score.

[0067] Object feature element score

[0068] The calculation of the object feature element score is the same as the calculation of the feature score, except that the object of the calculation is replaced by the object. In the calculation formula below, the superscript material is replaced with object accordingly.

[0069] In one embodiment, the Jaccard metric is also used to compare two segmented images of an object and calculate the similarity for each object class:

[0070]

[0071] Where sim is the similarity function. It is the original image (i.e., the image taken by a mobile phone or other digital camera) and the candidate pose x j The similarity index of the generated image (candidate image) to a certain object class; Img cam_seg The segmented image representing the original image. Indicates the candidate pose x j The segmented image of the generated candidate image.

[0072] Calculate the above similarity index for each object class, and then sum the similarity indices for each class using weighted averages to obtain the object feature score:

[0073]

[0074] Indicates the candidate pose x j The generated candidate image is compared with the feature score of a certain object in the original image. The above formula (5) is a weighted average based on the number of pixels occupied by a certain object class in the candidate image, where N total This represents the total number of pixels in the candidate image. This indicates the number of pixels of a certain type of object in the image.

[0075] Finally, the summation is performed on each object class calculated using formula (5) to obtain a candidate pose x. j Object feature score:

[0076]

[0077] score object (x j ) represents the candidate pose x j The overall object feature score.

[0078] Line feature element score

[0079] In one embodiment, a regional-based metric is used, specifically the boundary F1 metric, to compare line or boundary features.

[0080] Let B cam_seg (class) is the image Img of a certain class. cam_seg The boundary of (class), and let For a certain class of images The boundary of the image. The distance from a point (or pixel) in the image to the image Img. cam_seg (class) boundary or to image The distance to the boundary is expressed in pixels. In this embodiment, a distance of 5 pixels is set as the threshold (this invention is not limited to this, and other numbers of pixels can also be set as the threshold). The above boundary F1 metric does not consider the content of the segmented image exceeding the 5-pixel threshold. In addition, the accuracy of a certain class is calculated according to the following formula:

[0081]

[0082] The call (recall) to a class is defined as:

[0083]

[0084] in, It's Iverson bracket notation. Here, if... but otherwise And, d() represents the Euclidean distance in pixels. The calculation result of the above call represents a positive prediction of pixel loss. The numerical calculation of the boundary F1 of a certain class is as follows:

[0085]

[0086] in, Indicates the candidate pose x j The line feature element score for a specific class in the generated image is calculated. Finally, the total line feature score for the candidate image is obtained by averaging the line feature scores of all classes appearing in the generated candidate image.

[0087]

[0088] Where n_class represents the total number of classes, and score line (x j ) represents x j The total score of the line features of the candidate image at the location.

[0089] Feature point (edge) element score

[0090] In one embodiment, the Fast Library for Approximate Nearest Neighbors (FLANN) metric is used to compare feature points of two images. Methods for calculating this metric include, but are not limited to, using the FLANN function as follows:

[0091]

[0092] score edge (x j ) represents the pose x j The feature point scores of the candidate image and the original image.

[0093] The above describes the calculation of scores for the four individual elements of the candidate image at each location. Next, using these scores, the similarity of each individual element can be calculated using the following formula:

[0094]

[0095] In the above formula, * represents the corresponding element, namely, material, object, line, edge; σ is the standard deviation, and μ is the average value of the cumulative distribution function (CDF) mentioned above.

[0096] In one embodiment of the present invention, comparing the generated image with the original image involves comparing the two-dimensional generated image with the two-dimensional original image.

[0097] In another embodiment of the present invention, comparing the generated image with the original image includes first deconstructing the two-dimensional original image into three-dimensional data. Specifically, depth data of each feature element (e.g., object, line, and feature point) in the two-dimensional photograph is calculated based on artificial intelligence methods. Then, for each candidate pose, three-dimensional image data is generated using a BIM model. This three-dimensional image data includes not only pixels on the plane (i.e., a two-dimensional planar image) but also the depth of each feature on the planar image. Finally, the similarity between the three-dimensional data obtained by deconstructing the original image and the three-dimensional image data generated using the BIM model is calculated.

[0098] In the aforementioned embodiment that compares the generated two-dimensional image with the original two-dimensional image, since it is not necessary to deconstruct the original two-dimensional image into three-dimensional data, the computational cost is less than that of calculating the similarity between the generated three-dimensional image data and the original three-dimensional image, thus the positioning speed is faster.

[0099] In a preferred embodiment, a tightly coupled method is used for positioning. This contrasts with a loosely coupled method. Loosely coupled methods use data from a single sensor for positioning, such as IMU data or GNSS data alone, or images captured by a camera alone. In contrast, tightly coupled methods combine raw data from two or more sensors, along with other sensor parameters, for positioning. For example, they combine raw IMU data, GNSS data, and raw images captured by a camera, resulting in higher positioning accuracy.

[0100] The present invention also relates to a data processing apparatus (or central server) equipped with computer-readable software, which, when run by the data processing apparatus, enables the data processing apparatus to perform the visual positioning methods described in the above embodiments.

[0101] Reference above Figure 1a The outdoor photographs shown illustrate an example of outdoor positioning. This invention can also be applied to indoor positioning, such as positioning in large shopping malls.

[0102] The method of the present invention is a supplement to existing positioning methods, especially capable of improving positioning accuracy in urban areas without increasing investment in hardware facilities.

[0103] Although the present invention has been described through specific embodiments, those skilled in the art will understand that various changes and equivalent substitutions can be made to the invention without departing from its scope. Furthermore, various modifications can be made to the invention for specific situations or materials without departing from its scope. Therefore, the present invention is not limited to the specific embodiments disclosed, but should include all embodiments falling within the scope of the claims.

Claims

1. A visual positioning method, comprising: Receive the original image captured by the user and the associated image capture parameters and positioning request, wherein the image capture parameters include the initial position and orientation data of the user's shooting device; Determine several candidate poses within a certain range centered on the initial position; The user's location is determined using visual positioning technology and the candidate image parameters. The step of determining the user's location using visual positioning technology, along with the photographic parameters and candidate poses, includes: Using the BIM model, an image is provided for each candidate pose based on the pose data; Calculate the similarity between the generated image and the original image for each candidate pose; The position of the candidate pose containing the image with the highest similarity is determined as the user's position; in, Calculating the similarity between the generated image and the original image for each candidate pose involves using the following formula: prob(x j )=prob material (x j )·prob object (x j )·prob line (x j )·prob edge (x j ) (1) in, Where x indicates the location and attitude of the user's mobile phone, the subscript j represents the number of each candidate pose, lat represents latitude, lon represents longitude, alt represents altitude; ψ represents pitch angle, θ represents yaw angle, ... The value represents the roll angle; prob represents the similarity; material represents the feature material; object represents the feature object; line represents the feature line; and edge represents the feature point.

2. The visual positioning method according to claim 1, wherein, The method further includes sending the determined user location to the user.

3. The visual positioning method according to claim 1, wherein, Using the BIM model to provide an image for each candidate pose based on the pose data is: Using the BIM model, an image is generated for each candidate pose based on the pose data.

4. The visual positioning method according to claim 1, wherein, After receiving the raw image taken by the user, along with its associated shooting parameters and location request, it also includes: Extract multiple features from the original image, the multiple features including at least two of the following: material, object, line, and feature point; The calculation of the similarity between the generated image and the original image for each candidate pose includes: calculating the similarity between at least two features in the generated image and corresponding features in the original image.

5. The visual positioning method according to any one of claims 1-4, wherein, The image generated for each candidate pose using the BIM model based on the pose data is a two-dimensional image.

6. The visual positioning method according to any one of claims 1-4, wherein, Receiving raw images taken by the user includes receiving a photograph taken by the user.

7. The visual positioning method according to any one of claims 1-4, wherein, Receiving raw images taken by the user includes receiving two or more tightly coupled photographs taken by the user.

8. The visual positioning method according to any one of claims 1-4, in, Extracting features from the original image includes extracting tightly coupled features from the original image; Furthermore, calculating the similarity between the generated image and the original image for each candidate pose includes calculating the similarity between the tightly coupled features of the generated image and the tightly coupled features of the original image for each candidate pose.

9. A data processing apparatus having computer-readable software installed thereon, which, when executed by the data processing apparatus, enables the data processing apparatus to perform the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Image-based indoor positioning method and device

    CN111340882A