Visual Indoor Positioning Method and System Based on Prior Information of Architectural Floor Plans

By registering visual information with building floor plans and using deep learning models, the problem of coupling positioning and drawing construction in traditional visual SLAM technology is solved, and the equipment estimates its own position without fully exploring the building body, improving the adaptability and convenience of indoor positioning.

CN114708309BActive Publication Date: 2025-06-13GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210172955.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-22
Publication Date
2025-06-13
Estimated Expiration
2042-02-22

AI Technical Summary

Technical Problem

In traditional visual SLAM technology, positioning and mapping are coupled, making it difficult to estimate the position of the equipment in the building without fully exploring the building.

Method used

By registering the visual information currently observed by the device with the building plan, using K nearest neighbor distance histogram, point cloud feature extraction network, graph attention network, optimal transmission algorithm and self-supervised training method, an end-to-end deep learning model is built to realize that the device estimates its own position in the building without fully exploring the building.

Benefits of technology

The adaptability and convenience of indoor positioning related applications to unknown indoor environments is improved, and the equipment can obtain its own position relative to the building under initial conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708309B_ABST
    Figure CN114708309B_ABST
Patent Text Reader

Abstract

To address the deficiencies of the prior art, the present invention provides a vision-based indoor positioning method and system based on prior information of building floor plans, including: preprocessing of building floor plans, data downsampling, input tensor construction, feature extraction and feature matching, registration and pose calculation, and model training. The method described in the present invention involves an end-to-end deep learning model that innovatively introduces the K-nearest neighbor distance histogram, point cloud feature extraction network, graph attention network, optimal transport algorithm, and self-supervised training method, making the model have strong robustness and adaptability, with relatively small model parameters, easy to train and deploy; the method described in the present invention can estimate the position of a device in a building based solely on the building floor plan without fully exploring the building structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The method of the present invention relates to the technical field of indoor positioning, and more specifically to a visual indoor positioning method and system based on prior information of building floor plans. Background Art

[0002] The problem of indoor positioning has always been a fundamental problem in many real-world applications, such as AR house viewing, indoor navigation, on-site furniture layout design, building information systems, mobile robots, etc. Currently, there are mainly two types of technical solutions for indoor positioning. One is the technology based on external beacons, such as ultra-wideband wireless communication technology positioning (UWB), Bluetooth, positioning QR codes, etc. However, this type of technology relies on pre-arranging relevant beacons in the positioning scenario. The other is to achieve indoor positioning based on the Simultaneous Localization and Mapping (SLAM) technology. This technology can use sensors such as lidar and cameras to achieve the positioning of the device itself without external beacons, so it has relatively wide application value. The visual SLAM technology has many advantages such as rich observation information, rapid adaptation to environmental changes, and relatively fast repositioning speed, and has become an important research direction in the field of SLAM research. However, in traditional visual SLAM technology, positioning and mapping are two mutually coupled and simultaneous processes - positioning depends on the landmark information provided by the mapping process, and establishing a map requires first determining one's own current position to determine the position of the observed landmarks, and the map coordinate system is only related to the initial position. If the device needs to know its specific position in the building environment at the beginning of positioning, then prior building structure information needs to be introduced.

[0003] The building floor plan is a type of prior data that can intuitively and conveniently express the building structure information. The building floor plan can abstract relevant entities in the building space, such as doors, windows, walls, etc., into a two-dimensional representation form and retain a certain spatial structure relationship. These building floor plans are usually carefully drawn manually and have high accuracy and rich static structure information. The inventor believes that by registering the visual information currently observed by the device with the building floor plan and estimating the position parameters, it is possible to estimate the position of the device in the building without fully exploring the building, thereby improving the adaptability of indoor positioning-related applications to unknown indoor environments and the convenience of application use. Summary of the Invention

[0004] In order to solve the deficiencies of the prior art, the present invention provides a visual indoor positioning method and system based on prior information of building floor plans, which is used to solve at least one technical problem in the background art.

[0005] The technical solution adopted by the present invention is:

[0006] A visual indoor positioning method based on prior information of architectural floor plans, comprising:

[0007] Obtaining a set of architectural floor plan points with real dimensions according to the image data of the architectural floor plan to be registered;

[0008] Constructing a three-dimensional landmark point set, and downsampling the three-dimensional landmark point set and the architectural floor plan point set to obtain a downsampled three-dimensional landmark point set and a downsampled architectural floor plan point set;

[0009] After respectively extracting the N-dimensional feature vectors of the K-nearest neighbor distance histograms of the downsampled three-dimensional landmark point set and the downsampled architectural floor plan point set, obtaining a three-dimensional landmark input tensor corresponding to any point in the downsampled three-dimensional landmark point set and a floor plan point set input tensor of any point in the downsampled architectural floor plan point set;

[0010] Extracting the three-dimensional landmark feature vector of the three-dimensional landmark input tensor and the floor plan point set feature vector of the floor plan point set input tensor, and matching the three-dimensional landmark feature vector and the floor plan point set feature vector to obtain a matching matrix of any point in the architectural floor plan point set to the corresponding point in the three-dimensional landmark point set;

[0011] According to the matching matrix, obtaining the rigid body transformation R, t of the three-dimensional landmark point set to the architectural floor plan point set, where R is a rotation matrix and t is a translation vector;

[0012] Performing end-to-end deep learning using a self-supervised training method to obtain a model for indoor positioning.

[0013] The "obtaining a set of architectural floor plan points with real dimensions according to the image data of the architectural floor plan to be registered" includes:

[0014] Taking the image data of the wall layer of the architectural floor plan to be registered;

[0015] Extracting the edges of the wall layer image;

[0016] Taking the pixel coordinates of all points on the edges as the set of architectural floor plan points to be registered, and converting the set of architectural floor plan points to be registered into a set of architectural floor plan points with real dimensions according to the scale information of the architectural floor plan to be registered.

[0017] The process of "extracting the N-dimensional feature vectors of the K-nearest neighbor distance histograms of the downsampled three-dimensional landmark point set and the downsampled architectural floor plan point set" includes:

[0018] Query for the K points with the closest Euclidean distance to any point in the downsampled three-dimensional road landmark set and the downsampled building floor plan point set, and sort the Euclidean distance values of the points from the closest to the farthest;

[0019] Divide the range between 0 and the maximum distance value into M intervals, and calculate the number of points falling into each interval;

[0020] Divide the number of points in each interval by K to obtain the frequency of points appearing in each interval, thereby forming a frequency histogram, that is, the K-nearest neighbor distance histogram, which is represented by an N-dimensional vector.

[0021] The process of "obtaining the three-dimensional road landmark input tensor corresponding to any point in the downsampled three-dimensional road landmark set and the floor plan point set input tensor of any point in the downsampled building floor plan point set" includes:

[0022] Concatenate the N-dimensional feature vector of the K-nearest neighbor distance histogram with the three-dimensional coordinate vectors corresponding to each point in the downsampled three-dimensional road landmark set to obtain the three-dimensional road landmark input tensor;

[0023] Concatenate the N-dimensional feature vector of the K-nearest neighbor distance histogram with the three-dimensional coordinate vectors corresponding to each point in the downsampled building floor plan point set to obtain the floor plan point set input tensor.

[0024] The specific steps of "extracting the three-dimensional road landmark feature vector of the three-dimensional road landmark input tensor and the floor plan point set feature vector of the floor plan point set input tensor, and matching the three-dimensional road landmark feature vector with the floor plan point set feature vector" include:

[0025] The three-dimensional road landmark input tensor and the floor plan point set input tensor are subjected to feature extraction through a dynamic graph convolutional neural network, and a matrix composed of feature vectors of each point is output;

[0026] Construct a complete multi-graph through the point set, which includes intra-point set edges and inter-point set edges, where the intra-point set edges connect all vertices of the point clouds respectively, and the inter-point set edges connect a certain point in the point set with all vertices of another point set;

[0027] The complete multi-graph is input into a multi-layer graph attention network to achieve feature aggregation;

[0028] Through the feature aggregation, a three-dimensional road landmark tensor and a floor plan point set tensor are obtained, and after inner product calculation, a transmission cost matrix is obtained;

[0029] Obtain the matching matrix through the Sinkhorn optimal transport algorithm.

[0030] The specific steps of "obtaining the rigid body transformation R, t from the three-dimensional road punctuation set to the building floor plan point set according to the matching matrix, where R is the rotation matrix and t is the translation vector" include:

[0031] Generate a soft mapping matrix y' from the matching matrix P n×m through the following formula f :

[0032] y'=(y n×3 ) T ·P n×m ;

[0033] In the formula, y n×3 is the matrix composed of the building floor plan point set;

[0034] Substitute the soft mapping matrix into the ICP objective function shown in the following formula, and use the singular value decomposition method to calculate the rigid body transformation R, t;

[0035]

[0036] In the formula, x i is each point of the three-dimensional road punctuation set, y' (xi) is the corresponding point matched by x i , N is the number of x i , R * , t * is the final result of the rigid body transformation R, t.

[0037] The specific steps of "using the self-supervised training method for end-to-end deep learning to obtain a model for indoor positioning" include:

[0038] Take a random height within a certain range, generate a certain number of points vertically on the part of the building floor plan representing the wall, and form a three-dimensional point cloud corresponding to the building floor plan;

[0039] Crop and / or randomly rotate and / or translate and / or add noise points to the generated three-dimensional point cloud to achieve data augmentation;

[0040] Use the following loss function to train the parameters of each part of the deep learning model:

[0041] Loss = ||R T R g -I|| 2 +||t - t g || + λ||θ|| 2 ;

[0042] Among them, R g , t gThey are the true values of the rotation matrix and the translation vector from the point set of the building floor plan in the dataset to the landmark point set of ORB-SLAM, I is the identity matrix, and λ||θ|| 2 is the L2 regularization term, and θ is the model parameter.

[0043] A vision-based indoor positioning system based on prior information of building floor plans, comprising:

[0044] An acquisition module, connected to an external acquisition device, for acquiring image data of a building floor plan to be registered;

[0045] A processing unit, connected to the acquisition module, for obtaining the image data of the building floor plan to be registered and identifying the building floor plan to be registered according to the model for indoor positioning, and performing indoor positioning;

[0046] A display module, connected to the processing unit, for displaying information on indoor positioning.

[0047] A vision-based indoor positioning electronic device based on prior information of building floor plans, comprising:

[0048] A storage medium for storing a computer program;

[0049] A processing unit, for performing data exchange with the storage medium, and for, when performing indoor positioning, executing the computer program through the processing unit to perform the steps of the vision-based indoor positioning method as described above.

[0050] A computer-readable storage medium, in which a computer program is stored;

[0051] When the computer program runs, it executes the steps of the vision-based indoor positioning method as described above.

[0052] The beneficial effects of the present invention are:

[0053] The present invention provides a vision-based indoor positioning method based on prior information of building floor plans, which involves an end-to-end deep learning model. The model innovatively introduces the K-nearest neighbor distance histogram, the point cloud feature extraction network, the graph attention network, the optimal transport algorithm, and the self-supervised training method, making the model have strong robustness and adaptability, and having small model parameters, being easy to train and deploy; the method described in the present invention can estimate the position of the device in the building body only based on the building floor plan without fully exploring the building body.

[0054] The system described in the present invention uses data exchange between the processing unit and the storage module to identify the acquired building floor plan data, and can self-estimate the position of the device itself in the building body. Description of the Drawings

[0055] Figure 1 Schematic diagram of the method flow described in the present invention;

[0056] Figure 2 Schematic diagram of the registration result;

[0057] Figures 3(A) and 3(B) are schematic diagrams of the method for constructing the K-nearest neighbor distance histogram;

[0058] Figure 4 Schematic diagram of the feature extraction and feature matching model;

[0059] Figure 5 Schematic diagram of the generation of the virtual 3D landmark point set in self-supervised training;

[0060] Figure 6 It is the system block diagram of the system described in the present invention. Specific implementation manners

[0061] The present invention will be further described below with reference to the accompanying drawings.

[0062] The object of the present invention is to overcome the current coupling relationship between traditional visual SLAM positioning and mapping. On this basis, the prior information of the building floor plan is introduced to realize the visual indoor positioning function, make full use of the prior information brought by the building floor plan, and design an end-to-end deep learning model for registering the building floor plan to the 3D landmark point set based on the idea of feature matching; improve the application flexibility and convenience of indoor positioning devices, expand the application scenarios of indoor positioning, and obtain the pose of the device relative to the building under the initial conditions of the positioning system.

[0063] The present invention provides an embodiment:

[0064] As Figure 1 shown, a visual indoor positioning method based on the prior information of the building floor plan includes the following steps:

[0065] S1: Preprocessing of the building floor plan: Take the image data of the wall layer of the building floor plan to be registered, extract the edges of the image by the Canny edge detection algorithm, take the pixel coordinates of all edge points as the building floor plan point set, and convert it into a building floor plan point set with real size according to the scale information of the building floor plan;

[0066] S2: Data downsampling: Downsample the 3D landmark point set constructed by the ORB-SLAM system and the building floor plan point set by voxel filtering and uniform sampling;

[0067] S3: Input tensor construction: Extract the N-dimensional feature vectors of the K-nearest neighbor distance histograms of the downsampled 3D landmark point set and the building floor plan point set respectively, and splice the feature vectors with the 3D coordinate vectors of each point in the point set to form an input tensor in which each point is described by an (N + 3)-dimensional vector;

[0068] S4: Feature Extraction and Feature Matching: Use a Dynamic Graph Convolutional Neural Network (DGCNN) feature extraction network to extract features from the input tensors corresponding to the 3D road landmark set and the building floor plan point set respectively. Then, use a Graph Attention Network (GAT) and the Sinkhorn optimal transport method to match the feature vectors, obtaining a matching matrix for each point in the building floor plan point set to each point in the 3D road landmark set. This matching matrix describes the matching degree between each point;

[0069] S5: Registration and Pose Calculation: Introduce the Correspondence-ICP algorithm. According to the correspondence matching result, perform soft registration and calculate the rigid body transformation R, t from the road landmark set to the building floor plan, that is, the pose of the device relative to the building. Here, R is the rotation matrix and t is the translation vector;

[0070] S6: Model Training: Use the self-supervised training method to train the parameters of the end-to-end deep learning model composed of the processes described in S3, S4, and S5.

[0071] In the specific implementation process, mainly utilize the geometric features jointly represented by the local environment road landmark set constructed by ORB-SLAM and the building floor plan. Combining the idea of feature matching, convert the positioning problem into the registration problem between the building floor plan and the road landmark set; The three processes of input tensor construction, feature extraction and feature matching, and registration and pose calculation described in steps S3, S4, and S5 constitute an end-to-end deep learning network for registering the road landmark set and the building floor plan. This network can quickly and robustly calculate the rigid body transformation R, t based on the local environment road landmark set and the building floor plan features, thereby realizing the indoor positioning function, where Figure 2 is an example of the registration effect. After the right road landmark set goes through the implementation process of the present invention, the effect of aligning the road landmark set with the building floor plan shown on the left side of the figure is achieved;

[0072] More specifically, based on the above embodiments, each step in the above embodiments is further elaborated:

[0073] More specifically, step S1 includes the following steps:

[0074] S11: Extract the layer of the wall in the building floor plan CAD file and convert it into image data.

[0075] S12: Use the Canny edge detection algorithm to extract the edge part of the wall in the building floor plan image.

[0076] S13: Take the pixel coordinates of all edge points as the point set of the building floor plan.

[0077] S14: Convert it into a point set of the building floor plan with real dimensions according to the scale information of the building floor plan.

[0078] In the specific implementation process, converting the building floor plan into point set information with the same scale as the real world can make good use of its geometric relationship characteristics, which is an important step in converting the positioning problem into a point cloud registration problem.

[0079] As shown in Figure 3(A) and Figure 3(B), the method for constructing the K-nearest neighbor distance histogram described in step S3 includes the following steps:

[0080] S31: Query the K points with the closest Euclidean distance to each point in the point set, and sort the Euclidean distance values of each point from near to far.

[0081] S32: Divide M intervals between 0 and the maximum distance value, and calculate the number of points falling into each interval.

[0082] S33: Divide the number of points in each interval by K to obtain the frequency of points appearing in each interval, thereby constructing a frequency histogram, that is, the K-nearest neighbor distance histogram described in S3, which can be represented by an M-dimensional vector.

[0083] In the specific implementation process, the introduction of the K-nearest neighbor distance histogram is mainly to extract the characteristics of the local geometric structure of the point set with invariance, and the distance distribution is a common characteristic with invariance under rigid body transformation. Here, the method of frequency histogram is used to represent this characteristic, effectively improving the registration accuracy and robustness of the model.

[0084] As Figure 4 shown, the feature extraction and feature matching process described in step S4 includes the following steps:

[0085] S41: The input tensor undergoes feature extraction by a dynamic graph convolutional neural network (DGCNN), and outputs a matrix composed of feature vectors of each point.

[0086] S42: The point set can form a complete multi-graph, which includes edges within the point set and edges between point sets. Among them, the edges within the point set connect all vertices of the point cloud respectively, and the edges between point sets connect a certain point in the point set with all vertices of another point set.

[0087] S43: The complete multi-graph is input into a multi-layer graph attention network to achieve feature aggregation.

[0088] S44: The result of feature aggregation obtains two tensors. After inner product calculation, the transmission cost matrix C is obtained.

[0089] S45: Use the Sinkhorn optimal transport algorithm to calculate the point set matching matrix Pn×m 。

[0090] In the specific implementation process, after abstracting the original features, DGCNN extracts the key features in the input tensor, constructs a fully multi-graph structure for each point in the point set, inputs the corresponding key features into a multi-layer graph attention network to achieve feature aggregation, obtains a pair of aggregated tensors, and calculates the cost matrix through inner product calculation, which is used to measure the matching cost between points. The Sinkhorn optimal transport algorithm solves the optimal matching according to the matching cost, which is represented by a matching matrix P;

[0091] The process of the iterative closest point algorithm for the corresponding relationship described in step S5 includes the following steps:

[0092] S51: The matching matrix P obtained from S44 n×m , generates a soft mapping matrix y' through equation (1) f :

[0093] y' = (y n×3 ) T ·P n×m (1)

[0094] In the formula, y n×3 is the matrix composed of the points in the building floor plan point set;

[0095] S52: Substitute the soft mapping matrix into the ICP objective function shown in equation (2), and use the method of singular value decomposition (SVD) to calculate the rigid body transformation R, t;

[0096]

[0097] In the formula, x i is each point in the 3D landmark point set, y' (xi) is the corresponding point matched with x i , N is the number of x i , R * , t * is the final result of the rigid body transformation R, t;

[0098] In the specific implementation process, the iterative closest point algorithm for the corresponding relationship using soft registration and singular value decomposition makes the registration process differentiable, is easy to implement the end-to-end training of the model, and has a fast registration speed and accuracy;

[0099] The self-supervised training method of the deep learning model described in step S6 includes the following steps:

[0100] S61: As Figure 5As shown, a random height within a certain range is taken, and a certain number of points are vertically generated for the part of the building floor plan representing the wall surface to form a three-dimensional point cloud corresponding to the building floor plan;

[0101] S62: Perform a certain clipping, random rotation, translation, and addition of noise points on the generated three-dimensional point cloud to achieve data augmentation;

[0102] S63: Use the following loss function to train the parameters of each part of the deep learning model described in S4:

[0103] Loss=||R T R g -I|| 2 +||t - t g ||+λ||θ|| 2 (3)

[0104] where R g 、t g are respectively the true values of the rotation matrix and translation vector from the building floor plan point set in the dataset to the ORB - SLAM road point set, I is the identity matrix, λ||θ|| 2 is the L2 regularization term, and θ is the model parameter.

[0105] In the specific implementation process, a self - supervised training method is introduced. Only the dataset of the building floor plan needs to be provided to realize the training of the network, thereby reducing the data acquisition cost. Due to the data augmentation method described in S62, the trained model can well adapt to the real scene.

[0106] The present invention also provides an embodiment:

[0107] As Figure 6 , a vision - based indoor positioning system based on the prior information of the building floor plan, includes: a collection module 100, a processing unit 200, and a display module 300; wherein, the collection module 100 is connected to an external collection device, such as a camera, to collect the image data of the building floor plan to be registered; the processing unit 200 is connected to the collection module 100, obtains the image data of the building floor plan to be registered, and identifies the building floor plan to be registered according to the model for indoor positioning to perform indoor positioning; the display module 300 is connected to the processing unit 200 to display the information of indoor positioning; the collection module 100 can be a mobile phone, a camera, or other imaging devices; if the image data of the building floor plan to be registered drawn by a computer, it can directly perform data interaction with the processing unit 200; the processing unit 200 can be a computer, a tablet, or other intelligent devices; the display module 300 can be a mobile phone or a display device such as a monitor.

[0108] The present invention also provides an embodiment:

[0109] A computer program product includes a computer program carried on a computer-readable medium, and the computer program contains program code for performing the method as shown above. The computer program can be downloaded and installed from a network. When the computer program is executed by a CPU, the above functions defined in the system of the present invention are performed.

[0110] The present invention also provides an embodiment:

[0111] A computer-readable storage medium stores a computer program therein; when the computer program runs, it performs the steps of the visual indoor positioning method as described above.

[0112] In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0113] The above-disclosed are only several specific implementation scenarios of the present invention. However, the present invention is not limited thereto, and any change that can be conceived by those skilled in the art shall fall within the protection scope of the present invention.

Claims

1. A visual indoor positioning method based on the prior information of building floor plans, characterized in that, it includes: According to the image data of the building floor plan to be registered, obtain a building floor plan point set with real dimensions; Construct a three-dimensional landmark point set, and downsample the three-dimensional landmark point set and the building floor plan point set to obtain a downsampled three-dimensional landmark point set and a downsampled building floor plan point set; After respectively extracting the K-nearest neighbor distance histogram N-dimensional feature vectors of the downsampled three-dimensional landmark point set and the downsampled building floor plan point set, obtain the three-dimensional landmark input tensor corresponding to any point in the downsampled three-dimensional landmark point set and the floor plan point set input tensor of any point in the downsampled building floor plan point set; Extract the three-dimensional landmark feature vector of the three-dimensional landmark input tensor and the floor plan point set feature vector of the floor plan point set input tensor, and match the three-dimensional landmark feature vector with the floor plan point set feature vector to obtain a matching matrix from any point in the building floor plan point set to the corresponding point in the three-dimensional landmark point set; According to the matching matrix, obtain the rigid body transformation R, t from the three-dimensional landmark point set to the building floor plan point set, where R is the rotation matrix and t is the translation vector; Use the method of self-supervised training for end-to-end deep learning to obtain a model for indoor positioning; The specific steps of the "using the method of self-supervised training for end-to-end deep learning to obtain a model for indoor positioning" include: Take a random height within a certain range, generate a certain number of points perpendicular to the part of the building floor plan representing the wall surface, and form a three-dimensional point cloud corresponding to the building floor plan; Crop and / or randomly rotate and / or translate and / or add noise points to the generated three-dimensional point cloud to achieve data augmentation; Use the following loss function to train the parameters of each part of the deep learning model: Loss=||R T R g -I|| 2 +||t-t g ||+λ||θ|| 2 ; where R g and t g are the true values of the rotation matrix and the translation vector from the point set of the building floor plan in the dataset to the landmark point set of ORB-SLAM, I is the identity matrix, and λ||θ|| 2 is the L2 regularization term, and θ is the model parameter.

2. The visual indoor positioning method based on the prior information of building floor plans according to claim 1, characterized in that, The "According to the image data of the building floor plan to be registered, obtain a building floor plan point set with real dimensions" includes: Take the image data of the wall layer of the building floor plan to be registered; Extract the edges of the wall layer image; Take the pixel coordinates of all points on the edges as the building floor plan point set to be registered, and convert the building floor plan point set to be registered into a building floor plan point set with real dimensions according to the scale information of the building floor plan to be registered.

3. The visual indoor positioning method based on the prior information of building floor plans according to claim 1, characterized in that, The process of "extracting the K-nearest neighbor distance histogram N-dimensional feature vectors of the downsampled three-dimensional landmark point set and the downsampled building floor plan point set" includes: Query the K points with the closest Euclidean distance to any point in the downsampled three-dimensional landmark point set and the downsampled building floor plan point set, and sort the Euclidean distance values of each point from near to far; Divide M intervals between 0 and the maximum distance value, and calculate the number of points falling into each interval; The number of points in each interval is divided by K to obtain the frequency of point occurrences in each interval, thereby forming a frequency histogram, that is, the K-nearest neighbor distance histogram, which is represented by an N-dimensional vector.

4. A visual indoor positioning method based on prior information of architectural floor plans according to claim 1, characterized in that the process of "obtaining a 3D road sign input tensor corresponding to any point in the downsampled 3D road sign point set and a floor plan point set input tensor of any point in the downsampled architectural floor plan point set" includes: concatenating the N-dimensional feature vector of the K-nearest neighbor distance histogram with the 3D coordinate vectors corresponding to each point in the downsampled 3D road sign point set to obtain the 3D road sign input tensor; concatenating the N-dimensional feature vector of the K-nearest neighbor distance histogram with the 3D coordinate vectors corresponding to each point in the downsampled architectural floor plan point set to obtain the floor plan point set input tensor.

5. A visual indoor positioning method based on prior information of architectural floor plans according to claim 1, characterized in that the specific steps of "extracting the 3D road sign feature vector of the 3D road sign input tensor and the floor plan point set feature vector of the floor plan point set input tensor, and matching the 3D road sign feature vector with the floor plan point set feature vector" include: the 3D road sign input tensor and the floor plan point set input tensor undergo feature extraction through a dynamic graph convolutional neural network, and a matrix composed of point feature vectors is output; a complete multi-graph is formed by the point set, which includes intra-point set edges and inter-point set edges, where the intra-point set edges connect all vertices of the point clouds respectively, and the inter-point set edges connect a certain point in the point set with all vertices of another point set; the complete multi-graph is input into a multi-layer graph attention network to achieve feature aggregation; through the feature aggregation, a 3D road sign tensor and a floor plan point set tensor are obtained, and after inner product calculation, a transmission cost matrix is obtained; the matching matrix is obtained through the Sinkhorn optimal transport algorithm.

6. A visual indoor positioning method based on prior information of architectural floor plans according to claim 1, characterized in that the specific steps of "obtaining the rigid body transformation R, t of the 3D road sign point set to the architectural floor plan point set according to the matching matrix, where R is the rotation matrix and t is the translation vector" include: The matching matrix P n×m , generates a soft mapping matrix y' through the following formula f : y'=(y n×3 ) T ·P n×m ; where y n×3 is a matrix formed by the point set of the building floor plan; substituting the soft mapping matrix into the ICP objective function shown in the following formula, and using the method of singular value decomposition to calculate the rigid body transformation R, t; where x i is each point of the three-dimensional road landmark set, is the corresponding point matched with x i , N is the quantity of x i , R * , t * is the final result of the rigid body transformation R, t.

7. A visual indoor positioning system based on prior information of architectural floor plans, which applies the visual indoor positioning method based on prior information of architectural floor plans according to any one of claims 1-6, characterized in that it includes: a collection module, connected to an external collection device, for collecting image data of the architectural floor plan to be registered; a processing unit, connected to the collection module, for obtaining the image data of the architectural floor plan to be registered and identifying the architectural floor plan to be registered according to the model for indoor positioning, and performing indoor positioning; a display module, connected to the processing unit, for displaying information on indoor positioning.

8. A visual indoor positioning electronic device based on prior information of architectural floor plans, characterized in that, it includes: a storage medium for storing a computer program; a processing unit that exchanges data with the storage medium and, when performing indoor positioning, executes the computer program through the processing unit to perform the steps of the visual indoor positioning method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that: the computer-readable storage medium stores a computer program; when the computer program runs, it executes the steps of the visual indoor positioning method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method and system for determining at least one image feature in at least one image

    CN109547753A

  • Visual odometer method based on end-to-end semi-supervised generative adversarial network

    CN110335337A