An indoor floor plan positioning method, system, device and storage medium

By integrating multi-view depth features and semantic features in the indoor positioning method, and calculating the posterior probability distribution using Bayesian rules and histogram filters, the problems of positioning trajectory drift, insufficient adaptability of non-upright postures and low utilization of semantic topological relationships in the prior art are solved, and the indoor positioning effect with high accuracy, robustness and fast response are achieved.

CN119963651BActive Publication Date: 2025-06-20SEVNCE ROBOTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510455316.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-06-20
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing indoor positioning method based on floor plan lacks robustness in the fusion of spatiotemporal features of continuous image sequences, resulting in frequent positioning trajectory drifts, insufficient adaptability to non-upright camera postures, and failure to effectively mine the semantic topological relationships implicitly in the plan picture, resulting in low utilization of environmental constraint information and difficult to meet the millisecond-level response requirements in large-scale scenarios.

Method used

By acquiring indoor scene images from multiple perspectives, extracting monocular depth features and multi-view depth features, fusing depth features and semantic features, and using Bayesian rules and histogram filters to calculate the posterior probability distribution of the pose of the target to be positioned, and obtaining the positioning results.

Benefits of technology

It improves the accuracy, efficiency and robustness of indoor positioning in complex scenarios, reduces positioning blur and error, improves the adaptability to non-upright camera postures and the utilization of semantic topological relationships, and meets the millisecond-level response requirements in large-scale scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963651B_ABST
    Figure CN119963651B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of indoor positioning technology, and relates to an indoor floor plan positioning method, system, device and storage medium; wherein, an indoor floor plan positioning method includes: acquiring multi-view indoor scene images collected by a target to be positioned; extracting monocular depth features of the indoor scene images, and parsing the monocular depth features of the multi-view indoor scene images to obtain multi-view depth features; fusing the monocular depth features and the multi-view depth features and processing the fused depth features according to a preset rule to obtain a first posterior probability distribution of the pose of the target to be positioned; extracting semantic features of the indoor scene images, and processing the semantic features according to a preset rule to obtain a second posterior probability distribution of the pose of the target to be positioned; determining the positioning result of the target to be positioned according to the first posterior probability distribution and the second posterior probability distribution. The present invention can improve the accuracy, efficiency and robustness of indoor positioning in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of indoor positioning, and in particular, to an indoor floor plan positioning method, system, device, and storage medium. Background Art

[0002] In recent years, the use of robots with autonomous navigation capabilities has become increasingly common; the autonomous navigation ability of robots refers to the ability of robots to perceive in known or unknown environments using sensor data and autonomously plan paths based on the perception data. In indoor environments, precise positioning is crucial for many applications such as virtual reality (VR) and robot autonomous navigation.

[0003] Traditional technical solutions mainly rely on pre-built 3D environment models or high-precision sensor databases to achieve positioning. Such methods not only require a large amount of storage resources to maintain environmental data but also suffer from the problem of lagging model updates due to dynamic changes in the scene, significantly increasing the system operation and maintenance costs. In recent years, floor plan-based positioning methods have gradually become an alternative. Floor plan-based positioning methods achieve device positioning by extracting the geometric and semantic features of building floor plans. However, existing floor plan-based positioning methods still have significant defects: First, there is a lack of robust algorithms in the spatio-temporal feature fusion of continuous image sequences, resulting in frequent occurrence of positioning trajectory drift phenomena; second, the adaptability to non-vertical camera poses (such as tilting, pitching, etc.) is insufficient, restricting the compatibility ability for complex observation perspectives; third, the implicit semantic topological relationships in the floor plan (such as porch connectivity, functional area division, etc.) are not effectively mined, resulting in low utilization rate of environmental constraint information; fourth, existing solutions have deficiencies in aspects such as the balance between positioning accuracy improvement and real-time performance, and it is difficult to meet the millisecond-level response requirements in large-scale scenarios. Summary of the Invention

[0004] In order to improve the accuracy, efficiency, and robustness of indoor positioning in complex scenarios to meet the requirements of positioning accuracy and real-time performance for different indoor application scenarios, the present application provides an indoor floor plan positioning method, system, device, and storage medium.

[0005] An indoor floor plan positioning method provided by the present invention includes:

[0006] Obtain multi-view indoor scene images collected by a target to be positioned;

[0007] Extract the monocular depth features of the indoor scene images and analyze the monocular depth features of the multi-view indoor scene images to obtain multi-view depth features;

[0008] Fuse the monocular depth feature and the multi-view depth feature to obtain a fused depth feature, and process the fused depth feature according to a preset rule to obtain a first posterior probability distribution of the pose of the target to be located, where the pose represents the position and orientation of the target to be located in the indoor space;

[0009] Extract the semantic feature of the indoor scene image, and process the semantic feature according to a preset rule to obtain a second posterior probability distribution of the pose of the target to be located;

[0010] Weight the first posterior probability distribution and the second posterior probability distribution according to a preset ratio to obtain a weighted posterior probability distribution, and input the weighted posterior probability distribution into a histogram filter, and the histogram filter outputs a three-dimensional probability volume of the pose of the target to be located to obtain the positioning result of the target to be located.

[0011] Optionally, the extracting the monocular depth feature of the indoor scene image and parsing the monocular depth features of the indoor scene images of multiple perspectives to obtain multi-view depth features includes:

[0012] Extract the monocular depth features of the indoor scene images of multiple perspectives respectively;

[0013] Extract the column features of the monocular depth feature, and aggregate the column features of the monocular depth features of the indoor scene images of multiple perspectives to obtain cross-view features;

[0014] Calculate the depth value of each pixel point of the indoor scene floor plan according to the feature variance of the cross-view features to obtain multi-view depth features.

[0015] Optionally, the fusing the monocular depth feature and the multi-view depth feature to obtain a fused depth feature, and processing the fused depth feature according to a preset rule to obtain a first posterior probability distribution of the pose of the target to be located includes:

[0016] Obtain the pose information of the indoor scene image, where the pose information is used to represent the position and angle of the camera when shooting the indoor scene image;

[0017] Determine the relative pose between image frames of the multi-view indoor scene images according to the pose information of the indoor scene images of multiple perspectives;

[0018] Determine the reference weights of the multi-view depth feature and the monocular depth feature during feature fusion according to the relative pose between image frames, and fuse the multi-view depth feature and the monocular depth feature according to the reference weights to obtain a fused depth feature;

[0019] Process the fused depth feature using Bayes' rule to obtain a first posterior probability distribution of the pose of the target to be located.

[0020] Optionally, the calculation formula of the first posterior probability distribution of the pose of the target to be located is:

[0021] , ;

[0022] wherein, represents the first posterior probability distribution of the pose of the target to be located, represents the probability distribution of the depth of the predicted pixel points corresponding to the monocular depth features, represents the probability distribution of the depth of the predicted pixel points corresponding to the multi-view depth features, represents the reference weight of the monocular depth features, represents the upsampling operation, represents the reference weight of the multi-view depth features.

[0023] Optionally, extracting the semantic features of the indoor scene image and processing the semantic features according to a preset rule to obtain the second posterior probability distribution of the pose of the target to be located includes:

[0024] performing semantic segmentation on the indoor scene image to obtain the semantic features of the indoor scene image;

[0025] determining the semantic recognition labels of each pixel point of the indoor scene image according to the semantic features;

[0026] selecting a semantic likelihood field according to the semantic recognition labels of the pixel points and performing semantic ray projection according to the semantic likelihood field to determine the second posterior probability distribution of the pose of the target to be located.

[0027] Optionally, the method further includes:

[0028] obtaining a floor plan of the indoor scene, and the floor plan of the indoor scene is provided with reference semantic labels;

[0029] calculating the matching value between the semantic recognition label of the pixel point and the reference semantic label;

[0030] constructing a semantic likelihood field according to the matching value.

[0031] Optionally, the indoor scene image is an RGB image. After obtaining the multi-view indoor scene images and before extracting the monocular depth features of the indoor scene image, the method further includes:

[0032] adjusting the orientation of the indoor scene image to convert the indoor scene image into a gravity-aligned image.

[0033] To solve the above problems, the present invention also provides an indoor floor plan positioning system, and the system includes:

[0034] an acquisition module, configured to acquire multi-view indoor scene images collected by a target to be located;

[0035] A feature extraction module, configured to extract monocular depth features of indoor scene images, and parse the monocular depth features of multi-view indoor scene images to obtain multi-view depth features;

[0036] A first processing module, configured to fuse the monocular depth features and the multi-view depth features to obtain fused depth features, and process the fused depth features according to a preset rule to obtain a first posterior probability distribution of the pose of the target to be located, where the pose represents the position and orientation of the target to be located in the indoor space;

[0037] A second processing module, configured to extract semantic features of indoor scene images, and process the semantic features according to a preset rule to obtain a second posterior probability distribution of the pose of the target to be located;

[0038] An output module, configured to weight the first posterior probability distribution and the second posterior probability distribution according to a preset ratio to obtain a weighted posterior probability distribution, and input the weighted posterior probability distribution into a histogram filter, and the histogram filter outputs a three-dimensional probability volume of the pose of the target to be located to obtain the positioning result of the target to be located.

[0039] To solve the above problems, the present invention further provides an electronic device, where the electronic device includes:

[0040] At least one processor; and,

[0041] A memory communicatively connected to the at least one processor; wherein,

[0042] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the indoor floor plan positioning method described above.

[0043] To solve the above problems, the present invention further provides a computer-readable storage medium, where at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is executed by a processor in an electronic device to implement the indoor floor plan positioning method described above.

[0044] In summary, the present application includes the following beneficial technical effects:

[0045] When processing the depth images of indoor scenes from multiple perspectives, the monocular depth features of the indoor scene images from multiple perspectives are extracted respectively, and the multi-view depth features are extracted. The first posterior probability distribution of the pose of the target to be located is calculated based on the monocular depth features and the multi-view depth features, and the positioning result of the target to be located is calculated using the first posterior probability distribution, reducing positioning ambiguity and error; an efficient histogram filter is used to fuse the positioning information calculated from the depth features of the indoor scene images and the positioning information calculated from the semantic features to obtain the final positioning information of the target to be located; when calculating the positioning result of the target to be located, the semantic information of each pixel point in the indoor scene map is referred to, improving the accuracy, efficiency and robustness of indoor positioning in complex scenes. Brief Description of the Drawings

[0046] Figure 1 It is a schematic flowchart of the indoor floor plan positioning method provided by an embodiment of the present invention;

[0047] Figure 2 It is a system flowchart of the indoor floor plan positioning method provided by an embodiment of the present invention;

[0048] Figure 3 It is a flowchart of the steps of extracting the monocular depth features of the indoor scene images and parsing the monocular depth features of the indoor scene images from multiple perspectives to obtain multi-view depth features provided by an embodiment of the present invention;

[0049] Figure 4 It is a schematic structural diagram of an electronic device for implementing the indoor floor plan positioning method provided by an embodiment of the present invention.

[0050] Reference Numerals: 10, processor; 11, memory; 12, communication bus; 13, communication interface.

[0051] The realization, functional features and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. Detailed Embodiments

[0052] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as a limitation of the present invention.

[0053] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0054] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0055] Refer to Figure 1 and Figure 2 , which is a schematic flow chart of the indoor floor plan positioning method provided by an embodiment of the present invention. In this embodiment, the indoor floor plan positioning method includes:

[0056] S1. Obtain multi-view indoor scene images collected by the target to be located.

[0057] The indoor scene images are RGB images. In this embodiment, a wide-angle camera is used to obtain a sequence of indoor scene images. The horizontal field of view angle of the wide-angle camera is 108°, so as to ensure wide coverage of the surrounding environment. In addition, the target to be located can also use its own sensor system (wheel speed sensor) to obtain motion data to simulate the odometer function, which serves as an auxiliary for subsequent indoor floor plan positioning.

[0058] In the preferred implementation manner of this embodiment, after obtaining the multi-view indoor scene images and before extracting the monocular depth features of the indoor scene images, the following operations are performed: adjust the direction of the indoor scene images to convert the indoor scene images into gravity-aligned images.

[0059] When converting the indoor scene images into gravity-aligned images, first calculate the rotation matrix through the roll angle and the pitch angle to facilitate subsequent rotation of the indoor scene images from the current pitch attitude to the horizontal attitude.

[0060] After that, the pixel coordinates in the original image are converted to the coordinates in the gravity-aligned image using the homography matrix; the homography matrix is ​​a 3×3 matrix used in computer vision to describe the projection transformation relationship between two planes, representing the mapping relationship of the same plane under different viewing angles. The expression of the homography matrix is ,in, is the camera intrinsic parameter matrix, is the inverse matrix of the camera intrinsic parameter matrix, are the homogeneous image coordinates of the original pixel, are the corresponding pixel coordinates in the gravity-aligned image.

[0061] Then, the pixels that cannot be seen at the current pitch angle are filtered out, and the invisible pixels are masked out, finally obtaining a gravity-aligned image. It should be explained that after gravity alignment, the image is rotated to a perspective based on the horizontal ground, and some pixels originally at the top of the image (such as the ceiling or high content) may be moved out of the image frame and become invisible. The pixels that are rotated out are marked as invisible pixels (masked out), avoiding the network from trying to make meaningless depth estimates for these pixels.

[0062] After converting the indoor scene image into a gravity-aligned image, the buildings in the gravity-aligned image appear vertical (equivalent to a photo taken by the camera in a horizontal posture). The state of the tilted buildings in the image is corrected to ensure the accuracy and robustness of feature matching and semantic segmentation in the subsequent depth estimation process.

[0063] S2. Extracting monocular depth features of indoor scene images, and parsing monocular depth features of multi-view indoor scene images to obtain multi-view depth features.

[0064] Reference Figure 3 , extracting monocular depth features of indoor scene images, and parsing monocular depth features of multi-view indoor scene images to obtain multi-view depth features, including:

[0065] S21. Extract monocular depth features of multi-view indoor scene images respectively.

[0066] Specifically, ResNet (Residual Neural Network) and Attention mechanism networks can be used to extract features of single-frame indoor scene images. The Attention mechanism network can mask invisible pixels and output a depth hypothesis probability distribution. The expectation of the depth hypothesis probability distribution is used as the floor plan depth prediction value, and the floor plan depth prediction value is used to construct equiangular ray scanning positioning.

[0067] S22. Extract the column features of the monocular depth features, and aggregate the column features of the monocular depth features of multi-view indoor scene images to obtain cross-view features.

[0068] The cross-view features refer to the column features extracted from images of different perspectives, and these features can correspond to each other for subsequent depth estimation. Specifically, a predicted depth of each pixel point in this frame of the image can be obtained through a single-frame indoor scene image. Through multi-view depth estimation, the predicted depth of each pixel point in the reference-frame indoor scene image can be obtained (the reference frame and the single-frame RGB image are the same frame).

[0069] Inspired by multi-view stereo vision, the MVS network is used to estimate the depth of the planar graph from multiple frames of images. The full English name of the MVS network is Multi-View Stereo network, and the Chinese interpretation is multi-view stereo network. The multi-view stereo network is a computer vision technology that uses images from multiple perspectives to restore the three-dimensional structure of the scene. This technology is usually used to estimate depth information from multiple two-dimensional images and then reconstruct a three-dimensional model. The MVS network estimates the depth of each pixel point by analyzing the image features from different perspectives and calculating the pixel or feature point matching between views, and finally realizes the reconstruction of the three-dimensional scene.

[0070] In step S22, first extract the column features of the single-frame indoor scene image one by one, aggregate the column features of the indoor scene images of different views through plane sweep, construct a 2D cost distribution according to the cross-view feature variance, and then calculate the final depth through soft-argmin to supplement multi-view geometric clues for positioning. The soft-argmin algorithm is a differentiable optimal value estimation method that realizes continuous space regression by probabilistically discretizing candidate values.

[0071] Specifically, after obtaining the feature variances under different depth hypotheses, the feature variances directly correspond to the 2D cost distribution. For each pair of matching feature points, a cost value is calculated through the feature variance (calculated by means of weighted summation or taking the mean). The cost value is usually based on the difference between feature descriptions. The smaller the difference, the lower the cost, indicating that the matching is more reliable and the probability that this depth hypothesis is correct is higher.

[0072] S23. Calculate the depth value of each pixel point in the indoor scene planar graph according to the feature variance of the cross-view features to obtain multi-view depth features.

[0073] S3. Fuse the monocular depth features and the multi-view depth features to obtain fused depth features, and process the fused depth features according to a preset rule to obtain the first posterior probability distribution of the pose of the target to be located.

[0074] Among them, the pose of the target to be located represents the position and orientation of the target to be located in the indoor space. The target to be located can be an autonomous navigation robot, an autonomous navigation UAV, a VR device, etc. Specifically, the orientation of the target to be located in the indoor space refers to the front-facing direction or the moving direction of the target to be located. Fusing monocular depth features and multi-view depth features for depth estimation can make full use of multi-modal clues and reduce positioning ambiguity and errors.

[0075] Referring to Figure 2 , the specific steps of S3 include:

[0076] S31. Obtain the pose information of the indoor scene image, where the pose information is used to represent the position of the camera when shooting the indoor scene image and the angle between the lens of the camera and the horizontal ground when shooting the indoor scene image.

[0077] S32. Determine the relative pose between image frames of the multi-view indoor scene image according to the pose information of the multi-view indoor scene images.

[0078] S33. Determine the reference weights of the multi-view depth features and the monocular depth features during feature fusion according to the relative pose between image frames, and fuse the multi-view depth features and the monocular depth features according to the reference weights to obtain the fused depth features.

[0079] S34. Process the fused depth features using the Bayesian rule to obtain the first posterior probability distribution of the pose of the target to be located.

[0080] The calculation formula for the first posterior probability distribution of the pose of the target to be located is:

[0081] , ;

[0082] Among them, represents the first posterior probability distribution of the pose of the target to be located, represents the probability distribution of the depth of the predicted pixel point corresponding to the monocular depth feature, represents the probability distribution of the depth of the predicted pixel point corresponding to the multi-view depth feature, represents the reference weight of the monocular depth feature, represents the upsampling operation, and the upsampling operation ensures the effectiveness of the addition; represents the reference weight of the multi-view depth feature; the first posterior probability distribution of the pose of the target to be located

[0083] Assume that the current pose of the camera in the positioning model corresponding to the target to be located is , represents the Coordinate position, represents the coordinate position of the camera on the two-dimensional plane, represents the rotation angle of the camera relative to the reference direction (the axis of the global coordinate system); Based on the acquired data and the semantic plane occupancy map, the camera pose is located by fusing multi-modal features.

[0084] Monocular depth estimation is independent of camera motion, but is prone to scale ambiguity problems; multi-view stereo methods can provide the correct scale, but rely on sufficient baselines and camera overlaps. Based on these observations, we use an MLP network to make a soft selection from two predictions (single-frame indoor scene images and multi-frame indoor scene images). The full English name of MLP is Multi-Layer Perceptron, which is a multi-layer perceptron; the MLP network can perform non-linear feature transformation by stacking fully connected layers and is widely used in classification, regression, and deep learning feature encoding.

[0085] The MLP network adaptively weights and fuses the probability distributions of monocular depth estimation and multi-view depth estimation according to the relative pose between image frames, the average depth prediction value of monocular depth features, and multi-view depth features, improving the accuracy and reliability of depth estimation. Specifically, the predicted floor plan depth is used as the observation, and the following observation model is adopted:

[0086]

[0087] where is the pose at the floor plan ray, is the ray interpolated from the floor plan depth prediction, is a constant factor; represents the observation of the camera at time is the time index, and the value range of .

[0088] refers to the depth prediction at the camera pose at time t obtained by complementary fusion of monocular depth and multi-view depth according to the camera observation introduced above; the ray interpolated from this depth prediction can be compared with the floor plan rays at possible camera poses in the entire floor plan to obtain a probability distribution for each pose.

[0089] In this embodiment, the calculation method of the probability distribution is:

[0090] Use the relative pose between frames, that is, the self-motion as the transition model ;

[0091] Among them, and represent the self - motion of time , represents the transfer noise of time . The operator means applying the self - motion to the state. represents the time The self - motion at coordinate position translation amount, represents the time The self - motion at coordinate position translation amount, represents the time The self - motion relative to the reference direction (the axis of the global coordinate system) rotation angle change amount; represents the time The motion noise at the x - coordinate position translation amount, represents the time The motion noise at the y - coordinate position translation amount and represents the time The motion noise relative to the reference direction (the x - axis of the global coordinate system) rotation angle change amount.

[0092] Further assume that the transfer noise obeys a Gaussian distribution, and the transfer probability is expressed as , represents the transpose; is the inverse matrix of the covariance matrix, representing the uncertainty of the state transition error. The covariance matrix Σ is a symmetric positive - definite matrix, and its inverse matrix is used to measure the magnitude of the error vector; the state transition probability is from the current state and the self - motion to derive the probability distribution of the next - moment state ; is used to evaluate the uncertainty in the modeling state transition. It is usually modeled as a Gaussian distribution, whose mean is the predicted state , and the covariance is used to describe the uncertainty.

[0093] We model as the covariance of a Gaussian distribution, represents the variance in the coordinate direction, represents the variance in the the variance corresponding to the rotation angle of the axis); After applying Bayes' rule, we get , where is the normalization factor, is the pose space; is from the observed state of the camera and the self-motion at time to derive the probability distribution of the state at the next time .

[0094] S4. Extract the semantic features of the indoor scene image, and process the semantic features according to the preset rules to obtain the second posterior probability distribution of the pose of the target to be located.

[0095] Extracting the semantic features of the indoor scene image and processing the semantic features according to the preset rules to obtain the second posterior probability distribution of the pose of the target to be located, including:

[0096] S41. Perform semantic segmentation on the indoor scene image to obtain the semantic features of the indoor scene image;

[0097] Use the convolutional neural network (CNN) encoder-decoder architecture for semantic segmentation to identify semantic objects such as walls, doors, and windows in the input RGB image (indoor scene image). This semantic information is fused with geometric features to enhance the scene understanding and discrimination ability of the positioning.

[0098] S42. Determine the semantic recognition labels of each pixel point in the indoor scene image according to the semantic features;

[0099] S43. Select the semantic likelihood field according to the semantic recognition labels of the pixel points, and perform semantic ray projection according to the semantic likelihood field to determine the second posterior probability distribution of the pose of the target to be located.

[0100] In the preferred implementation manner of this embodiment, the indoor floor plan positioning method further includes:

[0101] S401. Obtain the indoor scene floor plan, and the indoor scene floor plan is provided with reference semantic labels;

[0102] S402. Calculate the matching value between the semantic recognition label of the pixel point and the reference semantic label;

[0103] S403. Construct the semantic likelihood field according to the matching value.

[0104] Reference semantic tags are known semantic information (such as walls, doors, windows, etc.) extracted from indoor floor plans, while semantic recognition tags are semantic information recognized from real-time captured indoor scene images through a semantic segmentation network. The calculation of the matching value can be achieved by comparing the consistency of the semantic recognition tags of each pixel with the reference semantic tags, using a semantic similarity metric. The higher the matching value, the more consistent the semantic information of the pixel is with the reference information in the floor plan, and thus the greater the contribution of the pixel to localization when constructing the semantic likelihood field.

[0105] The semantic likelihood field is a probability model that combines environmental semantic information and is used to enhance the localization and navigation capabilities of a robot or system. It provides richer constraints for localization by calculating the probability of observing specific semantic tags (such as walls, doors, windows, etc.) at a given position and orientation. Compared with traditional geometric information-based localization methods, the semantic likelihood field can more accurately identify and utilize key landmarks in the environment, thereby improving the localization accuracy and robustness, especially in complex environments or when the map is inaccurate.

[0106] Specifically, reference semantic tags are extracted from existing indoor scene floor plans. In this embodiment, the reference semantic tags include walls, doors, and windows because walls, doors, and windows are easy to automatically extract from the floor plan and are crucial for human localization, thereby improving the localization effect.

[0107] To enable the indoor scene floor plan with reference semantic tags to be read by the robot, the indoor scene floor plan is converted into an occupancy grid map. The occupancy grid is a two-dimensional representation of the world, and each cell in the occupancy grid has an occupancy probability, which is determined by its normalized gray value.

[0108] Combining semantic likelihood and geometric feature likelihood, and updating the posterior probability of the camera pose using Bayes' rule to enhance the robustness of localization in complex environments and when the map is inaccurate. Here, a joint method is used, which can incorporate semantic information using the likelihood field in the presence of semantic tags. More importantly, it can also use ray casting within the likelihood field to operate without distance measurements.

[0109] The MCL motion model is usually represented by the distribution where the full English name of MCL is Monte Carlo Localization, and the Chinese interpretation is Monte Carlo localization; represents the probability that the system transfers to state at the next time step given the current state and the odometry measurement . At time One of the possible states, is an index of the possible states, representing different state hypotheses; At time odometry measurement value, At time the current state, is the index of the current state.

[0110] The previous set of particles is propagated to the current set of particles using the odometry measurement value .

[0111]

[0112] Among them, is a normalization factor, is the set containing each cell in the map, At time one of the possible states. Here, it is allowed to consider the two likelihoods as independent, and the motion is defined in the same way as RMCL. The full English name of RMCL is Rao - Blackwellized Monte Carlo Localization, which is a localization algorithm based on the Monte Carlo method.

[0113] The RMCL algorithm combines the random sampling characteristics of the Monte Carlo method and the optimization effect of the Rao - Blackwell theorem to improve the accuracy and efficiency of localization. In RMCL, the algorithm maintains a set of particles representing the possible positions of the robot, updates the weights of these particles according to sensor data and control inputs, and then generates a new set of particles through a resampling step to reflect a more accurate localization probability distribution. The prior is the occupancy likelihood of the cell containing , that is, , represents the occupancy probability corresponding to the floor plan at the observation at time t; is a random variable representing the probability of observing the semantic information o in state ; specifically, represents the state at time t, usually including position and orientation (for example, the position and orientation of the robot in the environment); o represents the observed semantic information, such as the reference semantic labels recognized from an image (such as walls, doors, windows, etc.); v is the semantic feature representing the matching degree between the observed semantic information o and a specific semantic label v in state .

[0114] The likelihood field model calculates a distance map. For each cell​ , calculate the distance to the nearest occupied cell and store, for each cell in the map The set of distances from each cell to the nearest occupied cell is the matching value between the semantic recognition label of the pixel point and the reference semantic label.

[0115] In the formula,[[]]END]] represents the distance to the landmark nearest to the landmark in the set of landmarks that meet the conditions; represents finding the minimum value among all possible landmarks , that is, finding the landmark nearest to ; it should be noted that the landmark is the semantic recognition label of the pixel point, and the landmark is the reference recognition label (such as doors, windows, and walls).

[0116] represents the distance between the landmark and the landmark , usually using the Euclidean distance or other distance metrics;

[0117] represents the occupancy rate of the landmark ;

[0118] represents a threshold for filtering landmarks ; only when is satisfied, the landmark will be taken into account.

[0119] When an observation is received, estimate the end point and use it as the index of the distance map. Assuming a Gaussian error distribution, the weight of each particle can be estimated as

[0120]

[0121] represents the th observation at time ;

[0122] represents one of the possible states at time , where is the index of the state;

[0123] contains the set of each cell in the map;

[0124] Indicates an observation and calculates the distance to the nearest occupied cell;

[0125] Represents the standard deviation of the distance error, which defines the width of the Gaussian distribution.

[0126] Is part of the probability density function of the Gaussian distribution, used to calculate the given distance difference and the standard deviation when the probability.

[0127] For each reference semantic label (door, wall, window) present in the floor plan, we can calculate a distance map that stores the shortest distance to cells with the same label. Formally, for each map cell , we can estimate the distance to the nearest cell of each label as

[0128] , calculating the shortest distance among cells with the same semantic label.

[0129] Represents a landmark occupancy rate; Represents a threshold for filtering landmarks . Only when is the case, the landmark will be taken into account.

[0130] Are the distances to the nearest wall, door, and window respectively, thus constructing likelihood fields for three semantics.

[0131] When we receive the observation , we use the label to decide which semantic likelihood field to use for semantic ray projection.

[0132] In this embodiment, there are three semantic likelihood fields, namely door, wall, and window; then when the part of the semantic label (reference semantic label) in my observed image is a door, the semantic likelihood field of the door is used for semantic matching to obtain a matching probability; the matching process is to project using semantic rays, where the door pixel points are 1 and the others are 0, for equiangular ray semantic matching with a sliding window. Similarly, the other two semantic likelihood fields (wall / window) are matched in the same way.

[0133]

[0134] After obtaining the observation probability, then calculate the weights of each particle , and the calculation formula is , where is the weight at the previous moment. After normalization, the new weight is used for state estimation, that is , and the state is accurately estimated in the current positioning. The two work together to improve the performance of the positioning system, enabling the robot to accurately position and navigate in a complex environment.

[0135] At time , the weight of the -th particle;

[0136] At time , the weight of the -th particle;

[0137] Given the state and the set of each cell in the map , the probability of observing the landmark ;

[0138] Sum the observation probabilities of all particles for normalizing the weights; is the particle index, the total number of particles;

[0139] Given the odometry input , the current observation , the previous state and the set of each cell in the map , the probability that the system is in the state at time ;

[0140] By multiplying the state of each particle by its corresponding weight and summing, estimate the state of the system at time ;

[0141] Since a lower prior is more discriminative, associated with the prior of each label , not only because it is a parameter that requires less adjustment, but also because it implicitly makes observing rare landmarks more beneficial than observing common landmarks.

[0142] Semantic information fusion enhances the tolerance to environmental changes and map errors. When dealing with complex scenarios such as inaccurate maps, repeated scene structures, dynamic obstacles, and lighting changes, it enables stable and accurate positioning, maintains system performance, and reduces the risks of positioning failure and drift.

[0143] S5. Weight the first posterior probability distribution and the second posterior probability distribution according to a preset ratio to obtain a weighted posterior probability distribution, and input the weighted posterior probability distribution into a histogram filter. The histogram filter outputs a three-dimensional probability volume of the pose of the target to be located, and the positioning result of the target to be located is obtained.

[0144] The three-dimensional probability volume refers to the digital expression of the probability density field in three-dimensional space, quantifying the geometric / semantic existence probability of each voxel. Through the three-dimensional probability volume, the position of the target to be located in the indoor scene can be known, so as to facilitate the positioning of the target to be located when the indoor scene is complex.

[0145] The specific steps of fusing the probability likelihood distribution (the second posterior probability distribution) obtained by semantic ray projection into the posterior probability distribution (the first posterior probability distribution) obtained by Bayesian rule are as follows:

[0146] Sum the weights calculated by the Bayesian likelihood function and the weights calculated by semantic ray projection according to a certain ratio. The weight of the Bayesian likelihood function is and the weight of semantic ray projection is .

[0147] The expression of the weighted probability distribution is:

[0148]

[0149] represents the first posterior probability distribution, represents the second posterior probability distribution.

[0150] Adopt an efficient histogram filter to fuse the positioning information calculated by Bayesian rule and the positioning information calculated by semantic features, improving the positioning stability and accuracy. The filter represents the pose posterior as a three-dimensional probability volume, decomposes translation and rotation according to the relative pose between frames, and realizes efficient transfer and update through grouped convolution in different directions, quickly converging the pose estimation probability distribution.

[0151] Based on the same inventive concept, an embodiment of the present invention provides an indoor floor plan positioning system.

[0152] The indoor floor plan positioning system described in the present invention can be installed in an electronic device. According to the functions achieved, the indoor floor plan positioning system includes an acquisition module, a feature extraction module, a first processing module, a second processing module, and an output module.

[0153] The acquisition module is capable of acquiring indoor scene images with multiple perspectives collected by the target to be located; the feature extraction module is capable of extracting monocular depth features of the indoor scene images, and parsing the monocular depth features of the indoor scene images with multiple perspectives to obtain multi-view depth features;

[0154] The first processing module is capable of fusing the monocular depth features and the multi-view depth features to obtain fused depth features, and processing the fused depth features according to a preset rule to obtain a first posterior probability distribution of the pose of the target to be located, where the pose represents the position and orientation of the target to be located in the indoor space; the second processing module is capable of extracting semantic features of the indoor scene images, and processing the semantic features according to a preset rule to obtain a second posterior probability distribution of the pose of the target to be located; the output module is capable of weighting the first posterior probability distribution and the second posterior probability distribution according to a preset ratio to obtain a weighted posterior probability distribution, and inputting the weighted posterior probability distribution into a histogram filter, and the histogram filter outputs a three-dimensional probability volume of the pose of the target to be located to obtain the positioning result of the target to be located.

[0155] The module described in the present invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0156] The various change methods and specific examples in the indoor floor plan positioning method provided in the above embodiments are equally applicable to the indoor floor plan positioning system in this embodiment. Through the foregoing detailed description of the indoor floor plan positioning method, those skilled in the art can clearly know the implementation method of the indoor floor plan positioning system in this embodiment. For the sake of brevity of the specification, it will not be described in detail here.

[0157] This application also discloses an electronic device, such as Figure 4 shown, which is a schematic structural diagram of an electronic device for the indoor floor plan positioning method provided by an embodiment of the present invention. The electronic device may include at least one processor 10, a memory 11 communicatively connected to the at least one processor, a communication bus 12, and a communication interface 13, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as an indoor floor plan positioning method program.

[0158] Among them, in some embodiments, the processor 10 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. By running or executing programs or modules stored in the memory 11 (such as executing the indoor floor plan positioning method, etc.), and by calling the data stored in the memory 11, it performs various functions of the electronic device and processes data.

[0159] The memory 11 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical discs, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device, such as the mobile hard disk of the electronic device. In other embodiments, the memory 11 may also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can be used not only to store application software installed on the electronic device and various types of data, such as the code of the indoor floor plan positioning method program, etc., but also to temporarily store data that has been output or will be output.

[0160] The communication bus 12 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to enable connection communication between the memory 11 and at least one processor 10, etc.

[0161] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.

[0162] Figure 4 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 4 The shown structure does not constitute a limitation on the electronic device, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0163] For example, although not shown, the electronic device may further include a power source (such as a battery) for powering each component. Preferably, the power source may be logically connected to at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The electronic device may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0164] It should be understood that the embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0165] Furthermore, if the integrated module / unit of the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. The computer-readable storage medium may be volatile or non-volatile.

[0166] An embodiment of the present application provides a computer-readable storage medium, for example, including: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory). The computer-readable storage medium stores a computer program that can be loaded and executed by a processor to perform the indoor floor plan positioning method in the above embodiment.

[0167] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", "one implementation manner", "one preferred implementation manner" or "some examples", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0168] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and purposes of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for indoor plan location, characterized in that: The method comprises: Acquire multi-view indoor scene images collected from the target to be located; Extracting monocular depth features of indoor scene images, and analyzing monocular depth features of multi-view indoor scene images to obtain multi-view depth features; The monocular depth feature and the multi-view depth feature are fused to obtain a fused depth feature, and the fused depth feature is processed according to a preset rule to obtain a first posterior probability distribution of the posture of the target to be located, where the posture represents the position and direction of the target to be located in the indoor space; Extracting semantic features of the indoor scene image and processing the semantic features according to preset rules to obtain a second posterior probability distribution of the posture of the target to be located; The first posterior probability distribution and the second posterior probability distribution are weighted according to a preset ratio to obtain a weighted posterior probability distribution, and the weighted posterior probability distribution is input into a histogram filter, the histogram filter outputs a three-dimensional probability body of the posture of the target to be located, and a positioning result of the target to be located is obtained; The method of fusing the monocular depth feature and the multi-view depth feature to obtain a fused depth feature, and processing the fused depth feature according to a preset rule to obtain a first posterior probability distribution of the posture of the target to be located includes: Acquire the posture information of the indoor scene image, where the posture information is used to indicate the position and angle of the camera when taking the indoor scene image; Determine relative poses between image frames of the multi-view indoor scene images according to pose information of the multi-view indoor scene images; Determine the reference weights of multi-view depth features and monocular depth features during feature fusion according to the relative poses between image frames, and fuse the multi-view depth features and monocular depth features according to the reference weights to obtain fused depth features; The Bayesian rule is used to process the fused deep features to obtain the first posterior probability distribution of the posture of the target to be located; The step of extracting semantic features of the indoor scene image and processing the semantic features according to preset rules to obtain a second posterior probability distribution of the posture of the target to be located includes: Perform semantic segmentation on indoor scene images to obtain semantic features of indoor scene images; Determine the semantic recognition label of each pixel of the indoor scene image according to the semantic features; A semantic likelihood field is selected according to the semantic recognition label of the pixel point, and semantic ray projection is performed according to the semantic likelihood field to determine the second posterior probability distribution of the posture of the target to be located.

2. The indoor plan location method according to claim 1, characterized in that: The extracting of the monocular depth features of the indoor scene image and the parsing of the monocular depth features of the multi-view indoor scene image to obtain the multi-view depth features include: Extracting monocular depth features of multi-view indoor scene images respectively; Extracting column features of monocular depth features, and aggregating column features of monocular depth features of multi-view indoor scene images to obtain cross-view features; The depth value of each pixel in the indoor scene plan is calculated according to the feature variance of the cross-view features to obtain the multi-view depth features.

3. The indoor plan location method according to claim 1, characterized in that: The calculation formula of the first posterior probability distribution of the posture of the target to be located is: ; in, Represents the first posterior probability distribution of the posture of the target to be located, Represents the probability distribution of the predicted pixel depth corresponding to the monocular depth feature, Represents the probability distribution of the predicted pixel depth corresponding to the multi-view depth feature, represents the reference weight of the monocular depth feature, represents the upsampling operation, Represents the reference weight of multi-view depth features.

4. An indoor plan location method according to any one of claims 1 to 3, characterized in that: The method further comprises: Obtain an indoor scene plan map, which has a reference semantic label. Calculate the matching value between the semantic recognition label of the pixel and the reference semantic label; Construct a semantic likelihood field based on the matching values.

5. An indoor plan location method according to any one of claims 1 to 3, characterized in that: The indoor scene image is an RGB image. After acquiring the multi-perspective indoor scene image collected by the target to be located and before extracting the monocular depth feature of the indoor scene image, the method further includes: Adjust the orientation of the indoor scene image and convert it into a gravity-aligned image.

6. An indoor plan view positioning system, used to implement an indoor plan view positioning method according to any one of claims 1 to 5, characterized in that: include: An acquisition module is used to acquire multi-view indoor scene images collected by the target to be located; A feature extraction module is used to extract monocular depth features of indoor scene images and to analyze the monocular depth features of multi-view indoor scene images to obtain multi-view depth features; A first processing module is used to fuse the monocular depth feature and the multi-view depth feature to obtain a fused depth feature, and process the fused depth feature according to a preset rule to obtain a first posterior probability distribution of the posture of the target to be located, where the posture represents the position and direction of the target to be located in the indoor space; The second processing module is used to extract the semantic features of the indoor scene image and process the semantic features according to preset rules to obtain a second posterior probability distribution of the posture of the target to be located; The output module is used to weight the first posterior probability distribution and the second posterior probability distribution according to a preset ratio to obtain a weighted posterior probability distribution, and input the weighted posterior probability distribution into a histogram filter. The histogram filter outputs a three-dimensional probability body of the posture of the target to be located, and obtains the positioning result of the target to be located.

7. An electronic device, characterized in that: The electronic device comprises: at least one processor (10); and, a memory (11) communicatively connected to the at least one processor (10); The memory (11) stores a computer program executable by the at least one processor (10), and the computer program is executed by the at least one processor (10) so that the at least one processor (10) can execute the indoor floor plan positioning method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the indoor plan location method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Deep learning-based method for constructing three-dimensional semantic map of indoor environment

    CN110243370A

  • Indoor environment 3D semantic map construction method based on point cloud deep learning

    CN111798475A