Indoor planar graph positioning method, system and device and storage medium
By extracting and fusion of depth features and semantic features of indoor scene images, a three-dimensional probability body is generated, which solves the problems of insufficient robustness of indoor positioning methods and low semantic information utilization in the prior art, and achieves high precision, efficiency and robust indoor positioning.
Patent Information
- Application Number
- CN202510455316.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing indoor positioning method based on floor plan lacks robustness in the fusion of spatiotemporal features of continuous image sequences, resulting in frequent positioning trajectory drifts, insufficient adaptability to non-upright camera postures, failure to effectively mine the semantic topological relationships implicitly in the plan picture, resulting in low utilization of environmental constraint information and difficult to meet the millisecond-level response requirements in large-scale scenarios.
By acquiring indoor scene images from multiple perspectives, monocular depth features and multi-view depth features are extracted, depth features are fused and processed to obtain the first posterior probability distribution of the pose of the target to be positioned, combined with semantic feature processing to obtain the second posterior probability distribution, input the histogram filter through the weighted posterior probability distribution to generate the three-dimensional probability body of the target to be positioned, thereby realizing positioning.
It improves the accuracy, efficiency and robustness of indoor positioning in complex scenarios, reduces positioning blur and error, can better adapt to non-upright camera postures, effectively utilize the semantic topological relationships in floor plans, and meets the high-precision and real-time positioning requirements in large-scale scenarios.
Smart Images

Figure CN119963651A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of indoor positioning technology, and in particular to an indoor plan view positioning method, system, device and storage medium. Background Art
[0002] In recent years, the use of robots with autonomous navigation capabilities has become increasingly common; robot autonomous navigation capability refers to the ability of robots to use sensor data to perceive known or unknown environments and autonomously plan paths based on the perceived data. In indoor environments, accurate positioning is crucial for many applications such as virtual reality (VR) and robot autonomous navigation.
[0003] Traditional technical solutions mainly rely on pre-built 3D environmental models or high-precision sensor databases to achieve positioning. Such methods not only require a large amount of storage resources to maintain environmental data, but also have the problem of model update lag caused by dynamic changes in the scene, which significantly increases the system operation and maintenance costs. In recent years, floor plan-based positioning methods have gradually become alternative solutions. Floor plan-based positioning methods achieve device positioning by extracting geometric and semantic features of building floor plans. However, existing floor plan-based positioning methods still have significant defects: First, there is a lack of robust algorithms in the spatiotemporal feature fusion of continuous image sequences, resulting in frequent positioning trajectory drift; second, the lack of adaptability to non-upright camera postures (such as tilt, pitch, etc.) limits the compatibility with complex observation perspectives; third, the semantic topological relationships implicit in the floor plan (such as porch connectivity, functional area division, etc.) are not effectively mined, resulting in low utilization of environmental constraint information; fourth, the existing solutions have deficiencies in balancing positioning accuracy and real-time performance, and it is difficult to meet the millisecond-level response requirements in large-scale scenarios. Summary of the invention
[0004] In order to improve the accuracy, efficiency and robustness of indoor positioning in complex scenarios and to meet the requirements of different indoor application scenarios for positioning accuracy and real-time performance, the present application provides an indoor floor plan positioning method, system, device and storage medium.
[0005] The present invention provides an indoor plan location method, comprising: Acquire multi-view indoor scene images collected from the target to be located; Extracting monocular depth features of indoor scene images, and analyzing monocular depth features of multi-view indoor scene images to obtain multi-view depth features; The monocular depth feature and the multi-view depth feature are fused to obtain a fused depth feature, and the fused depth feature is processed according to a preset rule to obtain a first posterior probability distribution of the posture of the target to be located. The posture represents the position and direction of the target to be located in the indoor space. Extracting semantic features of the indoor scene image and processing the semantic features according to preset rules to obtain a second posterior probability distribution of the posture of the target to be located; The first posterior probability distribution and the second posterior probability distribution are weighted according to a preset ratio to obtain a weighted posterior probability distribution, and the weighted posterior probability distribution is input into a histogram filter, and the histogram filter outputs a three-dimensional probability body of the posture of the target to be located, thereby obtaining a positioning result of the target to be located.
[0006] Optionally, extracting the monocular depth features of the indoor scene image, and analyzing the monocular depth features of the indoor scene image from multiple perspectives to obtain the multi-view depth features includes: Extracting monocular depth features of multi-view indoor scene images respectively; Extracting column features of monocular depth features, and aggregating column features of monocular depth features of multi-view indoor scene images to obtain cross-view features; The depth value of each pixel in the indoor scene plan is calculated according to the feature variance of the cross-view features to obtain the multi-view depth features.
[0007] Optionally, fusing the monocular depth feature and the multi-view depth feature to obtain a fused depth feature, and processing the fused depth feature according to a preset rule to obtain a first posterior probability distribution of the posture of the target to be located, including: Acquire the posture information of the indoor scene image, where the posture information is used to indicate the position and angle of the camera when taking the indoor scene image; Determine relative poses between image frames of the multi-view indoor scene images according to pose information of the multi-view indoor scene images; Determine the reference weights of multi-view depth features and monocular depth features during feature fusion according to the relative poses between image frames, and fuse the multi-view depth features and monocular depth features according to the reference weights to obtain fused depth features; The Bayesian rule is used to process the fused deep features and obtain the first posterior probability distribution of the posture of the target to be located.
[0008] Optionally, the calculation formula of the first posterior probability distribution of the posture of the target to be located is: , ; in, Represents the first posterior probability distribution of the posture of the target to be located, Represents the probability distribution of the predicted pixel depth corresponding to the monocular depth feature, Represents the probability distribution of the predicted pixel depth corresponding to the multi-view depth feature, represents the reference weight of the monocular depth feature, represents the upsampling operation, Represents the reference weight of multi-view depth features.
[0009] Optionally, extracting semantic features of the indoor scene image and processing the semantic features according to preset rules to obtain a second posterior probability distribution of the posture of the target to be located includes: Perform semantic segmentation on indoor scene images to obtain semantic features of indoor scene images; Determine the semantic recognition label of each pixel of the indoor scene image according to the semantic features; A semantic likelihood field is selected according to the semantic recognition label of the pixel point, and semantic ray projection is performed according to the semantic likelihood field to determine the second posterior probability distribution of the posture of the target to be located.
[0010] Optionally, the method further comprises: Obtain an indoor scene plan map, which has a reference semantic label. Calculate the matching value between the semantic recognition label of the pixel and the reference semantic label; Construct a semantic likelihood field based on the matching values.
[0011] Optionally, the indoor scene image is an RGB image, and after acquiring the multi-view indoor scene image and before extracting the monocular depth feature of the indoor scene image, the method further includes: Adjust the orientation of the indoor scene image and convert it into a gravity-aligned image.
[0012] In order to solve the above problems, the present invention also provides an indoor plan positioning system, the system comprising: An acquisition module is used to acquire multi-view indoor scene images collected by the target to be located; A feature extraction module is used to extract monocular depth features of indoor scene images and to analyze the monocular depth features of multi-view indoor scene images to obtain multi-view depth features; A first processing module is used to fuse the monocular depth feature and the multi-view depth feature to obtain a fused depth feature, and process the fused depth feature according to a preset rule to obtain a first posterior probability distribution of the posture of the target to be located, where the posture represents the position and direction of the target to be located in the indoor space; The second processing module is used to extract the semantic features of the indoor scene image and process the semantic features according to preset rules to obtain a second posterior probability distribution of the posture of the target to be located; The output module is used to weight the first posterior probability distribution and the second posterior probability distribution according to a preset ratio to obtain a weighted posterior probability distribution, and input the weighted posterior probability distribution into a histogram filter. The histogram filter outputs a three-dimensional probability body of the posture of the target to be located, and obtains the positioning result of the target to be located.
[0013] In order to solve the above problem, the present invention further provides an electronic device, the electronic device comprising: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the indoor floor plan positioning method described above.
[0014] In order to solve the above problem, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is executed by a processor in an electronic device to implement the indoor floor plan positioning method described above.
[0015] In summary, this application includes the following beneficial technical effects: When processing depth images of indoor scenes from multiple perspectives, monocular depth features of the indoor scene images from multiple perspectives are extracted respectively, and multi-view depth features are extracted. The first a posteriori probability distribution of the posture of the target to be located is calculated based on the monocular depth features and the multi-view depth features, and the positioning result of the target to be located is calculated using the first a posteriori probability distribution to reduce positioning ambiguity and error. An efficient histogram filter is used to fuse the positioning information calculated by the depth features of the indoor scene image and the positioning information calculated by the semantic features to obtain the final positioning information of the target to be located. When calculating the positioning result of the target to be located, the semantic information of each pixel point in the indoor scene image is referred to to improve the accuracy, efficiency and robustness of indoor positioning in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A schematic diagram of a flow chart of an indoor plan location method provided by an embodiment of the present invention; Figure 2 A system flow chart of an indoor plan location method provided by an embodiment of the present invention; Figure 3 A flowchart of the steps of extracting monocular depth features of indoor scene images and parsing monocular depth features of multi-view indoor scene images to obtain multi-view depth features provided by an embodiment of the present invention; Figure 4 A schematic diagram of the structure of an electronic device for implementing the indoor plan view positioning method provided by an embodiment of the present invention.
[0017] Reference numerals: 10, processor; 11, memory; 12, communication bus; 13, communication interface.
[0018] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0019] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0020] In the description of the present invention, it is necessary to understand that the terms "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0021] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal connection between two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.
[0022] Reference Figure 1 and Figure 2 , is a flow chart of an indoor plan view positioning method provided by an embodiment of the present invention. In this embodiment, the indoor plan view positioning method includes: S1. Acquire multi-view indoor scene images of the target to be located.
[0023] The indoor scene image is an RGB image. In this embodiment, a wide-angle camera is used to obtain a sequence of indoor scene images. The horizontal field of view of the wide-angle camera is 108° to ensure wide coverage of the surrounding environment. In addition, the target to be located can also use its own sensor system (wheel speed sensor) to obtain motion data to simulate the odometer function, which assists in the subsequent indoor floor plan positioning.
[0024] In a preferred implementation of this embodiment, after acquiring multi-view indoor scene images and before extracting monocular depth features of the indoor scene images, the following operations are performed: adjusting the direction of the indoor scene images and converting the indoor scene images into gravity-aligned images.
[0025] When converting an indoor scene image to a gravity-aligned image, we first use the roll angle and pitch angle To calculate the rotation matrix , so as to subsequently rotate the indoor scene image from the current pitch posture to the horizontal posture.
[0026] After that, the pixel coordinates in the original image are converted to the coordinates in the gravity-aligned image using the homography matrix; the homography matrix is a 3×3 matrix used in computer vision to describe the projection transformation relationship between two planes, representing the mapping relationship of the same plane under different viewing angles. The expression of the homography matrix is ,in, is the camera intrinsic parameter matrix, is the inverse matrix of the camera intrinsic parameter matrix, are the homogeneous image coordinates of the original pixel, are the corresponding pixel coordinates in the gravity-aligned image.
[0027] Then, the pixels that cannot be seen at the current pitch angle are filtered out, and the invisible pixels are masked out, finally obtaining a gravity-aligned image. It should be explained that after gravity alignment, the image is rotated to a perspective based on the horizontal ground, and some pixels originally at the top of the image (such as the ceiling or high content) may be moved out of the image frame and become invisible. The pixels that are rotated out are marked as invisible pixels (masked out), avoiding the network from trying to make meaningless depth estimates for these pixels.
[0028] After converting the indoor scene image into a gravity-aligned image, the buildings in the gravity-aligned image appear vertical (equivalent to a photo taken by the camera in a horizontal posture). The state of the tilted buildings in the image is corrected to ensure the accuracy and robustness of feature matching and semantic segmentation in the subsequent depth estimation process.
[0029] S2. Extracting monocular depth features of indoor scene images, and parsing monocular depth features of multi-view indoor scene images to obtain multi-view depth features.
[0030] Reference Figure 3 , extracting monocular depth features of indoor scene images, and parsing monocular depth features of multi-view indoor scene images to obtain multi-view depth features, including: S21. Extract monocular depth features of multi-view indoor scene images respectively.
[0031] Specifically, ResNet (Residual Neural Network) and Attention mechanism networks can be used to extract features of single-frame indoor scene images. The Attention mechanism network can mask invisible pixels and output a depth hypothesis probability distribution. The expectation of the depth hypothesis probability distribution is used as the floor plan depth prediction value, and the floor plan depth prediction value is used to construct equiangular ray scanning positioning.
[0032] S22. Extract column features of monocular depth features, and aggregate column features of monocular depth features of multi-view indoor scene images to obtain cross-view features.
[0033] Cross-view features refer to column features extracted from images of different perspectives. These features can correspond to each other and are used for subsequent depth estimation. Specifically, a single-frame indoor scene image will get a predicted depth for each pixel in this frame. Multi-view depth estimation can get the predicted depth of each pixel in the reference frame indoor scene image (the reference frame and the single-frame RGB image are the same frame).
[0034] Inspired by multi-view stereo vision, the MVS network is used to estimate the depth of a plane image from multiple frames. The full name of the MVS network is Multi-View Stereo network, which means multi-view stereo network in Chinese. Multi-view stereo network is a computer vision technology that uses images from multiple perspectives to restore the three-dimensional structure of a scene. This technology is usually used to estimate depth information from multiple two-dimensional images and then reconstruct a three-dimensional model. The MVS network estimates the depth of each pixel by analyzing image features from different perspectives and calculating pixel or feature point matches between views, ultimately achieving reconstruction of the three-dimensional scene.
[0035] In step S22, the column features of a single frame of indoor scene images are first extracted one by one, and the column features of indoor scene images of different views are gathered through plane scanning. The 2D cost distribution is constructed according to the cross-view feature variance, and the final depth is calculated through soft-argmin to supplement the multi-view geometric clues for positioning. The soft-argmin algorithm is a differentiable optimal value estimation method that realizes continuous space regression by probabilistically discrete candidate values.
[0036] Specifically, after obtaining the feature variance under different depth assumptions, the feature variance directly corresponds to the 2D cost distribution. For each pair of matched feature points, a cost value is calculated by the feature variance (using weighted summation or mean calculation). The cost value is usually based on the difference between the feature descriptions. The smaller the difference, the lower the cost, indicating that the match is more reliable and the probability that the depth assumption is correct is higher.
[0037] S23. Calculate the depth value of each pixel point in the indoor scene plan view according to the feature variance of the cross-view feature to obtain a multi-view depth feature.
[0038] S3. Fusing the monocular depth feature and the multi-view depth feature to obtain a fused depth feature, and processing the fused depth feature according to a preset rule to obtain a first posterior probability distribution of the posture of the target to be located.
[0039] The posture of the target to be located indicates the position and direction of the target in the indoor space. The target to be located can be an autonomous navigation robot, an autonomous navigation drone, or a VR device. Specifically, the direction of the target to be located in the indoor space refers to the front direction or moving direction of the target to be located. Fusion of monocular depth features and multi-view depth features for depth estimation can make full use of multimodal clues and reduce positioning ambiguity and errors.
[0040] Reference Figure 2 , the specific steps of S3 include: S31. Acquire posture information of the indoor scene image, where the posture information is used to indicate the position of the camera when shooting the indoor scene image and the angle between the camera lens and the horizontal ground when shooting the indoor scene image.
[0041] S32. Determine the relative pose between image frames of the multi-view indoor scene image according to the pose information of the multi-view indoor scene image.
[0042] S33. Determine the reference weights of the multi-view depth features and the monocular depth features during feature fusion according to the relative postures between the image frames, and fuse the multi-view depth features and the monocular depth features according to the reference weights to obtain a fused depth feature.
[0043] S34. Using the Bayesian rule to process the fused deep features, a first posterior probability distribution of the posture of the target to be located is obtained.
[0044] The calculation formula of the first posterior probability distribution of the posture of the target to be located is: , ; in, Represents the first posterior probability distribution of the posture of the target to be located, Represents the probability distribution of the predicted pixel depth corresponding to the monocular depth feature, Represents the probability distribution of the predicted pixel depth corresponding to the multi-view depth feature, represents the reference weight of the monocular depth feature, represents the upsampling operation, which ensures the validity of the addition; Represents the reference weight of multi-view depth features; the first posterior probability distribution of the posture of the target to be located The expected value of is used to provide the final depth prediction.
[0045] Assume that the current position of the camera in the positioning model corresponding to the target to be positioned is , Represents the camera on a two-dimensional plane Coordinate location, Represents the camera on a two-dimensional plane Coordinate location, Represents the camera relative to the reference direction (the global coordinate system The rotation angle of the camera axis is determined by fusing multimodal features based on the collected data and the semantic plane occupancy map.
[0046] Monocular depth estimation is independent of camera motion, but is prone to scale ambiguity; multi-view stereo methods can provide the correct scale, but rely on sufficient baseline and camera overlap. Based on these observations, we use an MLP network to soft-select from two predictions (single-frame indoor scene image and multi-frame indoor scene image). MLP stands for Multi-Layer Perceptron, which is a multi-layer perceptron. MLP networks can perform nonlinear feature transformations by stacking fully connected layers, and are widely used in classification, regression, and deep learning feature encoding.
[0047] The MLP network adaptively weights and fuses the probability distribution of monocular depth estimation and multi-view depth estimation based on the relative pose between image frames, the average depth prediction value of monocular depth features and multi-view depth features, thereby improving the accuracy and reliability of depth estimation. Specifically, the predicted plane image depth is used as observation, and the following observation model is adopted: in It's a gesture The plane ray at is the ray interpolated from the depth prediction of the floor plan, is a constant factor; express Observation of the camera at all times; is the time index, The value range is .
[0048] It refers to the depth prediction of the camera pose at time t, obtained through camera observation based on the complementary fusion of monocular depth and multi-view depth introduced earlier. The ray interpolated from the depth prediction can be compared with the plane map ray of the possible camera poses in the entire plane map to obtain a probability distribution of each pose.
[0049] In this embodiment, the probability distribution is calculated as follows: Use the relative pose between frames, i.e. self-motion, as a transfer model ; in, and Indicates time The self-motion Indicates time The transfer noise, The operator represents the application of self-motion to a state. Indicates time Self-motion Coordinate position translation, Indicates time Self-motion Coordinate position translation, Indicates time The self-motion is relative to the reference direction (the The change in the rotation angle of the axis; Indicates time The amount of translation of motion noise in the x-coordinate position, Indicates time The translation amount and Indicates time The rotation angle change of the motion noise relative to the reference direction (the x-axis of the global coordinate system).
[0050] Further assume that the transfer noise Obeying Gaussian distribution, the transition probability is expressed as , represents transpose; is the inverse matrix of the covariance matrix, which represents the uncertainty of the state transfer error. The covariance matrix Σ is a symmetric positive definite matrix, and its inverse matrix Used to measure the size of the error vector; state transition probability From the current state and self-motion Derivation of the next state The probability distribution of Used to evaluate the uncertainty in modeling state transitions, it is usually modeled as a Gaussian distribution. The mean of the predicted state , covariance Used to describe uncertainty.
[0051] we will Modeled as the covariance of a Gaussian distribution, express The variance in the coordinate direction, express The variance in the coordinate direction, Represents the direction relative to the reference direction (the global coordinate system The variance corresponding to the rotation angle of the axis); after applying the Bayesian rule, we get ,in is the normalization factor, It is the posture space; is from The observation state and self-motion of the camera at each moment Derivation of the next state The probability distribution of .
[0052] S4. Extracting semantic features of the indoor scene image, and processing the semantic features according to preset rules to obtain a second posterior probability distribution of the posture of the target to be located.
[0053] Extracting semantic features of the indoor scene image and processing the semantic features according to preset rules to obtain a second posterior probability distribution of the posture of the target to be located, including: S41, performing semantic segmentation on the indoor scene image to obtain semantic features of the indoor scene image; The convolutional neural network (CNN) encoder-decoder architecture is used for semantic segmentation to identify semantic objects such as walls, doors, and windows in the input RGB image (indoor scene image). This semantic information is integrated with geometric features to enhance the positioning ability to understand and distinguish scenes.
[0054] S42, determining a semantic recognition label for each pixel of the indoor scene image according to the semantic feature; S43, selecting a semantic likelihood field according to the semantic recognition label of the pixel point, and performing semantic ray projection according to the semantic likelihood field to determine a second posterior probability distribution of the posture of the target to be located.
[0055] In a preferred implementation of this embodiment, the indoor plan location method further includes: S401, obtaining an indoor scene plan view, the indoor scene plan view having a reference semantic label; S402, calculating the matching value between the semantic recognition label of the pixel point and the reference semantic label; S403: Construct a semantic likelihood field according to the matching value.
[0056] The reference semantic label is the known semantic information (such as walls, doors, windows, etc.) extracted from the indoor floor plan, while the semantic recognition label is the semantic information recognized from the indoor scene image collected in real time through the semantic segmentation network. The matching value can be calculated by comparing the consistency of the semantic recognition label of each pixel with the reference semantic label, using the semantic similarity metric. The higher the matching value, the more consistent the semantic information of the pixel is with the reference information in the floor plan, so that when constructing the semantic likelihood field, the pixel contributes more to the positioning.
[0057] The semantic likelihood field is a probabilistic model that combines semantic information of the environment to enhance the positioning and navigation capabilities of robots or systems. It provides richer constraints for positioning by calculating the probability of observing specific semantic labels (such as walls, doors, windows, etc.) at a given position and orientation. Compared with traditional positioning methods based on geometric information, the semantic likelihood field can more accurately identify and utilize key landmarks in the environment, thereby improving positioning accuracy and robustness, especially in complex environments or when maps are inaccurate.
[0058] Specifically, reference semantic labels are extracted from an existing indoor scene plan. In this embodiment, the reference semantic labels include walls, doors, and windows, because walls, doors, and windows are easy to automatically extract from the plan and are crucial for human positioning, thereby improving the positioning effect.
[0059] In order to make the indoor scene floor plan with reference semantic labels readable by the robot, the indoor scene floor plan is converted into an occupancy grid map. The occupancy grid is a two-dimensional representation of the world. Each cell in the occupancy grid has an occupancy probability, which is determined by its normalized grayscale value.
[0060] Combining semantic likelihood with geometric feature likelihood, and combining Bayesian rule to update the camera pose posterior probability, enhances the robustness of positioning in complex environments and inaccurate maps. A joint method is used here, which can use the likelihood field to incorporate semantic information in the presence of semantic labels. More importantly, it can also use ray casting within the likelihood field to operate without distance measurement.
[0061] The MCL motion model is usually distributed by It means that MCL is the full name of Monte Carlo Localization in English, and its Chinese meaning is Monte Carlo Positioning; Indicates that given the current state and odometer measurements Under the condition of Transfer to state probability. In time One of the possible states is is the index of possible states, representing different state hypotheses; In time Odometer measurements, In time The current status of is the index of the current state.
[0062] The previous set of particles Using odometer measurements Propagate to the current set of particles .
[0063] in, is a normalization factor, is a collection containing every cell in the map, In time This allows the two likelihoods to be considered independent, and the motion The definition is the same as that of RMCL. The full name of RMCL is Rao-Blackwellized Monte Carlo Localization, which is a positioning algorithm based on the Monte Carlo method.
[0064] The RMCL algorithm combines the random sampling characteristics of the Monte Carlo method and the optimization effect of the Rao-Blackwell theorem to improve the accuracy and efficiency of positioning. In RMCL, the algorithm maintains a set of particles representing the possible positions of the robot, updates the weights of these particles based on sensor data and control inputs, and then generates a new set of particles through a resampling step to reflect a more accurate localization probability distribution. The prior is a set of particles that contain The occupancy probability of the cell is , represents the occupancy probability corresponding to the plane diagram at the observation point at time t; is a random variable, indicating that Under this condition, the probability of observing semantic information o; specifically, represents the state at time t, usually including position and direction (for example, the position and orientation of the robot in the environment); o represents the observed semantic information, such as the reference semantic labels identified from the image (such as walls, doors, windows, etc.); v is a semantic feature, which represents the state at state Under this condition, the matching degree between the observed semantic information o and a specific semantic label v.
[0065] The likelihood field model computes a distance map. For each cell , calculate the distance to the nearest occupied cell And store, for each cell in the map The set of distances to the nearest occupied cells is the matching value between the semantic recognition label of the pixel and the reference semantic label.
[0066] In the formula, Indicates that in the set of landmarks that meet the conditions, Nearest landmarks distance; Indicates that among all possible landmarks Find the minimum value in Nearest landmarks It should be noted that the landmark That is, the semantic identification label of the pixel point, the landmark These are reference identification tags (such as doors, windows, and walls).
[0067] Indicates landmarks and landmarks The distance between them is usually measured using Euclidean distance or other distance metrics; Indicates landmarks Occupancy rate; Represents a threshold used to filter landmarks Only when Time, Landmark will be taken into consideration.
[0068] When receiving the observation When , the endpoints are estimated and used as indices in the distance map. Assuming a Gaussian error distribution, each particle The weight of can be estimated as Indicates at time No. observations; Indicates at time One of the possible states of is the index of the state; A collection containing every cell in the map; Indicates observation and calculate the distance to the nearest occupied cell; Represents the standard deviation of the distance error, which defines the width of the Gaussian distribution.
[0069] is part of the probability density function of the Gaussian distribution and is used to calculate the difference in a given distance and standard deviation The probability of time.
[0070] For each reference semantic label (door, wall, window) present in the floor plan, we can compute a distance map that stores the shortest distance to cells with the same label. Formally, for each map cell , we can estimate the distance to the nearest cell for each label as , calculate the shortest distance among cells with the same semantic label.
[0071] Indicates landmarks Occupancy rate; Represents a threshold used to filter landmarks Only when Time, Landmark will be taken into consideration.
[0072] They are the distances to the nearest wall, door, and window, respectively, thereby constructing three semantic likelihood fields.
[0073] When we receive observations We use the tag To decide which semantic likelihood field to use for semantic ray projection.
[0074] In this embodiment, there are three semantic likelihood fields, namely door, wall and window; then when my observation image has a semantic label (Refer to the semantic label) For the door part, the semantic likelihood field of the door is used for semantic matching to obtain a matching probability; the matching process is to use semantic ray projection, where the door pixel is 1 and the others are 0, and perform equiangular ray semantic matching of the sliding window. Similarly, the other two semantic likelihood fields (wall / window) are matched in the same way.
[0075] After obtaining the observation probability, the weight of each particle is calculated , the calculation formula is ,in is the weight at the previous moment. After normalization, the new weight is used for state estimation, that is , and the state is accurately estimated during positioning. The two work together to improve the performance of the positioning system, enabling the robot to accurately locate and navigate in complex environments.
[0076] In time At that time, The weight of each particle; In time At that time, The weight of each particle; In a given state and a collection of each cell in the map Under the condition of probability; For all The observation probability of each particle is summed up to normalize the weight; is the particle index, The total number of particles; Given odometer input , current observation , Previous State and a collection of each cell in the map Under the condition of In state probability; By transforming the state of each particle Multiply by its corresponding weight And sum to estimate the system in time status.
[0077] Since lower priors are more discriminative, With each label's prior is relevant not only because it is one less parameter to tune, but also because it implicitly makes observing rare landmarks more beneficial than observing common ones.
[0078] Semantic information fusion enhances tolerance to environmental changes and map errors. When dealing with complex scenarios such as inaccurate maps, repeated scene structures, dynamic obstacles, and lighting changes, it can provide stable and accurate positioning, maintain system performance, and reduce the risk of positioning failure and drift.
[0079] S5. Weight the first posterior probability distribution and the second posterior probability distribution according to a preset ratio to obtain a weighted posterior probability distribution, and input the weighted posterior probability distribution into a histogram filter. The histogram filter outputs a three-dimensional probability body of the posture of the target to be located, and obtains a positioning result of the target to be located.
[0080] A three-dimensional probability volume refers to a digital expression of a probability density field in three-dimensional space, which quantifies the geometric / semantic existence probability of each voxel. Through the three-dimensional probability volume, we can know the position of the target to be located in the indoor scene, so as to facilitate the positioning of the target when the indoor scene is relatively complex.
[0081] The specific steps of semantic ray projection are as follows: The weight of the Bayesian likelihood function calculation and the weight of the semantic ray projection calculation are summed in a certain proportion. The weight of the Bayesian likelihood function is , the semantic ray projection weight is .
[0082] The expression of the weighted probability distribution is: represents the first posterior probability distribution, represents the second posterior probability distribution.
[0083] An efficient histogram filter is used to fuse the positioning information calculated by the Bayesian rule and the positioning information calculated by semantic features to improve positioning stability and accuracy. The filter represents the pose posterior as a three-dimensional probability body, decomposes the translation and rotation according to the relative pose between frames, and implements efficient transfer updates for group convolution in different directions, quickly converging the pose estimation probability distribution.
[0084] Based on the same inventive concept, an embodiment of the present invention provides an indoor plan positioning system.
[0085] The indoor plan view positioning system of the present invention can be loaded in an electronic device. According to the functions to be implemented, the indoor plan view positioning system includes an acquisition module, a feature extraction module, a first processing module, a second processing module and an output module. The acquisition module can acquire multi-view indoor scene images collected by the target to be located; the feature extraction module can extract the monocular depth features of the indoor scene images, and analyze the monocular depth features of the multi-view indoor scene images to obtain multi-view depth features; The first processing module can fuse the monocular depth feature and the multi-view depth feature to obtain the fused depth feature, and process the fused depth feature according to the preset rules to obtain the first posterior probability distribution of the posture of the target to be located, and the posture represents the position and direction of the target to be located in the indoor space; the second processing module can extract the semantic features of the indoor scene image, and process the semantic features according to the preset rules to obtain the second posterior probability distribution of the posture of the target to be located; the output module can weight the first posterior probability distribution and the second posterior probability distribution according to the preset ratio to obtain the weighted posterior probability distribution, and input the weighted posterior probability distribution into the histogram filter, and the histogram filter outputs the three-dimensional probability body of the posture of the target to be located to obtain the positioning result of the target to be located.
[0086] The module described in the present invention may also be referred to as a unit, which refers to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and is stored in a memory of the electronic device.
[0087] The various variations and specific examples of the indoor floor plan positioning method provided in the above embodiment are also applicable to the indoor floor plan positioning system of the present embodiment. Through the above detailed description of the indoor floor plan positioning method, those skilled in the art can clearly know the implementation method of the indoor floor plan positioning system in the present embodiment. For the sake of brevity of the specification, it will not be described in detail here.
[0088] The present application also discloses an electronic device, such as Figure 4 , which is a schematic diagram of the structure of an electronic device of an indoor floor plan positioning method provided by an embodiment of the present invention. The electronic device may include at least one processor 10, a memory 11 connected to the at least one processor for communication, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as an indoor floor plan positioning method program.
[0089] The processor 10 may be composed of an integrated circuit in some embodiments, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor 10 is the control core (Control Unit) of the electronic device, and uses various interfaces and lines to connect various components of the entire electronic device, and executes various functions of the electronic device and processes data by running or executing programs or modules stored in the memory 11 (for example, executing indoor floor plan positioning methods, etc.), and calling data stored in the memory 11.
[0090] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of an electronic device, such as a mobile hard disk of the electronic device. In other embodiments, the memory 11 may also be an external storage device of an electronic device, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit of the electronic device and an external storage device. The memory 11 can not only be used to store application software and various types of data installed in the electronic device, such as the code of the indoor floor plan positioning method program, but also can be used to temporarily store data that has been output or is to be output.
[0091] The communication bus 12 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize connection and communication between the memory 11 and at least one processor 10, etc.
[0092] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device and other electronic devices. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)), and optionally, the user interface may also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode, organic light-emitting diode) touch device, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device and to display a visual user interface.
[0093] Figure 4 Only an electronic device with components is shown, and those skilled in the art will understand that Figure 4The structure shown does not constitute a limitation on the electronic device, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0094] For example, although not shown, the electronic device may also include a power source (such as a battery) for supplying power to various components. Preferably, the power source may be logically connected to at least one processor 10 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include any components such as one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, and power status indicators. The electronic device may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.
[0095] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited by this structure.
[0096] Furthermore, if the module / unit integrated in the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile.
[0097] The embodiment of the present application provides a computer-readable storage medium, for example, including: any entity or device, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM) that can carry the computer program code. The computer-readable storage medium stores a computer program that can be loaded by a processor and execute the indoor plan location method of the above embodiment.
[0098] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", "an implementation", "a preferred implementation" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0099] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for indoor plan location, characterized in that: The method comprises: Acquire multi-view indoor scene images collected from the target to be located; Extracting monocular depth features of indoor scene images, and analyzing monocular depth features of multi-view indoor scene images to obtain multi-view depth features; The monocular depth feature and the multi-view depth feature are fused to obtain a fused depth feature, and the fused depth feature is processed according to a preset rule to obtain a first posterior probability distribution of the posture of the target to be located, where the posture represents the position and direction of the target to be located in the indoor space; Extracting semantic features of the indoor scene image and processing the semantic features according to preset rules to obtain a second posterior probability distribution of the posture of the target to be located; The first posterior probability distribution and the second posterior probability distribution are weighted according to a preset ratio to obtain a weighted posterior probability distribution, and the weighted posterior probability distribution is input into a histogram filter, and the histogram filter outputs a three-dimensional probability body of the posture of the target to be located, thereby obtaining a positioning result of the target to be located.
2. The indoor plan location method according to claim 1, characterized in that: The extracting of the monocular depth features of the indoor scene image and the parsing of the monocular depth features of the multi-view indoor scene image to obtain the multi-view depth features include: Extracting monocular depth features of multi-view indoor scene images respectively; Extracting column features of monocular depth features, and aggregating column features of monocular depth features of multi-view indoor scene images to obtain cross-view features; The depth value of each pixel in the indoor scene plan is calculated according to the feature variance of the cross-view features to obtain the multi-view depth features.
3. The indoor plan location method according to claim 1, characterized in that: The method of fusing the monocular depth feature and the multi-view depth feature to obtain a fused depth feature, and processing the fused depth feature according to a preset rule to obtain a first posterior probability distribution of the posture of the target to be located includes: Acquire the posture information of the indoor scene image, where the posture information is used to indicate the position and angle of the camera when taking the indoor scene image; Determine relative poses between image frames of the multi-view indoor scene images according to pose information of the multi-view indoor scene images; Determine the reference weights of multi-view depth features and monocular depth features during feature fusion according to the relative poses between image frames, and fuse the multi-view depth features and monocular depth features according to the reference weights to obtain fused depth features; The Bayesian rule is used to process the fused deep features and obtain the first posterior probability distribution of the posture of the target to be located.
4. The indoor plan location method according to claim 3, characterized in that: The calculation formula of the first posterior probability distribution of the posture of the target to be located is: , ; in, Represents the first posterior probability distribution of the posture of the target to be located, Represents the probability distribution of the predicted pixel depth corresponding to the monocular depth feature, Represents the probability distribution of the predicted pixel depth corresponding to the multi-view depth feature, represents the reference weight of the monocular depth feature, represents the upsampling operation, Represents the reference weight of multi-view depth features.
5. An indoor plan location method according to any one of claims 1 to 4, characterized in that: The step of extracting semantic features of the indoor scene image and processing the semantic features according to preset rules to obtain a second posterior probability distribution of the posture of the target to be located includes: Perform semantic segmentation on indoor scene images to obtain semantic features of indoor scene images; Determine the semantic recognition label of each pixel of the indoor scene image according to the semantic features; A semantic likelihood field is selected according to the semantic recognition label of the pixel point, and semantic ray projection is performed according to the semantic likelihood field to determine the second posterior probability distribution of the posture of the target to be located.
6. The indoor plan location method according to claim 5, characterized in that: The method further comprises: Obtain an indoor scene plan map, which has a reference semantic label. Calculate the matching value between the semantic recognition label of the pixel and the reference semantic label; Construct a semantic likelihood field based on the matching values.
7. An indoor plan location method according to any one of claims 1 to 4, characterized in that: The indoor scene image is an RGB image. After acquiring the multi-view indoor scene image and before extracting the monocular depth feature of the indoor scene image, the method further includes: Adjust the orientation of the indoor scene image and convert it into a gravity-aligned image.
8. An indoor plan view positioning system, used to implement an indoor plan view positioning method according to any one of claims 1 to 7, characterized in that: include: An acquisition module is used to acquire multi-view indoor scene images collected by the target to be located; A feature extraction module is used to extract monocular depth features of indoor scene images and to analyze the monocular depth features of multi-view indoor scene images to obtain multi-view depth features; A first processing module is used to fuse the monocular depth feature and the multi-view depth feature to obtain a fused depth feature, and process the fused depth feature according to a preset rule to obtain a first posterior probability distribution of the posture of the target to be located, where the posture represents the position and direction of the target to be located in the indoor space; The second processing module is used to extract the semantic features of the indoor scene image and process the semantic features according to preset rules to obtain a second posterior probability distribution of the posture of the target to be located; The output module is used to weight the first posterior probability distribution and the second posterior probability distribution according to a preset ratio to obtain a weighted posterior probability distribution, and input the weighted posterior probability distribution into a histogram filter. The histogram filter outputs a three-dimensional probability body of the posture of the target to be located, and obtains the positioning result of the target to be located.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor (10); and, a memory (11) communicatively connected to the at least one processor (10); The memory (11) stores a computer program executable by the at least one processor (10), and the computer program is executed by the at least one processor (10) so that the at least one processor (10) can execute the indoor floor plan positioning method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the indoor plan view positioning method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Deep learning-based method for constructing three-dimensional semantic map of indoor environment
CN110243370A
Indoor environment 3D semantic map construction method based on point cloud deep learning
CN111798475A
Real scene three-dimensional semantic reconstruction method and device based on deep learning and storage medium
CN113673400A
Self-adaptive target navigation method and system for service robot
CN114460943A
Object perception SLAM algorithm based on quadric surface initialization and joint data association
CN115239809A