Indoor positioning method and device based on lightweight map
Through lightweight map technology, the two-dimensional image is processed using preset indoor positioning models, and bird's eye view features are generated and matched, solving the positioning problem of complex indoor environments and realizing high-precision indoor positioning that does not rely on the network in a dynamic environment.
Patent Information
- Application Number
- CN202510336897.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-13
AI Technical Summary
The existing indoor positioning technology has low accuracy in dynamic environments and relies on network signals, which cannot effectively solve the positioning problem of complex indoor environments.
The lightweight map method is used to collect two-dimensional images through mobile devices, and semantic clue extraction and inverse perspective mapping are used to generate BEV features, and feature correction and matching are performed to obtain indoor positioning.
It realizes the ability to accurately locate indoors in a dynamic environment without relying on the network, and is suitable for a variety of indoor scenarios, including AR navigation, emergency response and smart home control.
Smart Images

Figure CN120141490A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and particularly to an indoor positioning method and apparatus based on a lightweight map. Background Art
[0002] In the face of unfamiliar and complex indoor environments, such as large apartments, office spaces, etc., people are prone to getting lost. Therefore, how to locate users in such indoor environments is particularly important.
[0003] In related positioning technologies, such as the positioning technology of GPS, in cities, especially in indoor environments, satellite signals are severely attenuated after being blocked and interfered layer by layer, so the effect is not good; solutions based on technologies such as WIFI, Bluetooth, infrared, and ultrasonic have the disadvantage of being affected by environmental factors (such as signal strength, transmission distance, device hardware support, etc.). For SLAM positioning technology, vision-based SLAM is affected by cumulative errors and is not suitable for long-term positioning; laser-based SLAM is limited by cost, and most mobile devices do not support it, and the results are greatly affected by the dynamic indoor environment. Existing methods based on image retrieval, 2D image matching with 3D scenes (feature descriptors, point clouds, 3D reconstruction, etc.) require the server side to store large-scale maps, cannot support device-side computing, and rely on network signals. In addition, existing vision-based positioning methods cannot well solve the problem of dynamic changes in indoor environments.
[0004] Therefore, there is an urgent need for an indoor positioning method that is not affected by dynamic environments and does not rely on network restrictions. Summary of the Invention
[0005] In view of this, the present application provides an indoor positioning method and apparatus based on a lightweight map, which can achieve indoor positioning through a lightweight map without being affected by dynamic environments and without relying on networks.
[0006] To solve the above technical problems, the technical solution of the present application is implemented as follows:
[0007] In one embodiment, an indoor positioning method based on a lightweight map is provided, which is applied to a mobile device. The method includes:
[0008] Collecting a two-dimensional image centered on the user; the user is the user holding the mobile device;
[0009] Inputting the two-dimensional image into a preset positioning model, and obtaining the indoor pose corresponding to the user based on the preset indoor positioning model and a preset lightweight map; the indoor pose includes two-dimensional coordinates and an orientation angle; the preset lightweight map is a map that discards visual appearance information and retains the spatial layout and relationships of indoor elements;
[0010] Output the indoor pose in the preset lightweight map;
[0011] Among them, the preset indoor positioning model first extracts semantic clues from the two-dimensional image to obtain semantic features, then performs inverse perspective mapping on the semantic features to generate BEV features, then performs feature correction based on the BEV features to obtain complete BEV features, and finally matches the complete BEV features with the preset lightweight map to obtain the indoor pose.
[0012] Among them, the preset indoor positioning model includes: a preset feature extraction sub-model, a preset feature conversion sub-model, a preset feature repair sub-model, and a preset feature matching sub-model; the preset feature extraction sub-model is used to extract semantic clues from the two-dimensional image to obtain semantic features; the preset feature conversion sub-model is used to perform inverse perspective mapping on the semantic features to generate BEV features; the preset feature repair sub-model performs feature correction based on the BEV features to obtain complete BEV features; the preset feature matching sub-model is used to match the complete BEV features with the preset lightweight map to obtain the indoor pose.
[0013] Among them, the semantic features include: layout elements and semantic elements; the layout elements are elements that build the room structure; the semantic elements are devices placed in the room.
[0014] The preset feature extraction sub-model is used to extract semantic clues from the two-dimensional image to obtain semantic features, including:
[0015] Obtain the layout elements through room layout estimation, and filter out the elements that affect the user's viewpoint and angle;
[0016] Obtain semantic elements through instance segmentation.
[0017] Among them, the preset feature conversion sub-model is used to perform inverse perspective mapping on the semantic features to generate BEV features, including:
[0018] The preset feature conversion sub-model uses inverse perspective mapping to lift the semantic features from the spatial two-dimensional perspective space to the three-dimensional world space, and then projects them into the BEV space to obtain BEV features.
[0019] Among them, when projecting into the BEV space, the method further includes:
[0020] If different types of BEVs are projected onto the same two-dimensional space position, the priority of semantic elements is greater than that of layout elements, and the priority of layout elements is greater than that of elements other than semantic elements and layout elements;
[0021] If BEVs of the same type are projected onto the same two-dimensional spatial position, the average feature in the corresponding grid cell is used.
[0022] Among them, the training of the preset feature repair sub-model includes:
[0023] Obtain a first training sample set; among them, the first training sample set is a set of multiple two-dimensional images and reference BEV views; the reference BEV view is determined according to the pose information of radar positioning and the preset lightweight map;
[0024] Establish an initial feature repair sub-model;
[0025] Train the initial feature repair sub-model based on the first training sample set until the similarity between the BEV feature output by the initial feature repair sub-model in the encoded and decoded BEV view and the reference BEV view meets the first preset condition, and obtain the preset feature repair sub-model.
[0026] Among them, determining the reference BEV view according to the pose information of radar positioning and the preset lightweight map includes:
[0027] Obtain pose information; the pose information is synchronously obtained by radar positioning when the camera captures images;
[0028] Adjust the preset lightweight map based on the orientation angle in the pose information; and crop a map of a preset size in the adjusted preset lightweight map based on the two-dimensional coordinates in the pose information;
[0029] Project rays from the pose information, hit the part within the visual range, and filter out invisible elements to obtain the reference BEV view.
[0030] Among them, the feature matching sub-model is used to match the complete BEV feature with the preset lightweight map to obtain the indoor pose, including:
[0031] The feature matching sub-model predicts the indoor pose based on Transformer;
[0032] Use the complete BEV feature as the input of the encoder of the Transformer;
[0033] Use the encoded preset lightweight map as the input of the decoder of the Transformer;
[0034] Output the lightweight map matching result through the Transformer; and decode the lightweight map matching result into the indoor pose.
[0035] Among them, the training of the preset feature matching sub-model includes:
[0036] Obtain a second training sample set; wherein, the second training sample set is a set of multiple two-dimensional images and reference indoor poses; the reference indoor poses are obtained by radar positioning;
[0037] Establish an initial feature matching sub-model;
[0038] Train the initial feature matching sub-model based on the second training sample set until the similarity between the indoor pose output by the initial feature matching sub-model and the corresponding reference indoor pose meets the second preset condition, and obtain the preset feature matching sub-model.
[0039] In another embodiment, an indoor positioning device based on a lightweight map is provided, which is applied to a mobile device. The device includes:
[0040] A storage unit for storing a preset indoor positioning model and a preset lightweight map; wherein, the preset indoor positioning model first extracts semantic clues from the two-dimensional image to obtain semantic features, then performs inverse perspective mapping on the semantic features to generate BEV features, then performs feature correction based on the BEV features to obtain complete BEV features, and finally matches the complete BEV features with the preset lightweight map to obtain the indoor pose; the preset lightweight map is a map that discards visual appearance information and retains the spatial layout and relationships of indoor elements;
[0041] An acquisition unit for acquiring two-dimensional images centered on the user; the user is the user holding the mobile device;
[0042] The positioning unit is configured to input the two-dimensional image into the preset positioning model, and obtain the indoor pose corresponding to the user based on the preset indoor positioning model and the preset lightweight map; the indoor pose includes two-dimensional coordinates and an orientation angle;
[0043] The output unit is configured to output the indoor pose in the preset lightweight map.
[0044] In another embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements an indoor positioning method based on a lightweight map.
[0045] In another embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it implements an indoor positioning method based on a lightweight map.
[0046] As can be seen from the above technical solutions, in the above embodiments, a two-dimensional image is collected by a mobile device, and the current indoor pose of the user can be obtained based on a preset indoor positioning model and displayed to the user through the lightweight map to achieve indoor positioning. When determining the indoor pose through the preset indoor positioning model, first, semantic clues are extracted from the two-dimensional image to obtain semantic features, then the inverse perspective mapping is performed on the semantic features to generate BEV features, then the feature correction is performed based on the BEV features to obtain complete BEV features, and finally, the complete BEV features are matched with the preset lightweight map to obtain the indoor pose. This solution can achieve indoor positioning through a lightweight map without being affected by the dynamic environment and without relying on the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 Schematic diagram of the lightweight map in the embodiment of the present application;
[0049] Figure 2 Schematic diagram of the structure of the preset indoor positioning model in the embodiment of the present application;
[0050] Figure 3 Schematic diagram of obtaining layout elements in the embodiment of the present application;
[0051] Figure 4 Schematic diagram of obtaining semantic elements in the embodiment of the present application;
[0052] Figure 5 Schematic diagram of feature conversion in the embodiment of the present application;
[0053] Figure 6 Schematic diagram of different types of BEV projections in the embodiment of the present application;
[0054] Figure 7 Schematic diagram of the process of obtaining the preset feature repair sub-model in the embodiment of the present application;
[0055] Figure 8 Schematic diagram of the process of obtaining the reference BEV view in the embodiment of the present application;
[0056] Figure 9 Schematic diagram of obtaining the first training sample set in the embodiment of the present application;
[0057] Figure 10 Schematic diagram of the process of obtaining the preset feature matching sub-model in the embodiment of the present application
[0058] Figure 11 This is a schematic diagram of the indoor positioning process based on a lightweight map in the embodiments of this application;
[0059] Figure 12 This is a schematic diagram of the structure of the indoor positioning device based on a lightweight map in the embodiments of this application;
[0060] Figure 13 This is a schematic diagram of the physical structure of the electronic device provided in the embodiments of the present invention. Detailed implementation manners
[0061] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0062] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe the order or sequence of the targets. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0063] Next, the technical solutions of the present invention will be described in detail with specific embodiments. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0064] For unfamiliar and complex indoor environments, such as large apartments, office spaces, etc., it is often very easy to get lost. Therefore, how to locate users in such indoor environments is particularly important.
[0065] In related positioning technologies, such as GPS positioning technology, in cities, especially in indoor environments, satellite signals are severely attenuated after being blocked and interfered layer by layer, so the effect is not good; for solutions based on WIFI, Bluetooth, infrared, ultrasonic technology, etc., each has the disadvantage of being affected by environmental factors (such as signal strength, transmission distance, device hardware support, etc.). For SLAM positioning technology, vision-based SLAM is affected by cumulative errors and is not suitable for long-term positioning; laser-based SLAM is limited by cost, and most mobile devices do not support it, and the results are greatly affected by the dynamic indoor environment. Existing methods based on image retrieval and 2D image matching of 3D scenes (feature descriptors, point clouds, 3D reconstruction, etc.) require the server side to store large spatial maps, cannot support device-side computing, and rely on network signals. In addition, existing vision-based positioning methods cannot well solve the problem of dynamic changes in the indoor environment.
[0066] Based on the above problems, the embodiments of the present application provide an indoor positioning method based on a lightweight map. By collecting two-dimensional images, the current indoor pose of the user can be obtained based on a preset indoor positioning model and displayed to the user through this lightweight map to achieve indoor positioning. When determining the indoor pose through the preset indoor positioning model, first extract semantic clues from the two-dimensional image to obtain semantic features, then perform inverse perspective mapping on the semantic features to generate a bird's-eye view (BEV) feature, then perform feature correction based on the BEV feature to obtain a complete BEV feature, and finally match the complete BEV feature with the preset lightweight map to obtain the indoor pose. This solution can achieve indoor positioning through a lightweight map without being affected by the dynamic environment and without relying on the network.
[0067] In the embodiments of the present application, the map view, that is, the lightweight map, is introduced. The lightweight map in the embodiments of the present application is shown in Figure 1 , Figure 1 is a schematic diagram of the lightweight map in the embodiments of the present application.
[0068] Figure 1 The lightweight map in Figure 1 discards the visual appearance information and only retains the spatial layout and relationships of indoor elements. For example, only the elements that build the room structure, such as walls, doors, and windows, electronic devices such as computers and TVs placed in the room, furniture such as sofas and tables, and the spatial layout and relationships of these elements, such as the specific positions of each element, and the front-back, up-down, left-right, etc. relationships of each element.
[0069] Compared with the realistic image map of the 3D reconstruction technology, the lightweight map here requires less storage space, causing less pressure on the storage space of the mobile device implementing this solution. Usually, there will be no problem of non - support for ordinary mobile devices. In addition, the mobile device used in the embodiments of the present application only needs to be able to collect images and does not require additional support for some positioning functions and networks.
[0070] Before performing indoor positioning based on the lightweight map in the embodiments of the present application, it is necessary to first generate a preset lightweight map for the determined indoor scene and establish a preset indoor positioning model based on the preset lightweight map. The following gives the specific implementation process:
[0071] Different lightweight maps are used for different environments. For any lightweight map, it can be an existing one or directly generated. In the embodiments of the present application, the generation method of the lightweight map is not restricted.
[0072] When the preset indoor positioning model in the embodiments of the present application determines the indoor pose for positioning, it first extracts semantic clues from the two - dimensional image to obtain semantic features, then performs inverse perspective mapping on the semantic features to generate BEV features, then performs feature correction based on the BEV features to obtain complete BEV features, and finally matches the complete BEV features with the preset lightweight map to obtain the indoor pose.
[0073] See Figure 2 , Figure 2 which is a schematic structural diagram of the preset indoor positioning model in the embodiments of the present application. Figure 2 The schematic structural diagram of the preset indoor positioning model in includes: a preset feature extraction sub - model, a preset feature conversion sub - model, a preset feature repair sub - model, and a preset feature matching sub - model.
[0074] Among them, the preset feature repair sub - model and the preset feature matching sub - model need to be trained based on the established initial model, and the preset feature extraction sub - model and the preset feature conversion sub - model can be directly established.
[0075] The preset feature extraction sub - model is used to extract semantic clues from the two - dimensional image to obtain semantic features; the preset feature conversion sub - model is used to perform inverse perspective mapping on the semantic features to generate BEV features; the preset feature repair sub - model is based on the BEV features to perform feature correction to obtain complete BEV features; the preset feature matching sub - model is used to match the complete BEV features with the preset lightweight map to obtain the indoor pose.
[0076] The implementation of the preset feature extraction sub - model for extracting semantic clues from the two - dimensional image to obtain semantic features is described as follows:
[0077] The semantic features obtained in the embodiments of this application include two parts, namely layout elements and semantic elements;
[0078] Among them, the layout elements are the structural elements for constructing the room structure; such as the walls, doors, and windows of the room, etc.; they can provide consistent and clear layout clues even in dynamic scenes, and the structural elements of the room structure here usually do not change;
[0079] The semantic elements are the devices placed in the room; such as electronic devices, furniture, etc.
[0080] See Figure 3 , Figure 3 which is a schematic diagram of obtaining layout elements in the embodiments of this application. Figure 3 In [reference], the layout elements are obtained by estimating the room layout of the two-dimensional image through a preset feature extraction sub-model. Specifically: use a CNN+LSTM deep network to predict the layout elements 302 (ceiling, wall, floor) for the two-dimensional image 301; use a detection model with 2D orthogonal transformation to detect the planar surfaces 303 representing the doors and windows on each wall, while minimizing the influence of the user's viewing point and angle, that is, filtering out the elements that affect the user's viewing point and angle; fuse the planar surfaces 303 of the doors and windows on each wall and the predicted layout elements 302 to obtain the layout element information 304.
[0081] See Figure 4 , Figure 4 which is a schematic diagram of obtaining semantic elements in the embodiments of this application. Figure 4 In [reference], the preset feature extraction sub-model obtains semantic elements through instance segmentation. Specifically: the bounding boxes 402 of the semantic elements in the two-dimensional image 401 can be predicted through a customized object detector, and then a segmentation mask 403 corresponding to the two-dimensional image can be generated through a box hint using an edge-based segmentation model, that is, the semantic elements in the two-dimensional image are obtained.
[0082] Here, the object detector can be trained on a subset of the supported categories.
[0083] The following is a specific description of the implementation of the preset feature transformation sub-model for generating BEV features by performing inverse perspective mapping on semantic features:
[0084] In order to restore the hidden spatial information in the 3D space, inverse perspective mapping is used to lift the features of each pixel from the 2D projection space (u, v) to the 3D world space (x, y, z), and then projected to the BEV space (x, y, 0). During the perspective mapping process, the internal and external parameters of the acquisition device for collecting the two-dimensional image need to be collected, and these parameters can be obtained from the acquisition device.
[0085] See Figure 5 , Figure 5Schematic diagram of feature conversion in the embodiments of the present application. Figure 5 The 2D perspective view features 501 of each pixel in Figure 5 are used to obtain the BEV features 503 of multiple pixels through the inverse perspective mapping 502; the BEV features 503 of multiple pixels are fused to obtain the fused BEV features 504.
[0086] When fusing BEV features, that is, when projecting into the BEV space, the same or different pixel features may fall into the same position in the final BEV space. To maximize retention and maintain consistency with the lightweight map, the following projection rules are given in the embodiments of the present application:
[0087] If different types of BEV are projected onto the same two-dimensional space position, the priority of semantic elements is greater than that of layout elements, and the priority of layout elements is greater than that of elements other than semantic elements and layout elements; that is to say, in this case, the priority of semantic elements is the highest, followed by the priority of layout elements, and finally the priority of elements other than layout elements and semantic elements is the lowest;
[0088] See Figure 6 , Figure 6 is the schematic diagram of different types of BEV projection in the embodiments of the present application. As Figure 6 in Figure 6 , if the layout element 601 and the semantic element 602 need to be projected to the same position (x, y), then the semantic element 602 is displayed at the position (x, y).
[0089] If the same type of BEV is projected onto the same two-dimensional space position, the average feature in the corresponding grid cell is used.
[0090] Due to occlusion, some unobservable parts of the 2D transmission image have no corresponding relationship in the BEV space, resulting in incomplete or inaccurate semantic features. By learning the true BEV representation of indoor elements within the field of view, a depth neural network is trained to obtain a preset feature repair sub-model to correct the BEV features, and then complete BEV features are obtained.
[0091] The training process of the preset feature repair sub-model is specifically as follows:
[0092] See Figure 7 , Figure 7 is the schematic diagram of the process for obtaining the preset feature repair sub-model in the embodiments of the present application. The specific steps are as follows:
[0093] Step 701, obtain the first training sample set; wherein, the first training sample set is a set of multiple two-dimensional images and reference BEV views; the reference BEV view is determined according to the pose information of radar positioning and the preset lightweight map.
[0094] The training of the preset feature repair sub-model requires a large amount of real (GT) data (refer to the BEV view). In related technologies, a radar (LiDAR) can be used to generate a BEV occupancy grid map. However, LiDAR scans every detail, so it is affected in dynamic scenarios. Although 3D segmentation can solve this problem, the computational efficiency is very low. If a preset lightweight map is used as GT data, the deep neural network can indeed repair incomplete parts, but this may also lead to unexpected "imagination" because the GT contains parts that are not observable from the user's perspective.
[0095] In the embodiments of the present application, a method for automatically generating GT data is proposed to prevent "imagination". By having a robot drive indoors to obtain the correspondence between two-dimensional images and poses, the present application can efficiently and accurately obtain the BEV view, and then efficiently and accurately train to obtain the preset feature repair sub-model. The specific acquisition process is as follows:
[0096] For the acquisition of the reference BEV view, refer to Figure 8 , Figure 8 which is a schematic diagram of the reference BEV view acquisition process in the embodiments of the present application. The specific steps are as follows:
[0097] Step 801: Obtain pose information; where the pose information is synchronously obtained by radar positioning when a camera captures an image.
[0098] The two-dimensional image collected here and the reference BEV view finally obtained through the pose information Figure 1 are in one-to-one correspondence as a set of training samples.
[0099] Step 802: Adjust the preset lightweight map based on the orientation angle in the pose information; and crop a map of a preset size in the adjusted preset lightweight map based on the two-dimensional coordinates in the pose information.
[0100] Step 803: Project rays starting from the pose information, hit the parts within the visual range, and filter out invisible elements to obtain the reference BEV view.
[0101] In this step, in order to simulate the real world observed from the user's perspective, the rays are emitted from the pose and hit the observable parts. Invisible parts, such as things blocked behind a wall, are filtered out to avoid unnecessary hallucinations.
[0102] In the embodiments of the present application, the neural network parameters involved in training the preset indoor positioning model are generated by automatically generating reference values in the samples, making the present solution practical and scalable when deployed in various indoor environments.
[0103] Refer to Figure 9 , Figure 9Schematic diagram for obtaining the first training sample set in the embodiments of the present application. Figure 9 The radar 901 in it obtains the pose information 903. At the same time, the camera 902 obtains the two-dimensional image 904. The obtained pose information 903 and the two-dimensional image 904 have a one-to-one correspondence.
[0104] Adjust the preset lightweight map 905 based on the orientation angle in the pose information 903; and crop the map with a preset size in the adjusted preset lightweight map based on the two-dimensional coordinates in the pose information, that is, the cropped lightweight map 906.
[0105] Project rays starting from the pose information to hit the part 907 within the visual range of the cropped lightweight map, and filter out invisible elements to obtain the reference BEV view 908.
[0106] Based on the one-to-one correspondence between the pose information 903 and the two-dimensional image 904, and the one-to-one correspondence between the reference BEV view 908 and the corresponding two-dimensional image 904, use the reference BEV view 908 and the two-dimensional image 904 as a group of first training samples.
[0107] Step 702, establish an initial feature repair sub-model.
[0108] The initial feature repair sub-model here can also be an encoder. By learning the true BEV representation of indoor elements within the field of view, establish a deep neural network model as the initial feature repair sub-model.
[0109] Step 703, train the initial feature repair sub-model based on the first training sample set until the similarity between the BEV features output by the initial feature repair sub-model in the decoded BEV view and the reference BEV view meets the first preset condition, and obtain the preset feature repair sub-model.
[0110] The first preset condition here can be set according to actual needs. For example, set a preset threshold. When the similarity is greater than the preset threshold, it is considered to meet the first preset condition, and it can be determined that the training is completed.
[0111] Here, a loss function can also be established. Calculate the loss function value based on the reference BEV view and the BEV view obtained by the model, and perform gradient iteration to obtain the model with the minimum error as the preset feature repair sub-model. The specific training process is not limited here.
[0112] The implementation of the feature matching sub-model for matching the complete BEV features with the preset lightweight map to obtain the indoor pose is specifically described as follows:
[0113] See Figure 10 , Figure 10This is a flowchart for obtaining a preset feature matching sub - model in an embodiment of the present application. The specific steps are as follows:
[0114] Step 1001: Obtain a second training sample set; where the second training sample set is a set of multiple two - dimensional images and reference indoor poses; the reference indoor pose is obtained through radar positioning.
[0115] When collecting two - dimensional images, collect the indoor pose at the same time, and use this indoor pose as the reference pose for training the initial feature matching sub - model
[0116] Step 1002: Establish an initial feature matching sub - model.
[0117] Step 1003: Train the initial feature matching sub - model based on the second training sample set until the similarity between the indoor pose output by the initial feature matching sub - model and the corresponding reference indoor pose meets the second preset condition, and then obtain the preset feature matching sub - model.
[0118] The second preset condition here can be that the similarity is greater than the second preset threshold, or a loss function can be established to determine the gap between the indoor position output by the initial feature matching sub - model and the reference indoor pose to determine the training end condition. The embodiments of the present application do not limit this.
[0119] In the embodiments of the present application, the structures of the initial feature matching sub - model and the preset feature matching sub - model are the same. Both are based on a module of a deep neural network (Transformer) to predict the positioning pose in an end - to - end manner; the structure of the Transformer includes an encoder and a decoder. The input of the encoder of the Transformer is the repaired, that is, the complete BEV feature. The input (query input) of the decoder of the Transformer is the encoded preset lightweight map, and the encoded preset lightweight map consists of learnable semantic embeddings and the poses of all indoor elements.
[0120] The Transformer uses the self - attention mechanism to capture the spatial relationship between the BEV feature and the map query elements respectively, and then uses cross - attention to learn the mutual relationship between the BEV feature and the map query elements.
[0121] Output a lightweight map matching result through the Transformer; and decode the lightweight map matching result into an indoor pose.
[0122] The feature matching sub - model is used to match the complete BEV feature with the preset lightweight map to obtain the indoor pose, specifically including:
[0123] The complete BEV features, i.e., the repaired BEV features, are used as the input of the encoder in the form of positional encoding and BEV feature tiling. The encoded preset lightweight map (semantic embedding information and the poses of indoor elements) is used as the input of the decoder, and mapping queries are performed to obtain query results, i.e., matching results. Then, the matching results are decoded into indoor poses through MPL.
[0124] The indoor positioning process based on the lightweight map in the embodiments of the present application is given below in conjunction with the accompanying drawings.
[0125] See Figure 11 , Figure 11 which is a schematic diagram of the indoor positioning process based on the lightweight map in the embodiments of the present application. The specific steps are as follows:
[0126] Step 1101, the mobile device collects two-dimensional images centered on the user; the user is the user holding the mobile device.
[0127] Step 1102, the mobile device inputs the two-dimensional images into a preset positioning model and obtains the indoor pose corresponding to the user based on the preset indoor positioning model and the preset lightweight map; wherein, the indoor pose includes two-dimensional coordinates and an orientation angle.
[0128] The preset lightweight map is a map that discards the visual appearance information and retains the spatial layout and relationships of indoor elements.
[0129] The preset indoor positioning model first extracts semantic clues from the two-dimensional images to obtain semantic features, then performs inverse perspective mapping on the semantic features to generate BEV features, then performs feature correction based on the BEV features to obtain complete BEV features, and finally matches the complete BEV features with the preset lightweight map to obtain the indoor pose.
[0130] Step 1103, the mobile device outputs the indoor pose in the preset lightweight map.
[0131] In the embodiments of the present application, only the mobile device needs to collect two-dimensional images centered on the user and upload them to the indoor positioning device based on the lightweight map deployed on the mobile device to obtain the indoor pose. This solution only requires visual information, does not rely on signals, and the established positioning model is robust to dynamic scenes (overcoming the challenges of dynamic scenes by using more robust layouts and semantic clues), can be applied to different indoor environments, and realizes indoor positioning based on the lightweight map.
[0132] Regarding the characteristics of the bird's-eye view based on the lightweight map in the established preset indoor positioning model, different types of element features are converted and projected onto the bird's-eye view, and the filtering of partially invisible elements and completely invisible elements is introduced to alleviate the occlusion problem. These implementations can improve the positioning accuracy through more accurate environmental feature modeling;
[0133] And by automatically generating reference values in the samples to train the neural network parameters involved in the preset indoor positioning model, the proposed solution is practical and scalable when deployed in various indoor environments.
[0134] The indoor positioning solution based on the lightweight map in the embodiments of the present application can have various applications, specifically as follows:
[0135] The first type of application: Indoor AR navigation;
[0136] A user using the positioning solution provided by this solution may get lost between different office buildings or floors, and the signal is insufficient for trilateration positioning. The user uses a mobile device to collect a two-dimensional image centered on themselves and inputs it into the indoor positioning device based on the lightweight map, and can input the indoor pose of the user on the preset lightweight map, so that the user or themselves are at the position shown on the map;
[0137] Furthermore, it can be applied to navigation. For example, the user can input the destination in the preset lightweight map, and then the indoor pose determined in the embodiments of the present application can be used to perform positioning navigation for the user.
[0138] The positioning function here can also be extended to indoor positioning scenarios that require it. For example, it can help emergency responders navigate in complex buildings, while helping dangerous personnel report their positions, and can also help Walmart / Amazon assistants with home deliveries, providing real-time tracking and monitoring for delivery personnel to ensure that customers are not worried about theft and privacy issues.
[0139] The second type of application: Information recommendation based on geographical location personalization;
[0140] When a customer is in a large shopping mall, the indoor position of the user can be located through the two-dimensional picture provided by the customer, and nearby stores or products can be recommended to the customer according to this indoor pose.
[0141] The third type of application: Smart home control;
[0142] The positioning solution in the embodiments of the present application enables the indoor automation system to trigger operations according to the indoor pose of the user, such as lighting, temperature, volume, etc. control, and these implementations can be achieved without signals or detection sensors.
[0143] Any combination of the above optional technical solutions can form an optional embodiment of the present disclosure, which will not be elaborated one by one here.
[0144] Based on the same inventive concept, an indoor positioning device based on a lightweight map is also provided in an embodiment of the present application, which is applied to a mobile device. Refer to Figure 12 , Figure 12 which is a schematic structural diagram of the indoor positioning device of the lightweight map in the embodiment of the present application. The device includes:
[0145] A storage unit 1201, configured to store a preset indoor positioning model and a preset lightweight map; wherein, the preset indoor positioning model first extracts semantic clues from a two-dimensional image to obtain semantic features, then performs inverse perspective mapping on the semantic features to generate BEV features, then performs feature correction based on the BEV features to obtain complete BEV features, and finally matches the complete BEV features with the preset lightweight map to obtain the indoor pose; the preset lightweight map is a map that discards visual appearance information and retains the spatial layout and relationships of indoor elements;
[0146] An acquisition unit 1202, configured to acquire two-dimensional images centered on the user; the user is the user holding the mobile device;
[0147] A positioning unit 1203, configured to input the two-dimensional image into the preset positioning model, and obtain the indoor pose corresponding to the user based on the preset indoor positioning model and the preset lightweight map; wherein, the indoor pose includes two-dimensional coordinates and an orientation angle;
[0148] An output unit 1204, configured to output the indoor pose in the preset lightweight map.
[0149] In another example,
[0150] The preset indoor positioning model includes: a preset feature extraction sub-model, a preset feature conversion sub-model, a preset feature repair sub-model, and a preset feature matching sub-model; the preset feature extraction sub-model is configured to extract semantic clues from a two-dimensional image to obtain semantic features; the preset feature conversion sub-model is configured to perform inverse perspective mapping on the semantic features to generate BEV features; the preset feature repair sub-model performs feature correction based on the BEV features to obtain complete BEV features; the preset feature matching sub-model is configured to match the complete BEV features with the preset lightweight map to obtain the indoor pose.
[0151] In another example,
[0152] The semantic features include: layout elements and semantic elements; the layout elements are elements for constructing the room structure; the semantic elements are devices placed in the room;
[0153] The preset feature extraction sub-model is used to extract semantic clues from the two-dimensional image to obtain semantic features, including:
[0154] Obtain layout elements through room layout estimation, and filter out elements that affect the user's viewing point and angle;
[0155] Obtain semantic elements through instance segmentation.
[0156] In another example,
[0157] The preset feature transformation sub-model is used to generate BEV features by performing inverse perspective mapping on the semantic features, including:
[0158] The preset feature transformation sub-model uses inverse perspective mapping to lift the semantic features from the two-dimensional perspective space to the three-dimensional world space, and then projects them into the BEV space to obtain BEV features.
[0159] In another example,
[0160] The positioning unit 1203 is specifically used when projecting into the BEV space. If different types of BEVs are projected onto the same two-dimensional space position, the priority of the semantic elements is greater than that of the layout elements, and the priority of the layout elements is greater than that of the elements other than the semantic elements and the layout elements; if the same type of BEV is projected onto the same two-dimensional space position, the average feature in the corresponding grid cell is used.
[0161] In another example,
[0162] The training of the preset feature repair sub-model in the storage unit 1201 includes:
[0163] Obtain the first training sample set; among them, the first training sample set is a set of multiple two-dimensional images and reference BEV views; the reference BEV view is determined according to the pose information of radar positioning and the preset lightweight map;
[0164] Establish an initial feature repair sub-model;
[0165] Train the initial feature repair sub-model based on the first training sample set until the similarity between the BEV feature output by the initial feature repair sub-model in the encoded and decoded BEV view and the reference BEV view meets the first preset condition, and obtain the preset feature repair sub-model.
[0166] In another example,
[0167] The storage unit 1201 is further configured to determine a reference BEV view according to the pose information obtained by radar positioning and a preset lightweight map, including: obtaining the pose information; the pose information is synchronously obtained by radar positioning when the camera captures images; adjusting the preset lightweight map based on the orientation angle in the pose information; and cropping a map with a preset size in the preset lightweight map adjusted by the angle based on the two-dimensional coordinates in the pose information; obtaining the reference BEV view by projecting rays starting from the pose information, hitting the parts within the visual range, and filtering out invisible elements.
[0168] In another example,
[0169] The positioning unit 1203 is specifically configured to, when obtaining the indoor pose by matching the complete BEV feature with the preset lightweight map, the feature matching sub-model predicts the indoor pose based on the Transformer; use the complete BEV feature as the input of the encoder of the Transformer; use the encoded preset lightweight map as the input of the decoder of the Transformer; output the lightweight map matching result through the Transformer; and decode the lightweight map matching result into the indoor pose.
[0170] In another example,
[0171] The training of the preset feature matching sub-model in the storage unit 1201 includes:
[0172] Obtaining a second training sample set; wherein, the second training sample set is a set of multiple two-dimensional images and reference indoor poses; the reference indoor poses are obtained by radar positioning;
[0173] Establishing an initial feature matching sub-model;
[0174] Training the initial feature matching sub-model based on the second training sample set until the similarity between the indoor pose output by the initial feature matching sub-model and the corresponding reference indoor pose meets the second preset condition, and obtaining the preset feature matching sub-model.
[0175] The units in the above embodiments can be integrated into one body or separately deployed; they can be combined into one unit or further split into multiple sub-units.
[0176] In another embodiment, an electronic device is further provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the indoor positioning method based on the lightweight map.
[0177] In another embodiment, a computer-readable storage medium is further provided, on which computer instructions are stored. When the instructions are executed by the processor, the indoor positioning method based on the lightweight map is implemented.
[0178] Figure 13 This is a schematic diagram of the physical structure of the electronic device provided by the embodiment of the present invention. As Figure 13 shown, the electronic device may include: a processor 1310, a communications interface 1320, a memory 1330, and a communication bus 1340. Among them, the processor 1310, the communications interface 1320, and the memory 1330 complete mutual communication through the communication bus 1340. The processor 1310 can call the logical instructions in the memory 1330 to execute the following methods:
[0179] Collect two-dimensional images centered on the user; the user is the user holding the mobile device;
[0180] Obtain the indoor pose corresponding to the user based on a preset indoor positioning model and a preset lightweight map; the indoor pose includes two-dimensional coordinates and an orientation angle; the preset lightweight map is a map that discards the visual appearance information and retains the spatial layout and relationships of indoor elements;
[0181] Output the indoor pose in the preset lightweight map;
[0182] Among them, the preset indoor positioning model first extracts semantic features from the two-dimensional image to obtain semantic features, then performs inverse perspective mapping on the semantic features to generate BEV features, then performs feature correction based on the BEV features to obtain complete BEV features, and finally matches the complete BEV features with the preset lightweight map to obtain the indoor pose.
[0183] In addition, when the logical instructions in the above-mentioned memory 1330 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.
[0184] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0185] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0186] The flowcharts and block diagrams in the drawings of this application illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments disclosed in this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or part of the code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in the order marked in different drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0187] Those skilled in the art can understand that the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recorded in this application. In particular, without departing from the spirit and teachings of this application, the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways, and all such combinations and / or combinations fall within the scope disclosed in this application.
[0188] In this text, specific embodiments are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea, and is not used to limit this application. For those skilled in the art, they can make changes in the specific implementation manners and application scope according to the idea, spirit and principle of the present invention. Any modifications, equivalent replacements, improvements, etc. made by them shall be included within the scope protected by this application.
Claims
1. A lightweight map-based indoor positioning method, characterized in that: Applied to a mobile device, the method comprises: Capturing a two-dimensional image with a user as the center; the user is a user holding the mobile device; The two-dimensional image is input into a preset positioning model, and the indoor posture corresponding to the user is obtained based on the preset indoor positioning model and a preset lightweight map; the indoor posture includes two-dimensional coordinates and a direction angle; the preset lightweight map is a map that discards visual appearance information and retains the spatial layout and relationship of indoor elements; Outputting the indoor posture in the preset lightweight map; Among them, the preset indoor positioning model first extracts semantic clues from the two-dimensional image to obtain semantic features, then performs inverse perspective mapping on the semantic features to generate BEV features, then performs feature correction based on the BEV features to obtain complete BEV features, and finally matches the complete BEV features with the preset lightweight map to obtain the indoor posture.
2. The method according to claim 1, characterized in that The preset indoor positioning model includes: a preset feature extraction sub-model, a preset feature conversion sub-model, a preset feature repair sub-model and a preset feature matching sub-model; the preset feature extraction sub-model is used to extract semantic clues from the two-dimensional image to obtain semantic features; the preset feature conversion sub-model is used to perform inverse perspective mapping on the semantic features to generate BEV features; the preset feature repair sub-model performs feature correction based on the BEV features to obtain complete BEV features; the preset feature matching sub-model is used to match the complete BEV features with the preset lightweight map to obtain the indoor posture.
3. The method according to claim 2, characterized in that The semantic features include: layout elements and semantic elements; the layout elements are elements for constructing the room structure; the semantic elements are devices placed in the room; The preset feature extraction sub-model is used to extract semantic clues from the two-dimensional image to obtain semantic features, including: Obtaining the layout elements through room layout estimation, and filtering out elements that affect the user's viewpoint and angle; Semantic elements are obtained through instance segmentation.
4. The method according to claim 3, characterized in that The preset feature conversion sub-model is used to perform inverse perspective mapping on the semantic features to generate BEV features, including: The preset feature conversion sub-model uses inverse perspective mapping to promote the semantic features from the two-dimensional perspective space to the three-dimensional world space, and then projects them into the BEV space to obtain the BEV features.
5. The method according to claim 4, characterized in that When projecting into the BEV space, the method further comprises: If different types of BEVs are projected to the same two-dimensional space position, the priority of the semantic element is greater than the priority of the layout element, and the priority of the layout element is greater than the priority of the elements other than the semantic element and the layout element; If BEVs of the same type are projected to the same 2D spatial location, the average features in the corresponding grid cells are used.
6. The method according to claim 2, characterized in that The training of the preset feature repair sub-model includes: Acquire a first training sample set; wherein the first training sample set is a collection of multiple groups of two-dimensional images and reference BEV views; the reference BEV view is determined according to the posture information of radar positioning and the preset lightweight map; Establishing the initial feature repair sub-model; The initial feature restoration sub-model is trained based on the first training sample set until the BEV features output by the initial feature restoration sub-model meet a first preset condition in the similarity between the encoded and decoded BEV view and the reference BEV view, thereby obtaining the preset feature restoration sub-model.
7. The method according to claim 6, characterized in that Determining a reference BEV view according to the radar positioning posture information and the preset lightweight map includes: Acquire posture information; the posture information is acquired synchronously through radar positioning when the camera collects images; Adjusting the preset lightweight map based on the orientation angle in the posture information; and cropping a map of a preset size in the preset lightweight map after the angle is adjusted based on the two-dimensional coordinates in the posture information; The reference BEV view is obtained by projecting rays starting from the pose information, hitting the parts within the visual range, and filtering the invisible elements.
8. The method according to claim 2, characterized in that: The feature matching sub-model is used to match the complete BEV feature with the preset lightweight map to obtain the indoor posture, including: The feature matching sub-model predicts the indoor posture based on Transformer; Using the complete BEV features as input to the encoder of the Transformer; Using the encoded preset lightweight map as input of the decoder of the Transformer; Outputting a lightweight map matching result through the Transformer; and decoding the lightweight map matching result into the indoor pose.
9. The method according to claim 8, characterized in that The training of the preset feature matching sub-model includes: Acquire a second training sample set; wherein the second training sample set is a collection of multiple groups of two-dimensional images and reference indoor postures; the reference indoor posture is acquired by radar positioning; Establishing an initial feature matching sub-model; The initial feature matching sub-model is trained based on the second training sample set until the similarity between the indoor posture output by the initial feature matching sub-model and the corresponding reference indoor posture meets a second preset condition, thereby obtaining a preset feature matching sub-model.
10. An indoor positioning device based on a lightweight map, characterized in that: Applied to a mobile device, the device comprises: A storage unit, for storing a preset indoor positioning model and a preset lightweight map; wherein the preset indoor positioning model first extracts semantic clues from the two-dimensional image to obtain semantic features, then performs inverse perspective mapping on the semantic features to generate BEV features, then performs feature correction based on the BEV features to obtain complete BEV features, and finally matches the complete BEV features with the preset lightweight map to obtain the indoor posture; the preset lightweight map is a map that discards visual appearance information and retains the spatial layout and relationship of indoor elements; A collection unit, used for collecting two-dimensional images with a user as the center; the user is a user holding the mobile device; The positioning unit is used to input the two-dimensional image into the preset positioning model, and obtain the indoor posture corresponding to the user based on the preset indoor positioning model and the preset lightweight map; the indoor posture includes two-dimensional coordinates and orientation angles; The output unit is used to output the indoor posture in the preset lightweight map.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 9 is implemented.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method described in any one of claims 1 to 9 is implemented.