An Automatic Indoor Map Construction Method Integrating Crowdsourced Trajectories and Mobile Phone Images
By integrating the crowd source trajectory and mobile phone images collected by smartphones, combining behavior recognition and visual scene understanding technology, efficient and low-cost automatic construction of indoor maps is achieved, solving the problems of lack of data and low mapping efficiency in the existing technology, and improving the accuracy and authenticity of the map.
Patent Information
- Application Number
- CN202411570254.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-11-05
AI Technical Summary
The existing indoor map construction technology has problems such as lack of data, low mapping accuracy and low efficiency. Especially under the conditions of complex signal environments and variable spatial layout, it is difficult to achieve efficient and low-cost indoor map construction.
An indoor map automatic construction method that integrates the crowd source trajectory and mobile phone images is adopted. The indoor navigation map is constructed through the multi-sensor data collected by the smartphone and the behavior recognition algorithm, and the indoor planar structure is extracted through visual scene understanding technology, and combined with navigation paths for fusion and optimization.
It realizes efficient, low-cost and dependency-free indoor map feature plane structure extraction, improves the authenticity and accuracy of the map, saves labor costs and reduces time and calculation amount.
Smart Images

Figure CN119437202B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of indoor map construction, and in particular to an automatic indoor map construction method that integrates crowdsourced trajectories and mobile phone images. Background Art
[0002] With the rapid advancement of urbanization construction, indoor spaces are continuously expanding, and research on indoor space information applications has received increasing attention from the academic and industrial communities. The market scale of indoor applications centered on location-based services is growing rapidly. As the data foundation for indoor location applications, due to the influence of complex signal environments, spatial layouts, and topological variability, indoor two-dimensional maps have problems such as data scarcity, low mapping accuracy, and low efficiency, becoming the main factors restricting the development of indoor location-based services (LBS). Traditional indoor two-dimensional plane map construction can be mainly divided into two methods based on data collection methods: manual mapping and professional sensors. The manual mapping method uses a tape measure, steel ruler, handheld laser rangefinder, total station, etc. as basic measurement instruments, supplemented by mapping software such as AutoCAD for construction. Although it can effectively ensure map accuracy, manual modeling requires a large amount of time and effort. The method based on professional sensors is to use a laser scanner or camera to collect indoor space information, thereby generating an indoor two-dimensional map. However, professional mapping equipment has high requirements for collectors and high hardware costs, making it difficult to popularize and promote in applications.
[0003] Currently, the hardware of smart phones is updated extremely fast, with a variety of built-in sensors (accelerometers, gyroscopes, magnetometers, barometers, GPS, etc.), making it not only have strong computing power but also be able to obtain multi-source data (indoor movement trajectories, behavior data, images, etc.), which is an ideal platform for indoor mapping. Scholars at home and abroad have explored various indoor mapping methods for different data sources. Among them, based on SLAM
[0004] (Simultaneous Localization and Mapping, SLAM) smartphone indoor mapping methods are gradually maturing. Scholars have also conducted a lot of exploration and promoted the development of SLAM technology in terms of mapping accuracy, scale, reliability, etc. However, such methods require users to always turn on the camera and aim it at the surrounding environment, which does not conform to the user's usage habits. In addition, the algorithm heavily relies on image feature points in the scene, and mis-matching situations are widespread. Moreover, it has high requirements for the memory and computing power of the mobile phone during use. In addition, the smartphone indoor mapping method based on crowdsourced data effectively reduces the cost and technical threshold of data collection. However, it lacks effective management of the time, efficiency, and quality of data collection. Due to the non-professional nature of data collection, it is difficult to guarantee the accuracy of map construction, and a large amount of energy is required for data screening and effective information extraction. Summary of the Invention
[0005] The object of the present invention is to provide an indoor map automatic construction method that fuses crowdsourced trajectories and mobile phone images. By using the crowdsourced trajectory data collected by the smartphone platform and collecting indoor images according to the formulated photographing rules, the accurate construction of the indoor map can be achieved, providing a low-cost, high-precision, and high-efficiency mapping method for indoor plane map construction.
[0006] To achieve the above object, the present invention provides an indoor map automatic construction method that fuses crowdsourced trajectories and mobile phone images, including the following steps:
[0007] S1. Use a smartphone to collect corridor, indoor images, and multi-sensor data;
[0008] S2. Construct an indoor navigation map;
[0009] S3. Extract the indoor plane structure;
[0010] S4. Fuse and optimize the navigation path and the indoor plane structure.
[0011] Preferably, S2 includes the following steps:
[0012] S21. Obtain crowdsourced trajectory data from the multi-sensor data collected by the smartphone platform through pedestrian dead reckoning;
[0013] S22. Obtain behavior landmarks through a behavior recognition algorithm, and cluster the behavior landmarks through a clustering algorithm based on Wi-Fi fingerprints;
[0014] S23. Calculate the relative distances between the behavior landmarks to obtain an indoor navigation map.
[0015] Preferably, S3 includes the following steps:
[0016] S31. Input the monocular image of the indoor space collected by the smart phone, and extract the three-dimensional layout structure of the monocular image through the indoor three-dimensional scene understanding network model;
[0017] S32. Map the indoor 3D scene to a 2D image and perform projective transformation;
[0018] S33. Restore the spatial structure ratio to obtain the indoor plane structure.
[0019] Preferably, S32 includes the following steps:
[0020] S321. Use the VPBC technology to solve the camera parameters based on the vanishing point. The camera principal point is located at the center of the image, and the form of the camera intrinsic matrix K is:
[0021]
[0022] f is the focal length, which can be obtained from the vanishing point;
[0023] S322. Assume that the coordinates of the vanishing point in the image coordinate system are vp0(x i0 , y i0 ), vp1(x i1 , y i1 ), vp2(x i2 , y i2 ). Since the centroid of the triangle formed by the three vanishing points is the intersection point p(ox, oy) of the camera optical axis and the imaging plane, the coordinates of vp0 in the camera coordinate system are (x c0 , y c0 , f), where x c0 = x i0 - ox, y c0 = -(y i0 - oy). The coordinates of the other two vanishing points can be obtained similarly. Using the orthogonal relationship between the vanishing points, the focal length and the rotation matrix are respectively:
[0024]
[0025] S323. Calculate the projection transformation matrix H ij , and the plane determined by the vanishing points vp i and vp j can be "orthographically projected" and corrected. The solution of H ij is as follows:
[0026] H ij = K·R ij ·K -1 ;
[0027] Use the transformation matrices H 01 and H of the front wall and the floor02 Reproject the 4 intersection points to obtain the coordinates of each point in the front view after distortion elimination.
[0028]
[0029] Preferably, the true proportional relationship of the length, width and height of the indoor space is:
[0030]
[0031] Preferably, S4 includes the following steps:
[0032] S41. Corridor construction: Match the corrected trajectory route with the corridor plane structure extracted from the mobile phone image to realize the map construction of the corridor area.
[0033] S42. Match the room plane structure extracted from the image with the inertial navigation data to realize the drawing of the room structure in the indoor map.
[0034] Preferably, the map accuracy is further optimized by constructing an energy function. The energy equation to be constructed is as follows:
[0035] min x ={x f ,R,t}E l (x)+E c (x)+E b (x);
[0036] Where E l (x) represents the complexity of the layout, E c (x) represents the closure, E b (x) represents the boundary similarity between adjacent segments, and {x f ,R,t} represents the rotation and displacement when all corridor paths and room layouts are connected.
[0037] Preferably, the conditions for restricting the uniqueness of the corridor and room layout are:
[0038]
[0039] Preferably, during the image acquisition process, the following rules are followed: For shooting the indoor corridor, all fields of view facing the turning points need to be photographed in clockwise order during the turning process, and for shooting the indoor room, it needs to be shot from both sides of the room towards the opposite side.
[0040] Therefore, the indoor map automatic construction method of the present invention adopting the above method of fusing multi-source trajectories and mobile phone images has the following beneficial effects:
[0041] (1) A method for extracting the planar structure of indoor map elements with high efficiency, low cost and no dependence is proposed. Using the pictures taken by mobile phones as the data source, it is not necessary to reconstruct the geometric structure of the indoor three-dimensional model, nor is it necessary to use professional measurement equipment. Only by using the monocular pictures of the indoor scene taken by a smart phone and through visual scene understanding technology, the acquisition of the planar structure of indoor map elements is realized, effectively improving the authenticity and accuracy of indoor map expression.
[0042] (2) Only through different types of data collected by smart phones, fusing crowd-sourced trajectories and standardized mobile image acquisition, and using a new geometric feature extraction method, the high-precision construction of indoor maps is realized. Giving full play to the complementary advantages of the overall and local, framework and details of the indoor space features, it is not necessary for users to walk through all positions in the indoor environment, effectively saving the labor cost of indoor map drawing and reducing the time and calculation amount.
[0043] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0044] Figure 1 It is a flowchart of the present invention;
[0045] Figure 2 It is a schematic diagram of relative distance estimation of the present invention;
[0046] Figure 3 It is the step detection and angle measurement results between behavior landmarks of the present invention;
[0047] Figure 4 It is the corridor image acquisition rule of the present invention, where a is a right-angle turn; b is a four-way corner; c is a T-shaped corner;
[0048] Figure 5 It is the room image acquisition rule of the present invention. Detailed Embodiment
[0049] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0050] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention pertains. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "comprising" or "including" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0051] Embodiment
[0052] Please refer to Figures 1-5 , the present invention provides an indoor map automatic construction method that fuses crowdsourced trajectories and mobile phone images. Through three modules: indoor navigation map construction, indoor plane structure extraction, and indoor map construction by fusing navigation paths and plane structures, only by using the crowdsourced trajectory data collected by the smartphone platform and collecting indoor images according to the formulated photographing rules, the precise construction of the indoor map can be achieved, providing a low-cost, high-precision, and high-efficiency mapping method for indoor plane map construction.
[0053] An indoor map automatic construction method that fuses crowdsourced trajectories and mobile phone images includes the following steps:
[0054] S1. Use a smartphone to collect corridor, indoor images and multi-sensor data.
[0055] S2. Construct an indoor navigation map.
[0056] S21. Obtain crowdsourced trajectory data from the multi-sensor data collected by the smartphone platform through pedestrian dead reckoning. There are many behavioral landmarks in the crowdsourcing data, and some of the behavioral landmarks are collected at the same nodes. In order to construct an indoor map, it is first necessary to cluster the behavioral landmarks collected at the same nodes.
[0057] S22. Obtain behavioral landmarks through a behavior recognition algorithm, and cluster the behavioral landmarks through a clustering algorithm based on Wi-Fi fingerprints.
[0058] When a new trajectory is obtained, use the behavior recognition algorithm to obtain behavioral landmarks, and at the same time extract the Wi-Fi fingerprint at the moment when the behavior occurs as the feature of the behavioral landmark. Obtain a sequence of behavioral landmarks {NAL1, NAL2...NAL m}。Among them, there are m behavior landmarks (NAL represents the behavior landmarks extracted from the newly uploaded trajectory). Use the Wi-Fi fingerprint-based clustering algorithm to cluster the m behavior landmarks, and there may be n behavior landmarks with the same Wi-Fi fingerprint characteristics, where (0 ≤ n ≤ m).
[0059] If this is the first trajectory, then the behavior landmarks included in this trajectory are added to the node database in sequence. If this is not the first trajectory data, then through the Wi-Fi feature-based clustering algorithm, a large number of pictures of existing garbage can be collected from this trajectory to construct a training dataset.
[0060] S23. Calculate the relative distances between the behavior landmarks to obtain an indoor navigation map.
[0061] After the behavior landmark clustering, all the behavior landmarks are clustered into different classes, and each class represents a node of the indoor map. If two nodes are directly adjacent, the distance between them can be obtained through dead reckoning. As Figure 4 (a) shows, in the figure is a trajectory collected by a smartphone (obtained through dead reckoning), and A, B, C, and D are four behavior landmarks (turns). Among them, AB, BC, and CD are adjacent behavior landmarks (known from the time sequence of landmark detection), and the relative distances between them can be directly obtained through step detection and step length estimation. For non-adjacent landmarks, such as A and C, the lengths of AB and BC and the angle information of angle B are required to calculate the relative distance between the landmarks. Figure 4 The distance and angle information between the behavior landmarks in (a) can be obtained from the inertial data acquired by the smartphone. The distance information can be obtained through step detection using the accelerometer data, and the angle information is obtained through the gyroscope and electronic compass. Figure 3 Shows the step detection results and angle change information between different behavior landmarks.
[0062] Based on the behavior landmark clustering and relative distance calculation, the points for constructing the indoor navigation map and the relative distances between all points are obtained, forming a relative distance matrix. Using multidimensional scaling technology, with the relative distance between points as the non-similarity measurement parameter, calculate their relative spatial relationship to construct an indoor navigation map.
[0063] S3. Extract the indoor plane structure.
[0064] S31. Input the monocular image of the indoor space collected by the smartphone, and extract the three-dimensional layout structure of the monocular image through the indoor three-dimensional scene understanding network model.
[0065] Train an initial framework through Convolutional Neural Networks (CNN), and use it as a feature vector to input into a structured support vector machine to improve the extraction accuracy of indoor three-dimensional space features. Among them, the indoor layout estimation based on CNN estimates an ordered set of indoor layout key points by directly training room corners (key points) and room types, and connects them in a specific order to obtain the indoor layout framework.
[0066] S32. Map the indoor 3D scene to a 2D image and perform a projection transformation.
[0067] Since the structural information of the length, width, and height of the three-dimensional space layout extracted from the mobile phone image is represented by pixel distances, during the process of mapping the indoor 3D scene to a 2D image, the image will be distorted to varying degrees, resulting in distortion of the spatial structure ratio. Therefore, it is necessary to eliminate the influence of image imaging distortion and restore the true proportional relationship among the length, width, and height in the indoor space.
[0068] Use the VPBC technology to solve the camera parameters based on the vanishing point. The camera principal point is located at the center of the image, and the form of the camera internal parameter matrix K is:
[0069]
[0070] f is the focal length, which can be obtained from the vanishing point: Let the coordinates of the vanishing point in the image coordinate system be vp0(x i0 ,y i0 ), vp1(x i1 ,y i1 ), vp2(x i2 ,y i2 ). Since the centroid of the triangle formed by the three vanishing points is the intersection point p(ox,oy) of the camera optical axis and the imaging plane, the coordinates of vp0 in the camera coordinate system are (x c0 ,y c0 ,f), where x c0 = x i0 - ox, y c0 = -(y i0 - oy). The coordinates of the other two vanishing points can be obtained similarly. Using the orthogonal relationship between the vanishing points, the focal length and the rotation matrix are respectively:
[0071]
[0072] Calculate the projection transformation matrix H ij , and the plane determined by the vanishing points vp i and vp j can be "orthographic projected" and corrected. The solution of H ij is as follows:
[0073] H ij = K·R ij ·K -1 ;
[0074] Using the transformation matrix H of the front wall and the floor 01 and H 02 Reproject the 4 intersection points to obtain the coordinates of each point in the rectified front view
[0075]
[0076] S33. Restore the spatial structure ratio to obtain the indoor plane structure.
[0077] Based on the re-projected coordinates, calculate the relative lengths of the length, width, and height respectively, and finally solve the true proportional relationship of the length, width, and height of the indoor space:
[0078]
[0079] Assume that the floor-to-floor height of the indoor room is the average value H, and the length, width, and height extracted from the image are calculated based on the building height H to ensure the scale consistency of the map.
[0080] S4. Integrate and optimize the navigation path and the indoor plane structure.
[0081] S41. Corridor construction: First, the following rules need to be followed during image acquisition: During the turning process, take pictures of the fields of view facing all the inflection points in clockwise order. As Figure 4 shown: When the user passes through a corner, take pictures on the corridor in one direction first, and then rotate to the other direction to take pictures.
[0082] The gyroscope and accelerometer record the angle and acceleration simultaneously. The fluctuations (up or down) of the gyroscope readings indicate that the user turns left or right at this position. During the two fluctuations, the user takes pictures of the corridor, and the corner position where the user is located can be explained from the inertial data. According to the image acquisition rules and the analysis of the shooting behavior, match the corrected trajectory with the corridor plane structure extracted from the mobile phone image to realize the map construction of the corridor area.
[0083] S42. Room construction: The rules for collecting room images are as Figure 5 shown. The user takes pictures from both sides of the room towards the opposite side. Then, by matching the room plane structure extracted from the image with the inertial navigation data, the drawing of the room structure in the indoor map is realized.
[0084] After the trajectory data is matched with the position of the captured image, the map accuracy is further optimized by constructing an energy function. The energy equation to be constructed is as follows:
[0085] min x ={x f ,R,t}E l (x)+E c (x)+E b (x);
[0086] where E l (x) represents the complexity of the layout, E c (x) represents the closeness, E b (x) represents the boundary similarity between adjacent segments, and {x f ,R,t} represents the rotation and displacement when all corridor paths and room layouts are connected.
[0087] The uniqueness of the corridor and room layouts is restricted by the following formula.
[0088]
[0089] Therefore, the present invention adopts the above-mentioned automatic indoor map construction method that fuses multi-source trajectories and mobile phone images, and proposes a high-efficiency, low-cost, and non-dependent indoor map element plane structure extraction method. Using the pictures taken by the mobile phone as the data source, it is not necessary to reconstruct the geometric structure of the indoor three-dimensional model, nor is it necessary to use professional measurement equipment. Only by using the monocular pictures of the indoor scene taken by the smart phone, through the visual scene understanding technology, the acquisition of the indoor map element plane structure is realized, effectively improving the authenticity and accuracy of the indoor map expression. Only by collecting different types of data through the smart phone, fusing multi-source trajectories and standardized mobile phone image acquisition, and using the new geometric feature extraction method, the high-precision construction of the indoor map is realized. Giving full play to the complementary advantages of the overall and local, and the frame and details of the indoor space features, it is not necessary for the user to walk through all positions in the indoor environment, effectively saving the labor cost of indoor map drawing and reducing the time and calculation amount.
[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for automatically constructing indoor maps by integrating multi-source trajectories and mobile phone images, characterized in that: The following steps are involved: S1, using smartphones to collect corridor, indoor images and multi-sensor data; S2, build indoor navigation map; S3, extracting indoor plane structure; S4, integrating and optimizing the navigation path and indoor plane structure; S2 includes the following steps: S21, obtaining multi-source trajectory data by using pedestrian dead reckoning to collect multi-sensor data from the smartphone; S22, obtaining behavior landmarks through a behavior recognition algorithm, and clustering the behavior landmarks through a clustering algorithm based on Wi-Fi fingerprints; S23, calculating the relative distances between the behavioral landmarks to obtain an indoor navigation map; S3 includes the following steps: S31, inputting a monocular image of an indoor space captured by a smartphone, and extracting a three-dimensional layout structure of the monocular image through an indoor three-dimensional scene understanding network model; S32, mapping the indoor 3D scene into the 2D image and performing projection transformation; S33, restore the spatial structure ratio and obtain the indoor plane structure; S4 includes the following steps: S41, corridor construction, matching the corrected trajectory route with the corridor plane structure extracted from the mobile phone image to realize map construction of the corridor area; S42, matching the room plane structure extracted from the image with the inertial navigation data to achieve drawing of the room structure in the indoor map; S32 includes the following steps: S321, using VPBC technology to extract and solve camera parameters based on vanishing point. The camera principal point is located at the center of the image. The camera intrinsic parameter matrix K is in the form of: f is the focal length, which can be found from the vanishing point; S322, set the vanishing point in the image coordinate system to be vp0(x i0 ,y i0 ), vp1(x i1 ,y i1 ), vp2(x i2 ,y i2 ), since the centroid of the triangle formed by the three vanishing points is the intersection point p(ox,oy) of the camera optical axis and the imaging plane, the coordinates of vp0 in the camera coordinate system are (x c0 ,y c0 ,f), where x c0 =x i0 -ox,y c0 =-(y i0 -oy), the coordinates of the other two vanishing points can be obtained in the same way. Using the orthogonal relationship between the vanishing points, the focal length and rotation matrix can be obtained as follows: S323, using camera parameters to calculate the projection transformation matrix H ij , you can use the vanishing point vp i and vp j The determined plane is corrected by "orthographic projection", H ij The solution is as follows: H ij =K·R ij ·K -1 ; Using the transformation matrix H of the front wall and floor 01 and H 02 Reproject the 4 intersection points to get the coordinates of each point in the front view after eliminating distortion 2. The method for automatically constructing an indoor map by integrating multi-source trajectories and mobile phone images according to claim 1, characterized in that: The actual proportions of the length, width, and height of the interior space are:
3. The method for automatically constructing an indoor map by integrating multi-source trajectories and mobile phone images according to claim 2, characterized in that: By constructing an energy function to further optimize and ensure map accuracy, the energy equation is proposed to be constructed as follows: min x ={x f ,R,t}E l (x)+E c (x)+E b (x); Where E l (x) represents the complexity of the layout, E c (x) represents the degree of closure, E b (x) represents the boundary similarity between adjacent segments, {x f ,R,t} represents the rotation and displacement of all corridor paths and room layouts when they are connected.
4. The method for automatically constructing an indoor map by integrating multi-source trajectories and mobile phone images according to claim 3, characterized in that: The conditions that restrict the uniqueness of corridor and room layouts are:
5. The method for automatically constructing an indoor map by integrating multi-source trajectories and mobile phone images according to claim 4, characterized in that: The following rules are followed during image acquisition: when shooting indoor corridors, you need to take pictures of all the view areas facing the turning points in a clockwise order during the turning process; when shooting indoor rooms, you need to shoot from both sides of the room towards the opposite side.
Citation Information
Patent Citations
Three-dimensional model dynamic updating method based on visual perception
CN114494582A
Pedestrian indoor positioning and AR navigation method based on computer vision and PDR
CN114739410A