A trajectory correction method based on an indoor location network model
By constructing an indoor location network model and combining it with map matching methods, the problem of insufficient indoor positioning accuracy was solved, the deviation of indoor positioning trajectory data was corrected, the data accuracy was improved, and reliable positioning data support was provided for subsequent analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Chinese People's Liberation Army Cyberspace Force Information Engineering University
- Filing Date
- 2021-09-02
- Publication Date
- 2026-07-17
AI Technical Summary
Existing indoor positioning technologies lack accuracy in complex environments, making it difficult to support the analysis and application of personnel activities. In particular, the accuracy of WiFi positioning is affected by signal strength fluctuations and indoor channel environment, and there is a lack of effective solutions for WiFi data correction and accuracy analysis with map assistance.
The trajectory correction method based on the indoor location network model constructs an indoor location network model and combines it with map matching methods to match indoor positioning trajectory data onto the indoor location network model. It uses Delaunay triangulation and Voronoi diagram to generate road network skeleton lines and employs a hidden Markov model for map matching to improve the accuracy of positioning data.
It improves the location accuracy of indoor positioning trajectory data, providing more reliable positioning data for indoor positioning data analysis and application, and supports subsequent hotspot detection and personnel behavior analysis.
Smart Images

Figure CN115752459B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of indoor positioning technology, specifically relating to a trajectory correction method based on an indoor location network model. Background Technology
[0002] Human activity (gathering locations, movement routes, etc.) is an important type of geospatial intelligence. With the development of GPS technology, outdoor human activity location information can be accurately acquired and used to generate activity intelligence. However, in indoor environments, tunnels, underground fortifications, and other locations, GPS signal attenuation makes accurate positioning impossible, thus failing to provide effective human location information. In these environments with obstructed GNSS signals, human activity locations are often acquired using technologies such as WiFi, Bluetooth, and UWB. However, these location information has low accuracy and is difficult to support the mining and analysis of activity intelligence in complex indoor environments.
[0003] Understanding the activities of people indoors, such as their distribution, density, and movement patterns, is crucial for public disaster prevention, optimization of public facilities, and commercial services, thus possessing significant civilian value. Statistics show that nearly 80% of modern life is conducted indoors. Indoor Location-Based Services (LBS) have already played a vital role in many specific indoor scenarios, such as warehouse logistics, disaster relief, and store recommendations, and are gradually becoming the foundation for the development of the Internet, the Internet of Things, artificial intelligence applications, smart logistics, and smart cities.
[0004] As is well known, among various indoor positioning methods and systems, WiFi fingerprint positioning technology has significant advantages in terms of hardware cost, real-time performance, acquisition rate, positioning coverage, and scalability. Furthermore, the high penetration rate of smart mobile devices allows for easy access to WiFi signals in public places such as schools, shopping malls, and train stations, enabling positioning needs to be met by utilizing existing public resources. However, indoor environments suffer from severe signal multipath propagation and reflection / refraction phenomena, and WiFi positioning technology is easily affected by signal strength fluctuations, often resulting in insufficient WiFi positioning accuracy. This makes it difficult to support the analysis and application of indoor human activity. Therefore, improving WiFi positioning accuracy has significant practical application implications.
[0005] Similar to outdoor positioning based on satellite positioning and communication base station positioning, indoor positioning also primarily employs wireless positioning technology, with wireless positioning systems based on signal sources such as WiFi, Bluetooth, and UWB widely used in practical applications. These wireless positioning systems estimate the spatial location of a target by estimating the signal travel time between the target and the signal source, or by matching and identifying the signal characteristics of the target's location, and record information such as the target's device identification and positioning time, providing an important data source for studying individual spatial behavior. However, the complex indoor channel environment poses significant challenges to indoor wireless positioning. Wireless positioning accuracy often depends on signal coverage and is also affected by device status (such as whether a mobile phone is powered on or in standby mode), crowd density, building structure, and other electromagnetic interference factors, making it difficult to achieve the accuracy measured under ideal conditions in practical applications. Therefore, improving the accuracy of positioning data and providing accuracy analysis is fundamental to downstream data analysis tasks.
[0006] In numerous studies on improving the accuracy of indoor positioning data, effectively utilizing maps as spatial constraints to enhance the performance of (positioning or correction) models has always been a crucial direction for researchers seeking improvement. These map constraint methods can be categorized into three types based on how they are used: ① Bayesian estimation. This method uses the geometric and topological information of indoor maps to reduce candidate locations for position estimation, transmitting high-confidence position estimates in WiFi tracks, and improving WiFi track position accuracy by solving for the optimal estimate. This method effectively handles non-Gaussian noise in WiFi positioning data, but suffers from reliance on built-in sensors and poor robustness. ② Ray tracing. This method addresses signal reflection and penetration phenomena caused by indoor spatial layouts, calculating non-line-of-sight signal attenuation based on map geometric contour information to correct distance estimates based on signal strength. However, in practical applications, signal strength is influenced by many factors and has high uncertainty, making it difficult to establish an accurate signal attenuation model. ③ Map matching. This method abstracts the indoor movable range into a topological network of discrete locations, using nearest neighbor and topological analysis methods to map the WiFi track to the sequence of most similar location nodes in the network. This method reduces the accuracy of indoor position representation, retaining only accessibility relationships for indoor navigation applications. For location correction of large-scale historical WiFi positioning data, map matching is more suitable for the application scenario of "post-processing" historical data.
[0007] In summary, while there is a wealth of research on the analysis and processing of spatiotemporal information such as indoor and outdoor positioning data, there is relatively little work focusing on the accuracy assessment and correction of map-assisted indoor WiFi positioning data. In conclusion, a complete solution for indoor map-assisted WiFi data correction and accuracy analysis currently exists. Summary of the Invention
[0008] This invention provides a trajectory correction method based on an indoor location network model to solve the problem of low accuracy in indoor positioning trajectories.
[0009] To solve the above-mentioned technical problems, the technical solutions included in this invention and their corresponding beneficial effects are as follows:
[0010] The present invention provides a trajectory correction method based on an indoor location network model, comprising the following steps: 1) acquiring indoor positioning trajectory data; 2) constructing an indoor location network model using indoor floor plan data; 3) using a map matching method to match the trajectory positioning points in the indoor positioning trajectory data to the indoor location network model, so as to achieve correction of the indoor positioning trajectory data.
[0011] The beneficial effects of the above technical solution are as follows: This invention addresses the post-processing application needs of massive indoor positioning trajectory data. It extracts structural information of the indoor environment based on indoor floor plan data to construct an indoor location network model. Then, the indoor location network model is used as a spatial constraint, combined with map matching methods, to effectively correct the indoor positioning trajectory data, improve the positional accuracy of the indoor positioning trajectory data, and provide more reliable positioning data for indoor positioning data analysis and application.
[0012] Furthermore, in order to construct an accurate indoor location network model, the method for constructing the indoor location network model in step 2) includes: 2.1) Extracting the boundary range of indoor buildings and extracting the nodes of indoor buildings that affect the spatial structure; 2.2) Constructing a Delaunay triangulation network using the Delaunay triangulation method based on the extracted nodes; 2.3) Classifying the triangular primitives according to the relationship between the triangular primitives in the Delaunay triangulation network and the neighboring triangles, and extracting the Voronoi road network skeleton lines of indoor buildings according to the classification results and the corresponding rules; 2.4) Constructing the indoor location network model using the extracted Voronoi road network skeleton lines.
[0013] Furthermore, in order to satisfy the Delaunay triangle network conditions to correctly represent the proximity relationships between polygonal buildings, the method also includes a node encryption process for the extracted nodes, and the node encryption process includes the following steps:
[0014] Extract the outline of the interior building, and then extract the nodes on the outline; for two nodes {P} i} and {P i+1}, when the length between two nodes satisfies |P i P i+1 |>W sets an encryption point between the two nodes, and the encryption point {Q k}satisfy:
[0015]
[0016] Among them, (X) i Y i ) is node {P i The coordinates of}; (X i+1 Y i+1 ) is node {P i+1 The coordinates of}; (X k Y k ) is the encryption point {P k The coordinates of}; the coefficient λ k for W sets the width value.
[0017] Furthermore, it also includes a step of smoothing the extracted Voronoi road network skeleton lines, using the Douglas-Peucker algorithm; the threshold of the Douglas-Peucker algorithm is selected as 1 / 2 encryption step size, and the encryption step size is the encryption step size of the node encryption process.
[0018] Furthermore, after step 2.3), the method also includes a step of densifying the open areas in the Voronoi road network. The method of densifying the road network includes the following steps: establishing a secondary buffer for the Voronoi road network skeleton line with a set sampling threshold as a parameter, and performing uniform point sampling at intervals of the set sampling threshold; and using the samples after trimming the sampling points that are not within the scope of the study as road network densification points for road network densification.
[0019] Accordingly, the method for constructing the indoor location network model in step 2.4) is as follows: establish full connections for all encrypted points and establish connections with the nearest main network node to form the edges of the indoor location network model, so as to construct the indoor location network model.
[0020] Furthermore, the map matching method used in step 3) is a map matching method based on a hidden Markov model.
[0021] Furthermore, when constructing the Hidden Markov Model, it is necessary to determine the observation probability and state transition probability in the Hidden Markov Model; a normal distribution is used to describe the observation probability, and the hidden state at time t is r. i The corresponding observed value is o t The conditional probability is:
[0022]
[0023] Where, σ t Let σ be the standard deviation of the location data. t =1.4826median(dt );d t The distance from the location point at time t to the candidate road segment r i The distance, using ||o t -x t,i || greatcircle That is, the observed value o t To the corresponding candidate road segment r i Matching point x on t,i The distance to the Earth's surface, take o. t In r i The vertical projection point on the x-axis is used as the matching point. t,i And matching point x t,i The calculation method is as follows:
[0024]
[0025] Where (x,y) is the matching point x t,i The coordinates of P(x) p ,y p ) and Q(x q ,y q ) are the observed values o t (x o ,y o Corresponding candidate road segment r i The two endpoints; k is the candidate road segment r i The slope, that is:
[0026] The state transition probability follows an exponential distribution:
[0027]
[0028] Among them, dist t For two adjacent positioning points (o) t ,o t+1 The distance on the Earth's surface and the matching points on the candidate road segments corresponding to these two positioning points. The difference in the actual path distance between them, that is:
[0029]
[0030] And the method for estimating β is as follows:
[0031]
[0032] Among them, |||| greatcircle Indicates the distance between two points on a great circle; || || route This represents the actual path distance between two points.
[0033] Furthermore, step 1) also includes a coordinate transformation step for the indoor positioning trajectory data, wherein the coordinate transformation formula is:
[0034]
[0035] Where (X,Y) are the coordinates before the coordinate transformation, and (X′,Y′) are the coordinates after the coordinate transformation; [a,b,d,e,x off ,y off [ ] represents the affine transformation parameters, and a, b are the deformation parameters in the x-direction, x off Let d be the translation in the x-direction, and e be the deformation parameters in the y-direction. off This represents the translation in the y-direction.
[0036] Furthermore, in order to filter out noisy data, the method also includes a data cleaning step for the indoor positioning trajectory data after coordinate transformation. The data cleaning method includes the following steps: calculating the actual distance between each positioning point based on the coordinates after coordinate transformation; calculating the time difference between each positioning point based on the time label of the positioning point; calculating the average moving speed between each positioning point using the actual distance between each positioning point and the time difference between each positioning point; and filtering out positioning data with an average moving speed below a set speed value to achieve data cleaning.
[0037] Furthermore, after step 3), there is also a step of evaluating the result of the correction, and the evaluation index used for evaluation includes at least one of the following indices: dynamic time normalization, longest common substring, Hausdorff distance, Fraser distance, edit distance, and symmetric trajectory segment distance. Attached Figure Description
[0038] Figure 1 This is a visualization of the WiFi trajectory data of 50 users between 13:00 and 13:30 in an example of the method of this invention;
[0039] Figure 2 This is a diagram showing the movement trajectory of a user on the first floor from 13:00 to 13:30 in an embodiment of the method of the present invention.
[0040] Figure 3-1 This is a frequency distribution histogram of indoor crowd movement in an embodiment of the method of the present invention;
[0041] Figure 3-2 This is a Gaussian kernel density map of indoor crowd movement in an embodiment of the method of the present invention;
[0042] Figure 4-1 This is a trajectory diagram of a user on the first floor before data cleaning in an embodiment of the method of the present invention;
[0043] Figure 4-2 This is a trajectory diagram of a user on the first floor after data cleaning in an embodiment of the method of the present invention;
[0044] Figure 5 This is a schematic diagram of the indoor partial Voronoi diagram road network of the present invention;
[0045] Figure 6 This is a flowchart of the method for constructing an indoor location network model according to the present invention;
[0046] Figure 7 This is a diagram showing the boundary range of an indoor building in an embodiment of the method of the present invention;
[0047] Figure 8 This is a direct mesh diagram of the indoor building outline nodes in the embodiment of the method of the present invention;
[0048] Figure 9-1 This is a schematic diagram of the indoor building outline nodes before encryption in an embodiment of the method of the present invention;
[0049] Figure 9-2 This is a schematic diagram of the encrypted outline nodes of the indoor building in the embodiment of the method of the present invention;
[0050] Figure 10 This is a mesh diagram of the indoor building outline nodes after densification in the method embodiment of the present invention;
[0051] Figure 11-1 This is a schematic diagram of the generation rules of type I triangular primitives and skeleton lines in the embodiment of the method of the present invention;
[0052] Figure 11-2 This is a schematic diagram of the generation rules of type II triangular primitives and skeleton lines in the embodiment of the method of the present invention;
[0053] Figure 11-3 This is a schematic diagram of the generation rules of type III triangular primitives and skeleton lines in the embodiment of the method of the present invention;
[0054] Figure 12 This is a schematic diagram of the preliminary results of the fusion of indoor skeleton lines and channel information in an embodiment of the method of the present invention;
[0055] Figure 13 This is a diagram illustrating the effect of straightening the indoor skeleton lines in an embodiment of the method of the present invention.
[0056] Figure 14 This is a schematic diagram representing the spatial range of the indoor road network skeleton line in an embodiment of the method of the present invention;
[0057] Figure 15 This is a sampling result diagram of the uniform points in the secondary buffer zone of the indoor road network skeleton line in an embodiment of the method of the present invention;
[0058] Figure 16This is a schematic diagram of all nodes included in the graph model of the indoor location network in the embodiment of the method of the present invention;
[0059] Figure 17 This is a schematic diagram of the graph model of the indoor location network formed in the embodiment of the method of the present invention;
[0060] Figure 18 This is a schematic diagram of a Markov chain;
[0061] Figure 19 This is a graph showing the relationships between the parameters of a Hidden Markov Model;
[0062] Figure 20 This is a flowchart of the map matching method based on the Hidden Markov Model in the embodiment of the method of the present invention;
[0063] Figure 21 This is a schematic diagram of the candidate road segment selection process in an embodiment of the method of the present invention;
[0064] Figure 22 This is a schematic diagram illustrating the process of generating a WiFi location point candidate path set in an embodiment of the method of the present invention;
[0065] Figure 23 This is a schematic diagram of the WiFi positioning data map matching model construction in an embodiment of the method of the present invention;
[0066] Figure 24-1 This is a diagram showing the indoor original WiFi positioning test trajectory in an embodiment of the method of the present invention;
[0067] Figure 24-2 This is a diagram showing the WiFi positioning trajectory after speed filtering preprocessing in an embodiment of the method of the present invention;
[0068] Figure 25-1 A schematic diagram of map matching results when the observation prediction factor = 0 in an embodiment of the method of the present invention;
[0069] Figure 25-2 A schematic diagram of map matching results when the observation prediction factor = 0.4 in an embodiment of the method of the present invention;
[0070] Figure 25-3 A schematic diagram of map matching results when the observation prediction factor = 0.6 in an embodiment of the method of the present invention;
[0071] Figure 25-4 A schematic diagram of map matching results when the observation prediction factor = 0.8 in an embodiment of the method of the present invention;
[0072] Figure 26 This is a line graph showing the map matching results and the similarity evaluation values of the indoor WiFi positioning test trajectory. Detailed Implementation
[0073] This invention addresses the post-processing application needs of massive indoor WiFi positioning trajectory data (hereinafter referred to as WiFi positioning data). It studies methods for WiFi trajectory correction and accuracy analysis, automatically extracting spatial structure information of the indoor environment based on indoor floor plans. This introduces spatial constraints for indoor WiFi trajectory correction, improving the location accuracy of WiFi positioning data and providing more reliable positioning data for downstream tasks such as trajectory-based hotspot detection and human behavior analysis. The following detailed description, along with examples, illustrates a trajectory correction method based on an indoor location network model.
[0074] Method Implementation Examples:
[0075] This embodiment uses WiFi location data obtained from a large shopping mall as an example (the half-hour trajectory data of 50 randomly selected users on the first floor of the mall is as follows). Figure 1 (As shown), the method of the present invention is described.
[0076] Step 1: Obtain WiFi location trajectory data and preprocess the WiFi location data.
[0077] In practical applications, WiFi positioning systems can acquire massive amounts of WiFi positioning data through passive acquisition. However, this data contains random location noise due to differences in the model of the acquiring device, the device's operating status, and the channel environment. Therefore, directly using this large-scale, high-noise data can lead to problems such as reduced data analysis reliability and high computational load. Preprocessing methods such as data extraction, coordinate transformation, and data cleaning are required based on the specific characteristics of indoor WiFi positioning trajectory data.
[0078] 1. Indoor WiFi positioning trajectory data extraction.
[0079] Based on the characteristics of indoor WiFi positioning trajectory data, it is usually necessary to extract and classify the data according to attributes such as space (floor), time, and user identifier (device MAC address) to construct the sample dataset required for subsequent analysis. The attribute table of WiFi positioning data is shown in Table 1. Wherein, FID is the data identifier, Rec_Time is the positioning time, Device_Mac is the MAC address (user identifier) of the positioning device, Building is the building number, Floor is the floor number, X is the x-coordinate of the positioning point in the relative coordinate system, Y is the y-coordinate of the positioning point in the relative coordinate system, and Received_IP is the IP address of the WiFi positioning data acquisition device.
[0080] Table 1. Attributes of the Original WiFi Location Dataset (Partial)
[0081]
[0082] The user's location sequence within a specific time period, i.e., the user trajectory, can be represented as:
[0083] Trajectory(Device_Mac)=series(Rec_Time,X(Rec_Time),Y(Rec_Time)) (1)
[0084] The original dataset was extracted according to building number and floor. Then, within each floor's data, it was grouped by time period (every half hour) to effectively organize the dataset. Next, a set of data was selected as a sample for experimental research. Based on the user identifier (Device_Mac) and the time label of the location data, the trajectory was initially reconstructed using the MovingPandas open-source framework. Figure 2 The image shows the movement trajectory of a user on the first floor from 13:00 to 13:30. The dotted symbols represent the location points of different users, and the line symbols represent the user's trajectory line. It is evident that the trajectory contains obvious "jump points" and instances of passing through walls, which does not conform to objective reality. The location data quality is poor, and gross error filtering is necessary before data analysis.
[0085] 2. WiFi positioning trajectory data coordinate transformation and data cleaning.
[0086] Since indoor WiFi positioning systems often use local custom coordinate systems (as shown in Table 1), the positioning information in the data differs from the actual geographic mapping. Therefore, coordinate transformation is necessary first to facilitate the calculation of physical properties such as speed, supporting further analysis and processing. The origin of the local coordinate system is established at the upper left corner of the building. Based on the data, the relative coordinates of the lower right corner of the building can be determined. Then, based on the building's actual geographic coordinates (latitude and longitude), an affine relationship for the transformation between the two sets of actual geographic coordinates and relative positioning coordinates is established using the correspondence between them, and the affine transformation parameters are determined:
[0087] [a,b,d,e,x off ,y off (2)
[0088] Where a and b are the deformation parameters in the x-direction, x off d represents the translation in the x-direction; d and e represent the deformation parameters in the y-direction. off This represents the translation in the y-direction.
[0089] Therefore, the formula for transforming positioning data coordinates is:
[0090]
[0091] Right now:
[0092]
[0093] To address the issue of "jump points"—recording errors—in WiFi location data, this embodiment employs a gross error filtering method. Since physical attributes such as the speed of human movement indoors limit these processes, the original location data can be cleaned based on general indoor movement patterns and data distribution.
[0094] After coordinate transformation, the actual distance between each location point can be calculated based on its latitude and longitude coordinates (since the indoor space is relatively small and the influence of factors such as the curvature of the Earth is negligible, the Euclidean distance is used to approximate the great circle distance). Then, the time difference between each location point is calculated based on its time stamp. This allows for the calculation of the average speed of movement between each location point. In this embodiment, it is assumed that the movement speed of people indoors generally does not exceed 2 m / s, and this is used as a speed threshold, combined with the speed distribution (e.g., ...). Figure 3-1 The image shows a frequency distribution histogram. Figure 3-2 The Gaussian kernel density plot (with the dashed line representing the speed filtering threshold) filters out location data with an average moving speed below 1.8 m / s. This retains approximately 85% of the location data while conforming to the general movement patterns of people indoors, effectively filtering out gross errors in the data.
[0095] After initial data cleaning, obvious erroneous "jump points" in the positioning data can be filtered out. For example... Figure 4-1 Figures 4-2 show a comparison of the trajectory of a user on the first floor before and after data cleaning. Figure 4-1 For the original trajectory, Figure 4-2 The image shows the trajectory after gross error filtering. As can be seen, the gross error filtering effectively reduces some "jump points" (shown by the dotted circles in the image), which helps analyze user movement and facilitates the determination of behavioral patterns such as dwell points, laying the foundation for further data correction.
[0096] Step 2: Construct an indoor location network model using indoor floor plan data.
[0097] In this embodiment, a Voronoi diagram (also known as a Thiessen polygon) is used to describe the indoor location network model. Compared with other methods, the Voronoi diagram can maximize the gaps between "obstacles," allowing the road network to be arranged at equal distances from each surrounding geographic unit. Therefore, the Voronoi diagram has applications in optimal route planning problems in various fields such as indoor navigation and robot automatic pathfinding. The indoor local Voronoi diagram road network generation effect is as follows: Figure 5As shown in the figure, the dark dots in the structural outline represent indoor obstacles, the sparse light dots within the outline represent reference points, and the dense light lines within the outline represent the Voronoi diagram generated by the final fitting.
[0098] In this embodiment, when constructing an indoor location network model based on a Voronoi diagram, the Delaunay triangulation constraint is first established, and then its dual graph—the Voronoi diagram—is extracted from it to generate the skeleton lines of the indoor road network. Next, road network densification is performed on some "open areas," and combined with the passage information of indoor building structures, finally fusion is used to form an indoor location network model, laying the foundation for subsequent map matching and other related work. The entire process is as follows: Figure 6 As shown below. A detailed explanation follows.
[0099] 1. Interior floor plan data processing.
[0100] To construct an indoor location network model using interior floor plan data (including room coordinates, geometry, and usage attributes), data preprocessing is necessary. Before extracting the skeleton lines using the Delaunay triangulation method, the boundary range of building elements needs to be extracted, which can be described using the minimum bounding rectangle or minimum convex hull. Furthermore, node densification of the buildings is required to meet the network construction requirements of the Delaunay triangulation.
[0101] 1.1) Extraction of the boundary range of the indoor building structure.
[0102] To facilitate the generation of Delaunay triangulation and the clipping of triangulation outside the study area, the boundary of the indoor buildings needs to be determined first. The boundary of the indoor buildings can be described using the minimum bounding rectangle or minimum convex hull method, such as... Figure 7 As shown. This embodiment uses the minimum convex hull to describe the boundary of the interior building structure, and uses this to trim the constructed Delaunay triangulation.
[0103] 1.2) Perform node densification processing on the outline nodes of indoor buildings.
[0104] After the boundary of the interior building is constructed, it is necessary to use the building element outline and boundary as constraint edges, and extract the key points that affect the spatial structure to construct a constrained Delaunay triangulation. Methods for automatically constructing Delaunay triangulations based on nodes include divide-and-conquer algorithms, point-by-point insertion methods, or triangulation growth methods. For planar buildings such as interiors, the number of nodes in the outline is often too small to satisfy the Delaunay triangle construction conditions. Directly constructing the triangulation often violates the nearest neighbor property of Delaunay, resulting in numerous elongated triangles, making the Delaunay mesh structure unable to accurately represent the proximity relationships between polygonal buildings. Figure 8 The diagram shows the case where a Delaunay triangulation is directly constructed on the first floor of an indoor building in the area studied in this embodiment. Therefore, before constructing the mesh, it is necessary to densify the nodes of the building outline to allow more closely spaced points to participate in the mesh construction. The method for densifying the indoor building outline nodes is as follows:
[0105] Extract the outline of the interior building, and then extract the nodes {P} of the outline. i}. When the length |P i P i+1 |>W (where W is the average width of the passageway or corridor between interior buildings, and its value can be determined according to the application), then for the encryption point {Q k},satisfy:
[0106]
[0107] in:
[0108]
[0109] like Figure 9-1 , 9-2 As shown, the sparse points in the left image are joints extracted from the outline of the interior building, while the dense points in the right image are the interior outline points after densification. This can be used to construct a Delaunay triangulation.
[0110] like Figure 10 The image shows a Delaunay triangulation generated after the boundary points of the building contour have been encrypted, and then clipped using the minimum convex hull of the building. Its effect is better than a triangulation generated directly from the building contour joints. Figure 8 It offers significant improvements, making it easier to trace and generate Voronoi road network skeleton lines (referred to as skeleton lines); and it avoids being too dense, keeping the computational load within a certain range, which is convenient for implementation.
[0111] 2. Generation of skeleton lines under Delaunay triangle constraints.
[0112] In step 1, this embodiment has already used the Delaunay triangulation algorithm to subdivide the indoor architectural space structure and encrypted the triangulation according to the actual spatial conditions. Here, its basic unit is called a triangular primitive. In order to facilitate the tracking and generation of skeleton lines, it is necessary to first classify the triangular primitives according to the spatial characteristics, then generate skeleton lines, and then smooth the skeleton lines to finally form a node-edge graph structure, which facilitates subsequent map matching and other related work.
[0113] 2.1) Classification of triangular primitives.
[0114] For the triangular primitives in the Delaunay triangulation, we examine their relationships with neighboring triangles and classify them accordingly. Class I primitives have a neighboring triangle on only one side, Class II primitives have neighboring triangles on two sides, and Class III primitives have neighboring triangles on all three sides. Observing Figure 11, we can see that Class I triangular primitives mostly appear at corridor entrances and exits, Class III triangular primitives mostly appear at corridor intersections or in "open areas," and Class II triangular primitives appear in all other locations, mostly in the middle of corridors and the four corners of rooms.
[0115] After classifying the triangular primitives, this embodiment extracts the original skeleton lines of the interior building structure according to the following rules:
[0116] ①Type I triangular primitives connect the midpoint of the unique common edge with its opposite vertex to form a skeleton line segment, such as... Figure 11-1 As shown;
[0117] ②Type II triangular primitives connect the midpoints of two common sides, such as... Figure 11-2 As shown;
[0118] ③Type III triangular primitives connect the centroid to the midpoint of the three common sides, such as... Figure 11-3 As shown.
[0119] 2.2) Smoothing of skeleton lines and fusion of channel information.
[0120] The previous step extracted the building's structural framework on the first floor and manually added information such as doors and passageways. For example... Figure 12 The image shows the skeleton lines of the interior building structure extracted semi-automatically, with long straight lines used to represent doorways.
[0121] Meanwhile, it was noted that the local skeleton lines originating from type I triangles were not smooth and required straightening; the skeleton lines connecting type II triangles in some corridor passages were jagged and needed smoothing. Therefore, this embodiment uses the Douglas-Peucker algorithm (DP algorithm) to smooth and straighten the extracted skeleton lines. Since the fundamental reason for the insufficient smoothness of the skeleton lines is that the encryption step size of the building outline densification points is not small enough, the parameters of the DP algorithm should be related to the encryption step size. Experimental testing showed that when the DP algorithm threshold was set to 1 / 4 encryption step size, the smoothing effect was not obvious; when the DP algorithm threshold was set to 1 / 2 encryption step size, the smoothing effect was relatively good; when the DP algorithm threshold was set to the encryption step size, too many geometric features of the skeleton lines were lost. Therefore, this embodiment uses 1 / 2 encryption step size as the threshold parameter for the DP algorithm. The straightening results in some areas are as follows: Figure 13 As shown in the figure. The curved segments are the original skeleton lines, and the straight segments are the skeleton lines that have been straightened.
[0122] 3. Indoor open area road network densification (i.e., road network densification processing) and output.
[0123] This embodiment uses a linear network structure to describe the planar structure of the indoor passable area, which inevitably leads to coverage issues. Since indoor WiFi positioning uses a fingerprinting method, the positioning points themselves exhibit a regular grid distribution. Therefore, considering both indoor pedestrian movement characteristics and WiFi positioning data characteristics, and combined with practical application needs, it is considered that the skeleton linear network can represent a spatial range of 2 meters around it. Therefore, the indoor spatial range that the road network structure obtained in step 2 can represent is as follows: Figure 14 As shown.
[0124] It can be observed that because the road network is represented using the framework lines of indoor building structures, many "open areas" are not fully represented. Therefore, it is necessary to densify the road network in some open areas. In this embodiment, a secondary buffer zone is established for the road network framework lines using a 2m parameter, and uniform point sampling is performed on it at 2m intervals. Figure 15 As shown. The outer layer of the skeleton lines consists of uniformly sampled points in a first-level buffer zone, and the next outermost layer consists of uniformly sampled points in a second-level buffer zone. Then, sampling points outside the scope of the study are trimmed off, and these are used as the road network densification points. Simultaneously, to facilitate the subsequent establishment of the indoor location network graphical model, the main road network composed of the skeleton lines needs to be resampled. To reflect the geometric importance of the main road network and its impact on subsequent road network matching, this embodiment uses a sampling interval of 0.25m. Figure 15 The skeleton lines, the uniform sampling points of the first-level buffer, the uniform sampling points of the second-level buffer, and all nodes of the graph model of the indoor location network, such as... Figure 16As shown. Finally, a full connection is established for the encrypted point, and a connection is established with the nearest main network node to form the edges of the graph model of the indoor location network, as shown. Figure 17 As shown in the diagram, the darker color represents the main network, and the lighter color represents the encrypted network. Finally, the indoor location network model needs to be topologically reconstructed and output as a graph structure of (nodes, edges) to facilitate subsequent map matching and other related tasks.
[0125] Step 3: Correction of WiFi indoor positioning trajectory based on map matching method.
[0126] Building upon the indoor location network model constructed in step two, map matching is performed to correct WiFi location points onto the location network model, avoiding illogical situations such as points penetrating walls, thereby improving the accuracy and reliability of the location data. Of course, even if the map's accuracy doesn't significantly surpass the accuracy of the location results, the process of mapping location measurements to the map's road network often transfers map information or knowledge to the location results, supporting information mining in other areas.
[0127] To improve the matching accuracy of low-frequency trajectories and suit the characteristics of indoor WiFi positioning data, this embodiment employs a map matching algorithm based on a Hidden Markov Model (HMM). By constructing an HMM map matching model for indoor WiFi positioning data and solving the model using the Viterbi algorithm, parameter estimation is performed based on sample data containing actual movement trajectories, thus completing the correction of indoor WiFi positioning data.
[0128] 1. Map matching method based on Hidden Markov Model.
[0129] A Hidden Markov Model (HMM) is a Bayesian network that uses a directed graph representation to describe a Markov process containing a random sequence of unobservable (hidden) states (such as...). Figure 18 (where X is the hidden state sequence and Y is the observation sequence). It is mainly used for time series data modeling and has wide applications in speech recognition, natural language processing, pattern recognition, and bioinformatics. The elements and symbol conventions of the Hidden Markov Model are explained below.
[0130] ① Let X be the set of all possible hidden states, denoted as X = {x1, x2, x3, ..., x...} n Let}, where n is the number of possible states. These states satisfy the Markov property and are the actual implicit states in the Markov model, which are usually not obtainable through direct observation. In the map matching problem, X represents the set of ground truth points on the map corresponding to the location points (set).
[0131] ② Let Y be the set of all possible observed states, denoted as Y = {y1, y2, y3, ..., y...} m}, where m is the number of possible observations. These are associated with hidden states in the model and can be obtained through direct observation. The number of observable states does not necessarily have to match the number of hidden states.
[0132] ③ Let the length of the hidden Markov chain be T. If S represents the state sequence and O represents the observation sequence, then:
[0133] S=(s1,s2,s3,…,s T ), O=(o1,o2,o3,…,o T (7)
[0134] ④ Let Π be the initial state matrix, representing the probability matrix of the hidden state at the initial time t=1, i.e., Π=(Π i ), where Π i =P(s1=x) i ).
[0135] ⑤ Let the hidden state transition probability matrix be A = [a ij ] n*n ,in:
[0136] a ij =P(i t+1 =x j |s t =x i ), i=1,2,…,n; j=1,2,…,n (8)
[0137] Equation (8) describes the transition probabilities between states in a hidden Markov model, where x represents the state at time t. i Under the given conditions, the state at time t+1 is x. j The probability of this is known. It's not hard to see that the sum of each row of matrix A is 1.
[0138] ⑥ Let the observation state probability matrix (representing the probability of a certain observation occurring in a certain state, also known as the emission matrix) be B = [b j (k)] n*m ,
[0139] in:
[0140] b j (k)=P(o t =y k |s t =x j ),k=1,2,…,m; j=1,2,…,n (9)
[0141] Equation (9) indicates that at time t, the hidden state is x. j Under the condition that the observed state is y k The probability of.
[0142] As shown above, a Hidden Markov Model (HMM) is determined by the state transition matrix A, the initial probability vector Π, and the observation probability matrix B. A and Π determine the state sequence, while B determines the observation sequence. These three are collectively known as the three essential elements of an HMM. The initial state probability vector Π and the state transition probability matrix A determine the hidden Markov chain, generating the unobservable state sequence. The observation probability matrix B determines how the observation sequence is generated from the state sequence, and together with the state sequence, it determines how the observation sequence is produced. The relationships between the parameters in the HMM are as follows: Figure 19 As shown.
[0143] At the same time, attention should be paid to two fundamental assumptions in Hidden Markov Models—the homogeneity assumption and the observation independence assumption:
[0144] ① The homogeneity assumption, which states that the state at any time t depends only on the previous state and is independent of other state values and observations, is a manifestation of the aftereffect-free property of all Markov processes. That is:
[0145] (i t |i t-1 ,o t-1 ,…,i1,o1)=P(i t |i t-1 )t=1,2,…,T (10)
[0146] ② The observation independence assumption states that an observation at any given time depends only on the current state value and is independent of the state values and observations at any other time. That is:
[0147] P(o t |i T ,o T i T-1 ,o T-1 ,…,i t+1 ,o t+1 i t i t-1 ,o t -1,…,i1,o1)=P(o t |i t (11)
[0148] Using the two fundamental assumptions of Hidden Markov Models (HMMs) as constraints makes problem solving easier. Currently, HMMs can mainly solve three types of problems: (a) Observation probability calculation problems: given a model λ=(A,B,Π) and an observation sequence O=(o1,o2,o3,…,o…), ... T(a) Calculate the probability P = (O|λ) of the observed sequence under model λ; (b) Model parameter learning problem: Given the observed sequence O = (o1, o2, o3, ..., o T (c) Prediction problem (also known as decoding problem): Given a model λ = (A,B,Π) and an observation sequence O = (o1,o2,o3,…,o…), estimate the parameters of the model λ = (A,B,Π) such that the probability of the observed sequence P = (O|λ) is maximized under this model, i.e., estimate the parameters using the maximum likelihood estimation method; T Given an observation sequence, find the state sequence I = (i1, i2, i3, ..., i...) that maximizes the conditional probability. T ).
[0149] As can be seen, map matching is a prediction problem. That is, given the state transition matrix A and observation probability matrix B of the Hidden Markov Model, the initial state matrix Π, and a set of observation sequences (i.e., the WiFi positioning trajectory to be corrected in this embodiment) O=(o1,o2,o3,…,o…), the prediction problem is solved by the Hidden Markov Model. T The process involves finding the state sequence with the highest probability under these conditions (the corresponding truth sequence of WiFi location data on the map in this embodiment). Therefore, the map matching method for indoor WiFi location point data in this embodiment follows the steps outlined below. Figure 20 As shown.
[0150] 2. Construction of Hidden Markov Models.
[0151] As introduced in Step 1, the map matching method based on the Hidden Markov Model (HMM) proceeds in three steps. The core of model construction lies in determining the state transition matrix A and the observation probability matrix B in λ = (A, B, Π). Prior to this, candidate road segments can be selected in the indoor location network model using some prior knowledge. For example, by setting the road network within a certain threshold range around the indoor WiFi location point as candidate road segments, the computational data can be kept from being too large, thus reducing the algorithm's runtime. The error region set in this paper is a circular area with a radius of 15m, which improves algorithm efficiency while preventing the location point data from only matching the nearest road segment. The candidate road segment selection process is as follows: Figure 21 As shown.
[0152] Therefore, we can obtain the candidate road segment set for all WiFi location points, that is, the set of all possible hidden states, as shown in Figure 22. Hollow circles represent road segments outside the set threshold and are not considered.
[0153] Therefore, in the problem studied in this invention, the hidden state of the Hidden Markov Model (HMM) is determined by the indoor crowd WiFi location points and the candidate road segment set of the current WiFi location point. The WiFi location point is considered as the observed value of the state variable in the HMM's hidden Markov chain, and the candidate road segment set corresponding to the WiFi location point is the hidden variable of the hidden Markov state. Then, based on the characteristics of the WiFi location data and the specific situation of the indoor location network model, a state observation probability model and a state transition probability model are established. According to the two basic assumptions of the Hidden Markov Model, the observed value of each hidden state in the HMM is only related to the hidden variable corresponding to that state and is independent of the state value and observed value at any other time. The probability of this process is determined by the observation probability matrix. Simultaneously, since the subsequent state variable is only affected by the previous state, that is, the actual moving road segment of the WiFi location point at the next time moment is only affected by the road segment matched by the WiFi location point at the previous time moment. The probability of this process is determined by the transition matrix. Therefore, in this embodiment, when considering the transition probability value, only the road segment connected to the previous state is considered; otherwise, a situation of matching road segments jumping may occur.
[0154] A schematic diagram of map matching model construction based on Hidden Markov Model for indoor WiFi positioning data is shown below. Figure 23 As shown in the diagram. Indoor WiFi positioning data corresponds to the observation sequence, and the indoor location network model corresponds to the set of all hidden states. The aim is to inversely calculate the corresponding state sequences after determining the observation probability matrix and state transition matrix.
[0155] 2.1) Construction of the observation probability model.
[0156] The observation probability refers to the probability that the candidate road segment corresponding to a WiFi location point is the actual location. In the map matching process, the observation probability at time t=1 is the initial probability.
[0157] There are two main factors affecting the observation probability: WiFi positioning error and the directional error between the WiFi observation point and the candidate road segment. Since indoor location network models are not directional, and people indoors may exhibit back-and-forth movement, this embodiment does not consider directional error. Common WiFi positioning errors follow the following three distributions:
[0158] ① Laplace distribution: X ~ L(μ,β), where μ and β are the location and scale parameters, respectively, and their expected value and variance are:
[0159] E(x)=μD(x)=2β 2 (12)
[0160] ② Normal distribution: X ~ N(μ,σ) 2 ), where μ and σ2 These are the expected value and the variance, respectively:
[0161]
[0162] ③ P-norm distribution: The expected value and variance of this distribution are (denoted by μ, β, σ):
[0163]
[0164] When p=1, the P-norm distribution follows a Laplace distribution; when p=2, the P-norm distribution follows a normal distribution. Since the P-norm distribution is typically used in theoretical research and is less common in practical applications, and since most map matching algorithms in recent years have used a normal distribution to simulate the observation probability model, and experiments have shown that the error of indoor WiFi positioning data also approximately follows a normal distribution, this paper uses a normal distribution to describe the observation probability model. Therefore, the hidden state at time t is r. i The corresponding observed value is o t The conditional probability is:
[0165]
[0166] Where, σ t Let σ be the standard deviation of WiFi location data. t =1.4826median(d t ). d t Let r be the WiFi location point at time t and the candidate road segment r. i The distance is expressed as ||o. t -x t,i || greatcircle That is, the observed value o t To the corresponding candidate road segment r i Matching point x on t,i (This article takes o) t In r i The distance on the Earth's surface (great circle distance) is calculated by using the vertical projection point on the Earth as the matching point. Since the research content of this embodiment is within a geographically small indoor area, the influence of the Earth's curvature can be ignored. Therefore, the Euclidean distance is used instead of the great circle distance to simplify the calculation.
[0167] Where, the matching point x t,i The calculation method is as follows: Let the observed value be o. t (x o ,y o Corresponding candidate road segment r i The two endpoints are P(x) p ,y p ) and Q(x q ,y q If x t,iThe coordinates are:
[0168]
[0169] Where k is the candidate road segment r i The slope, that is:
[0170] 2.2) Construction of the state transition probability model.
[0171] The state transition probability refers to the probability that a WiFi location point will move from one candidate road segment at time t to another candidate road segment at time t+1 in an indoor location network model. It represents the probability of transitions between adjacent states in a Hidden Markov Model. The most significant influencing factor in the state transition probability model is the difference between the distance between the two candidate points corresponding to two adjacent WiFi location points at sampling times and the distance between these two WiFi location points (the great circle). The smaller the distance between the two candidate points corresponding to adjacent WiFi location points on their respective candidate road segments, the greater the likelihood of ultimately selecting that road segment, and thus the higher the transition probability value of that WiFi location point on that candidate road segment. The state transition probability approximately follows an exponential probability distribution.
[0172]
[0173] Among them, dist t For two adjacent WiFi location points (o t ,o t+1 The distance on the Earth's surface and the matching points on the candidate road segments corresponding to these two positioning points. The difference in the actual path distance between them, that is:
[0174]
[0175] Similarly, since the problem studied in this paper is located in an indoor area with a relatively small geographical span, the Euclidean distance is used to approximate the great circle distance to simplify the calculation. The estimation method for parameter β in equation (17) is as follows:
[0176]
[0177] Among them, |||| greatcircle This represents the distance between two points on a great circle; |||| route This represents the actual path distance between two points.
[0178] 3. Model solution based on Viterbi algorithm.
[0179] For the Hidden Markov Model established in step 2, based on the known observation sequence (i.e., the trajectory of WiFi location points), the observation probability model of WiFi location data, and the state transition probability model, the Viterbi algorithm can be used to solve for the most likely hidden state sequence corresponding to a given observation sequence. That is: given the Hidden Markov Model λ=(A,B,Π), where A=[a…]… ij ] n*n Let B be the state transition matrix, and B = [b j (k)] n*m Let Π be the observation probability matrix, and Π = (Π i Let O be the initial state matrix and O be the observation sequence O = (o1, o2, o3, ..., o2). T Solve
[0180] The Viterbi algorithm is essentially a dynamic programming algorithm commonly used for optimal path search problems in graph models. Its basic process is as follows: starting from the initial state, at each state transition, the maximum probability of all possible paths to the next state is recorded. This maximum probability is then used as a benchmark to continue calculations, finally backtracking the entire optimal path to complete the maximum likelihood estimation of the hidden state sequence. Here, a path represents a state sequence in a Hidden Markov Model. It can be observed that compared to traversing all possible paths, the Viterbi algorithm eliminates paths with probabilities lower than the maximum probability at each state node, thus significantly reducing time complexity and improving computational efficiency. The main steps of the Viterbi algorithm are as follows:
[0181] ① Initialization. Define initial values. For the first WiFi location point, its probability value is the probability value of each state multiplied by the probability of observing the first value in each state, i.e., ∏. i b i (o1).
[0182] δ1(i)=∏ i b i (o1), i = 1, 2, ..., n (20)
[0183] ψ1(i)=0,i=1,2,…,n (21)
[0184] ②Recursion. For t=2,3,…,T, given the initial values in (1), calculate the maximum probability δ of each subsequent observation in each state in turn. t (i), and the previous state that reached the maximum probability value (Equation 23).
[0185] δ t(i)=max 1≤j≤n [δ t-1 (j)a ji ]b i (o t ), i = 1, 2, ..., n (22)
[0186] ψ t (i) = argmax 1≤j≤n [δ t-1 (j)a ji ], i = 1, 2, ..., n (23)
[0187] ③ Termination. Consider the case at t=T, and calculate δ at time T. T The maximum value P of (i) * ψ is the probability of the most likely hidden state sequence occurring. Calculate ψ at time T. t The maximum value of (i) is the most likely hidden state at time T. That is:
[0188] P * =max 1≤i≤n δ T (i) (24)
[0189]
[0190] ④ Optimal path backtracking. For t = T-1, T-2, ..., 1, after calculating the maximum probability value at time T, the optimal path can be backtracked.
[0191]
[0192] ⑤ Find the optimal path, which is the state sequence with the highest probability, and output the result.
[0193]
[0194] Thus, the design and solution of the map matching algorithm based on the Hidden Markov Model for indoor WiFi positioning data have been completed.
[0195] Step 4: Evaluation and accuracy analysis of indoor trajectory correction experiments.
[0196] Based on the Hidden Markov Model-based map matching algorithm and Viterbi model solving algorithm constructed in step three, as well as the indoor location network model completed in step two, map matching experiments were conducted on WiFi positioning data that had undergone gross error filtering preprocessing to correct any biases. To comprehensively test the algorithm's effectiveness, this embodiment starts with verifying the effectiveness of local map matching methods and extends this to verifying the map matching algorithm for the entire indoor location network model. The experiments range from map matching experiments with relatively high-precision experimental test trajectory data to bias correction experiments with general measured WiFi positioning data. The methods also extend from map matching of WiFi positioning sample data containing ground truth to verify the algorithm and infer model parameters, to map matching experiments with large batches of data to verify the algorithm's application value. This comprehensive and multi-faceted approach demonstrates the effectiveness of the algorithm proposed in this invention. Furthermore, a trajectory similarity evaluation index is introduced to quantitatively analyze the experimental results, enhancing persuasiveness and credibility.
[0197] 1. Accuracy analysis indicators based on trajectory similarity.
[0198] To quantify the correction effect of the HMM-based map matching method proposed in this invention on indoor WiFi positioning data, some classic trajectory similarity evaluation metrics are introduced to assess the effectiveness of the algorithm. The trajectory similarity evaluation metrics used in this embodiment include:
[0199] ① Dynamic Time Warping (DTW). The idea behind the DTW algorithm is to locally scale and align two trajectory sequences along the time axis to make their shapes as consistent as possible, thereby maximizing their similarity. The DTW algorithm maps points from two trajectories in a many-to-many manner, thus efficiently solving the problem of data inconsistency. Its dynamic programming algorithm is as follows:
[0200]
[0201] Here, Head(tr1) represents the first point of the trajectory; Rest(tr) represents the subsequence of points other than the first point. Dynamic time correction has no limit on the trajectory length and has good results, but it does not handle noise points, and outliers can have a significant impact on the evaluation results.
[0202] ② Longest Common Sub-Sequence (LCSS). The similarity metric based on the longest common sub-sequence is to calculate the maximum number of points that can be considered the same point, which is the number of pairs of points on two trajectories that satisfy the minimum distance threshold. The dynamic programming algorithm is as follows:
[0203]
[0204] Here, the parameter ε represents the minimum distance threshold; points corresponding to two trajectories that are less than this value are considered to be the same point. It can be seen that the longest common substring algorithm also has no limit on the trajectory length and handles noise points. Noise points, because they deviate from the correct trajectory and have no closely related points, are not included in the final result, making it more robust and the evaluation results more accurate. However, the parameter ε in the algorithm is difficult to estimate, and the algorithm is greatly affected by the parameter.
[0205] ③ Hausdorff Distance. The Hausdorff distance is the maximum distance between the closest points of two trajectories being evaluated. That is:
[0206] d H (tr1,tr2)=max {h(tr1,tr2),h(tr2,tr1)} (30)
[0207] Where h(tr1,tr2) is the one-way Hausdorff distance from tr1 to tr2, i.e.:
[0208]
[0209] ④Frechet Distance. The Frechet distance is the minimum length between corresponding points on two trajectories. Its dynamic programming calculation method is as follows:
[0210]
[0211] Where d(p,q) is the Euclidean distance between two trajectory points p and q, and tr(n-1) is the sub-trajectory of trajectory tr with a length of n-1. The Frescher distance satisfies the triangle inequality and has excellent similarity measurement results between trajectory segments, and is easy to calculate; however, because it is too sensitive to outliers, it is not suitable for comparing the similarity between overall trajectories, and is usually used in conjunction with other evaluation metrics.
[0212] ⑤ Edit Distance on Real Sequence (EDR). Edit distance (EDR) refers to the number of operations (insertion, deletion, or replacement) required to transform tr1 into tr2 given two trajectories tr1 and tr2 of length n and m, and a minimum distance matching threshold ε. Its dynamic programming algorithm is as follows:
[0213]
[0214] in, The edit distance of a trajectory provides a new approach to trajectory similarity measurement, taking into full account the geometric features of the trajectory. However, its drawback is also obvious—it is sensitive to noise points.
[0215] ⑥ Symmetrical Segment-Path Distance (SSPD). The idea behind symmetric segment-path distance comes from Hausdorff distance and segment-path distance (SPD). It evaluates trajectory similarity segment by segment and uses the average (rather than the maximum) distance between the nearest points of two trajectories to combat noise. Furthermore, to maintain symmetry, a symmetric averaging method is used, i.e.:
[0216]
[0217] in:
[0218]
[0219]
[0220] The advantage of symmetrical trajectory segment distance lies in its use of interpolation to complete the trajectory, giving it an advantage in global similarity evaluation; it also reduces sensitivity to noise by using symmetrical averaging and eliminates the need to set threshold parameters. However, its calculation method is relatively complex and computationally expensive.
[0221] In summary, this embodiment uses six evaluation metrics—dynamic time normalization, longest common substring, Hausdorff distance, Fraser distance, edit distance, and symmetric trajectory segment distance—to measure trajectory similarity, and uniformly describes them as "distance values" for easy comprehensive evaluation.
[0222] 2. The accuracy of WiFi trajectory correction was analyzed through experiments.
[0223] To further explore the correction effects of indoor location network models and Hidden Markov Model-based map matching algorithms on indoor WiFi positioning data, this paper conducts map matching experiments using collected test data containing real movement trajectories. The map matching results are evaluated based on metrics such as trajectory similarity to help optimize model parameters and improve map matching accuracy. Figure 24-1 , 24-2 The image shows the WiFi location tracking of a user on this floor (the data was collected by a smartphone in standby mode). Among other things, Figure 24-1 This shows the original WiFi location trajectory. Figure 24-2 The image shows the WiFi positioning trajectory after speed filtering preprocessing. As can be seen from the image, the WiFi positioning accuracy is poor.
[0224] The WiFi positioning data correction method proposed in this invention is introduced to perform map matching on the above trajectory. Figure 24-1 , 24-2 It can be observed that due to issues such as uneven time sampling in WiFi positioning methods, some data exhibits significant sparsity. Furthermore, after data cleaning and preprocessing, some positioning data may be lost. Therefore, an interpolation function is incorporated into the algorithm implementation, introducing a speculative algorithm to recover potentially lost WiFi positioning observation points and to reasonably interpolate data in low-frequency sampling segments to avoid affecting the matching results. Figure 25-1 , 25-2 Figures 25-3 and 25-4 show the map matching results for observation prediction factors of 0, 0.4, 0.6, and 0.8, respectively. The thin gray solid line represents the indoor location network model, the dark gray dashed line represents the preprocessed WiFi positioning trajectory, the thick black solid line represents the WiFi matching correction result, and the line connecting the positioning trajectory and the correction result is the matching line.
[0225] Observation reveals that the map matching effect of this group of WiFi location trajectory data is best when the observation prediction factor is 0.6. Therefore, it can be inferred that the observation prediction factor should be set around 0.6 when performing batch WiFi location data correction. To quantitatively analyze the trajectory matching results, the trajectory similarity evaluation index from the previous section is introduced. To facilitate visualization of the calculation results, based on the characteristics of each index, the calculation results are normalized to the (0,1) interval (e.g., the longest common substring is calculated by dividing its value by the total number of points on the trajectory, so that the calculated value falls within the (0,1) interval), and then visualized. Figure 26 As shown. By Figure 26 It can be seen that setting the observation inference factor around 0.6 has the best correction effect on this group of WiFi positioning data. The optimal matching result can restore the trajectory to about 70% similarity with the real trajectory, which greatly improves the accuracy and reliability of WiFi positioning data and can provide support for downstream correlation analysis applications based on indoor WiFi positioning data.
[0226] In summary, this invention addresses the needs of indoor positioning data analysis applications by introducing spatial constraints from the indoor environment to achieve large-scale WiFi trajectory correction, providing more accurate positioning data for downstream analysis tasks. Overall, this invention has the following characteristics:
[0227] ①Based on the data distribution characteristics and physical constraints such as the movement speed of people indoors, the WiFi positioning data is filtered out for gross errors, and the coordinate system of the positioning data is transformed by affine transformation, thus realizing the WiFi positioning data preprocessing.
[0228] ② For indoor positioning data post-processing, we conducted an in-depth study on indoor location network models based on Voronoi diagrams, proposed an automatic construction method for indoor location network models based on skeleton line generation under Delaunay triangle constraints and road network densification in open areas, and conducted experimental analysis on the spatial representation capability of the model.
[0229] ③ An indoor WiFi trajectory correction method based on Hidden Markov Model map matching was designed. The effectiveness of the correction algorithm was experimentally verified on an indoor WiFi positioning dataset by collecting ground truth samples of positioning data. The WiFi correction accuracy was analyzed using the trajectory similarity evaluation method.
[0230] ④ By analyzing the changes in crowd distribution hotspots and pedestrian movement characteristics before and after the application of the algorithm, the effectiveness and practicality of the algorithm were further explored, providing possible directions for the application of the research results in this paper.
Claims
1. A trajectory correction method based on an indoor location network model, characterized in that, Includes the following steps: 1) Obtain indoor positioning trajectory data; 2) Using indoor floor plan data, construct an indoor location network model in the following manner: Extract the boundary range of the indoor building and extract the nodes that influence the spatial structure of the indoor building; for two nodes on the building outline, { }and{ }, when the length between two nodes satisfies >When W, an encryption point is set between the two nodes, and the encryption point { }satisfy: , They are nodes { }、{ The coordinates of}; For encryption point { Coordinates; coefficients for , W represents the average width of the passageway or corridor between interior buildings. Construct a Delaunay triangulation based on the extracted nodes using the Delaunay triangulation method; Based on the relationship between the triangular primitives and their neighboring triangles in the Delaunay triangulation, the triangular primitives are classified into three categories: Category I (with a neighboring triangle on only one side), Category II (with neighboring triangles on two sides), and Category III (with neighboring triangles on all three sides). Based on the classification results, the Voronoi road network skeleton lines for indoor buildings are extracted using the corresponding rules: Category I triangular primitives connect the midpoint of the only common edge to its opposite vertex to form a skeleton line segment; Category II triangular primitives connect the midpoints of two common edges; and Category III triangular primitives connect the centroid to the midpoints of the three common edges. Construct an indoor location network model using road network skeleton lines; 3) A map matching method is used to match the trajectory positioning points in the indoor positioning trajectory data to the indoor location network model in order to correct the deviation of the indoor positioning trajectory data.
2. The trajectory correction method based on an indoor location network model according to claim 1, characterized in that, The boundary range of indoor buildings is extracted using the minimum bounding rectangle or the minimum convex hull.
3. The trajectory correction method based on an indoor location network model according to claim 1, characterized in that, It also includes a step of smoothing the extracted Voronoi road network skeleton lines, using the Douglas-Peucker algorithm; the threshold of the Douglas-Peucker algorithm is selected as 1 / 2 encryption step size, and the encryption step size is the encryption step size of the node encryption process.
4. The trajectory correction method based on an indoor location network model according to claim 1, characterized in that, It also includes the step of densifying the open areas in the Voronoi road network, and the method of densifying the road network includes the following steps: establishing a secondary buffer for the Voronoi road network skeleton line with a set sampling threshold as a parameter, and performing uniform point sampling at intervals of the set sampling threshold; and using the samples after trimming the sampling points that are not within the scope of the study as road network densification points for road network densification. Accordingly, the method for constructing an indoor location network model is as follows: establish full connections for all encrypted points and establish connections with the nearest main network node to form the edges of the indoor location network model, thereby constructing the indoor location network model.
5. The trajectory correction method based on an indoor location network model according to claim 1, characterized in that, The map matching method used in step 3) is a map matching method based on hidden Markov models.
6. The trajectory correction method based on an indoor location network model according to claim 5, characterized in that, When constructing the Hidden Markov Model, it is necessary to determine the observation probability and state transition probability in the Hidden Markov Model; The observation probability is described using a normal distribution, and the hidden state at time t is... The corresponding observed value is : in, To determine the standard deviation of the location data, ; The location point at time t is the candidate road segment. The distance, using , i.e., observed values To the corresponding candidate road segment matching points The distance to the Earth's surface, taken exist The vertical projection point on the top is used as the matching point. And matching points The calculation method is as follows: Where (x,y) is the matching point The coordinates; Observed values Corresponding candidate road segments The two endpoints; ; The state transition probability follows an exponential distribution: in, For two adjacent positioning points ( The distance on the Earth's surface and the matching points on the candidate road segments corresponding to these two positioning points ( The difference in actual path distance between ) is: Let be the scale parameter in the Laplace distribution, and The estimation method is as follows: in, This represents the distance between two points on a great circle. This represents the actual path distance between two points.
7. The trajectory correction method based on an indoor location network model according to claim 5, characterized in that, The Viterbi algorithm is used to solve the hidden Markov model.
8. The trajectory correction method based on an indoor location network model according to claim 1, characterized in that, Step 1) also includes a coordinate transformation step for the indoor positioning trajectory data, wherein the formula for coordinate transformation is: Where (X,Y) are the coordinates before the coordinate transformation, () represents the coordinates after coordinate transformation; Let be the affine transformation parameters, and for Orientation deformation parameters, for Translation amount in direction, for Deformation parameters in the direction, for The amount of translation in the direction.
9. The trajectory correction method based on an indoor location network model according to claim 8, characterized in that, It also includes a data cleaning step for the indoor positioning trajectory data after coordinate transformation, and the data cleaning method includes the following steps: Calculate the actual distance between each positioning point based on the transformed coordinates of each positioning point. Calculate the time difference between each location point based on the time stamp of the location point; The average moving speed between each positioning point is calculated using the actual distance between each positioning point and the time difference between each positioning point. Positioning data with an average moving speed below a set speed value are filtered out to achieve data cleaning.
10. The trajectory correction method based on an indoor location network model according to any one of claims 1 to 9, characterized in that, After step 3), there is also a step to evaluate the results of the correction, and the evaluation index used for evaluation includes at least one of the following indices: dynamic time normalization, longest common substring, Hausdorff distance, Fraser distance, edit distance, and symmetric trajectory segment distance.