High-precision indoor and outdoor positioning method based on WiFi FTM distance estimation

By employing a high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation, and utilizing data preprocessing, error modeling, and machine learning compensation, the problem of multipath interference error in indoor and outdoor positioning is solved, achieving high-precision and robust positioning results.

CN121334838APending Publication Date: 2026-01-13TONGJI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511305530.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing indoor positioning technologies suffer from problems such as difficulty in modeling multipath interference errors, low positioning accuracy, and poor environmental adaptability in complex and variable environments. In particular, there is a lack of high-precision and robust positioning solutions in both indoor and outdoor environments.

Method used

We employ a WiFi FTM-based distance estimation method, which dynamically adjusts error modeling and signal quality classification through data preprocessing, error modeling, machine learning compensation, and quality control to construct a high-precision indoor and outdoor positioning method. This method includes data filtering, multinomial fitting, machine learning model training, and adaptive stochastic model optimization.

Benefits of technology

It achieves high-precision indoor and outdoor positioning in complex environments, improving positioning accuracy by more than 64%, reducing positioning error by 49.37% in dynamic environments, achieving an accuracy of 0.82m under multipath interference conditions, and 1.24m under complex interference conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121334838A_ABST
    Figure CN121334838A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation. The method comprises the following steps: step 1, data preprocessing and coarse correction; step 2, modeling an error system; step 3, constructing a regression model and performing multipath compensation; and step 4, quality control optimization: on the basis of the distance observation value after multipath compensation, new feature engineering and target variables are reconstructed based on new observation data. According to the method, dynamic error compensation and quality control based on machine learning are provided, the problems that in the prior art, multi-path interference errors are difficult to model, the positioning precision is low, and the environmental adaptability is poor are solved, and the precision, usability, reliability and robustness of WiFi FTM indoor and outdoor scene positioning are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of positioning technology in building construction, and in particular to a high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation. Background Technology

[0002] With the continuous development of the Internet of Things (IoT), the demand for high-precision and high-reliability indoor positioning has surged for sophisticated tasks such as automated parking and intelligent factories. However, current indoor positioning technologies mainly suffer from the following limitations: Inertial navigation and vision calculate position using accelerometer, gyroscope, and camera data, but the cumulative error increases linearly over time (drifting by about 1-3 meters per hour), and they rely on high-performance GPUs, resulting in high power consumption; Ultra-wideband (UWB) technology uses nanosecond-level pulse signals to achieve centimeter-level ranging, possessing strong anti-interference, low latency, and centimeter-level positioning characteristics, but it relies on dedicated hardware and high-cost base stations, making it difficult to integrate into consumer-grade devices; On the other hand, cost-effective alternatives such as RFID, ZigBee, Bluetooth, WiFi fingerprint recognition, infrared, and acoustic technologies are low-cost and easy to deploy, but they typically only provide meter-level accuracy, are susceptible to environmental changes, and have limited resistance to interference and multipath effects. In particular, fingerprint recognition positioning requires frequent updates to the fingerprint database, which cannot meet the requirements of high-precision applications. Therefore, in complex and ever-changing indoor environments, there is currently no unified solution to achieve robust indoor navigation and positioning. There is an urgent need for advanced technologies that can overcome these limitations and meet the needs of the next generation to make up for the deficiencies.

[0003] The standard protocol (IEEE 802.11-2016) developed by the IEEE 802.11 Working Group in 2016 provides a precise time measurement (FTM) scheme based on round-trip time (RTT), achieving meter-level accuracy distance measurement through signal interaction between smart terminals and wireless access points. Currently, some commercial chipsets and products support this protocol, providing hardware support for the application of this technology and demonstrating broad application prospects. This scheme has also attracted increasing attention from researchers in recent years, achieving a series of forward-looking results in indoor applications. However, existing work largely relies on other information sources, has limited scalability of smart devices, and fails to address systematic errors caused by multipath interference and non-line-of-sight (NLOS) propagation, thus limiting positioning accuracy. Summary of the Invention

[0004] To address the aforementioned problems in existing technologies, the purpose of this invention is to provide a high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation, which solves the problems of difficulty in modeling multipath interference errors, low positioning accuracy, and poor environmental adaptability in existing technologies by providing dynamic error compensation and quality control based on machine learning.

[0005] To address the above problems, the present invention adopts the following technical solution: a high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation, the method comprising the following steps:

[0006] Step 1: Data preprocessing and coarse correction. Obtain the original observation data and compare it with the timestamp to remove redundant data. For data with failed or abnormal observations, outliers are removed using the Gaussian distribution 3σ criterion. For the original observation data, the observation epochs are matched and distinguished based on the preset sampling interval parameter and whether the user has received observations from the current access point.

[0007] Step 2: Error system modeling. Compare the observation data of each epoch with the access point (AP) with known coordinates and eliminate mismatched redundant observation data. Under a fixed environment and parameter configuration, calculate the average deviation between all observed distances and the actual geometric distance based on the sample data. The average deviation is determined as the initial hardware deviation. Then, use a piecewise polynomial fitting method to determine the piecewise threshold and polynomial order, and model the residual observation error data of different distance segments respectively.

[0008] Step 3: Regression model construction and multipath compensation. The observed data corrected in Step 2 are transformed into features. The features of parameter types are summarized in a sliding window to construct a feature set and identify multipath errors caused by multipath interference.

[0009] The feature samples are cleaned, and modeling is performed on the feature sample set to learn the mapping relationship. Through correlation analysis, feature parameters that are weakly correlated with the target error or redundant between features are removed, and feature engineering is performed to reduce the data dimensionality. The dimensionality-reduced observation data is standardized and divided into training and test sets. Model training and prediction are performed using the LightGBM regression model, and the optimal hyperparameter combination is determined through grid search and cross-validation. The regression model is initialized with the optimal hyperparameters, and the training set is used to expand the training. The trained model predicts the multipath error component in the current distance measurement value based on the input features and compensates for it to obtain a more accurate distance estimate.

[0010] Step 4: Quality Control Optimization. Based on the distance observations after multipath compensation, new feature engineering and target variables are reconstructed based on the new observation data; the impact of sample uncertainty and imbalance on model accuracy is balanced to determine the classification threshold; dimensionality reduction of feature engineering data is performed through correlation analysis, and the hyperparameters of the LightGBM classification model are tuned using grid search to determine the optimal parameter combination; a new balanced minority class sample dataset is synthesized using synthetic minority class oversampling technology; the balanced dataset is initialized using hyperparameters, and training is carried out on the training set; an early stopping mechanism is introduced, and model performance is monitored through the validation set, stopping training when performance no longer improves; a stochastic model is constructed based on the classification prediction results, and iterative optimization and least squares localization are performed by combining global and local tests.

[0011] Furthermore, the determination of the initial hardware deviation in step two specifically involves: keeping the environment and configuration parameters unchanged, and determining the initial hardware deviation based on large sample data;

[0012] The functional expression for FTM distance estimation is as follows:

[0013]

[0014] in and The AP captures the time of transmission of measurement frames and the arrival time of ACK frames. and For the terminal (STA), the arrival time of the measurement frame and the time of sending the ACK response are captured; within each burst cycle Mean of repeated measures This is the distance estimate output by the FTM protocol;

[0015] The processing of precise time measurement frames includes system latency. To measure the delay between the departure time of the request frame transmission command and its actual transmission time at the physical layer output, The delay between the time the request frame is received at the physical layer and the time indicated at the medium access control layer. The actual measurement caused by the system delay is the time delay between the reception time of the request frame and the transmission time of the corresponding frame at the physical layer.

[0016]

[0017] Excluding the influence of other error terms, assume A coarse correction for the initial bias; the initial bias is determined based on the mean calibration of multiple APs for each class:

[0018]

[0019] in, It is the number of APs. This is the total number of measurement points corresponding to each AP. The mean distance measured at each measurement point over a fixed duration is represented by the geometric distance between the measurement point and the AP. .

[0020] Furthermore, the method for fitting the residual error using a piecewise polynomial in step two is specifically as follows: the function expression of the continuous piecewise polynomial model is as follows:

[0021]

[0022] in, For residual observation error, and These are the polynomial coefficients. The segmentation threshold is determined based on the fitting accuracy and fitting efficiency indicators. The most reasonable segmentation threshold and polynomial order are selected.

[0023] Furthermore, step three involves retraining the regression model and predicting corrections, specifically as follows:

[0024] The regression model is initialized using the optimal hyperparameters, and training is performed using the training set. An early stopping mechanism is introduced, and the model performance is monitored using the validation set. Training is stopped when the performance no longer improves. The trained model is saved and the predictions are loaded when needed. When in use, the feature data at each epoch is input into the regression model. The predicted values ​​are de-standardized and output as multipath interference predictions. Further multipath compensation is performed on the distance estimate corrected in step 2 to obtain a higher accuracy distance estimate.

[0025] Furthermore, step four involves constructing new observation feature engineering and targets using a sliding window. Specifically, this involves: reconstructing the feature engineering based on the new observation data; and, considering the differences between non-line-of-sight (NLOS) and line-of-sight (LOS) signal estimations, classifying the compensated distance estimates into two quality categories based on error: high-quality LOS signals (marked as 0) and low-quality NLOS signals (marked as 1), as shown in the following formula:

[0026]

[0027] The data undergoes cleaning, dimensionality reduction, standardization, and dataset partitioning. It's important to note that only the model's input data needs standardization here.

[0028] Furthermore, the classification prediction results described in step four are specifically as follows: a synthetic minority class oversampling technique is used to balance the dataset by synthesizing new minority class samples; the classification model is initialized using optimal hyperparameters and trained on the training set; an early stopping mechanism is introduced to monitor model performance through the validation set and stop training when performance no longer improves; the trained model is saved and the prediction is loaded when needed to obtain the prediction results.

[0029] Furthermore, step four involves constructing a stochastic model for classification prediction, and iteratively optimizing the model and performing least-squares localization by combining global and local tests. Specifically:

[0030] The WLS estimate of the unknown parameters is:

[0031]

[0032] in For adaptive stochastic models,

[0033]

[0034] in, , Given coefficients; the stochastic model is further optimized by combining global and local tests. If a gross error is found in a certain observation, and the original classification result is LOS, it is reclassified as NLOS and reweighted; if the original classification is NLOS, the weights are further reduced based on the original weights. .

[0035] Compared with the prior art, the beneficial technical effects of the present invention are as follows:

[0036] 1. This application provides dynamic piecewise error modeling: Near-field and far-field error models are divided based on a distance threshold, and polynomials of different orders are used for fitting each model. This significantly corrects far-field errors without substantially compromising near-field measurement accuracy.

[0037] 2. This application provides a hybrid machine learning architecture: a regression model predicts multipath error, and a classification model is used to distinguish LOS / NLOS signals based on multipath compensation (oversampling is introduced to solve sample imbalance).

[0038] 3. This application provides an adaptive stochastic model: dynamically adjusting the weight coefficients based on the classification results. Attached Figure Description

[0039] Figure 1 This is a flowchart of a high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation, as described in an embodiment of this application. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Example 1

[0042] like Figure 1 As shown, this application provides a high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation, the method including the following steps:

[0043] Step 1: Data Preprocessing Stage

[0044] 1.1 Identification and Removal of Redundant and Gross Data: Signal conflicts lead to redundant observation information received at the same time. Timestamps and observation information are used for comparison and filtering. The FTM distance estimation is actually the average result of multiple bursts. However, when factors such as excessive distance, low power, temperature, humidity, and environmental changes cause observation failure or anomalies, the Gaussian distribution 3σ criterion is combined to remove outliers, which has a good and efficient ability to identify outlier data.

[0045] 1.2 Clock Synchronization Matching; Due to hardware limitations and sampling rate constraints, strict synchronous observation cannot be achieved in the original observation sequence. Interactive signals need to be queued when leaving or arriving at the transceiver, requiring clock matching for identification and location. For the original observation sequence, the matching and differentiation of observation epochs are based on pre-set sampling interval parameters and whether the user has received a new round of observations from the current AP.

[0046] Step 2: Systematic Error Modeling

[0047] 2.1 Compare the observation data of each epoch with the AP of known coordinates and eliminate redundant observations that do not match;

[0048] 2.2 With the environment and configuration parameters unchanged, the initial hardware deviation is determined based on large sample data;

[0049] The functional expression for FTM distance estimation is as follows:

[0050]

[0051] in and The AP captures the time of transmission of measurement frames and the arrival time of ACK frames. and The STA captures the arrival time of the measurement frame and the time to send the ACK response. Within each burst cycle... Mean of repeated measures This is the distance estimate output by the FTM protocol.

[0052] The processing of precise time measurement frames involves a certain system delay, mainly due to... It measures the delay between the start time of the request frame transmission command and its actual transmission time at the physical layer output. It is the delay between the time the request frame is received at the physical layer and the time it is indicated at the media access control layer. This is the time delay between the reception time of the request frame and the transmission time of the corresponding frame at the physical layer. The actual measurement caused by this system delay should be:

[0053]

[0054] This portion of the deviation is usually inseparable and can be approximated as a systematic error. Furthermore, it is not significantly correlated with the distance between the AP and STA, but may vary across different hardware models. Therefore, the influence of other error terms is excluded, i.e., it is assumed that… This allows for a rough correction of the initial bias. To avoid measurement errors from a single AP, the initial bias is determined based on the average calibration of multiple APs for each class.

[0055]

[0056] in, It is the number of APs. This is the total number of measurement points corresponding to each AP. The mean distance measured at each measurement point over a fixed duration is represented by the geometric distance between the measurement point and the AP. .

[0057] 2.3 Piecewise polynomial fitting residual error;

[0058] The continuous piecewise polynomial model consists of two parts, and its functional expression is as follows:

[0059]

[0060] in, For residual observation error, and These are the polynomial coefficients. The segmented threshold is used. To find the most reasonable segmented threshold and polynomial order in formula (4), the threshold and order are selected based on fitting accuracy and fitting efficiency indicators. In some preferred embodiments, the segmented polynomial fitting model can change the segmented threshold and fitting method according to the actual application scenario to achieve higher fitting accuracy.

[0061] Step 3: Machine Learning Compensation Phase

[0062] 3.1 Sliding window construction of observation feature engineering and target; transforming raw data into features can better represent the actual problem processed by the model and improve the accuracy for unknown data. For the user end, the obtained observation information is limited, so it is particularly important to extract feature parameters that are highly correlated with the target based on the known information. In addition to the FTM distance estimation corrected in step 2, External, backward difference Information such as the sliding window mean difference reflects the abrupt change characteristics and the variation characteristics relative to historical information in FTM distance estimation. In addition, there is a high correlation between observation residuals and observation errors themselves, and signal strength, as an auxiliary observation information of the user, can also characterize measurement accuracy. Therefore, based on these information and the Gaussian distribution characteristics of distance estimation, common parameter types such as mean, standard deviation, skewness, and kurtosis characteristics are summarized in the sliding window to identify multipath errors caused by multipath interference.

[0063] 3.2 Correlation analysis is used for feature engineering and data dimensionality reduction;

[0064] First, feature sample cleaning is performed. In practice, due to system defects or other factors, data loss or uncertainty may occur. For outlier records with missing features or NaN values, they are directly deleted. For other outliers exceeding the 3σ criterion, no processing is performed. Instead, mining and modeling are performed directly on the dataset to learn more comprehensive mapping relationships.

[0065] Secondly, feature engineering is used for data dimensionality reduction. Large datasets and excessively high dimensionality can slow down machine learning models and consume excessive hardware. Data dimensionality reduction methods include two approaches: one, such as principal component analysis, disrupts the original data structure to extract key features; the other involves correlation analysis, using certain rules to select data attributes for dimensionality reduction. However, in practical engineering, data inherently possesses significant physical meaning and research value. Therefore, feature engineering-based filtering dimensionality reduction is used without destroying the original data information. Specifically, this includes analyzing the correlation between feature parameters and the target data based on the Pearson correlation index, and performing cross-correlation analysis on the parameters. The absolute value of the correlation coefficient reflects the strength of the correlation; for correlation coefficients below a certain threshold, the feature parameter can be removed to achieve initial dimensionality reduction. Furthermore, the remaining feature parameters also exhibit correlation; iteratively removing parameters with strong cross-correlation achieves further dimensionality reduction. Dimensionality reduction through filtering significantly reduces model training and prediction time and computational resource consumption, while removing noise and redundant information from the data, reducing the risk of overfitting, and improving generalization ability.

[0066] 3.3 Standardization of dimensionality-reduced data and dataset partitioning;

[0067] Standardization transforms data according to its eigenvalue distribution into a distribution with a mean of 0 and a standard deviation of 1. This eliminates the influence of feature dimensions, ensuring a balanced contribution of different features to the model, thereby improving the efficiency and performance of model training. Splitting the dataset into training and validation sets avoids data leakage and ensures that model performance evaluation is independent of the training process. Furthermore, the introduction of a validation set helps optimize hyperparameters, compares the performance of different models during cross-validation to select the optimal model, and identifies whether the model is overfitting the training data.

[0068] 3.4 Grid-based parameter tuning determines the optimal parameter combination for the regression model;

[0069] The LightGBM regression model is chosen, and a search space for hyperparameters is defined. A coarse grid search is performed initially, followed by finer parameter range settings for a more refined search, thus reducing computational costs. K-fold cross-validation is employed to ensure the stability of model evaluation, while an early stopping mechanism is implemented to accelerate the model training process, ultimately yielding the optimal parameters. It should be noted that both the current regression and classification models are based on the LightGBM model, balancing efficiency and accuracy. These can be replaced with other machine learning models such as Random Forest, XGBoost, CatBoost, and neural networks, and the hyperparameter tuning methods can be changed to other approaches, such as random search and Bayesian optimization.

[0070] 3.5 Retrain the regression model and correct the predictions;

[0071] The regression model is initialized using optimal hyperparameters and trained using the training set. To prevent overfitting, an early stopping mechanism is introduced, monitoring model performance using the validation set and stopping training when performance no longer improves. The trained model is saved for later use and can be loaded for predictions when needed. In practical applications, feature data at each epoch is input into the regression model, and the predicted values ​​are destandardized before outputting the multipath interference prediction. Further multipath compensation is applied to the distance estimate corrected in step 2 to obtain a final, more accurate distance estimate.

[0072] Step 4: Quality Control Optimization Phase

[0073] 4.1 Constructing New Observation Feature Engineering and Targets using a Sliding Window: After multipath compensation, the original feature engineering of the high-precision distance estimate is destroyed, requiring reconstruction of the feature engineering based on new observation data. Considering that multipath interference cannot be completely compensated and residual estimation errors still affect positioning accuracy, the compensated distance estimates are divided into two quality categories based on error, namely high-quality LOS signals (marked as 0) and low-quality NLOS signals (marked as 1), as shown in the following formula.

[0074]

[0075] The same data cleaning, dimensionality reduction, standardization, and dataset partitioning are performed. It's important to note that here, only the model's input data needs to be standardized.

[0076] 4.2 Balancing the impact of sample uncertainty and imbalance on model accuracy to determine the classification threshold;

[0077] Based on threshold discrimination of signal quality, the choice of threshold determines the sample balance and uncertainty of the classification model. These two factors jointly affect the accuracy and usability of the classification model, especially when handling complex tasks. As the threshold increases, the samples gradually become imbalanced, skewed towards the LOS category; due to distance error correction, the residual error's relevance to the target decreases; the model's prediction accuracy, precision, and AUC gradually increase, meaning the overall model performance gradually improves with increasing threshold. However, recall and F1 score gradually decrease because, with increasing sample imbalance, the model favors the LOS category, ignoring the NLOS category, leading to incorrect estimations. The goal of this invention is to more accurately and comprehensively identify NLOS signals; therefore, various metrics, especially recall and F1 score, must be considered comprehensively. In some embodiments, the classification threshold can be dynamically adjusted according to the scenario, such as appropriately lowering the threshold under weak interference conditions and relaxing the threshold restriction under strong interference conditions to improve measurement utilization. Furthermore, a multi-classification model can also be used to refine signal quality.

[0078] 4.3 Correlation analysis was used to perform feature engineering and dimensionality reduction of the data, and gridded parameter tuning was used to determine the optimal parameter combination for the LightGBM classification model;

[0079] 4.4 Classification model training and prediction recognition;

[0080] To mitigate overfitting or underfitting caused by potentially imbalanced datasets, a synthetic minority class oversampling technique is employed. This technique synthesizes a new, balanced minority class dataset to improve the model's generalization ability and ensure accuracy and stability on imbalanced datasets. The classification model is initialized with optimal hyperparameters and trained on the training set. To prevent overfitting, an early stopping mechanism is introduced, monitoring model performance on the validation set and stopping training when performance no longer improves. The trained model is saved for later use and can be loaded for predictions when needed.

[0081] 4.5 Based on classification prediction, a stochastic model is constructed, and the model is optimized and least squares localized through iterative modeling by combining global and local tests;

[0082] The WLS estimate of the unknown parameters is:

[0083]

[0084] in For adaptive stochastic models,

[0085]

[0086] Here, , The coefficients are given empirically; the stochastic model is further optimized by combining global and local tests, making full use of redundant observation information to improve positioning accuracy and reliability. Specifically, if an observation is found to have gross errors after testing, and the original classification result is LOS, it is reclassified as NLOS and reweighted; if the original classification is NLOS, the weights are further reduced based on the original weights. .

[0087] Static experimental data show that the FTM ranging accuracy decreased from 2.81 m to 1.48 m, an improvement of 64.00%, with significant correction for larger measurement errors; it achieved a planar positioning accuracy of 0.82 m in a weak interference environment and 1.24 m in a complex interference environment, with the overall positioning accuracy decreasing from 1.94 m to 0.98 m, an improvement of 49.37%.

[0088] The dynamic experiment covered typical indoor scenarios. Using the precise coordinates of the starting point, waypoints, and ending point calibrated with a total station, a reference trajectory was generated using spline interpolation. The positioning error was measured by the closest distance between the positioning result and the interpolated point of the reference trajectory. The dynamic positioning performance was evaluated. The results showed that the estimated trajectory of this invention was the smoothest, had the highest degree of agreement with the reference estimate, and exhibited low positioning dispersion with no obvious outliers. The final statistical results showed that the positioning error decreased from the original 4.04 m to 0.76 m.

[0089] Example 1

[0090] This application provides a high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation. The deployment configuration in this example is as follows:

[0091] The underground parking garage is equipped with 18 access points (spaced approximately 10 meters apart), and each access point is a Google Pixel 6 Pro smartphone.

[0092] The sampling frequency is 5 Hz, and the communication protocol is IEEE 802.11ax.

[0093] Seven static points were randomly deployed within the site, and their precise coordinates were measured using a total station. For each point, approximately 5 minutes of observation data was collected for static positioning.

[0094] Specific implementation steps:

[0095] Step S1:

[0096] S11: Receives the raw observation data stream, parses the timestamp, AP identifier, distance estimate and signal strength information, performs clock matching based on the appearance of the pre-set sampling interval, and obtains all observation values ​​at each epoch.

[0097] S12: Based on whether there are duplicate AP identifiers and whether there are zero or negative observations in each epoch, redundant and abnormal observations are eliminated.

[0098] Step S2:

[0099] S21: Compare and analyze the observation information of each epoch with the prior AP list, and eliminate the observation values ​​that are unknown in the AP;

[0100] S22: Read the precise coordinates of the known AP and the prior initial hardware deviation. And correct the observed values. ;

[0101] S23: Based on the established segmentation thresholds The piecewise polynomial model that is fitted is

[0102]

[0103] Solving for the distance model bias for each observation Save as after correction .

[0104] Step S3:

[0105] S31: Achieve preliminary positioning for each epoch based on least squares and obtain the observation residuals. Based on each AP, statistics are compiled within the current epoch. Distance correction value per epoch Backward difference Sliding window mean difference Observation residuals and signal strength The mean, variance, kurtosis, and skewness of each value are calculated as feature parameters, and the deviation between the measured value and the true value is calculated as the target value. These values ​​are then stored in the regtrain.txt file.

[0106] S32: If the current stage is the regression model training phase, read the complete regtrain.txt file obtained from the large-sample data collection, set the initial number of each type of parameter, calculate the correlation between the feature parameters and the target value, and remove feature parameter columns with a correlation lower than 0.1; calculate and sort the cross-correlation of the feature parameters, and remove feature parameter columns with a cross-correlation higher than 0.6 and low correlation with the target value in order of priority, and save the numbers of the remaining parameters to the reglist.pkl file. If the current stage is the regression model prediction phase, read and filter reglist.pkl according to the feature parameters of each observation value in the current epoch in step S31.

[0107] S33: If the current stage is the regression model training phase, based on the feature engineering after dimensionality reduction in step S32, standardize the parameters and target respectively, save the standardized parameters as regscaler_x.pkl and regscaler_y.pkl, and randomly split the dataset into training and validation sets in a 4:1 ratio. If the current stage is the regression model prediction phase, read regscaler_x.pkl and standardize the features after dimensionality reduction in step S32.

[0108] S34: If the current stage is the regression model training phase, pre-select the hyperparameter tuning range, perform gridded hyperparameter tuning, and determine the optimal solution as the final regression model hyperparameters, i.e., colsample_bytree=0.9, max_depth=4, min_child_samples=20, n_estimators=100, num_leaves=10, subsample=0.8, learning_rate=0.3 in this example. If the current stage is the regression model prediction phase, skip this step.

[0109] S35: If the current stage is the regression model training phase, then based on the training set in S33, substitute the hyperparameters in S34 to retrain the regression model, and add a validation set to evaluate the model performance. Save the finally trained model as regclf.pickle. If the current stage is the regression model prediction phase, read regclf.pickle and input the standardized features from S33 for prediction. Further read regscaler_y.pkl to de-standardize the predicted values ​​to obtain the final predicted values. This corrects the observed value. .

[0110] Step S4:

[0111] S41: If the current stage is the classification model training phase, based on the corrected observations in step S35, re-statistically analyze the previous... Distance correction value per epoch Backward difference Sliding window mean difference Observation residuals and signal strength The mean, variance, kurtosis, and skewness are calculated as feature parameters. Considering practical positioning requirements, empirically, they are set to... If the absolute value of the deviation between the measured value and the true value exceeds the standard, it is classified as NLOS (marked as 1); otherwise, it is classified as LOS (marked as 0). The feature parameters and classification results are stored in the classtrain.txt file in sequence.

[0112] S42: If the current stage is the classification model training phase, read the complete classtrain.txt file obtained from the large-sample data collection, set the initial number of each class parameter, calculate the correlation between the statistical features and the target value, remove feature columns with a correlation lower than 0.05, and save the numbers of the remaining features to the classlist.pkl file. If the current stage is the classification model prediction phase, read classlist.pkl and filter it according to the feature parameters of each observation value in the current epoch in step S41.

[0113] S43: If the current stage is the classification model training phase, standardize the features based on the dimensionality reduction feature engineering in step S42, save the standardization parameters as classscaler_x.pkl, and randomly split the dataset into training and validation sets in a 4:1 ratio. If the current stage is the classification model prediction phase, read classscaler_x.pkl and standardize the features after dimensionality reduction in step S32.

[0114] S44: If the current stage is the classification model training phase, pre-select the hyperparameter tuning range, perform gridded hyperparameter tuning, and determine the optimal solution as the final classification model hyperparameters, i.e., colsample_bytree=0.9, max_depth=4, min_child_samples=20, n_estimators=100, num_leaves=10, subsample=0.8, learning_rate=0.3 in this example. Retrain the classification model based on the training set in S43, and add a validation set to evaluate the model performance. Save the finally trained model as classclf.pickle. If the current stage is the classification model prediction phase, read classclf.pickle and input the standardized features from S43 to perform prediction, obtaining the final classification result.

[0115] S45: If the current stage is positioning, the function model constructed based on the corrected observations in S35 is as follows.

[0116]

[0117] in It is the location estimation up to the first Geometric distances of each AP. Collect data from each epoch. The observations form a matrix equation:

[0118]

[0119] in .

[0120] An adaptive stochastic model is constructed based on the classification results in S44. ,

[0121]

[0122] Estimate the unknown parameters as follows The stochastic model was further optimized by combining overall and local tests. When a gross error was found in a certain observation, if the original classification result was LOS, it was reclassified as NLOS and reweighted; if the original classification was NLOS, the weights were further reduced based on the original weights. After multiple iterations, the final localization result is obtained.

[0123] result:

[0124] (1) After correction in step S23, the root mean square (RMS) of the observation error was improved from 7.29 m to 6.56 m. After correction, the mean error of 80% of the observations was reduced to 3.74 m, which is an improvement of 23.56% compared with 4.90 m before correction; and the mean error of 90% of the observations was reduced to 6.57 m, which is an improvement of 31.31% compared with 9.57 m before correction.

[0125] (2) After correction in step S35, the observation error RMS improved from 3.18 m to 1.07 m, an improvement of 66.52%. After correction, 90% of the error was reduced to 1.53 m, compared with 3.64 m before correction, an improvement of 57.98%; after correction, 95% of the error was reduced to 1.84 m, compared with 6.24 m before correction, an improvement of 70.51%.

[0126] (3) After classification in step S44, the recognition accuracy is divided into different indicators as accuracy rate. accuracy Recall rate F1 score AUC value This indicates that the model has good classification performance.

[0127] Under weak multipath interference conditions, the static positioning accuracy of this invention is 0.60 m, an improvement of 64.58% compared to the original observation of 1.68 m; under strong multipath interference conditions, the static positioning accuracy is 1.69 m, an improvement of 47.19% compared to the original observation of 3.21 m. Finally, the overall static positioning accuracy of this invention is 1.37 m, an improvement of 49.48% compared to the original observation of 2.72 m.

[0128] Example 2

[0129] This application provides a high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation. The deployment configuration in this example is as follows:

[0130] The underground parking garage is equipped with 18 access points (spaced approximately 10 meters apart), and each access point is a Google Pixel 6 Pro smartphone.

[0131] The sampling frequency is 5 Hz, and the communication protocol is IEEE 802.11 ax;

[0132] The small car is fixed on a mobile car and moves randomly around the parking lot while collecting dynamic observation data to verify dynamic positioning.

[0133] The implementation steps are the same as in Example 1.

[0134] Location results:

[0135] Spline interpolation is performed using the coordinates of reference points during the acquisition process to form a reference trajectory. The distance between the positioning result and the nearest reference trajectory interpolation point is considered as the positioning error. Ultimately, the dynamic positioning accuracy of this invention is 0.77 m, which is 81.01% higher than the original observation value of 4.04 m.

[0136] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation, characterized in that, The method includes the following steps: Step 1: Data preprocessing and coarse correction. Obtain the original observation data and compare it with the timestamp to remove redundant data. For data with failed or abnormal observations, outliers are removed using the Gaussian distribution 3σ criterion. For the original observation data, the observation epochs are matched and distinguished based on the preset sampling interval parameter and whether the user has received observations from the current access point. Step 2: Error system modeling. Compare the observation data of each epoch with the access point (AP) with known coordinates and eliminate redundant observation data that do not match. Under a fixed environment and parameter configuration, calculate the average deviation between all observation distances and the actual geometric distance based on the sample data. The average deviation is determined as the initial hardware deviation. Then, a piecewise polynomial fitting method is used to determine the piecewise threshold and polynomial order, and to model the residual observation error data for different distance segments respectively; Step 3: Regression model construction and multipath compensation. The observed data corrected in Step 2 are transformed into features. The features of parameter types are summarized in a sliding window to construct a feature set and identify multipath errors caused by multipath interference. The feature samples are cleaned, and modeling is performed on the feature sample set to learn the mapping relationship. Through correlation analysis, feature parameters that are weakly correlated with the target error or redundant between features are removed, and feature engineering is performed to reduce the data dimensionality. The dimensionality-reduced observation data is standardized and divided into training and test sets. Model training and prediction are performed using the LightGBM regression model, and the optimal hyperparameter combination is determined through grid search and cross-validation. The regression model is initialized with the optimal hyperparameters, and the training set is used to expand the training. The trained model predicts the multipath error component in the current distance measurement value based on the input features and compensates for it to obtain a more accurate distance estimate. Step 4: Quality Control Optimization. Based on the distance observations after multipath compensation, reconstruct new feature engineering and target variables; balance the impact of sample uncertainty and imbalance on model accuracy to determine the classification threshold; perform data dimensionality reduction for feature engineering through correlation analysis, and fine-tune the hyperparameters of the LightGBM classification model using grid search to determine the optimal parameter combination; synthesize a new balanced minority class dataset using synthetic minority oversampling technology, initialize the classification model with hyperparameters, and train it on the training set, introducing an early stopping mechanism to monitor model performance through the validation set, stopping training when performance no longer improves; construct a stochastic model based on the classification prediction results, and perform iterative optimization and least squares localization by combining global and local tests.

2. The high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation according to claim 1, characterized in that, Step two, which involves determining the initial hardware deviation, specifically involves keeping the environment and configuration parameters constant and determining the initial hardware deviation based on a large sample of data. The functional expression for FTM distance estimation is as follows: in and The AP captures the time of transmission of measurement frames and the arrival time of ACK frames. and For the terminal (STA), the arrival time of the measurement frame and the time of sending the ACK response are captured; within each burst cycle Mean of repeated measures This is the distance estimate output by the FTM protocol; The processing of precise time measurement frames includes system latency. To measure the delay between the departure time of the request frame transmission command and its actual transmission time at the physical layer output, The delay between the time the request frame is received at the physical layer and the time indicated at the medium access control layer. The actual measurement caused by the system delay is the time delay between the reception time of the request frame and the transmission time of the corresponding frame at the physical layer. Excluding the influence of other error terms, assume A coarse correction for the initial bias; the initial bias is determined based on the mean calibration of multiple APs for each class: in, It is the number of APs. This is the total number of measurement points corresponding to each AP. The mean distance measured at each measurement point over a fixed duration is represented by the geometric distance between the measurement point and the AP. .

3. The high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation according to claim 1, characterized in that, The method for fitting residual error using piecewise polynomials in step two is as follows: The function expression of the continuous piecewise polynomial model is as follows: in, For residual observation error, and These are the polynomial coefficients. The segmentation threshold is determined based on the fitting accuracy and fitting efficiency indicators. The most reasonable segmentation threshold and polynomial order are selected.

4. The high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation according to claim 1, characterized in that, Step 3, which involves retraining the regression model and making prediction corrections, specifically includes: Initialize the regression model using optimal hyperparameters, expand the training using the training set, introduce an early stopping mechanism, monitor model performance using the validation set, and stop training when performance no longer improves; save the trained model and load the predictions when needed; When in use, the feature data of each epoch is input into the regression model. After the predicted value is destandardized, the output is a multipath interference prediction. Further multipath compensation is performed on the distance estimate after the correction in step 2 to obtain a higher accuracy distance estimate.

5. A high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation according to claim 1, characterized in that, Step four involves constructing new observation feature engineering and targets using a sliding window. Specifically, this involves reconstructing the feature engineering based on the new observation data; and, considering the differences between non-line-of-sight (NLOS) and line-of-sight (LOS) signal estimations, classifying the compensated distance estimates into two quality categories based on error: high-quality LOS signals (marked as 0) and low-quality NLOS signals (marked as 1), as shown in the following formula: The data undergoes cleaning, dimensionality reduction, standardization, and dataset partitioning. It's important to note that only the model's input data needs standardization here.

6. A high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation according to claim 1, characterized in that, The classification prediction results described in step four specifically involve using a synthetic minority class oversampling technique to balance the dataset by synthesizing new minority class samples. The classification model is initialized with optimal hyperparameters and trained on the training set. An early stopping mechanism is introduced to monitor model performance through the validation set and stop training when performance no longer improves. The trained model is saved and the prediction is loaded when needed to obtain the prediction results.

7. A high-precision indoor and outdoor positioning method based on WiFi FTM distance estimation according to claim 1, characterized in that, Step four describes the construction of a stochastic model for classification prediction, followed by iterative model optimization and least-squares localization using both global and local tests. Specifically: The WLS estimate of the unknown parameters is: in For adaptive stochastic models, in, , Given coefficients; the stochastic model is further optimized by combining global and local tests. If a gross error is found in a certain observation, and the original classification result is LOS, it is reclassified as NLOS and reweighted; if the original classification is NLOS, the weights are further reduced based on the original weights. .

Citation Information

Cited By

  • UWB positioning error correction method for complex indoor environment

    CN121596203A