A multimodal indoor localization method and system based on deep generative learning networks

By combining deep generative learning networks with magnetic field fingerprints and inertial data, a multimodal indoor positioning method is developed, which resolves the contradiction between positioning accuracy and privacy protection in existing technologies, achieving high-precision indoor positioning and effective location privacy protection.

CN118382065BActive Publication Date: 2026-03-13YANSHAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-22
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing indoor positioning technologies struggle to effectively protect user location privacy while maintaining positioning accuracy. Existing privacy protection methods suffer from high computational overhead, low positioning accuracy, or are easily cracked.

Method used

A multimodal indoor positioning method based on deep generative learning networks is adopted. Data is collected by the magnetometer, accelerometer and gyroscope of mobile devices. The server uses variational autoencoder to extract sequence feature vectors and combines magnetic field fingerprints and inertial data to generate new modal features, thereby achieving high-precision positioning. At the same time, the weak position correlation of inertial data is used to protect privacy.

Benefits of technology

While ensuring sub-meter level positioning accuracy, it effectively prevents location privacy leakage, reduces computational overhead, and improves positioning accuracy, thus achieving efficient location privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118382065B_ABST
    Figure CN118382065B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of indoor positioning technology, specifically disclosing a multimodal indoor positioning method and system based on deep generative learning networks. The system includes a mobile device, a server, and a service device. It collects azimuth, angular velocity, and acceleration sequence data through the magnetometer, accelerometer, and gyroscope of the mobile device. The data is converted into a two-dimensional fingerprint image on the server, and a sequence feature vector is extracted from the fingerprint image. This extracted sequence feature vector is then concatenated with the fingerprint information collected by the mobile device and used as input to the convolutional neural network of the mobile device. The mobile device outputs the predicted position result through the convolutional neural network. This invention utilizes accelerometer and gyroscope data with weak global position correlation to generate position feature vectors in the cloud and achieves accurate global position estimation with magnetometer data on the user device. It prevents location privacy leakage while ensuring sub-meter level positioning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of indoor positioning technology, and in particular to a multimodal indoor positioning method and system based on deep generative learning networks. Background Technology

[0002] With the continuous improvement of global indoor positioning systems, outdoor positioning can achieve relatively accurate positioning in most situations. People can use global navigation systems such as GPS, BeiDou, and GLONASS to provide services. However, due to insufficient signal penetration, these systems suffer severe attenuation indoors, resulting in extremely poor positioning accuracy and rendering them unusable indoors. Based on this, research into indoor positioning has been ongoing for many years. Even now, indoor positioning services still have a huge market potential. Related reports indicate that the compound annual growth rate of the indoor positioning and navigation market is expected to reach 34.07% from 2022 to 2027.

[0003] While the market potential is enormous, indoor environments are far more complex than outdoor environments. The varying positioning conditions present significant uncertainties for indoor positioning scenarios, environments, and targets. To achieve accurate and reliable indoor positioning, numerous radio positioning technologies have emerged over the past two decades. These technologies can be categorized by signal type, including WiFi, Bluetooth, magnetic field, visible light, LoRa, ultra-wideband, RFID, acoustic signals, and visual imaging. Among these technologies, fingerprint positioning, which utilizes different signals, has garnered significant attention due to its low deployment complexity. Fingerprint positioning is an indoor positioning technology that uses signals as fingerprints and similarity techniques to match locations. The ubiquitous nature of magnetic field signals has catalyzed the development of magnetic field positioning technology. Magnetic field data is relatively stable and does not easily change over time. Using magnetic fields as the signal source for fingerprint positioning not only reduces maintenance costs but also generally provides higher accuracy than WiFi fingerprint positioning.

[0004] With the increasing prevalence of indoor positioning technology, privacy issues are receiving growing attention. Because sensitive information may be leaked through location data, there is a risk of tracking and surveillance. Privacy protection in positioning technology is a double-edged sword. On the one hand, Localization Servers (LS) can track a user's location in real time based on uploaded data and may share this data with third parties; on the other hand, if the database and algorithms are sent to the user, allowing them to calculate their own location, then LS will lose its privacy and may be subject to targeted attacks and abuse.

[0005] As indoor positioning technology has developed to its current stage, increasing privacy and security issues are the main challenges hindering market growth. Existing work generally focuses on improving the accuracy, stability, continuity, and maintainability of indoor positioning. However, a significant portion of this work uses collected user data without restraint, which to some extent compromises user privacy. Location privacy and security protection is difficult to achieve, and there is currently no ideal solution. Data encryption increases computational overhead and time costs; data perturbation reduces positioning accuracy; data anonymization is easily cracked; and using middleware requires complex deployment and implementation. While some efforts combine different technologies for data protection, these technologies cannot be arbitrarily stacked, and these efforts still cannot fully meet the needs of location privacy protection.

[0006] Based on the above facts, it is urgent to maintain or even improve positioning accuracy while also protecting location privacy. Summary of the Invention

[0007] To address the aforementioned technical problems, this invention provides a multimodal indoor positioning method and system based on deep generative learning networks.

[0008] To achieve the above objectives, the present invention is implemented according to the following technical solution:

[0009] One objective of this invention is to provide a multimodal indoor localization method based on deep generative learning networks, comprising:

[0010] S1. Acquire azimuth, angular velocity, and acceleration sequence data using the magnetometer, accelerometer, and gyroscope of the mobile device;

[0011] S2. The server acquires the angular velocity and acceleration sequence data collected by the mobile device and converts it into a two-dimensional fingerprint image;

[0012] S3. The server extracts sequence feature vectors from the two-dimensional fingerprint image using a variational autoencoder.

[0013] First, the two-dimensional fingerprint image is input into the VAE model, which consists of an encoder and a decoder. The encoder is responsible for mapping the input two-dimensional fingerprint image to latent feature vectors in the latent space. The encoder output includes the mean vector and variance vector in the latent space. Then, a latent feature vector is sampled from the mean and variance output by the encoder. Finally, the sampled latent feature vector is input into the decoder, which is responsible for decoding the latent feature vector into the original sequence feature vector. This sequence feature vector is used for subsequent sequence data processing and analysis.

[0014] S4. The server concatenates the extracted sequence feature vector with the fingerprint information collected by the mobile device as the input to the convolutional neural network of the mobile device. The mobile device outputs the predicted location result through the convolutional neural network and transmits the predicted location result to the service device.

[0015] First, the user requests location information via their mobile device, which uses built-in sensors to collect real-time sensor data and uploads the weakly correlated location data to the server. Second, the server receives the weakly correlated location data provided by the user, extracts features using a feature extractor, and returns them to the user. Then, the CSLoc system uses a deep generative network combined with magnetic field fingerprints and inertial data to generate new modal features, thereby achieving high-precision indoor positioning. The original sensor data is then stitched together with the features provided by the server as input to the location predictor.

[0016] Further, step S2 specifically includes:

[0017] The server converts the angular velocity and acceleration sequence data into a recursive graph. The formula for the recursive graph is:

[0018]

[0019]

[0020] Where: R represents the recursive graph, which is a standardized Euclidean error matrix; k represents the k-th dimension of the recursive graph; l s The sequence length is represented by ; i and j represent the data at index i and index j in r in the sequence, respectively; u and v represent the indices when the maximum difference is taken when the instantaneous data of the current dimension is subtracted.

[0021] Further, in step S3, the feature length of the extracted sequence features is set to 6, and it is converted from 1×6 sequence features to 2×3 sequence features.

[0022] Furthermore, in step S4, the generated 2×3 sequence features are concatenated with the azimuth, angular velocity and acceleration sequence data collected by the gyroscope of the mobile device to form a 2D feature matrix of size 5×3.

[0023] The second objective of this invention is to provide a multimodal indoor positioning system based on a deep generative learning network, comprising:

[0024] The mobile device, equipped with a built-in magnetometer, accelerometer, and gyroscope, is used to collect azimuth, angular velocity, and acceleration sequence data; the mobile device outputs the predicted position result through a built-in convolutional neural network.

[0025] The server communicates with the mobile device via a network. It acquires the angular velocity and acceleration sequence data collected by the mobile device through a built-in variational autoencoder and converts it into a two-dimensional fingerprint image. It then extracts the sequence feature vector from the two-dimensional fingerprint image and concatenates the extracted sequence feature vector with the fingerprint information collected by the mobile device as the input of the mobile device's convolutional neural network.

[0026] Service equipment communicates with mobile devices via a network to receive location predictions from the mobile devices.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] This invention utilizes accelerometer and gyroscope data with weak global position correlation to generate position feature vectors in the cloud and achieves accurate global position estimation with magnetometer data on the user device. While ensuring sub-meter level positioning accuracy, it makes full use of the uncertainty brought about by the cumulative error of accelerometer and gyroscope to prevent the leakage of location privacy. Attached Figure Description

[0029] Figure 1 This is a structural diagram of the multimodal indoor positioning system based on a deep generative learning network according to the present invention.

[0030] Figure 2 This is a network structure diagram of a variational autoencoder network.

[0031] Figure 3 This refers to the location information for collecting the dataset.

[0032] Figure 4 The features are represented by different original data.

[0033] Figure 5 For the accuracy of CNN and LSTM (l s =10).

[0034] Figure 6 The error is composed of different data for the predictor.

[0035] Figure 7 The localization error of CSLoc is given for different sequence lengths.

[0036] Figure 8 A comparison of errors for different numbers of layers in an autoencoder network.

[0037] Figure 9 This represents the positional error of different feature extractors.

[0038] Figure 10 Ablation experiment based on CSLoc magnetic field data. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. The specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention.

[0040] As we all know, inertial sensors only acquire angular velocity and acceleration data of a device in the object's coordinate system, which cannot determine the user's absolute position in space. IMU-based PDRs often determine the initial position and continuously calibrate using other data. This actually prompts us to think in the opposite direction: the relative position of inertial navigation data makes it difficult to leak user location privacy, so can we use this data to achieve a certain degree of privacy protection? For ease of description, we call this type of data weakly position-correlated modal data. Magnetic field data is the corresponding strongly position-correlated modal data.

[0041] Weakly correlated modal data can protect user location privacy but cannot provide precise positioning. The first challenge is to leverage the positional uncertainties of accelerometers and gyroscopes to achieve server-side location privacy while eliminating their accumulated errors for accurate location estimation.

[0042] Based on the human movement state during data acquisition, weakly position-correlated inertial data can be divided into static data and dynamic data. Most existing works tend to train models based on inertial data acquired when a moving individual stays at a specific location for a specific time. However, in reality, much more inertial data is generated during the movement of the individual, i.e., dynamic data. Overcoming the feature diversity of static and dynamic data to generate global position features is the second challenge.

[0043] To solve these problems, such as Figure 1 As shown, this embodiment proposes a multimodal indoor positioning system based on a deep generative learning network, including:

[0044] The mobile device, equipped with a built-in magnetometer, accelerometer, and gyroscope, collects position, angular velocity, and acceleration sequence data. It outputs predicted position results via a built-in convolutional neural network (CNN). The CNN predictor employs a conventional structure of 4 convolutional layers and 3 fully connected layers. Similar to the network layers in the feature extractor, BatchNorm and LeakyReLU are used after almost every layer. The convolutional layers increase the number of channels, while the fully connected layers compress features. The maximum number of convolutional channels in the predictor is 64. Except for the third layer, where the number of input and output channels remains the same, the number of channels in each subsequent layer doubles.

[0045] The server communicates with the mobile device via a network. It acquires angular velocity and acceleration sequence data collected by the mobile device using a built-in variational autoencoder (VAE), converts it into a two-dimensional fingerprint image, extracts sequence feature vectors from the fingerprint image, and then concatenates these feature vectors with the fingerprint information collected by the mobile device as input to the mobile device's convolutional neural network. Since the input recursive graph or the concatenated feature matrix is ​​insufficient to support a large convolutional kernel, the network structure used in this embodiment employs a 1×1 convolutional kernel in all convolutional layers. A 1×1 convolutional kernel increases the network depth without changing the input dimension. By simply increasing the number of channels, the expressive power, especially for nonlinear features, can be improved without increasing computational cost. No padding or step parameter is added to the convolutional layers. Both the encoder and decoder of the variational autoencoder (VAE) use 3 convolutional layers. The number of channels in the encoder's 3 convolutional layers increases exponentially, while the number of channels in the decoder's 3 convolutional layers decreases exponentially, with the smallest convolutional layer having 32 channels. To extract features, a fully connected layer is added after the convolutional layer in the encoder, and the same fully connected layer is added before the convolutional layer in the decoder to recover the data. The last layer in the decoder also uses Tanh activation.

[0046] Service equipment communicates with mobile devices via a network to receive location predictions from the mobile devices.

[0047] Regarding the optimizer and loss function, the CSLoc system uses the Adam optimizer for both the feature extraction model and the prediction model. The learning rate for the feature extractor is 0.001, and the learning rate for the predictor is 0.0001.

[0048] In this embodiment, the batch size is 16, the feature extractor is trained for a total of 400 rounds, and the predictor is trained for 200 rounds.

[0049] In this embodiment, the multimodal indoor positioning system based on deep generative learning networks is named the CSLoc system. Figure 1 As shown, the mobile device on the left is itself a data acquisition device, so it can use all fingerprint data, and CSLoc imposes almost no restrictions on it. CSLoc only allows the service device on the right to use location-weakly correlated modal data—data from accelerometers and gyroscopes. Device differences manifest on the network specifically as restrictions on input data.

[0050] Unlike traditional privacy protection methods, the CSLoc system in this embodiment attempts to prevent service providers from leaking user location privacy by directly utilizing the weak location correlation of partial modal data. Therefore, CSLoc focuses more on utilizing data rather than encrypting it. Compared to existing fingerprint-based indoor positioning systems, CSLoc's positioning work places greater emphasis on distinguishing between the mobile device and the server. We divide CSLoc into two parts: the convolutional neural network of the mobile device and the variational autoencoder network of the server device.

[0051] The goal of the CSLoc system is to improve positioning accuracy by obtaining effective features from existing modal data while restricting the input modalities. To achieve this goal, such as... Figure 2 As shown, a multimodal indoor localization method based on deep generative learning networks is proposed, including:

[0052] S1. Acquire azimuth, angular velocity, and acceleration sequence data using the magnetometer, accelerometer, and gyroscope of the mobile device;

[0053] S2. The server acquires the angular velocity and acceleration sequence data collected by the mobile device and converts it into a two-dimensional fingerprint image;

[0054] The server converts the angular velocity and acceleration sequence data into a recursive graph. The formula for the recursive graph is:

[0055]

[0056]

[0057] Where: R represents the recursive graph, which is a standardized Euclidean error matrix; k represents the k-th dimension of the recursive graph; l s The sequence length is represented by ; i and j represent the data at index i and index j in r in the sequence, respectively; u and v represent the indices when the maximum difference is taken when the instantaneous data of the current dimension is subtracted.

[0058] The recursive graph R is a transformed normalized Euclidean matrix used to calculate the relative distances between components, which can reduce the impact of sensor errors between smartphones. However, the recursive graph can only capture some features of the time series, and the subtraction operation involved in the formula may lead to the loss of some features. Although the recursive graph can provide some features, they are not distinguishable in terms of orientation, which means that relying solely on the features provided by the recursive graph is insufficient for high-precision indoor positioning. However, this weak correlation of location is exactly what we need. To enhance this weak correlation, this embodiment avoids using magnetic field data for recursive graph transformation, which is the basis for location privacy protection in this embodiment.

[0059] S3. The server extracts sequence feature vectors from the two-dimensional fingerprint image using a variational autoencoder.

[0060] First, the two-dimensional fingerprint image is input into the VAE model, which consists of an encoder and a decoder. The encoder maps the input two-dimensional fingerprint image to a latent feature vector in the latent space. This latent feature vector can be seen as an abstract representation of the input data. The encoder output includes the mean vector and variance vector in the latent space, which describe the distribution in the latent space. Next, a latent feature vector is sampled from the mean and variance output of the encoder. This sampling process introduces a degree of randomness, ensuring that each input can be mapped to a different latent feature vector. Finally, the sampled latent feature vector is input into the decoder, which decodes the latent feature vector into the original sequence feature vector. This sequence feature vector is used for subsequent sequence data processing and analysis.

[0061] S4. The server concatenates the extracted sequence feature vector with the fingerprint information collected by the mobile device as the input to the convolutional neural network of the mobile device. The mobile device outputs the predicted location result through the convolutional neural network and transmits the predicted location result to the service device.

[0062] First, the user requests location information via their mobile device, which uses built-in sensors to collect real-time sensor data. Weakly location-correlated data is then uploaded to the server. Next, the server receives the weakly location-correlated data from the user, extracts features using a feature extractor, and returns them to the user. Then, the CSLoc system uses a deep generative network combined with magnetic field fingerprints and inertial data to generate new modal features, achieving high-precision indoor positioning. The original sensor data is then concatenated with the features provided by the server, serving as input to the location predictor.

[0063] While fingerprint localization methods rely on pre-collected data, they can be limited by dynamically changing environments, requiring timely updates to the fingerprint database to maintain accuracy. To overcome the limitations of sequential data, image neural network processing techniques can be utilized, with recursive graphs being a key solution for processing time series data. In the location prediction process, the user requests location information via their mobile device's sensors. The server extracts features and returns them to the user. The user then concatenates the raw sensor data with the features as input to the location predictor, ultimately estimating the user's true location. Regarding optimizer selection, the Adam optimizer typically outperforms the SGD optimizer because Adam combines momentum and adaptive learning rate features, enabling efficient updates to the neural network weights.

[0064] The CSLoc system is essentially a feature-enhanced indoor positioning system. We utilize new features extracted from the recursive graph and combine them with the original point data to construct a new feature matrix. This step is actually completed locally after the user device obtains the features extracted by the server. The original instantaneous data contains only 3 modalities and a total of 9 data points. To avoid excessively large weights for the new features, this embodiment selects n=6 as the size of the new features. For ease of concatenation, this embodiment converts it from 1×6 to 2×3, resulting in a final 5×3 2D matrix as the input to the location predictor.

[0065] This embodiment uses an open-source dataset for simulation. This dataset collected 36,795 continuous samples from a 185-square-meter indoor environment at a data acquisition frequency of 10Hz, including WiFi fingerprint data, inertial data, geomagnetic field data, and orientation data. The dataset includes two connected corridors, a hall, and two rooms. Data acquisition devices included mobile phones and smartwatches with sensors. The acquisition program was a separately developed Android application, and the data was in triaxial format. The interval between acquisition points was 0.6 meters. The dataset label locations are as follows: Figure 3 As shown.

[0066] To verify the feasibility of this embodiment, this embodiment uses a portion of the data collected by a smartwatch. Compared to data collected by a smartphone, it lacks WiFi data but contains more modal information, such as angular velocity data measured by a gyroscope. In fact, the modal data involved in this embodiment refers to all modal data from the sensors collected by the dataset.

[0067] The dataset provides the following label format for static data:

[0068] {Timestamp of arrival at collection point A, timestamp of departure from collection point A, collection point A}.

[0069] For ease of processing, we provide the following dynamic data label format:

[0070] {Timestamp of leaving collection point A, Timestamp of arriving at collection point B, Collection point A, Collection point B}

[0071] To align with static data, we selected collection point B as the data location label. When processing dynamic data, we also need to know the dataset's path information. Based on the collection time intervals, we provide the dataset's collection path information, as detailed in Table 1.

[0072] Table 1 Data Collection Path

[0073]

[0074] The performance of this embodiment is evaluated using the above data, and the specific evaluation indicators are as follows:

[0075] This embodiment uses a method similar to existing work to evaluate positioning performance. The first method is the regression distance error, which is the Euclidean distance between the predicted result and the actual ground value, as shown in the following formula:

[0076]

[0077] Where (x, y) are the actual ground measurements. This is an estimated value calculated by the system.

[0078] The second type is localization accuracy, which is the classification performance of the prediction results, and the formula is as follows:

[0079]

[0080] Where r is the number of correctly predicted data in the test set, and c is the total number of data in the test set used for calculation.

[0081] Compared to the first type of fine-grained error, the second metric is more intuitive.

[0082] For the evaluation of privacy protection performance, we still used the Euclidean distance in formula (2), which is the positioning error that an attacker can achieve after obtaining all information and models except for magnetic field data.

[0083] To verify the effectiveness of dynamic data extraction and test the performance of using angular velocity, acceleration, magnetic field, and orientation data for localization, this embodiment provides the localization accuracy of the convolutional predictor under different modal combinations, such as... Figure 4 As shown in the figure, the localization accuracy decreases to some extent due to the introduction of dynamic data, regardless of the modality combination. However, the localization accuracy is not significantly different from that of static data, which essentially proves the accuracy and reliability of dynamic data segmentation. In the figure, localization using magnetic field data (MAG) alone is poor, while using only orientation data (ORIEN) or acceleration and gyroscope data (AC+GYRO) results in very poor localization accuracy. The former demonstrates that magnetic field data contains more features, while the latter proves that relying solely on data from inertial sensors cannot provide accurate localization. This also proves the assertion that inertial data has a weak correlation with position, while magnetic field data has a strong correlation with position. When these data are combined, better localization results can be achieved, demonstrating the necessity of multimodal data fusion.

[0084] As mentioned earlier, orientation data is calculated from multiple sensors, which may even include magnetometers. Furthermore, even if the differences are not significant, orientation data is slightly less expressive of features than inertial data. Therefore, in the subsequent algorithm, this embodiment does not use orientation data, but instead uses a combination of inertial and magnetic field data for positioning.

[0085] The predictor is crucial to the performance of the CSLoc localization system. To verify the predictor's performance, we used CNN and LSTM as the network models for the predictor, with the LSTM taking a sequence of data of length 10 as input. In this experiment, we used magnetic field data, acceleration data, and angular velocity data. Figure 5 A bar chart showing the classification accuracy of the test results is provided. When processing dynamic data, CNN's accuracy is approximately 45% higher than LSTM's, while the difference is negligible when processing static data. Since dynamic data only accounts for 30% of the total data, and static data has a compensating effect on the model, the accuracy of mixed data does not decrease significantly.

[0086] The essence of LSTM is to predict the next location of a user based on patterns in existing sequence data. For static data, since the user is stationary and the data changes slowly and regularly, LSTM can handle this well and provide relatively accurate localization results. However, for dynamic data, in the experimental scenario of this example, one second of sequence data is not long enough, and the user's movement is unpredictable, resulting in lower localization accuracy. Another possible explanation is the limited number of samples in dynamic sequence data; LSTM requires more data samples than CNN to achieve better localization accuracy.

[0087] After comparative analysis, we ultimately decided to continue using point-based convolutional neural networks as the predictor. In the following experiments, we consistently used a combination of static and dynamic data, with the predictor employing a CNN structure. Except for the GAN-based generator where static and dynamic data were input separately, all other cases used a mixed input approach. The former had two training sets and two test sets, while the latter had only one training set and one test set.

[0088] Conventional magnetic field fingerprinting methods typically use only static data for processing and calculation. We utilize the characteristics of the dataset to extract dynamic data and attempt to reduce the interference of dynamic data. Figure 6 The graph shows a comparison of regression localization errors when the predictor uses different data compositions. It can be seen from the graph that for the predictor using only traditional CNN convolutions, the localization performance on static data is better than on dynamic data. However, the problem arises when static data is mixed with dynamic data; the predictor's localization performance then deteriorates significantly. Figure 6As shown in the mix diagram, calculations show that the positioning accuracy using only static data (approximately 1.01m) is significantly lower than that using mixed data (approximately 1.26m), by about 25.3%.

[0089] and Figure 5 Accuracy comparisons reveal that regression methods produce more significant differences. This is because classification is a probabilistic, not deterministic, calculation; high probabilities can mask low-probability data distributions. Therefore, some studies calculate coordinate errors based on classification results instead of directly performing regression calculations. This approach typically yields lower localization errors.

[0090] In this experiment, dynamic data, which accounted for only 30% of the dataset, increased the positioning error. In real-world positioning, dynamic data often constitutes a larger proportion than static data, thus the expected positioning accuracy would be even lower. Therefore, this further demonstrates the necessity of processing dynamic data.

[0091] like Figure 7 We tested the localization error of CSLoc under different sequence lengths and ruled out the influence of outliers. The results show that a sequence length of 12 achieves relatively good accuracy. Errors increase with sequence lengths shorter or longer than 12. We speculate that this is related to the small number of data samples near each point in the original dataset. When the sequence length is too short, the recursive graph performance is affected; when the sequence length is too long, the sample size of the preprocessed dataset decreases significantly due to the limited number of samples, resulting in excessively reduced data input to the model and ultimately lower accuracy. In fact, excessively long sequence lengths also significantly reduce the real-time performance of localization, hindering real-time localization and tracking.

[0092] The figure also shows that the impact of sequence length on localization error is negligible, regardless of the sequence length. This is also due to the superiority of recursive graphs in processing short time-series data. In this embodiment, except for the experiment verifying the effect of sequence length on system performance, all other experiments used data with a sequence length of 12 as the generator input to achieve the best localization effect.

[0093] To extract more effective features, this embodiment tested autoencoders with different numbers of network layers for comparative experiments. The experimental results are as follows: Figure 8 As shown, similar to the predictor, the optimal number of convolutional layers for the autoencoder is also 3, significantly different from other layers. We believe this is because the number of parameters is too low before 3 layers, while after 3 layers, parameter updates in some parts of the network stagnate, resulting in the vanishing gradient phenomenon.

[0094] In the previous introduction to network parameters, this embodiment gave the minimum number of channels for the encoder and decoder in the autoencoder. When the convolutional layers reach 5, the maximum number of channels is 512. Another reason for the poor performance of 5 convolutional layers may be that the number of channels is too large. Since the input recurrence graph is actually only 6×12×12 in size, even with 5 fully connected layers, gradient updates may be insensitive without using residual networks and dropout.

[0095] Based on the test results and the above analysis, this embodiment ultimately selected a 3-layer convolutional structure to form the encoder and decoder. This reduces the model size while providing higher localization accuracy. The final autoencoder model size in this embodiment is approximately 2MB, which allows for localization of the autoencoder model. However, since training an autoencoder based on inertial data is an unsupervised learning process, and in a real-world production environment, the server can receive a large amount of real-time data to supplement and update the training model, placing the model on a server results in higher feature extraction capabilities and more frequent model updates.

[0096] like Figure 9 As shown, we present the cumulative distribution of location errors based on different feature extractors when using mixed data. Figure 9 The figures show the cumulative distribution of location errors when using VAE, GAN, and predictor-only methods. The CDF (Cumulative Distribution Function) graph clearly shows that while GAN networks increase the number of low-error data points, they have no significant impact on high-error data. VAE, however, significantly outperforms GAN networks and predictor-only methods, achieving an average localization error of 0.90m, which is better than the 1.25m localization error when using only predictor methods with the original data.

[0097] This embodiment attributes the performance improvement of GAN networks to their attempt to learn the entire recurrence graph input. VAE, on the other hand, aims to master the extraction of features from a recurrence graph. For high-precision training and testing data with similar features, GANs can achieve better fitting than classic convolutional networks. However, for outlier data, GANs are prone to generating features that deviate from the target data. This embodiment attributes the improved performance of VAE to the increased feature diversity of sequential data. Although these features cannot be directly used for indoor positioning, the features extracted from the recurrence graph complement existing 3D data, thus achieving higher accuracy.

[0098] In an extreme scenario, we assume that the network model and all data uploaded to the server are leaked, with only the magnetic field data remaining protected because it was not transmitted. To address this, we designed a set of control experiments, such as... Figure 10 As shown, even with the loss of magnetic field data, CSLoc's average error is 19m, far exceeding the average error of 0.9m. This essentially proves that CSLoc virtually eliminates the possibility of service providers leaking user locations. When using CSLoc, as long as local data leakage does not occur, user location privacy and security will be effectively guaranteed.

[0099] CSLoc can effectively protect user location information without losing magnetic field data, essentially based on the weak correlation of inertial data. This leverages the relativity of inertial data positioning, similar to PDR. While recursive graph processing of inertial data can improve CSLoc's positioning performance, the extracted features are not unique due to the absence of magnetic field data. This means that this method alone cannot achieve a significant improvement in positioning accuracy. Therefore, further enhancing data utilization without compromising user location privacy will be a key focus of this embodiment. Similarly, while weakly correlated location data does not imply no correlation, inertial data, though not revealing location information, does reveal some user behavior patterns. This can be addressed by combining it with other privacy protection measures, but is not the focus of this embodiment.

[0100] In addition, we compared the positioning accuracy with the latest indoor positioning-related work, as shown in Table 2.

[0101] Table 2 Comparison of accuracy for different positioning operations

[0102]

[0103] As shown in Table 2, the positioning accuracy of the present invention is improved by 38.89% compared with the traditional solution, and the final average positioning error is about 0.9m. On this basis, the present invention also makes full use of the uncertainty brought about by the cumulative error of the accelerometer and gyroscope to prevent the leakage of location privacy, thereby providing a relatively excellent privacy protection capability.

[0104] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.

Claims

1. A multimodal indoor localization method based on deep generative learning networks, characterized in that, include: S1. Acquire azimuth, angular velocity, and acceleration sequence data using the magnetometer, accelerometer, and gyroscope of the mobile device; S2. The server acquires the angular velocity and acceleration sequence data collected by the mobile device and converts it into a two-dimensional fingerprint image; step S2 specifically includes: The server converts the angular velocity and acceleration sequence data into a recursive graph. The formula for the recursive graph is: ; ; in: The recursive graph is represented by a standardized Euclidean error matrix. Representing the first recursive graph Dimension; Represents the length of the sequence; and The representative sequence represents respectively Chinese index and index Data at the location; and This represents the index at which the maximum difference is obtained when the instantaneous data of the current dimension is subtracted; S3. The server extracts sequence feature vectors from the two-dimensional fingerprint image using a variational autoencoder. First, the two-dimensional fingerprint image is input into the VAE model, which consists of an encoder and a decoder. The encoder is responsible for mapping the input two-dimensional fingerprint image to latent feature vectors in the latent space. The encoder output includes the mean vector and variance vector in the latent space. Then, a latent feature vector is sampled from the mean and variance output by the encoder. Finally, the sampled latent feature vector is input into the decoder, which is responsible for decoding the latent feature vector into the original sequence feature vector. This sequence feature vector is used for subsequent sequence data processing and analysis. S4. The server concatenates the extracted sequence feature vector with the fingerprint information collected by the mobile device as the input to the convolutional neural network of the mobile device. The mobile device outputs the predicted location result through the convolutional neural network and transmits the predicted location result to the service device. First, the user requests location information via their mobile device, which uses built-in sensors to collect real-time sensor data and uploads the weakly correlated location data to the server. Second, the server receives the weakly correlated location data provided by the user, extracts features using a feature extractor, and returns them to the user. Then, the CSLoc system uses a deep generative network combined with magnetic field fingerprints and inertial data to generate new modal features, thereby achieving high-precision indoor positioning. The original sensor data is then stitched together with the features provided by the server as input to the location predictor.

2. The multimodal indoor positioning method based on deep generative learning networks according to claim 1, characterized in that, In step S3, the feature length of the extracted sequence features is set to 6, and it is converted from 1×6 sequence features to 2×3 sequence features.

3. The multimodal indoor positioning method based on deep generative learning networks according to claim 2, characterized in that, In step S4, the generated 2×3 sequence features are concatenated with the azimuth, angular velocity and acceleration sequence data collected by the gyroscope of the mobile device to form a 2D feature matrix of size 5×3.

4. A multimodal indoor positioning system based on a deep generative learning network, used to execute the multimodal indoor positioning method based on a deep generative learning network as described in any one of claims 1-3; characterized in that, include: Mobile devices with built-in magnetometers, accelerometers, and gyroscopes are used to collect azimuth, angular velocity, and acceleration sequence data. Mobile devices output predicted location results through a built-in convolutional neural network; The server communicates with the mobile device via a network. It acquires the angular velocity and acceleration sequence data collected by the mobile device through a built-in variational autoencoder and converts it into a two-dimensional fingerprint image. It then extracts the sequence feature vector from the two-dimensional fingerprint image and concatenates the extracted sequence feature vector with the fingerprint information collected by the mobile device as the input of the mobile device's convolutional neural network. Service equipment communicates with mobile devices via a network to receive location predictions from the mobile devices.

Citation Information

Patent Citations

  • Wireless positioning method and system

    CN112312541A