Training method for a localization model, localization method, localization device and computer program product
Patent Information
- Application Number
- PCT/EP2026/059054
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure IMGF000023_0001 
Figure 00000039_0000 
Figure 00000039_0001
Abstract
Description
[0001] Training method for a localization model, localization method, localization device and computer program product
[0002] In the following, different aspects of embodiments in the context of indoor localization technology are described, wherein reference is made to appendices.
[0003] Appendix 1 refers to the background and the results obtained with certain embodiments using CNN (convolutional neural networks) and LSTM (long short-term memory model) networks in the field of indoor localization.
[0004] Appendix 2 refers to the data, devices and processes used in some embodiments.
[0005] Appendix 3 refers details of the models used in some embodiments.
[0006] Appendix 4 refers to an embodiment of an indoor Localization system, i.e. everGuide. In the following embodiments are described in the context of figures
[0007] Fig. 1 for Appendix 1 shows sample magnetic field data forming a specific pattern and sliding window size based on best performing sequence Length for the Talbot building; Fig.2 for Appendix 1 shows example models for the CSL building using the TSER approach with LSTM (left) and the TSR approach with CNN architecture (right)
[0030] ;
[0008] Fig.3 for Appendix 1 shows a comparison of models using a sequence length of 200 in the same Talbot building trial in regression mode. LSTM models perform much better here; Fig. 4 for Appendix 1 shows example TSC performance for different grid sizes with a sequence length of 200 on the same test set trial. Lower grid sizes lead to a lower accuracy but also reduce the MAE;
[0009] Fig. 5 shows the effect of a background calibration technique applied to magnetometer readings which are reported as uncalibrated by the smartphone operating system;
[0010] Fig. 6 shows the effect on values which are automatically pre-calibrated by the smartphone;
[0011] Fig. A2.1 of Appendix 2 shows a process per building, floor or other spatial boundary;Fig. A2.2 of Appendix 2 shows a close-up on the data flow from the data collection to the trained model;
[0012] Fig. A3.1 of Appendix 3 shows two example models;
[0013] Fig. A4.1 of Appendix 4 show a functional diagram of the indoor navigation.
[0014] The concept of indoor localization is known per se. The prior art and the resulting challenges are described in Appendix 1, section I. For example, it is known to use radio based data (e.g. Bluetooth) or camera data to help in the location of an object and / or a person in an indoor spatial area. Spatial area can imply a two-dimensional area ora three-dimensional space e.g. in a building.
[0015] The use of other field data, such as magnetic field data, has also been described (see Appendix 1 , Section I) in this context. But this has only been used insofar that magnetic field data is evaluated in this context using an offline generated grid map (for example using a probabilistic filter together with a step detector) or using offline generated trials collected from a reference sensor system (see Appendix 1 , Section II).
[0016] The embodiments described herein use a different approach, as claimed in the independent claims. Embodiments are described in the dependent claims.
[0017] It should be noted that the appendices refer to magnetic data for the localization. It is possible to use field data in general, e.g. acoustic field data, electromagnetic field data, radiation field data instead and / or magnetic field data. Magnetic and radiation field data are mainly due to natural processes, but it might be overlayed with man-made sources. In one step, a spatial representation (in 2D or 3D) of an indoor spatial area is generated using markers and / or comprises markers in the indoor spatial area, in particular markers being optical anchors, optical markers, acoustic markers, radio beacons and / or ranging radar / LIDAR data.
[0018] After the establishment of the spatial representation, field data is measured within the indoor spatial area using at least one field data sensor.
[0019] That field data is then augmented with positional data from the spatial representation previously obtained and / or with temporal data, to obtain a field data spatial representation. The field data spatial representations can e.g. be clustered as a set oftrials in which each trial comprises magnetic field information at several specific positions (e.g. along a path) within the indoor spatial area
[0020] Then a deep neural net model is computed and / or trained using the field data spatial representations generated.
[0021] With that method it is possible to generate a model which enables a device to determine its location within the indoor spatial area using field data, such as measure magnetic data, either only or in conjunction with other data.
[0022] In one embodiment the measurement of the field data can be performed e.g. by a robot which records the marker data to build the spatial representation while measuring the magnetic data at the same time.
[0023] In one embodiment, not only point magnetic data is measured, but magnetic data is measured in trials or along trails in the indoor spatial area. Therefore, the three-dimensional magnetic field vector is not only measured at discrete points, but in or along paths in the indoor spatial area. This is described e.g. in Appendix 1, Section III. In particular, a trail can contain an ordered list of points (spatial data) where at least some of the points are coupled with the magnetic data measured at those points. Using such magnetic field trails can both improve the (self-)positioning of the sensing device and the accuracy of the map enhanced by this data.
[0024] In one embodiment there is a post-training, re-training and / or an adaption period of the deep neural net model in which furtherfield data spatial representation data, in particular further data which has not been used in the generation of a previous version of the deep neural net model is used. This allows e.g. for an adaptation or improvement of the deep neural model during the normal use of the spatial area.
[0025] In a further embodiment the field data is measured through a crowd-measurement (e.g. users with mobile phones having magnetic sensors moving through or in the spatial area) and / or an on-the-fly measurement by non-specialized user equipment, in particular mobile devices using a localization app (see e.g. Appendix 4). This allows the use of “normal” equipment of users to improve the deep neural net model during the normal use of the building.In Appendix 2, the data, in particular the magnetic measured data, is described. Furthermore, some examples for the device used in the embodiments described herein are detailed. Then, an example for the training and the use of the embodiment is given. In Fig. A2.1 , two phases of the use are shown.
[0026] In the upper part of Fig. A2.1 , an embodiment of an initial training phase for obtaining an initial magnetic representation of the indoor spatial area is described. Here, for example, a (arbitrary, non-pre-determined) crowd of mobile devices (such e.g. as smartphone) comprising magnetic sensors is used. Initial measurements of the indoor spatial area are taken with robots, whereas other measurement methods are also possible.
[0027] A server distributes map data and marker data (i.e. a spatial representation) to the crowd and receives magnetic data from the crowd and the robots. The server is connected to a machine-learning (ML) backend which trains the deep neural net model.
[0028] The Lower part of Fig. A2.1 shows a continuous retraining phase, i.e. the phase in which real-time magnetic data is collected by at least one mobile device or low performance device moving in the indoor spatial area. The server distributes map and marker data enhanced accordingly to the measuring devices and the returned magnetic measurement data is used to further train the model. With this phase, it is possible to continuously improve and / or augment the model.
[0029] Fig. A2.2 in Appendix 2 shows a close up of the data flow from the data collection to the trained model.
[0030] In the upper third of Fig. A2.2, the units of the measuring devices (e.g. mobile devices) are shown. In the middle of Fig. A2.2, the server is shown. In the lower third of Fig. A2.2, the units of the machine-learning backend are shown.
[0031] The devices use 3-axial magnetic field measurements and device localization data in a locator unit. This generates spatially aware magnetic field data.
[0032] The Trial Divider in Fig. A2.2 is used to generate trials. If a user takes measurements for a Long time (e.g. 2 hours) while moving in the spatial area with a mobile device, it might be useful to divide the measurement results into shorter stretches (e.g. 30 to 50 m). This would later allow an easier anonymization.This data is received by the server which correlates it with previously received reference data (e.g. data measured with a robot or a smartphone).
[0033] The data is cleaned and aligned which can be based on confidence values of the corresponding localization for example. Another important step to align the data is to perform a calibration. As the data is gathered by potentially many different devices and / or at different times there are several reasons why the data can differ between the measurements at the same location even if the true magnetic field stays the same:
[0034] a) Due to magnetic material fixed to the sensor which is mostly the case in practical use cases like in smartphones
[0035] b) Due to hardware or quality differences in the sensor itself
[0036] c) Due to software-based adjustments of the sensor values made by the manufacturer of the sensor, the smartphone or the operating system
[0037] The problem a) is normally solved by an n-poses calibration algorithm where the sensor is turned in at least 6 different poses to determine the calibration as offset and sensitivity values for each axis. However, this is not practical for a crowd-based data gathering especially if it should be done in the background as this requires a user action. So a background calibration is used here where the current measurements are compared with a reference measurement (e.g. from a robot or another smartphone) at specific locations preferably where a marker is visible. At these locations a very precise localization and pose is ensured. Having at least one such measurement allows a precise determination of the offset as the difference between the reference and the trial measurement. Having more reference points also allowsto calculate a sensitivity value. Havingthe offset and / or sensitivity values every measurement in the trial can be calibrated which solves problem a) effectively. Figure 5 shows the effect of this background calibration technique applied to magnetometer readings which are reported as uncalibrated by the smartphone operating system. Figure 6 shows the effect on values which are automatically precalibrated by the smartphone. Using pre-calibrated values can have a positive effect especially in the beginning but the calibration is often Less accurate and causes more problems of type c).
[0038] Furthermore, a calibration might get lost during a longer measurement period, calibrations might be performed for the shorter stretches.Another option is to add the offset and / or sensitivity values to the data vector itself at the locations where they can be calculated (like marker Locations) instead of performing the actual calibration in this step. This data vector will be feed as input to the neural network Later. This allows the network not only to learn and counterpart the calibration from point a), but also to address the issues described in b) and c). The hardware and software measurement differences in b) and c) are mainly specific to device groups, Like a group of smartphone models by a manufacturer or specific to an operating system. This effect is shown in figures 5 and 6 as an example. By providing the network the prior knowledge of the reference measurement as offset and / or sensitivity values, the network can learn which group the device belongs to and how to handle and weight the differences in the measurements.
[0039] If the server has gathered sufficient data (i.e. according to some predetermined threshold) the data is forwarded to the machine-learning backend. As described in Appendix 1, section III, the data can be split up into chunks according to the sequence length of the measurement.
[0040] The chunks can be further processed in a classification mode (left half of the lower third in Fig. A2.2). This involves a fixed number of classes (like e.g. in a grid model). The model is trained to predict a class. As the class is a representation of spatial data, the spatial Location can be determined from that class.
[0041] Alternatively, the processing can take place in regression mode, resulting in the direct prediction of a spatial representation.
[0042] Appendix 3 described details of the deep neural net model, in particular which parts are fixed, and which parts are adaptable (flexible) during training.
[0043] Appendix 4 describes an embodiment of an app to be used in a localization system.APPENDIX 1
[0044] Comparing CNN and LSTM Networks for
[0045] Magnetic Localization of loT Devices and Pedestrian Tracking
[0046] Abstract — When outdoors, nowadays it is common to rely on Global Navigation Satellite Systems (GNSS) based navigation to find a location. However, for indoor environments no common solution exists, as GNSS positioning is not available indoors. While many substitute technologies rely on infrastructure installed in buildings, e.g., beacons, in this paper we use the magnetic field characteristics of buildings as a solution that is available everywhere. In the implementation, a Convolutional Neural Network (CNN) and a Long Short-Term Memory (LSTM) model are used for classification of the characteristics and regression, respectively. The approaches are evaluated against each other on a public dataset showing that the magnetic field can be a robust ubiquitous solution for indoor localization. The regression with LSTM shows the highest precision, while the error of a classification approach is constrained by the building boundaries and enables the usage of class confidence values for further processing.
[0047] Index Terms — indoor positioning, neural network, loT, magnetic field
[0048] I. INTRODUCTION
[0049] The precise localization of persons and devices already plays a significant role in applications. It ranges from locationbased services on smartphones like navigation, tracking of belongings via tags to the essential localization of robots and autonomous cars. However, GNSS signals are the main localization source, which normally require rather unobstructed line of sight to various satellites for a precise localization.
[0050] Especially in indoor environments, the GNSS accuracy is usually very low, so that no location fix can be calculated. This hinders the development and dissemination of applications. In addition to indoor navigation, which is especially important for blind users [1], loT and robot-based applications cannot count on a reliable position indoors in public, company, or health buildings.APPENDIX 1
[0051] Hospitals, for example, benefit greatly from indoor asset tracking as important and expensive mobile equipment often gets lost in the huge buildings. They also need to track contaminated beds and equipment. Mobile emergency buttons and air quality sensors could be used in a much more flexible way if they know their own indoor position.
[0052] However, there are solutions and products to address this, but their distribution is very limited. A main reason for this issue is infrastructure costs. Many solutions rely on radiobased beacons like Bluetooth Low Energy (BLE) [2], [3], WiFi [4] or recently Ultrawideband (UWB) [5]. However, due to the high number of required beacons, installation and maintenance costs, for example, to change batteries, are often a relevant obstacle [6].
[0053] Other systems use natural or artificial optical features in combination with cameras, significantly reducing infrastructure costs [7], [8]. This comes with the drawback that visual contact and also a camera are needed, which is not applicable to smartphones in the pocket or most loT and tracking devices.
[0054] Avery promising approach is localization using Earth’s magnetic field. Depending on the structure of the building, there are large metal parts in the walls and ceilings of the floors. These structures disturb the natural Earth’s magnetic field in combination with other metal structures in the building. Its distribution inside the building forms a specific pattern.
[0055] Even low-cost and energy-saving Micro-Electro-Mechanical Systems (MEMS) magnetometer sensors can measure the deviations. They are built into most smartphones and are easily applicable in loT devices. Compared to the radio-based solutions, besides the missing infrastructure costs, another advantage is that the magnetic field is much more static over time. On the other hand, it has a maximum of three features, as it is only a three-dimensional vector, and the values are not unique. Independent of the localization method, remote tracking of devices normally requires a server infrastructure, which also allows the usage of complex tracking algorithms even for loT devices.
[0056] The same value patterns typically appear at different locations within buildings and floors. Thus, patterns can only be detected overtime, respectively, the traveled distance. This makes magneticAPPENDIX 1
[0057] localization a challenging task, and the algorithms are often only applicable by fusion with other localization techniques. In this paper, we present a novel standalone solution based on CNN and LSTM deep neural network structures adapted to each building. We evaluate the approaches against each other in different buildings.
[0058] The remainder of this paper is structured as follows. In section II, current research in the field of indoor localization is discussed. Section III continues with the preparation of the dataset used, before the method is described in section IV. In section V, our evaluation metrics are introduced, which are used for the presentation of results in section VI. Finally, a summary and an outlook on future work are given in section VII.
[0059] II. RELATED WORK
[0060] A. Magnetic localization
[0061] The majority of the solutions for magnetic indoor localization rely on fingerprinting techniques building up a gridbased map of the whole magnetic vector or extracted features in offline mode first. In localization mode, the classification problem is often solved with traditional algorithms, such as k-nearest neighbors [9], support vector machines (SVMs)
[0010] ,
[0011] , and decision trees
[0012] for the task of identifying a localization by fingerprints
[0013] . They are often used for WiFi fingerprints but are mainly applicable to magnetic field data as well.
[0062] Another possibility is to use probabilistic filters, such as Bayes filters. In
[0014] , the magnetic field is fused with a step detector comprising a gyroscope and an accelerometer in a grid-based Bayes filter showing good results. However, the solution focuses on the one-dimensional case and requires a start position which is not practical for all scenarios. Also, setting up a comprehensive magnetic field fingerprinting map can be quite challenging.
[0063] While traditional methods are less computationally expensive, deep learning approaches have recently gained significant traction and shown promising results
[0013] ,
[0015] ,
[0016] . CNN and Recurrent Neural Network (RNN), especially LSTM architectures, are widely used in the reviewed literature,APPENDIX 1
[0064] sometimes even in combination. These architectures are particularly effective at capturing the spatial and temporal characteristics of sequential data, making them well-suited for this application. For instance,
[0017] ,
[0018] ,
[0019] , and
[0020] all employed variations of RNN respectively LSTM models in their work. The complexity of these models ranges from basic RNNs
[0017] to more sophisticated architectures like the hierarchical LSTM proposed by
[0020] .
[0015] ,
[0016] , and
[0013] utilized CNN-based architectures in their work.
[0065] Most of the reviewed papers directly use magnetic field patterns as input, while
[0016] and
[0021] incorporate the use of Recurrence Plots (RPs) as sequence fingerprints. In
[0013] the use of Temporal Convolutional Networks (TCNs) is introduced, which are specifically designed to handle sequential data by capturing long-range dependencies through a series of convolutions.
[0066] Recent research has also explored the combination of different neural network architectures to leverage their respective strengths. Both
[0022] and
[0023] use more advanced hybrid approaches that integrate CNNs with RNNs or LSTMs to build and process multichannel visual representations of magnetic data. These approaches combine the spatial processing capabilities of CNNs with the sequential learning of RNNs.
[0067] B. Dataset Selection
[0068] As the practical gathering of enough high-quality data to train the models is an arduous task, we focus on public available datasets. This also allows a comparison of our approaches with state-of-the-art algorithms. Due to the significant differences between datasets including the methods of data acquisition, the sensor frequency, and the total amount of data collected, we opted to focus on a single dataset.
[0069] All candidates offer at least 3 axis magnetometer data. The approach in
[0024] combines magnetic and WIFI data in 5 buildings with a recording frequency of 10Hz. In
[0025] Bluetooth data is added, covering about 2000m2on parts of three floors in a building. Additionally, the system in
[0026] covers around 260m2of space at 10Hz as well.APPENDIX 1
[0070] For our training and evaluation, we choose the MagPie dataset as provided in
[0027] . It includes magnetometer data from three differently sized buildings with a total of over 960m2floor space at a 50Hz sampling rate. This is advantageous as it provided multiple large locations with the same level of data quality, eliminating the need for data alignment. The MagPie dataset was collected using a smartphone for measuring the magnetometer data and a Lenovo tablet using the Tango API to collect the ground truth measurements.
[0071] There are separate measurements with the devices either handheld or attached to a robot. As the devices were mounted close to the engines of the robot, we assume corresponding artifacts in the magnetic measurements in the robot data. In order to avoid errors through this kind of artifacts, we only used the handheld data. MagPie already provides over 110 trials of walked trajectories for each building. This makes comparing different approaches easier and eliminates the need to create artificial trajectories. Also, it is close to a practical application of the proposed algorithm where a visual localization system, for example, can record magnetometer data trails of many users during their usage.
[0072] III. DATA PREPARATION
[0073] Like most existing magnetic solutions, we focus on data sequences instead of single data points as the magnetic field value is not unique at each position. The sequence length remains as a parameter for evaluation, as its effect is of high relevance for the performance and for the possible applications as well. Asliding window approach using overlapping windows is implemented (Figure 1). In contrast to consecutive, non-overlapping ones, this avoids losing information between windows, as the target position always corresponds to the last data point in the sequence. In this work we use the whole raw magnetic field vector B = [Bx,By,Bz] accepting that this does not lead to a fully rotation-invariant solution as presented in
[0014] . But at the same time, we do not lose information this way, as the vector direction can also include hints about the walking direction the network can learn.APPENDIX 1
[0074] For the classification approaches, we introduce an additional step of dividing the floor plan into discrete areas to gather the corresponding landmark indices of each point. This is performed using DBSCAN clustering to segment the points and then assigning each group of points to a specific landmark / square ID of the same size. This allows us to convert the positioning problem into a classification task, where each grid cell represents a distinct location class as the target.
[0075] To structure the MagPie dataset for training, we extract the data into CSV format by matching the ground-truth positional coordinates to the magnetometer data based on the timestamps. A trial number column is also added to the dataset to clearly identify and separate individual trials. This allows efficient splitting of the data into training, validation, and testing sets, ensuring that no data points from the same trials appear in multiple sets. This separation is important because the data is recorded at 50Hz, which can lead to consecutive measurements that are very similar. If data points from the same trial appear in both training and validation sets, the model might learn to memorize these points, leading to a low validation error but poor generalization performance in realworld applications. It also helps us to later evaluate the model based on individual trials.
[0076] For classification tasks, it is also important to ensure that all classes inside the validation and test sets have corresponding targets inside the train set. Otherwise, targets that were not part of the model’s training process would need to be predicted.
[0077] The inputs are of shape (# input sequences, sequence length, 3). For the regression case the target is an array of shape (# input sequences, 2) as we predict 2-dimensional coordinates in this work. For the classification case we get the output shape (# input sequences, 1). Both cases are further described in the following section.
[0078] IV. CONCEPT AND MODEL DEFINITION
[0079] Our method does not focus on a general model to solve the localization problem in every building. The magnetic distribution patterns can vary strongly between buildings, depending mainly on their construction method. So, we decided to keep the models as flexible as possible and train themAPPENDIX 1
[0080] individually for each building. We just set some very basic model structures for each approach. For each building, we identified the best model architecture and optimal training parameters by conducting a hyperparameter search using Bayesian optimization on an undersampled training set.
[0081] In our work, we focus on the usage of CNN and LSTM model architectures. In general, CNNs apply filters with different kernel sizes to the input signal detecting patterns in the signal which basically fits our problem. For the CNN variant, our model is formed by different one-dimensional convolution layers as shown in Figure 2. The number of filters, the kernel sizes as well as the number of layers are kept as a parameter for the hyperparameter tuning for each building. However, CNNs are most effective for local patterns where historical data correlations are not that relevant.
[0082] LSTMs on the other side are widely used especially for time series problem tasks. Every LSTM unit forms a memory cell to retain information overtime utilizing an input-gate, an output-gate, and a forget-gate. The architecture allows the network to remember information even over longer time periods. Our LSTM integration is quite similar to the CNN version where the number of layers and the LSTM units are individually set by the Bayesian optimizer for each building. A final dense layer before the output layer is optimized as well for both model types. All models utilize the Adam optimizer. To enhance generalization within a building, dropout layers with optimizable dropout rates and L2 regularization are incorporated across all architectures. A building-specific example model for both architectures is shown in Figure 2.
[0083] We investigated two main approaches to finally predict the 2D positions. Each of the approaches are tested with a CNN and an LSTM model structure.
[0084] A. Classification approach
[0085] The concept of the classification problem approach is based on a Time Series Classification (TSC) task which is used to predict class labels for unlabeled time series of raw datapoints. The key component of TSC is to identify relevant patterns in the input data that allow for furtherAPPENDIX 1
[0086] classification which is applicable to our problem. Hence, the localization space will be divided into discrete areas with landmark indices. The classification models employ sparse categorical crossentropy loss and a softmax activation.
[0087] B. Regression approach
[0088] In the other approach, we treat the prediction as a regression problem. The idea is based on Time Series Regression (TSR), which is used to predict continuous values based on time series data. The method is often employed for forecasting purposes. With Time Series Extrinsic Regression (TSER) as a specific type of TSR, we predict a continuous response variable that is not directly related to the time series being analyzed
[0028] ,
[0029] . Thereby, we directly predict the position vector P= [Px,Py] in our case. The models use a mean squared error loss function, where the output layer consists of two units to directly predict the x and y coordinates.
[0089] V. EVALUATION METRICS
[0090] For regression tasks, we use the Mean Absolute Error (MAE) between the ground truth and the predicted positional coordinates to assess and compare the performance of the final model. The MAE is given by the following formula:
[0091] M
[0092]
[0093] AE = -^1\yi-yi\ (1)
[0094] where y,are the ground truth coordinates and ytare the predicted coordinates.
[0095] For classification tasks, the evaluation is more complex. The typical metric is accuracy, which indicates the percentage of cases in which the model correctly predicts the landmark index. However, this metric has limitations:
[0096] 1) It does not really help with the comparison between models trained on different grid square sizes.
[0097] 2) It does not provide information on the distances between predicted and actual points when the model misclassifies a point, making it difficult to compare with the regression case.APPENDIX 1
[0098] To address these issues, we introduced a distance-based metric using an adjacency matrix that holds the distances between landmark IDs. This metric computes the Euclidean distance between the centroids of predicted and actual landmarks, allowing us to basically mimic the mean absolute error between the centroids of landmarks. All distances are reported in meters.
[0099] VI. RESULTS
[0100] A. Comparing LSTM and CNN in Regression Mode
[0101] We first examined the performance difference between the LSTM and CNN architectures using the same sequence lengths to have a basis for comparing both approaches. As shown in I for a sequence length of 50 samples, the LSTM architecture achieved an MAE of 3.12 m, whereas the CNN had an MAE of 4.41 m. When further increasing the sequence length to 100 and 200, the LSTM approach starts to significantly outperform the CNN by even reducing the MAE to 0.82 m compared to 2.31 m for a sequence length of 200. This can be seen even better when visually comparing the prediction performance, as shown in Figure 3, the LSTM model predicts the trajectory with much higher accuracy, closely following the ground truth. In contrast, the CNN model shows considerable fluctuations and deviations from the actual path, even though both models receive exactly the same information.
[0102] In other evaluation runs, we especially focused on the sequence lengths itself. While the error goes down when increasing the context from 50 the way up to 200 samples, beyond this, the performance started to decline, with the Mean Absolute Error (MAE) increasing slightly for longer sequences (e.g., 0.9 m at 250, 0.96 m at 300, and 1.07 m at 500) in the Talbot building. These results suggest that shorter sequences provided insufficient context for accurate position prediction, leading to higher errors. Longer sequences can make the predictions more accurate, but also introduce challenges that can lead to an increase in error. This likely comes from the fact that longer sequences introduce much more possible trajectories that could lead to the same point, which requires diverse training data that includes all those possible trajectories.APPENDIX 1
[0103] B. Comparing LSTM and CNN in Classification Mode
[0104] Unlike the regression case, there is no clear better architecture when comparing CNN and LSTM approaches. Their performance is nearly equal across different configurations. Interestingly, while CNNs seem to outperform LSTM for lower sequence lengths, LSTMs are slightly superior for larger sequence lengths. This trend is visible in Table II. Additionally, the grid size plays a role in the TSC case where smaller grid sizes lead to a worse class-accuracy but also a lower MAE, which is more important in this case. The contradiction is due to the trade-off between classification accuracy and spatial precision. With larger grid sizes, the model more often predicts the correct square, increasing accuracy. However, mistakes lead to bigger spatial errors. For example, with a 2-meter grid, a wrong prediction in the neighboring cell automatically causes a 2 m error as shown in Figure 4. On the other hand, smaller grid sizes, do not achieve particularly high accuracies.
[0105] Though, when predictions are incorrect, most of the time they are spatially closer to the true position, resulting in lower overall errors.
[0106] However, it is important to note that blindly decreasing the grid size is not a viable strategy, as besides the computational increase, the amount of available training data in each grid cell decreases. So, similar to the sequence length described in subsection Vl-A, the amount of data limits the minimal grid size and so effects the accuracy respectively the MAE. Regarding the given data for the Talbot building, we found an optimal grid size of 0.33 m with a sequence length of 200 samples leading to an accuracy of 0.21 and a very good MAE of 1.04 m.
[0107] Comparing the TSC and TSER, the regression approach generally outperformed the classification method in terms of positioning accuracy and precision. Our TSER method consistently achieved a lower MAE across various configurations. The regression approach offers a more straightforward implementation, as it does not require dividing the building into distinct squares. This simplicity extends to benchmarking, as there are fewer possible configurations, and the results are directly comparable via MAE. However, the classification approach has a unique advantage: Predictions are always constrained within the building’s boundaries, which is not guaranteed with regression.APPENDIX 1
[0108] Additionally, the TSC approach is able to output class confidences, which can be a useful information if the results will be fused with other location sources in a later processing stage. For both approaches, the sequence length is a crucial parameter whose optimal value can vary across buildings. Most of the time, the optimal performance was achieved with sequence lengths of around 200 datapoints for the CSL and Talbot buildings, which relates to 4s with a sensor frequency at 50Hz. However, larger buildings or datasets with more extended trajectories might benefit from even longer sequences, as seen in the Loomis building results, where performance continued to improve even up to a sequence length of 500.
[0109] D. Final Results and Comparison
[0110] As the magnetic localization is a not a classical machine learning research field and its performance depends on the data respectively the buildings, a comparison between solutions of different researchers is quite challenging. The differences in the rare public datasets and the fact that many state-of-the-art solutions use their own datasets makes it even more difficult. However,
[0022] compared their different approaches on the MagPie dataset. They also compared regression vs. classification approaches using a combination of CNN and Fully Connected Networks (FN) as well as Recurrent Neural Networks (RNN). Be aware that it stays unclear whether they also used the provided test dataset to evaluate their models or if they conducted their own test splits. Table III provides the comparison of the different approaches outlined in the paper against our approach. Note that
[0022] did not mention which metric they exactly used in their reported results. Thus, the comparison should therefore rather be regarded as an indication.
[0111] Comparably, our results show better overall performance for the Talbot and CSL buildings for both regression and classification, but worse overall performance for the Loomis building. It needs to be mentioned that
[0022] achieved the best results using RNN which require a starting point. Although they added noise of 2-3 meters to the starting point, the results are not fully comparable for the RNN version with our more challenging lost robot problem approach. On the other hand, as
[0022] differs in their approach and especially their RNN solution performs on the same level, we can assume that it might also perform better in other buildings. Despite this, the discrepancy for the Loomis building could also be due to their use of longer sequences (about 7 - 10 seconds) whichAPPENDIX 1
[0112] was not in our scope. If we just compare both approaches regarding the lost robot problem having no start position even our landmarkbased solution outperformed their CNN+FN approach in all three buildings. It is important to note that for our results in this section, we always report the best overall outcomes, independent of architecture, sequence length, and grid size.
[0113] VII. CONCLUSION
[0114] In this paper, we presented an approach based on Deep Learning for indoor localization, without the need of additional infrastructure. To this end, the specific magnetic field characteristics of buildings are exploited. For this purpose, the magnetic field inside a building is measured along with ground truth position data forming a set of single trails. For each building, we trained two different types of neural networks, a CNN and a LSTM, for two different localization tasks each. The first approach treats the localization task as a classification problem, while the other approach formulates the problem as a regression task.
[0115] We have then optimized the network architecture and model parameters for both network types and localization tasks on a per-building-level. The localization accuracy is thoroughly evaluated in terms of key parameters, network type, and task formulation. While treating the localization task as a
[0116] classification problem does not reach the high precision of the regression task, it has the advantage of constraining predictions within building boundaries. Our optimized models reach similar or even better results than the much more complex models described in
[0022] .
[0117] To enable a practical implementation, we plan to validate our findings on larger datasets that also include large vertical displacements (such as floor changes) and a wider range of building types and user movements. In the next step, the localization could then be integrated in a mobile indoor navigation app, such as everGuide [1] or an loT-asset tracking use-case. If required by the use-case, the precision could be increased by spatial context between individual predictions or by leveraging class probabilities of all landmark indices, instead of relying just on the best one. Additionally, more network architectures or modules used during the hyperparameter tuning, are worth evaluating in the future.APPENDIX 1
[0118] REFERENCES
[0119] in J. Wortmann, B. Schaeufele, K. Klipp, I. Radusch, K. Blali, and T. Jung, “Enhanced Accessibility for Mobile Indoor Navigation,” in 2024 14th International Conference on Indoor Positioning and Indoor Navigation (IPIN). IEEE, 2024.
[0120] [2] F. J. Aranda, F. Parralejo, T. Aguilera, F. J. Alvarez, and J. Torres-' Sospedra, “RSS Channel- Based Integration for BLE Fingerprinting Positioning,” in 2023 13th International Conference on Indoor Positioning and Indoor Navigation (IPIN), 2023, pp. 1-6.
[0121] [3] L. Bouse, S. A. King, and T. Chu, “Simplified Indoor Localization Using Bluetooth Beacons and Received Signal Strength Fingerprinting with Smartwatch,” Sensors, vol. 24, no. 7, 2024.
[0122] [Online]. Available: https: / / www.mdpi.eom / 1424-8220 / 24 / 7 / 2088
[0123] [4] J. Hu and C. Hu, “A WiFi Indoor Location Tracking Algorithm Based on Improved Weighted K Nearest Neighbors and Kalman Filter,” IEEE Access, vol. 11, pp. 32907-32918, 2023.
[0124] [5] J. Qu, “A review of UWB indoor positioning,” Journal of Physics: Conference Series, vol. 2669, no. 1, p. 012003, dec 2023. [Online]. Available: https: / / dx.doi.org / 10.1088 / 1742- 6596 / 2669 / 1 / 012003
[0125] [6] R. S. Naser, M. C. Lam, F. Qamar, and B. B. Zaidan, “SmartphoneBased Indoor Localization Systems: A Systematic Literature Review,” Electronics, vol. 12, no. 8, 2023. [Online]. Available: https : II\N\N\N. md pi .com / 2079-9292 / 12 / 8 / 1814
[0126] m Q. Zeng, B. Ou, R. Wang, H. Yu, J. Yu, and Y. Hu, “A Robust Indoor
[0127] Localization Method Based on DAT-SLAM and Template Matching Visual Odometry,” IEEE Sensors Journal, vol. 23, no. 8, pp. 8789-8796, 2023.
[0128] [8] F. Gao and J. Ma, “Indoor Location Technology with High Accuracy Using Simple Visual Tags,” Sensors, vol. 23, no. 3, 2023. [Online]. Available: https: / / www.mdpi.eom / 1424-8220 / 23 / 3 / 1597 [9] D. Li, B. Zhang, and C. Li, “A Feature-Scaling-Based k-Nearest Neighbor Algorithm for Indoor Positioning Systems,” IEEE Internet of Things Journal, vol. 3, no. 4, pp. 590-597, 2016. no] A. S. Abdou, M. A. Aziem, and A. Aboshosha, “An efficient indoor localization system based on Affinity Propagation and Support Vector Regression,” in 2016 Sixth International Conference on Digital Information Processing and Communications (ICDIPC), 2016, pp. 1-7.APPENDIX 1
[0129] [in A. Chriki, H. Touati, and H. Snoussi, “SVM-based indoor localization in Wireless Sensor Networks,” in 2017 13th International Wireless Communications and Mobile Computing Conference (IWCMC), 2017, pp. 1144-1149.
[0130]
[0012] J. Yim, “Introducing a decision tree-based indoor positioning technique,” Expert Syst. AppL, vol. 34, no. 2, p. 1296-1302, Feb. 2008. [Online]. Available: https: / / doi.Org / 10.1016 / j.eswa.2006.12.028
[0131] H3] G. Ouyang, K. Abed-Meraim, and Z. Ouyang, “Magnetic-field-based indoor positioning using temporal convolutional networks,” Sensors, vol. 23, no. 3, p. 1514, 2023.
[0132] H4] K. Klipp, H. Rose, J. Willaredt, O. Sawade, and I. Radusch, “Rotation-' invariant magnetic features for inertial indoor-localization," in 2018 International Conference on Indoor Positioning and Indoor Navigation (IPIN). IEEE, 2018, pp. 1-10.
[0133] ns] I. Ashraf, M. Kang, S. Hur, and Y. Park, “MINLOC: Magnetic field patterns-based indoor localization using convolutional neural networks,” lEEEAccess, vol. 8, pp. 66213-66227, 2020.
[0134]
[0016] M. Abid, P. Compagnon, and G. Lefebvre, “Improved CNN-based magnetic indoor positioning system using attention mechanism,” in 2021 International Conference on Indoor Positioning and Indoor Navigation (IPIN). IEEE, 2021, pp. 1-8.
[0135] in] H. J. Jang, J. M. Shin, and L. Choi, “Geomagnetic field based indoor localization using recurrent neural networks,” in GLOBECOM 2017- 2017 IEEE Global Communications Conference. IEEE, 2017, pp. 1-6.
[0136] ns] H. J. Bae and L. Choi, “Large-scale indoor positioning using geomagnetic field with deep neural networks,” in ICC 2019-2019 IEEE International Conference on Communications (ICC). IEEE, 2019, pp. 1-6.
[0137] [is] B. Bhattarai, R. K. Yadav, H.-S. Gang, and J.-Y. Pyun, “Geomagnetic field based indoor landmark classification using deep learning,” lEEEAccess, vol. 7, pp. 33943-33956, 2019.
[0138]
[0020] L. Wang, H. Luo, Q. Wang, W. Shao, and F. Zhao, “A hierarchical LSTM-based indoor geomagnetic localization algorithm,” IEEE Sensors Journal, vol. 22, no. 2, pp. 1227-1237, 2021.
[0139]
[0021] N. Lee, S. Ahn, and D. Han, “AMID: Accurate magnetic indoor localization using deep learning,” Sensors, vol. 18, no. 5, p. 1598, 2018.APPENDIX 1
[0140]
[0022] L. Antsfeld and B. Chidlovskii, ‘‘Magnetic field sensing for pedestrian and robot indoor positioning,” in 2021 International Conference on Indoor Positioning and Indoor Navigation (IPIN). IEEE, 2021, pp. 1-8.
[0141]
[0023] L. Fernandes, S. Santos, M. Barandas, D. Folgado, R. Leonardo, R. Santos, A. Carreiro, and H. Gamboa, “An infrastructure-free magnetic-based indoor positioning system with deep learning,” Sensors, vol. 20, no. 22, p. 6664, 2020.
[0142]
[0024] I. Ashraf, S. Din, M. U. Ali, S. Hur, Y. B. Zikria, and Y. Park, “MagWi: Benchmark Dataset for Long Term Magnetic Field and Wi-Fi Data
[0143] Involving Heterogeneous Smartphones, Multiple Orientations, Spatial Diversity and Multi-Floor Buildings,” IEEE Access, vol. 9, pp. 77976- 77996, 2021.
[0144]
[0025] Z. Toth and J. Tam' as, “Miskolc IIS hybrid IPS: Dataset for hybrid indoor' positioning,” in 2016 26th International Conference Radioelektronika (RADIOELEKTRONIKA), 2016, pp. 408-412.
[0145]
[0026] J. Torres-Sospedra, D. Rambla, R. Montoliu, O. Belmonte, and J. Huerta, “UJIIndoorLoc-Mag:
[0146] A new database for magnetic field-based localization problems,” in 2015 International conference on indoor positioning and indoor navigation (IPIN). IEEE, 2015, pp. 1-10.
[0147]
[0027] D. Hanley, A. B. Faustino, S. D. Zelman, D. A. Degenhardt, and T. Bretl, “MagPIE: A dataset for indoor positioning with magnetic anomalies,” in 2017 International Conference on Indoor Positioning and Indoor Navigation (IPIN). IEEE, 2017, pp. 1-8.
[0148]
[0028] D. Guijo-Rubio, M. Middlehurst, G. Arcencio, D. F. Silva, and A. Bagnall, “Unsupervised feature based algorithms for time series extrinsic regression,” Data Mining and Knowledge Discovery, pp. 1-45, 2024.
[0149]
[0029] N. Mohammadi Foumani, L. Miller, C. W. Tan, G. I. Webb, G. Forestier, and M. Salehi, “Deep learning for time series classification and extrinsic regression: A current survey,” ACM Computing Surveys, vol. 56, no. 9, pp. 1-45, 2024.
[0150]
[0030] P. Gavrikov, “visualkeras,” https: / / github.com / paulgavrikov / visualkeras, 2020.Appendix 2Data description
[0151] A Marker is e.g. an optical reference point within a building, whose detection in the camera image enables highly accurate localization. Acoustic markers or radio beacons can also be used as markers. The measurement device must have a sensor for the respective signal.
[0152] Map data refers to information from a high definition map of the building, in which the users try to localize themselves or a device is localized.
[0153] Magnetic field measurement refers to a set of measurements of the magnetic field strength along a sensor's X, Y and Z axes. Typically, such measurements will be taken using Low-cost MEMS magnetometers. Each individual measurement is associated with a timestamp whose time base is shared by all measurements but arbitrarily chosen per measuring device and is allowed to change after rebooting the measuring device. Measuring the magnetic field strength can be done using previously calibrated sensors or uncalibrated, which is noted in the respective measurement.
[0154] A spatial aware magnetic field measurement is a magnetic field measurement in combination with estimated positions and orientations of the measurement device. Like the magnetic field measurement, all estimated 2D or 3D poses (position and orientation) of the device include timestamps that share the same time base but may differ across devices and reboot. In addition, the time base of the magnetic field measurements and the one of the poses is the same.
[0155] All estimated poses are relative to the reference coordinate system of the high definition map of the building.
[0156] All estimated poses include an accuracy estimation, which allows to distinguish between high -precision pose estimations (usually in presence of a marker) and less precise estimations (usually in-between markers).
[0157] Inferred 2D pose refers to a timestamped estimate of the position and orientation of a device within a building. The position consists of a 2D position in relation to the coordinate systemdefined by the high definition map, along with floor information. It additionally includes a confidence indication that assesses the accuracy of the estimate. The orientation of the device describes the 1 D orientation in the plane of the floor.
[0158] Device description
[0159] Mobile Devices are devices with sufficient computational power to run the learned magnetic field neural networks for inference in terms of edge Al. An example of such a device are smartphones on which e.g. a localization app (such e.g. as everGuide, see Appendix 4) is running. These devices may come from different manufacturers and potentially run different operating systems.
[0160] Low Performance Devices are devices that do not have enough computing power to execute the learned neural networks for inference. However, they have a communication unit to exchange data with a server that executes the learned models on behalf of the devices. An example of such a device is an loT Asset Tracker.
[0161] A server refers to one or more devices that are located outside the building and outside the control of the users. The server defines interfaces for communication with other devices (mobile devices and low performance devices). The server runs the machine Learning (ML) backend, which is responsible fortraining, optimizing, and evaluating the magnetic field neural networks.
[0162] Process
[0163] The embodiment of a process is performed in an indoor spatial area, per building, floor or other spatial boundary, whatever makes the most sense for the specific use case. For ease of reading in the following section, we use the word “building” synonymously for all the previously enumerated cases.
[0164] The solution builds on the existence of a high accuracy localization system like the everGuide system which is a marker-based localization method used for indoor navigation in equipped buildings.Initially, the building is measured once with a dedicated mapping hardware (e.g. a robot). During this process, a highly accurate map is created in which previously installed optical anchors are precisely marked.
[0165] Additionally, an initial “spatial aware magnetic field measurement” is conducted by the mapping hardware. Because of the small number of data points and potential interferences by the mapping hardware itself, this initial measurement is not enough to train a magnetic field neural network.
[0166] Note that for buildings already equipped with the everGuide or a similar localization system, this initial step does not have to be repeated, because it has been done in the past already. After the initial measurement, the localization system is operational without the localization based on magnetic field neural networks. Users of the system can achieve high-precision localization at optical anchor points using the camera integrated into their devices. Between those anchor points, localization with reduced precision is possible as well. Now, the users of the localization system e.g. continuously collect magnetic field data through the MEMS magnetometer integrated into their devices, this is a form or real time magnetic data measurement. Together with the corresponding timestamped three-dimensional poses, this data is sent anonymously as “spatial aware magnetic field measurement” to the localizsation servers (e.g. everGuide servers). On these servers, that measurements are further processed and stored.
[0167] In this case, processing includes grouping data by building, floor, etc. and then splitting them into individual trials align pose with magnetic field measurements and perform re-calibration of the magnetic field data.
[0168] The re-calibration might be necessary as magnetometers often change their calibration values (sense and bias per axis) due to external magnetic influences. Sensor calibrations normally need a user interaction to turn the sensor in different directions at the same place which is often not achievable. Having a magnetometer measurement together with a complete pose estimation with a very high reported accuracy, we look for a reference measurement (e.g. from the initial robot measurement) at the same position. By rotating the vectors in the same direction, the difference of each axis between the reference measurement and the currentmeasurement is the offset value. We simplify the problem and assume the sense to be 1. The calibration values are calculated and applied to each trial individually.
[0169] In case the users’ devices have the necessary computing capacity, processing steps can also be executed on the devices themselves, before uploading the data to the server. The server then only needs to verify these results.
[0170] By combining a large volume of recorded data (crowdsourcing approach), machine learning models can be trained and optimized in the server's ML backend (see model details and accompanying paper). The quality of the trained models is automatically assessed. As soon as the quality is deemed sufficient, the operating phase follows (continuous-retrain-phase). In the operating phase, it is now possible to estimate the position of devices using only measured magnetic field data. This facilitates the integration of low-performance devices (e.g. asset trackers) that lack the necessary hardware, visual contact or computing power for localization through the localization system. To localize such a device, the device needs to be equipped with a magnetometer to perform a “ Magnetic field measurement”. This measurement is then sent to the server, which can subsequently estimate the device's position, based on the trained magnetic neural network models.
[0171] In addition, these trained models can be distributed by the server to all mobile devices, which are using the localization system such as e.g. the everGuide system. Since these devices are, by definition, powerful enough to run these models themselves, they can use the models to improve their pose estimation. This is especially relevant between optical anchor points. Just as in the initial phase, users continue to send their “spatial aware magnetic field measurement” to the server, where it is used to retrain the models. To prevent an inaccurate position estimation from the model on the user's device from degrading the subsequent training results, filtering and accuracy estimation are performed on the device side.Advantages of using the users of the localization system (e.g. the everGuide navigation system) to generate training data (Crowd-approach):
[0172] - Naturally moving users: Users behave naturally while moving across the building, which avoids (unintended) bias by people in controlled and hence artificial ground truth recordings. - Large variety of different magnetometer sensors, because of very large device base (which is very hard to get in controlled ground truth recording.
[0173] - Less expensive than controlled truth recordings because users are using the system anyway. -Very high precision at the optical anchor points, which allows calibrating the magnetometers. - Knowing the sensors orientation, not only position, which allows the use of magnetic field features that are not rotation invariant.
[0174] - Absolute coordinates (shared reference frame) of all users, which simplifies data aggregation across different trials and devices.
[0175] Challenges:
[0176] - Potentially more noise in the data compared to controlled ground truth recordings
[0177] - Due to the nature of the everGuide system, the sensor frequency of the magnetometer and the estimated positions may differ significantly.Appendix 3
[0178] Model Definition
[0179] Definition which parts of the models are considered as fixed parts, and which are subject to changes (flexible parts).
[0180] As the earth magnetic field distribution patterns of buildings or parts of buildings (like floors) can be very different due to their building structure or interieur, we adjust the model structure for each of these separate areas.
[0181] Every model, independent of regression or classification and CNN- or LSTM-based, is formed from a fixed part and a basic definition and a flexible part. Most parts of the model are flexible which means that their structure can adjust to the specific buildi ng structure. This is done using a form of an extensive hyperparameter tuning using a Bayesian optimization algorithm.
[0182] 1. Building the specific CNN Model structure
[0183] Input Layer -fixed
[0184] Shape: (sequencejength, numjeatures)
[0185] The sequence length is typically set between 50 and 300 samples which corresponds to 1-6 seconds assuming a 50 Hz sensor update rate and a normal walking speed. The optimal value of the sequence length can be optimized as it depends on the building structure, the amount of available data and the area size.
[0186] Convl D Layer - fixed but flexible parameters
[0187] Amount of convolution filters: flexible
[0188] Kernel size: flexible
[0189] Activation: 'relu'
[0190] kernel_regularizer: 12(0.001)
[0191] padding: 'same'
[0192] Batch normalization - fixedDropout -fixed
[0193] Dropout value: flexible
[0194] Start optimization-block [
[0195] Number of optimization block layers: flexible
[0196] ConvID Layer -fixed position in optimization-block but flexible parameters Amount of convolution filters: flexible
[0197] Kernel size: flexible
[0198] Activation: 'relu'
[0199] Kernel regularizer: 12(0.001)
[0200] Padding: 'same'
[0201] Batch normalization Layer -fixed position in optimization-block
[0202] Dropout Layer-fixed position in optimization-block but flexible parameter Dropout value: flexible
[0203] ] End optimization-block
[0204] Flatten Layer - fixed
[0205] Dense Layer - fixed position but flexible parameter
[0206] Units: flexible
[0207] Activation: 'relu'
[0208] Kernel regularizer: 12(0.001)
[0209] Batch normalization - fixed
[0210] Dropout -fixed
[0211] Dropout value: flexibleOutput-Dense Layer
[0212] Output vector length:
[0213] 2 -for regression 2d coordinate estimation
[0214] 3 - for regression 3d coordinate estimation
[0215] Number of classes / grid cells - for classification approach
[0216] 2. Building the specific LSTM Model structure
[0217] Input Layer -fixed
[0218] Shape: (sequencejength, numjeatures)
[0219] The sequence length is typically set between 50 and 300 samples which corresponds to 1-6 seconds assuming a 50 Hz sensor update rate and a normal walking speed. The optimal value of the sequence length can be optimized as it depends on the building structure, the amount of available data, the area size and the requirements on how fast a result is needed in real time applications.
[0220] LSTM Layer -fixed but flexible parameters
[0221] Amount of LSTM Units: flexible
[0222] Kernel regularizer: 12(0.001)
[0223] Batch normalization - fixed
[0224] Dropout -fixed
[0225] Dropout value: flexible
[0226] Start optimization-block [
[0227] Number of optimization block layers: flexible
[0228] LSTM Layer -fixed position in optimization-block but flexible parameters Amount of LSTM Units: flexibleKernel regularizer: 12(0.001)
[0229] Batch normalization - fixed position in optimization-block
[0230] Dropout -fixed position in optimization-block but flexible parameter Dropout value: flexible
[0231] ] End optimization-block
[0232] Batch normalization - fixed
[0233] Dropout -fixed
[0234] Dropout value: flexible
[0235] Output-Dense Layer
[0236] Output vector length:
[0237] 2 -for regression 2d coordinate estimation
[0238] 3 - for regression 3d coordinate estimation
[0239] number of classes / grid cells - for classification approachAppendix 4
[0240] High-precision indoor-navigation
[0241] Normally when a visitor enters a building, signage serves for orientation. In larger buildings howeverthese signs are so complex, that they can cost time and test the nerves of the visitor, and sometimes the visitor is so overburdened that he does not reach the desired destination.
[0242] Orientation usingthe smartphone
[0243] Conventional digital navigation on the smartphone as we know it from regular traffic does not function inside a building due to the weak GPS signals. Here our digital indoor-navigation-system helps out.
[0244] The requirements are high: above all it has to be precise in order to guide everyone reliably to their destination. Furthermore it has to be inexpensive to purchase and maintain. Hitherto beacon or WLAN based systems did not meet these requirements. They are expensive to install, require constant adjustments and are far too inaccurate.
[0245] The solution developed by Fraunhofer FOKUS offers an accuracy of less than one meter with minimal installation and maintenance costs. With this accuracy you can let yourself be navigated with closed eyes through doorframes right upto your work station.
[0246] High flexibility
[0247] The system is set up modularly so that it can be flexibly adapted to your requirements. Optical anchor points (modern signage), on request in your corporate design will be affixed to the ceiling. The camera of the smartphone automatically recognizes these on entering the building. This recognition runs totally in the background. No action by the user is required for the localization.Through the sensors in the smartphone, the system can also calculate a precise position when an optical anchor point is not determined.
[0248] Incase only a few markers can be affixed and you still require a highly precise navigation, we offer a foot sensor, that can calculate the exact position of the user at any given point in time. The sensor could for example be deposited with the doorman or at the check-out.
[0249] Maintenance-free infrastructure in your building
[0250] Through the high accuracy and the ability to point in the exact direction, our indoor-navigation-system can reliably direct your customers in your building. It is so precise that in field-trails blind persons could be directed safely. The indoor-navigation-system operates offline on every modern smartphone.
[0251] For you as building operator, only minimal costs are incurred in acquiring the system. Our localizing-infrastructure is maintenance-free after the initial installation and requires no further investment. Updates of room numbers for example or interesting navigation destinations can be carried out independently by you in the onsite database. Changes to the localizing-infrastructure are not required for this.
[0252] In case you already use an App in your building, our indoor-navigation-system can be integrated via SDK without any problem. In addition our indoor-navigation-app is available free of charge to you.
[0253] The advantages of the system at a glance
[0254] - highly accurate location
[0255] -inexpensive
[0256] -maintenance-free- includes accurate heading advice
[0257] -offline-capable
[0258] -robust and reliable
[0259] - barrier-free for blind and visually impaired persons
[0260] - can be integrated into every app and software environment
[0261] Extremely versatile
[0262] We offer an indoor-navigations-system which enables smart-phone users to exactly locate themselves in your building be it a hospital, railway station, airport, government agency or a shopping center or a large company complex.
[0263] The application areas go far beyond indoor navigation. All location-based services are feasible in your building:
[0264] - analysis of visitor streams
[0265] - optimization of routes
[0266] -site-specific contextual information
[0267] - barrier-free access for blind and visually impaired persons
[0268] - active asset tracking
[0269] - locating / heading to “Points of Interest”
[0270] - curated guided tours
Claims
Claims1. Training method for obtaining a localization model using measured field data with a) generating a spatial representation (2D, 3D) of an indoor spatial area using markers and / or comprising markers in the indoor spatial area, in particular the markers being optical anchors, optical markers, acoustic markers, radio beacons and / or ranging radar / LIDAR data,b) subsequently measuring field data within the indoor spatial area using at least one field data sensor,c) the field data obtained in step b) is augmented with positional data from the spatial representation previously obtained in step a) and / ortemporaldata, to obtain afield data spatial representation,d) computing and / or training a deep neural net model using the field data spatial representations generated in step c).
2. Training method according to claim 1 , where in the field data comprises magnetic data and / or radiation data.
3. Training method according to claim 1 or 2 , wherein there is a post-training, re-training and / or an adaption of the deep neural net model with further data obtained in step c), in particular further data which has not been used in the generation of a previous version of the deep neural net model.
4. Training method according to at least one of the preceding claims, wherein the field data in step b) is measured through a crowd-measurement and / or an on-the-fly measurement by non-specialized user equipment, in particular mobile devices using a navigation app or localization app.
5. Training method according to at least one of the preceding claims, wherein the deep neural net model obtained in step d) is further trained at a later stagewith field data spatial representations obtained according to step b) with temporal annotations of the field data,andaugmented with positional data obtained from the deep neural net model itself or from the deep neural model itself and the positional data obtained in step a).
6. Training method according to at least one of the preceding claims, wherein magnetic field data is measured with a magnetic sensor device, in particular with a MEMS.
7. Training method according to at least one of the preceding claims, wherein the spatial representation obtained in step a) comprises magnetic field data.
8. Training method according to at least one of the preceding claims, wherein the field data spatial representations are used to form a map.
9. Training method according to at least one of the preceding claims, wherein the field data spatial representations comprises trails or trials of field data measurements, in particular a sequence of temporally ordered representation of field data spatial representations.
10. Training method according to at Least one of the preceding claims, wherein a recalibration is performed to exclude or minimize the effect of magnetic material and / or radioactive material in the vicinity of the measurement device using at least one magnetic data measurement as a reference measurement, so that an offset from the measured magnetic data to that reference measurement is computed, in particular by taking reference measurements at defined locations, such as e.g. at marker locations to determine the effect of the magnetic material.
11. Training method according to at least one of the preceding claims, wherein a recalibration is performed to exclude or minimize the effect of magnetic material within the measurement device or the non-location-specific effect of magnetic material in the vicinity of the measurement sensor by taking reference measurements at defined locations, such as e.g. at marker locations to determine the effect of the magnetic material.
12. Training method according to at least one of the preceding claims, wherein a smartphone is used as a data sensor comprising a software controlling the magnetometer and its output and reference measurements are taken at defined Locations, such as e.g. at marker Locations to determine an offset and / or a sensitivitycoefficient in the data caused by the software and the reference measurements are used to counteract the offset and / or the sensitivity coefficient.
13. Training method according to at least one of the preceding claims, wherein the deep neural net model is a CNN or LSTM, with fixed and flexible parts where the model is automatically defined in a pre-training phase on a subset of data per indoor spatial area (e.g. a building, floor).
14. Training method according to at least one of the preceding claims, wherein a smartphone is used as a data sensor comprising a software controlling the magnetometer and its output and reference measurements are taken at defined locations, such as e.g. at marker locations to determine an offset and / or a sensitivity coefficient in the data caused by the device specific software and / or hardware and the offset and / or the sensitivity coefficient are used as input in the training, in particular so that the network can learn to distinguish between smartphone device groups or manufacturers.
15. Localization method for a device and / or a person in an indoor spatial area using the deep neural net model obtained by the training method of at least one of the claims 1 to 14, using a device providing at least - in particular only - measured field data as input data.
16. Localization method of claim 15, wherein the device and / or persons are movable within the indoor spatial area while taking field data measurements.
17. Localization method to claim 15 or 16, wherein a smartphone is used as a data sensor comprising a software controlling the magnetometer and its output, and wherein reference measurements are taken at defined locations, such as e.g. marker locations to determine an offset and / or a sensitivity coefficient in the data caused by the software and / or hardware and the reference measurements are used to counteract the offset and / or the sensitivity coefficient during the localization, in particular so that the network can distinguish between smartphone device groups or manufacturers, in particular to counteract static or dynamic sensor value differences caused by the software and / or hardware during the localization on the whole localization area.
18. Localization device configured to be used in connection with the localization method in at least one of the claims 15 to 17, the device being able to perform at least field data measurements, in particular magnetic field measurement, in the indoor spatial area.
19. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the steps of the method according to at least one of the claims 1 to 11 for training the deep neural net model or according to at Least one of the claims 15 to 19 for localizing a person and / or a device.