Indoor localization system
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE UNIV OF SYDNEY
- Filing Date
- 2025-05-21
- Publication Date
- 2026-08-06
Smart Images

Figure AU2025050531_06082026_PF_FP_ABST
Abstract
Description
" Indoor localization system"Cross-Reference to Related Applications
[0001] The present application claims priority from Australian Provisional Patent Application No 2025900267 filed on 31 January 2025, the contents of which are incorporated herein by reference in their entirety.Technical Field
[0002] This disclosure relates to locating a device in an indoor environment.Background
[0003] The widespread adoption of the global navigation satellite system (GNSS) has revolutionized localization and navigation services, enabling the development of various applications in vertical industries and everyday life. Although the GNSS can provide real-time localization with meter-level accuracy in most outdoor environments, it may be not applicable in indoor environments due to signal blockage. Nevertheless, there has been an increasing demand for indoor localization in different application scenarios, such as smart cities, intelligent plants, and Internet of Things (IoT). As an example, mobile users may quickly navigate within a specific indoor environment (such as a store or a hospital) and get to the desired location within the indoor localization system.
[0004] However, indoor localization is difficult as line-of- sight (LoS) paths may not be available between a device and a receiver. For example, a device and a receiver may communicate via radio signals, but these radio signals may be attenuated by walls within the indoor environment. Some indoor localization methods rely on resource-intensive and timeconsuming measurements. Further, they may raise potential issues on privacy, suffer from accumulated errors, may require additional hardware to process the signals, and may be highly sensitive to environmental changes within the indoor environment. In particular, some indoor localization methods struggle in new or unseen environments, which limits their generalisation ability.
[0005] Any discussion of documents, acts, materials, devices, articles, or the like which has been included in the present specification is not to be taken as an admission that any or all of these matters form part of the prior art base or were common general knowledge in the field relevant to the present disclosure as it existed before the priority date of each of the appended claims.
[0006] Throughout this specification, the word “comprise”, or variations such as “comprises” or “comprising”, will be understood to imply the inclusion of a stated element, integer or step, or group of elements, integers, or steps, but not the exclusion of any other element, integer or step, or group of elements, integers or steps.Summary
[0007] Disclosed herein are a system and methods for generally locating a device within an indoor environment. Embodiments of the system and methods may involve the use of a floor plan of the indoor environment to increase the localization accuracy. Further, some embodiments of the system and methods may utilise machine learning to increase the localization accuracy.
[0008] According to the present disclosure, there is provided a method for locating a device in an indoor environment, the method comprising:receiving distance measurements between the device and multiple access points within the indoor environment;determining an estimated location of the device from the distance measurements; for one or more of the multiple access points:generating a feature image from a floor plan of the indoor environment based on the estimated location of the device, the feature image corresponding to a physical region between the device and the access point; andapplying a trained machine learning model to the distance measurement between the device and the access point and the feature image to calculate an adjusted distance between the device and the access point; anddetermining an adjusted location of the device in the indoor environment based on the adjusted distance between the device and the one or more of the multiple access points.
[0009] It may be an advantage to apply a trained machine learning model to the feature image, as the feature image provides signal propagation information in scenarios with mixed line ofsight (LoS) and non-LoS (nLoS) paths. This provides a more accurate location of the device, especially in an indoor environment where the path between the device and the multiple access points is typically a non-LoS path.
[0010] In some embodiments, generating the respective feature image comprises cropping a rectangular section from the floor plan, wherein the estimated location of the device lies on an edge of the rectangular section.
[0011] In some embodiments, generating the respective feature image comprises rotating the rectangular section, such that each feature image has a same orientation.
[0012] In some embodiments, the trained machine learning model comprises a first sub-model and a second sub-model; and applying the trained machine learning model comprises applying the first sub-model to the respective feature image and applying the second sub-model to the respective distance measurement.
[0013] In some embodiments, the trained machine learning model comprises a third sub-model; applying the trained machine learning model comprises applying the third sub-model to an output of the first sub-model and an output of the second sub-model; and an output of the third sub-model corresponds to the estimated distance between the device and the respective access point.
[0014] In some embodiments, the first, second and third sub-models are one of: a vision transformer; and a neural network.
[0015] In some embodiments, the device is one of multiple devices within the indoor environment; and receiving distance measurements comprises receiving distance measurements between the multiple devices and receiving distance measurements between the multiple devices and the multiple access points.
[0016] In some embodiments, the method further comprises determining an estimated location of each of the multiple devices in the indoor environment based on the distance measurements.
[0017] In some embodiments, the method further comprises determining a graph based on the estimated locations, wherein nodes of the graph are indicative of the multiple devices and the multiple access points, and edges of the graph are indicative of the distance measurements.
[0018] In some embodiments, the method further comprises applying a trained graph model to the graph to determine improved estimated locations of the multiple devices, wherein generating the respective feature image is based on the corresponding improved estimated location of the device.
[0019] In some embodiments, the trained graph model is a graph neural network.
[0020] In some embodiments, the method further comprises receiving historical distance measurements between the multiple devices and the multiple access points.
[0021] In some embodiments, determining the adjusted location of the device comprises applying a Kalman filter to the historical distance measurements.
[0022] In some embodiments, the method further comprises: creating multiple historical graphs using the historical distance measurements, wherein nodes of each of the multiple historical graphs are indicative of the multiple devices and the multiple access points, and edges of each of the multiple historical graphs are indicative of the historical distance measurements; for one or more of the multiple historical graphs, applying the trained graph model to the historical graph to generate an output; and determining the improved estimated locations based on the multiple outputs.
[0023] In some embodiments, the method further comprises training, at least in part, a machine learning model using synthetic distance measurements to generate the trained machine learning model.
[0024] In some embodiments, the method further comprises generating the synthetic distance measurements by, for one or more of the multiple access points, applying a trained generator model to the feature image and a ground truth distance between the access point and one of the multiple devices.
[0025] In some embodiments, the trained generator model comprises a first sub-model and a second sub-model; and applying the trained generator model comprises applying the first submodel to the respective feature image and applying the second sub-model to the respective ground truth distance.
[0026] In some embodiments, the trained generator model comprises a third sub-model; applying the trained generator model comprises applying the third sub-model to an output of the first sub-model and an output of the second sub-model; and an output of the third sub-model corresponds to one of the synthetic distance measurements.
[0027] In some embodiments, the first, second and third sub-models are one of: a vision transformer; and a neural network.
[0028] In some embodiments, the method further comprises training, at least in part, a generator model using ground truth distances and a set of training feature images from a training floor plan to generate the trained generator model.
[0029] In some embodiments, the training floor plan is different to the floor plan.
[0030] According to the present disclosure, there is provided software that, when installed on a computer and executed by the computer, causes the computer to perform the method of any one of the previously described embodiments, or part thereof.
[0031] According to the present disclosure, there is provided a system for locating a device in an indoor environment, the system comprising one or more processors configured to perform the method of any one of the previously described embodiments.
[0032] Optional features provided in relation to the method equally apply as optional features to the software and the systems.Brief Description of Drawings
[0033] An example will be described with reference to the following drawings:Fig. 1 illustrates an example embodiment of a system for locating a device within an indoor environment.Fig. 2a illustrates an example embodiment of a method for locating a device within an indoor environment.Fig. 2b illustrates a different example embodiment of a method for locating a device within an indoor environment.Fig. 3 illustrates the framework of a preferred embodiment.Fig. 4 illustrates the structure of the graph neural network (GNN) in the preferred embodiment.Fig. 5 shows the image preparation of the feature images according to the preferred embodiment.Fig. 6a illustrates the structure of the floor-plan- aided neural network (FPDNN) in the preferred embodiment.Fig. 6a illustrates the structure of the trained generator model in the preferred embodiment.Fig. 7 illustrates the structure of the deep vision transformer (DeepVIT) in the preferred embodiment.Fig. 8 illustrates a flowchart for pre-training of the disclosed machine learning models. Fig. 9a shows the generation of synthetic round trip time (RTT) measurements by a trained generator model in an office environment.Fig. 9b shows the generation of synthetic received signal strength (RSS) measurements by a trained generator model in an office environment.Fig. 10a shows an office environment and the corresponding floor plan.Fig. 10b shows a shopping mall environment and corresponding floor plan.Fig. 10c shows a laboratory environment and corresponding floor plan.Fig. 11a shows Wi-Fi access point (AP) and line-tracking vehicles used in the experimental setup.Fig. 11b shows the mobile devices used in the experimental setup.Fig. 11c shows the estimated location of the device on the map in the experimental setup. Fig. 12 shows the locations of eight APs and the trajectory of the device, where the device moves along the trajectory in the real-world scenario.Fig. 13 presents Table 2, which shows a summary of the hyper-parameters of each block in the disclosed method.Fig. 14 shows the performance results of the disclosed method in comparison to other baseline localization methods.Fig. 15 shows the cumulative distribution functions (CDFs) of localization errors of the disclosed method in comparison to other baseline localization methods.Fig. 16a shows the locations of APs, ground truth, and estimated locations in the office environment for different variations of the disclosed method with the GNN (either with or without Kalman Filter).Fig. 16b shows the locations of APs, ground truth, and estimated locations in the office environment for different variations of the disclosed method with the GNN and FPDNN (either with or without Kalman Filter).Fig. 17 shows the CDFs of the location errors of the variations of the disclosed method of Figs. 16a and 16b.Fig. 18a shows the samples used in fine-tuning and testing in the experimental results described herein.Fig. 18b shows the locations of APs, ground truth, and estimated locations of the disclosed method with fine-tuning.Fig. 19a shows the mean average error (MAE), root mean square error (RMSE), median, and 90-th percentile CDF of variations of the disclosed method in comparison with the other baseline localization methods in the office environment.Fig. 19b shows the MAE, RMSE, median, and 90-th percentile CDF of further variations of the disclosed method in comparison with the other baseline localization methods in the office environment.Fig. 20a shows the MAE, RMSE, median, and 90-th percentile CDF of the further variations of Fig. 19a in comparison with the other baseline localization methods in the shopping mall environment.Fig. 20b shows the MAE, RMSE, median, and 90-th percentile CDF of the further variations of Fig. 19a in comparison with the other baseline localization methods in the laboratory environment.Fig. 21 shows the locations of two devices and locations of six APs, where the devices move along the trajectories.Fig. 22 shows the trajectory of a device estimated by the disclosed method in comparison to the trajectory estimated by other localization methods.Description of Embodiments
[0034] This disclosure provides a system and methods for increasing localization accuracy in an indoor environment. This disclosure may address some limitations of indoor localization,particularly in complex indoor environments with non-line-of- sight (nLoS) conditions, where distance measurements may be inaccurate due to multi-path propagation.
[0035] Based on an estimated location of the device, a floor-plan image between the device and an access point may be used to improve localization accuracy. For example, a floor-plan-aided machine learning model may be used to improve localization accuracy. In this disclosure, the use of floor-plan images in localization algorithms is investigated, and experimental results thereof are presented herein.
[0036] Floor-plan images can be obtained of most indoor scenarios, such as evacuation diagrams. The widespread use and mandatory updating of evacuation diagrams in most buildings ensure accessibility and up-to-date information on the floor plan. Floor plans can also be obtained from construction plans. As such, floor plans may be a useful tool in increasing localization accuracy due to the rich amount of information they may provide regarding the indoor environment. The locations of access points (such as Wi-Fi routers) within the indoor environment can also be done quickly through some simple measurements.
[0037] In this disclosure, a trained graph model that is scalable to the number of access points and devices may also be used for obtaining estimated) locations of devices. The disclosure may also provide a zero-shot learning framework that does not need real-world measurements in a new communication environment. This may be achieved by using a trained generator model to provide synthetic data samples in different scenarios where real-world samples are not available, for example. This may improve the generalization ability of the machine learning models described herein and may overcome a lack of training data. As will be discussed, experimental results show that the disclosed method can reduce localization errors by around 30% to 55 %.
[0038] It is noted that, while the disclosed system and methods are primarily directed towards localization of a device within an indoor environment, the disclosed system and methods are also applicable to an outdoor environment (such as, but not limited to, gardens, parks, outdoor living and community spaces, car parks, festivals, and outdoor arenas). Similar to indoor environments, some outdoor environments may not have LoS paths available between a device and a receiver. Given a floor plan of an outdoor environment, the floor plan may be used to improve localization accuracy. For example, machine learning may be utilised with the floor plan of the outdoor environment to improve localization accuracy. As such, this disclosure provides a system and methods for increasing localization accuracy in an outdoor environment. Given this, it can besaid that this disclosure provides a system and methods for increasing localization accuracy in an environment (either indoor or outdoor). Moreover, “indoor environment” and “outdoor environment” may be interchangeable in this disclosure. However, for explanatory purposes, the disclosed systems and methods are described for locating a device within an indoor environment.System
[0039] Fig. 1 illustrates an example embodiment of a system 100 for locating device 111 in indoor environment 105. Fig. 1 is one example of a configuration of system 100. However, system 100 is not strictly limited to this configuration, and this may be one possible embodiment of system 100. It is noted that system 100 of Fig. 1 is only meant to illustrate an example system that is capable of performing the disclosed methods.
[0040] Indoor environment 105 may be any type of indoor environment including, but not limited to, offices, education facilities, shopping centres, hospitals, libraries, museums and residential facilities. As depicted in Fig. 1, indoor environment 105 comprises walls that define different rooms and entryways. The disclosed methods are particularly advantageous in such an environment depicted in Fig. 1 as the walls may cause nLoS conditions within indoor environment 105. More specifically, there may not be a direct path between the device being located and a receiver within indoor environment 105 which is used for localization. Indoor environment 105 may also comprise other structures (such as furniture) or moving bodies (such as people). The disclosed methods are particularly advantageous as they can adapt to changing environments. However, it is noted that the disclosed methods are equally applicable to simple indoor environments, such as empty rooms without walls and / or where LoS conditions can be easily established. The disclosed methods may still be advantageous in such environments to mitigate localization errors, such as ranging errors.
[0041] As depicted in Fig. 1, indoor environment 105 comprises devices 111, 112, 113.Devices 111, 112, 113 may be smartphones, computers, tablets, or any other similar devices. Preferably, devices 111, 112, 113 are mobile devices and / or personal / user devices, such as a smartphone. Devices 111, 112, 113 may also be a field-programmable gate array (FPGA), an application specific integrated circuits (ASIC), or one or more single-board computers, such as a Raspberry Pi or an Arduino. Any of devices 111, 112, 113 may be located using the methods described herein. As will be discussed later in this disclosure, in some embodiments, the disclosed methods may comprise determining the location of each of devices 111, 112, 113simultaneously. Any of devices 111, 112, 113 may also be communicatively coupled with each other (e.g., device 111 may communicate with device 112. Each other permutation may be possible). While three devices are depicted in Fig. 1, it is noted that indoor environment 105 may comprise any number of devices.
[0042] As depicted in Fig. 1, indoor environment 105 comprises multiple access points 121, 122, 123. Multiple access points 121, 122, 123 may be communicatively coupled with any one or more of devices 111, 112, 113. For example, multiple access points 121, 122, 123 may communicate with any one or more of devices 111, 112, 113 using radio signals or other suitable electromagnetic signals. Multiple access points 121, 122, 123 may be routers, such as Wi-Fi routers, for example. As such, multiple access points 121, 122, 123 may communicate with any of devices 111, 112, 113 using a Wi-Fi network according to IEEE 802.11. In some examples, multiple access points 121, 122, 123 may be referred to as anchors. While three access points are depicted in Fig. 1, it is noted that indoor environment 105 may comprise any number of access points. Preferably, indoor environment 105 may comprises more than one access point.
[0043] One consideration for an indoor localization system is to choose a technology that is readily accessible at the user end. Wi-Fi may stand out among readily accessible technologies due to its wide deployment in indoor environments. For example, many electronic devices (such as personal electronic devices like smartphones) are Wi-Fi-enabled, which makes Wi-Fi technology an ideal candidate for indoor localization. Hence, preferably, multiple access points 121, 122, 123 are Wi-Fi routers, or other similar technologies. However, it is noted that multiple access points 121, 122, 123 may be other types of devices, such as, but not limited to, radio transmitters.
[0044] System 100 comprises device 150, which performs the methods disclosed herein.Device 150 may be smartphone, computer, tablet, a server device, or any other similar device. Device 110 may also be a field-programmable gate array (FPGA), an application specific integrated circuits (ASIC), or one or more single-board computers, such as a Raspberry Pi or an Arduino. Device 110 comprises processor 151, which may be configured to perform the methods described in this disclosure. Device 110 comprises memory 152, which comprises non-volatile memory 153 and / or volatile memory 154. Processor 151 may communicate with memory 152 by communicating with non-volatile memory 153 and / or volatile memory 154. Non-volatile memory 153 is a non-transitory computer readable medium and may be an optical disk drive, hard disk drive, solid-state drive, flash memory, storage server, cloud storage or anotherequivalent type of memory. Volatile memory 154 may be cache, RAM, or another equivalent type of memory.
[0045] Any of devices 111, 112, 113 may be equivalent or similar to device 150. For example, device 111 may be equivalent to device 150 and hence, device 111 may perform the methods disclosed herein for locating itself with indoor environment 105. However, device 150 may also be a server that is communicatively coupled to devices 111, 112, 113. If device 150 is a server or equivalent device, device 150 may be located remotely to indoor environment 105, or may be located within indoor environment 105.
[0046] Memory 152 may store data to be retrieved for later use. For example, memory 152 may store distance measurements between any of devices 111, 112, 113 and any of access points 121, 122, 123. Memory 152 may also store the estimated locations and the adjusted locations of any of devices 111, 112, 113. Memory 152 may also store a floor plan of indoor environment 105 (i.e., a floor plan image) and associated feature images. In essence, memory 152 may store any data used by processor 151 when performing the methods described herein. The data thereof may be stored in memory 152 in the form of a JSON format file, XML format file, or another equivalent data format file. The images may be stored in memory 152 in a Joint Photographic Experts Group (JPEG) format, RAW data format, or another equivalent data format. It is noted that the reference to “image” in this disclosure refers to “image data”, which may be two-dimensional image data such as an RGB image.
[0047] The methods described herein comprise applying one or more trained machine learning models. These machine learning models may be stored on memory 152 by storing the weights that form the respective models, for example. Memory 152 may also store any output values calculated by processor 151 applying one or more machine learning models, or any other variable or data necessary to perform such methods described herein.
[0048] Software, that is, an executable program stored on non-volatile memory 153 causes processor 151 to perform methods for locating any of devices 111, 112, 113. While the singular of “processor” is used herein, it is meant to also encompass multiple processors that are individually or together configured (e.g., programmed) to perform the methods disclosed herein. As such, processor 151 may refer to multiple central processing units (CPUs) and / or graphical processing units (GPUs) that are configured to collectively perform the methods disclosed herein.
[0049] Once executed, the software may cause processor 151 to receive distance measurements between any of devices 111, 112, 113 and the multiple access points 121, 122, 123 within indoor environment 105, determine an estimated location of any of devices 111, 112, 113, generate a feature image from a floor plan of indoor environment 105 based on the estimated location of any of devices 111, 112, 113, apply a trained machine learning model to the distance measurement and the feature image, and determine an adjusted location of any of devices 111, 112, 113.
[0050] Device 150 (more specifically, processor 151) may communicate with any of devices 111, 112, 113 via antenna 155. Antenna 155 may communicate with any of devices 111, 112, 113 through a wireless connection, such as by using a Wi-Fi network according to IEEE 802.11. In some examples, antenna 155 may communicate using a short-range wireless technology standard, such as Bluetooth, particularly if device 150 is located within indoor environment 105.MethodsMethod 1
[0051] Fig. 2a illustrates an example embodiment of method for locating a device (such as any of devices 111, 112, 113) in indoor environment 105, given as method 200. Fig. 2a is to be understood as a blueprint for a software program and may be implemented step-by-step, such that each step in Fig. 2a is represented by a function in a programming language, such as, but not limited to, Python, C++, or Java. The resulting source code is then compiled and stored as computer-executable instructions on non-volatile memory 153, which causes processor 151 (or multiple processors or a distributed computing architecture) to perform method 200. It is noted that a method for locating a device in an outdoor environment may be similar and / or equivalent to method 200.
[0052] For illustrative purposes, the example embodiment of method 200 will be explained by locating device 111 within indoor environment 105. However, it is noted that any of devices 111, 112, 113 may be located using method 200. Further, in some examples, the location of each of devices 111, 112, 113 may be located simultaneously.
[0053] Processor 151 first receives 210 distance measurements between device 111 and multiple access points 121, 122, 123. The distance measurements may be based oncommunication between device 111 and multiple access points 121, 122, 123. The distance measurements may be indicative of a distance between device 111 and multiple access points 121, 122, 123. However, there may be error (such as ranging error) due to nLoS conditions, for example. So, the distance measurements may not represent an accurate distance between device 111 and multiple access points 121, 122, 123. The distance measurements may be based on round trip time (RTT) measurements, received signal strength (RSS) measurements and / or channel state information (CSI).
[0054] Processor 151 then determines 220 an estimated location of device 111 from the distance measurements. More specifically, processor 151 determines 220 an estimated location of device 111 within indoor environment 105. Processor 151 may determine 220 the estimated location of device 111 by applying an algorithm, mathematical operation, or a machine learning model. For example, processor 151 may determine 220 the estimated location of device 111 by applying a triangulation algorithm, a trilateration method, or at least square algorithm to the distance measurements.
[0055] Next, for one or more of multiple access points 121, 122, 123, processor 151 generates 231, 232 a feature image from a floor plan of indoor environment 105 based on the estimated location of device 111. The floor plan may be an image, such as a two-dimensional image that represents indoor environment 105. The feature image corresponds to a physical region between device 111 and the corresponding access point 121, 122, 123. For example, the feature images may be a section of the floor plan, which is spanned by the distance between the estimated location of device 111 and the location of the corresponding access point 121, 122, 123. With reference to Fig. 2a, Processor 151 may generate 231, a feature image corresponding to the physical region between device 111 and access point 121. Similarly, processor 151 may also generate 232, a feature image corresponding to the physical region between device 111 and access point 122. The floor plan may be an evacuation diagram, blueprint, artistic representation of indoor environment 105, or another equivalent type of diagram representing the floor plan. Preferably, the floor plan is a “birds’ eye” view of indoor environment 105. Preferably, the floor plan represents obstructions within indoor environment 105 that may obstruct a direct path between device 111 and multiple access points 121, 122, 123 (or any of devices 111, 112, 113 and multiple access points 121, 122, 123), such as walls and furniture.
[0056] Processor 151 then applies 241, 242 a trained machine learning model to the distance measurement between device 111 and one of access points 121, 122, 123 and the correspondingfeature image to calculate an adjusted distance between device 111 and the corresponding access point 121, 122, 123. For example, with reference to Fig. 2a, processor 151 may apply 241 the trained machine learning model to a feature image corresponding to the physical region between device 111 and access point 121. Similarly, processor 151 may also apply 242 the trained machine learning model a feature image corresponding to the physical region between device 111 and access point 122. Preferably, the adjusted distance may represent a more accurate distance measurement between device 111 and the corresponding access point 121, 122, 123. In particular, the trained machine learning model may be trained to identify or recognise LoS and nLoS conditions from the floor plan (specifically, the corresponding feature image).
[0057] Finally, processor 151 determines 250 an adjusted location of device 111 in indoor environment 105 based on the adjusted distance between device 111 and one or more of multiple access points 121, 122, 123. Preferably, the adjusted location represents a more accurate location of device 111 compared to the estimated location based on the distance measurements alone. As such, the adjusted location may be referred to as an improved location or the like. Processor 151 may determine 250 the adjusted location of device 111 by applying an algorithm, mathematical operation or a machine learning model. For example, processor 151 may determine 250 the adjusted location of device 111 by applying a triangulation algorithm, a trilateration method, or at least square algorithm to the adjusted distances.
[0058] In some embodiments, processor 151 generates 231, 232 the respective feature image by cropping a rectangular section from the floor plan. The estimated location of device 111 (within the indoor environment) may lie on an edge of the rectangular section. Similarly, the location of the corresponding access point 121, 122, 123 may lie on another edge of the rectangular section. Preferably, the estimated location of device 111 and the location of the corresponding access point 121, 122, 123 lie on opposite edges of the rectangular section.
[0059] There may be one or more feature images corresponding to one or more of the multiple access points 121, 122, 123. As such, each of the one or more feature images may have a different orientation based on the respective locations of device 111 and the corresponding access point 121, 122, 123 within indoor environment 105. In some embodiments, processor 151 generates 231, 232 the respective feature image by rotating the rectangular section, such that each feature image has a same orientation. This may ensure that each of the feature image is compatible with the trained machine learning model and hence, can be used as input into the trained machine learning model.
[0060] As discussed earlier, indoor environment 105 may comprise multiple devices 111, 112, 113. Hence, in some embodiments, processor 151 receives 210 distance measurements between multiple devices 111, 112, 113 and / or receives 210 distance measurements between multiple devices 111, 112, 113 and multiple access points 121, 122, 123. Devices 111, 112, 113 may communicate with each other in a similar manner to which devices 111, 112, 113 may communicate with multiple access points 121, 122, 123 e.g., using radio signals according to a Wi-Fi protocol. Processor 151 may thus determine 220 an estimated location of each of the multiple devices 111, 112, 113 in indoor environment 105 based on the distance measurements.
[0061] In some embodiments, processor 151 may receive historical distance measurements between any of multiple devices 111, 112, 113, and the multiple access points 121, 122, 123. The historical distance measurements may be indicative of historical locations of any of the devices 111, 112, 113. Historical in the present context refers to a time before the present. For example, historical may refer to a time before a time corresponding to the location of any of devices 111, 112, 113 being determined by method 200. The historical distance measurements may comprise distance measurements at multiple historical time steps. The time steps may be any time interval, such as every few seconds, for example.
[0062] In some embodiments, processor 151 may continuously receive distance measurements between device 111 and multiple access points 121, 122, 123. In other words, in subsequent timesteps, processor 151 may receive further distance measurements between device 111 and multiple access points 121, 122, 123. As such, previously received distance measurements may become historical distance measurements. It may be an advantage to continuously receive distance measurements between device 111 and multiple access points 121, 122, 123 as indoor environment 105 may change with time (e.g., furniture may be moved). Moreover, processor 151 may repeat method 200 (or parts thereof) after determining 250 the adjusted location of device 111 or may iteratively apply method 200 such that the adjusted location in one iteration becomes the estimated location in the following iteration. This may be an advantage as it provides dynamic determination of device 111 in indoor environment 105. For example, this may be an advantage as device 111 may not be stationary. In other words, the location of device 111 may change with each time step.
[0063] In some embodiments, processor 151 may determine the adjusted location of device 111 by applying a sensor fusion algorithm to the historical distance measurements. For example, the processor 151 may determine the adjusted location of device 111 by applying a Kalman filter tothe historical distance measurements. Processor 151 may also determine the adjusted location of device 111 using a particle filter algorithm, which is based on the historical distance measurements. Processor 151 may also determine the adjusted location of device 111 based on the historical distance measurements and the (current) distance measurements received at 210. For example, processor 151 may determine the adjusted location of device 111 by applying a sensor fusion algorithm to the historical distance measurements and the (current) distance measurements received at 210.
[0064] The machine learning models described in this disclosure (including the trained machine learning model recited at 241, 242) are understood to be models, such as mathematical models, which receive input and generate an output based on the input. The machine learning models may be of an architecture, such as, but not limited to, a neural network, for example. In general, machine learning models are ‘trained’ to learn and recognise patterns in input and provide an output that is a prediction based on the training it has undergone. Training involves updating weights or parameters of the machine learning model, which define the machine learning models, to minimise a loss value, thereby creating a trained machine learning model (in other words, a machine learning model trained to generate an output). This may involve a gradient descent and backpropagation method.
[0065] The machine learning models recited in this disclosure may be stored on memory 152 storing the weights that define the model. As such, each of the machine learning models may be referred to as a “memory model,” given that it is defined by parameters (i.e., the weights) that can be stored on computer memory. In some embodiments, the machine learning models may be programmed on an integrated circuit, such as a field-programmable gate array (FPGA) or an NVIDIA processing unit. In such an embodiment, the handling modules (or processor(s) that may perform the method or part thereof) may not retrieve the parameters from memory 152. Instead, an input may be communicated to the integrated circuit, and the integrated circuit may apply the machine learning model to the input and generate an output, which is then communicated to the handling modules (or processor(s) that may perform the method or part thereof).
[0066] Integrated circuits, such as FPGAs, can be used where flexibility, speed, and parallel processing capabilities are desired. In such an embodiment, the integrated circuit may be part of device 150 of system 100 and may be considered as a “processor” or “processing unit”. Otherimplementations, such as application specific integrated circuits (ASIC) or neuromorphic architectures, are equally useable.
[0067] Applying any of the machine learning models in this disclosure (such as the trained machine learning model at 241, 242) may also be considered to be executing or evaluating the machine learning model. Applying, executing, or evaluating the machine learning models may involve calling an application programming interface (API) routine to send the input to a server and the server then performs the calculations according to the machine learning model and returns the results. In other examples, applying, executing, or evaluating may involve issuing a command to local hardware, such as a local chip, device, machine learning accelerator (e.g., a USB device designed to efficiently perform machine learning tasks or NVIDIA’ s Deep Learning Accelerator (DLA)), etc., that has the machine learning model stored thereon and provides a command interface to interact with the model. It is also possible to have a local copy of the machine learning model available so that the calculations are performed by the main processor of the local machine. Other local, remote, or distributed implementations (such as cloud computing environments) are equally useable.
[0068] In some examples, the machine learning models described herein may be a neural network. In further examples, these machine learning models may be a neural network comprising one or more convolutional layers. As such, the machine learning models may perform the methods described herein by creating feature maps using convolutional filters. Such a machine learning model is known as a convolutional neural network (CNN). A CNN is ideal for applications involving images, and the image as it accounts for the positioning and shape of objects captured in the image.
[0069] However, other types of machine learning models are equally applicable here, such as K nearest neighbour, decision tree, support vector machines, regression models and other artificial neural networks, such as long short-term memory (LSTM) networks or deep neural networks. The machine learning models described herein may also be a collective of different models. The machine learning models may also be based on a transformer model, which comprise encoder and / or decoder blocks and predictions by ‘tokenising’ the input (such as an image). As such, the machine learning models may also comprise self-attention. The machine learning models may output a numerical value.
[0070] In some embodiments, the machine learning models are a multimodal machine learning model, in which multiple inputs of different modalities (e.g., text, image data and audio data) are used to provide one or more generated outputs. An example of a multimodal machine learning model is an object detection model, which detects the location of a specific object (specified by input text, for example) in an image. This example model may generate output text that describes the location of the specified object in the image. Although the multimodal machine learning model can be evaluated on multiple input of different modalities, the multimodal machine learning model can also be evaluated on a single input and still generate an output based on the single input.
[0071] In some embodiments, the trained machine learning model comprises a first sub-model and a second sub-model. The first sub-model and the second sub-model may be a similar or different type of machine learning model. The first sub-model and the second sub-model may be a similar or different type of machine learning model architecture. In these embodiments, processor 151 may apply 241, 242 the trained machine learning model by applying the first submodel to the respective feature image of the corresponding access point 121, 122, 123 and applying the second sub-model to the respective distance measurement. Processor 151 may apply the first sub-model and the second sub-model simultaneously, rather than sequentially. For example, processor 151 may utilise parallel computing, such that each of the first sub-model and second sub-model are evaluated on different processing units.
[0072] The outputs of the first sub-model and the second sub-model may be combined to determine the adjusted distance between device 111 and the corresponding access point 121, 122, 123. Processor 151 may apply an algorithm, mathematical operation of a machine learning model to the outputs of the first sub-model and the second sub-model. For example, processor 151 may apply a fusion algorithm to the outputs of the first sub-model and the second sub-model to determine the adjusted distance.
[0073] In the embodiments described above, the trained machine learning model may comprise a third sub-model. The third sub-model may be a similar or different type of machine learning model compared to the first sub-model and the second sub-model. The third sub-model may be a similar or different type of machine learning model architecture compared to the first sub-model and the second sub-model. In these embodiments, processor 151 applies 241, 242 the trained machine learning model by applying the third sub-model to an output of the first sub-model and an output of the second sub-model. As such, an output of the third sub-model may correspond tothe estimated distance between device 111 and the respective access point 121, 122, 123. In some embodiments, the first, second and third sub-models are one of a vision transformer; and a neural network. Vision transformers may break down an input image into a series of patches, similar to how text is broken into tokens in Natural Language Processing (NLP). Each patch may then be serialized into a vector and mapped to a smaller dimension with a single matrix multiplication.
[0074] In embodiments where, processor 151 receives 210 distance measurements between multiple devices 111, 112, 113 and / or receives 210 distance measurements between multiple devices 111, 112, 113 and multiple access points 121 (and hence, processor 151 may determine an estimated location of each of the multiple devices 111, 112, 113 in indoor environment 105), processor 151 may determine a graph based on the estimated locations of the multiple devices 111, 112, 113. Nodes of the graph may be indicative of multiple devices 111, 112, 113 and multiple access points 121, 122, 123, and edges of the graph may be indicative of the distance measurements. For example, the nodes may correspond to the estimated locations of multiple devices 111, 112, 113 and the locations of multiple access points 121, 122, 123.
[0075] In the above embodiments, processor 151 may apply a trained graph model to the graph. The trained graph model may be a machine learning model which may receive a graph as input and may output one or more values (e.g., numerical values) representing a prediction or the like based on the information contained in the input graph. The trained graph model may also receive a graph as input and output a different graph that corresponds to a prediction based on the input graph. Similar to other machine learning models, graph models may comprise adjustable weights (parameters) that define the model. These weights may be modified during a training process. The trained graph model may be trained to improve the accuracy of the estimated locations by adjusting the estimated locations based on the received distance measurements, for example. In some examples, the trained graph model is a graph neural network.
[0076] The trained graph model may follow the following procedure or parts thereof when applying to the previously described graph: (i) Node Representation: Each node in the graph may be represented by a feature vector, which could be a set of attributes associated with the node; (ii) Message Passing: Nodes may exchange information with their neighbours, where the information a node receives may be determined by its position in the graph and its connections (i.e., the edges); (iii) Aggregation: The node may take all the information it has received from its neighbours and combines it in some way. This could be a simple operation like taking the sum oraverage, or a more complex operation; and (iv) Update: The node may update its own feature vector based on the aggregated information.
[0077] In the above embodiments, processor 151 may apply a trained graph model to the graph to determine improved estimated locations of the multiple devices 111, 112, In such embodiments, processor 151 may generate 231, 232 the respective feature image based on the corresponding improved estimated location of the device 111, rather than generating 231, 232 the respective feature image based on the estimated locations determined only from the distance measurements.
[0078] As discussed previously, processor 151 may receive historical distance measurements between the multiple devices 111, 112, 113 and the multiple access points 121, 122, 123. In these embodiments, processor 151 may also create multiple historical graphs using the historical distance measurements. For example, nodes of each of the multiple historical graphs may be indicative of the multiple devices 111, 112, 113 and the multiple access points 121, 122, 123, and edges of each of the multiple historical graphs may be indicative of the historical distance measurements. For example, the historical distance measurements may comprise distance measurements at multiple historical time steps and hence, processor 151 may create historical graphs of one or more or each of the multiple historical time steps.
[0079] In these embodiments, for one or more of the multiple historical graphs, processor 151 may apply the trained graph model to the historical graph to generate an output. Then processor 151 may determine the improved estimated locations based on the multiple outputs. Using the historical distance measurements provides an indication of the historical locations of any of devices 111, 112, 113. As such, historical distance measurements may be used to predict the movement of any of devices 111, 112, 113 and hence, predict the likely location of any of devices 111, 112, 113 in the present. Processor 151 may apply the trained graph model to the multiple historical graphs and the graph created from the (current) distance measurements and determined multiple outputs to thus, determine the improved estimated locations.
[0080] Training of machine learning models is generally difficult, as a large amount of training data may be required, such that the trained machine learning models provide accurate predictions. Obtaining large amounts of training data may also be difficult, especially for some specific machine learning tasks, as data may be generally unavailable, difficult, infeasible, or time-consuming to obtain. Moreover, some training data may lack variability (i.e., the trainingdata is similar), which may make it difficult for the trained machine learning models to be applicable to a variety of input data.
[0081] These issues may be addressed by training the machine learning model in this disclosure using synthetic data, which is data that may be generated in-silico, rather than corresponding to real-world measurements, for example. As such, in some embodiments, the trained machine learning model may be trained, at least partially, using synthetic distance measurements. For example, processor 151 may train, at least in part, a machine learning model using synthetic distance measurements to generate the trained machine learning model. It is noted that the trained machine learning model may be trained using a dataset comprising real-world training data (e.g., training data obtained from real-world measurements) and synthetic training data. For example, a machine learning model may be trained initially using the synthetic training data and then the (at least partially) trained machine learning model may be further trained (or fine-tuned) using real- world training data. The synthetic distance measurements may be generated from the floor plan of indoor environment 105.
[0082] In some embodiments, processor 151 may generate the synthetic distance measurements by, for one or more of the multiple access points 121, 122, 123, applying a trained generator model to the feature image and a ground truth distance between the access point 121, 122, 123 and one of the multiple devices 111, 112, 113. The trained generator model may be referred to as a synthetic data generator. In general, the trained generator model is a machine learning model trained to generate new data that resembles a given dataset. The trained generator model may learn the underlying patterns and structures of the input data (e.g., the feature images of the floor plan) and use this knowledge to create new, similar data. The trained generator model may be applied to an image (in other words, the image in input to the trained generator model) and outputs distance measurements, which is a prediction based on the input image (i.e., the feature image of the floor plan of indoor environment 105). The trained generator model may be trained in a similar manner described earlier in the disclosure.
[0083] The trained generator model may be similar or comprise similar features to the trained machine learning model applied to the feature images at 241, 242 of method 200. As such, some embodiments of the trained machine learning model may correspond to embodiments of the trained generator model. For example, the trained generator model comprises a first sub-model and a second sub-model. The first sub-model and the second sub-model may be a similar or different type of machine learning model. The first sub-model and the second sub-model may besimilar or different types of machine learning model architecture. Processor 151 may apply the trained generator model by applying the first sub-model to the respective feature image and applying the second sub-model to the respective ground truth distance. Similar to the embodiments of the trained machine learning, processor 151 may apply the first sub-model and the second sub-model of the trained generator model simultaneously using parallel computing, for example.
[0084] As another example, the trained generator model may comprise a third sub-model, similar to embodiments of the trained machine learning model. Processor 151 may apply the trained generator model by applying the third sub-model to an output of the first sub-model and an output of the second sub-model. The output of the third sub-model may correspond to one of the synthetic distance measurements. Similar to the trained machine learning model, alternatively or additionally, processor 151 may combine the outputs of the first sub-model and the second sub-model by applying an algorithm (such as a fusion algorithm) and mathematical operation (such as a sum, weighted sum, average, or the like). In some embodiments, the first, second and third sub-models are one of a vision transformer; and a neural network.
[0085] In some embodiments, the trained generator model may be trained using a training floor plan. For example, the training floor plan may be different to the floor plan i.e., the training floor plan represents a shopping centre, while the floor plan represents an office. As such, the trained generator model may be trained, at least in part, using ground truth distances and a set of training feature images from a training floor plan. For example, processor 151 may train, at least in part, using ground truth distances and a set of training feature images from a training floor plan. This may provide the machine learning models disclosed herein with greater variability and accommodate different floor plans.Method 2
[0086] Fig. 2b illustrates a different example embodiment of a method for locating a device (such as any of devices 111, 112, 113) in indoor environment 105, given as method 260. Fig. 2b is to be understood as a blueprint for a software program and may be implemented step-by-step, such that each step in Fig. 2b is represented by a function in a programming language, such as, but not limited to, Python, C++, or Java. The resulting source code is then compiled and stored as computer-executable instructions on non-volatile memory 153, which causes processor 151 (or multiple processors or a distributed computing architecture) to perform method 260.Embodiments of method 200 may also be embodiments of method 260. It is noted that a method for locating a device in an outdoor environment may be similar and / or equivalent to method 260.
[0087] Processor 151 receives 261 distance measurements between the multiple devices 111, 112, 113 and the multiple access points 121, 122, 123. The distance measurements may be distance measurements between the multiple devices 111, 112, 113, distance measurements between the multiple devices 111, 112, 113 and the multiple access points 121, 122, 123, or combination thereof. For example, device 111 may be connected to (or in communication with) each of the multiple access points 121, 122, 123, while device 112 may only be connected to access points 121, 122. Device 111 may also be connected to device 112, and processor 151 would hence receive 261 distance measurements corresponding to these connections. Preferably, processor 151 receives 261 distance measurements indicative of a device-to-device connection and distance measurements indicative of a device-to-access point connection.
[0088] Processor 151 then determines 262 an estimated location of any of devices 111, 112, 113 from the distance measurements. In the example above, processor 151 may determine 262 an estimated location of devices 111, 112. 262 may be similar to 220 of method 200. Next, processor 151 determines 263 a graph based on the estimated locations. Nodes of the graph are indicative of the multiple devices 111, 112, 113 and the multiple access points 121, 122, 123, and edges of the graph are indicative of the distance measurements. This is similar to the embodiments described in relation to method 200.
[0089] Processor 151 applies 264 a trained graph model to the graph to determine improved estimated locations of the multiple devices 111, 112, 113. Finally, processor 151 determines 265 an adjusted location of device 111 in indoor environment 105 based on the improved estimated locations of the multiple devices 111, 112, 113. For example, the improved estimated locations may correspond to the adjusted location of device 111. In another example, processor 151 may determine 265 the adjusted location of device 111 by applying an algorithm, mathematical operation or machine learning model to the improved estimated locations of the multiple devices 111, 112, 113.Mathematical formulation and preferred embodiment
[0090] The mathematical formulation of the disclosed method is now presented. A preferred embodiment of method 200 is also presented. However, it is noted that alterations and alternatives from the preferred embodiment are equally possible.Notations
[0091] In the mathematical formulation, uppercase letters, e.g., X, to represent given constant numbers or parameters. Lowercase letters, e.g., x, are used to represent scalar variables. A calligraphic uppercase letter, e.g., A, is used to denote a set, and | A | is its cardinality (i.e., the size of set containing unique elements). In particular, {xk}=1denotes a set with given elements, i.e., {xz. ={xl,...,xA.|. and R is the set of real numbers. An uppercase bold letter, e.g., A, is used to represent a matrix, and a lowercase bold letter, e.g., a, to denote a row vector.Superscript ’ denotes the transposition of a matrix or vector, e.g., A’ and a’, respectively. For Amatrices A and B, [A, B] is their horizontal concatenation, while = [A’, B’ ]’ is the Bvertical concatenation. Other notations will be specifically stated.Framework of the preferred embodiment
[0092] Fig. 3 illustrates the framework of the preferred embodiment. As depicted in Fig. 1, there are two blocks in the preferred embodiment as will be described in more detail below: 1) the trained graph model (e.g., graph-based pre-localization), and 2) the trained machine learning model (e.g., a floor plan aid neural network) for improving the localization accuracy. In other words, the preferred embodiment comprises applying a trained graph model to the graph to determine improved estimated locations of the multiple devices before applying method 200. As will be described below, the preferred embodiment comprises multiple devices and multiple access points, where the locations (more specifically, the adjusted locations) of the multiple devices are determined generally simultaneously.
[0093] Consider a localization system with M devices and K Wi-Fi access points (APs). Thus, there are a total of N = M + K nodes in the network. The indices of devices are denoted by m G M, where M > {1,...,, and the indices of APs are denoted by k G K, whereKD {M + 1,..., M + K. In some cases where this is no need to distinguish devices and APs, the index n is used to refer to the n -th node.
[0094] Time may be discretised into time intervals. In each time interval, a device may scan the nearby nodes (including both APs and devices) and broadcast an initial RTT request. The nodes may respond to the request and send the acknowledgment back to the device to establish the connection. Once the connection is established, the device starts measuring the distance to nearby nodes, including devices and APs. The measurement between the m -th device and the n -th node includes their RTT distance,, RSS,, and the index of the node. Thus, the information obtained by the m -th device is given by(1)Since a device cannot establish a link to itself, p^m= 0.
[0095] In the t -th time interval, all the devices upload their information to a server (e.g., device 150 of system 100). The information that can be used for localization is denoted by P^, wherePu(2,! • • • PSI ••• P2,i 1pW _ Pt2 P2,2 ••• P™?2 ••• PM, 2(2)P!1 P(2,! V ••• pJJv ••• p£wj
[0096] In the preferred embodiment, from P^, the trained graph model estimates the locations of devices (i.e., determined the improved estimated locations, which may be referred to as coarse location). Based on the coarse locations, the feature image (as referred to as a floor-plan image) from each AP to the m -th device is cropped from the floor plan, which is used by the trained machine learning to further improve localization accuracy. It is noted that the coarse locations obtained from the trained graph model help the trained machine learning model to identify the communication environment between the device and its neighbours. Thus, a better prelocalization algorithm helps to improve the final localization accuracy.GNN Pre-localization
[0097] In this section, constructing a graph representation of a wireless network is introduced (i.e., determining a graph based on the estimated locations). Then, a trained graph model for prelocalization to obtain the coarse locations of devices are developed. In the preferred embodiment, the trained graph model is a trained graph neural network (GNN). Hence, the trained graph model (and its untrained counterpart) will be referred to as GNN in the following description.Graph Representation
[0098] Fig. 4 illustrates the structure the GNN in the preferred embodiment. The graph representation at the t -th time interval is represented by G^. As depicted in Fig. 4, the circles are the APs, and the squares are the devices. Each solid line is an edge between a device and an AP, and the dashed line represents the edge between two devices. The node and edge features may be summarized as follows:• The node feature of an AP is the coordinate of it,which remain constant in the localization system.• The node feature of a device includes the estimated location of it. The estimated location may be updated by a GNN with L layers. The edge feature in the I -th layer of the GNN are denoted by (lg) e R2xl. The initial node feature(0) is the estimated location obtained from the trilateration method that only uses the RTT distances,■ • Letand M be the sets of APs and devices that can establish stable connections with the m -th device at the t -th time interval, respectively. The set of neighbour nodes of the m -th device is defined asNm>. The edge feature between the m -th device and the n -th node is defined as,n \, \ / n G N.
[0099] It is assumed that two devices can establish a stable connection between them when there is a LoS path. In nLoS cases, there is no edge between the two devices.Graph Neural Network
[0100] As previously discussed, the GNN may employed for the pre-localization of devices. The input of the GNN block is the initial graph representation at the t -th time interval, G^. The output of the GNN block is the coarse locations of devices,= (x, y m = 1, 2,... M.
[0101] As shown in Fig. 4, the GNN consists of four steps: 1) message-passing, 2) aggregation, 3) feature update, and 4) output.1. Message-passing: The GNN consists of Lglayers. The input of the m -th device is (0), which is also the initial feature of the device. In the I -th layer of the GNN, to update the feature of the m -th device, all its neighbour nodes generate messages based on their features,(lg- 1), and the edge feature between the m -th device and the n -th node, e^n, n G Nm. The message can be expressed asm'2, (lg) = < / > (h' '1( Y - o,e!i ’0M ), (3)where ^(-;0M) represents a feed forward neural network (FFNN) with parameters 0M.2. Aggregation: The m -th device aggregates the messages from all its neighbours to obtain aggregated information p^ lg). The aggregation can be expressed as) = AGG(mw(l ),ne N Y (4)In the disclosed GNN, the aggregation function is set to the mean of all the messages. 3. Feature update: The m -th device updates its feature by merging (lg-1) and P^( / g).The output of the I -th layer can be expressed askTJ1’"’ k - -^’kl.(5)where ^>(-;0F) is an FFNN with parameters 0F.4. Output: After Lgrounds of updates, the output from all the nodes, {h^ (Lg)}^=1, is obtained. The output of the GNN is the estimated locations of all the devices,Y('’ = hl(')(L ),m = l,2,..., Af.
[0102] To train the GNN, the ground truth locations of devices are used as labels, and the gap between labels and the outputs of the GNN is minimised, e.g.,LG™(6M.0F)=E(||Y<',’ -Y!;’|I2), <6)where is the ground truth of the location, and | |2is the Euclideandistance betweenand. With the loss function in Eq. (6), a stochastic gradient descent can be used to train the parameters 0Mand 0Fof the GNN. However, other algorithms may be used to train the GNN, such as an Adam optimiser.The trained machine learning model (FPDNN)
[0103] In this section, the trained machine learning model (with reference to method 200) is introduced to estimate the distance between a device and an AP better. In the preferred embodiment, the trained machine learning model is a neural network (specifically, a deep neural network). Hence, the trained machine learning model (or its untrained counterpart) may be referred to as a floor-plan- aided deep neural network (FPDNN), given that the trained machine learning model is applied to feature images of the floor plan.Image Preparation
[0104] As previously discussed, indoor localization is difficult due to the presence of nLoS paths. More specifically, the RTT distance is biased due to the multi-path effect, especially when the LoS path is blocked. To improve localization accuracy, the feature image (i.e., floor-plan sub-image) between a device and an AP is utilised to compensate for the nLoS conditions within the indoor environment. The images are cropped from the floor plan.
[0105] Fig. 5 shows the image preparation of the feature images according to the preferred embodiment. More specifically, Fig. 5 shows the procedure for preparing images between the m-th device and the k -th AP. Firstly, the coarse location of the device is estimated by the GNN block. Then, a rectangular image with size xffj between the device and the AP is croppedfrom the floor plan, whereis the coarse distance between the device and the AP, and H is the width of the image. It is noted that the selection of width H may depend on the positioning accuracy of the coarse location. Specifically, the width H of the cropped image may be set to 256, corresponding to approximately 4 m in the floor-plan image. In the example shown in Fig.5, the root mean square error (RMSE) of the coarse location estimated by the GNN block may less than 2 m. Thus, the redundancy is sufficient for extracting geometric information surrounding the device.
[0106] To regulate the cropped image, the image is rotated such that the direction from the AP to the device in the image is always pointing rightward, as shown in Fig. 5. However, other regulation methods may be used. For example, the image may be rotated such that the direction from the AP to the device in the image is always pointing leftward. The image between the m -th device and the k -th AP is denoted bye R" ' ”, which can be obtained from the following function,tWc(VY..A). (7)where A is the floor plan of the whole area, andIG(•,•,•) includes the image cropping and rotation. Therefore, the input features of FPDNN include 1) the collected information from the device, P^; 2) the images between the device and its neighbour APs, G.Structure of FPDNN
[0107] Fig. 6a illustrates the structure of FPDNN and shows how to estimate the distance between the m -th device and the k - AP. FPDNN consists of three blocks: the deep vision transformer (DeepVIT) block 610 and two FFNN blocks 621, 622. At the t -th time interval, the image feature and the measurementare the inputs of the DeepVIT block 610 and the first FFNN 621, respectively. The concatenation of their outputs serves as the input of the secondFFNN 622, which outputs the estimated distance, denoted. The true distance between the AP and the device, denoted by d^k, is used as the label to train FPDNN.
[0108] Deep VIT block
[0109] Fig. 7 illustrates the structure of Deep VIT. Deep VIT is a kind of transformer that assigns weighting parameters to different parts of the image. The parts with higher weighting parameters will have stronger impacts on the final output. In this way, the output can “pay attention” to the important parts of the image.
[0110] The input of the Deep VIT block is the floor-plan imageand the output of the Deep VIT is the hidden feature given by oT. There are three different layers in the Deep VIT block, including the splitting and flattening layer, the position embedding layer, and the transformer encoder layer.• Spliting and flatening-. As shown in Fig. 7, the input image I is split into / P2mini-images. Let I(G R / denote the f -th mini-image, where P is the side length of each mini-page. Then, each mini-image is vectorized to a vector s, with dimension D' = Px P. For instance, let I.., denote the element of I, at the position (z, j, k), and thenS'—[I 1,1,1 ’ I 1,1,2 ’ - ■ I 1,1, C ’ I 1,2,1 • ■ I 1, P, C ’ zOx( o)I I I I I 1Next, each vector szis compressed into the patch sfE with the size of IxD, where E G RD XDdenotes a trainable matrix. Finally, the output of the flatten layer, denoted as F, is expressed as follows:SjE s, E F =., (9)_S^E_where F eR^x£>• Position embedding: An extra learnable embedding sclassG R1XDis prepended to F, c classresulting in. Then, the position embedding, denoted as Eposis applied to obtainFthe input of the transformer encoder, UT. Specifically,c class UT+ E pos ', (10)F(vj'1+ljx£>where Epose R. Please refer to [8] for more details on position embedding. • Transformer encoder: The transformer encoder in the preferred embodiment is comprised of two norm layers, a re-attention layer, and an FFNN head. It employs reattention mechanisms to discern relationships among mini-images, thereby enhancing system performance. Given the input UT, the output of the transformer encoder can be expressed asOoT= fT(UT; θT), (11)where fTis a transformer encoder neural network, and θTis the corresponding trainable parameters.
[0111] FFNN layers
[0112] FFNN layers: The collected informationserves as input of the FFNN 1. The first FFNN takes the measured information,, as its input and output some hidden features according to(12)where θF1are the training parameters.
[0113] The input of the second FFNN is uF2= [oT,oF1]. It outputs the estimated distance between the m -th device and the k -th AP,(13)
[0114] To train FPDNN, the ground truth distances between the m -th device and the k -th AP d^kare used as the labels. The loss function is defined as the mean square error (MSE) between the output of FPDNN and the label,L(E,θT,θF1,θF2) = E (14)
[0115] The parameters of FPDNN are updated by the stochastic gradient descent algorithm. However, other training methods may be possible in other embodiments.The Least Square Block
[0116] After obtaining the estimated distances from the m -th device to all its neighbour APs, I } w ’ the least square algorithm is used to estimate the location of the device. The outputof the least square can be expressed as?), (15)where / LS(•) is the least square algorithm, and Y^) gym^ is the location of the m -th device.
[0117] In addition, the Kalman filter could be applied to improve the localization accuracy in the current time interval by using the historical trajectory of the device.Synthetic Data Generation in Virtual Environments
[0118] The flow chart of pre-training is illustrated in Fig. 8. Considering that data samples in a real- world environment may not be enough for the training, a virtual environment is established to generate synthetic data. In different real-world environments, the data samples may followdifferent distributions. The generator model is first trained in the scenario 1 by the corresponding real-world data and floor-plan image. Then, the generator model uses floor-plan images in different target scenarios to generate synthetic data samples. The GNN and FPDNN are pretrained offline using synthetic data samples. In this way, the localization algorithm can be implemented in unseen environments without real-world data samples.Synthetic Data Generation in One Virtual Environment
[0119] A data sample of the disclosed localization algorithm includes input information and the corresponding label. The input information of the t -th sample is composed of node and edge features of a graph,, and the floor-plan image between the device and its nearby APs, I^. The label is the ground truth location of the device,.
[0120] Some methods to generate synthetic data (1) are very sensitive to the dielectric properties of the reflective surfaces of the obstacles in the radio environment; 2) require detailed three-dimensional (3D) geometric information; and 3) require high computing time for generating a sample, especially when there are a large number of paths between a device and an AP.
[0121] In the preferred embodiments, a machine learning model is used to generate the RTT distance and RSS for a set of locations of the device. Using machine learning may overcome some of the limitations mentioned above. Given the locations of the m -th device and the k -th AP, the ground-truth distance d^ and the floor-plan imagebetween them can be obtained. As shown in Fig. 6b, the deep learning model is composed of a DeepVIT 660 and two FFNNs 671, 672 and can be expressed as[cU-zIl] (16)where φDis a DNN in Fig. 6b, and the corresponding training parameters are θD. The outputm kand ym kare the synthetic RTT distance and RSS, respectively.
[0122] The generator is trained in a supervised manner with the following loss functionL (0D) = E(C’ y- - / J:.’ i2), a?)where and are the RTT distance and RSS of the same location in the real-world scenario. For example, to generate synthetic data in the office scenario in Fig. 9a and 9b, real- world data samples are collected on the purple trajectory and then used as labelled samples. After training, the trained generator model is used to estimate the RTT distance and RSS at any location in the office.Synthetic Data Generation in Different Environments
[0123] In other environments, such as the shopping mall and laboratory, the corresponding floor-plan images are used to generate synthetic data samples. There is no need to fine-tune the trained generator model in new scenarios. Note that the motivation for using synthetic data is to improve the diversity of the training samples for the disclosed localization algorithm. The distribution of synthetic data samples in an unseen scenario does not need to be the same as the real-world data samples. The difference between them increases the diversity of data samples and may help improve the generalization ability of the localization algorithm.Wi-Fi Platform and Data Sets
[0124] In this section, the Wi-Fi Platform and data sets in each scenario is introduced.Wi-Fi Platform
[0125] The Wi-Fi platform consists of eight APs, two line-tracking vehicles, four mobile phones, and a laptop. The details of these devices are given as follows:• Access points'. As shown in Fig. 11, Google Nest Wi-Fi APs, which support both 2.4 GHz and 5 GHz with IEEE 802.11 me standard, are used. It is noted that RTT and RSS are obtained from reference signals with carrier frequency at 5 GHz. In each of the three scenarios, the locations of all eight APs are fixed, and the heights of the tripods are 1.5 m.• Line-tracking vehicles'. Makeblock Ultimate 2.0 programmable robot kits were used to build line-tracking vehicles. The line-tracking vehicles are programmed to follow the black line at a constant speed. Each vehicle was equipped with two mobile phones. Thehorizontal mobile phones (C2 and C3 in Fig. 11) measure the RTT and RSS from nearby APs. The vertical mobile phones (Cl and C4 in Fig. 11) measure the RTT and RSS between each other.• Mobile phones'. The four mobile phones are either Google Pixel 6 Pro or Google Pixel 6a with Android 13. To measure the RTT and RSS between Cl and C4, Wi-Fi NanScan was used. To measure the RTT and RSS from C2 or C3 to all the APs simultaneously, an application was developed because the available applications can only measure RTT and RSS from one AP to the mobile phone at a time. The sampling interval between the mobile phones and the APs is set to 200 ms, and the sampling interval between Cl and C4 is set to 1 s.• Server. A desktop is used as a server. The detailed specifications for the server are Windows 11 Pro, Intel(R) Core(TM) i9-12900KF central processing unit (CPU) @ 3.20GHz, NVIDIA GeForce RTX 3090, 64GB Memory, and 1TB SSD.Typical Scenarios
[0126] The data sets were collected in the following three scenarios on the campus of the University of Sydney. All these three scenarios mixed with both LoS and nLoS paths:• Office: The data samples are collected in the office of the Centre of loT and Telecommunications. This area consists of one long office corridor and an office room, i.e., a 60 x 20 m2rectangular area.• Laboratory: The data samples are collected in the Mechanical Engineering Laboratory.The lab has multiple cement posts, iron guardrails, and an elevator, covering a 25x9 m2rectangular area.• Shopping mall: The data samples are collected in the shopping mall in the Wentworth building at the university campus. The experiment is carried out in the mall hallway outside the shops, i.e., a bank, a chemist, and a coffee shop.
[0127] All the scenarios and their corresponding floor-plan images are provided in Figs. lOa-c, which shows different environments and corresponding floor plans. Fig. 10a shows an office environment and the corresponding floor plan. Fig. 10b shows a shopping mall environment and corresponding floor plan. Fig. 10c shows a laboratory environment and corresponding floor plan. It is noted that furniture is not included in the floor plans, and the evaluation results are obtained in real-world environments with furniture.Data Sets
[0128] The data set of each scenario consists of RTT distance, RSS, and floor-plan image.
[0129] Data calibration and processing
[0130] There are offset components between the RTT distance and the ground truth distance, which lead to inaccurate RTT measurement. These components are static and can be removed from the measured RTT distance. Specifically, the true distance and the measured RTT distance between all the APs and the device are compared. The results show that the offset components lead to a nearly constant value between the real distance and the measured distance. Then, the average offset between the device and each AP is removed from the raw RTT distance.
[0131] In the data sets, the raw data is provided as well as the data sets after calibration and processing. It is noted that since the heights of APs and devices are fixed, the RTT distance is converted from in 3D space into a two-dimensional (2D) plane, which can be easily used in 2D localization algorithms. The original RTT distance in the 3D space is persevered in the data sets.
[0132] Illustration of data set
[0133] As previously shown, the data samples are collected in three scenarios. For each scenario, four files were created: 1) the floor plan of the scenario, 2) measured information from each mobile phone to all the APs, 3) measured information between Cl and C4, and 4) the locations and indices of the APs.
[0134] The “office” is used as an example to illustrate the data sets. The locations of all the APs, the trajectory of the vehicles, and the layout of the office can be found in Fig. 12. More specifically, Fig. 12 shows the locations of eight APs and the trajectory of the device, where the device moves along the trajectory in the real-world scenario. As shown in Table 1 below, a sample between the device and all APs consists of a time stamp, RTT distance, RSS, and the ground truth location of the device.Data samples between a mobile phone and all APsTime AP 1 AP 1 AP 1 AP N AP N AP N Ground Truth stamp RTT RTT RSS RTT RTT RSS(ms) (m) std (dBm) (m) std (dBm) ^yjt}(m)0 1.17 0.056 -45 14.75 0.11 -69 [31,13,75]267.10 0.58 0.055 -46 14.27 0.055 -68 [31 / 03,13.75]Table 1: Illustration of Data Samples.Measurement Errors
[0135] RSS measurement errors
[0136] From the theoretical path-loss models of wireless channels, the distance between two devices can be estimated based on the RSS. The RSS could be inaccurate due to stochastic interference and noise. In addition, since the communication environment is dynamic, walls and obstacles may cause severe signal attenuation and contribute to additional estimation errors. As a result, the communication distance estimated from the RSS is inaccurate.
[0137] RTT measurement errors
[0138] Based on the assumption that RTT equals the propagation time, it is possible to estimate the distance between two devices based on the RTT. However, the assumption does not hold in practice. Specifically, the unstable and low-precision oscillators may cause clock measurement noise. In addition, due to the Media Access Control (MAC) processing delay, the multi-path effect, refractive indices of different propagation mediums, and other offset components caused by antennas and chips, the measured RTT is longer than the propagation time. It is worth noting that the offset components are static, and it is possible to remove this part from the raw data.Experimental Results
[0139] In this section, extensive experimental results are presented, which verify the accuracy of the disclosed methods, and the performance of the disclosed methods is further compared with some existing baseline methods in terms of localization error. The experimental results include various embodiments of the disclosed methods (methods 200, 260) to determine the performance of such embodiments. These results also include the performance of the preferred embodiment of method 200.Hyper-Parameters1. FPDNN: The hyper-parameters of each block, e.g., FFNN, DeepVIT, of the FPDNN according to method 200 are summarized in Table 2 presented in Fig. 13. MSE was used as the loss function, Adam as the optimizer, with a learning rate of 1 x 10-3, and a weight decay of 1 x 10-5During the training of the FPDNN, the epoch was set to 50, and the batch size was set to 32.2. GNN: The hyper-parameters of different steps in GNN, e.g., message-passing and aggregation, are summarized in Table 2. It is noted that the loss function, the training epochs, and the batch size are the same as FPDNN. The Adam optimizer was used with a learning rate 2 x 10-4.3. DNN in the trained generator model: It is noted that the DeepVIT block and the FFNN3 in the synthetic data generation DNN have the same structure as the DeepVIT and FFNN 1 in FPDNN, respectively. In addition, the loss function, the optimizer, the number of epochs, and the batch size are the same as FPDNN.Performance Metrics
[0140] The localization error of the m -th device at the t -th time slot is defined as= √((x̂m(t)- xm(t))2+ (ŷm(t)- ym(t))2). (18)
[0141] From the cumulative distribution function (CDF)of, the mean absolute error (MAE) (with legend “MAE"), the median of the errors (with legend “50% CDF"), and the 90-th percentile CDF (with legend “ 90% CDF") can be obtained. In addition, the RMSE of the errors is evaluated according to the following expression(19)√(1 / Nt)where Ntis the number of testing samples.Training of the trained generator model in the office
[0142] To train the generator model, samples in the office are collected. As shown in Fig. 12, a line-tracking robot moves along the trajectory (with legend “Trajectory") for five circles. 50,000data samples in this scenario were obtained, and each data sample consists of the RTT and RSS information from an AP to the device. The samples collected were used in the first four circles to train the data generator model. After the training, the trained generator model can output RTT and RSS information at any location in the office, as illustrated in Fig. 9a and 9b.Localization in The Office
[0143] Training with real-world data samples
[0144] The performance of the disclosed method is first evaluated when it is trained with real-world data samples in the office environment. Specifically, 80 % of the samples are used for training and the rest 20 % of the samples for testing. The estimated locations using other localization methods (denoted as Method 1, Method 2 and Method 3) were also tested to determine a baseline and their performance results, along with the performance results of the disclosed method, are shown in Fig. 14. The results indicate that the estimated locations of the disclosed method are well-aligned with the ground-truth trajectory in Fig. 13, and the localization errors are much lower than the other three baselines.
[0145] To further illustrate the localization errors of different methods, the CDFs of localization errors for the different methods, as well as the disclosed method, are provided in Fig.15. With the disclosed method, the median of localization errors is 0.2 m, and the localization errors are lower than 0.44 m with a probability of 90%. For the other three baseline methods, their medians are larger than 1.5 m, and the 90-th percentile CDFs are larger than 3 m.Zero-shot learning without real-world data samples
[0146] Figs. 16a and 16b show the locations of APs, ground truth, and estimated locations in the office environment for different variations of the disclosed method. Fig. 17 shows the CDFs of the localization errors achieved by different variations of the disclosed method. In Figs. 16a, 16b and 17, GNN and FPDNN are trained with the synthetic data samples. To validate the effectiveness of the trained generator model, the deployment of APs was changed. The locations of APs in Fig. 12 are different from the locations of APs in Fig. 12. Since there is no real-world sample in the new scenario, this approach may be referred to as zero-shot learning.
[0147] The estimated locations in Fig. 16a are obtained with the GNN (either with or without Kalman Filter). With the help of the Kalman Filter, better pre-localization accuracy can be obtained, and the corresponding results serve as one part of the input in the FPDNN. In Fig. 16b, the final results obtained from the FPDNN are shown. The results indicate that by combining the GNN, the FPDNN, and the Kalman filter, the estimated locations are close to the ground truth in Fig. 12. The RMSE of the estimated location with the GNN and Kalman filter is 1.33 m, and the location RMSE with the GNN+FPDNN is 1.15 m. After smoothing by the Kalman filter, the RMSE further decreases to 0.95 m.
[0148] The CDFs of the location errors are illustrated in Fig. 17. From the CDFs, the medians of localization errors achieved by the four approaches mentioned above are 0.82 m, 0.95 m, 1.05 m, and 1.26 m, respectively. With a probability of 90%, the estimation errors of the four approaches are lower than 1.39 m, 1.73 m, 2.06 m, and 2.46 m, respectively. By combining the GNN, the FPDNN, and the Kalman Filter, the RMSE is 0.95 m, even with no sample in the new scenario. These results indicate that with the help of synthetic data samples, the disclosed method outperforms the three existing baselines in unseen scenarios.
[0149] Fine-tuning with part of real-world data samples
[0150] The localization accuracy may be further improved if parts of the real-world data samples are used to fine-tune the pre-trained GNN and FPDNN. As shown in Fig. 18a, the samples used in fine-tuning and testing are in green and red, respectively. The testing results are shown in Fig. 18b. To better illustrate the gaps among different approaches, the MAE, RMSE, median, and 90-th percentile CDF are provided in Fig. 19a and 19b. The results indicate that by fine-tuning the GNN and FPDNN, it is possible to obtain 10% ~ 20 % performance gain compared with zero-shot learning approaches. For example, the localization RMSE is 1.15 m for the zero-shot GNN and FPDNN, while the RMSE decreases to 1 m after fine-tuning.Nevertheless, the localization errors achieved by the zero-shot learning approaches are lower than the baselines.
[0151] By comparing Figs. 19a and 19b, it can be seen that if only part of the environment information is present, there is a performance loss. If the GNN and FPDNN are trained with data samples obtained in the first four circles and tested with the data samples obtained in the last circle, the RMSE is 0.3 m (with GNN, FPDNN, and Kalman filter). If the GNN and FPDNN are trained with the samples in green and tested with red samples, then the RMSE is 0.87 m (withGNN, FPDNN, and Kalman filter). This is likely because the green samples do not provide the environment information in the right part of the office.Experimental Results in Di fferent Scenarios
[0152] This part illustrates how to use the disclosed method in unseen scenarios with no real-world data sample.
[0153] To obtain experimental results in different scenarios, the generator model and the disclosed localization method were pre-trained with real-world data samples in the office environment. In the other two scenarios, shopping mall and laboratory, real-world data samples were not used to fine-tune the generator model and localization algorithm. To improve localization accuracy, the trained generator model uses floor-plan images to generate synthetic data in the new scenarios. The newly generated data samples are used to fine-tune the localization algorithm. As there is no real-world data sample, this approach is also referred to as zero-shot learning.
[0154] As a baseline, the localization algorithm is also with real- world data samples in the new scenarios. Like Fig. 18a, the GNN and FPDNN are fine-tuned with one part of the samples and test them with the other part of the samples.
[0155] The MAE, RMSE, median, and 90-th percentile CDF obtained in the shopping mall are provided in Fig. 20a. Specifically, the RMSE achieved by zero-shot learning with (or without) the Kalman filter is 1.33 m (or 1.19 m). After fine-tuning, the RMSE is reduced to 1.25 m (or 1.05 m). The performance gap is around 6.0% ~11.8%. Similar performance gaps are observed when using the other performance metrics. Nevertheless, the zero-shot learning approaches can reduce localization errors by around 31.1% ~ 55.4 % compared with the other baseline localization methods.
[0156] In Fig. 20b, the localization errors in the laboratory are estimated. The performance gaps between zero-shot learning and fine-tuned approaches are around 7.0 % ~16.8%. Zeroshot learning approaches can reduce the localization errors by around 38.5 % ~ 58.7 % compared with the other baseline localization methods. It is noted that in Fig. 20b, the 90-th percentile CDFs achieved by the zero-shot learning (GNN+FPDNN+KF) and the fined-tunedapproach are 1.55 m and 1.61 m, respectively. This is likely because the MSE was used as the loss function to fine-tune the GNN and FPDNN, and this may result in higher peak errors.Multi-device Localization
[0157] This subsection evaluates the disclosed method in multi-device scenarios, where M = 2 devices are considered. In the testing described herein, device 1 is connected to all APs, and device 2 is connected to two of the APs. The gaps between the zero-shot learning and other baselines are similar to the single-user scenarios. The generator model was trained in the office, and it was used to generate synthetic data samples in the shopping mall. The localization algorithm is only fine-tuned with the synthetic data samples, and real-world data samples in the shopping mall are only used for testing. For the laboratory scenario, there is no stable connection between the two devices due to the blockages of walls and stairs. Thus, it is reduced to single-user localization. The testing samples of the two devices are shown in Fig. 21. More specifically, Fig. 21 shows the locations of two devices and locations of six APs, where the devices move along the trajectories.
[0158] The results are provided in Table 3 below. With the existing baselines, the location of device 1 is first estimated and then used to estimate the location of device 2. Thus, the estimation error of device 1 is accumulated in the estimation error of device 2. This issue may have more significant impacts on localization accuracy when there are more than two devices. Specifically, with the trilateration method, the RMSE gap between device 1 and device 2 is 1.18 m. By using the GNN, the estimated locations of the two devices can be obtained directly. With the GNN and Kalman filter, the RMSE gap between device 1 and device 2 is 0.53 m. This observation implies that the disclosed method can alleviate the accumulation of estimate errors. In addition, compared with the localization accuracy in the single-user scenario in Fig. 20a, device 1 achieves lower localization errors in the multi-user scenario. For example, the RMSE of device 1 is 18.8% lower than the single-user scenario. This result indicates that the RTT and RSS information of device 2 helps to improve the localization of device 1 with the disclosed method.Methods MAE RMSE 50% CDF 90% CDF Trilateration 2.19 2.64 2.02 3.57 With GNN 1.22 1.57 1.05 1.75 Device 1With GNN+KF 1.03 1.23 0.92 1.71With GNN+FPDNN 0.9 1.08 0.78 1.65With GNN+FPDNN+KF 0.79 0.9 0.71 1.37 Trilateration 2.86 3.82 2.40 5.05 With GNN 1.86 2.47 1.39 3.54 Device 2 With GNN+KF 1.42 1.76 1.19 2.52 With GNN+FPDNN 1.77 2.39 1.28 3.52With GNN+FPDNN+KF 1.47 1.7 1.37 2.47 Table 3: Estimation errors in the shopping mall with two devices. Complexity Analysis
[0159] Since the GNN and FPDNN are trained offline, the training complexity has no impact on the localization performance. The complexity of online inference is analysed. The number of on-device parameters and computational complexity, which are evaluated by the number of floating-point operations (FLOPs) for processing one data sample of the proposed GNN and FPDNN, are provided in Table 4 below. It can be seen that the numbers of FLOPs and on-device parameters of the proposed GNN model are 2,704 and 1,381, respectively. For the proposed FPDNN, the numbers of FLOPs and on-device parameters are 23,522,913 and 863,097, respectively.
[0160] The average computation time for inference using the GNN and FPDNN with the NVIDIA GeForce RTX 3090 is evaluated. The processing time of the GNN and FPDNN is 2 ms and 6 ms, respectively. The existing baselines require more computation time on the GPU than the CPU. This is because GPUs are optimized for specialized computations, such as deep learning. Thus, for comparison, their processing time on the Intel(R) Core(TM) i9-12900KF CPU is also evaluated. Specifically, the processing time of the other baseline localization methods are 2.94 ms, 0.43 ms, and 26 ms. Compared with the minimum sampling interval between devices and APs (200 ms), the processing time of all the localization algorithms is acceptable.On-device computation (FLOPs) On-device parameters GNN 2,704 1,381FPDNN 23,522,913 863,097Table 4: The on-device parameters and computational complexity for proposed GNN and FPDNN.Experimental results of a car park environment
[0161] The disclosed method was further tested in a car park using a floor plan of the car park. For this experiment, the disclosed method was compared to different localization methods todetermine the accuracy of the disclosed methods. The other localization methods do not utilise machine learning. The results of this experiment are shown visually in Fig. 22. Fig. 22 illustrates the trajectory of a device estimated by the disclosed method (e.g., trajectory 2210) in comparison to the trajectory estimated by another localization method (e.g., trajectory 2220). Each sub-figure corresponding to a different localization method. As demonstrated in Fig. 22, the disclosed method achieves significantly higher accuracy and robustness compared to the other localization methods. Specifically, it was determined that the root mean square error (RMSE) of the disclosed method is at least 35% lower than that of the other localization methods.
[0162] It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the above-described embodiments, without departing from the broad general scope of the present disclosure. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.
Claims
CLAIMS:
1. A method for locating a device in an indoor environment, the method comprising: receiving distance measurements between the device and multiple access points within the indoor environment;determining an estimated location of the device from the distance measurements; for one or more of the multiple access points:generating a feature image from a floor plan of the indoor environment based on the estimated location of the device, the feature image corresponding to a physical region between the device and the access point; andapplying a trained machine learning model to the distance measurement between the device and the access point and the feature image to calculate an adjusted distance between the device and the access point; anddetermining an adjusted location of the device in the indoor environment based on the adjusted distance between the device and the one or more of the multiple access points.
2. The method of claim 1, wherein generating the respective feature image comprises cropping a rectangular section from the floor plan, wherein the estimated location of the device lies on an edge of the rectangular section.
3. The method of claim 2, wherein generating the respective feature image comprises rotating the rectangular section, such that each feature image has a same orientation.
4. The method of any one of the preceding claims, whereinthe trained machine learning model comprises a first sub-model and a second submodel; andapplying the trained machine learning model comprises applying the first sub-model to the respective feature image and applying the second sub-model to the respective distance measurement.
5. The method of claim 4, whereinthe trained machine learning model comprises a third sub-model;applying the trained machine learning model comprises applying the third sub-model to an output of the first sub-model and an output of the second sub-model; andan output of the third sub-model corresponds to an estimated distance between the device and the respective access point.
6. The method of claim 5, wherein the first, second and third sub-models are one of: a vision transformer; anda neural network.
7. The method of any one of the preceding claims, whereinthe device is one of multiple devices within the indoor environment; andreceiving distance measurements comprises receiving distance measurements between the multiple devices and receiving distance measurements between the multiple devices and the multiple access points.
8. The method of claim 7, wherein the method further comprises determining an estimated location of each of the multiple devices in the indoor environment based on the distance measurements.
9. The method of claim 8, wherein the method further comprises determining a graph based on the estimated locations, wherein nodes of the graph are indicative of the multiple devices and the multiple access points, and edges of the graph are indicative of the distance measurements.
10. The method of claim 9, wherein the method further comprises applying a trained graph model to the graph to determine improved estimated locations of the multiple devices, wherein generating the respective feature image is based on the corresponding improved estimated location of the device.
11. The method of any one of claims 7 to 10, wherein the method further comprises receiving historical distance measurements between the multiple devices and the multiple access points.
12. The method of claim 10, wherein determining the adjusted location of the device comprises applying a Kalman filter to the historical distance measurements.
13. The method of claim 11 or 12, wherein the method further comprises:creating multiple historical graphs using the historical distance measurements, wherein nodes of each of the multiple historical graphs are indicative of the multiple devices and the multiple access points, and edges of each of the multiple historical graphs are indicative of the historical distance measurements;for one or more of the multiple historical graphs, applying the trained graph model to the historical graph to generate an output; anddetermining the improved estimated locations based on the multiple outputs.
14. The method of any one of the preceding claims, wherein the method further comprises training, at least in part, a machine learning model using synthetic distance measurements to generate the trained machine learning model.
15. The method of claim 14, wherein the method further comprises generating the synthetic distance measurements by, for one or more of the multiple access points, applying a trained generator model to the feature image and a ground truth distance between the access point and one of the multiple devices.
16. The method of claim 15, whereinthe trained generator model comprises a first sub-model and a second sub-model; and applying the trained generator model comprises applying the first sub-model to the respective feature image and applying the second sub-model to the respective ground truth distance.
17. The method of claim 16, whereinthe trained generator model comprises a third sub-model;applying the trained generator model comprises applying the third sub-model to an output of the first sub-model and an output of the second sub-model; andan output of the third sub-model corresponds to one of the synthetic distance measurements.
18. The method of any one of claims 15 to 17, wherein the method further comprises training, at least in part, a generator model using ground truth distances and a set of training feature images from a training floor plan to generate the trained generator model.
19. Software that, when installed on a computer and executed by the computer, causes the computer to perform the method of any one of the preceding claims, or part thereof.
20. A system for locating a device in an indoor environment, the system comprising one or more processors configured to perform the method of any one of claims 1 to 18.