System and method for WiFi-based indoor positioning via unsupervised domain adaptation
Through the Widora system, the synthetic data is generated and the neural network is trained using the variational automatic encoder, which solves the problem of positioning instability of the WiFi positioning system under different users and environment changes, and achieves higher positioning accuracy and environmental adaptability.
Patent Information
- Application Number
- CN202080070253.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-26
- Filing Date
- 2020-09-02
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-09-02
AI Technical Summary
The existing WiFi-based positioning system has unstable positioning accuracy when facing changes in different users and environments, and it is difficult to adapt to new users and environment changes without the need for additional labeling data.
Widora, a domain adaptive positioning system based on clustering assumptions, uses a variational automatic encoder to generate synthetic data and train the neural network through joint reconstruction-classification structure to adapt to new users and environmental changes, and realize unsupervised domain adaptation through data expanders and domain adaptive classifiers.
The positioning accuracy under different users and environment changes was improved, with the average F1 score increasing by 17.8%, the worst-case accuracy increased by 20.2%, and the positioning performance in a changing environment was 23.1% better than the existing technology.
Smart Images

Figure CN114514759B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to location-aware applications, and more particularly, to an electronic device for location-aware applications and an operating method thereof. Some embodiments relate to WiFi (also known as "Wi-Fi" in the communications industry). Background Art
[0002] Location-aware applications can provide convenience and infotainment. Current applications estimate a user's location with a granularity of more than one meter. Location sources are sometimes unavailable indoors (e.g., GPS) or coarse-grained (e.g., user registration). WiFi signals are sometimes used in positioning systems. WiFi characteristics are also sensitive to the following factors: 1) the body shapes of different users, and 2) objects in the background environment. Therefore, WiFi fingerprint-based systems are vulnerable to 1) new users with different body shapes, and 2) daily changes in the environment, such as opening / closing doors. Summary of the Invention
[0003] Technical Solution
[0004] A method for determining a location by a system including an electronic device having a WiFi transceiver and a neural network-based AI model is provided herein. The method includes: receiving, via the WiFi transceiver, a first fingerprint associated with a first person and a first environment, where the first person is not a registered user of the system; using a shared feature extraction layer of the neural network-based AI model and obtaining a first set of feature data based on an operation on the first fingerprint; using a reconstruction layer of the neural network-based AI model and determining a reconstruction loss associated with the first person and the first fingerprint based on the first set of feature data from the shared feature extraction layer; when the reconstruction loss is equal to or lower than a threshold, using a classification layer of the neural network-based AI model and classifying the first fingerprint based on the first set of feature data from the shared feature extraction layer to obtain a first estimated location tag of the first person, and outputting an indicator of the first estimated location tag of the first person in a visual or audible manner; when the reconstruction loss is higher than the threshold, performing registration of the first person.
[0005] A method executed by an electronic device is also provided herein, where the electronic device includes a neural network-based AI model. The method includes: jointly training a feature extraction layer, a reconstruction layer, and a classification layer of the neural network-based AI model based on a reconstruction loss and a clustering loss, where the reconstruction layer is configured to provide the reconstruction loss and the classification layer is configured to provide the clustering loss; processing a fingerprint to obtain an augmented fingerprint; performing an operation on the augmented fingerprint by the feature extraction layer to generate a code; classifying the code by the classification layer to obtain an estimated location tag; and providing an application output based on the estimated location tag.
[0006] The present disclosure also provides an electronic device, which includes: one or more memories, wherein the one or more memories include instructions; a WiFi transceiver; a neural network-based AI model; and one or more processors, wherein the one or more processors are configured to execute the instructions to: receive, via the WiFi transceiver, a first fingerprint associated with a first person and a first environment, wherein the first person is not a registered user of the electronic device, obtain a first set of feature data using a shared feature extraction layer operating on the first fingerprint, wherein the neural network-based AI model includes the shared feature extraction layer, determine a reconstruction loss associated with the first person and the first fingerprint using a reconstruction layer and based on the first set of feature data from the shared feature extraction layer, wherein the neural network-based AI model includes the reconstruction layer, when the reconstruction loss is at or below a threshold: classify the first fingerprint using a classification layer and based on the first set of feature data from the shared feature extraction layer to obtain a first estimated location tag of the first person; wherein the neural network-based AI model includes the classification layer and outputs an indicator of the first estimated location tag of the first person through a visual or audible device, and when the reconstruction loss is higher than the threshold, perform registration of the first person. In one embodiment, the one or more processors are further configured to execute the instructions to: perform the registration by: obtaining an identifier of the first person via a user interface of the electronic device, obtaining a ground truth location tag of the first person via the user interface, and obtaining a second fingerprint associated with the ground truth location tag; and providing an application output based on the first estimated location tag or the ground truth location tag.
[0007] The present disclosure also provides a machine-readable storage medium storing instructions that, when executed, cause at least one processor of an electronic device to execute any of the methods provided herein.
[0008] The present disclosure also provides an apparatus, including: an input end configured to receive a first plurality of fingerprints and a second plurality of fingerprints, wherein the first plurality of fingerprints includes a first fingerprint, and the first fingerprint is associated with a first person at a first location, and wherein the second plurality of fingerprints includes unlabeled data and a second fingerprint; a data augmenter including at least one first processor and at least one first memory, wherein the data augmenter is coupled to the input end and configured to generate variational data based on the first plurality of fingerprints; a feature extraction layer coupled to the data augmenter, including at least one second processor and at least one second memory, wherein the feature extraction layer is configured to provide a first encoding and a second encoding, the first encoding is associated with the variational data, and the second encoding is associated with the second plurality of fingerprints, wherein the feature extraction layer is arranged in a neural network-based AI model; a reconstruction layer coupled to the feature extraction layer, including at least one third processor and at least one third memory, wherein the reconstruction layer is configured to generate a reconstructed first fingerprint and a reconstructed second fingerprint, wherein the reconstruction layer is arranged in a neural network-based AI model; a classification layer including at least one fourth processor and at least one fourth memory coupled to the feature extraction layer, wherein the classification layer is configured to predict a second location of a second person, wherein a statistical performance measure for optimizing the prediction of the second location of the second person in a location marking domain, wherein the classification layer is arranged in a neural network-based AI model; and a controller including at least one fifth processor and at least one fifth memory, wherein the controller is configured to: perform a first update of the classification weights of the classification layer, the feature extraction weights of the feature extraction layer, and the reconstruction weights of the reconstruction layer based on a first difference between the reconstructed first fingerprint and the first fingerprint, and perform a second update of the feature extraction weights and the reconstruction weights based on a second difference between the reconstructed second fingerprint and the second fingerprint.
[0009] Other aspects will be set forth in part in the following description, and will in part be obvious from the description, or may be learned by practice of the embodiments disclosed in this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other aspects, features, and aspects of the embodiments of the present disclosure will become more apparent from the following description in conjunction with the accompanying drawings, in which:
[0011] Figure 1A is a block diagram of an apparatus that functions as a classification system and an application system according to an embodiment.
[0012] Figure 1B is a schematic diagram of a geographical area with an exemplary location marking according to an embodiment.
[0013] Figure 2 is according to an embodimentFigure 1A The classification system receives unlabeled data and Figure 1A The application system provides a logical flow of application results.
[0014] Figure 3 is according to an embodiment Figure 1A The classification system receives labeled data, augments the labeled data, updates weights, and Figure 1A The application system provides a logical flow of application results.
[0015] Figure 4 is according to an embodiment Figure 1A The classification system receives unlabeled data and labeled data, augments the labeled data, operates on a shared feature extraction layer, updates weights, and Figure 1A The application system provides a logical flow of application results.
[0016] Figure 5 is a block diagram showing an enlarged representation of Figure 1A the classification system, showing loss values from each of a feature extraction layer, a reconstruction layer, and a classification layer according to an embodiment.
[0017] Figure 6A and Figure 6B shows exemplary WiFi fingerprint data and augmented fingerprint data according to an embodiment.
[0018] Figure 7A shows an exemplary labeled WiFi fingerprint of label p1 according to an embodiment.
[0019] Figure 7B shows a first exemplary augmented WiFi fingerprint based on Figure 7A the WiFi fingerprint according to an embodiment.
[0020] Figure 7C shows a second exemplary augmented WiFi fingerprint based on Figure 7A the WiFi fingerprint according to an embodiment.
[0021] Figure 8 shows an exemplary set of feature data corresponding to Figure 7B the first augmented WiFi fingerprint according to an embodiment.
[0022] Figure 9A shows an exemplary labeled WiFi fingerprint of label p2 according to an embodiment.
[0023] Figure 9B shows a first exemplary augmented WiFi fingerprint based on Figure 9A the WiFi fingerprint according to an embodiment.
[0024] Figure 9C shows a WiFi fingerprint of a second exemplary expansion based on the Figure 9A WiFi fingerprint.
[0025] Figure 10 shows an exemplary set of feature data corresponding to the first expanded WiFi fingerprint of Figure 9B according to an embodiment.
[0026] Figure 11 shows an exemplary controller for an update layer according to an embodiment.
[0027] Figure 12A shows the first exemplary embodiment of the Figure 1A classification system and Figure 1A application system according to an embodiment.
[0028] Figure 12B shows the second exemplary embodiment of the Figure 1A classification system and Figure 1A application system according to an embodiment.
[0029] Figure 12C shows the third exemplary embodiment of the Figure 1A classification system and Figure 1A application system according to an embodiment.
[0030] Figure 12D shows the fourth exemplary embodiment of the Figure 1A classification system and Figure 1A application system according to an embodiment.
[0031] Figure 12E shows the fifth exemplary embodiment of the Figure 1A classification system and Figure 1A application system according to an embodiment.
[0032] Figure 12F shows the sixth exemplary embodiment of the Figure 1A classification system and Figure 1A application system according to an embodiment.
[0033] Figure 13 shows an exemplary logic flow for initializing the Figure 1A classification system, querying a user, and operating a layer using a set of feature data extracted based on an expanded fingerprint according to an embodiment.
[0034] Figure 14 shows the initialization according to an embodiment Figure 1AExemplary logic flow of a classification system that operates and trains layers using a set of feature data extracted from unlabeled fingerprints.
[0035] Figure 15 Illustrates the initialization according to an embodiment Figure 1A Exemplary logic flow of a classification system that operates and trains layers using a set of feature data extracted from unlabeled fingerprints, and obtains an application result.
[0036] Figure 16 Illustrates the initialization according to an embodiment Figure 1A Exemplary logic flow of a classification system that operates and retrains layers using a set of feature data extracted from unlabeled fingerprints, and obtains an application result.
[0037] Figure 17 Is a block diagram of an exemplary autoencoder according to an embodiment.
[0038] Figure 18 Is a block diagram of an exemplary feature extraction layer according to an embodiment.
[0039] Figure 19 Is a block diagram of an exemplary classification layer according to an embodiment.
[0040] Figure 20 Is a block diagram of an exemplary reconstruction layer according to an embodiment.
[0041] Figure 21 Is a block diagram of an electronic device that implements a device for location-aware applications according to an embodiment.
[0042] Figure 22 Illustrates an exemplary training algorithm according to an embodiment.
[0043] Figure 23 Illustrates an exemplary environment for evaluating an embodiment. Detailed Description
[0044] Aspects of the present disclosure address at least the above problems and / or disadvantages and provide at least the following advantages.
[0045] In a first aspect, a positioning method determined by a system including an electronic device having a WiFi transceiver and a neural network-based AI model is disclosed. The method includes: receiving, via the WiFi transceiver, a first fingerprint associated with a first person and a first environment, where the first person is not a registered user of the system; using a shared feature extraction layer of the neural network-based AI model and based on an operation on the first fingerprint, obtaining a first set of feature data; using a reconstruction layer of the neural network-based AI model and based on the first set of feature data from the shared feature extraction layer, determining a reconstruction loss associated with the first person and the first fingerprint; when the reconstruction loss is equal to or lower than a threshold, using a classification layer of the neural network-based AI model and based on the first set of feature data from the shared feature extraction layer, classifying the first fingerprint to obtain a first estimated position marker of the first person, and outputting an indicator of the first estimated position marker of the first person in a visual or audible manner; when the reconstruction loss is higher than the threshold, performing registration of the first person.
[0046] Optionally, the registration includes: obtaining, via a user interface of the system, an identifier of the first person, obtaining, via the user interface, a ground truth position marker of the first person, and obtaining, via the WiFi transceiver, a second fingerprint associated with the ground truth position marker; and providing an application output based on the first estimated position marker or the ground truth position marker.
[0047] Optionally, the method further includes training the weights of the system based on a third fingerprint associated with the first environment and a second person who is a first registered user of the system to obtain updated weights.
[0048] Optionally, the training weights includes: receiving, by the system, a third fingerprint associated with the first environment and the second person; augmenting the third fingerprint to form a first augmented fingerprint; updating the weights of the system based on the first augmented fingerprint.
[0049] Optionally, determining the reconstruction loss includes: receiving the first fingerprint via the WiFi transceiver; using the shared feature extraction layer based on the updated weights to determine a first set of feature data; based on the first set of feature data, determining a first reconstructed fingerprint of the first person; based on the first fingerprint and the first reconstructed fingerprint, determining the reconstruction loss of the first reconstructed fingerprint.
[0050] Optionally, the application output includes automatically turning off the electronic device based on identifying that the room is empty.
[0051] Optionally, the application output includes beamforming sound by the electronic device to the first estimated position marker of the first person.
[0052] Optionally, the application output includes tilting a display screen of the electronic device by the electronic device according to the first estimated position marker of the first person.
[0053] Optionally, the application output includes updating the augmented reality display for the first person based on the first estimated position marker of the first person.
[0054] Optionally, the application output includes providing an advertisement to the personal electronic device of the first person based on the first estimated position marker of the first person.
[0055] In a second aspect, a method performed by an electronic device is disclosed. The electronic device includes a neural network-based AI model, and the method includes: jointly training a feature extraction layer, a reconstruction layer, and a classification layer of the neural network-based AI model based on a reconstruction loss and a clustering loss, wherein the reconstruction layer is configured to provide the reconstruction loss, and the classification layer is configured to provide the clustering loss; processing a fingerprint to obtain an augmented fingerprint; performing an operation on the augmented fingerprint by the feature extraction layer to generate a code; classifying the code by the classification layer to obtain an estimated position marker; and providing an application output based on the estimated position marker.
[0056] Optionally, the application output includes automatically turning off the electronic device based on identifying that the room is empty.
[0057] Optionally, the application output includes the electronic device beamforming sound to the first estimated position marker of the first person.
[0058] Optionally, the application output includes the electronic device tilting the display screen of the electronic device according to the estimated position marker.
[0059] Optionally, the application output includes updating the augmented reality presentation for the person based on the estimated position marker.
[0060] Optionally, the application output includes providing an advertisement to the personal electronic device of the individual based on the estimated position marker.
[0061] In a third aspect, an electronic device includes: one or more memories, wherein the one or more memories include instructions; a WiFi transceiver; a neural network-based AI model; and one or more processors, wherein the one or more processors are configured to execute the instructions to: receive, via the WiFi transceiver, a first fingerprint associated with a first person and a first environment, wherein the first person is not a registered user of the electronic device, obtain a first set of feature data using a shared feature extraction layer that operates on the first fingerprint, wherein the neural network-based AI model includes the shared feature extraction layer, determine a reconstruction loss associated with the first person and the first fingerprint using a reconstruction layer and based on the first set of feature data from the shared feature extraction layer, wherein the neural network-based AI model includes the reconstruction layer, when the reconstruction loss is at or below a threshold: classify the first fingerprint using a classification layer and based on the first set of feature data from the shared feature extraction layer to obtain a first estimated location marker for the first person; wherein the neural network-based AI model includes the classification layer, and output an indicator of the first estimated location marker for the first person via a visual or audible device, when the reconstruction loss is above the threshold, perform registration of the first person, wherein the one or more processors are configured to perform registration by: obtaining an identifier of the first person via a user interface of the electronic device, obtaining a ground truth location marker of the first person through the user interface, and obtaining a second fingerprint associated with the ground truth location marker; and providing an application output based on the first estimated location marker or the ground truth location marker.
[0062] In a fourth aspect, a computer-readable storage medium stores instructions that are configured to cause a processor to perform the following operations: receive, via a WiFi transceiver, a first fingerprint associated with a first person and a first environment, where the first person is not a registered user of the electronic device; obtain a first set of feature data using a shared feature extraction layer that operates on the first fingerprint, where an AI model based on a neural network includes the shared feature extraction layer; use a reconstruction layer and based on the first set of feature data from the shared feature extraction layer, determine a reconstruction loss associated with the first person and the first fingerprint, where an AI model based on a neural network includes the reconstruction layer; when the reconstruction loss is equal to or lower than a threshold: use a classification layer and based on the first set of feature data from the shared feature extraction layer, classify the first fingerprint to obtain a first estimated location marker of the first person, where an AI model based on a neural network includes the classification layer, and output an indicator of the first estimated location marker of the first person via a visual or audible device; when the reconstruction loss is higher than the threshold, perform registration of the first person, where the one or more processors are configured to perform the registration by the following steps: obtain an identifier of the first person via a user interface of the electronic device, obtain a ground truth location marker of the first person via the user interface, and obtain a second fingerprint associated with the ground truth location marker; and provide an application output based on the first estimated location marker or the ground truth location marker.
[0063] In a fifth aspect, an apparatus includes: an input configured to receive a first plurality of fingerprints and a second plurality of fingerprints, wherein the first plurality of fingerprints includes a first fingerprint, and the first fingerprint is associated with a first person at a first location, wherein the second plurality of fingerprints includes unlabeled data and the second plurality of fingerprints includes a second fingerprint; a data augmenter including at least one first processor and at least one first memory, the data augmenter being coupled to the input, wherein the data augmenter is configured to generate variational data based on the first plurality of fingerprints; a feature extraction layer coupled to the data augmenter, including at least one second processor and at least one second memory, wherein the feature extraction layer is configured to provide a first encoding and a second encoding, the first encoding being associated with the variational data, and the second encoding being associated with the second plurality of fingerprints, wherein the neural network-based AI model includes the feature extraction layer; a reconstruction layer coupled to the feature extraction layer, including at least one third processor and at least one third memory, wherein the reconstruction layer is configured to produce a reconstructed first fingerprint and a reconstructed second fingerprint, wherein the neural network-based AI model includes the reconstruction layer; a classification layer coupled to the feature extraction layer, including at least one fourth processor and at least one fourth memory, wherein the classification layer is configured to predict a second location of a second person, wherein a statistical performance measure for optimizing the prediction of the second location of the second person in a location labeling domain, wherein the neural network-based AI model includes the classification layer; and a controller including at least one fifth processor and at least one fifth memory, wherein the controller is configured to: perform a first update of the classification weights of the classification layer, the feature extraction weights of the feature extraction layer, and the reconstruction weights of the reconstruction layer based on a first difference between the reconstructed first fingerprint and the first fingerprint, and perform a second update of the feature extraction weights and the reconstruction weights based on a second difference between the reconstructed second fingerprint and the second fingerprint.
[0064] Optionally, the second update is a joint update.
[0065] Optionally, the second update is based on a clustering loss metric, wherein the clustering loss metric promotes avoiding high-density regions when adjusting the classification weights.
[0066] Optionally, the controller is further configured to perform a third update of the classification weights based on a difference between the reconstructed second fingerprint and the second fingerprint.
[0067] Embodiments of the present disclosure provide a localization system called Widora based on domain adaptation using clustering assumptions. Widora is capable of 1) localizing different users with labeled data from only one or two example users, and 2) localizing the same user in a changing environment without labeling any new data. To achieve these, Widora integrates two main modules. It first employs a data augmenter that introduces data diversity using a variational autoencoder (VAE). Then, it trains a domain-adaptive classifier that uses a joint reconstruction-classification structure to adapt itself to newly collected unlabeled data. The embodiments of Widora provided herein are evaluated with real-world experiments. The results show that when testing on unlabeled users, Widora increases the average F1 score by 17.8% and improves the worst-case accuracy by 20.2%. Additionally, when applied to a changing environment, Widora's performance is 23.1% better than the prior art.
[0068] Location systems deliver intelligent environments that not only sense the location of a human body but also respond accordingly. They provide users with more information about their surroundings and, at the same time, help service providers deliver services / content in a location-aware manner. As a result, many mobile applications have been realized, including location-based social networks, point-of-interest (POI) recommendations, and local business / product searches. However, current sources of location information have many constraints and limitations. For example, the Global Positioning System (GPS) signal, which is currently the main source of location information, is not available indoors and is constantly affected outdoors due to the urban canyon effect. User check-ins, as well as locations embedded in tweets and other online posts, are not available all the time and have a very low resolution (house level or lower). None of them can support emerging applications such as mobile advertising targeting, cashierless shopping, indoor navigation / tracking, smart home automation, and location-based augmented and virtual reality (AR / VR), which require (sub-)meter-level location information at any time and anywhere.
[0069] To meet the usability and resolution requirements, WiFi signals have been studied. WiFi signals have a wavelength of 12.5 cm on the 2.4 GHz frequency channel (or 6.0 cm on the 5 GHz channel). They are affected by the position / movement of a person as small as the wavelength level (i.e., decimeter level). By using WiFi signals, it is possible to provide decimeter-level location information using dedicated or commercial off-the-shelf (COTS) devices.
[0070] Existing WiFi-based positioning systems can be divided into two major categories: anchor-assisted systems and fingerprinting systems. Anchor-assisted systems use WiFi access points (APs) as positioning anchors and calculate the distances and / or angles from the user (or his / her device) to these anchors. A major drawback is that these systems typically locate the device rather than the user. Therefore, they can hardly handle the situation where the user has multiple devices or is not carrying a device at this time. Another limitation is that the coordinates and orientations of the WiFi AP antennas need to be accurately measured in advance. This measurement process is troublesome and error-prone for users.
[0071] Fingerprinting systems sense the physical presence of a user by passively learning the changes in WiFi propagation caused by the human body. These systems do not assume user-device association and do not require pre-measurement of AP location / orientation. Therefore, they are more flexible, practical, and easier to deploy. To complete their positioning tasks, when the user is in different places, these systems first extract readings of WiFi propagation characteristics, such as received signal strength indicator (RSSI) and channel state information (CSI). These WiFi propagation characteristics, especially CSI, are very sensitive to human location. Therefore, by pairing CSI readings with location tags / coordinates, the embodiments provided herein establish sub-meter-level location fingerprints. Once deployed, the fingerprinting system listens passively to the WiFi signals in the air, extracts real-time CSI readings, and matches them with the fingerprints to estimate / predict the user's location in real time.
[0072] However, these CSI fingerprints are sensitive not only to location but also to human body shape and the surrounding environment. Therefore, a positioning system trained with fingerprints of some users may not work for other users with different body shapes. At the same time, even for the same user, daily changes in the environment, such as closing the blinds, may cause a significant deterioration in the positioning accuracy.
[0073] Among all the readings of WiFi propagation characteristics, CSI is the most informative and representative. For example, the information contained in a CSI reading is at least 60 times more than that contained in an RSSI reading (another widely used reading). Next, the reasons are explained by discussing some relevant details of the WiFi standard and its implementation on WiFi chips.
[0074] The physical layer of WiFi follows the IEEE 802.11 standard, which uses orthogonal frequency division multiplexing (OFDM) technology to divide the WiFi frequency channel into 64 or more frequency subcarriers. It improves the frequency efficiency by supporting the simultaneous transmission of multiple data signals on multiple orthogonal subcarriers. In addition, the standard also supports spatial multiplexing. The signal flow from one antenna on the transmitter side to another antenna on the receiver side is called a spatial stream, or simply a stream.
[0075] The complex signal transmitted on the i-th stream and the j-th subcarrier can be represented as ai,j. This signal propagates through the WiFi channel, where the channel response is described as h1,j. Being altered during propagation, the signal finally arrives at the receiver as bi,j. In the frequency domain, this propagation process is modeled as a modification of the transmitted signal, i.e.,
[0076] bi,j = hi,j · ai,j.
[0077] As previously mentioned, the WiFi signal can reach the receiver through multiple paths. Therefore, hj is actually a superposition representation of the channel responses on multiple paths. To reconstruct ai,j, the receiver must reverse the propagation process by eliminating hi,j from bi,j.
[0078] To estimate each hj, the WiFi chip executes certain algorithms, which report the CSI readings as the estimation results. Taking the Intel 5300 WiFi chip as an example, for each spatial stream, these chips estimate the complex-valued CSI values on 30 subcarriers, which altogether correspond to 60 floating-point values. Extracting only 1 real-valued RSSI has less information, which reflects the aggregated signal strength of all subcarriers. The CSI values contain more detailed information about the WiFi signal propagation and, therefore, can establish location fingerprints with much higher resolution compared to other WiFi characteristics.
[0079] While being sensitive to the human body position, the WiFi CSI fingerprint (as well as RSSI and other fingerprints) is also sensitive to 1) the human body shape and 2) changes in the physical environment. In one example, when standing at the same location, a tall person blocks the WiFi propagation path 0, while a short person does not block the WiFi propagation path 0. At the same time, a large-sized person absorbs more signal power on path 1, while a small-sized person absorbs less signal power. Therefore, different WiFi CSI readings can be observed when different users stand even at the same location. Another example of the impact of environmental changes occurs when the same user stands at the same location. When the door is closed, it creates a propagation path 0 by reflecting the WiFi signal. However, when the door is open, this path disappears. As a result, the CSI readings change with the environmental changes, even when the same user stands at the same location.
[0080] To address this inconsistency problem and provide robust positioning information for different users in a changing environment, the embodiments provided herein disclose a WiFi-based positioning system, Widora, which is domain-adaptive based on exploiting clustering assumptions and is capable of positioning new users in a changing environment without labeling any new data, thus saving a significant amount of system calibration time and effort.
[0081] Problem (Challenge) 1: It is impractical to collect WiFi fingerprints covering all possible user and environmental variations, let alone their combinations.
[0082] The implementation with the technical name "Widora" provided herein addresses Challenge 1 from a data perspective by augmenting the collected data. It uses a data augmenter to learn the statistics of the collected data and then generates a large amount of synthetic data based on these statistics. The data augmenter consists of multiple variational autoencoders (VAEs), and each variational autoencoder focuses only on the CSI readings at one location. The VAE first maps the CSI readings at its corresponding location into a latent space described by a normal distribution. From this latent space, the VAE generates synthetic CSI readings that are similar but different from the original CSI readings. Then, the locations associated with this VAE are used to automatically label these synthetic CSI readings, thus becoming synthetic fingerprints. By adjusting the standard deviation of the normal distribution (describing the latent space), the similarity between the synthetic fingerprints and the collected fingerprints is controlled. In this way, Widora can cover more people using fingerprints collected from only one or two users.
[0083] Problem (Challenge) 2: Once the system is deployed, it is cumbersome and sometimes even impossible to label new data (i.e., associate newly recorded WiFi CSI readings with locations).
[0084] To label WiFi CSI readings, the readings need to be associated with the true location (i.e., the location where the user stood when the readings were collected). However, when a system is deployed, it is difficult to collect this true information. First, it is cumbersome and unfriendly to require people to manually collect and label data for new users or environmental changes. Second, installing auxiliary systems (e.g., LIDARs) is expensive and brings more challenges in terms of synchronization, cross-modal alignment, etc.
[0085] Widora addresses Challenge 2 from a model perspective by extracting and learning from invariant features of the user / environment. Widora treats the example user (or original environment) as the source domain and the new user (or changed environment) as the unlabeled target domain. To adapt to the target domain, Widora includes a domain adaptation classifier to leverage the potentially large amount of unlabeled data. The domain adaptation classifier adopts a special neural network (NN) design with a Joint Classification-Reconstruction (JCR) structure, which consists of two paths: a classification path and a reconstruction path. On the one hand, the classification path learns from CSI fingerprints (i.e., labeled data in the source domain) and predicts the location for new CSI readings. On the other hand, the reconstruction path extracts a common representation of the labeled data in the source domain and the unlabeled data in the target domain by reproducing the labeled data in the source domain and the unlabeled data in the target domain with the same autoencoder (AE) structure. The classification path is configured to use the representation produced by the reconstruction path by sharing the same feature extraction layer. In this way, classification can adapt to new users and / or environmental changes using the representation updated by the unlabeled data. Widora outperforms the original version of the JCR structure proposed by M. Ghifary et al. as DRCN, and its predecessor WiDo proposed by X. Chen et al. mainly by applying the following three techniques: i) Although the reconstruction path in DRCN is only trained with target domain data, the reconstruction path in Widora (and WiDo) learns a common representation on both the source domain and the target domain. By doing so, the extracted features are more focused on the commonalities between the two domains and do not emphasize the domain differences; ii) Although the classification path in WiDo only focuses on the classification loss of the labeled data, the classification path in Widora uses an additional clustering loss to drive the decision boundary away from high-density regions and thus reduces the class ambiguity of the unlabeled data; iii) Instead of using an optional training algorithm as in DRCN and WiDo, Widora adopts a more efficient joint training algorithm.
[0086] Widora is different from domain adversarial neural networks (DANNs) and their variants. Although DANNs can leverage data without class labels, they still need to know whether the data points come from the source domain. DANNs are still limited to specific types of labeled data where the label is the domain rather than the class. Unfortunately, in practical applications, it is difficult for the system to collect these true domain labels (i.e., the identity of the user or environmental changes). In contrast to DANNs, Widora performs domain adaptation with completely unlabeled data that has neither class labels nor domain labels. Therefore, Widora can adapt to unknown users or unpredictable environmental changes without the need for manual intervention.
[0087] Experiments have been conducted on a positioning test platform established on COTS WiFi devices. The experiments were carried out in a furnished room ( Figure 23 ), where the positions in the room were scattered at a minimum interval of 70 cm. To evaluate the impact of user diversity, 9 volunteers with different heights, weights, and genders were invited to participate in the experiment. The results show that in the presence of unknown users, Widora is 17.7%, 14.9%, and 17.8% higher than the baseline AutoFi in terms of average recall rate, accuracy, and F1-score respectively, and the worst projection accuracy among different populations is also improved by 20.2%. At the same time, compared with WiDo, the average recall rate, accuracy, and F1-score of Widora are increased by 6.0%, 5.7%, and 6.0% respectively. Embodiments of WiDo are also introduced in this paper. For WiDo, neither augmented fingerprints nor clustering loss is used.
[0088] In addition, to test the robustness to environmental changes, different changes were introduced to the environment, such as closing the glass door, opening the metal shutters, and removing chairs from the room. CSI data were collected before and after each environmental change, and the data collected before (after) the change were processed as source (target) domain data. These source and target domain data were used to evaluate different positioning systems. Compared with the baselines AutoFi and Wido, the average positioning accuracy of Widora is increased by 23.1% and 15.6% respectively.
[0089] A fingerprint recognition system has been developed in combination with (deep) AI / ML models. The core idea is to record the WiFi propagation characteristics, label them with the positions of users / devices during the recording process, and use the WiFi characteristics and labels to train the AI / ML model. The combination of WiFi propagation characteristics and their corresponding position labels is simply referred to as WiFi positioning fingerprint or fingerprint.
[0090] Device-oriented fingerprints have been widely used. These fingerprints capture the way signal propagation looks when the device carried by the user is at a specific location. For example, Z. Yang et al. adopted the WiFi RSSI fingerprint between smartphones and WiFi APs in their LiFS system, which used an accelerometer integrated in the smartphone for automatic labeling of fingerprints. P. Wilk et al. developed a crowdsourcing mechanism to collect WiFi RSSI fingerprints. J. Yang et al. applied ensemble learning techniques to further improve the fidelity and robustness of crowdsourced RSSI fingerprints. To handle expired fingerprints, C. Chu et al. created a calibration method based on particle filters to update RSSI fingerprints. To obtain higher resolution and more informative WiFi features, for example, S. Sen et al. used WiFi CSI to develop the PinLoc system, which achieved 89% meter-level positioning accuracy using a clustering-based classification algorithm. Using WiFi CSI, X. Wang et al. combined offline training with online probability data fusion based on radial basis functions. H. Chen et al. designed Confi, which projects CSI into feature images and feeds them into a convolutional neural network (CNN) that models localization as an image classification problem. A. Zarell et al. have utilized WiFi multipath delay distribution to obtain better performance than RSSI. Drawbacks: However, these device-oriented fingerprints show several limitations, including the ambiguity of user devices and the burden of always carrying the device / sensor.
[0091] It has been demonstrated that device-free fingerprints more or less achieve the same performance as device-oriented fingerprints while providing user-centric and convenient positioning services. These device-free fingerprints describe how the user's location / movement changes the signal propagation between two fixed devices. By learning the association between RSSI / CSI changes and location, the system can subsequently directly locate the user's body. For example, L. Chang et al. adopted RSSI device-free fingerprints in the proposed system FitLoc, which projects the fingerprints into a latent space to capture environment-independent features. However, the implementation of Fitloc only works in the line-of-sight (LOS) scenario of dedicated devices. Using COTS devices, J. Xiao et al. designed a system called Pilot, which uses the maximum a posteriori probability (MAP) algorithm to estimate location from WiFi CSI fingerprints.
[0092] S. Shi et al. developed a Bayesian filtering method using CSI fingerprints and applied the algorithm to untrained users. K. Chen et al. addressed this problem from the perspective of computer vision and developed a convolutional neural network for localization and an interpretation framework for explaining the features learned by the CNN. To perform effective feature selection, T. Sanam et al. applied generalized inter-view and intra-view discriminant correlation analysis to extract distinguishable CSI features for localization. Disadvantage: The above device-free fingerprint recognition system does not fully consider the fingerprint inconsistency problem, so it may perform poorly in the case of new users or environmental changes.
[0093] To address this fingerprint inconsistency problem, several fingerprint database update methods have been proposed. For example, L. Chang et al. developed iUpdater, a Regularized Singular Value Decomposition (RSVD) method, to reconstruct the RSSI fingerprint matrix using several newly labeled fingerprints. X. Chen et al. developed AutoFi, which uses a polynomial source-target mapping function and an autoencoder to calibrate the localization profile in an unsupervised manner. The fingerprint update of AutoFi is triggered by a specific event where the region of interest is detected as empty. X. Rao et al. proposed a time-triggered update method that replaces expired fingerprints with new fingerprints using a fully connected NN. Disadvantage: These fingerprint update methods rely on certain (event- or time-based) indications of whether the environment / user has changed. In other words, these methods require inherent domain labels to distinguish CSI fingerprints before and after the environment / user change. However, these inherent domain labels are either inaccurate or unavailable in practical scenarios, which may lead to significant performance degradation.
[0094] Other CSI detection systems: In addition to localization, CSI data has also been used in other human-centered applications. W. Xi et al. proposed a CSI method for counting people by establishing a monotonic relationship with the percentage of non-zero elements in the extended CSI matrix.
[0095] Y. Zheng et al. developed Widar 3.0, which converts the Doppler frequency shift into a human coordinate velocity distribution map, which is then fed into a combined RNN-CNN for human pose recognition. A. Pokunuru et al. designed the user recognition system NeuralWave, which extracts gait-induced signatures from CSI data and uses a deep CNN to distinguish the gait biometrics of different people.
[0096] Korany et al. proposed XModal-ID, which learns the mapping from CSI readings to video clips in order to extract user identification information. To eliminate the environmental dependence in activity recognition, W. Jiang et al. proposed an adversarial network to extract the commonalities shared by different CSI domains.
[0097] In the embodiments disclosed in this paper, for fully unsupervised domain adaptation without location markers or domain markers, WiDo is based on a joint classification-reconstruction NN structure. This structure allows WiDo to extract general features from unlabeled data without knowing the domain.
[0098] In the embodiments disclosed in this paper, a clustering loss and a more efficient joint training algorithm are provided and labeled as Widora. While the clustering loss forces the decision boundary away from high-density regions, the joint training algorithm converges quickly with only a small amount of data. Therefore, Widora further improves the positioning performance on WiDo. Both methods are useful.
[0099] Since this disclosure allows various changes and many examples, the embodiments will be shown in the drawings and described in detail in the written description. However, this is not intended to limit this disclosure to the practice mode, and it should be understood that all changes, equivalents, and alternatives that do not depart from the essence and technical scope of this disclosure are included in this disclosure.
[0100] In the description of the embodiments, when the detailed explanation of the related art may unnecessarily obscure the essence of this disclosure, the detailed explanation of the related art is omitted. In addition, the numbers used in the description of the specification (e.g., first, second, etc.) are identification codes for distinguishing one element from another.
[0101] In addition, in this specification, it should be understood that when elements are "connected" or "coupled" to each other, the elements can be directly connected or coupled to each other, but they can also be connected or coupled to each other through intermediate elements therebetween, unless otherwise specified.
[0102] In this specification, regarding the elements represented as "units" or "modules", two or more elements can be combined into one element, or one element can be divided into two or more elements according to the subdivided functions. In addition, each element described below can additionally perform some or all of the functions performed by another element in addition to its own main function, and some of the main functions of each element can be completely performed by another component.
[0103] In addition, in this specification, "image" or "picture" can represent a still image, a moving image including a plurality of consecutive still images (or frames), or a video.
[0104] In addition, in this specification, a deep neural network (DNN) or a CNN is a representative example of an artificial neural network model that mimics the brain nerves, and is not limited to an artificial neural network model using an algorithm.
[0105] In addition, in this specification, a "parameter" is a value used during the operation of each layer that forms a neural network, and can include, for example, weights used when an input value is applied to an operation expression. Here, the parameter can be represented in matrix form. The parameter is a set of values as a training result, and can be updated with separate training data when necessary.
[0106] Throughout the disclosure, the expression "at least one of a, b, or c" means only a, only b, only c, both a and b, both a and c, both b and c, or all of a, b, and c.
[0107] As described above, the physical layer of WiFi follows the IEEE 802.11 standard, and the IEEE 802.11 standard uses orthogonal frequency division multiplexing (OFDM) technology to divide the WiFi frequency channel into 64 or more frequency subcarriers. This improves the frequency efficiency by supporting the simultaneous transmission of multiple data signals on multiple orthogonal subcarriers. In addition, the standard also supports spatial multiplexing. The separated data signals can be transmitted from multiple antennas and / or transmitted to multiple antennas. The signal flow between one antenna on the transmitter side and another antenna on the receiver side is called a spatial stream, or simply a stream.
[0108] As described above, the complex signal transmitted on the i-th stream and the j-th subcarrier can be expressed as ai,j. This signal propagates through the WiFi channel, and its channel response is described as h1,j. Changed during propagation, the signal finally arrives at the receiver as bi,j. In the frequency domain, this propagation process is modeled as a change in the transmitted signal, that is:
[0109] bi,j = hi,j · ai,j (1)
[0110] As previously mentioned, the WiFi signal can reach the receiver through multiple paths. Therefore, hj is actually a superimposed representation of the channel responses on multiple paths. To reconstruct ai,j, the receiver must reverse the propagation process by eliminating hi,j from bi,j. To estimate each hj, the WiFi chip executes certain algorithms, which report the CSI readings as the estimation result, such as the Intel 5300 WiFi chip. For each spatial stream, these chips estimate the complex-valued CSI values on 30 subcarriers, which in total correspond to 60 floating-point values. The CSI value contains more detailed information about the WiFi signal propagation than the RSSI.
[0111] In some embodiments, a user interacts with an electronic device. The electronic device can be a television (TV), and the electronic device can include a classification system and an application system. These can be collectively referred to as the "system" hereinafter.
[0112] The system can ask a person whether they are a new user. If a new user answers, for example, "Yes, I am", the system can record the user's data and perform various levels of updates to support the user. In some embodiments, the data is a user profile. The user's data can include name, gender, height, and weight. In some embodiments, the system can ask the user whether they wish to create a new profile. If a new user answers, for example, "This is my profile", the system accepts the profile.
[0113] The embodiments provided herein provide a device-free localization service.
[0114] In the setup phase, an owner can collect WiFi CSI data (fingerprints) and tag them with location markers.
[0115] After setup, if the system is applied to a new user, the localization accuracy of the new user may generally be lower than that of the owner.
[0116] If more data of a new user is collected, the localization accuracy of the new user increases. The new user does not need to tag the newly collected data.
[0117] In some embodiments, the tagged data passes through an augmenter. The tagged data and the untagged data pass through the same feature extractor. The tagged data passes through a classifier to generate an estimated location (e.g., a soft value or a tag). The untagged data first passes through a reconstructor.
[0118] Figure 1A The input WiFi fingerprint 1-1 provided to the classification system 1 (items 1-2) is shown. The classification system 1 provides an output Y1 (items 1-3) that can be a probability vector. Each element of the vector is associated with a location marker, such as Figure 1B the location marker. Y1 is input to the application system 1-4 that provides an application result E (items 1-5).
[0119] Figure 1B The environment 1-7 is shown. The environment can be, for example, a room with a table, chairs, doors, windows, etc. Inside the environment, locations are associated with a set of markers Omega; Omega includes p1, p2, p3, p4, p5, p6, p7, p8, and p9 (these are exemplary). The X coordinate 1-9 and the y coordinate 1-8 are location coordinate axes.
[0120] Classification system 1 can process fingerprints from registered users or guests (unknown users) and provide output Y1. Y1 can then be processed by an application system to provide application result E. Thus, Figure 1A and Figure 1B is an overview of the human position awareness service.
[0121] Figure 2 shows the exemplary logic 2-0 associated with Figure 1A - Figure 1B At 2-0, the system is initialized. At 2-1, the weights of the layer are trained based on registered users. The exemplary layer is shown in Figure 5 and described later. At 2-2, unlabeled data is input, and the shared feature extraction layer is operated to generate feature data (item 2-21). Generally, the feature data is represented by the label "Z" (ZA for augmentation and ZU for unlabeled). The feature data 2-21 is input to the reconstruction layer, and at 2-3, the reconstruction layer reconstructs the fingerprint associated with the feature data 2-21. The reconstruction layer has been trained previously (see, for example, Figure 11 ). It should be noted that generally there is no corresponding ground truth fingerprint for the feature data 2-21. The reconstruction layer evaluates the error measurement of the reconstructed fingerprint and generates a reconstruction loss. This loss can be based on conditional entropy. Due to the calculation of the loss, the weights can be updated even without a ground truth fingerprint.
[0122] If the reconstruction loss is higher than a predetermined threshold, the logic flows to 2-4, and the user interface is activated to register the unknown person. During the registration process, the ground truth position label of the unknown person (item 2-41) can be obtained.
[0123] If the reconstruction loss is lower than the predetermined threshold, the logic flows to 2-5. The classification layer operates on the feature data 2-21 to generate the value Y1 and the estimated position label (item 2-51). The estimated position label can be obtained from Y1 using the argmax function or the softmax function.
[0124] Finally, the logic (via 2-4 or 2-5) flows to 2-6, and the application process is operated based on the ground truth position label 2-41 or the estimated position label 2-51 to generate the application result E (item 2-61).
[0125] Additional logic flows are provided herein to describe further details and / or variations of the logic 2-0.
[0126] Figure 3 shows the exemplary logic 3-0. At 3-1, the system is trained based on registered users, such as a person who owns Figure 1A an electronic device.
[0127] At 3-2, labeled data is input. The labeled data is the fingerprint and tells where the person is atFigure 1B The ground truth location marker of the location in the environment. Data can be input through the dialogue process between the electronic device and the person. For example: 1. "I want to input data" (from the person); 2. "Where are you now?" (from the electronic device); 3. "I am sitting on the sofa" (from the person); 4. The electronic device collects WiFi fingerprints from one or more WiFi devices (such as access points) and associates the fingerprint with the person's location marker "sitting on the sofa".
[0128] In 3-21, the marked data is augmented (or autoencoded, see Figure 17 ) to generate augmented fingerprints. The augmented data can include determining the statistical moments of the marked fingerprints, such as the mean and variance, and generating random realizations from a distribution with the determined mean and variance. The random realizations are called augmented data. Several such random realizations are obtained from a single determined distribution (see Figure 6A , 6B , Figure 7A , Figure 7B , Figure 7C , Figure 9A , Figure 9B , Figure 9C ); thus, generally the augmented fingerprints include several fingerprint waveforms.
[0129] Then, in 3-4, the feature extraction layer operates on the augmented fingerprints and generates feature data (item 3-41). Each input fingerprint is associated with a feature data vector (or a set of features of the data). In some embodiments, the XA that generates ZA is a fingerprint that generates a feature vector, so an X of a marker can generate many ZA vectors. Generally, an X will be augmented into multiple XA. Then one XA will generate one ZA. Generally, there is a feature data vector for the marked fingerprint, and there is a feature data vector for each random realization that constitutes the augmented fingerprint. Therefore, Figure 8 corresponding to Figure 7B (see below).
[0130] The feature data 3-41 flows to the reconstruction layer and the classification layer. The reconstruction layer has the true marker of the marked fingerprint 3-21 and generates a loss for the training layer (updating the weights, item 3-51, and this action can be performed after 3-71 described below).
[0131] The classification layer operates on the feature data 3-41 in 3-6 and generates Y1 ( Figure 3 not shown in
[0132] ) and the estimated location marker 3-61.
[0133] Figure 4 shows the Figure 2 and Figure 3 There is some overlapping logical flow 4-0. At the beginning, the system has been trained (as shown in 4-1). More than one activity occurs in the logical flow 4-0 because unlabeled data (unlabeled fingerprints) are input (at 4-2) and labeled data are input (at 4-3). The labeled data are augmented at 4-4 to produce augmented fingerprints.
[0134] The feature extraction layer operates on the unlabeled data and the augmented fingerprints, generating a set of feature data ZU from the unlabeled data (item 4-51) and a set(s) of feature data ZA from the augmented fingerprints (item 4-52).
[0135] The operations of 4-6, 4-7 (producing label 4-71) and 4-8 (producing 4-81) are similar to Figure 3 the corresponding events in logical 3-0.
[0136] Figure 5 Further details of the training of classification system 1 ( Figure 5 item 5-3 in
[0137] are provided. Item 5-3 can be implemented as, for example, a neural network-based AI model. The labeled fingerprint f1 (item 5-1) is input to 5-3. Similarly, the unlabeled fingerprint f2 (item 5-2) is input to 5-3. Classification system 1 (item 5-3) processes these and produces a loss value L1 (item 5-4) that is used as training data (item 5-5).
[0138] Figure 5 An enlarged image depicting the internal functions of classification system 1 is shown in Figure 5 These functions can be implemented as software executed on a CPU, dedicated hardware, off-the-shelf software, or any combination configured according to Figure 17)), the feature extraction layer TF (item 5-32, also known as F), its output encodings ZA (item 5-33) and ZU (item 5-34). ZA and ZU are input into the classification layer TC (item 5-35, also known as ΘC) and the reconstruction layer TR (item 5-36, also known as ΘR). TC outputs a loss value LC (item 5-37), TR outputs a loss value LR (item 5-38), and TF outputs a loss value LF (item 5-39). These loss values are processed individually or jointly to train the layers. See Figure 10 。
[0139] Figure 6A Shows the fingerprint of the tag at position P1 and five random realizations based on the statistics of the fingerprint of the tag at position P1.
[0140] Figure 6B Shows the fingerprint of the tag at position P2 and five random realizations based on the statistics of the fingerprint of the tag at position P2.
[0141] In Figure 6A 、 Figure 6B 、 Figure 7A 、 Figure 7B 、 Figure 7C 、 Figure 9A 、 Figure 9B and Figure 9C the ordinate is the magnitude squared value of the CSI and the abscissa is the subcarrier index (frequency).
[0142] Figure 7A Shows the fingerprint of the tag at P1, Figure 7B and Figure 7C shows the random realizations generated based on the statistics (moments) of the fingerprint of the tag at P1 Figure 7A of.
[0143] Figure 8 Shows the group of feature data corresponding to Figure 7B the augmented data, in Figure 8 the ordinate is a dimensionless measure related to the information and the abscissa is the feature vector index.
[0144] Except for the tag p2, Figure 9A 、 Figure 9B and Figure 9C is similar to Figure 7A 、 Figure 7B and Figure 7C , Figure 10 is similar to Figure 8 。
[0145] Figure 11 Shows the value fed back for training.
[0146] With respect to Figure 5Similar to the description of , the labeled fingerprint f1 is input to the expander 11-1 which outputs the variational data 11-11 (also called the expanded fingerprint). The feature extraction layer TF outputs the codes ZA 11-21 and ZU 11-22. The unlabeled fingerprint f2 is directly input to the TF.
[0147] The controller 11-5 processes the loss values 11-31 (LC), 11-41 (LR) and 11-21 (LF) to generate feedback functions 11-51, 11-52 and 11-53. These functions are used to update TR (item 11-4), TC (item 11-3) and TF (item 11-2), respectively.
[0148] Figure 12A , Figure 12B , Figure 12C , Figure 12D , Figure 12E and Figure 12F Various implementations of the application system are shown.
[0149] Figure 12A An input fingerprint f1 or f2 (item 12-10) is shown as input to the classification system 1 (item 12-1). Figure 12B , Figure 12C , Figure 12D , Figure 12E and Figure 12F The classification system 1 is implied and not shown.
[0150] Figure 12A An exemplary embodiment is shown in which application system A (item 12-2) produces an output EA (item 12-21). In this example, EA is an event that estimates that no one is in the room and turns off the TV.
[0151] Figure 12B An exemplary embodiment is shown in which application system B (item 12-3) produces output EB (item 12-31). In this example, EB is the event of estimating the position of a person in the room and beamforming the sound from the TV to the person's position. Beamforming the sound means creating a constructive audio pattern of sound waves so that a person listening at a position corresponding to the estimated position marker experiences a pleasant sound effect.
[0152] Figure 12C An exemplary embodiment is shown in which an application system C (item 12-4) produces an output EC (item 12-41). In this example, EC is an event of estimating the position of a person in the room and tilting the TV to a position corresponding to the estimated position marker. Tilt means rotating the TV screen around the vertical axis and rotating the TV screen around the horizontal axis to point the screen at the person.
[0153] Figure 12DAn exemplary embodiment is shown in which application system D (item 12-5) generates output ED (item 12-51). In this example, ED is an event of estimating the position of a person in a room and updating an augmented reality display for the person at the position marked corresponding to the estimated position. The augmented reality display provides a three-dimensional image for a person wearing, for example, a special pair of display glasses or viewer in front of their eyes on their head.
[0154] Figure 12E An exemplary embodiment is shown in which application system E (item 12-6) generates output EE (item 12-61). In this example, EE is an event of estimating the position of a person in a room and updating an advertisement for the person at the position marked corresponding to the estimated position. In some embodiments, the advertisement display screen is configured by the application system. In some embodiments, the application system controls the display screen to automatically turn on / off based on whether there is a person in the room or whether there is a person in the geographical area of interest within the viewing distance of the display screen.
[0155] Figure 12F An exemplary embodiment is shown in which application system F (item 12-7) generates output EF (item 12-71). In this example, EF is an event of estimating the position of a person in an environment (such as a store) and performing a financial transaction for the person at the position corresponding to the estimated position marker.
[0156] Figure 13 An exemplary logic 13-0 involving querying a user is shown. At 13-1, layers TF, TC, and TR are initialized, for example, by randomization.
[0157] At 13-2, query whether the user is existing or new, obtain the profile of the new user, and conduct a conversation with the user to determine their position, associate the measured fingerprint with the reported position, and input the marked data. At 13-3, augment the marked data to generate an augmented fingerprint.
[0158] At 13-4, operate TF to generate ZA. Input ZA into TR and TC. Subsequent events can proceed as in the Figure 2 , Figure 3 or Figure 4 logic flow of
[0159] Figure 14 An exemplary logic 14-0 involving an unknown user is shown. At 14-1, for example, layers TF, TC, and TR have been trained based on the marked data. At 14-2, obtain a fingerprint without a marker.
[0160] At 14-3, process the unmarked data by TF to generate ZU. Input ZU into TR and TC (perhaps). Subsequent events can proceed as in the Figure 2 ,Figure 3 or Figure 4 as in the logic flow of Figure 2 The threshold event of 2-3 in
[0161] Figure 15 Exemplary logic 15-0 involving an unknown user is shown. At 15-1, for example, layers TF, TC, and TR have been trained based on labeled data and unlabeled data. At 15-2, a fingerprint is obtained without labeling.
[0162] At 15-3, the unlabeled data is processed by TF to produce ZU. ZU is input into TR. Subsequent events can proceed as Figure 2 shown (the threshold event of 2-3 in Figure 2 ). The estimated location label 2-51 is obtained based on Y115-31 (similar to Figure 2 the estimated label 2-51). At 15-4, the application process is run, producing the application result E item 15-41. For example, see the exemplary applications in Figure 12A , Figure 12B etc.
[0163] Figure 16 The logic 16-0 of
[0164] shows an example of a retraining event. At 16-1, the classification system 1 has been trained. At 16-2, unlabeled data is input. At 16-3, TF and TR are operated. The loss measurement at TR is high, higher than the retraining threshold.
[0164] At 16-4, starting with one of several weight initial conditions, the weights are retrained. The weight initial conditions are: i) existing weights, ii) random weights, iii) zero weights. In some embodiments, the randomized weights for initialization are the preferred embodiment.
[0165] In some embodiments, classification is then performed. For example, at 16-5, TC is operated using the retrained weights, and Y1 (item 16-51) (or location label) is obtained. At 16-6, the application process is then operated and the application result E (item 16-61) is produced.
[0166] Figure 17 An exemplary structure of an autoencoder (item 17-0) for performing data augmentation such as Figure 7B , 7C , 9B, and 9C is provided.
[0167] Data augmentation aims to augment the diversity and sample size of training data to improve the final performance of the learned model. It is commonly used in image classification tasks and has also shown effectiveness for several other applications including machine fault detection and recommendation.
[0168] There are several ways to implement data augmentation. General data augmentation methods include translation, rotation, shearing, and flipping of the original data. A generative model can be a data augmentation method. Data augmentation with GAN can be used in few-shot learning settings.
[0169] In some embodiments, data generated by a VAE model is used to expand the training dataset.
[0170] Figure 18 There is provided an exemplary structure of a feature extraction layer TF (Item 18-0) for extracting features as Figure 8 and Figure 10 shown. Further discussion of the feature extraction layer is given below.
[0171] Figure 19 There is provided an exemplary structure of a classification layer TC (Item 19-0) for generating a probability Y1. Further discussion of the classification layer is given below.
[0172] Figure 20 There is provided an exemplary structure of a reconstruction layer TR (Item 20-0) for reconstructing fingerprints. Further discussion of the reconstruction layer is given below.
[0173] Figure 21 is a block diagram of an electronic device 2100 that implements an electronic device for location-aware applications according to an embodiment.
[0174] Referring to Figure 21 , the electronic device 2100 includes a memory 2105, a processor 2110, an input interface 2115, a display 2120, and a communication interface 2125. Figure 1A The classification system 1 and / or the application system of
[0175] can be implemented as the electronic device 2100. The processor 2110 controls the electronic device 2100 overall. The processor 2110 executes one or more programs stored in the memory 2105.
[0176] The memory 2105 stores various data, programs, or applications for driving and controlling the electronic device 2100. The programs stored in the memory 2105 include one or more instructions. The programs (one or more instructions) or application programs stored in the memory 2105 can be executed by the processor 2110.
[0177] The processor 2110 can execute Figure 1A any one or any combination of the operations of the electronic device of
[0178] The input interface 2115 can receive user input and / or data such as 2D images. The input interface 2115 can include, for example, a touch screen, a camera, a microphone, a keyboard, a mouse, or any combination thereof.
[0179] The display 2120 can obtain data from, for example, the processor 2110 and can display the obtained data. The display 2120 can include, for example, a touch screen, a television, a computer monitor, etc.
[0180] The communication interface 2125 sends data to and receives data from other electronic devices and can include one or more components that enable communication to be performed via a local area network (LAN), a wide area network (WAN), a value-added network (VAN), a mobile wireless communication network, a satellite communication network, or a combination thereof.
[0181] A block diagram of the electronic device 2100 is provided as an example. Each component in the block diagram can be integrated, added, or omitted according to the specifications of the actually implemented electronic device 2100. That is, two or more components can be integrated into one component, or one component can be divided into two or more components as needed. In addition, the functions performed by the respective blocks are provided to illustrate embodiments of the present disclosure, and the operations or devices of the respective blocks do not limit the scope of the present disclosure.
[0182] Additional description is now provided to further explain Figure 1A to Figure 23 the location awareness embodiments described in
[0183] Suppose a WiFi device is using its NTX transmit antennas to send signals. Another WiFi device uses its NRX receive antennas to receive these signals. In this sense, there are a total of NSS = NTX * NRX spatial streams. Suppose each WiFi frequency channel is divided into NSC frequency subcarriers by OFDM. Let hi,j[t] denote the complex WiFi CSI value of the i-th stream on the j-th subcarrier collected at time t. Let hi[t] denote the complex vector of the CSI on the i-th stream collected at time t, i.e.,
[0184] hi[t] = (hi,1[t], hi,2[t], …, hi,NSC[T]). (2)
[0185] Let H[t] denote the complex vector of the CSI on all subcarriers and streams collected at time t:
[0186] H[t] = (h1[t], h2[t], …, hNSS[t])
[0187] = (h1,1[t], …, h1,NSC[t], h2,1[t], …, h2,NSC[t], …, hNSS,1[t], …, hNSS,NSC[t]). (3)
[0188] In one example, consider the Intel 5300 NIC, and take one transmit antenna and three receive antennas as an example. In other words, NTX = 1, NRX = 3, and NSC = 30, so H[t] has a total of 90 elements.
[0189] The data input to the positioning system at time t is denoted as x[t]. In this paper, the CSI amplitude is used as this data, that is
[0190] x[t] = (|h1,1[t]|,..., |h1,NSC[t]|, |h2,1[t]|,..., |h2,NSC[t]|,..., |hNSS,1[t]|,..., |hNSS,NSC[t]|). (4)
[0191] where |x| represents the magnitude (i.e., absolute value) of a complex number. Figure 1B Ω, "Omega" represents the finite set of positions supported by the proposed framework. y[t] is the Ω element that marks the true position of the user at time t. In addition, yt uses one - hot encoding to obtain the one - hot labeled vector y[t], where the dimension is the number of supported positions NΩ. Y = y[t] represents all the true position labels.
[0192] If an input data point (i.e., a CSI amplitude vector such as Figure 5 f1 or f2) is associated with a position label, it is called a labeled data point, denoted as xL[t] (e.g., f1). Otherwise, it is called an unlabeled data point, denoted as xU[t] (e.g., f2). Let XL = xL[t] represent all the collected labeled data points, and let XU = xU[t] represent all the collected unlabeled data points. Then the position fingerprint f[t] is defined as the tuple of the labeled data point and its position label, that is
[0193] F[t] = (xL[t]; y[t]) (5)
[0194] Let F = {f[t]} represent the set of all collected fingerprints. The implementation provided in this paper uses F and XU to train a classifier and accurately predict y[t] given x[t].
[0195] The data augmenter first receives the collected fingerprints F (i.e., the labeled data) and generates synthetic fingerprints. The synthetic fingerprint fS[t] is defined as:
[0196] fS[t] = (xS[t]; yS[t]) (6)
[0197] where xS[t] is a synthetic data point from the synthetic dataset xS, and yS[t]Ω is the synthetic location associated with the label of xS[t]. Additionally, let YS = yS[t] denote all synthetic labels, and let FS = fS[t] denote the set of all synthetic fingerprints.
[0198] The data augmenter can be considered as the first neural network, e.g., Neural Network 1.
[0199] Then the final output FA of the data augmenter is the disjoint union of FS and F, i.e.,
[0200] FA = F (disjoint union) FS (7)
[0201] These augmented fingerprints FA and the unlabeled data XU will be used to train the domain - adaptive classifier. To achieve domain - adaptive capabilities, the classifier uses a Joint Classification Reconstruction (JCR) structure. As Figure 3 shown, the classifier consists of two paths: a classification path and a reconstruction path. The classification path contains a feature extraction layer followed by a classification layer. This path takes as input data points (labeled or unlabeled) and predicts location labels. The reconstruction path shares the same feature extraction layer as the classification path and then diverges into separate reconstruction layers. This path takes as input data points (labeled or unlabeled) and reproduces it as similarly as possible.
[0202] The domain - adaptive classifier can be considered as the second neural network, e.g., Neural Network 2.
[0203] The first system can consist of only Neural Network 2.
[0204] In some embodiments, the system can include Neural Network 1 and Neural Network 2.
[0205] For the classification path, when the augmented fingerprint fA[t] has been created, this path passes the data part xA[t] of the augmented fingerprint through the feature extraction layer to compute the feature vector zA[t] (corresponding to ZA for example Figure 5 ).
[0206] In Figure 1A to Figure 21 , according to this feature vector ZA, the classification layer predicts the location probability vector ŷA[t] represented as Y1. Let ŶA = {ŷA[t]} denote the predicted probability vectors for all augmented data points in XA. The classification loss of the classifier is defined as:
[0207] LA(YA, ŶA) = ∑t lA(yt[t], ŷt[t]) (8)
[0208] where lA[t] is the cross - entropy loss for a single data point, as follows:
[0209] 1A(yA[t], y^A[t]) = -kρk[t] logρ^k[t] (9)
[0210] ρk[t] and ρ^k[t] are the k-th elements of yA[t] and y^A[t] respectively. For the classification path, when the unlabeled data point xU[t] arrives (unlabeled fingerprint), this data point passes through the layer in the same way as the labeled data point xA[t]. The feature vector zU[t] is extracted and the predicted y^U[t] is generated. Since there is no ground truth label for this unlabeled data point, the classification loss is not calculated as in Equation (9). Instead, the implementation calculates the conditional entropy as
[0211]
[0212] where represents the k-th element of y^U[t]. Moreover, the clustering loss of the domain adaptation classifier is defined as:
[0213] LU(Y^U) = 1U(y^U[t]) (11)
[0214] In the fields of domain adaptation and semi-supervised / unsupervised learning, the use of this clustering loss follows the clustering assumption. This clustering assumption states that the decision boundary should avoid high-density regions. By minimizing the clustering loss in Equation (11), the classifier is more reliable for unlabeled data, thus moving the decision boundary away from them.
[0215] For the reconstruction path, XA and XU will pass through it in the same process. The following implementation takes the unlabeled data point xU[t] as an example. If the unlabeled data point xU[t] arrives, the reconstruction path first passes xU[t] through the shared feature extraction layer to obtain the feature vector zU[t]. Given this zU[t], the reconstruction layer outputs x^U[t], which attempts to reproduce xU[t] as much as possible. Similarly, for the labeled data point xA[t], the intermediate feature vector zA[t] is generated and x^A[t] is reproduced based on this vector. The implementation uses X^A = {x^A[t]} and X^U = {x^U[t]} to represent the reconstructed data from the augmented and unlabeled data respectively. Then, the implementation defines the reconstruction loss of the domain adaptation classifier as:
[0216] LR = td(xA[t], x^A[t]) + td(xU[t], x^U[t]) (12)
[0217] where d(,) is the distance metric between two vectors. In this article, d(,) is implemented as the Euclidean distance.
[0218] By sharing the feature extraction layer between two paths, the classification layer uses the features representing both labeled and unlabeled data. In this way, the classifier can adapt itself to the unlabeled domain, where the unlabeled data comes from the unlabeled domain. The unlabeled data is used to adapt the classifier to the new domain. Thus, Challenge 2 is solved.
[0219] The augmenter ( Figure 17 ) generates synthetic fingerprints from a limited amount of collected (labeled) fingerprints.
[0220] The data augmenter implements NΩ variational autoencoders (VAEs). See Figure 17 , where each VAE is dedicated to one of the NΩ supported positions. The fingerprints collected at the k-th position are used to train the k-th VAE, and the k-th VAE generates synthetic fingerprints only for that position.
[0221] As Figure 17 shown, the VAE has an encoder-decoder structure.
[0222] The encoder of the VAE aims to encode the input into a normal distribution N(0,γkI), where γk is a pre-selected hyperparameter and I is the identity matrix. Accordingly, the decoder attempts to reproduce the input from this normal distribution. After that, the implementation feeds the samples extracted from N(0,γkI) into the decoder to generate a large amount of synthetic data that is similar but different from the training data.
[0223] In the training phase, the data augmenter divides the fingerprints into NΩ subsets according to the position labels of the fingerprints. Use Fp=k to represent the fingerprint subset at the k-th (k = 1, 2, ……, NΩ) position, and represent the data part of these fingerprints as Xp=k. Then Xp=k is fed into the k-th VAE for training. The k-th VAE takes a batch Xb from Xp=k and feeds this batch into its encoder. For clarity, the data flow is described by taking a single data point xk[t]∈Xb as an example.
[0224] From the 90-dimensional input xk[t], a 10-dimensional vector of the mean μk[t] and a 10-dimensional vector of the standard deviation σk[t] are extracted by the encoding. These two vectors define a 10-dimensional normal distribution, from which the latent vector ζk[t] is obtained. To maintain the backpropagation mechanism, reparameterization is applied to generate ζk[t]:
[0225] ζk[t] = μk[t] + εk[t] * σk[t], (13)
[0226] where εk[t] is obtained from a normal distribution N(0, I), and I is 10-dimensional. Equation (13) is implemented by inserting two element-wise arithmetic nodes in the network model. Recall that the task of the encoder is to match the distribution of ζk[t] with the normal distribution N(0, γkI). Therefore, the loss between the above distributions is defined as the Kullback-Leibler divergence.
[0227] IKLD[t] = D(N(μk[t], σk[t]), N(0, γkl)). (14)
[0228] Then, the output of the encoder (i.e., ζk[t]) is passed to the decoder, which generates the vector x^k[t] to reconstruct the input xk[t]. The reconstruction loss 1 is defined as the Euclidean distance between xk[t] and x^k[t]:
[0229] lB[t] = d(xk[t], x^k[t]) (15)
[0230] The loss function of the entire VAE model is defined as the sum of the distribution loss and the reconstruction loss, i.e.,
[0231] lVAE[t] = lKLD[t] + lB[t] (16)
[0232] The generation of synthetic data is only performed by the decoder. The decoder starts from a sampler that generates a 10-dimensional sample nk[t] from the normal distribution N(0, γkI) in each iteration. Then this nk[t] passes through the decoding layer to generate a 90-dimensional vector xS,p=k[t]. New synthetic data points are continuously generated until the size of the synthetic data XS,p=k = xS,p=k[t] at the k-th position is 10 times larger than the size of the original labeled data Xp=k.
[0233] The one-to-one VAE position mapping allows each VAE to automatically label the synthetic data points with the positions dedicated to that VAE. Once the k-th VAE generates a synthetic data point xS,p=k, the position vector k (the one-hot encoding of k) is assigned to it to create the synthetic fingerprint as follows:
[0234] fS[t] = (xS,p=k[t]; k) (17)
[0235] The synthetic data points are collected and labeled based on Equation (17) to obtain the synthetic fingerprint as FS = fS[t], and then, as in Equation (7), the disjoint union FA is output as the union of FS and F.
[0236] The feature extraction layer is the starting point of both the classification path and the reconstruction path. All augmented (labeled) and unlabeled data will pass through the feature extraction layer to generate their feature vectors. Next, the augmented data XA will be used as an example to describe how data flows through the feature extraction layer. The flow of unlabeled data XU is the same in the feature extraction layer.
[0237] Figure 18 The data flow from left to right is shown. The augmented data point xA[t]XA enters the first linear layer, which expands the 90-dimensional vector into a 360-dimensional vector. The second linear layer expands the 360-dimensional vector into a 480-dimensional vector. The third linear layer further expands the dimension from 480 to 600 and outputs the feature vector zA[t]. The above three layers calculate the linear combination of the CSI amplitudes on different subcarriers and streams to generate meaningful features. In addition, the implementation inserts sigmoid activation layers, batch normalization layers, and dropout layers (dropout rate 0.3) between the first and second linear layers and between the second and third linear layers. The sigmoid layer introduces non-linearity; the batch normalization layer speeds up the training; the dropout layer attempts to avoid overfitting.
[0238] After obtaining the feature vectors, the classification path and the reconstruction path diverge from each other.
[0239] After extracting the feature vectors, the classification path flows into the classification layer. Both ZA and ZU pass through these classification layers to generate the predictions Y^A and Y^U respectively. Figure 19 The detailed design of the classification layer is described.
[0240] The feature vector of the labeled data point is an example. Specifically, through the first and second linear layers respectively, the 600-dimensional feature vector zA[t] is shrunk to a 300-dimensional vector and then to a 100-dimensional vector. Then, this 100-dimensional vector is passed to the third linear layer to predict the NΩ-dimensional one-hot encoded position probability vector y^A[t]. There are sigmoid activation layers, batch normalization layers, and dropout layers (dropout rate 0.3) between the first and second linear layers. Sigmoid activation layers and batch normalization layers are present between the second and third linear layers. Given the vector y^A[t], the implementation calculates the classification loss lC[t] of a data point according to formula (9). Then, the aggregated classification loss LC of all augmented data (or a batch of augmented data) is calculated according to formula (8). Similarly, these classification layers predict the position probability vector y^U[t] given zU[t], also known as Y1. The conditional entropy is calculated using formula (10), and the aggregated clustering loss of all unlabeled data (or a batch of unlabeled data) is calculated using formula (11).
[0241] After coming out of the feature extraction layer ( Figure 5 、Figure 18 ), the reconstruction path flows into the reconstruction layer ( Figure 5 and Figure 20 ), which takes as input both the augmented and unlabeled data simultaneously. An example data stream of the unlabeled data is as follows. The stream of the augmented data is the same.
[0242] As Figure 20 shown, the 600 - dimensional feature vector zU[t] is shrunk to a 480 - dimensional vector and then to a 360 - dimensional vector by the first and second linear layers respectively. The third linear layer attempts to generate a 90 - dimensional vector x^U[t] from the 360 - dimensional vector that reproduces xU[t] as much as possible. The aggregated reconstruction loss is calculated for all data with and / or without labels (or a batch of unlabeled and / or augmented data points) using Equation (12).
[0243] The above - mentioned layers form a JCR structure, with the shared feature extraction layer as the connection point of the two paths.
[0244] In Figure 22 is shown the joint training algorithm (“Algorithm 1”).
[0245] Under this structure, the extracted feature represents both labeled and unlabeled data and helps the classifier to extend its coverage to the unlabeled domain (i.e., the unlabeled users in this paper). In this way, even when newly collected data cannot be labeled, the embodiments provided here still automatically learn to utilize the features of new users or changing environments. To train this JCR structure, a joint training algorithm is applied. This joint training algorithm has certain efficiency advantages over the alternating training algorithm used by WiDo. For clarity, the parameters of the feature extraction layer, classification layer, and reconstruction layer are denoted as ΘF(TF), ΘC(TC), and ΘR(TR) respectively, see Figure 5 .
[0246] Given the augmented fingerprint FA (including data points XA and their labels YA) and unlabeled data XU, for newly collected unlabeled data, the embodiment first samples a batch of labeled data points Xb (and the corresponding Yb from YA), and a batch of unlabeled data points Xb from XU. Then these two batches are removed from their original sets respectively. The embodiment passes XbA and XbU through the feature extraction layer to extract features ZbA and ZbU, see Figure 5 and Figure 18 . These features are fed into the Figure 19 classification layer to produce location predictions Y^b and Y^b. At the same time, these features pass through Figure 20The reconstruction layer to generate reconstructed data as X^bA and X^bU. Using these outputs, the implementation can calculate the loss. The implementation calculates the classification loss L(Y b,Y^b) using Equation (8). And calculates the clustering loss LU(Y^b) based on Equation (11). The implementation also obtains the reconstruction loss LR(Xb,X^b,Xb,X^b) according to Equation (12). According to these implementations, the aggregated loss L is constructed as
[0247] L = λ1 * LA + λ2 * LU + λ3 * LR, (18)
[0248] where λ1, λ2, and λ3 are weights that balance the losses. Then, the aggregated loss is backpropagated to update ΘF, ΘC, and ΘR. Specifically, ΘC is updated by λ1LA + λ2LU, ΘR is updated by λ2LU + λ3LR, and ΘF is updated by L.
[0249] In the iterations of Algorithm 1, refer to Figure 22 , ΘF is trained to capture features that can support two different tasks: (1) predicting locations, and (2) representing labeled and unlabeled data in the same embedding space. Since these two tasks are trained in a joint manner, the classifier can use the features representing labeled and unlabeled data points to predict locations. In this way, the implementation can adapt the prediction to newly collected unlabeled data.
[0250] The experimental environment and setup are shown in Figure 23 . The experiment is conducted in an office of 3m × 4m. It is expected that all tested positioning systems cover 8 locations (i.e., NΩ = 8), denoted as p0, p1, ……, p7, etc. Among them, p0 is a virtual location corresponding to the state where the room is empty. The remaining locations are distributed in the room with a minimum distance of 70cm. A Dell Latitude E7440 laptop with an Intel 5300WiFi chip is connected to a TP-Link AC1750 WiFi router to create a standard WiFi signal covering the entire room. The implementation installs Linux CSI on the laptop to extract CSI readings and feeds them to Widora (and other systems) that are also installed on the laptop.
[0251] To collect CSI readings, the implementation turns on 1 antenna (1 out of 3 antennas) on the WiFi router and enables all 3 antennas on the laptop. The Intel 5300WiFi chip measures the CSI on 30 subcarriers.
[0252] Therefore, in the experiment, the embodiment has NTX = 1, NRX = 3, NSS = 1×3 = 3, and NSC = 30. For training and testing, the CPU (Intel-i5 1.9GHz) utilization is less than 15% and 10% respectively, and the memory usage is less than 60MB. 4.9MB of disk space is required to store the trained model.
[0253] [Table 1]
[0254] Index Glass door Swivel chair Metal shutter Env1 Closed Inside Closed Env2 Open Inside Closed Env3 Open Outside Closed Env4 Open Outside Open
[0255] Table 1: Different environments [Table 2]
[0256] Index Action Source Target Change 1 Closed Env1 Env2 Change 2 Open Env2 Env3 Change 3 Open Env3 Env4
[0257] Table 2: Different environmental changes [Table 3]
[0258] Labeled data Unlabeled data Training ~5K from ID2 and ID4 ~1.08M from ID1 - ID9 Testing ~1.08M from ID1 - ID9 Not applicable
[0259] Table 3: Data separation of user diversity [Table 4]
[0260] Labeled data Unlabeled data Training ~20K from the source ~700K from both Testing ~720K from both Not applicable
[0261] Table 4: Data separation of environmental changes
[0262] Introducing user diversity: To record the CSI of different users, 9 volunteers with different genders (3 females and 6 males), heights (ranging from 155 cm to 186 cm), and weights (ranging from 45 kg to 88 kg) were recruited.
[0263] Introducing environmental changes: To study the impact of environmental changes, the embodiment first defined several environments in Table 1. Then, the embodiment deliberately introduced some common daily changes to the environment. For example, by the action of opening a glass door, Environment 1 is changed to Environment 2. The embodiment further considered that for Change 1, Environment 1 is the source domain and Environment 2 is the target domain. As an example, the environmental change is defined as a combination of 1) daily actions, 2) the source domain to which the action is applied, and 3) the resulting target domain. Therefore, the embodiment defined Change 1, Change 2, and Change 3 in Table 2.
[0264] Data collection: The embodiment includes the router periodically sending signals to the notebook at a rate of 100 packets per second. This corresponds to a traffic of 51.2 Kbps, consuming only a small part of the WiFi capacity. Now describe the concept of the iteration and operation of data collection.
[0265] One iteration of data collection: When a volunteer stands at a location for 60 seconds, the corresponding WiFi CSI readings are recorded and labeled with that location. This is one iteration of data collection. During the iteration, the volunteer is allowed to move freely and naturally (e.g., turn, stretch, put hands on the hips, etc.) as long as he / she stays at that location.
[0266] One run of data collection: The volunteer performs 5 iterations at the first location, moves to the second location, completes another 5 iterations, and repeats continuously until all locations are covered.
[0267] One data acquisition consists of 5×NΩ = 40 iterations, resulting in 100 * 60 * 40 = 240K CSI readings.
[0268] To collect data on user diversity, the environment is kept constant (i.e., Environment 1), and each volunteer is required to perform one data collection. Thus, data is collected in a total of 9 runs, obtaining 9 * 240K = 2.16M data points to evaluate user diversity.
[0269] To collect data on environmental changes, volunteers ID1, ID2, and ID3 are required to repeat 4 different runs under 4 different environmental settings (listed in Table 1) respectively. Thus, a total of 3 * 4 * 240K = 2.28M data points are collected to evaluate environmental changes.
[0270] Data separation for user diversity: To simulate the real-world challenges faced by the positioning system (i.e., Challenge 1 and Challenge 2), the data is separated as follows. First, in the actual scenario, it is only feasible to label a small number of users. To mimic this, the labeled data for training only comes from two volunteers, who are called example users. Second, even for these two example users, they can provide a limited number of labeled data points. To reproduce this, there are only 5K labeled CSI readings from these two example users. In addition, to simulate a large amount of potential unlabeled data for training, the remaining data is divided into 50% unlabeled training data (by not using their class labels or domain labels in training) and 50% test data. By selecting different example users, the scenario will have different data separation results. The specific data separation used in this paper is summarized in Table 3.
[0271] [Table 5]
[0272] System Data augmentation JCR structure Clustering loss Fingerprint update Widora 0 0 0 × WiDo 0 0 × × Only VAE 0 × × × AutoFi × × × 0
[0273] Table 5: Systems evaluated. Here, 0 indicates that the system includes a certain module, while × indicates that the module is not present.
[0274] [Table 6]
[0275] Recall Precision F1 score AutoFi 72.9% 75.8% 72.7% Only VAE 81.0% 81.1% 80.7% WiDo 84.6% 85.0% 84.5% Widora 90.6% 90.7% 90.5%
[0276] Table 6: Average recall, precision, and weighted F1-score of different systems across all locations and IDs.
[0277] Widora and WiDo have been disclosed in this paper.
[0278] In contrast to Widora, WiDo does not apply the clustering hypothesis used in Widora, thus resulting in relatively high class ambiguity in the target domain.
[0279] Only VAE only implements the data augmenter of Widora and only connects the augmenter to the classification path. The reconstruction path and the clustering loss are discarded. Only VAE is used to conduct ablation studies on the modules of Widora.
[0280] AutoFi is representative of existing fingerprint update methods (e.g., MSDFL, which uses a fully connected NN instead of the autoencoder in AutoFi). AutoFi neither requires class labels nor domain labels as input. Instead, it periodically detects whether the room is empty and uses this empty room state as an indicator of domain change. All new data collected after detecting the empty room state are automatically labeled with new domain labels. Then a polynomial mapping function and an autoencoder are applied to update the CSI fingerprint database.
[0281] Table 5 summarizes their differences.
[0282] The two smallest volunteers (i.e., ID2 and ID4) are example users in the evaluation (using their 5K data points as labeled training data). Separate the data as shown in Table 3. Widora and Only DAV are able to utilize both 5K labeled training data and 1.08M unlabeled training data. Note that AutoFi and Only VAE only use 5K labeled data as training data (Only VAE augments the training data to obtain 10 times more data points).
[0283] Table 6 lists the average recall, precision, and weighted F1-score (i.e., the harmonic mean of precision and recall weighted by the number of true labels for each classification) across all locations and IDs. The results show that AutoFi has the worst performance due to its inability to handle user diversity. The results show that the recall and precision of this method are as low as 72.9% and 75.8% respectively. By augmenting the limited labeled data, only VAE can cover more potential users, thereby increasing the recall and precision to 81.0% and 81.1% respectively. WiDo incorporates both the augmentation of labeled data and the domain adaptation of unlabeled data, thereby increasing the recall and precision to 84.6% and 85.0% respectively. By introducing the clustering loss, Widora further improves WiDo and achieves the best performance. Compared with the baseline AutoFi, Widora improves the recall by 17.7%, the precision by 14.9%, and the F1-score (absolute value) by 17.8% respectively. In addition, compared with WiDo, Widora also increases these metrics by 6.0%, 5.7%, and 6.0% respectively.
[0284] To better compare the positioning systems, consider the location confusion matrices of different systems (not shown). The true positive rates (TPRs) at different locations of AutoFi vary greatly, ranging from 52.3% at location p3 to 95.6% at location p2. Such a large gap between the best and the worst (i.e., 43.3%) indicates that AutoFi is overfitted to the data from the example users. The data augmenter helps only VAE increase the worst-case TPR to 62.2% at p1, while keeping the best TPR at 95.6% at p2. However, the best-worst TPR gap of 33.4% indicates that only VAE may also be overfitted. Alternatively, the design of the domain adaptation classifier assists WiDo and Widora in handling overfitting and increases the worst-case TPR to 80.4%, while maintaining the best-case TPR as high as 93.3%. With the help of the clustering loss, Widora further increases the best TPR to 99.0%, while keeping the same worst TPR as WiDo. This means that Widora not only provides the maximum performance in the worst case, but also performs well consistently across different locations.
[0285] [Table 7]
[0286]
[0287] Table 7: Average F1-scores across all locations under different environmental changes.
[0288] High robustness to user height: The implementation uses the average height of example users (i.e., ID2 and ID4) as the example height, and divides all users into three groups based on the absolute difference between all users and this example height. It is confirmed that the positioning accuracy decreases as the height difference between the test user and the example user increases. AutoFi has the largest degree of decrease, indicating that it cannot handle the change in user height. When applied to example users, the accuracy of AutoFi is as high as 96.2%. However, if the test user is about 10 cm taller than the example user, this figure drops below 69.1%. This is evidence that AutoFi is overfitted to the data of example users. Widora, WiDo, and VAE only are able to learn location features independent of user height, thus obtaining better and more robust performance than the baseline. In particular, Widora can still maintain its accuracy above 87.1% even when the height difference is greater than 20 cm.
[0289] Although Widora has a 5.8% better performance than WiDo for unlabeled users, WiDo has an accuracy 1.7% higher than Widora for example users. This indicates that to some extent, WiDo is still overfitted to example users. By utilizing the clustering loss, Widora improves the decision boundary and, as an unexpected benefit, further corrects the overfitting.
[0290] Robustness to user weight: The implementation takes the average weight of example users as the example weight and divides all users into three groups according to the absolute difference between all users and this example weight. The three groups are <5 kg, 5 kg to 20 kg, >20 kg. Generally, when the absolute weight difference between the example user and the test user increases, the positioning accuracy decreases. AutoFi is overfitted to example users. If the test user is 20 kg heavier than the example user, its accuracy drops to 66.2%. In contrast, Widora has the strongest robust performance in the presence of different user weights. Even when training on the lightest users and testing on the heaviest users, Widora can still maintain its accuracy above 87.9%. Compared with AutoFi and WiDo, the improvement rates are 21.7% and 6.4% respectively.
[0291] Accuracy in different environments
[0292] Refer to Table 2. The performance of the system was evaluated under several daily environmental changes. For each environmental change, data separation as shown in Table 4 was performed. Taking Change 1 as an example, the experiment selected 20K data in Environment 1 as labeled training data, and regarded the remaining training data in this environment as unlabeled training data. In addition, the experiment regarded all the training data from Environment 2 as unlabeled training data. The experiment further mixed the unlabeled data from the two environments into an unlabeled training set, where the domain label and location label were discarded. Then training and testing were carried out based on such data separation.
[0293] The experiment used the F1 score as a measure of the localization accuracy, and the F1 scores of different systems under different environmental changes are given in Table 7. The results show that Widora can achieve the highest accuracy under different environmental changes. Specifically, the performance of Widora is 23.1% higher than that of the baseline AutoFi. In addition, compared with WiDo, the accuracies of the proposed Widora for the three environmental changes are increased by 7.7%, 4.4% and 15.6% respectively. This shows that using the clustering hypothesis / loss does help Widora better solve the fingerprint inconsistency caused by environmental changes. By comparing WiDo with VAE only, another interesting finding can be obtained. Although WiDo achieved a high F1 score in the target domain, VAE only was better than WiDo in the source domain. This indicates that 1) VAE only is somewhat overfitted to the augmented data; and 2) the JCR structure of WiDo (and Widora) can correct this overfitting by learning more general features from the labeled and unlabeled data. In addition, VAE only is better than AutoFi in both the target domain and the source domain. This shows that data augmentation not only provides more data points to improve the source domain performance, but also expands the coverage of the labeled data to improve the accuracy of the target domain.
[0294] As disclosed herein, the problem of inconsistent device-free WiFi fingerprints across different users and / or different environments has been solved. To address this issue, Widora, a WiFi-based domain adaptation localization system, is described, which utilizes a VAE-based data augmenter and a domain adaptation classifier with a joint classification and reconstruction structure. Different from existing domain adversarial solutions, Widora does not require inherent labels to distinguish target domain data from source domain data, making it more practical. It is also disclosed herein that, compared with WiDo, Widora also adopts a clustering assumption, thereby reducing class ambiguity in the new / unseen domain. In addition, a joint training algorithm is disclosed for achieving a flexible balance between classifying labeled data and extracting domain-invariant features from unlabeled data. Experiments show that Widora not only provides a decimeter-level localization resolution that greatly improves accuracy (accuracy +20.2% in the worst case), but also has strong robustness in the case of various unlabeled users and daily environment changes.
[0295] The above-disclosed embodiments can be written as computer-executable programs or instructions storable in a medium.
[0296] The medium can store computer-executable programs or instructions continuously, or temporarily store computer-executable programs or instructions for execution or downloading. In addition, the medium can be any of various recording media or storage media that combine single or multiple hardwares, and the medium is not limited to a medium directly connected to a computer system, and can also be distributed over a network. Examples of the medium include magnetic media such as hard disks, floppy disks, and magnetic tapes configured to store program instructions, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as optical disk cartridges, ROMs, RAMs, and flash memories. Other examples of the medium include recording media and storage media managed by an application store that distributes applications or by websites, servers, etc. that provide or distribute various other types of software.
[0297] The models related to the above DNN or CNN can be implemented by software modules. When the DNN or CNN model is implemented via software modules (e.g., program modules including instructions), the DNN or CNN model can be stored in a computer-readable recording medium.
[0298] In addition, the DNN or CNN model can be integrated as part of the above device 100 in the form of a hardware chip. For example, the DNN or CNN model can be fabricated in the form of a dedicated hardware chip for AI, or can be fabricated as part of an existing general-purpose processor (e.g., CPU or application processor) or a graphics dedicated processor (e.g., GPU).
[0299] In addition, the DNN or CNN model can be provided in the form of downloadable software. A computer program product can include a product in the form of a software program that is electronically distributed by a manufacturer or an electronic marketplace (e.g., a downloadable application). For electronic distribution, at least a part of the software program can be stored in a storage medium or can be temporarily generated. In this case, the storage medium can be the server of the manufacturer or the electronic marketplace, or the storage medium of a relay server.
[0300] Although embodiments of the present disclosure have been described with reference to the accompanying drawings, those of ordinary skill in the art will understand that various changes in form and detail can be made without departing from the spirit and scope defined by the appended claims.
Claims
1. A method for position determination by an electronic device including a WiFi transceiver and a neural network-based AI model, the method comprising: Receiving, via the WiFi transceiver, a first fingerprint associated with a first person and a first environment, wherein the first person is an unregistered user; Using a shared feature extraction layer of the neural network-based AI model, obtaining a first set of feature data based on operating on the first fingerprint; Using a reconstruction layer of the neural network-based AI model, determining a reconstruction loss associated with the first person and the first fingerprint based on the first set of feature data from the shared feature extraction layer; When the reconstruction loss is equal to or lower than a threshold: Using a classification layer of the neural network-based AI model, classifying the first fingerprint based on the first set of feature data from the shared feature extraction layer, To obtain a first estimated position marker of the first person, and Outputting, via a visual or audible device, an indicator of the first estimated position marker of the first person; When the reconstruction loss is higher than the threshold, performing registration of the first person.
2. The method according to claim 1, wherein, Registration includes: Obtaining, via a user interface of the electronic device, an identifier of the first person; Obtaining, via the user interface, a ground truth position marker of the first person; And Obtaining, via the WiFi transceiver, a second fingerprint associated with the ground truth position marker; And Wherein, the method further includes: providing an application output based on the first estimated position marker or the ground truth position marker.
3. The method according to claim 1, further comprising training weights of the neural network-based AI model based on a third fingerprint associated with the first environment and a second person as a registered user to obtain updated weights.
4. The method according to claim 3, wherein Training weights includes: Receiving a third fingerprint associated with the first environment and the second person; Augmenting the third fingerprint to form a first augmented fingerprint; and Updating the weights based on the first augmented fingerprint.
5. The method according to claim 3, wherein, Determining the reconstruction loss includes: Receiving the first fingerprint via the WiFi transceiver; Based on the updated weights, using the shared feature extraction layer, determining the first set of feature data; Based on the first set of feature data, determining a first reconstructed fingerprint of the first person; and Based on the first fingerprint and the first reconstructed fingerprint, determining the reconstruction loss of the first reconstructed fingerprint.
6. The method according to claim 2, wherein, The application output includes at least one of the following: the electronic device automatically turning off based on identifying that the room is empty; the electronic device forming a sound beam to the first estimated position marker of the first person; The electronic device tilting a display screen of the electronic device according to the first estimated position marker of the first person; updating an augmented reality presentation to the first person based on the first estimated position marker of the first person; and providing an advertisement to a personal electronic device of the first person based on the first estimated position marker of the first person.
7. The method according to claim 1, the method further comprising: Jointly train the feature extraction layer, reconstruction layer, and classification layer of the neural network-based AI model based on the reconstruction loss and the clustering loss, wherein the reconstruction layer is configured to provide the reconstruction loss, and the classification layer is configured to provide the clustering loss; Process the fingerprint to obtain an augmented fingerprint; Operate the feature extraction layer on the augmented fingerprint to generate a code; Use the classification layer to classify the code to obtain an estimated position label; and Provide an application output based on the estimated position label.
8. The method according to claim 7, wherein The application output includes at least one of the following: the electronic device automatically shuts down based on identifying that the room is empty; the electronic device forms a sound beam to the estimated position label; The electronic device tilts the display screen of the electronic device according to the estimated position label; Update the augmented reality presentation for a person based on the estimated position label; And provide an advertisement for the personal electronic device of the person based on the estimated position label.
9. An electronic device for position determination, the electronic device comprising: One or more memories, wherein the one or more memories include instructions; A WiFi transceiver; a neural network-based AI model; and One or more processors, wherein the one or more processors are configured to execute the instructions to: Receive, via the WiFi transceiver, a first fingerprint associated with a first person and a first environment, wherein the first person is an unregistered user, Obtain a first set of feature data using a shared feature extraction layer that operates on the first fingerprint, wherein the neural network-based AI model includes the shared feature extraction layer, Use the reconstruction layer to determine a reconstruction loss associated with the first person and the first fingerprint based on the first set of feature data from the shared feature extraction layer, wherein the neural network-based AI model includes the reconstruction layer, When the reconstruction loss is equal to or lower than a threshold: Use the classification layer to classify the first fingerprint based on the first set of feature data from the shared feature extraction layer to obtain a first estimated position label of the first person, wherein the neural network-based AI model includes the classification layer, and Output an indicator of the first estimated position label of the first person through a visual or audible device; When the reconstruction loss is higher than the threshold, perform the registration of the first person.
10. The electronic device according to claim 9, wherein, The one or more processors are further configured to execute the instructions to: Perform the registration by: Obtaining an identifier of the first person via a user interface of the electronic device, Obtaining a ground truth position label of the first person via the user interface, Obtaining a second fingerprint associated with the ground truth position label, and Providing an application output based on the first estimated position label or the ground truth position label.
11. A machine-readable storage medium storing instructions that, when executed, cause at least one processor of an electronic device to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Feature extraction using multi-task learning
US20190156211A1