Large space positioning method and system adaptive to resistance unsupervised domain
Through the adversarial unsupervised domain adaptation method, aligning the distribution of the source domain and the target domain in the feature space, and building a positioning prediction network model is solved, which solves the problem of performance degradation of large-space positioning methods on the target domain, and realizes high-precision and low-cost positioning adaptation, which is suitable for immersive experiences in complex environments.
Patent Information
- Application Number
- CN202510541806.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing large-spatial positioning method has domain offset problems on the target domain, resulting in a decline in model performance. The unlabeled target domain data increases the cost of data labeling, making it difficult to maintain high-precision positioning in complex environments.
Adversarial unsupervised domain adaptation method is adopted to align the distribution of the source domain and the target domain in the feature space through adversarial training, and unsupervised adaptation is used to construct a positioning prediction network model, including multi-layer perceptrons, feature extractors, generators and domain discriminators, and optimize the parameters of the positioning prediction network model to reduce the impact of domain offset.
It improves the positioning accuracy and generalization ability of the model on the target domain, reduces the cost of data labeling, enhances the robustness of the model in dynamic and noisy environments, adapts to different application scenarios, and provides a smooth, safe and personalized immersive experience.
Smart Images

Figure CN120451482A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of artificial intelligence, computer vision, and computer graphics, and specifically relates to a large-space positioning method and system with adversarial unsupervised domain adaptation. Background Art
[0002] With the rapid development of virtual reality (VR) and augmented reality (AR) technologies, large-scale immersive tour experiences have become an integral part of diverse fields, including cultural tourism, education, and entertainment. In this context, precise positioning of users' locations is crucial. For large-scale immersive tour experiences, accurate positioning not only enhances the user experience but also effectively addresses operational challenges. First, by accurately identifying the locations of multiple users, collisions can be prevented, even in immersive VR tours. This is because in large-scale interactive experiences, multiple users may simultaneously explore or play in a relatively enclosed space. Without effective collision avoidance mechanisms, users could accidentally collide due to complete visual immersion, disrupting the experience and even causing injury. Therefore, through precise location tracking and real-time feedback, the system can issue warnings when users approach or automatically adjust their virtual walking paths to avoid direct contact. Secondly, providing differentiated AR or VR content based on the user's specific location is another key advantage of location-based positioning technology. Different geographic locations can trigger different storylines or quests, ensuring a unique experience for each user. For example, in the virtual reconstruction of historical sites, when a user walks up to a specific historical site, the system can automatically load the corresponding background introduction, character dialogue, or animation demonstration based on their location, greatly enhancing interactivity and participation. In addition, by analyzing the distribution of different user locations, it is possible to find out which areas or content are more popular with tourists. This not only helps park managers understand tourists' interests and behavior patterns, but also provides data support for subsequent content updates and service optimization. For example, if the data shows that a certain exhibition area is particularly popular, the park can add more similar exhibits or activities in the future; conversely, for less popular parts, improvements or replacements can be considered.
[0003] Although existing large-scale spatial localization methods have met the needs to a certain extent, they usually need to be adaptively adjusted to the new target environment to ensure the accuracy of localization. A major challenge is the domain shift problem, that is, the distribution difference between the source domain (training data) and the target domain (test data) may cause the performance of the model to degrade significantly on the target domain. In addition, the target domain data is often not labeled, and labeling this data may be very expensive or practically infeasible, which increases the difficulty of utilizing this data. Directly applying the model trained in the source domain to the target domain may also lead to insufficient generalization ability because the model fails to fully learn feature representations applicable to the new environment.
[0004] To solve these problems, this project proposes a large-space localization method and system with adversarial unsupervised domain adaptation. Summary of the Invention
[0005] In order to solve the problems existing in the prior art, the present invention provides a large-space positioning method and system for adversarial unsupervised domain adaptation. This method aligns the distribution of the source domain and the target domain in the feature space through adversarial training, thereby effectively reducing the impact of domain offset. At the same time, it can use unlabeled target domain data for unsupervised adaptation, avoiding the high cost of data labeling. Through feature alignment and adversarial training, not only the generalization ability of the model in the target domain is enhanced, but also the model has the ability to migrate and can adapt to new environments. This means that no matter in complex environments such as poor lighting conditions, changing obstacles or dense crowds, the system can maintain a high positioning accuracy and provide a smoother, safer and more personalized immersive experience.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] A large-space localization method for adversarial unsupervised domain adaptation, the method comprising:
[0008] Acquire and normalize multimodal data from a large indoor space to form domain data. The multimodal data includes Wi-Fi RSSI, Wi-Fi CSI, BLE, RFID, magnetometer data, accelerometer data, and ultra-wideband signals. The domain data is divided into source domain data and target domain data.
[0009] Build a positioning prediction network model;
[0010] The source domain data and the target domain data are input into the trained positioning prediction network model to complete positioning prediction in the target domain.
[0011] Preferably, the positioning prediction network model includes: a multi-layer perceptron, a feature extractor, a generator and a domain discriminator;
[0012] The multi-layer perceptron is used to extract shallow features of the input domain data;
[0013] The feature extractor is used to extract deep features of domain data and generate domain features;
[0014] The generator is used to predict the position of the object based on the extracted domain features;
[0015] The domain discriminator is used to distinguish whether the domain features come from the source domain or the target domain.
[0016] Preferably, a Bayesian optimization framework is used to train the positioning prediction network model by maximizing the log-likelihood of the true position, and to optimize the parameters θ of the positioning prediction network model. The optimization objective is:
[0017]
[0018] Among them, θ represents the optimal parameter, that is, the parameter that minimizes the objective function found through the optimization process, and θ is the positioning model Parameters used to predict the position of the object; θ f are the parameters of the feature extractor, which aims to map the input signal to the latent feature space; θ r are the parameters of the generator, which aims to predict the location of the object based on the extracted features; θ d It is the parameter of the domain discriminator, which is used to distinguish whether the feature comes from the source domain or the target domain; is the expectation, which represents the expected value of all input signals and corresponding positions; where Z s is the input signal of the source domain, which is the labeled dataset used for model training; Represents source domain features; It is the output of the domain discriminator for the source domain feature, indicating the probability that the discriminator predicts that the feature comes from the source domain; where Z t is the input signal of the target domain, which is the unlabeled dataset that the model needs to adapt to; is the feature extractor for the target domain data Z t The output of represents the feature representation of the feature extractor on the target domain; It is the output of the domain discriminator for the target domain feature, indicating the probability that the discriminator predicts that the feature comes from the target domain; where X s and Y s is the true position of the object in the source domain, which is used to calculate the generation loss and optimize the performance of the generator; and is the position of the object in the source domain predicted by the generator, which is used to compare with the true position and calculate the generation loss; μ is the weight coefficient, which is used to balance the importance of the feature consistency term, thereby controlling the weight of the feature consistency term in the total loss; It is a feature consistency term that measures the stability of the feature extractor on the target domain. By calculating the output consistency of the feature extractor on the target domain, it ensures the stability of the model in a dynamic environment. in is the expected value of the target domain data distribution, which represents the expected value of all target domain samples; ∈ is the noise added to the target domain data; ‖·‖ represents the Euclidean norm, which is used to measure the consistency of features.
[0019] Preferably, inputting the source domain data and the target domain data into the trained positioning prediction network model to complete positioning prediction in the target domain includes:
[0020] Inputting the source domain data into the trained positioning prediction network model to obtain the predicted location and the discrimination result of the domain discriminator;
[0021] Inputting the target domain data into the same trained positioning prediction network model to obtain the predicted location and the discrimination result of the domain discriminator;
[0022] Through the loss function, the network model adapts to the data distribution of the target domain and completes the positioning prediction in the target domain.
[0023] The present invention also provides a large-space positioning system for adversarial unsupervised domain adaptation, which is used to implement the above method. The system includes: a data acquisition module, a model construction module and a positioning prediction module;
[0024] The data acquisition module is used to acquire multimodal data of the indoor large space environment and normalize it to form domain data, wherein the multimodal data includes: Wi-Fi RSSI, Wi-Fi CSI, BLE, RFID, magnetometer data, accelerometer data, and ultra-wideband signal, and the domain data is divided into source domain data and target domain data;
[0025] The model building module is used to build a positioning prediction network model;
[0026] The positioning prediction module is used to input the source domain data and the target domain data into the trained positioning prediction network model to complete positioning prediction in the target domain.
[0027] Preferably, the positioning prediction network model includes: a multi-layer perceptron, a feature extractor, a generator and a domain discriminator;
[0028] The multi-layer perceptron is used to extract shallow features of the input domain data;
[0029] The feature extractor is used to extract deep features of domain data and generate domain features;
[0030] The generator is used to predict the position of the object based on the extracted domain features;
[0031] The domain discriminator is used to distinguish whether the domain features come from the source domain or the target domain.
[0032] Preferably, a Bayesian optimization framework is used to train the positioning prediction network model by maximizing the log-likelihood of the true position, and to optimize the parameters θ of the positioning prediction network model. The optimization objective is:
[0033]
[0034] Among them, θ represents the optimal parameter, that is, the parameter that minimizes the objective function found through the optimization process, and θ is the positioning model Parameters used to predict the position of the object; θ f are the parameters of the feature extractor, which aims to map the input signal to the latent feature space; θ r are the parameters of the generator, which aims to predict the location of the object based on the extracted features; θ d It is the parameter of the domain discriminator, which is used to distinguish whether the feature comes from the source domain or the target domain; is the expectation, which represents the expected value of all input signals and corresponding positions; where Z s is the input signal of the source domain, which is the labeled dataset used for model training; Represents source domain features; It is the output of the domain discriminator for the source domain feature, indicating the probability that the discriminator predicts that the feature comes from the source domain; where Z t is the input signal of the target domain, which is the unlabeled dataset that the model needs to adapt to; is the feature extractor for the target domain data Z t The output of represents the feature representation of the feature extractor on the target domain; It is the output of the domain discriminator for the target domain feature, indicating the probability that the discriminator predicts that the feature comes from the target domain; where X s and Y s is the true position of the object in the source domain, which is used to calculate the generation loss and optimize the performance of the generator; and is the position of the object in the source domain predicted by the generator, which is used to compare with the true position and calculate the generation loss; μ is the weight coefficient, which is used to balance the importance of the feature consistency term, thereby controlling the weight of the feature consistency term in the total loss; It is a feature consistency term that measures the stability of the feature extractor on the target domain. By calculating the output consistency of the feature extractor on the target domain, it ensures the stability of the model in a dynamic environment. in is the expected value of the target domain data distribution, which represents the expected value of all target domain samples; ∈ is the noise added to the target domain data; ‖·‖ represents the Euclidean norm, which is used to measure the consistency of features.
[0035] Preferably, the positioning prediction module includes: a source domain prediction unit, a target domain prediction unit and an adjustment unit;
[0036] The source domain prediction unit is used to input the source domain data into the trained positioning prediction network model to obtain the predicted location and the discrimination result of the domain discriminator;
[0037] The target domain prediction unit is used to input the target domain data into the same trained positioning prediction network model to obtain the predicted position and the discrimination result of the domain discriminator;
[0038] The adjustment unit is used to adapt the network model to the data distribution of the target domain through the loss function, and complete the positioning prediction on the target domain.
[0039] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the aforementioned method when executing the program.
[0040] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, the aforementioned method is implemented.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] 1. Feature alignment: Through adversarial training, the feature extractor learns domain-invariant feature representations, making the features of the source and target domains as consistent as possible in the feature space, thereby improving the positioning accuracy of the model in the target domain.
[0043] 2. Data utilization efficiency: Fully utilize unlabeled target domain data to achieve better model adaptability in the target domain, avoiding the high cost of labeled data.
[0044] 3. Model robustness: Adversarial training enhances the model’s robustness to noise and distribution changes, making the model more stable in dynamic and noisy environments.
[0045] 4. Flexibility: This method has wide applicability and can be applied to a variety of large-space positioning technologies to adapt to different application scenarios.
[0046] 5. Dynamic adaptability: Through feature alignment and adversarial training, the model can quickly adapt to new target environments, reducing the need for model retraining and improving adaptation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0048] Figure 1 This is a flowchart of a large-space localization method based on adversarial unsupervised domain adaptation according to an embodiment of the present invention;
[0049] Figure 2 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention.
[0050] Description of reference numerals:
[0051] 1010 , processor; 1020 , memory; 1030 , input / output interface; 1040 , communication interface; 1050 , bus. DETAILED DESCRIPTION
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0053] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0054] Example 1
[0055] like Figure 1 As shown, this embodiment provides a large-space localization method for adversarial unsupervised domain adaptation, the method comprising:
[0056] Acquire and normalize multimodal data from a large indoor space to form domain data. The multimodal data includes Wi-Fi RSSI, Wi-Fi CSI, BLE, RFID, magnetometer data, accelerometer data, and ultra-wideband signals. The domain data is divided into source domain data and target domain data.
[0057] Build a positioning prediction network model;
[0058] The source domain data and the target domain data are input into the trained positioning prediction network model to complete positioning prediction in the target domain.
[0059] In this embodiment, the method consists of three stages: "multimodal input data", "adversarial generative network" and "output data". The goal of this method is to analyze the signals recorded by the receiver (assuming that there are N receivers and one transmitter in the system, and the signals emitted by the transmitter interact with objects in the environment and are captured by the receiver), and to minimize the predicted position. The error between the true position (X, Y) and the true position (X, Y) is used to infer the true position of the signal source (e.g., RFID tag). To optimize the parameters θ of the positioning model, i.e., the positioning prediction network model (the generative adversarial network), this method uses a Bayesian optimization framework to train the model by maximizing the log-likelihood of the true position. The optimization objectives are as follows:
[0060]
[0061] Among them, θ represents the optimal parameter, that is, the parameter that minimizes the objective function found through the optimization process, and θ is the positioning model That is, the parameters of the adversarial generative network are used to predict the position of the object. Where Z is the signal recorded by the receiver, which contains the original information about the location of the signal source. Θ is the parameters of the positioning model, which are determined through an optimization process and used to adjust the model to minimize the error between the predicted position and the true position. and are the position coordinates of the signal source predicted by the model on the X-axis and Y-axis respectively. f θ is the parameter of the feature extractor, which aims to map the input signal to the latent feature space. r θ is the parameter of the generator, which aims to predict the location of the object based on the extracted features. d is the parameter of the domain discriminator, which is used to distinguish whether the feature comes from the source domain or the target domain. is the expectation, which represents the expected value of all possible input signals and corresponding positions. where Z s is the input signal of the source domain, which is the labeled dataset used for model training; Represents source domain features; It is the output of the domain discriminator for the source domain feature, indicating the probability that the discriminator predicts that the feature comes from the source domain. where Z t is the input signal of the target domain, which is the unlabeled dataset that the model needs to adapt to; is the feature extractor for the target domain data Z t The output of represents the feature representation of the feature extractor on the target domain. It is the output of the domain discriminator for the target domain feature, indicating the probability that the discriminator predicts that the feature comes from the target domain. where X s and Y s is the true position of the object in the source domain, which is used to calculate the generation loss and optimize the performance of the generator; and is the position of the object in the source domain predicted by the generator, which is used to compare with the true position to calculate the generation loss. μ is a weight coefficient used to balance the importance of the feature consistency term, thereby controlling the weight of the feature consistency term in the total loss. is a feature consistency term that measures the stability of the feature extractor on the target domain. By calculating the output consistency of the feature extractor on the target domain, the stability of the model in dynamic environments is ensured. in is the expected value of the target domain data distribution, representing the expected value of all possible target domain samples. ∈ is the noise added to the target domain data. ‖·‖ represents the Euclidean norm, which is used to measure the consistency of features.
[0062] The consistency item is added, which has the following advantages:
[0063] 1. Feature consistency: A feature consistency term is introduced to ensure the stability of the feature extractor in the target domain.
[0064] 2. Dynamic adaptability: Through feature consistency terms, the model’s sensitivity to noisy data is reduced and overall robustness is improved.
[0065] 3. Improved robustness: By introducing feature consistency terms, the model ensures stable feature extraction in the target domain, thereby better adapting to environmental changes.
[0066] In this embodiment, the multimodal input data:
[0067] (1) Because GPS signals are susceptible to signal reflection and multipath effects in large indoor spaces, resulting in reduced positioning accuracy, multiple positioning technologies are generally used, such as those based on radio frequency identification (RFID), ultra-wideband (UWB), and Bluetooth low energy (BLE). These technologies achieve high-precision large-space positioning by using data such as received signal strength (RSSI), channel state information (CSI), and phase. This method does not require the type or amount of data and supports multimodal input, including Wi-Fi RSSI, Wi-Fi CSI, BLE, RFID, magnetometer data, accelerometer data, and ultra-wideband signals.
[0068] (2) Normalize the original signal data to form domain data, which is divided into source domain data and target domain data. Source domain data is the labeled training data, and target domain data is the unlabeled test data. Target domain data is usually small in scale and has more noise, which makes the direct migration of the source domain model to the target domain less effective. This method inputs the source domain data into the network model to obtain the predicted position and the discrimination result of the domain discriminator, and then inputs the target domain data into the same network model to obtain the predicted position and the discrimination result of the domain discriminator. Through the loss function, the network model is adapted to the data distribution of the target domain, thereby improving the positioning prediction accuracy of the network model in the target domain.
[0069] The normalization process is as follows:
[0070] 1. Data preprocessing: For each modality of data X (m) , first perform preliminary preprocessing to remove outliers and noise and improve the quality of input data. m represents the mth mode, X (m) Represents the data of the mth mode.
[0071] 2. Feature extraction: Extract key features from the data of each modality (to reduce data dimensionality, extract the most representative features, and improve the efficiency of subsequent processing). For example, for Wi-Fi RSSI data, extract the mean and variance of the signal strength; for BLE data, extract the signal stability indicator.
[0072] 3. Adaptive normalization: Use adaptive methods to normalize the extracted features to make them comparable.
[0073]
[0074] Among them, X (m) represents the original data of the mth mode, and They are the mean and standard deviation obtained by adaptive calculation. They can analyze the real-time statistical characteristics of the data and dynamically calculate the mean and standard deviation to better adapt to data changes.
[0075] 4. Weight allocation: Assign different weights ω according to the importance and reliability of each modal data (m) ,The weight can be adjusted according to the actual application scenario and ,data characteristics.,For example, in areas with stronger Wi-Fi signals, ,a higher weight is assigned to Wi-Fi RSSI data.
[0076]
[0077] 5. Data fusion: The normalized and weighted data are fused into a unified feature vector as the input of the positioning model.
[0078]
[0079] Among them, M represents the total number of modes. The fused data Z integrates the information of multiple modes and can more comprehensively reflect the location characteristics of the signal source.
[0080] 6. Differences from existing technologies
[0081] (1) Adaptability: Existing technologies usually use fixed normalization parameters, while this method can better handle changes in data distribution by adaptively calculating the mean and standard deviation.
[0082] (2) Dynamic weight adjustment: In the data fusion stage, existing methods often simply fuse the data of each modality with equal weights. However, this method dynamically assigns weights according to the importance and reliability of the data, thereby improving the accuracy and robustness of the fusion results.
[0083] In this embodiment, (1) the “Generative Adversarial Network” is composed of a “Multi-layer Perceptron”, a “Feature Extractor”, a “Generator” and a “Domain Discriminator”.
[0084] (2) “Multi-layer Perceptron” is used to extract shallow features of input domain data.
[0085] Shallow features refer to basic features extracted directly from input data without deep processing. They primarily reflect the raw characteristics and primary information of the data, typically obtained from the raw data through simple mathematical transformations or linear operations. They can quickly capture the underlying patterns and structures in the data, providing a foundation for further deep feature extraction and analysis.
[0086] Shallow feature extraction can be achieved through a simple linear transformation and nonlinear activation function:
[0087] h=σ(W MLP z+b)
[0088] in, is the extracted shallow feature vector, and k is the dimension of the shallow feature vector. is the vector form of the input data, d is the dimension of the vector, σ is the nonlinear activation function (such as ReLU or Sigmoid, etc.), is the weight matrix of the MLP, is the bias vector.
[0089] The role of shallow features:
[0090] (1) Dimensionality reduction and feature representation: The original data is mapped to a lower-dimensional space through linear transformation to achieve data dimensionality reduction while retaining the main information of the data. This helps to reduce computational complexity and improve the training efficiency of the model.
[0091] (2) Provide a foundation for deep feature extraction: Shallow features serve as basic feature representations and provide input for further deep feature extraction and analysis. Deep feature extraction usually performs more complex nonlinear transformations on this basis to capture high-level semantic information and complex patterns in the data.
[0092] (3) Improve the generalization ability of the model: Shallow feature extraction can reduce the model's dependence on the original data distribution through simple mathematical transformations and nonlinear activation functions, and improve the model's generalization ability on different data sets.
[0093] (3) “Feature Extractor” is to extract the deep features of domain data and generate “domain features”. Specifically, the source domain data and target domain data Input to feature extractor Generate "domain features" and
[0094] in,
[0095] in, Represents the network model used by the Feature Extractor.
[0096] (4) The “generator” predicts the location of objects based on the extracted “domain features”. Objects are objects in a large indoor space, i.e., the targets to be located. For example, if there is a user wearing VR in a large space, then this user is the so-called “object”, or more precisely, the “target”. What we want to predict is the location of this target.
[0097] Specifically, the source domain features Input to the generator The location where the prediction was generated
[0098]
[0099] in, Represents the network model used by the "generator".
[0100] (5) “Domain discriminator” is used to distinguish whether the “domain features” come from the source domain or the target domain, that is, the source domain features and target domain features Enter the domain discriminator Generate discriminator results and in, in, Represents the network model used by the “domain discriminator”.
[0101] By distinguishing whether the “domain features” come from the source domain or the target domain, the feature extractor is forced to learn domain-invariant feature representations, that is, the features of the source domain and the target domain are as consistent as possible in the feature space. In this way, the model trained by the generator on the source domain can be directly applied to the target domain without retraining.
[0102] In this embodiment, the training process:
[0103] (1) Training objectives
[0104] 1) Feature Extractor: By minimizing the generation loss (Generation loss: used to optimize the generator Make the predicted position As close to the real position as possible (X s ,Y s )) to ensure the localization accuracy in the source domain. By maximizing the adversarial loss (adversarial loss: optimizing the feature extractor Domain Discriminator It is impossible to distinguish the features of the source domain and the target domain, thereby achieving feature alignment. ), ensuring feature alignment on the target domain.
[0105] 2) Generator: Improve the prediction accuracy on the source domain by minimizing the generation loss.
[0106] 3) Domain Discriminator: By minimizing the discriminator loss (discriminator loss: used to optimize the domain discriminator Make the discriminator result and As close to the real result as possible. ), improve the ability to distinguish the characteristics of the source domain and the target domain.
[0107] (2) Iterative Optimization
[0108] 1) Extract features of the source and target domains.
[0109] 2) Update the generator parameters to minimize the generation loss.
[0110] 3) Train the discriminator to distinguish source and target domain features.
[0111] 4) Update the feature extractor parameters to simultaneously minimize the generation loss and maximize the discriminator loss.
[0112] In this embodiment, “output data”: the generator outputs the “predicted location”, and the domain discriminator outputs the judgment result of whether it is the source domain or the target domain.
[0113] In this embodiment, the loss function is:
[0114] 1. Generate loss The generation loss measures the absolute error between the predicted position and the true position in the source domain. It is a standard MAE loss and is used to optimize the parameters θ of the generator. r .
[0115]
[0116] in, is the joint distribution The expectation of all possible signals and the expected value of the corresponding true position (X, Y). where X s and Y s is the true position of the object in the source domain, which is used to calculate the generation loss and optimize the performance of the generator; and is the position of the object in the source domain predicted by the generator, which is used to compare with the true position to calculate the generation loss. B is the batch size, which represents the number of samples used to calculate the loss. It is used for batch gradient descent to improve computational efficiency. in and is the true position of the b-th sample, which represents the labeled sample in the source domain dataset. and is the predicted position of the b-th sample, which represents the prediction result of the generator on the source domain data.
[0117] 2. Discriminator loss The discriminator loss measures the performance of the domain discriminator in distinguishing the source domain and target domain features, and is used to optimize the discriminator parameters θ d The discriminator improves its discrimination ability by minimizing the cross entropy loss.
[0118]
[0119] in, It is the distribution The expectation of all possible input signals expected value. where Z s is the input signal of the source domain, which is the labeled dataset used for model training; Represents source domain features; It is the output of the domain discriminator for the source domain feature, indicating the probability that the discriminator predicts that the feature comes from the source domain. where Z t is the input signal of the target domain, which is the unlabeled dataset that the model needs to adapt to; Represents the target domain characteristics; It is the output of the domain discriminator for the target domain feature, indicating the probability that the discriminator predicts that the feature comes from the target domain. Where B is the batch size, is the input signal of the b-th source domain sample, is the input signal of the b-th target domain sample.
[0120] 3. Feature Extractor Loss: The goal of the feature extractor is to simultaneously minimize the regression loss and maximize the discriminator loss (i.e., to deceive the discriminator). Through this adversarial training, the feature extractor learns feature representations that preserve localization information while being domain-invariant. Furthermore, considering the robustness of the model in dynamic and noisy environments, feature consistency loss, noise robustness loss, and prediction confidence loss are added.
[0121]
[0122] in, is the generation loss, which is used to measure the prediction accuracy of the model on the source domain. is the discriminator loss, which is used to optimize the discriminator parameters. is the adversarial loss, which measures the ability of the feature extractor to deceive the discriminator, in Feature consistency loss is used to measure the output stability of the feature extractor on the target domain; ∈ is the noise added to the target domain data. ‖·‖ represents the Euclidean norm, which is used to measure the consistency of features. in is the noise robust loss, which is used to measure the sensitivity of the model to noisy data. Var(·) is the variance, which is used to measure the sensitivity of the feature extractor to noise. in is the prediction confidence loss, which measures the model's emphasis on high-confidence predictions; Conf(·) is the prediction confidence, which measures the discriminator's confidence in the target domain features; MAE(·) is the mean absolute error, which measures the model's prediction accuracy. η1, η2, and η3 are weight coefficients used to control the weights of different items in the total loss. The advantages of adding three loss functions are:
[0123] 1. Multi-objective optimization: Maintaining the alignment of source and target domain features, combining adversarial training, generation loss, feature consistency, noise robustness, and prediction confidence to comprehensively measure the model's accuracy, stability, and dynamic adaptability.
[0124] 2. Dynamic adaptability: Through the feature consistency term, the stability of the feature extractor in the target domain is ensured, the model’s sensitivity to noisy data is reduced, and the overall robustness is improved.
[0125] 3. Robustness improvement: By introducing noise robustness terms, the model's sensitivity to noisy data is reduced, ensuring that the model's feature extraction in the target domain is stable, thereby better adapting to environmental changes.
[0126] 4. High-confidence prediction: By predicting the confidence item, the model pays more attention to high-confidence predictions, thereby improving positioning accuracy.
[0127] Example 2
[0128] The present invention also provides a large-space positioning system for adversarial unsupervised domain adaptation, the system being used to implement the method of embodiment 1, the system comprising: a data acquisition module, a model building module, and a positioning prediction module;
[0129] The data acquisition module is used to acquire multimodal data of large indoor spaces and normalize it to form domain data. The multimodal data includes Wi-Fi RSSI, Wi-Fi CSI, BLE, RFID, magnetometer data, accelerometer data, and ultra-wideband signals. The domain data is divided into source domain data and target domain data.
[0130] Model building module, used to build positioning prediction network model;
[0131] The positioning prediction module is used to input the source domain data and the target domain data into the trained positioning prediction network model to complete positioning prediction in the target domain.
[0132] In this embodiment, the positioning prediction network model includes: a multi-layer perceptron, a feature extractor, a generator, and a domain discriminator;
[0133] Multilayer perceptron, used to extract shallow features of input domain data;
[0134] Feature extractor, used to extract deep features of domain data and generate domain features;
[0135] Generator, used to predict the location of objects based on the extracted domain features;
[0136] Domain discriminator, used to distinguish whether the domain features come from the source domain or the target domain.
[0137] In this embodiment, a Bayesian optimization framework is used to train the positioning prediction network model by maximizing the log-likelihood of the true position, and to optimize the parameters θ of the positioning prediction network model. The optimization goal is:
[0138]
[0139] Among them, θ represents the optimal parameter, that is, the parameter that minimizes the objective function found through the optimization process, and θ is the positioning model Parameters used to predict the position of the object; θ f are the parameters of the feature extractor, which aims to map the input signal to the latent feature space; θ r are the parameters of the generator, which aims to predict the location of the object based on the extracted features; θ d It is the parameter of the domain discriminator, which is used to distinguish whether the feature comes from the source domain or the target domain; is the expectation, which represents the expected value of all input signals and corresponding positions; where Z s is the input signal of the source domain, which is the labeled dataset used for model training; Represents source domain features; It is the output of the domain discriminator for the source domain feature, indicating the probability that the discriminator predicts that the feature comes from the source domain; where Z t is the input signal of the target domain, which is the unlabeled dataset that the model needs to adapt to; is the feature extractor for the target domain data Z t The output of represents the feature representation of the feature extractor on the target domain; It is the output of the domain discriminator for the target domain feature, indicating the probability that the discriminator predicts that the feature comes from the target domain; where X s and Y s is the true position of the object in the source domain, which is used to calculate the generation loss and optimize the performance of the generator; and is the position of the object in the source domain predicted by the generator, which is used to compare with the true position and calculate the generation loss; μ is the weight coefficient, which is used to balance the importance of the feature consistency term, thereby controlling the weight of the feature consistency term in the total loss; It is a feature consistency term that measures the stability of the feature extractor on the target domain. By calculating the output consistency of the feature extractor on the target domain, it ensures the stability of the model in a dynamic environment. in is the expected value of the target domain data distribution, which represents the expected value of all target domain samples; ∈ is the noise added to the target domain data; ‖·‖ represents the Euclidean norm, which is used to measure the consistency of features.
[0140] In this embodiment, the positioning prediction module includes: a source domain prediction unit, a target domain prediction unit and an adjustment unit;
[0141] A source domain prediction unit, configured to input source domain data into the trained positioning prediction network model to obtain a predicted location and a discrimination result of a domain discriminator;
[0142] A target domain prediction unit, configured to input target domain data into the same trained positioning prediction network model to obtain a predicted location and a discrimination result of a domain discriminator;
[0143] The adjustment unit is used to adapt the network model to the data distribution of the target domain through the loss function, and complete the positioning prediction in the target domain.
[0144] Example 3
[0145] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in any of the above embodiments when executing the program.
[0146] Figure 2 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0147] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0148] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0149] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0150] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (e.g., USB (Universal Serial Bus), network cable, etc.) or a wireless method (e.g., mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).
[0151] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0152] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0153] The system of the above embodiment is used to implement the corresponding method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0154] Example 4
[0155] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method described in any of the above embodiments.
[0156] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0157] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0158] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Within the scope of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.
[0159] In addition, to simplify the description and discussion, and so as not to obscure the embodiments of the present disclosure, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring the embodiments of the present disclosure, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the purview of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0160] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0161] Therefore, the units of each example described in the embodiments of this application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0162] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A large-space localization method with adversarial unsupervised domain adaptation, characterized in that: The method comprises: Acquire and normalize multimodal data from a large indoor space to form domain data. The multimodal data includes Wi-Fi RSSI, Wi-Fi CSI, BLE, RFID, magnetometer data, accelerometer data, and ultra-wideband signals. The domain data is divided into source domain data and target domain data. Build a positioning prediction network model; The source domain data and the target domain data are input into the trained positioning prediction network model to complete positioning prediction in the target domain.
2. The method according to claim 1, characterized in that The positioning prediction network model includes: a multi-layer perceptron, a feature extractor, a generator and a domain discriminator; The multi-layer perceptron is used to extract shallow features of the input domain data; The feature extractor is used to extract deep features of domain data and generate domain features; The generator is used to predict the position of the object based on the extracted domain features; The domain discriminator is used to distinguish whether the domain features come from the source domain or the target domain.
3. The method according to claim 2, characterized in that The Bayesian optimization framework is used to train the positioning prediction network model by maximizing the log-likelihood of the true position and optimizing the parameters θ of the positioning prediction network model. The optimization goal is: Among them, θ represents the optimal parameter, that is, the parameter that minimizes the objective function found through the optimization process, and θ is the positioning model Parameters used to predict the position of the object; θ f are the parameters of the feature extractor, which aims to map the input signal to the latent feature space; θ r are the parameters of the generator, which aims to predict the location of the object based on the extracted features; θ d It is the parameter of the domain discriminator, which is used to distinguish whether the feature comes from the source domain or the target domain; is the expectation, which represents the expected value of all input signals and corresponding positions; where Z s is the input signal of the source domain, which is the labeled dataset used for model training; Represents source domain features; It is the output of the domain discriminator for the source domain feature, indicating the probability that the discriminator predicts that the feature comes from the source domain; where Z t is the input signal of the target domain, which is the unlabeled dataset that the model needs to adapt to; is the feature extractor for the target domain data Z t The output of represents the feature representation of the feature extractor on the target domain; It is the output of the domain discriminator for the target domain feature, indicating the probability that the discriminator predicts that the feature comes from the target domain; where X s and Y s is the true position of the object in the source domain, which is used to calculate the generation loss and optimize the performance of the generator; and is the position of the object in the source domain predicted by the generator, which is used to compare with the true position and calculate the generation loss; μ is the weight coefficient, which is used to balance the importance of the feature consistency term, thereby controlling the weight of the feature consistency term in the total loss; It is a feature consistency term that measures the stability of the feature extractor on the target domain. By calculating the output consistency of the feature extractor on the target domain, it ensures the stability of the model in a dynamic environment. in is the expected value of the target domain data distribution, which represents the expected value of all target domain samples; ∈ is the noise added to the target domain data; ‖·‖ represents the Euclidean norm, which is used to measure the consistency of features.
4. The method according to claim 3, characterized in that Inputting the source domain data and the target domain data into the trained positioning prediction network model to complete positioning prediction in the target domain includes: Inputting the source domain data into the trained positioning prediction network model to obtain the predicted location and the discrimination result of the domain discriminator; Inputting the target domain data into the same trained positioning prediction network model to obtain the predicted location and the discrimination result of the domain discriminator; Through the loss function, the network model adapts to the data distribution of the target domain and completes the positioning prediction in the target domain.
5. A large-space positioning system with adversarial unsupervised domain adaptation, the system being used to implement the method according to any one of claims 1 to 4, characterized in that: The system includes: a data acquisition module, a model building module and a positioning prediction module; The data acquisition module is used to acquire multimodal data of the indoor large space environment and normalize it to form domain data, wherein the multimodal data includes: Wi-Fi RSSI, Wi-Fi CSI, BLE, RFID, magnetometer data, accelerometer data, and ultra-wideband signal, and the domain data is divided into source domain data and target domain data; The model building module is used to build a positioning prediction network model; The positioning prediction module is used to input the source domain data and the target domain data into the trained positioning prediction network model to complete positioning prediction in the target domain.
6. The system according to claim 5, characterized in that The positioning prediction network model includes: a multi-layer perceptron, a feature extractor, a generator and a domain discriminator; The multi-layer perceptron is used to extract shallow features of the input domain data; The feature extractor is used to extract deep features of domain data and generate domain features; The generator is used to predict the position of the object based on the extracted domain features; The domain discriminator is used to distinguish whether the domain features come from the source domain or the target domain.
7. The system according to claim 6, characterized in that The Bayesian optimization framework is used to train the positioning prediction network model by maximizing the log-likelihood of the true position and optimizing the parameters θ of the positioning prediction network model. The optimization goal is: Among them, θ represents the optimal parameter, that is, the parameter that minimizes the objective function found through the optimization process, and θ is the positioning model Parameters used to predict the position of the object; θ f are the parameters of the feature extractor, which aims to map the input signal to the latent feature space; θ r are the parameters of the generator, which aims to predict the location of the object based on the extracted features; θ d It is the parameter of the domain discriminator, which is used to distinguish whether the feature comes from the source domain or the target domain; is the expectation, which represents the expected value of all input signals and corresponding positions; where Z s is the input signal of the source domain, which is the labeled dataset used for model training; Represents source domain features; It is the output of the domain discriminator for the source domain feature, indicating the probability that the discriminator predicts that the feature comes from the source domain; where Z t is the input signal of the target domain, which is the unlabeled dataset that the model needs to adapt to; is the feature extractor for the target domain data Z t The output of represents the feature representation of the feature extractor on the target domain; It is the output of the domain discriminator for the target domain feature, indicating the probability that the discriminator predicts that the feature comes from the target domain; where X s and Y s is the true position of the object in the source domain, which is used to calculate the generation loss and optimize the performance of the generator; and is the position of the object in the source domain predicted by the generator, which is used to compare with the true position and calculate the generation loss; μ is the weight coefficient, which is used to balance the importance of the feature consistency term, thereby controlling the weight of the feature consistency term in the total loss; It is a feature consistency term that measures the stability of the feature extractor on the target domain. By calculating the output consistency of the feature extractor on the target domain, it ensures the stability of the model in a dynamic environment. in is the expected value of the target domain data distribution, which represents the expected value of all target domain samples; ∈ is the noise added to the target domain data; ‖·‖ represents the Euclidean norm, which is used to measure the consistency of features.
8. The system according to claim 7, characterized in that The positioning prediction module includes: a source domain prediction unit, a target domain prediction unit and an adjustment unit; The source domain prediction unit is used to input the source domain data into the trained positioning prediction network model to obtain the predicted location and the discrimination result of the domain discriminator; The target domain prediction unit is used to input the target domain data into the same trained positioning prediction network model to obtain the predicted position and the discrimination result of the domain discriminator; The adjustment unit is used to adapt the network model to the data distribution of the target domain through the loss function, and complete the positioning prediction on the target domain.
9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 4 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 4 is implemented.