CSI-Based Adaptive Positioning Method, Electronic Device, and Storage Medium

By using adaptive convolutional neural network and fusion representation model in indoor positioning technology, the amplitude and phase information of CSI are processed and high-resolution fusion fingerprints are generated, which solves the problems of low positioning accuracy and environmental sensitivity in the existing technology, and achieves more efficient and robust indoor positioning.

CN119676824BActive Publication Date: 2025-06-17JIANGXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510187533.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-17
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The existing indoor positioning technology based on CSI has problems with insufficient fingerprint characteristics and environmental sensitivity, which leads to low positioning accuracy and difficulty in adapting to environmental changes.

Method used

Adaptive convolutional neural network (AdaptCNN) is used to combine fusion representation model, and the amplitude and phase information of CSI is used to generate high-resolution fusion fingerprints. The positioning model can adapt to environmental changes through unsupervised domain adaptation and meta-learning methods.

Benefits of technology

Improves positioning accuracy and robustness, reduces the demand for monitoring points and access points, reduces deployment and maintenance costs, and maintains efficient positioning in the event of environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119676824B_ABST
    Figure CN119676824B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of wireless network positioning, and specifically relates to an adaptive positioning method based on CSI, an electronic device, and a storage medium. In the present invention, a high-discrimination fusion fingerprint data is generated by a fusion representation model according to amplitude information data and phase information data, and a mapping between the high-discrimination fusion fingerprint data and corresponding physical positions in the environment is completed by a positioning model; the positioning model includes an adaptive convolutional neural network composed of a source network and a target network, the source network processes data from the source domain, and the target network processes data from the target domain. The positioning model uses an adaptive convolutional neural network composed of a source network and a target network, which improves the adaptability to the environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wireless network positioning, and particularly to an adaptive positioning method based on CSI, an electronic device, and a storage medium. Background Art

[0002] Indoor positioning has become a key technology in the development of intelligent environments and plays a crucial role in applications such as security monitoring, healthcare, human-computer interaction, and asset tracking. Traditional indoor positioning technologies mainly rely on active devices, such as smartphones or wearable sensors, which use signals such as Wi-Fi, Bluetooth, or RFID to determine positions. Although these methods have proven effective in many scenarios, they also have some limitations. Issues such as user compliance requirements to carry specific devices, privacy concerns related to tracking personal devices, and the need for additional hardware can hinder the widespread application and practical deployment of these methods.

[0003] To overcome these challenges, device-free passive localization (DFPL) has emerged as a promising alternative. DFPL systems do not require individuals to carry any electronic devices but rely on ambient signals present in the environment to detect and locate individuals. An important advancement in DFPL is the development of device-free passive Wi-Fi localization (DFPWL) technology. DFPWL uses channel state information (CSI) extracted from existing Wi-Fi infrastructure to detect and locate individuals without their active participation.

[0004] Although recent research on CSI-based DFPWL technology has demonstrated its potential, existing methods still have some limitations. First is the problem of insufficient fingerprint features. Many fingerprint-based positioning methods mainly focus on the amplitude or phase in CSI measurements. This limited focus may result in CSI fingerprints lacking sufficient discrimination ability to accurately distinguish different positions. Therefore, these methods usually require a large number of monitoring points (MPs) and access points (APs) to achieve acceptable positioning accuracy, which increases deployment complexity and maintenance costs. Additionally, relying on a large amount of infrastructure may also be an obstacle to implementation in environments where it is not possible to add a large number of MPs and APs. Second, environmental sensitivity is a major challenge faced by existing CSI-based positioning systems. CSI measurements are very sensitive to environmental changes, such as furniture rearrangement, door opening and closing, and the emergence of new obstacles. These changes may alter the signal propagation path and cause multipath effects. These environmental dynamics can significantly reduce positioning accuracy because the CSI fingerprints collected during the training phase may not match the changed environment during runtime. Existing positioning systems usually lack a mechanism to adapt to these changes, so their robustness and reliability are low in dynamic environments with constantly changing environmental conditions. Summary of the Invention

[0005] Based on this, the present invention provides a CSI-based adaptive positioning method, an electronic device, and a storage medium, which solve at least one problem in the prior art.

[0006] In a first aspect, the present invention provides a CSI-based adaptive positioning method, which includes the following steps:

[0007] Collect channel state information data of wireless signals in the environment, perform preprocessing, and obtain amplitude information data and phase information data;

[0008] Construct a fusion representation model, and use the amplitude information data and the phase information data to train the fusion representation model; wherein, the fusion representation model generates highly discriminative fusion fingerprint data according to the amplitude information data and the phase information data;

[0009] Construct a positioning model, and use the highly discriminative fusion fingerprint data and its coordinate labels to train the positioning model; wherein, the positioning model completes the mapping between the highly discriminative fusion fingerprint data and the corresponding physical position in the environment; the positioning model includes an adaptive convolutional neural network composed of a source network and a target network, the source network processes data from the source domain, and the target network processes data from the target domain; wherein, the data in the source domain is the channel state information data of wireless signals in the environment, and the data in the target domain is the channel state information data of wireless signals without labels;

[0010] Obtain the channel state information data of the wireless signal to be located, perform preprocessing, obtain amplitude information data and phase information data, and use the trained fusion representation model and the trained positioning model to generate highly discriminative fusion fingerprint data and the probability of predicting the target position.

[0011] It should be noted that the environment refers to the area where target positioning needs to be performed, such as indoors. The positioning model uses an adaptive convolutional neural network composed of a source network and a target network, which improves the adaptability to the environment.

[0012] In some optional embodiments, collecting the channel state information data of wireless signals in the environment includes:

[0013] Arrange a pair of Wi-Fi signal transceiver devices in the environment;

[0014] Divide the environment into multiple training points;

[0015] For each training point, collect wireless signal data through the Wi-Fi signal transceiver device, where the wireless signal data includes channel state information data.

[0016] It should be noted that the Wi-Fi signal transceiver device can be a wireless router and / or a computer with a wireless network card.

[0017] In some alternative embodiments, preprocessing the channel state information data of the wireless signal includes: denoising and phase calibration of the channel state information data of the wireless signal, and then splicing the obtained amplitude information data and phase information data into feature fingerprint data, and obtaining single-target fingerprint data after normalization processing.

[0018] In some alternative embodiments, the denoising is wavelet domain denoising (WDD).

[0019] In some alternative embodiments, the fusion representation model includes feature extractors for amplitude mode and phase mode, a feature fusion layer, and a decoding network. The feature extractors extract high-level feature maps from the amplitude mode and phase mode respectively and , and the feature fusion layer splices and along the channel dimension, as shown in Equation (1);

[0020]

[0021] where F represents the fused feature, and Concat(·) represents the splicing operation;

[0022] The decoding network reconstructs the output matrix from the fused feature .

[0023] In some alternative embodiments, the total loss of the fusion representation model is a weighted combination of the reconstruction loss and the triplet loss, as shown in Equation (4);

[0024]

[0025] where represents a hyperparameter; represents the reconstruction loss, which is calculated by Equation (2); represents the triplet loss, which is calculated by Equation (3);

[0026]

[0027]

[0028] where represents the number of training samples; represents the reconstruction output of the i th sample, represents the ground truth of the i th sample; represents the fused feature of the th sample; represents the fused feature of the positive sample; The fusion feature representing the negative sample; Represents a margin hyperparameter; Represents the positive part function; i Represents a positive integer.

[0029] In some alternative embodiments, in the localization model, a weight regularization term is used to minimize the difference between the transferable layer parameters in the source network and the target network, as shown in Equation (5);

[0030]

[0031] where represents the set of transferable layers; and respectively represent the weight parameters of the th layer in the source network and the target network; and respectively represent the layer-specific scaling and shift parameters learned during training; exp(·) represents the exponential function; represents the weight coefficient, calculated by Equation (7);

[0032]

[0033] where , (·) represents the meta-network; j and k respectively represent positive integers.

[0034] In some alternative embodiments, in the localization model, the maximum mean discrepancy is used to measure and minimize the distance between the feature representation distributions in the source network and the target network, and the domain discrepancy loss based on the maximum mean discrepancy is calculated according to Equation (9) ;

[0035]

[0036] where and respectively represent the number of training set samples in the source domain and the target domain; and respectively represent the th source sample and the th target sample's feature representation; and respectively represent the i' th source sample and the j' th target sample's feature representation; i' and j' respectively represent positive integers; the standard radial basis function kernel is used to calculate , as shown in Equation (10);

[0037]

[0038] Wherein, u represents or , v represents , or ; represents the width parameter of the kernel function.

[0039] In a second aspect, the present invention provides an electronic device, which includes:

[0040] At least one processor;

[0041] And a memory communicatively connected to the at least one processor;

[0042] Wherein, the memory stores instructions that, when executed by the at least one processor, implement the CSI-based adaptive positioning method as described above.

[0043] In a third aspect, the present invention provides a computer-readable storage medium that stores instructions that, when executed by a processor, implement the CSI-based adaptive positioning method as described above.

[0044] Due to the above technical solutions, the embodiments of the present invention have at least the following beneficial effects:

[0045] (1) Only one transmitter and one receiver are used, which not only reduces the deployment and maintenance costs, but also simplifies the actual implementation of the positioning system in various indoor environments;

[0046] (2) By adopting a fusion representation model, which combines the amplitude and phase information of CSI measurements, captures the often overlooked spatio-temporal redundancy and channel correlation, thus enhancing the discrimination ability of fingerprints, achieving more accurate positioning, and reducing the number of required monitoring points (MPs) and access points (APs), not only reducing the deployment and maintenance costs, but also simplifying the implementation process in various environments;

[0047] (3) By using an adaptive convolutional neural network with unsupervised domain adaptation and meta-learning dual-stream structure, the positioning model can adapt to the changes in the CSI link model and sample features caused by environmental changes. This adaptability ensures that the positioning model remains effective even when the environment changes significantly (such as furniture movement or occupancy pattern changes), and can maintain high positioning accuracy and reliability. Description of the Drawings

[0048] Figure 1Schematic diagram of the architecture of the fusion representation model in the embodiments of the present invention.

[0049] Figure 2 Schematic diagram of the architecture of LC-Net in the embodiments of the present invention.

[0050] Figure 3 Schematic diagram of the architecture of AdaptCNN in the embodiments of the present invention.

[0051] Figure 4 Schematic diagram of the layout of the meeting room in the embodiments of the present invention.

[0052] Figure 5 Images of the cumulative distribution functions of the positioning errors of different methods in a multi-day environment in the embodiments of the present invention, where Figure 5 (a) in is the image of the cumulative distribution function of the positioning error of the method of re-collecting and training (Baseline), Figure 5 (b) in is the image of the cumulative distribution function of the positioning error of the fine-tuning method, Figure 5 (c) in is the image of the cumulative distribution function of the positioning error of the re-training method, Figure 5 (d) in is the image of the cumulative distribution function of the positioning error of the method without adaptation, Figure 5 (e) in is the image of the cumulative distribution function of the positioning error of the AdaptCNN method.

[0053] Figure 6 Images of the cumulative distribution functions of the positioning errors of different systems in a multi-day environment in the embodiments of the present invention, where Figure 6 (a) in is the image of the cumulative distribution function of the positioning errors of different systems in the environment of the first day, Figure 6 (b) in is the image of the cumulative distribution function of the positioning errors of different systems in the environment of the second day, Figure 6 (c) in is the image of the cumulative distribution function of the positioning errors of different systems in the environment of the third day, Figure 6 (d) in is the image of the cumulative distribution function of the positioning errors of different systems in the environment of the fourth day, Figure 6 (e) in is the image of the cumulative distribution function of the positioning errors of different systems in the environment of the fifth day, Figure 6 (f) in is the image of the cumulative distribution function of the positioning errors of different systems in the environment of the sixth day. Detailed implementation manners

[0054] The concept of the present invention and the technical effects generated will be clearly and completely described below to fully elaborate the purpose, solution and effects of the present invention.

[0055] According to an embodiment of the present invention, an adaptive positioning method based on CSI is provided. In this method, only one wireless signal transmitter and one wireless signal receiver are used to collect the channel state information (CSI) amplitude and phase samples of the wireless signal in the environment, which are used to construct a fingerprint database.

[0056] Preprocess the CSI data to improve the quality and reliability of the input features. To reduce the noise and anomalies in the amplitude component of the CSI measurement, wavelet domain denoising (WDD) technology is used for denoising. Wavelet threshold denoising decomposes the signal into different frequency components and performs threshold processing on the wavelet coefficients, thus effectively distinguishing the real signal from the noise. This denoising step preserves the basic features of the signal while significantly reducing the high-frequency noise, thereby maintaining the integrity of the amplitude information crucial for positioning.

[0057] For the phase component, phase calibration is performed to obtain stable and reliable phase information, which can be effectively used for further analysis. The calibrated phase information data enhances the robustness of the system by providing additional discriminative features and complements the amplitude information data.

[0058] Generate two different feature matrices from the preprocessed amplitude and phase information, one for the amplitude information data and the other for the phase information data. These matrices capture different aspects of the signal features and are subsequently fused using a fusion representation model. This fusion process combines the complementary information of the amplitude and phase measurements to generate highly discriminative fused fingerprint features, called high discriminative fused fingerprint (HDFF). HDFF encapsulates the spatio-temporal redundancy and channel correlation inherent in the CSI data, provides a rich representation, and enhances the system's ability to distinguish different positions. These HDFFs are then used to train a convolutional neural network (CNN) positioning model. The CNN learns the complex mapping between the fused fingerprint and the corresponding physical location in the environment, laying the foundation for accurate positioning during the operation phase.

[0059] As Figure 1 shown, the fusion representation model in the offline training phase includes a feature extractor for each modality, a feature fusion layer, a decoding network, and a dedicated loss function designed to enhance the discriminative ability of the learned features. In the initial stage, for each modality (amplitude and phase), a convolutional neural network (CNN) is used as the feature extractor to obtain high-level feature representations from their respective input matrices, denoted as A (amplitude) and B (phase) respectively. The CNN architectures of the two feature extractors are the same, consisting of multiple convolutional layers and max-pooling layers with ReLU activation functions. The pooling window size is 2×2, and the stride is 2. The convolutional layers are responsible for learning the local patterns and feature hierarchies in the data, while the pooling layers reduce the computational complexity by reducing the spatial dimension and mitigate the risk of overfitting. The output of each modality feature extractor is a high-level feature map And respectively capture the significant features of the amplitude and phase information data. To effectively combine these features, a feature fusion layer is adopted to concatenate and along the channel dimension, as shown in Equation (1);

[0060]

[0061] where Concat(·) represents the concatenation operation.

[0062] The resulting fused feature has a dimension of 7 × 7 × 64 and encapsulates the comprehensive information of the amplitude and phase modalities. This fusion enhances the richness of the representation by integrating complementary features and can improve the discrimination ability of the localization model.

[0063] After the fusion, a decoder network is introduced, which reconstructs the output matrix from the fused feature . The goal of the decoder is to generate an output matrix of size 30×30×3, corresponding to the dimension of the original input matrix. The decoder network consists of convolutional layers and upsampling layers with ReLU activation functions, gradually increasing the spatial dimension. The last layer is a convolutional layer with a Sigmoid activation function, and the output range is [0, 1], which is suitable for the reconstruction task. According to the need, padding or cropping operations are applied to adjust the output size to match the required dimension.

[0064] To effectively train the fused representation model, a combination of reconstruction loss and triplet loss is adopted.

[0065] The reconstruction loss is calculated as the mean squared error (MSE) between the reconstructed output and the ground truth , as shown in Equation (2);

[0066]

[0067] where represents the number of training samples; represents the reconstructed output of the i th sample, and represents the ground truth of the i th sample. Since there is no fixed ground truth output for reconstruction, the input of either modality or is used as a reference . The reconstruction loss encourages the model to generate a ground truth output consistent with the original input data, thus ensuring that the fused features retain the key information from the input.

[0068] Triplet Loss Calculate on the fused features to enhance its discrimination ability. The triplet loss encourages the model to generate feature embeddings such that the distance between samples of the same class (i.e., similar locations) is closer, while the distance between samples of different classes (i.e., different locations) is farther. The triplet loss is defined as:

[0069]

[0070] where represents the fused feature (anchor) of the -th sample; represents the fused feature of the positive sample (i.e., the sample from the same location); represents the fused feature of the negative sample (i.e., the sample from a different location); represents a margin hyperparameter that defines the minimum desired separation between the positive and negative sample pairs (set to in this embodiment); represents the positive part function, ensuring that only positive values contribute to the loss.

[0071] The total loss used to train the fused representation model

[0072]

[0073] is a weighted combination of the reconstruction loss and the triplet loss, as shown in Equation (4); where the hyperparameter is used to balance the influence of the two loss terms. Since the main goal is to enhance the discriminative ability of the features to improve the positioning accuracy,

[0074] is set in this embodiment, thus emphasizing the triplet loss more during the training process. By integrating feature extraction, feature fusion, and a specialized loss function in this deep learning framework, the fused representation model can effectively combine data from the amplitude and phase modalities. The introduction of the triplet loss ensures that the learned features are highly discriminative, enabling the model to more accurately distinguish different locations. The fusion of amplitude and phase information captures the entire spectrum of CSI features, making full use of the complementarity of these two modalities. The amplitude information data provides information about signal strength and attenuation, while the phase information data reveals the signal propagation path and multipath effects. By combining these two aspects, the model can comprehensively understand the conditions of the wireless channel, resulting in more robust and accurate positioning results.

[0075] A convolutional neural network (CNN), called the localization model (LC-Net), is trained using high-discriminative fusion fingerprints (HDFF) from the fusion representation model and their corresponding coordinate labels. This approach enables the network to learn the complex spatial relationships and non-linear mappings inherent in the fusion fingerprints, thus enhancing the localization accuracy. As Figure 2 shown, LC-Net consists of three convolutional blocks, each containing a convolutional layer followed by a layer with the LeakyReLU activation function. A key design choice is not to use pooling layers in the network architecture. Although pooling layers are effective in reducing computational complexity and capturing abstract features in image classification tasks, they may lead to a loss of spatial resolution. In the localization task, preserving detailed spatial information is crucial because small differences in the fusion fingerprints correspond to different physical locations. After the convolutional blocks, LC-Net contains two fully connected layers that act as regression layers to estimate the exact position coordinates of the target. The output of the final fully connected layer provides the predicted coordinates, locating the exact position of the target in the indoor environment and achieving accurate localization.

[0076] During the training of LC-Net, the network weights are initialized using a Gaussian distribution with zero mean and a standard deviation suitable for the network size. Stochastic gradient descent (SGD) is used as the optimization algorithm, and the learning rate is set to 0.0001 to ensure stable convergence. The batch size is set to 25 to balance training efficiency and the accuracy of gradient estimation. Mean squared error (MSE) is used as the loss function to quantify the difference between the predicted coordinates and the ground truth labels. This training configuration effectively maps the complex fusion fingerprint features to exact position coordinates, enhancing the overall localization performance. By excluding pooling layers to preserve detailed spatial information and adopting a regression-based CNN model, high localization accuracy applicable to dynamic indoor environments (such as industrial environments where environmental conditions may change frequently) is achieved.

[0077] After training the localization model LC-Net using the initial fingerprint database, a robust mapping between the fusion fingerprint features and the corresponding coordinates is established. To further maintain high localization accuracy in dynamic environments, an adaptive convolutional neural network (AdaptCNN) is constructed, aiming to utilize the knowledge retained in the intermediate layers of LC-Net and adapt to the new fingerprint distribution through unsupervised domain adaptation methods. As Figure 3As shown, the AdaptCNN framework consists of a two-stream architecture, including a source stream and a target stream. Among them, the source stream processes data from the original fingerprint database (source domain), using the existing localization model trained in the offline stage; the target stream processes data from the target domain training set, and the CSI distribution of this data changes due to environmental changes. By jointly training these two streams, AdaptCNN automatically identifies and transfers the key layers and parameters from the source network to the target model using meta-learning methods. This process minimizes the differences between the corresponding layers while retaining the key knowledge obtained from the source domain. The meta-learning mechanism enables the model to adapt to new domains with little or no labeled data, even relying entirely on unlabeled data.

[0078] To prevent excessive deviation between the weights of the corresponding layers in the source network and the target network, a weight regularization term is introduced to constrain the parameters of the target network. This constraint helps reduce overfitting in the target domain, especially in cases where the number of samples is small, and encourages the target network to retain useful knowledge from the source network. Let denote the sample set in the source domain (original fingerprint database), denote the number of samples in this set, and the corresponding label is , where denotes the location information related to sample . Similarly, let denote the sample set from the target domain (new fingerprint distribution), denote the number of samples in this set, and these samples are unlabeled. The goal is to learn a target model that can accurately predict the location information of target domain samples despite the change in data distribution. Let and denote the parameters (weights and biases) of all layers in the source network and the target network respectively.

[0079] The weight regularization term is defined as minimizing the difference between the transferable layer parameters in the source network and the target network, and the expression of the regularization term is:

[0080]

[0081] where, denotes the set of transferable layers; and denote the weight parameters of the th layer in the source network and the target network respectively; and denote the layer-specific scaling and shifting parameters (hyperparameters) learned during training, which allow flexible transformation of the source network parameters to better match the target domain; Represents the weight coefficient specific to each layer, which is used to control the importance of each layer in the regularization term. It prioritizes certain layers according to the relevance of each layer to the target task; exp(·) represents the exponential function, which ensures that when the difference increases, the regularization penalty increases rapidly, thus encouraging the parameters of the target network to remain close to those of the source network.

[0082] Weight coefficient is dynamically determined by a compact meta-network, and the parameters of this meta-network are . This meta-network takes the attributes of the -th layer in the source network as input and outputs , as shown in Equation (6);

[0083]

[0084] To ensure that the coefficients form an effective probability distribution among the layers, a softmax function is applied for normalization, as shown in Equation (7);

[0085]

[0086] where, . This normalization ensures that .

[0087] To align the feature representations between the source domain and the target domain, the maximum mean discrepancy (MMD) is adopted to measure and minimize the distance between the feature representation distributions in the two domains. By reducing this difference, the knowledge transfer from the source domain to the target domain is promoted. The MMD measurement is based on the probability distribution distance between samples drawn from each distribution, and it is defined as the squared distance between the mean embeddings of the two distributions in the reproducing kernel Hilbert space (RKHS). The domain discrepancy loss based on MMD is shown in Equation (8);

[0088]

[0089] where, and respectively represent the feature representations (such as the output of a certain layer) of the -th source sample and the -th target sample; is the RKHS (reproducing kernel Hilbert space) mapping induced by the kernel function .

[0090] By expanding the squared norm, the domain discrepancy loss can be calculated as:

[0091]

[0092] Among them, and respectively represent the feature representations of the i' th source sample and the j' th target sample; i' and j' respectively represent positive integers, i' and j' are respectively used to distinguish the features of different input samples in the same layer of the same network; The standard radial basis function (RBF) kernel is used to calculate , as shown in Equation (10);

[0093]

[0094] Among them, u represents or , v represents , or ; is the width parameter of the kernel function, which is set to in this embodiment.

[0095] By minimizing , the feature representations of the source domain and the target domain are encouraged to be similar, thus promoting the transfer of knowledge.

[0096] The overall loss function for training the source network and the target network is:

[0097]

[0098] Among them, represents the loss related to the source stream, which is the mean squared error (MSE) between the predicted coordinates and the true coordinates of the source domain samples in this embodiment; represents the weight regularization term; represents the domain difference loss; and respectively represent the hyperparameters that balance the contributions of the regularization term and the domain difference term.

[0099] Parameter , and are jointly optimized by minimizing the loss function . The optimization process uses backpropagation. In this embodiment, the Adam optimizer is adopted, and the learning rate is set to . The weights of the target network are initialized using the weights of the source network from the offline training stage. The meta-network parameters of each layer and Initialize to and , so as to achieve identity mapping at the beginning.

[0100] The training process involves the following steps of alternately optimizing the source network, target network, and meta-network parameters:

[0101] Input the source domain dataset , target domain dataset , set the learning rate ;

[0102] Initialize , , and and meta-network parameters;

[0103] Repeat the following steps until the training is completed:

[0104] a. Randomly select mini-batch data and , each batch size is ;

[0105] b. Calculate the source loss on ;

[0106] c. Use and to calculate the weight regularization loss and on the previous two mini-batch datasets ;

[0107] d. Use and to calculate the domain difference loss function on the two mini-batch datasets ;

[0108] e. Calculate the total loss ;

[0109] f. Use backpropagation to minimize , update , and .

[0110] To evaluate the positioning performance of the CSI-based adaptive positioning method (the method of the embodiments of the present invention) in the embodiments of the present invention, a series of control experiments were conducted. The evaluation focused on the robustness to environmental changes, the effectiveness in adapting to new environmental conditions, and the overall efficiency of the system in terms of deployment and maintenance overhead.

[0111] AsFigure 4 As shown, experiments were conducted in a conference room of approximately 120 square meters, which was equipped with a large conference table, multiple stools, and cabinets, providing a typical cluttered indoor environment. Such an environment poses challenges to positioning due to multipath propagation and signal attenuation caused by obstacles.

[0112] In terms of experimental hardware, a TP-LINK WR841N router configured as a transmitter was used. This router operates in the 2.4 GHz band, representing the common standard Wi-Fi infrastructure in indoor environments. The receiver was a laptop equipped with an Intel 5300 network interface card (NIC) and running the Ubuntu 14.04 system.

[0113] In the conference room, 56 reference points (RPs) were selected and marked for training, and 20 test points (TPs) were used for evaluation. These points were evenly distributed in the open area of the conference room, and the positions of these points covered the entire area, ensuring that the collected data could represent various signal propagation conditions in the conference room.

[0114] To collect CSI data, the CSI Tools software suite was used, which can extract CSI measurement data from the Intel 5300 NIC. During the data acquisition process, the transmitter (router) used a single-antenna configuration and continuously sent packets at a rate of 100 packets per second. The receiver (laptop) was equipped with three antennas and received and recorded CSI data from multiple spatial streams.

[0115] At each training point, a human target was placed in a stationary position, and CSI data was collected for approximately 15 seconds, with about 1500 CSI samples collected at each point. Each CSI sample contains the amplitude and phase information of 30 subcarriers and three receiving antennas, and the total number of CSI amplitude samples for each packet is 30×3. The collected CSI samples were then processed and converted into two-dimensional matrices to capture the spatial and frequency-domain characteristics of the wireless channel. Based on these matrices, 50 fingerprint matrices were generated for each training point as the input to the fusion model in the subsequent data preprocessing stage.

[0116] To evaluate the adaptability of the method of the embodiments of the present invention under varying environmental conditions, experiments were conducted over six consecutive days. During this period, common changes in a realistic indoor environment were simulated, such as moving furniture (conference tables, stools, and cabinets), opening and closing doors, etc., to simulate the impact of actual changes on the CSI fingerprint distribution, which posed challenges to the adaptability of the positioning model. Importantly, unlabeled CSI samples were collected only when the environment changed. Specifically, the target was randomly placed at one of the 56 reference points, and 150 CSI samples were collected each time without annotating the coordinates of the reference points. Compared with recalibrating all training points, this method greatly reduced the workload and required almost no manual intervention.

[0117] In the experiment, the meta-network in AdaptCNN was configured as a single-layer fully connected neural network for each transmission channel between the source network (the original model) and the target model (the adapted model). Specifically, for each layer in the source network , the meta-network took the feature map as input and output a set of layer-specific weights . To ensure that these weights were properly normalized and had a positive effect on the adaptation process, the softmax function was applied to the output for normalization, thus imposing a constraint , where the weights within the index layer . This normalization ensured that the weights could be interpreted as a probability distribution over the features, facilitating the adaptation process. In addition, the ReLU6 activation function was adopted to ensure that the weights were strictly positive and bounded, preventing overgrowth that might lead to instability in the training process. The ReLU6 function limited the output value to a maximum of 6, thus providing an upper bound for the activation value. Therefore, the meta-network learned to adjust the feature representation from the source network to better align with the target domain, enabling the positioning system to maintain high accuracy even when the CSI fingerprint changed due to environmental variations.

[0118] On the first day, CSI data for all 56 training points were collected, and the positioning model was trained using this comprehensive dataset. The positioning model trained on the first day was called the baseline positioning model, which served as a reference point for the performance of the positioning model in the subsequent days. From the second day to the sixth day, environmental changes were introduced to simulate common dynamic conditions in an indoor environment. After each environmental change, additional unlabeled CSI fingerprint data were collected, and the existing model was adapted to the updated environment using the unlabeled CSI fingerprint data.

[0119] To evaluate the adaptability and effectiveness of AdaptCNN in the embodiments of the present invention, its performance was compared with that of several alternative methods:

[0120] No - adaptation: Directly use the localization model trained on the first day without any adaptation to the changes in CSI fingerprints, aiming to illustrate the impact of environmental changes on localization accuracy;

[0121] Re - training: On each day, use the target - domain training set collected after the environmental change and re - train the localization model from scratch;

[0122] Fine - tuning: Use the target - domain training set collected each day to fine - tune the existing model;

[0123] Re - collection and training (baseline method): Re - collect CSI fingerprints from all training points and re - train the model from scratch each day; This method represents an ideal scenario where the model is fully updated with comprehensive new data and serves as a performance benchmark.

[0124] Table 1 shows the average localization errors of different methods in a six - day experiment. It can be seen that the average localization error of the baseline method is about 55 cm; This performance benefits from the rich fingerprint features obtained by two - dimensional processing of CSI data and the model being re - trained with complete and updated data every day. In contrast, AdaptCNN achieved an average localization error of about 77 cm when adapting with only a small number of unlabeled CSI samples; This result demonstrates the strong robustness and effectiveness of AdaptCNN in maintaining localization accuracy while minimizing the re - calibration effort; Compared with the baseline method, the error increased slightly, but considering the significant reduction in data collection and annotation effort, this is acceptable. Other methods, such as fine - tuning and re - training, showed higher localization errors and larger performance fluctuations; These methods rely on smaller training datasets and are prone to overfitting and accuracy fluctuations, especially when significant changes occur in CSI fingerprints. The no - adaptation method performed the worst, highlighting the negative impact of environmental changes on the localization model and emphasizing the importance of using adaptive techniques in dynamic environments.

[0125] Table 1 Average localization errors (cm) of different methods in a multi - day environment

[0126]

[0127] Figure 5Shows the cumulative distribution function (CDF) of the positioning error of different methods over six days. It can be seen that although AdaptCNN significantly reduces the required human intervention, its average positioning error only increases by about 22 cm compared to the baseline method. Approximately 80% of the AdaptCNN positioning errors fall within 1 meter, exceeding the performance of the retraining and fine-tuning methods, which have a lower probability due to overfitting and the instability of small datasets. These results highlight the necessity and effectiveness of adapting the model to changes in the CSI fingerprint distribution to maintain acceptable positioning accuracy.

[0128] To further evaluate the performance of the method (ADCLoc) of the embodiments of the present invention, it was compared with existing positioning methods (CrossSense, TDLLoc, and MFFALoc). The same CSI dataset was used during the same six-day experiment, and these comparisons were carried out under consistent indoor environmental conditions to ensure a fair evaluation.

[0129] As shown in Table 2, on the first day, when the CSI fingerprint remained unchanged, ADCLoc achieved the lowest average positioning error of 53.8 cm, significantly outperforming CrossSense (152 cm), TDLLoc (111.8 cm), and MFFALoc (77.2 cm); this result indicates that ADCLoc has excellent positioning accuracy even in a static environment. From the second day to the sixth day, after introducing the CSI fingerprint changes caused by environmental changes, ADCLoc maintained an average positioning error of approximately 77 cm; in contrast, the positioning errors of other methods increased significantly, with the error of CrossSense averaging approximately 190 cm, TDLLoc approximately 145 cm, and MFFALoc approximately 116 cm; these increases indicate that other methods are less effective in adapting to environmental changes.

[0130] Table 2 Average positioning errors (cm) of CrossSense, TDLLoc, MFFALoc, and ADCLoc

[0131]

[0132] Figure 6 Shows the cumulative distribution function (CDF) of the positioning error during the multi-day experiment for different methods. It can be seen that ADCLoc is consistently superior to CrossSense, TDLLoc, and MFFALoc in terms of positioning performance and stability. The impact of CSI fingerprint changes is clearly reflected in the positioning models of other methods, resulting in a performance decline. In contrast, ADCLoc effectively adapts to these changes and maintains high-precision positioning.

[0133] These results demonstrate the robustness and adaptability of ADCLoc in dynamic environments. By leveraging two-dimensional feature processing, using triplet loss for feature learning, and implementing unsupervised domain adaptation through the AdaptCNN method, ADCLoc effectively addresses the challenges posed by environmental changes and CSI fingerprint variations. While maintaining high positioning accuracy, ADCLoc requires only minimal additional data collection and does not necessitate comprehensive retraining, which proves the feasibility of ADCLoc in practical applications.

[0134] To study the impact of model structure on positioning accuracy, five convolutional neural networks (CNNs) with different architectures were designed, labeled as Structure A to Structure E, as shown in Table 3. Starting from the basic LeNet-5 architecture (Structure A), the network was gradually modified by replacing larger convolutional kernels with smaller ones, increasing the number of channels in each layer, and removing pooling layers to preserve spatial information. Based on LeNet-5, modifying all convolutional layers to smaller convolutional kernels (3×3) and increasing the number of channels in the second layer to 32 resulted in Structure B; maintaining a 5×5 convolutional kernel in the first layer, changing the last two convolutional layers to smaller convolutional kernels (3×3), and further increasing the number of neurons in the last fully connected layer to 128 yielded Structure C; modifying all convolutional layers to smaller convolutional kernels (3×3) and changing the number of channels in the third convolutional layer to 64, and removing all pooling layers gave Structure D. Based on Structure D, increasing the number of channels in the first convolutional layer to 16 and modifying the number of neurons in the last fully connected layer to 256 resulted in Structure E. The performance of ADCLoc was evaluated for six days using each network structure, and the average positioning errors are presented in Table 4.

[0135] Table 3 Different positioning model structures

[0136]

[0137] Table 4 Average positioning errors (cm) under different model structures

[0138]

[0139] The analysis results show that the positioning accuracy gradually improves from Structure A to Structure E. Structure A is based on the standard LeNet-5 architecture and has the highest positioning error, indicating that this architecture may not be suitable for CSI-based positioning tasks. With the modification of the network, including using smaller convolutional kernels (such as 3×3), increasing the number of channels, and removing the pooling layer, the positioning accuracy is improved. It is worth noting that Structure E achieves the best and most stable performance, and the average positioning error is always lower than that of other structures; this structure uses small convolutional kernels, a larger number of channels ([16, 32, 64, 256]), and no pooling layer; the absence of the pooling layer helps to maintain the spatial resolution of the feature map, which is crucial for capturing the subtle differences in CSI fingerprints corresponding to different positions. These analysis results indicate that the choice of the positioning model architecture has a significant impact on the performance of ADCLoc. By optimizing the network structure, especially adopting a design similar to Structure E, the positioning accuracy and robustness of the system in a dynamic environment are improved.

[0140] As described above, these are only the preferred embodiments of the present invention. The present invention is not limited to the above-described embodiments. As long as it achieves the technical effects of the present invention by the same or equivalent means, it shall fall within the protection scope of the present invention. Within the protection scope of the present invention, various different modifications and variations can be made to its technical solutions and / or embodiments.

Claims

1. A CSI-based adaptive positioning method, characterized in that: The following steps are involved: Collect channel state information data of wireless signals in the environment, perform preprocessing, and obtain amplitude information data and phase information data; Construct a fusion representation model and use the amplitude information data and phase information data to train the fusion representation model; wherein the fusion representation model generates high-discrimination fusion fingerprint data according to the amplitude information data and the phase information data; the fusion representation model includes feature extractors of amplitude mode and phase mode, feature fusion layer and decoding network, and the feature extractor extracts high-level feature maps from amplitude mode and phase mode respectively and , the feature fusion layer will and Splicing is performed along the channel dimension, as shown in formula (1); in, F represents fusion features, Concat(·) represents concatenation operation; The decoding network is derived from the fusion features Reconstruct the output matrix ; A positioning model is constructed, and the positioning model is trained using high-resolution fused fingerprint data and its coordinate labels; wherein the positioning model completes the mapping between the high-resolution fused fingerprint data and the corresponding physical location in the environment; the positioning model includes an adaptive convolutional neural network composed of a source network and a target network, the source network processes data from a source domain, and the target network processes data from a target domain; wherein the data in the source domain is the channel state information data of the wireless signal in the environment, and the data in the target domain is the channel state information data of the wireless signal without labels; the convolution kernel size of all convolutional layers in the positioning model is 3×3, the number of channels is [16, 32, 64, 256], and there is no pooling layer; Acquire the channel state information data of the wireless signal to be located, perform preprocessing, obtain amplitude information data and phase information data, and use the trained fusion representation model and the trained positioning model to generate high-resolution fusion fingerprint data and predict the probability of the target location; Among them, fusion represents the total loss of the model It is a weighted combination of reconstruction loss and triplet loss, as shown in formula (4); in, represents a hyperparameter; represents the reconstruction loss, calculated by formula (2); represents the triplet loss, calculated by formula (3); in, Indicates the number of training samples; Indicates i The reconstructed output of samples is Indicates i The true value of samples; Indicates The fusion features of samples; Represents the fusion features of the positive sample; Represents the fusion features of negative samples; represents a margin hyperparameter; represents the positive partial function; i represents a positive integer; In the positioning model, use the weight regularization term Minimize the difference between the parameters of the transferable layers in the source network and the target network, as shown in Equation (5); in, Represents a collection of transferable layers; and Represents the source network and the target network respectively. The weight parameters of the layer; and denote the layer-specific scaling and shifting parameters learned during training, respectively; exp(·) denotes an exponential function; represents the weight coefficient, calculated by formula (7); in, , (·) indicates a meta-network; j and k Respectively represent positive integers; In the localization model, the maximum mean difference is used to measure and minimize the distance between the feature representation distributions in the source network and the target network. The domain difference loss based on the maximum mean difference is calculated according to formula (9): ; in, and Represents the number of training set samples in the source domain and the target domain respectively; and Respectively represent Source samples and Feature representation of target samples; and Respectively represent i' Source samples and j' Feature representation of target samples; i' and j' Represent positive integers respectively; use the standard radial basis function kernel to calculate , as shown in formula (10); in, u express or , v express , or ; Represents the width parameter of the kernel function.

2. The method according to claim 1, characterized in that The channel state information data of the wireless signal in the collection environment includes: Place a pair of Wi-Fi signal transceiver devices in the environment; Divide the environment into multiple training points; For each training point, wireless signal data is collected through a Wi-Fi signal transceiver, where the wireless signal data includes channel state information data.

3. The method according to claim 2, characterized in that Preprocessing the channel state information data of the wireless signal includes: denoising and phase calibration of the channel state information data of the wireless signal, then splicing the obtained amplitude information data and phase information data into feature fingerprint data, and obtaining single target fingerprint data after normalization.

4. The method according to claim 3, characterized in that Denoising is done in wavelet domain.

5. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions, which, when executed by at least one processor, implement the CSI-based adaptive positioning method as described in any one of claims 1-4.

6. A computer-readable storage medium, characterized in that: Instructions are stored, and when the instructions are executed by the processor, the CSI-based adaptive positioning method as described in any one of claims 1-4 is implemented.