Wireless signal positioning method based on convolutional neural network and self-distillation mechanism
By converting RSS data into images and using deep learning technology of CNN and self-distillation mechanisms, the problem of insufficient data effectiveness and noise immunity in existing indoor positioning technologies is solved, and higher positioning accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510294570.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
Existing indoor positioning technologies, especially methods based on received signal strength (RSS), have problems such as deterioration in data effectiveness, insufficient anti-noise capability and low data utilization efficiency, resulting in poor positioning errors and generalization capabilities.
The wireless signal positioning method based on convolutional neural network (CNN) and self-distillation mechanism is adopted to convert RSS data into equivalent images, and multi-level features are extracted through deep learning technology, combined with knowledge distillation technology to improve positioning accuracy and noise resistance.
It significantly improves the robustness and generalization ability of indoor wireless signal positioning, can maintain stable positioning effect under different signal-to-noise ratio environments, and reduces training cost and calculation complexity.
Smart Images

Figure CN120214689A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of wireless signal processing and deep learning. Specifically, it relates to a wireless signal positioning method and system based on a Convolution Neural Network (CNN) and a self-distillation mechanism, which is applicable to wireless signal positioning tasks in indoor environments. Background Art
[0002] Existing indoor positioning technologies mainly include: positioning methods based on signal fingerprints, Time-of-Arrival (ToA) ranging methods, Angle-of-Arrival (AoA) measurement methods, and positioning methods based on Received Signal Strength (RSS). Among them, RSS has become the mainstream indoor positioning method due to advantages such as easy acquisition and low device cost. However, traditional RSS positioning methods have the following challenges: 1) Data validity: With the changes in time and environment, the fingerprint database based on test data usually deteriorates or even fails accordingly, and the cost of re-measuring a new specific data is relatively high; 2) Insufficient noise resistance: In environments with different signal-to-noise ratios, that is, different levels of noise interference, the generalization ability of the model is poor, which easily leads to positioning errors; 3) Low data utilization efficiency: Traditional machine learning methods are difficult to fully explore the high-dimensional feature relationships of RSS data; Summary of the Invention
[0003] To solve the above problems, the present invention proposes a wireless signal positioning method based on CNN and a self-distillation mechanism. By converting RSS data into equivalent images and combining deep learning and knowledge distillation technologies, the positioning accuracy and noise resistance performance are improved. The purpose of the present invention is to provide a wireless signal positioning method and system based on CNN and a self-distillation mechanism to enhance the robustness and generalization ability of indoor wireless signal positioning.
[0004] The technical solution adopted by the present invention for solving its technical problems, a wireless signal positioning method based on a convolutional neural network and a self-distillation mechanism, includes the following steps: Divide the indoor area to be positioned into several square sub-areas, and use transmitting sensors and receiving sensors in the monitoring area to collect the environmental data of each square sub-area; the environmental data includes signal strength RSS data; Convert the RSS data into an equivalent grayscale image, and the input image size is 1×28×28. Construct a training set and a test set according to a proportion, and introduce additive white Gaussian noise at different signal-to-noise ratio levels to enhance the noise resistance ability of the model; splice the original data and the noise data together for training;
[0005] The CNN with three convolutional layers is adopted to extract multi-level features, and LeakyReLU is used as the activation function to enhance the sensitivity to negative value features. The feature vectors are processed by the fully connected layer, and the class probabilities are calculated through Softmax normalization to determine the target position;
[0006] Furthermore, in the CNN of the three convolutional layers, the input size of the first convolutional layer is 1×28×28, there are 32 convolutional kernels in total, with a size of 3×3, a stride of 1, and the padding keeps the size unchanged. LeakyReLU is used as the activation function, and the output size after passing through the pooling layer is 32×14×14;
[0007] Furthermore, the input size of the second convolutional layer is 32×14×14, there are 64 convolutional kernels in total, with a size of 3×3, a stride of 1, and the padding keeps the size unchanged. LeakyReLU is used as the activation function, and the output size after passing through the pooling layer is 64×7×7;
[0008] Furthermore, the input size of the third convolutional layer is 64×7×7, there are 128 convolutional kernels in total, with a size of 3×3, a stride of 1, and the padding keeps the size unchanged. LeakyReLU is used as the activation function, and the output size after passing through the pooling layer is 128×4×4;
[0008] After three convolutional layers, the feature map is flattened into a feature map of 1×2048, and then input into the fully connected layer, and the dimension is reduced to 1×1024 for further learning of high-dimensional features; after passing through the fully connected layer, 35 output neurons are output, and Softmax is used to calculate the probability of each category;
[0009] The output obtained by the model is split into the predicted output of the original data and the predicted output of the noise data respectively. According to the corresponding output and the true label, the loss of the original data is calculated using the cross-entropy loss; the difference between the output of the noise data and the output of the original data is calculated by KL divergence optimization, and the model parameters are updated by backpropagation.
[0010] The beneficial effects of the present invention are: (1) Utilize the powerful image feature extraction ability of CNN and effectively utilize the knowledge refinement from data to data, completely abandon the dependence on the auxiliary model, thereby significantly reducing the cost required for small networks with high training accuracy compared with traditional methods; (2) This technology can maintain a stable positioning effect in different signal-to-noise ratio environments, and has high computational efficiency, and is applicable to complex indoor environments; (3) It has shown significant advantages in terms of positioning accuracy and generalization ability. Especially under the environmental condition of SNR being -5dB, this technology can achieve a 100% positioning accuracy rate, and under the environmental condition of -15dB, it can achieve a 43% positioning accuracy rate; (4) Compared with the traditional KD technology that relies on an additional teacher model, the present invention adopts a self-distillation technology, eliminating the dependence on an external teacher model, which not only simplifies the training process but also significantly reduces the implementation cost. Description of the Drawings
[0011] Figure 1 is the flowchart of the indoor positioning method of the present invention; Figure 2 is the schematic diagram of data acquisition of the present invention; Figure 3 is the model structure diagram of the present invention; Figure 4 is the confusion matrix diagram of the test results of the present invention; Figure 3 The meanings of each annotation in: 1. Original image x raw , 2. Noisy image x dB , 3. Convolutional layer, 4. Batch normalization layer, 5. Fully connected layer, 6. Softmax function. Detailed Embodiments
[0012] The technical solution of the present invention will be further elaborated below in conjunction with the drawings and embodiments.
[0013] Embodiment 1: As Figure 1 shown, this embodiment proposes an adaptive stereo array indoor positioning method, which mainly includes the following steps: Step 1: As Figure 2 shown, divide the indoor area to be positioned with a size of A×B into several square grids of the same size. At the same time, assume that only one detected target can exist in each square grid, and define the position coordinates of the target at the center of the grid.
[0014] Step 2: Arrange a number of receiving sensor nodes with uniform distribution around a monitoring area, with fixed positions, and one transmitting sensor node. The transmitting sensor node has K variable transmitting positions, satisfying the realization of rich wireless links in the sensor network. When the target exists at a certain position in the monitoring area, the transmitting node transmits a radio signal, and all receiving nodes receive RSS information. The transmitting sensor node changes to the next selectable transmitting position and transmits a radio signal, and all receiving nodes receive RSS information; Among them, the transmitting sensor node changes to the next selectable transmitting position and transmits a radio signal, and all receiving nodes receive RSS information. All the collected RSS information is combined into an RSS matrix, which represents the position information of the current position. Among them, all the collected RSS information is combined into an RSS matrix, which represents the position information of the current position, and the corresponding data set G is obtained. Among them, the RSS value of the measurement area with the target is denoted as G target , and the RSS value collected when the monitoring area is empty is denoted as G null , and G target -G null is used as the calculation method to extract the pure RSS excluding environmental interference, denoted as ΔG. The specific formula of the final ΔG is as follows:
[0015] Step 3: Divide the collected data into a training set and a data set in a ratio of 5:1, and convert the data set into an image form for subsequent classification using CNN in the model. During the conversion of the positioning problem description, the RSS measurement value is analogized to the pixel points in the image. When the position of the target changes, any change in its position will result in the generation of an "image" with different features. Different RSS signal patterns can be used to map the position transformation of the target, that is, the image matrices related to different target positions have different patterns.
[0016] Step 4: As Figure 3 shown, in terms of input data, construct a data set D S ={(X1, l1),...,(X N , l N )}, where N is the total number of training samples, X i represents the i-th training sample, and l i is the corresponding hard label, specifically the position label in the present invention. Generate two branch instances from the same training data by applying a unified noise addition strategy. Specifically, different levels of additive white Gaussian noise vectors are introduced into the samples to generate two data branches: the original image x raw (1) without noise addition and the image x dB (2) processed by noise, denoted as x raw ={(x raw1 , l1),...,(x rawn , l n )}, x dB ={(x dB1 , l1),...,(x dBn , l n )};
[0017] Based on this, global feature extraction is performed through the CNN convolutional layer (3) to obtain a representation vector, which serves as the basis for the feature distribution. The Logits returned by the classifier of the fully connected layer are used as the output. Logits are unnormalized probability distributions, and their values may not be limited to the range of (0, 1), representing the scores or confidences of the model for various categories. In a classification task, each component of the Logits vector corresponds to a specific category, reflecting the prediction tendency of the model for that category;
[0018] In the CNN architecture, the core of the transfer mechanism between layers lies in that the output of the previous layer is used as the input of the next layer. Assume that M represents the feature map output by a certain layer. Then, for the features output by the convolutional layer, they will go through a non-linear activation mapping, and its mathematical expression is as follows: M = D S + f(b + W * D S ) where * represents the convolution operation, W represents the convolution kernel (i.e., the filter in the two-dimensional space), b is the corresponding bias term, and D S is used as the input data and is processed through feature extraction and fusion to obtain M. Here, the non-linear activation function f uses LeakyReLU to enhance the sensitivity of the model to small negative values;
[0019] The model consists of three convolutional layers, mainly used to mine and refine the features crucial for the localization task. The obtained feature maps will be flattened and directly connected to the fully connected layer. Among them, the number of neurons in the last fully connected layer is equal to the total number of position categories. To obtain the probability distribution of each position category, the Softmax activation function is used for normalization calculation to ensure that the sum of the probabilities of all categories is 1, and its mathematical expression is as follows: n is the total number of neurons in the output layer, and the output y of the k-th neuron k represents the probability that the input sample belongs to the k-th category. This mechanism enables the model to comprehensively consider all input information and thus make more accurate classification decisions;
[0020] Based on this, Softmax(6) converts Logits into class probabilities, representing the posterior distribution. At this time, the KL divergence is used to measure the consistency between the output P of the original image raw (the target probability of the soft label) and the output P of the noisy image dB and effectively quantifies the difference in their distributions. Finally, the KL divergence constraint is combined with the cross-entropy loss based on hard labels to form the overall optimization objective of the model, improving the generalization ability and classification accuracy of the model on the original data and noisy data.
[0021] In Figure 3For the right part, during the training phase, the cross-entropy loss function is used to measure the difference between the prediction result and the true label. The loss function is defined as follows: where P rawi = P{l = l i |x rawi} represents the posterior probability of predicting the true label l rawi given the input sample x i . Next, the model uses the KL divergence to quantify the deviation between posterior distributions. The posterior distribution vector can be expressed as: P(x) = {P(l = 1|x),..., P(l = L|x)} where L is the total number of classes, and each dimension represents the posterior probability of a certain class. To ensure that the posterior distributions of the model on the original image and the noisy image are consistent, the KL divergence is used as the consistency evaluation metric, and its calculation formula is as follows: The formula represents the KL distance from P raw to P dB , and P raw is used as the target probability distribution representing the soft label to guide the learning of the prediction probability P dB .
[0022] Finally, a loss function with feature consistency as the core is constructed for a single network: L net = αL CE + βD KL (P raw ||P dB ) where L CE is the cross-entropy loss, and α and β are loss weight hyperparameters. The Adam algorithm is used to optimize the entire network, and its core advantage is that it can assign an independent adaptive learning rate to each parameter. The stability and efficiency of the training process are improved by dynamically adjusting the learning rate.
[0023] Step 5: Use the test sets with different noise levels for verification, and the accuracy results are shown in Table 1. Noise level -15dB -10dB -5dB 0dB 5dB 10dB 15dB Accuracy rate 43.4% 83.4% 100% 100% 100% 100% 100% It can be seen that the accuracy of the predicted position is extremely high. Even in the case of severe noise interference, the accuracy can still remain above 40%, indicating that the positioning method provided by the present invention has high positioning accuracy. The confusion matrix of the positioning deviation of the test set with a noise level of -15dB is as Figure 4 shown.
Claims
1. A wireless signal positioning method based on a convolutional neural network (CNN) and a self-distillation mechanism, which performs positioning according to signal strength in an indoor environment; characterized in that: Convert the positioning problem into an image classification problem and improve the model generalization ability and robustness through autonomous learning, including the following steps: (1) Collecting wireless signal received strength (RSS) data and mapping the RSS data into equivalent grayscale image data for subsequent image classification processing; (2) Based on the RSS data, additive white Gaussian noise (AWGN) with different signal-to-noise ratio (SNR) levels is introduced to generate multiple noisy versions of the data and noisy images of the related data; (3) Using CNN to extract features, the CNN structure includes multiple convolutional layers and fully connected layers. The convolutional layers are used to extract global features and convert them into feature vectors for classifier input; (4) The feature vector is classified and normalized through a fully connected layer to generate the final category probability distribution; (5) Combined with the knowledge distillation (KD) technology, the KL divergence is used to measure the probability distribution P of the original RSS image output during the training process. raw The output probability distribution P of the noisy RSS image dB The consistency between them is calculated and CNN is optimized by minimizing the weighted combination of the cross entropy loss function and the KL divergence loss function. (6) The self-distillation mechanism is used to improve the model's noise resistance and generalization performance by self-learning soft target output; the Adam algorithm is used to optimize the entire network, and an independent adaptive learning rate is assigned to each parameter. The stability and efficiency of the training process are improved by dynamically adjusting the learning rate; (7) The positioning accuracy of the proposed method under different noise levels is verified through experiments, and the model structure is optimized to improve its adaptability in complex environments.
2. The wireless signal positioning method according to claim 1, wherein: Collecting RSS data includes the following steps: (1) Divide the monitoring area of size A×B into S square grids of the same size. At the same time, assume that there can be only one detected target in each square grid, define the target's position coordinates at the center of the grid, and record a digital label for each grid; (2) Arrange N receiving sensors around the monitoring area; (3) The first transmitting sensor node transmits a radio signal, and all receiving sensors receive it at the same time. Then the next transmitting node is switched to transmit, and all sensor nodes receive it. After traversing all transmitting nodes, the detection of a location is completed.
3. The wireless signal positioning method according to claim 1, characterized in that: The CNN structure includes three convolutional layers, which extract low-level, medium-level and high-level features respectively, and adopts LeakyReLU as the activation function to enhance the sensitivity to negative value features.
4. The wireless signal positioning method according to claim 1, characterized in that: The Softmax normalization process is used to convert the unnormalized probability distribution into class probabilities and ensure that the sum of the probabilities of all classes is 1.
5. The wireless signal positioning method according to claim 1, characterized in that: The calculation formula of the KL divergence is as follows: Among them, L is the total number of categories, P raw and P dB They represent the posterior probability distribution of the original image and the noisy image respectively.
6. The wireless signal positioning method according to claim 1, characterized in that: The total loss function is composed of cross entropy loss and KL divergence loss, and its calculation formula is as follows: L net =αL CE +βD KL (P raw ||P dB ) Among them, L CE is the cross entropy loss, and α and β are the loss weight hyperparameters.
7. The wireless signal positioning method according to claim 1, characterized in that: The Adam optimization algorithm dynamically adjusts the learning rate to suppress gradient oscillation and improve optimization efficiency.
8. A wireless signal positioning system based on CNN and self-distillation mechanism, characterized in that: include: (1) sensor array, used to collect wireless signals and calculate RSS data; (2) a data preprocessing module, which is used to convert the RSS data into equivalent image data and introduce additive white Gaussian noise with different signal-to-noise ratios; (3) CNN processing module, including convolutional layers and fully connected layers, used to extract features and perform classification tasks; (4) Soft target learning module, based on the self-distillation mechanism, uses KL divergence and cross entropy loss to jointly optimize CNN; (5) Optimization module: Adam algorithm is used to optimize weight parameters to improve convergence efficiency and noise resistance.
9. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is executed.
10. A wireless signal positioning device, comprising a processor and a memory, wherein the memory stores program instructions, and when the processor executes the program instructions, the method of claim 2 is executed.