CSI (Channel State Information) positioning method based on improved segmented averaging method and multi-scale feature extraction under multiple channels

By improving the segmented mean method and multi-scale feature extraction method, and combining 1D convolutional neural networks, long short-term memory networks and Transformers, the problems of deep learning models' dependence on training data and insufficient environmental adaptability in indoor positioning are solved, and high-precision and robust indoor positioning is achieved.

CN120991862APending Publication Date: 2025-11-21NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202511110469.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing deep learning models rely heavily on large-scale, high-quality training data for indoor positioning, consume a lot of computing resources, and are prone to overfitting and sensitivity to environmental noise in complex environments, resulting in unstable positioning accuracy and difficulty in adapting to different environments.

Method used

An improved segmented mean method is used to reduce the dimensionality of multi-channel CSI measurements. A multi-branch parallel network is constructed by combining 1D convolutional neural network, long short-term memory network and Transformer to extract multi-scale features, and high-precision positioning is achieved through a classification network.

Benefits of technology

It improves the model's adaptability and stability to complex environments, significantly enhances the accuracy and robustness of indoor positioning, and is suitable for a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120991862A_ABST
    Figure CN120991862A_ABST
Patent Text Reader

Abstract

The invention discloses a channel state information (CSI) positioning method based on improved piecewise averaging method (IPAA) dimensionality reduction and multi-scale feature extraction under multiple channels. In the data preprocessing stage, statistical features are extracted from obtained CSI measurement values by using an IPAA dimension reduction method, and data dimensions are reduced. Designing a parallel processing structure by using three deep neural networks, namely a 1D convolutional neural network (1DCNN), a long-short-term memory network (LSTM) and a Transform, extracting multi-scale features of the CSI measured value again, fusing the multi-scale features through a feature fusion strategy, and finally, performing offline classification learning on the fused features through a classification network, so as to obtain the multi-scale features of the CSI measured value. Therefore, a high-precision indoor positioning task is realized. The method has remarkable advantages in the aspects of positioning precision and calculation efficiency, and can effectively deal with a complex indoor environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the field of deep learning and indoor positioning, and particularly relates to a CSI positioning method based on an improved segmented mean method and multi-scale feature extraction in a multi-channel. BACKGROUND

[0002] With the large-scale popularization of intelligent terminal devices and the continuous evolution of Internet of Things technology, positioning services play an increasingly important role in modern society, especially in diversified and complex indoor scenarios, and the demand for high-precision positioning is increasingly urgent. In the fields of smart medical treatment, industrial production monitoring, public safety and other key application fields, accurate indoor positioning has become a core technology support. However, due to the serious shielding phenomenon of satellite signals inside buildings, the traditional GPS-based positioning system is difficult to meet the requirements of accuracy and real-time performance in indoor applications. Therefore, researchers gradually turn their attention to auxiliary positioning technology based on wireless communication signals, among which the Wi-Fi signal-based positioning method is widely concerned due to its strong device universality and low deployment cost.

[0003] In recent years, as a physical layer parameter that can be obtained in the Wi-Fi communication process, channel state information (CSI) has gradually become an important breakthrough in indoor positioning research because it can reflect the fading characteristics, multipath propagation and spatial variation of the channel in detail. CSI has higher spatial resolution and stability than received signal strength (RSSI) and can capture the slight changes in user location at the subcarrier level, providing a theoretical basis and data support for achieving centimeter-level precision positioning. Due to its high-dimensional and multi-channel structural characteristics, CSI provides a more detailed and comprehensive description of the indoor space environment.

[0004] Under the continuous promotion of new urbanization and smart city construction in China, the application demand of indoor positioning technology is growing, which promotes the two-way development of related theoretical research and engineering deployment. China has good conditions in terms of Wi-Fi infrastructure coverage, providing a technical basis for CSI collection and utilization. However, due to non-line-of-sight propagation, multipath interference and other problems in complex indoor environments, CSI data often has high dimension, strong noise and high dynamicity, which brings challenges to the modeling accuracy and robustness of traditional positioning algorithms. Therefore, domestic and foreign researchers have successively proposed various optimization strategies, from data preprocessing, time series modeling to deep learning structure innovation, constantly improving the performance of CSI positioning systems in terms of accuracy, robustness and computational efficiency.

[0005] With the deep integration of artificial intelligence, big data, and other technologies in the field of perception and positioning, CSI positioning algorithms based on deep learning are becoming a research hotspot. Relevant scientific research institutions and universities in China have successively proposed new algorithm architectures that fuse convolutional neural networks (CNN), long short-term memory networks (LSTM), and graph neural networks (GNN), and have made a series of progress in feature extraction, spatial modeling, and time series prediction. In addition, the construction of CSI public data sets and standardized verification mechanisms have gradually improved in China, promoting the engineering landing of algorithms. China is also actively involved in the development of global positioning standards, promoting the deployment and collaborative development of CSI positioning systems in practical scenarios such as rail transit, emergency dispatch, and smart buildings.

[0006] As the research on CSI indoor positioning gradually increases, it involves various application scenarios such as personnel positioning, item tracking, and smart home. Researchers have improved the accuracy and real-time performance of CSI positioning by improving algorithms, optimizing models, and combining other technologies. Some research has also explored positioning performance under different channel conditions and proposed corresponding solutions. Distance estimation methods typically use phase information in CSI to calculate the distance between the target and the base station. These methods have the advantage of being able to handle complex environmental changes and improve the robustness of positioning. The acquisition of CSI usually relies on multiple antenna technology, especially in MIMO (Multiple Input Multiple Output) systems, which can obtain detailed CSI data by simultaneously transmitting and receiving multiple signals. Research shows that algorithms combining machine learning and deep learning can effectively process these complex CSI data, thereby significantly improving positioning accuracy. When processing CSI data, due to its high-dimensional characteristics, how to efficiently model the relationship between CSI temporal characteristics and spatial information is also a major challenge currently facing.

[0007] With the advancement of deep learning technology, more and more indoor positioning methods have begun to use deep learning models. By training neural networks to process signal data, systems can automatically extract features, thereby improving positioning accuracy. Although deep learning models have shown strong modeling capabilities and high positioning accuracy in indoor positioning tasks, they still face many challenges in practical applications. First, these models have a high dependence on large-scale, high-quality training data, and the training process usually involves significant computational resource consumption. In addition, models are prone to overfitting in complex environments, and their sensitivity to environmental noise can also lead to unstable positioning accuracy. Therefore, how to ensure positioning accuracy while improving the adaptability and stability of models to different environments has become a key issue in current research. To address these challenges, researchers have begun to explore modeling methods with stronger structural expression capabilities, such as multi-scale neural networks, to enable models to extract effective features from local features, temporal information, and global dependencies in CSI information.

[0008] In view of the above limitations of the traditional model, researchers gradually explore new CSI positioning methods based on convolutional neural network (CNN) and multi-scale fusion. By reducing the dimension of the input data, introducing multiple channels, multiple antennas and multiple subcarrier information, the model can capture different features of the time series at the same time, thereby enhancing the learning ability of the positioning model for local patterns and global dependencies. This new type of model no longer relies on single feature input, but uses multi-dimensional feature fusion to improve positioning accuracy and robustness. Based on this concept, the present application proposes a CSI positioning method based on improved segmented mean method and multi-scale feature extraction in multiple channels, which fully excavates the local features of CSI data, dynamic information in time series, and global dependencies and important information, improves the expression ability of the model for complex signals and spatio-temporal relationships, thereby significantly improving the accuracy and robustness of indoor positioning, and providing an effective solution for building an efficient and low-error intelligent positioning system.

[0009] In the field of indoor positioning, the positioning method based on channel state information (CSI) has become a research hotspot due to its high resolution and stability. CN 110366108 A [1] fuses CSI amplitude value and received signal strength (RSSI), uses principal component analysis (PCA) dimension reduction and support vector machine (SVM) regression, which to some extent alleviates the influence of multipath effect, but does not involve deep feature time series and global correlation modeling; CN 112333642 B [2] converts CSI amplitude and phase information into an image, and realizes positioning by principal component analysis, wavelet denoising and convolutional neural network learning color features, focusing on local expression of image features, but lacking capture of multi-channel time series dynamics; CN 115908547 A [3] constructs a CSI grayscale image and uses a two-dimensional deep convolutional neural network combined with a probabilistic fine positioning method, which improves positioning accuracy, but relies on a single CNN structure and lacks sufficient mining of long-time series dependencies and global features. In contrast, the present application efficiently reduces the dimension of multi-channel CSI time series data by improving the segmented mean method, fuses 1D convolutional neural network, long short-term memory network and Transformer to construct a multi-branch parallel network, mines features from three dimensions of local, time series and global, and finally realizes high-precision positioning through a classification network, effectively making up for the shortcomings of existing methods in feature comprehensiveness, environmental adaptability and complex scene modeling capability.

[0010] [1] Yan Jun, Ma Chuanhui, Yang Mengwei, Kang Bin. Indoor positioning method based on channel state information and received signal strength [P]. China CN 110366108 A, 2019.10.22.

[0011] [2] Lu Duohua, Yan Jun, Zhu Weiping. Indoor positioning method based on channel state information [P]. China CN 112333642 B, 2022.08.30.

[0012] [3] Zhang, Z. C., Sun, L., Wu, L., Lu, L. J., Qian, P. Z., Xu, B. Y. A wireless positioning method based on deep learning [P]. China CN 115908547 A, 2023.04.04. SUMMARY

[0013] To solve the above problems, the application discloses a CSI positioning method based on improved piecewise average method (IPAA) and multi-scale feature extraction under multi-channel. The application uses IPAA to process the data dimension reduction of multi-channel CSI measurement value, uses 1D convolutional neural network (1DCNN), long short-term memory network (LSTM) and Transformer three kinds of model to carry out multi-scale feature extraction, so as to improve the positioning performance. The specific steps are as follows

[0014] Step 1: using IPAA method to process the obtained CSI amplitude measurement value.

[0015] Step 1-1: dimension reduction of CSI measurement value time sequence under each subcarrier

[0016] Step 1-1-1: under each CSI measurement value channel, for each subcarrier, the CSI amplitude measurement value collected is arranged in the order of data packet, and the CSI measurement value time sequence is constructed, and the training length is 1×D. Wherein D represents the number of data packets.

[0017] Step 1-1-2: using IPAA algorithm to reduce the dimension of CSI measurement value sequence with length of 1×D, and the size of the reduced sequence is 1×P, wherein P is the length of the reduced sequence.

[0018] Step 1-2: fusion of CSI reduced sequence

[0019] Step 1-2-1: the CSI time sequence reduced sequence of all subcarriers is spliced and fused according to the subcarrier label, and the CSI reduced sequence under one channel is obtained. The size is 1×(D×P).

[0020] Step 1-2-2: the CSI time sequence reduced sequence of all channels is spliced and fused according to the channel label, and the final CSI reduced sequence is obtained. The size is 1×(D×P×Q). Wherein Q is the number of CSI measurement value channels in the positioning system.

[0021] Step 2: using different deep learning network to carry out multi-scale feature extraction on CSI measurement value, and carrying out feature fusion on the extracted features.

[0022] Step 2-1: constructing 1DCNN feature extraction network

[0023] The 1D CNN feature extraction network is constructed, and the overall structure comprises two convolutional layers, two pooling layers, a flattening layer and two fully connected layers, and the output of the last fully connected layer is used as the output feature of the sub-feature extraction network.

[0024] Step 2-2: Constructing an LSTM feature extraction network

[0025] The LSTM feature extraction network comprises three layers, namely two LSTM layers and one fully connected layer. The model effectively models the context relationship in the sequence through the double-layer LSTM network, and the output of the last fully connected layer is used as the output feature of the sub-feature extraction network.

[0026] Step 2-3: Constructing a Transformer feature extraction network

[0027] The Transformer model comprises a position encoding layer, a self-attention layer, a feedforward network, a regularization layer, a max-pooling layer and a fully connected layer. The fully connected layer outputs the extracted features.

[0028] Step 2-4: Multi-scale feature fusion

[0029] The features extracted in steps 2-1, 2-2 and 2-3 are combined, and the fused features are used as the final result of multi-scale feature extraction and as the input of the classification network.

[0030] Step 3: Building a deep learning classification network, offline training the feature extraction network and the deep learning network in step 2 to learn the nonlinear relationship between the CSI features and the corresponding spatial position information, and obtaining an indoor positioning model.

[0031] Step 3-1: Dividing the data set into a training set and a validation set

[0032] Step 3-2: Building a classification network

[0033] The classification network comprises two fully connected layers, uses a softmax function, the loss function is cross-entropy, and the optimizer is Adam.

[0034] Step 3-3: Offline classification learning

[0035] The feature extraction network built in steps 2-1, 2-2 and 2-3 and the classification network built in step 3-2 are jointly trained to obtain a position estimation model.

[0036] The present application has the following beneficial effects:

[0037] (1) Strong feature extraction capability and good adaptability: The CSI indoor positioning method based on multi-scale deep neural network proposed in this invention integrates three model structures: 1DCNN, LSTM and Transformer. It can comprehensively mine the key features in CSI data from three dimensions: local, temporal and global, and can more fully describe the features. Finally, high-precision positioning is achieved through classification network, which effectively makes up for the shortcomings of existing methods in feature comprehensiveness, environmental adaptability and complex scene modeling capability.

[0038] (2) High positioning accuracy and strong modeling capability: The original CSI time series is efficiently reduced in dimensionality by using the IPAA algorithm, which preserves the channel change trend while reducing redundant calculations and improving feature representation capability. Then, multi-scale features are extracted by using a multi-channel CSI fusion strategy and a three-branch network structure, and a unified feature vector is formed by splicing and fusion to improve the offline learning performance of the model.

[0039] (3) Strong robustness and generalization ability, suitable for complex environments: The model of this invention adopts multi-source CSI channel information fusion and multi-branch structure design, which significantly enhances the system's ability to resist environmental changes, equipment differences and disturbances. Attached Figure Description

[0040] Figure 1 : Flowchart of the present invention;

[0041] Figure 2 : Schematic diagram of the IPAA structure of this invention;

[0042] Figure 3 : Schematic diagram of data feature fusion in this invention;

[0043] Figure 4 : Diagram of the 1DCNN structure of this invention;

[0044] Figure 5 : LSTM structure diagram of this invention;

[0045] Figure 6 : Transformer structure diagram of this invention;

[0046] Figure 7 : Classification learning structure diagram of this invention;

[0047] Figure 8 : Schematic diagram of the training results of this invention ((a) accuracy of training localization results, (b) loss function curve). Detailed Implementation

[0048] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. It should be noted that the terms "front," "rear," "left," "right," "up," and "down" used in the following description refer to directions in the accompanying drawings, and the terms "inner" and "outer" refer to directions toward or away from the geometric center of a specific component, respectively.

[0049] This example provides a CSI localization method based on IPAA dimensionality reduction and multi-scale feature extraction in a multi-channel environment. First, a dataset is constructed based on the collected raw CSI data. The CSI time-series amplitude data for each location point is dimensionality-reduced using IPAA, and then fused with the multi-channel dimensionality reduction results to form a one-dimensional vector representation. Subsequently, three models—1D Convolutional Neural Network (1DCNN), Long Short-Term Memory Network (LSTM), and Transformer—are used in parallel to extract multiple features from the time-series data. Next, a feature fusion strategy integrates various information into a unified feature vector. Finally, a classification network classifies the extracted features, and a fully connected layer completes the classification prediction of the location points. The flowchart of this invention is shown below. Figure 1 As shown.

[0050] Step 1: Use the IPAA method to reduce the dimension of the obtained CSI amplitude measurement values.

[0051] 1) Constructing CSI time series amplitude values

[0052] The raw CSI amplitude data is constructed into a time-series input sample. First, CSI amplitude data is collected at multiple locations, and at each location, the CSI amplitude of multiple channels formed by the combination of multiple transmit antennas and multiple receive antennas is recorded. Each channel contains complex CSI values ​​of multiple subcarriers, and multiple data packets are continuously collected to form the time series of the raw CSI data.

[0053] In constructing the CSI time series, this invention selects a construction length of 1×50, where the time dimension parameter 50 represents 50 consecutive data packets of CSI data collected at each location point, reflecting the dynamic changes of the channel within a short period. The CSI data at each location includes 6 channels (3 transmit antennas and 2 receive antennas), with each channel containing 30 subcarriers. From the 3000 collected CSI data packets, we divide them into 50 packets per time series, resulting in 60 time series. The final constructed CSI amplitude time series has a size of 1×50.

[0054] 2) Use the IPAA method to reduce the dimensionality of time series data.

[0055] To further compress the dimensionality of CSI time series data and extract its main trend features, this paper designs and implements an improved piecewise average method (IPAA), the structure of which is shown in the diagram below. Figure 2 As shown. It is an algorithm for dimensionality reduction of time series data. This method simplifies the data by reducing the number of sampling points while preserving the main trend or pattern of the data. The IPAA method divides the time series into multiple smaller time periods and replaces the data for the entire time period with a representative value (usually the average value of that period) within each period. This method helps reduce the complexity of the data while preserving the overall trend of the data as much as possible.

[0056] First, the constructed time series with a length of 1×50 is divided into segments. The sliding window size is set to M, and the sliding step size is N; in this paper, M is 8 and N is 2. Starting from the beginning of the sequence, the average of the 8 original data points in each segment is calculated using the in-window averaging method to obtain the representative value for that segment. The sliding window moves forward by the step size, repeating this averaging process until the entire sequence has been traversed. Finally, the means calculated for each segment are combined into a new dimensionality-reduced feature vector. The length of the dimensionality-reduced feature vector is 8.

[0057] The IPAA method is used to obtain the feature representation of a signal. This method effectively compresses the CSI signal while preserving feature-to-feature connections. IPAA can be expressed as:

[0058]

[0059] Among them, P i Let X represent the i-th sequence value after dimensionality reduction, M represent the size of the sliding window, N represent the sliding step size, and X represent the sliding window size. j This represents the j-th data value in the sequence.

[0060] 3) Data feature fusion

[0061] In machine learning and deep learning, the main purpose of data fusion is to extract valuable information from multiple data sources and then use this information to construct a more expressive feature set through related operations. Concatenation is a common method in data fusion. By connecting multiple feature vectors or data matrices along a certain dimension, it is very effective when processing multi-view data. Concatenation: The output sequence of the model is concatenated along the feature dimension to obtain a new sequence. Assume... and H1 is the feature matrix extracted by the model from the input sequence, where T corresponds to the sequence length and d1 is the dimension of the model output; H2 represents the temporal feature matrix extracted by another model from the same input sequence, where d2 is the dimension of the model output. H1 and H2 are the feature representations extracted by two different models from the same set of sequence data. The concatenated result is shown in the following equation:

[0062]

[0063] Where T is the sequence length, and d1 and d2 are the dimensions of the model output, respectively.

[0064] At each location point, CSI amplitude time series data were acquired from multiple antenna pairs. Each channel constitutes a separate CSI amplitude series to compress data dimensionality while preserving key temporal features. Each channel's data was first dimensionality-reduced using the IPAA method to generate a relatively low-dimensional representation. Subsequently, the dimensionality-reduced feature vectors from all channels were concatenated sequentially to form a unified feature vector, which serves as the final representation of the sample. A schematic diagram of the data feature fusion is shown below. Figure 3 As shown, this fusion vector not only preserves the channel difference information of multiple antenna pairs, but also effectively reduces data redundancy, making it easier to input into a deep learning network for feature learning and classification prediction.

[0065] In the data fusion stage, after reducing the dimensionality of the time series using the IPAA method, each time series has a length of 8. The data from 30 subcarriers are then fused to obtain a data set of length 240. Next, the CSI data from the six channels are fused to obtain a new processed dataset containing 67 locations, 60 data samples at each location, and each sample having a size of 1×1440. This dataset serves as the input dataset for our multi-scale neural network.

[0066] Step 2: Use different deep learning networks to extract multi-scale features from CSI measurements and then fuse the extracted features.

[0067] 1) Construct a 1DCNN feature extraction network

[0068] One-dimensional convolutional neural networks (1DCNNs) are deep learning models specifically designed for processing one-dimensional sequential data, widely used in time series analysis, signal processing, and sequence classification tasks. This network extracts local patterns and key features from sequences by sliding convolutional kernels along the time dimension, exhibiting strong temporal modeling capabilities. In its network structure, convolutional and pooling layers are stacked alternately, effectively compressing feature dimensions and enhancing the model's robustness. Finally, fully connected layers are used to achieve classification or regression tasks. 1DCNNs are simple in structure and computationally efficient.

[0069] In the invention application, the 1DCNN structure has 7 layers, and the structural diagram is as follows. Figure 4 As shown, the input layer receives an input of size (1440, 1), indicating that each sample has 1440 time steps, with one feature per step. The first layer is a convolutional layer with 64 filters and a kernel size of 3, resulting in an output size of (1438, 64), where 1438 is the time series length and 64 is the number of kernels. This layer extracts local time series features using 64 kernels of length 3. A pooling layer with a pooling size of 2 follows, with an input shape of (1438, 64) and an output of (719, 64), changing the time series length to 719. This pooling layer downsamples the time series by taking the maximum value from every two values, reducing computational cost and feature dimensionality. The third layer is also a convolutional layer with 128 filters and the same number of kernels as the first convolutional layer. It takes a shape of (719, 64) as input and outputs (717, 128), resulting in a time series length of 717 and 128 convolutional kernels. This layer uses 128 convolutional kernels to further extract deeper time series features. Next, a second pooling layer with a window size of 2 is applied. It takes an input of (717, 128) and outputs (358, 128), reducing the time series length to 358. This further downsampling reduces the time dimension while retaining important features. The fifth layer is a flattening layer with an input of (358, 128) and an output of 45760, representing the length of the flattened feature vector, containing all local feature information. This layer converts the 3D tensor into a 1D vector, preparing for the fully connected layer. Two more fully connected layers follow the flattening layer. The first fully connected layer takes an input of 45760 and outputs a size of 64. This layer reduces dimensionality and extracts high-level semantic features. The input to the last fully connected layer is a feature vector of length 64, and the output is a feature vector of length 32. This output serves as the final feature extracted by the sub-model.

[0070] 2) Construct an LSTM feature extraction network

[0071] Long Short-Term Memory (LSTM) networks are a special type of recurrent neural network (RNN) specifically designed to address the vanishing gradient and forgetting problems inherent in traditional RNNs for modeling long sequences. LSTMs effectively control the retention and discarding of information within a sequence by introducing gating mechanisms such as forget gates, input gates, and output gates, thereby enabling the modeling of long-term dependencies. Compared to traditional CNNs, LSTMs focus more on capturing the global dynamic evolution process over time, making them suitable for modeling temporal correlations and long-term dependency structures in CSI data.

[0072] like Figure 5As shown, the LSTM model consists of three layers: two LSTM layers and one fully connected layer. Its input layer receives an input of shape (1440, 1), representing a univariate time series with 1440 time steps, one feature per step. The first layer is an LSTM layer with a return sequence, meaning it returns the hidden state of each time step, not just the state of the last time step. Its output shape is (1, 128), where 128 is the number of LSTM units, indicating that each sample is compressed into a 128-dimensional feature vector after processing by this layer. This is a stateful time series encoder used to encode the features of each time step and return the output of each time step for the next layer to continue sequence modeling. Next is also an LSTM layer, receiving an input of shape (1, 128) and an output of shape (1, 64), where 64 is the dimension of the output features. This layer only returns the hidden state of the last time step in the sequence. Next, the network includes a fully connected layer with input (1, 64) that uses the ReLU activation function to compress the sequence encoding into a 32-dimensional feature vector. This is a compressed temporal representation extractor, yielding the final temporal feature vector.

[0073] 3) Constructing a Transformer feature extraction network

[0074] The Transformer is a deep learning model based on self-attention, initially proposed by Vaswani et al., and widely used in natural language processing, sequence modeling, and time series prediction. Compared to traditional recurrent neural networks, it completely abandons the recursive structure and uses parallel computing to process the entire sequence, significantly improving training efficiency and model expressive power. Its core lies in dynamically allocating information weights between different time steps through the self-attention mechanism, thereby capturing long-distance dependencies and global features. The Transformer structure consists of stacked multi-layer encoders and decoders, each layer containing a multi-head attention mechanism and a feedforward neural network, possessing powerful sequence modeling capabilities. When processing complex time-series data such as CSI channel state information, the Transformer can effectively model the nonlinear relationships between channel features and global dynamic features, exhibiting good generalization ability and potential for improving localization accuracy.

[0075] like Figure 6As shown, the Transformer model of this invention comprises six layers: a positional encoding layer, a self-attention layer, a feedforward network, a regularization layer, a max-pooling layer, and a fully connected layer. The input layer receives an input of shape (1440, 1), representing a sample containing data across 1440 time steps. Since the model itself lacks a concept of sequence order, positional encoding is needed to provide positional information. Therefore, an embedding layer is first used to generate positional codes, which are then added to the input data. The input of this layer is (1440, 1), and the output is (1440, 32). The output is the sum of the original data and the corresponding positional codes, representing 1440 time steps, with each time step having a feature dimension of 32. Then, a self-attention layer is passed through, with both its input and output being (1440, 32). This layer uses two parallel attention heads for computation, with each attention head having a key vector and a query vector of 16 dimensions. This multi-head self-attention mechanism captures the relationships between different time steps in the sequence while maintaining the time and feature dimensions unchanged. In the Transformer, the attention module performs multiple parallel computations. Each parallel computation is called an attention head. The attention module splits its parameter matrix of query, key, and value N times, passing each split independently through a separate attention head. Finally, all these identical attention computations are combined to produce the final attention score. This allows for a more nuanced capture and expression of the various connections and subtle differences between each word.

[0076] Following the attention layer is a feedforward network containing two fully connected layers and residual connections for further feature processing. Its input shape is (1440, 32). The first fully connected layer has 64 nodes, so the output shape is (1440, 64). The second fully connected layer restores the dimensions to 32, with an output shape of (1440, 32). The residual connections in this network add the original input to the final output, aiding in gradient backpropagation. The output of the feedforward network undergoes Dropout and layer normalization. Dropout randomly discards a portion of neurons, reducing the number of parameters that need updating in the network structure and helping to reduce overfitting. Next, it passes through a global average pooling layer with an input of (1440, 32) and an output shape of (1, 64). This layer averages all features across the time dimension to obtain global sequence features. Finally, the output is processed through a fully connected layer with a 32-dimensional output and a ReLU activation function to extract the final feature representation for fusion. The Transformer model effectively captures long-range dependencies in a sequence through its self-attention mechanism, making it particularly suitable for handling long sequences and complex temporal patterns.

[0077] The three feature extraction networks described above extract one-dimensional features from the data, and all three extract the same dimension, but they use different methods to extract these features. 1DCNN primarily captures local features in the data through convolutional and pooling layers; LSTM can capture long-term dependencies in the data; and Transformer processes sequential data through a self-attention mechanism, efficiently capturing global dependencies. Therefore, fusing the features extracted by these three networks will definitely result in a better overall expressive power than a single network.

[0078] In this invention, the feature fusion method uses concatenation, a common technique in deep learning frameworks. Its function is to merge multiple features into a higher-dimensional feature representation. The fusion layer concatenates the output features of three models—1DCNN, LSTM, and Transformer—on the same dimension. That is, after learning features at different levels, the outputs of the three networks are merged to form a vector containing information from all three networks. This approach fully combines the advantages of each model, enabling it to simultaneously utilize local features (from 1DCNN), long temporal dependencies (from LSTM), and global features (from Transformer). The concatenation process does not transform or weight the vector content; it directly concatenates along the feature dimension, forming a high-dimensional vector that comprehensively represents the three features. This fusion method is characterized by its simple structure and strong information integrity, helping the model make decisions based on more comprehensive features in subsequent classification stages.

[0079] Step 3: Build a deep learning classification network, and train the feature extraction network and deep learning network from Step 2 offline to learn the nonlinear relationship between CSI features and corresponding spatial location information, thereby obtaining an indoor positioning model.

[0080] 1) Construct a classification network model

[0081] In classification tasks, 1DCNN, LSTM, and Transformer models aim to extract different temporal features from the input sequence to achieve accurate location classification. After feature fusion, the fused feature sequence is input into a classification learning network for further classification learning. The classification learning network consists of two fully connected layers, as illustrated in the diagram below. Figure 7 As shown, the first fully connected layer is a feature fusion layer, using the ReLU activation function. The input size is 1×96 (the output of the concatenate layer), and the output is 1×64. The function of this layer is to perform a non-linear transformation on the fused features, enhancing their expressive power.

[0082] The second layer is the classification output layer. It takes input from the preceding fully connected layer and outputs the corresponding position and class. Its function is to generate the final class probability distribution. The output layer of the classification network structure is a softmax layer containing 67 neurons, and its form is as follows:

[0083]

[0084] z i It is the input to the Softmax function, representing the score of the i-th neuron in the fully connected layer of the model. The higher the score, the greater the likelihood that the model considers the sample to belong to that category. The input score z i The natural exponent is used, and the exponential function is used to map the score in the real number range to the positive number range, because the probability value must be non-negative. This involves summing the exponential scores of all 67 neurons. This sum serves as the denominator, acting as a normalization mechanism so that the sum of the outputs of the Softmax function (i.e., the probability values ​​for each class) equals 1. Softmax(z i ) represents the probability value of the i-th category after calculation by the Softmax function.

[0085] The softmax layer takes a 1×64-dimensional vector as input to the previous layer and outputs a probability value for each location's class. It maps the output to a probability distribution representing the predicted probability for each class. Assume the output sequence after model processing is H. final Its dimension is T×d, where T is the sequence length and d is the output dimension of the last layer of the model. The calculation formula for its classification layer is shown in Equation (4):

[0086] y = softmax(W class ·h final +b class (4)

[0087] Where: h class W is the final hidden state output by the model (usually the hidden state at the last time step or the state obtained from average pooling), representing the feature representation of the entire sequence; class and b class These are the weights and biases of the classification layer.

[0088] 2) Model training and classification learning

[0089] In classification tasks, the cross-entropy loss function is typically used to measure the difference between the probability distribution of the model output and the true label. When the label is an integer category number, sparse_categorical_crossentropy is used to calculate the loss. Its mathematical expression is the same as that of standard cross-entropy, but the input form is different, and its formula is shown in equation (5) below:

[0090]

[0091] Where y is the true label of the sample. In the case of an integer category number, it is a specific integer representing the category number. Let be the predicted probability of the model for the i-th class output, which is a vector of length C, satisfying . It is an indicator function, which is 1 when y = i and 0 otherwise; C is the total number of categories. This represents the relationship between the true label y and the predicted result. The smaller the loss, the closer the prediction is to the true label, and the better the prediction effect.

[0092] During training, backpropagation is used to calculate gradients, and the Adam optimizer is used to update the model parameters until the model can accurately classify different categories. The initial learning rate is set to 0.001, and the batch size for each training iteration is 64. The curves showing the training accuracy versus loss function of this invention are as follows. Figure 8 As shown. From the appendix Figure 8 As we can see from the graph, the accuracy versus loss function curves of this invention when the number of training samples is 60 show that the loss gradually converges and the accuracy continuously improves as the number of training rounds increases. When the number of training rounds reaches 64, the accuracy of the training set reaches 99.6%, and the accuracy of the validation set reaches 90.37%. The algorithm in this paper demonstrates good convergence and accuracy in offline training, fully verifying the effectiveness and robustness of the model.

[0093] The technical means disclosed in this invention are not limited to those disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features.

Claims

1. A CSI localization method based on an improved piecewise mean method and multi-scale feature extraction in a multi-channel environment, characterized in that, Includes offline and online phases. The offline phase specifically includes the following steps: Step 1: Dimensionality reduction of CSI measurements; The obtained CSI amplitude measurements are reduced in dimension using the IPAA method; Step 2: CSI measurement feature extraction; use different deep learning networks to extract multi-scale features from CSI measurements, and then fuse the extracted features; Step 3: Classification learning; Build a deep learning classification network, and train the feature extraction network and deep learning network from Step 2 offline to learn the non-linear relationship between CSI features and corresponding spatial location information to obtain an indoor positioning model; The online phase specifically includes the following steps: Step a: CSI data measurement preprocessing; dimensionality reduction of the obtained CSI amplitude measurements is performed using the method from step 1 of the offline learning phase; Step b: Location estimation; The location estimate is obtained by using the feature extraction network and classification network obtained in steps 2 and 3 of the offline stage.

2. The CSI localization method based on improved segmented mean method and multi-scale feature extraction under multi-channel conditions according to claim 1, characterized in that, Step 1, the dimensionality reduction of CSI measurements, includes the following steps: Step 1-1: Dimensionality reduction of CSI measurements over time for each subcarrier; Step 1-1-1: For each CSI measurement channel, for each subcarrier, construct a CSI measurement time series based on the order of the collected CSI amplitude measurements according to the data packets, with a training length of 1×D; Where D represents the number of data packets; Step 1-1-2: The IPAA algorithm is used to reduce the dimensionality of the CSI measurement sequence of length 1×D. The size of the reduced sequence is 1×P, where P is the length of the reduced sequence. IPAA is represented as: Steps 1-2: Fusing the CSI-reduced sequences; Step 1-2-1: For all subcarrier CSI time series after dimensionality reduction, the sequences are spliced ​​and fused according to the subcarrier labels to obtain a CSI dimensionality reduction sequence under one channel with a size of 1×(D×P). Step 1-2-2: After dimensionality reduction of the CSI time series for all channels, the sequences are spliced ​​and fused according to the channel labels to obtain the final CSI dimensionality-reduced sequence, which has a size of 1×(D×P×Q). Where Q represents the number of channels for CSI measurements in the positioning system.

3. The CSI localization method based on improved segmented mean method and multi-scale feature extraction under multi-channel conditions according to claim 1, characterized in that, Step 2, the CSI measurement feature extraction, includes the following steps: Step 2-1: Construct a 1DCNN feature extraction network; Construct a 1DCNN feature extraction network. The overall structure includes two convolutional layers, two pooling layers, a flattening layer, and two fully connected layers. The output of the last fully connected layer is used as the output feature of this sub-feature extraction network. Step 2-2: Construct an LSTM feature extraction network; The LSTM feature extraction network consists of three layers: two LSTM layers and one fully connected layer. The LSTM feature extraction network effectively models the contextual relationships in the sequence through the two LSTM layers, and the output of the last fully connected layer is used as the output feature of this sub-feature extraction network. Steps 2-3: Construct the Transformer feature extraction network; The Transformer model consists of a positional encoding layer, a self-attention layer, a feedforward network, a regularization layer, a max-pooling layer, and a fully connected layer; the fully connected layer outputs the extracted features. Step 2-4: Multi-scale feature fusion; Perform a union operation on the features extracted in steps 2-1, 2-2 and 2-3, and use the fused features as the final result of multi-scale feature extraction as the input to the classification network.

4. The CSI localization method based on improved segmented mean method and multi-scale feature extraction under multi-channel conditions according to claim 1, characterized in that, Step 3, classification learning, includes the following steps: Step 3-1: Divide the dataset into a training set and a validation set; Step 3-2: Construct a classification network The classification network consists of two fully connected layers, uses the softmax function, employs cross-entropy as the loss function, and uses Adam as the optimizer. Step 3-3: Offline classification learning; The feature extraction network built in Steps 2-1, 2-2, and 2-3 and the classification network built in Step 3-2 are jointly trained to obtain the location estimation model.

5. The CSI localization method based on improved segmented mean method and multi-scale feature extraction under multi-channel conditions according to claim 4, characterized in that, Step 3-2 specifically involves: The output layer of the classification network structure is a softmax layer containing 67 neurons, and its form is as follows: The softmax layer takes a 1×64-dimensional vector as input to the previous layer and outputs a probability value for each location's class. It maps the output to a probability distribution representing the predicted probability for each class. Assume the output sequence after model processing is H. final Its dimension is T×d, where T is the sequence length and d is the output dimension of the last layer of the model. The calculation formula for its classification layer is as follows: y =softmax(W class ·h final +b class ) Where: h class The final hidden state output from the model is usually the hidden state at the last time step or the state obtained through average pooling, representing the feature representation of the entire sequence; W class and b class These are the weights and biases of the classification layer.

6. The CSI localization method based on improved segmented mean method and multi-scale feature extraction under multi-channel conditions according to claim 4, characterized in that, Step 3-3 specifically involves: Model training and classification learning In classification tasks, the cross-entropy loss function is typically used to measure the difference between the probability distribution of the model's output and the true label. When the label is an integer category number, sparse_categorical_crossentropy is used to calculate the loss. Its mathematical expression is the same as the standard cross-entropy, but the input format is different, and its formula is shown below: Where y is the true label of the sample, which is an integer representing the category number; It is the predicted probability of the model for the i-th type of output, satisfying I(y=i) is an indicator function, which is 1 when y=i and 0 otherwise; C is the total number of categories. During training, the backpropagation algorithm is used to calculate the gradient, and the Adam optimizer is used to update the model's parameters until the model can accurately classify different categories.

Citation Information

Patent Citations

  • Indoor positioning method based on channel state information and received signal strength

    CN110366108A

  • Indoor positioning method based on channel state information

    CN112333642B

  • Wireless positioning method based on deep learning

    CN115908547A

Cited By

  • Operation determination system and operation determination method

    JP7880191B1