Non-line-of-sight estimation algorithm based on CNN-GRU and multi-head attention mechanism
By adopting a non-sight line-of-view judgment algorithm based on CNN-GRU and multi-head attention mechanism in indoor positioning, the problem of difficulty in accurately distinguishing the non-sight line-of-view propagation environment in the prior art is solved, and higher positioning accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510061551.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to accurately distinguish the distance-of-sight (LOS) and non-line-of-sight (NLOS) propagation environment in indoor positioning, resulting in low positioning accuracy and poor reliability.
A non-horizontal judgment algorithm based on recurrent neural network (RNN) and multi-head attention mechanism is adopted, and a multi-dimensional and timing characteristics of channel data are extracted by constructing a CNN-GRU neural network model and combining the multi-head attention mechanism.
It improves the accuracy and robustness of NLOS recognition, and improves the reliability and accuracy of indoor positioning.
Smart Images

Figure CN119992183A_ABST
Abstract
Description
Technical Field
[0003] The present invention relates to the field of indoor positioning technology, and specifically relates to a non-line-of-sight judgment algorithm based on a recurrent neural network and a multi-head attention mechanism. Background Art
[0004] With the rapid development of wireless communication technology, indoor positioning technology has been widely used in smart home, industrial automation, robot navigation and emergency rescue. Unlike the traditional global navigation satellite system (GNSS), due to the severe attenuation and multipath interference of satellite signals in indoor environments, indoor positioning technology needs to rely on short-range wireless communication technologies such as ultra-wideband (UWB), Bluetooth, Wi-Fi, etc. Among them, ultra-wideband technology has become one of the preferred solutions for indoor positioning due to its high-precision time resolution and anti-multipath interference capabilities.
[0005] However, in practical applications, the signal propagation path may be blocked by walls, furniture and other obstacles, resulting in non-line-of-sight (NLOS) propagation. NLOS propagation can cause significant changes in signal characteristics, such as signal strength reduction, arrival time (ToA) deviation and multipath effect enhancement, which in turn has a serious impact on positioning accuracy. Therefore, accurately distinguishing between line-of-sight (LOS) and non-line-of-sight (NLOS) propagation environments is of great significance for improving the reliability and accuracy of indoor positioning.
[0006] At present, NLOS identification mainly relies on the following methods:
[0007] 1. Traditional method based on empirical rules:
[0008] Traditional NLOS identification methods are usually based on the statistical characteristics of the signal, such as time of arrival (ToA), angle of arrival (AoA), and signal amplitude distribution, and are classified through predefined thresholds or empirical rules. Such methods may be effective in a single scenario, but they are more sensitive to noise and multipath effects in complex indoor environments, and have low identification accuracy.
[0009] 2. Classification method based on machine learning:
[0010] With the accumulation of data and the improvement of computing power, machine learning technology has been widely used in NLOS identification. For example, models such as support vector machine (SVM), random forest (RF) and extreme gradient boosting (XGBoost) perform classification by manually extracting channel features (such as channel impulse response CIR). These methods are smarter than traditional empirical rules, but their performance depends on artificial feature design and is difficult to cope with diverse and dynamically changing indoor channel environments.
[0011] 3. Recognition method based on deep learning:
[0012] Deep learning technology has made significant progress in the field of NLOS recognition in recent years. For example, convolutional neural networks (CNNs) improve recognition by extracting spatial features of channel data, and recurrent neural networks (RNNs) can model the time series characteristics of channel data.
[0013] Regarding the above methods, someone has applied for publication number CN114339649B named "A NLOS identification system with WKNN classification"; the solution adopted in this application is to use the weighted-K neighbor algorithm to predict the NLOS situation; in addition, the existing deep learning methods still have the following problems in NLOS identification: the CNN model is not able to capture the time dependency of sequence data; the RNN model has limited performance when processing long-term dependencies and has a large computational overhead; the lack of effective enhancement of global key features in channel data leads to the negative impact of redundant features on model performance. Therefore, in order to solve the NLOS identification problem in indoor positioning, a new deep learning method that combines local feature extraction, time dependency modeling and global feature enhancement is urgently needed to improve the recognition accuracy and robustness of the model. Based on this demand, the present invention designs a non-line-of-sight estimation algorithm based on CNN-GRU and multi-head attention mechanism, which can capture multi-dimensional features and time series features better than traditional algorithms, and has higher model accuracy. Summary of the invention
[0014] In order to capture multi-dimensional features and temporal features and make the model more accurate, the present invention designs a non-line-of-sight estimation algorithm based on CNN-GRU and multi-head attention mechanism.
[0015] The implementation of the algorithm comprises the following steps:
[0016] Step 1: Write data acquisition embedded software to collect channel impulse response (CIR) through the DWM1000 module on STM32;
[0017] Step 2: Preprocess the collected CIR data and convert it into a feature matrix suitable for input into the deep learning model;
[0018] Step 3: Based on the improved CNN-GRU neural network with multi-head attention mechanism, build the NLOS estimation model and use the public dataset for training and testing.
[0019] Step 4: Use the collected CIR data to input the trained model to predict NLOS during the positioning process.
[0020] The step 1 specifically includes:
[0021] (1) The channel impulse response (CIR) of a signal is an important parameter that describes the impact of a signal on a communication channel. CIR reflects the signal's attenuation, delay, distortion, and multipath effects. In UWB communication systems, CIR is usually expressed as the superposition of multiple signal propagation paths, each of which corresponds to a specific delay and attenuation.
[0022] (2) Assuming that there is Gaussian white noise in UWB transmission, the UWB signal transmission model is: Where q(t) is the repetition time T c A single Gaussian pulse with a symbol duration of T s By N c pulses, α j ∈{-1,1} is the polarization sequence used for spectrum formation. Among them, β l and τ l They represent the attenuation coefficient and delay time of the lth path respectively, and the received signal can be expressed as follows: Where v(t) is additive white Gaussian noise.
[0023] (3) The DWM1000 module communicates with the STM32 embedded platform through the SPI interface to initialize and configure the module. The CIR data is transmitted to the host computer through the serial port.
[0024] The step 2 specifically includes:
[0025] The collected CIR data is expressed as a time domain sequence CIR(t)=[c1, c2, ..., c N ], where c i is the signal strength of the i-th sampling point in the time domain, and N is the length of the data. The standardized CIR signal is reshaped into a feature matrix X suitable for deep learning input. Assume that the length of each input sequence is The feature dimension is D, then the shape of the input matrix is X∈R D×L In this embodiment, the CIR length N is 1016 and the feature dimension D is 4.
[0026] The step 3 specifically includes:
[0027] (1) Constructing 1D convolutional layer: First, the time domain features of the input CIR data are extracted through the 1D convolutional layer. The output of the 1D convolutional layer is the feature map after the convolution operation. Suppose the input data is X∈R N×L , where N is the number of samples and L is the length of each CIR. Using a convolution kernel size of k and a convolution output dimension of d, the convolution layer output is: Y conv=Conv1D(X, k) (shape is N×d), where the data will pass through multiple convolutional layers to fully extract the important features of the signal. In this embodiment, the data will pass through three different convolutional layers, and the parameters of the convolutional layers are:
[0028] convl: input channels are 4, output channels are 64, convolution kernel size is 9, padding is 1;
[0029] conv2: input channel is 64, output channel is 64, convolution kernel size is 7, padding is 1;
[0030] conv3: input channel is 64, output channel is 128, convolution kernel size is 5, padding is 1;
[0031] (2) The feature map output by the convolution layer is passed to the GRU layer, which is used to model the temporal dependencies in the CIR data. Let the convolution output be Y conv ∈R N×d , the output of the GRU layer is Y gru ∈R N×H , where H is the number of hidden layers in the GRU layer. Through the processing of the GRU layer, the model can capture the temporal information in the CIR data. The input dimension is 128, which is the number of channels output by the convolutional layer. The hidden layer dimension is 64, and a bidirectional GRU is used, so the actual output dimension is 128.
[0032] (3) Use the output of the GRU layer to generate queries, keys, and values. Use linear transformations to map the input matrix to different query, key, and value spaces:
[0033] Q=Y gru W Q , K=Y gru W K , V = Y gru W V
[0034] Among them, W Q , W K , W V is a learnable weight matrix corresponding to the linear transformation of query, key and value. The input dimension and model dimension are 128;
[0035] (4) Calculate the attention score using the scaled dot product formula:
[0036]
[0037] Among them, d kThe softmax function is used to calculate the matching degree of each query with all keys (i.e., attention weight). The query, key, and value are processed in parallel by multiple attention heads. Each head calculates the attention score independently, and the output of each head is obtained by weighted summation. The outputs of all heads are then concatenated and sent to the final linear transformation: Multihead(Q, K, V) = Concat(head1, head2, ..., head h )W O
[0038] Among them head i is the output of the i-th attention head, and Wo is the linear transformation matrix of the output. In this embodiment, the number of heads is 4;
[0039] (5) The multi-head attention output is reduced in dimension through the global average pooling (GAP) operation to obtain a vector of fixed dimension. The pooled vector is passed to the fully connected layer to output the NLOS prediction result.
[0040] The step 4 specifically includes:
[0041] After the training is completed in step 3, the CIR data preprocessed in step 2 is passed to the trained neural network model for testing to obtain the corresponding NLOS prediction value. The data is divided into static data and dynamic data to ensure the robustness of the neural network model. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a schematic diagram of a flow chart of an embodiment of the present invention;
[0043] Figure 2 Schematic diagram of a CNN-GRU algorithm based on multi-head attention improvement according to an embodiment of the present invention;
[0044] Figure 3 is a loss and accuracy change curve during the training process of the neural network model of an embodiment of the present invention; DETAILED DESCRIPTION
[0045] The present invention will be specifically described below in conjunction with the accompanying drawings and embodiments, and the technical solution adopted by the present invention will be introduced in detail and completely. It should be noted that the described embodiments only represent a part of the present invention, not all embodiments. Therefore, the application, use, etc. of the present invention in other embodiments are not limited to the embodiments described herein. This embodiment combines public data sets and field collection experiments to identify NLOS signals using deep learning.
[0046] Figure 1This is a flow chart of a non-line-of-sight judgment algorithm based on CNN-GRU and multi-head attention mechanism of the present invention. As shown in the figure, the present invention specifically includes the following steps:
[0047] Step 1: Write data acquisition embedded software to collect channel impulse response (CIR) through the DWM1000 module on STM32;
[0048] Step 2: Preprocess the collected CIR data and convert it into a feature matrix suitable for input into the deep learning model;
[0049] Step 3: Build an NLOS estimation model, build a CNN-GRU neural network prediction model based on the improved multi-head attention mechanism, and use public datasets for training and testing.
[0050] Step 4: Use the collected CIR data to input the trained model to predict NLOS during the positioning process.
[0051] In step 1, the CIR data collection specifically includes:
[0052] (1) The experiment was conducted in the 407 conference room of the information building of a certain school. The spatial layout of the conference room includes tables, chairs, cabinets, blackboards and other obstacles. Multiple groups of LOS and NLOS paths were set up. The obstacles mainly included human bodies, glass and wood materials. The scene complexity was high and was used to simulate the complex indoor environment. The experiment used the DWM1000 chip in UWB technology as the signal acquisition device, equipped with an STM32 microcontroller, to realize the real-time acquisition of the channel impulse response (CIR) signal of a single base station and a single tag. The collected CIR vector is:
[0053] CIR(t)=[c1,c2,…,c N ]
[0054] where c i is the signal strength of the i-th sampling point in the time domain, and N is the length of the data. In this embodiment, N is 1016.
[0055] (2) When collecting CIR data, the base station (Anchor) needs to be fixed in a stable position and ensure that the base station is connected to the STM32 module and can receive signals from the tag (Tag). The tag collects data at all predetermined positions, each of which is at the four corners of a square with a side length of 0.5m. The tag transmits signals through the DWM1000 module to ensure that when there are no obstructions (such as pedestrians) between the base station and the tag, the signal between the base station and the tag is line-of-sight (LOS) propagation; then, by placing pedestrians or other appropriate obstructions, a non-line-of-sight (NLOS) propagation scenario is created between the base station and the tag, so that the presence of obstructions can be artificially controlled to simulate different propagation conditions.
[0056] (3) When collecting data, first collect data in the LOS scenario, ensure that there are no obstructions between the base station and the tag, and after the tag is completely still and the signal is stable, start collecting CIR data, record the delay and amplitude information of each signal point, and repeat the collection five times to ensure that five independent LOS data samples are obtained, which will be used for subsequent analysis and annotation. Then, collect data in the NLOS scenario, ensure that there are pedestrians or other obstructions between the base station and the tag, and after the tag is completely still and the signal is stable in the new NLOS environment, start collecting CIR data, record the delay and amplitude information of each signal point, and repeat the collection five times to obtain five different NLOS data samples, which will be labeled as NLOS data in subsequent analysis.
[0057] (4) After data collection is completed, all NLOS data need to be manually labeled to ensure that each set of data can correctly distinguish between LOS and NLOS states and provide a high-quality data set for subsequent classification tasks.
[0058] In step 2, specifically including:
[0059] The collected CIR data is expressed as a time domain sequence CIR(t)=[c1, c2, ..., c N ], where c i is the signal strength of the i-th sampling point in the time domain, and N is the length of the data. The standardized CIR signal is reshaped into a feature matrix X suitable for deep learning input. Assume that the length of each input sequence is The feature dimension is D, then the shape of the input matrix is X∈R D×L .
[0060] In step 3, it specifically includes:
[0061] The model was trained on a public dataset, and its performance in actual scenarios was verified by self-collecting data. The initial training of the model was completed using the public eWINE UWB LOS and NLOS Data Set dataset, which contains seven different indoor environments: Office Area 1, Office Area 2, Small Apartment, Small Studio, Kitchen and Living Room, Bedroom and Kitchen. 3,000 LOS samples and 3,000 NLOS samples were collected for each scene, totaling 42,000 samples, including 21,000 LOS and NLOS samples. To prevent the model from relying on a specific location, the dataset randomizes the samples, does not carry positioning reference information, and is only used for training the LOS / NLOS channel detection model. The main features of the dataset include the channel impulse response (CIR) signal and its corresponding statistical characteristics (such as noise standard deviation, channel power, maximum noise value, etc.). The length of each CIR sample is 1016 and the time resolution is 1 nanosecond. The model structure is as follows Figure 2 As shown:
[0062] (1) Constructing 1D convolutional layer: First, the time domain features of the input CIR data are extracted through the 1D convolutional layer. The output of the 1D convolutional layer is the feature map after the convolution operation. Suppose the input data is X∈R N×L , where N is the number of samples and L is the length of each CIR. Using a convolution kernel size of k and a convolution output dimension of d, the convolution layer output is: Y conv =Conv1D(X, k) (shape is N×d), where the data will pass through multiple convolutional layers to fully extract the important features of the signal. In this embodiment, the data will pass through three different convolutional layers, and the parameters of the convolutional layers are:
[0063] conv1: input channels are 4, output channels are 64, convolution kernel size is 9, and padding is 1;
[0064] conv2: input channel is 64, output channel is 64, convolution kernel size is 7, padding is 1;
[0065] conv3: input channel is 64, output channel is 128, convolution kernel size is 5, padding is 1;
[0066] (2) The feature map output by the convolution layer is passed to the GRU layer, which is used to model the temporal dependencies in the CIR data. Let the convolution output be Y conv ∈R N×d , the output of the GRU layer is Y gru ∈R N×H, where H is the number of hidden layers in the GRU layer. Through the processing of the GRU layer, the model can capture the temporal information in the CIR data. The input dimension is 128, which is the number of channels output by the convolutional layer. The hidden layer dimension is 64, and a bidirectional GRU is used, so the actual output dimension is 128.
[0067] (3) Use the output of the GRU layer to generate queries, keys, and values. Use linear transformations to map the input matrix to different query, key, and value spaces:
[0068] Q=Y gru W Q , K=Y gru w K , V = Y gru W V
[0069] Among them, W Q , W K , W V is a learnable weight matrix corresponding to the linear transformation of query, key and value. The input dimension and model dimension are 128;
[0070] (4) Calculate the attention score using the scaled dot product formula:
[0071]
[0072] Among them, d k The softmax function is used to calculate the matching degree of each query with all keys (i.e., attention weight). The query, key, and value are processed in parallel by multiple attention heads. Each head calculates the attention score independently, and the output of each head is obtained by weighted summation. The outputs of all heads are then concatenated and sent to the final linear transformation: Multihead(Q, K, V) = Concat(head1, head2, ..., head h )W O
[0073] Among them head i is the output of the ith attention head, W O is the output linear transformation matrix. In this embodiment, the number of heads is 4;
[0074] (5) The multi-head attention output is reduced in dimension through the global average pooling (GAP) operation to obtain a vector of fixed dimension. The pooled vector is passed to the fully connected layer to output the NLOS prediction result.
[0075] The accuracy (ACC), positive prediction value (PPV), recall rate (Recall), and F1 score of the algorithm in this embodiment are shown in Table 1 below. The accuracy of the model on the public dataset is 87.73%. In contrast, the recognition accuracy of the traditional CNN-LSTM is 82.14%.
[0076] Table 1
[0077]
[0078] In step 4, specifically including:
[0079] After the training is completed in step 3, the CIR data preprocessed in step 2 is passed into the trained neural network model for testing to obtain the corresponding NLOS prediction value.
[0080] The training process of the model on the conference room dataset is as follows Figure 3 As shown: Figure 3 (a) is a loss change curve during the training process of the neural network model of an embodiment of the present invention; Figure 3 (b) is the training set accuracy change curve during the neural network model training process; Figure 3 (c) is the test set accuracy change curve during the network model training process;
[0081] After testing, the CNN-GRU neural network prediction model improved based on the multi-head attention mechanism in this embodiment has an accuracy of 98.89% on the local data set, and the accuracy of pedestrian dynamic data is 94.25%.
Claims
1. A non-line-of-sight estimation algorithm based on CNN-GRU and multi-head attention mechanism, It is characterized in that The method comprises the following steps: Step 1: Write data acquisition embedded software to collect channel impulse response (CIR) through the DWM1000 module on STM32; Step 2: Preprocess the collected CIR data and convert it into a feature matrix suitable for input into the deep learning model; Step 3: Based on the CNN-GRU neural network improved by the multi-head attention mechanism, a NLOS recognition model is constructed, and public datasets are used for training and testing; Step 4: Use the collected CIR data to input the trained model to predict NLOS during the positioning process.
2. According to claim 1, a non-line-of-sight recognition algorithm based on CNN-GRU and multi-head attention mechanism is characterized in that: The step 1 specifically includes: (1) The channel impulse response (CIR) of a signal is an important parameter that describes the impact of a signal on a communication channel. The CIR reflects the signal's attenuation, delay, distortion, and multipath effects. In UWB communication systems, the CIR is usually expressed as the superposition of multiple signal propagation paths, each of which corresponds to a specific delay and attenuation. (2) Assuming that there is Gaussian white noise in UWB transmission, the UWB signal transmission model is: Where q(t) is the repetition time T c A single Gaussian pulse with a symbol duration of T s By N c pulses, α j ∈{-1,1} is the polarization sequence used for spectrum formation. Among them, β l and τ l They represent the attenuation coefficient and delay time of the lth path respectively, and the received signal can be expressed as follows: Where v(t) is additive Gaussian white noise; (3) The DWM1000 module communicates with the STM32 embedded platform through the SPI interface, initializes and configures the module, and transmits the CIR data to the host computer through the serial port.
3. The non-line-of-sight estimation algorithm based on CNN-GRU and multi-head attention mechanism according to claim 1, characterized in that: The step 2 specifically includes: The collected CIR data is expressed as a time domain sequence CIR(t)=[c1, c2, ..., c N ], where c i is the signal strength of the i-th sampling point in the time domain, N is the length of the data, and the standardized CIR signal is reshaped into a feature matrix X suitable for deep learning input. Assume that the length of each input sequence is The feature dimension is D, then the shape of the input matrix is X∈R D×L .
4. The non-line-of-sight estimation algorithm based on CNN-GRU and multi-head attention mechanism according to claim 1, characterized in that: The step 3 specifically includes: (1) Constructing a 1D convolutional layer: First, the time domain features of the input CIR data are extracted through the 1D convolutional layer. The output of the 1D convolutional layer is the feature map after the convolution operation. Suppose the input data is X∈R N×L , where N is the number of samples, L is the length of each CIR, the convolution kernel size is k, and the convolution output dimension is d, then the convolution layer output is: Y conc =Conv1D(X, k) (shape is N×d), where the data will pass through multiple convolutional layers to fully extract the important features of the signal; (2) The feature map output by the convolutional layer is passed to the GRU layer. The GRU layer is used to model the temporal dependency in the CIR data. Let the convolution output be Y conv ∈R N×d , the output of the GRU layer is Y gru ∈R N×H , where H is the number of hidden layers of the GRU layer. Through the processing of the GRU layer, the model can capture the timing information in the CIR data; (3) Use the output of the GRU layer to generate queries, keys, and values, and map the input matrix to different query, key, and value spaces through linear transformation: Q=Y gru W Q ,K=Y gru W K ,V=Y gru W V Among them, W Q , W K , W V is a learnable weight matrix corresponding to the linear transformation of query, key and value respectively; (4) Calculate the attention score using the scaled dot product formula: Among them, d k The softmax function is used to calculate the matching degree of each query with all keys (i.e., attention weight). The query, key, and value are processed in parallel by multiple attention heads. Each head calculates the attention score independently and obtains the output of each head by weighted summation. The outputs of all heads are then concatenated and sent to the final linear transformation: Multihead(Q,K,V)=Concat(head1,head2,...,head h )W O Among them head i is the output of the i-th attention head, W O is the output linear transformation matrix; (5) The multi-head attention output is reduced in dimensionality through the global average pooling (GAP) operation to obtain a vector of fixed dimensionality. The pooled vector is passed to the fully connected layer to output the NLOS prediction result.
5. The non-line-of-sight estimation algorithm based on CNN-GRU and multi-head attention mechanism according to claim 1, characterized in that: The step 4 specifically includes: After the training is completed in step 3, the CIR data preprocessed in step 2 is passed into the trained neural network model for testing to obtain the corresponding NLOS prediction value.
Citation Information
Patent Citations
A NLOS recognition system based on WKNN classification
CN114339649B
Cited By
Distance measurement method based on multi-modal deep learning
CN121232218A