Multi-class non-line-of-sight signal identification system and method based on spatio-temporal characteristics
By constructing a spatiotemporal feature model, using feature extractor, attention mechanism and feature fusion device to extract and fuse the spatiotemporal features of channel impulse response data, the positioning accuracy problem of UWB indoor positioning system under NLOS occlusion is solved, and the accurate identification and positioning accuracy of multi-category NLOS signals are achieved.
Patent Information
- Application Number
- CN202510542322.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing UWB-based indoor positioning system has inaccurate positioning accuracy under NLOS occlusion, and it is difficult for existing methods to effectively identify the severity of the NLOS signal, resulting in errors in positioning solution.
A multi-category NLOS signal recognition method based on spatiotemporal features is designed. By constructing a spatiotemporal feature model, a feature extractor, attention mechanism and feature fusion device extract and fuse the spatiotemporal features of channel impulse response data, and a light attention mechanism is used to identify multi-category NLOS signal.
It realizes the rapid and accurate identification of multi-category NLOS signals on embedded devices with limited computing resources, improves indoor positioning accuracy and robustness, and can effectively evaluate the severity of NLOS signals.
Smart Images

Figure CN120416772A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of wireless positioning technology, and particularly relates to a multi-category non-line-of-sight signal recognition system and method based on spatio-temporal features. Background Art
[0002] With the development of Internet of Things technology, the demand for precise indoor positioning in various industries has increased. UWB signals contain line-of-sight (LOS) information of the indoor environment, and indoor positioning systems can locate and track objects based on the collected UWB signals. However, indoor positioning systems based on UWB are vulnerable to non-line-of-sight (NLOS) conditions, resulting in serious errors in positioning calculations. Therefore, it is crucial to accurately identify and mitigate indoor NLOS signals.
[0003] In the task of indoor NLOS scenario recognition, the manually designed UWB signal features can more accurately describe the characteristics of indoor NLOS scenarios. For example, the "Method for Identifying Line-of-Sight / Non-Line-of-Sight Paths Based on Channel Impulse Response Energy Distribution" (CN109151724A) applied by Central South University to this technology discriminates signal energy features through various machine learning methods to determine the LOS / NLOS binary classification type of the signal. However, the design of these features often requires a large amount of prior knowledge, and the process of feature extraction is relatively cumbersome. Classification techniques based on deep learning have gradually become a new hotspot in the field of NLOS recognition. Research on deep feature extraction and structure optimization of deep networks in deep learning has promoted the further improvement of NLOS scenario recognition accuracy. For example, the "Method for Identifying NLOS Signals of UWB Based on Deep Learning" (CN109151724A) applied by Beijing University of Posts and Telecommunications builds a two-stream deep neural network and trains the network using a common UWB measurement data set to identify LOS / NLOS signals collected in a real environment. However, it requires a large amount of data for training, and the process is cumbersome. In addition, most current NLOS classification methods only focus on the simple binary classification of LOS / NLOS, resulting in the inability to fully evaluate the severity of NLOS signals, and thus unable to utilize the effective components in NLOS signals to improve positioning accuracy. Therefore, it is urgent to improve the existing technology. Summary of the Invention
[0004] Aiming at the problem of positioning drift of indoor positioning systems based on UWB under NLOS occlusion, the invention provides a multi-category NLOS signal recognition method based on spatio-temporal features: extracting spatio-temporal features from channel impulse response data through a designed model, fusing key features of a lightweight attention mechanism, and performing multi-category NLOS signal recognition based on the fused features. The method adopted by the invention to solve its technical problems is:
[0005] Collect channel impulse response data separately in an indoor environment with multiple NLOS scenarios, and construct a training dataset, a validation dataset, and a test dataset based on the channel impulse response data of the indoor environment;
[0006] Construct a spatio-temporal feature model, which mainly includes four main functional parts: a feature extractor, an attention mechanism, a feature fuser, and a classification recognizer; among them, the feature extractor is used to deeply extract the spatial and temporal features of the channel impulse response data in the UWB signal; the attention mechanism is used to adjust the weight of each spatio-temporal feature of the channel impulse response data; the feature fuser is used to fuse the original deep spatio-temporal features processed by the feature extractor and the key features of the attention mechanism, and output more detailed fused features; the classification recognizer predicts the NLOS type of the UWB signal to be recognized based on the fused features;
[0007] Label the CIR data from the collected UWB dataset, and divide it into a training set, a validation set, and a test set according to a preset ratio, where the label is used to characterize the target non-line-of-sight category to which the CIR data belongs. The CIR dataset contains several scenarios, and each scenario contains several target data; input the CIR data in the training dataset into the initial network, and use the preset loss function to iteratively train the initial network until convergence to obtain a converged network. During the iteration process, continuously adjust the network parameters in the feature extractor, attention mechanism, feature fuser, and classification recognizer to make the overall loss function meet the accuracy requirements; input the validation dataset into the converged network to verify the NLOS recognition ability of the converged network. When the verification is passed, obtain the trained spatio-temporal feature model; input the CIR data in the test dataset into the trained spatio-temporal feature network to obtain the recognition result of the NLOS signal category output by the spatio-temporal feature model.
[0008] The feature extractor consists of a spatial feature extraction part based on a one-dimensional convolutional network and a temporal feature extraction part based on a long short-term memory network; respectively deeply extract the spatial and temporal features of the channel impulse response data of the UWB signal;
[0009] The one-dimensional convolutional network for extracting the spatial features of the channel impulse response data in the UWB signal consists of three serial batch normalization convolutional filters, and the three serial convolutional filters are the first convolutional filter, the second convolutional filter, and the third convolutional filter; the kernel sizes of the first convolutional filter to the third convolutional filter are K h ×K w ,2K h ×K w ,4K h ×K w ; The definitions of the above convolutional filters are as follows:
[0010]
[0011] Wherein, Z = [z1, z2,..., z L is the input channel impulse response signal, with a length L = 512, represents the output feature map after convolution operation; the feature map has a spatial dimension of H×W and channels C; V = [v1, v2,..., v C is the set of learned filter kernels, where each filter v c is defined as
[0012] Adding normalization for the overall regularization of the feature extractor is defined as follows:
[0013]
[0014] Wherein, γ BN is the learnable parameter, μ and σ 2 are the batch mean and variance, and ε is a small constant added to the variance for numerical stability; after the batch normalization operation, the output is passed through the rectified linear activation function; a pooling layer is added after the convolutional neural network module for spatial downsampling of the feature map;
[0015] The long short-term memory network for extracting the time features of the channel impulse response signal consists of long short-term memory units with 512 hidden units; after the deep spatial features extracted by the one-dimensional convolutional network in the feature extractor are flattened by the flattening layer, they are fed into the long short-term memory network for time feature extraction to generate the spatio-temporal feature sequence; among them, the calculation of a long short-term memory network memory block is as follows:
[0016] f t = σ(w f · [h t-1 , x t + b f ) (3)
[0017] i t = σ(w i · [h t-1 , x t + b i ) (4)
[0018]
[0019]
[0020] O t = σ(w o· [h t-1 , x t + b o ) (7)
[0021] h t = o f · tanh(C t ) (8)
[0022] In the formula, f t represents the activation vector of the forget gate, represents the activation vector of the input / update gate, C×t is the cell input activation vector, C t is the current cell memory, O t represents the activation vector of the output gate, h t is the current cell output; the bias vectors and weight matrices of the input gate (i), output gate (o), forget gate (f), and memory cell (c) are represented by b and w, h t-1 is the output of the previous cell, c t-1 is the memory of the previous cell, σ represents the sigmoid function, and "·" represents the Hadamard product;
[0023] The attention mechanism for adjusting the weights of each spatio-temporal feature of the channel impulse response data is mainly composed of an efficient channel attention mechanism; a global descriptor corresponding to the spatio-temporal features of the channel impulse response data is generated, and the formula is as follows:
[0024]
[0025] In the formula, y c represents the global descriptor of the c-th channel, c = 1, 2,..., C; x c (i, j) represents the value of the input feature X at the position (i, j) on the c-th channel; then, the local dependencies between channels are modeled through one-dimensional convolution, and the convolution kernel size k is adaptively determined according to the number of channels C:
[0026]
[0027] In the formula, ψ(C) is a function based on the number of channels; γ and b are hyperparameters used to adjust the range of the kernel size; |·| odd represents adjusting the result to the nearest odd number; next, the channel attention weights are generated, and the formula is as follows:
[0028] σ c = f conv1d (y, k) = Conv1D(y, W k ) (11)
[0029] s c = sigmoid(σc ) (12)
[0030] In the formula, f conv1d represents a one-dimensional convolution operation, and the convolution kernel weights are y = [y1, y2,..., y C is the global descriptor vector; σ c is the intermediate result after convolution; s c is the attention weight for the c-th channel, with a range between [0, 1]; sigmoid(·) is the activation function; finally, the weights generated above are weighted to the original input feature H:
[0031] X c ′ = s c ×X c (13)
[0032] In the formula, X c ′ represents the output feature after weighting for the c-th channel; X c represents the original input feature for the c-th channel;
[0033] The feature fuser fuses the original depth - spatio - temporal feature vector H lstm obtained from the feature extractor processing with the key feature vector H eca of the attention mechanism to generate a fused feature vector F concat = [H lstm , H eca , and then flattens it using a flattening layer; then two fully - connected layers are used, with a ReLU layer and a dropout layer embedded between the first fully - connected layer and the second fully - connected layer; further, the number of neurons in the first fully - connected layer in the feature fuser is much larger than that in the second fully - connected layer; the second fully - connected layer is defined as the output layer and consists of five fully - connected neurons;
[0034] The output of the classification recognizer is based on the output of the feature fuser, and the output of the softmax layer is defined as the output class probability to predict the NLOS scenario recognition result of the UWB signal to be recognized, which is achieved through the following formula:
[0035]
[0036] where and are the is - th and js - th input elements respectively, Softmax(·) is the softmax activation function, and exp(·) is the exponential function; according to 's output size, the LOS / NLOS matching result in five scenarios is recognized.
[0037] The present invention also designs a multi-class recognition system for UWB NLOS signals based on a spatio-temporal feature model, including:
[0038] Data acquisition module: used to respectively acquire corresponding channel impulse response data in an indoor environment with multiple NLOS scenarios, and construct a classification training data set based on the channel impulse response data collected in the multiple NLOS scenarios;
[0039] Model construction module: used to construct a spatio-temporal feature model, the spatio-temporal feature model includes a feature extractor, an attention mechanism, a feature fusion unit and a classification recognizer; among them, the feature extractor is used to extract time and space features for NLOS multi-class recognition from the channel impulse response data; the attention mechanism is used to further extract key features from the time and space features extracted by the feature extractor to improve the NLOS multi-class recognition accuracy; the feature fusion unit is used to fuse the original spatio-temporal features of the feature extractor and the key features of the attention mechanism; the classification recognizer judges the NLOS scenario type of the actual environment channel impulse response data based on the fused features output by the feature fusion unit; Model training module: used to perform classification recognition training on the spatio-temporal feature model using the training data set until the spatio-temporal feature model is trained to convergence, and the classification recognizer judges the NLOS scenario type of the unlabeled UWB signal to a certain accuracy rate, then the training ends; Classification recognition module: used to obtain the NLOS type of the channel impulse response data to be recognized, input the channel impulse response data to be recognized into the feature extractor, attention mechanism, feature fusion unit and classification recognizer of the trained spatio-temporal feature model, and output the NLOS scenario type of the channel impulse response data to be recognized through the classification recognizer.
[0040] The present invention also provides an electronic device, including a processor and a memory, wherein a computer program is stored in the memory, and the processor is used to execute the steps of the method described above by calling the computer program stored in the memory.
[0041] The beneficial effects of the present invention are mainly manifested in:
[0042] (1) The multi-class NLOS signal recognition method based on spatio-temporal features of the present invention, based on the relatively distinguishable channel impulse response data collected in multi-class NLOS scenarios, extracts deep spatio-temporal features for distinguishing similar NLOS scenarios with the help of a spatio-temporal feature model, and further uses an attention mechanism to extract key NLOS characterization features, so that the fused features output by the feature fusion unit have sufficient multi-class NLOS signal distinguishability, so that the classification recognizer can judge the NLOS signal type of the channel impulse response data to achieve accurate multi-class recognition of NLOS signals in the UWB positioning network under indoor scenarios.
[0043] (2) This method is based on a lightweight attention mechanism to assist in channel impulse response feature extraction, thus achieving a balance between high classification accuracy and low processing time, and enabling fast deployment on embedded devices with limited computing resources. Description of the Drawings
[0044] Figure 1 It is the flowchart of the multi-class NLOS signal recognition method based on deep learning in the preferred embodiment of this application.
[0045] Figure 2 It is the schematic diagram of the sub-process of step S3.
[0046] Figure 3 It is the comparison chart of the NLOS recognition and classification effects of the spatio-temporal feature model.
[0047] Figure 4 It is the comparison chart of the recognition results between the spatio-temporal feature model and other mainstream classification models.
[0048] Figure 5 It is the recognition effect diagram of the NLOS multi-class recognition system. Figure 6 It is the schematic diagram of the module structure of the NLOS multi-class recognition system. Detailed Embodiments
[0049] The present invention will be further described below in conjunction with specific embodiments.
[0050] Embodiment 1, a multi-class NLOS signal recognition method based on deep learning, the specific process can be as Figure 1 shown, including:
[0051] Step S1: Collect the channel impulse response data of UWB signals respectively in an indoor environment with multiple NLOS scenarios, and construct a training dataset based on the collected channel impulse response data.
[0052] Step S2: Construct a spatio-temporal feature model, which is composed of four parts: a feature extractor, an attention mechanism, a feature fuser, and a classification recognizer. Among them, the feature extractor is used to extract the time and space features for NLOS multi-class recognition from the channel impulse response data; the attention mechanism is used to further extract key features from the time and space features extracted by the feature extractor to improve the NLOS multi-class recognition accuracy; the feature fuser is used to fuse the original spatio-temporal features of the feature extractor and the key features of the attention mechanism; the classification recognizer judges the NLOS scenario type of the actual environment channel impulse response data based on the fused features output by the feature fuser.
[0053] Step S3: Select channel impulse response datasets with different occlusion types from the training dataset for spatio-temporal feature model classification learning and training until the spatio-temporal feature model is trained to convergence to obtain a converged network, and the classification recognizer reaches a certain accuracy rate for the NLOS scenario recognition of unlabeled UWB signals, then the training ends.
[0054] Step S4: Input the channel impulse response data to be recognized into the feature extractor, attention mechanism, feature fusion, and classification recognizer of the trained spatio-temporal feature model, and output the NLOS scenario type of the channel impulse response data to be recognized through the classification recognizer.
[0055] It can be understood that in the multi-class NLOS signal recognition method based on spatio-temporal features of this embodiment, since channel impulse response data with high distinguishability in multiple NLOS scenarios is collected, deep spatio-temporal features for distinguishing similar NLOS scenarios are extracted by means of a spatio-temporal feature model, and sufficient attention is given to important features based on the attention mechanism, so that the fusion features output by the feature fusion have sufficient multi-class NLOS signal distinguishability. Therefore, the classification recognizer can judge the NLOS signal type of the channel impulse response data to achieve accurate multi-class recognition of NLOS signals in the UWB positioning network under changing scenarios, greatly improving the robustness of the model and the accuracy of multi-class recognition results.
[0056] It can be understood that this method is based on deep learning and combines a lightweight and efficient channel attention mechanism to assist in channel impulse response feature extraction. There is no need to manually select and extract UWB signal features, which helps to identify deep spatio-temporal features with higher discrimination contributions, thus achieving a balance between high classification accuracy and low processing time, and enabling fast and effective discrimination of NLOS signals that are often confused by other methods on embedded devices with limited computing resources.
[0057] Specifically, in the step S1, the channel impulse response data of the UWB signal is collected respectively in an indoor environment with multiple NLOS scenarios. The channel impulse response data is a UWB signal feature used in the signal field to identify LOS signals and NLOS signals. The multiple NLOS scenarios refer to the scenarios that cause different degrees of blocking effects and multipath propagation effects on the UWB signal, including but not limited to wall occlusion, human body occlusion, glass occlusion, metal obstacles, and other common indoor obstacle occlusion scenarios. And, a training data set is constructed based on the collected indoor environment channel impulse response data for the subsequent classification and recognition training of the spatio-temporal feature model. Optionally, due to the large differences in the dielectric constants of multi-category NLOS scenarios, there are certain differences in the channel impulse response data under multi-category NLOS scenarios, which is beneficial to improving the robustness of the model. It can be understood that the present invention does not require extensive collection of UWB test data in the indoor environment, and only needs to collect a small amount of channel impulse response data under multi-category NLOS environments.
[0058] In the step S2, a spatio-temporal feature model is constructed. Among them, the spatio-temporal feature model mainly includes four main functional parts: a feature extractor, an attention mechanism, a feature fusion unit, and a classification and recognition unit. Among them, the feature extractor is used to deeply extract the spatial and temporal features of the channel impulse response data in the UWB signal; the attention mechanism is used to adjust the weight of each spatio-temporal feature of the channel impulse response data; the feature fusion unit is used to fuse the original deep spatio-temporal features processed by the feature extractor and the key features of the attention mechanism, and output more detailed fusion features; the classification and recognition unit predicts the NLOS type of the UWB signal to be recognized based on the fusion features to achieve multi-classification recognition of NLOS signals, providing support for accurately and effectively reducing ranging errors in the subsequent process. Among them, the feature extractor uses a spatio-temporal feature extraction neural network, preferably including a one-dimensional convolutional neural network with multi-serial batch normalization and a long short-term memory network.
[0059] As Figure 2 shown, in the step S3, channel impulse response data sets of different occlusion types are selected from the training data set for spatio-temporal feature model classification learning training until the spatio-temporal feature model is trained until convergence to obtain a converged network, and the classification and recognition unit can judge the NLOS scenario type of the unlabeled UWB signal, including the following content:
[0060] Step S31: Select the channel impulse response data of multi-class LOS / NLOS signals from the training dataset, and use the feature extractor to extract the spatio-temporal features thereof. During the iteration process, continuously adjust the network parameters of the feature extractor to obtain and label the spatio-temporal features of the channel impulse response data in the multi-class LOS / NLOS signals of the training dataset, so as to construct a spatio-temporal feature training dataset.
[0061] Step S32: Select the channel impulse response data of multi-class LOS / NLOS signals and the corresponding obtained spatio-temporal features from the training dataset, use the attention mechanism to adjust the weights of the spatio-temporal features of the channel impulse response data to obtain key features, and then fuse the original spatio-temporal features and the key features through the feature fuser, and add the corresponding NLOS signal type labels, thereby constructing a fused feature training dataset.
[0062] Step S33: Set the overall loss function of the spatio-temporal feature model, input the constructed spatio-temporal feature training dataset and the fused feature training dataset into the classification recognizer respectively. During the iteration process, continuously classify and identify the network parameters of the recognizer respectively, so that the overall loss function meets the accuracy requirements. Finally, the classification recognizer judges the NLOS signal type of the fused features, and finally obtains the spatio-temporal feature model with the attention mechanism and the spatio-temporal feature model without the attention mechanism respectively.
[0063] It can be understood that the attention mechanism is not unique, and preferably includes an efficient attention mechanism. In order to fully compare the effectiveness of the attention mechanism, other attention mechanisms can be used to compare with the attention mechanism selected in the present invention, and the comparison results are as Figure 3 shown. It can be found that the recognition accuracy rate of the NLOS signal type under the efficient channel attention mechanism is the highest, and it performs excellently in terms of recognition resource consumption and recognition speed.
[0064] Among them, when iteratively training the initial network, the global loss function adopted adds a regulation factor on the basis of the cross-entropy loss to give more attention to these multi-class NLOS signals with classification challenges, which is expressed as follows:
[0065]
[0066] In the formula, α is a balance factor used to balance the uneven proportion of positive and negative samples itself, γ is a focusing parameter used to control the contribution of easy-to-classify samples and difficult-to-classify samples to the loss, is the prediction probability of the i y th class.
[0067] In step S4, the channel impulse response data to be recognized is input into the feature extractor, attention mechanism, feature fusion unit, and classification and recognition unit of the trained spatio-temporal feature model, and the NLOS scenario type of the channel impulse response data to be recognized is output through the classification and recognition unit. Specifically, the channel impulse response data to be recognized is sequentially input into the trained feature extractor and attention mechanism to extract the spatio-temporal features and key features of the channel impulse response data to be recognized, and then fused by the feature fusion unit. The trained classification and recognition unit performs multi-class NLOS recognition based on the fused features and outputs the NLOS scenario type of the channel impulse response data to be recognized, so as to realize the detailed recognition of NLOS occlusion between base stations and between base stations and tags in the multi-base station indoor positioning system, which is beneficial to accurately evaluating the severity of NLOS and multipath effects.
[0068] The above steps S1 to S3 are offline training, and step S4 is online recognition. Of course, in other embodiments of the present invention, steps S1 to S3 can also adopt online training.
[0069] It can be understood that in order to prove the effectiveness of the UWB NLOS multi-class recognition method based on the spatio-temporal feature model of the present invention, the inventor of the present application compared this method with the current mainstream NLOS recognition model on the validation data set, and the experimental results are as Figure 4 shown. It can be found that the recognition prediction rate and recall rate for the five types of NLOS signals mostly exceed 95.5%, indicating that the method proposed by the present invention has a strong recognition and classification ability for NLOS signals.
[0070] In addition, as Figure 5 shown, another embodiment of the present invention further provides a UWB NLOS multi-class recognition system based on spatio-temporal features, preferably adopting the method described above. The system includes:
[0071] Data acquisition module: used to respectively collect the corresponding channel impulse response data in an indoor environment with multiple NLOS scenarios, and construct a classification training data set based on the channel impulse response data collected in the multiple NLOS scenarios.
[0072] Model construction module: used to construct a spatiotemporal feature model, which includes a feature extractor, an attention mechanism, a feature fusion mechanism, and a classification identifier. The feature extractor is used to extract temporal and spatial features for NLOS multi-category identification from the channel impulse response data; the attention mechanism is used to further extract key features from the temporal and spatial features extracted by the feature extractor to improve the accuracy of NLOS multi-category identification; the feature fusion mechanism is used to fuse the original spatiotemporal features of the feature extractor with the key features of the attention mechanism; the classification identifier determines the NLOS scenario type of the actual environment channel impulse response data based on the fused features output by the feature fusion mechanism.
[0073] Model training module: used to perform classification and recognition training on the spatiotemporal feature model using the training data set until the spatiotemporal feature model training converges to obtain a converged network, and the classification identifier reaches a certain accuracy rate for NLOS scene recognition of unlabeled UWB signals, and the training ends;
[0074] Classification and recognition module: used to obtain the NLOS type of the channel impulse response data to be identified, input the channel impulse response data to be identified into the feature extractor, attention mechanism, feature fusion device and classification identifier of the trained spatiotemporal feature model, and output the NLOS scenario type of the channel impulse response data to be identified through the classification identifier.
[0075] It can be understood that the UWB NLOS multi-classification recognition system based on the spatiotemporal feature model of this embodiment collects channel impulse response data with high discrimination in multiple NLOS scenarios, and extracts deep spatiotemporal features for distinguishing similar NLOS scenarios with the help of a deep learning network, and gives sufficient attention to important features based on a lightweight attention mechanism, so that the fused features output by the feature fuser have sufficient multi-category NLOS signal discrimination, so that the classification identifier can determine the NLOS signal type of the channel impulse response data, so as to realize multi-classification and accurate recognition of NLOS signals in the UWB positioning network under changing scenarios, thereby greatly improving the robustness of the model and the accuracy of the multi-classification recognition results.
[0076] Finally, it should be noted that the above examples are merely specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples and is subject to numerous variations. All variations that can be directly derived or conceived by a person of ordinary skill in the art from the disclosure of the present invention are considered to be within the scope of protection of the present invention.
Claims
1. A multi-class non-line-of-sight signal recognition method based on spatio-temporal features, characterized in that Including the following steps: Collect CIR data respectively in an indoor environment with multiple NLOS scenarios, and construct a training dataset, a validation dataset, and a test dataset based on the CIR data of the indoor environment; construct a spatio-temporal feature model, which mainly includes four main functional parts: a feature extractor, an attention mechanism, a feature fuser, and a classification recognizer; among them, the feature extractor is used for deep extraction of the spatial and temporal features of the CIR data in the UWB signal; the attention mechanism is used to adjust the weight of each spatio-temporal feature of the CIR data; the feature fuser is used to fuse the original deep spatio-temporal features obtained from the processing of the feature extractor and the key features of the attention mechanism, and output the fused features; the classification recognizer predicts the NLOS type of the UWB signal to be recognized based on the fused features, so as to realize the multi-classification recognition of non-line-of-sight signals, and provide support for accurately and effectively alleviating the ranging error in the follow-up. Label the CIR data from the collected UWB dataset, and divide it into a training set, a validation set, and a test set according to a preset ratio, where the label is used to represent the target non-line-of-sight category to which the CIR data belongs. The CIR dataset contains several scenarios, and each scenario contains several target data; input the CIR data in the training dataset into the initial network, and use the preset loss function to iteratively train the initial network until convergence to obtain a converged network. During the iteration process, continuously adjust the network parameters in the feature extractor, the attention mechanism, the feature fuser, and the classification recognizer, so that the overall loss function meets the accuracy requirements; input the validation dataset into the converged network to verify the NLOS recognition ability of the converged network. In the case of passing the verification, obtain the trained spatio-temporal feature model; input the CIR data in the test dataset into the trained spatio-temporal feature model to obtain the recognition result of the NLOS signal category output by the spatio-temporal feature model.
2. The multi-class non-line-of-sight signal recognition method based on spatio-temporal features according to claim 1, wherein The process of inputting the CIR data in the training dataset into the initial network, using the preset loss function to iteratively train the initial network until convergence to obtain a converged network, and continuously adjusting the network parameters in the feature extractor, the attention mechanism, the feature fuser, and the classification recognizer during the iteration process so that the overall loss function meets the accuracy requirements includes the following: Select the channel impulse response data of multi-class LOS / NLOS signals from the training dataset, and use the feature extractor to extract the spatio-temporal features thereof. During the iteration process, continuously adjust the network parameters of the feature extractor to obtain the spatio-temporal features of the channel impulse response data of multi-class LOS / NLOS signals in the training dataset and mark and distinguish them to construct a spatio-temporal feature training sample set. Select the channel impulse response data of multi-class LOS / NLOS signals and the corresponding spatio-temporal features from the training dataset, use the attention mechanism to adjust the weights of the spatio-temporal features of the channel impulse response data to obtain key features, and then fuse the original spatio-temporal features and the key features through the feature fuser, and add the corresponding NLOS signal type label to construct a fused feature training sample set. Set the overall loss function of the spatio-temporal feature model. Input the constructed spatio-temporal feature training sample set and the fused feature training sample set into the classification recognizer respectively. During the iteration process, continuously classify and identify the network parameters of the recognizer respectively, so that the overall loss function meets the accuracy requirements. The classification recognizer determines the NLOS signal type based on the fused features. Finally, obtain the spatio-temporal feature model with the attention mechanism and the spatio-temporal feature model without the attention mechanism respectively.
3. The multi-class non-line-of-sight signal recognition method based on spatio-temporal features according to claim 1, characterized in that The feature extractor is used for deep extraction of the spatial and temporal features of the CIR data in the UWB signal, including the following: The feature extractor consists of a spatial feature extraction part and a temporal feature extraction part. The spatial feature extraction part and the temporal feature extraction part respectively perform deep extraction of the spatial and temporal features of the CIR data of the UWB signal. The one-dimensional convolutional network for extracting the spatial features of CIR data in UWB signals consists of three serial batch normalization convolutional filters, namely the first convolutional filter, the second convolutional filter, and the third convolutional filter; the kernel sizes of the first convolutional filter to the third convolutional filter are K h ×K w , 2K h ×K w , 4K h ×K w ; the definitions of the above convolutional filters are as follows: where \(Z = [z_1, z_2, \ldots, z L \) is the input CIR signal with length \(L = 512\), denotes the output feature map after convolution; the feature map has spatial dimensions \(H\times W\) and channels \(C\); \(V = [v_1, v_2, \ldots, v C \) is the set of learned filter kernels, where each filter \(v c is defined as Normalize the input to reduce the number of training epochs, defined as follows: where γ BN is a learnable parameter, μ and σ 2 are the mean and variance of the batch, and ε is a parameter for maintaining numerical stability; after the batch normalization operation, the output is passed through a rectified linear activation function; a pooling layer is added after the convolutional neural network module for spatial downsampling of the feature map; the rectified linear activation function and the max pooling layer can introduce sparsity and enhance the nonlinearity for the spatio-temporal feature model; The deep spatial features extracted by the one-dimensional convolutional network are flattened through a flattening layer and then sent to the long short-term memory network for temporal feature extraction to generate a spatio-temporal feature sequence.
4. The multi-class non-line-of-sight signal recognition method based on spatio-temporal features according to claim 1, characterized in that, Include the following steps: The attention mechanism adjusts the weights of each spatio-temporal feature of the CIR data. First, it generates a global description of the channel dimension corresponding to the spatio-temporal features of the CIR data. The formula is as follows: where y c represents the global descriptor of the c-th channel, where c = 1, 2, …, C; x c (i, j) represents the value of the input feature X at the position (i, j) on the c-th channel; then, the local dependencies between channels are modeled by a one-dimensional convolution, and the convolution kernel size k is adaptively determined according to the number of channels C: where ψ(C) is a function based on the number of channels; γ and b are hyperparameters used to adjust the range of the kernel size; |·| odd represents adjusting the result to the nearest odd number; next, the channel attention weights are generated using a one-dimensional convolution operation, and the formula is as follows: σ c = f conv1d (y, k) = Conv1D(y, W k ) s c = sigmoid(σ c ) where f conv1d represents a one-dimensional convolution operation, and the convolution kernel weights are y = [y1, y2, …, y C is the global descriptor vector; σ c is the intermediate result after convolution; s c is the attention weight for the c-th channel, ranging between [0, 1]; sigmoid(·) is the activation function; finally, the weights generated above are weighted to the original input feature H: X c ′ = s c × X c where X c ′ represents the output feature after weighting for the c-th channel; X c represents the original input feature of the c-th channel.
5. The multi-category non-line-of-sight signal recognition method based on spatio-temporal features according to claim 1, wherein, Include the following steps: The feature fuser fuses the original depth spatio-temporal feature vector H obtained by processing from the feature extractor lstm with the key feature vector H of the attention mechanism eca to generate a fused feature vector F concat =[H lstm , H eca , and then flattens it using a flattening layer; further uses two fully connected layers, with a rectified linear unit layer and a dropout layer embedded between the first fully connected layer and the second fully connected layer; further, the number of neurons in the first fully connected layer in the feature fuser is much larger than the number of neurons in the second fully connected layer; the second fully connected layer is defined as the output layer and consists of five fully connected neurons; the output of the classification recognizer is based on the output of the feature fuser, the output of the softmax layer is defined as the output class probability, and the NLOS scenario recognition result of the UWB signal to be recognized is predicted, which is achieved through the following formula: wherein and are the is-th and js-th input elements respectively, Softmax(·) is the softmax activation function, and exp(·) is the exponential function; according to the output size of, identify the LOS / NLOS matching results in five scenarios.
6. A multi-class non-line-of-sight signal recognition system based on spatio-temporal features, which adopts the multi-class non-line-of-sight signal recognition method based on spatio-temporal features according to any one of claims 1 to 5, and is characterized in that, Include the following: Data acquisition module: Used to collect the corresponding CIR data respectively in the indoor environment with multiple NLOS scenarios, and construct a classification training data set based on the CIR data collected in the multiple NLOS scenarios. Model construction module: Used to construct a spatio-temporal feature network model. The spatio-temporal feature network model includes a feature extractor, an attention mechanism, a feature fuser, and a classification recognizer. Among them, the feature extractor is used to extract the temporal and spatial features for NLOS multi-class recognition from the CIR data; the attention mechanism is used to further extract key features from the temporal and spatial features extracted by the feature extractor to improve the accuracy of NLOS multi-class recognition; the feature fuser is used to fuse the original spatio-temporal features of the feature extractor and the key features of the attention mechanism; the classification recognizer determines the NLOS scene type of the actual environment CIR data based on the fused features output by the feature fuser. Model training module: Used to perform classification recognition training on the spatio-temporal feature network model using the training data set until the spatio-temporal feature network model is trained until convergence, and the classification recognizer determines the NLOS scene type of the unlabeled UWB signal to a certain accuracy rate, then the training ends. Classification recognition module: Used to obtain the NLOS type of the CIR data to be recognized. Input the CIR data to be recognized into the feature extractor, attention mechanism, feature fuser, and classification recognizer of the trained spatio-temporal feature network model, and output the NLOS scene type of the CIR data to be recognized through the classification recognizer.
Citation Information
Patent Citations
Non-line-of-sight signal identification method based on wavelet Gramer convolutional neural network
CN115496097A
UWB non-line-of-sight multi-classification identification method and system based on domain adversarial learning
CN118503857A
Method and system for identification and mitigation of errors in non-line-of-sight distance estimation
US20110177786A1