First-view sign language recognition method based on WiFi channel state information and SE-TCN

By using WiFi channel state information and SE-TCN, a method that employs 1×1 convolution and channel attention mechanisms to reduce model complexity, and combines TCN blocks with dilated causal convolution for sign language recognition, the problem of high model training overhead and difficult deployment in existing technologies is solved, achieving efficient sign language recognition on mobile devices.

CN120974309APending Publication Date: 2025-11-18JINAN UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511080936.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing sign language recognition technologies based on vision or sensors suffer from problems such as lighting changes and occlusion, and are costly. Methods based on millimeter-wave radar and ultra-wideband radar require dedicated hardware, which limits their widespread application. Meanwhile, sign language recognition technologies based on WiFi signals have high model training overhead and are difficult to deploy in deep learning.

Method used

We adopt a first-person sign language recognition method based on WiFi channel state information and SE-TCN. We use 1×1 convolution dimensionality reduction, S-ENet channel attention mechanism and TCN blocks with dilated causal convolution for sign language recognition. We combine batch normalization and dropout layers to optimize feature representation and reduce model complexity.

Benefits of technology

It achieves efficient recognition of sign language gestures on mobile devices, reduces the amount of network parameters and computation, improves recognition accuracy, adapts to complex scenarios, and has privacy protection advantages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974309A_ABST
    Figure CN120974309A_ABST
Patent Text Reader

Abstract

The invention discloses a first-view sign language recognition method based on WiFi channel state information and SE-TCN, and the method comprises the steps: carrying out the channel dimension reduction of data through 1 * 1 convolution, and reducing the subsequent network parameter quantity and calculation quantity; a channel attention mechanism based on S-ENet is adopted to carry out adaptive distribution on weights of different channels of CSI data, feature representation is optimized, better feature input is provided for a subsequent network, and the channel attention mechanism is realized through global pooling, 1 * 1 convolution and an activation function; performing time sequence feature modeling on the CSI data by adopting a TCN block based on expansion causal convolution, and capturing a time sequence dependency relationship of sign language actions; according to the classification task part, global representation of each channel is obtained through global average pooling, the feature size is further reduced, and then the features are mapped to final sign language action probability distribution through 1 * 1 convolution and a Softmax function.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of sign language recognition, and particularly relates to a first-view sign language recognition method based on WiFi channel state information and SE-TCN. BACKGROUND

[0002] Sign Language (SL) is an important way for the deaf community to communicate with the outside world, and its essence is a natural language based on hand gestures to express information. In recent years, Sign Language Recognition (SLR) has received extensive attention, mainly relying on visual or sensor technology for gesture capture and recognition. For example, computer vision-based methods use cameras or Kinect devices to track hand movements and extract gesture features. However, such methods have many limitations in practical applications, such as changes in lighting, occlusion problems, and privacy protection. In addition, sensors such as Leap Motion can accurately capture fine finger movements, but they are highly sensitive to operating distance and displacement, making it difficult to adapt to complex dynamic scenarios. In contrast, sign language recognition based on wireless signals has been extensively studied in recent years, especially methods such as millimeter wave (mmWave) radar and ultra-wideband (UWB) radar, which can achieve high recognition accuracy. However, these technologies often require specialized hardware devices, which are costly and complex to deploy, limiting their practical application. In this context, gesture perception technology based on WiFi signals has gradually become a popular research direction due to its low cost, strong environmental adaptability, lack of light interference, and privacy protection advantages.

[0003] In recent years, deep learning technology has provided a variety of solutions for human gesture recognition based on WiFi channel state information (CSI). CSI data has obvious time series and high-dimensional characteristics, with the time frame dimension reflecting the dynamic evolution of the action and the subcarrier dimension containing rich frequency domain information. Based on this, convolutional neural networks (CNN) that extract CSI spatial features through their local perception capabilities are widely used in this field, but traditional CNNs have limitations in modeling ability and difficulty in capturing long-range dependencies when processing time series information. Models based on the Transformer architecture [2] capture the relationship between each time step in the sequence data through attention mechanisms, with stronger long-range dependency modeling capabilities, and can more effectively process complex time series features when processing CSI data. However, due to the large number of model parameters and high computational complexity, the model training overhead is large and difficult to deploy on mobile devices.

[0004] [1] L. Jia et al.,"BeAware:Convolutional neural network(CNN)based user behavior understanding through WiFi channel state information," Neurocomputing, vol. 397, pp. 457-463, 2020.

[0005] [2] B. Li, W. Cui, W. Wang, L. Zhang, Z. Chen, and M. Wu,"Two-stream convolution augmented transformer for human activity recognition," Proceedings of the AAAI Conf. Artif. Intell., vol. 35, no. 1, pp. 286-293, 2021. SUMMARY

[0006] The present application aims to overcome the shortcomings and deficiencies of the prior art, and provide a first-view sign language recognition method based on WiFi channel state information and SE-TCN.

[0007] The purpose of the present application is achieved by the following technical solutions:

[0008] The first-view sign language recognition method based on WiFi channel state information and SE-TCN comprises the following steps:

[0009] S1, after decoding and preprocessing the CSI data collected from sign language actions, a two-dimensional matrix is obtained, wherein the first dimension refers to the time frame number of the CSI data, and the second dimension refers to the subcarrier number of the CSI data;

[0010] S2, the input layer performs batch preprocessing on the CSI data two-dimensional matrix;

[0011] S3, 1x1 convolution is performed on the preprocessed CSI data two-dimensional matrix to reduce the dimension; then the BatchNormalization batch normalization layer is used to optimize the feature distribution of the data, and the feature X' is output;

[0012] S4, then the feature X' is subjected to channel attention mechanism based on S-ENet to realize adaptive channel weight distribution, and different weights are given to the importance of each channel;

[0013] S5, multiply the generated channel dynamic weight s by each channel of the feature X', to obtain

[0014] S6, multi-channel feature matrix weighted by channel Temporal feature modeling is performed by a 3-layer TCN module; multi-channel feature matrix for the feature set;

[0015] S7, feature information modeled by 3-layer TCN The final classification task is performed by the classification layer: first, global average pooling is performed to compress the time dimension, obtaining the global feature representation of each channel, then 1x1 convolution is used to map C channel features to the final classification action number, and the classification output vector Z, finally each classification output Z of the classification output vector Z is normalized to a probability distribution to obtain the final prediction probability of each class. j

[0016] The step S4 is specifically as follows:

[0017] S4.1, the channel attention mechanism module first compresses the time dimension information by global pooling operation, thereby obtaining the global representation vector Z of each channel, and the formula is as follows:

[0018]

[0019] Wherein, Z c is the global representation of the cth channel of the representation vector Z, T is the time frame number of the CSI data, X' c (t) represents the feature information of the cth channel of the feature X' at time t, and C is the channel dimension of the feature X';

[0020] S4.2, the vector Z c uses two 1x1 convolution nonlinear transformations to generate the dynamic weight of the channel:

[0021] s = σ (W2δ (W1Z) ) ;

[0022] Wherein, is used to reduce the dimension of the representation vector Z, C out is the channel number of the representation vector Z, and r is the dimension reduction ratio; δ(·) is the ReLU activation function used to introduce nonlinear transformation; dimensioned back to the original channel number; σ(·) represents the sigmoid activation function, which is used to normalize the weight of each channel, and the finally output vector s represents the importance of each channel, i.e. weight.

[0023] The step S6 is specifically as follows:

[0024] S6.1, multi-channel feature matrix ​The output at time step t The calculation formula is as follows:

[0025]

[0026] Wherein, K is the size of the convolution kernel; f (l) (i) represents the convolution weight at the index i position in the lth layer convolution kernel; d (l) The convolution expansion factor of the lth layer is represented, which represents the interval between the convolution kernel elements; by using the expansion factor, the feature receptive field of the convolution network can be expanded under the condition of keeping low parameter amount, and long scale time dependent information can be captured;

[0027] S6.2, each layer of TCN needs to fill the left of the input feature, and the formula is:

[0028] P (l) =(K-1)·d (l) ;

[0029] P (l) , which is the number of fillings required by the lth layer.

[0030] In step S6, each layer of the TCN module includes a one-dimensional dilated causal convolution layer, a batch normalization layer, a ReLU nonlinear activation layer and a Dropout layer; the dilated causal convolution layer captures the long-range temporal dependence of the action features, while the batch normalization layer and the ReLU nonlinear activation layer are used to normalize and nonlinearly express the data, so that the data can learn the discriminative features of different sign language actions more robustly; the Dropout layer is regularized to prevent overfitting;

[0031] Three layers of TCN modules are sequentially arranged, and the output of the last layer of the three layers of TCN modules is taken as the input of the classification layer.

[0032] In step S7, the calculation formula of the prediction probability of each category is:

[0033]

[0034] Wherein, Presents the prediction probability of the jth sign language action, CLS represents the total number of predicted sign language action categories, j is the index of the sign language action category currently concerned, and the value range is from 1 to CLS; k is an index variable within the summation symbol in the Softmax formula, which is used to traverse all sign language action categories; And Z j And Z k The results of the exponential operation of the jth and kth sign language action category outputs are used to participate in the probability calculation of the final category output;

[0035] each Softmax output vector is a probability distribution vector of the Softmax output, with a dimension of CLS, vector each element in corresponds to the predicted probability of the lth sign language action category.

[0036] Meanwhile, the present application provides:

[0037] A server, comprising a processor and a memory, at least one program is stored in the memory, the program is loaded and executed by the processor to realize the above-mentioned first-view sign language recognition method based on WiFi channel state information and SE-TCN.

[0038] A computer-readable storage medium, at least one program is stored in the storage medium, the program is loaded and executed by the processor to realize the above-mentioned first-view sign language recognition method based on WiFi channel state information and SE-TCN.

[0039] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0040] 1. The present application uses 1*1 convolution to reduce the channel dimension of data, reduces the subsequent network parameter quantity and calculation amount.

[0041] 2. The present application uses a channel attention mechanism based on S-ENet to adaptively allocate the weights of different channels of CSI data, optimizes the feature representation, and provides better feature input for the subsequent network. The channel attention mechanism is realized through global pooling, 1*1 convolution and activation function.

[0042] 3. The present application uses a TCN block based on dilated causal convolution to model the time sequence features of CSI data, capture the time sequence dependence relationship of sign language actions; each TCN block further includes batch normalization, ReLU activation function to optimize feature representation and improve training stability; and Dropout to reduce model overfitting.

[0043] 4. The present application first uses global average pooling to obtain the global representation of each channel, and further reduces the feature size, then uses 1*1 convolution and Softmax function to map the feature to the final sign language action probability distribution. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 is a flowchart of the first-view sign language recognition method based on WiFi channel state information and SE-TCN.

[0045] Figure 2 is Figure 1 is a SE-TCN model structure diagram.

[0046] Figure 3 is an experimental scene diagram.

[0047] Figure 4 is a sign language action diagram.

[0048] Figure 5 is a ten-class sign language recognition result of the SE-TCN model. DETAILED DESCRIPTION

[0049] The application will be described in further detail below with reference to the embodiments and drawings, but the embodiments of the application are not limited thereto. As Figure 1 The first-view sign language recognition method based on WiFi channel state information and SE-TCN includes the following steps:

[0050] First, the system collects WiFi channel state information (CSI) and decodes and pre-processes it. Subsequently, the data first passes through an encoder based on SE-Net: the data will pass through 1x1 convolution and batch normalization for channel compression and feature standardization. Next, a channel attention module based on S-ENet is introduced, which dynamically weights the importance of each channel through global average pooling and nonlinear transformation, achieving the reinforcement of key channels. The weighted features are input into a three-layer stacked time series convolution network (TCN), which uses dilated causal convolution, ReLU activation, and Dropout mechanism to efficiently and robustly capture the time series dependent features of sign language actions. Finally, the data features pass through a classification layer, which uses global average pooling to extract global features and outputs the prediction results of 10-class sign language actions through 1x1 convolution and Softmax function.

[0051] The specific implementation is as follows:

[0052] The first-view sign language recognition method based on WiFi channel state information and SE-TCN includes the following steps:

[0053] S1, after decoding and preprocessing the collected CSI data of sign language actions, a two-dimensional matrix with a size of [80, 234] is obtained, where the first dimension refers to the time frame number of CSI data, and the second dimension refers to the subcarrier number of CSI data;

[0054] S2, the input layer batch pre-processes the CSI data two-dimensional matrix, adjusts the feature dimension, and makes the output data dimension [1, 234, 80];

[0055] S3, 1x1 convolution is used to reduce the dimension of the pre-processed CSI data two-dimensional matrix, so that the feature dimension is reduced to 64, so as to reduce the parameter quantity of the subsequent network layer; then the BatchNormalization batch normalization layer is used to optimize the feature distribution of the data, and the feature X' is output, and the dimension becomes [1, 64, 80];

[0056] S4, then the feature X' is subjected to channel attention mechanism based on S-ENet to realize adaptive channel weight distribution, and different weights are given to the importance of each channel;

[0057] S4.1, the channel attention mechanism module first compresses the time dimension information through global pooling operation, so as to obtain the global representation vector Z of each channel, and the formula is as follows:

[0058]

[0059] Wherein, Z c is the global representation of the cth channel of the representation vector Z, T is the time frame number of the CSI data (i.e. 80), X' c (t) represents the feature information of the cth channel of the feature X' at time t, and C is the channel dimension of the feature X' (i.e. 64);

[0060] S4.2, the vector Z c is subjected to nonlinear transformation of two 1x1 convolution to generate the dynamic weight of the channel:

[0061] s = σ (W2δ (W1Z) ) ;

[0062] Wherein, is used to reduce the dimension of the representation vector Z, C out is the channel number of the representation vector Z (i.e. 64), and r is the dimension reduction ratio; δ(·) is the ReLU activation function used to introduce nonlinear transformation; The dimension is increased to the original channel number; σ(·) represents the sigmoid activation function, which is used to normalize the weight of each channel, and the finally output vector s (64 dimensions) represents the importance of each channel, i.e. weight;

[0063] S5, the generated channel dynamic weight s is multiplied by each channel of the feature X', so as to realize channel-level adaptive adjustment:

[0064]

[0065] Wherein, is the output of the cth channel of the feature X' after weighting, s c is the weight of the cth channel of the feature X', and X' c is the feature information of the cth channel of the feature X';

[0066] The Encoder section uses this adaptive scaling to enhance the information of key channels while suppressing noise in irrelevant channels, improving the model's ability to represent the features of CSI signals and providing better feature input for the subsequent TCN module, thereby enabling better time-series modeling of CSI data.

[0067] S6. Multi-channel feature matrix after channel weighting Temporal feature modeling is performed using a 3-layer TCN module. Each TCN module includes a one-dimensional dilated causal convolution, a batch normalization layer, a ReLU nonlinear activation layer, and a Dropout layer; multi-channel feature matrix. Features A set;

[0068] S6.1, Multi-channel feature matrix The output after time step t The calculation formula is as follows:

[0069]

[0070] Where K is the size of the convolution kernel, set to 3; f (l) (i) represents the convolution weight at index i in the l-th convolutional kernel; d (l) represents the convolution dilation factor of the l-th layer, and represents the spacing between convolution kernel elements. By using the dilation factor, the convolutional network can expand the feature receptive field while maintaining a low number of parameters, capturing long-scale temporal dependency information. The dilation factors of the three-layer TCN are set to 1, 2, and 4.

[0071] S6.2 To ensure the causality of the model when modeling CSI time series, each TCN layer needs to left-padded the input features, and the formula is as follows:

[0072] P (l) = (K-1)·d (l) ;

[0073] P (l) That is, the number of fill layers required for layer l;

[0074] S7. Feature information after time-series modeling through 3 layers of TCN The final classification task is performed through a classification layer: first, global average pooling is used to compress the time dimension and obtain the global feature representation of each channel. Then, a 1×1 convolution is used to map the C (i.e., 64) channel features to the final classification action number (i.e., 10) and the classification output vector Z (i.e., 10 classification output vectors). Finally, Softmax is used to convert each classification output Z of the classification output vector Z into a single classification output vector. jThe normalized probability distribution is obtained, and the final prediction probability of each category is obtained, and the formula is:

[0075]

[0076] wherein, represents the prediction probability of the jth sign language action, CLS represents the total number of predicted sign language action categories (i.e., 10), j is the index of the sign language action category currently concerned, and the value range is from 1 to CLS; k is an index variable within the summation symbol in the Softmax formula, used to traverse all sign language action categories; and and Z j and Z k are the results of exponential operation, used to participate in the probability calculation of the final category output;

[0077] Each Softmax output vector is a probability distribution vector of the Softmax output, the dimension of the vector is CLS, and each element in the vector corresponds to the prediction probability of the lth sign language action category.

[0078] Figure 2 The sign language recognition model (SE-TCN model) of the application is as shown in the figure, and specifically includes the following main components: an encoder based on SE-Net, a time sequence convolution network layer, and a classification layer.

[0079] The encoder based on SE-Net: the encoder is composed of a 1x1 convolution and a channel attention mechanism: first, the data is subjected to preliminary feature dimension reduction and data standardization through 1x1 convolution and batch normalization, so as to reduce the complexity of the model; then, the channel attention mechanism is used, the weights of each subcarrier channel are dynamically learned and adjusted through global pooling and nonlinear convolution operation, important channel features are strengthened, and redundant irrelevant noise information is suppressed. The weight is multiplied by the channel dimension of the feature data to obtain the input of the next module:

[0080] Time sequence convolution network layer (TCN): each TCN module uses dilated causal convolution to capture the long-range temporal dependence of action features, while batch normalization and ReLU activation function are used to normalize and nonlinearly express the data, so that the data can learn the discriminative features of different sign language actions more robustly; each layer also uses Dropout regularization to prevent overfitting. The system has a total of 3 layers of time sequence convolution network, and the output of the last layer is used as the input of the classification module.

[0081] The classification layer: the input features are further compressed along the channel dimension by global average pooling, mapped to the final number of sign language recognition classes by 1x1 convolution, and the classification probability of sign language action is generated by the Softmax function to realize the probability distribution output of the final sign language action class.

[0082] The first view sign language recognition method based on WiFi channel state information and SE-TCN has an experimental scene diagram as Figure 3 .

[0083] The parameter details of the SE-TCN model are shown in Table 1.

[0084] Table 1

[0085]

[0086] As Figure 4 , 5 The SE-TCN model of the embodiment can effectively recognize (learn, help, read, like, I, good, friend, want, not, and yes) 10 different Chinese sign language actions, and the recognition accuracy of each action is more than 90%.

[0087] Through experimental comparison, the current deep learning methods for recognizing human action gestures based on CSI, such as CNN, LSTM, and CRNN, do not consider the network model parameter quantity and calculation amount, and under the same experimental settings, they cannot achieve a balance between accuracy and model lightweight, and the comprehensive effect is inferior to the model SE-TCN of the present application. The sign language recognition effect and lightweight index of different models are compared in Table 2.

[0088] Table 2

[0089]

[0090] Compared with other traditional neural network models in the Chinese sign language recognition task, SE-TCN can achieve high accuracy, and the parameter quantity and calculation amount of the model are much smaller than CNN, LSTM, and CRNN, which maintains high action recognition accuracy with low overhead, providing an effective solution for mobile terminal device deployment with limited computing power and storage capacity.

[0091] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited to the above embodiments, and any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application are equivalent replacement methods and are included in the protection scope of the present application.

Claims

1. A first-person sign language recognition method based on WiFi channel state information and SE-TCN, characterized in that, Includes the following steps: S1. After decoding and preprocessing the CSI data of the collected sign language actions, a two-dimensional matrix is ​​obtained, where the first dimension refers to the time frame number of the CSI data and the second dimension refers to the number of subcarriers of the CSI data. S2. The input layer performs batch preprocessing on the two-dimensional matrix of CSI data. S3, 1×1 convolution is used to reduce the dimensionality of the preprocessed CSI data two-dimensional matrix; then the feature distribution of the data is optimized by the batch normalization layer, and the feature X' is output. S4. Next, feature X' uses an S-ENet-based channel attention mechanism to achieve adaptive channel weight allocation, assigning different weights to each channel based on its importance. S5. Multiply the generated channel dynamic weights s by each channel of feature X' to obtain... S6. Multi-channel feature matrix after channel weighting Temporal feature modeling is performed using a 3-layer TCN module; Multi-channel feature matrix Features A set; S7. Feature information after time-series modeling through 3 layers of TCN The classification layer then performs the final classification task: First, global average pooling is used to compress the time dimension, obtaining the global feature representation for each channel. Next, a 1×1 convolution is used to map the C channel features to the final classification action and the classification output vector Z. Finally, Softmax is used to convert each classification output Z into a single classification output vector Z. j Normalize to a probability distribution to obtain the final predicted probability for each category.

2. The first-view sign language recognition method based on WiFi channel state information and SE-TCN according to claim 1, characterized in that, Step S4 is as follows: S4.1 The channel attention mechanism module first compresses the time dimension information through a global pooling operation to obtain the global representation vector Z for each channel, as shown in the following formula: Among them, Z c To represent the global representation of the c-th channel of vector Z, where T is the time frame number of the CSI data, X' c (t) represents the feature information of the c-th channel of feature X' at time t, where C is the channel dimension of feature X'; S4.2, Vector Z c The dynamic weights of the channels are generated using a nonlinear transformation with two 1×1 convolutions: s=σ(W2δ(W1Z)); in, Used to reduce the dimensionality of the representation vector Z, C out To characterize the number of channels in vector Z, r is the dimensionality reduction ratio; δ(·) is the ReLU activation function used to introduce nonlinear transformation; The number of channels is increased to the original number; σ(·) represents the sigmoid activation function, which is used to normalize the weights of each channel. The final output vector s represents the importance of each channel, i.e., its weight.

3. The first-view sign language recognition method based on WiFi channel state information and SE-TCN according to claim 1, characterized in that, Step S6 is as follows: S6.1, Multi-channel feature matrix The output after time step t The calculation formula is as follows: Where K is the size of the convolution kernel; f (l) (i) represents the convolution weight at index i in the l-th convolutional kernel; d (l) represents the convolution dilation factor of the l-th layer, and represents the spacing between convolution kernel elements; by using the dilation factor, the convolutional network can expand the feature receptive field while maintaining a low number of parameters, and capture long-scale time-dependent information. S6.2 Each layer of TCN needs to left-padded the input features, and the formula is as follows: P (l) =(K-1)·d (l) ; P (l) This is the number of fill layers required for layer l.

4. The first-view sign language recognition method based on WiFi channel state information and SE-TCN according to claim 1, characterized in that, In step S6, each TCN module includes a one-dimensional dilated causal convolutional layer, a batch normalization layer, a ReLU nonlinear activation layer, and a Dropout layer. Dilated causal convolutional layers capture the long-range temporal dependencies of action features, while batch normalization layers and ReLU nonlinear activation layers are used to normalize and nonlinearly represent the data, enabling the data to learn the discriminative features of different sign language actions more robustly. Dropout layer regularization prevents overfitting; The three TCN modules are set up sequentially, with the output of the last TCN module serving as the input to the classification layer.

5. The first-view sign language recognition method based on WiFi channel state information and SE-TCN according to claim 1, characterized in that, In step S7, the formula for calculating the predicted probability of each category is: in, This represents the predicted probability of the j-th sign language action, CLS represents the total number of predicted sign language action categories, j is the index of the currently interested sign language action category, and its value ranges from 1 to CLS; k is the index variable within the summation symbol in the Softmax formula, used to traverse all sign language action categories; and Output Z for the j-th and k-th sign language gesture categories respectively. j and Z k The result of the exponential operation is used to participate in the probability calculation of the final category output; Each Softmax output vector It is the probability distribution vector output by Softmax, with dimension CLS. Each element in The predicted probability corresponding to the l-th sign language movement category.

6. A server, the server comprising a processor and a memory, the memory storing at least one program, characterized in that, The program is loaded and executed by the processor to implement the first-view sign language recognition method based on WiFi channel state information and SE-TCN as described in any one of claims 1 to 5.

7. A computer-readable storage medium storing at least one program, the program being loaded and executed by a processor to implement the first-view sign language recognition method based on WiFi channel state information and SE-TCN as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Blood pressure estimation method based on attention mechanism and time domain convolutional neural network

    CN114118386A

  • Lightweight wireless gesture recognition method based on attention enhancement

    CN119150142A

  • Ultra-short-term wind power combined prediction method based on TCN-SENet-Transformer

    CN119357634A