A method for recognizing human activity through a wall based on CSI
By using an eight-antenna receiving array and an attention-enhanced neural network, the problems of low sampling resolution and insufficient feature extraction in human activity recognition in wall-penetrating scenarios are solved, achieving efficient and accurate human activity recognition through walls.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHWESTERN POLYTECHNICAL UNIV
- Filing Date
- 2026-03-10
- Publication Date
- 2026-06-23
AI Technical Summary
In scenarios involving penetration through walls, existing technologies suffer from low sampling resolution, insufficient feature extraction, difficulty in eliminating environmental influences, and limited data types, resulting in insufficient stability and effectiveness in human activity recognition.
An eight-antenna receiver array is used to separate the human body dynamic components by using the static component suppression coefficients of the reference channel and the monitoring channel. Combined with Butterworth bandpass filtering, PCA and short-time Fourier transform, the signal subspace and time spectrum are obtained, and feature learning is performed using an attention-enhanced neural network.
It improves signal acquisition efficiency and analysis accuracy, can accurately identify human activity, enhances recognition capabilities in wall-penetrating scenarios, adapts to different wall materials and diverse scenarios, and has efficient, accurate and reliable recognition effects.
Smart Images

Figure CN122268429A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless communication technology, specifically relating to a method for identifying human activity through walls based on CSI. Background Technology
[0002] With the widespread deployment of WiFi infrastructure in public workplaces and homes, understanding everyday human activity using WiFi is playing an increasingly important role in smart homes, medical monitoring, and public safety. However, this has led to growing concern about the stability and effectiveness of CSI-based human activity recognition. CSI, or Channel State Information, is a stable and feature-rich signal parameter; compared to received signal strength, CSI-based methods can achieve finer detection accuracy. Due to the complexity of multipath signals in wall-penetrating scenarios and the significant attenuation caused by wall materials, effectively acquiring, processing, and analyzing CSI to improve the efficiency and accuracy of human activity recognition has become a pressing technical challenge.
[0003] In existing technologies, the acquisition and analysis of through-wall signal CSI mainly face the following technical challenges: 1. Low sampling resolution: Existing technologies typically use commercial WiFi devices with only three antennas, which make it difficult to fully sample multipath signals in wall-penetrating scenarios, thus affecting spatial resolution and recognition accuracy.
[0004] 2. Insufficient feature extraction: Focusing only on the temporal or spatial features of CSI often ignores frequency domain features, making it difficult to accurately reflect changes caused by human activities.
[0005] 3. Difficult to eliminate environmental influences: The signal is greatly weakened after being absorbed by the wall, and the signal reflected by the human body is even weaker, and may even be drowned out by static environmental noise.
[0006] 4. Limited data types: Most existing methods use the amplitude or phase of CSI as data input to neural networks for learning, paying less attention to other possible data types.
[0007] Therefore, how to provide a CSI-based method for identifying human activity through walls is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0008] To overcome the shortcomings of existing technologies, this invention provides a through-wall human activity recognition method based on Channel State Information (CSI). The method includes: acquiring 5GHz band signals emitted by a signal source and calculating CSI; separating human dynamic components using a reference channel; segmenting the action segment and the stationary segment, and obtaining their signal subspace through eigenvalue decomposition; selecting two antennas for conjugate multiplication, and obtaining a time-spectrum graph through Butterworth bandpass filtering, PCA, and short-time Fourier transform; and inputting the signal subspace and time-spectrum graph into an attention-enhanced neural network for human activity recognition. This invention has advantages such as high signal acquisition efficiency, high analysis accuracy, and high reliability.
[0009] The technical solution adopted by this invention to solve its technical problem is as follows: Step 1: Transmit a 5GHz signal using a signal source, collect the signal, and calculate the CSI after frequency offset estimation and compensation; Step 2: Use one antenna of the eight-antenna receiving array as the reference channel and the other seven antennas as the monitoring channels. Reduce the static component by calculating the static component suppression coefficient and extract the human body dynamic component. Step 3: Set a time window to calculate the variance. Segment the segments with variance exceeding the set threshold into action segments and the rest into still segments. Based on this, extract the signal subspace by calculating its autocorrelation matrix and eigenvalue decomposition. Step 4: Calculate the ratio of the mean amplitude to the standard deviation of the signal on each antenna, select the two antennas with the smallest and largest ratios and multiply them by conjugate, filter out low-frequency interference and burst noise through Butterworth bandpass filter, and then obtain the time spectrum diagram through PCA and short-time Fourier transform. Step 5: Input the signal subspace and time-spectrum data into the neural network for feature learning, and add an attention module to focus on the spatiotemporal features of the signal subspace and the frequency domain features of the time-spectrum, thereby realizing human activity recognition.
[0010] Preferably, the signal acquisition is performed using a self-developed motherboard integrating Zynq UltraScale 19EG and ADRV9009.
[0011] Preferably, step 2 specifically comprises: Step 2-1: The transmitter uses a directional antenna to transmit signals, and the receiver uses one directional antenna to acquire signals from the reference channel, while the rest are omnidirectional antennas responsible for acquiring signals from the monitoring channel; the channel frequency responses of the two channels are expressed as follows:
[0012]
[0013] in, This represents the channel frequency response of the reference channel. This represents the channel frequency response of the monitoring channel. m Indicates the first m Subcarriers, R Indicates the reference channel. S Indicates the monitoring channel. , and Both are signal attenuation factors. P s It is a set of strong signals, including signals that propagate directly from the transmitter to the receiver and signals that are strongly reflected by the walls; P w This is a set of weak signals, including signals reflected from objects inside the room. , and Both indicate signal propagation delay. f m For the first m The frequency of each subcarrier This indicates packet detection delay. Indicates the sampling time offset. Indicates center frequency offset. , and These represent the phase errors caused by packet detection delay, sampling time offset, and center frequency offset, respectively. and It's noise; Step 2-2: Calculate the static component suppression coefficient antenna-by-antenna and subcarrier-by-subcarrier using the CSI extracted from the reference channel and monitoring channel.
[0014] in, For reference channel number k CSI of each subcarrier, For monitoring channel number n The antenna number k CSI of each subcarrier, To ensure numerical stability and prevent the denominator from being zero; The linear effect of the static component of the reference channel on the monitoring signal of the nth antenna is characterized, and the static interference is eliminated to the greatest extent by optimizing it using the least squares criterion. The phase error of the reference channel and the monitoring channel is the same, and the phase error can be eliminated in the process of calculating the static component suppression coefficient. Steps 2-3: Extracting the dynamic components of the human body using the static component suppression coefficient:
[0015] Obtain CSI caused solely by human reflexes; Steps 2-4: Multiply the phase conjugate of the reference channel CSI by the human activity component of the CSI:
[0016] in, This indicates a phase-taking operation, which ultimately yields human activity components without phase error.
[0017] Preferably, step 3 specifically comprises: Step 3-1: Use the Hampel filter to remove outliers:
[0018] in, , , The coefficient 1.4826 is used to convert MAD into a Gaussian equivalent standard deviation estimate. This represents the filtered signal value. Represents the original signal value. Indicates signal length. Indicates the index of the signal. This indicates the median operation. Indicates the index of the signal within the filter window; Step 3-2: Use the sym3 wavelet function to decompose the signal to the fifth level, and use a heuristic thresholding method for denoising:
[0019] in, This represents the denoised signal. This indicates the inverse discrete wavelet transform using the sym3 wavelet basis. These represent the fifth-level approximation coefficients, indicating the low-frequency component of the signal after five levels of wavelet decomposition. The j-th level detail coefficients represent the high-frequency detail components of the signal at different scales. , The threshold is represented and determined using the heuristic method HeurSure:
[0020] in, This represents the noise standard deviation estimate, based on the first-level detail coefficients. d Calculate the median of 1. This represents the Stein unbiased risk estimation threshold, used for adaptive threshold selection. Step 3-3: Set the window length, calculate the sliding variance, and distinguish between active and static segments based on the degree of signal fluctuation within the window:
[0021] in, This represents the variance of the signal within the window. This represents the set of the final selected action starting points. This represents the mean of the signal within the window. , Let represent the starting points of the actions at both ends, and Ω be the set of candidate action starting points. Describes a subset of Ω. x t Take the median CSI amplitude for each subcarrier of each antenna. d For window length, g The minimum interval between two action segments. K For the target number of action segments, set each selected starting point... Corresponding to the interval [ i , i + d -1] gives the action segment, and the remaining time points are considered static segments. The current maximum value is repeatedly taken. v ( i The starting point and its ± g Non-maximum suppression is applied to the neighborhood until K windows are selected or no more windows are available. Steps 3-4: Measure the degree of interference from human activity on the subcarrier amplitude based on information entropy, and select subcarriers that are sensitive to activity; First, quantify the CSI amplitude, and define the amplitude range [ A min , A max Divided into n There are several equally spaced intervals, each interval having a width of [missing value]. ;No. p The boundaries of each interval are defined as [ A min + ( p 1) A , A min + p A ],in p =1, 2, ..., n ,set up M The total number of CSI amplitude samples. M p To fall into the first p The number of samples in each interval; in the discrete case, the Shannon entropy is:
[0022] Steps 3-5: For single-frame input Perform CSI smoothing, and the smoothed matrix Use the eigenvalue decomposition of covariance to obtain the first... k 3D signal subspace:
[0023] in, l and q These represent the number of input antennas and the number of subcarriers, respectively. L and Q This represents the dimension of the virtual array after CSI smoothing. This represents the matrix after CSI smoothing. express Y The covariance matrix, This represents the matrix composed of the eigenvectors of R. Indicates the corresponding maximum k The eigenvector with eigenvalues, , U s It is a set of orthogonal bases for the signal subspace, which project data onto the signal subspace and stack them according to their real / imaginary parts:
[0024] in, Indicates taking the real part, This indicates taking the imaginary part.
[0025] Preferably, step 4 specifically comprises: Step 4-1: Calculate the conjugate multiplication of the CSI of a pair of antennas, and then... i The CSI of the antenna is represented as Among them, antenna 1's CSI is characterized by small amplitude and large variance, while antenna 2's CSI is characterized by large amplitude and small variance.
[0026] in, The sum of responses from all static paths. A collection representing dynamic paths. Indicates the first l Complex attenuation of a path, Indicates the first l Doppler frequencies of the paths; superscripts (1) and (2) represent antenna 1 and antenna 2, respectively; Step 4-2: A Butterworth bandpass filter is used to filter out static components, low-frequency interference, and burst noise. The lower cutoff frequency of the filter is determined by a trade-off between eliminating interference and losing low-frequency components due to the Doppler effect. The upper cutoff frequency is calculated from the speed of human activity in the experiment.
[0027] in, , , , , Indicates the lower cutoff frequency. This indicates the upper cutoff frequency, and this filter is applied to the multiplication sequence of each subcarrier. Step 4-3: Perform principal component analysis on all CSI subcarriers and select the first principal component that contains the power change caused by the target motion:
[0028] in, This is a matrix of subcarriers stacked over time. For row average, v 1 is the sample covariance The largest eigenvector, This is the time series of the first principal component; Step 4-4: Perform a short-time Fourier transform on the first principal component to obtain the time-spectrum diagram of the Doppler frequency shift:
[0029] in Use a Gaussian window; Finally, the non-overlapping time spectrograms of all CSI segments are stitched together to generate the complete time spectrogram.
[0030] Preferably, step 5 specifically comprises: Step 5-1: Use a 2D CNN to extract the spatiotemporal features of CSI, and introduce a spatiotemporal attention module to focus on the features to be extracted. The spatiotemporal attention module performs average pooling and max pooling along the channel dimension to generate compressed features. J avg and J max These two features are concatenated, and the spatiotemporal attention weight AST is generated through 3D convolution:
[0031] Finally, the attention weights are calculated. The refined features are obtained by multiplying them by the Hadamard product of the original features. , This represents the feature map output by a 2D CNN. Step 5-2: Compress information in the frequency domain using two-dimensional discrete cosine transform (DCT). The basis functions of 2D DCT are:
[0032] 2D DCT is written as:
[0033] in, , , It is a 2D DCT spectrum. It is input. H and E They are X Height and width; Step 5-3: Divide the input X into multiple parts along the frequency dimension. ,in For each part, assign the corresponding 2D DCT frequency components:
[0034] in , It corresponds to 2D index of frequency components It is compressed. dimensional vector; Indicates the first i Each segment in time t ,aisle m The values of all frequency points on the surface, Indicates the position of the 2D DCT basis functions ( t , m The value on ) corresponds to the frequency component ( u i , v i ); Step 5-4: The entire compression vector is obtained through concatenation:
[0035] Among them, the obtained vector ; The entire frequency domain attention framework is written as follows:
[0036] in, Indicates a fully connected layer. This represents the final output of the frequency domain attention; 1D CNN is used to extract frequency domain features from the time-spectrum graph, and a frequency domain attention framework is introduced to focus on the extracted features; Step 5-5: The classification module consists of a Flatten layer, a Concatenate layer, and a fully connected layer. The Flatten layer expands the spatiotemporal features of the signal subspace and the frequency domain features of the spectrogram into a one-dimensional vector. The Concatenate layer connects the two features to obtain a combined feature. The combined feature is input into the fully connected layer, and then the classification result is obtained through the softmax activation function. The process of classifying activities using the softmax activation function is shown in the following equation:
[0037] in, This is the output value of the fully connected layer, where G is the number of activity categories. This is the output of the classification module.
[0038] Preferably, the amplitude range [ A min , A max Divided into n =10 equally spaced intervals.
[0039] An electronic device includes: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to enable the electronic device to perform the above-described CSI-based through-wall human activity recognition method.
[0040] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described CSI-based method for identifying human activity through walls.
[0041] A chip includes a processor for retrieving and running a computer program from a memory, causing a device equipped with the chip to perform the aforementioned CSI-based through-wall human activity recognition method.
[0042] A computer program product includes a computer storage medium storing a computer program, the computer program including instructions executable by at least one processor, which, when executed by the at least one processor, implement the above-described CSI-based through-wall human activity recognition method.
[0043] The beneficial effects of this invention are as follows: (1) By establishing a reference channel and introducing a static component suppression coefficient, this invention achieves the separation of static environmental components and dynamic human body components, thus avoiding the masking of weak dynamic signals by strong static signals. The reference channel only reflects the characteristics of static environmental components, while the monitoring channel captures the superposition of dynamic and static signals, thereby effectively eliminating the influence of static components and highlighting the characteristics of dynamic components.
[0044] (2) By employing an eight-antenna receiving array, this invention significantly improves spatial resolution and signal acquisition capability, especially enabling the acquisition of larger-sized CSI signals, effectively enhancing the ability to distinguish multipath signals. In wall-penetrating scenarios, the multi-antenna structure helps to separate the target signal from the environmental noise components in multipath propagation, further improving the ability to capture dynamic features.
[0045] (3) This invention proposes a multi-dimensional information fusion method that combines the CSI signal subspace and time-spectrum graph caused by human activity, so as to extract the change features caused by human activity more fully. The received CSI data is processed to obtain its signal subspace and time-spectrum graph respectively. Then, spatiotemporal features are extracted from the signal subspace and frequency domain features are extracted from the time-spectrum graph, so that the system can perceive human activity more accurately. Attached Figure Description
[0046] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0047] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0048] This invention fully utilizes technologies in signal processing, wireless communication, data analysis, and deep learning. It details a method for efficiently acquiring and analyzing CSI (Communication Signal Indicator) data, a multi-faceted data feature extraction approach, and a neural network design that incorporates parallel learning of input features and an attention mechanism. Through these technical means, this invention can accurately and efficiently identify human activity in wall-penetrating scenarios. Furthermore, the method and technology of this invention possess high versatility and adaptability, easily adapting to walls of different materials or applicable to diverse scenarios such as monitoring and home environments, demonstrating broad application prospects. Therefore, this invention offers advantages such as high signal acquisition efficiency, high analysis accuracy, and strong reliability.
[0049] refer to Figure 1 A method for identifying through-wall human activity based on CSI, characterized by the following steps: S1. Transmit a 5GHz signal using a signal source, acquire the signal using a self-developed motherboard integrating Zynq UltraScale 19EG and ADRV9009, and calculate the CSI after frequency offset estimation and compensation. S2. Using one antenna of the eight-antenna receiving array as the reference channel and the other seven antennas as the monitoring channel, the static component is reduced to the greatest extent by calculating the static component suppression coefficient, and the human body dynamic component is extracted. S21. The transmitter uses a directional antenna to transmit signals. The receiver uses one directional antenna specifically for acquiring signals from the reference channel, while the rest are omnidirectional antennas responsible for acquiring signals from the surveillance channel. The channel frequency responses of the two channels can be expressed as:
[0050]
[0051] S22. Using the CSI extracted from the reference channel and the monitoring channel, calculate the static component suppression coefficient antenna by antenna and subcarrier:
[0052] S23. Extracting human dynamic components using the static component suppression coefficient:
[0053] By calculating this formula, we can obtain the CSI caused solely by human reflexes; S24. The human activity component obtained at this time still contains phase error. Multiply the phase conjugate of the reference channel CSI by the human activity component of the CSI:
[0054] in, This indicates a phase-taking operation, which ultimately yields human activity components without phase error.
[0055] S3. Set a time window to calculate the variance. Divide the segments with variance exceeding a certain threshold into action segments and the rest into still segments. Based on this, extract the signal subspace by calculating its autocorrelation matrix and eigenvalue decomposition. S31. Use the Hampel filter to remove outliers:
[0056] S32. Use the sym3 wavelet function to decompose the signal to the fifth level, and use a heuristic thresholding method for denoising:
[0057] in, , The threshold is represented and determined using a heuristic method (HeurSure):
[0058] S33. Set the window length, calculate the sliding variance, and distinguish between the action segment and the stationary segment by the degree of signal fluctuation within the window:
[0059] in, This represents the variance of the signal within the window. This represents the set of the final selected action starting points. This represents the mean of the signal within the window. , Let represent the starting points of the actions at both ends, and Ω be the set of candidate action starting points. Describes a subset of Ω. x t Take the median CSI amplitude for each subcarrier of each antenna. d For window length, g The minimum interval between two action segments. K For the target number of action segments, set each selected starting point... Corresponding to the interval [ i , i + d -1] gives the action segment, and the remaining time points are considered static segments. The current maximum value is repeatedly taken. v ( i The starting point and its ± g Non-maximum suppression is applied to the neighborhood until the neighborhood is full. K Until there are no selectable windows; S34. Based on information entropy, measure the interference level of human activity on the subcarrier amplitude and select subcarriers sensitive to activity. First, quantize the CSI amplitude and define the amplitude range […]. A min , A max Divided into n There are several equally spaced intervals, each interval having a width of [missing value]. . No. p The boundaries of each interval are defined as [ A min + ( p 1) A, A min + p A], where p =1, 2, ...,n, let... M The total number of CSI amplitude samples. M p To fall into the first p The number of samples in each interval. In the discrete case, the Shannon entropy is:
[0060] This method allows for the quantitative assessment of the sensitivity of different subcarriers to activity. Higher information entropy indicates greater interference from activity on that subcarrier, and thus greater sensitivity to activity. S35, For single-frame input Perform CSI smoothing, and the smoothed matrix Use the eigenvalue decomposition of covariance to obtain the first... k 3D signal subspace:
[0061] in, l and q These represent the number of input antennas and the number of subcarriers, respectively. L and Q This represents the dimension of the virtual array after CSI smoothing. This represents the matrix after CSI smoothing. express Y The covariance matrix, express R The matrix composed of eigenvectors, Indicates the corresponding maximum k The eigenvector with eigenvalues, , Represents all the eigenvalues obtained by performing eigenvalue decomposition on R. r There are eigenvectors, where Us is a set of orthogonal bases for the signal subspace. The data is projected onto the signal subspace and stacked according to their real / imaginary parts:
[0062] in, Indicates taking the real part, This indicates taking the imaginary part.
[0063] S4. Calculate the ratio of the mean amplitude to the standard deviation of the signal on each antenna. Select the two antennas with the smallest ratio (small amplitude and large variance) and the largest ratio (large amplitude and small variance) and perform conjugate multiplication. Use Butterworth bandpass filtering to filter out low-frequency interference and burst noise. Then, use PCA and short-time Fourier transform to obtain the time spectrum. S41. Calculate the conjugate multiplication of the CSI of a pair of antennas, and then... i The CSI of the antenna is represented as Among them, antenna 1's CSI is characterized by small amplitude and large variance, while antenna 2's CSI is characterized by large amplitude and small variance.
[0064] in, The sum of responses from all static paths. A collection representing dynamic paths. Indicates the first l Complex attenuation of a path, Indicates the first lThe Doppler frequencies of the paths, through the closely placed antennas 1 and 2 of the receiver, can be considered to indicate that the main multipath of the CSI signal of the two antennas is the same. ); S42. A Butterworth bandpass filter is used to filter out static components, low-frequency interference, and burst noise. The lower cutoff frequency of the filter is determined by a trade-off between fully eliminating interference and losing low-frequency components due to the Doppler effect. The upper cutoff frequency is calculated from the speed of human activity in the experiment.
[0065] in, , , , , Indicates the lower cutoff frequency. This indicates the upper cutoff frequency, and this filter is applied to the multiplication sequence of each subcarrier. S43. Perform principal component analysis on all CSI subcarriers and select the first principal component that contains the main and consistent power changes caused by the target motion:
[0066] S44. Perform a short-time Fourier transform on the first principal component to obtain the time-spectrum diagram of the Doppler frequency shift:
[0067] in A Gaussian window is used. Finally, the non-overlapping spectrograms of all CSI segments are stitched together to generate the complete spectrogram.
[0068] S5. Input the two types of data, signal subspace and time-spectrum graph, into the neural network for feature learning, and add an attention module to focus on the spatiotemporal features of the signal subspace and the frequency domain features of the time-spectrum graph, thereby realizing human activity recognition. S51. Use a 2D CNN to extract the spatiotemporal features of CSI, and introduce a spatiotemporal attention module to focus on the features to be extracted. The spatiotemporal attention module performs average pooling and max pooling along the channel dimension to generate compressed features. J avg and J max These two features are cascaded, with spatiotemporal attention weights. A ST Generated through 3D convolution:
[0069] Finally, the attention weights are calculated. The refined features are obtained by multiplying them by the Hadamard product of the original features. ,F This represents the feature map output by a 2D CNN. S52. Two-dimensional discrete cosine transform is used to compress information in the frequency domain. The basis functions of 2D DCT are:
[0070] Therefore, 2D DCT can be written as:
[0071] S53, Input X Divided into multiple parts along the frequency dimension ,in For each part, assign the corresponding 2D DCT frequency components:
[0072] S54. The entire compressed vector can be obtained through concatenation:
[0073] Among them, the obtained vector The entire frequency domain attention framework can be written as:
[0074] 1D CNN is used to extract frequency domain features from the time-spectrum graph, and the above-mentioned frequency domain attention module is introduced to focus on the extracted features; S55. The classification module consists of a Flatten layer, a Concatenate layer, and a fully connected layer. The Flatten layer expands the spatiotemporal features of the signal subspace and the frequency domain features of the spectrogram into a one-dimensional vector. The Concatenate layer concatenates the two features to obtain a combined feature. The combined feature is input into the fully connected layer, and then the classification result is obtained through the softmax activation function. The process of classifying activities using the softmax activation function is shown in the following equation:
[0075] in, It is the output value of the fully connected layer. G It is the number of activity categories. This is the output of the classification module.
[0076] Example: To verify the effectiveness and practical value of this invention, it was applied to a nighttime monitoring scenario in a nursing home. Elderly people frequently need to move between their bedrooms and bathrooms at night, and traditional video surveillance is limited by privacy concerns and cannot be widely deployed. This invention, by deploying WiFi transmitting and receiving devices on both sides of the wall between the bedroom and the corridor, utilizes CSI signals to achieve contactless detection and identification of the elderly's nighttime activities, ensuring both security and privacy.
[0077] In this embodiment, the invention was deployed in a room on the third floor of a nursing home in a certain city in April 2024. The system uses one transmitter and one eight-antenna receiver array. The collected data is separated into human activity components through a reference channel. The invention successfully distinguishes key activities such as getting up, falling, walking, and sitting down, achieving efficient signal acquisition, feature extraction, and analysis, significantly improving the efficiency and accuracy of human activity recognition. The following is a comparison of specific data before and after the application of the invention to demonstrate its beneficial effects: Table 1. Report on the Improvement of Nighttime Monitoring in Nursing Homes
[0078] As shown in the table above, this invention significantly improves the accuracy of activity recognition in this practical application scenario, particularly in the detection of falls, where the detection rate increases from 65% to 98%, effectively ensuring the safety of the elderly at night. Simultaneously, the false alarm rate is greatly reduced, minimizing unnecessary intervention by caregivers. Furthermore, since this method does not require a camera to collect video data, it can achieve high-precision monitoring while ensuring privacy.
[0079] This embodiment demonstrates the practicality and superiority of the invention in real-world applications, particularly in healthcare monitoring scenarios with high privacy requirements. Through efficient CSI signal acquisition, fusion of signal subspace and temporal spectrogram features, and attention-enhanced neural network classification, the invention not only improves the accuracy and robustness of wall penetration detection but also expands the system's application prospects in smart homes and public safety.
Claims
1. A method for identifying through-wall human activity based on CSI, characterized in that, Includes the following steps: Step 1: Transmit a 5GHz signal using a signal source, collect the signal, and calculate the CSI after frequency offset estimation and compensation; Step 2: Use one antenna of the eight-antenna receiving array as the reference channel and the other seven antennas as the monitoring channels. Reduce the static component by calculating the static component suppression coefficient and extract the human body dynamic component. Step 3: Set a time window to calculate the variance. Segment the segments with variance exceeding the set threshold into action segments and the rest into still segments. Based on this, extract the signal subspace by calculating its autocorrelation matrix and eigenvalue decomposition. Step 4: Calculate the ratio of the mean amplitude to the standard deviation of the signal on each antenna, select the two antennas with the smallest and largest ratios and multiply them by conjugate, filter out low-frequency interference and burst noise through Butterworth bandpass filter, and then obtain the time spectrum diagram through PCA and short-time Fourier transform. Step 5: Input the signal subspace and time-spectrum data into the neural network for feature learning, and add an attention module to focus on the spatiotemporal features of the signal subspace and the frequency domain features of the time-spectrum, thereby realizing human activity recognition.
2. The method for through-wall human activity recognition based on CSI according to claim 1, characterized in that, The signal acquisition was performed using a self-developed motherboard integrating Zynq UltraScale 19EG and ADRV9009.
3. The method for through-wall human activity recognition based on CSI according to claim 1, characterized in that, Step 2 specifically involves: Step 2-1: The transmitter uses a directional antenna to transmit signals, and the receiver uses one directional antenna to acquire signals from the reference channel, while the rest are omnidirectional antennas responsible for acquiring signals from the monitoring channel; the channel frequency responses of the two channels are expressed as follows: in, This represents the channel frequency response of the reference channel. This represents the channel frequency response of the monitoring channel. m Indicates the first m Subcarriers, R Indicates the reference channel. S Indicates the monitoring channel. , and Both are signal attenuation factors. P s It is a set of strong signals, including signals that propagate directly from the transmitter to the receiver and signals that are strongly reflected by the walls; P w This is a set of weak signals, including signals reflected from objects inside the room. , and Both indicate signal propagation delay. f m For the first m The frequency of each subcarrier This indicates packet detection delay. Indicates the sampling time offset. Indicates center frequency offset. , and These represent the phase errors caused by packet detection delay, sampling time offset, and center frequency offset, respectively. and It's noise; Step 2-2: Calculate the static component suppression coefficient antenna-by-antenna and subcarrier-by-subcarrier using the CSI extracted from the reference channel and monitoring channel. in, For reference channel number k CSI of each subcarrier, For monitoring channel number n The antenna number k CSI of each subcarrier, To ensure numerical stability and prevent the denominator from being zero; The linear effect of the static component of the reference channel on the monitoring signal of the nth antenna is characterized, and the static interference is eliminated to the greatest extent by optimizing it using the least squares criterion. The phase error of the reference channel and the monitoring channel is the same, and the phase error can be eliminated in the process of calculating the static component suppression coefficient. Steps 2-3: Extracting the dynamic components of the human body using the static component suppression coefficient: Obtain CSI caused solely by human reflexes; Steps 2-4: Multiply the phase conjugate of the reference channel CSI by the human activity component of the CSI: in, This indicates a phase-taking operation, which ultimately yields human activity components without phase error.
4. The method for through-wall human activity recognition based on CSI according to claim 3, characterized in that, Step 3 specifically involves: Step 3-1: Use the Hampel filter to remove outliers: in, , , The coefficient 1.4826 is used to convert MAD into a Gaussian equivalent standard deviation estimate. This represents the filtered signal value. Represents the original signal value. Indicates signal length. Indicates the index of the signal. This indicates the median operation. Indicates the index of the signal within the filter window; Step 3-2: Use the sym3 wavelet function to decompose the signal to the fifth level, and use a heuristic thresholding method for denoising: in, This represents the denoised signal. This indicates the inverse discrete wavelet transform using the sym3 wavelet basis. These represent the fifth-level approximation coefficients, indicating the low-frequency component of the signal after five levels of wavelet decomposition. The j-th level detail coefficients represent the high-frequency detail components of the signal at different scales. , The threshold is represented and determined using the heuristic method HeurSure: in, This represents the noise standard deviation estimate, based on the first-level detail coefficients. d Calculate the median of 1. This represents the Stein unbiased risk estimation threshold, used for adaptive threshold selection. Step 3-3: Set the window length, calculate the sliding variance, and distinguish between active and static segments based on the degree of signal fluctuation within the window: in, This represents the variance of the signal within the window. This represents the set of the final selected action starting points. This represents the mean of the signal within the window. , Let represent the starting points of the actions at both ends, and Ω be the set of candidate action starting points. Describes a subset of Ω. x t Take the median CSI amplitude for each subcarrier of each antenna. d For window length, g The minimum interval between two action segments. K For the target number of action segments, set each selected starting point... Corresponding to the interval [ i , i + d -1] gives the action segment, and the remaining time points are considered static segments. The current maximum value is repeatedly taken. v ( i The starting point and its ± g Non-maximum suppression is applied to the neighborhood until K windows are selected or no more windows are available. Steps 3-4: Measure the degree of interference from human activity on the subcarrier amplitude based on information entropy, and select subcarriers that are sensitive to activity; First, quantify the CSI amplitude, and define the amplitude range [ A min , A max Divided into n There are several equally spaced intervals, each interval having a width of [missing value]. ;No. p The boundaries of each interval are defined as [ A min + ( p 1) A , A min + p A ],in p =1, 2,..., n ,set up M The total number of CSI amplitude samples. M p To fall into the first p The number of samples in each interval; in the discrete case, the Shannon entropy is: Steps 3-5: For single-frame input Perform CSI smoothing, and the smoothed matrix Use the eigenvalue decomposition of covariance to obtain the first... k 3D signal subspace: in, l and q These represent the number of input antennas and the number of subcarriers, respectively. L and Q This represents the dimension of the virtual array after CSI smoothing. This represents the matrix after CSI smoothing. express Y The covariance matrix, This represents the matrix composed of the eigenvectors of R. Indicates the corresponding maximum k The eigenvector with eigenvalues, , U s It is a set of orthogonal bases for the signal subspace, which project data onto the signal subspace and stack them according to their real / imaginary parts: in, Indicates taking the real part, This indicates taking the imaginary part.
5. A method for identifying through-wall human activity based on CSI according to claim 4, characterized in that, Step 4 specifically involves: Step 4-1: Calculate the conjugate multiplication of the CSI of a pair of antennas, and then... i The CSI of the antenna is represented as Among them, antenna 1's CSI is characterized by small amplitude and large variance, while antenna 2's CSI is characterized by large amplitude and small variance. in, The sum of responses from all static paths. A collection representing dynamic paths. Indicates the first l Complex attenuation of a path, Indicates the first l Doppler frequencies of the paths; superscripts (1) and (2) represent antenna 1 and antenna 2, respectively; Step 4-2: A Butterworth bandpass filter is used to filter out static components, low-frequency interference, and burst noise. The lower cutoff frequency of the filter is determined by a trade-off between eliminating interference and losing low-frequency components due to the Doppler effect. The upper cutoff frequency is calculated from the speed of human activity in the experiment. in, , , , , Indicates the lower cutoff frequency. This indicates the upper cutoff frequency, and this filter is applied to the multiplication sequence of each subcarrier. Step 4-3: Perform principal component analysis on all CSI subcarriers and select the first principal component that contains the power change caused by the target motion: in, This is a matrix of subcarriers stacked over time. For row average, v 1 is the sample covariance The largest eigenvector, This is the time series of the first principal component; Step 4-4: Perform a short-time Fourier transform on the first principal component to obtain the time-spectrum diagram of the Doppler frequency shift: in Use a Gaussian window; Finally, the non-overlapping time spectrograms of all CSI segments are stitched together to generate the complete time spectrogram.
6. A method for through-wall human activity recognition based on CSI according to claim 5, characterized in that, Step 5 specifically involves: Step 5-1: Use a 2D CNN to extract the spatiotemporal features of CSI, and introduce a spatiotemporal attention module to focus on the features to be extracted. The spatiotemporal attention module performs average pooling and max pooling along the channel dimension to generate compressed features. J avg and J max These two features are concatenated, and the spatiotemporal attention weight AST is generated through 3D convolution: Finally, the attention weights are calculated. The refined features are obtained by multiplying them by the Hadamard product of the original features. , This represents the feature map output by a 2D CNN. Step 5-2: Compress information in the frequency domain using two-dimensional discrete cosine transform (DCT). The basis functions of 2D DCT are: 2D DCT is written as: in, , , It is a 2D DCT spectrum. It is input. H and E They are X Height and width; Step 5-3: Divide the input X into multiple parts along the frequency dimension. ,in For each part, assign the corresponding 2D DCT frequency components: in , It corresponds to 2D index of frequency components It is compressed. dimensional vector; Indicates the first i Each segment in time t ,aisle m The values of all frequency points on the surface, Indicates the position of the 2DDCT basis function ( t , m The value on ) corresponds to the frequency component ( u i , v i ); Step 5-4: The entire compression vector is obtained through concatenation: Among them, the obtained vector ; The entire frequency domain attention framework is written as follows: in, Indicates a fully connected layer. This represents the final output of the frequency domain attention; 1D CNN is used to extract frequency domain features from the time-spectrum graph, and a frequency domain attention framework is introduced to focus on the extracted features; Step 5-5: The classification module consists of a Flatten layer, a Concatenate layer, and a fully connected layer. The Flatten layer expands the spatiotemporal features of the signal subspace and the frequency domain features of the spectrogram into a one-dimensional vector. The Concatenate layer connects the two features to obtain a combined feature. The combined feature is input into the fully connected layer, and then the classification result is obtained through the softmax activation function. The process of classifying activities using the softmax activation function is shown in the following equation: in, This is the output value of the fully connected layer, where G is the number of activity categories. This is the output of the classification module.
7. The method for identifying through-wall human activity based on CSI according to claim 6, characterized in that, The amplitude range [ A min , A max Divided into n =10 equally spaced intervals.
8. An electronic device, characterized in that, include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.
10. A chip, characterized in that, include: A processor for retrieving and running a computer program from memory, causing a device on which the chip is mounted to perform the method as described in any one of claims 1 to 7.