Human behavior recognition method and system based on channel state information
By preprocessing and feature extraction of CSI data, combined with a random forest algorithm optimized by sparrow search, the problem of insufficient accuracy in CSI segment recognition in existing technologies has been solved, and high-precision human behavior recognition has been achieved.
Patent Information
- Application Number
- CN202511063174.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-31
Smart Images

Figure CN120561662B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of behavior recognition, and in particular to a method and system for human behavior recognition based on channel state information. Background Technology
[0002] Human behavior recognition has made significant progress in smart homes, health monitoring, and smart security. Driven by machine learning and artificial intelligence, research on human behavior recognition is extending to various fields, especially location services such as indoor navigation and positioning. However, current indoor navigation and positioning services are constrained by the interference of human behavior on wireless signal propagation and the dynamic changes in sensor posture. They often require users to maintain a certain posture or behavior, failing to adapt to changes in behavior, resulting in a less than ideal user experience and limiting the development and widespread application of the indoor navigation and positioning industry. Accurate identification of human behavior allows for targeted strategy optimization, further improving the system's robustness and positioning accuracy.
[0003] Currently, research on human behavior recognition can be categorized into four types based on device type: wearable sensors, computer vision, environmental sensors, and wireless signals. Wearable sensors suffer from poor portability and are prone to movement; computer vision requires bright, open environments and raises privacy concerns; environmental sensors suffer from low accuracy, high cost, and difficulty in deployment. With the development of wireless communication technology, non-contact sensing methods based on wireless signals have become a research hotspot due to their advantages such as ease of deployment, no need for worn sensors, lower environmental requirements, and no privacy concerns.
[0004] Wireless Fidelity (WiFi) signals have become the primary means of indoor network connectivity due to their convenience, ease of use, and low cost. The resulting Channel State Information (CSI) possesses fine-grained spatiotemporal resolution capabilities and has been widely applied in human behavior perception research. Compared to coarse-grained features such as Received Signal Strength (RSS), CSI can more accurately reflect minute perturbations during signal propagation, making it suitable for human action recognition in complex environments. Existing research has explored various types of action recognition based on CSI, such as behavior recognition systems for specific groups (children, the elderly, drivers) or specific scenarios (fall detection, conference rooms, corridors), achieving a certain level of recognition accuracy. However, existing methods often struggle to accurately capture CSI segments highly correlated with action changes during behavior recognition, leading to insufficient extraction of key behavioral information. Furthermore, CSI data often contains a large amount of redundant or invalid information, interfering with feature expression and severely impacting recognition accuracy. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for human behavior recognition based on channel state information, which achieves a recognition rate of 92.4%, which is superior to similar human behavior recognition systems.
[0006] To achieve the above objectives, the present invention provides a human behavior recognition method based on channel state information, comprising the following steps:
[0007] Raw CSI data is collected and preprocessed to obtain CSI data for human behavior recognition;
[0008] Motion capture is performed on CSI data used for human behavior recognition to obtain CSI data segments corresponding to the actions;
[0009] Extract the time-domain features, frequency-domain features, and time-frequency-domain features from the CSI data segments corresponding to the actions, and then normalize them to form a behavior feature sequence.
[0010] The behavioral feature sequence is input into the human behavior recognition model to perform human behavior recognition.
[0011] Preferably, the process of collecting and preprocessing raw CSI data to obtain CSI data for human behavior recognition includes the following steps:
[0012] Raw CSI data is collected through multiple antennas at the receiver of the CSI device;
[0013] The raw CSI data collected is processed using principal component analysis with maximum variance selection to select the optimal subcarrier, retaining the signal components that are more sensitive to motion, thereby reducing the data dimensionality.
[0014] Apply Hampshire filtering to CSI data to identify and remove outliers;
[0015] The wavelet thresholding method is used to decompose CSI data into sub-signals of different frequencies to achieve signal-to-noise separation. High and low frequency noise is removed by setting thresholds and threshold functions.
[0016] The CSI data is sorted in descending order within the moving window, the maximum and minimum values are removed, and the mean is calculated as the filtering result for the current time moment to further reduce noise and obtain CSI data for human behavior recognition.
[0017] Preferably, the CSI data used for human behavior recognition is used to perform motion capture to obtain CSI data segments corresponding to the actions, including the following steps:
[0018] The mean absolute deviation (MAD) of the sliding window is calculated on the amplitude, and then the average value of the window is superimposed for smoothing and denoising.
[0019] Calculate the average value of MAD and use multiple times the average value as the threshold for identifying moving and stationary parts;
[0020] By finding the minimum value before the first intersection point and the minimum value after the last intersection point of the threshold and the MAD curve, the start and end points of the action in a single antenna can be determined.
[0021] The intersection of the action regions identified by multiple antennas is taken as the final captured action region.
[0022] Preferably, the behavioral feature sequence is input into the human behavior recognition model for human behavior recognition, which includes the following iterative process:
[0023] Randomly select the first feature sequence from the behavioral feature sequence;
[0024] An initial analysis of the correlation between elements in the first feature sequence is performed to obtain the correlation results;
[0025] The second feature sequence is obtained by correcting the first feature sequence based on the association results;
[0026] The temporal and frequency characteristics of each element in the second feature sequence are combined with the association results to obtain the state corresponding to each element, and the corresponding behavior is found in the behavioral feature sequence based on the state corresponding to each element.
[0027] The iteration terminates when the set number of iterations is reached.
[0028] Preferably, the behavior of each iteration is compared with the actual behavior:
[0029] When the consistency between the iterative calculation behavior and the actual behavior reaches more than 95%, save the model parameters, correlation results, first feature sequence, and second feature sequence at this time.
[0030] When the consistency between the calculated behavior and the actual behavior in each iteration is less than 95%, the model parameters, association results, first feature sequence, and second feature sequence of this iteration are marked to avoid generating marked combinations in the next iteration.
[0031] Human behavior recognition systems based on channel state information include:
[0032] The data acquisition module collects and preprocesses raw CSI data to obtain CSI data for human behavior recognition;
[0033] The motion capture module performs motion capture on CSI data used for human behavior recognition to obtain CSI data segments corresponding to the actions.
[0034] The feature extraction module is used to extract time-domain features, frequency-domain features, and time-frequency-domain features from the CSI data segments corresponding to the action, and after normalization, form a behavior feature sequence.
[0035] The classification module is used to input behavioral feature sequences into the human behavior recognition model for human behavior recognition.
[0036] Therefore, the present invention employs the above-mentioned human behavior recognition method and system based on channel state information, achieving a recognition rate of 92.4%, which is superior to similar human behavior recognition systems. Attached Figure Description
[0037] Figure 1 This is a flowchart of the human behavior recognition method based on channel state information of the present invention;
[0038] Figure 2 This is a diagram showing the results of data preprocessing in Example 2; Figure 2 (a) is the original data; Figure 2 (b) is the preprocessing result;
[0039] Figure 3 This is a motion capture effect diagram of Example 2; Figure 3 (a) is the first antenna; Figure 3 (b) is the second antenna; Figure 3 (c) is the third antenna;
[0040] Figure 4 To improve the accuracy of human behavior recognition; Figure 4 (a) is the confusion matrix for human behavior recognition of 6 volunteers; Figure 4 (b) The accuracy of human behavior recognition for different volunteers;
[0041] Figure 5 To determine the recognition accuracy for different numbers of antennas;
[0042] Figure 6 Different recognition rates for different numbers of features;
[0043] Figure 7 Training and testing time for models with different numbers of features;
[0044] Figure 8 Recognition rate for different feature types. Detailed Implementation
[0045] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0046] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.
[0047] Channel Indicator (CSI) is information used to describe the channel state in wireless communication systems employing Multiple-Input Multiple-Output (MIMO) and Orthogonal Frequency Division Multiplexing (OFDM) technologies. It includes key characteristics such as channel fading, interference, and multipath effects. During wireless signal propagation, obstructions can cause signal reflection, leading to constantly changing signal paths. In such cases, CSI can be used to sense the continuously changing channel, achieving better communication performance and adaptability.
[0048] In the frequency domain, the received signal at a single antenna receiver at a certain moment is shown in equation (1). Y and X Indicates received and transmitted signals. N This represents Gaussian white noise. H The CSI matrix contains the CSI information of all subcarriers, as shown in equation (1).
[0049] (1);
[0050] (2);
[0051] The CSI value represents the value of a single subcarrier. Indicates subcarrier The center frequency, This represents the total number of subcarriers. The CSI value of a single subcarrier is a complex value, as shown in equation (3). Indicates the first The amplitude of each subcarrier, Indicates the first The phase of each subcarrier.
[0052] (3).
[0053] Example 1
[0054] like Figure 1 As shown, the human behavior recognition method based on channel state information includes the following steps:
[0055] Raw CSI data is collected and preprocessed to obtain CSI data that can be used for human behavior recognition;
[0056] Motion capture is performed on CSI data that can be used for human behavior recognition to obtain CSI data segments corresponding to the actions;
[0057] Extract the time-domain features, frequency-domain features, and time-frequency-domain features from the CSI data segments corresponding to the actions, and then normalize them to form a behavior feature sequence.
[0058] The behavioral feature sequence is input into the human behavior recognition model for human behavior recognition.
[0059] 1. Data Preprocessing
[0060] To improve data reliability and usability, the CSI signal was first subjected to dimensionality reduction. The complex amplitude signal was reduced to only a few principal components through principal component analysis, thus reducing the data dimensionality. Secondly, a series of data preprocessing methods, including Hampel filtering, wavelet thresholding denoising, and median averaging filtering, were used to perform multiple filtering processes on the amplitude signal, which effectively removed outliers, reduced high and low frequency noise, and improved the smoothness of the data.
[0061] (1) Data dimensionality reduction. Principal component analysis is used to select the optimal subcarrier by using the "maximum variance selection method" to retain the signal components that are more sensitive to action, thereby reducing the data dimensionality.
[0062] (2) Hampshire filtering. Perform Hampshire filtering on CSI data to identify and remove outliers.
[0063] (3) Wavelet threshold denoising method. The CSI data is decomposed into sub-signals of different frequencies to achieve signal-to-noise separation. High and low frequency noise is removed by setting thresholds and threshold functions.
[0064] (4) Median average filtering. The CSI data are sorted in descending order within the moving window, the maximum and minimum values are removed, and the average value is used as the filtering result at the current time to further reduce noise.
[0065] 2. Adaptive motion time window capture algorithm
[0066] This algorithm is used to capture the start and end points of actions in CSI amplitude values, thereby extracting amplitude information corresponding to the action state to improve action recognition accuracy. The specific steps are as follows:
[0067] (1) Calculate the mean absolute deviation (MAD) with a sliding window size of 100 on the amplitude, and then use the median average filter with a window size of 40 to smooth and denoise.
[0068] (2) Calculate the average value of MAD and use 1.5 times the average value as the threshold for recognizing moving and stationary parts;
[0069] (3) By finding the minimum value before the first intersection point and the minimum value after the last intersection point of the threshold and the MAD curve, the start and end points of the action in a single antenna are determined;
[0070] (4) Take the intersection of the action regions identified by multiple antennas as the final capture action region.
[0071] 3. Feature Extraction
[0072] The upper limit of machine learning tasks is determined by the data and features, making it crucial to select features with significant differences. Features that can be used for behavior recognition can be divided into: time domain, frequency domain, and time-frequency domain features, as shown in Table 1.
[0073] Table 1. Feature types that can be used for behavior recognition
[0074]
[0075] To eliminate the influence of dimensions, a Max-Min normalization process is performed to map the result value to the range [0,1].
[0076] 4. Random Forest Algorithm Based on Sparrow Search Optimization
[0077] Random forests (RF) build multiple decision trees by randomly selecting features and samples, and then use voting among these decision trees to obtain the final classification or regression result. The accuracy of a random forest model is related to the number of decision trees and the minimum number of leaves. Too few decision trees can easily lead to underfitting, while too many increase computational costs; too small or too large a minimum number of leaves can lead to underfitting or overfitting.
[0078] The Sparrow Search Algorithm (SSA) is a heuristic algorithm based on the swarm intelligence of sparrows, simulating their behavior of searching for food and clustering during foraging. Discoverers perform random searches, followers follow the optimal solution, and watchdogs protect the entire population from falling into local optima. During the search, the individual discovery and clustering phases alternate. In the discovery phase, individuals move randomly to find better solutions, while in the clustering phase, they aggregate to move closer to the optimal solution, achieving a global optimization effect. Compared to traditional algorithms, SSA has advantages such as simplicity, robustness, fewer control parameters, and high search efficiency.
[0079] In each iteration, the discoverer's position is updated as shown in equation (4). Indicates the current iteration number. Indicates the first The middle generation The position of an individual in the j-th dimension A random number between (0,1). Indicates the maximum number of iterations. These are random numbers that follow a standard normal distribution between [0,1]. L The dimension of an element that is all 1s is A row matrix. R This represents a warning value, a random number within the range [0,1]. ST This indicates the safety threshold.
[0080] (4);
[0081] The position update of the follower is shown in equation (4). Indicates the first The location of the least fit individual in the generation. Indicates the first The location of the individual with the best fitness in a generation. A This represents a 1xd row matrix with randomly selected elements of -1 or 1. = PD represents the number of discoverers in the population.
[0082] (5);
[0083] The location update of the vigilant is shown in (6). Indicates the first The location of the individual with the best fitness in a generation. The dimension of the elements that satisfy the standard normal distribution is Control step size matrix, A random number between [-1, 1] Set to a constant. For the current fitness of an individual, This indicates the current optimal fitness of the population. This indicates the current optimal fitness of the population.
[0084] (6).
[0085] Example 2
[0086] 1. Data Collection
[0087] The experimental setup was a conference room equipped with long desks, a whiteboard, air conditioning, a projector, lockers, and office chairs. One 3 dB gain antenna was installed at the transmitter, and three 3 dB gain antennas were installed at the receiver. The Linux CSI Tool was used to collect CSI data, with the packet transmission frequency set to 600 Hz.
[0088] Six volunteers (four men and two women) performed eight actions in a designated area: "still," "sitting," "falling," "walking," "running," "climbing stairs," "making a phone call," and "spinning." A total of 2,880 samples were collected as the CSI perception recognition dataset, which was used for training and testing at a ratio of 7:3.
[0089] 2. Data preprocessing results
[0090] Figure 2 From the raw data and the preprocessed results, Figure 2 (a) and Figure 2 As can be seen in (b), preprocessing greatly reduces outliers and high and low frequency noise in the data, while improving data smoothness.
[0091] 3. Motion capture results
[0092] A specific action region captured by each antenna subcarrier data, such as Figure 3 As shown in the figure, the highlighted part is the action area captured by the proposed algorithm. It can be seen from the figure that the proposed algorithm accurately captures the antenna amplitude change regions caused by the action and can effectively identify the initial amplitude segment of the action.
[0093] 4. Human behavior recognition results
[0094] Figure 4 (a) shows the confusion matrix after merging the data from six volunteers. Figure 4 (b) Demonstrates the accuracy of human behavior recognition across different volunteers. The human behavior recognition system proposed in Example 1 exhibits a high recognition rate in comprehensive behavior recognition, achieving a comprehensive recognition rate of 92.4% even when data from multiple individuals is combined. Analysis results indicate that the main factors affecting recognition accuracy are similar signal occlusion actions and misjudgments between actions caused by different individuals' movement habits. Experimental results show that the system can effectively recognize eight types of human behavior in typical environments with high accuracy.
[0095] Table 2 shows the comparative results of this invention with other recognition systems in recent years under similar experimental equipment, scenarios, and actions. The comparison results show that, in the case of 6 people and 8 actions, the average recognition rate of this invention is still superior to similar recognition systems trained with only a small number of volunteer samples. This indicates that this invention has better performance and advantages in human behavior recognition.
[0096] Table 2 Comparison of different human behavior recognition systems
[0097]
[0098] 5. The impact of different antenna numbers on recognition accuracy
[0099] Experimental observations revealed that during data dimensionality reduction, subcarriers of different antennas exhibit varying sensitivities to the same action. To further investigate the impact of different antenna fusion methods on recognition accuracy, seven sets of comparative experiments were conducted, such as... Figure 5 As shown in the figure, the experimental results show that the single-antenna accuracy of antenna 3 is better than that of antennas 1 and 2. Furthermore, as the number of antennas increases, the corresponding recognition accuracy also shows an upward trend, indicating that antenna fusion can improve recognition accuracy. This demonstrates that in the data dimensionality reduction stage, selecting an appropriate antenna combination can enhance the system's ability to recognize human behavior and improve accuracy.
[0100] 6. The impact of the number of features on recognition accuracy
[0101] To analyze the accuracy of the human behavior recognition system in this embodiment under different feature counts, 21 groups of data with different feature counts (1-15, 20, 25, 30, 35, 40, 43) extracted by the ReliefF algorithm were selected, and their accuracy was compared with that of the RF and SSA-RF classification algorithms. Figure 6 As shown, the recognition accuracy of the SSA-RF algorithm is significantly better than that of the RF algorithm, and the recognition accuracy increases with the increase of the number of features.
[0102] Figure 7 The values represent the training and testing times when different numbers of features are selected. When the number of features increases from 1 to 43, the model training time and model testing time only increase by 0.886 s and 0.088 s, respectively. This shows that the overall performance of the model is good, and the number of features has little impact on the training and testing times.
[0103] Depend on Figure 8 It can be seen that the accuracy of time-frequency domain features is higher than that of time-domain features and frequency-domain features, because time-frequency domain features take into account information from both the time and frequency domains; the recognition rate of features containing 43 mixed features is the highest, reaching over 90%.
[0104] Therefore, the present invention employs the above-mentioned human behavior recognition method and system based on channel state information, achieving a recognition rate of 92.4%, which is superior to similar human behavior recognition systems.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for human behavior recognition based on channel state information, characterized in that, Includes the following steps: Raw CSI data is collected and preprocessed to obtain CSI data for human behavior recognition; Motion capture is performed on CSI data used for human behavior recognition to obtain CSI data segments corresponding to the actions; Extract the time-domain features, frequency-domain features, and time-frequency-domain features from the CSI data segments corresponding to the actions, and then normalize them to form a behavior feature sequence. The behavioral feature sequence is input into the human behavior recognition model to perform human behavior recognition, including the following iterative process: Randomly select the first feature sequence from the behavioral feature sequence; An initial analysis of the correlation between elements in the first feature sequence is performed to obtain the correlation results; The second feature sequence is obtained by correcting the first feature sequence based on the association results; The temporal and frequency characteristics of each element in the second feature sequence are combined with the association results to obtain the state corresponding to each element, and the corresponding behavior is found in the behavioral feature sequence based on the state corresponding to each element. The iteration terminates when the set number of iterations is reached.
2. The human behavior recognition method based on channel state information according to claim 1, characterized in that, The process of acquiring and preprocessing raw CSI data to obtain CSI data for human behavior recognition includes the following steps: Raw CSI data is collected through multiple antennas at the receiver of the CSI device; The raw CSI data collected is processed using principal component analysis with maximum variance selection to select the optimal subcarrier, retaining the signal components that are more sensitive to motion, thereby reducing the data dimensionality. Apply Hampshire filtering to CSI data to identify and remove outliers; The wavelet thresholding method is used to decompose CSI data into sub-signals of different frequencies to achieve signal-to-noise separation. High and low frequency noise is removed by setting thresholds and threshold functions. The CSI data is sorted in descending order within the moving window, the maximum and minimum values are removed, and the mean is calculated as the filtering result for the current time moment to further reduce noise and obtain CSI data for human behavior recognition.
3. The human behavior recognition method based on channel state information according to claim 1, characterized in that, The process of motion capture from CSI data used for human behavior recognition to obtain corresponding CSI data segments includes the following steps: The mean absolute deviation (MAD) of the sliding window is calculated on the amplitude, and then the average value of the window is superimposed for smoothing and denoising. Calculate the average value of MAD and use multiple times the average value as the threshold for identifying moving and stationary parts; By finding the minimum value before the first intersection point and the minimum value after the last intersection point of the threshold and the MAD curve, the start and end points of the action in a single antenna can be determined. The intersection of the action regions identified by multiple antennas is taken as the final captured action region.
4. The human behavior recognition method based on channel state information according to claim 1, characterized in that, Compare the behavior of each iteration with the actual behavior: When the consistency between the iterative calculation behavior and the actual behavior reaches more than 95%, save the model parameters, correlation results, first feature sequence, and second feature sequence at this time. When the consistency between the calculated behavior and the actual behavior in each iteration is less than 95%, the model parameters, association results, first feature sequence, and second feature sequence of this iteration are marked to avoid generating marked combinations in the next iteration.
5. A human behavior recognition system based on channel state information, characterized in that, include: The data acquisition module collects and preprocesses raw CSI data to obtain CSI data for human behavior recognition; The motion capture module performs motion capture on CSI data used for human behavior recognition to obtain CSI data segments corresponding to the actions. The feature extraction module is used to extract time-domain features, frequency-domain features, and time-frequency-domain features from the CSI data segments corresponding to the action, and after normalization, form a behavior feature sequence. The classification module is used to input behavioral feature sequences into the human behavior recognition model for human behavior recognition, including the following iterative process: Randomly select the first feature sequence from the behavioral feature sequence; An initial analysis of the correlation between elements in the first feature sequence is performed to obtain the correlation results; The second feature sequence is obtained by correcting the first feature sequence based on the association results; The temporal and frequency characteristics of each element in the second feature sequence are combined with the association results to obtain the state corresponding to each element, and the corresponding behavior is found in the behavioral feature sequence based on the state corresponding to each element. The iteration terminates when the set number of iterations is reached.
Citation Information
Patent Citations
WiFi-based non-contact action recognition method and system and laboratory device
CN116304915A
Lightweight through-wall human body behavior recognition method and system based on MobileViT improvement
CN117272132A