Environment-independent human body action recognition method based on WiFi time reversal technology
Through time inversion technology and environment-independent feature extraction, combined with deep neural network, the stability problem of WiFi signal motion recognition in different environments is solved, and high-precision human body motion recognition across scenes is achieved.
Patent Information
- Application Number
- CN202510285674.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-11
AI Technical Summary
WiFi signals are highly dependent on environmental factors in human body movement recognition, resulting in limited cross-scene recognition capabilities and unstable recognition accuracy.
Time inversion technology is used to generate idealized inversion signals, and action recognition is performed by extracting amplitude difference, phase difference, spectrum difference and wavelet transformation difference as environmentally independent features.
It realizes stable and efficient action recognition in different environments, improves the accuracy and robustness of recognition, and is suitable for complex indoor scenarios.
Smart Images

Figure CN120296399A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for human action recognition based on WiFi, specifically, an environment-independent human action recognition method based on the time reversal technology of WiFi. Background Art
[0002] With the rapid development of technologies such as augmented reality (AR), smart home, and health monitoring, human action recognition technology is becoming increasingly important. By analyzing the movement patterns of the human body, human action recognition technology can achieve accurate understanding and prediction of user behavior, and is widely used in multiple fields such as virtual reality, gesture control, sports health monitoring, and smart home control. For example, in a smart home, action recognition can automatically adjust the home environment according to the user's activity pattern; in the field of health monitoring, by recognizing the user's movement state, remote health assessment can be carried out and personalized exercise programs can be provided; in augmented reality, real-time recognition of human actions provides a natural and immersive experience for interaction. The common point of these applications is that they all rely on efficient and accurate action recognition technology.
[0003] Currently, human action recognition mainly relies on three technical means: cameras, wearable devices, and wireless signals. Among them, camera technology analyzes actions by capturing images or videos of the human body, which has high accuracy and rich visual information, but also faces problems such as privacy leakage, environmental light changes, and occlusion; wearable devices collect the movement data of the human body through sensors (such as accelerometers, gyroscopes, etc.). Such devices can provide relatively accurate real-time data during movement, but their main disadvantage is that they need to be worn by users, are not very convenient to use, and are easily affected by the wearing position and sensor accuracy.
[0004] In recent years, WiFi signals have been widely regarded as one of the potential technologies for human action recognition, mainly using the wireless propagation characteristics of WiFi to infer human activities by analyzing changes such as reflection and diffraction of WiFi signals. Compared with traditional action recognition technologies based on cameras and wearable devices, WiFi action recognition technology has significant advantages. First, WiFi signals are ubiquitous, and almost all modern homes and offices are covered by WiFi networks, eliminating the need for additional hardware or wearable devices and greatly simplifying the user's usage threshold. Second, WiFi action recognition can monitor in a contactless manner, avoiding privacy concerns of camera technology and eliminating the inconvenience brought by wearable devices. Third, WiFi signals have strong capabilities in penetrating obstacles such as walls and furniture, enabling more stable recognition in complex environments.
[0005] However, although WiFi signals have shown great potential in human action recognition, their applications still face some challenges. Essentially, WiFi signals are affected by environmental factors, especially multipath effects and signal reflections, which may cause significant differences in WiFi signals in different environments. These environmental dependencies limit the generalization ability of WiFi action recognition in different scenarios. Specifically, the propagation of WiFi signals is affected by factors such as indoor structure, reflections and refractions of walls, furniture, and other objects. In different environments, even for the same action, due to changes in signal reflection paths and intensities, the perceived results of WiFi signals may vary significantly. Therefore, a WiFi action recognition system may perform well in some scenarios but be difficult to effectively recognize in others. This cross-scenario application problem has become a major obstacle in the popularization and practical application of WiFi action recognition technology. In addition, the multipath effect of WiFi signals makes them highly sensitive to the environment. Especially in complex indoor environments, the intensity and phase of signals are prone to change, thus affecting the accuracy of action recognition. In some cases, the reflections and refractions of WiFi signals may produce false action detection results, leading to instability and misrecognition of the recognition system. Therefore, how to overcome the strong environmental dependence of WiFi signals, especially to maintain the stability and accuracy of action recognition in different scenarios, has become an urgent problem to be solved in the practical application of WiFi action recognition technology. Summary of the Invention
[0006] Therefore, this patent aims to solve the difficulties of WiFi action recognition in cross-scenario applications by combining time reversal technology with environment-independent feature extraction. Specifically, by time reversal, the received WiFi signal is propagated backward to generate an idealized inverse signal, thereby eliminating the influence of static environmental factors (such as fixed objects like walls and furniture) on the signal. Subsequently, by comparing the difference between the idealized inverse signal and the actually back-propagated inverse signal, environment-independent dynamic features are extracted. Theoretically, if the human body is not moving, the idealized inverse signal and the actual inverse signal are exactly the same. Therefore, these features mainly reflect the changes in human actions rather than the signal fluctuations caused by the environmental structure, effectively overcoming the multipath effect and environmental dependence and achieving highly robust cross-scenario WiFi action recognition. This invention is realized through the following technical solutions:
[0007] This invention discloses an environment-independent human action recognition method based on WiFi time reversal technology. The method includes a training stage and a recognition stage; the training stage includes the following steps:
[0008] 1) Define multiple human actions to be recognized;
[0009] 2) Collect a batch of WiFi signal data for each human body movement and extract environment-independent features as the training set;
[0010] 3) Use the training set to train the action recognition model;
[0011] The recognition stage includes the following steps:
[0012] 4) Real-time collect the WiFi signals recording human activities;
[0013] 5) Extract environment-independent features from the collected WiFi signals;
[0014] 6) Input the features into the trained recognition model to obtain the action category.
[0015] As a further improvement, in step 2) of the present invention, the signal collection method is completed by the cooperation of the WiFi signal transceiver through multiple rounds of time reversal. In each round, the transmitter first sends a signal x. After being reflected by the human body, the receiver receives the signal s(t) and measures the channel state information H(t); then, the inverse channel state information H rev (t) is obtained by performing complex conjugation on H(t), and then it is multiplied by the original signal s(t) to obtain the theoretical inverse signal Finally, the inverse signal is sent back from the receiver to the transmitter to obtain the real inverse signal Complete the signal collection based on time reversal for one round; the theoretical inverse signals estimated in multiple rounds form a signal sample Similarly, the inverse signals collected in multiple rounds form a signal sample
[0016] As a further improvement, in step 2) of the present invention, the environment-independent features are extracted by calculating the difference between the theoretical inverse signal and the real inverse signal including amplitude difference, phase difference, spectrum difference, and wavelet transform difference.
[0017] As a further improvement, in step 3) of the present invention, the action recognition model is a four-branch deep neural network, which extracts deep dynamic features from the amplitude difference, phase difference, spectrum difference, and wavelet transform difference and maps them to the confidence of each action category. The action category with the highest confidence is the recognition result.
[0018] As a further improvement, in step 4) of the present invention, the real-time collection is realized through multiple rounds of time reversal, and the number of inversion rounds is the same as that in step 2).
[0019] As a further improvement, in step 5) of the present invention, the environment-independent feature extraction is consistent with the feature extraction in step 2), and the features include amplitude difference, phase difference, spectrum difference and wavelet transform difference.
[0020] As a further improvement, in step 6) of the present invention, the recognition model is a four-branch deep neural network trained in step 2), which takes the four features as input and outputs the confidence of each action defined in step 1).
[0021] The beneficial effects of the present invention are as follows:
[0022] The present invention proposes a motion recognition method that does not require carrying equipment, which can capture user dynamics imperceptibly through WiFi signals, and can assist in realizing applications such as smart home and augmented reality.
[0023] The present invention utilizes WiFi signals to realize motion recognition, and can realize wall recognition by virtue of the penetration of WiFi signals.
[0024] The present invention realizes action recognition through time reversal technology and extracting dynamic features independent of the environment, and can achieve cross-environment perception, that is, it only needs to collect data in one environment to train the recognition model, and after the training is completed, it can be deployed in any complex scene.
[0025] The present invention extracts four features, namely, amplitude difference, phase difference, spectrum difference and wavelet transform difference, to perform action recognition, covering amplitude and phase, time domain and frequency domain information. Comprehensive feature extraction makes the recognition accuracy higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 Flow chart of the environment-independent human action recognition method based on WiFi time reversal technology;
[0027] Figure 2 A round of time reversal flow chart;
[0028] Figure 3 Architecture diagram of the four-branch action recognition model. DETAILED DESCRIPTION
[0029] The present invention discloses an environment-independent human motion recognition method based on WiFi time reversal technology. Figure 1 As shown, the method includes a training phase and identification authentication; the training phase includes the following steps:
[0030] 1) The user defines several action categories that need to be recognized according to their needs, such as falling, walking, and sitting;
[0031] 2) The user collects a batch of signal samples for each action as a training set;
[0032] When collecting signal samples, the transmitter and receiver of the WiFi signal cooperate to collect signals through time reversal. Specifically, the collection of a signal sample is achieved through multiple rounds of time reversal. As Figure 2 shown, in each round, the transmitter first sends a signal x. After being reflected by the human body, the receiver receives the signal s(t), and the channel state information is estimated as H(t). Then, the inverse channel state information H rev (t) is obtained by performing the complex conjugate on H(t). If H(t) = a(t) + jb(t), then H rev (t) can be expressed as:
[0033] H rev (t) = a(t) - jb(t)
[0034] Through H rev (t), the theoretical inverse signal s rev (t) can be obtained, that is, the signal obtained when the signal s(t) is sent back from the receiver to the transmitter theoretically:
[0035]
[0036] Then, s(t) is actually sent back to the transmitter to obtain the real inverse signal Assume that the time reversal frequency is f Hz and each action lasts for T seconds. Then the signal sample composed of the theoretically obtained inverse signal and the signal sample composed of the real inverse signal both have a length of f * T in the time dimension.
[0037] Each round of time reversal can obtain one and one Then, by calculating the difference between these two signal samples, the environment-independent dynamic features can be obtained. Specifically, a total of four dynamic features can be obtained, including amplitude difference, phase difference, spectrum difference, and wavelet transform difference. The amplitude difference can be calculated by the following formula:
[0038]
[0039] where and represent the amplitudes of the real inverse signal and the theoretical inverse signal respectively. The phase difference can be calculated by the following formula:
[0040]
[0041] where and They represent the phases of the true inversion signal and the theoretical inversion signal respectively. The spectral difference can be calculated by the following formula:
[0042]
[0043] where STFT() represents calculating the short-time Fourier transform. The wavelet transform difference can be calculated by the following formula:
[0044]
[0045] where WT() represents calculating the wavelet transform.
[0046] These four features are concatenated together to obtain a feature vector f = [ΔA, ΔP, ΔI, ΔW]. Each feature vector is labeled for supervised training of the recognition model. For example, the feature vectors corresponding to falling, walking, and sitting can be labeled 0, 1, and 2 respectively. The labeled data is used as the training set for the subsequent action recognition model.
[0047] 3) Train the action recognition model in a supervised manner;
[0048] As Figure 3 shown, the action recognition model is a four-branch deep neural network. During training, each feature vector is input into the model by branch to obtain a confidence vector y for all action categories, which is compared with the true confidence vector Y to calculate the loss. The true confidence vector is the one-hot encoding of the label of the feature vector. For example, "0" is encoded as [1, 0, 0], "1" is encoded as [0, 1, 0], and "2" is encoded as [0, 0, 1]. The loss can be calculated by the following formula:
[0049] L = |y - Y| 2
[0050] Finally, the parameters in the model are updated based on the above loss by backpropagation to complete the training.
[0051] User authentication includes the following steps:
[0052] 4) Use the transceiver of WiFi to collect WiFi signals in real time to record the user's activity information;
[0053] When collecting data, the time reversal strategy described in step 2) is also used, that is, f*T rounds of time reversal are performed to obtain a theoretically inverted signal sample and an actually inverted signal sample
[0054] 5) Calculate the dynamic feature vector for the newly collected signal sample;
[0055] Extract the dynamic feature vector F of the unknown class from and using the feature extraction method described in step 2), where new F = [ΔA new , ΔP new , ΔI new , ΔW new .
[0056] 6) Input the newly received feature vector into the trained action recognition model to complete action recognition.
[0057] Input the newly calculated feature vector F new into the action recognition model trained in step 3) to obtain a confidence vector [p0, p1, p2]. The action class corresponding to the maximum confidence is the final recognition result. For example, if p0 is the largest, it is determined that the user has fallen.
Claims
1. An environment-independent human action recognition method based on WiFi time reversal technology, characterized in that The method described includes a training stage and a recognition stage; the training stage includes the following steps: 1) Define multiple human actions to be recognized; 2) Collect a batch of WiFi signal data for each human action and extract environment-independent features, which are used as the training set; 3) Use the training set to train the action recognition model; The recognition stage includes the following steps: 4) Real-time collect WiFi signals recording human activities; 5) Extract environment-independent features from the collected WiFi signals; 6) Input the features into the trained recognition model to obtain the action category.
2. The environment-independent human motion recognition method based on WiFi time reversal technology according to claim 1, characterized in that In the step 2), the signal collection method is completed by the cooperation of the WiFi signal transceiver through multiple rounds of time reversal. In each round, the transmitter first sends a signal x. After being reflected by the human body, the receiver receives the signal s(t) and measures the channel state information H(t). Then, the inverse channel state information H rev (t) is obtained by performing complex conjugation on H(t), and then it is multiplied by the original signal s(t) to obtain the theoretical inverse signal Finally, the inverse signal is sent back from the receiver to the transmitter to obtain the real inverse signal One round of signal acquisition based on time reversal is completed; the theoretical inverse signals estimated in multiple rounds form a signal sample Similarly, the inverse signals collected in multiple rounds form a signal sample 3. The environment-independent human motion recognition method based on the WiFi time reversal technology according to claim 2, characterized in that, In the said step 2), the environment-independent features are extracted by calculating the difference between the theoretical inversion signal and the true inversion signal , including amplitude difference, phase difference, spectral difference and wavelet transform difference.
4. The environment-independent human motion recognition method based on WiFi time reversal technology according to claim 1 or 2 or 3, characterized in that, In step 3), the action recognition model is a four-branch deep neural network, which extracts deep dynamic features from amplitude difference, phase difference, spectrum difference, and wavelet transform difference and maps them to the confidence of each action category. The action category with the highest confidence is the recognition result.
5. The environment-independent human action recognition method based on WiFi time reversal technology according to claim 4, characterized in that, In step 4), the real-time collection is achieved through multiple rounds of time reversal, and the number of reversal rounds is the same as the number of rounds described in step 2).
6. The environment-independent human motion recognition method based on WiFi time reversal technology according to claim 5, wherein In step 5), the extraction of environment-independent features is the same as the feature extraction described in step 2). The features include amplitude difference, phase difference, spectrum difference, and wavelet transform difference.
7. The environment-independent human motion recognition method based on WiFi time reversal technology according to claim 1 or 2 or 3 or 5 or 6, characterized in that In step 6), the recognition model is the four-branch deep neural network trained in step 2). Taking the four features as inputs, it outputs the confidence of each action defined in step 1).