A gesture action recognition method based on CSI signals
By collecting CSI data from LTE signals and utilizing the LS channel estimation algorithm and ResNet residual network model, the phase difference and amplitude difference of gesture actions are extracted. This solves the problem of recognition accuracy of existing wireless radio frequency signal gesture recognition methods at different starting positions, angles, and writing sizes, and realizes the universality and real-time performance of non-contact gesture recognition.
Patent Information
- Application Number
- CN202310822616.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-06
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2043-07-06
AI Technical Summary
Existing gesture recognition methods based on radio frequency signals have accuracy issues in recognizing different gesture starting positions, writing angles, and writing sizes, which limits their application in real-world scenarios and requires dedicated equipment or an increased number of devices to achieve recognition.
By collecting CSI data from LTE signals, and using the LS channel estimation algorithm and ResNet residual network model, the phase difference and amplitude difference of the CSI signal of the gesture are extracted to construct a contactless gesture recognition method. This method is applicable to common commercial devices such as LTE and 4G devices and can recognize gestures from different angles.
It achieves accurate gesture recognition at different device placement angles and positions, has a wide range of applications, runs in real time, and does not require device synchronization, making it suitable for natural interactions in daily life.
Smart Images

Figure CN117076998B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of mobile communication technology, in particular to a gesture action recognition method based on CSI signal. BACKGROUND
[0002] The existing gesture recognition technology can be divided into two categories: contact type and non-contact type. Among them, the existing traditional contact type gesture recognition method usually requires the user to hold or wear a special electronic sensing device, and this method can realize accurate gesture action recognition, but it is too dependent on special equipment, which hinders its large-scale deployment in daily life. Unlike the contact method, the non-contact gesture recognition method does not require the user to hold or wear any electronic sensing device, so that the entire interaction process can be completed in a natural and non-intrusive state.
[0003] Most of the existing gesture recognition methods based on wireless radio frequency signals construct the physical or statistical characteristics of wireless radio frequency signals. However, these characteristics are usually referred to as signal transceiver devices. When the starting position, writing direction and writing size of the gesture change relative to the signal transceiver device, the signal fluctuation pattern caused by the same gesture action will also change, thereby greatly affecting the recognition rate and hindering the application of the method in real scenes.
[0004] In summary, most of the existing non-contact gesture recognition methods based on wireless radio frequency signals cannot achieve accurate recognition of different gesture starting positions, writing angles and writing sizes. A few improved methods either need to construct a special machine learning classifier or achieve gesture recognition of different starting positions, writing angles and writing sizes by adding the number of transceiver devices, which limits the practicality and universality of gesture recognition technology. SUMMARY
[0005] In order to overcome the shortcomings of the above-mentioned prior art, the present application provides a gesture action recognition method based on CSI signal. The present application uses the change of channel state information CSI of LTE signal in wireless communication to reflect the instantaneous relative direction of the hand during movement, realizes position gesture recognition, and does not need sensing equipment.
[0006] The technical scheme adopted by the present application is:
[0007] A gesture action recognition method based on CSI signal, the method comprising the following steps:
[0008] Step 1: Collecting cell reference signal CRS data;
[0009] Step 2: Accurately channel pilot using LS channel estimation algorithm;
[0010] Step 3: According to the channel pilot complete channel matrix, using pilot average and interpolation, realize the CSI data acquisition of LTE signal;
[0011] Step 4: Collect the CSI data containing the gesture action of volunteers in the experimental environment;
[0012] Step 5: Use a phase difference correction algorithm to preprocess all the acquired CSI signal data;
[0013] Step 6: Based on the phase difference and amplitude of the CSI signal, extract the phase difference and amplitude difference of the CSI signal data of different gesture actions;
[0014] Step 7: Construct a Resnet residual network model and train it;
[0015] Step 8: Determine the evaluation index of the Resnet residual network model and evaluate it, and take the hit rate G as the evaluation index.
[0016] Further, in step 2, the process of accurately channel pilot using LS channel estimation algorithm is as follows:
[0017] The optimization criterion used by LS channel estimation algorithm is that the sum of square errors between actual received data and estimated received data is minimum, so the result of LS channel estimation can be expressed as:
[0018]
[0019] Where X represents the original transmitted signal vector, H represents the channel response vector, and Y represents the received signal vector.
[0020] For OFDM system, LS channel estimation can be performed on each subcarrier, let N be the number of subcarriers, X[k] and Y[k] be complex numbers, and X[k] have been amplitude normalized, then the LS channel estimation result on each subcarrier can be expressed as:
[0021]
[0022] LS channel estimation algorithm is relatively simple to implement and has low computational complexity. After obtaining the received reference signal, the channel pilot is accurately calculated using LS channel estimation algorithm, and the complete channel matrix is obtained according to the channel pilot.
[0023] Further, the process of step 5 is as follows:
[0024] If the absolute value of the phase difference difference between the current data packet and the next data packet exceeds the threshold e, the data is marked as a jump point, which can be expressed as:
[0025]
[0026] wherein denotes the phase difference of the current data packet, denotes the phase difference of the next data packet;
[0027] The data packets from 0 to t seconds are divided into three segments: 0 to t1 seconds are recorded as the first segment, t1 to t2 seconds are recorded as the second segment, and t2 to t seconds are recorded as the third segment, wherein t1 is the time corresponding to the first jump point, and t2 is the time corresponding to the second jump point; the phase difference corresponding to the time point of each of the three segments is subtracted by the average phase difference of each segment, and finally a continuous phase difference waveform is obtained, The corrected phase difference of the first segment, the second segment, and the third segment is respectively as follows:
[0028]
[0029]
[0030]
[0031] The corrected CSI phase difference of the data packets from 0 to t seconds is as follows:
[0032]
[0033] Further, the process of step 7 is as follows:
[0034] Residual unit y l may be expressed as
[0035] y l = h(x l ) + F(x l , W l ) x l+1
[0036] = f(y l )
[0037] wherein x l and x l+1 respectively represent the input and output of the lth residual unit, F is a residual function, represents the learned residual, W l represents the input of the weight layer, and h(x l ) = x l represents an identity mapping, f is a ReLU activation function, and the learned feature x L from the shallow layer l to the deep layer L is as follows:
[0038]
[0039] A Resnet network is built, and the main frame of the network is composed of four residual blocks, and each block contains a basic residual block.
[0040] The CSI data sample matrix of the corrected LTE signal is input, and the network is finally composed of two convolutional layers and a full connection layer to simultaneously classify the gesture action of two subtasks, and a maximum pooling operation is performed after the first layer, and an average pooling operation is performed after the last convolutional layer of the two subtasks. The detection result is obtained through the Softmax activation function.
[0041] The present application has the beneficial effects that: the present application provides a non-contact gesture recognition method, which utilizes the disturbance of wireless radio frequency signals generated during hand movement, measures and extracts the changes of CSI of LTE signals, and realizes gesture recognition. The technical scheme of the present application is based on prior knowledge to distinguish physical gestures. The present application does not require clock synchronization between the signal sending device and the signal receiving device, and can use most common commercial devices (such as LTE (Long Term Evolution), 4G) for sensing, and has wide application range. The technical scheme of the present application has no special requirements for the placement angle and position of the device, can recognize gestures made at different angles relative to the device, and can also run in real time, which is convenient and practical. BRIEF DESCRIPTION OF DRAWINGS
[0042] Fig. 1 is the gesture action recognition method based on CSI signal of the present application, and the working flow chart of the present application is shown in the figure;
[0043] Fig. 2 is the CSI data acquisition and detection working schematic diagram of the present application;
[0044] Fig. 3 is the hit rate obtained by using the CSI gesture recognition evaluation method of the present application. DETAILED DESCRIPTION
[0045] The present application discloses a gesture action recognition method based on CSI signal, and the present application will be further described in combination with specific cases.
[0046] Reference Figs. 1-3 A gesture action recognition method based on CSI signal, comprising the following steps:
[0047] Step 1: collect cell reference signal (CRS) data;
[0048] Step 2: accurately channel pilot by using LS channel estimation algorithm;
[0049] The optimization criterion used by the LS channel estimation is that the sum of square errors between the actual received data and the estimated received data is minimum, so the result of the LS channel estimation can be expressed as
[0050]
[0051] wherein X represents the original transmitted signal vector, H represents the channel response vector, and Y represents the received signal vector.
[0052] For the OFDM system, the LS channel estimation can be performed on each subcarrier, wherein N is the number of subcarriers, X[k] and Y[k] are complex numbers, and X[k] has been normalized in amplitude, and the LS channel estimation result on each subcarrier can be expressed as
[0053]
[0054] The LS channel estimation algorithm is relatively simple to implement and has low computational complexity. After the received reference signal is obtained, the channel pilot is accurately acquired by using the LS channel estimation algorithm, and the complete channel matrix is acquired according to the channel pilot.
[0055] Step 3: According to the complete channel matrix of the channel pilot, the pilot average and interpolation are used to realize the CSI data acquisition of the LTE signal.
[0056] Step 4: In the experimental environment, the CSI data containing the gesture actions of the volunteers is acquired, and the actions are divided into 0-9 ten handwritten digit corresponding gestures.
[0057] Step 5: A phase difference correction algorithm is used to pre-process all the acquired CSI signal data, and the process is as follows:
[0058] If the absolute value of the difference between the phase difference of the current data packet and the phase difference of the next data packet exceeds the threshold value e, the data is marked as a jump point, which can be expressed as:
[0059]
[0060] In the formula, the phase difference of the current data packet, the phase difference of the next data packet.
[0061] The data packets from 0 to t seconds are divided into three segments for processing: the segment from 0 to t1 seconds is recorded as l1, the segment from t1 to t2 seconds is recorded as l2, and the segment from t2 to t seconds is recorded as l3, wherein t1 is the time corresponding to the first jump point, and t2 is the time corresponding to the second jump point; the phase difference corresponding to the time points of the three segments is subtracted from the average phase difference of each segment, and a continuous phase difference waveform is finally obtained, The corrected phase differences of the l1 segment, the l2 segment, and the l3 segment are respectively as follows:
[0062]
[0063]
[0064]
[0065] The corrected CSI phase difference of the data packet from 0 to t seconds is as follows:
[0066]
[0067] Step 6: Based on the phase difference and amplitude of the CSI signal, the phase difference and amplitude difference of the CSI signal data of different gesture actions are extracted;
[0068] Step 7: Construct a Resnet residual network model and train it, the process is as follows:
[0069] Residual unit y l Can be expressed as
[0070] y l = h(x l ) + F(x l , W l ) x l+1
[0071] = f(y l )
[0072] Where x l and x l+1 respectively represent the input and output of the lth residual unit, F is the residual function, which represents the learned residual, and h(x l ) = x l represents the identity mapping, and f is the ReLU activation function. The learning feature x L from the shallow layer l to the deep layer L is:
[0073]
[0074] A Resnet network is built, which has four residual blocks in the main frame, and each block contains a basic residual block. The network has a total of 11 convolutional layers, including 9 shared convolutional layers and 2 independent convolutional layers.
[0075] The corrected CSI data sample matrix of the LTE signal is input, and the network is composed of two convolutional layers and a fully connected layer at the end. Two sub-tasks are performed simultaneously to classify gesture actions, and there is a max-pooling operation after the first layer, and an average-pooling operation after the last convolutional layer of the two sub-tasks. The detection result is obtained through the Softmax activation function.
[0076] Step 8: Determine the Resnet residual network model evaluation index and evaluate it to hit the rate H as the evaluation index.
[0077] The processing procedure of the present example is as follows:
[0078] 1. As shown in the accompanying drawings, the experimental platform of the present example is composed of two computers, both of which are equipped with USRPB210, one of which is used as LTE eNB and the other as LTE UE. Fig. 2
[0079] 2. The specific implementation site is an empty classroom, and volunteers make gestures in the activity area about 3 meters in front and back, as shown in the accompanying drawings. Fig. 2
[0080] 3. At each collection time, the hand is in the data collection point in the accompanying drawings, and the CSI data of the LTE signal is collected, with 30 seconds for each action. After the collection is completed, a plurality of.txt files can be obtained for each action; Fig. 2
[0081] 4. All the obtained CSI signal data are preprocessed by using a phase difference correction algorithm, and then the phase difference and amplitude difference of the CSI signal data of different gesture actions are extracted based on the phase difference and amplitude of the CSI signal;
[0082] 5. The processed CSI signal is input into the Resnet residual network, and the error is regressed in the gradient minimum direction to train the model;
[0083] 6. After the gesture recognition model is trained, the model is tested, and the hit rate of gesture recognition is calculated (as shown in the accompanying drawings); Fig. 3
[0084] The above has made further detailed introduction to the present application, and the specific implementation manner of the present application is not limited to the above description. In general, for the related content of the present application, any simple calculation or replacement within the protection scope of the claims of the present application is within the protection scope of the present application.
Claims
1. A gesture recognition method based on CSI signals, comprising the following steps: Step 1: Collect cell reference signal (CRS) data; Step 2: Accurately predict channel pilots using the LS channel estimation algorithm; Step 3: Based on the complete channel matrix of the channel pilots, use pilot averaging and interpolation to acquire CSI data of LTE signals; Step 4: Collect CSI data containing volunteer hand gestures in the experimental environment; Step 5: Preprocess all acquired CSI signal data using a phase difference correction algorithm, including: If the absolute value of the phase difference between the current data packet and the next data packet exceeds the threshold e, then the data packet is marked as a transition point, which can be expressed by the following formula: In the formula This indicates the phase difference of the current data packet. This indicates the phase difference of the next data packet; The data packets from 0 to t seconds are divided into three segments: 0 to t1 seconds are denoted as segment l1, t1 to t2 seconds as segment l2, and t2 to t seconds as segment l3, where t1 is the time corresponding to the first transition point and t2 is the time corresponding to the second transition point. The average phase difference of each segment is subtracted from the phase difference at each of the three time points to obtain a continuous phase difference waveform. The phase differences after correction for segments l1, l2, and l3 are shown in the following formula: The corrected CSI phase difference for data packets from 0 to t seconds is shown in the following formula: Step 6: Based on the phase difference and amplitude of the CSI signal, extract the phase difference and amplitude difference of the CSI signal data for different hand gestures; Step 7: Construct and train the ResNet residual network model; Step 8: Determine and evaluate the ResNet residual network model, using the hit rate G as the evaluation metric.
2. The gesture recognition method based on CSI signals according to claim 1, characterized in that... The process of step 2 is as follows: The optimization criterion used in the LS channel estimation algorithm is to minimize the sum of squared errors between the actual received data and the estimated received data. Therefore, the result of LS channel estimation can be expressed as: Where X represents the original transmitted signal vector, H represents the channel response vector, and Y represents the received signal vector; For OFDM systems, LS channel estimation can be performed on each subcarrier. Let N be the number of subcarriers, X[k] and Y[k] be complex numbers, and let X[k] be amplitude normalized. Then, the LS channel estimation result on each subcarrier can be expressed as: After acquiring the received reference signal, the LS channel estimation algorithm is used to accurately determine the channel pilot, and the complete channel matrix is determined based on the channel pilot.
3. The gesture recognition method based on CSI signals according to claim 1, characterized in that... The process of step 7 is as follows: residual unit y l It can be represented as: y l =h(x l )+F(x l ,W l )x l+1 =f(y l ) Where x l and x l+1 These represent the input and output of the l-th residual unit, respectively. F is the residual function, representing the learned residual, and W... l It is the input to the weight layer, while h(x) l )=x l Let f represent the identity mapping, be the ReLU activation function, and represent the learned features x from the shallow layer l to the deep layer L. L for: A ResNet network is constructed, whose main structure consists of four residual blocks, each containing a basic residual block. The network has a total of 11 convolutional layers, including 9 shared convolutional layers and 2 independent convolutional layers.