A gesture recognition method based on wireless sensing
By using BVP sample data and an integrated learning strategy based on wireless WiFi signals, combined with a base learner trained using the early stopping method, the problems of high sensor cost, privacy leakage, and strong environmental dependence in gesture recognition technology are solved, achieving efficient and secure gesture recognition.
Patent Information
- Application Number
- CN202510160233.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-02-13
AI Technical Summary
Existing gesture recognition technology has problems such as high sensor cost, poor user comfort, privacy leakage and strong environmental dependence.
BVP sample data based on wireless WiFi signals is used for posture recognition. An ensemble learning strategy is used to fuse base learners trained by early stopping method, including MLP, LSTM, BiLSTM, ResNet18, ResNet50 and ViT models. The posture recognition results are determined by soft voting method.
It reduces recognition costs, avoids privacy issues, improves security and recognition accuracy, and saves computing resources.
Smart Images

Figure CN119723674B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of wireless sensing technology and computer vision technology, and more particularly to a gesture recognition method based on wireless sensing. Background Art
[0002] Early human posture recognition was mainly based on wearable sensors. By selecting suitable sensors, such as accelerometers and gyroscopes, and placing them on the human body, human activity data was collected. Currently, mainstream human posture recognition mainly relies on devices such as cameras for image capture and analysis. Computer vision technology can be used to complete tasks such as face recognition and traffic violation identification.
[0003] Although the above recognition technologies can effectively capture posture, they have disadvantages such as relatively high sensor cost, poor user comfort, privacy leakage, and strong environmental dependence. Summary of the Invention
[0004] In order to overcome the defects of the above-mentioned prior art, such as high cost and poor safety, the present invention provides a gesture recognition method based on wireless sensing.
[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0006] In a first aspect, a gesture recognition method based on wireless sensing includes:
[0007] Receive BVP sample data;
[0008] A first posture recognition result is determined based on the BVP sample data using a posture recognition model; wherein the posture recognition model fuses the second posture recognition results output by at least two base learners based on the BVP sample data based on an integrated learning strategy to determine the first posture recognition result, and the base learners are trained based on the early stopping method.
[0009] In a second aspect, a computer program product comprises a computer program or computer executable instructions, wherein when the computer program or computer executable instructions are executed by a processor, the method of the first aspect is implemented.
[0010] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0011] This application discloses a posture recognition method based on wireless sensing, which uses BVP sample data based on wireless WiFi signals for posture recognition, reducing recognition costs, avoiding privacy issues of traditional methods, improving security, and having strong adaptability. In addition, through an integrated learning strategy that fuses multiple base learners trained based on the early stopping method, the accuracy of posture recognition results is improved while saving computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 Schematic diagram of the flow of the gesture recognition method based on wireless sensing in Example 1 of the present application;
[0013] Figure 2 Schematic diagram of the featureConv operation process in Example 1 of the present application;
[0014] Figure 3 This is a flow chart comparing the reshape operation and the featureConv operation in Example 1 of the present application;
[0015] Figure 4 This is a schematic diagram of the accuracy of each model on the training set with a training-to-test ratio of 0.8:0.2 in Example 2 of this application;
[0016] Figure 5 This is a schematic diagram of the accuracy of each model on the validation set with a training-to-test ratio of 0.8:0.2 in Example 2 of this application;
[0017] Figure 6 Schematic diagram of the loss values of each model on the training set at a training-to-test ratio of 0.8:0.2 in Example 2 of this application;
[0018] Figure 7 Schematic diagram of the loss values of each model on the validation set with a training-to-test ratio of 0.8:0.2 in Example 2 of this application;
[0019] Figure 8 This is a schematic diagram of the accuracy of each model on the training set with a training-to-test ratio of 0.9:0.1 in Example 2 of this application;
[0020] Figure 9 This is a schematic diagram of the accuracy of each model on the validation set with a training-to-test ratio of 0.9:0.1 in Example 2 of this application;
[0021] Figure 10 Schematic diagram of the loss values of each model on the training set at a training-to-test ratio of 0.9:0.1 in Example 2 of this application;
[0022] Figure 11 This is a schematic diagram of the loss values of each model on the validation set with a training-to-test ratio of 0.9:0.1 in Example 2 of this application. DETAILED DESCRIPTION
[0023] The terms "first", "second" etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable in appropriate circumstances, and this is merely a way of distinguishing the objects of the same attribute when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment. The term "determine" widely covers various actions, may include obtaining, calculating, computing, processing, deriving, investigating, searching (for example, searching in a table, a database or other data structure), ascertaining and similar actions, may also include receiving (for example, receiving information), accessing (for example, accessing data in a memory) and similar actions, may also include generating, creating, establishing and similar actions, and parsing, selecting, selecting and similar actions etc. The relevant definitions of other terms will be provided in the following description.
[0024] It should be noted that when an element is considered to be "connected" to another element, it can be directly connected to the other element or connected to the other element through an intervening element. In addition, the "connection" in the following embodiments should be understood as "electrical connection", "communication connection", etc., if there is transmission of electrical signals or data between the connected objects.
[0025] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting this patent;
[0026] In order to better illustrate this embodiment, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product size;
[0027] It is understandable to those skilled in the art that some well-known structures and descriptions thereof may be omitted in the drawings.
[0028] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.
[0029] Example 1
[0030] As mentioned in the background technology section, although existing gesture recognition technologies can effectively capture gestures, they have disadvantages such as relatively high sensor costs, poor user comfort, privacy leakage, and strong environmental dependence.
[0031] In view of this, this embodiment provides a posture recognition method based on wireless sensing, which can detect various postures and gestures of the human body without image capture, thereby solving the above-mentioned deficiencies.
[0032] First, it should be noted that the BVP (Body-coordinate Velocity Profile) sample data involved in this embodiment is extracted from CSI (Channel State Information). CSI is information used to describe the status of the communication channel. In wireless communication systems, the channel refers to the medium through which signals travel from the transmitter to the receiver, and its state is affected by various factors, such as multipath fading, interference, and attenuation. Channel state information includes a description of channel characteristics, such as fading, delay characteristics, and phase variation.
[0033] Existing research has found that neural network models can be used to extract various information from CSI data. However, these models suffer from limitations such as low information extraction rates and poor recognition accuracy. Background interference can cause the same gesture to have different CSI features in different environments, posing a greater challenge to model recognition. Furthermore, due to the complexity of gestures, a single gesture can have hundreds of dynamic contours, making it difficult to achieve accurate predictions using standard CSI data alone. However, in the human coordinate system, each gesture has a unique velocity distribution. BVP data based on the human coordinate system enables context-independent and cross-domain gesture recognition.
[0034] It should also be noted that BVP sample data is similar to a picture sequence, and each BVP sample is a A tensor of is the number of snapshots of BVP, is the dimension of the feature map.
[0035] Referring to FIG1 , this embodiment provides a gesture recognition method based on wireless sensing, including:
[0036] Receive BVP sample data;
[0037] A first posture recognition result is determined based on the BVP sample data using a posture recognition model; wherein the posture recognition model fuses the second posture recognition results output by at least two base learners based on the BVP sample data based on an integrated learning strategy to determine the first posture recognition result, and the base learners are trained based on the early stopping method.
[0038] The method described in this embodiment implements gesture recognition through BVP sample data based on wireless signals, avoiding the privacy issues of traditional methods, improving security, and having strong adaptability. In addition, the accuracy of gesture recognition results is improved by fusing multiple base learners through an ensemble learning strategy. At the same time, the early stopping method is used to train the base learners, saving computing resources.
[0039] As a non-limiting example, the at least two base learners can be neural network models with different model structures, or neural network models with the same model structure but different model parameters, or neural network models with the same model structure but different training methods and / or training data.
[0040] In some preferred embodiments, the base learner includes at least one of an MLP model, an LSTM model, a BiLSTM model, a ResNet18 model, a ResNet50 model, and a ViT model.
[0041] In some optional embodiments, when the base learner is a ResNet18 model, before inputting the BVP sample data into the ResNet18 model, a reshape operation is performed on the BVP sample data, including:
[0042] The BVP sample data is sequentially passed through the first deconvolution layer and the second deconvolution layer; wherein,
[0043] The first deconvolution layer includes a first number of convolution kernels with a first step length and a first size, a padding size of a first padding amount, and uses a ReLU activation function as an activation function; and
[0044] The second deconvolution layer includes a second number of convolution kernels with a second step size and a second size, a padding size of a second padding amount, and adopts a ReLU activation function as an activation function.
[0045] It should be noted that ResNet18 was originally a neural network used to solve image classification. The input of the original ResNet18 network is a The tensor is a three-channel sample; therefore, in order to use ResNet18 for the posture recognition problem of this embodiment, it is necessary to perform a reshape operation on the sample before the sample enters the ResNet18 network.
[0046] In some specific implementations, the 22×20×20 BVP sample dataset provided by Widar3.0 is used. This dataset contains 22 types of human gestures, namely: Push&Pull, Sweep, Clap, Slide, Draw-N(H), Draw-O(H), Draw-Rectangle(H), Draw-Triangle(H), Draw-Zigzag(H), Draw-Zigzag(V), Draw-N(V), Draw-O(V), Draw-1, Draw-2, Draw-3, Draw-4, Draw-5, Draw-6, Draw-7, Draw-8, Draw-9, Draw-0. The processing of the BVP sample data based on ResNet18 is as follows:
[0047] S11, reshape operation: the number of input channels is 22. First, a convolution kernel is used with a size of 3 and a convolution kernel of , a deconvolution layer with a step size of 2 and a padding of 3, and then a ReLU activation function; and then a convolution kernel with a kernel size of 3. , a deconvolution layer with a step size of 2 and a padding of 3, and finally a ReLU activation function. The number of output channels is 3. After this operation, the data is The tensor form becomes The tensor form can be used to directly input into the ResNet18 architecture.
[0048] S12, the first convolution layer conv1: the number of input channels is 3. First, it passes through a convolution kernel with a size of 64. , a convolutional layer with a stride of 2, followed by a Batch Normalization layer, and finally a ReLU activation function. The number of output channels is 64.
[0049] S13, the first residual block conv2_x: the number of input channels is 64. First, it undergoes a maximum downsampling pooling operation with a convolution kernel size of 3×3 and a stride of 2; then it stacks 2 Each layer of the residual block of the structure passes through a Batch Normalization layer and a ReLU activation function. The number of output channels is 64.
[0050] S14, the second residual block conv3_x: the number of input channels is 64. This layer is stacked 2 Each layer of the residual block of the structure passes through a Batch Normalization layer and a ReLU activation function. The number of output channels is 128.
[0051] S15, the third residual block conv4_x: the number of input channels is 128. This layer is stacked 2 Each layer of the residual block of the structure passes through a Batch Normalization layer and a ReLU activation function. The number of output channels is 256.
[0052] S16, the fourth residual block conv5_x: the number of input channels is 256. This layer is stacked 2 Each layer of the residual block of the structure passes through a Batch Normalization layer and a ReLU activation function. The number of output channels is 512.
[0053] S17, average pooling layer and fully connected layer average pool & fc: Finally, through an average pooling layer and a fully connected layer, the number of input channels is 512, and the number of output channels is 22, which is the number of categories corresponding to the human gesture classification in the Widar3.0 dataset.
[0054] The method based on ResNet18 combined with the reshape operation in the above embodiment is denoted as ResNet18(1.0). It can be seen that before the sample data enters the first convolutional layer conv1 of ResNet18, a reshape operation is performed on the sample. After this reshape operation, it is then input into the original ResNet18 network. After testing, ResNet18(1.0) achieved a prediction accuracy of 71.70% on the test set, exceeding the 71.00% prediction accuracy of the conventional ResNet18 model without the reshape operation.
[0055] Although the reshape operation can make the number of sample channels correspond to the network architecture of ResNet18, it will also cause the input samples to lose their order information, making it difficult to extract the temporal features in the samples and limiting the improvement of the model's recognition accuracy.
[0056] To this end, in some optional embodiments, when the base learner is a ResNet18 model, before the BVP sample data is input into the ResNet18 model, a featureConv operation is performed on the BVP sample data, referring to FIG2 , including:
[0057] Performing a dimension permutation on the BVP sample data, and subjecting the BVP sample data after the dimension permutation to convolution operations in sequence through a first convolutional layer and a second convolutional layer to obtain a first intermediate tensor;
[0058] After performing a second dimension permutation on the first intermediate tensor, the first intermediate tensor is again subjected to convolution operations through the first convolutional layer and the second convolutional layer in sequence to obtain a second intermediate tensor;
[0059] After performing three-dimensional permutation on the second intermediate tensor, a third intermediate tensor is obtained as the input of the ResNet18 model, and the number of input channels of the ResNet18 model is adapted to the dimension of the third intermediate tensor.
[0060] Furthermore, the first convolutional layer includes a third number of convolution kernels with a third step size and a third size, a padding size of a third padding amount, and adopts a ReLU activation function as an activation function; and
[0061] The second convolution layer includes a fourth number of convolution kernels with a fourth step size and a fourth size, a padding size of a fourth padding amount, and adopts a ReLU activation function as an activation function.
[0062] Furthermore, the number of input channels of the first convolutional layer is 20, the third number is 11, the third step size is 1, the third size is 7*7, and the third padding amount is 3;
[0063] The number of output channels of the second convolutional layer is 5, the fourth number is 5, the fourth step size is 1, the fourth size is 5*5, and the fourth padding amount is 2.
[0064] In some specific implementations, for the 22×20×20 BVP sample dataset provided by Widar3.0, the featureConv operation includes:
[0065] S21, will (Right now ) is first replaced by (Right now ).
[0066] S22. Perform the following convolution operation on the permuted sample: The number of input channels is 20. First, pass it through a convolution layer with 11 convolution kernels, a convolution kernel size of 7, a stride of 1, and a padding of 3. Then, pass it through a ReLU activation function. Then, pass it through a convolution layer with 5 convolution kernels, a convolution kernel size of 5, a stride of 1, and a padding of 2. Finally, pass it through a ReLU activation function. The number of output channels is 5.
[0067] S23, after the above convolution operation, the (Right now ) (i.e., the first intermediate tensor) is then permuted into (Right now ), and then repeat the operation of step S22.
[0068] S24, after the above two dimension permutation and convolution operations, we get (Right now ) size sample (i.e. the second intermediate tensor), and then replace the sample to its original appearance, i.e. ( ), and then input the sample (i.e., the third intermediate tensor) into the original ResNet18 network architecture, and change the number of input channels of the conv1 part in ResNet18 to 22.
[0069] The network architecture after the above-mentioned featureConv operation is ResNet18 (2.0). The comparison between the featureConv operation and the reshape operation in ResNet18 (1.0) is shown in Figure 3.
[0070] After testing, ResNet18 (2.0) after using the featureConv operation achieved a prediction accuracy of 72.98% on the test set, exceeding the recognition accuracy of 71.70% of ResNet18 (1.0), indicating that the featureConv operation did play a role.
[0071] It should be noted that due to the number of channels in the BVP sample data (22) It is more about expressing the order of actions rather than different performances of the same action. Therefore, when considering convolution, featureConv embeds the two information dimensions of a single image into the time dimension through a series of alternating transposed convolutions (i.e., the dimension permutation and convolution operations mentioned above). While preserving the order of the actions, it also extracts the features of a single frame of action, enriching the available information and improving recognition accuracy. In contrast, ResNet18 (1.0) with the reshape operation simply performs a standard convolution operation and changes the shape of the input tensor, without extracting the features of each time point and combining them with the temporal information.
[0072] It should be emphasized that even ResNet18 (1.0) has higher recognition accuracy than the existing technology.
[0073] In some preferred embodiments, the gesture recognition model uses soft voting for ensemble learning, including:
[0074] receiving a second posture recognition result output by each of the base learners;
[0075] Calculate the predicted probability corresponding to each second posture recognition result using softmax;
[0076] Grouping is performed based on whether the second posture recognition results are the same, and calculating the mean of the predicted probabilities corresponding to the second posture recognition results as the grouping probability;
[0077] The second gesture recognition result corresponding to the largest grouping probability is selected as the first gesture recognition result.
[0078] In some optional embodiments, the process of selecting the base learner in the gesture recognition model includes:
[0079] Receiving a second posture recognition result output by each candidate base learner on a test set; wherein the test set includes test BVP sample data and corresponding true category labels;
[0080] Determining a prediction accuracy of each candidate base learner on a test set according to the second posture recognition result and the true category label;
[0081] A preset number of the candidate base learners are selected according to the prediction accuracy and combined into the posture recognition model.
[0082] It should be understood that the candidate base learners refer to base learners used for training for screening as part of the integrated model (ie, the gesture recognition model).
[0083] In some specific implementation processes, based on the Widar3.0 dataset, the pseudo code for the soft voting algorithm implementation and accuracy verification of multiple candidate base learners is shown in Table 1.
[0084] Table 1 Soft voting pseudo code
[0085]
[0086] During the implementation, the most intuitive model selection strategy was adopted. Based on the performance of individual models on the test set, the six best-performing models were selected: ResNet18(1.0), ResNet18(2.0), ResNet50, MLP, BiLSTM, and ViT. These models were combined separately to obtain five different model combinations (i.e., posture recognition models after integrating multiple base learners), which were respectively recorded as Ensemble1, Ensemble2, Ensemble3, Ensemble4, and Ensemble5. Table 2 shows the performance of these five different model combinations on the test set.
[0087] Table 2 Model ensemble prediction accuracy
[0088]
[0089] The above-mentioned embodiment introduces ensemble learning to significantly improve the model's prediction accuracy, even exceeding 80%. Specifically, a soft voting method is used to calculate the predicted class probabilities of all models (i.e., base learners). These predicted probabilities are summed and averaged. The class with the highest probability among the obtained probabilities (i.e., the second gesture recognition result) is the final predicted class (i.e., the first gesture recognition result).
[0090] Those skilled in the art should understand that, based on the method described in this embodiment, the ensemble learning strategy may also be implemented using methods such as averaging, weighted averaging, and hard voting.
[0091] Furthermore, training the candidate base learner based on the early stopping method includes:
[0092] Establishing the initial candidate base learner, initializing the counter, optimal loss value, loss value change threshold and patience value;
[0093] Iteratively train the candidate base learners on the training set to update the model parameters, and
[0094] After every T training rounds:
[0095] Determining a validation loss value of the candidate base learner on a validation set;
[0096] According to whether the optimal loss value is an initialization value or whether the difference between the optimal loss value and the verification loss value meets the loss value change threshold, updating the optimal loss value based on the verification loss value and reinitializing the counter;
[0097] When the difference between the optimal loss value and the verification loss value does not meet the loss value change threshold, the counter is incremented to end the current round, and
[0098] Whether to end the training is determined according to whether the value of the counter meets the patience value; when the value of the counter meets the patience value, the training is stopped and the model parameters corresponding to the candidate base learner of the current round are saved.
[0099] It should be understood that the patience value is a preset threshold set for the counter.
[0100] Furthermore, different base learners may be trained on the training set with the same training-to-test ratio or with different training-to-test ratios, which may be set by those skilled in the art according to actual conditions.
[0101] In order to further alleviate the overfitting phenomenon in the model, this embodiment also explores the setting of the value of the model training round epoch, and each model needs to find the training round that best suits it.
[0102] The simplest approach is to set multiple epoch values that might meet the criteria, perform multiple training runs at these epoch values, compare the results, and select the epoch value with the highest accuracy on the test set as the final training round. However, this method is not only time-consuming, but the resulting epoch value may not be the optimal training round.
[0103] Therefore, this embodiment adopts the early stopping method to solve this problem, which can save time and find the training rounds that are most suitable for the base learner at one time.
[0104] In some specific implementations, the pseudo code for implementing the early stopping method used in the candidate base learner training process is shown in Table 3.
[0105] Table 3 Early stopping pseudo code
[0106]
[0107] It should be noted that the above-described embodiments employ early stopping to address overfitting and enhance the model's generalization capabilities. When training large deep learning models or large datasets, each additional training cycle can consume significant GPU hours, energy, and storage space. Early stopping ensures that only training that improves the model's generalization capabilities is performed, avoiding inefficient computation during the overfitting phase. This is highly beneficial for accelerating experimental iterations and optimizing resource utilization.
[0108] It should be emphasized that the computing resources required for ensemble learning (including CPU, GPU time, and memory, etc.) are usually large. This embodiment uses the early stopping method to further save computing resources while ensuring accuracy, significantly reducing time costs.
[0109] This embodiment also provides a computer-readable storage medium, on which is stored at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor, so that the processor performs some or all steps of the method provided in the embodiment of the present application.
[0110] It is understood that the storage medium may be transient or non-transient. Exemplarily, the storage medium includes, but is not limited to, a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code.
[0111] Exemplarily, the processor may be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA).
[0112] Exemplarily, the read-only memory includes but is not limited to MASK ROM, PROM, EPROM, EEPROM, Flash, etc.
[0113] Exemplarily, the random access memory includes but is not limited to DRAM, SRAM, SDRAM, DDR SDRAM, etc.
[0114] In some examples, a computer program product is provided, which can be implemented in hardware, software, or a combination thereof. As a non-limiting example, the computer program product can be embodied as the storage medium, or as a software product, such as an SDK (Software Development Kit).
[0115] As a non-limiting example, a computer program product is provided, comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform some or all of the steps of the method described in the embodiments of the present application.
[0116] In some examples, a computer program is provided, comprising a computer-readable code. When the computer-readable code is run in a computer device, a processor in the computer device executes the code to implement part or all of the steps in the method.
[0117] This embodiment also proposes an electronic device, including a memory and a processor, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and when the processor executes the at least one instruction, at least one program, code set or instruction set, it implements part or all of the steps of the method described in the embodiment of the present application.
[0118] In some examples, a hardware entity of the electronic device is provided, including: a processor, a memory and a communication interface; wherein the processor generally controls the overall operation of the electronic device; the communication interface is used to enable the electronic device to communicate with other terminals or servers through a network; the memory is configured to store instructions and applications executable by the processor, and can also cache data to be processed or processed by the processor and various modules in the electronic device (including but not limited to image data, audio data, voice communication data and video communication data), and can be implemented by flash memory (FLASH), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or random access memory (RAM).
[0119] A processor may include one or more processing elements. Thus, a processor may include one or more integrated circuits (ICs) configured to perform the functions of the processor. Furthermore, each integrated circuit may include circuits (e.g., a first circuit, a second circuit, and other circuits) configured to perform the functions of the processor.
[0120] Furthermore, data may be transmitted between the processor, the communication interface and the memory via a bus, which may include any number of interconnected buses and bridges, connecting various circuits of one or more processors and memories.
[0121] Example 2
[0122] In order to prove the technical effect of the method described in this application, this embodiment also conducted a comparative experiment on the Widar3.0 dataset between the method proposed in this application and the method based on the conventional network model. The deep learning environment used in the experiment is shown in Table 4.
[0123] Table 4 Deep learning environment
[0124]
[0125] First, we use the conventional method (fixed 100 epochs) to train multiple deep neural network models as base learners. The prediction accuracy is shown in Table 5:
[0126] Table 5 Prediction accuracy of conventional scheme
[0127]
[0128] Next, we trained and tested multiple models using the early stopping method, including seven models: MLP, LSTM, BiLSTM, ResNet18(1.0), ResNet18(2.0), ResNet50, and ViT. Figures 4-7 show the accuracy and loss of these seven models during training with a training-to-test ratio of 0.8:0.2, and Figures 8-11 show the accuracy and loss of these seven models during training with a training-to-test ratio of 0.9:0.1.
[0129] Referring to Table 6, the experimental results show that under the condition of a training-test ratio of 0.9:0.1, ResNet18(2.0) achieved a prediction accuracy of 74.74%, surpassing the highest prediction accuracy of 72.98% achieved by ResNet18(2.0) on this dataset before the introduction of the early stopping method, and achieving a new high in accuracy on the current test set. This shows that the introduction of the early stopping method has indeed alleviated the overfitting phenomenon in model training to a certain extent.
[0130] Table 6 Accuracy of each model on the test set (training-test ratio 0.9:0.1)
[0131]
[0132] Furthermore, the model training curves in Figures 4-11 show the training time for each model. The longest-lasting model took just over an hour to train, while the shortest took only 15 minutes. Compared to conventional model training strategies (such as training each base learner for 100 epochs), model training time was reduced from half a day to under an hour, with some models even achieving fitting after just a few epochs. Early stopping, which prematurely terminates unnecessary training iterations, significantly reduces the time and computing resources required for model training.
[0133] Furthermore, combining Tables 5 and 6 reveals that the prediction accuracy of ResNet50 increased from 68.56% to 71.40%, achieving the largest improvement among the seven models. After applying early stopping, the prediction accuracy of ResNet50 also exceeded the target accuracy of 71%. For deep neural network models, the more parameters they have, the more expressive the model is, allowing them to learn more complex functional relationships to adapt to the details of the training data. High-complexity models offer greater flexibility, enabling them to accurately fit specific features, including noise, during training. When the number of model parameters is too large, the model may have sufficient degrees of freedom to "memorize" individual examples or noise in the training set rather than learning the underlying general patterns. This over-adaptation to the training data increases the risk of the model failing to generalize well to new data, a phenomenon known as overfitting. Table 7 shows the parameter counts of the seven models mentioned above. ResNet50 has 23.58M parameters, making it the most complex model with the most parameters. Therefore, ResNet50 is more likely to overfit. When early stopping is applied to ResNet50, the overfitting effect is more obvious, and the model is improved.
[0134] Table 7 Model parameters
[0135]
[0136] Finally, a comparison between the method provided in the above embodiment and the conventional method is given, as shown in Table 8. The base learner combinations corresponding to different ensembles can be found in Table 2.
[0137] Table 8 Comparison with conventional methods
[0138]
[0139] It can be seen that the method provided in this application has achieved improvements over conventional methods, with recognition accuracy reaching up to 81.46%. Building on the improved accuracy of ResNet18 (2.0) using the featureConv operation, the accuracy of the ensemble model based on the ensemble learning strategy (i.e., the posture recognition model) has been significantly improved.
[0140] It can be understood that the options in the above embodiment 1 are also applicable to this embodiment, so they will not be described again here.
[0141] The same or similar reference numerals correspond to the same or similar components;
[0142] The terms used in the drawings to describe positional relationships are for illustrative purposes only and are not to be construed as limiting the present application.
[0143] It should be noted that, unless there is any conflict, the embodiments and features in the embodiments of this application can be combined with each other.
[0144] In different specific implementations, the method or system described in this application can be implemented in software, hardware or a combination thereof. In addition, the order of the steps of the method can be changed, and various elements can be added, reordered, combined, omitted, modified, etc.
[0145] Obviously, the above embodiments of the present application are merely examples for clearly illustrating the present application, and are not intended to limit the implementation methods of the present application, and are not intended to limit the present application. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. Each discrete structural / functional module or unit can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part, and the structure and function of the discrete components can be implemented as a combined structure or component. It is not necessary and impossible to enumerate all the implementation methods here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the claims of the present application.
Claims
1. A gesture recognition method based on wireless sensing, characterized in that: include: Receive BVP sample data; Determining a first posture recognition result based on the BVP sample data using a posture recognition model; wherein the posture recognition model determines the first posture recognition result by fusing second posture recognition results output by at least two base learners based on the BVP sample data based on an ensemble learning strategy, wherein the base learners are trained based on an early stopping method; When at least one of the base learners is a ResNet18 model, before inputting the BVP sample data into the ResNet18 model, performing a featureConv operation on the BVP sample data, including: Performing a dimension permutation on the BVP sample data, and subjecting the BVP sample data after the dimension permutation to convolution operations in sequence through a first convolutional layer and a second convolutional layer to obtain a first intermediate tensor; After performing a second dimension permutation on the first intermediate tensor, the first intermediate tensor is again subjected to convolution operations through the first convolutional layer and the second convolutional layer in sequence to obtain a second intermediate tensor; After performing three-dimensional permutation on the second intermediate tensor, a third intermediate tensor is obtained as the input of the ResNet18 model, and the number of input channels of the ResNet18 model is adapted to the dimension of the third intermediate tensor.
2. The method for gesture recognition based on wireless sensing according to claim 1, characterized in that: The base learner includes at least one of an MLP model, an LSTM model, a BiLSTM model, a ResNet18 model, a ResNet50 model, and a ViT model.
3. The method for gesture recognition based on wireless sensing according to claim 2, characterized in that: When at least one of the base learners is a ResNet18 model, before inputting the BVP sample data into the ResNet18 model, a reshape operation is performed on the BVP sample data, including: The BVP sample data is sequentially passed through the first deconvolution layer and the second deconvolution layer; wherein, The first deconvolution layer includes a first number of convolution kernels with a first step length and a first size, a padding size of a first padding amount, and uses a ReLU activation function as an activation function; and The second deconvolution layer includes a second number of convolution kernels with a second step size and a second size, a padding size of a second padding amount, and adopts a ReLU activation function as an activation function.
4. The method for gesture recognition based on wireless sensing according to claim 3, characterized in that: The first convolutional layer includes a third number of convolution kernels with a third step size and a third size, a padding size of a third padding amount, and adopts a ReLU activation function as an activation function; and The second convolution layer includes a fourth number of convolution kernels with a fourth step size and a fourth size, a padding size of a fourth padding amount, and adopts a ReLU activation function as an activation function.
5. The method for gesture recognition based on wireless sensing according to claim 4, characterized in that: The number of input channels of the first convolutional layer is 20, the third number is 11, the third step size is 1, the third size is 7*7, and the third padding amount is 3; The number of output channels of the second convolutional layer is 5, the fourth number is 5, the fourth step size is 1, the fourth size is 5*5, and the fourth padding amount is 2.
6. A gesture recognition method based on wireless sensing according to any one of claims 1 to 5, characterized in that: The posture recognition model uses soft voting for ensemble learning, including: receiving a second posture recognition result output by each of the base learners; Calculate the predicted probability corresponding to each second posture recognition result using softmax; Grouping is performed based on whether the second posture recognition results are the same, and calculating the mean of the predicted probabilities corresponding to the second posture recognition results as the grouping probability; The second gesture recognition result corresponding to the largest grouping probability is selected as the first gesture recognition result.
7. The method for gesture recognition based on wireless sensing according to claim 6, characterized in that: The selection process of the base learner in the posture recognition model includes: Receiving a second posture recognition result output by each candidate base learner on a test set; wherein the test set includes test BVP sample data and corresponding true category labels; Determining a prediction accuracy of each candidate base learner on a test set according to the second posture recognition result and the true category label; A preset number of the candidate base learners are selected according to the prediction accuracy and combined into the posture recognition model.
8. The method for gesture recognition based on wireless sensing according to claim 7, characterized in that: Training the candidate base learner based on the early stopping method includes: Establishing the initial candidate base learner, initializing the counter, optimal loss value, loss value change threshold and patience value; Iteratively train the candidate base learners on the training set to update the model parameters, and After every T training rounds: Determining a validation loss value of the candidate base learner on a validation set; According to whether the optimal loss value is an initialization value or whether the difference between the optimal loss value and the verification loss value meets the loss value change threshold, updating the optimal loss value based on the verification loss value and reinitializing the counter; When the difference between the optimal loss value and the verification loss value does not meet the loss value change threshold, the counter is incremented to end the current round, and Whether to end the training is determined according to whether the value of the counter meets the patience value; when the value of the counter meets the patience value, the training is stopped and the model parameters corresponding to the candidate base learner of the current round are saved.
9. A computer program product comprising a computer program or computer executable instructions, characterized in that When the computer program or computer executable instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.