A method and device for single-sample action recognition based on WiFi

By generating virtual action data and training a deep learning model in two stages, single-sample action recognition of the WiFi action recognition method was achieved, which solved the problem of high sample collection and training overhead in the existing technology and improved the scalability and recognition accuracy of the system.

CN118277853BActive Publication Date: 2026-07-28ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2024-03-27
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing WiFi action recognition methods require a large number of samples to achieve ideal performance, and the model needs to be retrained from scratch when faced with new actions, resulting in high data collection and training costs and poor scalability.

Method used

By collecting WiFi channel state information of basic actions, preprocessing it to generate Doppler spectrum data, combining it with physical modeling to generate virtual action data, and using a deep learning model for two-stage training, single-sample action recognition is achieved.

Benefits of technology

When recognizing new actions, only a single sample data is needed for model fine-tuning, which reduces the overhead of sample collection and training, and improves the scalability and recognition accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118277853B_ABST
    Figure CN118277853B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for single-sample action recognition based on WiFi, which physically models the WiFi action recognition scene, generates virtual action data from basic actions, thereby enriching the data set and reducing the cost of collecting real data. Based on supervised learning and meta-learning, a single-sample action recognition framework is designed and implemented. The framework uses a traditional supervised learning mechanism for first-stage training on a virtual action data set, uses a single-sample meta-learning mechanism for second-stage training on a basic action data set, and uses a single-sample meta-learning mechanism for model fine-tuning on a single-sample data set of a new action, thereby obtaining a model capable of accurately recognizing the new action. The training of the two stages only needs to be performed at the initial deployment, and if there is a need to change the action type subsequently, only the model fine-tuning needs to be performed by using a single-sample data of a new action, thereby significantly reducing the model training cost, being highly scalable, and being applicable to actual application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless signal human motion perception, specifically, it is a method and device for single-sample motion recognition based on WiFi. Background Technology

[0002] Currently, WiFi infrastructure has permeated countless households, and its potential in the sensing field has been extensively explored. Motion recognition is a crucial application of WiFi sensing and a core component of many human-computer interaction applications. Compared to traditional sensing solutions based on cameras, sonar, or wearable devices, WiFi-based motion recognition offers numerous advantages, such as: no need for wearable sensors, reduced privacy concerns, and immunity to light interference. Therefore, WiFi motion recognition technology has broad application prospects in health monitoring, human-computer interaction, and smart homes.

[0003] However, existing WiFi action recognition methods suffer from the following problems. On the one hand, most recognition algorithms in the field of WiFi action recognition still rely on traditional supervised learning, requiring a large number of samples for each action to achieve satisfactory performance; however, collecting WiFi action data is time-consuming and labor-intensive, and obtaining a large number of high-quality samples is costly. On the other hand, in real-world scenarios, adding new actions is a common requirement, but traditional action recognition methods have poor scalability, requiring the model to be retrained from scratch in such cases, resulting in significant overhead. Summary of the Invention

[0004] This invention addresses the problems existing in the prior art by providing a WiFi-based single-sample action recognition method and device. Specifically, it trains an action classification model based on collected WiFi data of basic actions and virtual action data generated based on physical modeling. For any new type of action, only a single sample of each type of action needs to be used to fine-tune the model, thereby achieving WiFi action recognition for new actions.

[0005] Specifically, this method collects WiFi channel state information (CSI) when a human performs basic actions, preprocesses it to obtain Doppler spectrum data, generates virtual action data using physical modeling, uses the virtual action data and basic action data to train a deep learning model in two stages, then collects single-sample CSI data of new actions and preprocesses it to convert it into Doppler spectrum data, and uses the single-sample data of new actions to fine-tune the model after the two-stage training to obtain a model that can accurately identify new actions.

[0006] This invention is achieved through the following technical solution:

[0007] A WiFi-based single-sample action recognition method includes the following steps:

[0008] S1. Obtain Channel State Information (CSI) when a human body performs a preset basic action in a WiFi environment;

[0009] S2. Perform preprocessing on the CSI data obtained in step S1, such as cropping, filtering, and time-frequency conversion, to obtain Doppler spectrum data;

[0010] S3. Based on the Doppler spectrum data obtained in step S2, generate virtual action Doppler spectrum data by physically modeling the WiFi action recognition scenario;

[0011] S4. Construct an action recognition model, using the basic action data and virtual action data described in steps S2 and S3 as input, and after two stages of training, obtain model parameters that can accurately recognize basic actions.

[0012] S5. Acquire CSI data of the human body when performing new actions in a WiFi environment. Only one data point is required for each type of action.

[0013] S6. Use the preprocessing method described in step S2 to convert the CSI data obtained in step S5 into Doppler spectrum data of the new action.

[0014] S7. Input the Doppler spectrum data of the new action obtained in step S6 into the action recognition model in step S4, and fine-tune the model parameters based on the two-stage training described in step S4 to obtain a model that can accurately recognize the new action and realize single-sample action recognition.

[0015] S8. Collect CSI data in real time when the user performs the new action involved in step S5, and input it into the recognition model obtained in step S7 after preprocessing as described in step S2, and output the recognized action type.

[0016] As a further improvement, in step S2 of this invention, the CSI data obtained in step S1 is preprocessed by cropping, filtering, and time-frequency conversion to obtain Doppler spectrum data, which is used to eliminate noise and extract environment-independent WiFi data features. Specifically:

[0017] The original CSI data is cropped to a uniform length; for each receiver, the antenna with the highest mean variance ratio is selected as the reference signal, and the reference signal is multiplied by the other subcarriers of that receiver using conjugate multiplication; low-pass and high-pass filters are used to remove static and high-frequency components respectively; principal component analysis (PCA) is used to extract the first principal component to reduce the remaining noise interference;

[0018] Using the sliding window technique, a short-time Fourier transform (STFT) is performed within a selected time window to extract Doppler spectrum data from the CSI data; the amplitude of the STFT is the Doppler spectrum, expressed as:

[0019] Where f is the frequency, t is the time, x(t) is the CSI data, and w(t) is the window function;

[0020] On the Doppler spectrum S(f,t), a cutoff frequency of [-60Hz, 60Hz] is set to obtain the Doppler spectrum within the frequency range that best represents the Doppler frequency shift, which is then used for further action recognition.

[0021] As a further improvement, in step S3 of this invention, based on the Doppler spectrum data obtained in step S2, virtual action Doppler spectrum data is generated by physically modeling the WiFi action recognition scenario to expand the training dataset, specifically:

[0022] In the WiFi motion recognition scenario, the WiFi signal emitted by the WiFi transmitter is reflected by a moving human target and reaches the signal receiver. According to the Doppler effect, human motion causes a difference between the frequency of the WiFi signal received by the receiver and the frequency of the signal emitted by the transmitter. This difference is called the Doppler frequency shift, expressed as:

[0023]

[0024] Where f0 is the frequency of the WiFi signal transmitted by the transmitter, v is the speed of the target, c is the speed of light, θ is the angle between the target's direction of motion and the line connecting the transceiver, and α... T Let α be the departure angle (AoD). R Angle of arrival (AoA);

[0025] In the scenario described, if the basic action is accelerated by a factor of n, then the position at time t after acceleration is the same as the position at time nt before acceleration, satisfying s acc (t)=s(nt), θ acc (t) = θ(nt), Where X acc X represents the physical quantity after acceleration; the velocity v after acceleration is obtained according to the aforementioned formula. acc (t) = n·v(nt);

[0026] Substituting the given relationship into the Doppler frequency shift formula, we obtain the Doppler frequency shift f at time t after acceleration. acc (t) = n·f(nt);

[0027] The original Doppler spectral data S(f,t), where f∈[-F,F], t∈[0,T], is used to determine the accelerated Doppler spectral data based on the Doppler frequency shift equation: Where f∈[-nF,nF],

[0028] For any two basic actions A and B, their Doppler spectra S A (f,t), S B The Doppler spectrum data for generating virtual action A+B using (f,t)(f∈[-F,F],t∈[0,T]) is:

[0029]

[0030] As a further improvement, in step S4 of this invention, an action recognition model is constructed using the basic action data and virtual action data described in steps S2 and S3 as input. After two stages of training, model parameters that can accurately recognize basic actions are obtained, specifically:

[0031] This system employs a deep neural network architecture, using Doppler spectrum data as input to output predicted action recognition type labels. The deep neural network consists of a feature extractor and a classifier. The feature extractor contains three convolutional layers using two-dimensional convolutions to extract features in both the time and frequency domains. After each convolutional layer, batch normalization (BN) and rectified linear unit (ReLU) activation functions are used to prevent network distribution shift and overfitting. After the final convolutional layer and the corresponding BN and ReLU, the extracted two-dimensional features are flattened into one dimension and fed into the subsequent classifier. The classifier contains two linear fully connected functions, with a ReLU function between them to increase the non-linearity of the deep neural network.

[0032] The parameters of the deep neural network are randomly initialized, and the network is trained using virtual action data to obtain the parameters of the feature extractor and classifier after the first stage of training.

[0033] The deep neural network is constructed using the feature extractor parameters trained in the first stage and the randomly initialized classifier parameters. A series of single-sample tasks are sampled from basic action data, and the network is trained using a single-sample meta-learning mechanism to obtain the feature extractor and classifier parameters trained in the second stage.

[0034] As a further improvement, in step S7 of this invention, the Doppler spectrum data of the new action obtained in step S6 is input into the action recognition model in step S4. Based on the model parameters trained in the two stages described in step S4, fine-tuning is performed to obtain a model capable of accurately recognizing the new action, thus achieving single-sample action recognition. Specifically:

[0035] The deep neural network is constructed using the feature extractor parameters trained in the second stage of step S4 and the randomly initialized classifier parameters. The network is then fine-tuned using a new Doppler dataset containing only a single sample for each action class, employing a single-sample meta-learning mechanism, to obtain a model that can achieve single-sample action recognition on new actions.

[0036] The present invention also discloses a WiFi-based single-sample action recognition device, comprising:

[0037] The basic motion data acquisition module is used to acquire the channel state information (CSI) when a human body performs a preset basic motion in a WiFi environment;

[0038] The basic motion data preprocessing module is used to preprocess the channel state information (CSI) obtained by the basic motion data acquisition module to obtain the Doppler spectrum data of the basic motion.

[0039] The virtual motion generation module is used to generate Doppler spectrum data for virtual motions based on the Doppler spectrum data obtained by the basic motion data preprocessing module.

[0040] The two-stage training module of the model is used to build an action recognition model. It uses the basic action data and virtual action data obtained by the basic action data preprocessing module and the virtual action generation module as input. After two stages of training, the model parameters that can accurately recognize basic actions are obtained.

[0041] The single-sample data acquisition module is used to acquire CSI data of the human body when performing new actions in a WiFi environment. Only one data point is required for each action type.

[0042] The single-sample data preprocessing module is used to preprocess the CSI data obtained by the single-sample data acquisition module to obtain the Doppler spectrum data of the new action;

[0043] The model fine-tuning module is used to input the Doppler spectrum data of the new action obtained by the single sample data preprocessing module into the action recognition model obtained by the two-stage training module. Based on the model parameters of the two-stage training module, the module is fine-tuned to obtain a model that can accurately recognize the new action.

[0044] The real-time new action recognition module is used to collect CSI data when the user performs a new action involved in the single sample data acquisition module. After preprocessing and converting it into Doppler spectrum data, it is input into the recognition model obtained by the model fine-tuning module and outputs the recognized action type.

[0045] The beneficial effects of this invention are as follows:

[0046] Current WiFi action recognition technologies often rely on traditional supervised learning to classify actions, which suffers from problems such as requiring a large number of samples, high data acquisition costs, and the need to retrain the model from scratch when changing action types, resulting in high training costs and poor scalability. This invention uses existing technology to collect CSI data and denoise it, then selects the environment-independent Doppler spectrum as samples to better reflect human dynamic behavior and reduce interference from environmental factors. To address the high cost of sample acquisition, this invention designs a method based on signal propagation laws to expand the training dataset. Specifically, it physically models the WiFi action recognition scene, generating virtual action data from basic actions, thereby enriching the dataset and reducing the cost of collecting real data. To address the scalability issue, this invention designs and implements a single-sample action recognition framework based on supervised learning and meta-learning. Specifically, the framework employs a traditional supervised learning mechanism for the first stage of training on a virtual action dataset, a single-sample meta-learning mechanism for the second stage of training on a basic action dataset, and a single-sample meta-learning mechanism for fine-tuning the model on a single-sample dataset of new actions, thereby obtaining a model capable of accurately recognizing new actions. In the single-sample action recognition framework designed in this invention, the two stages of training only need to be performed during initial deployment. If there is a need to change the action type subsequently, such as adding a new action type, it is only necessary to reuse the single-sample data of the new action for model fine-tuning, without having to train from scratch. Compared with existing technologies, the framework designed in this invention significantly reduces model training overhead, has strong scalability, and is suitable for practical application scenarios. Attached Figure Description

[0047] Figure 1 This is a flowchart of the present invention;

[0048] Figure 2 This is a schematic diagram of the experimental equipment layout for implementing the present invention. Detailed Implementation

[0049] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments:

[0050] The purpose of this invention is to address the problems of excessive sample acquisition overhead and excessive training overhead when facing new actions in existing WiFi action recognition technologies, and to propose a single-sample action recognition method based on WiFi. Figure 1 This is a flowchart of the present invention;

[0051] The specific implementation method of the present invention is as follows:

[0052] S1. Collect WiFi CSI signals of personnel performing pre-set basic actions in a WiFi sensing environment; both the WiFi transmitter and receiver use laptops with Intel 5300 wireless network cards, with one antenna for each transmitter and three antennas for each receiver.

[0053] S2. Trim the CSI data of the basic actions described in step S1 to a uniform time length; select the antenna with the highest average variance ratio among the three antennas connected to the same receiver as the reference signal, and multiply the CSI data of the other antennas at the receiver with the reference signal by conjugate to eliminate the phase shift caused by the clock asynchrony between the transmitter and receiver; use a low-pass filter with a cutoff frequency of 2Hz to eliminate low-frequency reflections caused by static paths in the environment; use a high-pass filter with a cutoff frequency of 60Hz to remove high-frequency noise; use principal component analysis on the filtered CSI data to extract the first principal component that best reflects the user's action characteristics; perform a short-time Fourier transform on the denoised CSI data, calculate the square of the amplitude of the short-time Fourier transform result, and retain the part with frequencies in the range of [-60Hz, 60Hz] to obtain Doppler spectrum data;

[0054] S3. By modeling and analyzing the signal transmission of WiFi action recognition scenarios, the formula for speeding up actions at the Doppler spectrum level as described in the invention is obtained; for the Doppler spectrum data of each basic action, the formula is used to obtain data after speeding up twice; the Doppler spectrum data of each pair of basic actions after speeding up twice are spliced ​​together to obtain the Doppler spectrum data of virtual actions.

[0055] S4. Construct a deep neural network model for action recognition. The model consists of a feature extractor and a classifier. The feature extractor contains three layers of two-dimensional convolutions, followed by a batch normalization function (BN) and a linear rectified function (ReLU) after each convolution. The classifier contains two linear fully connected functions and a ReLU function between them. The input Doppler spectrum data is used to extract two-dimensional features from the feature extractor. The two-dimensional features are flattened into one dimension and then input into the classifier. The classifier outputs the probability corresponding to each category, thereby inferring the action category corresponding to the WiFi data.

[0056] The action recognition model is constructed using randomly initialized feature extractors and classifiers. The number of output features of the last linear layer of the classifier is set to be the same as the number of virtual action categories. The Doppler spectrum data of the virtual actions is used as the training set, and the model parameters are optimized by minimizing the cross-entropy loss function to obtain the initial training parameters of the feature extractor and classifier. The feature extractor will continue to be used in the next stage.

[0057] An action recognition model is constructed using a feature extractor obtained after initial training and a randomly initialized classifier. The number of output features of the last linear layer of the classifier is set to a natural number n greater than 2 and not exceeding the number of basic action categories. A series of subtasks T are obtained by sampling from the basic action Doppler spectrum dataset. i Each subtask is supported by a set S i and query set Q i It consists of n basic actions and supports a set S. i Each action category contains only one sample;

[0058] In each subtask T i In, the support set S is used i To optimize the classifier's parameters φ, the optimized parameters φ for the subtask are calculated by minimizing the cross-entropy loss function. i And record it, without directly updating the model parameters; on each subtask, based on the optimized parameters φ i Calculate query set Q i The loss is calculated by summing the query set losses for all subtasks, which is then used as the cumulative loss for all subtasks. This cumulative loss is minimized to optimize the parameters of the feature extractor and classifier. The resulting trained feature extractor will be used in subsequent stages.

[0059] S5. Collect WiFi CSI signals when the user performs new actions in a WiFi-aware environment. Only one data point is needed for each type of action.

[0060] S6. Using the preprocessing methods described in step S2, such as cropping, conjugate multiplication, filtering, principal component analysis, and short-time Fourier transform, convert the CSI data of the new action described in step S5 into Doppler spectrum data.

[0061] S7. Using the trained feature extractor described in step S4 and the randomly initialized classifier, an action recognition model is constructed. The number of output features of the last linear layer of the classifier is set to be the same as the number of virtual action categories. Using the Doppler spectrum data of the new action described in step S6, the classifier parameters are optimized by minimizing the loss function, while the feature extractor parameters remain unchanged. That is, the model is fine-tuned using a single sample dataset of the new action. After fine-tuning, a model that can accurately recognize the new action is obtained.

[0062] S8. Real-time acquisition of the CSI when the user performs the new action involved in step S5, and input into the recognition model obtained by single-sample fine-tuning in step S7 after preprocessing described in step S2. The model can output the corresponding action type more accurately.

[0063] The WiFi-based single-sample action recognition method of the present invention was experimentally verified as follows:

[0064] In a laboratory setting, a 2m x 2m WiFi sensing area is constructed using one transmitter and three receivers, such as... Figure 2 As shown; the WiFi data packet transmission rate is 1000Hz; in step S1, WiFi data of 20 preset basic actions are collected, including gesturing the numbers "0" to "9", etc.; in step S3, Doppler spectrum data of 380 virtual actions are generated by arranging the 20 basic actions in pairs; in step S5, individual samples of 6 new actions are collected, including "push and pull", "sweep from left to right", "swipe from right to left", "high five", "draw a zigzag", and "draw a triangle"; in the feature extractor of the action recognition model, the first convolutional layer uses a 3×5 convolutional kernel, and the second and third convolutional layers both use a 4×4 convolutional kernel, and the sliding stride of the three convolutional layers is (2,2).

[0065] Using the single-sample action recognition method of the present invention, the recognition accuracy on the six new actions reaches 93%. Moreover, when the new action changes, the method of the present invention does not need to retrain the model from scratch, but only needs to fine-tune the model with a single sample, that is, execute steps S5 to S8. The WiFi-based single-sample action recognition method of the present invention reduces the sample collection overhead and training overhead while ensuring the accuracy of action recognition, effectively improving scalability and making it suitable for application deployment in real-world scenarios.

[0066] The above description is not intended to limit the present invention. It should be noted that, for those skilled in the art, various changes, modifications, additions or substitutions can be made without departing from the essential scope of the present invention, and these improvements and refinements should also be considered within the scope of protection of the present invention.

Claims

1. A single-sample action recognition method based on WiFi, characterized in that: Includes the following steps: S1. Obtain Channel State Information (CSI) when a human body performs a preset basic action in a WiFi environment; S2. Preprocess the Channel State Information (CSI) obtained in step S1 to obtain Doppler spectrum data; S3. Based on the Doppler spectrum data obtained in step S2, generate Doppler spectrum data for virtual actions; this is used to expand the training dataset, specifically: By physically modeling the WiFi action recognition scenario, the basic actions are spliced ​​at the Doppler spectrum level, and the different spliced ​​basic actions are combined to obtain the Doppler spectrum data corresponding to the virtual actions. S4. Construct an action recognition model. Using the basic action data and virtual action data described in steps S2 and S3 as input, train the model parameters to accurately recognize basic actions. Specifically, use a deep neural network architecture, take Doppler spectrum data as input, and output the predicted type label for action recognition. The deep neural network includes two parts: a feature extractor and a classifier. The feature extractor is mainly composed of convolutional layers, and the classifier is mainly composed of fully connected layers. The parameters of the feature extractor and classifier are randomly initialized, virtual action data is used as the training set, and the model is trained using traditional supervised learning methods to obtain the initial trained parameters of the feature extractor and classifier. Using the pre-trained feature extractor parameters and reconnecting a randomly initialized classifier, a deep neural network is constructed; using basic action data as the training set, a series of single-sample tasks are sampled from it, and the network is trained using a single-sample meta-learning mechanism to obtain the trained feature extractor and classifier parameters. The feature extractor will be used in subsequent stages; S5. Acquire CSI data of the human body when performing new actions in a WiFi environment. Only one data point is required for each type of action. S6. Use the preprocessing method described in step S2 to convert the CSI data described in step S5 into Doppler spectrum data of the new action; S7. Input the Doppler spectrum data of the new action obtained in step S6 into the action recognition model in step S4, and fine-tune it based on the model parameters described in step S4 to obtain a model that can accurately recognize the new action. S8. Collect CSI data in real time when the user performs the new action involved in step S5, and input it into the recognition model obtained in step S7 after preprocessing as described in step S2, and output the recognized action type.

2. The WiFi-based single-sample action recognition method according to claim 1, characterized in that, In step S2, the Channel State Information (CSI) obtained in step S1 is preprocessed to obtain Doppler spectrum data, which is used to eliminate noise and extract environment-independent WiFi data features. Specifically: The original CSI data is cropped to a uniform length; the antenna with the highest mean variance ratio is selected as the reference signal, and the CSI measurement of the reference signal is adjusted. Then, the reference signal is multiplied by the other subcarriers of the receiver using conjugate multiplication; low-pass and high-pass filters are used to remove static and high-frequency components respectively; principal component analysis (PCA) is used to extract the first principal component to reduce the remaining noise interference. Using the sliding window technique, a short-time Fourier transform (STFT) is performed within a selected time window to extract Doppler spectrum data from the CSI data for further action recognition.

3. The WiFi-based single-sample action recognition method according to claim 1, characterized in that, In step S7, the Doppler spectrum data of the new action obtained in step S6 is input into the action recognition model of step S4. Based on the model parameters described in step S4, fine-tuning is performed to obtain a model capable of accurately recognizing the new action. Specifically: Using the feature extractor parameters trained in step S4, and reconnecting a randomly initialized classifier, a deep neural network is constructed. Using a new Doppler dataset containing only a single sample for each action class, the network is fine-tuned using a single-sample meta-learning mechanism to obtain a model that can achieve single-sample action recognition on new actions.

4. A WiFi-based single-sample action recognition device, characterized in that, include: The basic motion data acquisition module is used to acquire the channel state information (CSI) when a human body performs a preset basic motion in a WiFi environment; The basic motion data preprocessing module is used to preprocess the channel state information (CSI) obtained by the basic motion data acquisition module to obtain the Doppler spectrum data of the basic motion. The virtual motion generation module is used to generate Doppler spectrum data for virtual motions based on the Doppler spectrum data obtained from the basic motion data preprocessing module; it is also used to expand the training dataset, specifically: By physically modeling the WiFi action recognition scenario, the basic actions are spliced ​​at the Doppler spectrum level, and the different spliced ​​basic actions are combined to obtain the Doppler spectrum data corresponding to the virtual actions. The two-stage training module is used to construct the action recognition model. It uses basic action data and virtual action data obtained from the basic action data preprocessing module and the virtual action generation module as input. After two stages of training, it obtains model parameters that can accurately recognize basic actions. Specifically, it uses a deep neural network architecture, takes Doppler spectrum data as input, and outputs the predicted type label of action recognition. The deep neural network includes two parts: a feature extractor and a classifier. The feature extractor is mainly composed of convolutional layers, and the classifier is mainly composed of fully connected layers. The parameters of the feature extractor and classifier are randomly initialized, virtual action data is used as the training set, and the model is trained using traditional supervised learning methods to obtain the initial trained parameters of the feature extractor and classifier. Using the pre-trained feature extractor parameters and reconnecting a randomly initialized classifier, a deep neural network is constructed; using basic action data as the training set, a series of single-sample tasks are sampled from it, and the network is trained using a single-sample meta-learning mechanism to obtain the trained feature extractor and classifier parameters. The feature extractor will be used in subsequent stages; The single-sample data acquisition module is used to acquire CSI data of the human body when performing new actions in a WiFi environment. Only one data point is required for each action type. The single-sample data preprocessing module is used to preprocess the CSI data obtained by the single-sample data acquisition module to obtain the Doppler spectrum data of the new action; The model fine-tuning module is used to input the Doppler spectrum data of the new action obtained by the single sample data preprocessing module into the action recognition model obtained by the two-stage training module. Based on the model parameters of the two-stage training module, the module is fine-tuned to obtain a model that can accurately recognize the new action. The real-time new action recognition module is used to collect CSI data when the user performs a new action involved in the single sample data acquisition module. After preprocessing and converting it into Doppler spectrum data, it is input into the recognition model obtained by the model fine-tuning module and outputs the recognized action type.