Metaaction-based few-sample wireless gesture recognition method
By extracting the Doppler spectrum and trajectory curvature characteristics of meta-action gestures, the network is trained using the weighted fusion method, which solves the problem of many samples and insufficient recognition stability in wireless gesture recognition, and achieves efficient and low-cost gesture recognition.
Patent Information
- Application Number
- CN202510703195.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing wireless gesture recognition technology requires a large number of training samples, and it is difficult to effectively distinguish gestures with similar motion trajectories, resulting in high system deployment costs and insufficient recognition stability.
By collecting meta-action gesture data, the Doppler spectrum features and trajectory curvature features are extracted, and the network training is performed using a weighted fusion method, and efficient wireless gesture recognition is achieved using a small number of samples.
Effectively distinguishing gestures under a small number of samples are achieved, reducing system deployment costs and improving identification stability.
Smart Images

Figure CN120234618A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human-computer interaction technology, and particularly to a few-shot wireless gesture recognition method based on meta-actions. Background Art
[0002] Gesture recognition technology, as an important development direction in the field of human-computer interaction, its core value lies in achieving a natural and intuitive operation experience. The current mainstream technical solutions mainly include three implementation paths: sensor-based contact recognition, computer vision-based optical recognition, and wireless signal-based non-contact recognition. Each type of solution faces different technical bottlenecks in practical applications.
[0003] Sensor-based contact recognition solutions usually require users to wear special devices integrated with sensors such as accelerometers and gyroscopes, and gesture parsing is achieved by collecting three-dimensional space motion data. Although this technical route has high measurement accuracy, the physical contact characteristic leads to a decrease in user comfort and a relatively high device maintenance cost. Computer vision-based optical recognition solutions use cameras to capture hand images and analyze gesture features through image processing algorithms. Its non-contact characteristic improves the user experience, but changes in environmental lighting conditions will significantly affect the image quality, resulting in insufficient recognition stability.
[0004] In recent years, the developed wireless signal recognition technology realizes recognition by analyzing the modulation characteristics of gesture actions on radio frequency signals, and has the advantages of both non-contact and strong environmental adaptability. However, this technology requires a large number of training samples to be collected for each gesture category, and the data collection process is affected by multiple factors such as signal interference and device layout, resulting in a high system deployment cost. In addition, for gesture types with similar motion trajectories, existing algorithms are difficult to effectively distinguish, which limits the reliable application of the technology in practical scenarios. Summary of the Invention
[0005] In view of this, the present invention provides a few-shot wireless gesture recognition method based on meta-actions. According to the characteristics of gesture motion trajectories, the meta-action gesture features are combined to obtain new gesture Doppler spectrogram features and trajectory curvature features. On the premise of ensuring the consistency of physical laws, the problem that wireless gesture recognition requires gesture data of different categories is overcome. By obtaining the Doppler spectrogram features and trajectory curvature features of meta-actions, the problem of difficult distinction of gestures with similar Doppler spectrogram features is effectively solved. After the gesture Doppler spectrogram features and gesture trajectory curvature features are weighted and fused, network training is carried out to achieve efficient wireless gesture recognition by collecting a small number of samples.
[0006] Therefore, the present invention provides the following technical solutions: The present invention provides a few-shot wireless gesture recognition method based on meta-actions, including: an offline training stage and an online recognition stage; in the offline training stage, millimeter-wave radar devices are used to collect meta-action gesture data to form a gesture action set, and each piece of meta-action gesture data in the gesture action set is preprocessed to eliminate the interference of meta-action gesture signal noise; for the preprocessed meta-action gesture signals, the first gesture Doppler spectrogram features and the first gesture trajectory curvature features are extracted; based on the first gesture Doppler spectrogram features and the first gesture trajectory curvature features, synthesis is performed to obtain the second gesture Doppler spectrogram features and the second gesture trajectory curvature features; a weighted fusion method is adopted to fuse the second gesture Doppler spectrogram features and the second gesture trajectory curvature features; the fusion features of each gesture in the gesture action set are sent into a perception network for training; in the online recognition stage, by collecting gesture data, a gesture sample to be recognized is constructed, the noise interference of the gesture signal to be recognized is removed, the range-time image of the gesture sample to be recognized is obtained, the Doppler spectrogram features and the trajectory curvature features of the gesture sample to be recognized are extracted, and the two features of the gesture are effectively fused by using a weighted fusion method and input into the trained perception network to achieve wireless gesture recognition.
[0007] Further, preprocessing each piece of meta-action gesture data in the gesture action set includes: obtaining raw radar data of multiple channels, and performing frame-by-frame processing on the radar data; performing time-domain windowing on the signal of each receiving channel within each frame, and eliminating the DC component and static interference in the signal by the method of removing the mean value.
[0008] Further, extracting the first gesture Doppler spectrogram features includes: based on the preprocessed meta-action gesture signals, performing fast Fourier transform on them in the range dimension to obtain the range-time image of the meta-action gesture; performing fast Fourier transform on the range-time image along the slow-time dimension to obtain the Doppler spectrogram of the meta-action gesture; by performing dimensionality reduction processing on the Doppler spectrogram of the meta-action gesture, converting the two-dimensional matrix into a one-dimensional vector, while retaining the Doppler spectrogram waveform signal of the gesture; traversing the obtained waveform signal to locate the first sign change point from negative to positive , and ensuring to meet a condition: there is a sign change point from positive to negative before the sign change point to determine the integrity of the first lower envelope of the curve; performing reverse order processing on the one-dimensional vector to find the sign change point ; restoring the data to the original order, and setting the Doppler data of the gesture before the point and after the point to zero, and only retaining the Doppler signal between the two points, so as to extract the first gesture Doppler spectrogram features.
[0009] Further, by performing dimensionality reduction on the Doppler spectrogram of the meta-action gesture, the two-dimensional matrix of is converted into a one-dimensional vector, while retaining the waveform signal of the Doppler spectrogram of the gesture, including: calculating the column sums of the first rows and the last rows of the two-dimensional matrix respectively to obtain and , ; selectively retaining the energy values by comparing the numerical magnitudes of and to obtain a one-dimensional vector; performing median filtering and smoothing on the one-dimensional vector to obtain a waveform signal reflecting the motion characteristics of the gesture.
[0010] Further, extracting the curvature feature of the first gesture trajectory, including: performing a fast Fourier transform on the meta-action gesture signal in the distance dimension to obtain the distance-time image of the meta-action gesture; analyzing the distance information of the gesture relative to the radar at a certain moment in the distance-time image, so as to obtain the distance values corresponding to the gesture at different time points, forming a position sequence : ; where represents the distance value of the meta-action gesture relative to the radar at time point ; using the position sequence to calculate the curvature of the meta-action gesture trajectory; Based on the obtained sign-changing points and sign-changing point on the time axis, extracting the curvature feature of the meta-action gesture, retaining the curvature data of the meta-action gesture trajectory between the two sign-changing points, and obtaining the curvature feature of the meta-action gesture trajectory.
[0011] Further, using the discrete points of the gesture trajectory, calculating the curvature using the following formula : ; where and represent the first derivative and the second derivative of the trajectory respectively, that is, the velocity and the acceleration; obtaining a curvature sequence vector with a length of through curvature calculation : ; where represents the curvature value of the meta-action gesture at time point .
[0012] Further, based on the first gesture Doppler spectrogram feature and the first gesture trajectory curvature feature, synthesis is performed to obtain the second gesture Doppler spectrogram feature and the second gesture trajectory curvature feature, including: performing time normalization on the first gesture Doppler spectrogram feature and the first gesture trajectory curvature feature; Generate the second gesture Doppler spectrogram feature, including: performing an acceleration process on the normalized first gesture Doppler spectrogram feature to simulate the speed change characteristics of different gestures in actual motion; synthesizing and splicing the Doppler spectrograms of the meta-action gestures along the time axis according to the composition order and motion characteristics of the target gesture; after completing the splicing of the Doppler spectrograms, performing time normalization processing on the generated second gesture Doppler spectrogram to make its time length consistent with that of the meta-action gesture, thereby obtaining the second gesture Doppler spectrogram feature; Generate the second gesture trajectory curvature feature, including: analyzing the motion trajectory of the meta-action gesture in space and extracting its key features; performing an acceleration process on the first gesture trajectory curvature feature by compressing the time axis of the gesture trajectory curvature map; combining multiple meta-action gesture trajectory curvature maps according to the time sequence and spatial relationship to form a second gesture trajectory curvature map, and performing normalization processing on it to make its time length consistent with that of the meta-action gesture, thereby obtaining the second gesture trajectory curvature feature.
[0013] Further, the input of the perception network is a two-dimensional feature matrix, and its label is the number of gesture categories. The difference between the model prediction and the true label is quantified by calculating the loss value, where the loss function is defined as: ; where represents the total number of categories of the labeled samples, represents the one-hot encoded label of the sample, represents the predicted output probability of the perception network; the gesture feature matrix passes through the convolutional layer Conv, activation layer ReLU, and fully connected layer Fc in the perception network respectively to train the perception network model.
[0014] Advantages and positive effects of the present invention: In the present invention, by collecting several groups of meta-action gestures, new gesture data is combined and generated to expand the gesture action set. By analyzing the time sequence and spatial relationship (such as connection points, motion directions) between the meta-action gestures, these meta-action gesture features are combined according to specific rules to form new complex gesture features, realizing the diversity of expanding the gesture action set. At the same time, each gesture is represented by Doppler spectrogram features and trajectory curvature features, and the network is trained and learned by weighted fusion of the two features of each gesture, achieving efficient wireless gesture recognition using a small number of samples. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 is the workflow diagram in the embodiments of the present invention; Figure 2 is the non-gesture signal generated by the raising and lowering hand actions in the embodiments of the present invention; Figure 3 is the gesture recognition network structure in the embodiments of the present invention; Figure 4 is the gesture expansion combination in the embodiments of the present invention. Detailed implementation manners
[0017] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0018] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0019] The present invention proposes a few-shot wireless gesture recognition method based on meta-actions. A meta-action refers to a simple basic gesture unit with general characteristics and is also the basic component that constitutes complex gestures. According to the characteristics of the gesture writing trajectory, it is found that some complex gestures can be synthesized through the combination and extension of meta-actions. Specifically, the principle of synthesizing gestures is based on the spatio-temporal relationship of meta-actions: First, a complex gesture is decomposed into multiple meta-action sequences, and each meta-action represents a basic motion unit of the gesture; Second, by analyzing the motion characteristics and spatial relationships between meta-actions, these meta-actions are combined according to specific rules to form a complete new gesture feature. For example, the motion characteristics and writing trajectory of the gesture "h" can be obtained by splicing the gesture "1" and the gesture "n". Specifically, the gesture "1" represents a straight-line motion vertically downward, and the gesture "n" represents a continuous curve motion. The two are combined through a connection point in space to form the complete motion characteristics and trajectory characteristics of the gesture "h".
[0020] The present invention comprehensively considers the extensibility of the Doppler spectrogram features and trajectory curvature features of gesture samples obtained based on wireless signals, uses several groups of meta-action gesture samples to combine to obtain new gesture data, and expands the gesture action set on the premise of ensuring the consistency of physical laws. The key of the present invention is that, first, by eliminating the influence of static interference and non-gesture signals generated by raising and lowering hands, the first gesture Doppler spectrogram features and the first gesture trajectory curvature features are extracted; Second, based on the principle of synthesizing new gestures by meta-actions, on the premise of ensuring the consistency of physical laws, the meta-action gesture features are combined to obtain the second gesture Doppler spectrogram features and gesture trajectory curvature features of a new category; Finally, the network model is trained by weighted fusion of the second gesture Doppler spectrogram features and the second gesture trajectory curvature features, and efficient wireless gesture recognition is realized using a small number of samples.
[0021] As Figure 1 shown, in an embodiment of the present invention, a few-shot wireless gesture recognition method based on meta-actions, the overall process is divided into two parts: offline training and online recognition.
[0022] In the offline training stage, first, meta-action gesture samples are collected and noise interference is removed to generate a Range-Time-Map (RTM), and then the first gesture Doppler spectrogram features and the first gesture trajectory curvature features are extracted; Second, based on the first gesture features, second gesture feature samples are synthesized to expand the gesture action set; Finally, the second gesture Doppler spectrogram features and the second gesture trajectory curvature features are fused to train the network model.
[0023] In the online recognition stage, gesture data is collected to construct a gesture sample to be recognized, noise interference of the gesture signal to be recognized is removed, the RTM image of the gesture sample to be recognized is obtained, two features of the gesture sample to be recognized are extracted, and after feature fusion, they are input into the trained network model to complete gesture recognition.
[0024] In this embodiment, the system configuration for implementing the few-shot wireless gesture recognition method is as follows: The system operates on a Texas Instrument IWR6843 FMCW radar transceiver device; the transceiver device operates at 60 GHz and the bandwidth is set to 3.06 GHz, capable of providing a range resolution of 4.9 cm; the transceiver device uses a 3-transmit and 4-receive antenna, and with the TDM-MIMO mode, it can form a 12-element virtual array, providing an angular resolution of 29°.
[0025] (1). Offline training for synthesizing new gestures based on meta-action gestures specifically includes the following steps: S11. Use a millimeter-wave radar device to collect meta-action gesture data, and after preprocessing, eliminate the interference of the meta-action gesture signal noise; In specific implementation, first, initialize the gesture recognition system based on preset radar parameters (including the number of sampling points, antenna configuration, range resolution, Doppler resolution, etc.); second, obtain radar raw data from multiple channels and process the radar data frame by frame; finally, perform time-domain windowing (such as Hamming window) on the signals of each receiving channel within each frame, and eliminate the DC component and static interference in the signal by the method of removing the mean value.
[0026] S12. For the preprocessed meta-action gesture signal, extract the first gesture Doppler spectrogram feature and the first gesture trajectory curvature feature of the meta-action; Among them, extracting the first gesture Doppler spectrogram feature includes: Based on the preprocessed meta-action gesture signal, first, perform a range fast Fourier transform (Range-FFT) on it to obtain the RTM image of the meta-action gesture; second, perform a fast Fourier transform (Doppler-FFT) on the RTM image along the slow time dimension to obtain the Doppler spectrogram of the meta-action gesture, and by capturing the frequency change characteristics of the gesture action, reflect the dynamic characteristics of the gesture.
[0027] Although the Doppler spectrogram based on meta-action gestures eliminates the influence of static interference, non-gesture signals such as raising and lowering the hand still exist when making gestures in the air. Since the Doppler spectrogram can intuitively reflect the change of the target speed over time, when the experimenter performs the hand-raising operation, the speed of the hand will first accelerate and then decelerate until the speed drops to zero, and at the same time the hand gradually approaches the radar; subsequently, the experimenter completes the gesture action in the air; after the gesture ends, when the experimenter lowers the hand, the speed of the hand will also first accelerate and then decelerate until the speed drops to zero, and at the same time the hand gradually moves away from the radar. As Figure 2 shown, the hand-raising action corresponds to the first lower envelope waveform in the Doppler spectrogram, while the hand-lowering action corresponds to the last upper envelope waveform. In order to remove the non-gesture action interference generated during the hand-raising and hand-lowering processes, it is necessary to find the end point of the first lower envelope waveform and the start point of the last upper envelope waveform in the gesture Doppler spectrogram, and extract the meta-action gesture Doppler spectrogram features by retaining the signals between these two points.
[0028] By performing dimensionality reduction on the Doppler spectrogram of the meta-action gesture, the two-dimensional matrix is converted into a one-dimensional vector while retaining the waveform information of the gesture Doppler spectrogram. First, the column sums of the first rows and the last rows of the two-dimensional matrix are obtained respectively to get and ( ), that is, the Doppler energy of each frame is summarized; secondly, by comparing the numerical sizes of and to selectively retain the energy values to obtain a one-dimensional vector; finally, by performing median filtering and smoothing on the one-dimensional vector, a waveform signal reflecting the gesture motion characteristics is obtained.
[0029] Traverse the obtained waveform signal. First, locate the first sign-changing point from negative to positive , and ensure that satisfies a condition: there is a sign-changing point from positive to negative before the sign-changing point to determine the integrity of the first lower envelope of the curve; secondly, reverse the order of the one-dimensional vector, and use the same method as above to find the sign-changing point ; finally, restore the data to the original order, and set the gesture Doppler data before the point and after the point to zero, and only retain the Doppler signals between the two points, so as to extract the first gesture Doppler spectrogram features.
[0030] Among them, extracting the curvature feature of the first gesture trajectory includes: performing a fast Fourier transform on the distance dimension of the meta-action gesture signal to obtain the RTM image of the meta-action gesture, which can visually reflect the energy distribution of the gesture changing with time in the distance dimension. By analyzing the distance information of the gesture relative to the radar at a certain moment in the RTM image, the distance values corresponding to the gesture at different time points are obtained, forming a position sequence : ; Among them, represents the distance value of the meta-action gesture relative to the radar at time point . Using the position sequences of these gestures at different time points, the curvature of the meta-action gesture trajectory is calculated.
[0031] The curvature calculation method is estimated based on the changes between consecutive points. Specifically, using the discrete points of the gesture trajectory, the following formula is used to calculate the curvature: ; Among them, and respectively represent the first derivative and the second derivative of the trajectory, which are the velocity and the acceleration.
[0032] Through curvature calculation, a curvature sequence vector with a length of is obtained : ; Among them, represents the curvature value of the meta-action gesture at time point .
[0033] Based on the positions of the obtained sign-changing points and the sign-changing point on the time axis, the curvature feature of the meta-action gesture is extracted, and the curvature data of the meta-action gesture trajectory between the two sign-changing points is retained to obtain the first gesture trajectory curvature feature.
[0034] S13. Synthesize based on the first gesture Doppler spectrogram feature and the first gesture trajectory curvature feature to obtain the second gesture Doppler spectrogram feature and the second gesture trajectory curvature feature; S111. Perform time normalization processing on the first gesture feature for synthesizing the second gesture; S112. Generate the second gesture Doppler spectrogram feature; By analyzing the motion characteristics of gestures, it is found that some gestures show the characteristic of moving in the opposite direction relative to the radar's motion trajectory, that is, their Doppler spectrograms are symmetric on the frequency axis. For example, by performing a frequency-axis flipping process on the Doppler spectrogram of the primitive gesture "n" collected, a Doppler spectrogram with mirror symmetry is obtained, that is, a Doppler spectrogram of the new gesture "u" is generated. Some complex gestures can be composed of combinations of simple primitive gesture features. For example, the gesture "m" can be composed of two sets of the gesture "n"; the gesture "h" can be composed of the gesture "1" and the gesture "n", etc. Based on the above findings, the specific method for generating the second gesture Doppler spectrogram feature is as follows: First, accelerate the normalized Doppler spectrogram feature of the primitive gesture to simulate the speed change characteristics of different gestures in actual motion. Then, according to the composition order and motion characteristics of the target gesture, splice the Doppler spectrograms of the primitive gestures along the time axis. For example, splicing the Doppler spectrograms of two sets of the gesture "n" along their motion trajectories can obtain the Doppler spectrogram of the new gesture "m"; splicing the Doppler spectrograms of the gesture "1" and the gesture "n" along their motion can obtain the Doppler spectrogram of the new gesture "h". After completing the splicing of the Doppler spectrograms, perform time normalization on the generated new gesture Doppler spectrogram to make its time length consistent with that of the primitive gesture, and obtain the second gesture Doppler spectrogram feature.
[0035] S113. Combine the primitive gesture trajectory curvature features according to the writing characteristics of the gesture in space to generate the second gesture trajectory curvature feature.
[0036] Specifically, first, analyze the motion trajectory of the primitive gesture in space and extract its key features; second, accelerate the primitive gesture trajectory curvature feature by compressing the time axis of the primitive gesture trajectory curvature map; finally, combine multiple primitive gesture trajectory curvature maps according to the time sequence and spatial relationship to form the second gesture trajectory curvature map, and perform normalization on it to make its time length consistent with that of the primitive gesture, and obtain the second gesture trajectory curvature feature.
[0037] S14. Use a weighted fusion method to fuse the generated second gesture Doppler spectrogram feature and the second gesture trajectory curvature feature; Construct a gesture action set by combining meta-action gestures and synthesized new gestures. Each gesture contains two key features: Doppler spectrogram feature and gesture trajectory curvature feature. The Doppler spectrogram feature is extracted through the Doppler effect of radar signals, reflecting the velocity distribution characteristics of gesture targets and effectively characterizing the dynamic motion patterns of gestures. The gesture trajectory curvature feature is obtained by calculating the curvature change of the gesture motion path, characterizing the spatial motion characteristics of gestures and effectively distinguishing gestures with similar Doppler features but different spatial trajectories. To make full use of the complementarity of these two features, in the embodiments of the present invention, a weighted fusion method is adopted to fuse the gesture Doppler spectrogram feature and the trajectory curvature feature. By assigning corresponding weight coefficients to each feature and adjusting the weight ratio in combination with the motion characteristics of the gesture, weighted fusion of the features is achieved.
[0038] S15. Send the fusion features of each gesture in the gesture action set into the perception network for training; Among them, the input of the network is a two-dimensional feature matrix, and its label is the number of gesture categories. The difference between the model prediction and the true label is quantified by calculating the loss value. The loss function is defined as: ; where represents the total number of categories of labeled samples, represents the one-hot encoded label of the sample, represents the predicted output probability of the perception network. As Figure 3 shown, the gesture feature matrix passes through the convolutional layer Conv, the activation layer ReLU, and the fully connected layer Fc respectively to train the perception network model.
[0039] (2). Use the trained perception network for online gesture recognition, which specifically includes the following steps: S21. Construct a gesture sample to be recognized by actually collecting gesture data; S22. Remove the noise interference of the gesture signal to be recognized, obtain the RTM image of the gesture sample to be recognized, and extract the Doppler spectrogram feature and the trajectory curvature feature of the gesture sample to be recognized; S23. Adopt a weighted fusion method to effectively fuse the two features of the gesture and input them into the perception network to achieve wireless gesture recognition.
[0040] For ease of understanding, the above-mentioned few-shot wireless gesture recognition method based on meta-action is described below with a specific example.
[0041] A few-shot wireless gesture recognition method based on meta-action includes the following steps: S31. Preprocessing of meta-action gesture signals: Based on the writing trajectory characteristics of numbers and letters, four primitive motion gestures, namely "1", "c", "3", and "n", are selected for synthesizing new gestures. The primitive motion gesture data for 4 seconds is collected by a radar device, with the number of sampling points being 64 and the range resolution being 4.9 cm. In this embodiment, the collected radar data is processed frame by frame, 60 frames of data are retained, and the direct current component and static interference in the signal are eliminated by the method of removing the mean value.
[0042] S32. Extract the Doppler spectrogram features of the first gesture: Perform a fast Fourier transform on the preprocessed primitive motion gesture data in the range dimension to obtain the RTM image of the primitive motion gesture, and then perform a fast Fourier transform on it along the slow time dimension to obtain the Doppler spectrogram of the primitive motion gesture. Perform dimensionality reduction processing on the Doppler spectrogram of the primitive motion gesture signal, reduce the two-dimensional Doppler matrix to a one-dimensional vector, and at the same time retain the key waveform information. In this embodiment, 1 second corresponds to 15 frames, 1 frame corresponds to 96 chirps, and the signal for 4 seconds is retained, that is, 60 frames of data are collected. Therefore, the size of the gesture Doppler matrix is 96×60. By calculating the column sum of the first 48 rows of the two-dimensional matrix, and calculating the column sum of the last 48 rows to obtain Extract effective features by comparing the front and back parts, that is, compare and If , it indicates that the main energy is concentrated in the positive frequency band. Retain the value of and take its opposite number. If , it indicates that the main energy is concentrated in the negative frequency band. Retain the value of to obtain a one-dimensional vector. After smoothing the obtained one-dimensional vector, display its curve features through an image. Traverse the data of the one-dimensional vector to locate the first sign change point from negative to positive, and at the same time ensure that there is a sign change point from positive to negative before this sign change point to confirm the integrity of the first lower envelope in the curve. Mark this sign change point as the starting point T of the gesture signal. Subsequently, reverse the curve image of the one-dimensional vector and use the same method to locate the sign change point and mark the ending point F of the gesture signal. Finally, reverse the data back to the original order to complete the accurate marking of the starting point T and ending point F of the gesture signal, and then extract the Doppler spectrogram features of the first gesture between the two points.
[0043] S33. Extract the trajectory curvature features of the first gesture: Save the obtained RTM image obtained by performing fast Fourier transform on the distance dimension based on the meta-action gesture data. By analyzing the RTM image of the meta-action gesture in detail, key features that can reflect the gesture trajectory are extracted from the image. Based on these features, the curvature information of the meta-action gesture is further calculated. Specifically, by analyzing the distance information of the gesture relative to the radar at each moment in the RTM image, the distance values corresponding to the gesture at different time points are obtained. In this embodiment, a total of 60 frames of distance information of the meta-action gesture are extracted to construct a complete position sequence: ; where represents the frame in the distance value of the meta-action gesture relative to the radar. Based on the position sequence of the meta-action gesture at different time points, the curvature of the meta-action gesture movement trajectory is calculated. The calculation of the curvature depends on the analysis of the change relationship between consecutive points. Specifically, through the discrete point data of the gesture trajectory, the curvature is calculated using the following formula: ; where and respectively represent the velocity and acceleration of the position of the meta-action gesture trajectory, and then a curvature sequence vector with a length of 60 is obtained, and each element is the curvature value at this time point: ; where represents the frame in the curvature value of the meta-action gesture trajectory, then represents the meta-action gesture trajectory curvature vector. Using the starting point and the ending point of the meta-action gesture signal at the position on the time axis, the influence of non-gesture signals generated by the raising and lowering of the hand is eliminated, and the curvature characteristics of the meta-action gesture trajectory are extracted.
[0044] S34. Combine the first gesture features to obtain the second gesture features: Perform time normalization processing on the extracted first gesture features to unify them to the same time scale (4 seconds) to ensure the time consistency of the data. According to the motion characteristics of the gesture, accelerate the Doppler spectrogram features of the meta-action gesture to simulate the speed change characteristics of different gestures in actual motion, and combine the Doppler spectrograms of the meta-action gesture in sequence to generate new gesture data, such as Figure 4As shown, a new gesture "u" is obtained by flipping the Doppler spectrogram of the meta-action gesture "n", and a new gesture "1 (up)" is obtained by flipping the Doppler spectrogram of the meta-action gesture "1", where the "1 (up)" gesture is used to assist in synthesizing new gestures. By accelerating and combining the Doppler spectrograms of the meta-action gestures "1" and "n", a new gesture "h" is obtained; by accelerating and combining the Doppler spectrograms of the meta-action gestures "c" and "1 (up)", a new gesture "d" is obtained; by accelerating and combining the Doppler spectrograms of the meta-action gestures "1 (up)" and "3", a new gesture "B" is obtained; and by accelerating and combining the Doppler spectrograms of the meta-action gestures "n" and "n", a new gesture "m" is obtained. Similarly, by analyzing the spatial motion trajectories of the meta-action gestures and combining the trajectory curvature characteristics of the meta-action gestures, a second gesture trajectory curvature characteristic is obtained. In this experiment, only 4 types of meta-action gesture data need to be collected to generate 9 types of gesture data (including the 4 collected meta-action gestures and the 5 synthesized new gestures), effectively expanding the gesture action set.
[0045] S35. Fuse two features of the gesture: Based on these two features, namely the second gesture Doppler spectrogram and the second gesture trajectory curvature, weighted fusion is performed. The gesture Doppler spectrogram feature is two-dimensional matrix data, and the gesture trajectory curvature feature is a one-dimensional vector: 1) Doppler spectrogram feature matrix , where is the number of time frames (i.e., 60 frames), is the number of chirps per frame (i.e., 96), which is the number of frequency units obtained after fast Fourier transform; 2) Gesture trajectory curvature feature vector , where is the number of time frames (i.e., 60 frames).
[0046] First, the gesture trajectory curvature feature vector is extended to have the same dimension as the Doppler spectrogram feature matrix, and then a two-dimensional matrix is obtained. The extended curvature matrix and the Doppler spectrogram matrix are weighted and fused: ; where, is the fused feature matrix, and are the weight coefficients of the Doppler feature and the curvature feature respectively, is 0.7, is 0.3, where the dot product means multiplying the matrix elements term by term.
[0047] S36. Train the network model: Use the gesture feature matrix after weighted fusion as the input of the perception network. There are a total of 9 gestures in the gesture dataset (4 meta-action gestures and 5 synthesized new gestures), and its label is the number of gesture categories. Calculate the loss value to quantify the difference between the model prediction and the true label. The loss function is defined as follows: ; where represents the total number of categories of the labeled samples, represents the one-hot encoded label of the sample, represents the predicted output probability of the perception network. Pass the gesture data through the convolutional layer Conv, activation layer ReLU, and fully connected layer Fc respectively, and finally output the training result.
[0048] S37. Online gesture recognition: Construct a gesture dataset to be recognized by actually collecting the nine gestures of "1", "c", "n", "3", "B", "d", "h", "m", and "u". Remove the noise interference of the gesture signal to be recognized, obtain the RTM image of the gesture sample to be recognized, extract the Doppler spectrogram features and trajectory curvature features of the gesture sample to be recognized, fuse the two features, and input them into the trained perception network for gesture recognition.
[0049] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A few-shot wireless gesture recognition method based on meta-actions, characterized in that Including: An offline training stage and an online recognition stage; In the offline training stage, use a millimeter-wave radar device to collect primitive action gesture data, form a gesture action set, and preprocess each piece of primitive action gesture data in the gesture action set to eliminate the interference of primitive action gesture signal noise; For the preprocessed primitive action gesture signals, extract the first gesture Doppler spectrogram features and the first gesture trajectory curvature features; Based on the first gesture Doppler spectrogram features and the first gesture trajectory curvature features, perform synthesis to obtain the second gesture Doppler spectrogram features and the second gesture trajectory curvature features; Adopt a weighted fusion method to fuse the second gesture Doppler spectrogram features and the second gesture trajectory curvature features; Send the fusion features of each gesture in the gesture action set into a perception network for training; In the online recognition stage, by collecting gesture data, construct a gesture sample to be recognized, remove the noise interference of the gesture signal to be recognized, obtain the range-time image of the gesture sample to be recognized, extract the Doppler spectrogram features and the trajectory curvature features of the gesture sample to be recognized, and use a weighted fusion method to effectively fuse the two features of the gesture, and input them into the trained perception network to achieve wireless gesture recognition.
2. The few-shot wireless gesture recognition method based on meta-action according to claim 1, characterized in that Preprocess each piece of primitive action gesture data in the gesture action set, including: Obtain the raw radar data of multiple channels and process the radar data frame by frame; Perform time-domain windowing on the signal of each receiving channel within each frame, and eliminate the DC component and static interference in the signal by the method of removing the mean value.
3. A few-shot wireless gesture recognition method based on meta-actions according to claim 1, characterized in that, Extract the first gesture Doppler spectrogram features, including: Based on the preprocessed primitive action gesture signals, perform fast Fourier transform on the range dimension to obtain the range-time image of the primitive action gesture; perform fast Fourier transform on the range-time image along the slow-time dimension to obtain the Doppler spectrogram of the primitive action gesture; By performing dimensionality reduction on the Doppler spectrogram of the meta-action gesture, convert the two-dimensional matrix of into a one-dimensional vector while retaining the waveform signal of the Doppler spectrogram of the gesture; Traverse the obtained waveform signal to locate the first sign change point from negative to positive , and ensure that meets a condition: there is a sign change point from positive to negative before the sign change point to determine the integrity of the first lower envelope of the curve; reverse the one-dimensional vector to find the sign change point ; restore the data to the original order, and set the gesture Doppler data before the point and after the point to zero, and only retain the Doppler signal between the two points, so as to extract the first gesture Doppler spectrogram feature.
4. A few-shot wireless gesture recognition method based on meta-actions according to claim 3, characterized in that, By performing dimensionality reduction on the Doppler spectrogram of the meta-action gesture, the two-dimensional matrix of is converted into a one-dimensional vector, while retaining the waveform signal of the Doppler spectrogram of the gesture, including: For the first rows and the last rows of the two-dimensional matrix, calculate the column sums respectively to obtain and , ; By comparing and the numerical magnitudes of, selectively retain the energy values to obtain a one-dimensional vector; Through median filtering and smoothing processing on the one-dimensional vector, obtain a waveform signal reflecting the motion characteristics of the gesture.
5. The few-shot wireless gesture recognition method based on meta-action according to claim 3, characterized in that Extract the first gesture trajectory curvature features, including: Perform fast Fourier transform on the range dimension of the primitive action gesture signal to obtain the range-time image of the primitive action gesture; By analyzing the distance information of the gesture relative to the radar at a certain moment in the distance-time image, the distance values corresponding to the gesture at different time points are obtained, forming a position sequence : ; Among them, represents the distance value of the meta-action gesture relative to the radar at time point ; Using the position sequence , calculate the curvature of the meta-action gesture trajectory; Based on the obtained sign-changing points and the sign-changing points at their positions on the time axis, extract the curvature features of the primitive action gesture, retain the curvature data of the primitive action gesture trajectory between the two sign-changing points, and obtain the curvature features of the primitive action gesture trajectory.
6. A few-shot wireless gesture recognition method based on meta-actions according to claim 5, characterized in that, Using the discrete points of the gesture trajectory, calculate the curvature using the following formula : ; where and respectively represent the first derivative and the second derivative of the trajectory, which are the velocity and the acceleration A curvature sequence vector of length is obtained through curvature calculation : ; where represents the curvature value of the meta-action gesture at time point .
7. A few-shot wireless gesture recognition method based on meta-actions according to claim 1, characterized in that, Based on the first gesture Doppler spectrogram features and the first gesture trajectory curvature features, perform synthesis to obtain the second gesture Doppler spectrogram features and the second gesture trajectory curvature features, including: Perform time normalization processing on the first gesture Doppler spectrogram features and the first gesture trajectory curvature features; Generate the second gesture Doppler spectrogram features, including: perform acceleration processing on the normalized first gesture Doppler spectrogram features to simulate the speed change characteristics of different gestures in actual motion; according to the composition order and motion characteristics of the target gesture, synthesize and splice the Doppler spectrograms of the primitive action gestures along the time axis; after completing the splicing of the Doppler spectrograms, perform time normalization processing on the generated second gesture Doppler spectrogram to make its time length consistent with the primitive action gesture, and obtain the second gesture Doppler spectrogram features; Generate the second gesture trajectory curvature feature, including: analyzing the motion trajectory of the meta-action gesture in space and extracting its key features; accelerating the first gesture trajectory curvature feature by compressing the time axis of the gesture trajectory curvature graph; combining multiple meta-action gesture trajectory curvature graphs according to the time sequence and spatial relationship to form the second gesture trajectory curvature graph, and normalizing it to make its time length consistent with the meta-action gesture, thereby obtaining the second gesture trajectory curvature feature.
8. A few-shot wireless gesture recognition method based on meta-actions according to any one of claims 1 to 7, characterized in that, The input of the perception network is a two-dimensional feature matrix, whose label is the number of gesture categories. The difference between the model prediction and the true label is quantified by calculating the loss value, where the loss function is defined as: ; where represents the total number of categories of labeled samples, represents the one-hot encoded label of the sample, represents the predicted output probability of the perception network; the gesture feature matrix passes through the convolutional layer, activation layer and fully connected layer in the perception network respectively to train the perception network model.
Citation Information
Patent Citations
Gesture recognition method based on millimeter-wave radar
CN111476058A
Multi-scale feature fusion gesture recognition method based on FMCW millimeter wave radar
CN113837131A
Millimeter wave radar gesture recognition method and system based on multi-head self-attention mechanism
CN115877376A
Millimeter wave radar gesture recognition method and system based on multi-domain spectrogram and multi-resolution fusion
CN116184394A
Millimeter wave radar gesture recognition method based on multi-dimensional dense gesture spectrogram
CN118731877A
Cited By
Non-contact personnel identity and posture recognition method based on motion instability compensation
CN122110050A