Dynamic gesture recognition method of multi-feature fusion network under complex scene interference

Through a multi-feature fusion network, combined with the DSCC module and the Transformer module, the problem of insufficient feature extraction and high model complexity of FMCW radar gesture recognition in complex scenarios is solved, and high accuracy and robust gesture recognition is achieved to adapt to multipath effect and complex background interference.

CN120496183AActive Publication Date: 2025-08-15GUANGDONG UNIV OF PETROCHEMICAL TECH

Patent Information

Application Number
CN202510595762.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing gesture recognition method based on FMCW radar has problems such as insufficient feature extraction, high model complexity, and weak generalization ability in complex scenarios, and it is difficult to adapt to multipath effect and complex background interference, resulting in insufficient recognition accuracy and robustness.

Method used

Using a multi-feature fusion network, by obtaining distance-time graphs and Doppler-time graphs, combining lightweight DSCC modules and Transformer modules, the spatial displacement characteristics and velocity change information of gestures are extracted and integrated, and the channel attention model is used to enhance feature expression, realizing global feature extraction and classification.

Benefits of technology

It significantly improves the accuracy and robustness of gesture recognition, enhances the adaptability and generalization ability of the model in complex environments, and improves the recognition effect, especially in four non-ideal environments, with an identification accuracy of 97.40%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496183A_ABST
    Figure CN120496183A_ABST
Patent Text Reader

Abstract

The invention relates to a dynamic gesture recognition method for a multi-feature fusion network under complex scene interference, and the method comprises the steps: obtaining a baseband intermediate frequency signal, carrying out the processing of the baseband intermediate frequency signal, and obtaining a distance-time diagram and a Doppler-time diagram; inputting the distance-time graph and the Doppler-time graph into a gesture recognition model to obtain a gesture recognition result; the gesture recognition model is obtained by training a lightweight neural network model by using a training set; and extracting space displacement features and gesture speed change information of the distance-time chart and the Doppler-time chart based on a DSCC module in the gesture recognition model, obtaining a feature map, combining the feature map with position codes, inputting the feature map into a Transform module for global feature extraction, and obtaining a gesture recognition result. According to the method, through extraction and fusion of multiple features, the distance, speed and time information of the target gesture is expressed more comprehensively, and the richness and expression ability of the gesture features in a complex scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of wireless communication and artificial intelligence technology, and in particular to a dynamic gesture recognition method using a multi-feature fusion network under interference from complex scenes. Background Art

[0002] In recent years, driven by wireless communications and artificial intelligence (AI), human-computer interaction (HCI) has become deeply integrated into everyday life, becoming one of the key technologies underpinning the development of modern society. As a core research area in HCI, hand gesture recognition (HGR) has garnered widespread attention due to its intuitive and natural interaction. This technology has achieved breakthroughs in a variety of application areas, including intelligent driving, smart homes, assistance for the disabled, and sign language interpretation, demonstrating its enormous application value and development potential.

[0003] Currently, gesture recognition systems fall into four main categories: those based on wearable sensors, computer vision, Wi-Fi, and FMCW radar. Wearable sensor-based systems rely on data gloves, which capture hand motion information using sensors such as accelerometers and gyroscopes. For example, in 2016, Zheng Y and his research team innovatively combined bending sensors with resistive force sensors to successfully develop a functional sensor glove, providing a method to visualize and quantify coordination abnormalities between joints. Kanokoda et al. utilized artificial neural networks for real-time gesture prediction. In 2019, Glauser et al. addressed the spatial resolution limitations of traditional sensors and proposed an innovative design based on stretchable array sensors. This design, using a dense array of regional stretchable sensors, enables comprehensive capture of hand surface deformation. However, these systems suffer from issues such as fragility, limited functionality, high cost, and inconvenience in wearing. Computer vision-based systems collect images using imaging devices such as cameras and extract hand features from them. Wu J et al. proposed an innovative dual-branch convolutional neural network architecture that improves the accuracy and robustness of gesture recognition by fusing features from depth and optical flow data. In 2018, MR Islam et al. used a deep convolutional neural network (DCNN) and a multi-class support vector machine (SVM) to recognize gestures for 26 alphabetic symbols. In 2019, Nguyen's team designed a neural network model based on positive matrix manifold learning and applied it to gesture recognition using human skeleton data collected by depth sensing cameras. However, these systems are susceptible to lighting and visual blind spots, and are not conducive to user privacy protection. Wi-Fi-based gesture recognition systems use Wi-Fi signals to recognize human gestures in other applications. In 2013, Pu et al. designed the WiSee system, which uses Doppler shift to sense gestures in home environments. In 2021, Wang et al. proposed a WiDG system based on CSI and deep learning, capable of recognizing the numbers 0-9 with an accuracy of 97.2% and 95.3% in both through-wall and wall-free scenarios, respectively. However, the limited detection range of Wi-Fi systems limits their widespread application. FMCW radar-based systems use radar technology to collect gesture signals and classify and recognize them using machine learning or deep learning methods. In 2019, Choi et al. used Google's 60GHz FMCW radar (Soli) combined with a long short-term memory (LSTM) network to successfully recognize 10 gestures with an accuracy of 99.10%, and the accuracy of gesture recognition for new participants was 98.48%. In 2021, a study introduced a reusable LSTM (RLSTM) network based on range-Doppler angle trajectory, using a 77GHz FMCW MIMO radar, achieving an average accuracy of 99%.FMCW radar offers advantages such as small size, light weight, strong anti-interference capabilities, and high resolution. Operating in a contactless mode, it is unaffected by line-of-sight, lighting, and inclement weather, significantly improving user experience and privacy. In summary, gesture recognition systems based on FMCW radar excel in performance, applicability, user experience, and privacy. They can simultaneously acquire multi-dimensional motion parameters such as target distance, speed, and angle through a fixed-slope frequency-modulated signal, providing richer information dimensions for gesture recognition. This system offers significant advantages over wearable sensors, computer vision, and Wi-Fi systems, and holds significant potential for development in intelligent technology and human-computer interaction.

[0004] A review of existing FMCW radar-based gesture recognition methods has revealed some promising results in the field. However, numerous challenges remain in feature extraction and gesture classification. First, when selecting gesture features as model input, some methods rely solely on micro-Doppler features, extract three-dimensional features based on range, Doppler, and angle, or extract features based on range, Doppler, and azimuth angle separately. However, these single features are susceptible to subjective influences and interference, and their signal representation capabilities are insufficient, making them inadequate for more complex scenarios. Second, while existing methods have achieved significant results in classification accuracy, they still face several key challenges in millimeter-wave radar-based gesture recognition. Most recognition systems rely on traditional convolutional modules to construct network models. While effective, this architecture leads to exponential growth in model complexity and parameter count as convolutional neural networks continue to deepen. Furthermore, as network depth increases, frequent downsampling often loses significant detail, potentially negatively impacting classification accuracy. Finally, if a dataset is collected only in a single environment, it lacks environmental diversity, resulting in weak generalization of the trained model. It also fails to fully reflect the multipath effects and complex background interference found in real-world scenarios, limiting the model's robustness. Such a dataset cannot fully validate the model's adaptability in real-world applications, potentially leading to a significant gap between experimental performance and real-world performance. Summary of the Invention

[0005] To address the aforementioned issues in the existing technology, this invention proposes a dynamic gesture recognition method using a multi-feature fusion network in complex scene interference. By processing radar signals, the RTM and DTM of gestures are extracted and fused, improving feature richness and signal expressiveness. Multi-scene data collection captures the specificity of each scene. To overcome the low model efficiency, a lightweight neural network model, Multi-DSCC-Transformer, is designed.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] A dynamic gesture recognition method using a multi-feature fusion network under complex scene interference includes:

[0008] Acquire a baseband intermediate frequency signal, process the baseband intermediate frequency signal, and acquire a range-time graph and a Doppler-time graph;

[0009] The distance-time graph and the Doppler-time graph are input into a gesture recognition model to obtain a gesture recognition result. The gesture recognition model is obtained by training a lightweight neural network model using a training set, where the training set includes: a distance-time spectrogram, a Doppler-time spectrogram, and a gesture label. The spatial displacement features and gesture velocity change information of the distance-time graph and the Doppler-time graph are extracted based on a DSCC module in the gesture recognition model to obtain a feature graph. The feature graph is combined with a position code and input into a Transformer module for global feature extraction to obtain the gesture recognition result.

[0010] Optionally, obtaining the baseband intermediate frequency signal includes:

[0011] The baseband intermediate frequency signal is generated using an FMCW radar device:

[0012] Acquire a detection signal whose frequency changes linearly with time, split the detection signal, amplify one detection signal and radiate it into space to detect the target, and use the other detection signal as a reference signal;

[0013] When the radiation signal detects the target, the signal is returned and mixed with the reference signal to generate an intermediate frequency signal containing target information. The intermediate frequency signal is filtered to obtain the baseband intermediate frequency signal.

[0014] Optionally, acquiring the distance-time graph and the Doppler-time graph includes:

[0015] Rearranging the baseband intermediate frequency signal into a two-dimensional data matrix; wherein the horizontal axis of the two-dimensional data matrix represents a slow time dimension, and the vertical axis represents a fast time dimension;

[0016] Performing a fast Fourier transform on the fast time dimension to obtain a distance-time matrix, and slightly scaling the distance-time matrix to obtain the distance-time graph;

[0017] Performing short-time Fourier transform on the slow time dimension, coherently superimposing the short-time Fourier transform results of each interval to obtain the Doppler-time map.

[0018] Optionally, the gesture recognition model includes:

[0019] A DSCC module is configured to extract spatial displacement features and gesture velocity change information from the distance-time graph and the Doppler-time graph to generate the feature graph; the feature graph includes a first feature graph and a second feature graph;

[0020] A position encoding module, configured to perform weighted feature fusion on the first feature map and the second feature map, and introduce position encoding;

[0021] Transformer module, used to extract global features of dynamic gestures using position encoding results;

[0022] The fully connected module is used to process the global features, perform gesture classification, and obtain the gesture recognition result.

[0023] Optionally, the DSCC module includes: two parallel convolution submodules, for respectively characterizing the spatial displacement characteristics of gesture motion and capturing gesture speed change information;

[0024] The convolution submodule includes: a depthwise separable convolution layer, the depthwise separable convolution layer and a pooling layer arranged alternately, the pooling layer at the end is sequentially connected to multiple standard convolution layers, and the standard convolution layer at the end is connected to a CAM unit; wherein, a batch normalization layer and a nonlinear layer are added after each convolution layer;

[0025] The depthwise separable convolutional layer is used to extract features from the range-time map and the Doppler-time map;

[0026] The pooling layer is used to reduce the spatial dimension of the output result of the depth-wise separable convolutional layer;

[0027] The standard convolution layer is used to perform secondary feature extraction on the output result of the pooling layer;

[0028] The CAM unit is used to perform channel attention enhancement on the output result of the standard convolutional layer.

[0029] Optionally, the Transformer module includes: a plurality of encoding submodules for extracting global features of dynamic gestures using position encoding results;

[0030] The encoding submodule includes: a dropout layer, a layer normalization layer, a multi-head self-attention layer, a feedforward neural network layer and a normalization layer connected in sequence; wherein, a residual connection is added after the multi-head self-attention layer and the feedforward neural network layer;

[0031] The multi-head self-attention layer is used to perform a linear transformation on the normalized processing result to generate a query vector, a key vector and a value vector, perform a dot product on the query vector and the key vector, and calculate the attention weight.

[0032] Optionally, generating the query vector, key vector, and value vector includes:

[0033]

[0034] Among them, Q is the query vector, K is the key vector, V is the value vector, T is the transpose of the vector, and d k is the dimension of the key vector and query vector.

[0035] Optionally, calculating the attention weight includes:

[0036]

[0037] Among them, Weigh ts is the weight.

[0038] Optionally, the fully connected module includes:

[0039] An adaptive average pooling layer for reducing the dimension of the global features;

[0040] The fully connected layer is used to map the global features after dimensionality reduction to the label space to obtain the gesture recognition result.

[0041] The beneficial effects of the present invention are:

[0042] The present invention designs a lightweight gesture recognition model Multi-DSCC-Transformer, which adopts a depth-separable convolutional structure to greatly reduce the complexity of the model. Two identical DSCC modules extract gesture features from the distance-time and Doppler-time feature maps respectively, and fuse the features to increase the diversity of gesture information. On this basis, the channel attention model CAM is used to effectively learn the important parts of the features, and the Transformer network is further introduced to extract the global features of dynamic gestures. By extracting and fusing multiple features, the present invention more comprehensively expresses the distance, speed and time information of the target gesture, improves the richness and expressiveness of gesture features in complex scenes, enhances the robustness of the model, makes it more stable in complex environments, and thus enhances the recognition effect.

[0043] The present invention can significantly improve the performance and practicality of the model. Experimental data are collected in four non-ideal environments with various interferences, noise and debris. These environments can effectively simulate real environments, which not only enriches the diversity of data samples but also increases the learning difficulty of the model, thereby enhancing its adaptability and generalization ability in complex environments. On the other hand, it also increases the learning difficulty of the model for gesture recognition, thereby improving the adaptability and generalization ability of the model in complex scenarios, laying a solid foundation for the deployment and promotion of gesture recognition systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 This is a block diagram of an FMCW radar system according to an embodiment of the present invention;

[0046] Figure 2 This is a flow chart of raw echo data preprocessing according to an embodiment of the present invention;

[0047] Figure 3 This is a diagram of the DSCC-transformer network structure according to an embodiment of the present invention;

[0048] Figure 4 This is a structural diagram of the channel attention mechanism (CAM) according to an embodiment of the present invention;

[0049] Figure 5 DSCC block structure diagram of an embodiment of the present invention;

[0050] Figure 6 This is a Transformer structure diagram of an embodiment of the present invention;

[0051] Figure 7 Schematic diagram of experimental gesture design for an embodiment of the present invention; (a) represents left-right swing, (b) represents right-left swing, (c) represents push, (d) represents pull, (e) represents clockwise rotation, (f) represents counterclockwise rotation, (g) represents double-click, and (h) represents applause;

[0052] Figure 8 Schematic diagram of the data collection environment of an embodiment of the present invention; wherein (a) represents different collection environments, and (b) represents different collection locations;

[0053] Figure 9 Schematic diagrams of confusion matrices for various features of an embodiment of the present invention; (a) is the confusion matrix of RTM, (b) is the confusion matrix of DTM, and (c) is the confusion matrix of the fusion of RTM and DTM;

[0054] Figure 10 : The model loss and validation curves of the embodiment of the present invention; wherein (a) is the validation accuracy curve, and (b) is the training loss curve;

[0055] Figure 11 Schematic diagram of gesture accuracy in different scenarios according to an embodiment of the present invention;

[0056] Figure 12 Schematic diagram of the accuracy of different distances and angles according to an embodiment of the present invention. DETAILED DESCRIPTION

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0058] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0059] Against the backdrop of today's rapid technological advancements, millimeter-wave radar is increasingly being used in complex scenarios. However, the numerous interference factors in complex scenarios and the inherently complex nature of the models present unprecedented challenges to the practical implementation of gesture recognition methods. To effectively address this challenge, this paper proposes a dynamic gesture recognition method in complex scene interference using a multi-feature fusion network. This algorithm collects data from four interfering environments to generate a self-constructed dataset, significantly increasing the complexity and richness of the data. Furthermore, the use of advanced clutter filtering (MTI) significantly improves the signal-to-noise ratio (SNR). Furthermore, multiple key modalities, such as the range-time map (RTM) and Doppler-time map (DTM), are successfully acquired. Modal fusion further enhances the richness and expressiveness of the features. In terms of gesture recognition model construction, this paper proposes a lightweight neural network model Multi-DSCC-Transformer: the DSCC module (Depthwise Separable Convolution with CAM Channel Attention Mechanism) uses depthwise separable convolution to reduce the size of model parameters, and uses CAM (channel attention module) to enhance the feature expression of each channel to improve the feature expression ability of the convolutional neural network. The Transformer module captures the global features of dynamic gestures through superposition and uses a multi-head attention mechanism to focus on basic features. A large number of experimental results show that the gesture recognition algorithm based on multimodal fusion combined with the Multi-DSCC-Transformer model exhibits excellent generalization and adaptability, and has extremely strong robustness even in complex scenarios. Specifically, this method achieves a recognition accuracy of up to 97.40% for 8 common gestures, providing a highly promising solution for gesture recognition applications of millimeter-wave radar in complex scenarios.

[0060] This embodiment discloses a method for dynamic gesture recognition using a multi-feature fusion network under complex scene interference, including: obtaining a baseband intermediate frequency signal, processing the baseband intermediate frequency signal to obtain a range-time graph and a Doppler-time graph; inputting the range-time graph and the Doppler-time graph into a gesture recognition model to obtain a gesture recognition result; the gesture recognition model is obtained by training a lightweight neural network model using a training set, the training set including a range-time spectrogram, a Doppler-time spectrogram, and gesture labels; extracting spatial displacement features and gesture velocity change information from the range-time graph and the Doppler-time graph using a DSCC module in the gesture recognition model to obtain a feature graph; and inputting the feature graph into a Transformer module in combination with a position code for global feature extraction to obtain a gesture recognition result.

[0061] Furthermore, obtaining the baseband intermediate frequency signal includes: generating the baseband intermediate frequency signal using an FMCW radar device; obtaining a detection signal whose frequency changes linearly with time, splitting the detection signal, power-amplifying one detection signal and radiating it into space to detect the target, and using the other detection signal as a reference signal; when the radiated signal detects the target, the signal is returned and mixed with the reference signal to generate an intermediate frequency signal containing target information, filtering the intermediate frequency signal, and obtaining the baseband intermediate frequency signal.

[0062] Specifically, a simplified block diagram of the FMCW radar used in the present invention is shown in FIG. Figure 1 As shown in the figure, the system primarily consists of a transmitter, a receiver, an analog-to-digital converter, and an antenna. Its basic operating principle is that a voltage-controlled oscillator generates a signal with a frequency that varies linearly with time. This signal is split into two paths by a power divider: one path is amplified by a power amplifier and radiated into space by the transmitting antenna for target detection; the other path is retained as a reference signal. When the transmitted signal encounters a target, some of its energy is reflected back and captured by the receiving antenna. The reflected echo signal is initially amplified by a low-noise amplifier to improve signal quality. The amplified echo signal is then mixed with the reference signal to generate an intermediate frequency (IF) signal containing target information. The mixed signal is filtered to extract the baseband IF signal, which contains the target's range and velocity. By analyzing and processing this IF signal, the target's range and relative velocity can be accurately calculated, enabling target detection and tracking.

[0063] For FMCW radar, the transmitted signal can be expressed as:

[0064]

[0065] Where: A tx is the amplitude of the transmitted signal, f c is the carrier frequency (starting frequency), α is the slope of frequency modulation (frequency change rate), t is time, is the initial phase.

[0066] The received signal is:

[0067]

[0068] Where: A rx is the amplitude of the received signal, which is usually smaller than the amplitude of the transmitted signal, and τ is the time delay of the signal traveling to and from the target. Where R is the target distance and c is the speed of light.

[0069] The intermediate frequency signal formed by combining the transmit signal and the receive signal can be expressed as:

[0070]

[0071] Where: A if is the amplitude of the intermediate frequency signal, f if It is a medium frequency, usually determined by the distance and speed of the target. is the phase of the intermediate frequency signal.

[0072] Furthermore, obtaining the distance-time diagram and the Doppler-time diagram includes: rearranging the baseband intermediate frequency signal into a two-dimensional data matrix; wherein the horizontal axis of the two-dimensional data matrix represents the slow time dimension, and the vertical axis represents the fast time dimension; performing a fast Fourier transform on the fast time dimension to obtain a distance-time matrix, slightly scaling the distance-time matrix to obtain the distance-time diagram; performing a short-time Fourier transform on the slow time dimension, coherently superimposing the short-time Fourier transform results of each interval to obtain the Doppler-time diagram.

[0073] Specifically, by processing the intermediate frequency signal collected by the FMCW radar, the range-time map (RTM) and Doppler-time map (DTM) can be generated. These spectrograms provide key feature information for gesture recognition. Figure 2 The preprocessing process for raw echo data is detailed. First, the raw echo data is rearranged into a two-dimensional data matrix, with the horizontal axis representing the slow time dimension and the vertical axis representing the fast time dimension. Next, a fast Fourier transform (FFT) is performed on the fast time dimension to obtain a range-time matrix. This matrix is slightly rescaled to produce a range-time spectrogram (RTM) to capture target variations in the range dimension. To eliminate static background interference, a fourth-order Butterworth high-pass filter with a cutoff frequency of 0.0075 Hz is used as a moving target indicator (MTI), effectively filtering out irrelevant static signals. To further process the arm-radar distance variations caused by different gestures, the system selects a range between 0.2 meters and 1.2 meters. Within each interval, a short-time Fourier transform (STFT) is performed along the slow time dimension with a window length of 0.2 seconds and a 95% overlap ratio to capture the dynamic characteristics of the gesture. The STFT results from each interval are coherently superimposed to generate a Doppler-time map, reflecting the changes in the gesture in the velocity dimension. Combining range-time maps with Doppler-time maps enables highly accurate gesture classification. Because each data type contains unique and important features, efficiently fusing multi-dimensional features is crucial for improving gesture classification accuracy. This multimodal fusion approach not only enhances the model's discriminative capabilities but also provides reliable technical support for gesture recognition in complex scenarios.

[0074] Furthermore, the gesture recognition model includes: a DSCC module, which is used to extract spatial displacement features and gesture speed change information from distance-time graphs and Doppler-time graphs to generate feature graphs; the feature graphs include: a first feature graph and a second feature graph; a position encoding module, which is used to perform weighted feature fusion on the first feature graph and the second feature graph, and introduce position encoding; a Transformer module, which is used to extract global features of dynamic gestures using the position encoding results; and a fully connected module, which is used to process global features, perform gesture classification, and obtain gesture recognition results.

[0075] Specifically, to meet the needs of millimeter-wave radar gesture recognition in complex interference scenarios, the present invention proposes a Muti-DSCC-Transformer hybrid network, whose core architecture is as follows: Figure 3 As shown in the figure. First, the two input distance-time map (RTM) and Doppler-time map (DTM) of the model are respectively input into two depth-separable convolution modules to perform dual-path local feature extraction of dynamic gestures, and then the CAM (channel attention module) is introduced to enhance the feature expression of each channel. Secondly, in order to solve the problem of lack of position information in the parallel input mechanism of Transformer, position encoding is added in the dual-branch weighted feature fusion stage to ensure that the input and output feature order of the MHA mechanism remains consistent. Then, the Transformer architecture is used to extract global features and enhance spatiotemporal features. Finally, the fully connected (FC) module is used to process the output features and classify gestures.

[0076] Furthermore, the DSCC module includes: two parallel convolution sub-modules, which are used to respectively characterize the spatial displacement characteristics of gesture movement and capture the gesture speed change information; the convolution sub-module includes: depthwise separable convolution layer, depthwise separable convolution layer and pooling layer are alternately arranged, the pooling layer at the end is connected to multiple standard convolution layers in sequence, and the standard convolution layer at the end is connected to the CAM unit; the depthwise separable convolution layer is used to extract features from the distance-time map and the Doppler-time map; the pooling layer is used to reduce the spatial dimension of the output result of the depthwise separable convolution layer; the standard convolution layer is used to perform secondary feature extraction on the output result of the pooling layer; the CAM unit is used to perform channel attention enhancement on the output result of the standard convolution layer.

[0077] Specifically, this module is a deep feature extractor for lightweight design, which uses a symmetrical dual-branch structure to process the input image. The specific structure is as follows: Figure 3 shown.

[0078] like Figure 5As shown in Figure 2, the inputs of the DSCC module are RTM and DTM, which respectively characterize the spatial displacement characteristics of gesture motion and capture the change in gesture speed. The input dimensions are both 64×64×3. Each convolution module consists of two depthwise separable convolution layers, three standard convolution layers, and two pooling layers. The convolution kernel sizes are 3×3 and 1×1, and the pooling kernel size is 2×2. The convolution layer is followed by a batch normalization layer and a nonlinear layer. After feature extraction, the outputs are input into the channel attention module (CAM), as shown in Figure 2. Figure 4 As shown, the features of important channels are enhanced. The feature maps after channel attention enhancement are then weighted fused. The self-attention mechanism of the Transformer module is global and has no concept of order. In order to cope with the defects of the Transformer network's position information processing mechanism and capture the positional relationship of elements in the sequence, the position information must be explicitly added before entering the Transformer network. The position encoding used in the present invention is absolute position encoding, which can adapt to input sequences of different lengths. During the training process, the model can learn the relationship between the position encoding and the sequence elements. When encountering new sequences of different lengths, these learned relationships can still be used for reasoning, and to a certain extent, the structural information of the sequence can be restored, thereby improving the performance and stability of the model.

[0079] Furthermore, the Transformer module includes: multiple encoding sub-modules, which are used to extract global features of dynamic gestures using position encoding results; the encoding sub-module includes: a Dropout layer, a layer normalization, a multi-head self-attention layer, a feedforward neural network layer and a normalization layer connected in sequence; wherein, residual connections are added after the multi-head self-attention layer and the feedforward neural network layer; the multi-head self-attention layer is used to perform a linear transformation on the normalization processing results, generate a query vector, a key vector and a value vector, perform a dot product on the query vector and the key vector, and calculate the attention weight.

[0080] The present invention uses Transformer encoder to extract global features of dynamic gestures, such as Figure 6 As shown in the figure, the Transformer encoder is composed of multiple identical layers stacked together. Each layer mainly includes layer normalization, a multi-head self-attention layer, and a feedforward neural network layer. Among them, the most critical is the multi-head attention layer. Its operation process is as follows: First, the input vector is linearly transformed to generate a query vector (Q), a key vector (K), and a value vector (V). The calculation formula for these three vectors is:

[0081]

[0082] Secondly, the attention weight is calculated by performing a dot product between the query vector and all key vectors, and then the weight is normalized by the Softmax function. The calculation formula is:

[0083]

[0084] Finally, the weight coefficients are dot-producted with the value vector to obtain the attention values for different samples, and thus the output of the self-attention mechanism. The multi-head self-attention mechanism is the result of concatenating multiple self-attention mechanisms. First, the input sequence is added to the positional encoding and passed through a dropout layer. It is then processed through multiple block modules, each of which contains a multi-head attention and feedforward network with residual connections. Finally, the output is obtained through a final normalization layer.

[0085] Furthermore, the fully connected module includes: an adaptive average pooling layer for reducing the dimension of the global features; a fully connected layer for mapping the global features after dimensionality reduction to the label space to obtain gesture recognition results.

[0086] Specifically, in the final stage, an adaptive average pooling layer and a fully connected layer are used to output gesture labels. The adaptive average pooling layer plays a crucial role in reducing the dimensionality of the feature vector, which helps reduce the number of model parameters. The fully connected layer maps the feature vector to the label space, enabling dynamic gesture classification.

[0087] Experimental analysis and discussion

[0088] Hardware Platform: This example uses TI's AWR 1843BOOST millimeter-wave radar, operating at 77-81 GHz and with a maximum bandwidth of 4 GHz. It is paired with a DCA 1000 high-speed data acquisition card to collect gesture radar echoes. Table 1 shows the radar configuration parameters.

[0089] Table 1 Radar parameter configuration

[0090]

[0091]

[0092] Data collection: This embodiment designs eight dynamic gestures for data collection, such as Figure 7 (a)-(h) are shown. To improve the robustness of the model, data collection was performed in four non-ideal environments, including an open laboratory, an office, a conference room, and a crowded laboratory. A total of five volunteers participated in the data collection process. The distances and angles of the collection positions 1-6 were (0.6m, 0°), (0.6m, 45°), (0.8m, 0°), (0.8m, -30°), (1.0m, 0°), and (1.0m, 30°), respectively. Figure 8As shown in (a)-(b), each volunteer performed each gesture 48 times, generating a total of 1920 radar echo data files for dynamic gestures. Signal processing was performed on the collected data to obtain the RTM and DTM for each gesture. The resulting dataset was then partitioned into training, test, and validation sets in a 6:2:2 ratio.

[0093] Analysis of experimental results: In order to comprehensively evaluate the comprehensive performance of the recognition method proposed in this invention, this invention carries out systematic experimental verification from three aspects: feature verification, model performance comparison and method robustness.

[0094] (1) Feature Verification:

[0095] Through experimental analysis and comparison of range-time map (RTM), Doppler-time map (DTM) and fusion features, the present invention deeply analyzes the limitations of single features and verifies the reliability and advantages of fusion features. Figure 9 The confusion matrix generated for all features of the model shows that using a single feature for gesture classification results in significant misidentification and confusion. For example, the RTM confusion matrix shows that gesture 4 has a 100% recognition accuracy, while gesture 2 has a only 40% recognition accuracy. The DTM confusion matrix shows that gestures 3 and 4 both achieve high recognition accuracy, while gestures 2 and 5 have less than ideal recognition accuracy. This highlights the limitations and inadequacies of single features in representing gestures in complex and noisy scenes. Specifically, while RTM and DTM demonstrate high recognition accuracy for specific gestures, they each have significant drawbacks. For example, RTM primarily characterizes gestures based on the change in distance between the target object and the radar, resulting in poor performance for gestures involving rapid speed changes. DTM, on the other hand, is limited in its ability to capture distance information for static or slowly moving gestures, failing to provide sufficient discriminative information, resulting in reduced recognition rates. These shortcomings demonstrate that a single feature cannot fully capture the dynamic characteristics of gestures, limiting classification performance.

[0096] In order to overcome the limitations of a single feature, the present invention effectively integrates the features of RTM and DTM. The fusion feature can fully utilize the complementary advantages of the two, thereby significantly improving the accuracy and robustness of gesture recognition. Figure 9 (a) Figure 9 (b) and Figure 9 (c) By comparison, the fused features perform well in reducing misidentification and confusion. The confusion matrix clearly shows that the fused features significantly reduce the errors in gesture recognition, further verifying the effectiveness of the feature fusion strategy.

[0097] Model performance analysis:

[0098] In order to verify the overall performance of the recognition method proposed in the present invention, the present invention first conducted a comparative experiment with different learning rates, and then used four neural network models, including the classic neural network model VGG19, ResNet50, MobileNetV3 and the self-designed neural network model CNN, to compare and verify its performance.

[0099] Comparison of different learning rates:

[0100] During the training process of deep learning models, the learning rate and decay strategy are key parameters that affect network performance. They directly determine the convergence speed and final performance of the model. In order to test the effect of different learning rates on the Multi-DSCC-Transformer model, parameter optimization was performed and different learning rates were tested: 1e-3, 1e-4, 1e-5, 2e-4 and 2e-5. The training loss curve and validation accuracy curve of the model were then observed. Figure 10 As shown in (a)-(b), a learning rate of 1e-5 leads to the highest accuracy, the fastest convergence of the loss and accuracy curves, and the most stable. Therefore, the learning rate of the experiment of the present invention is 1e-5.

[0101] Model Accuracy Analysis: Model accuracy is an important metric for evaluating classification model performance. This paper compares different neural network models with the designed Multi-DSCC-transformer, not only comparing the accuracy of each model but also analyzing its time complexity (flops) and parameter count (params). This can be seen in Table 2.

[0102] The recognition accuracy of the model in the present invention is 97.40%, which is significantly better than other comparison models. Compared with the CNN model with the highest accuracy (92.97%) among the comparison models, our model has improved by 4.43%, showing stronger feature extraction and classification capabilities. In terms of time complexity, the model in the present invention has 42.79GFLOPs, second only to MobileNet V3-small's 4.04GFLOPs, and far lower than VGG19 (206.61GFLOPs), ResNet50 (19.22GFLOPs) and MobileNet V3-large (15.24GFLOPs). Although the model FLOPs is higher than MobileNet V3-small, its recognition accuracy is improved by 18.23% compared with the latter, significantly compensating for the slight increase in computational complexity. The model in this paper has 2.24M parameters, which is significantly lower than VGG19 (73.67M), ResNet50 (21.29M), and MobileNet V3-large (8.42M), and only slightly higher than MobileNet V3-small (3.05M) and CNN (0.96M). Despite this, the Multi-DSCC-transformer model achieves a better balance between parameter count and computational complexity while maintaining the highest recognition accuracy.

[0103] Table 2 Gesture recognition results of different models

[0104]

[0105] Model robustness verification (different environments, different locations):

[0106] Different environments: Environmental factors in different scenarios will significantly affect the propagation characteristics of millimeter wave signals, including the reflection, absorption and scattering of signals by obstacles. This difference leads to significant changes in the intensity, clutter interference and other aspects of the gesture signals collected in different scenarios, which increases the difficulty of gesture recognition. Therefore, it is crucial to verify the generalization ability and robustness of the proposed method in different scenarios. To this end, we conducted experiments in four scenarios: crowded laboratories, offices, conference rooms, and empty laboratories to comprehensively evaluate the scenario adaptability of the method. Experimental results show that there are significant differences in recognition accuracy in different scenarios. As Figure 11As shown, the conference room achieved the highest recognition accuracy, reaching 98.95%, while the crowded laboratory scene achieved the lowest recognition accuracy, at 90.63%. Comparing the recognition results in the crowded and open laboratory environments revealed that factors such as scene size, interference, and clutter significantly impacted the recognition results. Specifically, the conference room achieved the highest recognition accuracy due to its relatively large size and low interference, resulting in more stable signal propagation. In contrast, the crowded laboratory, with its confined space and numerous obstacles and people, led to complex signal reflection and scattering, increased clutter interference, and reduced recognition accuracy. The recognition accuracy rates for the office and open laboratory scenes were 96.85% and 92.71%, respectively. Although the office scene was smaller, it had relatively neatly arranged objects and less interference, resulting in higher recognition accuracy. However, the open laboratory, despite its larger size, lacked sufficient reflective objects, potentially resulting in insufficient signal strength, affecting recognition accuracy. These experimental results validate the adaptability and robustness of the proposed method across various scenarios and provide important insights for future optimization of gesture recognition technology.

[0107] Different positions: The different positions of the experimental gestures relative to the radar will cause the attenuation of the electromagnetic signal and the difference in multipath reflection, which will cause the amplitude and phase of the received signal to change, that is, the difference in distance and angle, and ultimately lead to differences in gesture features. This feature difference has an important impact on the accuracy of the gesture recognition method, so it needs further verification. In order to evaluate the performance of gesture recognition at different distances and angles, we designed 9 experimental positions, namely (0.6m, 45°), (0.6m, 30°), (0.6m, 0°), (0.8m, 45°), (0.8m, 30°), (0.8m, 0°), (1.0m, 45°), (1.0m, 30°), (1.0m, 0°). Corresponding to each experimental position, we collected 20 sets of data for each gesture to comprehensively evaluate the recognition performance. Figure 12 As shown, there are significant differences in recognition accuracy at different locations. Specifically, at (0.6m, 0°), the recognition accuracy reaches a maximum of 97.60%, while at (1.0m, 45°), the recognition accuracy is the lowest, at only 88.75%. Further analysis reveals that at the same angle, recognition accuracy gradually decreases with increasing distance; and at the same distance, recognition accuracy drops significantly as the angle increases from 0° to 45°. The main reasons for these differences include the following: the radar's field of view, signal transmission power, and receiving sensitivity limit its ability to capture gestures at long distances and specific angles. Obstacles, clutter, and other interference sources in complex environments further weaken signal quality, especially at long distances and non-optimal angles. These analyses further verify the adaptability and robustness of the proposed method at different locations, providing an important reference for gesture recognition in complex scenarios.

[0108] To address the interference challenges faced by millimeter-wave radar dynamic gesture recognition in complex scenarios, this paper proposes a dynamic gesture recognition method using a multi-feature fusion network. First, by collecting radar data in four different scenarios, the data diversity and complexity are enhanced, better simulating the challenges of real-world environments. Second, the gesture data is filtered and denoised, significantly improving the signal-to-noise ratio (SNR). Furthermore, feature maps such as the range-time map (RTM) and Doppler-time map (DTM) are extracted and analyzed in depth. These feature maps comprehensively reflect the gesture's motion state in the distance, velocity, and time dimensions. To further reduce model complexity, the paper designs a lightweight Multi-DSCC–Transformer model. This model combines depthwise separable convolution (DSCC) and Transformer modules, achieving high recognition accuracy while maintaining low memory requirements. By comparing the recognition performance of different features, the paper verifies the correctness of feature extraction and the necessity of multimodal feature fusion. This example also conducted comparative experiments with five different neural network models. The results showed that the proposed Multi-DSCC–Transformer method outperformed other models in recognition accuracy while maintaining the lowest memory requirements. The experiments further demonstrated that the proposed model exhibited strong robustness in terms of position, angle, and scene variations, demonstrating its adaptability in complex environments. This approach lays the foundation for subsequent continuous gesture recognition based on millimeter-wave radar.

[0109] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A dynamic gesture recognition method using a multi-feature fusion network under complex scene interference, characterized by: include: Acquire a baseband intermediate frequency signal, process the baseband intermediate frequency signal, and acquire a range-time graph and a Doppler-time graph; Inputting the distance-time graph and the Doppler-time graph into a gesture recognition model to obtain a gesture recognition result; The gesture recognition model is obtained by training a lightweight neural network model using a training set, wherein the training set includes: a distance-time spectrogram, a Doppler-time spectrogram, and a gesture label; The DSCC module in the gesture recognition model extracts the spatial displacement features and gesture velocity change information of the distance-time graph and the Doppler-time graph to obtain a feature graph. The feature graph is combined with the position code and input into the Transformer module for global feature extraction to obtain the gesture recognition result.

2. The dynamic gesture recognition method under complex scene interference using a multi-feature fusion network according to claim 1 is characterized in that: Acquiring the baseband intermediate frequency signal includes: The baseband intermediate frequency signal is generated using an FMCW radar device: Acquire a detection signal whose frequency changes linearly with time, split the detection signal, amplify one detection signal and radiate it into space to detect the target, and use the other detection signal as a reference signal; When the radiation signal detects the target, the signal is returned and mixed with the reference signal to generate an intermediate frequency signal containing target information. The intermediate frequency signal is filtered to obtain the baseband intermediate frequency signal.

3. The dynamic gesture recognition method under complex scene interference using a multi-feature fusion network according to claim 1 is characterized in that: Acquiring the range-time graph and the Doppler-time graph includes: Rearranging the baseband intermediate frequency signal into a two-dimensional data matrix; wherein the horizontal axis of the two-dimensional data matrix represents a slow time dimension, and the vertical axis represents a fast time dimension; Performing a fast Fourier transform on the fast time dimension to obtain a distance-time matrix, and slightly scaling the distance-time matrix to obtain the distance-time graph; Performing short-time Fourier transform on the slow time dimension, coherently superimposing the short-time Fourier transform results of each interval to obtain the Doppler-time map.

4. The dynamic gesture recognition method under complex scene interference using a multi-feature fusion network according to claim 1 is characterized in that: The gesture recognition model includes: A DSCC module is configured to extract spatial displacement features and gesture velocity change information from the distance-time graph and the Doppler-time graph to generate the feature graph; the feature graph includes a first feature graph and a second feature graph; A position encoding module, configured to perform weighted feature fusion on the first feature map and the second feature map, and introduce position encoding; Transformer module, used to extract global features of dynamic gestures using position encoding results; The fully connected module is used to process the global features, perform gesture classification, and obtain the gesture recognition result.

5. The dynamic gesture recognition method under complex scene interference using a multi-feature fusion network according to claim 4 is characterized in that: The DSCC module includes: two parallel convolution submodules for respectively characterizing the spatial displacement characteristics of gesture motion and capturing gesture speed change information; The convolution submodule includes: a depthwise separable convolution layer, the depthwise separable convolution layer and a pooling layer arranged alternately, the pooling layer at the end is sequentially connected to multiple standard convolution layers, and the standard convolution layer at the end is connected to a CAM unit; wherein, a batch normalization layer and a nonlinear layer are added after each convolution layer; The depthwise separable convolutional layer is used to extract features from the range-time map and the Doppler-time map; The pooling layer is used to reduce the spatial dimension of the output result of the depth-wise separable convolutional layer; The standard convolution layer is used to perform secondary feature extraction on the output result of the pooling layer; The CAM unit is used to perform channel attention enhancement on the output result of the standard convolutional layer.

6. The dynamic gesture recognition method under complex scene interference using a multi-feature fusion network according to claim 4 is characterized in that: The Transformer module includes: multiple encoding submodules for extracting global features of dynamic gestures using position encoding results; The encoding submodule includes: a dropout layer, a layer normalization layer, a multi-head self-attention layer, a feedforward neural network layer and a normalization layer connected in sequence; wherein, a residual connection is added after the multi-head self-attention layer and the feedforward neural network layer; The multi-head self-attention layer is used to perform a linear transformation on the normalized processing result to generate a query vector, a key vector and a value vector, perform a dot product on the query vector and the key vector, and calculate the attention weight.

7. The dynamic gesture recognition method using a multi-feature fusion network under complex scene interference according to claim 6 is characterized in that: Generating the query vector, key vector, and value vector includes: Among them, Q is the query vector, K is the key vector, V is the value vector, T is the transpose of the vector, and d k is the dimension of the key vector and query vector.

8. The method for dynamic gesture recognition using a multi-feature fusion network under complex scene interference according to claim 6 is characterized in that: Calculating the attention weight includes: Among them, Weigh ts is the weight.

9. The method for dynamic gesture recognition using a multi-feature fusion network under complex scene interference according to claim 4 is characterized in that: The fully connected module includes: An adaptive average pooling layer for reducing the dimension of the global features; The fully connected layer is used to map the global features after dimensionality reduction to the label space to obtain the gesture recognition result.

Citation Information

Patent Citations

  • Gesture segmentation network device and method based on multi-branch cascade Transformer

    CN115393950A

  • Millimeter wave radar dynamic gesture recognition method applied to interference environment

    CN116794602A

  • Multi-feature lightweight gesture recognition method based on millimeter wave radar

    CN117935368A

  • Millimeter wave radar gesture recognition method based on lightweight neural network

    CN118587769A

  • Gesture recognition method and apparatus

    US20230333209A1

Cited By

  • Gesture detection method and device based on application scene and storage medium

    CN121214500A

  • Gesture detection method and device based on application scenario, and storage medium

    CN121214500B