Millimeter wave radar gesture key point detection method, system and device based on double-flow depth fusion and medium

Through the dual-stream parallel network and cross-attention fusion mechanism, the problem of insufficient feature information in the detection of key points of gestures in millimeter-wave radar is solved, and high-precision and low-cost dynamic gesture recognition is achieved, which is suitable for fields such as virtual reality and intelligent driving.

CN120491097APending Publication Date: 2025-08-15XIDIAN UNIV +1

Patent Information

Application Number
CN202510843839.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the existing millimeter-wave radar gesture key point detection technology, the feature information is insufficient and the gesture key point regression accuracy is limited, especially in the recognition of complex dynamic gestures.

Method used

The dual-stream parallel network architecture is adopted to extract spatiotemporal and spectral features from the distance-Doppler graph and microDoppler spectrum respectively, and the cross-doppler depth fusion mechanism is used to perform timing regression by combining the key point detection decoder, cross-modal feature fusion and Hungarian algorithm optimal matching are designed.

Benefits of technology

It significantly improves the accuracy and robustness of gesture key point detection, realizes high-precision key point detection, reduces hardware costs, and maintains detection effect under changes in lighting conditions to ensure user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491097A_ABST
    Figure CN120491097A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of key point detection, and particularly relates to a millimeter wave radar gesture key point detection method, system and device based on double-flow depth fusion and a medium, and the method comprises the steps: carrying out the preprocessing of hand motion data, obtaining a distance-Doppler three-dimensional space-time tensor, a micro-Doppler spectrogram and a key point truth value sequence, distance-Doppler three-dimensional space-time tensor and micro-Doppler spectrogram features are extracted, cross-modal feature fusion is carried out, a key point detection decoder is designed, a gesture key point prediction coordinate sequence is estimated based on modal fusion features, and a distance-Doppler three-dimensional space-time tensor and micro-Doppler spectrogram features are obtained; screening the first m gesture key point prediction coordinate sequences with the highest confidence score, performing optimal matching with the key point true value sequence through a Hungary algorithm to obtain overall regression loss, and performing iterative training to complete a gesture key point detection model; systems, devices, and media for implementing the methods thereof; the method has the comprehensive application advantages of high precision, high robustness and low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of key point detection, and specifically relates to a millimeter wave radar gesture key point detection method, system, equipment and medium based on dual-stream deep fusion. Background Art

[0002] With the development of virtual reality, intelligent driving, and other fields, contactless gesture interaction has become a key technology for improving user experience and safety. However, traditional solutions, such as depth cameras or standard cameras, are susceptible to interference from ambient lighting and obstructions, and carry high computational costs and the risk of leaking personal privacy. Solutions such as inertial sensors require users to wear the device, compromising ease of use.

[0003] In contrast, millimeter-wave radar sensing technology, with its unique advantages, offers a completely new approach to gesture recognition, particularly for refined keypoint detection. Millimeter-wave radar uses frequency-modulated continuous wave signals to measure a target's distance, angle, and speed. Its advantages include all-weather operation, unaffected by lighting conditions, and signal penetration, making it effectively immune to obstruction by clothing or thin obstacles. Furthermore, since it doesn't capture visual images, it fundamentally protects user privacy.

[0004] In millimeter-wave radar gesture recognition research, detecting gesture keypoints by directly estimating the spatial coordinates of each hand joint is a more challenging task than simple gesture classification. Although some work has begun to address this issue, it still has many drawbacks.

[0005] For example, the patent application with publication number CN118409311A discloses a millimeter-wave radar hand key point detection system and method based on unsupervised feature learning. This solution only uses a single radar heat map feature as the input of the model, and fails to simultaneously utilize multi-dimensional radar features containing different physical information (such as micro-Doppler maps containing fine speed information). As a result, the feature information of the input model is insufficient, making it difficult to fully characterize complex dynamic gestures.

[0006] In addition, the gesture key point detection method proposed in the paper EgoHand: Ego-centric Hand Pose Estimation and Gesture Recognition with Head-mounted Millimeter-wave Radar and IMUs (Lv Y, Zhang T, Song Y, et al. EgoHand: Ego-centric Hand Pose Estimation and Gesture Recognition with Head-mounted Millimeter-wave Radar and IMUs[J]. arXiv preprint arXiv: 2501.13805, 2025.), although verifying the feasibility of key point detection, only performs simple channel splicing on the two feature maps of the millimeter-wave radar before inputting into the network, and fails to deeply explore the rich multi-radar feature fusion and interactive decoding mechanism, resulting in limited accuracy for refined gesture key point regression. Summary of the Invention

[0007] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a millimeter-wave radar gesture key point detection method based on dual-stream deep fusion. A dual-stream parallel network is used to extract two millimeter-wave radar features, decoupling gesture motion into spatiotemporal features and spectral features. A cross-attention deep fusion mechanism is then introduced, and a dedicated decoder with a "temporal mask" is used for temporal regression to improve the detection accuracy and robustness of dynamic gestures. This aims to solve the problems in the prior art of insufficient feature information and limited gesture key point regression accuracy caused by the use of a single millimeter-wave radar feature or a simple feature fusion method.

[0008] In order to achieve the above object, the technical solution adopted by the present invention is:

[0009] The millimeter-wave radar gesture key point detection method based on dual-stream deep fusion includes the following steps:

[0010] Step 1: The millimeter-wave radar is used to collect the original hand motion signal, and the RGB camera is used to synchronously capture multiple frames of hand images. The collected data is then preprocessed to obtain the range-Doppler three-dimensional space-time tensor X RDM , Micro-Doppler spectrum X MD And the key point true value sequence {H gt} (m) ;

[0011] Step 2: Extract the range-Doppler three-dimensional space-time tensor X RDMFeatures, get the range-Doppler space-time characteristics

[0012] Step 3: Extract the micro-Doppler spectrum X MD Spectral characteristics, get micro-Doppler spectrum characteristics

[0013] Step 4: transform the range-Doppler spatiotemporal characteristics and micro-Doppler spectrum characteristics Perform cross-modal feature fusion to obtain cross-modal fusion features

[0014] Step 5: Design a key point detection decoder based on modality fusion features Estimated gesture keypoint prediction coordinate sequence {H topk} (m) ;

[0015] Step 6: Predict the coordinate sequence of the top m gesture key points by filtering the ones with the highest confidence scores The key point truth sequence {H gt} (m) Perform the optimal matching of the Hungarian algorithm to obtain the overall regression loss L reg ;

[0016] Step 7: Repeat steps 2 to 6 for N, N ≥ 1 rounds to complete the training of the gesture key point detection model.

[0017] The specific steps of data preprocessing in step 1 are as follows:

[0018] First, apply a window function to each frame of the original hand motion signal and perform distance FFT and Doppler FFT operations to obtain the range-Doppler map of n frames. Then the continuous n frames of range-Doppler map Superposition to form the range-Doppler three-dimensional space-time transition tensor Among them, H is the distance resolution dimension, W is the velocity resolution dimension; at the same time, the range-Doppler three-dimensional space-time transition tensor X' RDM Perform short-time Fourier transform (STFT) in the time domain to generate the corresponding micro-Doppler transition spectrum Among them, T is the number of time frames, T = n, F is the frequency resolution; finally, the range-Doppler three-dimensional space-time transition tensor X' RDM and micro-Doppler transition spectrum X' MD Through vector mean elimination processing, the range-Doppler three-dimensional space-time tensor X is obtained RDM and Micro Doppler Spectrum X MD ;

[0019] Use MediaPipe to process n frames of hand images and generate 3D coordinate labels for a hand key points:

[0020]

[0021] Among them, x, y, z represent the 3D coordinates of the key points of the hand;

[0022] 3D coordinate labels and the range-Doppler three-dimensional space-time tensor X RDM and Micro Doppler Spectrum X MD Timing alignment, unified to the m-frame key point true value sequence through linear interpolation

[0023] The specific steps of step 2 are as follows:

[0024] Construct a hybrid architecture that combines a three-dimensional convolutional neural network (3D-CNN) with a temporal convolutional network (TCN):

[0025] Range-Doppler three-dimensional space-time tensor X RDM First, the preliminary features are extracted through the three-dimensional convolutional neural network (3D-CNN). The formula is:

[0026]

[0027] in, is the convolution kernel weight of this layer, is bias;

[0028] And through normalization and nonlinear activation, we get the nonlinear normalized spatiotemporal features

[0029]

[0030] Then, the nonlinear normalized spatiotemporal features W' are downsampled in the spatial dimension to obtain high-level semantic features H (1) <H,W (1) <W;

[0031] Then the high-level semantic features W1 are globally averaged in the spatial dimension to obtain the time series The formula is:

[0032]

[0033] Finally, the time series S is fed into a l-layer temporal convolutional network (TCN) to output the range-Doppler spatiotemporal features. The formula is:

[0034]

[0035] in, D' is the spatiotemporal feature dimension.

[0036] The specific steps of step 3 are as follows:

[0037] First, multi-layer two-dimensional convolution is used to extract the micro-Doppler spectrum X MD The local time-frequency texture feature U1' is as follows:

[0038]

[0039] in, is the first layer convolution weight, is the bias, C is the output channel;

[0040] And through normalization and nonlinear activation, the nonlinear normalized spectral feature U(t,f,c) is obtained;

[0041] The maximum pooling in frequency dimension is adopted, and the formula is:

[0042] Φ'(t,f',c)=max{U(t,2f',c),U(t,2f'+1,c)}

[0043] Where, f'=0...[F / 2]-1,

[0044] Then, a 1×1 convolution kernel is used to expand the number of channels after pooling to obtain the spectrum extension feature.

[0045] Finally, the spectrum expansion feature Φ is averaged in the frequency dimension

[0046] The final output is the micro-Doppler spectrum characteristics

[0047] The specific steps of step 4 are as follows:

[0048] Range-Doppler spatiotemporal characteristics and micro-Doppler spectrum characteristics Do linear mapping to get key features Value Features Query Features Among them, D is the fusion attention feature dimension; then the similarity score is performed And normalize the similarity to get the attention weight Then aggregate the value feature V RDM Get aggregated attention features Finally, residual connections and layer normalization are introduced to output cross-modal fusion features

[0049]

[0050] The specific steps of step 5 are as follows:

[0051] First, define a set of learnable pose query feature vectors D is the fusion attention feature dimension, where each pose query feature q i Random initialization and joint optimization with the training process, the number of queries is designed to be I to cover the redundant candidate poses in different interaction scenarios;

[0052] The key point detection decoder is a multi-layer decoder stacking structure. Specifically, each decoding layer cascades the self-attention module and the cross-attention module, with a total of L layers stacked. The calculation process of the l∈L layer is:

[0053]

[0054] in, For learnable position encoding, in MultiHeadAttn(·), the input parameters will undergo a linear transformation. is the temporal mask matrix:

[0055]

[0056] Among them, Δ represents the threshold of the number of frames before and after attention;

[0057] After L layers of decoding, the posture decoding feature vector is obtained Obtain candidate key point coordinates H through MLP mapping cand With confidence S cand :

[0058]

[0059] Among them, σ is the Sigmoid function, a represents the number of key points of the hand, 3 is the 3D coordinate dimension of the key points of the hand, and by selecting the confidence S cand The coordinates of the first m, m≤I candidate key points H cand , as the final gesture key point prediction coordinate sequence {H topk} (m) .

[0060] The specific steps in step 6 are as follows:

[0061] First, define the gesture key point prediction coordinate sequence {H topk} (m) With the key point true value sequence {H gt} (m) Any pair of gesture key point prediction coordinates H topkand the key point truth value H gt The key point loss between It calculates the L2 distance between a corresponding keypoints:

[0062]

[0063] in, and Predict the coordinates H for any pair of gesture key points topk and the key point truth value H gt The 3D coordinates of the jth key point in ;

[0064] Then, the Hungarian algorithm is used to find the coordinates H that can make m predict the key points of the gesture topk and the key point truth value H gt The matching permutation with the smallest total loss:

[0065]

[0066] in, is the set of all possible permutations of the set {1,2,…,m}, is the predicted coordinate sequence of gesture key points {H topk} (m) The predicted coordinates of the i-th gesture key point in , is the key point true value sequence {H gt} (m) The σ(i)th one in S m Permutation of the i-th item in the set;

[0067] Finally, by arrangement The minimum total loss obtained is the overall regression loss L of the gesture key point detection model reg :

[0068]

[0069] The millimeter-wave radar gesture key point detection system based on dual-stream deep fusion includes:

[0070] The data preprocessing module collects the original hand motion signal through the millimeter wave radar and uses the RGB camera to synchronously capture multiple frames of hand images. The collected data is then preprocessed to obtain the range-Doppler three-dimensional space-time tensor X RDM , Micro-Doppler spectrum X MD And the key point true value sequence {H gt} (m) ;

[0071] Dual-stream feature extraction module, including a range-Doppler map spatiotemporal feature extraction submodule and a micro-Doppler feature extraction submodule;

[0072] The range-Doppler image spatiotemporal feature extraction submodule extracts the range-Doppler three-dimensional spatiotemporal tensor X RDM Features, get the range-Doppler space-time characteristics

[0073] The micro-Doppler feature extraction submodule extracts the micro-Doppler spectrum X MD Spectral characteristics, get micro-Doppler spectrum characteristics

[0074] Cross-modal feature fusion module combines distance-Doppler spatiotemporal features and micro-Doppler spectrum characteristics Perform cross-modal feature fusion to obtain cross-modal fusion features

[0075] Key point detection decoder module, design key point detection decoder based on modal fusion features Estimated gesture keypoint prediction coordinate sequence {H topk} (m) ;

[0076] Model training and testing module, predicting coordinate sequences by screening the top m gesture key points with the highest confidence scores With the key point true value sequence {H gt} (m) Perform the optimal matching of the Hungarian algorithm to obtain the overall regression loss L reg , and perform N, N ≥ 1 rounds of iteration to complete the training of the gesture key point detection model.

[0077] Millimeter-wave radar gesture key point detection equipment based on dual-stream deep fusion, including:

[0078] Memory: used for storing a computer program for implementing a millimeter-wave radar gesture key point detection method based on dual-stream deep fusion;

[0079] Processor: used to implement a millimeter-wave radar gesture key point detection method based on dual-stream deep fusion when executing the computer program.

[0080] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a millimeter-wave radar gesture key point detection method based on dual-stream deep fusion.

[0081] Compared with the prior art, the present invention has the following beneficial effects:

[0082] 1. The present invention significantly improves the accuracy of gesture key point detection based on millimeter wave radar. This effect is due to the dual-stream parallel network architecture designed by the present invention, which extracts complementary spatiotemporal features and speed features from the range-Doppler map and micro-Doppler spectrum through two different structures: "3D convolution + temporal convolution network" and "2D convolution + frequency pooling". Furthermore, the cross-attention fusion mechanism adopted by the present invention can deeply fuse these two features. These technical features together ensure that the model can obtain more comprehensive and profound feature information, so that in the test of the present invention, the mean error (MPJPE) of key point detection can reach 42.8 mm, which is significantly better than the existing technology and can accurately track subtle finger displacements.

[0083] 2. While achieving high precision, the present invention has achieved the application advantages of a simple system, low cost and strong practicality. The key point detection decoder architecture designed by the present invention uses the learnable "posture query feature vector" as a dynamic probe to actively capture the hand posture; in addition, a "temporal mask" mechanism is introduced during decoding, which strictly limits the attention to adjacent time frames, thereby ensuring the smoothness and coherence of the output trajectory at the algorithm level and effectively suppressing motion noise. Through this technical feature, the system can work stably without relying on additional hardware such as IMU for compensation, which directly brings the application advantages of simplified hardware structure and cost savings. In addition, the present invention only uses visual tools to generate labels during the training phase, and does not require a camera during application. This technical feature makes it unrestricted by lighting conditions and can be deployed in scenes with large light ratios or dim lighting. It can also thoroughly protect user privacy, further enhancing the practicality of the solution.

[0084] In summary, the present invention provides a method, system, device, and readable storage medium for millimeter-wave radar gesture keypoint detection based on dual-stream deep fusion. By designing a dual-stream parallel network to extract complementary radar features, employing a cross-attention mechanism for deep fusion, and combining it with a keypoint detection decoder with a "temporal mask" for temporal regression, the present invention not only significantly improves keypoint detection accuracy but also enables a system design that requires no additional hardware and preserves user privacy, achieving the combined advantages of high precision, high robustness, and low cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] Figure 1 Flow chart of the detection method of the present invention.

[0086] Figure 2 Schematic diagram of the network structure of the gesture key point detection model of the present invention.

[0087] Figure 3 This figure shows part of the data set used in the verification experiment of this invention. DETAILED DESCRIPTION

[0088] The present invention will be described in detail below with reference to the accompanying drawings.

[0089] like Figure 1 As shown in FIG, the millimeter wave radar gesture key point detection method based on dual-stream deep fusion includes the following steps:

[0090] Step 1: Data preprocessing

[0091] First, data is collected synchronously: the millimeter-wave radar collects the original hand motion signal, and the RGB camera simultaneously captures multiple frames of hand images;

[0092] Then, the collected data is preprocessed: a window function is applied to each frame of the original hand motion signal and distance FFT and Doppler FFT operations are performed to obtain a 25-frame range-Doppler map. Then the 25 consecutive frames of range-Doppler map Superposition to form the range-Doppler three-dimensional space-time transition tensor Among them, H is the distance resolution dimension, W is the velocity resolution dimension; at the same time, the range-Doppler three-dimensional space-time transition tensor X' RDM Perform short-time Fourier transform (STFT) in the time domain to generate the corresponding micro-Doppler transition spectrum Among them, T is the number of time frames, T = n, F is the frequency resolution; finally, the range-Doppler three-dimensional space-time transition tensor X' RDM and micro-Doppler transition spectrum X' MD Through vector mean elimination processing, the range-Doppler three-dimensional space-time tensor X is obtained RDM and Micro Doppler Spectrum X MD , to reduce the impact of static clutter.

[0093] Use MediaPipe to process n frames of hand images and generate 3D coordinate labels for 21 hand key points:

[0094]

[0095] Among them, x, y, z represent the 3D coordinates of the key points of the hand;

[0096] Finally, the 3D coordinate label and the range-Doppler three-dimensional space-time tensor X RDM and Micro Doppler Spectrum X MD Timing alignment, unified to 5-frame key point true value sequence through linear interpolation

[0097] Step 2: Extract the range-Doppler three-dimensional space-time tensor X RDM feature;

[0098] Construct a hybrid architecture that combines a three-dimensional convolutional neural network (3D-CNN) with a temporal convolutional network (TCN):

[0099] Hand motion in range-Doppler three-dimensional space-time tensor X RDM The spectral response characteristics that move over time are shown in the 3D convolutional neural network (3D-CNN). It can extract local features in the three dimensions of time, distance, and speed at the same time. The distance-Doppler three-dimensional space-time tensor X RDM First, preliminary features are extracted using a three-dimensional convolutional neural network (3D-CNN) with a convolution kernel size of 3×3×3, input channel 1, output channel A1=64, stride (1,1,1), and padding (1,1,1). The formula is:

[0100]

[0101] in, is the convolution kernel weight of this layer, is bias;

[0102] Through normalization and nonlinear activation, nonlinear normalized spatiotemporal features are obtained to improve the stability and nonlinear expression ability of network training. In this embodiment, Batch Normal (BN) is used for normalization and ReLU is used for nonlinear activation:

[0103] W'(t,h,w,a)=RELU(BN(W”(t,h,w,a)))

[0104] in,

[0105] Subsequently, a three-dimensional maximum pooling operation with a kernel of (1, 2, 2) and a step size of (1, 2, 2) is used to downsample the nonlinear normalized spatiotemporal features W' in the spatial dimension to obtain high-level semantic features. H (1) <H,W (1) <W;

[0106] To fully explore the temporal dependency of gestures, the present invention introduces a temporal convolutional network (TCN) consisting of multiple 1D dilated convolutions on the output channel. Specifically, the high-level semantic features W1 are globally averaged in the spatial dimension to obtain the time series The formula is:

[0107]

[0108] Finally, the time series S is fed into a l-layer temporal convolutional network (TCN), where each layer contains a temporal dilation convolution with a kernel width of 3, where the dilation factor increases exponentially, and the dilation rate of the l-th layer is d = 2 r,r=0,1,2,3, output range-Doppler spatiotemporal characteristics The formula is:

[0109]

[0110] in, D' is the spatiotemporal feature dimension, D'=256.

[0111] Step 3: Extract the micro-Doppler spectrum X MD Spectral characteristics;

[0112] The goal of this step is to convert the micro-Doppler spectrum X MD Compressed into a compact time-spectrum representation to be compared with the range-Doppler spatiotemporal characteristics Time alignment and subsequent fusion. Hand micro-motion in micro-Doppler spectrum X MD The local time-frequency energy ridges (slopes, curves, etc.) are shown in the image. Using two-dimensional convolution can capture these textures in both time and frequency directions. Therefore, we first use two-dimensional convolution to extract the micro-Doppler spectrum X. MD The local time-frequency texture feature U1', convolution kernel size: 3×3; input channel 1; output channel C1 = 32; step size (1, 1); padding (1, 1), the formula is:

[0113]

[0114] in, is the first layer convolution weight, is the bias, C1 is the output channel. And through normalization and nonlinear activation, the nonlinear normalized spectrum feature is obtained:

[0115] U1(t,f,c)=RELU(BN(U1'(t,f,c)))

[0116] in, BN is batch normalization, RELU is the activation function;

[0117] A second convolution layer is then used to integrate information across time frames and frequency channels at a higher level, which helps capture complex patterns such as high-frequency energy changes at the moment of kneading. The convolution kernel size is 3×3; input channels C1 = 32; output channels C2 = 64; stride (1, 1); padding (1, 1). The formula is:

[0118]

[0119] in,

[0120] In order to retain the most significant Doppler energy peak in each time frame, the frequency dimension is compressed to reduce redundancy. The maximum pooling in the frequency dimension is adopted, the convolution kernel size is 1×2, and the step size is (1,2). The formula is:

[0121] Φ'(t,f',c')=max{U2(t,2f',c'),U2(t,2f'+1,c')}

[0122] Where, f'=0...[F / 2]-1,

[0123] Then, a 1×1 convolution kernel is used to expand the number of channels after pooling C2=64 to C m =128, get the spectrum expansion feature d=1...128, so as to match the range-Doppler space-time characteristics The number of channels is compatible in subsequent fusions and improves expression capacity.

[0124] Finally, the spectrum expansion feature Φ is averaged in the frequency dimension

[0125] The final output is the micro-Doppler spectrum characteristics Provides compact and efficient time-spectrum representation for the cross-modal fusion module.

[0126] Step 4: transform the range-Doppler spatiotemporal characteristics and micro-Doppler spectrum characteristics Perform cross-modal feature fusion;

[0127] Micro-Doppler spectrum characteristics Contains only time-velocity information, lacks spatial positioning; range-Doppler spatiotemporal characteristics Contains distance, speed and spatial position information, but lacks macroscopic changes in speed. Cross-modal feature fusion combines the advantages of "micro-Doppler is sensitive to macroscopic changes in time-speed" and "range-Doppler is accurate in positioning space-speed-distance" to obtain a richer and more robust gesture representation. The inputs of this step are the distance-Doppler spatiotemporal features. As key and value features in the attention mechanism, micro-Doppler spectrum features As the query feature in the attention mechanism, in order to make the query feature have the same feature dimension as the key feature and value feature in the attention calculation, the distance-Doppler spatiotemporal feature and micro-Doppler spectrum characteristics Do linear mapping to get key features Value Features Query Features Then perform similarity scoring And normalize the similarity to get the attention weight Then aggregate the value feature V RDM Get aggregated attention features Finally, in order to maintain the micro-Doppler spectrum characteristics The information is used to stabilize the training, and the residual connection and layer normalization are introduced to output cross-modal fusion features.

[0128]

[0129] Step 5: Design a key point detection decoder based on modality fusion features Estimated gesture keypoint prediction coordinate sequence {H topk} (m) ;

[0130] The keypoint detection decoder is based on the Transformer multi-layer decoder architecture and implements end-to-end keypoint regression through a learnable query vector. The specific steps are as follows:

[0131] First, define a set of learnable pose query feature vectors Where each pose query feature q i Random initialization and joint optimization during training are used to capture the potential characteristic patterns of different hand postures. The number of queries is designed to be 50 to cover the redundant candidate postures in different interaction scenarios.

[0132] The key point detection decoder is a multi-layer decoder stacking structure. Specifically, each decoding layer cascades the self-attention module and the cross-attention module, with a total of 6 layers (L=6). The calculation process of the lth layer is:

[0133]

[0134] in, To provide temporal sequence information for learnable position encoding, in MultiHeadAttn(·), the input parameters will undergo a linear transformation. A temporal mask matrix designed for this invention ensures that each temporal query only incorporates information from adjacent frames, aiming to capture local correlations in continuous actions while suppressing interference from irrelevant information at distant moments:

[0135]

[0136] Among them, the threshold Δ=3 means that the query frame t only focuses on the posture key features and posture value features within the previous and next 3 frames;

[0137] After 6 layers of decoding, the posture decoding feature vector is obtained Obtain candidate key point coordinates H through MLP mapping cand With confidence S cand:

[0138]

[0139] Among them, σ is the Sigmoid function, 21 represents the 21 key points of the hand, 3 is the 3D coordinate dimension of the key points of the hand, and by selecting the confidence S cand The coordinates of the top 5 candidate key points H cand , as the final gesture key point prediction coordinate sequence {H topk} (5) .

[0140] The network structure of the gesture key point detection model in steps 2 to 5 is as follows Figure 2 shown.

[0141] Step 6: Predict the coordinate sequence of the top 5 gesture key points with the highest confidence scores {H topk} (5) The key point truth sequence {H gt} (5) Perform the Hungarian algorithm optimal matching, the formula is as follows:

[0142]

[0143] Get the regression loss L reg ;

[0144] First, define the gesture key point prediction coordinate sequence {H topk} (5) With the key point true value sequence {H gt} (5) Any pair of gesture key point prediction coordinates H topk and the key point truth value H gt The key point loss between It computes the L2 distance between 21 corresponding keypoints:

[0145]

[0146] in, and Predict the coordinates H for any pair of gesture key points topk and the key point truth value H gt The 3D coordinates of the jth key point in .

[0147] Then, the Hungarian algorithm is used to find the coordinates H that can predict the five pairs of gesture key points. topk and the key point truth value H gt The matching permutation with the smallest total loss

[0148]

[0149] in, is the set of all possible permutations of the set {1,2,…,5}, is the predicted coordinate sequence of gesture key points {H topk} (m) The predicted coordinates of the i-th gesture key point in , is the key point true value sequence {H gt} (5) The σ(i)th one in , where σ(i) is Permutes the i-th item in the set.

[0150] Finally, by arrangement The minimum total loss obtained is the overall regression loss L of the gesture key point detection model reg :

[0151]

[0152] Step 7: Repeat steps 2 to 6 for 300 rounds to complete the training of the gesture key point detection model.

[0153] The millimeter-wave radar gesture key point detection system based on dual-stream deep fusion includes:

[0154] The data preprocessing module collects the original hand motion signal through the millimeter wave radar and uses the RGB camera to synchronously capture multiple frames of hand images. The collected data is then preprocessed to obtain the range-Doppler three-dimensional space-time tensor X RDM , Micro-Doppler spectrum X MD And the key point true value sequence {H gt} (m) , used to implement step 1 of the gesture key point detection method of the present invention;

[0155] Dual-stream feature extraction module, including a range-Doppler map spatiotemporal feature extraction submodule and a micro-Doppler feature extraction submodule;

[0156] The range-Doppler image spatiotemporal feature extraction submodule extracts the range-Doppler three-dimensional spatiotemporal tensor X RDM Features, get the range-Doppler space-time characteristics Used to implement step 2 of the gesture key point detection method of the present invention;

[0157] The micro-Doppler feature extraction submodule extracts the micro-Doppler spectrum X MD Spectral characteristics, get micro-Doppler spectrum characteristics Used to implement step 3 of the gesture key point detection method of the present invention;

[0158] Cross-modal feature fusion module combines distance-Doppler spatiotemporal features and micro-Doppler spectrum characteristics Perform cross-modal feature fusion to obtain cross-modal fusion features Used to implement step 4 of the gesture key point detection method of the present invention;

[0159] Key point detection decoder module, design key point detection decoder based on modal fusion features Estimated gesture keypoint prediction coordinate sequence {H topk} (m) , used to implement step 4 of the gesture key point detection method of the present invention;

[0160] Model training and testing module, predicting coordinate sequences by screening the top m gesture key points with the highest confidence scores With the key point true value sequence {H gt} (m) Perform the optimal matching of the Hungarian algorithm to obtain the overall regression loss L reg , and perform N, N ≥ 1 rounds of iteration to complete the gesture key point detection model training, which is used to implement steps 6 and 7 of the gesture key point detection method described in the present invention.

[0161] Millimeter-wave radar gesture key point detection equipment based on dual-stream deep fusion, including:

[0162] Memory: used for storing a computer program for implementing a millimeter-wave radar gesture key point detection method based on dual-stream deep fusion;

[0163] Processor: used to implement a millimeter-wave radar gesture key point detection method based on dual-stream deep fusion when executing the computer program.

[0164] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a millimeter-wave radar gesture key point detection method based on dual-stream deep fusion.

[0165] Verification experiment

[0166] Data source:

[0167] The gesture dataset was collected using TI's IWR1443 millimeter-wave radar. It includes 10 subjects, 30 refined gestures (such as numbers 1-9, index finger and thumb pinch, three-finger zoom, index finger click, pinch knob, and thumbs up), and 100 repetitions of each gesture. Each repetition collects 25 frames of RDM sequence and the corresponding micro-Doppler spectrum. Part of the dataset is shown below. Figure 3 As shown in the figure, 70% of the gesture dataset is divided into a training set, 20% is divided into a validation set, and 10% is divided into a test set.

[0168] During the training phase, the preprocessed training set is first aligned and fed into the two-stream network. The gesture keypoint detection model is trained using the Adam optimizer (initial learning rate 1e–3, weight decay 1e–4). The overall loss function designed in step 6 guides the optimization of model parameters. The batch size is set to 16, and training is performed for N = 300 epochs.

[0169] During training, a cosine annealing learning rate schedule is used, and early stopping is performed after 10 consecutive rounds of validation set performance improvement to avoid overfitting. To enhance model robustness, data augmentation such as mild time shifts and spectral noise are only applied to the input sequence during training.

[0170] During the testing phase, the main indicators include: MPJPE (Mean Per Joint Position Error), which reflects the average positioning error of key points; RMSE (Root Mean Square Error), which measures the overall prediction deviation;

[0171] index The present invention Comparative Example 1 Comparative Example 2 MPJPE(mm) 42.8 62.7 78.5 RMSE(mm) 48.5 68.2 77.4 Maximum error of each key point (mm) 85.6 98.3 103.3 Minimum error of each key point (mm) 2.1 16.5 17.2

[0172] Comparative Example 1 shows the EgoHand in this paper. Its best performance of 62.7mm was achieved after integrating an additional hardware IMU. In experiments using radar data alone, the error was 76.4mm. The method proposed in this paper achieved an error of 42.8mm, significantly outperforming EgoHand's radar-only solution and demonstrating the algorithmic superiority of the dual-stream deep fusion architecture.

[0173] Comparative Example 2 is a patent application with publication number CN118409311A: This patent application describes in its document that "the comprehensive error can reach about 3 cm", which is the result of testing in a small number of specific data sets constructed by the invention. However, in the data set constructed in the present invention with a larger amount of data and richer action categories, the comprehensive error of the patented method is 78.5 mm. The method of the present invention is also significantly superior to the patent in terms of test error.

Claims

1. A millimeter-wave radar gesture key point detection method based on dual-stream deep fusion, characterized in that: The steps include: Step 1: The millimeter-wave radar is used to collect the original hand motion signal, and the RGB camera is used to synchronously capture multiple frames of hand images. The collected data is then preprocessed to obtain the range-Doppler three-dimensional space-time tensor X RDM , Micro-Doppler spectrum X MD And the key point true value sequence {H gt } (m) ; Step 2: Extract the range-Doppler three-dimensional space-time tensor X RDM Features, get the range-Doppler space-time characteristics Step 3: Extract the micro-Doppler spectrum X MD Spectral characteristics, get micro-Doppler spectrum characteristics Step 4: transform the range-Doppler spatiotemporal characteristics and micro-Doppler spectrum characteristics Perform cross-modal feature fusion to obtain cross-modal fusion features Step 5: Design a key point detection decoder based on modality fusion features Estimated gesture keypoint prediction coordinate sequence {H topk } (m) ; Step 6: Predict the coordinate sequence of the top m gesture key points by filtering the ones with the highest confidence scores The key point truth sequence {H gt } (m) Perform the optimal matching of the Hungarian algorithm to obtain the overall regression loss L reg ; Step 7: Repeat steps 2 to 6 for N, N ≥ 1 rounds to complete the training of the gesture key point detection model.

2. The gesture key point detection method according to claim 1, characterized in that: The specific steps of data preprocessing in step 1 are as follows: First, apply a window function to each frame of the original hand motion signal and perform range FFT and Doppler FFT operations to obtain the range-Doppler graph of n frames. Then the continuous n frames of range-Doppler map Superposition to form the range-Doppler three-dimensional space-time transition tensor Among them, H is the distance resolution dimension, W is the velocity resolution dimension; at the same time, the range-Doppler three-dimensional space-time transition tensor X' RDM Perform short-time Fourier transform (STFT) in the time domain to generate the corresponding micro-Doppler transition spectrum Among them, T is the number of time frames, T = n, F is the frequency resolution; finally, the range-Doppler three-dimensional space-time transition tensor X' RDM and micro-Doppler transition spectrum X' MD Through vector mean elimination processing, the range-Doppler three-dimensional space-time tensor X is obtained RDM and Micro Doppler Spectrum X MD ; Use MediaPipe to process n frames of hand images and generate 3D coordinate labels for a hand key points: Among them, x, y, z represent the 3D coordinates of the key points of the hand; 3D coordinate labels and the range-Doppler three-dimensional space-time tensor X RDM and Micro Doppler Spectrum X MD Timing alignment, unified to the m-frame key point true value sequence through linear interpolation 3. The gesture key point detection method according to claim 1, characterized in that: The specific steps of step 2 are as follows: Construct a hybrid architecture that combines a three-dimensional convolutional neural network (3D-CNN) with a temporal convolutional network (TCN): Range-Doppler three-dimensional space-time tensor X RDM First, the preliminary features are extracted through the three-dimensional convolutional neural network (3D-CNN). The formula is: in, is the convolution kernel weight of this layer, is bias; And through normalization and nonlinear activation, we get the nonlinear normalized spatiotemporal features Then, the nonlinear normalized spatiotemporal features W' are downsampled in the spatial dimension to obtain high-level semantic features Then the high-level semantic features W1 are globally averaged in the spatial dimension to obtain the time series The formula is: Finally, the time series S is fed into a l-layer temporal convolutional network (TCN) to output the range-Doppler spatiotemporal features. The formula is: in, D' is the spatiotemporal feature dimension.

4. The gesture key point detection method according to claim 1, characterized in that: The specific steps of step 3 are as follows: First, multi-layer two-dimensional convolution is used to extract the micro-Doppler spectrum X MD The local time-frequency texture feature U1' is as follows: in, is the first layer convolution weight, is the bias, C is the output channel; And through normalization and nonlinear activation, the nonlinear normalized spectral feature U(t,f,c) is obtained; The maximum pooling in frequency dimension is adopted, and the formula is: Φ'(t,f',c)=max{U(t,2f',c),U(t,2f'+1,c)} Where, f'=0...[F / 2]-1, Then, a 1×1 convolution kernel is used to expand the number of channels after pooling to obtain the spectrum extension feature. Finally, the spectrum expansion feature Φ is averaged in the frequency dimension The final output is the micro-Doppler spectrum characteristics 5. The gesture key point detection method according to claim 1, characterized in that: The specific steps of step 4 are as follows: Range-Doppler spatiotemporal characteristics and micro-Doppler spectrum characteristics Do linear mapping to get key features Value Features Query Features Among them, D is the fusion attention feature dimension; then the similarity score is performed And normalize the similarity to get the attention weight Then aggregate the value feature V RDM Get aggregated attention features Finally, residual connections and layer normalization are introduced to output cross-modal fusion features 6. The gesture key point detection method according to claim 1, characterized in that: The specific steps of step 5 are as follows: First, define a set of learnable pose query feature vectors D is the fusion attention feature dimension, where each pose query feature q i Random initialization and joint optimization with the training process, the number of queries is designed to be I to cover the redundant candidate poses in different interaction scenarios; The key point detection decoder is a multi-layer decoder stacking structure. Specifically, each decoding layer cascades the self-attention module and the cross-attention module, with a total of L layers stacked. The calculation process of the l∈L layer is: (Self-Attention) (cross-attention) in, For learnable position encoding, in MultiHeadAttn(·), the input parameters will undergo a linear transformation. is the temporal mask matrix: Among them, Δ represents the threshold of the number of frames before and after attention; After L layers of decoding, the posture decoding feature vector is obtained Obtain candidate key point coordinates H through MLP mapping cand With confidence S cand : Among them, σ is the Sigmoid function, a represents the number of key points of the hand, 3 is the 3D coordinate dimension of the key points of the hand, and by selecting the confidence S cand The coordinates of the first m, m≤I candidate key points H cand , as the final gesture key point prediction coordinate sequence {H topk } (m) .

7. The gesture key point detection method according to claim 1, characterized in that: The specific steps in step 6 are as follows: First, define the gesture key point prediction coordinate sequence {H topk } (m) With the key point true value sequence {H gt } (m) Any pair of gesture key point prediction coordinates H topk and the key point truth value H gt The key point loss between It calculates the L2 distance between a corresponding keypoints: in, and Predict the coordinates H for any pair of gesture key points topk and the key point truth value H gt The 3D coordinates of the jth key point in ; Then, the Hungarian algorithm is used to find the coordinates H that can make m predict the key points of the gesture topk and the key point truth value H gt The matching permutation with the smallest total loss: in, is the set of all possible permutations of the set {1,2,…,m}, is the predicted coordinate sequence of gesture key points {H topk } (m) The predicted coordinates of the i-th gesture key point in , is the key point true value sequence {H gt } (m) The σ(i)th one in , where σ(i) is Permutation of the i-th item in the set; Finally, by arrangement The minimum total loss obtained is the overall regression loss L of the gesture key point detection model reg :

8. A gesture key point detection system based on the method according to any one of claims 1 to 7, characterized in that: include: The data preprocessing module collects the original hand motion signal through the millimeter wave radar and uses the RGB camera to synchronously capture multiple frames of hand images. The collected data is then preprocessed to obtain the range-Doppler three-dimensional space-time tensor X RDM , Micro-Doppler spectrum X MD And the key point true value sequence {H gt } (m) ; Dual-stream feature extraction module, including a range-Doppler map spatiotemporal feature extraction submodule and a micro-Doppler feature extraction submodule; The range-Doppler image spatiotemporal feature extraction submodule extracts the range-Doppler three-dimensional spatiotemporal tensor X RDM Features, get the range-Doppler space-time characteristics The micro-Doppler feature extraction submodule extracts the micro-Doppler spectrum X MD Spectral characteristics, get micro-Doppler spectrum characteristics Cross-modal feature fusion module combines distance-Doppler spatiotemporal features and micro-Doppler spectrum characteristics Perform cross-modal feature fusion to obtain cross-modal fusion features Key point detection decoder module, design key point detection decoder based on modal fusion features Estimated gesture keypoint prediction coordinate sequence {H topk } (m) ; Model training and testing module, predicting coordinate sequences by screening the top m gesture key points with the highest confidence scores With the key point true value sequence {H gt } (m) Perform the optimal matching of the Hungarian algorithm to obtain the overall regression loss L reg , and perform N, N ≥ 1 rounds of iteration to complete the training of the gesture key point detection model.

9. Millimeter-wave radar gesture key point detection device based on dual-stream deep fusion, characterized by: include: Memory: used to store a computer program for implementing the method for detecting key points of millimeter-wave radar gestures based on dual-stream deep fusion as described in any one of claims 1 to 7; Processor: used to implement the millimeter wave radar gesture key point detection method based on dual-stream deep fusion as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the millimeter-wave radar gesture key point detection method based on dual-stream deep fusion as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Millimeter wave radar hand key point detection system and method based on unsupervised feature learning

    CN118409311A

Cited By

  • Non-contact personnel identity and posture recognition method based on motion instability compensation

    CN122110050A