A Millimeter-Wave Radar Gesture Recognition Method and System Based on Multi-Domain Spectral Map and Multi-Resolution Fusion

The millimeter-wave radar gesture recognition method, which integrates multi-domain spectral maps and multi-resolution data, generates multi-dimensional radar feature spectral maps and uses convolutional networks for feature fusion. This solves the problems of low efficiency, strong environmental dependence, and privacy leakage in existing gesture recognition technologies, and achieves high-precision non-contact gesture recognition.

CN116184394BActive Publication Date: 2026-03-13CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-06
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing gesture recognition technologies suffer from low recognition efficiency, strong environmental dependence, privacy risks, and insufficient recognition accuracy. In particular, non-contact gesture recognition performs poorly at long distances and in complex environments.

Method used

A millimeter-wave radar gesture recognition method based on multi-domain spectral mapping and multi-resolution fusion is adopted. Frequency, distance, velocity and horizontal angle spectral maps are generated through radar signal processing. Feature extraction and fusion are performed using a two-dimensional convolutional network and a multi-resolution fusion module to construct a gesture recognition network and achieve real-time classification of gestures.

Benefits of technology

It achieves efficient and accurate non-contact gesture recognition in complex environments, improves recognition accuracy and generalization ability, protects user privacy, and is not limited by lighting or usage scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116184394B_ABST
    Figure CN116184394B_ABST
Patent Text Reader

Abstract

This invention discloses a millimeter-wave radar gesture recognition method and system based on multi-domain spectral mapping and multi-resolution fusion. Targeting radar echo signals of human gestures, the method processes the signals using short-time Fourier transform, pulse compression, two-dimensional fast Fourier transform, and minimum variance distortion-free response beamforming algorithms to generate four types of radar spectra with different physical meanings and complementary features. A gesture recognition network based on a two-dimensional convolutional network and a multi-resolution fusion module is then constructed to achieve human gesture recognition. This network divides the radar multi-domain spectral maps into four different radar feature spectra. This method enables a more comprehensive and sufficient representation of radar multi-domain features under limited data conditions. The two-dimensional convolutional network has strong feature extraction capabilities, and the multi-resolution fusion module can acquire multi-resolution feature tensors, resulting in a high human gesture recognition rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of radar signal processing technology, and in particular to a millimeter-wave radar gesture recognition method and system based on multi-domain spectral mapping and multi-resolution fusion. Background Technology

[0002] Gesture recognition is an important research direction in the field of radar signal processing and applications, and it is widely used in smart homes, autonomous driving, and smart healthcare. According to reports, the global gesture recognition market reached 89.957 billion yuan (RMB) in 2021, and is projected to reach 305.106 billion yuan by 2027, growing at a compound annual growth rate of 22.9%. Therefore, research on gesture recognition technology has significant economic benefits.

[0003] Gesture recognition, as a contactless interaction method, allows control of smart electrical devices without touching them. Compared to traditional keyboards or touchscreen interfaces, which inevitably lead to device wear and tear and pose a risk of disease transmission, gesture recognition technology allows for command input without direct touch or other contact with the device. This contactless approach reduces device wear and tear, while also minimizing the spread of viruses and bacteria and lowering the risk of infection. Therefore, research into gesture recognition technology can effectively extend device lifespan, reduce the risk of user infection through contact, significantly improve the user experience, and protect users' physical and mental health.

[0004] Currently, gesture recognition technology is divided into two categories: contact and non-contact. Contact gesture recognition technology mainly utilizes wearable devices, which are equipped with attitude sensors such as accelerometers and gyroscopes to monitor the user's hand posture information in real time, and then perform gesture recognition based on the collected posture data. This method is characterized by its insensitivity to environmental influences and high recognition accuracy, but it requires the user to wear the detection device at all times, resulting in limited application scenarios and inconvenience in carrying it, which greatly affects the user experience. Non-contact gesture recognition technology mainly includes technologies based on visual sensors, ultrasound, and Wi-Fi. Gesture recognition based on visual sensors is limited by lighting conditions and is not suitable for extreme environments, and there is a risk of leaking personal privacy data. Ultrasonic sensors have limited detection distance and cannot meet the needs of long-distance gesture recognition. Wi-Fi is severely affected by environmental noise interference, resulting in low recognition accuracy. Therefore, in recent years, research in related fields has gradually focused on gesture recognition technology based on invisible imaging.

[0005] Therefore, a gesture recognition method with high recognition efficiency is needed. Summary of the Invention

[0006] In view of this, the purpose of this invention is to provide a millimeter-wave radar gesture recognition method based on multi-domain spectral image and multi-resolution fusion. This method uses radar signals to realize gesture recognition. This method is a millimeter-wave radar gesture recognition method that integrates multi-domain spectral image features and multi-resolution fusion, which has sufficient feature expression, high efficiency and high recognition accuracy.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] The present invention provides a millimeter-wave radar gesture recognition method based on multi-domain spectral mapping and multi-resolution fusion, comprising the following steps:

[0009] (1) Identify the targets to be detected within the radar detection area;

[0010] (2) Collect radar echo signals returned by the target, generate multiple radar spectra with different physical meanings and complementary features for the same target and the same gesture, and construct a dataset based on different targets and different gestures;

[0011] (3) Design a gesture recognition network based on the parallel input of multi-domain spectrograms, which is a spatiotemporal feature extraction and multi-domain spectrogram complementary feature fusion network in series with a two-dimensional convolutional network and a multi-resolution fusion module. Then, use the dataset to train and test the network, and design feature extraction and gesture recognition modules based on this network.

[0012] (4) In the feature extraction and gesture recognition module, the target gesture is classified according to the input fusion features and the target gesture action is judged in real time; if a preset gesture is detected, proceed to step (5); if no preset gesture is detected, return to step (4); the preset gestures include five gestures: waving to the left, waving to the right, waving upward, waving downward and pushing forward.

[0013] (5) Send the recognized gesture information.

[0014] Furthermore, the radar spectrum in step (2) includes frequency, range, velocity and horizontal angle spectra.

[0015] Furthermore, step (2) specifically involves:

[0016] (21) Preprocess the raw echo signals acquired by the radar;

[0017] (22) Perform short-time Fourier transform on the preprocessed signal and stack it in the slow time dimension to obtain the time-frequency characteristic expression of the echo signal;

[0018] (23) The preprocessed signal is pulse compressed, stacked in the slow time dimension, and the distance feature expression is obtained by accumulating through multiple receiving channels;

[0019] (24) Perform a two-dimensional fast Fourier transform on the preprocessed signal and stack them sequentially in the slow time dimension to obtain the velocity feature expression;

[0020] (25) The preprocessed signal is processed by the minimum variance distortionless response beamforming algorithm and stacked sequentially in the slow time dimension to obtain the horizontal angle feature representation.

[0021] Furthermore, step (21) specifically involves:

[0022] Static clutter suppression is performed on the original echo signal. Since the phase of the moving target in the echo signal differs at different time intervals, the current echo signal is canceled out with the echo signal at an interval of τ to obtain the target echo signal. The calculation formula is as follows:

[0023]

[0024] in, Indicates the slow time index within the frequency modulation period. for The radar echo signal at any given moment.

[0025] Furthermore, step (22) specifically involves:

[0026] First, the preprocessed signal is segmented, and windowing is applied to each segment. Then, a short-time Fourier transform is performed on the windowed data. Finally, the results of processing each segment are stacked in the slow-time dimension to obtain the frequency characteristic expression of the echo signal, which is calculated using the following formula:

[0027]

[0028] in, It is the target echo signal. For window functions, To obtain the short-time Fourier transform result, This indicates the slice length near the analysis time point t.

[0029] Furthermore, step (23) specifically involves:

[0030] First, the preprocessed signal is split into multiple complete chirped signals. Then, pulse compression is performed on each chirped signal to obtain the range profile of a single chirped signal. Finally, the range profiles of each chirped signal are stacked in a slow time dimension according to time, and the range feature expression of the echo signal is obtained by accumulating through multiple receiving channels. The calculation method is as follows:

[0031]

[0032] in, Representing the A chirp signal in The amplitude of the point; The window function is represented, and the window function is a Hamming window; Represents the number of sampling points; This represents the m-th chirp signal. The values ​​of each sampling point;

[0033] Secondly, if the target is far from the radar, the echo signal energy may be low, resulting in indistinct differences between the feature information and the background. To address this, a multi-channel data accumulation method is used, as detailed below:

[0034]

[0035] in, This represents the distance spectrum obtained from the Cth receiving channel;

[0036] Finally, the distance spectra of these different receiving channels are accumulated to form the final distance feature expression.

[0037] Furthermore, step (24) specifically involves:

[0038] First, the original signal is windowed, and a Fast Fourier Transform (FFT) is performed in the fast time dimension to obtain the range profile. Then, the range profile is windowed again and a FFT is performed in the slow time dimension to obtain the velocity spectrum. Finally, these velocity spectra are stacked in chronological order in the slow time dimension to obtain the velocity feature representation. The calculation method is as follows:

[0039]

[0040] in, This represents the amplitude at point s in the k-th row of the distance spectrum. This represents the amplitude in the k-th row and mi-th column of the f-th frame. The number of chirped signals contained in each frame.

[0041] Furthermore, step (25) specifically involves:

[0042] First, the original signal is windowed, and then a fast Fourier transform is performed in the fast time dimension to obtain the range image. The range image is then processed using the minimum variance distortion-free response beamforming algorithm to obtain a distortion-free output of the target azimuth signal. The calculation method is as follows:

[0043] ,

[0044] in, Let represent the echo signal energy at position θ in the m-th chirp signal, where n is the number of array elements. The autocorrelation matrix representing the distance spectrum of the m-th chirped signal is:

[0045]

[0046] For the guide vector, that is:

[0047]

[0048] Finally, by stacking these horizontal angle spectra in chronological order along the slow time dimension, the horizontal angle feature representation can be obtained.

[0049] Furthermore, step (3) specifically involves:

[0050] First, the multi-class feature spectra are uniformly reshaped into tensors of the same size, and then mapped to fixed-size feature tensors using a linear projection network. The obtained feature tensors are input into a two-dimensional convolutional network structure to abstract the spectra features. Then, the abstracted features are input into a multi-resolution fusion module to generate multi-resolution feature tensors. The multi-resolution feature tensors of the multi-domain spectra are fused to obtain complementary fused features from the multi-domain spectra. Finally, logistic regression is used to classify the features.

[0051] The present invention provides a millimeter-wave radar gesture recognition system based on multi-domain spectral and multi-resolution fusion, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the above-mentioned method.

[0052] The beneficial effects of this invention are as follows:

[0053] This invention provides a millimeter-wave radar gesture recognition method and system based on multi-domain spectral and multi-resolution fusion. The method acquires radar-detected targets within the detection area, collects radar echo signals returned by the targets, and uploads them to the system's gesture representation module. For the same target and the same gesture, four types of radar spectral maps with different physical meanings and complementary features are generated, including frequency, range, velocity, and horizontal angle spectral maps. A dataset is constructed based on different targets and different gestures. A gesture recognition network is established, and the dataset is used to train and test the network. A feature extraction and gesture recognition module is designed based on this network. In the feature extraction and gesture recognition module, the target gestures are classified according to the input multi-resolution fusion features, and the target's current gesture action is determined in real time. Gesture information is sent to a host computer via a wireless data transmission module. The host computer controls external devices to perform corresponding functions based on different gesture information, thereby achieving real-time gesture recognition.

[0054] This invention targets radar echo signals of human gestures. By processing the signals through short-time Fourier transform, pulse compression, two-dimensional fast Fourier transform, and minimum variance distortion-free response beamforming algorithm, four types of radar spectra with different physical meanings and complementary features are generated. A gesture recognition network based on spatiotemporal feature extraction and multi-domain spectra complementary feature fusion based on a two-dimensional convolutional network and a multi-resolution fusion module is constructed to achieve gesture recognition.

[0055] This method utilizes an intelligent recognition network that employs the layer-by-layer abstraction concept found in CNNs. This network reshapes the feature spectrum into tensors of uniform size and maps them to fixed-size feature tensors using a linear projection network. The acquired feature tensors are then input into a two-dimensional convolutional network structure to abstract the spectral features. These abstract features are then fed into a multi-resolution fusion module for multi-resolution feature tensor fusion, which is used in a classifier to achieve gesture recognition. By combining the multi-resolution fusion module with the two-dimensional convolutional network module, the method enables multi-resolution feature fusion of any type or different types of features, significantly enhancing the network's feature fusion and generalization capabilities. This method achieves more comprehensive and sufficient multi-domain feature representation of radar under limited data conditions. The two-dimensional convolutional network possesses strong feature extraction capabilities, and the multi-resolution fusion module can acquire multi-resolution feature tensors, resulting in a high human gesture recognition rate.

[0056] This invention utilizes millimeter-wave radar for gesture recognition, offering advantages such as protecting user privacy, being unaffected by ambient light, and not being limited by usage scenarios. Employing a multi-domain spectral map generation method that generates complementary multi-feature representations enables a more comprehensive and sufficient expression of gesture features even with limited data. Multi-channel accumulation is used to improve the imaging quality of the radar spectral map, thereby increasing the gesture recognition rate. A composite neural network consisting of a parallel two-dimensional convolutional network and a multi-resolution module is designed to achieve multi-resolution radar multi-domain spectral map feature fusion. Features are extracted from different domains using a multi-channel parallel input method, and the multi-resolution fusion module fuses gesture features from different resolutions, resulting in better recognition capabilities than a two-dimensional convolutional network. This effectively improves the network's generalization ability and recognition accuracy.

[0057] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0058] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the following figures are provided for illustration:

[0059] Figure 1 The flowchart shows a millimeter-wave radar gesture recognition method based on multi-domain spectral mapping and multi-resolution fusion.

[0060] Figure 2 Schematic diagram of millimeter-wave radar installation;

[0061] Figure 3 This is a schematic diagram of radar multi-domain spectral characterization technology.

[0062] Figure 4 This is a millimeter-wave radar gesture recognition network structure based on multi-domain spectral mapping and multi-resolution fusion;

[0063] Figure 5 The confusion matrix obtained for testing the network is represented by A1, A2, A3, A4, and A5, respectively, where waving to the left, waving to the right, waving upwards, waving downwards, and pushing forward.

[0064] Figure 6 This is a block diagram of a non-line-of-sight gesture recognition system. Detailed Implementation

[0065] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.

[0066] Example 1

[0067] As an environmental data acquisition sensor, radar not only effectively avoids infringing on user privacy but also possesses unique advantages such as high penetration capability, high range resolution, high velocity resolution, and high angular resolution. It is particularly advantageous in large-scale application scenarios for detecting and tracking human hand movements, and has broad application prospects. The human gesture recognition method based on millimeter-wave radar mainly includes three steps: feature representation, feature extraction, and classification. First, the millimeter-wave radar echo signal is processed to generate a radar image containing multi-dimensional feature information. Then, gesture feature information contained in the radar image is extracted manually or automatically through neural networks. Finally, a classifier is designed to classify gestures based on the feature information, thereby recognizing and classifying hand gestures.

[0068] Radar human gesture recognition often utilizes methods such as time-frequency analysis, pulse compression, two-dimensional Fourier transform, and minimum variance distortion-free beamforming to generate time-spectrum maps, range-spectrum maps, velocity-spectrum maps, and horizontal angle-spectrum maps to achieve multi-dimensional representation of gesture features. The time-spectrum map can be viewed as a power spectrum sequence that changes over time, reflecting the Doppler characteristics of the human gesture target and the positional parameters of the scattering center. Different signal processing methods can generate feature spectra maps with different physical meanings. Feature spectra maps in different domains differ in their geometric details and semantic information representation of the same gesture, but they also exhibit significant complementarity in feature representation. Existing research often uses single-class spectra maps for feature representation. However, to fully explore and utilize the complementary features between radar spectra maps in different domains, this method utilizes multi-domain radar spectra maps to more fully represent gesture features in more dimensions, thereby improving the recognition accuracy of human gestures.

[0069] Feature extraction methods in radar human gesture recognition mainly include manual extraction and automatic extraction via neural networks. Manual feature extraction involves manually designing feature extraction methods to annotate effective information from raw data and use it as the basis for gesture judgment. However, in practice, manual feature extraction requires a high level of professional knowledge from practitioners and is difficult to capture effective discrimination information from the raw radar spectrum, resulting in low efficiency and high operational complexity.

[0070] This embodiment utilizes radar signal processing and deep learning technologies to achieve human gesture recognition. It employs methods such as STFT, pulse compression, 2D-FFT, and MVDR to extract frequency, distance, speed, and horizontal angle features of different target gestures, generating radar feature spectra in four different domains. A gesture recognition network based on a cascaded two-dimensional convolutional network and a multi-resolution fusion module performs feature extraction, multi-resolution feature fusion, multi-domain feature fusion, and gesture classification, ultimately achieving human gesture recognition. This method features user privacy protection, is unaffected by ambient light, is not limited by usage scenarios, and provides sufficient feature representation, achieving a superior gesture recognition rate.

[0071] like Figure 1 As shown, this embodiment provides a non-line-of-sight human gesture recognition method based on multi-domain feature fusion using millimeter-wave radar, including the following steps:

[0072] (1) A millimeter-wave radar is installed in the detection area, and there are a certain number of targets within the radar coverage area;

[0073] (2) Real-time monitoring of targets within the detection area is performed using millimeter-wave radar sensors, and radar echo signals returned by the targets are collected and uploaded to the system's gesture representation module. Four types of radar spectra with different physical meanings and complementary features are generated for the same target and the same gesture. The radar spectra include time-frequency spectrum, time-range spectrum, time-velocity spectrum, and time-angle spectrum. Figure 4 The combination of spectrograms is used to construct datasets based on different gestures for different targets.

[0074] This embodiment includes time-frequency spectrum, time-distance spectrum, time-velocity spectrum, and time-angle spectrum. Figure 4 The combination of various spectra forms this feature information, which is temporal in nature and can represent target information (frequency, range, velocity, angle) at different times. This overcomes the shortcomings of non-temporal feature information, such as range-angle, which cannot identify the specific distance and angle of a target at each point in time. Especially for easily confused gestures, using range-angle spectra can result in very similar imaging results, making differentiation impossible. However, in the temporal dimension, they exhibit different characteristics. Range-Doppler only indicates the existence of a target object at that distance, but not its movement. However, by stacking these spectra in the temporal dimension, it can be observed that the target object is moving forward and approaching the radar.

[0075] The four radar spectra possess both independent and complementary features. When a single type of spectra is insufficient to distinguish certain gestures, radar spectra from other radar domains are used to supplement the gesture features, resulting in a more comprehensive and complete representation. For example, waving to the left and waving to the right exhibit extremely similar features in the range spectra. Adding horizontal angle features as a discrimination criterion significantly improves the recognition rate of confused gestures, as the features of these two gestures are completely opposite in the horizontal angle spectra. This channel-level fusion allows for a more comprehensive and complete representation of gesture characteristics, thereby enhancing gesture recognition accuracy.

[0076] (3) Construct a gesture recognition network. The gesture recognition network is a gesture recognition network based on the parallel input of multi-domain spectrograms, which is a spatiotemporal feature extraction and multi-domain spectrogram complementary feature fusion based on a two-dimensional convolutional network and a multi-resolution fusion module. The network is then trained and tested using a dataset. Based on this gesture recognition network, there are feature extraction and gesture recognition modules.

[0077] (4) In the feature extraction and gesture recognition module, the target gesture is classified according to the input fusion features and the target gesture action is judged in real time; if a preset gesture is detected, proceed to step (5); if no preset gesture is detected, return to step (4); the preset gestures include five gestures: waving to the left, waving to the right, waving upward, waving downward and pushing forward.

[0078] (5) Send gesture information to the host computer through the wireless data transmission module. The host computer controls the external device to perform the corresponding functions according to different gesture information, so as to realize real-time gesture recognition.

[0079] like Figure 2 As shown, in step (1) above, in order to enable the radar to achieve the best measurement effect, the millimeter-wave radar is installed on the wall at a height of 1.0m to 2.0m; the angle with the vertical direction is about -25° to 25°; in this embodiment, it is preferred to be 2.0m above the ground and tilted downward at 25°. The millimeter-wave radar is the Texas Instruments IWR 6843 ISK FMWC millimeter-wave radar.

[0080] like Figure 3 As shown, the specific steps of step (2) above are as follows:

[0081] (21) Preprocess the raw echo signals acquired by the radar;

[0082] (22) Perform short-time Fourier transform on the preprocessed signal and stack it in the slow time dimension to obtain the time-frequency characteristic expression of the echo signal;

[0083] (23) The preprocessed signal is pulse compressed, stacked in the slow time dimension, and the distance feature expression is obtained by accumulating through multiple receiving channels;

[0084] (24) Perform a two-dimensional fast Fourier transform on the preprocessed signal and stack them sequentially in the slow time dimension to obtain the velocity feature expression.

[0085] (25) The preprocessed signal is processed by the minimum variance distortionless response beamforming algorithm and stacked sequentially in the slow time dimension to obtain the horizontal angle feature representation.

[0086] The specific steps (21) are as follows:

[0087] Static clutter suppression is performed on the original echo signal. Here, a moving target indication algorithm is used. The phase of the moving target in the echo signal is different at different times. The target echo signal is obtained by canceling the current echo signal with the echo signal at an interval of time τ. The calculation formula is as follows:

[0088]

[0089] in, Indicates the slow time index within the frequency modulation period. for Radar echo signal at any given moment;

[0090] The specific steps (22) are as follows:

[0091] First, the preprocessed signal is segmented, and windowing is applied to each segment. Then, a short-time Fourier transform is performed on the windowed data. Finally, the results of processing each segment are stacked in the slow-time dimension to obtain the frequency characteristic expression of the echo signal, which can be calculated using the following formula:

[0092]

[0093] in, It is the target echo signal. For window functions, To obtain the short-time Fourier transform result, This indicates the slice length near the analysis time point t.

[0094] The specific steps (23) are as follows:

[0095] First, the preprocessed signal is split into multiple complete chirped signals. Then, pulse compression is performed on each chirped signal to obtain the range profile of a single chirped signal. Finally, the range profiles of each chirped signal are stacked in a slow time dimension according to time to obtain the range feature representation of the echo signal. The calculation method is as follows:

[0096]

[0097] in, Representing the A chirp signal in The amplitude of the point; The window function used in this method is the Hamming window. Represents the number of sampling points; This represents the m-th chirp signal. The value of each sampling point.

[0098] Secondly, if the target is far from the radar, the echo signal energy may be low, resulting in indistinct differences between the target's features and the background. To enhance the target's feature information, a multi-channel data accumulation method is used, the specific implementation of which is as follows:

[0099]

[0100] in, This represents the distance spectrum obtained from the Cth receiving channel.

[0101] Finally, the distance spectra of these different receiving channels are accumulated to form the final distance feature expression.

[0102] The specific steps (24) are as follows:

[0103] First, the original signal is windowed, and a Fast Fourier Transform (FFT) is performed in the fast time dimension to obtain the range profile. Then, the range profile is windowed again and a FFT is performed in the slow time dimension to obtain the velocity spectrum. Finally, these velocity spectra are stacked in chronological order in the slow time dimension to obtain the velocity feature representation. The calculation method is as follows:

[0104]

[0105] in, This represents the amplitude at point s in the k-th row of the distance spectrum. This represents the amplitude of the k-th row and mi-th column of the f-th frame. The number of chirped signals contained in each frame.

[0106] The specific steps (25) are as follows:

[0107] First, the original signal is windowed, and then a fast Fourier transform is performed in the fast time dimension to obtain the range image. The range image is then processed using the minimum variance distortion-free response beamforming algorithm to obtain a distortion-free output of the target azimuth signal. The calculation method is as follows:

[0108]

[0109] in, This represents the echo signal energy at position θ in the m-th chirped signal. n is the number of array elements. The autocorrelation matrix representing the distance spectrum of the m-th chirped signal is:

[0110]

[0111] The guide vector, i.e.:

[0112]

[0113] Finally, stacking these horizontal angle spectra in chronological order along the slow time dimension yields the horizontal angle feature representation.

[0114] The specific steps (3) above are as follows:

[0115] like Figure 4 As shown, Figure 4The present invention relates to a millimeter-wave radar gesture recognition network structure based on multi-domain spectral mapping and multi-resolution fusion; the gesture recognition network includes a linear projection layer, a feature abstraction layer, a multi-resolution feature fusion layer, a semantic information extraction layer, a complementary feature fusion layer, and a classifier.

[0116] The linear projection layer is used to map the multi-domain feature spectrum into a feature tensor of fixed size;

[0117] The feature abstraction layer is used to perform feature abstraction on the radar multi-domain spectrum.

[0118] The multi-resolution feature fusion layer is used to fuse feature tensors at different resolutions;

[0119] The semantic information extraction layer is used to obtain deeper semantic information, reduce the feature map size, and improve the system running speed.

[0120] The complementary feature fusion layer is used to fuse radar feature tensors from different domains;

[0121] The classifier is used to classify the fused features to obtain a classification result;

[0122] The gesture recognition network provided in this embodiment uses a parallel two-dimensional convolutional network module and a multi-resolution fusion mechanism for feature extraction and fusion. For radar multi-domain feature spectra, the multi-domain feature spectra include time-frequency spectra, time-range spectra, time-velocity spectra, and time-angle spectra. A linear projection layer is used to map the multi-domain feature spectra into feature tensors of a fixed size. The obtained H×W×C feature tensors are input into the feature abstraction layer to realize feature abstraction of radar feature spectra. The multi-resolution feature fusion layer fuses feature tensors at different resolutions to obtain multi-resolution fused features. Then, the multi-resolution features are input into the semantic information extraction layer.

[0123] In this embodiment, deeper semantic information is obtained through cross-processing of convolution and normalization. This operation can reduce the feature map size to improve the system running speed and ensure the real-time performance of the recognition system. Finally, it is cross-fused with features from other domains in the complementary feature fusion layer, that is, features from different domains are stacked in the channel dimension so that the fused features contain different physical meanings, so as to express the gesture features more comprehensively.

[0124] The fused features are then input into a logistic regression model for classification, ultimately achieving human gesture recognition.

[0125] exist Figure 4In the network structure, the linear projection module aims to reduce the dimensionality of high-dimensional features, while the two-dimensional convolutional network in the feature abstraction module is responsible for abstracting features from the radar multi-domain spectral map. In other words, the multi-resolution feature fusion layer module fuses feature tensors at different resolutions. Compared to the semantic features of a single-resolution feature tensor, the feature tensor processed by the multi-resolution feature fusion module contains both spatial texture features of gestures and deeper semantic features, effectively improving the recognition rate of easily confused gestures.

[0126] In this embodiment, the complementary feature fusion layer is used to cross-fuse multi-domain spectral features. Its function is to fuse radar feature tensors from different domains, and finally, the fused features are classified by a classifier to obtain the classification result, thereby realizing human gesture recognition.

[0127] Current technologies often use single-class spectrograms for gesture recognition. However, when classifying easily confused gestures, the similarity of their features can lead to confusion. For example, waving to the left and waving to the right have extremely similar features in the distance spectrogram. Adding horizontal angle features as a criterion can significantly improve the recognition rate of confused gestures, because the features of these two types of gestures in the horizontal angle spectrogram are completely opposite.

[0128] like Figure 5 As shown, Figure 5 To test the network, a confusion matrix was obtained using 23,040 radar spectra as training data, including five hand gestures: waving left, waving right, waving upwards, waving downwards, and pushing forward. Then, 5,760 radar spectra were used to test the network, resulting in the confusion matrix. The vertical axis represents the network's recognition result, and the horizontal axis represents the actual hand gesture category. A1, A2, A3, A4, and A5 represent waving left, waving right, waving upwards, waving downwards, and pushing forwards, respectively. The confusion matrix shows that the recognition rate for each gesture exceeds 90%, with an overall recognition rate of 96.0%.

[0129] Example 2

[0130] The millimeter-wave radar gesture recognition system based on multi-domain spectral and multi-resolution fusion provided in this embodiment includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the above-described method.

[0131] like Figure 6 As shown, Figure 6 The diagram shows a non-line-of-sight human gesture recognition system, including a millimeter-wave radar data acquisition module, a gesture recognition network, a wireless data transmission module, and a host computer platform.

[0132] The millimeter-wave radar data acquisition module is used to acquire the target's raw radar echo data;

[0133] The gesture recognition network is used to fuse multiple types of radar spectra to obtain fused features and to classify target gestures based on the fused features; the gesture recognition network includes a gesture representation module and a feature extraction and gesture recognition module.

[0134] The gesture representation module is used to generate a radar multi-domain spectral map of the target gesture;

[0135] The feature extraction and gesture recognition module is used to extract and fuse the multi-domain spectral features and multi-resolution features of the target and classify the target gestures to achieve human gesture recognition.

[0136] The wireless data transmission module is used to transmit data and instructions between the radar module and the host computer platform for communication.

[0137] The host computer platform is used to display the gesture recognition results and control external devices to perform corresponding functions when a corresponding gesture occurs, so as to facilitate users to achieve contactless intelligent control of devices.

[0138] In this embodiment, the gesture representation module generates four types of radar spectra with different physical meanings and complementary features for the same gesture on the same target. The radar spectra include time-frequency spectrum, time-range spectrum, time-velocity spectrum, and time-angle spectrum. Figure 4 The system combines different spectral maps and constructs datasets based on different targets and gestures. The four types of radar spectral maps have both independent features and complementary features. That is, when a single type of spectral map cannot distinguish certain types of gestures, radar spectral maps from other radar domains are used to supplement the features of that type of gesture to more fully and comprehensively complete the expression of gesture features, thereby improving the accuracy of gesture recognition.

[0139] The feature extraction and gesture recognition module in this embodiment includes a linear projection layer, a feature abstraction layer, a multi-resolution feature fusion layer, a semantic information extraction layer, a complementary feature fusion layer, and a classifier. It uses a parallel two-dimensional convolutional network module and a multi-resolution fusion mechanism for feature extraction and fusion. For radar multi-domain feature spectra, the linear projection layer maps the multi-domain feature spectra into feature tensors of a fixed size. The obtained H×W×C feature tensors are input into the feature abstraction layer to realize feature abstraction of the radar feature spectra. The multi-resolution feature fusion layer fuses feature tensors at different resolutions to obtain multi-resolution fused features. Then, the multi-resolution features are input into the semantic information extraction layer.

[0140] The above-described embodiments are merely preferred embodiments provided to fully illustrate the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are all within the scope of protection of the present invention. The scope of protection of the present invention is defined by the claims.

Claims

1. A method for gesture recognition based on multi-domain spectrogram and multi-resolution fusion, characterized in that: The method comprises the following steps: (1) determining a target to be detected in a radar detection area; (2) collecting radar echo signals returned by the target, generating a plurality of radar spectrum graphs for the same target and the same gesture, the radar spectrum graphs being radar spectrum graphs containing different physical meanings and having complementary features, and constructing a data set according to different targets and different gestures; the plurality of radar spectrum graphs comprise a time-frequency spectrum graph, a time-distance spectrum graph, a time-velocity spectrum graph and a time-angle spectrum graph; (3) constructing a gesture recognition network, the gesture recognition network being a composite neural network composed of a parallel two-dimensional convolution network and a multi-resolution module, for fusing the plurality of radar spectrum graphs to obtain fused features and classifying target gestures according to the fused features, and training and testing the gesture recognition network by using the data set; the gesture recognition network comprises a linear projection layer, a feature abstraction layer, a multi-resolution feature fusion layer, a semantic information extraction layer, a complementary feature fusion layer and a classifier; (4) inputting the fused features into the gesture recognition network to classify the target gesture and determine the gesture action of the target; if a preset gesture is detected, step (5) is performed; if the preset gesture is not detected, step (4) is performed again; the preset gesture comprises a left hand waving gesture, a right hand waving gesture, an upward hand waving gesture, a downward hand waving gesture and a forward pushing gesture; (5) sending the recognized gesture information.

2. The multi-domain spectrogram and multi-resolution fusion based millimeter wave radar gesture recognition method of claim 1, wherein: Step (2) specifically comprises: (21) preprocessing the original echo signals collected by the radar; (22) performing short-time Fourier transform on the preprocessed signals, and stacking in the slow time dimension to obtain time-frequency feature expression of the echo signals; (23) performing pulse compression processing on the preprocessed signals, stacking in the slow time dimension, and obtaining distance feature expression through multi-receiving channel accumulation; (24) performing two-dimensional fast Fourier transform on the preprocessed signals, and stacking in the slow time dimension in order to obtain velocity feature expression; (25) performing processing on the preprocessed signals by using the minimum variance distortionless response beam forming algorithm, and stacking in the slow time dimension in order to obtain horizontal angle feature expression.

3. The multi-domain spectrogram and multi-resolution fusion based millimeter wave radar gesture recognition method of claim 2, wherein: Step (21) specifically comprises: The original echo signals are subjected to static clutter suppression, and the phases of the moving targets in the echo signals are different at different times, so that the current echo signals and the echo signals with an interval time τ are cancelled to obtain target echo signals, and the calculation formula is as follows: wherein, denotes a slow-time index within a frequency modulation period, is a radar echo signal at time instant 4. The multi-domain spectrogram and multi-resolution fusion based millimeter wave radar gesture recognition method of claim 2, wherein: Step (22) specifically comprises: First, the preprocessed signals are segmented, and each segment of the signals is subjected to windowing operation, then the windowed data is subjected to short-time Fourier transform, and finally the results of processing each segment of the signals are stacked in the slow time dimension to obtain frequency feature expression of the echo signals, and the calculation formula is as follows: wherein is the target echo signal, is a window function, is the resulting short-time Fourier transform result, denotes the slice length around the analysis time point t.

5. The multi-domain spectrogram and multi-resolution fusion based millimeter wave radar gesture recognition method of claim 2, wherein: Step (23) specifically comprises: First, the preprocessed signals are split, i.e. the echo signals are cut into a plurality of complete chirp signals, then the pulse compression processing is performed on each chirp signal to obtain the range image of the single chirp signal, and finally the range images of the chirp signals are stacked in the slow time dimension according to time, and the distance feature expression of the echo signals is obtained through multi-receiving channel accumulation, and the calculation method is as follows: in, Representing the A chirp signal in The amplitude of the point; The window function is represented, and the window function is a Hamming window; Represents the number of sampling points; This represents the m-th chirp signal. The value of each sampling point; Second, if the target is far away from the radar, the echo signal energy will be low, resulting in the difference between the feature information and the background not being obvious. The method of accumulating multi-channel data is used, and the specific implementation method is as follows: wherein the distance profile obtained for the Cth reception channel; Finally, the range profiles of different receiving channels are accumulated to form the final range feature expression.

6. The multi-domain spectrogram and multi-resolution fusion based millimeter wave radar gesture recognition method of claim 2, wherein: Step (24) is specifically: First, the original signal is windowed, and the range image is obtained after fast Fourier transform in the fast time dimension. Then, the range image is windowed and fast Fourier transform is performed in the slow time dimension to obtain the velocity spectrum. Finally, the velocity spectrum is stacked in the slow time dimension according to the time sequence to obtain the velocity feature expression, and the calculation method is as follows: wherein, represents the amplitude at the point s of the kth row of the distance profile, represents the amplitude of the kth row m-i column of the fth frame, the number of chirp signals contained per frame.

7. The multi-domain spectrogram and multi-resolution fusion based millimeter wave radar gesture recognition method of claim 2, wherein: Step (25) is specifically: First, the original signal is windowed, and the range image is obtained after fast Fourier transform in the fast time dimension. Then, the range image is processed using the minimum variance distortionless response beamforming algorithm to obtain the distortionless output of the target azimuth signal, and the calculation method is as follows: , wherein, represents the echo signal energy of the θ position in the mth chirp signal, n is the number of array elements, represents the autocorrelation matrix of the range profile of the mth chirp signal, that is: is the steering vector, i.e. Finally, the horizontal angle spectrum is stacked in the slow time dimension according to the time sequence to obtain the horizontal angle feature expression.

8. The multi-domain spectrogram and multi-resolution fusion based millimeter wave radar gesture recognition method of claim 1, wherein: Step (3) is specifically: First, the multi-class feature spectrum is unified and reshaped into a tensor of the same size, and a linear projection network is used to map it to a fixed-size feature tensor. The obtained feature tensor is input into a two-dimensional convolution network structure to realize the abstraction of the spectrum feature, and then the abstracted feature is input into a multi-resolution fusion module to generate a multi-resolution feature tensor. The multi-resolution feature tensors of the multi-domain spectrum are fused to obtain the complementary fusion features of the multi-domain spectrum. Finally, the features are classified using logistic regression. 9.A millimeter wave radar gesture recognition system based on multi-domain spectrogram and multi-resolution fusion, comprising a memory, a processor and a computer program stored in the memory and capable of running on the processor, characterized in that, The processor executes the program to realize the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Hand gesture motion detection method based on millimeter wave radar

    CN109188414A

  • Radar human body posture recognition method and system based on multi-class spectrogram fusion and hierarchical learning

    CN111368930A