Dynamic Gesture Radar Recognition Method and Device
By using convolutional neural network to perform multi-level fusion and feature processing in the multi-channel radar micro-Doppler gesture recognition method, the problems of poor recognition effect and long training time are solved, and higher recognition accuracy and shorter training time are achieved.
Patent Information
- Application Number
- CN202010291849.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-04-14
AI Technical Summary
The gesture recognition method based on multi-channel radar microDoppler has poor recognition effect, and the prior art has poor single-layer fusion in neural networks, and the training time is long.
The convolutional neural network model is adopted to generate a normalized multi-resolution time-frequency map by collecting and preprocessing time-domain echo signals, and achieve multi-level fusion and efficient identification through feature fusion networks and classification networks.
Improves the accuracy of dynamic gesture radar recognition methods, reduces training time, and achieves higher gesture recognition accuracy through multi-layer fusion and learnable weighting coefficients.
Smart Images

Figure CN111680539B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction, and in particular, to a method and device for dynamic gesture radar recognition. Background Art
[0002] Using radar to observe targets has the advantages of long range, all-day and all-weather operation. The radar micro-Doppler effect refers to the additional Doppler effect caused by the vibration, rotation, etc. of the target relative to the centroid on the basis of the Doppler effect caused by the centroid motion of the target. The variation law of the micro-Doppler frequency shift of the target with time, that is, the micro-Doppler time-frequency diagram, can reflect the motion posture and structural information of the target, and thus be used as a feature for target recognition in application scenarios such as aircraft recognition, human posture recognition, and gesture recognition.
[0003] Classical target recognition algorithms based on micro-Doppler features usually consist of two steps: feature extraction and classification. That is, the algorithm first extracts empirical features specific to the data or mathematical transformation features without data specificity from the Doppler time-frequency diagram, and then inputs the above features into classical classifiers (such as support vector machines, naive Bayes classifiers, nearest neighbor classifiers) to obtain recognition results. Convolutional neural network is a pattern recognition algorithm that has emerged in recent years. It has achieved higher recognition accuracy than classical recognition algorithms in various recognition tasks in fields such as vision and speech, and does not require domain experts to select empirical features, featuring fast technology development speed and good universality.
[0004] A radar system with a transceiver antenna located at the same position and only one pair of transceiver antennas is called a single-channel radar; a radar system with multiple receiving (or transmitting) antennas arranged at a certain spacing in space is called a multi-channel radar. The significant drawback of a single-channel radar is that it can only obtain radial Doppler information and has no perception ability for the transverse Doppler component perpendicular to the radial direction. Therefore, it cannot obtain complete target motion information. For gesture recognition tasks, hand movements are usually not limited to the radial direction, and the inability to perceive the transverse Doppler component will greatly limit the gesture recognition accuracy. A multi-channel radar has multiple transceiver paths and can observe the target from different angles simultaneously, thus obtaining the radial and transverse motion information of the target at the same time. Among them, a key issue is how to fuse data from different nodes or transceiver paths. Combining the algorithm advantages of convolutional neural networks and the radar system advantages of multi-channel radars is a promising development direction.
[0005] The document "Z. Chen, G. Li, F. Fioranelli, and H. Griffiths, 'Dynamic hand gesture classification based on multistatic radar micro-Doppler signatures using convolutional neural network,' IEEE Radar Conf., Boston, MA, 2019." discloses a gesture recognition method based on multi-channel radar micro-Doppler. The multi-channel radar used in this method has a transmitting antenna at the center and four receiving antennas at the four corners. This method inputs the micro-Doppler time-frequency diagrams of the signals obtained by each receiving antenna into a convolutional neural network to achieve the classification of six gestures. Aiming at the problem of how to fuse data from different receiving antennas, this method designs a multi-input single-output model, in which multiple input branches perform single-layer information fusion at a pre-specified site. To find the optimal fusion site, this method trains the convolutional neural network multiple times and traverses all possible single-layer fusion sites. The prominent disadvantages of this method are, firstly, traversing all fusion sites consumes a large amount of training time; secondly, performing single-layer fusion at a certain position in the neural network has a poor recognition effect on gestures.
[0006] Aiming at the problem of poor recognition effect of the gesture recognition method based on multi-channel radar micro-Doppler in the related technology, no effective solution has been proposed yet. Summary of the Invention
[0007] The main purpose of the present invention is to provide a dynamic gesture radar recognition method and device to solve the problem of poor recognition effect of the gesture recognition method based on multi-channel radar micro-Doppler.
[0008] To achieve the above object, according to one aspect of the present invention, a method for dynamic gesture radar recognition is provided. The method includes: collecting a preset number of time-domain echo signals as sample data, where the sample data includes the normalized time-frequency map of each type of dynamic gesture and the type of the dynamic gesture; preprocessing the sample data to obtain a sample-normalized time-frequency map; modeling according to the sample-normalized time-frequency map to obtain the model parameters of the dynamic gesture recognition model, where the dynamic gesture recognition model is a convolutional neural network model, the convolutional neural network model has N inputs and one output, the N inputs respectively correspond to the normalized multi-resolution time-frequency maps of N time-domain echo signals, the N inputs of the convolutional neural network model are fused into one through a feature fusion network, and then one output is obtained through a classification network, and the one output is used to represent the type with the highest probability in the probability distributions of dynamic gestures in various gesture categories; recognizing the dynamic gesture to be recognized according to the dynamic gesture recognition model to obtain the type of the dynamic gesture to be recognized.
[0009] Further, recognizing the dynamic gesture to be recognized according to the dynamic gesture recognition model includes: obtaining the time-domain echo signal of the dynamic gesture to be recognized; performing data processing on the time-domain echo signal to obtain a normalized multi-resolution time-frequency map; inputting the normalized multi-resolution time-frequency map into the convolutional neural network model for recognition to obtain the type of the dynamic gesture to be recognized.
[0010] Further, recognizing the dynamic gesture to be recognized according to the dynamic gesture recognition model to obtain the type of the dynamic gesture to be recognized includes: recognizing the dynamic gesture to be recognized according to the dynamic gesture recognition model to obtain the type to which the dynamic gesture to be recognized belongs and the probability corresponding to this type; determining the type corresponding to the maximum probability as the type of the dynamic gesture to be recognized.
[0011] Further, the feature fusion network has N inputs and one output, and includes a processing module, a fusion module, and a weighted fusion node. Each processing module has one input and one output, each fusion module has N inputs and one output, each weighted fusion node has two inputs and one output, and the input-output relationship is: z l =x l,1 ·w l +x l,2 ·(1 - w l ), where x l,1 is the first input of the weighted fusion node V l , x l,2 is the second input of the weighted fusion node V l , z l is the output of the weighted fusion node V lOutput
[0012] Further, the first dimension, the second dimension, and the third dimension of the input of each branch of the convolutional neural network model respectively correspond to the independent variables t, f, and k of the normalized multi-resolution time-frequency map According to the probability distribution of each gesture type output by the classification network, the cross-entropy loss function between the probability distribution and each type of dynamic gesture is determined, and the cross-entropy loss function is minimized by the stochastic gradient descent method to obtain the model parameters of the dynamic gesture recognition model.
[0013] To achieve the above object, according to another aspect of the present invention, there is also provided a dynamic gesture radar recognition device, which includes: an acquisition unit for acquiring a preset number of time-domain echo signals as sample data, wherein the sample data includes the normalized time-frequency map of each type of dynamic gesture and the type of the dynamic gesture; a processing unit for preprocessing the sample data to obtain a sample-normalized time-frequency map; a modeling unit for modeling according to the sample-normalized time-frequency map to obtain the model parameters of the dynamic gesture recognition model, wherein the dynamic gesture recognition model is a convolutional neural network model, the convolutional neural network model has N inputs and one output, the N inputs respectively correspond to the normalized multi-resolution time-frequency maps of N time-domain echo signals, the N inputs of the convolutional neural network model are fused into one through a feature fusion network, and then one output is obtained through a classification network, and the one output is used to represent the type with the highest probability in the probability distribution of each gesture type of the dynamic gesture; an identification unit for identifying the dynamic gesture to be identified according to the dynamic gesture recognition model to obtain the type of the dynamic gesture to be identified.
[0014] Further, the identification unit includes: an acquisition module for acquiring the time-domain echo signal of the dynamic gesture to be identified; a processing module for performing data processing on the time-domain echo signal to obtain a normalized multi-resolution time-frequency map; a first identification module for inputting the normalized multi-resolution time-frequency map into the convolutional neural network model for identification to obtain the type of the dynamic gesture to be identified.
[0015] Further, the identification unit includes: a second identification module for identifying the dynamic gesture to be identified according to the dynamic gesture recognition model to obtain the type to which the dynamic gesture to be identified belongs and the probability corresponding to the type; a determination module for determining the type corresponding to the maximum probability as the type of the dynamic gesture to be identified.
[0016] To achieve the above object, according to another aspect of the present invention, there is also provided a storage medium including a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the dynamic gesture radar recognition method of the present invention.
[0017] To achieve the above object, according to another aspect of the present invention, there is also provided a device including at least one processor, and at least one memory and a bus connected to the processor. Among them, the processor and the memory complete communication with each other through the bus, and the processor is used to call program instructions in the memory to execute the dynamic gesture radar recognition method of the present invention.
[0018] The present invention collects a preset number of time-domain echo signals as sample data. Among them, the sample data includes the normalized time-frequency map of each type of dynamic gesture and the type of this dynamic gesture; preprocesses the sample data to obtain the sample-normalized time-frequency map; models according to the sample-normalized time-frequency map to obtain the model parameters of the dynamic gesture recognition model. Among them, the dynamic gesture recognition model is a convolutional neural network model. The convolutional neural network model has N inputs and one output. The N inputs respectively correspond to the normalized multi-resolution time-frequency maps of N time-domain echo signals. The N inputs of the convolutional neural network model are fused into one through the feature fusion network, and then one output is obtained through the classification network. One output is used to represent the type with the largest probability in the probability distribution of each gesture type of the dynamic gesture; recognizes the dynamic gesture to be recognized according to the dynamic gesture recognition model to obtain the type of the dynamic gesture to be recognized, solves the problem of poor recognition effect of the gesture recognition method based on multi-channel radar micro-Doppler, and further achieves the effect of improving the accuracy of the dynamic gesture radar recognition method. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0020] Figure 1 is a flowchart of the dynamic gesture radar recognition method according to an embodiment of the present invention;
[0021] Figure 2 is a schematic diagram of the convolutional neural network of this embodiment;
[0022] Figure 3 is a schematic diagram of the feature fusion network of this embodiment;
[0023] Figure 4 is a schematic diagram of the dynamic gesture radar recognition device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0025] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0026] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data used may be interchanged under appropriate circumstances for the embodiments of the present application described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] The embodiment of the present invention provides a dynamic gesture radar recognition method.
[0028] Figure 1 is a flowchart of the dynamic gesture radar recognition method according to the embodiment of the present invention, as Figure 1 shown, the method includes the following steps:
[0029] Step S102: Collect a preset number of time-domain echo signals as sample data, where the sample data includes the normalized time-frequency diagrams of each type of dynamic gesture and the type of the dynamic gesture;
[0030] Step S104: Preprocess the sample data to obtain a sample-normalized time-frequency diagram;
[0031] Step S106: Model according to the sample-normalized time-frequency diagram to obtain the model parameters of the dynamic gesture recognition model. The dynamic gesture recognition model is a convolutional neural network model. The convolutional neural network model has N inputs and one output. The N inputs respectively correspond to the normalized multi-resolution time-frequency diagrams of N time-domain echo signals. The N inputs of the convolutional neural network model are fused into one path through a feature fusion network, and then one output is obtained through a classification network. The one output is used to represent the type with the highest probability in the probability distribution of each gesture type of the dynamic gesture;
[0032] Step S108: Identify the dynamic gesture to be identified according to the dynamic gesture recognition model to obtain the type of the dynamic gesture to be identified.
[0033] This embodiment uses the collected time-domain echo signals of a preset number as sample data. Among them, the sample data includes the normalized time-frequency diagrams of each type of dynamic gesture and the type of the dynamic gesture; preprocess the sample data to obtain the sample-normalized time-frequency diagrams; model according to the sample-normalized time-frequency diagrams to obtain the model parameters of the dynamic gesture recognition model. Among them, the dynamic gesture recognition model is a convolutional neural network model. The convolutional neural network model has N inputs and one output. The N inputs respectively correspond to the normalized multi-resolution time-frequency diagrams of N time-domain echo signals. The N inputs of the convolutional neural network model are fused into one through the feature fusion network, and then passed through the classification network to obtain one output. One output is used to represent the type with the highest probability in the probability distribution of each gesture type of the dynamic gesture; identify the dynamic gesture to be identified according to the dynamic gesture recognition model to obtain the type of the dynamic gesture to be identified, which solves the problem of poor recognition effect of the gesture recognition method based on multi-channel radar micro-Doppler, and further achieves the effect of improving the accuracy of the dynamic gesture radar recognition method. The main difference from the above-mentioned paper is that the above-mentioned paper performs micro-Doppler feature fusion at a certain layer of the neural network; this patent adaptively performs micro-Doppler feature fusion at multiple layers of the neural network, thereby obtaining a higher gesture recognition accuracy rate.
[0034] The technical solution of this embodiment establishes a convolutional neural network model through a certain amount of sample data for identifying dynamic gestures. Since the designed convolutional neural network model has a classification network, the probability that the dynamic gesture belongs to each type can be obtained, and the type corresponding to the maximum probability is output, reducing the calculation process and improving the recognition accuracy.
[0035] Optionally, identifying the dynamic gesture to be identified according to the dynamic gesture recognition model includes: obtaining the time-domain echo signal of the dynamic gesture to be identified; performing data processing on the time-domain echo signal to obtain a normalized multi-resolution time-frequency diagram; inputting the normalized multi-resolution time-frequency diagram into the convolutional neural network model for identification to obtain the type of the dynamic gesture to be identified.
[0036] The time-domain echo signal can be obtained by a multi-channel radar that transmits radar signals and collects the time-domain echo signals of the dynamic gesture area. The multi-channel radar includes at least one transmitting antenna with different positions and multiple receiving antennas with different positions. The normalized multi-resolution time-frequency map can be obtained through the following steps: perform short-time Fourier transform on N channels of time-domain echo signals to calculate a set of multi-resolution time-frequency maps, take the logarithm of the above set of multi-resolution time-frequency maps, and normalize the maximum value of the multi-resolution time-frequency map after taking the logarithm to 0 dB. Set a threshold, and cut off the part less than the threshold in the multi-resolution time-frequency map after normalizing the maximum value to obtain a set of normalized multi-resolution time-frequency maps. Since a multi-channel radar is used, the lateral and radial motion information of the gesture can be obtained simultaneously, which is beneficial to distinguishing different gestures with subtle differences.
[0037] Optionally, the dynamic gesture to be recognized is recognized according to the dynamic gesture recognition model, and the types of the dynamic gesture to be recognized obtained include: recognizing the dynamic gesture to be recognized according to the dynamic gesture recognition model to obtain the type to which the dynamic gesture to be recognized belongs and the probability corresponding to this type; determining the type corresponding to the maximum probability as the type of the dynamic gesture to be recognized. The output of the model can directly be the type corresponding to the maximum probability, or the various types and probabilities can be output together.
[0038] Optionally, the feature fusion network has N inputs and one output, and is composed of a processing module, a fusion module, and a weighted fusion node. Each processing module has one input and one output, each fusion module has N inputs and one output, and each weighted fusion node has two inputs and one output, and the input-output relationship is: z l =x l,1 ·w l +x l,2 ·(1 - w l ), where x l,1 is the first input of the weighted fusion node V l , x l,2 is the second input of the weighted fusion node V l , and z l is the output of the weighted fusion node V l .
[0039] Optionally, the first dimension, second dimension, and third dimension of the input of each branch of the convolutional neural network model respectively correspond to the independent variables t, f, and k of the normalized multi-resolution time-frequency map . Determine the cross-entropy loss function between the probability distribution and each type of dynamic gesture according to the probability distribution of each gesture type output by the classification network, and perform minimization calculation on the cross-entropy loss function through the stochastic gradient descent method to obtain the model parameters of the dynamic gesture recognition model.
[0040] This embodiment also provides a preferred implementation.
[0041] The purpose of this embodiment is to overcome the above-mentioned disadvantages and deficiencies of the prior art, and propose a gesture recognition method based on multi-channel radar micro-Doppler and multi-layer fusion neural network. By using the micro-Doppler information of multi-channel radar, a convolutional neural network with multi-level fusion sites and learnable fusion weighting coefficients is proposed, which realizes the efficient fusion of multi-receiving antenna data in a short training time, thereby achieving high-accuracy gesture recognition. Compared with the gesture recognition methods in the related art, the radar micro-Doppler features contained in the convolutional neural network of this embodiment are fused at all levels of the network, which is called multi-layer fusion, and the fusion weighting coefficients can be autonomously learned.
[0042] (1) Collect the time-domain echo signals of known gesture types:
[0043] Use a multi-channel radar to collect the time-domain echo signals of a known gesture type. The multi-channel radar includes M T (M T ≥1) transmitting antennas with different positions and M R (M R >1) receiving antennas with different positions, and obtain N = M T ·M R channels of time-domain echo signals s n (t), n = 1, 2,..., N, where t is the acquisition time; when the radar used is a continuous wave radar, the above time-domain echo signals s n (t), n = 1, 2,..., N refer to the demodulated baseband signals; when the radar used is a frequency-modulated continuous wave radar or a pulse radar, the above time-domain echo signals s n (t), n = 1, 2,..., N refer to the time-domain signals changing along the slow time in a certain range unit or the weighted sum of the time-domain signals changing along the slow time in multiple range units; in this embodiment, a frequency-modulated continuous wave radar is used, including M T =1 transmitting antenna and M R =4 receiving antennas, and N = 4 channels of time-domain echo signals are obtained.
[0044] (2) Generate a normalized multi-resolution time-frequency map, including the following steps:
[0045] (2a) For the N channels of time-domain echo signals s n (t), n = 1, 2,..., N, use the short-time Fourier transform to calculate a set of multi-resolution time-frequency maps S n,k (t, f):
[0046]
[0047] where the independent variable t is time, n is the sequence number of the time-domain echo signal, f is the Doppler frequency of the time-domain echo signal, |·| represents taking the modulus of a complex number, j is the imaginary unit, π is the pi, k is the window function sequence number, and h k (t) are K short-time Fourier transform window functions with the same form but different window lengths; in an embodiment of the present invention, the short-time Fourier transform is calculated using K = 3 Blackman windows with different window lengths;
[0048] (2b) Take the logarithm of the above set of multi-resolution time-frequency diagrams S n,k (t, f), and normalize the maximum value of the multi-resolution time-frequency diagram S n,k (t, f) after taking the logarithm to 0 dB, and set a threshold θ th , and cut off the part less than the threshold θ n,k in the multi-resolution time-frequency diagram S th after normalizing the maximum value, to obtain a set of normalized multi-resolution time-frequency diagrams X n,k (t, f):
[0049]
[0050] In an embodiment of the present invention, an appropriate threshold θ is selected according to the radar background noise level th ;
[0051] (3) Construct a training data set:
[0052] Sample a gesture of a known class C multiple times, and repeat the above steps (1) and (2) to obtain multiple sets of normalized multi-resolution time-frequency diagrams where i = 1, 2, 3... is the sequence number of multiple samplings of a known gesture, n = 1, 2,..., N is the sequence number of the time-domain echo signal, k = 1, 2,..., K is the subscript of K short-time Fourier transform window functions with the same form but different window lengths, and the i-th set of normalized multi-resolution time-frequency diagrams and the sample label y i = C form an ordered pair to form a training sample; traverse all known gesture classes, repeat this step, obtain all training samples, and form all training samples into a training data set; in an embodiment of the present invention, a training data set containing 10 known gestures is constructed, and the above 10 gestures include: waving left, waving right, waving up, waving down, drawing a circle clockwise, drawing a circle counterclockwise, check mark, cross mark, cross, waving, and in the experiment, the above 10 gestures of multiple experimenters are collected;
[0053] (4) Construct a multi-locus fusion convolutional neural network:
[0054] Figure 2 is a schematic diagram of the convolutional neural network of this embodiment. The multi-site fusion convolutional neural network of this embodiment has N inputs and one output. The N inputs of the multi-site fusion convolutional neural network correspond to the normalized multi-resolution time-frequency diagrams of the N echo signals. The input of the nth channel (n=1, 2, ..., N) is in the form of a three-dimensional tensor, the first dimension, the second dimension and the third dimension of which correspond to the normalized multi-resolution time-frequency graph The subscripts t, f and k of the multi-site fusion convolutional neural network are first fused into one channel through the feature fusion network, and then passed through the classification network to obtain one channel output of the multi-site fusion convolutional neural network, that is, the probability distribution of the gesture to be recognized belonging to each gesture category; thereafter, the category with the largest probability in the probability distribution of each gesture category is taken as the recognition result; in one embodiment of the present invention, the specific form of the above-mentioned classification network is a global average pooling layer, a fully connected layer 1, a fully connected layer 2 and a Softmax layer connected in sequence.
[0055] Figure 3 is a schematic diagram of the feature fusion network of this embodiment, such as Figure 3 As shown, the feature fusion network has N inputs and one output, which is composed of a processing module P l,n , n=1,2,...,N+1,l=1,2,...,L,fusion module F l , l=1,2,...,L+1 and the fusion coefficients are w l , 0≤w l ≤1, l=1, 2, ..., L weighted fusion node V l , l=1,2,...,L, where n is the input number of the feature fusion network, l is the layer number, and L (L≥1) is the number of layers of the feature fusion network; any of the above processing modules has one input and one output, any fusion module has N inputs and one output, and any weighted fusion node V l , l = 1, 2, ..., L has two inputs and one output and the relationship between the output and the input is z l =x l,1 ·w l +x l,2 ·(1-w l ), where x l,1 is the weighted fusion node V l The first input, x l,2 is the weighted fusion node V l The second input, z l is the weighted fusion node V l The above processing modules and fusion modules are all convolutional neural networks including convolution, activation, batch normalization and other operations, among which the processing modules P with the same layer number ll,n , where \(n = 1, 2, \ldots, N + 1\) are structurally identical to each other and share weights; in an embodiment of the present invention, the number of layers \(L\) of the feature fusion network is 8, and the specific form of the processing module is a sequentially connected convolutional layer, a batch normalization layer, and a max pooling layer, and the specific form of the fusion module is a sequentially connected concatenation layer, a \(1\times1\) convolutional layer, and a batch normalization layer, where the concatenation layer concatenates the tensors of \(N = 4\) inputs along the third dimension;
[0056] The connection relationships of the above-mentioned processing modules, fusion modules, and weighted fusion nodes are as follows: for any \(n = 1, 2, \ldots, N\), the input \(n\) passes through the processing module \(P\) in sequence 1,n , \(P\) 2,n , \(\ldots\), \(P\) L,n ; for any \(l = 1, 2, \ldots, L\), the outputs of \(P\) l,1 , \(P\) l,2 , \(\ldots\), \(P\) l,N jointly serve as the input of the fusion module \(F\) l+1 ; in particular, the inputs 1, 2, \(\ldots\), \(N\) jointly serve as the input of the fusion module \(F1\); the output of the fusion module \(F1\) is the input of the processing module \(P\) 1,N+1 ; for any \(l = 1, 2, \ldots, L\), the first input of the weighted fusion node \(V\) l is the output of the processing module \(P\) l,N+1 , and the second input of the weighted fusion node \(V\) l is the output of the fusion module \(F\) l+1 ; for any \(l = 1, 2, \ldots, L - 1\), the output of the weighted fusion node \(V\) l is the input of the processing module \(P\) l+1,N+1 ; in particular, the output of the weighted fusion node \(V\) L is the output of the feature fusion network.
[0057] (5) Train the convolutional neural network in step (4) above:
[0058] Input the normalized multi-resolution time-frequency maps in the training set generated in step (3) into the multi-site fusion convolutional neural network constructed in step (4), where \(i = 1, 2, 3, \ldots\) is the training sample serial number, \(n = 1, 2, \ldots, N\) is the echo signal serial number, and \(k = 1, 2, \ldots, K\) is the subscript of the short-time Fourier transform window functions with the same form but different window lengths. The input of the \(n\)th branch of the multi-site fusion convolutional neural network is The first dimension, the second dimension, and the third dimension of the input of any branch respectively correspond to the independent variables \(t\), \(f\), and \(k\) of the normalized multi-resolution time-frequency map . The probability distribution of each gesture type is output by the classification network, and a probability distribution and the known gesture type \(y\) are introduced iThe cross - entropy loss function between them is minimized by using the stochastic gradient descent method to obtain the trained convolutional neural network;
[0059] The training of the multi - point fusion convolutional neural network constructed in step (4) above can be selected in one of the following two modes:
[0060] i. Learnable fusion coefficient mode: The fusion coefficient w l , 0 ≤ w l ≤ 1, l = 1, 2,..., L can be optimized synchronously with other weights in the convolutional neural network. This mode can obtain the best recognition rate, but the training stability and training speed are lower than those of mode ii;
[0061] ii. Fixed fusion coefficient mode: The fusion coefficient is fixed as w l = 1 / 2, l = 1, 2,..., L. This mode can be regarded as a simplified version of mode i. The recognition rate of mode ii is slightly lower than that of mode i, but the training stability and training speed are better;
[0062] In an embodiment of the present invention, training is performed separately in the above two modes to obtain two trained convolutional neural networks.
[0063] (6) Recognize the gesture type:
[0064] In an embodiment of the present invention, steps (1) and (2) are repeated multiple times to obtain multiple groups of normalized multi - resolution time - frequency maps of the gesture to be recognized, which constitute a test set. The normalized multi - resolution time - frequency maps of each gesture in the test set are respectively input into the trained convolutional neural network in step (5) to output the type of the gesture to be recognized, complete gesture recognition, and count the recognition accuracy rate.
[0065] The advantages of the gesture recognition method based on radar micro - Doppler in this embodiment are as follows:
[0066] 1. The multi - channel radar used in this embodiment can simultaneously obtain the radial and transverse motion information of the gesture, which is conducive to distinguishing different gestures with subtle differences.
[0067] 2. This embodiment uses a convolutional neural network for data processing, without the need for professionals in the field to select empirical features, with fast technology development speed, good universality, and high recognition accuracy.
[0068] 3. This embodiment uses the multi - resolution time - frequency map as the input of the convolutional neural network, which provides more information than the single - resolution time - frequency map and helps to improve the recognition rate.
[0069] 4. The convolutional neural network proposed in this embodiment has multiple fusion sites and the fusion coefficients are learnable, enabling efficient fusion of multi-receive antenna data in a relatively short training time, thereby achieving high-accuracy gesture recognition.
[0070] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0071] The embodiment of the present invention provides a dynamic gesture radar recognition device, which can be used to execute the dynamic gesture radar recognition method of the embodiment of the present invention.
[0072] Figure 4 is a schematic diagram of the dynamic gesture radar recognition device according to the embodiment of the present invention, as Figure 4 shown, the device includes:
[0073] An acquisition unit 10, configured to acquire a preset number of time-domain echo signals as sample data, where the sample data includes the normalized time-frequency map of each type of dynamic gesture and the type of the dynamic gesture;
[0074] A processing unit 20, configured to preprocess the sample data to obtain a sample-normalized time-frequency map;
[0075] A modeling unit 30, configured to perform modeling according to the sample-normalized time-frequency map to obtain the model parameters of the dynamic gesture recognition model. The dynamic gesture recognition model is a convolutional neural network model. The convolutional neural network model has N inputs and one output. The N inputs respectively correspond to the normalized multi-resolution time-frequency maps of N time-domain echo signals. The N inputs of the convolutional neural network model are fused into one through a feature fusion network, and then one output is obtained through a classification network. The one output is used to represent the type with the highest probability in the probability distribution of each gesture type of the dynamic gesture;
[0076] An identification unit 40, configured to identify the dynamic gesture to be identified according to the dynamic gesture recognition model to obtain the type of the dynamic gesture to be identified.
[0077] This embodiment employs an acquisition unit 10 for acquiring a preset number of time-domain echo signals as sample data. The sample data includes the normalized time-frequency diagrams of each type of dynamic gesture and the type of the dynamic gesture. A processing unit 20 is used to preprocess the sample data to obtain a sample-normalized time-frequency diagram. A modeling unit 30 is used to perform modeling based on the sample-normalized time-frequency diagram to obtain the model parameters of the dynamic gesture recognition model. The dynamic gesture recognition model is a convolutional neural network model. The convolutional neural network model has N inputs and one output. The N inputs respectively correspond to the normalized multi-resolution time-frequency diagrams of N time-domain echo signals. The N inputs of the convolutional neural network model are fused into one through a feature fusion network and then passed through a classification network to obtain one output. The one output is used to represent the type with the highest probability in the probability distribution of each gesture type of the dynamic gesture, thus solving the problem of poor recognition effect of the gesture recognition method based on multi-channel radar micro-Doppler, and further achieving the effect of improving the accuracy of the dynamic gesture radar recognition method.
[0078] Optionally, the recognition unit 40 includes: an acquisition module for acquiring the time-domain echo signal of the dynamic gesture to be recognized; a processing module for performing data processing on the time-domain echo signal to obtain a normalized multi-resolution time-frequency diagram; and a first recognition module for inputting the normalized multi-resolution time-frequency diagram into the convolutional neural network model for recognition to obtain the type of the dynamic gesture to be recognized.
[0079] Optionally, the recognition unit 40 includes: a second recognition module for recognizing the dynamic gesture to be recognized according to the dynamic gesture recognition model to obtain the type to which the dynamic gesture to be recognized belongs and the probability corresponding to this type; and a determination module for determining the type corresponding to the maximum probability as the type of the dynamic gesture to be recognized.
[0080] The dynamic gesture radar recognition device includes a processor and a memory. The above acquisition unit, processing unit, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions.
[0081] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and the accuracy of the dynamic gesture radar recognition method can be improved by adjusting the kernel parameters.
[0082] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0083] An embodiment of the present invention provides a storage medium, on which a program is stored, and when the program is executed by a processor, the dynamic gesture radar recognition method is implemented.
[0084] An embodiment of the present invention provides a processor, which is used to run a program, wherein when the program runs, the dynamic gesture radar recognition method is executed.
[0085] An embodiment of the present invention provides a device, which includes at least one processor, at least one memory connected to the processor, and a bus; wherein, the processor and the memory complete communication with each other through the bus; the processor is used to call program instructions in the memory to execute the above-mentioned dynamic gesture radar recognition method. The device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0086] This application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with the following method steps: collecting a preset number of time-domain echo signals as sample data, wherein the sample data includes the normalized time-frequency diagrams of each type of dynamic gesture and the type of the dynamic gesture; preprocessing the sample data to obtain a sample-normalized time-frequency diagram; modeling according to the sample-normalized time-frequency diagram to obtain the model parameters of the dynamic gesture recognition model, wherein the dynamic gesture recognition model is a convolutional neural network model, the convolutional neural network model has N inputs and one output, the N inputs respectively correspond to the normalized multi-resolution time-frequency diagrams of N time-domain echo signals, the N inputs of the convolutional neural network model are fused into one through a feature fusion network, and then one output is obtained through a classification network, and one output is used to represent the type with the highest probability in the probability distribution of each gesture type of the dynamic gesture; recognizing the dynamic gesture to be recognized according to the dynamic gesture recognition model to obtain the type of the dynamic gesture to be recognized.
[0087] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0088] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0089] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0090] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0091] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0092] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0093] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage, or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0094] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0095] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0096] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A dynamic gesture radar recognition method, characterized in that including: collecting a preset number of time-domain echo signals as sample data, where the sample data includes the normalized time-frequency maps of each type of dynamic gesture and the type of the dynamic gesture; preprocessing the sample data to obtain a sample-normalized time-frequency map; modeling according to the sample-normalized time-frequency map to obtain the model parameters of the dynamic gesture recognition model, where the dynamic gesture recognition model is a convolutional neural network model, the convolutional neural network model has N inputs and one output, the N inputs respectively correspond to the normalized multi-resolution time-frequency maps of N time-domain echo signals, the N inputs of the convolutional neural network model are fused into one through a feature fusion network, and then one output is obtained through a classification network, and the one output is used to represent the type with the largest probability in the probability distribution of each gesture type of the dynamic gesture; recognizing the dynamic gesture to be recognized according to the dynamic gesture recognition model to obtain the type of the dynamic gesture to be recognized; The feature fusion network has N inputs and one output, and consists of a processing module, a fusion module, and a weighted fusion node. Each processing module has one input and one output, each fusion module has N inputs and one output, and each weighted fusion node has two inputs and one output. The input-output relationship is as follows: , where is the first input of the weighted fusion node , is the second input of the weighted fusion node , is the output of the weighted fusion node , is the fusion coefficient, ; Among them, the connection relationships of the processing module, the fusion module, and the weighted fusion node are as follows: For any , the input n sequentially passes through the processing module ; For any , 's outputs are jointly used as the input of the fusion module , and the inputs 1, input 2,..., input N are jointly used as the input of the fusion module ; The output of the fusion module is the input of the processing module ; For any , the first input of the weighted fusion node is the output of the processing module , and the second input of the weighted fusion node is the output of the fusion module ; The output of the weighted fusion node is the output of the feature fusion network.
2. The method according to claim 1, characterized in that, recognizing the dynamic gesture to be recognized according to the dynamic gesture recognition model includes: obtaining the time-domain echo signal of the dynamic gesture to be recognized; performing data processing on the time-domain echo signal to obtain a normalized multi-resolution time-frequency map; inputting the normalized multi-resolution time-frequency map into the convolutional neural network model for recognition to obtain the type of the dynamic gesture to be recognized.
3. The method according to claim 1, characterized in that, recognizing the dynamic gesture to be recognized according to the dynamic gesture recognition model to obtain the type of the dynamic gesture to be recognized includes: recognizing the dynamic gesture to be recognized according to the dynamic gesture recognition model to obtain the type to which the dynamic gesture to be recognized belongs and the probability corresponding to the type; determining the type corresponding to the maximum probability as the type of the dynamic gesture to be recognized.
4. The method according to claim 3, wherein The first dimension, second dimension, and third dimension of the input of each branch of the convolutional neural network model respectively correspond to the independent variables t, f, and k of the normalized multi-resolution time-frequency map According to the probability distribution of each gesture type output by the classification network, the cross-entropy loss function between the probability distribution and each type of dynamic gesture is determined. The cross-entropy loss function is minimized by the stochastic gradient descent method to obtain the model parameters of the dynamic gesture recognition model.
5. A dynamic gesture radar recognition device, characterized in that, including: a collecting unit for collecting a preset number of time-domain echo signals as sample data, where the sample data includes the normalized time-frequency maps of each type of dynamic gesture and the type of the dynamic gesture; a processing unit for preprocessing the sample data to obtain a sample-normalized time-frequency map; a modeling unit for modeling according to the sample-normalized time-frequency map to obtain the model parameters of the dynamic gesture recognition model, where the dynamic gesture recognition model is a convolutional neural network model, the convolutional neural network model has N inputs and one output, the N inputs respectively correspond to the normalized multi-resolution time-frequency maps of N time-domain echo signals, the N inputs of the convolutional neural network model are fused into one through a feature fusion network, and then one output is obtained through a classification network, and the one output is used to represent the type with the largest probability in the probability distribution of each gesture type of the dynamic gesture; a recognition unit for recognizing the dynamic gesture to be recognized according to the dynamic gesture recognition model to obtain the type of the dynamic gesture to be recognized; Among them, the feature fusion network has N inputs and one output, and is composed of a processing module, a fusion module, and a weighted fusion node. Each processing module has one input and one output, each fusion module has N inputs and one output, and each weighted fusion node has two inputs and one output. The input-output relationship is as follows: , where is the first input of the weighted fusion node , is the second input of the weighted fusion node , is the output of the weighted fusion node , is the fusion coefficient ; Among them, the connection relationships of the processing module, the fusion module, and the weighted fusion node are as follows: For any , the input n sequentially passes through the processing module ; For any , 's outputs are jointly used as the input of the fusion module , and the inputs 1, input 2,..., input N are jointly used as the input of the fusion module ; The output of the fusion module is the input of the processing module ; For any , the first input of the weighted fusion node is the output of the processing module , and the second input of the weighted fusion node is the output of the fusion module ; The output of the weighted fusion node is the output of the feature fusion network.
6. The device according to claim 5, characterized in that the recognition unit includes: an obtaining module for obtaining the time-domain echo signal of the dynamic gesture to be recognized; a processing module for performing data processing on the time-domain echo signal to obtain a normalized multi-resolution time-frequency map; The first recognition module is configured to input the normalized multi-resolution time-frequency map into the convolutional neural network model for recognition to obtain the type of the dynamic gesture to be recognized.
7. The device according to claim 5, characterized in that, The recognition unit includes: The second recognition module is configured to recognize the dynamic gesture to be recognized according to the dynamic gesture recognition model to obtain the type to which the dynamic gesture to be recognized belongs and the probability corresponding to this type; The determination module is configured to determine the type corresponding to the maximum probability as the type of the dynamic gesture to be recognized.
8. A storage medium, characterized in that, The storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the dynamic gesture radar recognition method according to any one of claims 1 to 4.
9. A device, characterized in that, The device includes at least one processor, and at least one memory and a bus connected to the processor. Wherein, the processor and the memory complete communication with each other through the bus, and the processor is configured to call the program instructions in the memory to execute the dynamic gesture radar recognition method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Human body recognition method based on multi-base radar micro-Doppler and convolutional neural network
CN108872984A