Voice command recognition method based on millimeter wave sensing

By acquiring and processing echo data of voice commands using millimeter-wave radar, and combining deep convolutional adversarial networks and the vocal cord vibration recognition network BAC-Net, the problem of voice command recognition in complex acoustic environments has been solved, achieving efficient and accurate voice command recognition and privacy protection.

CN119724184BActive Publication Date: 2026-01-06XIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411952201.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2026-01-06
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

Existing speech recognition technologies struggle to effectively recognize voice commands in complex acoustic environments with multiple sound sources, noise, and far-field interference. Furthermore, traditional microphone methods cannot be applied in environments with through-wall sensing or when wearing masks, raising privacy concerns.

Method used

Echo data from voice commands are acquired using millimeter-wave radar. Through preprocessing, short-time Fourier transform, and Doppler spectrum analysis, and combined with an improved deep convolutional adversarial network, data augmentation is performed to construct a vocal cord vibration recognition network, BAC-Net. This network integrates Bayesian optimization and convolutional block attention modules to focus on key frequency bands and micro-Doppler features.

Benefits of technology

It achieved a voice command recognition accuracy of up to 94% in an indoor environment, reduced costs, avoided multi-sensor synchronization issues, and protected the privacy of the subjects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724184B_ABST
    Figure CN119724184B_ABST
Patent Text Reader

Abstract

The application discloses a voice command recognition method based on millimeter wave sensing, first, a millimeter wave radar data acquisition platform is built to collect echo data of different voice commands; then, the collected echo data is preprocessed; the preprocessed echo data on each distance unit is subjected to short-time Fourier transform to extract Doppler information of a target, and then, Doppler spectra of all distance units are incoherently superimposed to generate a micro-Doppler spectrum graph; data augmentation is performed by using an improved deep convolutional generative adversarial network, finally, a vocal cord vibration recognition network BAC-Net is constructed, and through training of different voice command words, an optimal model is obtained, so that different voice commands can be recognized. The application solves the limitations of the prior art in processing independent voice commands and detecting weak sounds (such as patient voice), and overcomes the difficulties brought by limited data quantity to model design and expansion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless sensing technology, specifically relating to a voice command recognition method based on millimeter-wave sensing. Background Technology

[0002] With the rapid development of internet technology and the increasing number of smart devices, voice assistants have been widely applied in fields such as autonomous driving, smart homes, and smart healthcare. However, existing voice assistants mainly rely on microphones to collect voice signals. While they perform well in ideal, noise-free environments, multiple sound sources, noise, and far-field interference in complex acoustic environments can easily lead to signal distortion, making traditional voice recognition technology difficult to meet practical needs. Furthermore, microphone methods cannot be applied in scenarios such as through-wall sensing. To improve voice recognition performance, researchers have begun exploring the possibilities of non-acoustic sensors, proposing both contact and non-contact methods. Contact methods use wearable devices to capture vibrations of the vocal cords and surrounding tissues to generate voice signals, but prolonged skin contact can be inconvenient. Therefore, research has gradually shifted towards non-contact voice sensing. Non-contact sensing was initially based on computer vision, and its principle is similar to human language communication. However, this method has privacy concerns and is limited by ambient light conditions, making it difficult to apply when users are wearing masks.

[0003] In recent years, millimeter-wave radar has been widely used in various human-computer interaction scenarios due to its advantages such as low cost, small size, high resolution, and privacy protection. Millimeter-wave radar transmits radio signals through an antenna and receives the reflected radio signals after phase modulation caused by glottal vibration. The collected echo signals show a high similarity to the speech signals recorded by a microphone. Its operating frequency range is 76-81 GHz, corresponding to a wavelength of approximately 4 mm, providing finer-grained sensing capabilities. This technology can achieve millimeter-level accuracy in resolving distance changes using phase information, can penetrate certain materials such as clothing, masks, and plastics, and is unaffected by harsh environmental conditions such as ambient light, dust, and rain / fog. Furthermore, users do not need to wear additional electronic devices, making it non-invasive. Zhang et al. proposed AmbiEar, the first millimeter-wave (mmWave) speech recognition method applicable to non-line-of-sight scenarios. Experiments show that AmbiEar achieves a word recognition accuracy of 87.21% in non-natural language scenarios, reducing the recognition error by 35.1%. Fan et al. proposed a novel millimeter-wave-based speech recognition solution aimed at addressing the problems of multi-source interference and environmental noise. This method accurately extracts multimodal features of lip movements and vocal cord vibrations using single-channel millimeter-wave signals. Experimental results in a real-world environment show that the method achieves an average speech recognition accuracy of 92.8%, with particularly outstanding performance in vowel and consonant recognition. Wu et al. provided a comprehensive technical review of the most advanced wireless voice sensing technologies, demonstrating the promising future of wireless sensing-based voiceprint recognition. Existing methods primarily focus on voice security, authentication, and speech recovery, with limited applications in speech recognition. Summary of the Invention

[0004] The purpose of this invention is to provide a voice command recognition method based on millimeter-wave sensing, which solves the limitations of existing technologies in processing independent voice commands and detecting weak sounds (such as patient voices), while overcoming the difficulties brought about by limited data volume in model design and expansion.

[0005] The technical solution adopted in this invention is a voice command recognition method based on millimeter-wave sensing, characterized by the following steps:

[0006] Step 1: Build a millimeter-wave radar data acquisition platform to collect echo data from different voice commands;

[0007] Step 2: Preprocess the echo data acquired in Step 1;

[0008] Step 3: Perform short-time Fourier transform on the preprocessed echo data of each range cell to extract the Doppler information of the target. Then, perform incoherent superposition of the Doppler spectra of all range cells to generate a micro-Doppler spectrum.

[0009] Step 4: Generate micro-Doppler spectra on all receiving antennas in Step 3, augment the data using an improved deep convolutional adversarial network, construct a dataset, and divide it into training, validation, and test sets in a ratio of 7:1.5:1.5.

[0010] Step 5: Construct the vocal cord vibration recognition network BAC-Net, integrating Bayesian optimization and convolutional block attention modules to focus on key frequency bands and micro-Doppler features, thereby improving hyperparameter tuning efficiency and recognition performance. Use the training set obtained in Step 4 as the input to the vocal cord vibration recognition network BAC-Net. The network obtains the optimal model by training on different speech command words, thereby recognizing different speech commands.

[0011] The invention is further characterized in that,

[0012] Step 1 is implemented in the following steps:

[0013] Step 1.1: Build a millimeter-wave radar data acquisition platform, including hardware architecture and software configuration. The hardware requires the installation of the Texas Instruments (TI) AWR1642BOOST evaluation module development board and DCA1000EVM data acquisition card, including the motherboard. The motherboard is equipped with two transmit antennas and four receive antennas. TI's official mmWave Studio 02.01.01.00 is installed on the motherboard, and the relevant millimeter-wave radar parameters are set as shown in the table below:

[0014]

[0015] Step 1.2: Design and set up the experimental scenario. The radar is placed on a tripod 1-1.2m above the ground and 0.5-0.8m radially away, 10-30cm in front of the target to be tested.

[0016] Step 1.3: To avoid errors that may be caused by signal segmentation, this invention uses a method of recording voice commands one by one to collect echo signals of different voice command words in an indoor scene. Each subject needs to say the required 10 command words in sequence. During the recording of each command word, the subject repeats the same command word 5 times according to the gesture guidance of the data collector. Data is automatically collected within a fixed 5-second interval. Each category has an average of 50 (5*10) sets of words. A multi-channel fusion method is used to integrate the signals of the 4 receiving antennas in the millimeter-wave radar. Therefore, each category has 200 (50*4) sets of words, which enhances the ability to perceive vocal cord micro-vibrations.

[0017] Step 2 is implemented in the following steps:

[0018] Step 2.1: Echo data of different voice command words were acquired in Step 1. This data is presented in a three-dimensional matrix. Formal representation, with the following dimensions:

[0019]

[0020] in, For fast sampling point indexing in the time dimension, For the Chirp index on the slow time dimension, For the index of the receive channel, The data type representing the matrix is ​​in complex form, meaning each data point can be represented as... It contains the amplitude and phase information of the signal. This represents the number of sampling points in the fast time dimension. Each fast time sampling point corresponds to a signal sample captured by the analog-to-digital converter (ADC) in one sampling period. Fast time is typically related to the distance information of the target. This represents the number of linear frequency modulated chirs in the slow time dimension. A chirp is a single frequency-modulated signal transmitted by the radar to obtain target velocity information. Indicates the number of receive channels;

[0021] Step 2.2: Perform a Fast Fourier Transform (FFT) on the echo data obtained in Step 2.1 in the fast time dimension to extract the target's range information, transforming the original ADC data from the time domain to the range domain. The specific calculation formula is as follows:

[0022]

[0023] For distance cell indexing, and sampling points in the fast time dimension One-to-one correspondence, ultimately generating a distance-time matrix. ;

[0024] Step 2.3: The distance-time matrix from Step 2.2 Static clutter is filtered out using the phasor averaging algorithm for all Chirp. Take the phasor mean, calculate the static clutter component, and subtract it from the original matrix:

[0025]

[0026] To filter out static clutter and improve the display of dynamic targets, a clearer dynamic target signal is obtained;

[0027] Step 2.4, regarding step 2.3 Remove the DC component; the formula for calculating the DC component is:

[0028]

[0029] from Subtract the calculated DC component from the result to eliminate background interference:

[0030]

[0031] Finally, the matrix with DC component removed is obtained. .

[0032] Step 3 is implemented in the following steps:

[0033] Step 3.1: Obtain the echo data matrix from Step 2.4. For each distance unit Perform a short-time Fourier transform on the target to extract its Doppler information:

[0034]

[0035] Distance unit ,frequency Receiving channel Micro-Doppler spectrum on;

[0036] Step 3.2: Based on the Doppler spectrum of each range cell obtained in Step 3.1, perform incoherent superposition of the Doppler spectra of all range cells to generate a micro-Doppler spectrum.

[0037]

[0038] Represents frequency and receiving channel Micro Doppler spectra on the surface It is the square of the amplitude of the micro-Doppler spectrum, used for incoherent accumulation to generate the final micro-Doppler spectrum.

[0039] Step 4 is implemented in the following steps:

[0040] Step 4.1: Construct an improved deep convolutional adversarial network;

[0041] Step 4.2: Input the micro-Doppler spectra generated in step 3.2 onto the improved DCGAN in step 4.1 to obtain the amplified micro-Doppler images.

[0042] Step 4.3: Merge the augmented data obtained in Step 4.2 with the original data, and divide the dataset according to the ratio of training set: test set: validation set = 7:1.5:1.5.

[0043] The specific structure of the improved deep convolutional adversarial network in step 4.1 is as follows:

[0044] The improved Deep Convolutional Adversarial Network (DCGAN) is used to augment the data, generating augmented data with the same number of micro-Doppler images per class as the original. The original DCGAN uses a sigmoid cross-entropy loss function, judging whether the input sample is correctly classified, but not penalizing the classification of generated samples. This leads to the gradient vanishing problem during GAN training. The improved DCGAN, however, uses a least squares loss function, with the corresponding discriminator loss function being:

[0045]

[0046] The loss function of the generator is:

[0047]

[0048] This represents the expected value or mean value. This indicates that the discriminator is effective against real samples. The output probability, This indicates that the discriminator evaluates the generated samples. The output probability, This represents a sample taken from the real data distribution. This represents a noise vector sampled from a noise distribution, with the first term... This indicates that the discriminator hopes to analyze real samples. The output should be as close to 1 as possible, the second term This indicates that the discriminator wants to analyze the generated samples. The output is made as close to 0 as possible. This network still provides learning error even when the discriminator misclassifies a sample, while penalizing correctly classified samples based on their distance from the decision boundary. This pulls generated samples far from the decision boundary closer to it, resulting in more gradients during generator updates, effectively mitigating the vanishing gradient problem. Furthermore, by pulling generated samples far from the decision boundary closer, the generated samples more closely resemble the distribution of real samples, effectively improving the quality of generated images and producing more realistic sample data.

[0049] The specific model structure of the improved deep convolutional adversarial network in step 4.1 is as follows:

[0050] The generator first receives 100-dimensional random Gaussian noise, fully connected to an 8×8×512 dimension, and then performs four alternating upsampling and convolution operations. The upsampling uses a transposed convolution kernel of size 4x4 with a stride of 2, and the number of kernels is 512, 256, 128, and 64 respectively. Finally, the generated micro-Doppler image is output after passing through a convolutional layer with a kernel size of 3x3 and a stride of 1. The discriminator receives the generated sample and the original sample data, and performs four convolution operations. The number of kernels is 64, 128, 256 and 512 respectively. Each time, a convolutional kernel with a kernel size of 4x4, a stride of 2 and a padding of 1 is used to extract image features. Then, a convolutional layer with a kernel size of 4x4 and a stride of 1 is used to compress the feature map into a 1x1x1 output. The Sigmoid activation function is used to output the real or fake image judgment result. Through continuous iterative optimization between the generator and the discriminator, the two reach Nash equilibrium, that is, the generated image that meets the requirements is obtained.

[0051] Step 5 is implemented in the following steps:

[0052] Step 5.1: The vocal cord vibration recognition network BAC-Net is based on a two-dimensional convolutional neural network (2D-CNN). It utilizes a Bayesian optimization algorithm to dynamically adjust the model's key hyperparameters, thereby efficiently finding the optimal hyperparameter combination suitable for a specific RGB image classification task. Based on this, a Convolutional Block Attention (CBAM) module is added after each convolutional operation. An optimization algorithm dynamically determines whether to introduce the CBAM module after each convolutional layer and optimizes the channel and spatial weight calculation details within the CBAM. This module combines channel attention and spatial attention mechanisms to dynamically adjust the weights of the feature map, guiding the network to focus more on key features, thereby improving the model's classification performance and generalization ability. Key hyperparameters include convolutional kernel size, number of channels per layer, optimizer selection, Dropout rate, and learning rate. Furthermore, a pruning mechanism is introduced; in each round of trials, hyperparameter combinations with a classification accuracy below 10% are directly removed to reduce the interference of ineffective combinations on the overall optimization.

[0053] Step 5.2: Input the training set from the dataset constructed in Step 4 into the BAC-Net constructed in Step 5.1 for training. This network integrates Bayesian optimization and convolutional block attention modules, focuses on key frequency bands and micro-Doppler features, improves hyperparameter tuning efficiency and recognition performance. The network obtains the optimal model for the validation set by training on different voice command words, and finally uses the test set for detection to achieve the recognition and classification of different voice commands, specifically including "hello", "goodbye", "yes", "no", "eat", "drink", "toilet", "thank you", "sleep" and "help".

[0054] The beneficial effects of this invention are that the millimeter-wave sensing-based voice command recognition method, abbreviated as mmMedVoice, aims to overcome the limitations of existing millimeter-wave radar-based voice recognition technologies in processing independent voice commands and detecting weak sounds (such as patient voices), while also overcoming the difficulties in model design and expansion caused by limited data. This method employs a 77GHz multi-input multi-output frequency-modulated continuous-wave millimeter-wave radar to sense human vocal cord vibration in a non-contact manner, and studies the correlation between speech and the micro-motion characteristics of vocal cord vibration. By sensing the phase change of the vocal cord vibration signal, the frequency-time spectrum is accurately extracted from the millimeter-wave single-channel signal, realizing the extraction of echo signals from the minute vibrations of the human vocal cords. This method not only effectively reduces costs and avoids synchronization problems between multiple sensors, but also achieves efficient monitoring while protecting the privacy of the subjects. This method includes: building a millimeter-wave radar data acquisition platform to collect echo data; preprocessing the echo data by generating a range-time matrix through Fast Fourier Transform (FFT), and using a phasor mean algorithm to filter out static clutter and remove DC components; extracting Doppler information for each range cell through Short-Time Fourier Transform (SFT), and incoherently superimposing the Doppler spectra of all range cells to generate a micro-Doppler spectrum; using an improved deep convolutional adversarial network (DAN) to amplify the micro-Doppler data, constructing and partitioning the dataset; and proposing a vocal cord vibration recognition network, BAC-Net, which integrates Bayesian optimization and convolutional block attention modules to focus on key frequency bands and micro-Doppler features, improving hyperparameter tuning efficiency and recognition performance. Comprehensive experiments show that, in an indoor environment, the accuracy reaches 94%, surpassing the most advanced methods and six classic deep learning models. Attached Figure Description

[0055] Figure 1 This is a framework diagram of the voice command recognition method based on millimeter-wave sensing of the present invention;

[0056] Figure 2 This is the micro-Doppler image corresponding to the "drink water" command word provided in the embodiments of the present invention;

[0057] Figure 3 This is the micro-Doppler image corresponding to the command word "eat" provided in the embodiments of the present invention;

[0058] Figure 4(a) is a real sample of the "Hello" command word provided in the embodiment of the present invention;

[0059] Figure 4(b) shows the generated sample after DCGAN enhancement of Figure 4(a);

[0060] Figure 5 This is a diagram of the optimized neural network model provided in an embodiment of the present invention;

[0061] Figure 6 This is the confusion matrix of the detection results of the BAC-Net model provided in this embodiment of the invention. Detailed Implementation

[0062] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0063] This invention relates to a voice command recognition method based on millimeter-wave sensing, combined with Figure 1 The specific steps are as follows:

[0064] Step 1: Build a millimeter-wave radar data acquisition platform to collect echo data from different voice commands;

[0065] Step 1 is implemented in the following steps:

[0066] Step 1.1: Build a millimeter-wave radar data acquisition platform, including hardware architecture and software configuration. The hardware requires the installation of the Texas Instruments (TI) AWR1642BOOST evaluation module development board and DCA1000EVM data acquisition card, including the motherboard. The motherboard is equipped with two transmit antennas and four receive antennas. TI's official mmWave Studio 02.01.01.00 is installed on the motherboard, and the relevant millimeter-wave radar parameters are set as shown in the table below:

[0067]

[0068] Step 1.2: Design and set up the experimental scenario. The radar is placed on a tripod 1-1.2m above the ground and 0.5-0.8m radially away, 10-30cm in front of the target to be tested.

[0069] Step 1.3: To avoid errors that may be caused by signal segmentation, this invention uses a method of recording voice commands one by one to collect echo signals of different voice command words in an indoor scene. Each subject needs to say the required 10 command words in sequence. During the recording of each command word, the subject repeats the same command word 5 times according to the gesture guidance of the data collector. Data is automatically collected within a fixed 5-second interval. Each category has an average of 50 (5*10) sets of words. A multi-channel fusion method is used to integrate the signals of the 4 receiving antennas in the millimeter-wave radar. Therefore, each category has 200 (50*4) sets of words, which enhances the ability to perceive vocal cord micro-vibrations.

[0070] Step 2: Preprocess the echo data acquired in Step 1;

[0071] Step 2 is implemented in the following steps:

[0072] Step 2.1: Echo data of different voice command words were acquired in Step 1. This data is presented in a three-dimensional matrix. Formal representation, with the following dimensions:

[0073]

[0074] in, For fast sampling point indexing in the time dimension, For the Chirp index on the slow time dimension, For the index of the receive channel, The data type representing the matrix is ​​in complex form, meaning each data point can be represented as... It contains the amplitude and phase information of the signal. This represents the number of sampling points in the fast time dimension. Each fast time sampling point corresponds to a signal sample captured by the analog-to-digital converter (ADC) in one sampling period. Fast time is typically related to the distance information of the target. This represents the number of linear frequency modulated chirs in the slow time dimension. A chirp is a single frequency-modulated signal transmitted by the radar to obtain target velocity information. Indicates the number of receive channels;

[0075] Step 2.2: Perform a Fast Fourier Transform (FFT) on the echo data obtained in Step 2.1 in the fast time dimension to extract the target's range information, transforming the original ADC data from the time domain to the range domain. The specific calculation formula is as follows:

[0076]

[0077] For distance cell indexing, and sampling points in the fast time dimension One-to-one correspondence, ultimately generating a distance-time matrix. ;

[0078] Step 2.3: The distance-time matrix from Step 2.2 Static clutter is filtered out using the phasor averaging algorithm for all Chirp. Take the phasor mean, calculate the static clutter component, and subtract it from the original matrix:

[0079]

[0080] To filter out static clutter and improve the display of dynamic targets, a clearer dynamic target signal is obtained;

[0081] Step 2.4, regarding step 2.3 Remove the DC component; the formula for calculating the DC component is:

[0082]

[0083] from Subtract the calculated DC component from the result to eliminate background interference:

[0084]

[0085] Finally, the matrix with DC component removed is obtained. .

[0086] Step 3: Perform short-time Fourier transform on the preprocessed echo data of each range cell to extract the Doppler information of the target. Then, perform incoherent superposition of the Doppler spectra of all range cells to generate a micro-Doppler spectrum.

[0087] Step 3 is implemented in the following steps:

[0088] Step 3.1: Obtain the echo data matrix from Step 2.4. For each distance unit Perform a short-time Fourier transform on the target to extract its Doppler information:

[0089]

[0090] Distance unit ,frequency Receiving channel Micro-Doppler spectrum on;

[0091] Step 3.2: Based on the Doppler spectrum of each range cell obtained in Step 3.1, perform incoherent superposition of the Doppler spectra of all range cells to generate a micro-Doppler spectrum.

[0092]

[0093] Represents frequency and receiving channel Micro Doppler spectra on the surface It is the square of the amplitude of the micro-Doppler spectrum, used for incoherent accumulation to generate the final micro-Doppler spectrum.

[0094] Step 4: Generate micro-Doppler spectra on all receiving antennas in Step 3, augment the data using an improved deep convolutional adversarial network, construct a dataset, and divide it into training, validation, and test sets in a ratio of 7:1.5:1.5.

[0095] Step 4 is implemented in the following steps:

[0096] Step 4.1: Construct an improved deep convolutional adversarial network;

[0097] The specific structure of the improved deep convolutional adversarial network in step 4.1 is as follows:

[0098] The improved Deep Convolutional Adversarial Network (DCGAN) is used to augment the data, generating augmented data with the same number of micro-Doppler images per class as the original. The original DCGAN uses a sigmoid cross-entropy loss function, judging whether the input sample is correctly classified, but not penalizing the classification of generated samples. This leads to the gradient vanishing problem during GAN training. The improved DCGAN, however, uses a least squares loss function, with the corresponding discriminator loss function being:

[0099]

[0100] The loss function of the generator is:

[0101]

[0102] This represents the expected value or mean value. This indicates that the discriminator is effective against real samples. The output probability, This indicates that the discriminator evaluates the generated samples. The output probability, This represents a sample taken from the real data distribution. This represents a noise vector sampled from a noise distribution, with the first term... This indicates that the discriminator hopes to analyze real samples. The output should be as close to 1 as possible, the second term This indicates that the discriminator wants to analyze the generated samples. The output is made as close to 0 as possible. This network still provides learning error even when the discriminator misclassifies a sample, while penalizing correctly classified samples based on their distance from the decision boundary. This pulls generated samples far from the decision boundary closer to it, resulting in more gradients during generator updates, effectively mitigating the vanishing gradient problem. Furthermore, by pulling generated samples far from the decision boundary closer, the generated samples more closely resemble the distribution of real samples, effectively improving the quality of generated images and producing more realistic sample data.

[0103] The specific model structure of the improved deep convolutional adversarial network in step 4.1 is as follows:

[0104] The generator first receives 100-dimensional random Gaussian noise, fully connected to an 8×8×512 dimension, and then performs four alternating upsampling and convolution operations. The upsampling uses a transposed convolution kernel of size 4x4 with a stride of 2, and the number of kernels is 512, 256, 128, and 64 respectively. Finally, the generated micro-Doppler image is output after passing through a convolutional layer with a kernel size of 3x3 and a stride of 1. The discriminator receives the generated sample and the original sample data, and performs four convolution operations. The number of kernels is 64, 128, 256 and 512 respectively. Each time, a convolutional kernel with a kernel size of 4x4, a stride of 2 and a padding of 1 is used to extract image features. Then, a convolutional layer with a kernel size of 4x4 and a stride of 1 is used to compress the feature map into a 1x1x1 output. The Sigmoid activation function is used to output the real or fake image judgment result. Through continuous iterative optimization between the generator and the discriminator, the two reach Nash equilibrium, that is, the generated image that meets the requirements is obtained.

[0105] Step 4.2: Input the micro-Doppler spectra generated in step 3.2 onto the improved DCGAN in step 4.1 to obtain the amplified micro-Doppler images.

[0106] Step 4.3: Merge the augmented data obtained in Step 4.2 with the original data, and divide the dataset according to the ratio of training set: test set: validation set = 7:1.5:1.5.

[0107] Step 5: Construct the vocal cord vibration recognition network BAC-Net, integrating Bayesian optimization and convolutional block attention modules to focus on key frequency bands and micro-Doppler features, thereby improving hyperparameter tuning efficiency and recognition performance. Use the training set obtained in Step 4 as the input to the vocal cord vibration recognition network BAC-Net. The network obtains the optimal model by training on different speech command words, thereby recognizing different speech commands.

[0108] Step 5 is implemented in the following steps:

[0109] Step 5.1: The vocal cord vibration recognition network BAC-Net is based on a two-dimensional convolutional neural network (2D-CNN). It utilizes a Bayesian optimization algorithm to dynamically adjust the model's key hyperparameters, thereby efficiently finding the optimal hyperparameter combination suitable for a specific RGB image classification task. Based on this, a Convolutional Block Attention (CBAM) module is added after each convolutional operation. An optimization algorithm dynamically determines whether to introduce the CBAM module after each convolutional layer and optimizes the channel and spatial weight calculation details within the CBAM. This module combines channel attention and spatial attention mechanisms to dynamically adjust the weights of the feature map, guiding the network to focus more on key features, thereby improving the model's classification performance and generalization ability. Key hyperparameters include convolutional kernel size, number of channels per layer, optimizer selection, Dropout rate, and learning rate. Furthermore, a pruning mechanism is introduced; in each round of trials, hyperparameter combinations with a classification accuracy below 10% are directly removed to reduce the interference of ineffective combinations on the overall optimization.

[0110] Step 5.2: Input the training set from the dataset constructed in Step 4 into the BAC-Net constructed in Step 5.1 for training. This network integrates Bayesian optimization and convolutional block attention modules, focuses on key frequency bands and micro-Doppler features, improves hyperparameter tuning efficiency and recognition performance. The network obtains the optimal model for the validation set by training on different voice command words, and finally uses the test set for detection to achieve the recognition and classification of different voice commands.

[0111] Example 1

[0112] This invention relates to a voice command recognition method based on millimeter-wave sensing, combined with Figure 1 The specific steps are as follows:

[0113] Step 1: Build a millimeter-wave radar data acquisition platform to collect echo data from different voice commands;

[0114] Step 2: Preprocess the echo data acquired in Step 1;

[0115] Step 3: Perform short-time Fourier transform on the preprocessed echo data of each range cell to extract the Doppler information of the target. Then, perform incoherent superposition of the Doppler spectra of all range cells to generate a micro-Doppler spectrum.

[0116] Step 4: Generate micro-Doppler spectra on all receiving antennas in Step 3, augment the data using an improved deep convolutional adversarial network, construct a dataset, and divide it into training, validation, and test sets in a ratio of 7:1.5:1.5.

[0117] Step 5: Construct the vocal cord vibration recognition network BAC-Net, integrating Bayesian optimization and convolutional block attention modules to focus on key frequency bands and micro-Doppler features, thereby improving hyperparameter tuning efficiency and recognition performance. Use the training set obtained in Step 4 as the input to the vocal cord vibration recognition network BAC-Net. The network obtains the optimal model by training on different speech command words, thereby recognizing different speech commands.

[0118] Example 2

[0119] This invention relates to a voice command recognition method based on millimeter-wave sensing, combined with Figure 1 The specific steps are as follows:

[0120] Step 1: Build a millimeter-wave radar data acquisition platform to collect echo data from different voice commands;

[0121] Step 1 is implemented in the following steps:

[0122] Step 1.1: Build a millimeter-wave radar data acquisition platform, including hardware architecture and software configuration. The hardware requires the installation of the Texas Instruments (TI) AWR1642BOOST evaluation module development board and DCA1000EVM data acquisition card, including the motherboard. The motherboard is equipped with two transmit antennas and four receive antennas. TI's official mmWave Studio 02.01.01.00 is installed on the motherboard, and the relevant millimeter-wave radar parameters are set as shown in the table below:

[0123]

[0124] Step 1.2: Design and set up the experimental scenario. The radar is placed on a tripod 1-1.2m above the ground and 0.5-0.8m radially away, 10-30cm in front of the target to be tested.

[0125] Step 1.3: To avoid errors that may be caused by signal segmentation, this invention uses a method of recording voice commands one by one to collect echo signals of different voice command words in an indoor scene. Each subject needs to say the required 10 command words in sequence. During the recording of each command word, the subject repeats the same command word 5 times according to the gesture guidance of the data collector. Data is automatically collected within a fixed 5-second interval. Each category has an average of 50 (5*10) sets of words. A multi-channel fusion method is used to integrate the signals of the 4 receiving antennas in the millimeter-wave radar. Therefore, each category has 200 (50*4) sets of words, which enhances the ability to perceive vocal cord micro-vibrations.

[0126] Step 2: Preprocess the echo data acquired in Step 1;

[0127] Step 3: Perform short-time Fourier transform on the preprocessed echo data of each range cell to extract the Doppler information of the target. Then, perform incoherent superposition of the Doppler spectra of all range cells to generate a micro-Doppler spectrum.

[0128] Step 3 is implemented in the following steps:

[0129] Step 3.1: Obtain the echo data matrix from Step 2.4. For each distance unit Perform a short-time Fourier transform on the target to extract its Doppler information:

[0130]

[0131] Distance unit ,frequency Receiving channel Micro-Doppler spectrum on;

[0132] Step 3.2: Based on the Doppler spectrum of each range cell obtained in Step 3.1, perform incoherent superposition of the Doppler spectra of all range cells to generate a micro-Doppler spectrum.

[0133]

[0134] Represents frequency and receiving channel Micro Doppler spectra on the surface It is the square of the amplitude of the micro-Doppler spectrum, used for incoherent accumulation to generate the final micro-Doppler spectrum.

[0135] Step 4: Generate micro-Doppler spectra on all receiving antennas in Step 3, augment the data using an improved deep convolutional adversarial network, construct a dataset, and divide it into training, validation, and test sets in a ratio of 7:1.5:1.5.

[0136] Step 5: Construct the vocal cord vibration recognition network BAC-Net, integrating Bayesian optimization and convolutional block attention modules to focus on key frequency bands and micro-Doppler features, thereby improving hyperparameter tuning efficiency and recognition performance. Use the training set obtained in Step 4 as the input to the vocal cord vibration recognition network BAC-Net. The network obtains the optimal model by training on different speech command words, thereby recognizing different speech commands.

[0137] Example 3

[0138] This invention relates to a voice command recognition method based on millimeter-wave sensing, combined with Figure 1 The specific steps are as follows:

[0139] Step 1: Build a millimeter-wave radar data acquisition platform to collect echo data from different voice commands;

[0140] Step 1 is implemented in the following steps:

[0141] Step 1.1: Build a millimeter-wave radar data acquisition platform, including hardware architecture and software configuration. The hardware requires the installation of the Texas Instruments (TI) AWR1642BOOST evaluation module development board and DCA1000EVM data acquisition card, including the motherboard. The motherboard is equipped with two transmit antennas and four receive antennas. TI's official mmWave Studio 02.01.01.00 is installed on the motherboard, and the relevant millimeter-wave radar parameters are set as shown in the table below:

[0142]

[0143] Step 1.2: Design and set up the experimental scenario. The radar is placed on a tripod 1-1.2m above the ground and 0.5-0.8m radially away, 10-30cm in front of the target to be tested.

[0144] Step 1.3: To avoid errors that may be caused by signal segmentation, this invention uses a method of recording voice commands one by one to collect echo signals of different voice command words in an indoor scene. Each subject needs to say the required 10 command words in sequence. During the recording of each command word, the subject repeats the same command word 5 times according to the gesture guidance of the data collector. Data is automatically collected within a fixed 5-second interval. Each category has an average of 50 (5*10) sets of words. A multi-channel fusion method is used to integrate the signals of the 4 receiving antennas in the millimeter-wave radar. Therefore, each category has 200 (50*4) sets of words, which enhances the ability to perceive vocal cord micro-vibrations.

[0145] Step 2: Preprocess the echo data acquired in Step 1;

[0146] Step 2 is implemented in the following steps:

[0147] Step 2.1: Echo data of different voice command words were acquired in Step 1. This data is presented in a three-dimensional matrix. Formal representation, with the following dimensions:

[0148]

[0149] in, For fast sampling point indexing in the time dimension, For the Chirp index on the slow time dimension, For the index of the receive channel, The data type representing the matrix is ​​in complex form, meaning each data point can be represented as... It contains the amplitude and phase information of the signal. This represents the number of sampling points in the fast time dimension. Each fast time sampling point corresponds to a signal sample captured by the analog-to-digital converter (ADC) in one sampling period. Fast time is typically related to the distance information of the target. This represents the number of linear frequency modulated chirs in the slow time dimension. A chirp is a single frequency-modulated signal transmitted by the radar to obtain target velocity information. Indicates the number of receive channels;

[0150] Step 2.2: Perform a Fast Fourier Transform (FFT) on the echo data obtained in Step 2.1 in the fast time dimension to extract the target's range information, transforming the original ADC data from the time domain to the range domain. The specific calculation formula is as follows:

[0151]

[0152] For distance cell indexing, and sampling points in the fast time dimension One-to-one correspondence, ultimately generating a distance-time matrix. ;

[0153] Step 2.3: The distance-time matrix from Step 2.2 Static clutter is filtered out using the phasor averaging algorithm for all Chirp. Take the phasor mean, calculate the static clutter component, and subtract it from the original matrix:

[0154]

[0155] To filter out static clutter and improve the display of dynamic targets, a clearer dynamic target signal is obtained;

[0156] Step 2.4, regarding step 2.3 Remove the DC component; the formula for calculating the DC component is:

[0157]

[0158] from Subtract the calculated DC component from the result to eliminate background interference:

[0159]

[0160] Finally, the matrix with DC component removed is obtained. .

[0161] Step 3: Perform short-time Fourier transform on the preprocessed echo data of each range cell to extract the Doppler information of the target. Then, perform incoherent superposition of the Doppler spectra of all range cells to generate a micro-Doppler spectrum.

[0162] Step 3 is implemented in the following steps:

[0163] Step 3.1: Obtain the echo data matrix from Step 2.4. For each distance unit Perform a short-time Fourier transform on the target to extract its Doppler information:

[0164]

[0165] Distance unit ,frequency Receiving channel Micro-Doppler spectrum on;

[0166] Step 3.2: Based on the Doppler spectrum of each range cell obtained in Step 3.1, perform incoherent superposition of the Doppler spectra of all range cells to generate a micro-Doppler spectrum.

[0167]

[0168] Represents frequency and receiving channel Micro Doppler spectra on the surface It is the square of the amplitude of the micro-Doppler spectrum, used for incoherent accumulation to generate the final micro-Doppler spectrum.

[0169] Step 4: Generate micro-Doppler spectra on all receiving antennas in Step 3, augment the data using an improved deep convolutional adversarial network, construct a dataset, and divide it into training, validation, and test sets in a ratio of 7:1.5:1.5.

[0170] Step 5: Construct the vocal cord vibration recognition network BAC-Net, integrating Bayesian optimization and convolutional block attention modules to focus on key frequency bands and micro-Doppler features, thereby improving hyperparameter tuning efficiency and recognition performance. Use the training set obtained in Step 4 as the input to the vocal cord vibration recognition network BAC-Net. The network obtains the optimal model by training on different speech command words, thereby recognizing different speech commands.

[0171] Example 4

[0172] This invention relates to a voice command recognition method based on millimeter-wave sensing, combined with Figure 1 The specific steps are as follows:

[0173] Step 1: Build a millimeter-wave radar data acquisition platform to collect echo data from different voice commands;

[0174] Step 1 is implemented in the following steps:

[0175] Step 1.1: Build a millimeter-wave radar data acquisition platform, including hardware architecture and software configuration. The hardware requires the installation of the Texas Instruments (TI) AWR1642BOOST evaluation module development board and DCA1000EVM data acquisition card, including the motherboard. The motherboard is equipped with two transmit antennas and four receive antennas. TI's official mmWave Studio 02.01.01.00 is installed on the motherboard, and the relevant millimeter-wave radar parameters are set as shown in the table below:

[0176]

[0177] Step 1.2: Design and set up the experimental scenario. The radar is placed on a tripod 1-1.2m above the ground and 0.5-0.8m radially away, 10-30cm in front of the target to be tested.

[0178] Step 1.3: To avoid errors that may be caused by signal segmentation, this invention uses a method of recording voice commands one by one to collect echo signals of different voice command words in an indoor scene. Each subject needs to say the required 10 command words in sequence. During the recording of each command word, the subject repeats the same command word 5 times according to the gesture guidance of the data collector. Data is automatically collected within a fixed 5-second interval. Each category has an average of 50 (5*10) sets of words. A multi-channel fusion method is used to integrate the signals of the 4 receiving antennas in the millimeter-wave radar. Therefore, each category has 200 (50*4) sets of words, which enhances the ability to perceive vocal cord micro-vibrations.

[0179] Step 2: Preprocess the echo data acquired in Step 1;

[0180] Step 2 is implemented in the following steps:

[0181] Step 2.1: Echo data of different voice command words were acquired in Step 1. This data is presented in a three-dimensional matrix. Formal representation, with the following dimensions:

[0182]

[0183] in, For fast sampling point indexing in the time dimension, For the Chirp index on the slow time dimension, For the index of the receive channel, The data type representing the matrix is ​​in complex form, meaning each data point can be represented as... It contains the amplitude and phase information of the signal. This represents the number of sampling points in the fast time dimension. Each fast time sampling point corresponds to a signal sample captured by the analog-to-digital converter (ADC) in one sampling period. Fast time is typically related to the distance information of the target. This represents the number of linear frequency modulated chirs in the slow time dimension. A chirp is a single frequency-modulated signal transmitted by the radar to obtain target velocity information. Indicates the number of receive channels;

[0184] Step 2.2: Perform a Fast Fourier Transform (FFT) on the echo data obtained in Step 2.1 in the fast time dimension to extract the target's range information, transforming the original ADC data from the time domain to the range domain. The specific calculation formula is as follows:

[0185]

[0186] For distance cell indexing, and sampling points in the fast time dimension One-to-one correspondence, ultimately generating a distance-time matrix. ;

[0187] Step 2.3: The distance-time matrix from Step 2.2 Static clutter is filtered out using the phasor averaging algorithm for all Chirp. Take the phasor mean, calculate the static clutter component, and subtract it from the original matrix:

[0188]

[0189] To filter out static clutter and improve the display of dynamic targets, a clearer dynamic target signal is obtained;

[0190] Step 2.4, regarding step 2.3 Remove the DC component; the formula for calculating the DC component is:

[0191]

[0192] from Subtract the calculated DC component from the result to eliminate background interference:

[0193]

[0194] Finally, the matrix with DC component removed is obtained. .

[0195] Step 3: Perform short-time Fourier transform on the preprocessed echo data of each range cell to extract the Doppler information of the target. Then, perform incoherent superposition of the Doppler spectra of all range cells to generate a micro-Doppler spectrum.

[0196] Step 3 is implemented in the following steps:

[0197] Step 3.1: Obtain the echo data matrix from Step 2.4. For each distance unit Perform a short-time Fourier transform on the target to extract its Doppler information:

[0198]

[0199] Distance unit ,frequency Receiving channel Micro-Doppler spectrum on;

[0200] Step 3.2: Based on the Doppler spectrum of each range cell obtained in Step 3.1, perform incoherent superposition of the Doppler spectra of all range cells to generate a micro-Doppler spectrum.

[0201]

[0202] Represents frequency and receiving channel Micro Doppler spectra on the surface It is the square of the amplitude of the micro-Doppler spectrum, used for incoherent accumulation to generate the final micro-Doppler spectrum.

[0203] Step 4: Generate micro-Doppler spectra on all receiving antennas in Step 3, augment the data using an improved deep convolutional adversarial network, construct a dataset, and divide it into training, validation, and test sets in a ratio of 7:1.5:1.5.

[0204] Step 4 is implemented in the following steps:

[0205] Step 4.1: Construct an improved deep convolutional adversarial network;

[0206] The specific structure of the improved deep convolutional adversarial network in step 4.1 is as follows:

[0207] The improved Deep Convolutional Adversarial Network (DCGAN) is used to augment the data, generating augmented data with the same number of micro-Doppler images per class as the original. The original DCGAN uses a sigmoid cross-entropy loss function, judging whether the input sample is correctly classified, but not penalizing the classification of generated samples. This leads to the gradient vanishing problem during GAN training. The improved DCGAN, however, uses a least squares loss function, with the corresponding discriminator loss function being:

[0208]

[0209] The loss function of the generator is:

[0210]

[0211] This represents the expected value or mean value. This indicates that the discriminator is effective against real samples. The output probability, This indicates that the discriminator evaluates the generated samples. The output probability, This represents a sample taken from the real data distribution. This represents a noise vector sampled from a noise distribution, with the first term... This indicates that the discriminator hopes to analyze real samples. The output should be as close to 1 as possible, the second term This indicates that the discriminator wants to analyze the generated samples. The output is made as close to 0 as possible. This network still provides learning error even when the discriminator misclassifies a sample, while penalizing correctly classified samples based on their distance from the decision boundary. This pulls generated samples far from the decision boundary closer to it, resulting in more gradients during generator updates, effectively mitigating the vanishing gradient problem. Furthermore, by pulling generated samples far from the decision boundary closer, the generated samples more closely resemble the distribution of real samples, effectively improving the quality of generated images and producing more realistic sample data.

[0212] Step 4.2: Input the micro-Doppler spectra generated in step 3.2 onto the improved DCGAN in step 4.1 to obtain the amplified micro-Doppler images.

[0213] Step 4.3: Merge the augmented data obtained in Step 4.2 with the original data, and divide the dataset according to the ratio of training set: test set: validation set = 7:1.5:1.5.

[0214] Step 5: Construct the vocal cord vibration recognition network BAC-Net, integrating Bayesian optimization and convolutional block attention modules to focus on key frequency bands and micro-Doppler features, thereby improving hyperparameter tuning efficiency and recognition performance. Use the training set obtained in Step 4 as the input to the vocal cord vibration recognition network BAC-Net. The network obtains the optimal model by training on different speech command words, thereby recognizing different speech commands.

[0215] Example 5

[0216] This invention relates to a voice command recognition method based on millimeter-wave sensing, combined with Figure 1 The specific steps are as follows:

[0217] Step 1: Build a millimeter-wave radar data acquisition platform to collect echo data from different voice commands;

[0218] Step 1 is implemented in the following steps:

[0219] Step 1.1: Build a millimeter-wave radar data acquisition platform, including hardware architecture and software configuration. The hardware requires the installation of the Texas Instruments (TI) AWR1642BOOST evaluation module development board and DCA1000EVM data acquisition card, including the motherboard. The motherboard is equipped with two transmit antennas and four receive antennas. TI's official mmWave Studio 02.01.01.00 is installed on the motherboard, and the relevant millimeter-wave radar parameters are set as shown in the table below:

[0220]

[0221] Step 1.2: Design and set up the experimental scenario. The radar is placed on a tripod 1-1.2m above the ground and 0.5-0.8m radially away, 10-30cm in front of the target to be tested.

[0222] Step 1.3: To avoid errors that may be caused by signal segmentation, this invention uses a method of recording voice commands one by one to collect echo signals of different voice command words in an indoor scene. Each subject needs to say the required 10 command words in sequence. During the recording of each command word, the subject repeats the same command word 5 times according to the gesture guidance of the data collector. Data is automatically collected within a fixed 5-second interval. Each category has an average of 50 (5*10) sets of words. A multi-channel fusion method is used to integrate the signals of the 4 receiving antennas in the millimeter-wave radar. Therefore, each category has 200 (50*4) sets of words, which enhances the ability to perceive vocal cord micro-vibrations.

[0223] Step 2: Preprocess the echo data acquired in Step 1;

[0224] Step 2 is implemented in the following steps:

[0225] Step 2.1: Echo data of different voice command words were acquired in Step 1. This data is presented in a three-dimensional matrix. Formal representation, with the following dimensions:

[0226]

[0227] in, For fast sampling point indexing in the time dimension, For the Chirp index on the slow time dimension, For the index of the receive channel, The data type representing the matrix is ​​in complex form, meaning each data point can be represented as... It contains the amplitude and phase information of the signal. This represents the number of sampling points in the fast time dimension. Each fast time sampling point corresponds to a signal sample captured by the analog-to-digital converter (ADC) in one sampling period. Fast time is typically related to the distance information of the target. This represents the number of linear frequency modulated chirs in the slow time dimension. A chirp is a single frequency-modulated signal transmitted by the radar to obtain target velocity information. Indicates the number of receive channels;

[0228] Step 2.2: Perform a Fast Fourier Transform (FFT) on the echo data obtained in Step 2.1 in the fast time dimension to extract the target's range information, transforming the original ADC data from the time domain to the range domain. The specific calculation formula is as follows:

[0229]

[0230] For distance cell indexing, and sampling points in the fast time dimension One-to-one correspondence, ultimately generating a distance-time matrix. ;

[0231] Step 2.3: The distance-time matrix from Step 2.2 Static clutter is filtered out using the phasor averaging algorithm for all Chirp. Take the phasor mean, calculate the static clutter component, and subtract it from the original matrix:

[0232]

[0233] To filter out static clutter and improve the display of dynamic targets, a clearer dynamic target signal is obtained;

[0234] Step 2.4, regarding step 2.3 Remove the DC component; the formula for calculating the DC component is:

[0235]

[0236] from Subtract the calculated DC component from the result to eliminate background interference:

[0237]

[0238] Finally, the matrix with DC component removed is obtained. .

[0239] Step 3: Perform short-time Fourier transform on the preprocessed echo data of each range cell to extract the Doppler information of the target. Then, perform incoherent superposition of the Doppler spectra of all range cells to generate a micro-Doppler spectrum.

[0240] Step 4: Generate micro-Doppler spectra on all receiving antennas in Step 3, augment the data using an improved deep convolutional adversarial network, construct a dataset, and divide it into training, validation, and test sets in a ratio of 7:1.5:1.5.

[0241] Step 5: Construct the vocal cord vibration recognition network BAC-Net, integrating Bayesian optimization and convolutional block attention modules to focus on key frequency bands and micro-Doppler features, thereby improving hyperparameter tuning efficiency and recognition performance. Use the training set obtained in Step 4 as the input to the vocal cord vibration recognition network BAC-Net. The network obtains the optimal model by training on different speech command words, thereby recognizing different speech commands.

[0242] Step 5 is implemented in the following steps:

[0243] Step 5.1: The vocal cord vibration recognition network BAC-Net is based on a two-dimensional convolutional neural network (2D-CNN). It utilizes a Bayesian optimization algorithm to dynamically adjust the model's key hyperparameters. A Block Convolutional Attention (CBAM) module is added after each convolutional operation. The optimization algorithm dynamically determines whether to introduce the CBAM module after each convolutional layer and optimizes the channel and spatial weight calculation details within the CBAM. This module combines channel attention and spatial attention mechanisms to dynamically adjust the weights of the feature maps, guiding the network to focus more on key features, thereby improving the model's classification performance and generalization ability. Key hyperparameters include convolutional kernel size, number of channels per layer, optimizer selection, Dropout rate, and learning rate. Furthermore, a pruning mechanism is introduced; in each round of trials, hyperparameter combinations with a classification accuracy below 10% are directly removed to reduce the interference of ineffective combinations on the overall optimization.

[0244] Step 5.2: Input the training set from the dataset constructed in Step 4 into the BAC-Net constructed in Step 5.1 for training. This network integrates Bayesian optimization and convolutional block attention modules, focuses on key frequency bands and micro-Doppler features, improves hyperparameter tuning efficiency and recognition performance. The network obtains the optimal model for the validation set by training on different voice command words, and finally uses the test set for detection to achieve the recognition and classification of different voice commands, specifically including "hello", "goodbye", "yes", "no", "eat", "drink", "toilet", "thank you", "sleep" and "help". Figure 2 This is the micro-Doppler image corresponding to the command "drink water"; Figure 3 Figure 4(a) is the micro-Doppler image corresponding to the command word "eat"; Figure 4(b) is the real sample of the command word "hello"; Figure 4(a) is the generated sample of Figure 4(a) after being enhanced by the improved DCGAN.

[0245] Example 6

[0246] Figure 5 This is a diagram of the optimized neural network model provided in this embodiment of the invention. The optimizer uses RootMean Square Propagation with a learning rate of 0.00012. The dropout of the first and second fully connected layers is 0.4. The epoch is 150, the batch size is 16, and there are 4 convolutional blocks with a kernel size of 3*3. The number of channels in each layer is 64, 32, 45, and 46, respectively.

[0247] Figure 6 This is the confusion matrix of the BAC-Net model detection results provided in this embodiment of the invention. The numbers 0-9 represent the 10 speech command words for classification. As can be seen, the values ​​on the main diagonal are mostly above 0.90, indicating that the model has a high recognition accuracy in most categories and can effectively distinguish between them, especially categories 0, 4, 6, 7, and 8, with a recognition rate of 95% or higher. The recognition rate for category 1 is 91%, with 3% of samples misclassified as category 2, indicating some confusion between categories 1 and 2. The recognition rate for category 3 is 90%, with 6% of samples misclassified as category 5; simultaneously, the recognition rate for category 5 is 87%, with 3% of samples misclassified as category 3. This indicates difficulty in distinguishing between categories 3 and 5, possibly due to the similarity of their features. The recognition rates for categories 7 and 9 are 97% and 93%, respectively. 4% of samples in category 9 are misclassified as category 8. Overall, the model performs well in recognizing most categories, but confusion exists between some similar categories (such as category 3 and category 5).

[0248] Table 1 compares the accuracy of the method of this invention with that of traditional deep learning methods.

[0249] Table 1. Accuracy comparison between the method of this invention and traditional deep learning methods.

[0250]

[0251] Table 2 compares the accuracy of the method of this invention with that of traditional millimeter-wave radar-based speech methods.

[0252]

Claims

1. A voice command recognition method based on millimeter wave sensing, characterized by, Specifically, the following steps are implemented: Step 1, build a millimeter wave radar data acquisition platform, collect echo data of different voice commands; Step 2, pre-process the echo data collected in step 1; Step 2 is implemented according to the following steps: Step 2.1, Collecting echo data for different speech command words by step 1, the data is in a three-dimensional matrix form representation, the dimensions are: wherein, is the index of the sampling point in the fast time dimension, is the Chirp index in the slow time dimension, is the index of the receiving channel, The data type of the matrix is complex, i.e. each data point can be represented as , which contains the amplitude and phase information of the signal, represents the number of sampling points in the fast time dimension, represents the number of linear frequency modulation Chirps in the slow time dimension, a Chirp is a single frequency modulated signal sent by the radar to obtain the velocity information of the target, represents the number of receiving channels; Step 2.2, on the echo data obtained in step 2.1, perform fast Fourier transform on the fast time dimension to extract the distance information of the target, and transform the original ADC data from the time domain to the distance domain, the specific calculation formula is: To distance unit index, with fast time dimension's sampling point One-to-one correspondence, finally generates distance-time matrix ; Step 2.

3. Calculate the distance-time matrix for step 2.2 The static clutter is filtered out using the phasor mean algorithm, and all Chirp The phasor mean is taken, the static clutter component is calculated, and is subtracted from the original matrix: The distance-time matrix after filtering out static clutter is obtained, and clearer dynamic target signals are obtained. Step 2.

4. To the solution of Step 2.3 Remove the DC component, the formula for the DC component is: Subtracting the calculated DC component from the signal in order to eliminate background interference: DC = 0.5 * (max + min) a distance-time matrix with removed direct current component ; Step 3, perform short-time Fourier transform on the pre-processed echo data in each distance unit to extract the Doppler information of the target, and then perform non-coherent superposition on the Doppler spectrum of all distance units to generate a micro-Doppler spectrum map; The step 3 is implemented according to the following steps: Step 3.1, Perform a short-time Fourier transform on the range-time matrix obtained in step 2.4 to extract Doppler information of the target: For each range cell Step 3.2, Perform a long-time Fourier transform on the Doppler information obtained in step 3.1 to extract the radial velocity of the target: representative distance unit , frequency , receive channel on micro-doppler spectrum; Step 3.2, according to the Doppler spectrum of each distance unit obtained in step 3.1, perform non-coherent superposition on the Doppler spectrum of all distance units to generate a micro-Doppler spectrum map: representative frequency and received channel micro-doppler spectrogram, is the amplitude square of the micro-doppler spectrum, used for non-coherent stacking to generate the final micro-doppler spectrogram; Step 4, use the improved deep convolutional generative adversarial network to perform data augmentation, and construct training set, validation set and test set; Step 5, build a vocal cord vibration recognition network BAC-Net, train different voice command words to get the optimal model, and then recognize different voice commands. 2.The millimeter-wave-sensor-based voice command recognition method of claim 1, wherein, The step 1 is implemented according to the following steps: Step 1.1, build a millimeter wave radar data acquisition platform, including a mainboard, the mainboard is equipped with two transmitting antennas and four receiving antennas, the mainboard is installed with mmWave Studio 02.01.01.00, and the millimeter wave radar related parameters are set as follows: Step 1.2, design and arrange the experimental scene, the radar is placed on a tripod with a radial distance of 1-1.2m from the ground and a radial distance of 0.5-0.8m, and placed in front of the measured target 10-30cm apart; Step 1.3, collect echo signals of different voice command words in indoor scene by recording voice commands one by one, each subject needs to say 10 required command words in turn, and in the recording process of each command word, the subject repeats the same command word 5 times according to the gesture guidance of the collector, and the data is automatically collected in fixed 5 second intervals, there are an average of 50 groups of words in each category, and a multi-channel fusion method is used to integrate the signals of the four receiving antennas in the millimeter wave radar, so there are 200 groups of words in each category. 3.The millimeter-wave-sensor-based voice command recognition method of claim 2, wherein, The step 4 is implemented according to the following steps: Step 4.1, build an improved deep convolutional generative adversarial network DCGAN; Step 4.2, input all the micro-Doppler spectrum maps on the receiving antennas generated in step 3.2 into the improved deep convolutional generative adversarial network DCGAN of step 4.1 to obtain the augmented micro-Doppler map; Step 4.3, merge the enhanced data obtained in step 4.2 with the original data, and divide the data set according to the training set:test set:validation set=7:1.5:1.

5. 4.The millimeter-wave-sensor-based voice command recognition method of claim 3, wherein, The specific structure of the improved deep convolutional generative adversarial network in step 4.1 is as follows: The improved deep convolutional generative adversarial network DCGAN is used to augment the data, and the same number of enhanced data as the original number of each type of micro-Doppler graph is generated. The loss function of the generator is: denotes the expected value or mean, denotes the output probability of the discriminator for real samples , denotes the output probability of the discriminator for generated samples , denotes a sample drawn from the real data distribution, denotes a noise vector drawn from the noise distribution, first term denotes that the discriminator wants the output for real samples to be as close to 1 as possible, second term denotes that the discriminator wants the output for generated samples to be as close to 0 as possible. 5.The millimeter wave sensor-based voice command recognition method of claim 4, wherein, The specific model structure of the improved deep convolutional generative adversarial network in step 4.1 is: The generator first receives 100-dimensional random Gaussian distribution noise, and is fully connected to 8x8x512 dimensions. Then, 4 alternating upsampling and convolution operations are performed, where the upsampling uses a transpose convolution kernel size of 4x4 and a step size of 2, and the number of convolution kernels is 512, 256, 128, and 64, respectively. Finally, a convolution kernel size of 3x3 and a step size of 1 are used for convolution to output the generated micro-Doppler image. The discriminator receives the generated sample and the original sample data, and performs 4 convolution operations with 64, 128, 256, and 512 convolution kernel numbers, respectively. Each time, a convolution kernel size of 4x4, a step size of 2, and a padding of 1 are used to extract image features. Then, a convolution kernel size of 4x4 and a step size of 1 are used to compress the feature map to 1x1x1 output, and a Sigmoid activation function is used to output the true or false judgment result of the image. Through the continuous iteration and optimization between the generator and the discriminator, they reach Nash equilibrium, i.e. the generated image meets the requirements. 6.The millimeter-wave-sensor-based voice command recognition method of claim 5, wherein, The step 5 is implemented according to the following steps: Step 5.1, the vocal cord vibration recognition network BAC-Net is based on a two-dimensional convolutional neural network 2D-CNN, and uses a Bayesian optimization algorithm to dynamically adjust the key hyperparameters of the model. A convolution block attention module CBAM is added after each convolution operation, and the optimization algorithm is used to dynamically determine whether the CBAM module is introduced after each convolution. The details of the channel and spatial weight calculation of the CBAM module are optimized, wherein the key hyperparameters include convolution kernel size, number of channels in each layer, optimizer selection, Dropout rate, and learning rate. A pruning mechanism is introduced, and in each round of test, the hyperparameter combination with a classification accuracy lower than 10% is directly eliminated to reduce the interference of invalid combinations on the overall optimization. Step 5.2, the training set in the data set constructed in step 4 is sent to the BAC-Net constructed in step 5.1 for training. The network obtains the optimal model of the validation set through training of different speech command words, and finally uses the test set for detection to realize recognition and classification of different speech commands.

Citation Information

Patent Citations

  • Lip language recognition method based on broadband multi-channel millimeter wave radar

    CN111856422A

  • Multimodal speech recognition system and method based on millimeter wave radar

    CN116416996A