A multi-modal speech recognition system and method based on millimeter-wave radar

Through a multimodal speech recognition system based on millimeter wave radar, the lip movement and vocal cord vibration characteristics are extracted using the FM signal and the feature fusion is performed, which solves the problem of low speech recognition accuracy in multi-sound sources and high-noise scenarios, and achieves efficient speech recognition effect.

CN116416996BActive Publication Date: 2025-08-05NANJING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310469259.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-27
Publication Date
2025-08-05
Estimated Expiration
2043-04-27

AI Technical Summary

Technical Problem

The existing speech recognition technology has low accuracy, poor real-time performance, low robustness in multi-sound sources and high-ambient noise scenarios, and the multi-sensor-based method has privacy problems or increases hardware costs.

Method used

A multimodal speech recognition system based on millimeter wave radar is adopted to extract lip motion and vocal cord vibration characteristics by transmitting frequency modulated continuous wave signals, and feature fusion is used to achieve complementarity and enhancement of lip motion and vocal cord vibration characteristics.

Benefits of technology

It improves the accuracy and robustness of speech recognition, reduces resource consumption, and realizes effective speech recognition in multi-user and high-noise interference scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116416996B_ABST
    Figure CN116416996B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal speech recognition system and method based on a millimeter-wave radar. The system includes a feature extraction module and a multi-modal fusion and recognition module. The feature extraction module uses the millimeter-wave radar to transmit a frequency-modulated continuous wave signal and extracts lip movement features and vocal cord vibration features from the reflected signal. The multi-modal fusion and recognition module is used to fuse the lip movement features and vocal cord vibration features and perform speech recognition. By fusing the lip movement feature and vocal cord vibration feature technologies, the present invention achieves the complementary and enhanced effects of the two features, further improving the accuracy of speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of wireless intelligent perception of the Internet of Things, and particularly relates to a multi-modal speech recognition system and method based on a millimeter-wave radar. Background Art

[0002] Nowadays, with the practical application of speech recognition technology, voice assistants are widely used in various application scenarios, bringing more convenience to our lives. In particular, speech recognition in voice assistants is used to improve the human-computer interaction efficiency in fields such as intelligent driving, smart home, and smart healthcare. For example, for intelligent meeting minutes, a voice assistant can be deployed in the meeting room to convert human voices into text. In addition, for safe driving, a voice assistant can also be used in intelligent driving to recognize the driver's instructions without manual touch. Currently, the speech recognition technology in voice assistants mainly collects speech signals of human subjects based on microphones. Speech recognition based on microphones works well in the absence of other speech interference and environmental noise. However, in the presence of multiple sound sources and environmental noise, the performance of microphone-based speech recognition will drop sharply. For example, in the field of intelligent driving, when multiple passengers in the car speak simultaneously, there will be a situation where multiple voices are mixed together in the voice assistant, resulting in the driver being unable to effectively interact with the voice assistant. Therefore, essentially new methods are needed to collect speech-related signals from human subjects to ensure the performance of speech recognition.

[0003] There are mainly two methods for speech recognition technology to improve the recognition performance in the presence of multiple sound sources and environmental noise:

[0004] 1. Speech recognition based on a single sensor and single modality; using a single sensor and single modality (such as microphones, WIFI, RFID, and cameras, etc.) to collect speech-related signals. Since speech signals are wideband signals and are easily interfered by environmental noise, generally, speech recognition based on a single sensor and single modality cannot achieve good performance in the presence of multiple sound sources and environmental noise.

[0005] 2. Multi-modal fusion speech recognition based on multiple sensors; this method simultaneously uses different modality sensors to collect speech-related signals (such as audio, video, and wireless signals). However, existing methods based on multi-modal fusion either have privacy issues (fusion of cameras and microphones), limited sensing range (fusion of ultrasonic waves and microphones), or increase additional hardware costs, and it is difficult to achieve synchronization in time and space.

[0006] Therefore, based on the above considerations, it is necessary to propose an innovative speech recognition system that can be applied to multiple sound sources and high ambient noise scenarios. By utilizing the macro and micro features of the same target simultaneously from a single sensor, and considering the correlation and complementarity of the features in the two dimensions, a feature fusion network framework based on TransFuser is proposed, thereby improving the accuracy, robustness, and security of speech recognition in low signal-to-noise ratio conditions and avoiding infringement of user privacy. Summary of the Invention

[0007] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a multimodal speech recognition system and method based on millimeter-wave radar to solve the existing problems of low speech recognition accuracy, poor real-time performance, and low robustness in scenarios such as multi-person scenarios and high ambient noise.

[0008] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0009] A multimodal speech recognition system based on millimeter wave radar of the present invention comprises: a feature extraction module and a multimodal fusion and recognition module;

[0010] The feature extraction module uses a millimeter wave radar to transmit a frequency modulated continuous wave (FMCW) signal and extracts lip movement features and vocal cord vibration features from the reflected signal;

[0011] The multimodal fusion and recognition module is used to fuse lip movement features and vocal cord vibration features and perform speech recognition.

[0012] Furthermore, the millimeter-wave radar transmitting end in the feature extraction module transmits a frequency-modulated continuous wave signal, and the signal characteristics are: each group consists of M frames of Chirp signals, the period of each frame of Chirp signal is T, and the Chirp interval time is T interval Starting frequency f c , each group of transmitted signals contains M frames and lasts for T frame ; Receive all echo signals of the millimeter-wave radar, mix each chirp signal of the echo signal with the chirp signal of the transmitted signal to obtain the demodulated intermediate frequency signal:

[0013]

[0014] Where A represents the signal gain, B represents the Chirp signal bandwidth, d represents the distance between the target and the radar, λ represents the wavelength, and c represents the speed of light.

[0015] The sampling rate is f adc For the intermediate frequency signal S IF (t) Downsampling is performed, and the sampling point is N.

[0016] Further, the frequency-modulated continuous-wave transmission signal S tx (t) and the echo signal S rx (t) are expressed as follows:

[0017]

[0018]

[0019] where α represents the signal path loss, A represents the signal gain, j represents the imaginary unit, f c represents the starting frequency of the signal, B represents the bandwidth of the Chirp signal, T represents the period of the Chirp signal, t represents time, τ represents the delay time of the echo signal, and φ0 represents the initial phase value.

[0020] Further, the extraction of the lip movement features specifically includes:

[0021] Detect the location of the user who emits the voice, extract one frame of the Chirp signal from each group of signals, perform the N-point discrete Fourier transform (DFT) algorithm on the sampling points of each Chirp signal, and determine the location of the lips of the user who emits the voice by detecting the peak position of the discrete Fourier transform; extract the signal phase change related to lip movement at the location of the user who emits the voice as: Δφ(t) = 4πΔd(t) / λ, and splice the target peak phase signal Δφ(t) detected for each frame of the Chirp signal; filter out the high-frequency signals through low-pass filtering with the cut-off frequency f stop , downsample at the sampling rate of f lip , and perform differencing on the downsampled signal to obtain the signal S l+d (t) related to the lip movement of the user who emits the voice, which is expressed as:

[0022] S l+d (t) = S l (t) + S d (t)

[0023] where S l (t) represents the lip movement signal of the user who emits the voice, and S d (t) represents the dynamic interference signal; determine whether there is voice activity by performing the voice activity detection algorithm; and perform filtering through the dynamic interference removal algorithm; finally, obtain the feature L p related to lip movement, which is expressed as:

[0024]

[0025] where STFT represents the short-time Fourier transform.

[0026] Further, the discrimination by the voice activity detection algorithm is specifically:

[0027] (11) Lip movement pre - detection: Considering that the user's speech activity contains both lip movement and vocal cord vibration, a threshold - based energy detection algorithm is used to estimate the energy intensity of the lip movement signal within a window; specifically, the lip movement signal is segmented by a Δτ time window, and then the energy value E l (t, t + Δτ) within the window is calculated, and a threshold E th is set to determine whether lip movement occurs;

[0028] (12) Vocal cord vibration verification: The vocal cord vibration signal within the Δτ time when lip movement occurs is further segmented, slid and segmented by a time window of dt / 2, and the signal energy value

[0029] (13) Decision discrimination: The lip movement energy feature E l (Δτ) and the vocal cord vibration energy feature E s (Δτ) within the Δτ time are combined into a new feature vector: E(Δτ) = CONCAT(E l (Δτ), E s (Δτ)), where CONCAT represents feature vector concatenation, and finally SVM is used to classify and discriminate the concatenated feature vector.

[0030] Furthermore, the specific filtering of the dynamic interference removal algorithm is as follows:

[0031] (21) The lip movement signal S l+d (R) containing body interference at the known Range Bin R is used to estimate the interference signal generated by body movement which is expressed as:

[0032]

[0033] where α i represents the weight coefficient of the i - th Range Bin, and

[0034] (22) The differential algorithm is used to remove the dynamic interference signal: The short - time Fourier transform (STFT) is respectively performed on the lip movement signal S l+d containing interference and the interference signal generated by body movement estimated in step (21). After removing the interference brought by body movement in the frequency domain, an interference - free estimated value of the lip movement signal is obtained, that is An estimated value of the lip movement signal without low - frequency dynamic interference is obtained.

[0035] Furthermore, the specific extraction of the vocal cord vibration feature includes:

[0036] Locate the position of the voice user by the change of the signal phase Δφ caused by vocal cord vibration. Extract one frame of Chirp signal from each group of signals, perform the N-point discrete Fourier transform (DFT) algorithm on the sampling points of each Chirp signal, and determine the vocal cord vibration position R of the voice user by detecting the peak position of the discrete Fourier transform s ; Then combine the signals at the R s positions of all frames, extract the vocal cord vibration signal by phase difference, and perform high-pass filtering to remove low-frequency interference signals and noise to obtain the interference-free vocal cord vibration signal S vib ; Determine whether there is voice activity by performing a voice activity detection algorithm; Finally, perform high-frequency signal estimation on the formant signal of the voice extracted from the vocal cords by the vocal cord vibration voice enhancement method to obtain the vocal cord enhanced voice signal L s .

[0037] Furthermore, the vocal cord vibration voice enhancement method is used to obtain the enhanced voice signal L s ; The specific process is as follows:

[0038] (31) Generate the enhanced voice signal spectrum: Perform short-time Fourier transform (STFT) on the input vocal cord vibration signal S vib , and then output the vocal cord vibration voice spectrogram containing high-frequency signals through the designed generator network GenNet(·): L vib = GenNet(STFT(S vib ));

[0039] (32) The inverse generator network generates the enhanced voice signal spectrum: Use the inverse network structure InvGenNet(·) of the generator to perform inverse generation and inverse Fourier transform (ISTFT) on the vocal cord vibration voice spectrogram L output in step (31), that is vib By satisfying the consistency constraint: To ensure the accuracy of the output result of the generator network; To ensure the accuracy of the output result of the generator network;

[0040] (33) Discriminate the output result: Input the generated vocal cord vibration voice spectrogram L vib (false sample) and the real voice signal spectrogram (true sample) collected by the microphone into the discriminator network DesNet(·) for recognition until the discriminator cannot distinguish between true and false samples, which indicates that the generated vibration voice spectrogram L vib contains high-frequency signal components, that is: L s = opt(L vib ).

[0041] Furthermore, the multimodal fusion and recognition module extracts the lip movement signal Lp Fuse the voice - band enhanced speech signal L s and use the lip motion feature encoder to encode L p as Encoder(L p ), and use the vocal cord vibration feature encoder to encode L s as Encoder(L s ); then fuse the encoded lip motion feature Encoder(L p ) and the vocal cord vibration feature Encoder(L s ) using the TransFuser structure to obtain the fused feature F fusion , which is expressed as:

[0042] F fusion = TransFuser[Encoder(L p ), Encoder(L s )]

[0043] Among them, the feature encoder Encoder(·) consists of a position encoder, regularization, and a multi - input attention sub - module att(·), and is specifically expressed as:

[0044]

[0045] Among them, Q, K, and V respectively represent the search matrix, the key - value matrix, and the eigenvalue matrix, represents the scale factor; finally, use the speech recognition method to classify and recognize the fused feature.

[0046] Furthermore, the multi - modal fusion and recognition module uses a multi - modal feature fusion algorithm based on the TransFuser structure for feature fusion, specifically as follows:

[0047] (41) Obtain the denoised lip motion feature and the vocal cord vibration feature after vocal cord signal enhancement;

[0048] (42) Lip motion feature encoding: Given the lip motion spectrum feature Segmented into 2 - D plane slices according to the time dimension The output feature of the encoder based on the attention mechanism is: Q L , K L , V L , and output the lip motion encoded feature through regularization and forward propagation;

[0049] (43) Vocal cord vibration feature encoding: For the input vocal cord vibration signal x=(x1,..., x i , x N ), the pre - encoded feature is y=(y1,..., yi , y N ), where Finally, the output feature of the encoder based on the attention mechanism is: Q S , K S , V S , and the vocal cord vibration coding feature is output through regularization and forward propagation;

[0050] (44) Cross-attention mechanism: For the features Q L , K L , V L output by the lip movement feature encoder and the features Q S , K S , V S output by the vocal cord vibration feature encoder, the cross-attention mechanism is used to exchange the lip movement features and the vocal cord vibration features, that is, the cross lip movement feature is: A L = att((K L , V L ), Q S ), and the vocal cord vibration feature is: A S = att((K S , V S ), Q L );

[0051] (45) Fusion attention mechanism: The output result of the cross-attention mechanism is used as the input of the fusion attention mechanism, that is, the output feature representation of the fusion attention mechanism: A F = att((K F , V F ), Q S ) + att((K F , V F ), Q L ), where K F represents the vector splicing of the feature K L output by the lip movement encoder and the feature K S output by the lip movement encoder in step (44), that is: K F = Concat(K L , K S ); Similarly, V F = Concat(V L , V S ).

[0052] Furthermore, the speech recognition method specifically includes: a decoder and a linear mapping; the multi-modal fusion features are recognized through a speech recognition network, and the fused feature F fusion is received, and the decoder Decoder(·) is used for Ffusion Perform parsing to obtain the decoded phonetic symbol-related feature F symbol , that is: F symbol = Decoder(F fusion ); Output the phonetic symbol recognition result through linear mapping Softmax.

[0053] A multi-modal speech recognition method based on a millimeter-wave radar according to the present invention, based on the above system, the steps are as follows:

[0054] 1) Use the millimeter-wave radar to continuously transmit FMCW signals and receive echo signals;

[0055] 2) Use the received echo signals to determine the location of the user who emits the voice, respectively extract the corresponding vocal cord vibration feature signals and lip movement feature signals and perform preprocessing;

[0056] 3) Use the voice activity detection algorithm to filter out non-effective voice activity signals;

[0057] 4) Remove noise interference from the extracted lip movement feature signals, and perform vibration speech enhancement algorithm processing on the extracted vocal cord vibration feature signals;

[0058] 5) Fuse the lip movement feature signals and vocal cord vibration feature signals;

[0059] 6) Perform speech recognition on the fused features.

[0060] Advantages of the present invention:

[0061] 1. By adopting the dynamic interference removal algorithm, lip movement feature signals with higher signal-to-noise ratio are extracted, thereby improving the accuracy and robustness of speech recognition;

[0062] 2. By adopting the voice activity detection algorithm at the front end, the effect of filtering out non-voice activities is achieved, reducing unnecessary resource consumption such as backend speech enhancement, feature fusion and speech recognition, thereby improving resource utilization rate and system processing speed;

[0063] 3. By fusing the lip movement feature and vocal cord vibration feature technologies, the complementary and enhanced effects of the two features are achieved, further improving the accuracy of speech recognition.

[0064] 4. The present invention simultaneously senses the features of two modalities of the same target through a single sensor and fuses them, realizing effective speech recognition of users in scenarios such as multi-users and high noise interference. Brief Description of the Drawings

[0065] Figure 1 is the schematic diagram of the system of the present invention.

[0066] Figure 2 It is a flowchart of voice activity detection.

[0067] Figure 3 It is a schematic diagram of the enhancement of voice with vocal cord vibration.

[0068] Figure 4 It is a structural diagram of a lip movement feature encoder.

[0069] Figure 5 It is a structural diagram of a vocal cord vibration feature encoder.

[0070] Figure 6 It is a schematic diagram of the structure principle of a multi-modal fusion and recognition module. Specific implementation mode

[0071] For the convenience of understanding by those skilled in the art, the present invention will be further described below in conjunction with embodiments and the accompanying drawings. The content mentioned in the implementation mode does not limit the present invention.

[0072] Refer to Figures 1-6 As shown, a multi-modal voice recognition system based on a millimeter-wave radar according to the present invention includes: a feature extraction module and a multi-modal fusion and recognition module;

[0073] The feature extraction module uses a millimeter-wave radar to transmit a frequency-modulated continuous wave (FMCW) signal and extracts lip movement features and vocal cord vibration features from the reflected signal;

[0074] Among them, the millimeter-wave radar transmitter in the feature extraction module transmits a frequency-modulated continuous wave signal, and the signal features are: each group consists of M frames of Chirp signals, the period of each frame of Chirp signal is T, and the Chirp interval time T interval Starting frequency f c , each group of transmitted signals contains M frames, and the duration is T frame ; receive all the echo signals of the millimeter-wave radar, and mix each Chirp signal of the echo signal with the Chirp signal of the transmitted signal to obtain a demodulated intermediate-frequency signal:

[0075]

[0076] In the formula, A represents the signal gain, B represents the Chirp signal bandwidth, d represents the distance between the target and the radar, λ represents the wavelength, and c represents the speed of light;

[0077] Downsample the intermediate-frequency signal S adc at a sampling rate of f IF (t), and the sampling points are N.

[0078] Specifically, the frequency-modulated continuous wave transmitted signal S tx(t) and the echo signal S rx The functional expression of (t) is:

[0079]

[0080]

[0081] where α represents the signal path loss, A represents the signal gain, j represents the imaginary unit, f c represents the starting frequency of the signal, B represents the bandwidth of the Chirp signal, T represents the period of the Chirp signal, t represents time, τ represents the delay time of the echo signal, and φ0 represents the initial phase value.

[0082] Specifically, the extraction of lip movement features specifically includes:

[0083] Detect the location of the user who emits the voice, extract one frame of Chirp signal from each group of signals, perform the N-point discrete Fourier transform (DFT) algorithm on the sampling points of each Chirp signal, and determine the location of the lips of the user who emits the voice by detecting the peak position of the discrete Fourier transform; extract the signal phase change related to lip movement at the location of the user who emits the voice as: Δφ(t) = 4πΔd(t) / λ, and splice the target peak phase signal Δφ(t) detected for each frame of Chirp signal; filter out high-frequency signals through low-pass filtering with the cut-off frequency f stop downsample at the sampling rate of f lip differentiate the downsampled signal to obtain the signal S l+d (t) related to the lip movement of the user who emits the voice, expressed as:

[0084] S l+d (t) = S l (t) + S d (t)

[0085] where S l (t) represents the lip movement signal of the user who emits the voice, and S d (t) represents the dynamic interference signal; determine whether there is voice activity by performing a voice activity detection algorithm; and filter through a dynamic interference removal algorithm; finally obtain the lip movement related feature L p expressed as:

[0086]

[0087] where STFT represents the short-time Fourier transform.

[0088] Specifically, the discrimination by the voice activity detection algorithm is specifically:

[0089] (11) Lip movement pre-detection: Considering that the user's speech activity includes simultaneous lip movement and vocal cord vibration, a threshold-based energy detection algorithm is used to estimate the energy intensity of the lip movement signal within a window; specifically, the lip movement signal is segmented by a Δτ time window, and then the energy value E l (t, t + Δτ) within the window is calculated, and a threshold E th is set to determine whether lip movement occurs;

[0090] (12) Vocal cord vibration verification: The vocal cord vibration signal within the Δτ time when lip movement occurs is further segmented, and it is segmented by a time window of dt / 2 and slides, and the signal energy value is calculated

[0091] (13) Decision discrimination: The lip movement energy feature E l (Δτ) and the vocal cord vibration energy feature E s (Δτ) within the Δτ time are combined into a new feature vector: E(Δτ) = CONCAT(E l (Δτ), E s (Δτ)), where CONCAT represents feature vector concatenation, and finally SVM is used to classify and discriminate the concatenated feature vector.

[0092] Specifically, the filtering of the dynamic interference removal algorithm is as follows:

[0093] (21) The lip movement signal S l+d (R) containing body interference at a known Range Bin of R is used to estimate the interference signal generated by body movement which is expressed as:

[0094]

[0095] where α i [[ID=​37]] represents the weight coefficient of the i-th Range Bin, and

[0096] (22) The differential algorithm is used to remove the dynamic interference signal: The short-time Fourier transform (STFT) is respectively performed on the lip movement signal S l+d containing interference and the interference signal generated by body movement estimated in step (21), and the interference brought by body movement is removed in the frequency domain to obtain an estimated value of the lip movement signal without interference that is An estimated value of the lip movement signal without low-frequency dynamic interference is obtained.

[0097] Specifically, the extraction of vocal cord vibration features specifically includes:

[0098] The position of the voice user is located by the change in the signal phase Δφ caused by vocal cord vibration. One frame of Chirp signal is extracted from each group of signals, and the N-point discrete Fourier transform (DFT) algorithm is executed on the sampling points of each Chirp signal. The position R of the vocal cord vibration of the voice user is determined by detecting the peak position of the discrete Fourier transform. s ; Then the signals at the R s positions of all frames are combined, the phase difference is used to extract the vocal cord vibration signal, and high-pass filtering is performed to remove low-frequency interference signals and noise, obtaining an interference-free vocal cord vibration signal S vib ; The presence of speech activity is discriminated by executing a speech activity detection algorithm; Finally, the high-frequency signal of the formant signal of the voice extracted from the vocal cords is estimated by a vocal cord vibration speech enhancement method, and the vocal cord enhanced speech signal L s .

[0099] Specifically, the vocal cord vibration speech enhancement method is used to obtain the enhanced speech signal L s ; The specific process is as follows:

[0100] (31) Generate the enhanced speech signal spectrum: Perform short-time Fourier transform (STFT) on the input vocal cord vibration signal S vib , and then output the vocal cord vibration speech spectrogram containing high-frequency signals through the designed generator network GenNet(·): L vib = GenNet(STFT(S vib ));

[0101] (32) The inverse generator network generates the enhanced speech signal spectrum: Use the inverse network structure InvGenNet(·) of the generator to perform inverse generation and inverse Fourier transform (ISTFT) on the vocal cord vibration speech spectrogram L output in step (31), that is vib , and ensure the accuracy of the output result of the generator network by satisfying the consistency constraint: ; to ensure the accuracy of the output result of the generator network;

[0102] (33) Discriminate the output result: Input the generated vocal cord vibration speech spectrogram L vib (false sample) and the real speech signal spectrogram (true sample) collected by the microphone into the discriminator network DesNet(·) for recognition until the discriminator cannot distinguish between true and false samples, indicating that the generated vibration speech spectrogram L vib contains high-frequency signal components, that is: L s = opt(L vib ).

[0103] The multimodal fusion and recognition module is used to fuse lip movement features and vocal cord vibration features and perform speech recognition;

[0104] Among them, the multi-modal fusion and recognition module fuses the extracted lip movement signal L p and the vocal cord enhanced speech signal L s and encodes L p into Encoder(L p ) using the lip movement feature encoder, and encodes L s into Encoder(L s ) using the vocal cord vibration feature encoder; then the encoded lip movement feature Encoder(L p ) and the vocal cord vibration feature Encoder(L s ) are fused using the TransFuser structure to obtain the fused feature F fusion , which is expressed as:

[0105] F fusion = TransFuser[Encoder(L p ), Encoder(L s )]

[0106] Among them, the feature encoder Encoder(·) consists of a position encoder, regularization, and a multi-input attention sub-module att(·), and is specifically expressed as:

[0107]

[0108] Among them, Q, K, and V respectively represent the search matrix, the key value matrix, and the eigenvalue matrix, represents the scale factor; finally, the fused feature is classified and recognized using a speech recognition method.

[0109] Specifically, the multi-modal fusion and recognition module uses a multi-modal feature fusion algorithm based on the TransFuser structure for feature fusion, as follows:

[0110] (41) Obtain the denoised lip movement features and the vocal cord vibration features enhanced by the vocal cord signal;

[0111] (42) Lip movement feature encoding: Given the lip movement spectrum features Split into 2D plane slices according to the time dimension The output features of the encoder based on the attention mechanism are: Q L , K L , V L , and the lip movement encoded features are output through regularization and forward propagation;

[0112] (43)Vocal cord vibration feature encoding: For the input vocal cord vibration signal x = (x1,..., x i , x N ), the pre-encoded feature is y = (y1,..., y i , y N ), where, Finally, the output feature of the encoder based on the attention mechanism is: Q S , K S , V S , and the vocal cord vibration encoding feature is output through regularization and forward propagation;

[0113] (44)Cross-attention mechanism: For the features Q L , K L , V L output by the lip movement feature encoder and the features Q S , K S , V S output by the vocal cord vibration feature encoder, the cross-attention mechanism is used to exchange the lip movement features and the vocal cord vibration features, that is, the cross lip movement feature is: A L = att((K L , V L ), Q S ), and the vocal cord vibration feature is: A S = att((K S , V S ), Q L );

[0114] (45)Fusion attention mechanism: The output result of the cross-attention mechanism is used as the input of the fusion attention mechanism, that is, the output feature representation of the fusion attention mechanism: A F = att((K F , V F ), Q S ) + att((K F , V F ), Q L ), where, K F represents the vector concatenation of the feature K L output by the lip movement encoder and the feature K S output by the lip movement encoder in step (44), that is: K F = Concat(K L , K S ); Similarly, V F = Concat(V L , V S ).

[0115] Specifically, the speech recognition method specifically includes: a decoder and a linear mapping; identifying the features after multi-modal fusion through a speech recognition network, and receiving the fused feature F fusion , and using the decoder Decoder(·) to parse F fusion to obtain the phonetic symbol-related features F symbol after decoding, that is: F symbol = Decoder(F fusion ); outputting the phonetic symbol recognition result through the linear mapping Sofimax.

[0116] In addition, the present invention also provides a multi-modal speech recognition method based on a millimeter-wave radar. Based on the above system, the steps are as follows:

[0117] 1) Continuously transmit FMCW signals using the millimeter-wave radar and receive the echo signals;

[0118] 2) Determine the location of the user who emits the speech using the received echo signals, respectively extract the corresponding vocal cord vibration feature signals and lip movement feature signals and perform preprocessing;

[0119] 3) Filter out non-effective speech activity signals using a speech activity detection algorithm;

[0120] 4) Remove noise interference from the extracted lip movement feature signals and perform a vibration speech enhancement algorithm on the extracted vocal cord vibration feature signals;

[0121] 5) Fuse the lip movement feature signals and the vocal cord vibration feature signals;

[0122] 6) Perform speech recognition on the fused features.

[0123] The specific application scenarios of the present invention are numerous. The above description is only the preferred implementation mode of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements can be made, and these improvements should also be regarded as the protection scope of the present invention.

Claims

1. A multimodal speech recognition system based on millimeter wave radar, characterized in that: include: Feature extraction module and multimodal fusion and recognition module; The feature extraction module uses the millimeter wave radar to transmit a frequency modulated continuous wave signal and extracts the lip movement features and vocal cord vibration features from the reflected signal; The multimodal fusion and recognition module is used to fuse lip movement features and vocal cord vibration features and perform speech recognition; The millimeter wave radar transmitting end in the feature extraction module transmits a frequency modulated continuous wave signal. The signal characteristics are: each group consists of M frames of Chirp signals, the period of each frame of Chirp signal is T, and the Chirp interval time is T interval Starting frequency f c , each group of transmitted signals contains M frames and lasts for T frame ; Receive all echo signals of the millimeter-wave radar, mix each chirp signal of the echo signal with the chirp signal of the transmitted signal to obtain the demodulated intermediate frequency signal: Where A represents the signal gain, B represents the Chirp signal bandwidth, d represents the distance between the target and the radar, λ represents the wavelength, and c represents the speed of light. The sampling rate is f adc For the intermediate frequency signal S IF (t) Downsampling is performed, and the sampling point is N; The extraction of lip movement features specifically includes: The position of the user making the speech is detected, a frame of Chirp signal is extracted from each group of signals, an N-point discrete Fourier transform algorithm is performed on the sampling point of each Chirp signal, and the position of the lips of the user making the speech is determined by detecting the peak position of the discrete Fourier transform; the phase change of the signal related to the lip movement at the position of the user making the speech is extracted as: Δφ(t) = 4πΔd(t) / λ, and the target peak phase signal Δφ(t) detected by each frame of Chirp signal is spliced; the cutoff frequency f is used to determine the position of the lips of the user making the speech. stop The low-pass filter removes high-frequency signals and lip The sampling rate is downsampled, and the downsampled signal is differentially obtained to obtain the signal S related to the lip movement of the user who sends the voice l+d (t), expressed as: S l+d (t)=S l (t)+S d (t) Among them, S l (t) represents the lip movement signal of the user who issues the speech, S d (t) Express the dynamic interference signal; determine whether there is voice activity by executing the voice activity detection algorithm; filter it by the dynamic interference removal algorithm; and finally obtain the lip movement related feature L p Expressed as: Among them, STFT stands for short-time Fourier transform; The extracting of vocal cord vibration characteristics specifically includes: The position of the user making the speech is located by the change in the signal phase Δφ caused by the vocal cord vibration. A frame of Chirp signal is extracted from each group of signals. The N-point discrete Fourier transform algorithm is performed on the sampling point of each Chirp signal. The position of the vocal cord vibration R of the user making the speech is determined by detecting the peak position of the discrete Fourier transform. s ; Then R of all frames s The signal combination of the position, phase difference extraction of vocal cord vibration signal, high-pass filtering to remove low-frequency interference signals and noise, and obtain the interference-free vocal cord vibration signal S vib ; By executing the voice activity detection algorithm to determine whether there is voice activity; finally, the high-frequency signal of the voice formant signal extracted by the vocal cords is estimated by the vocal cord vibration voice enhancement method to obtain the vocal cord enhanced voice signal L s .

2. The multimodal speech recognition system based on millimeter wave radar according to claim 1, characterized in that: The frequency modulated continuous wave transmission signal S tx (t) and the echo signal S rx The function expression of (t) is: Where α represents the signal path loss, j represents the imaginary unit, and f c represents the starting frequency of the signal, T represents the Chirp signal period, t represents time, τ represents the delay time of the echo signal, and φ0 represents the initial phase value.

3. The multimodal speech recognition system based on millimeter wave radar according to claim 1, characterized in that: The voice activity detection algorithm performs the following judgment: (11) Lip movement pre-detection: Considering that the user's speech activity contains both lip movement and vocal cord vibration, a threshold-based energy detection algorithm is used to estimate the energy intensity of the lip movement signal within the window. Specifically, the lip movement signal is segmented into a Δτ time window, and then the energy value E in the window is calculated. l (t,t+Δτ), set the threshold E th Determine whether lip movement occurs; (12) Vocal cord vibration verification: The vocal cord vibration signal within the Δτ time period that produces lip movement is further segmented, and the signal energy value is calculated by sliding the segmentation with a time window of dt / 2. (13) Decision judgment: The lip motion energy feature E within Δτ time l (Δτ) and vocal cord vibration energy characteristics E s (Δτ) is combined into a new eigenvector: E(Δτ) = CONCAT(E l (Δτ),E s (Δτ)), where CONCAT represents feature vector concatenation, and finally SVM is used to classify and discriminate the concatenated feature vectors.

4. The multimodal speech recognition system based on millimeter wave radar according to claim 1, characterized in that: The dynamic interference removal algorithm performs filtering specifically as follows: (21) It is known that the lip movement signal containing body interference at Range Bin R is S l+d (R), estimate the interference signal generated by body motion Expressed as: Among them, α i represents the weight coefficient of the i-th Range Bin, and (22) Differential algorithm to remove dynamic interference signal: respectively, the lip motion signal S containing interference l+d and the interference signal generated by the body motion estimated in step (21) Perform a short-time Fourier transform to remove the interference caused by body motion in the frequency domain to obtain an interference-free lip motion signal estimate Right now Obtain an estimate of the lip motion signal without low-frequency dynamic interference.

5. The multimodal speech recognition system based on millimeter wave radar according to claim 1, characterized in that: The vocal cord vibration speech enhancement method is used to obtain an enhanced speech signal L s The specific process is as follows: (31) Generate enhanced speech signal spectrum: Input vocal cord vibration signal S vib Perform short-time Fourier transform, and then output the vocal cord vibration speech spectrum containing high-frequency signals through the designed generator network GenNet(·): L vib =GenNet(STFT(S vib )); (32) The inverse generator network generates the enhanced speech signal spectrum: the inverse network structure InvGenNet(·) of the generator is used to transform the vocal cord vibration speech spectrum L output in step (31) into vib Perform inverse generation and inverse Fourier transform, i.e. By satisfying the consistency constraints: To ensure the accuracy of the generator network output results; (33) Determine the output result: generate the vocal cord vibration speech spectrum L vib The real speech signal spectrum graph collected by the microphone is input into the discriminator network DesNet(·) for recognition, until the discriminator cannot distinguish the true and false samples, which means that the vibration speech spectrum graph L generated by the generator is vib Contains high-frequency signal components, namely: L s =opt(L vib ).

6. The multimodal speech recognition system based on millimeter wave radar according to claim 1, characterized in that: The multimodal fusion and recognition module extracts the lip motion signal L p Harmony vocal cord enhanced speech signal L s Fusion, using lip motion feature encoder to L p Encoded as Encoder(L p ), using the vocal cord vibration feature encoder to convert L s Encoded as Encoder(L s ); Then the encoded lip motion feature Encoder(L p ) and vocal fold vibration characteristics Encoder(L s ) Use the TransFuser structure to perform feature fusion and obtain the fused feature F fusion , expressed as: F fusion =TransFuser[Encoder(L p ),Encoder(L s )] Among them, the feature encoder Encoder(·) consists of a position encoder, a regularization, and a multi-input attention submodule att(·), which is specifically expressed as: Among them, Q, K and V represent the search matrix, key value matrix and eigenvalue matrix respectively. represents the scale factor; finally, the speech recognition method is used to classify and recognize the fused features.

7. A multimodal speech recognition method based on millimeter wave radar, based on the system according to any one of claims 1 to 6, characterized in that: Here are the steps: 1) Use millimeter wave radar to continuously transmit FMCW signals and receive echo signals; 2) Using the received echo signal to determine the location of the user who made the speech, the corresponding vocal cord vibration feature signal and lip movement feature signal are extracted and preprocessed; 3) Using voice activity detection algorithm to filter out invalid voice activity signals; 4) performing noise interference removal on the extracted lip movement feature signal and performing vibration speech enhancement algorithm processing on the extracted vocal cord vibration feature signal; 5) Fusing the lip movement feature signal and the vocal cord vibration feature signal; 6) Perform speech recognition on the fused features.

Citation Information

Patent Citations

  • Anti-noise voiceprint recognition method based on millimeter wave sensing vibration signal

    CN117116269A