A millimeter wave voice signal processing method and system based on model and data hybrid drive
Through the combination of virtual vector compensation algorithm and multi-stage learning network, the problem of low signal-to-noise ratio and insignificant phase in millimeter wave speech signal processing is solved, and the recovery of high-quality speech signals is achieved, especially the clarity improvement in noise environments.
Patent Information
- Application Number
- CN202510703194.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The existing millimeter wave speech signal processing technology has problems such as low signal-to-noise ratio, insignificant initial phase and difficulty in dealing with amplitude and phase nonlinear noise, resulting in low speech clarity.
Using a hybrid model and data driver method, the observed phase is enhanced through a virtual vector compensation algorithm, and the multi-stage learning decomposition speech denoising task is used to optimize amplitude denoising and phase, and frequency compensation is performed in combination with attention mechanism.
It significantly improves the clarity and anti-interference ability of the voice signal, especially in high noise environments, and the output voice spectrum characteristics are close to the original voice.
Smart Images

Figure CN120236600B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of wireless communications and signal processing technology, and in particular to a millimeter wave voice signal processing method and system based on hybrid model and data driving. Background Art
[0002] Mobile phone voice communication is a widely used method of information transmission. Technologies for voice signal processing have important applications in areas such as smart homes and remote conferencing. Traditional methods for collecting voice signals include microphones and optical sensors. Microphones are susceptible to interference from ambient noise, while optical sensors are limited in darkness or obscured conditions. In contrast, millimeter-wave signals have the ability to sense micro-vibrations, enabling contactless signal collection by detecting weak vibrations on the device surface. They are also insensitive to ambient noise and lighting conditions.
[0003] However, existing millimeter-wave speech signal processing technology has the following problems: (1) The vibration signal generated by the mobile phone speaker is weak, resulting in a low signal-to-noise ratio of the millimeter-wave echo signal and unclear parsed speech; (2) The initial phase of the observation signal is affected by the randomness of the target position, resulting in unclear phase changes, further reducing speech clarity; (3) Existing denoising methods have difficulty in processing nonlinear noise in amplitude and phase at the same time, and cannot meet the needs of high-quality speech signal processing. Summary of the Invention
[0004] In view of this, the present invention provides a millimeter-wave speech signal processing method and system based on a hybrid drive of model and data, aiming to overcome the problem that the existing millimeter-wave radar-based speech signal processing method is difficult to achieve high-quality speech signal processing. First, the present invention enhances the observed phase of the echo signal through a virtual vector compensation algorithm, and amplifies the millimeter-wave phase change caused by speech vibration. Secondly, in response to the problem of low signal-to-noise ratio of millimeter-wave echo signals, the present invention decomposes the complex speech denoising task into two sub-tasks of amplitude denoising and phase optimization based on multi-stage learning based on the potential correlation between amplitude and phase, gradually realizes speech denoising, and uses a frequency conversion module based on the attention mechanism to further compensate for the frequency distortion part of the speech, thereby improving the quality of the speech.
[0005] To this end, the present invention provides the following technical solutions:
[0006] In one aspect, the present invention provides a millimeter wave speech signal processing method based on a hybrid model and data drive, comprising:
[0007] Collect and pre-process radar echo signals; and perform phase demodulation on the pre-processed echo signals;
[0008] Speech denoising is performed using the millimeter wave speech denoising algorithm, including:
[0009] The radar echo signal is processed using a model-driven virtual vector compensation algorithm; the model-driven virtual vector compensation algorithm amplifies the phase of the observation signal by adding a virtual vector to compensate the observation signal based on the constant value of the static path vector;
[0010] Perform short-time Fourier transform on the radar echo signal processed by the virtual vector compensation algorithm to convert the phase-enhanced millimeter-wave voice signal from the time domain to the time-frequency domain;
[0011] A data-driven multi-stage millimeter-wave speech denoising network is used to perform amplitude denoising and phase optimization on radar echo signals after short-time Fourier transform; the multi-stage millimeter-wave speech denoising network includes: an amplitude denoising network for denoising speech amplitude and a phase denoising network for optimizing speech phase; the amplitude denoising network includes: an amplitude encoder, a frequency conversion module, a first extended convolution module and an amplitude decoder; the amplitude encoder maps the input amplitude spectrum to a feature space, and the frequency conversion module compensates for the lost frequency components in the speech, the first extended convolution module extracts features through a multi-scale receptive field, and the amplitude decoder reconstructs the amplitude spectrum; the phase denoising network includes: a complex encoder, a second extended convolution module, a real decoder and an imaginary decoder; the complex encoder encodes the input complex spectrum, the second extended convolution module extracts sequence features, and the real decoder and the imaginary decoder reconstruct the real and imaginary parts of the complex spectrum respectively; the complex spectrum is composed of the denoised amplitude spectrum and the original phase spectrum;
[0012] The denoised speech signal is calculated through short-time inverse Fourier transform to restore the speech information.
[0013] Furthermore, the radar echo signal is processed using a model-driven virtual vector compensation algorithm, including:
[0014] Radar observation vector over a period of time Averaging is performed to initially obtain an original static vector An approximate estimate of
[0015] In the IQ plane, a triangle is constructed to describe the effect of the added virtual vector on the observed phase. 、 and are the original static vector, the virtual vector and the rotated static vector, which are the three sides of the triangle respectively. Rotation Angle after forming ,horn yes and The angle between
[0016] according to , the static vector after rotation Calculated as: ;in, Represents the original static vector The model, express Angle with the I axis;
[0017] According to the law of cosines, the virtual vector The amplitude of is calculated as: ;in, Represents a virtual vector The amplitude, Represents the original static vector The amplitude of Represents the static vector after rotation Amplitude;
[0018] according to ,horn Calculated as , virtual vector Phase angle Expressed as: ;
[0019] For a given phase shift , corresponding to the phase of the virtual vector , is the phase of the static vector;
[0020] The obtained virtual vector is used to compensate the observed signal to enhance the phase change of the observed signal: ;in, represents the observed signal after adding the virtual vector.
[0021] Furthermore, in the amplitude denoising network, the amplitude encoder includes: a first two-dimensional convolution layer, a first batch of normalization layers, a first ReLU nonlinear activation layer and a first maximum pooling layer; the convolution kernel size in the first two-dimensional convolution layer and the first maximum pooling layer is 2×2, and the step size is 1; the frequency conversion module includes: a second two-dimensional convolution layer, a first one-dimensional convolution layer and a third two-dimensional convolution layer; the convolution kernel size of the second two-dimensional convolution layer is 1×1, and the step size is 1; the convolution kernel size of the first one-dimensional convolution layer is 9, and the step size is 1; the third two-dimensional convolution layer The convolution kernel size of the layer is 2×2, and the stride is 1; the first extended convolution module includes the second one-dimensional convolution layer and the first two-dimensional extended convolution layer; the convolution kernel size of the second one-dimensional convolution layer is 1, and the stride is 1; the convolution kernel size of the first two-dimensional extended convolution layer is 1, and the stride is 1, and the extended layer is [1, 2, 4, 8, 16, 32]; the amplitude decoder includes: the first two-dimensional inverted convolution layer, the second batch normalization layer and the second ReLU nonlinear activation layer; the convolution kernel size of the first two-dimensional inverted convolution layer is 2×2, and the stride is 2.
[0022] Furthermore, in the phase denoising network, the complex encoder includes: a fourth two-dimensional convolutional layer, a third batch normalization layer, a third ReLU nonlinear activation layer and a second maximum pooling layer; the convolution kernel size in the fourth two-dimensional convolutional layer and the second maximum pooling layer is 2×2, and the step size is 1; the second extended convolution module includes a third one-dimensional convolutional layer and a second two-dimensional extended convolutional layer; the convolution kernel size of the third one-dimensional convolutional layer is 1, and the step size is 1; the convolution kernel size of the second two-dimensional extended convolution layer is 1, and the step size is 1, and the extended layer is [1, 2, 4, 8, 16, 32]; the real part decoder includes: a second two-dimensional inverted convolutional layer, a fourth batch normalization layer and a fourth ReLU nonlinear activation layer; the convolution kernel size of the second two-dimensional inverted convolution layer is 2×2, and the step size is 2; the imaginary part decoder includes: a third two-dimensional inverted convolutional layer, a fifth batch normalization layer and a fifth ReLU nonlinear activation layer; the convolution kernel size of the third two-dimensional inverted convolution layer is 2×2, and the step size is 2.
[0023] Furthermore, the data-driven multi-stage millimeter wave speech denoising network training process is as follows:
[0024] pass Loss training magnitude denoising network: ;
[0025] in, and Represent the mean square error and the amplitude spectrum of the training speech labels respectively.
[0026] pass The loss jointly trains the amplitude denoising network and the phase denoising network: ;
[0027] in, The loss function of the phase denoising network is as follows:
[0028] ;
[0029] in, and Represent the real and imaginary parts of the complex spectrum of the speech label respectively.
[0030] Furthermore, collecting and preprocessing radar echo signals include:
[0031] Use millimeter-wave radar to collect vibration information from the target mobile phone;
[0032] Beamforming is performed on the echo signals from multiple receiving antennas of the millimeter-wave radar.
[0033] Furthermore, phase demodulation is performed on the pre-processed echo signal, including:
[0034] The intermediate frequency signal is organized into a matrix with distance on the horizontal axis and sampling points on the vertical axis. Fast Fourier transform is further performed along the distance dimension to obtain the spectrum S.
[0035] The range bin of the vibration target is extracted from S and the phase of the vibration signal is extracted to obtain the demodulated noisy speech P.
[0036] In another aspect, the present invention further discloses a millimeter wave voice signal processing system based on a hybrid model and data drive, comprising:
[0037] Acquisition and preprocessing module, used for acquiring and preprocessing radar echo signals;
[0038] A phase demodulation module is used to perform phase demodulation on the echo signal preprocessed by the acquisition and preprocessing module;
[0039] The speech denoising module is used to perform speech denoising on the phase-demodulated echo signal obtained by the phase demodulation module using the millimeter-wave speech denoising algorithm, including:
[0040] The model-driven submodule processes the radar echo signal using a model-driven virtual vector compensation algorithm. The model-driven virtual vector compensation algorithm amplifies the phase of the observation signal by adding a virtual vector to compensate the observation signal based on the constant value of the static path vector.
[0041] The first conversion submodule performs a short-time Fourier transform on the radar echo signal processed by the virtual vector compensation algorithm, and converts the phase-enhanced millimeter-wave voice signal from the time domain to the time-frequency domain;
[0042] The data-driven submodule utilizes a data-driven multi-stage millimeter-wave speech denoising network to perform amplitude denoising and phase optimization on the radar echo signal after short-time Fourier transform; the multi-stage millimeter-wave speech denoising network includes: an amplitude denoising network for denoising the speech amplitude and a phase denoising network for optimizing the speech phase; the amplitude denoising network includes: an amplitude encoder, a frequency conversion module, a first extended convolution module and an amplitude decoder; the amplitude encoder maps the input amplitude spectrum to the feature space, and the frequency conversion module compensates for the lost frequency components in the speech, the first extended convolution module extracts features through a multi-scale receptive field, and the amplitude decoder reconstructs the amplitude spectrum; the phase denoising network includes: a complex encoder, a second extended convolution module, a real decoder and an imaginary decoder; the complex encoder encodes the input complex spectrum, the second extended convolution module extracts sequence features, and the real decoder and the imaginary decoder reconstruct the real and imaginary parts of the complex spectrum respectively; the complex spectrum is composed of the denoised amplitude spectrum and the original phase spectrum.
[0043] The signal recovery module is used to calculate the denoised speech signal through short-time inverse Fourier transform and restore the speech information.
[0044] This invention achieves significant technological advancement in the field of millimeter-wave voice signal processing through a hybrid model and data-driven technical solution:
[0045] In the present invention, through the application of the virtual vector compensation algorithm, the observed phase of the echo signal is effectively enhanced, the phase change of the millimeter wave signal caused by the voice vibration is amplified, and the sensitivity of phase detection is significantly improved compared with the traditional method. This technical means enables the accurate capture of voice vibration information in a low signal-to-noise ratio environment, laying a solid foundation for subsequent processing. At the same time, the multi-stage millimeter wave voice denoising network of the present invention achieves comprehensive optimization of the voice signal through an innovative two-stage processing architecture of amplitude denoising and phase denoising. Among them, the amplitude denoising stage effectively suppresses the interference of environmental noise, while the phase denoising stage is specifically optimized for the phase distortion problem unique to millimeter wave signals. In particular, through the design of the frequency conversion module, the frequency distortion of the voice signal is successfully compensated, so that the spectral characteristics of the output voice are closer to the original voice.
[0046] The present invention gives full play to the advantages of model-driven and data-driven technologies through the organic combination of both. The model-driven part ensures the accuracy of demodulation based on physical principles, while the data-driven part adapts to complex actual environments through deep learning networks. This hybrid drive architecture not only improves the robustness of the system, but also enhances the adaptability of the algorithm to different scenarios. Compared with the existing technology, the sequential processing strategy of the present invention realizes the gradual optimization of millimeter wave voice signals. First, the signal quality is improved by virtual vector compensation, and then further denoising and compensation are performed through a multi-stage network. This staged processing method avoids the information loss problem caused by one-time processing in traditional methods. The final output voice signal has significant improvements in clarity, naturalness and intelligibility, especially in high-noise environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0048] Figure 1 This is an overall flow chart of the millimeter wave voice signal processing method according to an embodiment of the present invention;
[0049] Figure 2 This is a flowchart of the millimeter wave speech denoising algorithm according to an embodiment of the present invention;
[0050] Figure 3 Schematic diagram of observed phase difference caused by initial phase difference of observed signals in an embodiment of the present invention;
[0051] Figure 4 A vector diagram of a model-driven virtual vector compensation algorithm according to an embodiment of the present invention; Figure 4 (a) represents the original vector space, and (b) represents the new vector space obtained by the compensation algorithm;
[0052] Figure 5 This is the amplitude denoising network workflow in an embodiment of the present invention;
[0053] Figure 6 This is the phase denoising network workflow in an embodiment of the present invention;
[0054] Figure 7 Schematic diagram of performance comparison of speech enhancement methods in embodiments of the present invention. DETAILED DESCRIPTION
[0055] The present invention proposes a millimeter-wave voice signal processing method and system driven by a hybrid model and data to solve the problem of demodulated voice distortion caused by low signal-to-noise ratio during millimeter-wave voice signal processing and improve the clarity of millimeter-wave voice.
[0056] The core concept of this invention is to first enhance the observed phase of the echo signal through a virtual vector compensation algorithm, amplifying the phase change of the millimeter-wave signal caused by speech vibration. A model-driven virtual vector compensation algorithm is then used to enhance the observed phase of the demodulated millimeter-wave echo signal. Then, to further improve speech clarity, a data-driven multi-stage millimeter-wave speech denoising network is used to sequentially perform amplitude and phase denoising on the speech signal, and a frequency conversion module is used to compensate for speech frequency distortion. Sequentially executing the virtual vector compensation algorithm and the multi-stage millimeter-wave speech denoising algorithm progressively denoises the millimeter-wave signal, reducing noise interference in the demodulated speech and improving speech clarity.
[0057] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0058] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0059] like Figure 1 As shown, the millimeter wave speech signal processing method in the embodiment of the present invention is generally divided into a signal acquisition and preprocessing stage, a signal phase demodulation stage, and a millimeter wave speech denoising algorithm, and finally the content of the speech is determined based on the result of the algorithm recovery. Among them:
[0060] S1, radar echo signal acquisition and preprocessing;
[0061] When the target is making a voice call, the millimeter-wave radar is aimed at the target phone to collect vibration information.
[0062] Perform beamforming processing on the echo signals from multiple receiving antennas of the millimeter-wave radar: ;in, Indicates the The sampling data of the signal received by each antenna, Indicates the antennas in a specific beam direction or angle The weight of Indicates the total number of antennas, is the echo signal obtained after beamforming processing.
[0063] S2. performing phase demodulation on the pre-processed echo signal;
[0064] The intermediate frequency signal is formed into a matrix with the horizontal axis being the distance (range, the physical distance between the target and the radar measured by the time difference of electromagnetic wave propagation) and the vertical axis being the sampling points. Fast Fourier Transform (FFT) is further performed along the range dimension to obtain the spectrum. : ;in, Fast Fourier transform is a key technology in radar signal processing, which is mainly used to extract target distance information from the original echo signal. Extract the range bin of the vibration target (range bin, refers to the discrete distance unit divided by Range FFT) and extract the phase of the vibration signal to obtain the demodulated noisy speech : ;in, and Represent the real and imaginary parts of the extracted range bin respectively.
[0065] S3, using millimeter wave speech denoising algorithm to perform speech denoising;
[0066] This embodiment proposes a high-performance millimeter wave speech denoising algorithm, such as Figure 2 As shown, the specific steps include:
[0067] S31. Processing radar echo signals using a model-driven virtual vector compensation algorithm;
[0068] The position of the target is random, such as Figure 3 As shown in the figure, when the initial phase of the received signal is near the Q axis of the IQ plane, the phase change of the observed signal will not be obvious, which will make it difficult to recover clear voice information. represents the radar observation vector, represents all the primitive static vectors in the environment, is the dynamic vector caused by voice vibration, such as Figure 4 (a) Because the initial phase of the echo signal depends on the target's position and is random, the observed effects of the same target vibration at different locations can vary significantly. A better observed phase can enhance the interpreted speech content. To address this issue, this paper proposes a model-driven virtual vector compensation algorithm to enhance the observed effect of the vibration signal and improve speech accuracy.
[0069] The radar echo phase enhancement is realized based on the model-driven virtual vector compensation algorithm: According to the constant value of the static path vector, the observation signal can be compensated by adding a virtual vector and the phase of the observation signal can be amplified. Figure 4 As shown in Figure 1, when a new virtual vector is introduced, the original static vector and the new virtual vector together form a new static vector. The dynamic vector value remains unchanged, but rotates around the new static vector. From the perspective of vector transformation, adding such a virtual vector actually transforms the original IQ vector space into a new I*Q* vector space. After the introduction of the virtual vector, the static vector is transformed by Rotate to .
[0070] The steps to implement the model-driven virtual vector compensation algorithm are as follows:
[0071] In order to perform quantitative analysis, in this embodiment, a triangle is constructed to describe the effect of the added virtual vector on the observed phase. 、 and They are the original static vector, the virtual vector and the rotated static vector, which are the three sides, angles and yes and The angle between Rotation Angle after forming ,horn yes and The angle between Figure 4 (b)
[0072] To determine , first of all, the observation vector within a period of time Averaging is performed to initially obtain a static vector An approximate estimate of . , for a given rotation angle , Calculated as: ;in, represents the modulus of the original static vector, express The angle with the I axis. For the constructed triangle, the static vector is known 、 Amplitude and phase shift , from which the virtual vector can be calculated The amplitude and phase of the virtual vector The amplitude of is calculated as: ;in, Represents a virtual vector The amplitude of Represents the original static vector The amplitude of Represents the static vector after rotation Amplitude.
[0073] according to ,horn Calculated as , virtual vector Phase angle Expressed as: ;
[0074] For a given phase shift , corresponding to the phase of the virtual vector , is the phase of the original static vector.
[0075] Add a virtual vector to amplify the phase change of the observed signal: Use the obtained virtual vector to compensate the observed signal to enhance the phase change of the observed signal: ;in, represents the observed signal after adding the virtual vector.
[0076] S32. Perform short-time Fourier transform on the radar echo signal processed by the virtual vector compensation algorithm to convert the phase-enhanced millimeter-wave voice signal from the time domain to the time-frequency domain: ;in, is the short-time Fourier transform, is the radar echo signal in the time-frequency domain.
[0077] S33. Use a data-driven multi-stage millimeter-wave speech denoising network to perform amplitude denoising and phase optimization on radar echo signals after short-time Fourier transform.
[0078] Even after enhancing the phase of the observed signal, the radar echo is still affected by device and environmental noise. In the presence of complex nonlinear noise, both the amplitude and phase of the speech are distorted. To obtain clear speech, this embodiment proposes a data-driven, multi-stage learning speech enhancement network to sequentially eliminate speech amplitude and phase noise.
[0079] (1) Speech amplitude denoising.
[0080] Through the amplitude denoising network Amplitude To perform denoising, the network denoising process designed in this embodiment is as follows Figure 5 As shown in Figure 1, it consists of an amplitude encoder, a frequency conversion module, an extended convolution module, and an amplitude decoder. The amplitude encoder maps the input amplitude spectrum to the feature space, and the frequency conversion module compensates for the lost frequency components in the speech. The extended convolution module further extracts important features through multi-scale receptive fields, and the amplitude decoder reconstructs the amplitude spectrum. The calculation process of the amplitude denoising network is:
[0081] ;
[0082] in, and are respectively represented as the mapping function and parameters of the amplitude denoising network, The detailed structure of the amplitude denoising network is shown in Table 1.
[0083] Table 1
[0084]
[0085] (2) Speech phase denoising.
[0086] After amplitude denoising, the denoised amplitude spectrum and the original phase spectrum are combined into a complex spectrum, and then the real and imaginary parts are taken as follows: ; ;
[0087] in, and denote the real and imaginary parts of the complex spectrum, respectively. and Represents the real and imaginary part operations, respectively. yes In the speech phase denoising stage, the trained phase denoising network is used for phase denoising, and the real and imaginary parts are concatenated as the input of the phase denoising network.
[0088] The denoising process of the phase denoising network is as follows Figure 6As shown in Figure 1, it consists of a complex encoder, a dilated convolution module, a real decoder, and an imaginary decoder. The complex encoder encodes the input complex spectrum, and then the dilated convolution module is used to better extract sequence features. Finally, the real decoder and imaginary decoder reconstruct the real and imaginary parts of the complex spectrum, respectively.
[0089] The calculation process of the phase denoising network is as follows: ;
[0090] in, and are respectively represented as the mapping function and parameters of the phase denoising network, It is represented by the real and imaginary parts of the complex spectrum after phase denoising. The detailed structure of the phase denoising network is shown in Table II.
[0091] Table 2
[0092]
[0093] The training process of the data-driven multi-stage millimeter wave speech denoising network is as follows:
[0094] pass Loss training magnitude denoising network: ;
[0095] in, and They represent the mean square error and the amplitude spectrum of the training speech label respectively, and MSE represents the mean square error loss function.
[0096] pass The loss jointly trains the amplitude denoising network and the phase denoising network: ;
[0097] in, The loss function of the phase denoising network is as follows: ;in, and Represent the real and imaginary parts of the complex spectrum of the speech label respectively.
[0098] Speech denoising is performed based on a data-driven multi-stage millimeter wave speech denoising network, specifically including the following steps:
[0099] S331. Convert the speech training samples collected by the millimeter-wave radar into the time-frequency domain to extract the amplitude spectrum, and use the samples and labels to train the amplitude denoising network proposed in the present invention until convergence;
[0100] S332, using the trained converged amplitude denoising network to perform amplitude denoising on the speech training sample, and then combining it with the original phase to form a new complex spectrum;
[0101] S333, jointly training the amplitude denoising network and the phase denoising network using the combined complex spectrum samples and speech complex spectrum labels as well as the amplitude spectrum samples and labels until convergence;
[0102] S334: Input the speech sample to be denoised into the amplitude denoising network and the phase denoising network that have been trained and converged in sequence to achieve speech denoising.
[0103] S4. Calculate the denoised speech signal through short-time inverse Fourier transform to restore the speech information.
[0104] The real and imaginary parts after phase denoising are combined into a complex spectrum and the denoised speech signal is calculated through short-time inverse Fourier transform. : ;in, Represents short-time inverse Fourier transform.
[0105] In the above-mentioned embodiment, the application of the virtual vector compensation algorithm effectively enhances the observed phase of the echo signal, amplifies the phase change of the millimeter-wave signal caused by voice vibration, and significantly improves the sensitivity of phase detection compared to traditional methods. This technical means enables accurate capture of voice vibration information even in a low signal-to-noise ratio environment, laying a solid foundation for subsequent processing. At the same time, the multi-stage millimeter-wave voice denoising network in the above-mentioned embodiment achieves comprehensive optimization of the voice signal through an innovative two-stage processing architecture of amplitude denoising and phase denoising. Among them, the amplitude denoising stage effectively suppresses environmental noise interference, while the phase denoising stage is specifically optimized for the phase distortion problem unique to millimeter-wave signals. In particular, through the design of the frequency conversion module, the frequency distortion of the voice signal is successfully compensated, making the spectral characteristics of the output voice closer to the original voice.
[0106] Corresponding to the millimeter wave speech signal processing method in the above embodiment, this embodiment proposes a millimeter wave speech signal processing system driven by a hybrid model and data, which includes:
[0107] Acquisition and preprocessing module, used for acquiring and preprocessing radar echo signals;
[0108] A phase demodulation module is used to perform phase demodulation on the echo signal preprocessed by the acquisition and preprocessing module;
[0109] The speech denoising module is used to perform speech denoising on the phase-demodulated echo signal obtained by the phase demodulation module using the millimeter-wave speech denoising algorithm, including:
[0110] The model-driven submodule processes the radar echo signal using a model-driven virtual vector compensation algorithm. The model-driven virtual vector compensation algorithm amplifies the phase of the observation signal by adding a virtual vector to compensate the observation signal based on the constant value of the static path vector.
[0111] The first conversion submodule performs a short-time Fourier transform on the radar echo signal processed by the virtual vector compensation algorithm, and converts the phase-enhanced millimeter-wave voice signal from the time domain to the time-frequency domain;
[0112] The data-driven submodule utilizes a data-driven multi-stage millimeter-wave speech denoising network to perform amplitude denoising and phase optimization on the radar echo signal after short-time Fourier transform; the multi-stage millimeter-wave speech denoising network includes: an amplitude denoising network for denoising the speech amplitude and a phase denoising network for optimizing the speech phase; the amplitude denoising network includes: an amplitude encoder, a frequency conversion module, a first extended convolution module and an amplitude decoder; the amplitude encoder maps the input amplitude spectrum to the feature space, and the frequency conversion module compensates for the lost frequency components in the speech, the first extended convolution module extracts features through a multi-scale receptive field, and the amplitude decoder reconstructs the amplitude spectrum; the phase denoising network includes: a complex encoder, a second extended convolution module, a real decoder and an imaginary decoder; the complex encoder encodes the input complex spectrum, the second extended convolution module extracts sequence features, and the real decoder and the imaginary decoder reconstruct the real and imaginary parts of the complex spectrum respectively; the complex spectrum is composed of the denoised amplitude spectrum and the original phase spectrum.
[0113] The signal recovery module is used to calculate the denoised speech signal through short-time inverse Fourier transform and restore the speech information.
[0114] For ease of understanding, the millimeter wave voice signal processing system of the present invention is described below by taking the mobile phone call voice signal processing in an indoor scenario as an example.
[0115] The millimeter-wave radar device used in the system is a radar transceiver with an initial frequency of 77 GHz and a bandwidth set to 3.52 GHz. The model-driven virtual vector compensation algorithm is implemented in MATLAB, and the data-driven multi-stage speech millimeter-wave denoising algorithm is implemented in Python, running on a computer with an Intel i7-11700 CPU, an RTX 3070 Ti GPU, and 32 GB of RAM.
[0116] The millimeter-wave voice signal processing system is used to monitor mobile phone call targets in daily indoor environments. The millimeter-wave radar is used to obtain the target's incoming call voice content, including:
[0117] The acquisition and pre-processing module uses a millimeter-wave radar to aim at the target mobile phone to collect vibration information; it performs beamforming processing on the echo signals of multiple receiving antennas of the millimeter-wave radar;
[0118] The phase demodulation module performs phase demodulation on the echo signal preprocessed by the acquisition and preprocessing module;
[0119] The speech denoising module uses the millimeter wave speech denoising algorithm to perform speech denoising on the phase-demodulated echo signal obtained by the phase demodulation module, including:
[0120] The model-driven submodule processes the radar echo signal using a model-driven virtual vector compensation algorithm. The model-driven virtual vector compensation algorithm amplifies the phase of the observation signal by adding a virtual vector to compensate the observation signal based on the constant value of the static path vector.
[0121] The first conversion submodule performs short-time Fourier transform on the radar echo signal processed by the virtual vector compensation algorithm, and converts the phase-enhanced millimeter-wave voice signal from the time domain to the time-frequency domain;
[0122] The data-driven submodule uses a data-driven multi-stage millimeter-wave speech denoising network to perform amplitude denoising and phase optimization on the radar echo signal after short-time Fourier transform;
[0123] The signal recovery module is used to calculate the denoised speech signal through short-time inverse Fourier transform and restore the speech information.
[0124] The time domain and frequency domain effects of the speech obtained in this embodiment are as follows Figure 7 As shown in the figure, the coarse speech without noise removal using the method of the present invention is severely affected by noise and differs significantly from the original speech. The speech enhanced using the method of the present invention has a high degree of similarity to the original clean speech in both the time domain and the time-frequency domain. These results confirm the excellent speech enhancement effect of the present invention.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A millimeter wave speech signal processing method based on model and data hybrid drive, characterized in that: include: Collect and pre-process radar echo signals, and perform phase demodulation on the pre-processed echo signals; Speech denoising is performed using the millimeter wave speech denoising algorithm, including: The radar echo signal is processed using a model-driven virtual vector compensation algorithm; the model-driven virtual vector compensation algorithm amplifies the phase of the observation signal by adding a virtual vector to compensate the observation signal based on the constant value of the static path vector; Perform short-time Fourier transform on the radar echo signal processed by the virtual vector compensation algorithm to convert the phase-enhanced millimeter-wave voice signal from the time domain to the time-frequency domain; A data-driven multi-stage millimeter wave speech denoising network is used to perform amplitude denoising and phase optimization on radar echo signals after short-time Fourier transform; the multi-stage millimeter wave speech denoising network includes: an amplitude denoising network for denoising speech amplitude and a phase denoising network for optimizing speech phase; the amplitude denoising network includes: an amplitude encoder, a frequency conversion module, a first extended convolution module and an amplitude decoder; the amplitude encoder maps the input amplitude spectrum to a feature space, and the frequency conversion module compensates for the lost frequency components in the speech, the first extended convolution module extracts features through a multi-scale receptive field, and the amplitude decoder reconstructs the amplitude spectrum; the phase denoising network includes: a complex encoder, a second extended convolution module, a real decoder and an imaginary decoder; the complex encoder encodes the input complex spectrum, the second extended convolution module extracts sequence features, and the real decoder and the imaginary decoder reconstruct the real and imaginary parts of the complex spectrum respectively; the complex spectrum is composed of the amplitude denoised amplitude spectrum and the original phase spectrum; The real and imaginary parts after phase denoising are combined and the denoised speech signal is calculated through short-time inverse Fourier transform to restore the speech information.
2. The millimeter wave speech signal processing method based on model and data hybrid drive according to claim 1 is characterized in that: The radar echo signal is processed using a model-driven virtual vector compensation algorithm, including: Radar observation vector over a period of time Averaging is performed to initially obtain an original static vector An approximate estimate of In the IQ plane, a triangle is constructed to describe the effect of the added virtual vector on the observed phase. 、 and are the original static vector, the virtual vector and the rotated static vector, which are the three sides of the triangle respectively. Rotation Angle after forming ,horn yes and The angle between according to , the static vector after rotation Calculated as: ;in, Represents the original static vector The model, express Angle with the I axis; According to the law of cosines, the virtual vector The amplitude of is calculated as: ;in, Represents a virtual vector The amplitude, Represents the original static vector The amplitude, Represents the static vector after rotation Amplitude; according to ,horn Calculated as , virtual vector Phase angle Expressed as: ; is the phase of the static vector; For a given rotation angle , corresponding to the phase of the virtual vector ; The obtained virtual vector is used to compensate the observed signal to enhance the phase change of the observed signal: ;in, represents the observed signal after adding the virtual vector.
3. The millimeter wave speech signal processing method based on model and data hybrid drive according to claim 1, characterized in that: In the amplitude denoising network, the amplitude encoder includes: the first two-dimensional convolution layer, the first batch of normalization layers, the first ReLU nonlinear activation layer and the first maximum pooling layer; the convolution kernel size in the first two-dimensional convolution layer and the first maximum pooling layer is 2×2, and the step size is 1; the frequency conversion module includes: the second two-dimensional convolution layer, the first one-dimensional convolution layer and the third two-dimensional convolution layer; the convolution kernel size of the second two-dimensional convolution layer is 1×1, and the step size is 1; the convolution kernel size of the first one-dimensional convolution layer is 9, and the step size is 1; the convolution kernel size of the third two-dimensional convolution layer is 1×1, and the step size is 1. The convolution kernel size is 2×2 and the stride is 1; the first extended convolution module includes the second one-dimensional convolution layer and the first two-dimensional extended convolution layer; the convolution kernel size of the second one-dimensional convolution layer is 1 and the stride is 1; the convolution kernel size of the first two-dimensional extended convolution layer is 1 and the stride is 1, and the extension layer is [1, 2, 4, 8, 16, 32]; the amplitude decoder includes: the first two-dimensional inverted convolution layer, the second batch normalization layer and the second ReLU nonlinear activation layer; the convolution kernel size of the first two-dimensional inverted convolution layer is 2×2 and the stride is 2.
4. The millimeter wave speech signal processing method based on model and data hybrid drive according to claim 1, characterized in that: In the phase denoising network, the complex encoder includes: a fourth two-dimensional convolutional layer, a third batch normalization layer, a third ReLU nonlinear activation layer, and a second maximum pooling layer; the convolution kernel size in the fourth two-dimensional convolutional layer and the second maximum pooling layer is 2×2, with a stride of 1; the second extended convolution module includes a third one-dimensional convolutional layer and a second two-dimensional extended convolutional layer; the convolution kernel size of the third one-dimensional convolutional layer is 1, with a stride of 1; the convolution kernel size of the second two-dimensional extended convolution layer is 1, with a stride of 1, and the extended layer is [1, 2, 4, 8, 16, 32]; the real part decoder includes: a second two-dimensional inverted convolutional layer, a fourth batch normalization layer, and a fourth ReLU nonlinear activation layer; the convolution kernel size of the second two-dimensional inverted convolution layer is 2×2, with a stride of 2; the imaginary part decoder includes: a third two-dimensional inverted convolutional layer, a fifth batch normalization layer, and a fifth ReLU nonlinear activation layer; the convolution kernel size of the third two-dimensional inverted convolution layer is 2×2, with a stride of 2.
5. The millimeter wave speech signal processing method based on model and data hybrid drive according to claim 1, characterized in that: The data-driven multi-stage millimeter wave speech denoising network training process is as follows: pass Loss training magnitude denoising network: ; in, and Represent the mean square error and the amplitude spectrum of the training speech label respectively; pass The loss jointly trains the amplitude denoising network and the phase denoising network: ; in, The loss function of the phase denoising network is as follows: ;in, and Represent the real and imaginary parts of the complex spectrum of the speech label respectively.
6. The millimeter wave speech signal processing method based on model and data hybrid drive according to claim 1, characterized in that: Acquisition and pre-processing of radar echo signals, including: Use millimeter-wave radar to collect vibration information from the target mobile phone; Beamforming is performed on the echo signals from multiple receiving antennas of the millimeter-wave radar.
7. The millimeter wave speech signal processing method based on model and data hybrid drive according to claim 6, characterized in that: Perform phase demodulation on the pre-processed echo signal, including: The intermediate frequency signal is organized into a matrix with distance on the horizontal axis and sampling points on the vertical axis. Fast Fourier transform is further performed along the distance dimension to obtain the spectrum S. The range bin of the vibration target is extracted from S and the phase of the vibration signal is extracted to obtain the demodulated noisy speech P.
8. A millimeter wave speech signal processing system based on a hybrid model and data drive, characterized in that: include: Acquisition and preprocessing module, used for acquiring and preprocessing radar echo signals; The phase demodulation module is used to perform phase demodulation on the echo signal preprocessed by the acquisition and preprocessing module; The speech denoising module is used to perform speech denoising on the phase-demodulated echo signal obtained by the phase demodulation module using the millimeter-wave speech denoising algorithm, including: The model-driven submodule processes the radar echo signal using a model-driven virtual vector compensation algorithm. The model-driven virtual vector compensation algorithm amplifies the phase of the observation signal by adding a virtual vector to compensate the observation signal based on the constant value of the static path vector. The first conversion submodule performs a short-time Fourier transform on the radar echo signal processed by the virtual vector compensation algorithm, and converts the phase-enhanced millimeter-wave voice signal from the time domain to the time-frequency domain; A data-driven submodule utilizes a data-driven multi-stage millimeter-wave speech denoising network to perform amplitude denoising and phase optimization on the radar echo signal after short-time Fourier transform; the multi-stage millimeter-wave speech denoising network includes: an amplitude denoising network for denoising the speech amplitude and a phase denoising network for optimizing the speech phase; the amplitude denoising network includes: an amplitude encoder, a frequency conversion module, a first extended convolution module and an amplitude decoder; the amplitude encoder maps the input amplitude spectrum to the feature space, and the frequency conversion module compensates for the lost frequency components in the speech, the first extended convolution module extracts features through a multi-scale receptive field, and the amplitude decoder reconstructs the amplitude spectrum; the phase denoising network includes: a complex encoder, a second extended convolution module, a real decoder and an imaginary decoder; the complex encoder encodes the input complex spectrum, the second extended convolution module extracts sequence features, and the real decoder and the imaginary decoder reconstruct the real and imaginary parts of the complex spectrum respectively; the complex spectrum is composed of the amplitude denoised amplitude spectrum and the original phase spectrum. The signal recovery module is used to combine the real part and the imaginary part after phase denoising to calculate the denoised speech signal through short-time inverse Fourier transform and restore the speech information.
Citation Information
Patent Citations
Complex field speech enhancement method and system based on generative adversarial network and medium
CN110739002A
Audio enhancement method and apparatus, and electronic device and readable storage medium
WO2023226839A1