Pure voice generation method, device and equipment based on noise Mel spectrum

By using feature enhancement models, amplitude prediction models and phase prediction models in speech synthesis technology, the Mel spectrogram of noise pollution is solved, and the problems of insufficient noise robustness and poor generation quality in the prior art are achieved, and high-quality pure speech generation is achieved.

CN120108374APending Publication Date: 2025-06-06PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337951.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient noise robustness, imbalance of amplitude and phase modeling, contradiction between generation quality and complexity, lack of dedicated mechanisms for noise processing, and poor adaptability of multiple scenarios, resulting in poor speech synthesis quality in noise environments.

Method used

By inputting the Mel spectrogram with noise-existing characteristics to the enhancement model, the clear spectrum features are extracted using convolutional neural network and self-attention mechanism to generate the Mel spectrogram with enhanced speech characteristics; then input it into the amplitude prediction model and the phase prediction model, and the clean amplitude spectrum and the continuous smooth phase spectrum are extracted respectively, and a pure speech is generated.

Benefits of technology

Effectively remove noise and generate clear and natural speech waveforms, improve the sound quality and naturalness of the generated waveforms, obtain high-quality speech, and significantly improve the performance of speech synthesis in noise environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108374A_ABST
    Figure CN120108374A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence and the field of medical health, and discloses a pure speech generation method and device based on a noise Mel spectrum, computer equipment and a medium. Using a convolutional neural network and a self-attention mechanism to extract voice-related clear spectrum features from the Mel-frequency spectrogram with noise, and generating a Mel-frequency spectrogram with enhanced voice features; inputting the voice feature enhanced Mel spectrogram into an amplitude prediction model to extract related amplitude spectrum information, and generating a clean amplitude spectrum; inputting the voice feature enhanced Mel spectrogram into a phase prediction model, and generating a phase spectrum with a continuous and smooth phase by using a phase decoupling mechanism and a recurrent neural network; and combining the amplitude spectrum with the phase spectrum to generate pure voice. According to the scheme, the noise in the input Mel spectrogram can be effectively removed, and a clear and natural voice waveform is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and medical health, and in particular to a method, device, computer equipment and medium for generating pure speech based on noisy mel spectrum. Background Art

[0002] With the rapid development of speech technology, speech synthesis and enhancement technology plays an increasingly important role in many fields (such as speech recognition, automatic speech generation, voice communication, etc.). Speech synthesis and enhancement technology can support functions such as disease auxiliary diagnosis, health management, and remote consultation. In these applications, the vocoder is responsible for converting the Mel spectrogram into an audible sound waveform. However, the existing technology has serious deficiencies in processing noise-contaminated inputs: including insufficient robustness to noise, imbalance in amplitude and phase modeling, contradiction between generation quality and complexity, lack of dedicated mechanisms for noise processing, and poor adaptability to multiple scenarios. These problems limit the effectiveness of existing vocoders in practical applications, especially the quality of speech synthesis in noisy environments. Summary of the invention

[0003] The present invention provides a method, device, computer equipment and medium for generating pure speech based on noisy mel spectrum, aiming to solve the problem of speech synthesis quality in a noisy environment.

[0004] In a first aspect, a method for generating clean speech based on noisy mel spectrum is provided, comprising the following steps:

[0005] The noisy Mel-spectrogram is input into the feature enhancement model, and the convolutional neural network and self-attention mechanism are used to extract clear spectral features related to speech from the noisy Mel-spectrogram to generate a Mel-spectrogram with enhanced speech features;

[0006] The mel-spectrogram with enhanced speech features is input into the amplitude prediction model to extract relevant amplitude spectrum information and generate a clean amplitude spectrum;

[0007] The mel-spectrogram with enhanced speech features is input into the phase prediction model, and a phase spectrum with continuous and smooth phase is generated by using the phase decoupling mechanism and recurrent neural network.

[0008] Combining the magnitude spectrum with the phase spectrum generates pure speech.

[0009] In a second aspect, a clean speech generation device based on noisy mel spectrum is provided, comprising:

[0010] The feature enhancement module is used to input the noisy Mel-spectrogram into the feature enhancement model, extract the clear spectral features related to speech from the noisy Mel-spectrogram using the convolutional neural network and the self-attention mechanism, and generate a Mel-spectrogram with enhanced speech features;

[0011] The amplitude prediction module is used to input the Mel-spectrogram with enhanced speech features into the amplitude prediction model to extract relevant amplitude spectrum information and generate a clean amplitude spectrum;

[0012] The phase prediction module is used to input the Mel-spectrogram with enhanced speech features into the phase prediction model, and generate a phase spectrum with continuous and smooth phase by using the phase decoupling mechanism and recurrent neural network;

[0013] The waveform reconstruction module is used to combine the amplitude spectrum and the phase spectrum to generate pure speech.

[0014] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the above-mentioned method for generating clean speech based on noisy Mel-spectrogram are implemented.

[0015] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for generating clean speech based on noisy Mel-spectrogram are implemented.

[0016] In the scheme implemented by the above-mentioned method, device, computer equipment and medium for generating pure speech based on noisy Mel-spectrogram, by inputting the Mel-spectrogram with noise into the feature enhancement model, using the convolutional neural network and the self-attention mechanism to extract clear spectral features related to speech from the Mel-spectrogram with noise, a Mel-spectrogram with enhanced speech features is generated; the Mel-spectrogram with enhanced speech features is input into the amplitude prediction model to extract relevant amplitude spectrum information to generate a clean amplitude spectrum; the Mel-spectrogram with enhanced speech features is input into the phase prediction model, using the phase decoupling mechanism and the recurrent neural network to generate a phase spectrum with continuous and smooth phase; the amplitude spectrum and the phase spectrum are combined to generate pure speech. In the scheme provided by the present invention, the noise in the input Mel-spectrogram can be effectively removed by the feature enhancement model to generate a clear and natural speech waveform, and the combination of the amplitude prediction model and the phase prediction model can greatly improve the sound quality and naturalness of the generated waveform, thereby obtaining high-quality speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0018] Figure 11 is a schematic diagram of an application environment of a method for generating clean speech based on noisy mel spectrum in one embodiment of the present invention;

[0019] Figure 2 It is a flow chart of a method for generating clean speech based on noisy mel spectrum in one embodiment of the present invention;

[0020] Figure 3 is a schematic diagram of a process of feature enhancement in one embodiment of the present invention;

[0021] Figure 4 is a schematic diagram of a flow chart of amplitude prediction in one embodiment of the present invention;

[0022] Figure 5 is a schematic diagram of a process of phase prediction in one embodiment of the present invention;

[0023] Figure 6 is a schematic diagram of a process of combining an amplitude spectrum and a phase spectrum in one embodiment of the present invention;

[0024] Figure 7 is a schematic diagram of a clean speech generation device based on noisy mel spectrum in one embodiment of the present invention;

[0025] Figure 8 is a schematic diagram of a structure of a computer device in one embodiment of the present invention;

[0026] Fig. 9 It is another structural schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] The method for generating pure speech based on noisy mel spectrum provided by the embodiment of the present invention is applied in the following aspects: Figure 1In the application environment, the client communicates with the server through the network. The client can record the audio signal and send the audio signal to the server to view the data results returned by the server. The server can convert the audio signal into a mel-spectrogram, input the noisy mel-spectrogram into the feature enhancement model, use the convolutional neural network and self-attention mechanism to extract the clear spectral features related to speech from the noisy mel-spectrogram, and generate a mel-spectrogram with enhanced speech features; input the mel-spectrogram with enhanced speech features into the amplitude prediction model to extract the relevant amplitude spectrum information and generate a clean amplitude spectrum; input the mel-spectrogram with enhanced speech features into the phase prediction model, use the phase decoupling mechanism and recurrent neural network to generate a phase spectrum with continuous and smooth phase; combine the amplitude spectrum with the phase spectrum to generate pure speech.

[0029] In the solution provided by the present invention, the noise in the input mel-spectrogram can be effectively removed through the feature enhancement model to generate a clear and natural speech waveform. The combination of the amplitude prediction model and the phase prediction model can greatly improve the sound quality and naturalness of the generated waveform, thereby obtaining high-quality speech. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.

[0030] In the technical solution of the present invention, clean speech refers to the original speech signal that is not interfered by noise, that is, the audio data containing only the speaker's voice. The signal of clean speech has clear speech characteristics (such as fundamental frequency, formant, etc.) in the time domain and frequency domain, and does not contain any background noise or interference components. The clean amplitude spectrum refers to the energy distribution of the clean speech signal in the frequency domain.

[0031] See also Figure 2 As shown, Figure 2 A flow chart of a method for generating clean speech based on noisy mel spectrum provided in an embodiment of the present invention includes the following steps:

[0032] S10. Input the noisy Mel-spectrogram into the feature enhancement model, use the convolutional neural network and self-attention mechanism to extract clear spectral features related to speech from the noisy Mel-spectrogram, and generate a Mel-spectrogram with enhanced speech features.

[0033] like Figure 3 As shown, in a specific embodiment, step S10 specifically includes:

[0034] S11, inputting the noisy Mel-spectrogram into the feature enhancement model;

[0035] S12, extracting amplitude spectrum and phase spectrum from the noisy Mel spectrum map;

[0036] S13. The amplitude spectrum and phase spectrum are transformed through convolutional neural networks, and the long-distance dependencies in the frequency and time domains are captured through the attention mechanism to extract clear spectral features related to speech and generate a Mel-spectrogram with enhanced speech features.

[0037] In this embodiment, the feature enhancement model adopts an architecture combining a multi-scale convolutional neural network and a multi-head self-attention mechanism, and the specific design is as follows:

[0038] Input preprocessing: The dimension of the noise Mel spectrum is (Batch_size, 80, T), where 80 is the number of Mel frequency bands and T is the number of time frames. Normalization is required before input, using global mean-variance normalization to eliminate the distribution differences of different noise intensities.

[0039] The multi-scale convolutional neural network contains three parallel convolution branches, each using different convolution kernels (3×3, 5×5, 7×7) to capture the diversity of local frequency domain features. Each branch contains four layers of convolution, the activation function is LeakyReLU (α=0.2), and the output is fused by channel splicing.

[0040] Self-attention mechanism: Multi-head self-attention (4 heads) is applied in the time-frequency dimension, and position encoding is introduced when calculating the attention weight to preserve the temporal information. The attention output is added to the multi-scale convolutional neural network features, and the gradient disappearance is alleviated through residual connection.

[0041] Output layer: A 1×1 convolution is used to restore the number of channels to 80 and generate an enhanced Mel-spectrogram.

[0042] This embodiment can effectively remove noise in the Mel-spectrogram through the feature enhancement model to generate a clear and natural speech waveform.

[0043] S20, inputting the Mel-spectrogram with enhanced speech features into the amplitude prediction model to extract relevant amplitude spectrum information and generate a clean amplitude spectrum.

[0044] like Figure 4 As shown, in a specific embodiment, step S20 specifically includes:

[0045] S21, inputting the Mel-spectrogram with enhanced speech features into the amplitude prediction model;

[0046] S22. The Mel-spectrogram with speech feature enhancement extracts acoustic features through Fourier transform, and extracts relevant amplitude spectrum information from the acoustic features to generate a clean amplitude spectrum.

[0047] In this embodiment, the amplitude prediction model is provided with a Transformer encoder and residual block: it processes time series features, captures the dynamic changes of speech, and the residual block enhances the learning ability of complex spectral features. The amplitude prediction model focuses on recovering a clean amplitude spectrum from the enhanced Mel spectrum. The enhanced Mel spectrum is converted into a complex spectrum through Fourier transform, and the amplitude spectrum and phase spectrum are extracted as intermediate features. For example: Transformer encoder: includes 5 downsampling blocks, each block consists of 2 layers of convolution (kernel size 3×3, step size 1) and maximum pooling (2×2), and the number of channels increases from 64 to 512 layer by layer. Transformer decoder: upsampling through transposed convolution, each layer is jump-connected to the corresponding feature of the encoder, and the last layer outputs the predicted amplitude spectrum. Loss function: the mean square error loss of the amplitude spectrum is used, the weight is set to 0.7, and the energy distribution in the low-frequency area is optimized.

[0048] S30, inputting the Mel-spectrogram with enhanced speech features into the phase prediction model, and using the phase decoupling mechanism and recurrent neural network to generate a phase spectrum that is phase continuous and smooth.

[0049] like Figure 5 As shown, in a specific embodiment, step S30 specifically includes:

[0050] S31, inputting the Mel-spectrogram with enhanced speech features into the phase prediction model;

[0051] S32, through the phase decoupling mechanism, the phase modeling is decomposed into global phase prediction and local phase fine-tuning, and phase macro information processing and phase detail information adjustment are performed respectively;

[0052] S33. Through the recurrent neural network, the temporal correlation of the phase is captured to generate a phase spectrum that is continuous and smooth.

[0053] In this embodiment, the phase prediction model adopts a combination of a phase decoupling mechanism and a recurrent neural network. Global phase prediction: Generate a phase basis vector through a fully connected layer to characterize the overall rhythm of the speech. Local phase fine-tuning: Capture the phase continuity between frames and output the phase residual. Recurrent neural network design: The input is the concatenated features of the enhanced Mel spectrum and the global phase, and the time step is aligned to T frames. The output layer uses the Tanh activation function to limit the phase residual to the range of [-π,π]. Loss function: Phase continuity loss, calculates the smoothness of the first-order derivative of the phase difference between adjacent frames, and the weight is set to 0.3. Phase truth loss: The cosine similarity is used to measure the consistency between the predicted phase and the true phase, and the weight is set to 0.5.

[0054] S40, combining the amplitude spectrum and the phase spectrum to generate pure speech.

[0055] like Figure 6As shown, in a specific embodiment, step S40 specifically includes:

[0056] S41, combining the predicted amplitude spectrum and phase spectrum to generate a complex spectrum;

[0057] S42, performing inverse fast Fourier transform on the complex spectrum to generate a time domain waveform to obtain pure speech.

[0058] In this embodiment, in the synthesis process of the amplitude spectrum and the phase spectrum, complex spectrum reconstruction is first performed: the predicted amplitude spectrum and phase spectrum are combined into a complex form; then an inverse fast Fourier transform is performed: an inverse fast Fourier transform is performed on the complex spectrum to generate a time domain waveform with a sampling rate of 16kHz.

[0059] In a specific embodiment, the method for generating clean speech based on noisy mel-spectrogram further includes using training data to jointly train a feature enhancement model, an amplitude prediction model, and a phase prediction model; the loss function used is a multi-task loss function that combines amplitude prediction loss, phase prediction loss, and waveform adversarial loss. The training data contains noises of various types and intensities.

[0060] In this embodiment, the three models of feature enhancement, amplitude prediction, and phase prediction are jointly trained, wherein the amplitude prediction loss uses the mean square error to measure the difference between the predicted amplitude and the true amplitude; the phase prediction loss uses the cosine similarity loss to optimize the phase continuity; and the waveform adversarial loss optimizes the authenticity of the generated waveform through a generative adversarial network discriminator.

[0061] In summary, the technical solution of the above embodiment significantly improves the performance of speech synthesis in a noisy environment in multiple dimensions through innovative designs such as feature enhancement, amplitude-phase decoupling prediction, and multi-task joint training: it has a significant effect on suppressing sudden noise and improving noise robustness. The problem of spectrum over-smoothing is reduced through the amplitude prediction model. Phase continuity loss effectively reduces phase jumps, and the clarity of speech segmentation is improved.

[0062] In a specific embodiment, a method for generating clean speech based on noisy mel-spectrogram is provided for application in a medical scenario: speech monitoring and communication enhancement of patients in an intensive care unit (ICU).

[0063] In the intensive care unit (ICU), patients often have weak voice signals due to intubation, ventilator use or weakness, and their voice signals are seriously polluted by equipment noise (such as ventilator beeps, monitor alarms, etc.). Medical staff need to judge the patient's pain level or emergency needs through the patient's voice (such as groans, short words), but traditional voice acquisition systems have difficulty separating effective voice from background noise, which may lead to the omission of key information and delay treatment.

[0064] After the noisy Mel-spectrogram collected by the ICU bedside device is input into the feature enhancement model, the convolutional neural network and self-attention mechanism are used to separate the low-frequency noise of the ventilator (such as periodic airflow sound) and the high-frequency features of the patient's voice (such as intermittent words) to generate an enhanced Mel-spectrogram.

[0065] Medical value: Effectively suppress ventilator noise, preserve the spectral characteristics of the patient's weak voice, and provide high-quality input for subsequent processing.

[0066] The enhanced Mel-spectrogram is input into the amplitude prediction model to extract acoustic features through Fourier transform and predict the pure amplitude spectrum. For example, the amplitude information of words such as "pain" or "water" spoken intermittently by the patient is accurately restored to avoid amplitude distortion caused by noise interference.

[0067] The phase prediction model uses a phase decoupling mechanism and a recurrent neural network (RNN) to decompose the phase of the patient's speech into global (such as breathing rhythm) and local (such as syllable boundaries) information, generating a continuous and smooth phase spectrum to ensure that the synthesized speech is natural and understandable.

[0068] After the recovered amplitude spectrum is combined with the phase spectrum, the time domain waveform is generated through inverse Fourier transform, and a clear speech signal is output. The system can be further integrated with the ICU monitoring platform.

[0069] When high-frequency keywords (such as "help" and "pain") are detected, the nurse station will automatically trigger an alarm; the synthesized voice is transmitted to the medical staff's headphones in real time to reduce environmental noise interference and improve response efficiency.

[0070] The technical advantages and medical value of this technical solution are:

[0071] Multi-noise joint training: Typical ICU noises (such as equipment alarms and the sound of people walking) are introduced during model training to enhance the adaptability to complex scenarios.

[0072] End-to-end closed-loop management: From noise suppression to speech synthesis, a "collection-enhancement-analysis-response" closed loop is formed to facilitate precise care.

[0073] Balancing privacy and efficiency: Patient voice is processed locally before being transmitted, avoiding the risk of leakage of original noisy data and complying with medical data security regulations.

[0074] This technical solution can be extended to scenarios such as postoperative voice rehabilitation training (such as voice enhancement for laryngeal cancer patients) and remote home monitoring (such as reporting of voice symptoms by patients with chronic diseases), promoting the full penetration of voice enhancement technology in the medical field.

[0075] like Figure 7 As shown, in a specific embodiment, a clean speech generation device based on noise mel spectrum is provided, comprising:

[0076] The feature enhancement module 10 is used to input the noisy Mel-spectrogram into the feature enhancement model, extract clear spectral features related to speech from the noisy Mel-spectrogram using a convolutional neural network and a self-attention mechanism, and generate a Mel-spectrogram with enhanced speech features.

[0077] In a specific embodiment, the feature enhancement module 10 is specifically used for:

[0078] Input the noisy Mel-spectrogram into the feature enhancement model;

[0079] Extract amplitude and phase spectra from noisy mel-spectrograms;

[0080] The amplitude spectrum and phase spectrum are transformed through convolutional neural networks, and the long-distance dependencies in the frequency and time domains are captured through the attention mechanism to extract clear spectral features related to speech and generate a Mel-spectrogram with enhanced speech features.

[0081] The amplitude prediction module 20 is used to input the mel-spectrogram with enhanced speech features into the amplitude prediction model to extract relevant amplitude spectrum information and generate a clean amplitude spectrum.

[0082] In a specific embodiment, the amplitude prediction module 20 is specifically used for:

[0083] The mel-spectrogram with enhanced speech features is input into the amplitude prediction model;

[0084] The Mel-spectrogram with speech feature enhancement extracts acoustic features through Fourier transform, and extracts relevant amplitude spectrum information from the acoustic features to generate a clean amplitude spectrum.

[0085] The phase prediction module 30 is used to input the Mel-spectrogram with enhanced speech features into the phase prediction model, and generate a phase spectrum that is continuous and smooth by using a phase decoupling mechanism and a recurrent neural network.

[0086] In a specific embodiment, the phase prediction module 30 is specifically used for:

[0087] The mel-spectrogram enhanced with speech features is input into the phase prediction model;

[0088] Through the phase decoupling mechanism, phase modeling is decomposed into global phase prediction and local phase fine-tuning, and phase macro information processing and phase detail information adjustment are performed respectively;

[0089] Through the recurrent neural network, the temporal correlation of the phase is captured and a phase-continuous and smooth phase spectrum is generated.

[0090] The waveform reconstruction module 40 is used to combine the amplitude spectrum and the phase spectrum to generate pure speech.

[0091] In a specific embodiment, the waveform reconstruction module 40 is specifically used for:

[0092] Combining the predicted amplitude spectrum with the phase spectrum to generate a complex spectrum;

[0093] The complex spectrum is subjected to inverse fast Fourier transform to generate a time domain waveform and obtain pure speech.

[0094] In a specific embodiment, the apparatus for generating clean speech based on noisy mel-spectrogram further includes a model training module 10', which is used to jointly train the feature enhancement model, the amplitude prediction model, and the phase prediction model using training data; the loss function used is a multi-task loss function combining the amplitude prediction loss, the phase prediction loss, and the waveform adversarial loss. The training data contains noises of various types and intensities.

[0095] For the specific definition of the clean speech generation device based on the noise Mel spectrum, please refer to the definition of the clean speech generation method based on the noise Mel spectrum above, which will not be repeated here. Each module in the above-mentioned clean speech generation device based on the noise Mel spectrum can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0096] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, the functions or steps of a server-side method for generating pure speech based on noise Mel spectrum are implemented.

[0097] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Fig. 9As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, the functions or steps of the client side of a method for generating clean speech based on noise mel spectrum are implemented.

[0098] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:

[0099] The noisy Mel-spectrogram is input into the feature enhancement model, and the convolutional neural network and self-attention mechanism are used to extract clear spectral features related to speech from the noisy Mel-spectrogram to generate a Mel-spectrogram with enhanced speech features;

[0100] The mel-spectrogram with enhanced speech features is input into the amplitude prediction model to extract relevant amplitude spectrum information and generate a clean amplitude spectrum;

[0101] The mel-spectrogram with enhanced speech features is input into the phase prediction model, and a phase spectrum with continuous and smooth phase is generated by using the phase decoupling mechanism and recurrent neural network.

[0102] Combining the magnitude spectrum with the phase spectrum generates pure speech.

[0103] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0104] The noisy Mel-spectrogram is input into the feature enhancement model, and the convolutional neural network and self-attention mechanism are used to extract clear spectral features related to speech from the noisy Mel-spectrogram to generate a Mel-spectrogram with enhanced speech features;

[0105] The mel-spectrogram with enhanced speech features is input into the amplitude prediction model to extract relevant amplitude spectrum information and generate a clean amplitude spectrum;

[0106] The mel-spectrogram with enhanced speech features is input into the phase prediction model, and a phase spectrum with continuous and smooth phase is generated by using the phase decoupling mechanism and recurrent neural network.

[0107] Combining the magnitude spectrum with the phase spectrum generates pure speech.

[0108] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0109] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0110] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0111] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A method for generating pure speech based on noisy mel spectrum, characterized in that: The following steps are involved: The noisy Mel-spectrogram is input into the feature enhancement model, and the convolutional neural network and self-attention mechanism are used to extract clear spectral features related to speech from the noisy Mel-spectrogram to generate a Mel-spectrogram with enhanced speech features; The mel-spectrogram with enhanced speech features is input into the amplitude prediction model to extract relevant amplitude spectrum information and generate a clean amplitude spectrum; The mel-spectrogram with enhanced speech features is input into the phase prediction model, and a phase spectrum with continuous and smooth phase is generated by using the phase decoupling mechanism and recurrent neural network. Combining the magnitude spectrum with the phase spectrum generates pure speech.

2. The method for generating clean speech based on noisy mel-spectrogram according to claim 1, characterized in that: The step of inputting the noisy Mel-spectrogram into the feature enhancement model, extracting clear spectral features related to speech from the noisy Mel-spectrogram using a convolutional neural network and a self-attention mechanism, and generating a Mel-spectrogram with enhanced speech features includes: Input the noisy Mel-spectrogram into the feature enhancement model; Extract amplitude and phase spectra from noisy mel-spectrograms; The amplitude spectrum and phase spectrum are transformed through convolutional neural networks, and the long-distance dependencies in the frequency and time domains are captured through the attention mechanism to extract clear spectral features related to speech and generate a Mel-spectrogram with enhanced speech features.

3. The method for generating clean speech based on noisy mel-spectrogram according to claim 1, characterized in that: The step of inputting the Mel-spectrogram with enhanced speech features into the amplitude prediction model to extract relevant amplitude spectrum information and generate a clean amplitude spectrum comprises: The mel-spectrogram with enhanced speech features is input into the amplitude prediction model; The Mel-spectrogram with speech feature enhancement extracts acoustic features through Fourier transform, and extracts relevant amplitude spectrum information from the acoustic features to generate a clean amplitude spectrum.

4. The method for generating clean speech based on noisy mel-spectrogram according to claim 1, characterized in that: The step of inputting the Mel-spectrogram with enhanced speech features into the phase prediction model and using the phase decoupling mechanism and the recurrent neural network to generate a phase spectrum with continuous and smooth phase comprises: The mel-spectrogram enhanced with speech features is input into the phase prediction model; Through the phase decoupling mechanism, phase modeling is decomposed into global phase prediction and local phase fine-tuning, and phase macro information processing and phase detail information adjustment are performed respectively; Through the recurrent neural network, the temporal correlation of the phase is captured and a phase-continuous and smooth phase spectrum is generated.

5. The method for generating clean speech based on noisy mel-spectrogram according to claim 1, characterized in that: The step of combining the amplitude spectrum with the phase spectrum to generate pure speech comprises: Combining the predicted amplitude spectrum with the phase spectrum to generate a complex spectrum; The complex spectrum is subjected to inverse fast Fourier transform to generate a time domain waveform and obtain pure speech.

6. The method for generating clean speech based on noisy mel-spectrogram according to claim 1, characterized in that: It also includes using training data to jointly train the feature enhancement model, amplitude prediction model, and phase prediction model; the loss function used is a multi-task loss function that combines amplitude prediction loss, phase prediction loss, and waveform adversarial loss.

7. The method for generating clean speech based on noisy mel-spectrogram according to claim 6, characterized in that: The training data has noises of various types and intensities added thereto.

8. A clean speech generation device based on noisy mel spectrum, characterized in that: include: The feature enhancement module is used to input the noisy Mel-spectrogram into the feature enhancement model, extract the clear spectral features related to speech from the noisy Mel-spectrogram using the convolutional neural network and the self-attention mechanism, and generate a Mel-spectrogram with enhanced speech features; The amplitude prediction module is used to input the Mel-spectrogram with enhanced speech features into the amplitude prediction model to extract relevant amplitude spectrum information and generate a clean amplitude spectrum; The phase prediction module is used to input the Mel-spectrogram with enhanced speech features into the phase prediction model, and generate a phase spectrum with continuous and smooth phase by using the phase decoupling mechanism and recurrent neural network; The waveform reconstruction module is used to combine the amplitude spectrum and the phase spectrum to generate pure speech.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for generating clean speech based on noisy Mel-spectrogram are implemented as claimed in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for generating clean speech based on noisy Mel-spectrogram are implemented as claimed in any one of claims 1 to 7.