A training method for a voice noise reduction model and a voice enhancement method

A multi-stage speech enhancement model addresses low signal-to-noise challenges by tailoring processing stages to improve speech quality and intelligibility through frequency spectrum and masking techniques.

CN116137153BActive Publication Date: 2025-07-15INST OF ACOUSTICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111353720.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-16
Publication Date
2025-07-15
Estimated Expiration
2041-11-16

AI Technical Summary

Technical Problem

Existing single-channel speech enhancement methods struggle with low signal-to-noise ratios, leading to inadequate noise reduction performance, especially in challenging environments.

Method used

A multi-stage noise reduction model is trained with frequency spectrum and masking processes to enhance speech quality, utilizing different modules based on signal-to-noise ratios to optimize noise reduction.

Benefits of technology

The model effectively improves speech enhancement by adapting processing stages to varying noise levels, enhancing speech quality and intelligibility, especially in low signal-to-noise conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116137153B_ABST
    Figure CN116137153B_ABST
Patent Text Reader

Abstract

The present application provides a training method for a voice noise reduction model and a voice enhancement method. The voice noise reduction model includes: a first enhancement module and a second enhancement module. The first enhancement module is used to perform noise reduction processing on the input spectrum and output the spectrum; the second enhancement module is used to perform noise reduction processing on the input spectrum and output a complex mask. The processing order of the first enhancement module and the second enhancement module is determined according to the signal-to-noise ratio of the sound channel. Among them, when the signal-to-noise ratio of the sound channel is less than a preset value, the first enhancement module is first used for processing to restore the voice harmonics, and then the second enhancement module is used for processing to enhance the noise reduction performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech enhancement technology, and in particular, to a method for training a speech noise reduction model and a speech enhancement method. Background Art

[0002] In the application scenarios of speech, the perceived speech usually contains environmental interference from noise sources. For example, in scenarios such as automatic speech recognition, telecommunication systems, and hearing assistant devices, the noise in the speech will affect the actual application of the speech. The purpose of speech enhancement is to extract useful speech information from the noisy speech, thereby improving the speech quality and intelligibility.

[0003] Since the mono channel lacks spatial information, the enhancement of mono speech is a challenging topic. Especially in the case of low signal-to-noise ratio, the existing speech enhancement methods have low noise reduction performance for mono speech. Summary of the Invention

[0004] This application provides a method for training a speech noise reduction model with multi-stage noise reduction and a speech enhancement method. For mono speech, two processing links of spectrum processing and masking processing are designed to improve the performance of speech enhancement in a mono environment.

[0005] In a first aspect, this application provides a method for training a speech noise reduction model.

[0006] The method includes: obtaining a speech training set corresponding to a sound channel; the speech training set includes a plurality of noisy speech samples and a plurality of clean speech samples, and the plurality of noisy speech samples and the plurality of clean speech samples correspond one by one; using the speech training set to determine a speech noise reduction model corresponding to the sound channel; the speech noise reduction model includes: an analysis filter, a first enhancement module, a second enhancement module, and a synthesis filter module;

[0007] Wherein, the using the speech training set to determine the speech noise reduction model corresponding to the sound channel includes:

[0008] The analysis filter converts the input noisy speech sample into a first Fourier spectrum; when the signal-to-noise ratio of the sound channel is less than a preset value, the first enhancement module outputs a second Fourier spectrum based on the first Fourier spectrum, splices the first Fourier spectrum and the second Fourier spectrum into a third Fourier spectrum, the second enhancement module outputs a first complex mask based on the third Fourier spectrum, and converts the first complex mask into a fourth Fourier spectrum; when the signal-to-noise ratio of the sound channel is not less than the preset value, the second enhancement module outputs a second complex mask based on the first Fourier spectrum, converts the second complex mask into a fifth Fourier spectrum, splices the first Fourier spectrum and the fifth Fourier spectrum into a sixth Fourier spectrum, and the first enhancement module outputs a seventh Fourier spectrum based on the sixth Fourier spectrum; the synthesis filter module converts the fourth Fourier spectrum or the seventh Fourier spectrum into clean speech; and updates the speech denoising model according to the clean speech and the clean speech sample corresponding to the input noisy speech sample.

[0009] In the above solution, two stages of noise reduction links are designed in the speech denoising model according to the different signal-to-noise ratios of the sound channel, including a module for outputting a spectrum and a module for outputting a mask. Combining the advantages of the spectrum mapping and mask methods can enhance the effect of speech denoising. Especially in the case of a low signal-to-noise ratio, first use the module for outputting a spectrum to process to restore the speech components, and then use the module for outputting a mask to process to enhance the noise reduction performance, thereby improving the quality of speech enhancement.

[0010] In a possible implementation manner, the first enhancement module includes a plurality of first sub-modules, and the second enhancement module includes a plurality of second sub-modules; wherein, the number of the first sub-modules and the number of the second sub-modules are determined according to the signal-to-noise ratio of the sound channel; the first sub-module is used for performing noise reduction processing on the input Fourier spectrum and outputting a noise-reduced Fourier spectrum; the second sub-module is used for performing noise reduction processing on the input Fourier spectrum and outputting a noise-reduced complex mask.

[0011] When the signal-to-noise ratio of the sound channel is less than the preset value, the Fourier spectrum input to the nth first sub-module is determined according to the Fourier spectrum output by the analysis filter and the Fourier spectrum output by the (n - i)th first sub-module, and the Fourier spectrum input to the mth second sub-module is determined according to the Fourier spectrum output by the analysis filter, the Fourier spectrum output by each first sub-module, and the Fourier spectrum corresponding to the complex mask output by the (m - j)th second sub-module. Wherein, n ∈ [1, N], N is the number of the first sub-modules, i ∈ [1, n], m ∈ [1, M], M is the number of the second sub-modules, and j ∈ [1, m].

[0012] When the signal-to-noise ratio of the sound channel is not less than the preset value, the Fourier spectrum input to the m-th second sub-module is determined according to the Fourier spectrum output by the analysis filter and the Fourier spectrum output by the (m-j)-th second sub-module. The Fourier spectrum input to the n-th first sub-module is determined according to the Fourier spectrum output by the analysis filter, the Fourier spectrum corresponding to the complex mask output by each second sub-module, and the Fourier spectrum output by the (n-i)-th first sub-module. Wherein, n ∈ [1, N], N is the number of first sub-modules, i ∈ [1, n], m ∈ [1, M], M is the number of second sub-modules, and j ∈ [1, m].

[0013] In a possible implementation manner, both the first sub-module and the second sub-module include: an encoder, a time-domain recurrent neural network, a frequency-domain recurrent neural network, and a decoder; wherein, the encoder is configured to encode the input Fourier spectrum to obtain a first high-dimensional Fourier spectrum; the time-domain recurrent neural network is configured to determine a second high-dimensional Fourier spectrum according to the feature data of each sub-band in the first high-dimensional Fourier spectrum; the frequency-domain recurrent neural network is configured to determine a third high-dimensional Fourier spectrum according to the feature data of each time point in the second high-dimensional Fourier spectrum; the decoder in the first sub-module is configured to decode the third high-dimensional Fourier spectrum and output the denoised Fourier spectrum; the decoder in the second sub-module is configured to decode the third high-dimensional Fourier spectrum and output the denoised complex mask.

[0014] In a possible implementation manner, the method for determining the voice denoising model corresponding to the sound channel by using the voice training set further includes: determining a time-domain loss value between the clean voice and the clean voice sample according to a time-domain loss function; determining a frequency-domain loss value between the fourth Fourier spectrum and the eighth Fourier spectrum according to a frequency-domain loss function, or determining a third loss value between the seventh Fourier spectrum and the eighth Fourier spectrum according to the frequency-domain loss function, where the eighth Fourier spectrum is determined according to the clean voice sample; updating the voice denoising model according to the time-domain loss value and the frequency-domain loss value, or updating the voice denoising model according to the time-domain loss value and the third loss value.

[0015] In a second aspect, the present application provides a voice enhancement method.

[0016] The method includes: obtaining a voice denoising model corresponding to a sound channel; using the voice denoising model to perform denoising processing on the target voice transmitted by the sound channel to obtain a clean voice.

[0017] In a third aspect, the present application provides a training device for a voice denoising model.

[0018] The training device includes: an acquisition module for acquiring a voice training set corresponding to a sound channel; the voice training set includes a plurality of noisy voice samples and a plurality of clean voice samples, and the plurality of noisy voice samples and the plurality of clean voice samples correspond one by one; a training module for determining a voice noise reduction model corresponding to the sound channel by using the voice training set; the voice noise reduction model includes: an analysis filter, a first enhancement module, a second enhancement module, and a synthesis filter module;

[0019] Among them, the training module is specifically configured to: convert an input noisy voice sample into a first Fourier spectrum by using the analysis filter; when the signal-to-noise ratio of the sound channel is less than a preset value, the first enhancement module outputs a second Fourier spectrum based on the first Fourier spectrum, splices the first Fourier spectrum and the second Fourier spectrum into a third Fourier spectrum, the second enhancement module outputs a first complex mask based on the third Fourier spectrum, converts the first complex mask into a fourth Fourier spectrum, and the synthesis filter module converts the fourth Fourier spectrum into a clean voice; when the signal-to-noise ratio of the sound channel is not less than the preset value, the second enhancement module outputs a second complex mask based on the first Fourier spectrum, converts the second complex mask into a fifth Fourier spectrum, splices the first Fourier spectrum and the fifth Fourier spectrum into a sixth Fourier spectrum, the first enhancement module outputs a seventh Fourier spectrum based on the sixth Fourier spectrum, and the synthesis filter module converts the seventh Fourier spectrum into a clean voice; update the voice noise reduction model according to the clean voice and the clean voice sample corresponding to the input noisy voice sample.

[0020] In a possible implementation manner, the first enhancement module includes a plurality of first sub-modules, and the second enhancement module includes a plurality of second sub-modules; wherein, the number of the first sub-modules and the number of the second sub-modules are determined according to the signal-to-noise ratio of the sound channel; the first sub-module is configured to perform noise reduction processing on an input Fourier spectrum and output a noise-reduced Fourier spectrum; the second sub-module is configured to perform noise reduction processing on an input Fourier spectrum and output a noise-reduced complex mask;

[0021] Among them, when the signal-to-noise ratio of the sound channel is less than a preset value, the Fourier spectrum input to the nth first sub-module is determined according to the Fourier spectrum output by the analysis filter and the Fourier spectrum output by the (n-i)th first sub-module, and the Fourier spectrum input to the mth second sub-module is determined according to the Fourier spectrum output by the analysis filter, the Fourier spectrum output by each first sub-module, and the Fourier spectrum corresponding to the complex mask output by the (m-j)th second sub-module;

[0022] When the signal-to-noise ratio of the sound channel is not less than the preset value, the Fourier spectrum input to the m-th second sub-module is determined according to the Fourier spectrum output by the analysis filter and the Fourier spectrum output by the (m - j)-th second sub-module; the Fourier spectrum input to the n-th first sub-module is determined according to the Fourier spectrum output by the analysis filter, the Fourier spectrum corresponding to the complex mask output by each second sub-module, and the Fourier spectrum output by the (n - i)-th first sub-module;

[0023] n ∈ [1, N], where N is the number of first sub-modules, i ∈ [1, n], m ∈ [1, M], where M is the number of second sub-modules, and j ∈ [1, m].

[0024] In a possible implementation manner, both the first sub-module and the second sub-module include: an encoder, a time-domain recurrent neural network, a frequency-domain recurrent neural network, and a decoder; wherein, the encoder is configured to encode the input Fourier spectrum to obtain a first high-dimensional Fourier spectrum; the time-domain recurrent neural network is configured to determine a second high-dimensional Fourier spectrum according to the feature data of each sub-band in the first high-dimensional Fourier spectrum; the frequency-domain recurrent neural network is configured to determine a third high-dimensional Fourier spectrum according to the feature data of each time point in the second high-dimensional Fourier spectrum; the decoder in the first sub-module is configured to decode the third high-dimensional Fourier spectrum and output the denoised Fourier spectrum;

[0025] The decoder in the second sub-module is configured to decode the third high-dimensional Fourier spectrum and output the denoised complex mask.

[0026] In a possible implementation manner, the training module is further configured to:

[0027] Determine the time-domain loss value between the clean speech and the clean speech sample according to the time-domain loss function;

[0028] Determine the frequency-domain loss value between the fourth Fourier spectrum and the eighth Fourier spectrum according to the frequency-domain loss function, or determine the third loss value between the seventh Fourier spectrum and the eighth Fourier spectrum according to the frequency-domain loss function, where the eighth Fourier spectrum is determined according to the clean speech sample; update the speech denoising model according to the time-domain loss value and the frequency-domain loss value, or update the speech denoising model according to the time-domain loss value and the third loss value.

[0029] Fourthly, the present application provides a voice enhancement device.

[0030] The voice enhancement device includes: an acquisition module for acquiring a voice noise reduction model corresponding to a voice channel; and a processing module for performing noise reduction processing on target voice transmitted by the voice channel by using the voice noise reduction model to obtain clean voice.

[0031] In a fifth aspect, the present application provides a computing device. The computing device includes: a processor and a memory, where the processor is configured to execute a computer program stored in the memory to execute the training method in the foregoing first aspect and its optional embodiments, or execute the voice enhancement method in the foregoing second aspect.

[0032] In a sixth aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium includes instructions that, when run on a computer, cause the computer to execute the training method in the foregoing first aspect and its optional embodiments, or execute the voice enhancement method in the foregoing second aspect.

[0033] In a seventh aspect, the present application provides a computer program product. The computer program product includes program code that, when the computer runs the computer program product, causes the computer to execute the training method in the foregoing first aspect and its optional embodiments, or execute the voice enhancement method in the foregoing second aspect.

[0034] Any of the foregoing provided devices, computer storage media, or computer program products are all used to execute the methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding solutions in the corresponding methods provided above, which will not be elaborated here. Description of the Drawings

[0035] Figure 1 is a schematic structural diagram of a multi-stage voice enhancement model provided by an embodiment of the present application;

[0036] Figure 2 is a schematic structural diagram of a noise reduction module in a multi-stage voice enhancement model provided by an embodiment of the present application;

[0037] Figure 3 is a schematic structural diagram of a sub-module in a multi-stage voice enhancement model provided by an embodiment of the present application;

[0038] Figure 4 is a flowchart of a method for training a voice enhancement model provided by an embodiment of the present application;

[0039] Figure 5 is a schematic structural diagram of a training device for a voice enhancement model provided by an embodiment of the present application;

[0040] Figure 6 is a flowchart of a voice enhancement method provided by an embodiment of the present application;

[0041] Figure 7 It is a schematic structural diagram of a voice enhancement device provided by an embodiment of the present application;

[0042] Figure 8 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0043] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0044] In the description of the embodiments of the present application, words such as "exemplary", "for example", or "for instance" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for instance" is intended to present relevant concepts in a specific manner.

[0045] In the description of the embodiments of the present application, the term "and / or" is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, B exists alone, and A and B exist simultaneously. In addition, unless otherwise specified, the meaning of the term "plural" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals.

[0046] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise particularly emphasized in other ways.

[0047] In the application scenarios of voice, the commonly used voice enhancement methods usually include spectral subtraction, Wiener filtering, and statistics-based methods. Due to the powerful capabilities of deep neural networks, voice enhancement methods based on deep neural networks have also gradually emerged.

[0048] For voice enhancement methods based on deep neural networks, from the time-frequency domain perspective, such methods can be divided into two categories: mask-based and mapping-based. For the former, such as the ideal binary mask or the ideal ratio mask, the energy distribution relationship between the clean and noise components is analyzed. For the latter, the log power spectrum or the magnitude spectrum is used as the mapping target.

[0049] In practical applications, when a speech enhancement method based on a deep neural network is applied to a terminal device (such as electronic devices like mobile phones and headphones), due to the different computing capabilities of the terminal devices, the final noise reduction effect is also different.

[0050] Figure 1 It is a schematic structural diagram of a speech enhancement model for multi-stage noise reduction provided by an embodiment of the present application.

[0051] As Figure 1 shown, the speech enhancement model 100 includes: an analysis filter 101, a noise reduction module 102, and a synthesis filter 103.

[0052] The analysis filter 101 is used to process the input noisy speech x and convert it into a Fourier spectrum X. Specifically, the analysis filter 101 uses the short-time Fourier transform with a frame length of 512 and a frame shift of 128. The noisy speech is represented as x ∈ R^(1×L), and the output is the Fourier spectrum X ∈ R^(2×F×T), where L is the number of speech sampling points, F is the number of Fourier frequency points 256, and T is the number of frames.

[0053] The noise reduction module 102 is used to perform noise reduction processing on the Fourier spectrum output by the analysis filter 101 and output a Fourier spectrum or a complex mask. Among them, the noise reduction module 102 can be configured according to the signal-to-noise ratio of the channel transmitting the noisy speech to output a Fourier spectrum or a complex mask. For example, when the signal-to-noise ratio of the channel is less than a preset value, the noise reduction module 102 is configured to output a Fourier spectrum. For example, when the signal-to-noise ratio of the channel is not less than the preset value, the noise reduction module 102 is configured to output a complex mask to avoid losing information in the speech in an environment with a low signal-to-noise ratio.

[0054] In one example, the noise reduction module 102 may include a first enhancement module 1021 that outputs a Fourier spectrum and a second enhancement module 1022 that outputs a complex mask. Among them, the front-back relationship of the first enhancement module 1021 and the second enhancement module 1022 in processing data can be determined according to the signal-to-noise ratio of the channel.

[0055] Specifically, when the signal-to-noise ratio of the channel is less than the preset value, the first enhancement module 1021 processes the Fourier spectrum output by the analysis filter 101 and outputs an enhanced Fourier spectrum. The second enhancement module 1022 performs noise reduction processing based on the Fourier spectrum output by the analysis filter 101 and the Fourier spectrum output by the first enhancement module 1021 and outputs a complex mask.

[0056] Specifically, when the signal-to-noise ratio of the sound channel is not less than a preset value, the second enhancement module 1022 performs noise reduction processing based on the Fourier spectrum output by the analysis filter 101 and outputs a complex mask. The first enhancement module 1021 performs noise reduction processing based on the Fourier spectrum output by the analysis filter 101 and the Fourier spectrum corresponding to the complex mask output by the second enhancement module 1022, and outputs an enhanced Fourier spectrum. When the signal-to-noise ratio is low, first using the second enhancement module to process the Fourier spectrum and output a complex mask can avoid losing information in the spectrum and improve the quality of the noise reduction processing.

[0057] In one example, the first enhancement module 1021 and the second enhancement module 1022 can be configured with multiple sub-modules. As Figure 2 shown in the schematic diagram of the noise reduction module 102 in the case where the signal-to-noise ratio is less than the preset value, the first enhancement module 1021 can include N first sub-modules arranged in sequence, and the second enhancement module 1022 can include M second sub-modules arranged in sequence.

[0058] Each first sub-module processes the input Fourier spectrum in sequence. Among them, the Fourier spectrum input to each first sub-module is obtained by splicing the Fourier spectrum output by the analysis filter 101, the Fourier spectra output by each previous first sub-module, and / or the Fourier spectra corresponding to the complex masks output by each previous second sub-module.

[0059] Similarly, the Fourier spectrum input to each second sub-module is obtained by splicing the Fourier spectrum output by the analysis filter 101, the Fourier spectra corresponding to the complex masks output by each previous second sub-module, and / or the Fourier spectra output by each previous first sub-module.

[0060] Among them, splicing refers to splicing multiple Fourier spectra in the channel dimension. For example, as Figure 2 shown in the schematic diagram of the noise reduction module 102 in the case where the signal-to-noise ratio is less than the preset value, the Fourier spectrum input to the second first sub-module can be obtained by splicing the Fourier spectrum output by the analysis filter 101 and the Fourier spectrum output by the first first sub-module in the channel dimension. For another example, as Figure 2 shown, the Fourier spectrum input to the second second sub-module can be obtained by splicing the Fourier spectrum output by the analysis filter 101, the Fourier spectra output by the first first sub-module to the Nth first sub-module, and the Fourier spectrum corresponding to the complex mask output by the first second sub-module in the channel dimension. Figure 2 The process of splicing and the process of converting the complex mask into a spectrum are not shown in

[0061] Specifically, the number of the first sub-modules and the number of the second sub-modules can be determined according to the computing power of the terminal device. For example, the corresponding relationship between the main frequency of the processor in the terminal device and the number of sub-modules can be preset in advance, and the number of sub-modules of the first enhancement module 1021 and the number of sub-modules of the second enhancement module 1022 can be determined according to this corresponding relationship.

[0062] Specifically, the first sub-module and the second sub-module can adopt the same structure. As Figure 3 shown, the first sub-module and the second sub-module may include: an encoder 301, a time-domain recurrent neural network 302, a frequency-domain recurrent neural network 303, and a decoder 304.

[0063] The encoder 301 is used to encode the input Fourier spectrum to obtain a first high-dimensional Fourier spectrum. Specifically, the encoder 301 can be constructed by a convolutional neural network, and specifically may include three layers of complex convolutional layers. The number of channels of each convolutional layer can be designed to be 64, the size of the convolutional kernel of each channel is (5, 2), and the stride of each convolutional layer is (2, 1), (2, 1), and (1, 1). The Fourier spectrum input to the encoder 301 can be expressed as H ∈ R^(C×F'×T), where C represents the number of channels, F' represents the number of Fourier frequency points after downsampling, and F' = 256 / 4 = 64.

[0064] The time-domain recurrent neural network 302 models along the time axis of the first high-dimensional Fourier spectrum. The time-domain recurrent neural network 302 is used to determine a second high-dimensional Fourier spectrum according to the feature data of each sub-band in the first high-dimensional Fourier spectrum. The feature data of each sub-band can be expressed as H 1,f ∈ R^(C×T), f ∈ [1, F'].

[0065] The frequency-domain recurrent neural network 303 models along the frequency axis of the second high-dimensional Fourier spectrum. The frequency-domain recurrent neural network 303 is used to determine a third high-dimensional Fourier spectrum according to the feature data of each time point in the second high-dimensional Fourier spectrum. The feature data of each time point can be expressed as H 1,t ∈ R^(C×F'), t ∈ [1, T].

[0066] The decoder 304 in the first sub-module is used to decode the third high-dimensional Fourier spectrum and output the denoised Fourier spectrum. The decoder 304 in the second sub-module is used to decode the third high-dimensional Fourier spectrum and output the denoised complex mask. Specifically, the decoder 304 in each sub-module adopts a convolutional neural network and includes three layers of complex transposed convolutional layers. The third high-dimensional Fourier spectrum input to the decoder 304 can be expressed as H' ∈ R^(C×F'×T), the output Fourier spectrum is expressed as X ∈ R^(2×F×T), and the output complex mask M ∈ R^(2×F×T).

[0067] The synthesis filter 103 is used to perform noise reduction processing on the Fourier spectrum output by the noise reduction module 102 or the Fourier spectrum corresponding to the complex mask, so as to obtain the clean speech corresponding to the noisy speech. The synthesis filter 103 uses the short-time inverse Fourier transform, with a frame length of 512 and a frame shift of 128. The input of the synthesis filter 103 is the enhanced Fourier spectrum X∈R^(2×F×T), and the output is the enhanced where L is the number of speech sampling points, F is the number of Fourier frequency points 256, and T is the number of frames.

[0068] Figure 4 This is a method for training a speech enhancement model provided by an embodiment of the present application. This method is used to train Figure 1 the speech enhancement model 100 shown. As Figure 4 shown, this method includes the following steps S401 - step S402.

[0069] In step S401, a speech training set corresponding to the vocal tract is obtained.

[0070] The speech training set includes a plurality of noisy speech samples and a plurality of clean speech samples, and the plurality of noisy speech samples and the plurality of clean speech samples correspond one by one.

[0071] In step S402, a speech enhancement model corresponding to the vocal tract is determined by using the speech training set. Among them, for the specific structure of the speech enhancement model, please refer to the structure shown in Figure 1 which will not be elaborated here. Among them, the processing order of each sub-module in the speech enhancement model for the spectrum can be determined according to the signal-to-noise ratio of the vocal tract, and the number of each sub-module in the speech enhancement model can be determined according to the computing power of the device applying the speech enhancement model.

[0072] Specifically, step S402 for determining the speech enhancement model corresponding to the vocal tract by using the speech training set specifically includes:

[0073] Step S4021. Input the noisy speech into the speech enhancement model to obtain the clean speech output by the speech enhancement model.

[0074] Among them, the analysis filter converts the input noisy speech sample into a first Fourier spectrum;

[0075] When the signal-to-noise ratio of the vocal tract is less than a preset value, the first enhancement module outputs a second Fourier spectrum based on the first Fourier spectrum, splices the first Fourier spectrum and the second Fourier spectrum into a third Fourier spectrum, the second enhancement module outputs a first complex mask based on the third Fourier spectrum, and converts the first complex mask into a fourth Fourier spectrum;

[0076] When the signal-to-noise ratio of the sound channel is not less than the preset value, the second enhancement module outputs a second complex mask based on the first Fourier spectrum, converts the second complex mask into a fifth Fourier spectrum, and splices the first Fourier spectrum and the fifth Fourier spectrum into a sixth Fourier spectrum. The first enhancement module outputs a seventh Fourier spectrum based on the sixth Fourier spectrum;

[0077] The synthesis filter module converts the fourth Fourier spectrum or the seventh Fourier spectrum into clean speech;

[0078] Step S4022. Determine the time-domain loss value between the clean speech and the clean speech sample according to the time-domain loss function. When the signal-to-noise ratio of the sound channel is less than the preset value, determine the frequency-domain loss value between the fourth Fourier spectrum and the eighth Fourier spectrum according to the frequency-domain loss function. When the signal-to-noise ratio of the sound channel is not less than the preset value, determine the frequency-domain loss value between the seventh Fourier spectrum and the eighth Fourier spectrum according to the frequency-domain loss function, where the eighth Fourier spectrum is determined according to the clean speech sample

[0079] Step S4023. Update the speech noise reduction model according to the time-domain loss value and the frequency-domain loss value. Updating the speech noise reduction model includes updating the network parameters of each first sub-module and each second sub-module in the speech noise reduction model.

[0080] Specifically, the following formula can be used to determine the first loss value L according to the time-domain loss value and the frequency-domain loss value, and the speech noise reduction model is updated by using the gradient descent method based on the first loss value L.

[0081] L = L_audio + L_spectral

[0082] Where is the clean speech, y is the clean speech sample, and are the real part and the imaginary part of the spectrum corresponding to the clean speech respectively, |Y r | and |Y i | are the real part and the imaginary part of the spectrum corresponding to the clean speech sample respectively.

[0083] Figure 5 A training device for a speech noise reduction model provided by an embodiment of the present application.

[0084] As Figure 5 shown, the training device 500 includes:

[0085] An acquisition module 501, configured to acquire a speech training set corresponding to a sound channel; the speech training set includes a plurality of noisy speech samples and a plurality of clean speech samples, and the plurality of noisy speech samples and the plurality of clean speech samples correspond one by one;

[0086] A training module 502 is configured to determine a voice noise reduction model corresponding to the sound channel by using the voice training set; the voice noise reduction model includes: an analysis filter, a first enhancement module, a second enhancement module, and a synthesis filter module.

[0087] Specifically, the training module 502 is configured to:

[0088] Convert an input noisy voice sample into a first Fourier spectrum by using the analysis filter;

[0089] When the signal-to-noise ratio of the sound channel is less than a preset value, the first enhancement module outputs a second Fourier spectrum based on the first Fourier spectrum, splices the first Fourier spectrum and the second Fourier spectrum into a third Fourier spectrum, the second enhancement module outputs a first complex mask based on the third Fourier spectrum, converts the first complex mask into a fourth Fourier spectrum, and the synthesis filter module converts the fourth Fourier spectrum into clean voice;

[0090] When the signal-to-noise ratio of the sound channel is not less than the preset value, the second enhancement module outputs a second complex mask based on the first Fourier spectrum, converts the second complex mask into a fifth Fourier spectrum, splices the first Fourier spectrum and the fifth Fourier spectrum into a sixth Fourier spectrum, the first enhancement module outputs a seventh Fourier spectrum based on the sixth Fourier spectrum, and the synthesis filter module converts the seventh Fourier spectrum into clean voice;

[0091] Update the voice noise reduction model according to the clean voice and the clean voice sample corresponding to the input noisy voice sample.

[0092] Figure 6 This is a voice enhancement method provided by an embodiment of the present application, which is applied to a terminal device.

[0093] As Figure 6 shown, the method includes the following steps S601 - step S602.

[0094] In step S601, obtain a voice noise reduction model corresponding to the sound channel. Among them, the voice noise reduction model corresponding to the sound channel can be trained and obtained by using the Figure 4 method described above, which will not be elaborated here.

[0095] In step S602, perform noise reduction processing on the target voice transmitted by the sound channel by using the voice noise reduction model to obtain clean voice.

[0096] Figure 7 This is a schematic structural diagram of a voice enhancement device provided by an embodiment of the present application.

[0097] As Figure 7 shown, the voice enhancement device 700 includes:

[0098] An acquisition module 701, configured to acquire a voice noise reduction model corresponding to a sound channel;

[0099] A processing module 702, configured to perform noise reduction processing on the target voice transmitted by the sound channel by using the voice noise reduction model to obtain clean voice.

[0100] Figure 8 FIG. is a schematic hardware structure diagram of a computing device 800 provided in an embodiment of the present application.

[0101] The computing device 800 may be the above-mentioned model training device or the above-mentioned terminal device. Refer to Figure 8 , the computing device 800 includes a processor 801, a memory 802, a communication interface 803, and a bus 804. The processor 801, the memory 802, and the communication interface 803 are connected to each other through the bus 804. The processor 801, the memory 802, and the communication interface 803 may also be connected in other connection manners except the bus 604.

[0102] Among them, the memory 802 may be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, optical memory, hard disk, etc.

[0103] Among them, the processor 801 may be a general-purpose processor, and the general-purpose processor may be a processor that executes specific steps and / or operations by reading and executing the content stored in a memory (such as the memory 802). For example, the general-purpose processor may be a central processing unit (CPU). The processor 801 may include at least one circuit to execute Figure 4 or Figure 6 all or part of the steps of the method provided in the embodiment shown.

[0104] Among them, the communication interface 803 includes interfaces such as input / output (I / O) interfaces, physical interfaces, and logical interfaces for implementing the interconnection of components inside the computing device 800, as well as interfaces for implementing the interconnection between the computing device 800 and other devices (such as other computing devices or terminal devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.

[0105] Among them, the bus 804 can be of any type and is a communication bus for implementing the interconnection of the processor 801, the memory 802, and the communication interface 803, such as a system bus.

[0106] The above components can be separately provided on independent chips, or at least partially or entirely provided on the same chip. Whether to separately provide each component on different chips or integrate them on one or more chips often depends on the needs of product design. The embodiments of the present application do not limit the specific implementation forms of the above components.

[0107] Figure 8 The illustrated computing device 800 is merely exemplary. During implementation, the computing device 800 may further include other components, which will not be listed one by one herein.

[0108] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0109] It should be understood that the various numerical numbers involved in the embodiments of the present application are only for convenience of description and are not used to limit the scope of the embodiments of the present application. It should be understood that in the embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the sequence of execution, and the execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0110] The specific embodiments described above further elaborate on the purpose, technical solution and beneficial effects of the present application. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present application shall be included in the protection scope of the present application.

Claims

1. A training method for a voice noise reduction model, characterized in that The method includes: Obtaining a speech training set corresponding to a sound channel; the speech training set includes a plurality of noisy speech samples and a plurality of clean speech samples, and the plurality of noisy speech samples and the plurality of clean speech samples correspond to each other one by one; Determining a speech noise reduction model corresponding to the sound channel by using the speech training set; the speech noise reduction model includes: an analysis filter, a first enhancement module, a second enhancement module, and a synthesis filter module; Wherein, the determining the speech noise reduction model corresponding to the sound channel by using the speech training set includes: The analysis filter converts the input noisy speech sample into a first Fourier spectrum; When the signal-to-noise ratio of the sound channel is less than a preset value, the first enhancement module outputs a second Fourier spectrum based on the first Fourier spectrum, splices the first Fourier spectrum and the second Fourier spectrum into a third Fourier spectrum, the second enhancement module outputs a first complex mask based on the third Fourier spectrum, converts the first complex mask into a fourth Fourier spectrum, and the synthesis filter module converts the fourth Fourier spectrum into clean speech; When the signal-to-noise ratio of the sound channel is not less than the preset value, the second enhancement module outputs a second complex mask based on the first Fourier spectrum, converts the second complex mask into a fifth Fourier spectrum, splices the first Fourier spectrum and the fifth Fourier spectrum into a sixth Fourier spectrum, the first enhancement module outputs a seventh Fourier spectrum based on the sixth Fourier spectrum, and the synthesis filter module converts the seventh Fourier spectrum into clean speech; Updating the speech noise reduction model according to the clean speech and the clean speech sample corresponding to the input noisy speech sample.

2. The method according to claim 1, characterized in that, The first enhancement module includes a plurality of first sub-modules, and the second enhancement module includes a plurality of second sub-modules; The number of the first sub-modules and the number of the second sub-modules are determined according to the signal-to-noise ratio of the sound channel; The first sub-module is used to perform noise reduction processing on the input Fourier spectrum and output a noise-reduced Fourier spectrum; the second sub-module is used to perform noise reduction processing on the input Fourier spectrum and output a noise-reduced complex mask; Wherein, when the signal-to-noise ratio of the sound channel is less than a preset value, the Fourier spectrum input to the nth first sub-module is determined according to the Fourier spectrum output by the analysis filter and the Fourier spectrum output by the (n - i)th first sub-module, and the Fourier spectrum input to the mth second sub-module is determined according to the Fourier spectrum output by the analysis filter, the Fourier spectrum output by each first sub-module, and the Fourier spectrum corresponding to the complex mask output by the (m - j)th second sub-module; When the signal-to-noise ratio of the sound channel is not less than the preset value, the Fourier spectrum input to the mth second sub-module is determined according to the Fourier spectrum output by the analysis filter and the Fourier spectrum output by the (m - j)th second sub-module, and the Fourier spectrum input to the nth first sub-module is determined according to the Fourier spectrum output by the analysis filter, the Fourier spectrum corresponding to the complex mask output by each second sub-module, and the Fourier spectrum output by the (n - i)th first sub-module; n ∈ [1, N], where N is the number of the first sub - modules, i ∈ [1, n], m ∈ [1, M], where M is the number of the first sub - modules, j ∈ [1, m].

3. The method according to claim 2, characterized in that, Both the first sub - module and the second sub - module include: an encoder, a time - domain recurrent neural network, a frequency - domain recurrent neural network, and a decoder; Among them, the encoder is used to encode the input Fourier spectrum to obtain a first high - dimensional Fourier spectrum; The time - domain recurrent neural network is used to determine a second high - dimensional Fourier spectrum according to the feature data of each sub - band in the first high - dimensional Fourier spectrum; The frequency - domain recurrent neural network is used to determine a third high - dimensional Fourier spectrum according to the feature data of each time point in the second high - dimensional Fourier spectrum; The decoder in the first sub - module is used to decode the third high - dimensional Fourier spectrum and output the denoised Fourier spectrum; The decoder in the second sub - module is used to decode the third high - dimensional Fourier spectrum and output the denoised complex mask; 4. The method according to claim 1, characterized in that The step of determining the voice denoising model corresponding to the sound channel using the voice training set further includes: Determining a time - domain loss value between the clean voice and the clean voice sample according to a time - domain loss function; Determining a frequency - domain loss value between the fourth Fourier spectrum and the eighth Fourier spectrum according to a frequency - domain loss function, or determining a third loss value between the seventh Fourier spectrum and the eighth Fourier spectrum according to the frequency - domain loss function, where the eighth Fourier spectrum is determined according to the clean voice sample; Updating the voice denoising model according to the time - domain loss value and the frequency - domain loss value, or updating the voice denoising model according to the time - domain loss value and the third loss value.

5. A voice enhancement method, characterized in that, The method includes: Obtaining a voice denoising model corresponding to the sound channel, where the voice denoising model is obtained according to the method described in any one of claims 1 - 4; Using the voice denoising model to perform denoising processing on the target voice transmitted by the sound channel to obtain a clean voice.

6. A training device for a voice noise reduction model, characterized in that, The training device includes: An acquisition module, configured to acquire a voice training set corresponding to the sound channel; the voice training set includes a plurality of noisy voice samples and a plurality of clean voice samples, and the plurality of noisy voice samples and the plurality of clean voice samples are in one - to - one correspondence; A training module, configured to determine a voice denoising model corresponding to the sound channel using the voice training set; the voice denoising model includes: an analysis filter, a first enhancement module, a second enhancement module, and a synthesis filter module; Among them, the training module is specifically configured to: Convert the input noisy voice sample into a first Fourier spectrum using the analysis filter; When the signal - to - noise ratio of the sound channel is less than a preset value, the first enhancement module outputs a second Fourier spectrum based on the first Fourier spectrum, splices the first Fourier spectrum and the second Fourier spectrum into a third Fourier spectrum, the second enhancement module outputs a first complex mask based on the third Fourier spectrum, converts the first complex mask into a fourth Fourier spectrum, and the synthesis filter module converts the fourth Fourier spectrum into a clean voice; When the signal-to-noise ratio of the sound channel is not less than the preset value, the second enhancement module outputs a second complex mask based on the first Fourier spectrum, converts the second complex mask into a fifth Fourier spectrum, splices the first Fourier spectrum and the fifth Fourier spectrum into a sixth Fourier spectrum, the first enhancement module outputs a seventh Fourier spectrum based on the sixth Fourier spectrum, and the synthesis filter module converts the seventh Fourier spectrum into clean speech; Update the speech noise reduction model according to the clean speech and the clean speech sample corresponding to the input noisy speech sample.

7. The training device according to claim 6, characterized in that, The first enhancement module includes a plurality of first sub-modules, and the second enhancement module includes a plurality of second sub-modules; The number of first sub-modules and the number of second sub-modules are determined according to the signal-to-noise ratio of the sound channel; The first sub-module is used to perform noise reduction processing on the input Fourier spectrum and output the noise-reduced Fourier spectrum; The second sub-module is used to perform noise reduction processing on the input Fourier spectrum and output the noise-reduced complex mask; Wherein, when the signal-to-noise ratio of the sound channel is less than the preset value, the Fourier spectrum input to the nth first sub-module is determined according to the Fourier spectrum output by the analysis filter and the Fourier spectrum output by the (n - i)th first sub-module, and the Fourier spectrum input to the mth second sub-module is determined according to the Fourier spectrum output by the analysis filter, the Fourier spectrum output by each first sub-module, and the Fourier spectrum corresponding to the complex mask output by the (m - j)th second sub-module; When the signal-to-noise ratio of the sound channel is not less than the preset value, the Fourier spectrum input to the mth second sub-module is determined according to the Fourier spectrum output by the analysis filter and the Fourier spectrum output by the (m - j)th second sub-module, and the Fourier spectrum input to the nth first sub-module is determined according to the Fourier spectrum output by the analysis filter, the Fourier spectrum corresponding to the complex mask output by each second sub-module, and the Fourier spectrum output by the (n - i)th first sub-module; n ∈ [1, N], N is the number of first sub-modules, i ∈ [1, n], m ∈ [1, M], M is the number of first sub-modules, j ∈ [1, m].

8. The training device according to claim 7, wherein Both the first sub-module and the second sub-module include: an encoder, a time-domain recurrent neural network, a frequency-domain recurrent neural network, and a decoder; Wherein, the encoder is used to encode the input Fourier spectrum to obtain a first high-dimensional Fourier spectrum; The time-domain recurrent neural network is used to determine a second high-dimensional Fourier spectrum according to the feature data of each sub-band in the first high-dimensional Fourier spectrum; The frequency-domain recurrent neural network is used to determine a third high-dimensional Fourier spectrum according to the feature data of each time point in the second high-dimensional Fourier spectrum; The decoder in the first sub-module is used to decode the third high-dimensional Fourier spectrum and output the noise-reduced Fourier spectrum; The decoder in the second sub-module is used to decode the third high-dimensional Fourier spectrum and output the noise-reduced complex mask.

9. The training device according to claim 6, characterized in that The training module is further used for: Determine the time-domain loss value between the clean speech and the clean speech sample according to the time-domain loss function; Determine the frequency-domain loss value between the fourth Fourier spectrum and the eighth Fourier spectrum according to the frequency-domain loss function, or determine the third loss value between the seventh Fourier spectrum and the eighth Fourier spectrum according to the frequency-domain loss function, where the eighth Fourier spectrum is determined according to the clean speech sample; Update the speech noise reduction model according to the time-domain loss value and the frequency-domain loss value, or update the speech noise reduction model according to the time-domain loss value and the third loss value.

10. A voice enhancement device, characterized in that, The speech enhancement device includes: An acquisition module, configured to acquire a speech noise reduction model corresponding to a sound channel, where the speech noise reduction model is obtained according to the method of any one of claims 1 to 4; A processing module, configured to use the speech noise reduction model to perform noise reduction processing on the target speech transmitted by the sound channel to obtain clean speech.

Citation Information

Patent Citations

  • Speech enhancement model training and application method, device and equipment, equipment and storage medium

    CN113436643A

  • Two-dimensional smoothing of post-filter masks

    US20210241783A1