Diffusion speech enhancement system and method using noise level alignment
By aligning the ratio of environmental and Gaussian noise using a new stochastic differential equation, the proposed method improves speech enhancement performance and clarity in noisy conditions, addressing the limitations of existing models.
Patent Information
- Application Number
- PCT/KR2024/013364
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-26
- Filing Date
- 2024-09-04
- Publication Date
- 2026-01-02
AI Technical Summary
Existing diffusion speech enhancement models fail to account for the changing path of Gaussian noise during the noise removal process, leading to inconsistent noise removal performance and potential distortion of speech signals.
A new stochastic differential equation is proposed to maintain a constant ratio between environmental noise and Gaussian noise throughout the diffusion process, using a conditional diffusion forward processing unit and a backward processing unit to align noise levels, thereby improving speech enhancement performance.
The solution results in more uniform noise distribution and reduced performance degradation, enhancing speech clarity and reducing distortion in noisy environments, particularly in video conferencing and voice communication systems.
Smart Images

Figure KR2024013364_02012026_PF_FP_ABST
Abstract
Description
Diffusion speech enhancement system and method utilizing noise level alignment
[0001] The present invention relates to a diffusion speech enhancement system and method utilizing noise level alignment, and more particularly, to a diffusion speech enhancement system and method utilizing noise level alignment, which can improve the performance of diffusion speech enhancement by proposing a new stochastic differential equation that aligns the ratio of the size of the environmental noise remaining in the noise removal process to the size of the Gaussian noise added for the diffusion process so as to be constant.
[0002] The content described in this section merely provides background information for one embodiment of the present invention and does not constitute prior art.
[0003]
[0004] Since the COVID-19 pandemic, demand for remote conferencing platforms such as Zoom, Microsoft Teams, and Google Meet has skyrocketed. Even after the pandemic, demand for these video conferencing systems has persisted. According to Global Market Insights Inc., a company that provides international market research and management consulting, the video conferencing market was worth approximately $25 billion in 2022 and is projected to grow to approximately $95 billion by 2032. The company predicts that while North America has led the growth of the video conferencing market to date, led by companies such as Microsoft and Zoom, the Asia-Pacific region will also drive future growth. In particular, the Asia-Pacific region is expected to grow by approximately 15% annually until 2032. South Korea is also expected to contribute significantly to this market growth, ranking third in the video conferencing market in the Asia-Pacific region by 2024.
[0005]
[0006] Noise cancellation technology plays a crucial role in these video conferencing platforms. Many video conferencing software utilizes technology that emphasizes the user's voice and reduces background noise. This allows users to participate in meetings even in noisy environments and helps them focus better by reducing unnecessary breathing noises. This improves the efficiency of remote work and education, and can also lead to increased productivity and economic benefits.
[0007]
[0008] Furthermore, voice enhancement technologies analyze and adjust the frequency components of speech signals using various algorithms. Voice enhancement includes noise suppression, echo cancellation, voice separation, and filtering techniques for sound quality enhancement. These technologies are becoming more sophisticated with the advancement of machine learning and artificial intelligence, and are being used in an increasing variety of fields. Existing diffusion enhancement models have been implemented by focusing on the magnitude and consistency of environmental noise that changes over time during the learning and inference processes. However, the magnitude change path of Gaussian noise, which changes incidentally during the process of controlling the magnitude change path of environmental noise, has not been considered. Korean Patent Publication Nos. 10-2663669, 10-2518874, and 10-2288051 are disclosed as prior art documents.
[0009]
[0010] The background technology described above is technical information that the inventor possessed for the purpose of deriving the present invention or acquired in the process of deriving the present invention, and cannot necessarily be said to be publicly known technology disclosed to the general public prior to the application for the present invention.
[0011] The present invention has been proposed to solve the above problems of the existing proposed methods, and the purpose of the present invention is to provide a diffusion speech enhancement system and method utilizing noise level alignment, which comprises a conditional diffusion forward processing unit that processes input speech data as speech data containing noise by adding conditional information so that the size of the environmental noise and Gaussian noise included in the speech can always be maintained at a constant ratio over time as a noise level alignment process for speech enhancement, and a conditional diffusion backward processing unit that restores clean speech through processing in a reverse process using the speech data containing noise processed and output by the conditional diffusion forward processing unit and the conditional information as input, thereby proposing a new stochastic differential equation that aligns the ratio of the size of the environmental noise remaining in the noise removal process and the size of the Gaussian noise added for the diffusion process so that the performance of the diffusion speech enhancement can be improved.
[0012]
[0013] In addition, the present invention contributes to improving noise removal performance by performing diffusion processing by setting an appropriate size of Gaussian noise according to an environmental noise level, and provides a diffusion speech enhancement system and method utilizing noise level alignment, which enables the diffusion process to become more precise through noise level alignment, thereby reducing distortion of an existing speech signal in the process of removing noise.
[0014]
[0015] In addition, the present invention provides a diffusion speech enhancement system and method utilizing noise level alignment, which implement a generative model that sets the path of the diffusion forward process and the diffusion backward process through a formula of a stochastic differential equation that sets Gaussian noise appropriate for the size of environmental noise that changes over time, thereby making the noise distribution more uniform than the existing model in the speech enhancement processing through diffusion processing, thereby improving the noise removal performance as a result, and showing less performance degradation than the existing model when the number of backward diffusion operations is reduced.
[0016]
[0017] However, the technical problems to be solved by the present invention are not limited to the technical problems described above, and other technical problems may exist.
[0018] A diffusion speech enhancement system utilizing noise level alignment according to the characteristics of the present invention to achieve the above-mentioned purpose is as follows:
[0019] As a diffusion speech enhancement system utilizing noise level alignment,
[0020] A conditional diffusion forward processing unit that processes the input speech data as speech data containing noise and outputs it by adding conditional information so that the size of the environmental noise and Gaussian noise included in the speech can always be maintained at a constant ratio over time as a process of noise level alignment for speech enhancement; and
[0021] The reverse process of noise level alignment for voice enhancement is characterized by including a conditional diffusion reverse process section that restores clean voice through reverse process processing of voice data containing noise and conditional information processed and output by the conditional diffusion forward process section.
[0022]
[0023] Preferably, the conditional diffusion front process section is
[0024] The path of the diffusion forward process can be set by applying a stochastic differential equation of conditional information so that the size of the environmental noise and Gaussian noise contained in the voice can always be maintained in a constant ratio over time.
[0025]
[0026] More preferably, the conditional diffusion reverse process section is
[0027] Clean speech is restored through reverse processing using noise-containing speech data and conditional information processed and output in the above conditional diffusion forward process unit as input, and a diffusion reverse path can be set according to the path setting of the diffusion forward process that applies the probability differential equation of the conditional information of the above conditional diffusion forward process unit.
[0028]
[0029] A diffusion speech enhancement method utilizing noise level alignment according to the characteristics of the present invention to achieve the above-mentioned purpose is as follows.
[0030] As a diffusion speech enhancement method utilizing noise level alignment,
[0031] (1) A conditional diffusion front end is a process of noise level alignment for voice enhancement, a step of adding conditional information so that the size of the environmental noise and Gaussian noise included in the voice can always be maintained at a constant ratio over time, and processing the input voice data as voice data containing noise and outputting it; and
[0032] (2) The conditional diffusion backward process section is a reverse process of noise level alignment for voice enhancement, and its configuration is characterized by including a step of restoring clean voice through processing of voice data containing noise and conditional information processed and output by the conditional diffusion forward process section as input.
[0033]
[0034] Preferably, the conditional diffusion front process section in step (1) is
[0035] The path of the diffusion forward process can be set by applying a stochastic differential equation of conditional information so that the size of the environmental noise and Gaussian noise contained in the voice can always be maintained in a constant ratio over time.
[0036]
[0037] More preferably, the conditional diffusion reverse process section in step (2) is
[0038] Clean speech is restored through reverse processing using noise-containing speech data and conditional information processed and output in the above conditional diffusion forward process unit as input, and a diffusion reverse path can be set according to the path setting of the diffusion forward process that applies the probability differential equation of the conditional information of the above conditional diffusion forward process unit.
[0039]
[0040] The present invention is characterized in that it is a computer program stored in a computer-readable recording medium for executing a diffusion speech enhancement method utilizing noise level alignment according to the characteristics of the present invention to achieve the above-mentioned purpose on a computer.
[0041] According to the diffusion speech enhancement system and method utilizing noise level alignment proposed in the present invention, the system comprises a conditional diffusion forward processing unit that processes input speech data as speech data containing noise by adding conditional information so that the size of environmental noise and Gaussian noise included in the speech can always be maintained at a constant ratio over time as a process of noise level alignment for speech enhancement, and a conditional diffusion backward processing unit that restores clean speech through processing in a backward process with the speech data containing noise processed and output by the conditional diffusion forward processing unit and the conditional information as inputs, thereby proposing a new stochastic differential equation that aligns the ratio of the size of the remaining environmental noise in the noise removal process to the size of the Gaussian noise added for the diffusion process so that it is constant, thereby improving the performance of diffusion speech enhancement.
[0042]
[0043] In addition, according to the diffusion speech enhancement system and method utilizing noise level alignment of the present invention, by performing diffusion processing by setting an appropriate size of Gaussian noise according to an environmental noise level, it contributes to improving noise removal performance, and since the diffusion process becomes more precise through noise level alignment, it is possible to reduce distortion of an existing speech signal in the process of removing noise.
[0044]
[0045] In addition, according to the diffusion speech enhancement system and method utilizing noise level alignment of the present invention, by implementing a generative model that sets the path of the diffusion forward process and the diffusion backward process through a formula of a stochastic differential equation that sets Gaussian noise appropriate for the size of environmental noise that changes over time, the noise distribution is more uniform than that of the existing model in the speech enhancement processing through diffusion processing, and as a result, the noise removal performance is improved, and when the number of backward diffusion operations is reduced, the performance degradation is less than that of the existing model.
[0046]
[0047] In addition, the various advantageous advantages and effects of the present invention are not limited to the above-described contents, and will be more easily understood in the process of explaining specific embodiments of the present invention.
[0048] FIG. 1 is a diagram illustrating the configuration of a diffusion speech enhancement system utilizing noise level alignment according to an embodiment of the present invention as a functional block.
[0049] FIG. 2 is a diagram illustrating a processing configuration of a conditional diffusion front end of a diffusion speech enhancement system utilizing noise level alignment according to an embodiment of the present invention as a functional block.
[0050] FIG. 3 is a diagram illustrating a processing configuration of a conditional diffusion reverse process unit of a diffusion speech enhancement system utilizing noise level alignment according to an embodiment of the present invention as a functional block.
[0051] FIG. 4 is a diagram illustrating a diffusion speech enhancement model learning process in a diffusion speech enhancement system utilizing noise level alignment according to an embodiment of the present invention.
[0052] FIG. 5 is a diagram illustrating a process of performing diffusion speech enhancement using a learned model in a diffusion speech enhancement system utilizing noise level alignment according to an embodiment of the present invention.
[0053] FIG. 6 is a diagram illustrating a flow of a diffusion speech enhancement method utilizing noise level alignment according to an embodiment of the present invention.
[0054] <Explanation of symbols>
[0055] 100: Diffusion speech enhancement system utilizing noise level alignment according to one embodiment of the present invention
[0056] 110: Conditional diffusion forward process
[0057] 120: Conditional diffusion reverse process
[0058] S110: A conditional diffusion forward process is a process of noise level alignment for voice enhancement, which adds conditional information so that the size of the environmental noise and Gaussian noise included in the voice can always be maintained at a constant ratio over time, and processes the input voice data as voice data containing noise and outputs it.
[0059] S120: A step in which the conditional diffusion backward process is the reverse process of noise level alignment for voice enhancement, and restores clean voice by processing voice data containing noise and conditional information processed and output by the conditional diffusion forward process into the reverse process.
[0060] Below, with reference to the attached drawings, embodiments of the present invention are described in detail so that those skilled in the art can easily implement them. However, the present invention may be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, irrelevant parts have been omitted for clarity of description, and similar reference numerals have been used throughout the specification to indicate similar elements.
[0061]
[0062] Throughout the specification, when a part is said to be "connected" to another part, this includes not only the case where it is "directly connected" but also the case where it is "indirectly connected" with another element in between. Furthermore, when a part is said to "include" a component, this should be understood to mean that, unless specifically stated to the contrary, it may include other components rather than excluding them, and does not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0063]
[0064] The following examples are provided as detailed explanations to aid understanding of the present invention and do not limit the scope of the invention. Therefore, inventions with the same scope and function as the present invention are also within the scope of the present invention.
[0065]
[0066] In addition, each configuration, process, procedure or method included in each embodiment of the present invention may be shared within a scope that is not technically inconsistent with each other.
[0067]
[0068] Generally, speech enhancement technologies improve the quality and intelligibility of speech signals through various algorithms. The primary goal is to reduce noise and enhance speech clarity, thereby enhancing the overall listening experience. These speech enhancement techniques can be broadly categorized into noise suppression and echo cancellation. Traditional noise suppression techniques include spectral subtraction, wavelet transform, and adaptive filtering. Spectral subtraction is a method to restore the original signal by subtracting an estimate of the average noise spectrum from a noisy signal. Wavelet transform decomposes a signal into frequency and time information and removes noise from various frequency components. Adaptive filtering uses a filter that adapts to noise that changes in real time depending on the environment. While these methodologies have the advantage of being computationally efficient and can be applied to speech enhancement in real time even on low-end systems, they also present the challenge of being difficult to apply in various noisy environments.
[0069]
[0070] Furthermore, echo cancellation has traditionally used linear predictive coding (LPC), which predicts future values of a speech signal to eliminate echo. This approach represents the speech signal as a linear predictive model, which is then used to remove echo components. While this method is efficient and simple, making it suitable for real-time processing, it has limitations in removing high-frequency echoes.
[0071]
[0072] To overcome the limitations of these traditional speech enhancement techniques, speech enhancement techniques utilizing deep learning have been developed. AI-powered speech enhancement techniques can be broadly categorized into discriminative and generative models. Discriminative models use convolutional or recurrent neural networks to estimate the spectrum of clean speech from the spectrum of noisy speech, while generative models primarily utilize variable autoencoders or generative adversarial networks. Recent speech enhancement models tend to use structures based on transformers for discriminative models, while those based on diffusion are used for generative models. While transformer models are primarily used in natural language processing, they can also be applied to speech signal processing. The attention structure of the transformer architecture is particularly effective in learning long-term dependencies. Furthermore, speech enhancement techniques utilizing diffusion models, inspired by models originally designed for image generation and transformation, gradually transform noisy speech data through a specific stochastic process, effectively removing noise from the speech signal.
[0073]
[0074] In addition, the diffusion model is a probabilistic model used for data generation and transformation, and its research began due to its high utilization in the field of image generation. Recently, it has also been utilized in various application fields such as voice synthesis and enhancement. The model models the data transformation process as a Markov chain, and consists of a forward process that gradually adds noise to the data and a backward process that removes noise and restores the original data. The forward process gradually adds noise to the existing data x0 to obtain x T It is a process of converting to . The distribution of data at each time step of the discrete-time diffusion model is In the continuous-time diffusion model, the forward process is described by stochastic differential equations. Expressed as . Solve the probability differential equation x t Distribution of can be obtained.
[0075]
[0076] The reverse process is to compute the data x containing noise T It is the process of restoring the original data x0. In the reverse process of the discrete-time diffusion model, the distribution of data is learned using the learned model. It is expressed as follows. In the continuous-time diffusion model, the reverse process is Expressed as
[0077]
[0078] The learning process of the diffusion model is to learn a probability distribution for each time step, and the goal of learning is to minimize the difference between the original data and the restored data after removing noise. In the discrete-time diffusion model, , in the continuous time model We use a loss function such as . Here, is the noise component predicted by the model, is the distribution slope predicted by the model. In this way, the model learns how to restore the original data from noisy data.
[0079]
[0080] Conditional diffusion also incorporates conditional information into the original diffusion model and utilizes this information during the generation process. Conditional information is used at all stages of the diffusion process, allowing data with specific characteristics to be generated. While its basic structure is similar to diffusion, it includes an additional module that encodes the conditional information and incorporates it into the generation process. Conditional diffusion techniques have been used in models that generate images based on specific class labels or text descriptions, and in speech synthesis models that transform timbres based on information about a specific speaker. More recently, it has also been used in speech enhancement models, generating clean speech by specifying noisy speech as a condition.
[0081]
[0082] The forward and backward processes of conditional diffusion are forms of the existing diffusion process that include conditional information. In the forward process, x t is x t-1 In addition, it is also calculated by being influenced by the conditional information y. In the backward process, the model receives the conditional information y as an additional input when estimating and removing noise. When applying conditional diffusion to speech enhancement, the forward process gradually adds noise to the original data x0, which includes not only the Gaussian noise added in the diffusion process but also the gradually increasing environmental noise. x T is a form in which Gaussian noise is added to the speech containing environmental noise. In the reverse process, x containing both environmental noise and Gaussian noise T Starting from , a speech containing only environmental noise is given as conditional information y, and a noise prediction model is used to generate a clean speech by gradually reducing both environmental noise and Gaussian noise.
[0083]
[0084] Furthermore, existing speech noise removal systems utilizing diffusion technology employ continuous-time noise scheduling using stochastic differential equations. "Speech Enhancement with Score-Based Generative Models in the Complex STFT Domain," published in 2022, follows the "variance-divergent stochastic differential equations" proposed in "Score-Based Generative Modeling Through Stochastic Differential Equations," the paper that first introduced continuous-time diffusion technology.
[0085]
[0086] Afterwards, in the paper 'Reducing the Prior Mismatch of Stochastic Differential Equations for Diffusion-based Speech Enhancement' published in 2023, additional performance improvement is shown in the direction of reducing the prior mismatch by changing the path of change in the magnitude of environmental noise with the newly proposed 'Brown bridge stochastic differential equation with exponential diffusion coefficient'.
[0087]
[0088] Thus, existing diffusion enhancement models have been implemented with a focus on the magnitude and consistency of environmental noise, which changes over time during the learning and inference processes. However, the process of controlling the magnitude change path of environmental noise has not taken into account the change path of Gaussian noise, which changes incidentally during the process.
[0089]
[0090] In the present invention, we propose a new stochastic differential equation that arranges the ratio of the size of the remaining environmental noise in the noise removal process and the size of the Gaussian noise added for the diffusion process to be constant, thereby improving the performance of diffusion speech enhancement. The differential equation is When expressed as, If is calculated as It is expressed as , and as f(t) changes, the corresponding g(t) is calculated and can be applied flexibly to various f(t). The solution of the differential equation calculated in this way is x t It can ensure that the ratio of existing voice noise and Gaussian noise remains constant even when time (t) changes.
[0091]
[0092] FIG. 1 is a diagram illustrating the configuration of a diffusion speech enhancement system utilizing noise level alignment according to an embodiment of the present invention as a functional block. As illustrated in FIG. 1, a diffusion speech enhancement system (100) utilizing noise level alignment according to an embodiment of the present invention may be configured to include a conditional diffusion forward processing unit (110) that processes input speech data as speech data containing noise by adding conditional information so that the sizes of environmental noise and Gaussian noise included in the speech can always be maintained at a constant ratio over time as a process of noise level alignment for speech enhancement, and outputs the input speech data as speech data containing noise, and a conditional diffusion backward processing unit (120) that restores a clean speech through a reverse process using the speech data containing noise processed and output by the conditional diffusion forward processing unit (110) and the conditional information as inputs, as a reverse process of noise level alignment for speech enhancement. Hereinafter, a specific configuration of a diffusion speech enhancement system utilizing noise level alignment according to an embodiment of the present invention will be described in detail with reference to the attached drawings.
[0093]
[0094] FIG. 2 is a diagram illustrating a processing configuration of a conditional diffusion forward process unit of a diffusion speech enhancement system using noise level alignment according to an embodiment of the present invention as a functional block, FIG. 3 is a diagram illustrating a processing configuration of a conditional diffusion backward process unit of a diffusion speech enhancement system using noise level alignment according to an embodiment of the present invention as a functional block, FIG. 4 is a diagram illustrating a diffusion speech enhancement model learning process in a diffusion speech enhancement system using noise level alignment according to an embodiment of the present invention, and FIG. 5 is a diagram illustrating a process of performing diffusion speech enhancement with a learned model in a diffusion speech enhancement system using noise level alignment according to an embodiment of the present invention.
[0095]
[0096] The conditional diffusion forward process unit (110) is a process for noise level alignment for voice enhancement. It is a configuration that processes the input voice data as voice data containing noise by adding conditional information so that the size of the environmental noise and Gaussian noise included in the voice can always be maintained at a constant ratio over time and outputs the input voice data. This conditional diffusion forward process unit (110) can set the path of the diffusion forward process by applying a stochastic differential equation of the conditional information so that the size of the environmental noise and Gaussian noise included in the voice can always be maintained at a constant ratio over time.
[0097]
[0098] The conditional diffusion backward process unit (120) is a reverse process of noise level alignment for voice enhancement, and is configured to restore clean voice through reverse process with the conditional information and voice data containing noise processed and output by the conditional diffusion forward process unit (110) as input. The conditional diffusion backward process unit (120) restores clean voice through reverse process with the conditional information and voice data containing noise processed and output by the conditional diffusion forward process unit (110) as input, and the diffusion backward path can be set according to the path setting of the diffusion forward process that applies the probability differential equation of the conditional information of the conditional diffusion forward process unit (110).
[0099]
[0100] This diffusion speech enhancement system (100) can be implemented as a conditional diffusion speech enhancement model, and can be trained by considering that the size ratio of environmental noise and Gaussian noise is constant, so that the noise distribution is more uniform than that of existing models, and as a result, the noise removal performance can be improved. In addition, it can be functioned so that when the number of backward diffusion operations is reduced, the performance degradation is less than that of existing models.
[0101]
[0102] In addition, since the diffusion voice enhancement system (100) of the present invention is a method that only changes the design of the stochastic differential equation without significantly modifying the existing diffusion model structure, if it is a platform where diffusion enhancement was previously used, it can be used immediately by retraining only the stochastic differential equation code in the code for learning the model.
[0103]
[0104] In this way, the diffusion speech enhancement system (100) including the conditional diffusion forward process unit (110) and the conditional diffusion backward process unit (120) performs noise level alignment in the following manner, as illustrated in FIGS. 2 and 3, respectively. First, in the conditional diffusion forward process unit (110), the environmental noise and Gaussian noise included in the speech gradually increase over time (t) in the forward process of the conditional diffusion. At this time, the path of the forward process is set so that the magnitudes of the two noises increasing over time (t) always maintain a constant ratio. After setting the forward process, the backward process suitable for the forward process can be mathematically solved to set both the forward and backward processes of the diffusion. Below is This is an example of use when set to . The formula for the forward process for the set f(t) is Here, x0 is a clean voice, and y is a voice containing noise. is equal to the average of x t The distribution of Follows. is a constant that determines the initial size of Gaussian noise, and since it is a fixed value in the forward and backward processes, when y is fixed, x t The environmental noise and Gaussian noise included in the noise always maintain a constant ratio regardless of time (t).
[0105]
[0106] In addition, we train a diffusion model that fits the proposed formula, The reverse process follows to remove noise from the speech ( Starting from unit-size Gaussian noise, we restore a clean speech x0.
[0107]
[0108] FIG. 4 illustrates a process of learning a diffusion speech enhancement model in a diffusion speech enhancement system utilizing noise level alignment according to an embodiment of the present invention, and FIG. 5 illustrates a process of performing diffusion speech enhancement with a learned model in a diffusion speech enhancement system utilizing noise level alignment according to an embodiment of the present invention. As illustrated in FIGS. 4 and 5, specific examples of a diffusion speech enhancement system (100) utilizing noise level alignment according to an embodiment of the present invention are as follows.
[0109]
[0110] First, as illustrated in Fig. 4, the process of learning diffusion speech enhancement is performed. The input consists of a spectrogram X extracted by applying a short-time Fourier transform to a clean speech, a spectrogram Y extracted by applying a short-time Fourier transform to a speech containing environmental noise, and a scalar t extracted from a uniform distribution between 0 and 1. At this time, X_t, which is one of the inputs of the score model in step ① of Fig. 4, is calculated using a formula for X, Y, and t, and the formula used in the calculation process uses the formula of the proposed conditional information stochastic differential equation to make the size ratio of the environmental noise and Gaussian noise constant regardless of time (t). Here, X_t, Y, and t are input to the score model, and the value s output from the score model is input to the loss function as X_t, X, and t to obtain the loss. The score model can be trained in the direction of minimizing the loss using the gradient descent method through backpropagation from the loss.
[0111]
[0112] In addition, as shown in Fig. 5, the process of performing diffusion speech enhancement through the learned score model starts from Y, which is a spectrogram extracted by applying a short-time Fourier transform to a speech containing environmental noise, and the enhancement is performed. At this time, Gaussian noise is added to Y to calculate X_T. T is generally set to 1, and the size of the Gaussian noise added in the process is an equation calculated from the stochastic differential equation of the proposed conditional information. . X_T is input as the initial input of the reverse sampling process represented by process ② of Fig. 5. In this reverse sampling process, X_t, t, the learned score model, and Y are input to calculate the next step input, X_(t-Δt). In the process of calculation, The reverse process formula proposed by is used, and the reverse process is repeated a total of N times to obtain X_0. N can be freely selected within the range of natural numbers, and Δt is calculated according to N. μ_0 excluding the Gaussian noise part of X_0 becomes the final output. μ_0 is in the form of a spectrogram, and a clean voice can be restored by applying the reverse process of the short-time Fourier transform.
[0113]
[0114] A diffusion speech enhancement system (100) utilizing noise level alignment according to an embodiment of the present invention proposes a methodology that can further develop existing speech noise removal technology, and further improves the performance of speech noise removal technology utilizing diffusion by transforming the stochastic differential equation used in existing diffusion-based image generation into a form suitable for the domain of speech noise removal. Furthermore, it reveals that the design of the stochastic differential equation is also an important factor in designing the diffusion model, and thus suggests a method that can serve as a foundation for future research in the related field. The diffusion speech enhancement model of the present invention is trained by considering that the size ratio of environmental noise and Gaussian noise is constant, so that the noise distribution is more uniform than that of existing models, and as a result, the noise removal performance is improved. In addition, it exhibits a lesser performance degradation than existing models when the number of backward diffusions is reduced. By utilizing these characteristics, in an environment supported by sufficient specifications using a single model, high enhancement performance can be achieved with a sufficient number of diffusions, and in an environment with insufficient computational resources, the number of diffusions can be reduced to maintain computational speed while reducing the degradation of enhancement performance. In particular, since the present invention only modifies the stochastic differential equation design without significantly modifying the model structure, it can be used immediately on platforms where diffusion enhancement was previously used by simply changing the stochastic differential equation code in the model training code. In other words, the performance enhancement of voice enhancement technology can provide users of voice communication services with a better listening experience, reduce user fatigue, and enhance the accuracy of communication.
[0115]
[0116] In addition, the present invention relates to a technology related to diffusion speech enhancement that exhibits excellent performance in noise suppression and echo cancellation, and the technology is used to improve the quality of audio signals containing noise that occur during calls or video conferences, and speech enhancement mainly focuses on improving the clarity of voice signals and reducing noise and distortion to improve the listening experience. This speech enhancement technology can be applied to various fields, and in addition to calls or video conferences, it can be applied to hearing aids and hearing assistance devices to amplify voices and reduce background noises to provide better speech recognition to people with hearing impairments, and it can also remove background noise in in-vehicle call systems and voice control systems to enable clear voice communication and use of voice recognition systems even while driving.
[0117]
[0118] In addition, the present invention contributes to improving noise removal performance by performing diffusion processing by setting an appropriate size of Gaussian noise according to the environmental noise level, and in particular, since the diffusion process becomes more precise through noise level alignment, it can function to reduce distortion of the existing voice signal during the noise removal process. That is, in the existing diffusion enhancement model, there was a problem that the degree to which environmental noise and Gaussian noise increase with time (t) was different, so that in some sections, excessively large Gaussian noise was applied compared to the environmental noise, causing distortion, and in some sections, excessively small Gaussian noise was applied, failing to properly remove the environmental noise. The present invention solves the problem by proposing a formula for setting Gaussian noise appropriate for the size of environmental noise that changes with time (t), and in addition, speech enhancement based on a generative model utilizing diffusion can overcome the fundamental limitations of speech enhancement using a discriminative model. Until recently, most discriminative models, which have been used at the forefront of voice enhancement technology, perform voice enhancement by estimating a 'mask' by inputting a voice containing environmental noise and multiplying it by the input. The output of such mask-based voice enhancement has the advantage of less voice distortion because it is calculated by reusing the voice before the environmental noise is removed. However, since it does not generate a clean voice as is, it has limitations in performance. If a generative model is utilized, voice enhancement can be performed by directly generating a clean voice, thereby increasing the performance limit. Therefore, the present invention contributes to increasing the performance of voice enhancement using such a generative model, and in particular, as the clarity and quality of the voice signal are improved, it can provide a better listening experience to users of voice communication services.In particular, the present invention can be used in any field where conditional diffusion is applied, in which the size of existing noise and Gaussian noise is adjusted by a formula for time (t), such as removing noise from brainwave / sensor data that has a one-dimensional signal form, such as voice, as well as removing environmental noise from voice, and removing noise from images / videos, etc., by extending the fact that spectrograms have a two-dimensional form.
[0119]
[0120] FIG. 6 is a diagram illustrating a flowchart of a diffusion speech enhancement method utilizing noise level alignment according to an embodiment of the present invention. As illustrated in FIG. 6, the diffusion speech enhancement method utilizing noise level alignment according to an embodiment of the present invention can be implemented by including a step (S110) in which a conditional diffusion forward processing unit processes input speech data as speech data containing noise by adding conditional information so that the sizes of environmental noise and Gaussian noise included in the speech can always be maintained at a constant ratio over time as a process of noise level alignment for speech enhancement, and a step (S120) in which a conditional diffusion backward processing unit restores a clean speech by processing the speech data containing noise processed and output by the conditional diffusion forward processing unit and the conditional information in a backward process as a reverse process of noise level alignment for speech enhancement.
[0121]
[0122] In step S110, the conditional diffusion forward process unit (110) processes the input speech data as speech data containing noise by adding conditional information so that the size of the environmental noise and Gaussian noise included in the speech can always be maintained at a constant ratio over time, as a process of noise level alignment for speech enhancement, and outputs the input speech data. The conditional diffusion forward process unit (110) in step S110 can set the path of the diffusion forward process by applying a stochastic differential equation of the conditional information so that the size of the environmental noise and Gaussian noise included in the speech can always be maintained at a constant ratio over time.
[0123]
[0124] In step S120, the conditional diffusion backward process unit (120) restores a clean voice through backward process with the input of the conditional diffusion forward process unit (110) processed and outputted voice data containing noise and conditional information as the reverse process of noise level alignment for voice enhancement. In this step S120, the conditional diffusion backward process unit (120) restores a clean voice through backward process with the input of the conditional diffusion forward process unit (110) processed and outputted voice data containing noise and conditional information, and the diffusion backward path may be set according to the path setting of the diffusion forward process that applies the stochastic differential equation of the conditional information of the conditional diffusion forward process unit (110).
[0125]
[0126] This diffusion speech enhancement system (100) can be implemented as a conditional diffusion speech enhancement model, and can be trained by considering that the size ratio of environmental noise and Gaussian noise is constant, so that the noise distribution is more uniform than that of existing models, and as a result, the noise removal performance can be improved. In addition, it can be functioned so that when the number of backward diffusion operations is reduced, the performance degradation is less than that of existing models.
[0127]
[0128] In addition, since the diffusion voice enhancement system (100) of the present invention is a method that only changes the design of the stochastic differential equation without significantly modifying the existing diffusion model structure, if it is a platform where diffusion enhancement was previously used, it can be used immediately by retraining only the stochastic differential equation code in the code for learning the model.
[0129]
[0130] In this way, the diffusion speech enhancement system (100) including the conditional diffusion forward process unit (110) and the conditional diffusion backward process unit (120) performs noise level alignment in the following manner, as illustrated in FIGS. 2 and 3, respectively. First, in the conditional diffusion forward process unit (110), the environmental noise and Gaussian noise included in the speech gradually increase over time (t) in the forward process of the conditional diffusion. At this time, the path of the forward process is set so that the magnitudes of the two noises increasing over time (t) always maintain a constant ratio. After setting the forward process, the backward process suitable for the forward process can be mathematically solved to set both the forward and backward processes of the diffusion. Below is This is an example of use when set to . The formula for the forward process for the set f(t) is Here, x0 is a clean voice, and y is a voice containing noise. is equal to the average of x t The distribution of Follows. is a constant that determines the initial size of Gaussian noise, and since it is a fixed value in the forward and backward processes, when y is fixed, x t The environmental noise and Gaussian noise included in the noise always maintain a constant ratio regardless of time (t).
[0131]
[0132] In addition, we train a diffusion model that fits the proposed formula, The reverse process follows to remove noise from the speech ( Starting from unit-size Gaussian noise, we restore a clean speech x0.
[0133]
[0134] FIG. 4 illustrates a process of learning a diffusion speech enhancement model in a diffusion speech enhancement system utilizing noise level alignment according to an embodiment of the present invention, and FIG. 5 illustrates a process of performing diffusion speech enhancement with a learned model in a diffusion speech enhancement system utilizing noise level alignment according to an embodiment of the present invention. As illustrated in FIGS. 4 and 5, specific examples of a diffusion speech enhancement system (100) utilizing noise level alignment according to an embodiment of the present invention are as follows.
[0135]
[0136] First, as illustrated in Fig. 4, the process of learning diffusion speech enhancement is performed. The input consists of a spectrogram X extracted by applying a short-time Fourier transform to a clean speech, a spectrogram Y extracted by applying a short-time Fourier transform to a speech containing environmental noise, and a scalar t extracted from a uniform distribution between 0 and 1. At this time, X_t, which is one of the inputs of the score model in step ① of Fig. 4, is calculated using a formula for X, Y, and t, and the formula used in the calculation process uses the formula of the proposed conditional information stochastic differential equation to make the size ratio of the environmental noise and Gaussian noise constant regardless of time (t). Here, X_t, Y, and t are input to the score model, and the value s output from the score model is input to the loss function as X_t, X, and t to obtain the loss. The score model can be trained in the direction of minimizing the loss using the gradient descent method through backpropagation from the loss.
[0137]
[0138] In addition, as shown in Fig. 5, the process of performing diffusion speech enhancement through the learned score model starts from Y, which is a spectrogram extracted by applying a short-time Fourier transform to a speech containing environmental noise, and the enhancement is performed. At this time, Gaussian noise is added to Y to calculate X_T. T is generally set to 1, and the size of the Gaussian noise added in the process is an equation calculated from the stochastic differential equation of the proposed conditional information. . X_T is input as the initial input of the reverse sampling process represented by process ② of Fig. 5. In this reverse sampling process, X_t, t, the learned score model, and Y are input to calculate the next step input, X_(t-Δt). In the process of calculation, The reverse process formula proposed by is used, and the reverse process is repeated a total of N times to obtain X_0. N can be freely selected within the range of natural numbers, and Δt is calculated according to N. μ_0 excluding the Gaussian noise part of X_0 becomes the final output. μ_0 is in the form of a spectrogram, and a clean voice can be restored by applying the reverse process of the short-time Fourier transform.
[0139]
[0140] Embodiments of the present invention can also be implemented in the form of a computer program stored in a computer-readable recording medium to execute a diffusion speech enhancement method utilizing noise level alignment according to an embodiment of the present invention as described above on a computer. Here, the computer-readable medium can be any available medium that can be accessed by a computer, and includes both volatile and nonvolatile media, removable and non-removable media. In addition, the computer-readable medium can include both computer storage media and communication media. Computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, or other transport mechanism, and includes any information delivery media.
[0141]
[0142] As described above, the diffusion speech enhancement system and method utilizing noise level alignment according to an embodiment of the present invention comprises a conditional diffusion forward processing unit that processes input speech data as speech data containing noise by adding conditional information so that the ratio of the sizes of environmental noise and Gaussian noise included in the speech can always be maintained constant over time as a process of noise level alignment for speech enhancement, and a conditional diffusion backward processing unit that restores clean speech through processing in a backward process with the speech data containing noise processed and output by the conditional diffusion forward processing unit and the conditional information as inputs, thereby suggesting a new stochastic differential equation that aligns the ratio of the size of the remaining environmental noise in the noise removal process to the size of the Gaussian noise added for the diffusion process so that it is constant, thereby improving the performance of the diffusion speech enhancement, and in particular, by performing diffusion processing by setting an appropriate size of Gaussian noise according to the environmental noise level, it contributes to improving the noise removal performance, and the diffusion process is further improved through noise level alignment. Since it becomes more precise, it can reduce the distortion of the existing voice signal in the process of removing noise, and also, by implementing a generative model that sets the path of the diffusion forward process and the diffusion backward process through a formula of a stochastic differential equation that sets Gaussian noise appropriate for the size of the environmental noise that changes over time, the noise distribution is more uniform than the existing model in the process of voice enhancement processing through diffusion processing, and as a result, the noise removal performance is improved, and when the number of backward diffusions is reduced, it can show an aspect of less performance degradation than the existing model.
[0143]
[0144] The foregoing description of the present invention is for illustrative purposes only, and those skilled in the art will readily appreciate that the present invention can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single entity may be implemented in a distributed manner, and similarly, components described as distributed may be implemented in a combined manner.
[0145]
[0146] The scope of the present invention is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present invention.
Claims
1. A diffusion voice enhancement system (100) utilizing noise level alignment, A conditional diffusion forward processing unit (110) that processes and outputs input voice data as voice data containing noise by adding conditional information so that the size of the environmental noise and Gaussian noise included in the voice can always be maintained at a constant ratio over time as a process of noise level alignment for voice enhancement; and A diffusion speech enhancement system utilizing noise level alignment, characterized in that it includes a conditional diffusion reverse process unit (120) that restores clean speech through reverse process processing of speech data containing noise and conditional information processed and output by the conditional diffusion forward process unit (110) as inputs in the reverse process of noise level alignment for speech enhancement.
2. In the first paragraph, the conditional diffusion front process unit (110) is A diffusion speech enhancement system utilizing noise level alignment, characterized in that the path of the diffusion forward process is set by applying a stochastic differential equation of conditional information so that the size of environmental noise and Gaussian noise included in the speech can always be maintained in a constant ratio over time.
3. In the second paragraph, the conditional diffusion reverse process unit (120) is A diffusion speech enhancement system utilizing noise level alignment, characterized in that clean speech is restored through reverse processing using noise-containing speech data and conditional information processed and output by the above conditional diffusion forward process unit (110) as input, and a diffusion reverse path is set according to the path setting of the diffusion forward process that applies the probability differential equation of the conditional information of the above conditional diffusion forward process unit (110).
4. A diffusion voice enhancement method utilizing noise level alignment, (1) A conditional diffusion front process unit (110) is a process for noise level alignment for voice enhancement, a step for processing input voice data as voice data containing noise by adding conditional information so that the size of the environmental noise and Gaussian noise included in the voice can always be maintained at a constant ratio over time; and (2) A diffusion speech enhancement method utilizing noise level alignment, characterized in that the conditional diffusion reverse process unit (120) is a reverse process of noise level alignment for speech enhancement, and includes a step of restoring clean speech through reverse process processing of speech data containing noise and conditional information processed and output by the conditional diffusion forward process unit (110) as input.
5. In the fourth paragraph, the conditional diffusion front process section (110) in the step (1) is A diffusion speech enhancement method utilizing noise level alignment, characterized in that the path of the diffusion forward process is set by applying a stochastic differential equation of conditional information so that the size of environmental noise and Gaussian noise included in the speech can always be maintained in a constant ratio over time.
6. In the fifth paragraph, the conditional diffusion reverse process unit (120) in the step (2) is A diffusion speech enhancement method utilizing noise level alignment, characterized in that clean speech is restored through reverse processing using noise-containing speech data and conditional information processed and output by the above conditional diffusion forward process unit (110) as input, and a diffusion reverse path is set according to the path setting of the diffusion forward process that applies the probability differential equation of the conditional information of the above conditional diffusion forward process unit (110).
7. A computer program stored in a computer-readable recording medium for executing a diffusion speech enhancement method utilizing noise level alignment of any one of claims 4 to 6 on a computer.
Citation Information
Patent Citations
Video sound effect generation method and device, computer equipment and storage medium
CN118018800A