Audio enhancement method and electronic device

The diffusion model-based audio enhancement method addresses the challenge of degraded audio quality in diverse environments by processing audio in chunks and adjusting signal strength, enhancing voice clarity and accuracy in real-time across streaming and non-streaming scenarios.

WO2026089258A1PCT designated stage Publication Date: 2026-04-30SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/013069
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-24
Filing Date
2025-08-27
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing audio enhancement models are optimized for non-streaming environments and fail to provide optimal performance in diverse, real-world streaming environments, leading to degraded voice quality and clarity due to factors like distortion and network issues.

Method used

An audio enhancement method using a diffusion model that processes audio in chunks, applying a window function to adjust signal strength in overlapping regions and utilizing a generative artificial intelligence model to improve audio quality in both streaming and non-streaming environments.

Benefits of technology

Enhances audio quality in real-time by removing noise and distortion, improving voice clarity and accuracy in voice recognition, and supporting various environments, including streaming scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025013069_30042026_PF_FP_ABST
    Figure KR2025013069_30042026_PF_FP_ABST
Patent Text Reader

Abstract

An audio enhancement method is provided. The method may comprise the steps of: acquiring a plurality of first chunks by dividing input audio; acquiring first features by converting the first chunks into a frequency domain; acquiring, on the basis of the first features, enhanced second features by using a diffusion model, the diffusion model sequentially processing the first features and, when each feature is processed, processing an extended feature including a region partially overlapping that of the preceding feature; acquiring second chunks by converting each of the second features into a time domain; adjusting the signal strength of an overlapping region between adjacent second chunks by applying a window function to the second chunks; and generating enhanced audio by overlap-and-add of the second chunks.
Need to check novelty before this filing date? Find Prior Art

Description

Audio enhancement methods and electronic devices

[0001] The present disclosure relates to a method and electronic device for enhancing audio using an improved superposition-addition method and a diffusion model.

[0002] Mobile devices are used under demanding conditions where voice quality and clarity can degrade due to various factors, such as different types of distortion, network issues, environmental conditions, and low-quality devices. In such situations, multiple sound quality issues may occur in combination, and the audio may contain various types of distortion. While various models have been proposed for voice enhancement, most are trained assuming non-streaming environments to achieve optimal performance. However, since real-world applications require application across diverse environments, including streaming environments, real-time audio enhancement technology for streaming environments is required.

[0003] According to one aspect of the present disclosure, an audio enhancement method may be provided. The method may include: a step of obtaining a plurality of first fragments by dividing input audio; a step of obtaining first features by converting the first fragments into the frequency domain; a step of obtaining enhanced second features using a diffusion model based on the first features, wherein the diffusion model processes the first features sequentially, and when processing each feature, processes an extended feature that includes a region partially overlapping with the previous feature; a step of obtaining second fragments by converting each of the second features into the time domain; a step of adjusting the signal strength of the overlapping region between adjacent second fragments by applying a window function to the second fragments; and a step of generating enhanced audio by overlapping and adding the second fragments.

[0004] According to one aspect of the present disclosure, an electronic device for enhancing audio may be provided. The electronic device may include a communication interface, at least one processor, and a memory for storing instructions. By executing the instructions by the at least one processor, the electronic device may divide input audio to obtain a plurality of first chunks, convert the first chunks into the frequency domain to obtain first features, and obtain enhanced second features using a spreading model based on the first features, wherein the spreading model processes the first features sequentially, and when processing each feature, processes an extended feature that includes a region partially overlapping with the previous feature, and convert the second features into the time domain to obtain second chunks, and apply a window function to the second chunks to adjust the signal strength of the overlapping region between adjacent second chunks, and generate enhanced audio by overlapping-adding the second chunks.

[0005] According to one aspect of the present disclosure, a computer-readable recording medium may be provided having a program recorded thereon for executing any one of the methods described above and below for enhancing electronic device audio.

[0006] FIG. 1 is a drawing for explaining how an electronic device according to one embodiment of the present disclosure improves audio.

[0007] FIG. 2 is a flowchart illustrating the operation of an electronic device according to one embodiment of the present disclosure to enhance audio.

[0008] FIG. 3 is a drawing for explaining a diffusion model according to one embodiment of the present disclosure.

[0009] FIG. 4 is a flowchart illustrating an audio enhancement method using a diffusion model according to one embodiment of the present disclosure.

[0010] FIG. 5 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to process overlapping pieces.

[0011] FIG. 6a is a flowchart illustrating an audio enhancement method using a diffusion model according to one embodiment of the present disclosure.

[0012] Figure 6b shows the algorithm described in Figure 6a in pseudocode.

[0013] FIG. 7 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to process overlapping pieces.

[0014] FIG. 8a is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to extend a feature to create features having an overlapping region.

[0015] FIG. 8b is a diagram illustrating the operation of an electronic device according to one embodiment to extend a feature to create features having an overlapping area.

[0016] FIG. 9a is a drawing illustrating a policy applied by an electronic device according to one embodiment of the present disclosure when extending a feature.

[0017] FIG. 9b is a drawing illustrating a policy applied by an electronic device according to one embodiment of the present disclosure when extending a feature.

[0018] FIG. 9c is a drawing illustrating a policy applied by an electronic device according to one embodiment of the present disclosure when extending a feature.

[0019] FIG. 9d is a drawing illustrating a policy applied by an electronic device according to one embodiment of the present disclosure when extending a feature.

[0020] FIG. 10 is a drawing for explaining the audio enhancement results of an electronic device according to one embodiment of the present disclosure.

[0021] FIG. 11 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure.

[0022] The terms used in this specification will be briefly explained, and the present disclosure will be described in detail. In the present disclosure, the expression "at least one of a, b, or c" may refer to "a," "b," "c," "a and b," "a and c," "b and c," "all of a, b, and c," or variations thereof.

[0023] The terms used in this disclosure have been selected to be as widely used and general as possible, taking into account their functions within this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been selected at the applicant's discretion, and in such cases, their meanings will be described in detail in the relevant explanatory sections. Therefore, terms used in this disclosure should be defined not merely by their names, but based on their meanings and the overall content of this disclosure.

[0024] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as generally understood by those skilled in the art as described in this specification. Additionally, terms including ordinal numbers, such as "first" or "second," used in this specification may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another.

[0025] When a part of a specification is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, terms such as "part" or "module" as used in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware or software, or as a combination of hardware and software.

[0026] Embodiments of the present disclosure are described below with reference to the attached drawings so that those skilled in the art can easily implement the invention. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present disclosure in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.

[0027] The present disclosure will be described below with reference to the attached drawings.

[0028] FIG. 1 is a drawing for explaining how an electronic device according to one embodiment of the present disclosure improves audio.

[0029] In one embodiment, an electronic device can enhance audio using a diffusion model (100). The diffusion model may be a generative artificial intelligence model that utilizes a diffusion process. The diffusion model may be trained through a forward diffusion process that progressively adds noise and a backward diffusion process that predicts and removes noise, and the trained diffusion model can generate enhanced audio through a backward diffusion process that generates initial noise and predicts and removes noise from the initial noise. In this case, the diffusion model can generate enhanced audio (120) by referencing the input audio (110).

[0030] The audio enhancement method of the present disclosure uses an improved overlap-add method different from a general overlap-add method. An electronic device can divide input audio (110) into a plurality of chunks and perform processing by a diffusion model (100) on a chunk-by-chunk basis. For overlap between adjacent chunks, the electronic device uses features processed in the previous chunk as context information when processing the current chunk. In other words, when the diffusion model (100) processes the current chunk, it loads a portion of the features processed in the previous chunk and processes them together with the features of the current chunk. During the diffusion process using the diffusion model (100), the electronic device can adjust the trade-off between the quality of the enhanced audio and the amount of computation by changing the length of the overlapping area between features in various ways.

[0031] In one embodiment, when audio fragments of input audio (110) are input to a diffusion model (100), the diffusion model (100) outputs enhanced audio fragments. As the electronic device processes some of the features of the previous fragment together with the current fragment, an overlapping area occurs between adjacent fragments, and the electronic device can adjust the signal strength of the overlapping area by applying a window function. The electronic device can generate enhanced audio (120) by applying a window function to the enhanced audio fragments, overlapping them with the audio stream, and adding them.

[0032] In one embodiment, the electronic device may be of various types of devices capable of outputting audio. For example, the electronic device may be implemented as an electronic device of various types and forms, including a speaker. The electronic device may include, but is not limited to, smartphones, tablet PCs, laptop PCs, artificial intelligence speakers, etc. Alternatively, the electronic device may be any type of device capable of outputting audio through a speaker (e.g., a Bluetooth speaker, a TWS (True Wireless Stereo) headset, etc.) connected to the electronic device wirelessly or via a wire. In one embodiment, the electronic device may be a device that performs audio enhancement processing and provides it to a user device. In this case, the electronic device and the user device may be configured in the form of a server-client device.

[0033] In one embodiment, the electronic device generating enhanced audio (120) can be applied to a streaming environment or a non-streaming environment.

[0034] For example, audio enhancement can be applied to situations where a user answers or makes a voice or video call via an electronic device (e.g., a smartphone or TWS headset). In this process, audio from both the transmitting and receiving ends is processed in real time by a diffusion model within the electronic device to improve voice quality and clarity.

[0035] For example, audio enhancement can be applied to voice or video recording. Audio is processed by diffusion models within electronic devices, enabling more accurate audio transcription.

[0036] For example, audio enhancement can be applied to voice assistants. When a user speaks a wake-up command or a task command to a voice assistant, the accuracy of Automatic Speech Recognition (ASR) models can be improved by enhancing audio quality using an in-device diffusion model.

[0037] In addition, electronic devices can provide enhanced audio by improving audio in various environments, such as restoring old audio recorded on low-quality devices, or when the sound quality of audio played on the internet is low or the internet transmission status is poor.

[0038] The specific operations of the electronic device that enhance audio will be described in more detail through the drawings and descriptions below.

[0039] FIG. 2 is a flowchart illustrating the operation of an electronic device according to one embodiment of the present disclosure to enhance audio.

[0040] In operation S210, the electronic device can divide the input audio to obtain a plurality of first chunks.

[0041] In one embodiment, the electronic device may acquire input audio. The input audio is raw audio data and may contain noise (e.g., ambient noise). The input audio may be, for example, speech, but is not limited thereto.

[0042] The electronic device performs audio enhancement processing on the input audio. To process the input audio in real time and stream the enhanced audio, the electronic device can divide the input audio into multiple segments.

[0043] In one embodiment, an audio segment may include one or more frames. A single frame may include multiple samples. A sample may refer to an individual data point representing the sound level (amplitude) at a specific moment. The number of samples may be determined by the sampling rate. For example, if the sampling rate is 48 kHz (48,000 Hz), this means there are 48,000 samples per second. A single frame may include, for example, 128 to 512 samples. In this case, the frame may represent audio of 2.6 to 10.6 ms.

[0044] In one embodiment, the input audio may be divided into a plurality of separated first fragments. In one embodiment, the input audio may be divided into a plurality of first fragments that include overlapping regions.

[0045] In operation S220, the electronic device can convert the first pieces into the frequency domain to obtain the first features.

[0046] In one embodiment, the electronic device may perform a time-frequency domain transformation by applying a Short-Time Fourier Transform (STFT) to an audio fragment. Through a data processing process including the time-frequency domain transformation, the electronic device may extract first features from the first fragments to be used as input data for a diffusion model. The data processing process for obtaining the first features may include various processes for processing the audio data by the diffusion model in addition to the time-frequency domain transformation. For example, the data processing process may include spectrogram generation, normalization, etc., but is not limited thereto.

[0047] In operation S230, the electronic device can obtain enhanced second features using a diffusion model based on the first features.

[0048] In one embodiment, an electronic device can enhance audio using an improved overlap-add method. To this end, the diffusion model processes first features sequentially, and when processing each feature, it can process an extended feature that includes a region partially overlapping with the previous feature.

[0049] The diffusion model referred to in this disclosure means a generative artificial intelligence model that utilizes a diffusion process. The diffusion model is not limited to the architecture of a specific generative artificial intelligence model. In other words, the audio enhancement method of this disclosure is applicable to any architecture of a generative artificial intelligence model that includes a diffusion process. The diffusion model can be trained through a forward diffusion process that progressively adds noise and a backward diffusion process that predicts and removes noise. Then, an audio enhancement operation can be performed using the trained diffusion model. The diffusion model can acquire second features corresponding to the enhanced audio fragment through a backward diffusion process that generates initial noise and predicts and removes noise from the initial noise.

[0050] When the diffusion model processes each first feature, it includes an extended feature that includes a region partially overlapping with the previous feature, so the enhanced second features may also include a region overlapping between adjacent features.

[0051] A detailed explanation of the audio enhancement process using the diffusion model will be described in detail in the subsequent description of the drawings.

[0052] In operation S240, the electronic device can obtain second pieces by converting each of the second features into the time domain.

[0053] In one embodiment, the electronic device may perform a frequency-time domain transformation by performing an Inverse STFT on the second feature. The electronic device may convert the data back into an audio fragment by undergoing an inverse transformation process corresponding to the data processing process (e.g., spectrogram generation, normalization, etc.) performed in operation S220. Since the second feature is a feature corresponding to the enhanced audio output from the diffusion model, the second fragment refers to the enhanced audio fragment.

[0054] Since the diffusion model processes each feature by including a region that partially overlaps with the previous feature, the second features, which are the output of the diffusion model, and the second audio fragments transformed from them may have overlapping regions between adjacent data.

[0055] In operation S250, the electronic device can apply a window function to the second pieces to adjust the signal strength of the overlapping area between adjacent second pieces.

[0056] In one embodiment, the second fragments have overlapping regions. The electronic device may apply a window function to mitigate discontinuities occurring in the overlapping regions of the second fragments. The window function can minimize discontinuous changes that may occur when adding the overlapping regions by adjusting the signal strength of the overlapping regions between adjacent second fragments. The window function may be, for example, a Hanning window, but is not limited thereto. For example, when combining signal fragments, various window functions (e.g., Hamming window, Blackman window, etc.) that can maintain the central part of the signal and smoothly reduce the boundary regions may be applied.

[0057] In operation S260, the electronic device can generate enhanced audio by overlapping and adding the second pieces.

[0058] In one embodiment, the electronic device can generate enhanced audio by overlapping and adding second pieces to which a window function is applied. Enhanced audio may refer to audio from which noise (e.g., ambient noise, etc.) has been removed from raw audio data, which is input data.

[0059] In one embodiment, the electronic device can generate enhanced audio while streaming. In one embodiment, the electronic device may also generate enhanced audio in a non-streaming manner.

[0060] In one embodiment, if the electronic device includes a speaker, the electronic device may output enhanced audio through the speaker. In one embodiment, the electronic device may transmit enhanced audio to an external device (e.g., a speaker or another electronic device including a speaker) connected to the electronic device via wired or wireless connection, so that enhanced audio is output through the external device.

[0061] FIG. 3 is a drawing for explaining a diffusion model according to one embodiment of the present disclosure.

[0062] In the following drawings, audio pieces will be simply referred to as 'chunks' for the sake of convenience of explanation.

[0063] In one embodiment, the diffusion model may be a generative model comprising an audio information generator (300) capable of performing a forward diffusion process and a backward diffusion process. The audio information generator (300) may function to process the forward diffusion process and the backward diffusion process of the diffusion model so that the diffusion model trains and infers audio information.

[0064] The audio information generator (300) may be implemented using a neural network architecture for inferring audio by predicting and removing noise, or through a variation of said neural network architecture. For example, the audio information generator (300) may be implemented based on a U-Net architecture. However, the audio enhancement method of the present disclosure provides compatibility with any diffusion model architectures. In other words, the diffusion model is not limited to the architecture of a specific generative artificial intelligence model. The audio information generator (300) may include an attention module for applying an attention mechanism that merges features corresponding to the input audio (310) into the audio with noise added in stages. For example, the audio information generator (300) may include one or more cross-attention modules.

[0065] In one embodiment, the electronic device can train a diffusion model. During the training process of the diffusion model, the diffusion model can learn audio features through a forward diffusion process that adds noise to the original audio at time steps and a backward diffusion process that restores the original audio by denoising the noise added to the audio at time steps. The diffusion model can be trained using a training dataset consisting of pairs of clean audio without noise and audio with added noise.

[0066] In one embodiment, the electronic device may train a diffusion model using a training dataset. In one embodiment, the electronic device may receive a trained diffusion model from an external device (e.g., a server).

[0067] After the diffusion model is prepared to perform inference through training and performance verification, it can perform the audio enhancement work of the present disclosure.

[0068] In one embodiment, the electronic device may perform an inference task using a trained diffusion model. An inference task refers to improving the input audio (310) to generate enhanced audio (320). The diffusion model may generate enhanced audio (320) based on the input audio (310).

[0069] In the inference process of the diffusion model, the diffusion model may generate initial noise (330). The initial noise may be Gaussian noise sampled from a standard normal distribution N(0,1).

[0070] The diffusion model can perform a back-diffusion process that repeats the prediction and removal of noise from the initial noise (330) at each time step. Through the repeated time steps, the noise is gradually removed, and at the final time step, features representing enhanced audio with all or almost all of the noise removed can be obtained.

[0071] In one embodiment, the input audio (310) is divided into a plurality of fragments (first fragments) and can be converted into input data (first features) of a diffusion model through a feature extraction operation. The feature extraction operation may include time-frequency domain transformation. The first features may be sequentially processed by the audio information generator (300) of the diffusion model to become output data (second features). The output data of the diffusion model is a sequential sequence of second features and may be data including overlapping regions between adjacent second features.

[0072] The second features can be converted into the final result, enhanced audio (320), through a feature restoration operation. The feature restoration operation may include frequency-time domain transformation, and a window function application process may be performed to process overlapping signals for generating enhanced audio (320).

[0073] The audio enhancement process, which is the inference task of the diffusion model, is described further with reference to Fig. 4.

[0074] FIG. 4 is a flowchart illustrating an audio enhancement method using a diffusion model according to one embodiment of the present disclosure.

[0075] In one embodiment, the audio enhancement operation of the electronic device may include an outer loop (400) and an inner loop (410).

[0076] The outer loop (400) represents sequentially processing first features corresponding to different pieces. For example, when input audio is divided into first pieces (piece A, piece B, ...) which are sequential pieces, each of the first features (feature A, feature B, ...) corresponding to the first pieces corresponds to each cycle of the outer loop (400).

[0077] In the outer loop (400), the electronic device identifies whether there is available data, the available data may mean whether there is data remaining to be processed for audio enhancement work.

[0078] The electronic device may collect data from the input audio if available data exists. The collected data may refer to segments of the input audio. To distinguish them from the enhanced segments after audio enhancement processing is completed, the input segments are referred to as the first segments.

[0079] When the first fragment is collected, the electronic device can perform a feature extraction operation to obtain the first feature. The feature extraction operation may include data processing steps to use the first fragment as input data for the diffusion data. For example, the electronic device can obtain data in the frequency domain by performing a time-frequency domain transformation on the first fragment. The time-frequency domain transformation may be an STFT transformation. The electronic device can obtain the first feature by performing additional processing (e.g., normalization, etc.) based on the data in the frequency domain. When the first feature is obtained in the outer loop (400), the electronic device can assign a time step variable of the diffusion model.

[0080] The inner loop (410) represents processing the time steps of the back-diffusion process, which is the inference process of the diffusion model. For example, if the number of time steps of the diffusion model is S, the inner loop (410) is repeated S times. When the iteration of the inner loop (410) ends, the next feature can be processed again through the outer loop (400).

[0081] The electronic device may perform a back-diffusion process to process features using a diffusion model. The back-diffusion process is a process of gradually removing noise from a time step t=S to t=0. One back-diffusion process may be performed at each time step. For example, in the back-diffusion process of the first time step of the inner loop (410), the diffusion model [processes] data (e.g., latent data) z containing initial noise. s It can generate (initial time step: t=S). The diffusion model uses data z s Predicting noise in and z s The data for the next stage can be obtained by removing the noise predicted from. The inner loop (410) is repeated as many times as the number of time steps, and the data at any time step t between t=S and t=0 is z tWhen one inner loop (410) cycle is completed, one second feature, which is an enhanced feature corresponding to one first feature, is obtained.

[0082] When feature processing (inverse diffusion) is performed, the electronic device may perform a feature restoration operation to restore the features back into audio data. The feature restoration operation may be the inverse transformation process of the feature extraction operation and may include a frequency-time domain transformation. The frequency-time domain transformation may be ISTFT. Since the fragment obtained through feature restoration represents enhanced audio, it may be referred to as an enhanced fragment or a second fragment.

[0083] The electronic device may apply a window function to the second piece and add it to the output. Assuming that each sample of the audio contains a timestamp, each sample may be added to the stream of the output audio signal at positions corresponding to each timestamp. In this case, since pieces are added at each time step of the inner loop (410), the electronic device may process the pieces added at each time step to generate the output audio. For example, the electronic device may generate the output audio by applying normalization, weighted average, range adjustment after weighted average, etc., but is not limited thereto.

[0084] In one embodiment, among the processes included in the inner loop (410), such as 'feature processing', 'feature restoration', 'window application', and 'add to output', the remaining operations excluding feature processing may be performed only once. In other words, only the reverse diffusion process of processing features is repeated in the inner loop (410) for the number of time steps, and the {feature restoration-window application-add to output} operation may be performed once on the data obtained at the final time step of the inner loop (410).

[0085] When the number of iterations in the inner loop (410) is completed for the number of time steps, the next sequence of features can be processed through the outer loop (400). That is, after one cycle of the inner loop (410) is completed, the outer loop (400) operates to proceed with the next cycle of the inner loop (410). The next sequence of features can be processed again through the inner loop (410).

[0086] The electronic device uses an improved overlap-add method in the process of generating enhanced audio, which is the output audio. To this end, each feature processed in the operations of the outer loop (400) and the inner loop (410) using a diffusion model may include a region that partially overlaps with the previous feature. This is described further with reference to FIG. 5.

[0087] FIG. 5 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to process overlapping pieces.

[0088] In one embodiment, an electronic device can generate enhanced audio using an improved superposition-addition method. As with general signal processing types, in audio enhancement processing, diffusion models work better when historical data is present. To this end, the electronic device can add a portion of the previous fragment when processing the fragment.

[0089] Typically, the number of time steps in a diffusion process is designed to decrease (e.g., from t=S to t=0), so the processing of the axis representing the diffusion time steps is done in order from bottom to top. In Figure 5, only five time steps of the diffusion process are illustrated as examples for convenience of explanation. However, the number of diffusion time steps is not limited to this example.

[0090] In one embodiment, the electronic device can process data equal to the size of the piece for the initial piece.

[0091] In the initial piece, piece A (510), only the size of the piece of data can be processed. For example, piece A (510) is processed at each time step t=5, 4, 3, 2, 1. Once the diffusion processes for piece A (510) are completed, the next piece in sequence (e.g., piece B (520), piece C (530), ...) can be processed.

[0092] For each piece after the initial piece, the electronic device can process data of an extended size that includes an area overlapping with a part of the previous piece. The extended data may be a part of the previous piece. Since the previous piece is related to the audio history or context, the part of the previous piece may be referred to as 'context'. In other words, for each piece after the initial piece, data of an extended size 'chunk+context' can be processed.

[0093] The electronic device identifies whether there is a previous piece, and if a previous piece exists, it can take a portion of the previous piece and append it to the current piece. For example, when piece B (520) is processed after piece A (510) has been processed, context A (512), which is a part of the previous piece A (510), is added and processed. That is, extended data equal to context A (512) + piece B (520) is processed. And, when piece C (530) is processed after piece B (520) has been processed, context B (522), which is a part of the previous piece B (520), is added and processed. That is, extended data equal to context B (522) + piece C (530) is processed.

[0094] Extended data is processed at each time step as many times as there are time steps. For example, extended data of context A (512) + fragment B (520) may be processed at time steps t=5, 4, 3, 2, 1, and then extended data of context B (522) + fragment C (530) may be processed at time steps t=5, 4, 3, 2, 1.

[0095] Meanwhile, in FIG. 5, for convenience of explanation, it was described that the fragments are processed to have overlapping regions, but the fragments undergo a feature extraction process to be processed in the diffusion model. Therefore, the description of the fragments in the above embodiment may actually include transforming the fragments to process features.

[0096] In one embodiment, features processed by a diffusion model can be restored back into fragments. The restored fragments may be referred to as enhanced fragments. Since the enhanced fragments include overlapping regions, the electronic device can adjust the signal strength of the overlapping regions to add the overlapping signals. For example, the electronic device can apply a window function (540) to each enhanced fragment. By applying the window function (540), the electronic device can maintain the central part of the signal and smoothly reduce the boundary regions. The electronic device can generate enhanced audio by adding the enhanced fragments to which the window function (540) has been applied.

[0097] FIG. 6a is a flowchart illustrating an audio enhancement method using a diffusion model according to one embodiment of the present disclosure.

[0098] As previously described in the description of Figures 4 and 5, the diffusion model utilizes context to proceed with inference by referencing previous fragments as history. However, adding context to each fragment increases the amount of computation at the diffusion time steps corresponding to each fragment. To reduce the amount of computation at each time step of the diffusion model, the electronic device can adjust the time steps to which the overlapping region is applied. For example, the electronic device can reduce the amount of data processed in the inference of the diffusion model by processing the original fragment in the initial time steps and processing the overlapping region by applying context in further time steps. Typically, since the number of time steps of the diffusion process is designed to decrease, the electronic device can process fragments without overlapping in the initial time steps. Then, as the time steps decrease, the electronic device can process extended fragments (i.e., extended features) by creating an overlapping region at that time step. Alternatively, the electronic device can process a piece containing an overlapping region from an initial time step, and as the time step decreases, it can process an extended piece (i.e., an extended feature) by increasing the size of the overlapping region at that time step.

[0099] In one embodiment, the audio enhancement operation of the electronic device may include an outer loop (600) and an inner loop (610). Since the operation of the outer loop (600) corresponds to the operation of the outer loop (400) of FIG. 4, FIG. 6a describes only the operation of the inner loop (610) which includes additional operations.

[0100] In the inner loop (610), the electronic device determines whether to change the feature size. For example, the electronic device may determine whether to extend the feature at each time step based on a defined policy. The defined policy may be a predefined in various ways of determining which time step among the time steps included in the inner loop (610) the feature will be extended.

[0101] In one embodiment, the defined policy may be to change the feature size at time step n among the time steps between t=S and t=0. In this case, the electronic device may use a diffusion model to decrease the time steps starting from time step t=S, and then extend the feature size when time step t=n is reached. The context representing the extended feature may be a part of the feature processed in the previous cycle of the outer loop (600). The electronic device may load the context, which is a part of the previous feature, from memory and append it to the current feature. The electronic device may store information about the extended feature in memory and perform a reverse diffusion process to process the extended feature. When feature processing (reverse diffusion) is performed, the electronic device may perform 'feature restoration', 'window application', and 'add to output' operations. The 'feature restoration', 'window application', and 'add to output' operations may be performed only once at the final time step.

[0102] In other words, the electronic device can reduce the amount of computation by processing extended-size features in some time steps rather than all time steps based on a defined policy.

[0103] Figure 6b shows the algorithm described in Figure 6a in pseudocode.

[0104] In FIG. 6b, each line of the algorithm corresponds to each operation of the outer loop (600) or inner loop (610) of FIG. 6a.

[0105] For example, the operation of identifying whether data is available and collecting data in the outer loop (600) of FIG. 6a corresponds to the algorithm loading data into the data variable in hop-sized chunks, data←collect_data(hop); and the operation of extracting features in the outer loop (600) of FIG. 6a corresponds to the algorithm extracting features from the data and assigning them to the feature variable, feature←feature_extract(data); and the inner loop (610) of FIG. 6a corresponds to the for loop of the algorithm.

[0106] The modify_context() function is a function that controls the extension of the feature size by changing the context. There may be various policies for extending the feature size. Examples of specific policies for changing the feature size are described in the description of Figures 9a through 9d.

[0107] The get_memory() function extracts data from a feature matrix stored in memory. The context added by the algorithm to extend the feature size is extracted from previous data fragments and may be stored in memory.

[0108] The add_to_output() function can add each sample of data to the stream of the output audio signal at positions corresponding to each timestamp, assuming that each sample of data has a timestamp.

[0109] FIG. 7 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to process overlapping pieces.

[0110] In one embodiment, the electronic device processes data of the size of the piece for the initial piece, and can process data of an extended size for each subsequent piece.

[0111] In the initial piece, piece A (710), only the size of the piece of data can be processed. For example, piece A (710) is processed without context at each time step t=5, 4, 3, 2, 1. Once the diffusion processes for piece A (710) are completed, the next piece in sequence (e.g., piece B (720)) can be processed.

[0112] For each piece after the initial piece, the electronic device can process data of an extended size that includes an area overlapping with a part of the previous piece. The extended data, which is a part of the previous piece, may be referred to as the context. To reduce the computational load, the electronic device can gradually increase the size of the context in time steps.

[0113] For example, regarding the internal loop cycle of piece B (720), at the first time step t=5, only piece B (720) is processed. Subsequently, as time steps progress (i.e., as time steps decrease), the electronic device may add the context A (712) of the previous piece A (710) to the current piece B (720). In this case, the electronic device may cause the size of the overlapping area to gradually increase at each time step. In the exemplary case of FIG. 7, the data processed at each time step of the piece after the initial piece is summarized as follows.

[0114] Time step t=5: Piece B(720), Total size: 5

[0115] Time step t=4: 1 / 4 (recent) of context A (712) + fragment B (720), total size: 6

[0116] Time step t=3: 2 / 4 (recent) of context A (712) + fragment B (720), total size: 7

[0117] Time step t=2: 3 / 4 (recent) of context A (712) + fragment B (720), total size: 8

[0118] Time step t=1: Context A(712) + Fragment B(720), Total size: 9

[0119] Meanwhile, in FIG. 7, for convenience of explanation, it was explained that the size of each fragment is 5 and the size of the context is 4, but it is not limited to the example described above. Also, for convenience of explanation, it was explained that the fragments are processed to have overlapping regions, but the fragments undergo a feature extraction process to be processed in the diffusion model. Therefore, the description of the fragments in the above embodiment may actually include transforming the fragments to process features.

[0120] In one embodiment, the electronic device can obtain enhanced fragments by restoring features processed by a diffusion model into fragments. Since the enhanced fragments include overlapping regions, the electronic device can adjust the signal strength of the overlapping regions to add the overlapping signals. For example, the electronic device can apply a window function (730) to each enhanced fragment. By applying the window function (730), the electronic device can maintain the central part of the signal and smoothly reduce the boundary regions. The electronic device can generate enhanced audio by adding the enhanced fragments to which the window function (730) has been applied.

[0121] In one embodiment, the maximum latency of the audio enhancement operation is context + fragment. This is because the data collected at time t is output when the interval to which the data belongs ends. In the example of FIG. 7, since the context size = 4 and the fragment size = 5, the maximum latency can be 9. Based on the example illustrated in FIG. 7, the data output time and latency according to the data collection time are shown as in Table (740).

[0122] Referring to the table (740), fragment A (710) is collected at t=1 and output at t=6, so the delay time is 5. In the case of fragment B (720), the data collection time may vary depending on the size of context A (712), but the maximum delay time is 9 when collected at t=2 and output at t=11. In other words, the delay time of the audio enhancement operation can be determined by the fragment size and the context size.

[0123] In order to provide streaming while performing audio enhancement processing on an electronic device, a short latency is required. Generally, for a system processing an audio stream to avoid causing inconvenience to the user, it is appropriate for the latency to be 50-60ms or less. In the present disclosure, the maximum latency is determined by the sum of the fragment length and the context length. Accordingly, in one embodiment, the electronic device may set the fragment length to 42ms and the context length to 16ms.

[0124] FIG. 8a is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to extend a feature to create features having an overlapping region.

[0125] In one embodiment, the electronic device enhances audio through a back-diffusion process that gradually removes noise from time step t=S to t=0. In this case, the electronic device can process extended features including regions that partially overlap with the previous feature at each time step.

[0126] In FIG. 8a, the back-diffusion processes of the current piece h (800) are performed sequentially. The sequential execution of the back-diffusion processes means that a single back-diffusion process is performed at each time step while the time step decreases from t=S to t=0. Below, the intermediate time steps k+1, k, and k-1 included in the time steps from t=S to t=0 are described as examples (S > k > 0).

[0127] At time step k+1, when the electronic device processes piece h (800), which is the current piece, the electronic device performs a single back-diffusion process for the feature (802) of step k+1.

[0128] Next, at time step k, the electronic device performs a single back-diffusion process on the feature (804) of step k. The feature (804) of step k includes the feature (802) of step k+1 and the context (812) of step k+1.

[0129] Next, at time step k-1, piece h (800) performs a single back-diffusion process on the feature (806) of step k-1. The feature (806) of step k-1 includes the feature (804) of step k and the context (814) of step k.

[0130] FIG. 8b is a diagram illustrating the operation of an electronic device according to one embodiment to extend a feature to create features having an overlapping area.

[0131] Referring to FIG. 8b, the extension of features for piece h (800) is described. The operations for extending features of piece h (800) may have been performed in the same way on the previous piece, piece h-1 (810), and may also be performed in the same way on the next piece, piece h+1. In other words, feature extension can be applied in the same manner to the remaining pieces, excluding the initial piece, and FIG. 8b is described only with respect to piece h (800) for the sake of convenience of explanation.

[0132] Referring to FIG. 8b, all back-diffusion processes corresponding to the previous piece, piece h-1 (810), are performed sequentially. For example, a single back-diffusion process is performed at each time step from t=S to t=0 of piece h-1 (810), and subsequently, back-diffusion processes of the current piece, piece h (800), are performed sequentially (S > k > 0).

[0133] For example, at time step k+1 of fragment h (800), a single backdiffusion process is performed on the features (802) of step k+1. After the backdiffusion process of time step k+1 is performed, the context (812) of step k+1 is added. The context (812) of step k+1 can be obtained by loading from memory a portion of the features processed in the backdiffusion process of the previous fragment, fragment h-1 (810).

[0134] Next, at time step k, a single back-diffusion process is performed on the feature (804) of step k. After the back-diffusion process of time step k is performed, the context (814) of step k is added. The context (814) of step k can be obtained by loading from memory a portion of the feature processed in the back-diffusion process of the previous piece, piece h-1 (810).

[0135] Next, a single backdiffusing process is performed for each time step k-1, k-2, ..., 0. The feature length at each single time step can be extended or maintained.

[0136] When all back-diffusion processes corresponding to piece h (800) are performed sequentially, back-diffusion processes are performed in the same way for the next piece, piece h+1.

[0137] In the manner described above, when an electronic device performs an inference task using a diffusion model, if the time step of the diffusion process decreases, the size of the region partially overlapping with the previous feature can be increased to process the extended feature at that time step.

[0138] FIG. 9a is a drawing illustrating a policy applied by an electronic device according to one embodiment of the present disclosure when extending a feature.

[0139] In one embodiment, the electronic device may apply the aforementioned feature extension to all time steps of the diffusion model. However, if the feature is extended by adding context at all time steps, the amount of computation at the diffusion time steps increases.

[0140] Accordingly, in another embodiment, the electronic device may adjust the time steps to which the overlapping region is applied in order to reduce the amount of computation at each time step of the diffusion model. That is, the electronic device may adjust feature extension using various predefined policies to determine which time step among the time steps the feature will be extended. The electronic device may reduce the amount of computation by processing extended-size features at some time steps rather than all time steps based on the defined policies.

[0141] For example, the electronic device may apply policy A (910), which is an example of a defined policy. Referring to policy A (910), the left column of the table indicates the diffusion time step number. The right column of the table indicates whether to extend the context. In the table, (+) indicates a time step to extend the feature by adding context, and (-) indicates a time step to retain the feature without extending the context.

[0142] As the context is extended at earlier time stages, the diffusion model can utilize more information for inference, which can lead to improved performance. However, there is a trade-off in that the amount of computation increases as the amount of data to be processed grows.

[0143] The electronic device can apply various policies to control the performance of optimal computation while minimizing quality degradation of the audio enhancement process.

[0144] In policy A (910), context extension is applied in the initial time steps S, S-1, S-2, ..., the feature is maintained in the remaining intermediate time steps, and context extension is applied in the final time steps 1, 0 so that the feature size can be increased.

[0145] However, policy A (910) is merely an example showing that context extension is applied only at the beginning and end of the time stage, and does not limit the context extension to specific time stages.

[0146] FIG. 9b is a drawing illustrating a policy applied by an electronic device according to one embodiment of the present disclosure when extending a feature.

[0147] In one embodiment, the electronic device may apply policy B (920), which is an example of a defined policy. Referring to policy B (920), the electronic device may progressively increase the feature size in the last time steps. For example, the feature may be maintained in the initial time steps S, S-1, S-2, ..., and then the feature size may be increased by applying context extension in the last time steps 2, 1, 0.

[0148] The electronic device can reduce the amount of data processed in the inference of the diffusion model by maintaining the feature length during the early stages of the diffusion process where the influence of context extension is not significant, and then applying context extension to process the overlapping regions during the later stages of the diffusion process where context extension becomes important.

[0149] However, policy B (920) is merely an example where context extension is applied only in the latter part of the time step, and does not limit the context extension to a specific time step.

[0150] FIG. 9c is a drawing illustrating a policy applied by an electronic device according to one embodiment of the present disclosure when extending a feature.

[0151] In one embodiment, the electronic device may apply policy C (930), which is an example of a defined policy. Referring to policy C (930), the electronic device may progressively increase the feature size at preset time intervals. For example, the electronic device may increase the feature size by applying context extension at preset time intervals of 3, such as S, S-3, ..., 0, and maintain the feature size at the remaining time steps. In this case, maintaining the feature size may include maintaining the feature size that was already extended in the previous step.

[0152] The electronic device can reduce the amount of data processed in the inference of the diffusion model by applying context extension only at pre-set time steps, rather than at all time steps.

[0153] However, policy C (930) is merely an example of context extension being applied at pre-set time intervals, and does not require that the context be extended at a specific time step or limit the length of the specific time interval at which context extension is applied.

[0154] FIG. 9d is a drawing illustrating a policy applied by an electronic device according to one embodiment of the present disclosure when extending a feature.

[0155] In one embodiment, the electronic device may apply policy D (940), which is an example of a defined policy. Referring to policy D (940), the electronic device may progressively increase the feature size in the initial time steps. For example, context extension may be applied in the initial time steps S, S-1, S-2, S-3 to increase the feature size, and the extended feature size may be maintained in the subsequent time steps.

[0156] The electronic device can reduce the amount of data processed in the inference of the diffusion model by gradually increasing the feature size in the initial time steps rather than increasing it all at once (e.g., see the example shown in Fig. 5). Additionally, sufficient performance of audio enhancement can be achieved by applying context extension in the initial time steps.

[0157] However, policy D (940) is merely an example where context extension is applied only at the beginning of the time step, and does not limit the context extension to a specific time step.

[0158] FIG. 10 is a drawing for explaining the audio enhancement results of an electronic device according to one embodiment of the present disclosure.

[0159] For real-time streaming while generating enhanced audio, it is necessary to divide the audio into multiple fragments and concatenate them for continuous playback whenever each audio fragment is processed. Additionally, to enable the diffusion model to function better, it may be necessary to include overlapping regions between adjacent audio fragments. Each audio fragment contains one or more frames, and each frame may contain multiple samples of audio depending on the sampling rate.

[0160] The first spectrogram (1010) illustrates the result of processing audio using a general overlapping method (e.g., overlap-save, overlap-add, etc.). Referring to the first spectrogram (1010), artifacts may occur whenever frames in the overlapping area are connected and added.

[0161] In one embodiment, the electronic device can reduce artifacts that occur when connecting frames during voice enhancement in real-time streaming by applying a window function. A second spectrogram (1020) illustrates the result of processing audio using the improved overlap-addition method of the present disclosure. Referring to the second spectrogram (1020), it can be seen that no artifacts occur even when connecting and adding frames in the overlapping area. The window function may be, for example, a Hanning window, but is not limited thereto.

[0162] In one embodiment, the electronic device can apply the window function in two steps.

[0163] First, the electronic device can apply a window function at the data block level, as described in the examples above. The size of the data block (e.g., 50-60ms) is determined by the fragment + context and may also be referred to as the extended fragment or extended feature. Since an audio fragment contains one or more frames, the size of the audio fragment is a multiple of the frame, and the context is also a multiple of the frame. Therefore, the size of the data block can also be a multiple of the frame.

[0164] Secondly, electronic devices can apply a window function at the frame level, which is a smaller unit. The frame size (e.g., 5-10ms) can be determined by the input unit used for STFT transformation. When performing STFT transformation, if long-duration data, such as an entire data block, is treated as a single frame, there is a trade-off where more frequencies can be extracted but resolution is reduced. Since it has been experimentally proven that using a short window is advantageous for the performance of the diffusion model, electronic devices can apply a window function at the frame level, which is a short time unit.

[0165] FIG. 11 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure.

[0166] In one embodiment, the electronic device (2000) may include a communication interface (2100), a memory (2200), and a processor (2300).

[0167] The communication interface (2100) can perform data communication with other electronic devices under the control of the processor (2300).

[0168] The communication interface (2100) can perform data communication between an electronic device (2000) and another electronic device (e.g., a server, an external electronic device, etc.) by using at least one of data communication methods including, for example, a wired LAN (e.g., Ethernet), a wireless LAN (e.g., Wi-Fi), a cellular network (e.g., 4G, 5G, etc.), Bluetooth, BLE (Bluetooth Low Energy), ZigBee, infrared communication (IrDA, infrared Data Association), NFC (Near Field Communication), RF communication, and various other types of known wireless / wired communication technologies. The communication interface (2100) may include a communication circuit designed to use the aforementioned communication methods.

[0169] The electronic device (2000) can transmit and receive data for generating enhanced audio with another electronic device using a communication interface (2100). For example, it can transmit and receive input data (e.g., original audio) and output data (e.g., enhanced audio) of a diffusion model to another electronic device, and receive a diffusion model for audio enhancement from another electronic device.

[0170] The memory (2200) may include various types of memory. The memory (2200) may include a main memory that stores data currently being processed in the electronic device (2000). For example, the main memory may include volatile memory such as RAM (Random Access Memory) or SRAM (Static Random Access Memory), but is not limited thereto. The memory (2200) may include a secondary memory that permanently stores large amounts of data (e.g., programs, system files, etc.). For example, the secondary memory may include non-volatile memory including at least one of a hard disk drive (HDD), a solid-state drive (SSD), an optical drive (e.g., CD), a flash drive, ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), and PROM (Programmable Read-Only Memory), but is not limited thereto.

[0171] The memory (2200) may store one or more instructions and programs that cause the electronic device (2000) to operate to enhance audio. For example, the memory (2200) may store data for performing training and inference operations of a diffusion model.

[0172] The processor (2300) can control the overall operations of the electronic device (2000). The processor (2300) may include a processing circuit. For example, the processor (2300) can control the overall operations of the electronic device (2000) processing and enhancing audio by executing one or more instructions of a program stored in memory (2200). There may be one or more processors (2300).

[0173] The processor (2300) may be composed of at least one of, for example, a Central Processing Unit (CPU), a Microprocessor, a Graphic Processing Unit (GPU), ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), an Application Processor (AP), a Neural Processing Unit (NPU), or an AI-dedicated processor designed with a hardware structure specialized for processing AI models, but is not limited thereto.

[0174] In one embodiment, there may be one or more processors (2300). If there is one or more processors (2300), the operations of the present disclosure may be performed by one or more processors individually or collectively by executing instructions and / or programs stored in memory (2200). If the method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one processor (2300) or by a plurality of processors (2300).

[0175] For example, when the first, second, and third operations are performed by a method according to one embodiment, the first, second, and third operations may all be performed by a first processor, or some of the first to third operations may be performed by a first processor (e.g., a general-purpose processor) and the remaining operations may be performed by a second processor (e.g., an AI-dedicated processor). Here, operations for training / inference of an AI model may be performed by an AI-dedicated processor, which is an example of a second processor. However, the embodiments of the present disclosure are not limited thereto.

[0176] One or more processors according to the present disclosure may be implemented as a single-core processor or as a multi-core processor. When a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by a single core or by a plurality of cores included in one or more processors.

[0177] In one embodiment, the electronic device may include a microphone and a speaker. The electronic device (2000) may acquire or output audio through the microphone and the speaker. For example, the electronic device (2000) may receive a user's voice or audio through the microphone. However, the voice or audio to which audio enhancement processing is applied does not necessarily have to be acquired through the microphone. The electronic device (2000) may perform audio enhancement on audio acquired in any manner and output the enhanced voice or audio through the speaker.

[0178] In one embodiment, the electronic device may include an input / output interface. The input / output interface may include configurations for input / output methods such as a 3.5mm audio jack or a USB audio interface, but is not limited thereto. The electronic device (2000) may be connected to external devices such as a microphone or a speaker through the input / output interface.

[0179] The present disclosure provides an audio enhancement method that enhances audio using an improved superposition-addition method and is compatible with both streaming and non-streaming of the enhanced audio. The technical problems to be solved by the present disclosure are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by those skilled in the art from the description in this specification.

[0180] According to one aspect of the present disclosure, an audio enhancement method may be provided.

[0181] The above method may include the step of dividing the input audio to obtain a plurality of first pieces.

[0182] The above method may include the step of converting the first pieces into the frequency domain to obtain first features.

[0183] The above method may include the step of obtaining enhanced second features using a diffusion model based on the first features.

[0184] The above diffusion model processes the first features sequentially, and when processing each feature, it can process an extended feature that includes a region partially overlapping with the previous feature.

[0185] The above method may include the step of converting each of the second features into a time domain to obtain second pieces.

[0186] The above method may include the step of applying a window function to the second pieces to adjust the signal strength of the overlapping area between adjacent second pieces.

[0187] The above method may include the step of generating enhanced audio by overlapping and adding the second pieces.

[0188] Each audio piece of the first pieces or the second pieces may include one or more frames, and each frame may include multiple samples of audio according to the sampling rate.

[0189] The step of acquiring the enhanced second features using the above diffusion model may include a step of acquiring the enhanced second features through a reverse diffusion process that proceeds in time steps and repeats the process of predicting and removing noise as the time steps decrease.

[0190] The step of acquiring the enhanced second features using the above diffusion model may include, when the time step of the diffusion process decreases, increasing the size of the region partially overlapping with the previous feature to process the extended feature at that time step.

[0191] The step of acquiring the enhanced second features using the above diffusion model may include a step of determining whether to extend the features at each time step based on a defined policy.

[0192] The policy defined above may be to gradually increase the feature size in the final time steps.

[0193] The policy defined above may be to gradually increase the feature size at preset time intervals.

[0194] The policy defined above may be to gradually increase the feature size during the initial time steps.

[0195] The step of generating the above-mentioned enhanced audio may be to generate the enhanced audio while streaming it.

[0196] The above window function may be a Hanning window function.

[0197] According to one aspect of the present disclosure, an electronic device for enhancing audio may be provided.

[0198] The electronic device may include a communication interface, at least one processor, and a memory for storing instructions.

[0199] By executing the above instructions by the at least one processor, the electronic device can divide the input audio to obtain a plurality of first pieces.

[0200] By executing the above instructions by the at least one processor, the electronic device can convert the first pieces into the frequency domain to obtain the first features.

[0201] By executing the above instructions by the at least one processor, the electronic device can obtain enhanced second features using a diffusion model based on the first features.

[0202] The above diffusion model processes the first features sequentially, and when processing each feature, it can process an extended feature that includes a region partially overlapping with the previous feature.

[0203] By executing the above instructions by the at least one processor, the electronic device can obtain the second pieces by converting each of the second features into the time domain.

[0204] By executing the above instructions by the at least one processor, the electronic device can apply a window function to the second pieces to adjust the signal strength of the overlapping area between adjacent second pieces.

[0205] By executing the above instructions by the at least one processor, the electronic device can generate enhanced audio by overlapping and adding the second pieces.

[0206] Each audio piece of the first pieces or the second pieces may include one or more frames, and each frame may include multiple samples of audio according to the sampling rate.

[0207] By executing the above instructions by the at least one processor, the electronic device can acquire the enhanced second features through a reverse diffusion process that proceeds in time steps and repeats the process of predicting and removing noise as the time steps decrease.

[0208] By executing the above instructions by the at least one processor, the electronic device can process the extended feature at the corresponding time step by increasing the size of the region partially overlapping with the previous feature when the time step of the diffusion process is reduced.

[0209] By executing the above instructions by the at least one processor, the electronic device can determine whether to extend a feature at each time step based on a defined policy.

[0210] The policy defined above may be to gradually increase the feature size in the final time steps.

[0211] The policy defined above may be to gradually increase the feature size at preset time intervals.

[0212] The policy defined above may be to gradually increase the feature size during the initial time steps.

[0213] By executing the above instructions by the at least one processor, the electronic device can generate the enhanced audio while streaming.

[0214] Meanwhile, embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules executed by a computer. A computer-readable medium may be any available medium accessible by a computer and includes both volatile and non-volatile media, and both removable and non-removable media. Additionally, a computer-readable medium may include computer storage media and communication media. Computer storage media include both volatile and non-volatile, removable and non-removable media implemented by any method or technique for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include other data of modulated data signals, such as computer-readable instructions, data structures, or program modules.

[0215] Additionally, computer-readable storage media may be provided in the form of non-transitory storage media. Here, 'non-transitory storage media' simply means that it is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily. For example, 'non-transitory storage media' may include a buffer in which data is stored temporarily.

[0216] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0217] The foregoing description of the present disclosure is for illustrative purposes only, and those skilled in the art will understand that other specific forms can be easily modified without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.

[0218] The scope of the present disclosure is defined by the claims set forth below rather than by the detailed description above, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts thereof should be interpreted as being included within the scope of the present disclosure.

Claims

1. Regarding audio enhancement methods, A step of splitting the input audio to obtain a plurality of first chunks; A step of converting the above first pieces into the frequency domain to obtain first features; A step of obtaining enhanced second features using a diffusion model based on the first features, wherein the diffusion model sequentially processes the first features, and when processing each feature, processes an extended feature that includes a region partially overlapping with the previous feature; A step of obtaining second fragments by converting each of the above second features into a time domain; A step of applying a window function to the second pieces to adjust the signal strength of the overlapping area between adjacent second pieces; and A method comprising the step of generating enhanced audio by overlapping and adding the second pieces.

2. In Paragraph 1, A method in which each audio piece of the first pieces or the second pieces comprises one or more frames, and each frame comprises a plurality of samples of audio according to a sampling rate.

3. In Paragraph 2, The step of obtaining the enhanced second features using the above diffusion model is, A method comprising the step of obtaining the enhanced second features through a reverse diffusion process that proceeds in time steps and repeats the process of predicting and removing noise as the time steps decrease.

4. In Paragraph 3, The step of obtaining the enhanced second features using the above diffusion model is, A method comprising the step of processing extended features in the time step of the diffusion process by increasing the size of the area partially overlapping with the previous feature when the time step of the diffusion process is reduced.

5. In Paragraph 4, The step of obtaining the enhanced second features using the above diffusion model is, A method comprising a step of determining whether to extend a feature at each time step based on a defined policy.

6. In Paragraph 5, The policy defined above is a method of gradually increasing the feature size at preset time intervals.

7. In Paragraph 5, The policy defined above is a method of gradually increasing the feature size in initial time steps.

8. In an electronic device for enhancing audio, Communication interface; At least one processor; and It includes memory for storing instructions, By executing the above instructions by the at least one processor, the electronic device, The input audio is split to obtain multiple first chunks, and The above first pieces are converted into the frequency domain to obtain first features, and Enhanced second features are obtained using a diffusion model based on the first features, wherein the diffusion model processes the first features sequentially, and when processing each feature, processes an extended feature that includes an area partially overlapping with the previous feature. Each of the above second features is converted into the time domain to obtain second pieces, and By applying a window function to the second pieces above, the signal strength of the overlapping area between adjacent second pieces is adjusted, and An electronic device that generates enhanced audio by overlapping and adding the above second pieces.

9. In Paragraph 8, An electronic device wherein each audio piece of the first pieces or the second pieces comprises one or more frames, and each frame comprises a plurality of samples of audio according to a sampling rate.

10. In Paragraph 9, By executing the above instructions by the at least one processor, the electronic device, An electronic device that obtains the enhanced second features through a reverse diffusion process that proceeds in time steps and repeats the process of predicting and removing noise as the time steps decrease.

11. In Paragraph 10, By executing the above instructions by the at least one processor, the electronic device, An electronic device that processes extended features in a time step by increasing the size of a region partially overlapping with the previous feature when the time step of the diffusion process decreases.

12. In Paragraph 11, By executing the above instructions by the at least one processor, the electronic device, An electronic device that determines whether to extend a feature at each time step based on a defined policy.

13. In Paragraph 12, The above-defined policy is an electronic device that gradually increases the feature size at preset time intervals.

14. In Paragraph 12, The above-defined policy is an electronic device that gradually increases the feature size in initial time steps.

15. A computer-readable recording medium having a program for executing the method of any one of paragraphs 1 through 7 on a computer.

Citation Information

Patent Citations

  • A diffusion-weighted imaging method

    CN107240125B

  • Terminal device, information providing system, processing method of terminal device and program

    JP2018106100A

  • Voice data generation device

    JP2022083706A

  • Device and method for a bandwidth extension of an audio signal

    KR1020110007083A

  • Bandwidth extension encoder, bandwidth extension decoder and phase vocoder

    KR1020120031957A