Audio separation method and electronic device for performing the same

By extracting and identifying audio and background sound features and combining target information for audio separation, the problems of artifacts and unclear sound quality caused by background sound mixing are solved, achieving a clearer audio separation effect.

CN120752701APending Publication Date: 2025-10-03SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480012857.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-26
Filing Date
2024-01-24
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing audio separation technologies fail to effectively process background sounds, resulting in artifacts and unclear sound quality in the separated audio.

Method used

By extracting audio and background sound features from the sound source, identifying the degree of correlation between the two, and generating adjusted audio features based on control parameters, audio separation is performed in combination with target information, and background sound control is achieved using an audio separation system.

Benefits of technology

Effectively remove or reduce background sounds, improve the clarity and quality of separated audio, and avoid the generation of artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120752701A_ABST
    Figure CN120752701A_ABST
Patent Text Reader

Abstract

There is provided a method of separating at least one candidate audio from a sound source comprising the at least one candidate audio and a background sound using an audio separation system, the method comprising the steps of: extracting a first audio feature from the sound source; extracting a background sound feature from the sound source, wherein the background sound feature is used for identifying a correlation degree between the first audio feature and the background sound; generating a second audio feature based on the first audio feature, the background sound feature and a background control parameter, wherein the background control parameter is configured to control the background sound; and generating at least one separated audio based on the target information corresponding to the one candidate audio, the first audio feature, and the second audio feature in which the background sound is adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an audio processing method and an electronic device for performing the method. Specifically, the present disclosure relates to a method for performing audio separation on a sound source and an electronic device for performing the method. Background Art

[0002] Audio separation technology is a technique in which one or more candidate audio signals are separated from a sound source. For example, audio features are typically extracted from the sound source using a convolution block, and a spectrum corresponding to each candidate audio signal is generated from the audio features using an upconvolution block, resulting in the final separated audio signal. This process requires target information corresponding to each candidate audio signal, and typically uses visual information (e.g., an image containing all or part of the face of the speaker corresponding to the candidate audio signal) or audio information (e.g., pre-recorded speech of the speaker corresponding to the candidate audio signal) as the target information.

[0003] Existing audio separation techniques focus solely on extracting features from candidate audio, without considering the background sounds included in the sound source. Consequently, the separated audio often includes background sounds. In this case, a separate model for extracting background sound characteristics is required to adjust the background sounds in the separated audio (e.g., reduce their volume or remove them). However, because two inference processes are performed during the separation of candidate audio from the sound source and subsequent removal of background sounds from the separated audio, artificial sounds (e.g., artifacts) may be incorporated into the final separated audio, potentially resulting in unnatural or unclear audio. Summary of the Invention

[0004] Technical Solutions

[0005] According to one aspect of the present disclosure, a method is provided for separating one or more candidate audios in a sound source including one or more candidate audios and background sounds by using an audio separation system, the method comprising: extracting a first audio feature from the sound source, extracting a background sound feature from the sound source, the background sound feature identifying a degree of association between the first audio feature and the background sound, generating a second audio feature based on the first audio feature, the background sound feature, and a background sound control parameter, the background sound control parameter being configured to control the background sound, and generating one or more separated audios based on target information corresponding to the one or more candidate audios, the first audio feature, and the second audio feature in which the background sound is adjusted.

[0006] According to another aspect of the present disclosure, a computer-readable recording medium having a computer program recorded thereon is provided. When the computer program is executed by one or more computing devices, the one or more computing devices perform a method for separating one or more candidate audios in a sound source including one or more candidate audios and background sounds by using an audio separation system. The method includes: extracting a first audio feature from the sound source, extracting a background sound feature from the sound source, the background sound feature identifying a degree of association between the first audio feature and the background sound, generating a second audio feature based on the first audio feature, the background sound feature, and a background sound control parameter, the background sound control parameter being configured to control the background sound, and generating one or more separated audios based on target information corresponding to the one or more candidate audios, the first audio feature, and the second audio feature in which the background sound is adjusted.

[0007] According to another aspect of the present disclosure, an electronic device is provided that includes one or more processors and a memory storing a program for separating one or more candidate audios from a sound source including one or more candidate audios and background sounds by using an audio separation system. When executed by the one or more processors, the program can cause the electronic device to perform operations including: extracting a first audio feature from the sound source, extracting a background sound feature from the sound source, the background sound feature identifying a degree of association between the first audio feature and the background sound, generating a second audio feature based on the first audio feature, the background sound feature, and a background sound control parameter, the background sound control parameter being configured to control the background sound, and generating one or more separated audios based on target information corresponding to the one or more candidate audios, the first audio feature, and the second audio feature in which the background sound is adjusted. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a diagram for describing a situation in which an audio separation method can be applied according to an embodiment of the present disclosure.

[0009] Figure 2 An audio separation system according to an embodiment of the present disclosure is shown.

[0010] Figure 3a is a diagram for describing the operation of a background sound feature control module according to an embodiment of the present disclosure.

[0011] Figure 3b An example of a background sound control function is shown.

[0012] Figure 4 is a diagram for describing a method of training an audio separation system according to an embodiment of the present disclosure.

[0013] Figure 5 An audio separation system according to an embodiment of the present disclosure is shown.

[0014] Figure 6 is a diagram for describing the operation of a background sound feature control module according to an embodiment of the present disclosure.

[0015] Figure 7 is a flowchart of an audio separation method according to an embodiment of the present disclosure.

[0016] Figure 8 is a block diagram schematically illustrating a configuration of an electronic device for performing an audio separation method according to an embodiment of the present disclosure.

[0017] Figure 9a is a diagram for describing an application example of the audio separation system according to an embodiment of the present disclosure.

[0018] Figure 9b is a diagram for describing an application example of the audio separation system according to an embodiment of the present disclosure.

[0019] Figure 10a is a diagram for describing an application example of the audio separation system according to an embodiment of the present disclosure.

[0020] Figure 10b is a diagram for describing an application example of the audio separation system according to an embodiment of the present disclosure.

[0021] Figure 10c is a diagram for describing an application example of the audio separation system according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] The following detailed description is provided to help the reader obtain a comprehensive understanding of the methods, devices and / or systems described herein. However, after understanding the disclosure of the application, various changes, modifications and equivalents of the methods, devices and / or systems described herein will be apparent. For example, the order of operations described herein is merely an example and is not limited to those set forth herein, but can be changed as will be apparent after understanding the disclosure of the application, except for operations that must occur in a specific order. In addition, for the purpose of improving clarity and brevity, descriptions of features known after understanding the disclosure of the application may be omitted.

[0023] As used herein, the expression "at least one of a, b, or c" may mean only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or variations thereof.

[0024] Although the terms used herein are selected from commonly used terms currently in use, taking into account their functions in the embodiments of the present disclosure, these terms may be different depending on the intentions of those skilled in the art, precedents, or the emergence of new technologies. In addition, in certain cases, terms are arbitrarily selected by the applicant of the present disclosure, in which case the meanings of these terms will be described in detail in the corresponding parts of the specific embodiments. Therefore, the terms used herein should be understood based on the meaning of the terms and content throughout the present disclosure, rather than just the designation of the terms.

[0025] Although terms such as "first," "second," and the like may be used to describe various components, such components should not be limited by these terms. These terms are merely used to distinguish one component from another. For example, a first component may be referred to as a second component, and a second component may be referred to as a first component in a similar manner without departing from the scope of the embodiments of the present disclosure.

[0026] It should be understood that when a specific component is referred to as being “coupled to” or “connected to” another component, the component may be directly coupled to or connected to the other component, but it may also be understood that there are different components between them. On the other hand, it should be understood that when a specific component is referred to as being “directly coupled to” or “directly connected to” another component, there are no other components between them.

[0027] A singular expression may also include a plural meaning as long as it is not inconsistent with the context. All terms used herein, including technical and scientific terms, may have the same meaning as commonly understood by those skilled in the art.

[0028] As used herein, terms such as “includes,” “including,” or “having” specify the presence of stated features, quantities, stages, operations, components, parts, or a combination thereof, but do not preclude the presence or addition of one or more other features, quantities, stages, operations, components, parts, or a combination thereof.

[0029] As used herein, components represented as, for example, “…device”, “…unit”, “…module”, etc. may represent units in which two or more components are combined into one component or one component is divided into two or more components according to their functions. In addition, each component described below may additionally perform some or all of the functions undertaken by other components in addition to its main function, and some of the main functions of each component may be performed separately by other components. Depending on the embodiment, components represented as, for example, “…device”, “…unit”, “…module”, etc. may be implemented by hardware, software, or a combination of hardware and software. For example, in the case of hardware, these components may be implemented by circuits, processors, etc. In the case of software, these components may be implemented as software codes, computer programs, and / or instructions, which may be implemented by or in a processor or other circuits.

[0030] In the present disclosure, functions related to artificial intelligence are performed by a processor and a memory. The processor may include one or more processors. In this case, the one or more processors may be a general-purpose processor (such as a central processing unit (CPU), an application processor (AP), or a digital signal processor (DSP)), a dedicated graphics processor (such as a graphics processing unit (GPU) or a visual processing unit (VPU)), or a dedicated artificial intelligence processor (such as a neural processing unit (NPU)). The one or more processors perform control to process input data according to predefined operating rules or an artificial intelligence model stored in the memory. Alternatively, in the case where the one or more processors are dedicated artificial intelligence processors, the dedicated artificial intelligence processor can be designed to have a hardware structure specifically for processing a specific artificial intelligence model.

[0031] Predefined operating rules or artificial intelligence models can be generated through a training process. Here, being generated through a training process can mean generating predefined operating rules or artificial intelligence models that are configured to perform desired characteristics (or purposes) by training a basic artificial intelligence model using a learning algorithm that utilizes a large amount of training data. The training process can be performed by the device itself on which the artificial intelligence according to the present disclosure is executed, or by a separate server and / or system. Examples of learning algorithms may include, for example, supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning, but are not limited thereto.

[0032] The artificial intelligence model may include multiple neural network layers. Each neural network layer has multiple weight values, and the neural network arithmetic operation is performed via an arithmetic operation between the arithmetic operation result of the previous layer and the multiple weight values. The multiple weight values ​​in each neural network layer can be optimized by training the artificial intelligence model. For example, the multiple weight values ​​can be refined to reduce or minimize the loss or cost value obtained by the artificial intelligence model during the training process. The artificial neural network may include, for example, a deep neural network (DNN), and may include, for example, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q network (DQN), etc., but is not limited thereto.

[0033] In the present disclosure, a machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" refers to a tangible device and does not include signals (e.g., electromagnetic waves). The term "non-transitory storage medium" does not distinguish between cases where data is semi-permanently stored in the storage medium and cases where data is temporarily stored. For example, a non-transitory storage medium may include a buffer in which data is temporarily stored.

[0034] According to an embodiment of the present disclosure, the methods according to various aspects disclosed herein may be included in a computer program product and then provided. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an app store, or distributed directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable application) may be temporarily stored in a machine-readable storage medium, such as a memory on a manufacturer's server, an app store's server, or a relay server.

[0035] Hereinafter, the embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to allow those skilled in the art to easily implement the embodiments of the present disclosure. However, the present disclosure can be embodied in many different forms and should not be construed as limited to the embodiments of the present disclosure set forth herein.

[0036] In the present disclosure, the term "background sound" can be understood as audio other than candidate audio among multiple audios included in a sound source. For example, in the case of separating the voice of each of two specific speakers from a sound source including the voice of the speaker, the voices of other people around the speaker, ambient noise from the surrounding environment (such as the sound of car horns or the shouting of a crowd), background music used to create the atmosphere of the scene where the speaker appears, etc. can be background sounds.

[0037] It is important to note that candidate audio and background sound are not distinguished by their type. That is, the candidate audio and background sound can be the same type of sound, or different types of sound. For example, the candidate audio may be the voice of a specific speaker, and the background sound may be the voice of another person around the speaker. For example, the candidate audio may be the singing voice of a singer, and the background sound may be the accompaniment of the song sung by the singer.

[0038] Figure 1 is a diagram for describing a situation in which an audio separation method according to an embodiment of the present disclosure can be applied.

[0039] exist Figure 1 In the example, a war movie is reproduced on the screen, and a scene of a first person 11 and a second person 12 having a conversation on the battlefield is displayed. In this scene, not only the voices of the first person 11 and the second person 12 are output, but also background sounds such as the painful groans of other nearby injured people, gunshots, sound effects of nearby bomb explosions, or music used to create a more intense and dramatic atmosphere for the scene may be output. In the example case where the volume of such background sounds is much louder than the volume of the voices of the first person 11 and the second person 12, the audience may have difficulty fully understanding the content of the conversation between the first person 11 and the second person 12.

[0040] like Figure 1 As shown, in the example case where the volume 21 of the voice of the first person 11, the volume 22 of the voice of the second person 12, and the volume 23 of the background sound can be adjusted individually, by reducing the volume of the background sound in the scene, the audience can better understand the conversation between the first person 11 and the second person 12. In order to adjust the volume individually, it is necessary to separate the voice of the first person 11 and the voice of the second person 12 from the sound source mixed with various audios, and adjust the amplitude (i.e., volume) of the background sound included in the separated voice of each person.

[0041] Hereinafter, a method of separating one or more candidate audios from a sound source including the one or more candidate audios and background sounds, and adjusting the amplitude of the background sounds included in the separated audios will be described in detail.

[0042] Figure 2 An audio separation system 200 according to an embodiment of the present disclosure is shown.

[0043] In an embodiment of the present disclosure, audio separation system 200 may receive inputs including a sound source 201, target information 203, and background sound control parameters 202. Sound source 201 may include one or more candidate audio tracks (Audio #1, ..., Audio #N) and background sounds. Target information 203 may correspond to each candidate audio track. Audio separation system 200 may output separated audio tracks 204 and 205, with the amplitude (i.e., volume) of the background sounds adjusted according to background sound control parameters 202. Audio separation system 200 may be an artificial intelligence model trained using training data.

[0044] The type of candidate audio can be determined in various ways. In embodiments of the present disclosure, the candidate audio may be human speech. For example, the candidate audio may be the speech of a speaker, the singing of a singer, etc. In embodiments of the present disclosure, the candidate audio may be sound produced by a musical instrument or machine. For example, the candidate audio may be a piano piece played by an orchestra, the accompaniment of a song sung by a singer, etc.

[0045] Depending on the type of training data, in the example case where multiple audios are to be separated, the types of candidate audios may be the same or different from each other. In an embodiment of the present disclosure, the audio separation system 200 can separate multiple candidate audios of the same type from the sound source 201. For example, the audio separation system 200 can separate the speech of each speaker from a conversation between two speakers. In an embodiment of the present disclosure, the audio separation system 200 can separate multiple candidate audios of different types from the sound source 201. For example, the audio separation system 200 can separate the singer's singing voice and the song's accompaniment from a song sung by a singer. Figure 4 The process of training the audio separation system 200 is described.

[0046] The type of background sound can be determined differently. In an embodiment of the present disclosure, the type of background sound can be the same as the type of candidate audio. For example, in the case where the candidate audio is human voice, the background sound can be the voice of a speaker other than the speaker corresponding to the candidate audio. In an embodiment of the present disclosure, the type of background sound can be different from the type of candidate audio. For example, in the case where the candidate audio is human voice, the background sound can be the noise that occurs around the speaker corresponding to the candidate audio, such as the sound of a car horn.

[0047] The target information 203 corresponding to the candidate audio can be information used to specify or distinguish the candidate audio from the multiple audio sources included in the sound source 201. For example, the target information 203 can be understood as additional information used to specify which audio source the audio separation system 200 needs to separate from the multiple audio sources included in the sound source 201. In embodiments of the present disclosure, the target information 203 can be visual information. For example, if the candidate audio source is human speech, the target information 203 can be an image that includes all or part of the face of the speaker corresponding to the candidate audio source. For example, the target information 203 can be an image that includes the mouth area of ​​the speaker corresponding to the candidate audio source. However, the present disclosure is not limited to this. In embodiments of the present disclosure, the target information 203 can be audio information separated from the candidate audio source. For example, if the candidate audio source is human speech, the target information 203 can be the pre-recorded speech of the speaker corresponding to the candidate audio source. As another example, if the candidate audio source is piano music, the target information 203 can be a pre-recorded piano music.

[0048] Target information 203 corresponding to the candidate audio can be determined based on the candidate audio and the original data. For example, when separating the voices of two speakers appearing in a video including audio and images, the image including the target speaker can be used as target information 203. As another example, when separating the voice of a speaker appearing in the video and the voice of another speaker not appearing in the video, the image including the speaker appearing in the video and the pre-recorded voice of the speaker not appearing in the video can each be used as target information 203.

[0049] The background sound control parameter 202 may be a parameter α for determining the volume of the background sound to be included in the separated sound source 201. The background sound control parameter 202 may be determined by the user. For example, the background sound control parameter 202 may be determined by the user adjusting the volume of the background sound, such as Figure 1As shown. In embodiments of the present disclosure, the background sound control parameter 202 can be a real number between 0 and 1. For example, the audio separation system 200 can output separated audio 204 with the background sound removed if the background sound control parameter 202 is 0, and can output separated audio 205 with the background sound intact if the background sound control parameter 202 is 1. In other words, when the background sound control parameter 202 is closer to 0, the audio separation system 200 can remove more background sound, and when the background sound control parameter 202 is closer to 1, the audio separation system 200 can remove less background sound. In other words, the amount (or amplitude) of the background sound removed from the candidate audio is based on the background sound control parameter. For example, the amount (or amplitude) of the background sound removed from the candidate audio can be proportional to the background sound control parameter. However, the present disclosure is not limited thereto, and therefore, the background sound control parameter 202 can be identified in various ways. For example, the background sound control parameter 202 can include a numeric range other than 0 and 1. In another example, the background sound control parameter 202 can include a non-numeric identifier.

[0050] In an embodiment of the present disclosure, the audio separation system 200 may generate a spectrum of the sound source 201. For example, the audio separation system 200 may generate the spectrum of the sound source 201 by applying a short-time Fourier transform (STFT) to the sound source 201. A spectrum is a visual representation of the sound source 201. For example, a spectrum may represent differences in amplitude based on changes in time and frequency in the time-frequency domain, represented as differences in color or density between pixels.

[0051] In an embodiment of the present disclosure, the audio separation system 200 may extract the first audio feature h from the sound source 201 by processing the spectrum of the sound source 201 using the audio feature extraction module 210. i In an embodiment of the present disclosure, the audio feature extraction module 210 may include a plurality of convolution blocks, each of which may include one or more convolution layers, batch normalization, and an activation function. The activation function may include, but is not limited to, ReLu or sigmoid. In the audio feature extraction module 210, the feature extracted by one convolution block may be transferred to the next convolution block, and thus, the first audio feature h may be extracted by each convolution block. i Here, i is a natural number, and h i represents the first audio feature extracted by the i-th convolutional block. Figure 2 As shown, in the case where the audio feature extraction module 210 includes eight convolution blocks, the convolution blocks can respectively extract first audio features corresponding to the convolution blocks (for example, h1, h2, . . . , h8).

[0052] In an embodiment of the present disclosure, the audio separation system 200 can extract background sound features h from the sound source 201 by processing the spectrum of the sound source 201 using the background sound analysis module 220. i c In an embodiment of the present disclosure, the background sound analysis module 220 may include a plurality of convolution blocks, each of which may include one or more convolution layers, batch normalization, and an activation function (e.g., ReLu or sigmoid). In an embodiment of the present disclosure, the background sound analysis module 220 may be implemented in the same structure as the audio feature extraction module 210. In the background sound analysis module 220, features extracted by one convolution block may be transferred to the next convolution block, and thus, background sound features h may be extracted by each convolution block. i c Here, i is a natural number, and h i c represents the background sound features extracted by the i-th convolutional block. Figure 2 As shown, in the case where the background sound analysis module 220 includes eight convolution blocks, the convolution blocks can respectively extract the background sound features corresponding to the convolution blocks (ie, h1 c 、h2 c 、……、h8 c ).

[0053] Background sound characteristics h i c can have any real number value and can be understood as indicating the corresponding background sound feature h i c The corresponding first audio feature h i For example, it can be understood that as the background sound feature h i c The size of the corresponding background sound feature h i c The corresponding first audio feature h i It contributes more to the generation of background sound, and as the background sound feature h i c The size of the corresponding background sound feature h i c The corresponding first audio feature h i The contribution to generating background sound is less. For example, in the background sound feature h1 c In the case of increase, the corresponding background sound feature h1 c The corresponding first audio feature h1 contributes more to the generation of background sound or may have a higher influence. cThe size of the corresponding background sound feature h1 c The corresponding first audio feature h1 contributes less to generating the background sound or may have a lower impact.

[0054] In an embodiment of the present disclosure, the audio separation system 200 may process the first audio feature h by using the background sound feature control module 230. i To generate the second audio feature h in which the background sound is adjusted i '. The following will refer to Figure 3a and Figure 3b The operation of the background sound feature control module 230 is described.

[0055] In an embodiment of the present disclosure, the audio separation system 200 can extract target features from the target information 203 by processing the target information 203 using the target feature extraction module 240. The target feature extraction module 240 can be implemented in a form suitable for processing the target information 203. For example, if the target information 203 is visual information, the target feature extraction module 240 can be a visual feature network (e.g., a convolutional neural network) capable of extracting visual features from the visual information. As another example, if the target information 203 is audio information, the target feature extraction module 240 can be a speech embedding network (e.g., a convolutional neural network) capable of generating an audio embedding from the audio information.

[0056] In an embodiment of the present disclosure, the audio separation system 200 can generate a spectrum of separated audio by processing the first audio feature, the second audio feature, and the target feature using the audio generation module 250. In an embodiment of the present disclosure, the audio generation module 250 may include a plurality of up-convolution blocks, each of which may include one or more transposed convolution layers, cascades, and convolution blocks. In the audio generation module 250, the features generated by one up-convolution block may be transferred to the next up-convolution block, and thus, a spectrum of separated audio may be generated. Each up-convolution block processes the features generated by the previous up-convolution block and the second audio feature h i ' to generate features to be passed to the next upconvolution block. Here, the j-th upconvolution block generates the first audio feature h extracted by using the features generated by the (j-1)-th upconvolution block and by processing the first audio feature h extracted by the convolution block corresponding to the j-th upconvolution block (for example, the (n-j+1)-th convolution block when the number of convolution blocks and upconvolution blocks is n). i The generated second audio feature h i ' to generate features to be passed to the next up-convolution block. The first up-convolution block generates the first audio feature h extracted by the last convolution block by using iFeatures obtained by concatenating with the target feature and second audio features generated by processing the first audio features extracted by the last convolution block generate features to be transferred to the second up-convolution block.

[0057] Where the first audio feature h generated by the corresponding convolution block i Generate the second audio feature h i The operation of transferring the background sound directly from the convolution block to the corresponding upconvolution block can be understood as a skip connection. In other words, the audio separation system 200 can adjust the amplitude of the background sound included in the separated audio by adjusting the size of the features directly transferred from the convolution block to the upconvolution block via the background sound analysis module 220 and the background sound feature control module 230.

[0058] In an embodiment of the present disclosure, the audio separation system 200 may generate separated audio by applying an inverse STFT (ISTFT) to the spectrum of the separated audio.

[0059] Figure 3a is a diagram for describing the operation of the background sound feature control module 230 according to an embodiment of the present disclosure, and Figure 3b An example of a background sound control function is shown.

[0060] like Figure 3a As shown, in an embodiment of the present disclosure, the background sound feature control module 230 may be based on the background sound control parameter 202 (eg, α) and the background sound feature h i c , by using the first audio feature h i The degree of association with background sound scales the first audio feature h i To generate the second audio feature h in which the background sound is adjusted i '. Here, the first audio feature h i The degree of association with background sound can be understood as the first audio feature h in the case where the separated audio is generated by the audio generation module 250. i The contribution of background sounds to include in the separated audio.

[0061] For example, in the case where the background sound control parameter 202 is 1 (α=1), this means that the user does not want to adjust the background sound, and therefore, the background sound feature control module 230 may not scale the first audio feature h i , so that the first audio feature h i is transmitted to the up-convolution block as is. In other words, the background sound feature control module 230 may scale the first audio feature h by 1. i .

[0062] For example, in the case where the background sound control parameter 202 is 0 (α=0), this means that the user wants to remove the background sound, and therefore, as the background sound feature h i c As the size of the first audio feature h increases, the background sound feature control module 230 can scale the first audio feature h by a smaller scaling factor. i , so that the first audio feature h contributes a lot to generating background sound i The size of is reduced and then passed to the up-convolution block. In other words, as the background sound feature h i c As the size of increases, the scaling factor can approach 0.

[0063] For example, in the case where the background sound control parameter 202 is 0 (α=0), this means that the user wants to remove the background sound, and therefore, as the background sound feature h i c As the size of the first audio feature h decreases, the background sound feature control module 230 may scale the first audio feature h with a larger scaling factor. i , so that the first audio feature h contributes little to generating background sound i is transmitted to the up-convolution block as completely as possible. In other words, as the background sound feature h i c As the size of is reduced, the scaling factor can be close to 1.

[0064] In an embodiment of the present disclosure, the background sound feature control module 230 can scale the first audio feature h according to Equation 1. i To generate the second audio feature h in which the background sound is adjusted i '.

[0065] Equation 1

[0066]

[0067] Here, h i represents the first audio feature, h i c is the background sound feature, h i ' represents the second audio feature, α represents the background sound control parameter, and f(x) represents the background sound control function.

[0068] like Figure 3b As shown, in an embodiment of the present disclosure, the background sound control function f(x) may be any function that is symmetric about x=0 (i.e., symmetric about the y-axis) and approaches 0 as the absolute value of x (i.e., |x|) increases, and satisfies f(0)=1. For example, the background sound control function f(x) may be the function of Equation 2.

[0069] Equation 2

[0070]

[0071] When the background sound control parameter 202 is 1 (α=1), h i '=f(0)×h i =h i , and therefore, the background sound feature control module 230 does not scale the first audio feature h i In other words, the background sound feature control module 230 scales the first audio feature h by 1. i .

[0072] When the background sound control parameter 202 is close to 0 (α is close to 0) and |h i c | is large enough, f(h i c ×(1-α)) is close to 0, and therefore, the background sound feature control module 230 scales the background sound feature h with a scaling factor close to 0. i c The corresponding first audio feature h i .

[0073] In|h i c | is sufficiently small, even when the background sound control parameter 202 is close to 0 (α is close to 0), (h i c ×(1-α)) is also close to 1, and therefore, the background sound feature control module 230 scales the background sound feature h with a scaling factor close to 1. i c The corresponding first audio feature h i .

[0074] Figure 4 is a diagram for describing a method of training the audio separation system 200 according to an embodiment of the present disclosure.

[0075] In an embodiment of the present disclosure, training of the audio separation system 200 can be performed by comparing a target result 405 with an audio separation result 404 inferred by the audio separation system 200 for a training sound source 401. Training can be understood herein as updating weight values ​​in the audio feature extraction module 210, the background sound analysis module 220, the background sound feature control module 230, and the audio generation module 250 included in the audio separation system 200. A specific loss function (e.g., mean square error or cross-entropy error) can be used to compare the inferred audio separation result 404 with the target result 405.

[0076] In an embodiment of the present disclosure, the training sound source 401 may include one or more training candidate audios.

[0077] In the embodiment of the present disclosure, the training candidate audio can be determined according to the purpose of the audio separation system 200. For example, Figure 1 As shown, in the case where the goal is to separate the voice of each character in a movie in order to adjust the volume of each character's voice and background sound, the training candidate audio can be determined to be human voice. As another example, in the case where the goal is to separate the singing voice of a singer and the accompaniment of the song sung by the singer in a video of the singer's live performance including the shouts of the audience, the training candidate audio can be determined to be human voice and sounds produced by musical instruments.

[0078] In an embodiment of the present disclosure, the training sound source 401 may include training background sound.

[0079] In an embodiment of the present disclosure, the training background sound may be determined depending on the purpose of the audio separation system 200. For example, if the purpose is to separate a singer's singing voice in a video of the singer's live performance, the training background sound may be determined as a sound produced by a musical instrument.

[0080] In the embodiment of the present disclosure, the training background sound may be any noise, for example, white noise.

[0081] In an embodiment of the present disclosure, training sound source 401 may be generated by mixing one or more training candidate audios with training background sounds at a specific volume ratio. For example, training sound source 401 may be a sound source in which the speech of two speakers is mixed with white noise at a volume ratio of 2:1. However, the present disclosure is not limited thereto, and therefore, the ratio may be different from 2:1.

[0082] In embodiments of the present disclosure, training target information 403 may be determined to correspond to candidate training audio. For example, if the candidate training audio is human speech, training target information 403 may be an image that includes all or part of the speaker's face (e.g., the mouth area), or pre-recorded speech of the speaker. As another example, if the candidate training audio is the sound of a musical instrument, training target information 403 may be a pre-recorded sound produced by the instrument.

[0083] In an embodiment of the present disclosure, the audio separation system 200 may be trained for multiple situations corresponding to multiple background sound parameters 402. Here, a separate target result may be used for each situation. Figure 4As shown, the audio separation system 200 can be trained for a first case where the background sound parameter 402 is 0 (e.g., α=0), a second case where the background sound parameter 402 is 0.5 (e.g., α=0.5), and a third case where the background sound parameter 402 is 1 (e.g., α=1). Here, the target result 405 for the first case can be a training candidate audio with the background sound removed. , the target result 405 in the second case can be the training candidate audio with the background sound removed and the average of the training candidate audio x including background sound , the target result 405 of the third case may be the training candidate audio x including background sound.

[0084] In an embodiment of the present disclosure, a training sound source 401 may be generated by mixing random noise with original audio including one or more candidate audios, and a target result 405 may be generated by removing noise from the training sound source 401 using a denoising model. For example, the denoising model may be an artificial intelligence model.

[0085] Figure 5 An audio separation system 500 according to an embodiment of the present disclosure is shown, and Figure 6 Detailed description is provided for describing the operation of the background sound feature control module 530 according to an embodiment of the present disclosure.

[0086] When separating one or more audios from a sound source 501 including multiple audios, the candidate audios and background sounds can be determined based on the design of the training method (e.g., the setting of training candidate audios, training background sounds, or target results) and the audio separation system (e.g., the setting of the network structure or candidate audios and target information). In addition, depending on the design of the training method and the audio separation system, the background sounds can be classified into two or more types. For example, for a sound source including a singer's singing voice, the accompaniment of a song, and ambient noise (e.g., the shouting or applause of the audience), the singer's singing voice and the accompaniment of the song can be set as candidate audios, and the ambient noise can be set as background sounds. Alternatively, the singer's singing voice can be set as candidate audio, and the accompaniment of the song and the ambient noise can be set as background sounds. For example, in the case where the accompaniment of the song and the ambient noise are set as background sounds, Figure 5 The illustrated audio separation system 500 may be used to obtain an audio separation result in which two or more background sounds may be adjusted independently.

[0087] and Figure 2 Compared with the audio separation system 200, Figure 5 The audio separation system 500 is configured to receive background sound control parameters 502 (β) as an additional input and includes an additional background sound feature extraction module 520B. Figure 5The audio separation system 500 is an example of a case where two background sounds can be adjusted, but is not limited thereto. It will be apparent to those skilled in the art that the audio separation system 500 can be modified to receive more background sound control parameters and include more background sound feature extraction modules to adjust more background sounds. The operation of the audio separation system 500 will be described below without repeating the operation of the audio separation system 200 described above.

[0088] In an embodiment of the present disclosure, audio separation system 500 may receive input including a sound source 501 including one or more candidate audios (Audio #1, ..., Audio #N) and background sound, target information 503 corresponding to each candidate audio, and background sound control parameters 502 (e.g., α and β), and output separated audios 504 and 505, wherein the amplitude of the background sound is adjusted according to background sound control parameters 502 (e.g., α and β). Audio separation system 500 may be an artificial intelligence model trained using training data.

[0089] The background sound control parameter α is a parameter used to determine the volume of the first background sound to be included in the separated sound sources 504 and 505. The background sound control parameter α can be determined by the user. In an embodiment of the present disclosure, the background sound control parameter α can be a real number between 0 and 1. For example, the audio separation system 500 can output the separated audio 504 with the first background sound removed based on the background sound control parameter α being 0, and output the separated audio 505 including the first background sound (for example, the background sound is not removed) based on the background sound control parameter α being 1. In other words, the audio separation system 500 can remove more of the first background sound as the background sound control parameter α approaches 0, and can remove less of the first background sound as the background sound control parameter α approaches 1.

[0090] The background sound control parameter β is a parameter for determining the volume of the second background sound to be included in the separated sound sources 504 and 505. The background sound control parameter β can be determined by the user. In an embodiment of the present disclosure, the background sound control parameter β can be a real number between 0 and 1. For example, the audio separation system 500 can output the separated audio 504 with the second background sound removed based on the background sound control parameter β being 0, and output the separated audio 505 including the second background sound based on the background sound control parameter β being 1. In other words, the audio separation system 500 can further remove the second background sound as the background sound control parameter β approaches 0, and can remove less of the second background sound as the background sound control parameter β approaches 1.

[0091] In an embodiment of the present disclosure, the audio separation system 500 may extract the first background sound feature h from the sound source 501 by processing the spectrum of the sound source 501 using the first background sound analysis module 520A. i c1 In an embodiment of the present disclosure, the first background sound analysis module 520A may include a plurality of convolution blocks, each of which may include one or more convolution layers, batch normalization, and an activation function (e.g., ReLu or sigmoid). In an embodiment of the present disclosure, the first background sound analysis module 520A may be implemented in the same structure as the audio feature extraction module 510. In the first background sound analysis module 520A, features extracted by one convolution block may be transferred to the next convolution block, and thus, the first background sound feature h may be extracted by each convolution block. i c1 Here, i is a natural number, and h i c1 represents the first background sound feature extracted by the i-th convolutional block. Figure 5 As shown, in the case where the first background sound analysis module 520A includes eight convolution blocks, the convolution blocks can respectively extract the first background sound features corresponding thereto (ie, h1 c1 、h2 c1 、……、h8 c1 ).

[0092] The first background sound feature h i c1 can have any real number value and can be understood as indicating the first audio feature h corresponding to it i That is, it can be understood that as the first background sound feature h i c1 The size of the first audio feature h i The contribution to generating the first background sound is greater, and as the first background sound feature h i c1 The size of the first audio feature h is reduced, and the corresponding i The contribution to the generation of the first background sound is smaller.

[0093] In an embodiment of the present disclosure, the audio separation system 500 may extract the second background sound feature h from the sound source 501 by processing the spectrum of the sound source 501 using the second background sound analysis module 520B. i c2In an embodiment of the present disclosure, the second background sound analysis module 520B may include a plurality of convolution blocks, each of which may include one or more convolution layers, batch normalization, and an activation function (e.g., ReLu or sigmoid). In an embodiment of the present disclosure, the second background sound analysis module 520B may be implemented in the same structure as the audio feature extraction module 510. In the second background sound analysis module 520B, features extracted by one convolution block may be transferred to the next convolution block, and thus, second background sound features h may be extracted by each convolution block. i c2 Here, i is a natural number, and h i c2 represents the second background sound feature extracted by the i-th convolutional block. Figure 5 As shown, in the case where the second background sound analysis module 520B includes eight convolution blocks, the convolution blocks can respectively extract the second background sound features corresponding thereto (ie, h1 c2 、h2 c2 、……、h8 c2 ).

[0094] The second background sound feature h i c2 can have any real number value and can be understood as indicating the first audio feature h corresponding to it i That is, it can be understood that as the second background sound feature h i c2 The size of the first audio feature h i The contribution to generating the second background sound is greater, and as the second background sound feature h i c2 The size of the first audio feature h is reduced, and the corresponding i The contribution to generating the second background sound is smaller.

[0095] In an embodiment of the present disclosure, the audio separation system 500 can process the first audio feature h by using the background sound feature control module 530. i To generate a second audio feature h in which the first background sound and the second background sound are adjusted i '.

[0096] like Figure 6 As shown, the background sound feature control module 530 can be based on the background sound control parameters α and β, the first background sound feature h i c1 and the second background sound feature h i c2 , by using the first audio feature h iThe first audio feature h is scaled by the degree associated with the first background sound and the second background sound, respectively. i , to generate a second audio feature h in which the first background sound and the second background sound are adjusted i '. Here, the first audio feature h i The degree of being associated with the first background sound and the second background sound, respectively, can be understood as the first audio feature h being associated with the first background sound and the second background sound, respectively, when the separated audio is generated by the audio generation module 550. i Contributions to the first background sound and the second background sound to be included in the separated audio.

[0097] In an embodiment of the present disclosure, the background sound feature control module 530 can scale the first audio feature h according to Equation 3. i To generate a second audio feature h in which the first background sound and the second background sound are adjusted i '.

[0098] Equation 3

[0099]

[0100] Here, h i represents the first audio feature, h i c1 represents the first background sound feature, h i c2 represents the second background sound feature, h i ' represents the second audio feature, α and β represent background sound control parameters, and f(x) represents the background sound control function.

[0101] In an embodiment of the present disclosure, the background sound control function f(x) may be any function that is symmetric about x=0 (i.e., symmetric about the y-axis) and approaches 0 as the absolute value of x (i.e., |x|) increases and satisfies f(0)=1. For example, the background sound control function f(x) may be a function of Equation 2, such as Figure 3b shown.

[0102] Figure 7 is a flowchart of an audio separation method 700 according to an embodiment of the present disclosure.

[0103] Figure 7 The audio separation method 700 is used to separate one or more candidate audios from a sound source including one or more candidate audios and background sounds by using the audio separation system 200 or 500. The audio separation method 700 may be Figure 8 Executed by electronic device 800.

[0104] In operation 710, the method may include extracting a first audio feature h from a sound source.i .

[0105] In an embodiment of the present disclosure, operation 710 may include generating a spectrum of the sound source by applying STFT to the sound source, and processing the spectrum of the sound source by using the audio feature extraction module 210 or 510 including a plurality of convolution blocks to extract a first audio feature h from the sound source. i .

[0106] In operation 720, the method may include extracting background sound features h from the sound source. i c .

[0107] In an embodiment of the present disclosure, operation 720 may include generating a spectrum of the sound source by applying STFT to the sound source, and processing the spectrum of the sound source by using the background sound analysis module 220, 520A, or 520B including a plurality of convolution blocks to extract background sound features h from the sound source. i c .

[0108] In an embodiment of the present disclosure, operation 720 may include extracting a first background sound feature h from a sound source. i c1 , and extracting the second background sound feature h from the sound source i c2 .

[0109] In operation 730, the method may include: i and background sound to generate the second audio feature h i '. For example, based on the background sound control parameter α and the background sound feature h i c , according to the first audio feature h i The first audio feature h is processed based on the degree of association with the background sound i , to generate a second audio feature h in which the background sound is adjusted i '.

[0110] In an embodiment of the present disclosure, operation 730 may include: i c Determine a scaling factor, and scale the first audio feature h by using the scaling factor i To generate the second audio feature h in which the background sound is adjusted i '.

[0111] In the embodiment of the present disclosure, based on the background sound control parameter α and the background sound feature h i cDetermining the scaling factor may include determining the scaling factor to be 1 based on the background sound control parameter α being 1, and determining the scaling factor to be 0 based on the background sound control parameter α being 0. i c As the size of the background sound feature h increases, the scaling factor is determined to be a smaller number, and the background sound control parameter α is 0. i c As the size of is reduced, the scaling factor is determined to be a larger number. Here, the scaling factor is greater than 0 but less than or equal to 1.

[0112] In the embodiment of the present disclosure, based on the background sound control parameter α and the background sound feature h i c When determining the scaling factor, you can use Determine the scaling factor. Here, f(x) satisfies f(0)=1, is symmetrical about x=0, and is a background sound control function that approaches 0 as the absolute value of x increases.

[0113] In an embodiment of the present disclosure, the background sound control function may be the function of Equation 2.

[0114] In an embodiment of the present disclosure, operation 730 may include: determining the background sound feature h based on the first background sound control parameter α, the second background sound control parameter β, and the first background sound feature h. i c1 and the second background sound feature h i c2 Determine a scaling factor, and scale the first audio feature h by using the scaling factor i To generate the second audio feature h in which the background sound is adjusted i '.

[0115] In the embodiment of the present disclosure, based on the first background sound control parameter α, the second background sound control parameter β, the first background sound feature h i c1 and the second background sound feature h i c2 Determining the scaling factor may include based on Determine the scaling factor. Here, f(x) satisfies f(0)=1, is symmetrical about x=0, and is a background sound control function that approaches 0 as the absolute value of x increases.

[0116] In operation 740 , the method may include generating one or more separated audios by using target information corresponding to the candidate audio, the first audio feature, and the second audio feature in which the background sound is adjusted.

[0117] In an embodiment of the present disclosure, operation 740 may include extracting a target feature from target information corresponding to the candidate audio by using a target feature extraction module 240, processing the target feature, the first audio feature, and the second audio feature by using an audio generation module 250 including a plurality of up-convolution blocks to generate a spectrum of the separated audio, and generating the separated audio by applying ISTFT to the spectrum of the separated audio.

[0118] Figure 8 8 is a block diagram schematically illustrating a configuration of an electronic device 800 for performing an audio separation method according to an embodiment of the present disclosure.

[0119] Figure 8 The electronic device 800 shown in the figure can be a display device (e.g., a smartphone, tablet, or television) for reproducing video, or can be a separate server connected to the display device via wired or wireless communication. The audio separation method described in the present disclosure can be performed by the display device for reproducing video, can be performed by a separate server connected to the display device, or can be performed by the display device and the server in collaboration with each other.

[0120] refer to Figure 8 In an embodiment of the present disclosure, the electronic device 800 may include a communication interface 810, an input / output interface 820, an audio output unit 830, a processor 840, and a memory 850. However, the components of the electronic device 800 are not limited to the above examples, and the electronic device 800 may include more or fewer components than the above components. In an embodiment of the present disclosure, some or all of the communication interface 810, the input / output interface 820, the audio output unit 830, the processor 840, and the memory 850 may be implemented as a single chip, and the processor 840 may include one or more processors.

[0121] For example, the electronic device 800 may receive a sound source including one or more candidate audio signals and background sounds from an external device via the communication interface 810. Here, the sound source may be received in a form including only an audio signal, or may be received in a form included in a video. The received sound source may be stored in the memory 850. At the same time, the electronic device 800 may separate the one or more candidate audio signals from the sound sources previously stored in the memory 850.

[0122] In addition, the electronic device 800 may receive an input for setting the volume of the background sound from the user through the input / output interface 820. In this case, the processor 840 may determine the background sound control parameters α and β based on the input volume of the background sound.

[0123] In addition, the electronic device 800 may separate one or more candidate audios from a sound source using the audio separation system 200 or 500 stored in the memory 850 through the processor 840 .

[0124] For example, upon completion of audio separation, the electronic device 800 may output separated audio (Audio#1, . . . , Audio#N) in which background sound is adjusted through the audio output unit 830. Here, the user may adjust the volume of each separated audio separately.

[0125] In an embodiment of the present disclosure, the communication interface 810 is a component for transmitting and receiving signals (e.g., control commands or data) to and from external devices in a wired or wireless manner, and may include a communication chipset that supports various communication protocols. The communication interface 810 may receive signals from the outside and output them to the processor 840, or transmit signals output from the processor 840 to the outside. According to an embodiment of the present disclosure, the communication interface 810 may receive sound sources including one or more candidate audio and background sounds from the outside.

[0126] In an embodiment of the present disclosure, the input / output interface 820 may include an input interface (e.g., a touch screen, hard buttons, or a microphone) for receiving control commands or information from a user, and an output interface (e.g., a display panel) for indicating the result of an operation performed according to the user's control or the status of the electronic device 800. According to an embodiment of the present disclosure, the input / output interface 820 may display a video being reproduced and receive input from the user for adjusting the volume of background sound included in a sound source.

[0127] In an embodiment of the present disclosure, the audio output unit 830 is a component for outputting an audio signal, and may be an output device (e.g., a built-in speaker) built into the electronic device 800 to directly reproduce a sound corresponding to the audio signal, may be an interface (e.g., a 3.5mm port, a 4.4mm port, an RCA port, a USB port) for allowing the electronic device 800 to send and receive audio signals to a wired audio reproduction device (e.g., a speaker, a sound bar, earphones, or a headset), or may be an interface (e.g., a Bluetooth module or a wireless local area network (WLAN) module) for allowing the electronic device 800 to send and receive audio signals to a wireless audio reproduction device (e.g., a wireless earphone, a wireless headset, or a wireless speaker).

[0128] In an embodiment of the present disclosure, the processor 840 is a component configured to control a series of processes to enable the electronic device 800 to operate, and may include one or more processors. In this case, the one or more processors may be a general-purpose processor (such as a central processing unit (CPU), an application processor (AP), or a digital signal processor (DSP)), a dedicated graphics processor (such as a graphics processing unit (GPU) or a visual processing unit (VPU)), or a dedicated artificial intelligence processor (such as a neural processing unit (NPU)). For example, in the case where the one or more processors are dedicated artificial intelligence processors, the dedicated artificial intelligence processor may be designed to have a hardware structure specifically for processing a specific artificial intelligence model.

[0129] In an embodiment of the present disclosure, the processor 840 may write data to the memory 850 or read data stored in the memory 850, and specifically, may execute a program stored in the memory 850 to process data according to predefined operating rules or artificial intelligence models. Therefore, the processor 840 may perform the operations described herein, and unless otherwise specified, the operations described herein to be performed by the electronic device 800 may be performed by the processor 840.

[0130] In embodiments of the present disclosure, the memory 850 is a component for storing various programs or data and may include storage media such as read-only memory (ROM), random access memory (RAM), a hard disk, a compact disc ROM (CD-ROM), or a digital video disk (DVD), or a combination of these media. The memory 850 may not be a separate component and may be included in the processor 840. The memory 850 may include volatile memory, nonvolatile memory, or a combination of volatile and nonvolatile memory. The memory 850 may store programs for performing operations according to the embodiments described herein. For example, the memory 850 may store programs corresponding to an audio separation system. The memory 850 may provide the processor 840 with the data stored therein in response to a request from the processor 840.

[0131] Figure 9a 、 Figure 9b 、 Figure 10a 、 Figure 10b and Figure 10c 2 is a diagram for describing an application example of the audio separation system 200 or 500 according to an embodiment of the present disclosure.

[0132] Figure 9aA scene depicts a video captured by an electronic device including a camera (e.g., a smartphone or tablet). Assume that the video includes not only the voice of a person 911 in the video but also the voice of a camera operator 912, who is not in the video, such as "I'm filming," and ambient noise such as the voices of others around person 911. Audio separation system 200 can be used to separate the voice of person 911 in the video from the voice of camera operator 912, while also adjusting the volume of the ambient noise included in the separated voices.

[0133] like Figure 9b As shown, audio separation system 200 can separate the voice of person 911 and the voice of cameraman 912 from the sound source by processing the voice of person 911 and the voice of cameraman 912 as candidate audio and processing the ambient noise as background sound. Here, an image including the face of person 911 and the pre-recorded voice of cameraman 912 can be used as target information for the voice of person 911 and the voice of cameraman 912. The background sound control parameter α can be determined based on the volume of the ambient noise adjusted by the user through volume adjustment interface 923.

[0134] In the case where the user adjusts the volume of the voice of the person 911 , the voice of the cameraman 912 , and the ambient noise respectively through the volume adjustment interfaces 921 , 922 , and 923 , the volume of each sound may be adjusted individually.

[0135] Figure 10a A scene of a video of a singer's live performance is shown. Assume that the performance video contains the singing voice of singer 1010, the accompaniment of the song sung by singer 1010, and the shouts of the audience.

[0136] like Figure 10b As shown, the audio separation system 200 can be used to separate the singing voice of the singer 1010 and the accompaniment of the song in the performance video, and adjust the volume of the shouting of the audience to be included in the separated sounds.

[0137] The audio separation system 200 can separate the singer's 1010 singing voice and the song's accompaniment from the sound source by processing the singer's 1010 singing voice and the song's accompaniment as candidate audio and processing the audience's shouting as background sound. Here, an image including the singer's 1010 face and pre-recorded music can be used as target information for the singer's 1010 singing voice and the song's accompaniment, respectively. The background sound control parameter α can be determined based on the volume of the audience's shouting adjusted by the user through the volume adjustment interface 1023.

[0138] At the same time, if Figure 10c As shown, the audio separation system 500 can be used to separate only the singing voice of the singer 1010 in the performance video, and individually adjust the volume of the accompaniment of the song and the shouts of the audience included in the separated voice.

[0139] The audio separation system 500 can separate the singer's 1010 singing voice from the sound source by processing the singer's 1010 singing voice as the candidate audio, the song's accompaniment as the first background sound, and the audience's shouting as the second background sound. Here, an image including the singer's 1010 face can be used as target information for the singer's 1010 singing voice. The background sound control parameter α can be determined based on the volume of the song's accompaniment adjusted by the user through the volume adjustment interface 1022, and the background sound control parameter β can be determined based on the volume of the song's accompaniment adjusted by the user through the volume adjustment interface 1023.

[0140] In the case where the user adjusts the volume of the singer's 1010 singing voice, the accompaniment of the song, and the audience's shouting respectively through the volume adjustment interfaces 1021 , 1022 , and 1023 , the volume of each sound can be adjusted individually.

[0141] According to an embodiment of the present disclosure, a method for separating one or more candidate audios from a sound source including one or more candidate audios and background sounds by using an audio separation system includes: extracting a first audio feature from the sound source, extracting a background sound feature from the sound source, the background sound feature identifying a degree of association between the first audio feature and the background sound, generating a second audio feature based on the first audio feature, the background sound feature and a background sound control parameter, the background sound control parameter being configured to control the background sound, and generating one or more separated audios based on target information corresponding to the one or more candidate audios, the first audio feature and the second audio feature in which the background sound is adjusted.

[0142] In an embodiment, extracting the first audio feature from the sound source may include generating a spectrum of the sound source by applying a short-time Fourier transform (STFT) to the sound source, and processing the spectrum of the sound source by using an audio feature extraction module including a plurality of convolution blocks to extract the first audio feature from the sound source.

[0143] In an embodiment, extracting background sound features from a sound source may include generating a spectrum of the sound source by applying a short-time Fourier transform (STFT) to the sound source, and processing the spectrum of the sound source by using a background sound analysis module including a plurality of convolution blocks to extract background sound features from the sound source.

[0144] In an embodiment, the method may include obtaining a scaling factor based on the background sound control parameter and the background sound feature, and generating the second audio feature by scaling the first audio feature using the scaling factor.

[0145] In an embodiment, obtaining a scaling factor based on a background sound control parameter and a background sound feature may include obtaining a scaling factor of 1 based on the background sound control parameter being 1, obtaining a scaling factor of a smaller number as the size of the background sound feature increases based on the background sound control parameter being 0, and obtaining a scaling factor of a larger number as the size of the background sound feature decreases, wherein the scaling factor is greater than 0 but less than or equal to 1.

[0146] In an embodiment, obtaining the scaling factor based on the background sound control parameter and the background sound feature may include obtaining the scaling factor according to a background sound control function f(x), which is symmetric about x=0, approaches 0 as the absolute value of x increases, and satisfies f(0)=1.

[0147] The background sound control function is .

[0148] In an embodiment, obtaining a scaling factor based on the background sound control parameter and the background sound feature includes: Get the scaling factor, where α is the background sound control parameter, and h i c It is the background sound feature.

[0149] In an embodiment, the method includes extracting a target feature from target information corresponding to one or more candidate audios by using a target feature extraction module, processing the target feature, the first audio feature, and the second audio feature by using an audio generation module including a plurality of up-convolution blocks to generate a spectrum of one or more separated audios, and generating the one or more separated audios by applying an inverse STFT (ISTFT) to the spectrum of the one or more separated audios.

[0150] In an embodiment, with respect to a training sound source including one or more training candidate audios and training background sounds, the audio separation system can be trained by comparing the audio separation results inferred by the audio separation system with target results. Training can be performed for multiple cases corresponding to multiple background sound parameters, and a separate target result can be used for each case.

[0151] In an embodiment, extracting the background sound feature from the sound source may include: extracting a first background sound feature from the sound source; and extracting a second background sound feature from the sound source.

[0152] In an embodiment, generating the second audio feature based on the first audio feature may include: obtaining a scaling factor based on the first background sound control parameter, the second background sound control parameter, the first background sound feature, and the second background sound feature, and generating the second audio feature in which the background sound is adjusted by scaling the first audio feature using the scaling factor.

[0153] In an embodiment, obtaining the scaling factor based on the first background sound control parameter, the second background sound control parameter, the first background sound feature, and the second background sound feature comprises: Get the scaling factor, where α is the first background sound control parameter, β is the second background sound control parameter, and h i c1 is the first background sound feature, and h i c2 It is the second background sound feature.

[0154] In an embodiment, the target information corresponding to the one or more candidate audios may be one of visual information or audio information separated from the one or more candidate audios.

[0155] In an embodiment, each of the one or more candidate audios may be a human voice or a sound produced by a musical instrument or a machine. The background sound may include at least one of a human voice, environmental noise, or background music.

[0156] According to an embodiment, a computer-readable recording medium has a computer program recorded thereon. When the computer program is executed by one or more computing devices, it can cause the one or more computing devices to perform a method for separating one or more candidate audios in a sound source including one or more candidate audios and background sounds by using an audio separation system. The method includes: extracting a first audio feature from the sound source, extracting a background sound feature from the sound source, the background sound feature identifying a degree of correlation between the first audio feature and the background sound, generating a second audio feature based on the first audio feature, the background sound feature and a background sound control parameter, the background sound control parameter being configured to control the background sound, and generating one or more separated audios based on target information corresponding to the one or more candidate audios, the first audio feature and the second audio feature in which the background sound is adjusted.

[0157] According to an embodiment, an electronic device may include one or more processors and a memory storing a program for separating one or more candidate audios from a sound source including one or more candidate audios and background sounds by using an audio separation system. When the program is executed by the one or more processors, the electronic device may perform operations including: extracting a first audio feature from the sound source, extracting a background sound feature from the sound source, the background sound feature identifying a degree of association between the first audio feature and the background sound, generating a second audio feature based on the first audio feature, the background sound feature, and a background sound control parameter, the background sound control parameter being configured to control the background sound, and generating one or more separated audios based on target information corresponding to the one or more candidate audios, the first audio feature, and the second audio feature in which the background sound is adjusted.

[0158] While the present disclosure has been described with reference to the embodiments thereof, it will be apparent to those skilled in the art that various changes and modifications can be made therein without departing from the spirit and scope of the disclosure as set forth in the following claims.

Claims

1. A method for separating one or more candidate audios from a sound source including one or more candidate audios and background sounds by using an audio separation system, the method comprising: extracting a first audio feature from the sound source; extracting a background sound feature from the sound source, the background sound feature identifying a degree of correlation between the first audio feature and the background sound; generating a second audio feature based on the first audio feature, the background sound feature, and a background sound control parameter configured to control the background sound; as well as One or more separated audios are generated based on target information corresponding to the one or more candidate audios, the first audio feature, and the second audio feature in which the background sound is adjusted.

2. The method according to claim 1, wherein Extracting the first audio feature from the sound source includes: generating a spectrum of the sound source by applying a short-time Fourier transform (STFT) to the sound source, and The frequency spectrum of the sound source is processed by using an audio feature extraction module including a plurality of convolution blocks to extract the first audio feature from the sound source.

3. The method according to claim 1 or 2, wherein: Extracting the background sound feature from the sound source includes: generating a spectrum of the sound source by applying a short-time Fourier transform (STFT) to the sound source, and The frequency spectrum of the sound source is processed by using a background sound analysis module including a plurality of convolution blocks to extract the background sound features from the sound source.

4. The method according to any one of claims 1 to 3, further comprising: obtaining a scaling factor based on the background sound control parameter and the background sound characteristic, and The second audio feature is generated by scaling the first audio feature using the scaling factor.

5. The method according to any one of claims 1 to 4, wherein Obtaining the scaling factor based on the background sound control parameter and the background sound feature includes: Based on the background sound control parameter being 1, the scaling factor is obtained to be 1, Based on the background sound control parameter being 0, as the size of the background sound feature increases, the scaling factor is obtained to be a smaller number, and as the size of the background sound feature decreases, the scaling factor is obtained to be a larger number, and The scaling factor is greater than 0 but less than or equal to 1.

6. The method according to any one of claims 1 to 5, wherein Obtaining the scaling factor based on the background sound control parameter and the background sound feature includes: The scaling factor is obtained according to a background sound control function f(x), which is symmetric about x=0, approaches 0 as the absolute value of x increases, and satisfies f(0)=1.

7. The method according to any one of claims 1 to 6, wherein Obtaining the scaling factor based on the background sound control parameter and the background sound feature includes: according to To obtain the scaling factor, Wherein, α is the background sound control parameter, and h i c is the background sound feature.

8. The method according to any one of claims 1 to 7, further comprising: extracting target features from target information corresponding to the one or more candidate audios by using a target feature extraction module, generating a spectrum of the one or more separated audios by processing the target feature, the first audio feature, and the second audio feature using an audio generation module including a plurality of upconvolution blocks, and The one or more separated audio frequencies are generated by applying an inverse STFT (ISTFT) to the frequency spectra of the one or more separated audio frequencies.

9. The method according to any one of claims 1 to 8, wherein With respect to a training sound source including one or more training candidate audios and training background sounds, the audio separation system is trained by comparing an audio separation result inferred by the audio separation system with a target result, wherein training is performed for a plurality of situations corresponding to a plurality of background sound parameters, and Here, a separate target outcome is used for each case.

10. The method according to any one of claims 1 to 9, wherein Extracting the background sound feature from the sound source includes: extracting a first background sound feature from the sound source; and extracting a second background sound feature from the sound source; and Among them, generating the second audio feature based on the first audio feature includes: obtaining the scaling factor based on the first background sound control parameter, the second background sound control parameter, the first background sound feature and the second background sound feature; and generating the second audio feature in which the background sound is adjusted by scaling the first audio feature using the scaling factor.

11. The method according to any one of claims 1 to 10, wherein Obtaining the scaling factor based on the first background sound control parameter, the second background sound control parameter, the first background sound feature, and the second background sound feature includes: according to To obtain the scaling factor, Wherein, α is the first background sound control parameter, β is the second background sound control parameter, and h i c1 is the first background sound feature, and h i c2 It is the second background sound feature.

12. The method according to any one of claims 1 to 11, wherein The target information corresponding to the one or more candidate audios is one of visual information or audio information separated from the one or more candidate audios.

13. The method according to any one of claims 1 to 12, wherein Each of the one or more candidate audios is a human voice or a sound produced by a musical instrument or a machine, and The background sound includes at least one of human voice, environmental noise or background music.

14. A computer-readable recording medium having a computer program recorded thereon, wherein when the computer program is executed by one or more computing devices, the one or more computing devices are caused to execute the method according to any one of claims 1 to 13.

15. An electronic device (800), comprising: one or more processors (840); and A memory (850) storing a program for separating one or more candidate audios from a sound source including one or more candidate audios and background sounds by using an audio separation system. Wherein, when executed by the one or more processors (840), the program causes the electronic device (800) to perform operations including: extracting a first audio feature from the sound source; extracting a background sound feature from the sound source, the background sound feature identifying a degree of correlation between the first audio feature and the background sound; generating a second audio feature based on the first audio feature, the background sound feature, and a background sound control parameter configured to control the background sound; and One or more separated audios are generated based on target information corresponding to the one or more candidate audios, the first audio feature, and the second audio feature in which the background sound is adjusted.