Method for separating sound source from audio signal, and electronic device for performing same

The method and device automatically extract and separate sound sources from mixed audio signals by detecting single sound source segments and generating embeddings, enhancing accuracy and efficiency in complex environments.

WO2025183447A1PCT designated stage Publication Date: 2025-09-04SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/002663
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-21
Filing Date
2025-02-26
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing sound source separation technologies require prior input of specific sound source characteristics and struggle to accurately identify the separated sound sources in complex audio environments, limiting their efficiency and accuracy.

Method used

A method and electronic device that perform fast scanning on an audio signal to extract embeddings of primary sound sources without prior information, using a fast scanning module to detect single sound source segments, generate embeddings, and separate sound sources based on these embeddings, with additional audio signal input for enhanced accuracy.

Benefits of technology

Enables accurate and efficient separation of sound sources from mixed audio signals by identifying primary sound sources automatically, improving separation accuracy and applicability in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025002663_04092025_PF_FP_ABST
    Figure KR2025002663_04092025_PF_FP_ABST
Patent Text Reader

Abstract

This method for separating a sound source from an audio signal may comprise the steps of: acquiring an audio signal including a sound in which a plurality of sound sources are generated; acquiring, on the basis of a single sound source segment including only a sound in which one from among the plurality of sound sources is generated, an embedding corresponding to at least one primary sound source from among the plurality of sound sources; and separating the at least one primary sound source from the audio signal on the basis of the acquired embedding.
Need to check novelty before this filing date? Find Prior Art

Description

Method for separating a sound source from an audio signal and an electronic device for performing the same

[0001] The present disclosure relates to a method for separating sound sources from an audio signal containing sounds generated by multiple sound sources, and more particularly, to a method for quickly extracting features of a main sound source when an audio signal is given, and performing segmentation on the main sound source based on the extracted features.

[0002] Technologies that separate multiple sound sources from an audio signal are used in diverse fields, including speech recognition, noise reduction, and music analysis. These technologies require prior input of the characteristics of a specific sound source (e.g., timbre, frequency pattern, etc.) to isolate the sound. However, these technologies face the problem of not being able to clearly identify which sound source the separated sound corresponds to.

[0003] These limitations hinder the accuracy and efficiency of sound source separation, limiting its applicability in complex audio environments. Therefore, there is a need for technology that can separate multiple sound sources from an audio signal without prior information and clearly identify the relationship between the separated sounds and the sound sources.

[0004] According to one aspect of the present disclosure, a method for separating a sound source from an audio signal may include the steps of: obtaining an audio signal including sounds generated by a plurality of sound sources; obtaining an embedding corresponding to at least one primary sound source among the plurality of sound sources based on a single sound source segment including a first sound generated by one of the plurality of sound sources; and separating the at least one primary sound source from the audio signal based on the obtained embedding.

[0005] According to one aspect of the present disclosure, an electronic device includes a memory storing a program or at least one instruction and at least one processor operably coupled to the memory, wherein the at least one processor executes the program stored in the memory or the at least one instruction, thereby enabling the electronic device to obtain an audio signal including sounds generated by a plurality of sound sources, obtain an embedding corresponding to at least one primary sound source among the plurality of sound sources based on a single sound source segment including a first sound generated by one of the plurality of sound sources, and then separate the at least one primary sound source from the audio signal based on the obtained embedding.

[0006] According to one aspect of the present disclosure, a computer-readable recording medium may have stored thereon a program for executing at least one of the embodiments of the disclosed method on a computer.

[0007] According to one aspect of the present disclosure, a computer program may be stored on a medium for performing at least one of the embodiments of the disclosed method on a computer.

[0008] The above and other aspects, features and advantages of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings.

[0009] FIG. 1 illustrates modules included in an electronic device for performing a sound source separation process according to one embodiment of the present disclosure.

[0010] FIG. 2 illustrates detailed configurations included in a fast scanning module according to one embodiment of the present disclosure.

[0011] FIG. 3 illustrates a process in which a fast scanning module according to one embodiment of the present disclosure performs fast scanning on an audio signal.

[0012] FIG. 4 illustrates detailed configurations included in a sound source separation module according to one embodiment of the present disclosure.

[0013] FIG. 5 illustrates a process in which a sound source separation module according to one embodiment of the present disclosure separates a sound source from an audio signal based on a fast scanning result.

[0014] FIG. 6 illustrates a method of connecting the sound of a target sound source to an audio signal to increase sound source separation accuracy according to one embodiment of the present disclosure.

[0015] FIG. 7 illustrates a method of adding the sound of a target sound source to an audio signal to increase sound source separation accuracy according to one embodiment of the present disclosure.

[0016] FIG. 8 illustrates modules included in an electronic device for performing a sound source separation process according to one embodiment of the present disclosure.

[0017] FIG. 9 illustrates a process of an electronic device according to one embodiment of the present disclosure matching a person in a video to a sound source separated from an audio signal.

[0018] FIG. 10 illustrates UI screens indicating sound source separation results displayed on a screen of an electronic device according to one embodiment of the present disclosure.

[0019] FIG. 11 illustrates components included in an electronic device according to one embodiment of the present disclosure.

[0020] FIGS. 12 to 18 illustrate a method of generating a digital zoom image using generative AI according to embodiments of the present disclosure.

[0021] In describing the present disclosure, descriptions of technical details that are well-known in the technical field to which the present disclosure pertains and are not directly related to the present disclosure may be omitted. This is to avoid obscuring the gist of the present disclosure by omitting unnecessary explanations and to convey the gist more clearly. Furthermore, the terms described below are defined based on their functions in the present disclosure and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the contents throughout this specification.

[0022] For the same reason, some components in the attached drawings are exaggerated, omitted, or schematically depicted. Furthermore, the dimensions of each component do not entirely reflect its actual size. Identical or corresponding components in each drawing are assigned the same reference numbers.

[0023] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. The disclosed embodiments are provided to ensure that the disclosure of the present disclosure is complete and to fully inform those skilled in the art of the present disclosure of the scope of the disclosure. An embodiment of the present disclosure may be defined according to the claims. Like reference numerals denote like elements throughout the specification. In addition, when describing an embodiment of the present disclosure, if a detailed description of a related function or configuration is determined to unnecessarily obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, the terms described below are terms defined in consideration of the functions of the present disclosure and may vary depending on the intention or custom of the user or operator. Therefore, the definitions should be made based on the contents throughout this specification.

[0024] In one embodiment, each block of the flowchart diagrams and combinations of the flowchart diagrams can be performed by computer program instructions. The computer program instructions can be installed on a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, and the instructions, when executed by the processor of the computer or other programmable data processing apparatus, can create means for performing the functions described in the flowchart block(s). The computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing apparatus to implement the functions in a particular manner, and the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). The computer program instructions can also be installed on a computer or other programmable data processing apparatus.

[0025] Additionally, each block in the flowchart diagram may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specified logical function(s). In one embodiment, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may be executed substantially simultaneously or, depending on the function, may be executed in reverse order.

[0026] The term '~unit' or '~module' used in one embodiment of the present disclosure may represent software or a hardware component such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the '~unit' or '~module' may perform a specific role. Meanwhile, the '~unit' or '~module' is not limited to software or hardware. The '~unit' or '~module' may be configured to be in an addressable storage medium and may be configured to play one or more processors. In one embodiment, the '~unit' or '~module' may include components such as software components, object-oriented software components, class components, and task components, processes, functions, properties, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided by a particular component or a particular "part" or "module" may be combined or separated into additional components to reduce their number. Furthermore, in one embodiment, a "part" or "module" may include one or more processors.

[0027] Below, the meanings of terms used in this disclosure are explained.

[0028] A "sound source" can be the source that generates sound or something corresponding to it. For example, a physical entity such as a speaker, a musical instrument, an animal, a machine, or the environment can be considered a sound source. In the present disclosure, a "sound source" is not limited to a physical source, but can also be considered a unit of analysis for distinguishing a specific sound in signal processing.

[0029] "Sound" refers to a physical wave generated from a sound source and transmitted through a medium (air, water, etc.), or its corresponding counterpart. Sound may manifest as a speaker's speech, the sound of a musical instrument, animal sounds, background noise, etc., and in the present disclosure, it may be sound data expressed as an audio signal or its corresponding counterpart. As a result of a sound source, sound may be subject to separation, analysis, or transformation during signal processing.

[0030] An 'audio signal' may be data converted from sound generated by a sound source into an electrical signal, or may be a corresponding data. In the present disclosure, an audio signal may include sounds generated from multiple sound sources.

[0031] "Sound source separation" can be the process of extracting or separating sounds corresponding to individual sound sources from a mixed audio signal, or its corresponding counterpart. For example, techniques for separating the voices of individual speakers from an audio signal containing a conversation between multiple speakers, or extracting the sounds of specific instruments from a music recording, can be considered sound source separation. Terms such as "source separation," "sound source extraction," "speech separation," or "speaker separation" may also be used instead of "sound source separation."

[0032] An "embedding" can be a vector representing the characteristics of a sound source, sound, or audio signal, or its corresponding vector. For example, an audio signal of a specific sound can be converted into a vector containing the characteristics of that sound, and this converted vector can be an embedding. Therefore, an embedding corresponding to a specific sound source can include the characteristics of that sound source (e.g., timbre, speech pattern, frequency pattern, etc.). Terms such as "feature vector" may also be used instead of "embedding."

[0033] "Fast scanning" refers to the task of selecting a primary sound source from an audio signal mixed with multiple sound sources and obtaining an embedding corresponding to the primary sound source, or its corresponding counterpart. In other words, fast scanning refers to the task of finding a single sound source segment in an audio signal and extracting the features of the sound contained in the single sound source segment, or its corresponding counterpart. For example, an electronic device can generate an embedding (i.e., a feature vector) by performing an embedding transformation on the sound contained in a single sound source segment. The embedding thus generated can include information about the features of the sound source.

[0034] A 'single sound source segment' may be a frame containing only sound generated from a single sound source among frames obtained by dividing an audio signal into a certain time length, or may correspond thereto. In other words, a single sound source segment may be a frame in which no sound from other sound sources exists during the time period, and only the sound from a specific sound source is active, or may correspond thereto. In the present disclosure, a single sound source segment may be used to extract characteristics of sound sources and select a main sound source in the fast scanning process described above. Instead of a 'single sound source segment', terms such as 'single sound source frame' or 'exclusive sound source segment' may also be used.

[0035] "Diarization" can be the process of representing active sections of a sound source based on the time axis, or its corresponding function. In other words, diarization can be the process of dividing an audio signal into time intervals and identifying each interval so that it corresponds to a specific sound source or speaker, or its corresponding function. For example, in an audio signal containing a conversation between multiple speakers, diarization can be the process of distinguishing when each speaker spoke by time interval. Instead of "diarization," terms such as "speaker diarization" or "sound source diarization" can also be used.

[0036] Embodiments of the present disclosure relate to a method for separating a sound source from an audio signal including sounds generated by a plurality of sound sources, wherein by performing fast scanning on the audio signal, an embedding corresponding to a primary sound source among the plurality of sound sources is obtained, and the primary sound source can be separated from the audio signal using the obtained embedding.

[0037] 1. Overall configuration and operation for performing the sound source separation process

[0038] FIG. 1 illustrates modules included in an electronic device for performing a sound source separation process according to one embodiment of the present disclosure. Referring to FIG. 1, the electronic device may include a fast scanning module (100) and a sound source separation module (200).

[0039] In some embodiments, the modules (100, 200) of FIG. 1 may be software configurations that are classified based on their functions or roles, and are implemented by the processor (1120) of the electronic device (1100) described below with reference to FIG. 11 executing a program or command stored in the memory (1130). In some embodiments, the modules (100, 200) of FIG. 1 may be virtual configurations for which no matching hardware device actually exists.

[0040] In other words, the operations performed by the processor (1120) of the electronic device (1100) of FIG. 11 by executing a program or command stored in the memory (1130) may be classified into a plurality of groups by function or purpose, and the subjects performing the operations included in each classified group may be expressed as the modules (100, 200) of FIG. 1.

[0041] Accordingly, the operations performed by the modules (100, 200) can be seen as actually being performed by the processor (1120) of the electronic device (1100) of FIG. 11 executing a program or command stored in the memory (1130).

[0042] These contents can be equally applied to the modules (110, 120, 130, 140) illustrated in FIG. 2, the modules (210, 220, 230) illustrated in FIG. 4, and the modules (100, 200, 800) illustrated in FIG. 8.

[0043] In one embodiment, an audio signal (10) may include sounds generated by multiple sound sources. For example, the audio signal (10) may include speech generated by multiple speakers. In some embodiments, the audio signal (10) may include sounds generated by various types of sound sources, such as musical instrument sounds or animal sounds, in addition to the voices of speakers. An electronic device according to one embodiment of the present disclosure may separate sounds included in the audio signal (10) by sound source.

[0044] When an audio signal (10) is input to a fast scanning module (100), the fast scanning module (100) can output embeddings (E1, E2, E3) corresponding to a primary sound source. The primary sound source may be a portion of a plurality of sound sources selected based on importance, or a corresponding portion thereof. A method for selecting a primary sound source, and further, a specific method for performing fast scanning, will be described in detail below with reference to FIGS. 2 and 3.

[0045] The sound source separation module (200) can separate main sound sources from an audio signal (10) based on the embeddings (E1, E2, E3) received from the fast scanning module (100). The sound source separation result (20) output from the sound source separation module (200) may include the following data.

[0046] - Information about the list of major sound sources

[0047] - Audio signal of sound corresponding to the main sound source

[0048] - Information about the section where the sound corresponding to the main sound source is activated

[0049] 2. Detailed configuration and operation of the fast scanning module

[0050] An electronic device according to one embodiment of the present disclosure, when given an audio signal (10), can select a primary sound source from among a plurality of sound sources by analyzing only a portion of the audio signal (10) and obtain an embedding corresponding to the primary sound source. This process of selecting a primary sound source and obtaining an embedding corresponding to the primary sound source by analyzing a portion of the audio signal (10) is referred to as fast scanning in the present disclosure. A process by which an electronic device according to one embodiment of the present disclosure performs fast scanning can be briefly summarized as follows.

[0051] - Detecting a single sound source section from an audio signal

[0052] - Convert sounds included in a single sound source section to embedding

[0053] - Classify embeddings by sound source by performing clustering

[0054] - Selecting some of the multiple sound sources as the main sound source

[0055] An electronic device according to one embodiment of the present disclosure can automatically and quickly obtain information about sound sources through fast scanning.

[0056] Previously, to isolate a specific sound source from an audio signal, information about that sound source had to be provided separately. However, an electronic device according to one embodiment of the present disclosure can perform sound source separation based on information about the sound source (e.g., embedding of the main sound source) acquired through fast scanning, even without providing information about the sound source to be separated.

[0057] Hereinafter, a process in which an electronic device according to one embodiment of the present disclosure performs fast scanning on an audio signal (10) will be described in detail with reference to FIGS. 2 and 3.

[0058] FIG. 2 illustrates detailed configurations included in a fast scanning module according to one embodiment of the present disclosure. FIG. 3 illustrates a process in which a fast scanning module according to one embodiment of the present disclosure performs fast scanning on an audio signal.

[0059] Referring to FIG. 2, as a non-limiting example, the fast scanning module (100) may include a single sound source section detection module (110), an embedding transformation module (120), a clustering module (130), and a main sound source selection module (140). The operation of each module will be described with reference to FIG. 3.

[0060] (1) Extracting embeddings from a single sound source section

[0061] A single sound source section detection module (110) can detect a single sound source section from an audio signal (10). As defined above, a single sound source section may be a section (e.g., frame) that contains only sound generated from a single sound source, or its corresponding section.

[0062] According to one embodiment of the present disclosure, an electronic device can process an audio signal (10) by dividing it into a plurality of frames. In FIG. 3, the length (window size) of each frame is set to 2 seconds, but in some embodiments, the length of the frame may be set to various lengths depending on given conditions or needs. In addition, the stride, which indicates the degree to which consecutive frames overlap, may also be set to various lengths (e.g., 1 second).

[0063] In Fig. 3, it is assumed that the single sound source section detection module (110) performs fast scanning on the first 10 seconds of the audio signal (10). The section where fast scanning is performed is divided into the first frame (F1) to the fifth frame (F5) in 2-second units. Fig. 3 illustrates the sounds included in the audio signal (10) separated by sound source for the section where fast scanning is performed.

[0064] The single sound source section detection module (110) can detect frames in which only the sound of one sound source is activated as single sound source sections by detecting whether sound sources are activated for each of the frames (F1, F2, F3, F4, F5). Since only the third sound source (SS3) is activated in the second frame (F2), the single sound source section detection module (110) can detect the second frame (F2) as a single sound source section. Similarly, since only the second sound source (SS2) is activated in the third frame (F3), and only the first sound source (SS1) is activated in the fourth frame (F4) and the fifth frame (F5), the single sound source section detection module (110) can detect the third frame (F3), the fourth frame (F4), and the fifth frame (F5) as single sound source sections.

[0065] When the independent sound source sections are detected, the embedding transformation module (120) can obtain an embedding that includes the features of the sounds activated in each independent sound source section. In other words, the embedding transformation module (120) can generate a vector (i.e., embedding) that includes the features of the sound source that generated the sound included in the independent sound source section. For example, the embedding transformation module (120) can generate a feature vector (i.e., embedding) by converting the audio signal of the sound included in the independent sound source section into the time-frequency domain and then performing embedding transformation using a neural network-based model.

[0066] Referring to FIG. 3, the embedding transformation module (120) can generate a first embedding (311) by performing embedding transformation on the sound (310) of the third sound source (SS3) activated in the second frame (F2). Similarly, the embedding transformation module (120) can extract a second embedding (321), a third embedding (331), and a fourth embedding (341) from the sound (320) included in the third frame (F3), the sound (330) included in the fourth frame (F4), and the sound (340) included in the fifth frame (F5), respectively.

[0067] Embeddings (311, 321, 331, 341) extracted from individual sound source sections (F2, F3, F4, F5) in this way include the acoustic characteristics of the sound source that generated the sound included in each section, and thus can be used to distinguish the sound source.

[0068] (2) Embedding clustering and determining the main sound source

[0069] The clustering module (130) performs clustering on the embeddings (311, 321, 331, 341) extracted from the individual sound source sections (F2, F3, F4, F5), thereby classifying the embeddings (311, 321, 331, 341) into multiple clusters and determining the corresponding embedding for each sound source.

[0070] For example, the clustering module (130) can compare all extracted embeddings (311, 321, 331, 341) with each other and classify embeddings with a similarity higher than a certain standard into the same cluster. In Fig. 3, as a result of the clustering module (130) performing clustering, the first embedding (311) is classified into the first cluster, the second embedding (321) is classified into the second cluster, and the third embedding (331) and the fourth embedding (341) are classified into the third cluster.

[0071] The clustering module (130) can assign sound sources to each classified cluster. Accordingly, an embedding corresponding to each sound source can be determined. In FIG. 3, a third sound source (SS3) is assigned to a first cluster, a second sound source (SS2) is assigned to a second cluster, and a first sound source (SS1) is assigned to a third cluster. Accordingly, the first embedding (311) corresponds to the third sound source (SS3), the second embedding (321) corresponds to the second sound source (SS2), and the third embedding (331) and the fourth embedding (341) correspond to the first sound source (SS1). The embedding corresponding to each sound source can be used in a later sound source separation process.

[0072] The primary sound source selection module (140) may select one or more sound sources from among a plurality of sound sources that generate sounds included in the audio signal (10) as the primary sound source. According to one embodiment of the present disclosure, the primary sound source selection module (140) may determine the primary sound source based on the length of the sound generated by each sound source within a section in which fast scanning is performed. In one embodiment, the primary sound source selection module (140) may identify an active segment for each sound source based on the clustering result, and determine at least one sound source among the plurality of sound sources as the primary sound source based on the length of the active segment. For example, the primary sound source selection module (140) may determine a preset number of sound sources as the primary sound sources in the order of the length of the active segment.

[0073] Comparing the lengths of the activation sections of each sound source in FIG. 3, it can be seen that the activation section of the first sound source (SS1) is the longest, and the activation section of the second sound source (SS2) is the shortest. The main sound source selection module (140) can select the first sound source (SS1) as the main sound source if one sound source must be selected as the main sound source, and can select the first sound source (SS1) and the third sound source (SS3) as the main sound sources if two sound sources must be selected as the main sound sources. In some embodiments of the present disclosure, the first sound source (SS1), the second sound source (SS2), and the third sound source (SS3) are all selected as the main sound sources. As illustrated in FIG. 2, the fast scanning module (100) can output an embedding (E1) corresponding to a first sound source (SS1), an embedding (E2) corresponding to a second sound source (SS2), and an embedding (E3) corresponding to a third sound source (SS3).

[0074] Referring to FIGS. 2 and 3, either the third embedding (331) extracted from the fourth frame (F4) or the fourth embedding (341) extracted from the fifth frame (F5) can be used as the embedding (E1) corresponding to the first sound source (SS1). Similarly, the second embedding (321) extracted from the third frame (F3) can be used as the embedding (E2) corresponding to the second sound source (SS2). In addition, the first embedding (311) extracted from the second frame (F2) can be used as the embedding (E3) corresponding to the third sound source (SS3).

[0075] In some embodiments, the primary sound source is selected based on the length of the activation interval. In some embodiments, the primary sound source selection module (140) may also select the primary sound source based on various other criteria.

[0076] 3. Detailed configuration and operation of the sound source separation module

[0077] Returning to FIG. 1 again, the sound source separation module (200) can separate main sound sources from the audio signal (10) based on the embeddings (E1, E2, E3) corresponding to the main sound sources. According to one embodiment of the present disclosure, the sound source separation module (200) can select some or all of the main sound sources as target sound sources and separate the target sound sources from the audio signal (10).

[0078] The following explains how some or all of the main sound sources detected as a result of fast scanning are selected as target sound sources.

[0079] According to one embodiment of the present disclosure, the electronic device can select target sound sources in descending order of the length of their activation intervals. For example, if the sound source is a speaker, the electronic device can select a certain number of speakers as target sound sources (target speakers) in descending order of the number of utterances.

[0080] According to one embodiment of the present disclosure, a user may select at least one of the primary sound sources as a target sound source. For example, if a list of primary sound sources detected through fast scanning is displayed on the screen of an electronic device, the user may select a target sound source from the displayed list.

[0081] At this time, the electronic device may display a screen that allows the user to identify the main sound source. For example, the electronic device may display a screen that allows the user to identify a single sound source section (a section in which the main sound source is activated alone) corresponding to each main sound source on a time axis. In some embodiments, the electronic device may allow the user to listen to the sound generated by each sound source for the main sound sources. In some embodiments, the electronic device may display a screen that allows the user to view a video corresponding to a single sound source section corresponding to the main sound source.

[0082] The sound source separation module (200) can separate the selected target sound source from the audio signal (10) according to the method described above.

[0083] (1) Process of separating target sound source

[0084] FIG. 4 illustrates detailed components included in a sound source separation module according to one embodiment of the present disclosure. FIG. 5 illustrates a process of a sound source separation module according to one embodiment of the present disclosure separating a sound source from an audio signal based on a fast scanning result.

[0085] Referring to FIG. 4, the sound source separation module (200) may include a sound separation module (210), an embedding conversion module (220), and an embedding matching module (230). The operation of each module will be described with reference to FIG. 5.

[0086] The sound separation module (210) can separate sounds included in an audio signal (10). As described above, the audio signal (10) can be processed frame by frame, and thus the sound separation module (210) can separate sounds frame by frame.

[0087] When a target frame segmented from an audio signal (10) is input to a sound separation module (210), the sound separation module (210) can separate sounds (510, 520, 530) included in the target frame. Referring to FIG. 5, the sound separation module (210) separates a first sound (510), a second sound (520), and a third sound (530) from the target frame.

[0088] The embedding transformation module (220) can generate a vector (embedding) including the characteristics of the separated sound by performing an embedding transformation on the sound separated from the audio signal (10). In FIG. 5, the embedding transformation module (220) can convert the first sound (510), the second sound (520), and the third sound (530) into a first embedding (511), a second embedding (521), and a third embedding (531), respectively. The embedding corresponding to each sound can include the characteristics of the sound source that generated the sound.

[0089] When the embedding conversion is completed, the embedding matching module (230) compares the embeddings (511, 521, 531) converted from the separated sounds (510, 520, 530) with the embeddings of the main sound source (or target sound source) to determine whether there is a match, and can separate the sound source based on the matching result.

[0090] As previously explained, some of the main sound sources can be selected as target sound sources, and in FIG. 5, the first sound source (SS1) and the third sound source (SS3) are selected as target sound sources. Accordingly, the embedding matching module (230) can determine whether there is a match by comparing each of the first embedding (511) to the third embedding (531) with the embedding (E1) of the first sound source (SS1) and the embedding (E3) of the third sound source (SS3). The embedding matching module (230) can determine that two embeddings are matched if the similarity between them is above a certain standard.

[0091] Referring to FIG. 5, the embedding matching module (230) can determine that the first embedding (511) matches the embedding (E1) of the first sound source (SS1), and that the third embedding (531) matches the embedding (E3) of the third sound source (SS3). Accordingly, the first sound (510) corresponds to the first sound source (SS1), and the third sound (530) corresponds to the third sound source (SS3).

[0092] The embedding matching module (230) can perform diarization on target sound sources based on the matching results. For example, the embedding matching module (230) can represent the active section of the sound generated by the target sound source on the time axis by matching the sound corresponding to the target sound source to all frames constituting the audio signal (10).

[0093] (2) Method for improving separation accuracy (connection or summation of sounds generated by target sound sources)

[0094] As described above, the sound source separation module (200) can separate the target sound source from the audio signal based on the embedding of the target sound source. According to embodiments of the present disclosure, in order to increase the separation accuracy, the sound generated by the target sound source can be added to the input of the sound source separation module (200). Specific embodiments will be described with reference to FIGS. 6 and 7.

[0095] FIG. 6 illustrates a method for connecting the sound of a target sound source to an audio signal to improve sound source separation accuracy, according to one embodiment of the present disclosure. FIG. 7 illustrates a method for adding the sound of a target sound source to an audio signal to improve sound source separation accuracy, according to one embodiment of the present disclosure.

[0096] As described above, the sound source separation module (200) according to one embodiment of the present disclosure can separate the sound of the target sound source from the audio signal when it receives an audio signal and an embedding (target embedding) of the target sound source. In order to separate the sound source, the sound source separation module (200) determines whether the embedding of the sound included in the audio signal matches the embedding of the target sound source. However, if the sound source is separated only based on the result of comparing the embeddings in this way, the separation accuracy may decrease. Therefore, the electronic device according to one embodiment of the present disclosure can additionally input an audio signal of a sound generated by the target sound source to the sound source separation module (200) so that the sound generated by the target sound source can also be referred to when separating the sound source.

[0097] In Fig. 6, the first sound source (SS1) is selected as the target sound source, and therefore, the embedding (E1) of the first sound source (SS1) is input to the sound source separation module (200).

[0098] A sound (62) generated by a first sound source (SS1), which is a target sound source, is concatenated in front of a target frame (61) segmented from an audio signal, and the concatenated signal can be input to a sound source separation module (200). The sound (62) generated by the target sound source can also be concatenated behind the target frame (61).

[0099] The sound source separation module (200) can accurately separate only the sound of the target sound source from the target frame (61) by also referring to the audio signal (62) of the sound generated by the first sound source (SS1), which is the target sound source.

[0100] The sound generated by the target sound source used in this embodiment can be acquired during the process of the electronic device performing fast scanning. That is, the audio signal of the sound included in the single sound source section detected during the process of the electronic device performing fast scanning can be connected before or after the target frame (61). In addition, the embedding (E1) of the first sound source (SS1) acquired during the process of the electronic device performing fast scanning can be input as the target embedding to the sound source separation module (200).

[0101] In Fig. 7, the first sound source (SS1) is selected as the target sound source, and therefore, the embedding (E1) of the first sound source (SS1) is input to the sound source separation module (200).

[0102] The sound (72) generated by the first sound source (SS1), which is the target sound source, is summed in the target frame (71) segmented from the audio signal, and the summed signal can be input to the sound source separation module (200). Specifically, the electronic device can sum the audio signal of the sound (72) generated by the first sound source (SS1) to the audio signal of the target frame (71) in the same time interval, and then input the summed audio signal to the sound source separation module (200).

[0103] The sound source separation module (200) can accurately separate only the sound of the target sound source from the target frame (71) by also referring to the audio signal (72) of the sound generated by the first sound source (SS1), which is the target sound source.

[0104] The sound generated by the target sound source used in this embodiment can be acquired during the process of the electronic device performing fast scanning. That is, the audio signal of the sound included in the single sound source section detected during the process of the electronic device performing fast scanning can be added to the target frame (71). In addition, the embedding (E1) of the first sound source (SS1) acquired during the process of the electronic device performing fast scanning can be input as a target embedding to the sound source separation module (200).

[0105] 4. How to match separated audio sources (speakers) to characters in a video

[0106] An electronic device according to one embodiment of the present disclosure can separate speech from an audio signal included in a video and match the separated speech to a person (i.e., a speaker) included in the video.

[0107] According to one embodiment of the present disclosure, an electronic device can match a voice separated from an audio signal to a speaker appearing in a video based on a single sound source section detected as a result of performing fast scanning. Specifically, the electronic device according to one embodiment of the present disclosure can match a person whose lip shape has been detected in a single sound source section to the speaker who generated the voice included in the single sound source section. Hereinafter, with reference to FIGS. 8 and 9, a method for an electronic device according to one embodiment of the present disclosure to match a separated voice to a speaker in a video will be described in detail.

[0108] (1) Audio-visual matching

[0109] FIG. 8 illustrates modules included in an electronic device for performing a sound source separation process according to one embodiment of the present disclosure. FIG. 9 illustrates a process of matching a character in a video to a separated sound source by an electronic device according to one embodiment of the present disclosure.

[0110] Referring to FIG. 8, an electronic device according to one embodiment of the present disclosure may include a fast scanning module (100), a sound source separation module (200), and an audio-visual matching module (800). The detailed configuration and operation of the fast scanning module (100) and the sound source separation module (200) are as described above with reference to FIGS. 1 to 7.

[0111] The audio-visual matching module (800) can match a voice included in a video to a speaker appearing in the video based on audio signals and video signals included in the video. Specifically, if there is a speaker whose lip movement is detected in a video section corresponding to a single sound source section, the audio-visual matching module (800) can determine the speaker as the speaker who generated the voice included in the single sound source section. Accordingly, the output (80) of the audio-visual matching module (800) can display the results of segmenting multiple voices, along with pictures (81, 82, 83) of speakers matching each voice.

[0112] Below, the specific operation of the audio-visual matching module (800) is described with reference to FIG. 9.

[0113] As illustrated in Fig. 9, a video input to an electronic device may include an audio signal (10) and a video signal (90). In Fig. 9, two people (speaker A, speaker B) appear in the video.

[0114] The fast scanning module (100) of the electronic device can detect single sound source sections by performing fast scanning on the audio signal (10). In Fig. 9, the electronic device performed fast scanning on five frames (F1, F2, F3, F4, F5), and as a result, detected the second frame (F2) to the fifth frame (F5) as single sound source sections.

[0115] The speaker who generated the voice included in the second frame (F2) is called the third speaker (i.e., speaker 3), the speaker who generated the voice included in the third frame (F3) is called the second speaker (i.e., speaker 2), and the speaker who generated the voice included in the fourth frame (F4) and the fifth frame (F5) is called the first speaker (i.e., speaker 1).

[0116] The audio-visual matching module (800) can analyze the lip motion of people included in a video section corresponding to a single sound source section, and select a person (speaker) corresponding to the voice of the single sound source section based on the analysis result.

[0117] Referring to FIG. 9, the audio-visual matching module (800) can analyze frames (91) included in the same section (2 - 4 sec) as the second frame (F2), among the frames included in the video signal (90), to find a speaker matching the voice included in the second frame (F2).

[0118] According to one embodiment of the present disclosure, the audio-visual matching module (800) can analyze lip movements of persons (i.e., speaker A, speaker B) included in frames (91) of a video signal corresponding to a second frame (F2). For example, the audio-visual matching module (800) can find the area around the mouth (Ra, Rb) of the persons in each frame, and compare the area around the mouth (Ra, Rb) of the frames (91) to determine whether there was a change in the lip shape of the persons (speaker A, speaker B) during the section corresponding to the second frame (F2).

[0119] The audio-visual matching module (800) can determine a person with a change in lip shape as the speaker who generated the voice included in the second frame (F2). In FIG. 9, the audio-visual matching module (800) calculates the probability of matching each of the people included in the video (i.e., speaker A, speaker B) with the third speaker (i.e., speaker 3) corresponding to the second frame (F2), and if the results are 0.1 for speaker A and 0.9 for speaker B, it can be determined that speaker B corresponds to the third speaker (i.e., speaker 3) who generated the voice of the second frame (F2).

[0120] (2) Example of displaying audio source separation results on the video playback screen

[0121] FIG. 10 illustrates UI screens indicating sound source separation results displayed on a screen of an electronic device according to one embodiment of the present disclosure.

[0122] Referring to Fig. 10, the first screen (1010) displays an area for displaying a video and buttons for playing or stopping the video.

[0123] The electronic device can separate sound sources from audio signals included in a video when receiving a request for sound source separation from a user, or automatically. The sound source separation result is displayed in the first area (1021) of the second screen (1020). The first area (1021) displays a photo for identifying the speaker corresponding to the separated sound source (voice), and buttons (back button, previous button) for skipping each speaker's voice or returning to the previous voice.

[0124] An electronic device according to one embodiment of the present disclosure performs segmentation for each separated speaker and displays the result on a screen, thereby enabling a user to easily identify the segment in which each speaker spoke.

[0125] The third screen (1030) displays a first timeline (1032) indicating the section where the first speaker (1031) spoke. When a user selects a photo of the first speaker (1031), the first timeline (1032) may be displayed on the screen as shown.

[0126] The user can grasp the entire speech segment of the first speaker (1031) through the first timeline (1032). In addition, the user can easily move between the speech segments of the first speaker (1031) by selecting the back or previous button.

[0127] Similarly, a second timeline (1042) indicating the section where the second speaker (1041) spoke is displayed on the fourth screen (1040). When a user selects a photo of the second speaker (1041), the second timeline (1042) may be displayed on the screen as described above.

[0128] The user can obtain a comprehensive overview of the second speaker's (1041) speech segments through the second timeline (1042). Furthermore, the user can easily navigate between the second speaker's (1041) speech segments by selecting the back or previous button.

[0129] 5. Overall configuration and operation of the electronic device

[0130] Below, an electronic device for performing the operations described above will be described. An electronic device according to one embodiment of the present disclosure may be a device with a photographing function and a computational processing function, such as a smartphone or digital camera. In other embodiments, the electronic device may be any type of device (e.g., a laptop or cloud server) capable of receiving video or audio files and performing a sound source separation process, even without a photographing function. The configuration of an electronic device according to one embodiment of the present disclosure will be described in detail below with reference to FIG. 11.

[0131] FIG. 11 illustrates components included in an electronic device according to one embodiment of the present disclosure. Referring to FIG. 11, an electronic device (1100) according to one embodiment of the present disclosure may include an input / output interface (1110), a processor (1120), and a memory (1130).

[0132] The input / output interface (1110) may include an input interface (e.g., a touch screen, a keyboard, a microphone, etc.) for receiving commands or information from a user, and an output interface (e.g., a display panel, a speaker, etc.) for displaying the result of an operation according to a user's command or the status of the electronic device (1100). According to one embodiment of the present disclosure, the electronic device (1100) may receive an input (e.g., a sound source separation request) from a user through the input / output interface (1110), and when an operation is completed, may output the result of performing the operation (e.g., a sound source separation result) through the input / output interface (1110).

[0133] The processor (1120) controls a series of processes to operate the electronic device (1100) according to the embodiments described in the present disclosure, and may be composed of one or more processors. The one or more processors included in the processor (1120) may be circuitry such as a System on Chip (SoC), an Integrated Circuit (IC), etc. The one or more processors included in the processor (1120) may be a general-purpose processor such as a Central Processing Unit (CPU), a Micro Processor Unit (MPU), an Application Processor (AP), a Digital Signal Processor (DSP), a graphics-only processor such as a Graphics Processing Unit (GPU), a Vision Processing Unit (VPU), an artificial intelligence-only processor such as a Neural Processing Unit (NPU), or a communication-only processor such as a Communication Processor (CP). When the one or more processors included in the processor (1120) are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0134] The processor (1120) can write data to the memory (1130) or read data stored in the memory (1130), and in particular, process data according to predefined operation rules or artificial intelligence models by executing a program or at least one instruction stored in the memory (1130). Accordingly, the processor (1120) can perform operations described in the embodiments of the present disclosure, and operations described as being performed by the electronic device (1100) or modules included in the electronic device (1100) in the present disclosure can be regarded as being performed by the processor (1120) unless otherwise specifically described.

[0135] The memory (1130) is a configuration for storing various programs or data, and may be configured as a storage medium such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD, or a combination of storage media. The memory (1130) may not exist separately and may be configured to be included in the processor (1120). The memory (1130) may be configured as a volatile memory, a non-volatile memory, or a combination of volatile memory and non-volatile memory. A program or at least one instruction for performing operations according to embodiments described below may be stored in the memory (1130). The memory (1130) may also provide stored data to the processor (1120) at the request of the processor (1120).

[0136] The embodiments described above with reference to FIGS. 1 to 10 can be performed by an electronic device (1100).

[0137] 6. Describe the process by referring to the flowcharts.

[0138] Hereinafter, with reference to the flowcharts of FIGS. 12 to 18, a method for an electronic device according to embodiments of the present disclosure to separate a sound source from an audio signal will be described. The steps included in the flowcharts of FIGS. 12 to 18 can be performed by the electronic device (1100) of FIG. 11, and therefore, the contents previously described with reference to FIGS. 1 to 11 may be equally applied to FIGS. 12 to 18, even if omitted below.

[0139] Referring to FIG. 12, in steps 801 to 1201, the electronic device can obtain an audio signal containing sounds generated by multiple sound sources.

[0140] In step 1202, the electronic device can obtain an embedding corresponding to at least one primary sound source among the multiple sound sources based on a single sound source section that includes only sounds generated by one of the multiple sound sources. The detailed steps included in step 1202 are illustrated in FIG. 13.

[0141] Referring to FIG. 13, at step 1301, the electronic device can divide the audio signal into multiple frames.

[0142] At step 1302, the electronic device can determine each frame in which only the sound generated by one sound source among the multiple frames is activated as a single sound source section.

[0143] At step 1303, the electronic device can obtain an embedding that includes the characteristics of the sound activated in each individual sound source section.

[0144] At step 1304, the electronic device can determine the corresponding embedding for each sound source by performing clustering on the embeddings. For example, the electronic device can classify similar embeddings into the same cluster and assign sound sources to each cluster.

[0145] At step 1305, the electronic device can identify the activation section for each sound source based on the clustering result.

[0146] At step 1306, the electronic device may determine at least one sound source among the multiple sound sources as the primary sound source based on the length of the activation interval. For example, the electronic device may determine a certain number of sound sources in descending order of the length of the activation interval as the primary sound source.

[0147] Returning to Figure 12, at step 1203, the electronic device can isolate at least one primary sound source from the audio signal based on the acquired embedding. Detailed steps included in step 1203 are illustrated in Figures 14 and 15, respectively.

[0148] Referring to FIG. 14, at step 1401, the electronic device can separate the sound included in the audio signal into frames of a preset length.

[0149] At step 1402, the electronic device can obtain an embedding corresponding to the isolated sound.

[0150] At step 1403, the electronic device can determine whether the embedding corresponding to the isolated sound matches the embedding corresponding to at least one primary sound source. For example, the electronic device can determine that embeddings with a similarity level above a certain threshold are matched.

[0151] At step 1404, the electronic device may perform segmentation on at least one primary sound source based on the result of the match determination. In this case, segmentation may be an operation that indicates the active section of the sound generated by the primary sound source based on the time axis, or a corresponding operation.

[0152] Referring to Figure 15, at step 1501, the electronic device may select one of the embeddings corresponding to at least one primary sound source as the target embedding. For example, the electronic device may select the target embedding based on the length of the activation interval, or the user may select one of the primary sound sources as the target embedding.

[0153] At step 1502, the electronic device can separate sounds matching the target embedding from the audio signal. The detailed steps involved in step 1502 are illustrated in FIGS. 16 and 17, respectively.

[0154] Referring to FIG. 16, at step 1601, the electronic device can obtain a segmented target frame from an audio signal.

[0155] At step 1602, the electronic device can concatenate the audio signal of the sound generated by the sound source corresponding to the target embedding with the audio signal of the target frame before or after the audio signal. The concatenated audio signal may be the audio signal of the sound used to generate the target embedding.

[0156] At step 1603, the electronic device can input the connected audio signal and target embedding into the audio source separation module.

[0157] At step 1604, the electronic device can obtain separated sound from the sound source separation module.

[0158] Referring to FIG. 17, at step 1701, the electronic device can obtain a segmented target frame from an audio signal.

[0159] At step 1702, the electronic device can sum the audio signal of the sound generated by the sound source corresponding to the target embedding to the audio signal of the target frame in the same time interval. The connected audio signal may be the audio signal of the sound used to generate the target embedding.

[0160] At step 1703, the electronic device can input the summed audio signal and target embedding into a sound source separation module.

[0161] At step 1704, the electronic device can obtain separated sound from the sound source separation module.

[0162] According to one embodiment of the present disclosure, an electronic device may separate a primary sound source from an audio signal and then match the separated sound source to a character appearing in a video. Figure 18 illustrates steps for matching the separated sound source to a character appearing in a video. The steps of Figure 18 may be performed subsequent to step 1203 of Figure 12.

[0163] Referring to FIG. 18, in step 1801, the electronic device can analyze lip movements of at least one person included in a video segment corresponding to a single sound source segment.

[0164] At step 1802, the electronic device can select a person corresponding to at least one primary sound source based on the analysis results. For example, the electronic device can determine that the person whose lip shape has been detected is the speaker who generated the voice included in the single sound source section.

[0165] According to the embodiments described above, an electronic device can quickly determine the primary sound source by performing fast scanning on an audio signal, extract the characteristics of the primary sound source, and isolate the primary sound source from the audio signal based on the extracted characteristics. Therefore, user convenience is expected to be enhanced, as there is no need to separately input information about the target sound source for separation.

[0166] A method for separating a sound source from an audio signal according to one embodiment of the present disclosure may include the steps of: obtaining an audio signal including sounds generated by a plurality of sound sources; obtaining an embedding corresponding to at least one primary sound source among the plurality of sound sources based on a single sound source segment including only sounds generated by one of the plurality of sound sources; and separating the at least one primary sound source from the audio signal based on the obtained embedding.

[0167] According to one embodiment, the step of obtaining an embedding corresponding to the at least one main sound source may include the steps of dividing the audio signal into a plurality of frames, determining frames in which only a sound generated by one sound source among the plurality of frames is activated as a single sound source section, obtaining an embedding including a characteristic of a sound activated in each of the single sound source sections, and determining an embedding corresponding to each sound source by performing clustering on the embeddings.

[0168] According to one embodiment, the step of obtaining an embedding corresponding to the at least one main sound source may further include the step of identifying an active segment for each sound source based on the clustering result, and the step of determining at least one sound source among the plurality of sound sources as the main sound source based on the length of the active segment.

[0169] According to one embodiment, the step of separating the at least one main sound source may include the steps of separating a sound included in the audio signal for each frame of a preset length, obtaining an embedding corresponding to the separated sound, determining whether the embedding corresponding to the separated sound matches the embedding corresponding to the at least one main sound source, and performing diarization on the at least one main sound source based on a result of determining whether the embedding matches.

[0170] According to one embodiment, the segmentation may be an operation of indicating an active section of sound in which at least one main sound source is generated based on a time axis.

[0171] According to one embodiment, the step of separating the at least one main sound source may include the step of selecting one of the embeddings corresponding to the at least one main sound source as a target embedding and the step of separating a sound matching the target embedding from the audio signal.

[0172] According to one embodiment, the step of separating a sound matching the target embedding may include the steps of obtaining a target frame segmented from the audio signal, concatenating an audio signal of a sound in which a sound source corresponding to the target embedding is generated before or after an audio signal of the target frame, inputting the concatenated audio signal and the target embedding into a sound source separation module, and obtaining a separated sound from the sound source separation module.

[0173] According to one embodiment, the step of separating a sound matching the target embedding may include the steps of obtaining a target frame segmented from the audio signal, summing an audio signal of a sound generated by a sound source corresponding to the target embedding to an audio signal of the target frame in the same time interval, inputting the summed audio signal and the target embedding to a sound source separation module, and obtaining a separated sound from the sound source separation module.

[0174] According to one embodiment, the at least one main sound source is a speaker, and the method may further include a step of analyzing lip motion of at least one person included in a video section corresponding to the single sound source section, and a step of selecting a person corresponding to the at least one main sound source based on the analysis result.

[0175] According to one embodiment, the step of selecting the person may determine the person whose lip shape change is detected as the speaker who generated the speech included in the single sound source section.

[0176] An electronic device according to one embodiment of the present disclosure includes a memory in which a program or at least one instruction is stored, and at least one processor operably coupled to the memory, wherein the at least one processor executes the program stored in the memory or the at least one instruction, thereby enabling the electronic device to obtain an audio signal including sounds generated by a plurality of sound sources, obtain an embedding corresponding to at least one primary sound source among the plurality of sound sources based on a single sound source segment including only sounds generated by one of the plurality of sound sources, and then separate the at least one primary sound source from the audio signal based on the obtained embedding.

[0177] According to one embodiment, the electronic device may determine an embedding corresponding to each sound source by dividing the audio signal into a plurality of frames, determining frames in which only a sound generated by one sound source among the plurality of frames is activated as a single sound source section, obtaining an embedding including a characteristic of a sound activated in each of the single sound source sections, and then performing clustering on the embeddings.

[0178] According to one embodiment, the electronic device may, upon obtaining an embedding corresponding to the at least one main sound source, determine an active segment for each sound source based on the clustering result, and then determine at least one sound source among the plurality of sound sources as the main sound source based on the length of the active segment.

[0179] According to one embodiment, the electronic device may separate the sound included in the audio signal for each frame of a preset length in separating the at least one main sound source, obtain an embedding corresponding to the separated sound, determine whether the embedding corresponding to the separated sound matches the embedding corresponding to the at least one main sound source, and then perform diarization on the at least one main sound source based on a result of determining whether the embedding matches.

[0180] According to one embodiment, the segmentation may be an operation of indicating an active section of sound in which at least one main sound source is generated based on a time axis.

[0181] According to one embodiment, the electronic device may select one of the embeddings corresponding to the at least one main sound source as a target embedding, and then separate a sound matching the target embedding from the audio signal, in separating the at least one main sound source.

[0182] According to one embodiment, the electronic device may obtain a target frame segmented from the audio signal in separating a sound matching the target embedding, concatenate an audio signal of a sound generated by a sound source corresponding to the target embedding before or after the audio signal of the target frame, input the concatenated audio signal and the target embedding into a sound source separation module, and then obtain a separated sound from the sound source separation module.

[0183] According to one embodiment, the electronic device may obtain a target frame segmented from the audio signal in order to separate a sound matching the target embedding, sum an audio signal of a sound generated by a sound source corresponding to the target embedding to the audio signal of the target frame in the same time interval, input the summed audio signal and the target embedding to a sound source separation module, and then obtain a separated sound from the sound source separation module.

[0184] According to one embodiment, the at least one main sound source is a speaker, and the electronic device analyzes lip motion of at least one person included in a video section corresponding to the single sound source section, and then selects the person corresponding to the at least one main sound source based on the analysis result.

[0185] Various embodiments of the present disclosure may be implemented or supported by one or more computer programs, and the computer programs may be formed from computer-readable program code and embodied in a computer-readable medium. In the present disclosure, "application" and "program" may refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in computer-readable program code. "Computer-readable program code" may include various types of computer code, including source code, object code, and executable code. "Computer-readable medium" may include various types of media that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), a hard disk drive (HDD), a compact disc (CD), a digital video disc (DVD), or various types of memory.

[0186] Additionally, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, a 'non-transitory storage medium' is a tangible device and may exclude wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. Meanwhile, this 'non-transitory storage medium' does not distinguish between cases where data is permanently stored in the storage medium and cases where it is temporarily stored. For example, a 'non-transitory storage medium' may include a buffer where data is temporarily stored. A computer-readable medium may be any available medium that can be accessed by a computer, and may include both volatile and non-volatile media, and removable and non-removable media. A computer-readable medium includes a medium on which data can be permanently stored and a medium on which data can be stored and later overwritten, such as a rewritable optical disk or an erasable memory device.

[0187] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0188] The above description of the present disclosure is for illustrative purposes only, and those skilled in the art will appreciate that the present disclosure can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present disclosure. For example, suitable results can be achieved even if the described techniques are performed in a different order than the described method, and / or components of the systems, structures, devices, circuits, etc. described are combined or combined in a different form than the described method, or are replaced or substituted by other components or equivalents. Therefore, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. For example, each component described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined form.

[0189] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.

Claims

1. A method for separating a sound source from an audio signal, A step of obtaining an audio signal including sounds generated by multiple sound sources; A step of obtaining an embedding corresponding to at least one primary sound source among the plurality of sound sources based on a single sound source segment that includes only a sound generated by one of the plurality of sound sources; and A method comprising a step of separating at least one main sound source from the audio signal based on the obtained embedding.

2. In paragraph 1, The step of obtaining an embedding corresponding to at least one main sound source is: A step of dividing the above audio signal into a plurality of frames; A step of determining each frame in which only the sound generated by one of the plurality of frames is activated as a single sound source section; A step of obtaining an embedding including the characteristics of the sound activated in each of the above-mentioned single sound source sections; and A method characterized by including a step of determining an embedding corresponding to each sound source by performing clustering on the above embeddings.

3. In either paragraph 1 or paragraph 2, The step of obtaining an embedding corresponding to at least one main sound source is: A step of confirming an active segment for each sound source based on the clustering result; and A method characterized by further comprising a step of determining at least one sound source among the plurality of sound sources as a main sound source based on the length of the activation section.

4. In any one of paragraphs 1 to 3, The step of separating at least one main sound source is, A step of separating sound included in the audio signal by frame of a preset length; A step of obtaining an embedding corresponding to the separated sound; A step of determining whether the embedding corresponding to the separated sound matches the embedding corresponding to at least one main sound source; and A method characterized by comprising a step of performing diarization on at least one main sound source based on the result of determining whether or not there is a match.

5. In any one of paragraphs 1 to 4, The above segmentation is, A method characterized in that the operation is to indicate an activation section of a sound in which at least one main sound source is generated based on a time axis.

6. In any one of paragraphs 1 to 5, The step of separating at least one main sound source is, A step of selecting one of the embeddings corresponding to at least one of the main sound sources as a target embedding; and A method characterized by comprising a step of separating a sound matching the target embedding from the audio signal.

7. In any one of paragraphs 1 to 6, The step of separating the sound matching the target embedding is as follows: A step of obtaining a segmented target frame from the above audio signal; A step of concatenating an audio signal of a sound generated by a sound source corresponding to the target embedding to the front or back of the audio signal of the target frame; A step of inputting the above-mentioned connected audio signal and the target embedding into a sound source separation module; and A method characterized by comprising a step of obtaining a sound separated from the sound source separation module.

8. In any one of paragraphs 1 to 7, The step of separating the sound matching the target embedding is as follows: A step of obtaining a segmented target frame from the above audio signal; A step of summing the audio signal of the sound generated by the sound source corresponding to the target embedding to the audio signal of the target frame in the same time interval; A step of inputting the above-mentioned summed audio signal and the target embedding into a sound source separation module; and A method characterized by comprising a step of obtaining a sound separated from the sound source separation module.

9. In any one of paragraphs 1 to 8, At least one of the primary sound sources is a speaker, A step of analyzing the lip motion of at least one person included in a video section corresponding to the above-mentioned single sound source section; and A method characterized by further comprising a step of selecting a person corresponding to at least one main sound source based on the analysis results.

10. In electronic devices (1100), a memory (1130) in which a program or at least one instruction is stored; and comprising at least one processor (1120) operably coupled to the memory (1130); The electronic device (1100) executes the program stored in the memory (1130) or the at least one instruction by the at least one processor (1120). Acquire an audio signal containing sounds generated by multiple sound sources, After obtaining an embedding corresponding to at least one primary sound source among the plurality of sound sources based on a single sound source segment that includes only the sound generated by one of the plurality of sound sources, An electronic device that separates at least one main sound source from the audio signal based on the obtained embedding.

11. In paragraph 10, In obtaining an embedding corresponding to at least one of the above main sound sources, The electronic device (1100) executes the program stored in the memory (1130) or the at least one instruction by the at least one processor (1120). Divide the above audio signal into multiple frames, Among the above multiple frames, only the frames in which the sound generated by one sound source is activated are determined as a single sound source section, After obtaining an embedding that includes the characteristics of the sound activated in each of the above-mentioned single sound source sections, An electronic device characterized in that it determines an embedding corresponding to each sound source by performing clustering on the above embeddings.

12. In either of paragraphs 10 or 11, In obtaining an embedding corresponding to at least one of the above main sound sources, The electronic device (1100) executes the program stored in the memory (1130) or the at least one instruction by the at least one processor (1120). Based on the above clustering results, the active segment is confirmed for each sound source, An electronic device characterized in that, based on the length of the above-mentioned activation section, at least one sound source among the plurality of sound sources is determined as the main sound source.

13. In any one of paragraphs 10 to 12, In isolating at least one main sound source, The electronic device (1100) executes the program stored in the memory (1130) or the at least one instruction by the at least one processor (1120). Separate the sound contained in the audio signal by frames of a preset length, Obtain an embedding corresponding to the above separated sound, After determining whether the embedding corresponding to the separated sound matches the embedding corresponding to at least one main sound source, An electronic device characterized in that it performs diarization on at least one main sound source based on the result of determining whether there is a match.

14. In any one of paragraphs 10 to 13, In isolating at least one main sound source, The electronic device (1100) executes the program stored in the memory (1130) or the at least one instruction by the at least one processor (1120). After selecting one of the embeddings corresponding to at least one of the main sound sources as the target embedding, An electronic device characterized in that it separates a sound matching the target embedding from the audio signal.

15. In any one of paragraphs 10 to 14, At least one of the primary sound sources is a speaker, The electronic device (1100) executes the program stored in the memory (1130) or the at least one instruction by the at least one processor (1120). After analyzing the lip motion of at least one person included in the video section corresponding to the above-mentioned single sound source section, An electronic device characterized in that, based on the analysis results, a person corresponding to at least one main sound source is selected.

Citation Information

Patent Citations

  • Speaker separation system and method using voice feature vectors

    KR1020160013592A

  • Apparatus for electronic stability control in a vehicle and control method thereof

    KR1020230045347A

  • Thin film forming method and transistor and capacitor manufactured using the same

    KR102731905B1

  • Voice separation device, voice separation method, voice separation program, and voice separation system

    WO2020039571A1

  • KR20220103507A