Device for processing video and operation method thereof

By using an artificial intelligence model to extract audio-related information from image signals and mixed audio signals and apply it to the sound source separation model, the problem of difficulty in separating multiple sound sources in the prior art is solved, and efficient sound source separation performance is achieved.

CN120051826APending Publication Date: 2025-05-27SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073453.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-20
Filing Date
2023-10-17
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively separate multiple sound sources from mixed audio signals, especially in the overlapping sound sources and noise environments, resulting in a degradation of separation performance.

Method used

By generating audio-related information from the image signal and the mixed audio signal using a first artificial intelligence (AI) model, indicating the degree of overlap of multiple sound sources, and applying this information to the second AI model, at least one sound source is separated from the mixed audio signal.

Benefits of technology

It realizes efficient separation of multiple sound sources in mixed audio signals in complex environments, and improves separation performance, especially in sound source overlap and noise environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051826A_ABST
    Figure CN120051826A_ABST
Patent Text Reader

Abstract

An electronic device for processing a video including an image signal and a mixed audio signal includes: a memory configured to store at least one program for processing the video; and at least one processor configured to: generate audio related information from the image signal and the mixed audio signal by using a first artificial intelligence (AI) model, the audio related information indicating a degree of overlap of a plurality of sound sources included in the mixed audio signal; and separating at least one of the plurality of sound sources included in the mixed audio signal from the mixed audio signal by applying the audio-related information to the second AI model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an apparatus for processing a video and an operating method of the apparatus, and more particularly, to an apparatus for separating at least one sound source from a plurality of sound sources included in a mixed audio signal for a video including an image signal and a mixed audio signal and an operating method of the apparatus. Background Art

[0002] As the environment for watching videos gradually becomes more diverse, the methods that enable users to interact with videos are increasing. For example, when playing a video on a screen that supports touch input (e.g., a smart phone or tablet), the user can use a method of zooming in on a portion of the video (e.g., an area where a specific character appears) through touch input to enable the focus of the video to be fixed to a specific character. Then, intuitive feedback can be provided to the user by reflecting the fixed focus in the audio output. To do this, it is necessary to first separate the sound source (voice) from the video. Summary of the invention

[0003] According to one aspect of the present disclosure, an electronic device for processing a video including an image signal and a mixed audio signal, the electronic device comprising: at least one processor; and a memory configured to store at least one program for processing the video. By executing the at least one program, the at least one processor is configured to: generate audio-related information from the image signal and the mixed audio signal by using a first artificial intelligence (AI) model, the audio-related information indicating the degree of overlap of multiple sound sources included in the mixed audio signal; and separate at least one of the multiple sound sources included in the mixed audio signal from the mixed audio signal by applying the audio-related information to a second AI model.

[0004] In an embodiment, the audio-related information includes a map indicating the degree of overlap of the plurality of sound sources, and each interval of the map has a probability value corresponding to the degree to which one of the plurality of sound sources overlaps with another sound source in the time-frequency domain.

[0005] In one embodiment, the first AI model includes: a first sub-model, configured to generate multiple mouth movement information representing temporal pronouncing information of multiple speakers corresponding to multiple sound sources from an image signal; and a second sub-model, configured to generate audio-related information from a mixed audio signal based on the multiple mouth movement information.

[0006] In one embodiment, the first AI model is trained by comparing training audio-related information estimated from a training image signal and a training audio signal with a true value, and wherein the true value is generated by a product operation between a plurality of probability maps generated from a plurality of spectrograms generated based on each of a plurality of separate training sound sources included in the training audio signal.

[0007] In an embodiment, each of the plurality of probability maps is represented by MaxClip(log(1+||F|| 2 ), 1) generates, where ||F|| 2 is the size of the corresponding spectrogram from among the plurality of spectrograms, and MaxClip(x,1) is a function that outputs x when x is less than 1 and outputs 1 when x is equal to or greater than 1.

[0008] In one embodiment, the second AI model includes an input layer, an encoder including multiple feature layers, and a bottleneck layer, and wherein applying audio-related information to the second AI model includes at least one of: applying audio-related information to the input layer, applying audio-related information to each of the multiple feature layers included in the encoder, or applying audio-related information to the bottleneck layer.

[0009] In one embodiment, the at least one processor is further configured to: generate information related to the number of speakers included in the mixed audio signal from the mixed audio signal or from the mixed audio signal and visual information by using a third AI model; generate the audio related information from the image signal and the mixed audio signal based on the information related to the number of speakers by using the first AI model; and separate at least one sound source of the multiple sound sources included in the mixed audio signal from the mixed audio signal by applying the information related to the number of speakers and the audio related information to the second AI model, wherein the visual information includes at least one key frame included in the image signal, and wherein at least one key frame includes a facial area, which includes the lips of at least one speaker corresponding to the at least one sound source included in the mixed audio signal.

[0010] In an embodiment, the speaker quantity related information included in the mixed audio signal includes at least one of first speaker quantity related information about the mixed audio signal or second speaker quantity related information about the visual information.

[0011] In an embodiment, the first speaker quantity related information includes a probability distribution of the number of speakers corresponding to a plurality of sound sources included in the mixed audio signal, and wherein the second speaker quantity related information includes a probability distribution of the number of speakers included in the visual information.

[0012] In one embodiment, applying information related to the number of speakers to the second AI model includes applying information related to the number of speakers to the input layer, applying information related to the number of speakers to each of a plurality of feature layers included in the encoder, or applying information related to the number of speakers to at least one of the bottleneck layers.

[0013] In one embodiment, the at least one processor is further configured to: obtain multiple mouth movement information associated with multiple speakers from the image signal; and separate at least one sound source from the multiple sound sources included in the mixed audio signal by applying the obtained multiple mouth movement information to the second AI model.

[0014] In an embodiment, the electronic device further includes: an input / output interface configured to display a screen on which a video is played back and to receive an input from a user for selecting at least one speaker from a plurality of speakers corresponding to a plurality of sound sources included in a mixed audio signal; and an audio output interface configured to output at least one sound source corresponding to at least one speaker selected from a plurality of sound sources included in the mixed audio signal.

[0015] In an embodiment, the at least one processor is further configured to: display on a screen a user interface for adjusting the volume of at least one sound source corresponding to the selected at least one speaker, and receive an adjustment of the volume of the at least one sound source from a user; and adjust the volume of the at least one sound source output through the audio output interface based on the adjustment of the volume of the at least one sound source.

[0016] According to one aspect of the present disclosure, a method for processing a video including an image signal and a mixed audio signal includes: generating audio-related information from the image signal and the mixed audio signal by using a first artificial intelligence (AI) model, the audio-related information indicating a degree of overlap of multiple sound sources included in the mixed audio signal; and separating at least one of the multiple sound sources included in the mixed audio signal from the mixed audio signal by applying the audio-related information to a second AI model.

[0017] In an embodiment, the audio-related information includes a map indicating the degree of overlap of the plurality of sound sources, and each interval of the map has a probability value corresponding to the degree to which one of the plurality of sound sources overlaps with another sound source in the time-frequency domain.

[0018] In one embodiment, the generation of audio-related information includes: generating a plurality of pieces of mouth movement information representing temporal pronunciation information of a plurality of speakers corresponding to a plurality of sound sources included in a mixed audio signal from an image signal; and generating audio-related information from the mixed audio signal based on the plurality of pieces of mouth movement information.

[0019] In one embodiment, the second AI model includes an input layer, an encoder including multiple feature layers, and a bottleneck layer, and wherein applying audio-related information to the second AI model includes at least one of: applying audio-related information to the input layer, applying audio-related information to each of the multiple feature layers included in the encoder, or applying audio-related information to the bottleneck layer.

[0020] In one embodiment, the method also includes: generating information related to the number of speakers included in the mixed audio signal from the mixed audio signal or from the mixed audio signal and visual information by using a third AI model; generating the audio related information from the image signal and the mixed audio signal based on the information related to the number of speakers by using the first AI model; and separating at least one sound source of the multiple sound sources included in the mixed audio signal from the mixed audio signal by applying the information related to the number of speakers and the audio related information to the second AI model, wherein the visual information includes at least one key frame included in the image signal, and wherein at least one key frame includes a facial area, which includes the lips of at least one speaker corresponding to the at least one sound source included in the mixed audio signal.

[0021] In one embodiment, applying information related to the number of speakers to the second AI model includes applying information related to the number of speakers to an input layer, applying information related to the number of speakers to each of a plurality of feature layers included in an encoder, or applying information related to the number of speakers to at least one of the bottleneck layers.

[0022] According to one aspect of the present disclosure, a non-transitory computer-readable recording medium storing a computer program for processing a video including an image signal and a mixed audio signal is provided. When the computer program is executed by at least one processor, the at least one processor may perform: generating audio-related information from the image signal and the mixed audio signal by using a first artificial intelligence (AI) model, the audio-related information indicating a degree of overlap of multiple sound sources included in the mixed audio signal; and separating at least one sound source of the multiple sound sources included in the mixed audio signal from the mixed audio signal by applying the audio-related information to a second AI model. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent through the following description in conjunction with the accompanying drawings, in which:

[0024] Figure 1 is a diagram for describing a case where a sound source is separated for each speaker included in a video and playback of the video is controlled according to a result of the separation according to an embodiment of the present disclosure;

[0025] Figure 2 is a schematic block diagram of a structure of an electronic device for processing a video according to an embodiment of the present disclosure;

[0026] Figure 3 is a diagram expressing a program that allows an electronic device to execute a method of processing a video according to an embodiment of the present disclosure as a plurality of artificial intelligence (AI) models;

[0027] Figure 4a The structure and input-output relationship of the sound source characteristic analysis model according to an embodiment of the present disclosure are shown;

[0028] Figure 4b The structure and input-output relationship of the first sub-model included in the sound source characteristic analysis model according to an embodiment of the present disclosure are shown;

[0029] Figure 4c shows the structure and input-output relationship of the second sub-model included in the sound source characteristic analysis model according to an embodiment of the present disclosure;

[0030] Figure 5 is a view for explaining an exemplary structure of a sound source separation model and a method of applying audio-related information to the sound source separation model according to an embodiment of the present disclosure;

[0031] Figure 6a and Figure 6b shows the input-output relationship of the speaker quantity analysis model according to an embodiment of the present disclosure;

[0032] Figure 7 is a diagram for explaining a method of applying information related to the number of speakers to a sound source separation model according to an embodiment of the present disclosure;

[0033] Figure 8 is a view for explaining a method of training a sound source characteristic analysis model according to an embodiment of the present disclosure;

[0034] Fig. 9 is a flowchart of a method for processing a video according to an embodiment of the present disclosure;

[0035] Fig.10a and Fig.10b Experimental results on the performance of an electronic device and method for processing a video according to an embodiment of the present disclosure are shown;

[0036] Fig.11a and Fig.11b The experimental results comparing the performance of the electronic device and method for processing video according to the embodiments of the present disclosure with the related art are shown;

[0037] Fig.12A method for controlling the playback of a video based on the result of processing the video according to an embodiment of the present disclosure is shown;

[0038] Fig.13 A method for controlling the playback of a video based on the result of processing the video according to an embodiment of the present disclosure is shown;

[0039] Fig.14 A method for controlling the playback of a video based on the result of processing the video according to an embodiment of the present disclosure is shown; and

[0040] Fig.15 A method of controlling the playback of a video based on a result of processing the video according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0041] In the following description of the present disclosure, descriptions of technologies that are well known in the art and not directly related to the present disclosure are omitted. This is to clearly convey the gist of the present disclosure by omitting any unnecessary explanations. In addition, the terms used below are defined in consideration of the functions in the present disclosure and may have different meanings depending on the intentions, habits, etc. of the user or operator. Therefore, the terms should be defined based on the descriptions throughout the specification.

[0042] For the same reason, some elements in the accompanying drawings are exaggerated, omitted or schematically shown. In addition, the actual size of each element is not necessarily represented in the accompanying drawings. In the accompanying drawings, the same or corresponding elements are represented by the same reference numerals.

[0043] With reference to the embodiments of the present disclosure described in detail below with reference to the accompanying drawings, the advantages and features of the present disclosure and the methods for achieving the advantages and features will become apparent. However, this is not intended to limit the present disclosure to the disclosed embodiments, and all changes, equivalents and substitutes that do not depart from the spirit and technical scope are included in the present disclosure. The disclosed embodiments are provided so that the present disclosure will be thorough and complete, and the scope of the present disclosure will be fully conveyed to those of ordinary skill in the art. The embodiments of the present disclosure may be defined according to the claims. The same reference numerals in the accompanying drawings represent the same elements. In the description of the embodiments of the present disclosure, when it is considered that some detailed explanations of the relevant functions or configurations may unnecessarily obscure the subject matter of the present disclosure, these detailed explanations are omitted. In addition, the terms used below are defined in consideration of the functions in the present disclosure, and may have different meanings according to the intentions, habits, etc. of the user or operator. Therefore, the terms should be defined based on the descriptions throughout the specification.

[0044] Each block of the flowchart illustration and the combination of blocks in the flowchart illustration can be implemented by computer program instructions. The computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, and the instructions executed by the processor of the computer or other programmable data processing device can generate a device for performing the functions specified in the flowchart block. The computer program instructions can also be stored in a computer-usable or computer-readable memory, which can instruct the computer or other programmable data processing device to function in a specific manner, and the instructions stored in the computer-usable or computer-readable memory can produce an article of manufacture including an instruction device for performing the functions specified in the flowchart block. The computer program instructions can be installed on a computer or other programmable data processing device.

[0045] In addition, each frame of the flow chart can represent a module, a fragment or a part of a code, which includes one or more executable instructions for implementing a specified logical function. According to an embodiment of the present disclosure, the functions mentioned in the block may also occur out of order. For example, two frames shown in succession can actually be executed substantially simultaneously, or can be executed in reverse order according to the function.

[0046] The term "unit" or "device" used herein may represent a software element or a hardware element (e.g., a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC)), and performs a specific function. The term "unit" or "device" may be configured to be included in an addressable storage medium or to reproduce one or more processors. According to an embodiment of the present disclosure, the term "unit" or "device" may include, for example, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcodes, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided by a specific component or a specific "unit" may be combined to reduce the number, or may be divided into additional components. According to an embodiment of the present disclosure, a "unit" or "device" may include one or more processors.

[0047] The term "coupling" and its derivatives refer to any direct or indirect communication between two or more elements, whether or not these elements are in physical contact with each other. The terms "send", "receive" and "communication" and their derivatives cover both direct and indirect communication. The terms "include" and "comprise" and their derivatives refer to including but not limited to this. The term "or" is an inclusive term, meaning "and / or". The phrase "associated with ... " and its derivatives refer to including, being included in ..., interconnected with ..., including, being included in ..., connected to or with ..., connected to or with ..., coupled to or with ..., can communicate with ..., collaborate with ..., interlace, juxtapose, be close to, be bound to or with ..., have, have the property of ..., have to or with ..., etc. The term "controller" refers to any device, system or part thereof that controls at least one operation. Such a controller can be implemented in hardware or a combination of hardware and software and / or firmware. The function associated with any particular controller can be centralized or distributed, whether local or remote. The phrase "at least one of" when used with a list of items means that different combinations of one or more of the listed items may be used, and that only one of the items in the list may be required. For example, "at least one of A, B, and C" includes any of the following combinations: A, B, C, A and B, A and C, B and C, and A and B and C, and any variations thereof. Similarly, the term "set" means one or more. Thus, a set of items may be a single item or a collection of two or more items.

[0048] Embodiments of the present disclosure will now be described more fully with reference to the accompanying drawings.

[0049] Figure 1 is a diagram for describing a case where a sound source is separated for each speaker included in a video and playback of the video is controlled according to a result of the separation according to an embodiment of the present disclosure.

[0050] Figure 1 The first screen 100a corresponds to a scene of the video being played back, and Figure 1 The second screen 100b is an image obtained by enlarging a partial area 10 of the first screen 100a. Two persons 1 and 2 appear on the first screen 100a. Hereinafter, a person appearing in a video is referred to as a "speaker" included in the video.

[0051] When both the first speaker 1 and the second speaker 2 speak in the video, the voices of the two speakers 1 and 2 may be mixed and output. Hereinafter, voice and sound source are used with the same meaning.

[0052] like Figure 1As shown, in order to make the voice of the first speaker 1 emphasized and output or only the voice of the first speaker 1 output when the partial area 10 on the first screen 100a is enlarged and the second screen 100b is displayed, it is necessary to separate the voice of the first speaker 1, the various sound sources, and the audio signal mixed with the voice of the second speaker 2 and (in some cases) the voice of the third speaker that does not appear on the first screen 100a. Therefore, a method of separating the various sound sources from the mixed audio signal included in the video will now be described in detail, and then an embodiment of controlling video playback according to the result of the separation will now be described.

[0053] Figure 2 is a schematic block diagram of a structure of an electronic device 100 for processing a video according to an embodiment of the present disclosure.

[0054] Figure 2 The electronic device 200 shown in the figure may be a display device (e.g., a smart phone or a tablet) that plays a video, or may be a separate server connected to the display device via wired or wireless communication. The method for processing a video according to the embodiment introduced in the present disclosure may be performed by an electronic device including a display that displays a video, may be performed by a separate server connected to the electronic device including the display, or may be performed jointly by an electronic device including the display and the server (the process included in the method is performed separately by two devices).

[0055] In one embodiment, Figure 2 The electronic device 200 is an electronic device including a display, and performs the video processing method described herein. However, as described above, the embodiments of the present disclosure are not limited thereto, and it is apparent that there is a separate server connected to the electronic device including the display, or the server can perform some or all of the processes. Therefore, it should be explained that, among the operations performed by the electronic device 200 in the embodiments described below, the remaining operations except the operation of displaying a video on the display can be performed by a separate device such as a server, even if no further description is given.

[0056] refer to Figure 2 According to an embodiment of the present disclosure, the electronic device 200 may include a communication interface 210, an input / output interface 220, an audio output interface 230, a processor 240, and a memory 250. However, the components of the electronic device 200 are not limited to the above examples, and the electronic device 200 may include more or fewer components than the aforementioned components. According to an embodiment of the present disclosure, some or all of the communication interface 210, the input / output interface 220, the memory 250, the processor 240, and the memory 250 may be implemented as a single chip, and the processor 240 may include one or more processors.

[0057] The communication interface 210 is a component for transmitting and receiving signals (e.g., control commands and data) with an external device by wire or wireless, and may be configured to include a communication chipset supporting various communication protocols. The communication interface 210 may receive a signal from an external source and output the signal to the processor 240, or may transmit a signal output by the processor 240 to an external source.

[0058] The input / output interface 220 may include an input interface (e.g., a touch screen, a hard button, or a microphone) for receiving a control command or information from a user and an output interface (e.g., a display panel) for displaying an execution result of an operation under the control of the user or a state of the electronic device 200. According to an embodiment, the input / output interface 220 may display a video currently being played back, and may receive an input from the user for zooming in on a partial area of ​​the video or selecting a specific speaker or a specific sound source included in the video.

[0059] The audio output interface 230, which is a component for outputting an audio signal included in a video, may be an output device (e.g., an embedded speaker) that is built into the electronic device 200 and is capable of directly reproducing the sound corresponding to the audio signal, may be an interface (e.g., a 3.5 mm terminal, a 4.4 mm terminal, an RCA terminal, or a USB) that allows the electronic device 200 to send and receive audio signals to and from a wired audio playback device (e.g., a speaker, a sound bar, earphones, or a headset), or may be an interface (e.g., a Bluetooth module or a wireless LAN (WLAN) module) that allows the electronic device 200 to send and receive audio signals to and from a wireless audio playback device (e.g., a wireless earphone, a wireless headset, or a wireless speaker).

[0060] The processor 240 is a component that controls a series of processes so that the electronic device 200 operates according to the embodiments described below, and may include one or more processors. One or more processors may be a general-purpose processor such as a central processing unit (CPU), an application processor (AP), or a digital signal processor (DSP), a graphics-only processor such as a graphics processing unit (GPU) or a visual processing unit (VPU), or an artificial intelligence (AI)-only processor such as a neural processing unit (NPU). In an embodiment, when one or more processors are only AI processors, only the AI ​​processor may be designed in a hardware structure dedicated to processing a specific AI model.

[0061] The processor 240 may write data to the memory 250 or read data stored in the memory 250, and specifically, may execute a program stored in the memory 250 to process data according to a predefined operation rule or AI model. Therefore, the processor 240 may perform the operations described in the following embodiments, and unless otherwise specified, the operations described in the following embodiments as being performed by the electronic device 200 may be considered to be performed by the processor 240.

[0062] The memory 250 as a component for storing various programs or data may be composed of a storage medium, such as a read-only memory (ROM), a random access memory (RAM), a hard disk, a compact disk (CD)-ROM, and a digital versatile disk (DVD) or a combination thereof. The memory 250 may not exist alone, but may be included in the processor 240. The memory 250 may be implemented as a volatile memory, a nonvolatile memory, or a combination of a volatile memory and a nonvolatile memory. The memory 250 may store a program for performing the operation according to an embodiment of the present disclosure to be described later. In response to a request from the processor 240, the memory 250 may provide the stored data to the processor 240.

[0063] A method of separating individual sound sources from a mixed audio signal included in a video will now be described in detail, and then an embodiment of controlling video playback according to a result of the separation will be described.

[0064] Figure 3 A diagram expressing a program that allows an electronic device to execute a method of processing a video according to an embodiment of the present disclosure as a plurality of AI models. Figure 3 The models 310, 320, and 330 shown in FIG. 2 can be obtained by classifying operations performed by the processor 240 executing the program 300 stored in the memory 250 according to functions. Therefore, the following description is given as being performed by Figure 3 The operations performed by the illustrated models 310 , 320 , and 330 may be considered to be actually performed by the processor 240 .

[0065] Functions related to the AI ​​model according to an embodiment of the present disclosure may be operated by the processor 240 and the memory 250. The processor 240 may control input data to be processed according to a predefined operation rule or the AI ​​model stored in the memory 250.

[0066] The predefined operating rules or AI models are characterized in that they are created by learning. Here, creation by learning means training a basic AI model using multiple training data through a learning algorithm, so that a predefined operating rule or AI model set to perform the desired characteristics (or desired purpose) is created. This learning can be performed in the device itself that executes the AI ​​according to the present disclosure, or it can be performed by a separate server and / or system. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0067] The AI ​​model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values, and the neural network operation is performed by the operation between the operation result of the previous layer and the multiple weight values. The multiple weight values ​​of the multiple neural network layers can be optimized by the learning results of the AI ​​model. In one embodiment, the multiple weight values ​​can be updated so that the loss value or cost value obtained from the AI ​​model is reduced or minimized during the learning process. The artificial neural network may include a deep neural network (DNN), such as a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN) or a deep Q network, but the embodiments of the present disclosure are not limited thereto.

[0068] Reference Figure 3 According to an embodiment, the program 300 for processing a video may include a sound source characteristic analysis model 310, a sound source separation model 320, and a speaker quantity analysis model 330. The sound source characteristic analysis model 310, the sound source separation model 320, and the speaker quantity analysis model 330 may be referred to as a first AI model, a second AI model, and a third AI model, respectively.

[0069] According to an embodiment, the sound source characteristic analysis model 310 generates audio-related information indicating the degree of overlap of multiple sound sources included in the mixed audio signal from an image signal and a mixed audio signal both included in the video. In this case, the image signal refers to an image (i.e., one or more frames) including a facial area including the lips of a speaker appearing in the video. The facial area including the lips may include the lips and a facial part within a certain distance from the lips. For example, the image signal is an image including a facial area including the lips of a first speaker appearing in the video, and may include 64 frames of 88×88 size. As will be described later, the sound source characteristic analysis model 310 may generate the speaker's mouth movement information representing the speaker's temporal pronunciation information by using the image signal.

[0070] According to an embodiment, the image signal may be a plurality of images including a facial region including lips of a plurality of speakers appearing in the video. For example, when two speakers appear in the video, the image signal may include a first image including a facial region including lips of the first speaker and a second image including a facial region including lips of the second speaker.

[0071] According to an embodiment, the image signal may be an image containing a facial region including lips of as many speakers as the number of speakers determined based on the speaker number distribution information generated by the speaker number analysis model 330. In other words, the number of images included in the image signal may be determined based on the speaker number distribution information generated by the speaker number analysis model 330. In an embodiment, when the probability that three speakers exist is highest according to the speaker number distribution information, the image signal may include three images corresponding to the three speakers.

[0072] According to an embodiment, the audio-related information includes a graph indicating the degree of overlap of multiple sound sources. The graph indicates the degree to which multiple sound sources (corresponding to multiple speakers) overlap each other in the time-frequency domain as a probability. In other words, each bin of the graph in the audio-related information has a probability value corresponding to the degree to which the voices (sound sources) of speakers speaking simultaneously in the corresponding time-frequency domain overlap each other. The term 'probability' means that the degree to which multiple sound sources overlap each other is expressed as a value between 0 and 1. For example, the value in a specific interval can be determined based on the volume of each sound source in the corresponding time-frequency domain.

[0073] Will refer to it later Figure 4a , Figure 4b and Figure 4c The detailed structure and detailed operation of the sound source characteristic analysis model 310 are described.

[0074] According to an embodiment, the sound source separation model 320 separates at least one of the multiple sound sources included in the mixed audio signal from the mixed audio signal by using the audio-related information. According to an embodiment, the sound source separation model 320 may separate a sound source corresponding to the target speaker from the multiple sound sources included in the mixed audio signal by further using an image signal corresponding to the target speaker in addition to the audio-related information.

[0075] The sound source separation model 320 may include Figure 7 The input layer 321 contains multiple feature layers Figure 7 The encoder 322 and Figure 7 As will be described later, audio related information may be applied to Figure 7 The input layer 321, including Figure 7Each of the multiple feature layers in the encoder 322 or Figure 7 In other words, the processor 240 may apply the audio related information to at least one of the bottleneck layers 323. Figure 7 The input layer 321, including Figure 7 Each of the multiple feature layers in the encoder 322 or Figure 7 At least one of the bottleneck layers 323. Figure 5 Describe how audio related information is applied to the sound source separation model 320.

[0076] Any neural network model capable of separating individual sound sources from an audio signal in which a plurality of sound sources are mixed may be employed as the sound source separation model 320. For example, a model such as "VISUAL VOICE" may be employed, which operates by receiving speaker's lip movement information, speaker's facial information, and a mixed audio signal and separating the target speaker's voice (sound source). The detailed operation of "VISUAL VOICE" is disclosed in the following document: GAO, Ruohan; Kristen Grauman. Visualvoice: Audiovisual Speech Separation with Cross-Modal Consistency. In: 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2021. Pages 15490-15500 (GAO, Ruohan; GRAUMAN, Kristen. Visualvoice: Audio-visual speech separation with cross-modal consistency. In: 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2021. p. 15490-15500).

[0077] Conventional sound source separation models (such as "visual speech") operate using a method in which an encoder within the model directly analyzes the area where multiple voices overlap and separates the sound sources, without the need for additional information generated by analyzing the mixed audio signal separately. Therefore, in traditional sound source separation technology, when multiple voices included in the mixed audio signal overlap in the time-frequency domain, the separation performance is sharply reduced according to the characteristics of the multiple voices. For example, when the voices (sound sources) of two male speakers are mixed, the two voices overlap in a significant part of the time-frequency domain, and are therefore more difficult to accurately separate from each other than when the voice (sound source) of a male speaker and the voice (sound source) of a female speaker are mixed. In an actual use environment, the characteristics of the sound sources included in the mixed audio signal and the number of speakers corresponding to them may not be specified, and the separation performance will still decrease due to various noises that occur when the surrounding environment changes over time.

[0078] According to an embodiment of the present disclosure, the electronic device 200 can generate audio-related information indicating the degree of overlap of multiple sound sources included in the mixed audio signal from the image signal and the mixed audio signal included in the video, and by applying the audio-related information to the sound source separation model 320, the sound source can be separated in an actual use environment mixed with various noises, or even for an audio signal mixed with a sound source (speech) having similar characteristics, excellent separation performance can be achieved because the audio-related information is used as an attention map in the sound source separation model 320. The attention map refers to information indicating the relative importance of individual data (i.e., each component of the vector) within the input data (i.e., the vector) to better achieve the target task.

[0079] According to an embodiment, the speaker number analysis model 330 may generate information related to the number of speakers included in the mixed audio signal from the mixed audio signal or from the mixed audio signal and the visual information. The information related to the number of speakers includes the probability distribution of the number of speakers. In an embodiment, the information related to the number of speakers may be an N-dimensional feature vector, which includes the probabilities that the number of speakers included in the mixed sound source is 0, 1, ... and N. According to an embodiment, the information related to the number of speakers may include at least one of the first speaker number related information about the mixed audio signal or the second speaker number related information about the visual information. In an embodiment, the first speaker number related information may include the probability distribution of the number of speakers corresponding to the multiple sound sources included in the mixed audio signal, and the second speaker number related information may include the probability distribution of the number of speakers included in the visual information. According to an embodiment, the speaker number analysis model 330 is an optional component and may be omitted accordingly.

[0080] Conventional sound source separation models, such as the above-mentioned "visual speech", are trained to separate the speech (sound source) of a single target speaker from a mixed audio signal or are trained to separate the speech (sound source) of a specific number of target speakers (people). Therefore, when a conventional sound source separation model is used and the mixed audio signal includes speech (sound sources) corresponding to multiple speakers, each speaker requires a separate model to separate all separate speech (sound sources). In addition, when the number of speakers corresponding to the multiple sound sources included in the mixed audio signal is different from the number of trained target speakers, the separation performance decreases rapidly. On the other hand, the electronic device 200 according to an embodiment of the present disclosure can separate all multiple separate sound sources included in the mixed audio signal by using only a single model by applying information related to the number of speakers to the sound source separation model 320, and provide excellent separation performance regardless of the number of speakers.

[0081] Will refer to it later Figure 6a , Figure 6b and Figure 7 The detailed operation of the speaker quantity analysis model 330 is described.

[0082] Figure 4a , Figure 4b and Figure 4c Detailed description of the structure and operation of the sound source characteristic analysis model according to the embodiment. Figure 4a 2 shows the structure and input-output relationship of the sound source characteristic analysis model 310 according to an embodiment of the present disclosure, Figure 4b shows the structure and input-output relationship of the first sub-model 311 included in the sound source characteristic analysis model 310 according to an embodiment of the present disclosure, and Figure 4c The structure and input-output relationship of the second sub-model 312 included in the sound source characteristic analysis model 310 according to an embodiment of the present disclosure are shown.

[0083] According to an embodiment, the sound source characteristic analysis model 310 receives an image signal and a mixed audio signal and outputs audio related information. The operation of the sound source characteristic analysis model 310 may be performed by a first sub-model 311 and a second sub-model 312 .

[0084] According to an embodiment, the first submodel 311 generates temporal pronunciation information of multiple speakers corresponding to multiple sound sources included in the mixed audio signal from the image signal. For example, the first submodel 311 can extract features of the mouth movements of the speakers included in the image signal from the image signal including 64 frames of 88×88 size. In this case, the features of the mouth movements represent the temporal pronunciation information of the speakers. According to an embodiment, the first submodel 311 may include a first layer consisting of a three-dimensional (3D) convolution (Conv) layer and a pooling (Pool) layer, a second layer (shuffle) consisting of a two-dimensional (2D) convolution layer, a third layer for combining and resizing feature vectors, a fourth layer consisting of a one-dimensional (1D) convolution layer, and a fifth layer consisting of a fully connected layer.

[0085] According to an embodiment, the second sub-model 312 generates audio-related information from the mixed audio signal based on the temporal pronunciation information of the multiple speakers. In an embodiment, the mixed audio signal may be converted into a vector in the time-frequency domain by a short-time Fourier transform (STFT) and may be input to the second sub-model 312. Figure 4a In the example of , the vector in the time-frequency domain into which the mixed audio signal is converted may be 2×256×256 dimensional data having a real part and an imaginary part. According to an embodiment, the second sub-model 312 may include a first layer consisting of a 2D convolution layer and a rectified linear unit (ReLU) activation function, a second layer (encoder) consisting of a 2D convolution layer to downscale the feature vector, a third layer that concatenates the reduced result with the temporal pronunciation information generated in the first sub-model 311 (i.e., the feature vector of the speaker's mouth movement), a fourth layer (decoder) consisting of a 2D convolution layer to upscale the feature vector, and a fifth layer consisting of a 2D convolution layer and a ReLU activation function. As Figure 4c As shown, the temporal pronunciation information generated in the first sub-model 311 is applied between the encoder and the decoder of the second sub-model 312. Figure 4a , Figure 4b and Figure 4c In the example of , the audio-related information generated in the second sub-model 312 may be data having a dimension of 1×256×256.

[0086] Figure 5 is a view for explaining an exemplary structure of a sound source separation model and a method of applying audio-related information to the sound source separation model according to an embodiment of the present disclosure.

[0087] Reference Figure 5According to an embodiment of the present disclosure, the sound source separation model 320 may include an input layer 321, an encoder 322 including a plurality of feature layers, and a bottleneck layer 323. The sound source separation model 320 may further include a decoder for upscaling, and an output layer consisting of a 2D convolution layer and a ReLU activation function. According to an embodiment, the input layer 321 may be composed of a 2D convolution layer and a ReLU activation function, the encoder 322 may be composed of a plurality of convolution layers for downscaling, and the bottleneck layer 323 may be composed of a 2D convolution layer and a ReLU activation function.

[0088] According to an embodiment, the audio-related information may be applied to at least one of the input layer 321, each of the plurality of feature layers included in the encoder 322, or the bottleneck layer 323. In other words, the audio-related information may be applied to the sound source separation model 320, such as Figure 5 , as shown by arrows 510, 520, and 530 shown in . Arrow 510 indicates that the audio-related information is concatenated with the mixed audio signal and applied to the input layer 321, arrow 520 indicates that the audio-related information is applied to each of the multiple feature layers included in the encoder 322, and arrow 530 indicates that the audio-related information is concatenated with the output of the encoder 322 and applied to the bottleneck layer 323, and then sent to the decoder.

[0089] In order to apply audio-related information to the sound source separation model 320, the audio-related information needs to be spliced ​​with the data input to the corresponding layer (i.e., the output of the previous layer). Taking the case of arrow 510 as an example, the mixed audio signal is converted to the time-frequency domain and the mixed audio signal is input to the input layer 321. Therefore, in order to apply audio-related information to the input layer 321, the result of the time-frequency domain conversion of the mixed audio signal needs to be spliced ​​with the audio-related information. In the above example, a vector with a dimension of 2×256×256 as a result of the time-frequency domain conversion of the mixed audio signal is spliced ​​with the audio-related information with a dimension of 1×256×256, and finally a vector with a dimension of 3×256×256 is input to the input layer 321.

[0090] Figure 6a and Figure 6b The input-output relationship of the speaker quantity analysis model 330 according to an embodiment of the present disclosure is shown. Figure 7 is a diagram for explaining a method of applying information related to the number of speakers to a sound source separation model according to an embodiment of the present disclosure.

[0091] refer to Figure 6aAccording to an embodiment of the present disclosure, the speaker quantity analysis model 330 can generate first speaker quantity related information about the mixed audio signal from the mixed audio signal. As described above, the speaker quantity related information includes the probability distribution of the speaker quantity. For example, the speaker quantity related information can be an N-dimensional feature vector, which includes the probabilities that the number of speakers included in the mixed sound source is 0, 1,... and N. When the first speaker quantity related information generated from the mixed audio signal is applied to the sound source separation model 320, for example, even in the case where the number of speakers changes within a unit section of the mixed audio signal or the number of speakers may not be specified, such as in the case where only the first speaker speaks, the first speaker and the second speaker speak simultaneously from a certain point in time, and then only the first speaker speaks again, robust separation performance can be achieved.

[0092] refer to Figure 6b According to an embodiment of the present disclosure, the speaker number analysis model 330 may generate first speaker number related information about the mixed audio signal and second speaker number related information about the visual information from the mixed audio signal and the visual information. The visual information may be at least one key frame among a plurality of frames included in the image signal. The key frame may include a facial region, the facial region including the lips of at least one speaker corresponding to at least one sound source included in the mixed audio signal. Since the first speaker number related information is generated based on the mixed audio signal, and the second speaker number related information is generated based on the visual information, speaker number distribution information that considers both the number of speakers appearing in the video and the number of speakers not appearing in the video may be obtained. Therefore, when the first speaker number related information and the second speaker number related information are applied to the sound source separation model 320, the sound source generated in the area that does not appear in the video may be separated. For example, even when the video recorded by the smart phone includes the voice of the shooter, the voice of the speaker appearing in the video and the voice of the shooter may be separated.

[0093] Reference Figure 7 , the information related to the number of speakers can be applied to at least one of the input layer 321 included in the sound source separation model 320, each of the plurality of feature layers included in the encoder 322, or the bottleneck layer 323. In other words, the information related to the number of speakers can be applied to the sound source separation model 320, such as by Figure 7, as shown by arrows 710, 720, and 730 shown in FIG. Arrow 710 indicates that information related to the number of speakers is concatenated with the mixed audio signal and applied to the input layer 321, arrow 720 indicates that information related to the number of speakers is applied to each of a plurality of feature layers included in the encoder 322, and arrow 730 indicates that information related to the number of speakers is concatenated with the output of the encoder 322 and applied to the bottleneck layer 323, and then sent to the decoder.

[0094] Figure 8 is a view for explaining a method of training a sound source characteristic analysis model according to an embodiment of the present disclosure.

[0095] Reference Figure 8 According to an embodiment of the present disclosure, the sound source characteristic analysis model 310 can be trained by comparing the training audio related information estimated from the training image signal and the training audio signal with the ground truth. The ground truth can be generated by a product operation between a plurality of probability maps generated from a plurality of spectrograms, which are generated based on each of a plurality of individual training sound sources included in the training audio signal.

[0096] In detail, the training individual sound source can be converted into a spectrogram in the time-frequency domain by short-time Fourier transform (STFT), and the corresponding probability map can be generated by applying the operation of equation 1 to each spectrogram. The true value can be generated by a multiplication operation between corresponding components of the generated probability map.

[0097] [Equation 1]

[0098] MaxClip(log(1+||F|| 2 ), 1)

[0099] where ||F||, which is the norm operation on the spectrogram F, 2 Indicates the size of the spectrogram, and MaxClip(x,1) is a function that outputs x when x is less than 1 and outputs 1 when x is equal to or greater than 1.

[0100] Fig. 9 is a flowchart of a method for processing a video according to an embodiment of the present disclosure. The method 900 for processing a video may be performed by the electronic device 200.

[0101] refer to Fig. 9In operation 910, information related to the number of speakers included in the mixed audio signal may be generated from the mixed audio signal or from the mixed audio signal and the visual information by using the speaker number analysis model 330. Information related to the number of speakers includes a probability distribution of the number of speakers. For example, information related to the number of speakers may be an N-dimensional feature vector including probabilities that the number of speakers included in the mixed sound source is 0, 1, ..., and N. According to an embodiment, the information related to the number of speakers may include at least one of first information related to the number of speakers about the mixed audio signal or second information related to the number of speakers about the visual information. As described above, operation 910 may be omitted according to the implementation method.

[0102] In operation 920, audio-related information indicating the degree of overlap of multiple sound sources included in the mixed audio signal is generated from the image signal and the mixed audio signal by using the first AI model. The audio-related information includes a graph indicating the degree of overlap of multiple sound sources corresponding to multiple speakers. Each interval of the graph has a probability value corresponding to the degree to which one of the multiple sound sources overlaps with another sound source in the time-frequency domain. In other words, each interval of the graph in the audio-related information has a probability value corresponding to the degree to which the voices (sound sources) of the speakers speaking simultaneously in the corresponding time-frequency domain overlap with each other. The term 'probability' means that the degree to which multiple sound sources overlap with each other is expressed as a value between 0 and 1.

[0103] In operation 930, at least one of the plurality of sound sources included in the mixed audio signal is separated from the mixed audio signal by applying the audio-related information to the second AI model. When operation 910 is not omitted, in operation 930, at least one of the plurality of sound sources included in the mixed audio signal is separated from the mixed audio signal by applying the audio-related information and the speaker number-related information to the second AI model.

[0104] Fig.10a and Fig.10b Experimental results on the performance of an electronic device and method for processing a video according to an embodiment of the present disclosure are shown.

[0105] Fig.10a shows the qualitative comparison results of the mixed audio signal, the true value and the audio-related information, and Fig.10b is a table showing the results of quantitative comparison. Fig.10a In the embodiment, by using each sound source included in the mixed audio signal, Figure 8 The procedure shown in generates true values.

[0106] refer to Fig.10a and Fig.10bIn all cases including the case where there are many areas where multiple sound sources overlap (Case 1), the case where there are few overlapping areas (Case 3), and the case where the overlapping areas are between Case 1 and Case 3 (Case 2), it can be seen that the audio-related information is almost the same as the true value and quite accurately reflects in which area of ​​the time-frequency domain the multiple sound sources overlap within the mixed audio signal. The accuracy of the audio-related information estimation of the sound source separation model 320 is measured to be 97.27%, and the F-score (F-measure) is 0.77.

[0107] Fig.11a and Fig.11b Experimental results comparing the performance of the electronic device and method for processing video according to the embodiments of the present disclosure with related technologies are shown.

[0108] Fig.11a shows a qualitative comparison of the results of separating the sound source of the first speaker from the mixed audio signal, and Fig.11b is a table showing a quantitative comparison. Fig.11a and Fig.11b In , “visual speech” is used as a relative for comparison.

[0109] refer to Fig.11a and Fig.11b , in the sound source classification according to the present disclosure, the audio-related information shown on the left is applied to the sound source separation model, and in "visual speech", the audio-related information is not applied to the sound source separation model. When the sound source of the first speaker is separated from the mixed audio signal in the first time-frequency region 1110, the sound source of the second speaker is included in regions 1111 and 1112 in "visual speech", but, in the separation result according to the present disclosure, the sound source of the second speaker is rarely included. When the sound source of the first speaker is separated from the mixed audio signal in the second time-frequency region 1120, the sound source of the first speaker is included in region 1121 in "visual speech", but, in the separation result according to the present disclosure, the sound source of the first speaker is rarely included. As a result of quantitative comparison, the separation performance of the separation result according to the present disclosure is improved by about 1.5 dB compared with the separation performance of "visual speech".

[0110] Fig.12 A method of controlling the playback of a video based on a result of processing the video according to an embodiment of the present disclosure is shown.

[0111] According to an embodiment of the present disclosure, the electronic device 200 may control separate sound sources to be allocated to a plurality of speakers.

[0112] refer to Fig.12, in the example, two speakers 1 and 2 included in screen 1200 output corresponding voices (sound sources), a first sound source voice #1 corresponds to the first speaker 1, and a second sound source voice #2 corresponds to the second speaker 2.

[0113] The electronic device 200 executes Fig. 9 Operations 910, 920, and 930 separate the first sound source voice #1 and the second sound source voice #2 from the mixed sound source.

[0114] Since the first speaker 1 is located on the right side and the second speaker 2 is located on the left side in the image, the electronic device 200 can match the first sound source voice #1 and the second sound source voice #2 with the first speaker 1 and the second speaker 2, respectively, and then control the left speaker 1210L to amplify and output the second sound source voice #2 and control the right speaker 1210R to amplify and output the first sound source voice #1. Since the sound source is output according to the position of the speaker, the user can better feel the sense of presence or 3D sound.

[0115] Fig.13 A method of controlling the playback of a video based on a result of processing the video according to an embodiment of the present disclosure is shown.

[0116] According to an embodiment, when the electronic device 200 receives a user input for selecting a speaker from among a plurality of speakers while playing a video, the electronic device 200 may control the output of a sound source corresponding to the selected speaker among a plurality of sound sources to be emphasized.

[0117] refer to Fig.13 In the example, two speakers 1 and 2 included in the playback screen 1300 output corresponding voices, a first sound source voice #1 corresponds to the first speaker 1, and a second sound source voice #2 corresponds to the second speaker 2.

[0118] The electronic device 200 executes Fig. 9 Operations 910, 920, and 930 separate the first sound source voice #1 and the second sound source voice #2 from the mixed sound source.

[0119] When the user selects the first speaker 1 by the selection means 1320 (such as a finger or a mouse cursor), the electronic device 200 can control the first sound source voice #1 corresponding to the first speaker 1 to be emphasized and output. Therefore, the first sound source voice #1 is amplified and output from both the left speaker 1310L and the right speaker 1310R.

[0120] exist Fig.13In the example, both the first sound source voice #1 and the second sound source voice #2 are output through the two speakers 1310L and 1310R, but only the first sound source voice #1 is amplified and output. However, when the first speaker 1 is selected, the electronic device 200 can also control to output only the first sound source voice #1 except the second sound source voice #2.

[0121] Fig.14 A method of controlling the playback of a video based on a result of processing the video according to an embodiment of the present disclosure is shown.

[0122] According to an embodiment, when the electronic device 200 receives a user input of enlarging a region displaying one of multiple speakers while playing a video, the electronic device 200 may control the output of a sound source corresponding to an object included in the enlarged region among multiple sound sources to be emphasized.

[0123] Reference Fig.14 , playback screen 1400 is obtained by Fig.13 The playback screen 1300 is a screen obtained by enlarging an area including the first speaker 1. In the example, two speakers 1 and 2 included in the entire screen before enlarging output corresponding sound sources, the first sound source voice #1 corresponds to the first speaker 1, and the second sound source voice #2 corresponds to the second speaker 2.

[0124] The electronic device 200 executes Fig. 9 Operations 910, 920, and 930 separate the first sound source voice #1 and the second sound source voice #2 from the mixed sound source.

[0125] When the user zooms in the area including the first speaker 1 through a touch input or the like and thus displays the playback screen 1400 (such as a finger or a mouse cursor), the electronic device 200 can control the first sound source voice #1 corresponding to the first speaker 1 to be emphasized and output. Therefore, only the first sound source voice #1 is output from both the left speaker 1410L and the right speaker 1410R.

[0126] exist Fig.14 , only the first sound source voice #1 except the second sound source voice #2 is output through the two speakers 1410L and 1410R. However, the electronic device 200 can control the output of both the first sound source voice #1 and the second sound source voice #2 through the left speaker 1410L and the right speaker 1410R, but only amplify and output the first sound source voice #1.

[0127] Fig.15 A method of controlling the playback of a video based on a result of processing the video according to an embodiment of the present disclosure is shown.

[0128] According to an embodiment, the electronic device 200 may display a screen on which a video is played back, may receive a user input for selecting at least one speaker from among a plurality of speakers corresponding to a plurality of sound sources included in a mixed audio signal, may display a user interface for adjusting the volume of at least one sound source corresponding to the selected at least one speaker on the screen, and receive an adjustment of the volume of at least one sound source from the user. Based on the adjustment of the volume of at least one sound source, the volume of at least one sound source output through the audio output interface may be adjusted.

[0129] refer to Fig.15 , playback screen 1500 is a screen on which a video directly shot by a user is played back. In the example, the user's voice, such as "I'm shooting", is recorded when shooting a video, and the first sound source voice #1 corresponds to the first speaker 1 included in the video, and the second sound source voice #2 corresponds to the second speaker (i.e., the shooter) not included in the video.

[0130] The electronic device 200 executes Fig. 9 Operations 910, 920, and 930 separate the first sound source voice #1 and the second sound source voice #2 from the mixed sound source.

[0131] After the user allows the interfaces 1520, 1530, and 1540 for selecting a speaker and adjusting the volume of the corresponding sound source to be displayed on the playback screen 1500 through a touch input or the like, the user can increase the volume of the first sound source voice #1 corresponding to the first speaker 1. In response to a user input for reducing the volume of the second sound source voice #2 corresponding to the second speaker outside the video, the electronic device 200 can control the first sound source voice #1 corresponding to the first speaker 1 to be emphasized and output, and control the second sound source voice #2 corresponding to the second speaker to be output at a low volume.

[0132] According to one aspect of the present disclosure, an electronic device 200 for processing a video including an image signal and a mixed audio signal, the electronic device 200 includes: at least one processor 240; and a memory 250, which is configured as at least one program for processing the video. By executing the at least one program, the at least one processor 240 is configured to: generate audio related information from the image signal and the mixed audio signal by using a first artificial intelligence (AI) model 310, the audio related information indicating the degree of overlap of multiple sound sources included in the mixed audio signal; and separate at least one sound source from the multiple sound sources included in the mixed audio signal by applying the audio related information to a second AI model 320.

[0133] In an embodiment, the audio-related information includes a map indicating the degree of overlap of the plurality of sound sources, and each interval of the map has a probability value corresponding to the degree to which one of the plurality of sound sources overlaps with another sound source in the time-frequency domain.

[0134] In an embodiment, the first AI model 310 includes: a first sub-model 311, which is configured to generate multiple mouth movement information representing temporal pronunciation information of multiple speakers corresponding to multiple sound sources from an image signal; and a second sub-model 312, which is configured to generate audio-related information from a mixed audio signal based on the multiple mouth movement information.

[0135] In an embodiment, the first AI model 310 is trained by comparing training audio-related information estimated from a training image signal and a training audio signal with a true value, and wherein the true value is generated by a product operation between a plurality of probability maps generated from a plurality of spectrograms generated based on each of a plurality of separate training sound sources included in the training audio signal.

[0136] In an embodiment, each of the plurality of probability maps is represented by MaxClip(log(1+||F|| 2 ), 1) generates, where ||F|| 2 is the size of the corresponding spectrogram from among the plurality of spectrograms, and MaxClip(x,1) is a function that outputs x when x is less than 1 and outputs 1 when x is equal to or greater than 1.

[0137] In an embodiment, the second AI model 320 includes an input layer, an encoder including multiple feature layers, and a bottleneck layer, and wherein applying audio-related information to the second AI model 320 includes applying audio-related information to the input layer, applying audio-related information to each of the multiple feature layers included in the encoder, or applying audio-related information to at least one of the bottleneck layers.

[0138] In an embodiment, at least one processor 240 is further configured to: generate information related to the number of speakers included in the mixed audio signal from the mixed audio signal or from the mixed audio signal and visual information by using a third AI model; generate audio related information from the image signal and the mixed audio signal based on information related to the number of speakers by using the first AI model 310; and separate at least one sound source of the multiple sound sources included in the mixed audio signal from the mixed audio signal by applying information related to the number of speakers and audio related information to the second AI model 320, wherein the visual information includes at least one key frame included in the image signal, and wherein at least one key frame includes a facial area, which includes the lips of at least one speaker corresponding to at least one sound source included in the mixed audio signal.

[0139] In an embodiment, the speaker quantity related information included in the mixed audio signal includes at least one of first speaker quantity related information about the mixed audio signal or second speaker quantity related information about the visual information.

[0140] In an embodiment, the first speaker quantity related information includes a probability distribution of the number of speakers corresponding to a plurality of sound sources included in the mixed audio signal, and wherein the second speaker quantity related information includes a probability distribution of the number of speakers included in the visual information.

[0141] In an embodiment, applying information related to the number of speakers to the second AI model 320 includes applying information related to the number of speakers to an input layer, applying information related to the number of speakers to each of a plurality of feature layers included in an encoder, or applying information related to the number of speakers to at least one of the bottleneck layers.

[0142] In an embodiment, at least one processor 240 is further configured to: obtain multiple mouth movement information associated with multiple speakers from the image signal; and separate at least one sound source from the multiple sound sources included in the mixed audio signal by applying the obtained multiple mouth movement information to the second AI model 320.

[0143] In an embodiment, the electronic device further includes: an input / output interface configured to display a screen on which a video is played back and to receive an input from a user for selecting at least one speaker from a plurality of speakers corresponding to a plurality of sound sources included in a mixed audio signal; and an audio output interface configured to output at least one sound source corresponding to at least one speaker selected from a plurality of sound sources included in the mixed audio signal.

[0144] In an embodiment, the at least one processor 240 is further configured to: display on a screen a user interface for adjusting the volume of at least one sound source corresponding to the selected at least one speaker, and receive an adjustment of the volume of the at least one sound source from a user; and adjust the volume of the at least one sound source output through the audio output interface based on the adjustment of the volume of the at least one sound source.

[0145] According to one aspect of the present disclosure, a method for processing a video including an image signal and a mixed audio signal includes: generating audio-related information from the image signal and the mixed audio signal by using a first artificial intelligence (AI) model, the audio-related information indicating a degree of overlap of multiple sound sources included in the mixed audio signal; and separating at least one of the multiple sound sources included in the mixed audio signal from the mixed audio signal by applying the audio-related information to a second AI model 320.

[0146] In an embodiment, the audio-related information includes a map indicating the degree of overlap of the plurality of sound sources, and each interval of the map has a probability value corresponding to the degree to which one of the plurality of sound sources overlaps with another sound source in the time-frequency domain.

[0147] In an embodiment, the generation of audio-related information includes: generating a plurality of pieces of mouth movement information representing temporal pronunciation information of a plurality of speakers corresponding to a plurality of sound sources included in a mixed audio signal from an image signal; and generating audio-related information from the mixed audio signal based on the plurality of pieces of mouth movement information.

[0148] In an embodiment, the second AI model 320 includes an input layer, an encoder including multiple feature layers, and a bottleneck layer, and wherein applying audio-related information to the second AI model 320 includes applying audio-related information to the input layer, applying audio-related information to each of the multiple feature layers included in the encoder, or applying audio-related information to at least one of the bottleneck layers.

[0149] In an embodiment, the method also includes: generating information related to the number of speakers included in the mixed audio signal from the mixed audio signal or from the mixed audio signal and visual information by using a third AI model; generating audio related information from the image signal and the mixed audio signal based on information related to the number of speakers by using the first AI model 310; and separating at least one sound source of the multiple sound sources included in the mixed audio signal from the mixed audio signal by applying information related to the number of speakers and audio related information to the second AI model 320, wherein the visual information includes at least one key frame included in the image signal, and wherein the at least one key frame includes a facial area, which includes the lips of at least one speaker corresponding to the at least one sound source included in the mixed audio signal.

[0150] In an embodiment, applying information related to the number of speakers to the second AI model 320 includes applying information related to the number of speakers to an input layer, applying information related to the number of speakers to each of a plurality of feature layers included in an encoder, or applying information related to the number of speakers to at least one of the bottleneck layers.

[0151] The machine-readable storage medium may be provided as a non-transitory storage medium. A "non-transitory storage medium" is a tangible device and simply means that it does not contain signals (e.g., electromagnetic waves). The term does not distinguish between a case where data is semi-permanently stored in a storage medium and a case where data is temporarily stored. For example, a non-transitory recording medium may include a buffer in which data is temporarily stored.

[0152] According to an embodiment of the present disclosure, a method according to various disclosed embodiments can be provided by being included in a computer program product. A computer program product as a commodity can be traded between a seller and a buyer. The computer program product is distributed in the form of a device-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or can be distributed directly and online (e.g., downloaded or uploaded) through an application store or between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) can be at least temporarily stored in a device-readable storage medium, such as a memory of a manufacturer's server, an application store's server, or a relay server, or can be temporarily generated.

Claims

1. An electronic device (200) for processing a video including an image signal and a mixed audio signal, the electronic device include: at least one processor (240); and A memory (250) configured to store at least one program for processing the video; Wherein, by executing the at least one program, the at least one processor (240) is configured to: generating audio-related information from the image signal and the mixed audio signal by using a first artificial intelligence (AI) model (310), the audio-related information indicating a degree of overlap of a plurality of sound sources included in the mixed audio signal; as well as By applying the audio-related information to a second AI model (320), at least one sound source among the plurality of sound sources included in the mixed audio signal is separated from the mixed audio signal.

2. The electronic device according to claim 1, in, The audio-related information includes a graph indicating a degree of overlap of the plurality of sound sources, and Each interval of the graph has a probability value corresponding to the degree to which one of the plurality of sound sources overlaps with another sound source in the time-frequency domain.

3. The electronic device according to claim 1 or 2, in, The first AI model (310) includes: A first sub-model (311) configured to generate, from the image signal, a plurality of pieces of mouth movement information representing temporal pronunciation information of a plurality of speakers corresponding to the plurality of sound sources; and The second sub-model (312) is configured to generate the audio-related information from the mixed audio signal based on the multiple mouth movement information.

4. The electronic device according to any one of claims 1 to 3, in, The first AI model (310) is trained by comparing training audio related information estimated from training image signals and training audio signals with true values, The true value is generated by a product operation between a plurality of probability maps generated from a plurality of spectrograms, the plurality of spectrograms being generated based on each of a plurality of separate training sound sources included in the training audio signal.

5. The electronic device according to any one of claims 1 to 4, in, Each of the plurality of probability maps is represented by MaxClip(log(1+||F|| 2 ), 1) Generate, Among them, ||F|| 2 is the size of the corresponding spectrogram among the plurality of spectrograms, and MaxClip(x,1) is a function that outputs x when x is less than 1 and outputs 1 when x is equal to or greater than 1.

6. The electronic device according to any one of claims 1 to 5, in, The second AI model (320) includes an input layer (321), an encoder (322) including a plurality of feature layers, and a bottleneck layer (323), Wherein, applying the audio-related information to the second AI model includes at least one of the following: applying the audio-related information to the input layer (321), applying the audio-related information to each of the multiple feature layers included in the encoder (322), or applying the audio-related information to the bottleneck layer (323).

7. The electronic device according to any one of claims 1 to 6, in, The at least one processor (240) is further configured to: generating information related to the number of speakers included in the mixed audio signal from the mixed audio signal or from the mixed audio signal and visual information by using a third AI model (330); generating the audio related information from the image signal and the mixed audio signal based on the speaker quantity related information by using the first AI model (310); and By applying the speaker number related information and the audio related information to the second AI model (320), at least one sound source of the plurality of sound sources included in the mixed audio signal is separated from the mixed audio signal, The visual information includes at least one key frame included in the image signal, and The at least one key frame includes a facial region, and the facial region includes lips of at least one speaker corresponding to at least one sound source included in the mixed audio signal.

8. The electronic device according to any one of claims 1 to 7, in, The speaker quantity-related information included in the mixed audio signal includes at least one of first speaker quantity-related information about the mixed audio signal or second speaker quantity-related information about the visual information.

9. The electronic device according to any one of claims 1 to 8, in, The first speaker quantity related information includes a probability distribution of the number of speakers corresponding to the plurality of sound sources included in the mixed audio signal, The second information related to the number of speakers includes a probability distribution of the number of speakers included in the visual information.

10. The electronic device according to any one of claims 1 to 9, in, The second AI model (320) includes an input layer (321), an encoder (322) including a plurality of feature layers, and a bottleneck layer (323), Wherein, applying the information related to the number of speakers to the second AI model (320) includes at least one of the following: applying the information related to the number of speakers to the input layer (321), applying the information related to the number of speakers to each of the multiple feature layers included in the encoder (322), or applying the information related to the number of speakers to the bottleneck layer (323).

11. The electronic device according to any one of claims 1 to 10, in, The at least one processor (240) is further configured to: obtaining a plurality of pieces of mouth movement information associated with the plurality of speakers from the image signal; and At least one sound source among the plurality of sound sources included in the mixed audio signal is separated from the mixed audio signal by applying the obtained plurality of pieces of mouth movement information to the second AI model.

12. The electronic device according to any one of claims 1 to 11, further comprising: include: an input / output interface (220) configured to display a screen on which the video is played back, and receive an input from a user for selecting at least one speaker from a plurality of speakers corresponding to the plurality of sound sources included in the mixed audio signal; and An audio output interface (230) is configured to output at least one sound source corresponding to the at least one speaker selected from the plurality of sound sources included in the mixed audio signal.

13. The electronic device according to any one of claims 1 to 12, in, The at least one processor (240) is further configured to: displaying a user interface on the screen, the user interface being used to adjust the volume of at least one sound source corresponding to the selected at least one speaker, and receiving an adjustment of the volume of the at least one sound source from the user; and Based on the adjustment of the volume of the at least one sound source, the volume of the at least one sound source output through the audio output interface (230) is adjusted.

14. A method (900) for processing a video comprising an image signal and a mixed audio signal, the method include: generating audio-related information from the image signal and the mixed audio signal by using a first artificial intelligence (AI) model (310), the audio-related information indicating a degree of overlap of a plurality of sound sources included in the mixed audio signal; and By applying the audio-related information to a second AI model (320), at least one sound source among the plurality of sound sources included in the mixed audio signal is separated from the mixed audio signal.

15. A computer-readable recording medium storing a computer program which, when executed by at least one processor, causes the at least one processor to perform the method of claim 14.