Method for removing target sound from input audio, and electronic device therefor
The electronic device employs an AI model to remove target sounds and their harmonic components from input audio, enhancing audio quality by addressing residual noise issues.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-10-17
- Publication Date
- 2026-04-23
AI Technical Summary
Existing electronic devices struggle to effectively remove target sounds, such as system notifications, from input audio, leading to residual components that degrade audio quality.
An electronic device uses an artificial intelligence model to detect and remove target sounds from input audio, followed by estimating the shape of the residual sound to eliminate harmonic components, thereby enhancing audio quality.
The method effectively removes target sounds and their harmonic components, resulting in improved audio quality by minimizing residual noise.
Smart Images

Figure KR2025016482_23042026_PF_FP_ABST
Abstract
Description
Method for removing a target sound from input audio and electronic device thereof
[0001] The present disclosure relates to a method for removing a target sound from input audio and an electronic device thereof.
[0002] Target sounds, such as system sounds (e.g., notifications, warnings) generated by electronic devices, are output for the purpose of conveying the device's operating status to the user. Since these target sounds can act as unnecessary components in input audio, there is a trend among electronic devices to remove them from input audio to provide users with high-quality video. Consequently, there is an increasing need for technology that precisely detects target sounds within input audio and minimizes residual components attributable to them by removing or attenuating them.
[0003] A method for removing a target sound from an input audio, performed by an electronic device according to one embodiment of the present disclosure, may include the step of obtaining a first output audio by removing the target sound from the input audio using an artificial intelligence model.
[0004] A method for removing a target sound from input audio, performed by an electronic device according to one embodiment of the present disclosure, may include the step of obtaining residual audio containing the target sound based on the difference between the input audio and the first output audio.
[0005] A method for removing a target sound from input audio, performed by an electronic device according to one embodiment of the present disclosure, may include the step of estimating the shape of the target sound based on the residual audio.
[0006] A method for removing a target sound from an input audio, performed by an electronic device according to one embodiment of the present disclosure, may include the step of determining a harmonic component of the target sound in the first output audio based on the shape of the estimated target sound.
[0007] A method for removing a target sound from an input audio, performed by an electronic device according to one embodiment of the present disclosure, may include the step of obtaining a second output audio by reducing the magnitude of the harmonic component in the first output audio.
[0008] An electronic device according to one embodiment of the present disclosure may include: a memory for storing instructions; and at least one processor operably coupled to the memory and comprising processing circuitry.
[0009] An electronic device according to one embodiment of the present disclosure can obtain a first output audio by removing the target sound from the input audio using an artificial intelligence model, by having at least one processor execute the instructions individually or collectively.
[0010] An electronic device according to one embodiment of the present disclosure can obtain residual audio containing a target sound based on the difference between the input audio and the first output audio by having at least one processor execute the instructions individually or collectively.
[0011] An electronic device according to one embodiment of the present disclosure can estimate the shape of the target sound based on the residual audio by having at least one processor execute the instructions individually or collectively.
[0012] An electronic device according to one embodiment of the present disclosure can determine harmonic components of the target sound in the first output audio based on the shape of the estimated target sound by having at least one processor execute the instructions individually or collectively.
[0013] An electronic device according to one embodiment of the present disclosure can obtain a second output audio by reducing the magnitude of the harmonic component in the first output audio by having at least one processor execute the instructions individually or collectively.
[0014] The present disclosure may be understood by the combination of the following detailed description and the accompanying drawings, where reference numerals denote structural elements.
[0015] FIG. 1 is a diagram schematically illustrating a method for removing a target sound from input audio according to one embodiment of the present disclosure.
[0016] FIG. 2 is a flowchart illustrating the process of removing a target sound from input audio using an electronic device according to one embodiment of the present disclosure.
[0017] FIG. 3 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure removing a target sound from input audio.
[0018] FIG. 4a is a flowchart illustrating the process of an electronic device according to one embodiment of the present disclosure acquiring a first output audio using an artificial intelligence model.
[0019] FIG. 4b is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure acquiring a first output audio using an artificial intelligence model.
[0020] FIG. 5a is a diagram illustrating the process of generating a training dataset for an artificial intelligence model according to one embodiment of the present disclosure.
[0021] FIG. 5b is a diagram illustrating the operation of performing pre-training of an artificial intelligence model according to one embodiment of the present disclosure.
[0022] FIG. 6 is a diagram illustrating the process of an electronic device acquiring residual audio according to one embodiment of the present disclosure.
[0023] FIG. 7 is a flowchart illustrating the process of an electronic device according to one embodiment of the present disclosure estimating the shape of a target sound.
[0024] FIG. 8a is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure estimating the shape of a target sound.
[0025] FIG. 8b is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure estimating the shape of a target sound.
[0026] FIG. 8c is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure estimating the shape of a target sound.
[0027] FIG. 9 is a flowchart illustrating the process of an electronic device according to one embodiment of the present disclosure searching for harmonic components in a first output audio.
[0028] FIG. 10a is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure searching for harmonic components in a first output audio.
[0029] FIG. 10b is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure searching for harmonic components in a first output audio.
[0030] FIG. 11 is a flowchart illustrating a method for an electronic device according to one embodiment of the present disclosure to determine a harmonic region in a first output audio.
[0031] FIG. 12a is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure determining a harmonic region within a first output audio.
[0032] FIG. 12b is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure determining a harmonic region within a first output audio.
[0033] FIG. 13a is a flowchart illustrating the process of an electronic device according to one embodiment of the present disclosure acquiring a second output audio with reduced harmonic component magnitude.
[0034] FIG. 13b is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to obtain a second output audio in which the magnitude of the harmonic component is reduced.
[0035] FIG. 14a is a flowchart illustrating the process of an electronic device according to one embodiment of the present disclosure acquiring a second output audio with reduced harmonic component magnitude.
[0036] FIG. 14b is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to obtain a second output audio in which the magnitude of the harmonic component is reduced.
[0037] FIG. 15 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure removing a target sound from input audio.
[0038] FIG. 16a is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure removing a target sound from input audio.
[0039] FIG. 16b is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to remove a target sound and harmonic components caused by the target sound from input audio.
[0040] FIG. 17a is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure removing a target sound from input audio.
[0041] FIG. 17b is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to remove a target sound and harmonic components caused by the target sound from input audio.
[0042] FIG. 18 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure removing a target sound from input audio.
[0043] FIG. 19 is a block diagram schematically illustrating the configuration of an electronic device according to one embodiment of the present disclosure.
[0044] The terms used in this disclosure will be briefly explained, and an embodiment of this disclosure will be described in detail.
[0045] Throughout this disclosure, unless specifically stated otherwise, “or” is inclusive and not exclusive. Accordingly, “A or B” may mean “A, B, or both” unless clearly indicated otherwise by the context.
[0046] In the present disclosure, the expression “at least one of a, b, or c” may refer to “a”, “b”, “c”, “a and b”, “a and c”, “b and c”, “a, b, and c all”, or variations thereof.
[0047] The terms used in this disclosure have been selected to be as widely used as possible, taking into account the functions in the embodiments of this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description section of the relevant embodiments of this disclosure. Therefore, the terms used in this disclosure should be defined not merely by their names, but based on their meanings and the content throughout this disclosure.
[0048] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as generally understood by those skilled in the art as described in this specification.
[0049] Throughout this disclosure, when a part is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, terms such as "...part," "module," etc., as used in this disclosure refer to a unit that processes at least one function or operation, and may be implemented in hardware or software, or as a combination of hardware and software.
[0050] The expression “configured to” as used in this disclosure may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware. Instead, in some situations, the expression “system configured to” may mean that the system is “capable of” together with other devices or components. For example, the phrase “a processor configured (or set) to perform A, B, and C” may mean a dedicated processor for performing said operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or an application processor) capable of performing said operations by executing one or more software programs stored in memory.
[0051] In addition, when a component is described in the present disclosure as being “connected” or “connected” to another component, it should be understood that the component may be directly connected to or directly connected to the other component, but unless otherwise specifically stated, it may also be connected or connected through another component in between.
[0052] It should be understood that the blocks in each flowchart and combinations of flowcharts can be executed by one or more computer programs containing computer-executable instructions. One or more computer programs may be stored all in a single memory or may be partitioned and stored in multiple different memories.
[0053] All functions or operations described in this document may be processed by a single processor or a combination of multiple processors. A single processor or a combination of processors is a circuitry that performs processing and may include circuitry such as an AP (Application Processor), CP (Communication Processor), GPU (Graphical Processing Unit), NPU (Neural Processing Unit), MPU (Microprocessor Unit), SoC (System on Chip), IC (Integrated Chip), etc.
[0054] Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0055] The predefined rules of operation or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a predefined rules of operation or artificial intelligence models configured to perform desired characteristics (or objectives) are created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.
[0056] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above.
[0057] Embodiments of the present disclosure are described below with reference to the attached drawings so that those skilled in the art can easily implement them. However, an embodiment of the present disclosure may be implemented in various different forms and is not limited to the embodiment described herein. Furthermore, in order to clearly explain an embodiment of the present disclosure in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the present disclosure are denoted by similar reference numerals.
[0058] Embodiments of the present disclosure will be described in detail below with reference to the drawings.
[0059] FIG. 1 is a diagram schematically illustrating a method for removing a target sound (101) from an input audio (110) according to one embodiment of the present disclosure.
[0060] Referring to FIG. 1, an electronic device (1000) according to one embodiment of the present disclosure may be a device for processing input audio (110). For example, the electronic device (1000) may be a device for removing a specific sound from the input audio (110). The specific sound may be a sound to be removed set by a user or a system. In the following, the sound to be removed may be referred to as a target sound (101).
[0061] In one embodiment of the present disclosure, the electronic device (1000) may be a device for playing audio. For example, the electronic device (1000) may be a device capable of playing audio data stored in digital form to an audio output unit, such as a speaker or earphones, through a digital-to-analog converter and an amplifier.
[0062] In one embodiment of the present disclosure, the electronic device (1000) may be a device that provides audio along with video simultaneously. For example, the electronic device (1000) may play a file (e.g., a video file) in which video data and audio data are recorded together, and at the same time, may output video to a screen and process audio data to play sound through a sound output unit. For example, the electronic device (1000) may be implemented as an electronic device (1000) of various shapes, such as a mobile device, a smartphone, a monitor, a laptop computer, a tablet PC, a wearable device, a head-mounted display (HMD) device, digital signage, etc. For example, the electronic device (1000) may be a device that captures video and provides the captured video along with recorded audio.
[0063] In one embodiment of the present disclosure, an electronic device (1000) may acquire input audio (110). For example, the input audio (110) may be audio included in a captured image captured by the electronic device (1000) or an external device. In this case, the input audio (110) may include all sounds recorded together with the image (e.g., voice, background sound, ambient noise, etc.). For example, the input audio (110) may be audio recorded by the electronic device (1000) or an external device. In this case, the input audio (110) may include sounds recorded independently through a device such as a microphone (e.g., voice, background sound, ambient noise, etc.).
[0064] In one embodiment of the present disclosure, the target sound (101) may refer to a specific sound to be removed from the original audio. The target sound (101) may include sounds that can be defined according to the user's intent, such as speech, specific background music, or mechanical noise. The target sound (101) may be defined as a set of sounds having identifiable characteristics. The target sound (101) may have a unique pattern that is distinguished from other audio components in terms of frequency and temporal characteristics.
[0065] In one embodiment of the present disclosure, the target sound (101) may be a system sound of the electronic device (1000) (or external device). The system sound may refer to a sound generated by the electronic device (1000) (or external device) itself, such as an alert sound, a camera sound, etc. For example, the target sound (101) may include a system alert sound, a vibration alert sound, such as a video recording start sound, a video recording end sound, or a photo shooting sound. However, examples of the target sound (101) are not limited thereto, and any sound that can be removed by having identifiable characteristics may be included.
[0066] In one embodiment of the present disclosure, an electronic device (1000) can obtain a first output audio (120) in which the target sound (101) in the input audio (110) is removed or attenuated based on input audio (110) and original target sound (200).
[0067] In one embodiment of the present disclosure, the original target sound (200) may correspond to an acoustic sample obtained by independently playing the target sound under a specific acoustic environment (e.g., anechoic chamber, indoor, outdoor, etc.) or specific recording conditions (e.g., fixed microphone position, certain distance, fixed device setting) and collecting the played signal through an actual recording device.
[0068] In one embodiment of the present disclosure, an electronic device (1000) can obtain information of the original target sound (200). For example, the information of the original target sound (200) may be stored in the memory of the electronic device (1000) or in a predefined database, and the electronic device (1000) may obtain the information of the original target sound (200) by loading or referencing the information of the original target sound (200) stored in the memory or database. Alternatively, for example, the electronic device (1000) may receive information of the original target sound (200) from an external server or an external electronic device.
[0069] In one embodiment of the present disclosure, the information of the acquired original target sound (200) may be data in the form of a spectrogram. For example, the information of the original target sound (200) may include time information, frequency information, and energy (or intensity) information.
[0070] In one embodiment of the present disclosure, an electronic device (1000) inputs input audio (110) and original target sound (200) into an artificial intelligence model to obtain a first output audio (120) in which the target sound (101) in the input audio (110) is removed or attenuated within the input audio (110).
[0071] In one embodiment of the present disclosure, the first output audio (120) (or input audio (110)) may include not only the target sound (101) but also the harmonic components (102) of the target sound (101). The harmonic components (102) of the target sound (101) may be due to waveform distortion that occurs when the fundamental frequency (f0) of the target sound (101) passes through a non-linear medium or a non-linear signal path during the actual acoustic output process. For example, waveform distortion may occur due to the non-linear driving characteristics of the speaker, the saturation operation of the amplifier circuit, or the non-linear acoustic radiation characteristics caused by the housing and duct structure in which the speaker is mounted. Due to such waveform distortion, harmonic components (102) corresponding to integer multiples (2f0, 3f0, ...) of the fundamental frequency (f0) of the target sound (101) may be generated. When harmonic components (102) caused by the target sound (101) remain in the audio, reverberation or residual sound may be perceived in the actual audio.
[0072] In one embodiment of the present disclosure, an artificial intelligence model may be stored (or mounted) in an electronic device (1000), and the artificial intelligence model may be a lightweight model for removing a target sound (101). When the target sound (101) is modified, the lightweight artificial intelligence model may not be able to detect the modified part and thus may not be able to remove it from the input audio (110). For example, the lightweight artificial intelligence model may not detect the harmonic component (102) of the target sound (101) and thus may not be able to remove it from the input audio (110), and the harmonic component (102) of the target sound (101) may remain in the first output audio (120).
[0073] In one embodiment of the present disclosure, the electronic device (1000) can obtain a second output audio (130) in which a harmonic component (102) of a target sound (101) is removed from a first output audio (120).
[0074] In one embodiment of the present disclosure, an electronic device (1000) can estimate the shape of a target sound (101) within an input audio (110) based on residual audio obtained by removing a first output audio (120) from an input audio (110). Based on the estimated shape of the target sound (101), the electronic device (1000) can search for a harmonic component (102) caused by the target sound (101) within the first output audio (120). By removing or attenuating the harmonic component (102) found within the first output audio (120), the electronic device (1000) can generate a second output audio (130) in which the harmonic component (102) caused by the target sound (101) is removed or attenuated.
[0075] According to one embodiment of the present disclosure, an electronic device (1000) can obtain a final output audio (i.e., a second output audio (130)) in which the target sound (101) is more effectively removed or attenuated from an input audio (110) by first removing the target sound (101) using an artificial intelligence model and then additionally removing or attenuating the harmonic component (102) caused by the target sound (101).
[0076] FIG. 2 is a flowchart illustrating a method by which an electronic device (1000) according to one embodiment of the present disclosure removes a target sound from an input audio (110). FIG. 3 is a diagram illustrating an operation by which an electronic device (1000) according to one embodiment of the present disclosure removes a target sound from an input audio (110).
[0077] Referring to FIG. 2, a method for an electronic device (1000) to remove a target sound (101) from an input audio (110) may include steps S210 to S250. In one embodiment of the present disclosure, steps S210 to S250 may be executed by at least one processor (1920, see FIG. 19) included in the electronic device (1000). A method for removing a target sound (101) from an input audio (110) is not limited to that illustrated in FIG. 2 and, in one or more embodiments, may further include steps not illustrated in FIG. 2.
[0078] Hereinafter, with reference to FIG. 2 and FIG. 3 together, the function and / or operation of the electronic device (1000) of the present disclosure removing or attenuating the target sound (101) and the harmonic component (102) of the target sound (101) from the input audio (110) will be described in detail.
[0079] In step S210 of FIG. 2, the electronic device (1000) can obtain a first output audio (120) by removing the target sound (101) from the input audio (110) using an artificial intelligence model (115).
[0080] Referring together to operation 1 of the embodiment illustrated in FIG. 3, the electronic device (1000) can input an original target sound (200) corresponding to the input audio (110) and the target sound (101) into an artificial intelligence model (115). The artificial intelligence model (115) can remove the target sound (101) from the input audio (110) based on the input audio (110) and the original target sound (200). In one embodiment of the present disclosure, the artificial intelligence model (115) may include a lightweight model stored in the electronic device (1000).
[0081] The original target sound (200) may correspond to an acoustic sample obtained by independently playing the target sound under a specific acoustic environment (e.g., anechoic room, indoor, outdoor, etc.) or specific recording conditions (e.g., fixed microphone position, certain distance, fixed device setting) and collecting the played signal through an actual recording device.
[0082] An electronic device (1000) can acquire an audio signal corresponding to an input audio (110). The acquired audio signal may be an analog signal. The electronic device (1000) can convert the acquired audio signal from an analog signal into a digital signal. The electronic device (1000) can divide the converted digital signal into short frames that are easy to process in the time domain. The divided frames may be configured to overlap each other at a certain ratio. For example, the divided frames may be generated by shifting 20ms of length frames by 10ms. In this case, signal distortion that may occur at the boundary between frames can be minimized. However, the embodiment is not limited thereto. The electronic device (1000) can convert the divided frames into a spectrogram in the time-frequency domain through a Fast Fourier Transform (FFT). The spectrogram may include frequency information and amplitude information of the audio signal corresponding to the input audio (110).
[0083] In one embodiment of the present disclosure, an electronic device (1000) may input a spectrogram converted from an input audio (110) and a spectrogram corresponding to an original target sound (200) to an artificial intelligence model (115). Based on the spectrogram of the input original target sound (200), the artificial intelligence model (115) may estimate the region (and / or ratio) in which the actual target sound (101) exists within the input audio (110) and output a first output audio (120) in which the estimated target sound (101) within the input audio (110) is removed or attenuated.
[0084] In step S220 of FIG. 2, the electronic device (1000) can obtain residual audio (125) containing the target sound (101) by using the difference between the input audio (110) and the first output audio (120). The electronic device (1000) can obtain residual audio (125) containing the target sound (101) based on the difference between the input audio (110) and the first output audio (120).
[0085] Referring together with operation ② of the embodiment illustrated in FIG. 3, the electronic device (1000) can subtract the first output audio (120) from the input audio (110) to generate residual audio (125).
[0086] In one embodiment of the present disclosure, the electronic device (1000) can generate residual audio (125) in a time domain difference manner. That is, the electronic device (1000) can generate residual audio (125) by subtracting the audio waveform itself between the input audio (110) and the first output audio (120).
[0087] Alternatively, in one embodiment of the present disclosure, the electronic device (1000) may generate residual audio (125) using a frequency domain difference method. That is, the electronic device (1000) may generate residual audio (125) by converting each of the input audio (110) and the first output audio (120) into spectrograms and then subtracting complex values in each frequency-time bin. The complex values may include both amplitude information and phase information of the audio signal.
[0088] The electronic device (1000) can generate residual audio (125) containing only the separated target sound (101) through the artificial intelligence model (115). Through this, the electronic device (1000) can obtain information regarding the target sound (101) in the environment where the actual audio was recorded through the residual audio (125).
[0089] In step S230 of FIG. 2, the electronic device (1000) can estimate the shape of the target sound (101) based on the residual audio (125).
[0090] Referring together to operation ③ of the embodiment illustrated in FIG. 3, the electronic device (1000) can estimate the shape of the target sound (101) based on the residual audio (125) from which the target sound (101) removed through the artificial intelligence model (115) has been separated. The shape of the target sound (101) can collectively refer to characteristics that make the corresponding acoustic signal distinguishable from other sounds, such as the time domain waveform, frequency spectrum distribution, phase characteristics, and amplitude modulation pattern of the target sound (101). For example, the shape of the target sound (101) may correspond to a pattern representing the distribution characteristics of the target sound (101) in the time-frequency domain. In this case, the shape of the target sound (101) may include temporal changes by frequency band expressed in a spectrogram.
[0091] In one embodiment of the present disclosure, an electronic device (1000) may obtain a pattern map (300) corresponding to the shape of a target sound (101). The pattern map (300) may correspond to a map representing the distribution of frequency components over time associated with the target sound (101) in the time-frequency domain. The horizontal axis (x-axis) of the pattern map (300) may represent time information, and the vertical axis (y-axis) of the pattern map (300) may represent frequency information. The size of the pattern map (300) may be determined based on residual audio (125) by the time interval in which the target sound (101) was generated and the frequency interval of the target sound (101).
[0092] In step S240 of FIG. 2, the electronic device (1000) can determine (e.g., search) the harmonic components of the target sound (101) within the first output audio (120) based on the shape of the estimated target sound (101).
[0093] Referring together with operation ④ of the embodiment illustrated in FIG. 3, the electronic device (1000) can search for harmonic components (102) of the target sound (101) within the first output audio (120) using a pattern map (300) corresponding to the shape of the target sound (101).
[0094] The electronic device (1000) can estimate a target interval in which a target sound (101) exists within the input audio (110) based on residual audio (125). The target interval may refer to a temporal interval, but the embodiment is not limited thereto and may include a frequency interval as well as a temporal interval. The electronic device (1000) can search for a harmonic component (102) by using a filter corresponding to the pattern map (300) and moving the filter within the target interval by a predetermined (e.g., predetermined) frequency interval.
[0095] For example, a filter corresponding to a pattern map (300) may include a pattern region where harmonic components (102) exist and a surrounding region other than the pattern region. The electronic device (1000) can calculate an attention score corresponding to the ratio of the average of the first spectrogram magnitude of the region matching the pattern region and the average of the second spectrogram magnitude of the region matching the surrounding region in each of the window regions that are sequentially matched with the filter corresponding to the pattern map (300) within the target interval.
[0096] The electronic device (1000) can identify (or determine) an area among the window regions that contains a harmonic component (102) based on the attention score of each of the window regions. The electronic device (1000) can determine at least one window region among the window regions as at least one harmonic region that contains a harmonic component (102). For example, the electronic device (1000) can determine at least one window region among the window regions where the attention score corresponds to a higher preset ratio as at least one harmonic region.
[0097] In step S250 of FIG. 2, the electronic device (1000) can obtain a second output audio (130) by reducing the magnitude of the harmonic component (102) in the first output audio (120).
[0098] Referring together with operation ⑤ of the embodiment illustrated in FIG. 3, the electronic device (1000) can reduce the magnitude in at least one harmonic region determined in the first output audio (120).
[0099] In one embodiment of the present disclosure, the electronic device (1000) can change the first spectrogram magnitude to the average of the second spectrogram magnitude. That is, the electronic device (1000) can change the magnitude in the area matching the pattern area to the average magnitude in the area matching the surrounding area. The operation of the electronic device (1000) changing the magnitude in the area matching the pattern area to the average magnitude in the area matching the surrounding area will be examined in detail later with reference to FIG. 13a and FIG. 13b.
[0100] Alternatively, in one embodiment of the present disclosure, the electronic device (1000) may reduce the first spectrogram size based on the magnitude of the target sound (101). For example, the electronic device (1000) may estimate the magnitude of the target sound (101) based on residual audio (125). The electronic device (1000) may reduce the magnitude in the area matching the pattern area based on the estimated magnitude of the target sound (101). For example, the electronic device (1000) may reduce the magnitude in the area matching the pattern area by "1 / magnitude of the estimated target sound (101)". The operation of the electronic device (1000) reducing the first spectrogram size based on the magnitude of the target sound (101) will be examined in detail later with reference to FIGS. 14a and FIGS. 14b.
[0101] According to one embodiment of the present disclosure, an electronic device (1000) can remove the harmonic components (102) caused by the remaining target sound (101) after applying the original target sound (200) and input audio (110) to an artificial intelligence model (115) to remove the target sound (101) within the input audio (110). At this time, the electronic device (1000) can more accurately estimate the shape of the actual recorded target sound (101) by estimating the shape of the target sound (101) based on the residual audio (125), even if the recording environment of the original target sound (200) and the recording (or filming) environment of the input audio (110) are different from each other. Accordingly, the electronic device (1000) can also more accurately search for the harmonic components caused by the target sound (101) within the input audio (110) (or the first output audio (120)) based on the more accurately estimated shape of the target sound (101). The electronic device (1000) can provide audio or video of improved quality by more effectively removing not only the target sound (101) but also the harmonic components (102) caused by the target sound (101), thereby removing sounds that are recorded regardless of the user's intention.
[0102] FIG. 4a is a flowchart illustrating a method in which an electronic device (1000) according to one embodiment of the present disclosure obtains a first output audio (120) using an artificial intelligence model (115). FIG. 4b is a diagram illustrating an operation in which an electronic device (1000) according to one embodiment of the present disclosure obtains a first output audio (120) using an artificial intelligence model (115).
[0103] Step S410 illustrated in FIG. 4a is an operation that embodies the operation of step S210 of FIG. 2. In one embodiment of the present disclosure, step S410 may be executed by at least one processor (1920, see FIG. 19) included in an electronic device (1000). After step S410 of FIG. 4a is performed, step S220 of FIG. 2 may be performed.
[0104] In step S410 of FIG. 4a, an electronic device (1000) according to one embodiment of the present disclosure can obtain a first output audio (120) in which the target sound (101) is removed from the input audio (110) by applying an original target sound (200) corresponding to the input audio (110) and the target sound (101) to an artificial intelligence model (115).
[0105] Referring together with FIG. 4b, in one embodiment of the present disclosure, an electronic device (1000) may input an input audio (110) and an original target sound (200) to an artificial intelligence model (115). For example, the electronic device (1000) may input a spectrogram of the input audio (110) and a spectrogram of the original target sound (200) to the artificial intelligence model (115). Based on the original target sound (200), the artificial intelligence model (115) may output a first output audio (120) in which the target sound (101) within the input audio (110) is removed or attenuated from the input audio (110). For example, the artificial intelligence model (115) may output a spectrogram of the first output audio (120).
[0106] The electronic device (1000) can obtain a first output audio (120) in which the target sound (101) within the input audio (110) is removed or attenuated through an artificial intelligence model (115). For example, the electronic device (1000) can obtain a spectrogram of the first output audio (120) through an artificial intelligence model (115).
[0107] According to one embodiment of the present disclosure, an artificial intelligence model (115) may be stored (or mounted) in an electronic device (1000), and the artificial intelligence model (115) may be a lightweight model for removing a target sound (101). The harmonic component (102) of the target sound (101) is a component modified from the target sound (101), and the lightweight model may not recognize the harmonic component (102) of the target sound (101) as a target for removal from the input audio (110). Accordingly, the harmonic component (102) of the target sound (101) may remain in the first output audio (120) output by the lightweight model.
[0108] FIG. 5a is a diagram illustrating a method for generating a training dataset of an artificial intelligence model (115) according to one embodiment of the present disclosure. FIG. 5b is a diagram illustrating an operation for performing prior training of an artificial intelligence model (115) according to one embodiment of the present disclosure.
[0109] Referring to FIG. 5a, in one embodiment of the present disclosure, the artificial intelligence model (115) may be a pre-trained model. The artificial intelligence model (115) may be a model that has completed pre-training using a training dataset. In one embodiment of the present disclosure, the training dataset may include a training original target sound (510), a training original audio (530), and a plurality of training synthetic audios (540_1 to 540_n).
[0110] A plurality of synthetic audio samples for training (540_1 to 540_n) may be generated based on a plurality of target sound samples (520_1 to 520_n) and original audio for training (530). The plurality of target sound samples (520_1 to 520_n) may be audio data obtained by collecting (e.g., recording) the target sound under various environmental conditions. For example, the plurality of target sound samples (520_1 to 520_n) may include a first target sound sample (520_1) obtained by collecting the target sound in a first environment, a second target sound sample (520_2) obtained by collecting the target sound in a second environment, and an nth target sound sample (520_n) obtained by collecting the target sound in an nth environment. The first environment, the second environment, ..., and the nth environment may be various different environments. For example, various environments may include indoor environments, outdoor environments, quiet environments, environments with a lot of background noise, environments with a lot of reverberation, and environments played by various playback devices, but the embodiments are not limited thereto.
[0111] A plurality of synthetic audio samples for training (540_1 to 540_n) may be data obtained by synthesizing a plurality of target sound samples (520_1 to 520_n) to original audio samples for training (530) that do not include the target sound. For example, the original audio samples for training (530) may refer to an audio signal obtained when the target sound is not played while collecting ambient sounds (e.g., recording). For example, a plurality of synthetic audio for learning (540_1 to 540_n) may include a first synthetic audio for learning (540_1) obtained by synthesizing a first target sound sample (520_1) with the original audio for learning (530), a second synthetic audio for learning (540_2) obtained by synthesizing a second target sound sample (520_2) with the original audio for learning (530), ..., and an nth synthetic audio for learning (540_n) obtained by synthesizing an nth target sound sample (520_n) with the original audio for learning (530).
[0112] Referring to FIG. 5b, a training dataset (550) can be configured for pre-training an artificial intelligence model (115). The training dataset (550) may consist of input variables and labels. The labels may correspond to correct data for the input variables. For example, a plurality of training synthetic audios (540_1 to 540_n), including first to nth training synthetic audios (540_1 to 540_n), may be set as the first input variable. The training original target sound (510) may be set as the second input variable. The training original audio (530) may be set as the label. The training dataset (550) may be configured based on a combination of the first input variable, the second input variable, and the label.
[0113] The artificial intelligence model (115) may be a model that has been pre-trained based on a training dataset (550). In this case, the pre-training process may include steps of optimizing parameters so that the artificial intelligence model (115) learns the mapping relationship between input variables (e.g., a first input variable and a second input variable) and labels.
[0114] For example, the first input variable and the second input variable can be transmitted to the input layer of the artificial intelligence model (115). The artificial intelligence model (115) can calculate a predicted value based on the first and second input variables, and optimize the performance of the artificial intelligence model (115) by repeatedly updating the weight and bias parameters of the artificial intelligence model (115) in a direction that minimizes the error between the calculated predicted value and the label. Through this, the artificial intelligence model (115) can learn the relationship between the original target sound and the target sound in the actual collection environment.
[0115] Through this, the artificial intelligence model (115) can be trained to output audio with the target sound removed from the input audio (i.e., the first output audio) based on the input audio and the target sound collected in the actual collection environment (e.g., the original target sound).
[0116] FIG. 6 is a drawing illustrating a method for an electronic device (1000) according to one embodiment of the present disclosure to acquire residual audio (125).
[0117] Referring to FIG. 6, in one embodiment of the present disclosure, an electronic device (1000) can obtain residual audio (125) by removing a first output audio (120) from an input audio (110). The first output audio (120) may be audio in which a target sound (101) predicted by an artificial intelligence model (115) within the input audio (110) is removed or attenuated. By obtaining residual audio (125) by removing the first output audio (120) from the input audio (110), the electronic device (1000) can obtain a distribution of spectrograms regarding the target sound (101) predicted by the artificial intelligence model (115).
[0118] Meanwhile, in one embodiment of the present disclosure, the artificial intelligence model (115) fails to predict the harmonic component (102) caused by the target sound (101) in the input audio (110), so the harmonic component (102) caused by the target sound (101) may remain in the first output audio (120) without being removed. Since the harmonic component (102) caused by the target sound (101) is distributed in a frequency band different from the frequency band of the target sound (101), it may be difficult to remove it through the artificial intelligence model (115).
[0119] The distribution of the spectrogram regarding the target sound (101) in the actual recording environment may differ somewhat from the distribution of the spectrogram regarding the original target sound (200). That is, the distribution of the spectrogram regarding the target sound (101) in the input audio (110) may differ somewhat from the distribution of the spectrogram regarding the original target sound (200). According to one embodiment of the present disclosure, the electronic device (1000) can obtain the distribution of the spectrogram regarding the target sound (101) in the actual recording environment within the input audio (110) by obtaining residual audio (125). Through this, the electronic device (1000) can estimate the shape of the target sound (101) actually recorded in the input audio (110) based on the distribution of the spectrogram regarding the target sound (101) in the actual recording environment.
[0120] The shape of the harmonic component (102) attributable to the target sound (101) may be similar to the shape of the target sound (101). According to one embodiment of the present disclosure, by estimating the shape of the target sound (101) in the actual recording environment based on residual audio (125), a harmonic component (102) similar to the shape of the target sound (101) actually recorded in the first output audio (120) (or input audio (110)) can be more accurately retrieved.
[0121] FIG. 7 is a flowchart illustrating a method for an electronic device (1000) according to an embodiment of the present disclosure to estimate the shape of a target sound. FIG. 8a is a diagram illustrating an operation in which an electronic device (1000) according to an embodiment of the present disclosure estimates the shape of a target sound (811). FIG. 8b is a diagram illustrating an operation in which an electronic device (1000) according to an embodiment of the present disclosure estimates the shape of a target sound (812). FIG. 8c is a diagram illustrating an operation in which an electronic device (1000) according to an embodiment of the present disclosure estimates the shape of a target sound (813).
[0122] Step S710 illustrated in FIG. 7 is an operation that embodies the operation of step S230 of FIG. 2. In one embodiment of the present disclosure, step S710 may be executed by at least one processor (1920, see FIG. 19) included in an electronic device (1000). Step S710 of FIG. 7 may be performed after step S220 of FIG. 2 has been performed. After step S710 of FIG. 7 has been performed, step S240 of FIG. 2 may be performed.
[0123] In step S710 of FIG. 7, an electronic device (1000) according to one embodiment of the present disclosure may obtain a pattern map corresponding to the shape of a target sound (101). The electronic device (1000) may estimate the shape of the target sound (101) within the input audio (110) based on the residual audio (125) obtained by removing the first output audio (120) from the input audio (110). For example, the electronic device (1000) may generate a pattern map corresponding to the shape of the target sound (101) estimated based on the residual audio (125).
[0124] Referring together to FIGS. 8a through 8c, in one embodiment of the present disclosure, pattern maps (821, 822, 823) may correspond to a map representing the distribution of frequency components over time of a target sound (811, 812, 813) in the time-frequency domain. The horizontal axis (x-axis) of the pattern maps (821, 822, 823) may represent time information. The vertical axis (y-axis) of the pattern maps (821, 822, 823) may represent frequency information. The time axis unit represented by the horizontal axis and the frequency axis unit represented by the vertical axis of the pattern maps (821, 822, 823) may vary depending on parameters set during the process of performing a Fast Fourier Transform (FFT). For example, the unit interval of the time axis may vary by the hop size of the window frame, the sampling frequency, etc. For example, the unit interval of the frequency axis can vary depending on the sampling frequency, the number of FFT points, etc.
[0125] In one embodiment of the present disclosure, the size of a cell determined by the unit interval of the time axis and the unit interval of the frequency axis in the pattern map (821, 822, 823) may correspond to the size of the time-frequency bin in the spectrogram. A cell in the pattern map (821, 822, 823) may refer to a minimum interval determined by the unit interval of the time axis and the frequency axis on the pattern map. However, the embodiment is not limited thereto, and the size of the cell in the pattern map (821, 822, 823) may be set differently from the size of the time-frequency bin in the spectrogram.
[0126] In one embodiment of the present disclosure, each cell of the pattern map (821, 822, 823) may be recorded with a binary value regarding the presence or absence of an acoustic signal in a corresponding time interval and a corresponding frequency band. For example, in each cell of the pattern map (821, 822, 823), if the energy of the acoustic signal is greater than or equal to a predefined threshold, the value of the corresponding cell may be set to 1 (corresponding to light brightness in FIG. 8a to 8c), and if the energy of the acoustic signal is less than a predefined threshold, the value of the corresponding cell may be set to 0 (corresponding to dark brightness in FIG. 8a to 8c).
[0127] Cells marked with 1 in the pattern map (821, 822, 823) may indicate that the target sound (811, 812, 813) exists in the corresponding time-frequency range. The temporal / spatial distribution of the cells marked with 1 in the pattern map (821, 822, 823) may form a specific type of pattern. Through this, the electronic device (1000) can estimate the shape of the target sound (811, 812, 813) based on the pattern map (821, 822, 823). The shape of the target sound (811, 812, 813) may refer to characteristics in the time-frequency domain that allow the target sound (811, 812, 813) to be distinguished from other sounds. For example, the form of the target sound (811, 812, 813) may include a temporal change pattern of a specific frequency band. The form of the target sound (811, 812, 813) may include a temporal change pattern of the frequency distribution, such as a pattern in which a specific frequency band appears or disappears along the time axis.
[0128] In one embodiment of the present disclosure, the electronic device (1000) can search for an area where the target sound (811, 812, 813) exists in the spectrogram of the first output audio (120) based on the shape of the target sound (811, 812, 813) estimated through the pattern map (821, 822, 823). For example, the electronic device (1000) can search for an area where the target sound (811, 812, 813) exists within the spectrogram of the first output audio (120) by using a filter (831, 832, 833) corresponding to the pattern map (821, 822, 823). Filters (831, 832, 833) corresponding to pattern maps (821, 822, 823) can be implemented as a structure that performs matching operations between comparison targets based on patterns on the temporal-spatial distribution of the pattern maps (821, 822, 823). For example, filters (831, 832, 833) can be implemented as two-dimensional coefficient matrices.
[0129] In one embodiment of the present disclosure, the size of the filter (831, 832, 833) may be set based on the pattern map (821, 822, 823). For example, the filter (831, 832, 833) may be configured to have dimensions equal to the number of rows and columns of the pattern map (821, 822, 823), thereby allowing each element of the filter (831, 832, 833) to be mapped 1:1 to each cell of the pattern map (821, 822, 823) so that operations can be performed. As the filter (831, 832, 833) is determined based on the cell arrangement of the pattern map (821, 822, 823), a matching operation based on the shape of the target sound (811, 812, 813) in the pattern map (821, 822, 823) can be performed using the filter (831, 832, 833).
[0130] FIG. 8a illustrates, by way of example, a first target sound (811) corresponding to a video recording start sound, a first pattern map (821) corresponding to the first target sound (811), and a first filter (831). The video recording start sound is a sound that is played at the time when recording begins, and is a signal sound intended to notify the user that recording has started. For example, the video recording start sound may be a sound that is played at the time when the user presses the recording button. Alternatively, for example, the video recording start sound may be a sound that is played at the time when the device automatically starts video recording (e.g., scheduled recording, sensor detection, event trigger, etc.).
[0131] FIG. 8a illustrates, as an example, that the first target sound (811) corresponding to the video recording start sound is implemented in a form close to a beep, but the embodiment is not limited thereto.
[0132] The temporal range of the first pattern map (821) can be determined by the temporal interval in which the first target sound (811) occurs within a spectrogram containing the first target sound (811), and the frequency range of the first pattern map (821) can be determined by the frequency interval of the first target sound (811) within a spectrogram containing the first target sound (811). The size of the first filter (831) can be determined in correspondence with the temporal range and frequency range of the first pattern map (821).
[0133] FIG. 8b illustrates, by way of example, a second target sound (812) corresponding to a video recording end sound, a second pattern map (822) corresponding to the second target sound (812), and a second filter (832). The video recording end sound is a sound that is played when recording is stopped or recording is terminated, and is a signal sound intended to notify the user that recording has ended. For example, the video recording end sound may be a sound that is played when the user presses the recording stop button or the recording end button. Alternatively, for example, the video recording end sound may be a sound that is played when the device automatically stops or terminates video recording (e.g., scheduled recording, sensor detection, event trigger, etc.).
[0134] FIG. 8b illustrates, by way of example, that the second target sound (812) corresponding to the video recording end sound is implemented in the form of two consecutive tones or melodies (e.g., two lowering tones), but the embodiment is not limited thereto.
[0135] The temporal range of the second pattern map (822) can be determined by the temporal interval in which the second target sound (812) occurs within the spectrogram containing the second target sound (812), and the frequency range of the second pattern map (822) can be determined by the frequency interval of the second target sound (812) within the spectrogram containing the second target sound (812). The size of the second filter (832) can be determined in correspondence with the temporal range and frequency range of the second pattern map (822).
[0136] FIG. 8c illustrates, by way of example, a third target sound (813) corresponding to a photo shutter sound, a third pattern map (823) corresponding to the third target sound (813), and a third filter (833). The photo shutter sound is a sound that is played at the time a still image is taken, and is a signal sound intended to make the user aware of the time of shooting. For example, the photo shutter sound may be a sound that is played at the time the user presses the shutter button. Or, for example, the photo shutter sound may be a sound that is played at the time the device automatically takes a still image (e.g., scheduled shooting, sensor detection, event tree, etc.).
[0137] FIG. 8c illustrates, as an example, that a third target sound (813) corresponding to a photo shutter sound is implemented in the form of a shutter sound having a relatively wideband frequency component, but the embodiment is not limited thereto.
[0138] The temporal range of the third pattern map (823) can be determined by the temporal interval in which the third target sound (813) occurs within the spectrogram containing the third target sound (813), and the frequency range of the third pattern map (823) can be determined by the frequency interval of the third target sound (813) within the spectrogram containing the third target sound (813). The size of the third filter (833) can be determined in correspondence with the temporal range and frequency range of the third pattern map (823).
[0139] FIG. 9 is a flowchart illustrating a method for an electronic device (1000) according to one embodiment of the present disclosure to search for harmonic components within a first output audio (120). FIG. 10a is a diagram illustrating an operation for an electronic device (1000) according to one embodiment of the present disclosure to search for harmonic components within a first output audio (120). FIG. 10b is a diagram illustrating an operation for an electronic device (1000) according to one embodiment of the present disclosure to search for harmonic components within a first output audio (120).
[0140] Steps S910 and S920 illustrated in FIG. 9 are operations that embody the operation of step S240 of FIG. 2. In one embodiment of the present disclosure, steps S910 and S920 may be executed by at least one processor (1920, see FIG. 19) included in an electronic device (1000). Step S910 of FIG. 9 may be performed after step S230 of FIG. 2 has been performed. Step S250 of FIG. 2 may be performed after step S920 of FIG. 9 has been performed.
[0141] In step S910 of FIG. 9, an electronic device (1000) according to one embodiment of the present disclosure can estimate a target section in which a target sound (101) exists in a spectrogram of an input audio (110) based on residual audio (125).
[0142] Referring to FIG. 10a and FIG. 10b, in one embodiment of the present disclosure, an electronic device (1000) may determine a target interval (1010) based on a temporal interval in which the target sound (101) is presumed to exist within the input audio (110). For example, the target interval (1010) may be determined as a temporal interval in which the target sound (101) is presumed to exist within the residual audio (125). However, the embodiment is not limited thereto, and the target interval (1010) may be determined to include additional time before and / or after the temporal interval in which the target sound (101) is presumed to exist. For example, the target interval (1010) may be set as a interval between a point in time that is a predetermined (e.g., predetermined) time earlier than the time when the target sound (101) starts and a point in time that is a predetermined (e.g., predetermined) time later than the time when the target sound (101) ends.
[0143] FIGS. 10a and FIGS. 10b show a target interval (1010) determined based on residual audio (125) within the first output audio (120). FIGS. 10a and FIGS. 10b exemplarily illustrate that the target interval (1010) is determined to be the same as the temporal interval in which the target sound (101) exists within the residual audio (125).
[0144] In step S920 of FIG. 9, an electronic device (1000) according to one embodiment of the present disclosure can search for harmonic components by using a filter (1020) corresponding to a pattern map and moving the filter (1020) within a target interval (1010) by a predetermined (e.g., predetermined) frequency interval (1030). The filter (1020) corresponding to the pattern map may be implemented as a structure that performs a matching operation between comparison targets based on a temporal-spatial distribution on the pattern map. For example, the filter (1020) corresponding to the pattern map may be a structure in which each element of the filter (1020) is mapped 1:1 to each cell of the pattern map. For example, the filter (1020) may be implemented as a two-dimensional coefficient matrix.
[0145] Referring together with FIG. 10a, in one embodiment of the present disclosure, an electronic device (1000) can sequentially search the entire frequency band within a target interval (1010) using a filter (1020). For example, the electronic device (1000) can sequentially search the target interval (1010) by moving the filter (1020) within the target interval (1010) from an upper frequency (or maximum frequency) to a lower frequency (or minimum frequency) by a predetermined (e.g., predetermined) frequency interval (1030). Or, for example, the electronic device (1000) can sequentially search the target interval (1010) by moving the filter (1020) within the target interval (1010) from a lower frequency (or minimum frequency) to an upper frequency (or maximum frequency) by a predetermined (e.g., predetermined) frequency interval (1030).
[0146] Referring to FIG. 10b, in one embodiment of the present disclosure, the electronic device (1000) may sequentially search only some frequency bands (1015) within a target interval (1010) using a filter (1020). For example, the electronic device (1000) may search only some frequency bands (1015) above the frequency band of the target sound (101), but the range of some frequency bands (1015) is not limited thereto.
[0147] Referring together to FIG. 10a and FIG. 10b, in one embodiment of the present disclosure, the structure of the filter (1020) is determined based on the cell arrangement of the pattern map, so that the electronic device (1000) can perform a matching operation in the target section (1010) based on the shape of the target sound (101) in the pattern map using the filter (1020). The electronic device (1000) can perform the matching operation by moving the filter (1020) within the target section (1010) from an upper frequency to a lower frequency by a predetermined (e.g., predetermined) frequency interval (1030). Based on the result of the matching operation, the electronic device (1000) can determine (or determine) whether there is a harmonic component of the target sound (101) in the section (or region) that matches the filter (1020) within the target section (1010).
[0148] Areas that are sequentially matched with the filter (1020) within the target section (1010) can be defined as window areas (1050). The electronic device (1000) can match the shape of the spectrogram in each window area (1050) with the filter (1020). Based on the degree of matching between each window area (1050) and the filter (1020), the electronic device (1000) can determine (or identify, determine) the presence or absence of harmonic components caused by the target sound (101). A method for determining the presence or absence of harmonic components caused by the target sound (101) will be described later with reference to FIGS. 11 to 12b. FIGS. 10a and 10b exemplarily illustrate that three harmonic components (1041, 1042, 1043) are detected within the target section (1010) of the first output audio (120).
[0149] According to one embodiment of the present disclosure, an electronic device (1000) can search for harmonic components (1041, 1042, 1043) of a target sound (101) in a first output audio (120) by using a filter (1020) corresponding to a pattern map that reflects the shape of an actual recorded target sound. As the harmonic components (1041, 1042, 1043) of the target sound (101) are similar to the shape of the target sound (101), the harmonic components (1041, 1042, 1043) of the target sound (101) in the first output audio (120) can be detected more accurately.
[0150] Meanwhile, FIG. 10a and FIG. 10b illustrate, by way of example, that the target interval (1010) is equal to the temporal length of the filter (1020), but the embodiment is not limited thereto. For example, if the target interval (1010) is determined to include additional time before and / or after the temporal interval where the target sound (101) is presumed to exist, the target interval (1010) may be longer than the temporal length of the filter (1020). In this case, the electronic device (1000) can search for harmonic components (1041, 1042, 1043) by using the filter (1020) corresponding to the pattern map to move the filter (1020) within the target interval (1010) by a predetermined (e.g., predetermined) frequency interval (1030) (i.e., move along the y-axis) and at the same time move it by a predetermined (e.g., predetermined) time interval (i.e., move along the x-axis).
[0151] According to one embodiment of the present disclosure, an electronic device (1000) determines a target interval (1010) based on a temporal interval in which a target sound (101) is presumed to exist, and by searching for harmonic components (1041, 1042, 1043) of the target sound (101) only within the target interval (1010), it is possible to prevent audio signals other than the target sound (101) (e.g., voice, background sound, ambient noise, etc.) from being removed or attenuated.
[0152] FIG. 11 is a flowchart illustrating a method by which an electronic device (1000) according to one embodiment of the present disclosure determines a harmonic region within a first output audio. FIG. 12a is a diagram illustrating an operation by which an electronic device (1000) according to one embodiment of the present disclosure determines a harmonic region within a first output audio. FIG. 12b is a diagram illustrating an operation by which an electronic device (1000) according to one embodiment of the present disclosure determines a harmonic region within a first output audio.
[0153] Steps S1110 and S1120 illustrated in FIG. 11 are operations that embody the operation of step S240 of FIG. 2. In one embodiment of the present disclosure, steps S1110 and S1120 may be executed by at least one processor (1920, see FIG. 19) included in an electronic device (1000). Step S1110 of FIG. 11 may be performed after step S230 of FIG. 2 is performed. Step S250 of FIG. 2 may be performed after step S1120 of FIG. 11 is performed.
[0154] In step S1110 of FIG. 11, an electronic device (1000) according to one embodiment of the present disclosure can calculate an attention score corresponding to the ratio of the average of the first spectrogram magnitude of the area matching the pattern area and the average of the second spectrogram magnitude of the area matching the surrounding area in each of the window areas that are sequentially matched with the filter in the target interval.
[0155] Referring to FIG. 12a and FIG. 12b, in one embodiment of the present disclosure, a filter (1220) corresponding to a pattern map (1211) may include a pattern area (1221) and a surrounding area (1222). The pattern area (1221) may be an area where harmonic components exist. The pattern area (1221) may correspond to an area within the pattern map (1211) where the target sound is presumed (or determined, identified). The surrounding area (1222) may be the remaining area within the filter (1220) excluding the pattern area (1221). The surrounding area (1222) may correspond to an area within the pattern map (1211) where the target sound is presumed (or determined, identified).
[0156] In one embodiment of the present disclosure, the electronic device (1000) can calculate an attention score for each of the window regions that are sequentially matched with the filter (1220). The attention score can be calculated using the following Equation 1.
[0157]
[0158] Here, the average of the first spectrogram size refers to the average of the spectrogram size in the area (1231, 1241) (or referred to as the pattern matching area (1231, 1241)) that matches the pattern area (1221) of the filter (1220) within the window area. The average of the second spectrogram size refers to the average of the spectrogram size in the area (1232, 1242) (or referred to as the surrounding matching area (1232, 1242)) that matches the surrounding area (1222) of the filter (1220) within the window area.
[0159] FIGS. 12a and 12b illustrate, for example, one window area (1230, 1240) that matches the filter (1220).
[0160] As shown in FIG. 12a, it can be seen that the average spectrogram size in the area (1231) that matches the pattern area (1221) (i.e., the pattern matching area (1231)) is higher than the average spectrogram size in the area (1232) that matches the surrounding area (1222) (i.e., the surrounding matching area (1232)). Accordingly, the attention score in the window area (1230) shown in FIG. 12a may be relatively high.
[0161] As shown in FIG. 12b, it can be seen that the average spectrogram size in the area (1241) that matches the pattern area (1221) (i.e., the pattern matching area (1241)) and the average spectrogram size in the area (1242) that matches the surrounding area (1222) (i.e., the surrounding matching area (1242)) are similar to each other. In the window area (1240) shown in FIG. 12b, since the spectrogram size in the area (1241) that matches the pattern area as well as the spectrogram size in the area (1242) that matches the surrounding area are formed to be high, the attention score in the window area (1240) shown in FIG. 12b may be relatively low.
[0162] In step S1120 of FIG. 11, an electronic device (1000) according to one embodiment of the present disclosure may determine at least one window region among the window regions as at least one harmonic region containing a harmonic component based on the attention score of each of the window regions.
[0163] In one embodiment of the present disclosure, the electronic device (1000) may determine one or more window regions corresponding to an upper preset ratio (e.g., upper preset percentage range) as one or more harmonic regions based on the attention scores of the window regions.
[0164] Alternatively, in one embodiment of the present disclosure, the electronic device (1000) may determine one or more window regions in which the attention score is greater than or equal to a preset threshold score as one or more harmonic regions based on whether each window region is greater than or equal to a preset threshold score.
[0165] For example, the electronic device (1000) may determine the area (1231) that matches the pattern area (1221) in the window area (1230) shown in FIG. 12a as a harmonic area based on the fact that the attention score in the window area (1230) shown in FIG. 12a is relatively high. In other words, the electronic device (1000) may determine that the area (1231) that matches the pattern area (1221) in the window area shown in FIG. 12a contains harmonic components.
[0166] For example, the electronic device (1000) may determine that the window area (1240) shown in FIG. 12b does not contain a harmonic region based on the fact that the attention score in the window area (1240) shown in FIG. 12b is relatively low. In other words, the electronic device (1000) may determine that the window area (1240) shown in FIG. 12b does not contain harmonic components.
[0167] FIG. 13a is a flowchart illustrating a method for an electronic device (1000) according to one embodiment of the present disclosure to obtain a second output audio with reduced harmonic component magnitude. FIG. 13b is a diagram illustrating an operation for an electronic device (1000) according to one embodiment of the present disclosure to obtain a second output audio (1320) with reduced harmonic component magnitude.
[0168] Step S1310 illustrated in FIG. 13a is an operation that embodies the operation of step S250 of FIG. 2. In one embodiment of the present disclosure, step S1310 may be executed by at least one processor (1920, see FIG. 19) included in an electronic device (1000). Step S1310 of FIG. 13a may be performed after step S240 of FIG. 2 has been performed.
[0169] In step S1310 of FIG. 13a, an electronic device (1000) according to one embodiment of the present disclosure may change the first spectrogram magnitude of the area matching the pattern area among the target sections to the average of the second spectrogram magnitude of the area matching the surrounding area among the target sections.
[0170] FIG. 13b illustrates, by way of example, a window area (1310) determined as a harmonic region as the attention score is greater than or equal to a preset threshold score. Hereinafter, the window area (1310) of FIG. 13b will be described as a harmonic region (1310). Referring together with FIG. 13b, in one embodiment of the present disclosure, the harmonic region (1310) may include a region (1311) that matches the pattern region of the filter (or is also referred to as the pattern matching region (1311)) and a region (1312) that matches the surrounding region of the filter (or is also referred to as the surrounding matching region (1312)).
[0171] In one embodiment of the present disclosure, the electronic device (1000) may obtain information regarding a second spectrogram size, which is the average of the spectrogram sizes in the surrounding matching region (1312). The electronic device (1000) may change the spectrogram size in the pattern matching region (1311) to the second spectrogram size. That is, the electronic device (1000) may change the spectrogram size in the pattern matching region (1311) to the average of the spectrogram sizes in the surrounding matching region (1312). The electronic device (1000) may generate a second output audio (1320) by changing the magnitude of each time-frequency bin corresponding to the pattern matching region (1311) in the first output audio (e.g., harmonic region (1310)) to the second spectrogram size.
[0172] According to one embodiment of the present disclosure, an electronic device (1000) can determine a pattern matching area (1311) within a harmonic region (1310) where harmonic components caused by a target sound are presumed to exist. The electronic device (1000) can remove or attenuate harmonic components within the harmonic region (1310) by changing the spectrogram size in the pattern matching area (1311) where harmonic components are presumed to exist to the average of the spectrogram sizes in the surrounding matching area (1312). At this time, by changing the spectrogram size of the pattern matching area (1311) to the average of the spectrogram sizes in the surrounding matching area (1312), harmonic components within the harmonic region (1310) are removed or attenuated, and at the same time, the sense of dissonance between the audio components in the area where harmonic components are removed or attenuated and audio components other than the harmonic components caused by the target sound (e.g., voice, background sound, ambient noise, etc.) can be reduced.
[0173] The electronic device (1000) can output a final audio (i.e., the second output audio (1320)) in which both the target sound and the harmonic components caused by the target sound are removed or attenuated from the original audio by first generating a first output audio in which an audio component corresponding to the target sound is removed or attenuated from the first output audio. By doing so, the electronic device (1000) can provide the user with a final audio (i.e., the second output audio (1320)) in which the degree of perception of the target sound is reduced by effectively removing the target sound within the input audio. The electronic device (1000) can provide audio or video with improved quality by removing sounds that are recorded regardless of the user's intention.
[0174] FIG. 14a is a flowchart illustrating a method for an electronic device (1000) according to one embodiment of the present disclosure to obtain a second output audio with reduced harmonic component magnitude. FIG. 14b is a diagram illustrating an operation for an electronic device (1000) according to one embodiment of the present disclosure to obtain a second output audio (1420) with reduced harmonic component magnitude.
[0175] Steps S1410 and S1420 illustrated in FIG. 14a are operations that embody the operation of step S250 of FIG. 2. In one embodiment of the present disclosure, steps S1410 and S1420 may be executed by at least one processor (1920, see FIG. 19) included in an electronic device (1000). Step S1410 of FIG. 14a may be performed after step S240 of FIG. 2 has been performed.
[0176] In step S1410 of FIG. 14a, an electronic device (1000) according to one embodiment of the present disclosure can estimate the magnitude of a target sound (1402) based on residual audio (1401).
[0177] Referring together with FIG. 14b, in one embodiment of the present disclosure, an electronic device (1000) can obtain a distribution of spectrograms regarding a target sound (1402) predicted by an artificial intelligence model by removing a first output audio from an input audio to obtain residual audio (1401). The electronic device (1000) can estimate the average spectrogram size of the target sound (1402) through the residual audio (1401). For example, the electronic device (1000) can obtain (e.g., calculate) the average spectrogram size in the interval where the target sound (1402) exists within the residual audio (1401). The electronic device (1000) can estimate the average spectrogram size in the interval where the target sound (1402) exists, obtained through the residual audio (1401), as the size of the target sound (1402).
[0178] In step S1420 of FIG. 14a, an electronic device (1000) according to one embodiment of the present disclosure can reduce the first spectrogram magnitude of a region (1411) that matches a pattern region among the target intervals based on the magnitude of the target sound (1402).
[0179] FIG. 14b illustrates, by way of example, a window area (1410) determined as a harmonic region where the attention score is greater than or equal to a preset threshold score. Hereinafter, the window area (1410) of FIG. 14b will be described as a harmonic region (1410). Referring together with FIG. 14b, in one embodiment of the present disclosure, the harmonic region (1410) may include a region (1411) that matches the pattern region of the filter (or is also referred to as the pattern matching region (1411)) and a region (1412) that matches the surrounding region of the filter (or is also referred to as the surrounding matching region (1412)).
[0180] In one embodiment of the present disclosure, the electronic device (1000) may reduce the spectrogram magnitude in the pattern matching region (1411) based on the magnitude of the target sound (1402) estimated through the residual audio (1401). For example, the electronic device (1000) may reduce the spectrogram magnitude in the pattern matching region (1411) by applying a scaling factor of “1 / magnitude of target sound”. The electronic device (1000) may generate a second output audio (1420) by reducing the magnitude of each time-frequency bin corresponding to the pattern matching region (1411) in the first output audio by a factor of “1 / magnitude of target sound”.
[0181] According to one embodiment of the present disclosure, an electronic device (1000) can determine a pattern matching region (1411) within a harmonic region (1410) where harmonic components by a target sound (1402) are presumed to exist. The electronic device (1000) can reduce the magnitude of the harmonic components in proportion to the magnitude of the target sound (1402) by applying a scaling factor corresponding to the inverse of the magnitude of the target sound (1402) to the spectrogram magnitude in the pattern matching region (1411) where harmonic components are presumed to exist. For example, the magnitude of the target sound (1402) and the magnitude of the harmonic components by the target sound can generally be proportional to each other. If the magnitude of the target sound (1402) is large, the magnitude of the harmonic components by the target sound (1402) may also be large, and if the magnitude of the target sound (1402) is small, the magnitude of the harmonic components by the target sound (1402) may also be small. According to one embodiment of the present disclosure, the electronic device (1000) determines the degree of reduction of the magnitude of the harmonic component in proportion to the magnitude of the target sound (1402), thereby allowing large harmonic components to be attenuated to a relatively large degree and small harmonic components to be attenuated to a relatively small degree. Accordingly, by attenuating the harmonic component in proportion to its inherent magnitude, the electronic device (1000) can reduce the sense of dissonance between the audio component in the region where the harmonic component is removed or attenuated and the audio component other than the harmonic component caused by the target sound (e.g., voice, background sound, ambient noise, etc.).
[0182] The electronic device (1000) can first generate a first output audio in which an audio component corresponding to the target sound (1402) is removed from the original audio, and secondly generate a second output audio (1420) in which an audio component corresponding to the harmonic component of the target sound (1402) is removed or attenuated from the first output audio, thereby outputting a final audio (i.e., the second output audio (1420)) in which both the target sound (1402) and the harmonic component of the target sound are removed or attenuated from the original audio. Through this, the electronic device (1000) can provide the user with a final audio (i.e., the second output audio (1420)) in which the degree of perception of the target sound (1402) is reduced by effectively removing the target sound (1402). The electronic device (1000) can provide audio or video with improved quality by removing sounds that are recorded regardless of the user's intention.
[0183] FIG. 15 is a diagram illustrating the operation of an electronic device (1000) according to one embodiment of the present disclosure removing a target sound from an input audio (110a).
[0184] Referring to FIG. 15, in one embodiment of the present disclosure, the input audio (110a) may include a left input audio (110L) and a right input audio (110R). The input audio (110a) may be stereophonic audio. Stereophonic audio may be an audio signal that is recorded and played back using a plurality of audio channels including a 'left' channel and a 'right' channel. Stereophonic audio can reproduce the directionality and depth of sound compared to monophonic audio using a single channel, thereby providing three-dimensional sound.
[0185] In one embodiment of the present disclosure, an electronic device (1000) can acquire (e.g., record) left-side input audio (110L) through a left-side microphone of the electronic device (1000) (or external device), and can acquire (e.g., record) right-side input audio (110R) through a right-side microphone of the electronic device (1000) (or external device) positioned at a certain distance from the left-side microphone.
[0186] Alternatively, in one embodiment of the present disclosure, the electronic device (1000) may obtain stereo audio converted from mono audio by software-converting (or rendering) mono audio. For example, the electronic device (1000) may form left input audio (110L) and right input audio (110R) by artificially applying a time delay or frequency filtering based on mono audio to form a difference between left and right channels.
[0187] In one embodiment of the present disclosure, the original target sound (200a) may include a left original target sound (200L) and a right original target sound (200R). For example, the left original target sound (200L) may be obtained by acquiring the target sound through a left microphone, and the right original target sound (200R) may be obtained by acquiring the target sound through a right microphone. Alternatively, for example, the left original target sound (200L) and the right original target sound (200R) may be obtained by software-converting (or rendering) the original target sound in a single-channel (mono) format.
[0188] In one embodiment of the present disclosure, an electronic device (1000) may obtain a first output audio (120a) comprising a first left output audio (120L) and a first right output audio (120R). The electronic device (1000) may obtain the first left output audio (120L) by removing a target sound (101L) from a left input audio (110L) using an artificial intelligence model (115). The target sound (101L) included in the left input audio (110L) may be referred to as the left target sound (101L). The electronic device (1000) may obtain the first right output audio (120R) by removing a target sound (101R) from a right input audio (110R) using an artificial intelligence model (115). The target sound (101R) included in the right input audio (110R) can be referred to as the right target sound (101R).
[0189] In one embodiment of the present disclosure, an electronic device (1000) may obtain residual audio (125a) including left residual audio (125L) and right residual audio (125R). The electronic device (1000) may obtain left residual audio (125L) including left target sound (101L) by utilizing the difference between left input audio (110L) and first left output audio (120L). The electronic device (1000) may obtain right residual audio (125R) including right target sound (101R) by utilizing the difference between right input audio (110R) and first right output audio (120R).
[0190] In one embodiment of the present disclosure, the electronic device (1000) can estimate the shape of a left target sound (101L) based on a left residual audio (125L). The electronic device (1000) can estimate the shape of a right target sound (101R) based on a right residual audio (125R). For example, the electronic device (1000) can obtain a left pattern map corresponding to the shape of the left target sound (101L) and a right pattern map corresponding to the shape of the right target sound (101R).
[0191] In one embodiment of the present disclosure, the electronic device (1000) can search for harmonic components of the left target sound (101L) within the first left output audio (120L) based on the shape of the estimated left target sound (101L) (e.g., a left pattern map). The electronic device (1000) can search for harmonic components of the right target sound (101R) within the first right output audio (120R) based on the shape of the estimated right target sound (101R) (e.g., a right pattern map).
[0192] In one embodiment of the present disclosure, the electronic device (1000) may obtain a second output audio (130a) comprising a second left output audio (130L) and a second right output audio (130R). The electronic device (1000) may obtain the second left output audio (130L) by reducing the magnitude of a harmonic component found in the first left output audio (120L). The electronic device (1000) may obtain the second right output audio (130R) by reducing the magnitude of a harmonic component found in the first right output audio (120R).
[0193] According to one embodiment of the present disclosure, an electronic device (1000) can obtain a second left output audio (130L) in which a left target sound (101L) and a harmonic component (102L) caused by the left target sound (101L) are all removed or attenuated from a left input audio (110L), and can obtain a second right output audio (130R) in which a right target sound (101R) and a harmonic component (102R) caused by the right target sound (101R) are all removed or attenuated from a right input audio (110R). Accordingly, the electronic device (1000) can obtain a final output audio in a stereo format (i.e., the second output audio (130a)) in which a target sound is effectively removed or attenuated from a stereo input audio (110a).
[0194] FIG. 16a is a diagram illustrating the operation of an electronic device (1000) according to one embodiment of the present disclosure removing a target sound from an input audio (1610). FIG. 16b is a diagram illustrating the operation of an electronic device (1000) according to one embodiment of the present disclosure removing a target sound and harmonic components caused by the target sound from an input audio (1610).
[0195] Referring to FIG. 16a, in one embodiment of the present disclosure, an electronic device (1000) may include a "Single Take" shooting function. The Single Take shooting function is a function that automatically generates various types of content (e.g., photos, videos, GIFs, timelapses, etc.) using artificial intelligence technology for a set (e.g., predetermined) shooting duration by having the user perform a shooting start input only once. For example, when the electronic device (1000) receives the user's shooting start input, it may continuously record a video for a set (e.g., predetermined) time (e.g., a 3 to 15 second interval after receiving the input). The electronic device (1000) may generate various types of content based on data included in the recorded video. For example, the electronic device (1000) may provide a best shot by selecting a photo determined to be the best moment. For example, the electronic device (1000) may provide a short video focusing on dynamic moments or key scenes. For example, the electronic device (1000) can convert a short video of a specific action being repeated into a GIF file. For example, the electronic device (1000) can provide a time-lapse by compressing a specific scene. Through this, the electronic device (1000) can provide multiple optimized results with a single shooting action of the user, without the user needing to select a shooting mode or adjust the timing in advance.
[0196] In one embodiment of the present disclosure, an electronic device (1000) may acquire input audio (1610) including a video recording start sound (1601) and a video recording end sound (1602). For example, after receiving a user's input to start a single take recording, the electronic device (1000) may start video recording at a time when a preset time has elapsed, and the video recording start sound (1601) may be played at the time when recording begins. In this case, the video recording start sound (1601) may be recorded in the original video (1605) of the single take recording. For example, after receiving a user's input to start a single take recording, the electronic device (1000) may end video recording at a time when a preset time has elapsed, and the video recording end sound (1602) may be played at the time when recording ends. In this case, the video recording end sound (1602) may be recorded in the original video (1605) of the single take recording.
[0197] Referring to FIG. 16b, in one embodiment of the present disclosure, the input audio (1610) may include a first target sound (1601) corresponding to a video recording start sound and a second target sound (1602) corresponding to a video recording end sound.
[0198] In one embodiment of the present disclosure, an electronic device (1000) can obtain a first output audio (1620) by removing a first target sound (1601) from an input audio (1610) using an artificial intelligence model (115). For example, the electronic device (1000) can input an original target sound (hereinafter referred to as the first original target sound) corresponding to the input audio (1610) and the first target sound (1601) into the artificial intelligence model (115). The electronic device (1000) can obtain a first output audio (1620) in which the first target sound (1601) is removed from the input audio (1610) based on the input audio (1610) and the first original target sound through the artificial intelligence model (115).
[0199] In one embodiment of the present disclosure, an electronic device (1000) can obtain a second output audio (1620) by removing a second target sound (1602) from an input audio (1610) using an artificial intelligence model (115). For example, the electronic device (1000) can input an original target sound (hereinafter referred to as the second original target sound) corresponding to the input audio (1610) and the second target sound (1602) into the artificial intelligence model (115). The electronic device (1000) can obtain a first output audio (1620) in which the second target sound (1602) is removed from the input audio (1610) based on the input audio (1610) and the second original target sound through the artificial intelligence model (115).
[0200] In one embodiment of the present disclosure, an electronic device (1000) can obtain a second output audio (1630) in which the first harmonic component (1603) (hereinafter referred to as the first harmonic component (1603)) caused by the first target sound (1601) is removed or attenuated by reducing the magnitude of the harmonic component (1603) caused by the first target sound (1601) in the first output audio (1620). For example, the electronic device (1000) can obtain residual audio containing the first target sound (1601) by using the difference between the first output audio (1620) and the input audio (1610). Based on the residual audio, the electronic device (1000) can estimate the shape of the first target sound (1601). The electronic device (1000) can search for a first harmonic component (1603) in the first output audio (1620) based on the shape of the first target sound (1601), and can reduce the magnitude of the searched first harmonic component (1603).
[0201] In one embodiment of the present disclosure, an electronic device (1000) can obtain a second output audio (1630) in which the second harmonic component (1604) (hereinafter referred to as the second harmonic component (1604)) caused by the second target sound (1602) is removed or attenuated by reducing the magnitude of the harmonic component (1604) caused by the second target sound (1602) in the first output audio (1620). For example, the electronic device (1000) can obtain residual audio containing the second target sound (1602) by using the difference between the first output audio (1620) and the input audio (1610). Based on the residual audio, the electronic device (1000) can estimate the shape of the second target sound (1602). The electronic device (1000) can search for a second harmonic component (1604) in the first output audio (1620) based on the shape of the second target sound (1602), and can reduce the magnitude of the searched second harmonic component (1604).
[0202] Through this, the electronic device (1000) can obtain a final output audio (i.e., a second output audio (1630)) in which the video recording start sound (corresponding to the first target sound (1601)) and the harmonic component (1603) caused by the video recording start sound in the input audio (1610) are removed or attenuated. The electronic device (1000) can obtain a final output audio (i.e., a second output audio (1630)) in which the video recording end sound (corresponding to the second target sound (1602)) and the harmonic component (1604) caused by the video recording end sound in the input audio (1610) are removed or attenuated. Accordingly, even if the video recording start / end sound is recorded in the input audio (1610), the electronic device (1000) can estimate the form of the actual video recording start / end sound within the input audio (1610) and more accurately search for and reduce the harmonic components (1603, 1604) caused by the video recording start / end sound based on the form of the actual video recording start / end sound. By effectively removing the video recording start / end sound within the input audio (1610), the electronic device (1000) can provide the user with a final audio (i.e., the second output audio (1630)) in which the degree of perception of the video recording start / end sound is reduced.
[0203] FIG. 17a is a diagram illustrating the operation of an electronic device (1000) according to one embodiment of the present disclosure removing a target sound from input audio. FIG. 17b is a diagram illustrating the operation of an electronic device (1000) according to one embodiment of the present disclosure removing a target sound and harmonic components caused by the target sound from input audio.
[0204] Referring to FIG. 17a, in one embodiment of the present disclosure, the electronic device (1000) may include a "motion photo" shooting function. The motion photo shooting function is a function that records a short video before and after the photo at the time of taking the photo. For example, when the electronic device (1000) receives user input pressing the shutter button, it can automatically record a short video (e.g., a video of 2 to 3 seconds) from just before pressing the shutter button until immediately after pressing the shutter button. By automatically recording the short video, the electronic device (1000) can provide a video that captures the movement or dynamism of the subject, which cannot be conveyed by a still photograph. The electronic device (1000) can select a good moment from the automatically recorded short video even if the moment the user presses the shutter button is not perfect, thereby reducing the shooting failure rate.
[0205] In one embodiment of the present disclosure, an electronic device (1000) may acquire input audio (1710) including a photo shutter sound (1701). For example, after receiving a user's motion photo shooting start input, the electronic device (1000) may start video recording at a time earlier than when the motion photo shooting start input is received and end video recording at a time later than when the motion photo shooting start input is received. Accordingly, in the video (1705) automatically recorded during motion photo shooting, the photo shutter sound (1701) played at the time the user presses the shutter button for motion photo shooting may be recorded.
[0206] Referring to FIG. 17b, in one embodiment of the present disclosure, the input audio (1710) may include a target sound corresponding to a photo taking sound (1701).
[0207] In one embodiment of the present disclosure, an electronic device (1000) can obtain a first output audio (1720) by removing a photo shutter sound (1701) from an input audio (1710) using an artificial intelligence model (115). For example, the electronic device (1000) can input the input audio (1710) and the original photo shutter sound (1701) corresponding to the photo shutter sound into the artificial intelligence model (115). The electronic device (1000) can obtain a first output audio (1720) in which the photo shutter sound (1701) is removed from the input audio (1710) based on the input audio (1710) and the original photo shutter sound (1701) through the artificial intelligence model (115).
[0208] In one embodiment of the present disclosure, the electronic device (1000) can obtain a second output audio (1730) in which the harmonic component (1702) caused by the photo-taking sound (1701) is removed or attenuated by reducing the magnitude of the harmonic component (1702) caused by the photo-taking sound (1701) in the first output audio (1720). For example, the electronic device (1000) can obtain residual audio containing the photo-taking sound (1701) by using the difference between the first output audio (1720) and the input audio (1710). Based on the residual audio, the electronic device (1000) can estimate the shape of the photo-taking sound (1701) within the input audio (1710). The electronic device (1000) can search for a harmonic component (1702) of the photographic sound (1701) in the first output audio (1720) based on the shape of the photographic sound (1701), and can reduce the magnitude of the searched harmonic component (1702).
[0209] Through this, the electronic device (1000) can obtain a final output audio (i.e., a second output audio (1730)) in which the photo-taking sound (1701) and the harmonic components (1702) caused by the photo-taking sound (1701) within the input audio (1710) are removed or attenuated. Accordingly, even if the photo-taking sound (1701) is recorded in the input audio (1710), the electronic device (1000) can estimate the shape of the actual photo-taking sound (1701) within the input audio (1710) and more accurately search for and reduce the harmonic components (1702) caused by the photo-taking sound (1701) based on the shape of the actual photo-taking sound (1701). The electronic device (1000) can provide the user with a final audio (i.e., a second output audio (1730)) in which the degree of perception of the photo-taking sound (1701) is reduced by effectively removing the photo-taking sound (1701) within the input audio (1710).
[0210] FIG. 18 is a diagram illustrating the operation of an electronic device (1000) according to one embodiment of the present disclosure removing a target sound from input audio (1810).
[0211] Referring to FIG. 18, in one embodiment of the present disclosure, an electronic device (1000) may store input audio (1810) including one or more target sounds in a memory (1930). The electronic device (1000) may load (or acquire) input audio (1810) including one or more target sounds (1801, 1802) stored in the memory (1930) and process the input audio (1810).
[0212] In one embodiment of the present disclosure, an electronic device (1000) may receive user input requesting the removal of target tones (1801, 1802) from a previously stored input audio (1810). For example, the electronic device (1000) may select an input audio (1810) stored within the electronic device (1000) and receive user input regarding a request to remove target tones (1801, 1802) from the selected input audio (1810). Based on receiving user input regarding a request to remove target tones (1801, 1802) from the input audio (1810), the electronic device (1000) may remove or attenuate the target tones (1801, 1802) and harmonic components (1803, 1804) caused by the target tones (1801, 1802) from the selected input audio (1810). The electronic device (1000) can generate a final output audio (e.g., a second output audio (1830)) in which the target sound (1801, 1802) and the harmonic components (1803, 1804) of the target sound (1801, 1802) are removed or attenuated from the selected input audio (1810), and the generated final output audio can be stored back in memory (1930).
[0213] Alternatively, in one embodiment of the present disclosure, the electronic device (1000) may receive a user input requesting a final captured image (1806) (or final recorded image) from which target sounds (1801, 1802) have been removed from the original captured image (1805) during image capture (or voice recording). The electronic device (1000) may temporarily store the captured (or recorded) input audio (1810) (or original audio) in memory (1930) when image capture ends. The electronic device (1000) can load the input audio (1810) (or original audio) that has been temporarily stored and remove or attenuate the target tones (1801, 1802) and the harmonic components (1803, 1804) caused by the target tones (1801, 1802) within the input audio (1810) (or original audio) to generate the final output audio (i.e., the second output audio (1830)). Afterward, the electronic device (1000) can overwrite the memory (1930) location where the input audio (1810) (or original audio) is stored with the final output audio (i.e., the second output audio (1830)). However, the embodiments are not limited thereto, and the electronic device (1000) may form a final captured image (1806) in which the target sound (1801, 1802) and the harmonic components (1803, 1804) caused by the target sound (1801, 1802) are removed or attenuated in real time while capturing the image (or recording the voice). In this case, when the image capturing ends, the electronic device (1000) may store the final captured image (1806) (or the second output audio (1830)) in which the target sound (1801, 1802) and the harmonic components (1803, 1804) caused by the target sound (1801, 1802) are removed or attenuated in the memory (1930). The method for removing or attenuating the target sound (1801, 1802) and the harmonic components (1803, 1804) caused by the target sound (1801, 1802) from the original audio (1810) has been described in detail above, so it will be omitted below.
[0214] FIG. 19 is a block diagram schematically illustrating the configuration of an electronic device according to one embodiment of the present disclosure.
[0215] Referring to FIG. 19, an electronic device (1000) according to one embodiment of the present disclosure may include an input / output interface (1910), a processor (1920), and a memory (1930).
[0216] The input / output interface (1910) may include an input interface (e.g., touch screen, keyboard, microphone, etc.) for receiving commands or information from a user, and an output interface (e.g., display panel, speaker, etc.) for displaying the result of an operation according to the user's command or the status of the electronic device (1000). According to one embodiment of the present disclosure, the electronic device (1000) may receive input (e.g., a request for sound source separation) from a user through the input / output interface (1910), and when the operation is completed, may output the result of the operation (e.g., a result of sound source separation) through the input / output interface (1910).
[0217] A processor (1920) is a component that controls a series of processes to enable an electronic device (1000) to operate according to the embodiments described in this disclosure, and may be composed of one or more processors. One or more processors included in the processor (1920) may be circuitry such as a System on Chip (SoC) or an Integrated Circuit (IC). One or more processors included in the processor (1920) may be a general-purpose processor such as a Central Processing Unit (CPU), a Micro Processor Unit (MPU), an Application Processor (AP), or a Digital Signal Processor (DSP); a graphics-dedicated processor such as a Graphic Processing Unit (GPU) or a Vision Processing Unit (VPU); an artificial intelligence-dedicated processor such as a Neural Processing Unit (NPU); or a communication-dedicated processor such as a Communication Processor (CP). If one or more processors included in the processor (1920) are artificial intelligence-dedicated processors, said artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0218] The processor (1920) can write data to memory (1930) or read data stored in memory (1930), and in particular, can process data according to a predefined operation rule or artificial intelligence model by executing a program or at least one instruction stored in memory (1930). Accordingly, the processor (1920) can perform the operations described in the embodiments of the present disclosure, and the operations described in the present disclosure as being performed by modules included in the electronic device (1000) can be seen as being performed by the processor (1920) unless otherwise specified.
[0219] Memory (1930) is a configuration for storing various programs or data and may be composed of a storage medium or a combination of storage media such as ROM, RAM, hard disk, CD-ROM, and DVD. Memory (1930) may not exist separately but may be configured to be included in the processor (1920). Memory (1930) may be composed of volatile memory, non-volatile memory, or a combination of volatile and non-volatile memory. Memory (1930) may store a program or at least one instruction for performing operations according to the embodiments described below. Memory (1930) may provide stored data to the processor (1920) upon the request of the processor (1920).
[0220] The embodiments described above with reference to FIGS. 1 to 18 can be performed by an electronic device (1000).
[0221] In order to solve the above-described technical problem, in one embodiment of the present disclosure, a method is provided for removing a target sound (101) from an input audio (110) which is performed by an electronic device (1000).
[0222] In one embodiment of the present disclosure, the method may include the step (S210) of obtaining a first output audio (120) by removing a target sound (101) from an input audio (110) using an artificial intelligence model (115).
[0223] In one embodiment of the present disclosure, the method may include the step (S220) of obtaining residual audio (125) containing a target sound (101) based on the difference between input audio (110) and a first output audio (120).
[0224] In one embodiment of the present disclosure, the method may include the step (S230) of estimating the shape of the target sound (101) based on residual audio (125).
[0225] In one embodiment of the present disclosure, the method may include the step (S240) of determining a harmonic component (102) of the target sound (101) in the first output audio (120) based on the shape of the estimated target sound (101).
[0226] In one embodiment of the present disclosure, the method may include the step (S250) of obtaining a second output audio (130) by reducing the magnitude of a harmonic component (102) in a first output audio (120).
[0227] In one embodiment of the present disclosure, the step (S230) of estimating the shape of the target sound (101) may include the step (S710) of obtaining a pattern map (300) corresponding to the shape of the target sound (101).
[0228] In one embodiment of the present disclosure, the method may be characterized in that the pattern map (300) is a map representing the distribution of frequency components over time associated with a target sound in the time-frequency domain.
[0229] In one embodiment of the present disclosure, the step (S240) of determining the harmonic component (102) may include the step (S910) of estimating the target interval where the target sound (101) exists in the spectrogram of the input audio (110) based on the residual audio (125).
[0230] In one embodiment of the present disclosure, the step (S240) of searching for a harmonic component (102) may include the step (S920) of searching for the harmonic component (102) by moving a filter within a target interval by a predetermined (e.g., predetermined) frequency interval. The filter may correspond to a pattern map (300).
[0231] In one embodiment of the present disclosure, the method may be characterized in that the filter includes a pattern region in which a harmonic component (102) exists and a surrounding region other than the pattern region.
[0232] In one embodiment of the present disclosure, the step (S240) of determining the harmonic component (102) may include the step (S1110) of calculating an attention score based on the ratio of the average of the first spectrogram magnitude of the window area matching the pattern area and the average of the second spectrogram magnitude of the window area matching the surrounding area in each window area that is matched with the filter (e.g., sequentially matched) among the target intervals.
[0233] In one embodiment of the present disclosure, the step (S240) of determining the harmonic component (102) may include the step (S1120) of determining at least one window region among the window regions as at least one harmonic region containing the harmonic component (102) based on the attention score of each window region.
[0234] In one embodiment of the present disclosure, the step of determining at least one window region as at least one harmonic region (S1120) may include determining at least one window region among the window regions as at least one harmonic region where the attention score corresponds to a higher preset ratio.
[0235] In one embodiment of the present disclosure, the step (S250) of obtaining (e.g., generating) a second output audio (130) by reducing the magnitude of a harmonic component (102) may include the step (S1310) of changing a first spectrogram magnitude within at least one harmonic region to the average of a second spectrogram magnitude.
[0236] In one embodiment of the present disclosure, the step (S250) of obtaining (e.g., generating) a second output audio (130) by reducing the magnitude of a harmonic component (102) may include the step (S1410) of estimating the magnitude of a target sound (101) based on residual audio (125).
[0237] In one embodiment of the present disclosure, the step (S250) of generating a second output audio (130) by reducing the magnitude of a harmonic component (102) may include the step (S1420) of reducing a first spectrogram magnitude within at least one harmonic region based on the magnitude of a target sound (101).
[0238] In one embodiment of the present disclosure, the step (S210) of obtaining a first output audio (120) may include the step (S410) of obtaining a first output audio (120) in which the target sound (101) is removed from the input audio (110) by applying an original target sound (200) corresponding to the input audio (110) and the target sound (101) to an artificial intelligence model (115).
[0239] In one embodiment of the present disclosure, the method may be characterized in that the artificial intelligence model (115) includes a lightweight model stored in an electronic device (1000).
[0240] In one embodiment of the present disclosure, the method may be characterized in that the target sound (101) includes a system notification sound comprising at least one of a video recording start sound, a video recording end sound, or a photo recording sound.
[0241] In one embodiment of the present disclosure, the method may include an input audio (110), a left input audio (110L) and a right input audio (110R).
[0242] In order to solve the above-described technical problem, an electronic device (1000) is provided in one embodiment of the present disclosure. In one embodiment of the present disclosure, an electronic device (1000) for removing a target sound (101) from an input audio (110) may be provided.
[0243] In one embodiment of the present disclosure, the electronic device (1000) may include a memory (1930) for storing instructions; and at least one processor (1920) operably coupled to the memory (1930) and comprising processing circuitry.
[0244] In one embodiment of the present disclosure, an electronic device (1000) can obtain a first output audio (120) by removing a target sound (101) from an input audio (110) using an artificial intelligence model (115) by having at least one processor (1920) execute instructions individually or collectively.
[0245] In one embodiment of the present disclosure, an electronic device (1000) can obtain residual audio (125) containing a target sound (101) based on the difference between an input audio (110) and a first output audio (120) by having at least one processor (1920) execute instructions individually or collectively.
[0246] In one embodiment of the present disclosure, the electronic device (1000) can estimate the shape of the target sound (101) based on residual audio (125) by having at least one processor (1920) execute instructions individually or collectively.
[0247] In one embodiment of the present disclosure, the electronic device (1000) can determine the harmonic components (102) of the target sound (101) in the first output audio (120) based on the shape of the estimated target sound (101) by having at least one processor (1920) execute instructions individually or collectively.
[0248] In one embodiment of the present disclosure, an electronic device (1000) can obtain a second output audio (130) by reducing the magnitude of a harmonic component (102) in a first output audio (120) by having at least one processor (1920) execute instructions individually or collectively.
[0249] In one embodiment of the present disclosure, the electronic device (1000) can obtain a pattern map (300) corresponding to the shape of the target sound (101) by having at least one processor (1920) execute instructions individually or collectively.
[0250] In an electronic device (1000) according to one embodiment of the present disclosure, the pattern map (300) may be characterized as a map showing the distribution of frequency components over time associated with a target sound (101) in the time-frequency domain.
[0251] In one embodiment of the present disclosure, the electronic device (1000) can estimate a target interval in which a target sound (101) exists in a spectogram of an input audio (110) based on residual audio (125) by having at least one processor (1920) execute instructions individually or collectively.
[0252] In one embodiment of the present disclosure, an electronic device (1000) can search for harmonic components (102) by moving a filter within a target interval by a predetermined frequency interval, by having at least one processor (1920) execute instructions individually or collectively. The filter may correspond to a pattern map (300).
[0253] In an electronic device (1000) according to one embodiment of the present disclosure, the filter may be characterized by including a pattern region in which a harmonic component (102) exists and a surrounding region other than the pattern region.
[0254] In one embodiment of the present disclosure, an electronic device (1000) can calculate an attention score based on the ratio of the average of the first spectrogram magnitude of the window area matching the pattern area and the average of the second spectrogram magnitude of the window area matching the surrounding area in each window area that matches the filter (e.g., sequentially) among the target intervals by having at least one processor (1920) execute instructions individually or collectively.
[0255] In one embodiment of the present disclosure, the electronic device (1000) can determine at least one window region among the window regions as at least one harmonic region containing a harmonic component (102) based on the attention score of each window region by having at least one processor (1920) execute instructions individually or collectively.
[0256] In one embodiment of the present disclosure, the electronic device (1000) can change a first spectrogram magnitude in at least one harmonic region to the average of a second spectrogram magnitude in a target region by having at least one processor (1920) execute instructions individually or collectively.
[0257] In one embodiment of the present disclosure, the electronic device (1000) can estimate the magnitude of a target sound (101) based on residual audio (125) by having at least one processor (1920) execute instructions individually or collectively.
[0258] In one embodiment of the present disclosure, the electronic device (1000) can reduce a first spectrogram magnitude in at least one harmonic region based on the magnitude of the target sound (101) by having at least one processor (1920) execute instructions individually or collectively.
[0259] In one embodiment of the present disclosure, an electronic device (1000) can obtain a first output audio (120) by applying an original target sound (200) corresponding to an input audio (110) and a target sound (101) to an artificial intelligence model (115) by having at least one processor (1920) execute instructions individually or collectively.
[0260] In order to solve the above-described technical problem, in one embodiment of the present disclosure, at least one computer-readable recording medium is provided on which a program is recorded for implementing a method for removing a target sound (101) from an input audio (110) using the electronic device (1000) described above.
[0261] A device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory storage medium' simply means that it is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily. For example, a 'non-transitory storage medium' may include a buffer in which data is stored temporarily.
[0262] According to one embodiment of the present disclosure, the method according to the various embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0263] According to one embodiment of the present disclosure, a non-transient computer-readable medium may be provided for storing instructions that cause the electronic device to obtain a first output audio (120) by removing a target sound (101) from an input audio (110) using an artificial intelligence model (115) when executed by at least one processor of the electronic device. According to one embodiment of the present disclosure, a non-transient computer-readable medium may be provided for storing instructions that cause the electronic device to obtain a residual audio (125) containing the target sound (101) based on the difference between the input audio (110) and the first output audio (120) when executed by at least one processor of the electronic device. According to one embodiment of the present disclosure, a non-transient computer-readable medium may be provided for storing instructions that cause the electronic device to estimate the shape of the target sound (101) based on the residual audio (125) when executed by at least one processor of the electronic device. According to one embodiment of the present disclosure, a non-transient computer-readable medium may be provided for storing instructions that cause the electronic device to determine a harmonic component (102) of the target sound (101) in the first output audio (120) based on the shape of the estimated target sound (101) when executed by at least one processor of the electronic device. According to one embodiment of the present disclosure, a non-transient computer-readable medium may be provided for storing instructions that cause the electronic device to obtain the first output audio (120) by removing the target sound (101) from the input audio (110) using an artificial intelligence model (115) when executed by at least one processor of the electronic device.According to one embodiment of the present disclosure, a non-transient computer-readable medium may be provided for storing instructions that, when executed by at least one processor of an electronic device, cause the electronic device to obtain a second output audio (130) by reducing the magnitude of a harmonic component (102) in a first output audio (120).
Claims
1. A method for removing a target sound (101) from an input audio (110), which is performed by an electronic device (1000), A step (S210) of obtaining a first output audio (120) by removing the target sound (101) from the input audio (110) using an artificial intelligence model (115); A step (S220) of obtaining residual audio (125) containing the target sound (101) based on the difference between the input audio (110) and the first output audio (120); Step (S230) of estimating the shape of the target sound (101) based on the above residual audio (125); A step (S240) of determining a harmonic component (102) by the target sound (101) within the first output audio (120) based on the shape of the estimated target sound (101); and A method comprising the step (S250) of obtaining a second output audio (130) by reducing the magnitude of the harmonic component (102) in the first output audio (120).
2. In Paragraph 1, The step (S230) of estimating the shape of the target sound (101) above is, A method comprising the step (S710) of obtaining a pattern map (300) corresponding to the shape of the target sound (101).
3. In Paragraph 2, The above pattern map (300) is a map showing the distribution of frequency components over time associated with the target sound (101) in the time-frequency domain.
4. In any one of paragraphs 2 to 3, The step (S240) of determining the above harmonic component (102) is, A step (S910) of estimating a target interval in which the target sound (101) exists in the spectrogram of the input audio (110) based on the above residual audio (125); and A method comprising the step (S920) of searching for the harmonic component (102) by moving the filter within the target section by a predetermined frequency interval, wherein the filter corresponds to the pattern map (300).
5. In Paragraph 4, The filter above includes a pattern region where the harmonic component (102) exists and a surrounding region other than the pattern region, The step (S240) of determining the above harmonic component (102) is, A step (S1110) of calculating an attention score based on the ratio of the average of the first spectrogram magnitude of the window area matching the pattern area and the average of the second spectrogram magnitude of the window area matching the surrounding area in each window area matching the filter among the target intervals; and A method further comprising the step (S1120) of determining at least one window region among the window regions as at least one harmonic region containing the harmonic component (102) based on the attention score of each of the window regions.
6. In Paragraph 5, The step (S1120) of determining the above at least one window region as the above at least one harmonic region is, A method comprising the step of determining at least one window region among the above window regions as the at least one harmonic region, wherein the attention score corresponds to a higher preset ratio.
7. In either Paragraph 5 or Paragraph 6, The step (S250) of obtaining the second output audio (130) by reducing the magnitude of the harmonic component (102) is, A method comprising the step (S1310) of changing the first spectrogram magnitude within the above at least one harmonic region to the average of the second spectrogram magnitude.
8. In either of Paragraphs 5 and 6, The step (S250) of obtaining the second output audio (130) by reducing the magnitude of the harmonic component (102) is, A step (S1410) of estimating the magnitude of the target sound (101) based on the above residual audio (125); A method comprising the step (S1420) of reducing the first spectrogram magnitude within the at least one harmonic region based on the magnitude of the target sound (101).
9. An electronic device (1000) for removing a target sound (101) from an input audio (110), Memory (1930) for storing instructions; and It includes at least one processor (1920) operably coupled to the memory (1930) and comprising a processing circuitry; and By having at least one processor (1920) execute the instructions individually or collectively, the electronic device (1000) By using an artificial intelligence model (115) to remove the target sound (101) from the input audio (110), a first output audio (120) is obtained, and Based on the difference between the input audio (110) and the first output audio (120), a residual audio (125) containing the target sound (101) is obtained, and Estimating the shape of the target sound (101) based on the above residual audio (125), and Based on the shape of the estimated target sound (101) above, a harmonic component (102) by the target sound (101) is determined within the first output audio (120), and An electronic device (1000) that obtains a second output audio (130) by reducing the magnitude of the harmonic component (102) in the first output audio (120).
10. In Paragraph 9, By having at least one processor (1920) execute the instructions individually or collectively, the electronic device (1000) further, A pattern map (300) corresponding to the shape of the target sound (101) is obtained, and The above pattern map (300) is an electronic device (1000) that is a map showing the distribution of frequency components over time associated with the target sound (101) in the time-frequency domain.
11. In Paragraph 10, By having at least one processor (1920) execute the instructions further, either alone or in cooperation, the electronic device (1000) is, Based on the above residual audio (125), the target interval where the target sound (101) exists in the spectrogram of the input audio (110) is estimated, and The electronic device (1000) searches for the harmonic component (102) by moving the filter within the above target section by a predetermined frequency interval, wherein the filter corresponds to the pattern map (300).
12. In Paragraph 11, The filter above includes a pattern region where the harmonic component (102) exists and a surrounding region other than the pattern region, By having at least one processor (1920) execute the instructions further, either alone or in cooperation, the electronic device (1000) is, In each window region among the target intervals that matches the filter, an attention score is calculated based on the ratio of the average of the first spectrogram magnitude of the window region that matches the pattern region and the average of the second spectrogram magnitude of the window region that matches the surrounding region. An electronic device (1000) that determines at least one window region among the window regions as at least one harmonic region containing the harmonic component (102) based on the attention score of each of the above window regions.
13. In Paragraph 12, By having at least one processor (1920) execute the instructions further, either alone or in cooperation, the electronic device (1000) is, An electronic device (1000) that changes the first spectrogram magnitude within the above at least one harmonic region to the average of the second spectrogram magnitude in the above target region.
14. In Paragraph 12, By having at least one processor (1920) execute the instructions further, either alone or in cooperation, the electronic device (1000) is, Based on the above residual audio (125), the magnitude of the target sound (101) is estimated, and An electronic device (1000) that reduces the first spectrogram magnitude within the at least one harmonic region based on the magnitude of the target sound (101).
15. A computer-readable recording medium having at least one program recorded thereon for implementing the method described in any one of claims 1 through 8.
Citation Information
Patent Citations
Polishing apparatus for substrate
KR1020250139116A
Voice Signal Enhancing Method and Device
US20200342892A1
Method and apparatus for speech source separation based on a convolutional neural network
US20220223144A1
Method and apparatus combining separation and classification of audio signals
US20230215423A1
Deep learning based voice extraction and primary-ambience decomposition for stereo to surround upmixing with dialog-enhanced center channel
US20240267701A1