Electronic device, operating method thereof, and storage medium

The use of AI models in electronic devices for audio signal processing addresses the challenge of enhancing user satisfaction by effectively separating and combining audio signals, improving sound quality and reducing loss.

WO2025249984A1PCT designated stage Publication Date: 2025-12-04SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/095153
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-11
Filing Date
2025-04-01
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing electronic devices face challenges in efficiently separating and combining audio signals to enhance user satisfaction due to the increasing complexity of sound processing technologies.

Method used

The electronic device employs artificial intelligence models, such as deep neural networks, to identify and process audio sources based on setting values, applying filters and weights to generate improved audio signals by separating and mixing audio components effectively.

Benefits of technology

This approach enhances sound quality and reduces sound loss by accurately identifying and processing audio sources, resulting in improved audio output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025095153_04122025_PF_FP_ABST
    Figure KR2025095153_04122025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is an electronic device. The electronic device comprises a memory for storing instructions and at least one processor, wherein the instructions, when executed by the at least one processor, cause the electronic device to: identify a first audio signal including a plurality of audio sources; identify, from the first audio signal, a first audio source corresponding to at least one type; signal-process the first audio source on the basis of a set value corresponding to each of the at least one type; generate a second audio signal on the basis of the signal-processed first audio source and at least a portion of remaining signals other than the first audio source in the first audio signal; and output the second audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic devices and their operating methods and storage media

[0001] Embodiments of the present disclosure relate to an electronic device for processing sound, a method of operating the same, and a storage medium.

[0002] The variety of services and additional features offered through electronic devices, such as smartphones, is steadily increasing. To enhance the utility of these devices and satisfy the diverse needs of users, telecommunications service providers and electronic device manufacturers are competitively developing electronic devices that offer diverse features and differentiate themselves from competitors. Consequently, the various functions offered through electronic devices are also becoming increasingly sophisticated. With the advancement of sound processing technology (sound separation and mixing technology), there is a growing demand for methods that efficiently separate and combine sounds to enhance user satisfaction.

[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.

[0004] An electronic device according to one embodiment of the present disclosure may include a memory storing instructions and at least one processor. According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to identify a first audio signal including a plurality of audio sources, and to identify a first audio source corresponding to at least one type from the first audio signal.

[0005] In one embodiment, the instructions may cause the electronic device to signal process the first audio source based on a setting value corresponding to each of the at least one types.

[0006] In one embodiment, the instructions may cause the electronic device to generate a second audio signal based on at least a portion of a signal of the first audio signal other than the first audio source and the processed first audio source, and to output the second audio signal.

[0007] A method of operating an electronic device according to one embodiment of the present disclosure may include an operation of identifying a first audio signal including a plurality of audio sources, and identifying a first audio source corresponding to at least one type from the first audio signal.

[0008] According to one embodiment, the operating method may include an operation of signal processing the first audio source based on a setting value corresponding to each of the at least one type.

[0009] According to one embodiment, the operating method may include generating a second audio signal based on at least a portion of a signal remaining from the first audio signal excluding the first audio source and the signal-processed first audio source, and outputting the second audio signal.

[0010] A storage medium storing computer-readable instructions according to one embodiment of the present disclosure, wherein the instructions, when executed by at least one processor of an electronic device, cause the electronic device to identify a first audio signal including a plurality of audio sources, and to identify a first audio source corresponding to at least one type from the first audio signal.

[0011] In one embodiment, the instructions may cause the electronic device to signal process the first audio source based on a setting value corresponding to each of the at least one types.

[0012] In one embodiment, the instructions may cause the electronic device to generate a second audio signal based on at least a portion of a signal of the first audio signal other than the first audio source and the processed first audio source, and to output the second audio signal.

[0013] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0014] FIG. 1 is a block diagram of an electronic device within a network environment, according to one embodiment.

[0015] FIG. 2 is a block diagram of configurations of an electronic device according to one embodiment.

[0016] FIG. 3 is a flowchart illustrating an operating method of an electronic device according to one embodiment.

[0017] FIG. 4 is a flowchart illustrating a method for determining a type corresponding to an audio source according to one embodiment.

[0018] FIG. 5 is a flowchart illustrating a method for providing a second audio signal according to one embodiment.

[0019] FIG. 6 is a flowchart illustrating a method for performing signal processing on a first audio signal according to one embodiment.

[0020] FIG. 7 is a flowchart illustrating a method for verifying the remaining signals according to one embodiment.

[0021] Figure 8 is a diagram for explaining a second artificial intelligence model according to one embodiment.

[0022] FIG. 9 is a drawing for explaining a method of providing a user interface (UI) according to one embodiment.

[0023] Figure 10 is a flowchart illustrating a method for checking weights according to one embodiment.

[0024] Figure 11 is a flowchart illustrating a method for checking weights according to one embodiment.

[0025] Figure 12 is a flowchart illustrating a method for checking weights according to one embodiment.

[0026] FIG. 13a is a diagram for explaining a first artificial intelligence model and a third artificial intelligence model according to one embodiment.

[0027] FIG. 13b is a diagram for explaining a method for confirming a second audio signal according to one embodiment.

[0028] Fig. 14 is a diagram for explaining an implementation example of a first artificial intelligence model according to one embodiment.

[0029] Fig. 15 is a drawing for explaining an implementation example of a first artificial intelligence model according to one embodiment.

[0030] Fig. 16 is a diagram for explaining an implementation example of a third artificial intelligence model according to one embodiment.

[0031] Fig. 17 is a diagram for explaining a fifth artificial intelligence model according to one embodiment.

[0032] FIGS. 18a, 18b, and 18c are drawings for explaining a filter according to one embodiment.

[0033] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in many different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0034] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to one embodiment. Referring to FIG. 1 , in the network environment (100), the electronic device (101) may communicate with the electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (104) or the server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0035] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0036] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0037] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0038] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0039] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0040] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0041] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0042] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0043] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0044] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0045] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0046] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0047] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0048] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0049] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0050] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0051] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0052] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0053] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0054] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0055] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0056] In the detailed description below, reference numerals in the drawings may be used interchangeably or omitted for components that can be easily understood through the preceding embodiments, and their detailed descriptions may also be omitted. An electronic device according to an embodiment disclosed in this document may be implemented by selectively combining components of different embodiments, and components of one embodiment may be replaced by components of another embodiment. For example, it should be noted that the present invention is not limited to specific drawings or embodiments.

[0057] FIG. 2 is a block diagram of electronic device configurations according to one embodiment.

[0058] According to FIG. 2, according to one embodiment, an electronic device (200, e.g., electronic device (101) of FIG. 1) may include a memory (210, e.g., memory (130) of FIG. 1) for storing instructions and at least one processor (220, or processor (120) of FIG. 1)).

[0059] According to one embodiment, the memory (210) may have at least a portion of the same or similar configuration as the memory (130) of FIG. 1. For example, the memory (210) may be configured to temporarily or permanently store digital data and may include at least a portion of the configuration and / or functions of the memory (130) of FIG. 1.

[0060] The memory (210) according to one embodiment can store various instructions that can be executed by at least one processor (220). In addition, the memory (210) can store at least a portion of the program (140) of FIG. 1. Such instructions can include control commands such as logical operations and data input / output that can be recognized and executed by the processor (220). There is no limitation on the type and / or amount of data that the memory (210) can store, but this document will describe the configuration and function of the memory related to the operation of the processor (220) that performs the method and the method of confirming a user command according to various embodiments. The memory (210) can store various information, and the various information stored by the memory (210) will be described in detail below.

[0061] According to one embodiment, at least one processor (220, hereinafter, processor) may have at least a portion of the same or similar configuration as the processor (120) of FIG. 1. According to one embodiment, the processor (220) may include one or more processors.

[0062] According to one embodiment, the processor (220) may perform various operations by executing instructions stored in the memory (210).

[0063] According to one embodiment, the processor (220) may identify an audio source corresponding to at least one type from a first audio signal including a plurality of audio sources. According to one example, the at least one type may include various types including a speech type, a music type, an alarm type, or a siren type. According to one example, the first audio signal may include a plurality of audio sources. According to one example, the processor (220) may identify an audio source corresponding to at least one type from among the plurality of audio sources included in the first audio signal. According to one example, the first audio signal may include an audio source corresponding to at least one type and remaining signals (or unclassified signals or unidentified signals) excluding the identified audio sources.

[0064] In one example, the processor (220) can identify at least one audio source from the first audio signal. In one example, the audio source corresponding to at least one type includes a feature (e.g., a waveform or a main frequency band, etc.) corresponding to each type, and the processor (220) can identify at least one audio source based on this. In one example, the processor (220) can identify at least one audio source from the first audio signal using a designated sound source separation algorithm. In one example, the processor (220) can identify at least one audio source from the first audio signal using an artificial intelligence model. This will be described in detail with reference to FIG. 4.

[0065] For example, when at least one audio source is identified, the processor (220) can identify the type corresponding to each identified audio source. For example, the processor (220) can identify the type corresponding to at least one audio source using an artificial intelligence model capable of classifying the type corresponding to the input audio source. This will be described in detail with reference to FIGS. 4 and 8.

[0066] According to one embodiment, the processor (220) may signal-process an audio source corresponding to at least one identified type based on a setting value corresponding to at least one type. According to one example, the setting value may include at least one of a weight setting value (or weight) corresponding to at least one type and a filter setting value corresponding to at least one type. According to one example, when an audio source (or first audio source) corresponding to a first type and an audio source (or second audio source) corresponding to a second type are identified from a first audio signal, the processor (220) may obtain a signal-processed audio source by applying a first weight corresponding to the first type to the audio source corresponding to the first type and applying a second weight corresponding to the second type to the audio source corresponding to the second type. In the embodiments described below, identifying two types of audio sources is described as an example, but is not limited thereto, and three or more types of audio sources may be identified.

[0067] For example, the processor (220) may signal-process an audio source using a filter (or filter setting value) corresponding to at least one type. For example, when an audio source corresponding to a first type is identified, the processor (220) may filter the audio source corresponding to the first type through a filter corresponding to the first type. For example, there may be a main frequency band corresponding to each type of audio, including a music type or a speech type, and a filter corresponding to a type may provide an audio signal with improved sound quality by filtering a signal in the main frequency band corresponding to the type. This will be described in detail with reference to FIGS. 5 and 6.

[0068] According to one embodiment, the processor (220) may provide a second audio signal based on at least a portion of the remaining signals excluding at least one audio source identified among the first audio signals and the signal-processed audio sources. According to one embodiment, it may be assumed that an audio source corresponding to the first type and an audio source corresponding to the second type are identified from the first audio signal. According to one embodiment, the processor (220) may identify the remaining signals excluding the audio source corresponding to the first type and the audio source corresponding to the second type from the first audio signal. According to one embodiment, the processor (220) may mix each audio source to which a setting value has been applied and at least a portion of the remaining signals to generate, output, or identify the second audio signal. According to one embodiment, the processor (220) may signal-process the remaining signals using the setting value corresponding to the remaining signals. According to one example, the processor (220) may mix at least a portion of the signal-processed remaining signals and the signal-processed audio sources to generate, output, or identify the second audio signal.

[0069] FIG. 3 is a flowchart illustrating an operating method of an electronic device according to one embodiment.

[0070] Referring to FIG. 3, according to one embodiment, the operating method may, in operation 301, identify an audio source corresponding to at least one type from a first audio signal including a plurality of audio sources. According to one example, an electronic device (e.g., the electronic device (200) of FIG. 2) may identify an audio source corresponding to a first type and an audio source corresponding to a second type from the first audio signal using a designated sound source separation algorithm or an artificial intelligence model. According to one example, the electronic device may identify the remaining signal together with the audio source corresponding to the first type and the audio source corresponding to the second type.

[0071] According to one embodiment, the operating method may signal-process an audio source corresponding to at least one identified type based on a setting value corresponding to at least one type in operation 303. According to one embodiment, the electronic device may signal-process an audio source based on at least one of a weight corresponding to the first type or a filter setting value corresponding to the first type as a setting value corresponding to the audio source corresponding to the first type. According to one embodiment, when an audio source corresponding to the first type is identified from the first audio signal, the electronic device may signal-process the audio source corresponding to the first type using the weight corresponding to the identified first type and the filter corresponding to the first type. For example, when an audio source corresponding to the second type is identified, the electronic device may signal-process the audio source corresponding to the second type using the weight corresponding to the identified second type and the filter corresponding to the second type.

[0072] According to one embodiment, the operating method may provide a second audio signal based on at least a portion of the remaining signals excluding the identified audio source from the first audio signal and the signal-processed audio source, in operation 305. According to one embodiment, the electronic device may identify the remaining signals excluding the audio source corresponding to the first type and the audio source corresponding to the second type from the first audio signal. According to one embodiment, the electronic device may generate, output, or identify the second audio signal by mixing the signal-processed audio source corresponding to the first type and the signal-processed audio source corresponding to the second type with the remaining signals. Alternatively, according to one embodiment, the electronic device may generate, output, or identify the second audio signal using only at least a portion of the remaining signals.

[0073] According to the above-described example, the electronic device can improve the sound quality of an existing audio signal and reduce sound loss in the mixing process of an audio source by performing signal processing on an audio source corresponding to a specified type and mixing the remaining signals with the signal on which signal processing has been performed.

[0074] FIG. 4 is a flowchart illustrating a method for determining a type corresponding to an audio source according to one embodiment.

[0075] Referring to FIG. 4, according to one embodiment, in operation 401, the operating method can identify at least one audio source included in a first audio signal (e.g., the first audio signal of FIG. 3) based on inputting the first audio signal into a first artificial intelligence model.

[0076] According to one embodiment, the first artificial intelligence model may be an artificial intelligence model trained to separate at least one audio source included in an audio signal. According to one embodiment, the first artificial intelligence model may be a model composed of an artificial neural network. For example, the first artificial intelligence model may be a model implemented with a deep neutral network (DNN). According to one embodiment, the first artificial intelligence model may be a model trained such that the sum of the output value of the remaining signal and the output value of the at least one audio source is identical to or similar to the output value of the first audio signal. According to one embodiment, the first artificial intelligence model may be a model trained such that the sum of the input audio signal, the output value of the at least one audio source, and the remaining signal is maintained. The first artificial intelligence model will be described in detail with reference to FIGS. 13a, 14, and 15.

[0077] According to one embodiment, an electronic device (e.g., electronic device (200) of FIG. 2) can input a first audio signal into a first artificial intelligence model to identify at least one audio source and remaining signals excluding at least one audio source.

[0078] According to one embodiment, in operation 403, the operating method may respectively identify a type (e.g., a type of FIG. 3) corresponding to each of the identified at least one audio source. According to one embodiment, the electronic device may identify each type for at least one audio source included in the first audio signal. For example, the electronic device may identify a first audio source included in the first audio signal as a first type (e.g., a music type). For example, the electronic device may identify a second audio source included in the first audio signal as a second type (e.g., a voice type). According to one embodiment, the electronic device may identify audio sources that are not identified as a specified type as remaining signals. This will be described in detail with reference to FIG. 7.

[0079] According to one embodiment, the electronic device can identify the type corresponding to at least one audio source using a third artificial intelligence model trained to classify the type of the input audio source. According to one embodiment, the third artificial intelligence model can be implemented as a classifier. For example, the third artificial intelligence model can be implemented as a deep neural network (DNN). According to one embodiment, the third artificial intelligence model can be implemented using various types of algorithms, including a support vector machine (SVM) or a linear regression algorithm. The third artificial intelligence model will be described in detail with reference to FIGS. 13A and 16 .

[0080] FIG. 5 is a flowchart illustrating a method for providing a second audio signal according to one embodiment.

[0081] Referring to FIG. 5, according to one embodiment, in operation 501, the operating method may signal process an audio source (e.g., an audio source of FIG. 3) corresponding to at least one identified type based on a filter and weight (e.g., a weight of FIG. 3) corresponding to the identified type (e.g., a type of FIG. 4).

[0082] According to one embodiment, the filter corresponding to the identified type may be a filter capable of filtering a signal in a main frequency band corresponding to the identified type. According to one embodiment, the weight corresponding to the identified type may be a value set based on a signal-to-noise ratio (SNR) corresponding to the audio source. According to one embodiment, the weight may be obtained based on a user input, or the weight may be obtained based on image content corresponding to the first audio. A method for verifying the weight will be described in detail with reference to FIGS. 10 to 12.

[0083] According to one embodiment, when a first type of first audio source included in a first audio signal is identified, an electronic device (e.g., the electronic device (200) of FIG. 2) may filter the first audio source through a filter corresponding to the first type, and apply a weight corresponding to the first type to the filtered first audio source, thereby performing signal processing on the first audio source. This will be described in detail with reference to FIG. 6.

[0084] According to one embodiment, in operation 503, the method may provide an audio signal in which at least a portion of a remaining signal (e.g., the remaining signal of FIG. 3) and a signal-processed audio source are mixed, as a second audio signal (e.g., the second audio signal of FIG. 3). According to one embodiment, the electronic device may identify, as the remaining signal, a signal excluding at least one type of audio source identified among the first audio signal.

[0085] According to one embodiment, the electronic device may generate, output, or verify a second audio signal by mixing the signal-processed audio source and at least a portion of the remaining signals. For example, the electronic device may generate, output, or verify a second audio signal that minimizes output loss that may occur by separating the first audio signal by mixing the signal-processed audio source and the remaining signals. However, the present invention is not limited thereto, and for example, the electronic device may generate, output, or verify the second audio signal using only a portion of the remaining signals.

[0086] FIG. 6 is a flowchart illustrating a method for performing signal processing on a first audio signal according to one embodiment.

[0087] Referring to FIG. 6, according to one embodiment, the operating method may, in operation 601, identify a first type corresponding to a first audio source among at least one audio source (e.g., at least one audio source of FIG. 4), and identify a first audio source filtered through a first filter corresponding to the first type.

[0088] According to one embodiment, an electronic device (e.g., the electronic device (200) of FIG. 2) may identify a first type as a type corresponding to a first audio source through a third artificial intelligence model (e.g., the third artificial intelligence model of FIG. 4). According to one embodiment, the electronic device may perform filtering on the first audio source using a first filter corresponding to the first type. For example, if the first type is a music type, the first filter may filter the first audio signal so that the proportion of a signal in a main frequency band corresponding to an audio signal of the music type increases. The filter will be described in detail with reference to FIGS. 18A to 18C.

[0089] According to one embodiment, the method of operation may signal process the first audio source based on a first weight corresponding to the first type and a filtered first audio source, in operation 603.

[0090] According to one embodiment, when filtering is performed on a first audio source, the electronic device may perform signal processing on the first audio source by multiplying the filtered first audio source by a set first weight. For example, the electronic device may obtain the first weight based on a user input. The electronic device may determine the proportion of the first audio source within the first audio signal (or the SNR value corresponding to the first audio source) as the first weight. The electronic device may signal process the first audio source by multiplying the filtered first audio source by the first weight. For example, the electronic device may determine a second audio signal (e.g., the second audio signal of FIG. 3) based on the signal-processed first audio source.

[0091] FIG. 7 is a flowchart illustrating a method for verifying the remaining signals according to one embodiment.

[0092] Referring to FIG. 7, according to one embodiment, the operating method may, in operation 701, determine an audio type (e.g., the type of FIG. 6) corresponding to the second audio source as a second type based on determining that an output value of a second audio source among at least one audio source (e.g., the at least one audio source of FIG. 6) is greater than or equal to a first value.

[0093] According to one embodiment, an electronic device (e.g., the electronic device (200) of FIG. 2) may input a first audio signal to a first artificial intelligence model (e.g., the first artificial intelligence model of FIG. 4) to identify at least one audio source and signals remaining from the at least one audio source. According to one embodiment, the electronic device may compare an output value of a second audio source among the at least one audio source with a first value to identify whether the output value of the second audio source is greater than or equal to the first value. If it is identified that the output value of the second audio source is greater than or equal to the first value, the electronic device may identify an audio type corresponding to the second audio source as the second type.

[0094] According to one embodiment, the electronic device may input a second audio source to a third artificial intelligence model (e.g., the third artificial intelligence model of FIGS. 13a-13b) to identify an audio type corresponding to the second audio source as the second type. According to one embodiment, the third artificial intelligence model may be a model trained to output a type corresponding to the audio source. According to one embodiment, when the input audio source is classified as an audio source of a type, if an output value of the input audio source is greater than or equal to a specified value, the third artificial intelligence model may identify it as an audio source corresponding to the type.

[0095] According to one embodiment, the operating method may, at operation 703, identify the second audio source as a remaining signal (e.g., the remaining signal of FIG. 3) based on determining that the output value of the second audio source is less than the first value. According to one embodiment, the electronic device may compare the output value of the second audio source among at least one audio source with the first value to determine whether the output value of the second audio source is less than the first value. According to one embodiment, the electronic device may also identify audio sources less than a specified output value as the remaining signal, separately from the remaining signals output through the first artificial intelligence model.

[0096] In one embodiment, the electronic device may input the second audio source to a third artificial intelligence model (e.g., the third artificial intelligence model of FIGS. 13a-13b ) to identify the second audio source as a residual signal. In one embodiment, even if the second type corresponding to the second audio source is identified, the third artificial intelligence model may identify the output value of the second audio source as a residual signal if it is less than a specified value.

[0097] Figure 8 is a diagram for explaining a second artificial intelligence model according to one embodiment.

[0098] Referring to FIG. 8, according to one embodiment, an electronic device (e.g., the electronic device 200 of FIG. 2) may include a second artificial intelligence model (800). According to one embodiment, the second artificial intelligence model (800) may be a model trained to output an audio source (or at least one type of noise signal) of at least one noise type (a first noise type, a second noise type, ..., or an n-th noise type) included in an input residual signal when the residual signal is input. According to one embodiment, the second artificial intelligence model (800) may be a neural network model implemented with a DNN. According to one embodiment, the at least one noise type may include various types including a babble noise type, a white noise type, a wind noise type, or an engine noise type. According to one embodiment, the second artificial intelligence model (800) may be trained to distinguish and output at least one type of noise signal included in an input audio signal based on a feature corresponding to each noise type.

[0099] According to one embodiment, the electronic device can identify a noise signal corresponding to at least one type from the first audio signal (e.g., the first audio signal of FIG. 3) based on inputting the remaining signals among the first audio signals to the second artificial intelligence model (800). According to one embodiment, the electronic device can identify a noise signal corresponding to a white noise type included in the first audio signal and a noise signal of a machine sound type.

[0100] According to one embodiment, even if the second type (e.g., the type of FIG. 7) corresponding to the second audio source is identified, if the output value of the second audio source is less than a specified value, the electronic device may identify it as a remaining signal or include it in the remaining signal. According to one embodiment, the remaining signal including the second audio source may be input to the second artificial intelligence model (800). The second artificial intelligence model (800) may separate the second audio source whose output value is less than a specified value from among the signals included in the remaining signal.

[0101] According to the above-described example, the electronic device can not only classify the main audio source (e.g., voice type or music type) included in the audio signal, but also classify a plurality of different types of noise signals included in the audio signal and provide the classified information to the user. The electronic device can provide various information to the user by identifying not only the audio source of the specified type, but also the audio source of the specified noise type (or the noise signal of the specified type).

[0102] FIG. 9 is a drawing for explaining a method of providing a user interface (UI) according to one embodiment.

[0103] Referring to FIG. 9, according to one embodiment, an electronic device (e.g., electronic device (200) of FIG. 2) may provide a UI (910) including information about each of at least one type of audio source identified. According to one embodiment, the electronic device may further include a display (900, e.g., display module (160) of FIG. 1). The electronic device may provide the UI (910) through the display (900).

[0104] According to one embodiment, an electronic device may identify, from a first audio signal (e.g., the first audio signal of FIG. 3), an audio source corresponding to at least one type (e.g., an audio source corresponding to at least one type of FIG. 5) and noise corresponding to at least one type (e.g., a noise type of FIG. 8). According to one embodiment, the electronic device may identify information about the audio source corresponding to the identified at least one type and information about the audio source corresponding to at least one noise type. According to one embodiment, the information about the audio source corresponding to at least one type may include information related to adjustment of an output value of the audio source corresponding to at least one type. According to one embodiment, the information about the audio source corresponding to at least one noise type may include information related to adjustment of an output value of the audio source corresponding to at least one noise type.

[0105] For example, it can be assumed that a first audio signal includes an audio source of a user voice type, an audio source of a speech type, an audio source of a music type, an audio source of a wind type, and an audio source of an unclassified noise type. The electronic device can provide a UI to guide adjustment of an output corresponding to each identified type of audio. For example, when a user input corresponding to a button (911) corresponding to a user voice type is received, the UI can include information for adjusting an output value corresponding to the user voice type within a specified range. According to one embodiment, the UI can include a button (911) corresponding to a user voice type, a button (912) corresponding to a speech type, a button (913) corresponding to a music type, a button (914) corresponding to a wind type, and a button (915) corresponding to an unclassified noise type. The electronic device can change an output value for an audio source based on a user input corresponding to different types of buttons included in the UI.

[0106] According to the above-described example, the electronic device can provide a second audio signal (e.g., the second audio signal of FIG. 3) based on a user's request by providing information about various types of noise signals as well as information about the main audio source (e.g., voice type or music type) included in the audio signal.

[0107] Figure 10 is a flowchart illustrating a method for checking weights according to one embodiment.

[0108] Referring to FIG. 10, according to one embodiment, the operating method may, in operation 1001, determine a first weight (e.g., a weight of FIG. 6) corresponding to a first type based on an output value of a first audio source and an output value of the remaining audio signals (e.g., the first audio signal of FIG. 3) excluding the first audio source (e.g., the audio source of FIG. 6).

[0109] According to an example, an electronic device (e.g., the electronic device 200 of FIG. 2) may identify a first audio source (e.g., a music type audio source) corresponding to a first type among a first audio signal. According to an embodiment, the electronic device may compare an output value of the first audio source with an output value of the first audio signal to identify a first weight corresponding to the first type. For example, the first weight may be a ratio of the first audio source to the first audio signal. However, the present invention is not limited thereto, and for example, the first weight may be a ratio of the first audio source to the remaining signals excluding the first audio signal.

[0110] According to one embodiment, the operating method may signal-process an audio source corresponding to at least one type identified based on the identified weights in operation 1003. According to one example, when a first weight of a first type corresponding to a first audio source is identified, the electronic device may apply the first weight to the first audio source and identify a second audio signal (e.g., the second audio signal of FIG. 3) based on the first audio source to which the first weight is applied.

[0111] Figure 11 is a flowchart illustrating a method for checking weights according to one embodiment.

[0112] Referring to FIG. 11, according to one embodiment, the operating method may, in operation 1101, identify context information corresponding to video content corresponding to a first audio signal. According to one embodiment, the context information corresponding to the video content may be information about a type corresponding to an object (e.g., a person or a piano) included in the video content. According to one embodiment, the electronic device (e.g., the electronic device (200) of FIG. 2) may identify an object included in the video content by using an image analysis algorithm for identifying an object included in the video content.

[0113] In one embodiment, the electronic device can identify an audio type corresponding to the identified object. In one example, a memory (e.g., memory (130) of FIG. 1) may store information about an object corresponding to at least one type (e.g., type of FIG. 4). In one embodiment, the type corresponding to a person may be 'speech'. In one embodiment, the type corresponding to a piano may be 'music'. For example, if a video content includes a person and a piano, the electronic device can identify the speech type and the music type as context information corresponding to the first audio signal.

[0114] In one embodiment, an electronic device can identify contextual information corresponding to video content using a designated image recognition algorithm. For example, if the video content is "content featuring children singing," the electronic device can identify "voice type" and "music type" as contextual information corresponding to the video content.

[0115] According to one embodiment, the operating method may, in operation 1103, determine a weight corresponding to at least one type based on the identified context information. According to one embodiment, when a type corresponding to video content is determined based on the identified context information, the electronic device may update the weight so that the weight of the audio source corresponding to the determined type increases. For example, when the context information is determined to be a voice type and a music type, the weight may be updated so that the weight for the audio source corresponding to the determined voice type and music type increases.

[0116] In one embodiment, the electronic device may update the weights to decrease the weights corresponding to audio sources other than those corresponding to the identified type. For example, if the music type is not identified among the context information, the electronic device may decrease the weight corresponding to the music type.

[0117] Figure 12 is a flowchart illustrating a method for checking weights according to one embodiment.

[0118] Referring to FIG. 12, according to one embodiment, the operating method may update a weight corresponding to each of at least one type (e.g., a type of FIG. 4) based on receiving a user input for changing at least one of the set values ​​(e.g., a weight of FIG. 3) corresponding to each of at least one type, in operation 1201.

[0119] According to one embodiment, an electronic device (e.g., the electronic device (200) of FIG. 2) may include an input module (e.g., the input module (150) of FIG. 1). According to one embodiment, the electronic device may receive, through the input module, a user input for changing a weight corresponding to each of at least one type. According to one embodiment, the electronic device may change the weight corresponding to each type based on the received input. For example, if a user input for increasing a weight of a music type is received, the electronic device may increase the weight corresponding to the music type.

[0120] Referring to FIG. 12, according to one embodiment, the operating method may signal-process an audio source corresponding to at least one type identified based on an updated weight in operation 1203. According to one embodiment, when the updated weight is identified, the electronic device may perform signal processing on the audio source by applying the identified weight to the corresponding audio source. According to one embodiment, the electronic device may identify a second audio signal (e.g., the second audio signal of FIG. 3) based on the signal-processed audio source.

[0121] FIG. 13a is a diagram for explaining a first artificial intelligence model and a third artificial intelligence model according to one embodiment.

[0122] Referring to FIG. 13A, according to one embodiment, an electronic device (e.g., the electronic device 200 of FIG. 2) may include a first artificial intelligence model (1310) and a third artificial intelligence model (1320). According to one embodiment, the first artificial intelligence model (1310) may be a model trained to separate an input audio signal into at least one audio source when the audio signal is input. According to one embodiment, the first artificial intelligence model (1310) may be a model trained based on machine learning. According to one embodiment, the first artificial intelligence model (1310) may be trained to separate audio sources based on unique characteristics of each type (e.g., the types of FIG. 6). According to one embodiment, the types of the audio sources may be various types including a speech type, a music type, an alarm type, or a siren type. According to one embodiment, the first artificial intelligence model (1310) may be implemented as an artificial intelligence model based on a DNN (Deep Neural Network), but may also be implemented as a DNN and various types of models.

[0123] According to one embodiment, the electronic device may input a first audio signal (e.g., the first audio signal of FIG. 3) to a first artificial intelligence model (1310) to identify a plurality of audio sources (a first audio source, a second audio source, ..., or an n-th audio source). According to one embodiment, the plurality of audio sources may be audio signals included in the first audio signal. According to one example, the first artificial intelligence model (1310) may also output the remaining signals excluding the plurality of audio sources from among the first audio signal.

[0124] According to one embodiment, the first artificial intelligence model (1310) may be trained so that the sum of the magnitudes of the input audio signal and the magnitudes of the output audio signal are equal. According to one embodiment, the cost function of the first artificial intelligence model (1310) may include a difference value between the sum of the magnitudes of the output signal and the sum of the magnitudes of the input signal. Accordingly, the first artificial intelligence model (1310) may be trained so that the difference value between the sum of the magnitudes of the output signal and the sum of the magnitudes of the input signal is 0. For example, the input signal may be a first audio signal, and the output signal may include a plurality of audio sources separated from the first audio signal and the remaining signal. According to one embodiment, the first artificial intelligence model (1310) may be implemented as a model including a plurality of modules, or may be implemented as a single model. This will be described in detail with reference to FIGS. 14 and 15.

[0125] In one embodiment, the first AI model (1310) may be a model trained to output an audio source corresponding to at least one type. For example, the first AI model (1310) may be a model trained to output an audio source of the music type as the first audio source, and an audio source of the voice type as the second audio source.

[0126] According to one embodiment, the third artificial intelligence model (1320) may be a model trained to classify the type of an input audio signal when multiple audio sources are input. According to one embodiment, the electronic device may identify a type corresponding to at least one audio source using the third artificial intelligence model (1320) trained to classify the type of the input audio source. According to one embodiment, the third artificial intelligence model (1320) may be implemented as a classifier. For example, the third artificial intelligence model (1320) may be implemented as a DNN. According to one embodiment, the third artificial intelligence model (1320) may be implemented based on various types of algorithms, including a support vector machine (SVM) or a linear regression algorithm.

[0127] According to one embodiment, when multiple audio sources are input, the third artificial intelligence model (1320) can classify the multiple input audio sources and check the output of the classified audio sources to identify the type of the audio source. For example, when the first audio source separated through the first artificial intelligence model (1310) is input to the third artificial intelligence model (1320), the third artificial intelligence model (1320) can identify the type of the first audio source. Even when the type of the first audio source is identified, if the output value of the first audio source is identified as being less than a specified value, the third artificial intelligence model (1320) can identify the first audio source as the remaining signal (e.g., the remaining signal of FIG. 3). According to one embodiment, the third artificial intelligence model (1320) may be a model that only classifies the multiple input audio sources.

[0128] In one embodiment, the third artificial intelligence model (1320) may include a second artificial intelligence model. For example, it may be assumed that the third artificial intelligence model (1320) includes a second artificial intelligence model (e.g., the second artificial intelligence model of FIG. 8). When the remaining signal is input to the third artificial intelligence model (1320) along with multiple audio sources, the third artificial intelligence model (1320) may perform classification on the multiple audio sources and classify the input remaining signal as at least one type of noise signal. In one embodiment, the third artificial intelligence model (1320) may be a model implemented with multiple classification models. This will be described in detail with reference to FIG. 17.

[0129] According to the example described above, the first artificial intelligence model (1310) performs a function of separating an audio signal into a plurality of audio sources and a remaining signal, and the third artificial intelligence model (1320) can classify the plurality of audio sources by corresponding types or classify the remaining signal into at least one type of noise signal.

[0130] FIG. 13b is a diagram for explaining a method for confirming a second audio signal according to one embodiment.

[0131] Referring to FIG. 13b, according to one embodiment, an electronic device (e.g., electronic device (200) of FIG. 2) may include a third artificial intelligence model (1320), an input receiving module (1330), a context recognition module (1340), and a mixing module (1350).

[0132] According to one embodiment, the input receiving module (1330) may be a module that receives a user input. According to one embodiment, based on the provision of a UI (e.g., UI (910) of FIG. 9) through a display (e.g., display module (160) of FIG. 1) as illustrated in FIG. 9, the input receiving module (1330) may identify a user input corresponding to at least one type of audio source (e.g., at least one type of audio source of FIG. 3) or noise corresponding to at least one type (e.g., noise signal of FIG. 8). According to one embodiment, the input receiving module (1330) may receive information on the type of at least one audio source included in the first audio signal through the third artificial intelligence model (1320). The input receiving module (1330) may provide a UI based on the information on the received type.

[0133] According to one embodiment, when a user input for adjusting an output value corresponding to each type of audio source or noise signal is confirmed through the input receiving module (1330), the electronic device can adjust the output value of each audio source or noise signal based on the received input. According to one embodiment, the electronic device can also confirm a relative output value of each type of audio source or noise signal based on the received input, and confirm a weight based on the confirmed relative output value. According to one embodiment, the input receiving module (1330) can transmit the confirmed weight or the confirmed output value to the mixing module (1350).

[0134] According to one embodiment, the context recognition module (1340) may identify context information (e.g., context information of FIG. 11) corresponding to the video content from the video content corresponding to the first audio signal (e.g., the first audio signal of FIG. 3). According to one embodiment, the context information may be information about a type corresponding to an object (e.g., a person or a piano) included in the video content. According to one embodiment, the context recognition module (1340) may identify the type of the object included in the video content using an image analysis algorithm for identifying the object included in the video content. The context recognition module (1340) may also identify the context information based on the content of the video content.

[0135] According to one embodiment, the context recognition module (1340) may identify a weight corresponding to at least one type (e.g., the type of FIG. 4) based on the identified context information, and transmit the identified weight to the mixing module (1350). According to one embodiment, the context recognition module (1340) may receive information on the type of at least one audio source included in the first audio signal through the third artificial intelligence model (1320). The context recognition module (1340) may identify a weight corresponding to at least one type based on the information on the received type and the identified context information.

[0136] In one embodiment, the context recognition module (1340) may transmit the identified context information itself to the mixing module (1350). In this case, the mixing module (1350) may also identify a weight corresponding to at least one type based on the received context information.

[0137] According to one embodiment, the mixing module (1350) may receive at least one type of audio source included in the first audio signal from the third artificial intelligence model (1320). According to one embodiment, the mixing module (1350) may perform signal processing on the at least one type of audio source to generate, output, or confirm a second audio signal. For example, the mixing module (1350) may apply a set filter value and a set weight corresponding to each of the at least one type of audio source to confirm the signal-processed audio source. According to one embodiment, the mixing module (1350) may perform signal processing on the audio source based on the weight received from at least one of the input receiving module (1330) or the context recognition module (1340). According to one embodiment, the mixing module (1350) may mix each of the signal-processed audio sources to generate, output, or confirm a second audio signal.

[0138] According to one embodiment, the mixing module (1350) may check the signal-processed remaining signal by applying at least one of a set filter value and a set weight corresponding to the remaining signal (e.g., the remaining signal of FIG. 3) and the remaining signal. According to one embodiment, the mixing module (1350) may generate, output, or check the second audio signal by applying a set filter value and a set weight to the received remaining signal when a remaining signal including at least one type of noise signal is received from at least one of the second artificial intelligence model (e.g., the second artificial intelligence model of FIG. 8) or the third artificial intelligence model (1320).

[0139] Fig. 14 is a diagram for explaining an implementation example of a first artificial intelligence model according to one embodiment.

[0140] Referring to FIG. 14, according to one embodiment, a first artificial intelligence model (e.g., the first artificial intelligence model of FIG. 4) may include a plurality of modules (1410, 1420, and 1430). According to one example, the first artificial intelligence model may be a model that performs encoding on an input audio signal (e.g., the first audio signal of FIG. 3), masks the encoded audio signal, and decodes the masked audio signal to output at least one type of audio source (a first audio source, a second audio source, ..., or an n-th audio source). In this case, the first artificial intelligence model may output a remaining signal (e.g., the remaining signal of FIG. 3) together with the at least one type of audio source.

[0141] According to one embodiment, the encoder (1410) may perform encoding on an input audio signal. According to one embodiment, the first module (1420) may separate the encoded audio signal into information corresponding to a plurality of audio sources by masking the encoded audio signal with a plurality of masks. According to one embodiment, the masks may be an algorithm (or a value output from the algorithm) for estimating a magnitude spectra corresponding to each of a plurality of audio sources included in the audio signal in order to separate the audio signal including the plurality of audio sources. According to one embodiment, the decoder (1430) may perform decoding on the plurality of masked sources. According to one embodiment, the first artificial intelligence model may be trained to satisfy the following mathematical expression 1.

[0142]

[0143] Referring to mathematical expression 1, as an example, when separating the input first audio signal into N audio sources, may be a mask value corresponding to the nth source (or encoded audio source) included in the nth audio signal. For example, the size of n may be less than or equal to N. Among the first audio signals, the remaining signals may be signals excluding N separated audio sources. According to mathematical expression 1, the first artificial intelligence model can be trained so that the size of the input signal and the size of the output signal are maintained even when the audio signal is separated by maintaining the sum of the entire mask as 1.

[0144] According to one embodiment, in order to perform learning in which the size of the input signal and the size of the output signal are maintained, the cost function of the first artificial intelligence model includes: ) may be included. In one example, Fn(x) may be implemented as various types of functions, such as sigmoid(x), hypertangent(x), exponential(x), and log(x), as examples of cost functions. By including the above-described cost function, the first artificial intelligence model may be trained so that the sum of all masks approaches 1. In one embodiment, the first artificial intelligence model may be trained using a machine learning method.

[0145] According to one embodiment, the first artificial intelligence model may separate an audio signal into multiple audio sources by masking the encoded audio signal with masks corresponding to multiple sources when the audio signal is encoded through the encoder (1410) and performing decoding on each of the multiple masked sources. In this case, the sum of the masks corresponding to each of the multiple sources is trained to remain 1, thereby preventing signal loss in the process of separating the audio signal into multiple audio sources.

[0146] Fig. 15 is a drawing for explaining an implementation example of a first artificial intelligence model according to one embodiment.

[0147] Referring to FIG. 15, according to one embodiment, a first artificial intelligence model (e.g., the first artificial intelligence model of FIG. 4) may be implemented as a fourth artificial intelligence model (1510).

[0148] According to one embodiment, the fourth artificial intelligence model (1510) may be a model implemented as a single artificial intelligence model, unlike FIG. 14. According to one embodiment, the fourth artificial intelligence model (1510) may be an end-to-end (E2E) model trained to output at least one audio source when an audio source is input. According to one embodiment, the fourth artificial intelligence model (1510) may be a model that outputs at least one type of audio source (a first audio source, a second audio source, ..., or an n-th audio source) when an audio signal is input. In this case, the fourth artificial intelligence model (1510) may output the remaining signal (e.g., the remaining signal of FIG. 3) together with the at least one type of audio source.

[0149] According to one embodiment, the fourth artificial intelligence model (1510) can be trained to satisfy the following mathematical expression 2.

[0150]

[0151] Referring to mathematical expression 2, as an example, when separating the input first audio signal into N audio sources, , may be the magnitude of the nth audio signal. For example, the magnitude of n may be less than or equal to N. In one embodiment, Sin may be the magnitude of the input signal. Silver may be the size of the remaining signal excluding the N separated audio sources among the input first audio signals. According to mathematical expression 2, the fourth artificial intelligence model (1510) can be trained to maintain the size of the input signal and the size of the output signal even when the audio signal is separated into multiple audio sources.

[0152] According to one embodiment, in order to perform learning in which the size of the input signal and the size of the output signal are maintained, the cost function of the first artificial intelligence model includes: ) may be included. In one embodiment, Fn(x) may be implemented as various types of functions, such as sigmoid(x), hypertangent(x), exponential(x), or log(x), as an example of a cost function. By including the above-described cost function, the fourth artificial intelligence model (1510) may be trained so that the difference between the size of the input signal and the size of the output signal approaches 0. In one example, the fourth artificial intelligence model (1510) may be trained using a machine learning method.

[0153] Fig. 16 is a diagram for explaining an implementation example of a third artificial intelligence model according to one embodiment.

[0154] Referring to FIG. 16, according to one embodiment, the third artificial intelligence model (e.g., the third artificial intelligence model (1320) of FIG. 13A) may be a model implemented with a plurality of detection modules (1610, 1620, ..., 1630). According to one embodiment, the plurality of detection modules (1610, 1620, ..., 1630) including a first detection module, a second detection module, and an n-th detection module may be modules trained to detect an audio source of a type corresponding to each detection module. According to one embodiment, each of the plurality of audio sources separated from the fourth artificial intelligence model (1510) may be input to each of the plurality of detection modules (1610, 1620, ..., 1630).

[0155] According to one embodiment, the plurality of detection modules (1610, 1620, ..., 1630) may be implemented as classifiers. According to one embodiment, for example, each of the plurality of detection modules (1610, 1620, ..., 1630) may be a neural network model implemented as a DNN. For example, the first detection module (1610) may detect an audio source of a first type (e.g., a music type). According to one embodiment, when a first audio source is input, the first detection module (1610) may determine whether the first audio source is an audio source of the first type. According to one embodiment, the first detection module (1610) may be a module that outputs (or bypasses) the first audio source as a first output signal (1611) when the first audio source is determined to be an audio source of the first type.

[0156] According to one embodiment, the plurality of detection modules (1610, 1620, ..., 1630) may classify the input audio source as a remaining signal (e.g., the remaining signal of FIG. 3) when the size of the input audio source is less than a specified value even when the type of the input audio source is confirmed. For example, the second detection module (1620) may be a module that detects an audio source of a second type (e.g., a voice type). The second detection module (1620) may confirm the second audio source as a remaining signal (e.g., the remaining signal of FIG. 3) when the size of the second audio source is less than a first value even when the input second audio source is confirmed as the second type. According to one embodiment, when the size of the second audio source is greater than or equal to the first value, the second detection module (1620) may output the second audio source as a second output signal (1621). If the size of the second audio source is less than the first value, the second detection module (1620) may output the second audio source as a remaining signal (e.g., the nth output signal (1631)).

[0157] In one embodiment, the third artificial intelligence model may include a second artificial intelligence model (e.g., the second artificial intelligence model of FIG. 8). For example, if the n-th detection module (1630) is implemented as the second artificial intelligence model, the n-th detection module (1630) may be a model trained to output at least one type of noise signal included in the input remaining signal when the remaining signal is input. In one embodiment, the n-th detection module (1630) may be a neural network model implemented as a DNN. In one embodiment, the at least one type of noise signal (or, an audio source of at least one type of noise) may include various types including a babble noise type, a white noise type, a wind noise type, or an engine noise type. In one embodiment, the n-th detection module (1630) may be trained to distinguish and output at least one type of noise signal included in the input remaining signal based on features corresponding to each noise type.

[0158] According to one embodiment, the nth detection module (1630) may output the remaining signals received from detection modules other than the nth detection module (1630). For example, if the signal output from the second detection module (1620) is identified as the remaining signal, the nth detection module (1630) may output the signal output from the second detection module (1620) together. According to one embodiment, if the signal output from the second detection module (1620) is identified as the remaining signal, the nth detection module (1630) may classify it as one of at least one type of noise signal. According to one embodiment, the nth output signal (1631) may include at least one audio signal corresponding to each of at least one type of noise signal.

[0159] In one embodiment, the first artificial intelligence model (e.g., the first artificial intelligence model of FIG. 4), the second artificial intelligence model, the third artificial intelligence model, the fourth artificial intelligence model (e.g., the fourth artificial intelligence model (1510) of FIG. 15), and the fifth artificial intelligence model described below may be artificial intelligence models included in an electronic device (e.g., the electronic device (200) of FIG. 2). According to one embodiment, the first artificial intelligence model, the second artificial intelligence model, the third artificial intelligence model, the fourth artificial intelligence model, and the fifth artificial intelligence model described below may be artificial intelligence models included in an external electronic device (e.g., the electronic device (104) of FIG. 1 or the server (108) of FIG. 1). According to one embodiment, data (e.g., at least one type of noise signal) output from an artificial intelligence model included in the external electronic device may be transmitted to the electronic device via a communication module (e.g., the communication module (190) of FIG. 1).

[0160] Fig. 17 is a diagram for explaining a fifth artificial intelligence model according to one embodiment.

[0161] According to FIG. 17, according to one embodiment, an electronic device (e.g., the electronic device (200) of FIG. 2) may include a fifth artificial intelligence model (1700). According to one example, the fifth artificial intelligence model (1700) may be a model trained to output a second audio signal (e.g., the second audio signal of FIG. 3) when a first audio signal (e.g., the first audio signal of FIG. 3) is input. According to one embodiment, the fifth artificial intelligence model (1700) may be a single neural network model implemented to output a second audio signal corresponding to improved sound quality than the sound quality of the first audio signal by using the first audio signal and a designated weight when the first audio signal is input.

[0162] According to one embodiment, the fifth artificial intelligence model (1700) may be implemented as a machine learning model. According to one embodiment, the model may be trained so that a target value (e.g., an ideal second audio signal) and a set weight value (e.g., a weight value of FIG. 3) for the second audio signal are input as labels, and the difference between the target value for the second audio signal and the output second audio signal is reduced. Accordingly, a prediction error for the second audio signal that is ultimately output can be directly reflected in the model. In addition, since the separation operation of the audio source and the mixing operation of the audio source are not performed as separate operations, sound quality deterioration due to leakage components can be minimized.

[0163] According to one embodiment, when the weight values ​​obtained through the input reception module (1330) and the context recognition module (1340) are input to the embedding model (1710), the embedding model (1710) can convert the input weight values ​​into embeddings corresponding to the weight values ​​based on the input weight values. According to one embodiment, the embedding model (1710) can provide the embeddings corresponding to the weight values ​​to the fifth artificial intelligence model (1700). The fifth artificial intelligence model (1700) can output a second audio signal using the received embeddings. According to one embodiment, the embedding model (1710) can be implemented as a neural network model including a plurality of layers.

[0164] Figures 18a to 18c are drawings for explaining a filter according to one embodiment.

[0165] Referring to FIGS. 18A to 18C , according to one embodiment, an electronic device (e.g., the electronic device (200) of FIG. 2 ) may perform signal processing on an audio source corresponding to at least one type using a filter (e.g., the filter of FIG. 6 ) corresponding to each of at least one type. Signal processing may also be performed on an audio source corresponding to at least one noise type (e.g., the noise type of FIG. 8 ).

[0166] In one embodiment, the filter may be implemented as a filter that filters so that the output ratio of a signal corresponding to a designated frequency band is increased, as in the graph illustrated in FIG. 18a. For example, a filter corresponding to wind noise may be implemented as a filter corresponding to the graph illustrated in FIG. 18a. In one embodiment, the filter may be implemented as a filter that filters so that the output ratio of a signal in a low frequency band is increased and signals in a band above a designated frequency band are blocked, as in the graph illustrated in FIG. 18b. For example, a filter corresponding to an audio source of the user voice type may be implemented as a filter corresponding to the graph illustrated in FIG. 18b. In one embodiment, the filter may be implemented as a filter corresponding to various proportions for each frequency band, as in the graph illustrated in FIG. 18c. For example, the filter may be a filter that filters so that an output value in a low frequency band is output with a relatively higher proportion than an output value in a high frequency band. In this case, the filter corresponding to the graph illustrated in Fig. 18c may be a filter corresponding to the remaining signal (e.g., the remaining signal of Fig. 3).

[0167] An electronic device (101; 200) according to one embodiment of the present disclosure may include a memory (130; 210) storing instructions and at least one processor (120; 220). According to one embodiment, the instructions, when executed by the at least one processor (120; 220), may cause the electronic device (101; 200) to identify an audio source corresponding to at least one type from a first audio signal including a plurality of audio sources.

[0168] According to one embodiment, the instructions may cause the electronic device (101; 200) to signal process an audio source corresponding to the at least one identified type based on a setting value corresponding to the at least one type.

[0169] In one embodiment, the instructions may cause the electronic device (101; 200) to provide a second audio signal based on at least a portion of the remaining signal of the first audio signal excluding the identified audio source and the signal-processed audio source.

[0170] In one embodiment, the instructions may cause the electronic device (101; 200) to identify at least one audio source included in the first audio signal based on inputting the first audio signal to the first artificial intelligence model (1310).

[0171] In one embodiment, the instructions may cause the electronic device (101; 200) to identify a type corresponding to each of the identified audio sources.

[0172] According to one embodiment, the first artificial intelligence model (1310) may be a model trained such that the sum of the output value of the remaining signal and the output value of the at least one audio source is equal to the output value of the first audio signal.

[0173] According to one embodiment, the instructions may cause the electronic device (101; 200) to signal process an audio source corresponding to at least one identified type based on a filter and a weight corresponding to the identified type, respectively.

[0174] In one embodiment, the instructions cause the electronic device (101; 200) to provide an audio signal in which at least a portion of the remaining signal and the signal-processed audio source are mixed as the second audio signal.

[0175] In one embodiment, the instructions may cause the electronic device (101; 200) to identify a first audio source filtered through a first filter corresponding to the first type based on identifying a first type corresponding to the first audio source among the at least one audio source.

[0176] In one embodiment, the instructions may cause the electronic device (101; 200) to signal process the first audio source based on the first weight corresponding to the first type and the filtered first audio source.

[0177] In one embodiment, the instructions may cause the electronic device (101; 200) to identify a second type corresponding to the second audio source based on determining that an output value of the second audio source among the at least one audio source is greater than or equal to a first value.

[0178] In one embodiment, the instructions may cause the electronic device (101; 200) to identify the second audio source as the remaining signal based on determining that an output value of the second audio source is less than the first value.

[0179] In one embodiment, the instructions may cause the electronic device (101; 200) to identify an audio source corresponding to at least one noise type from the first audio signal based on inputting the remaining signal of the first audio signal to a second artificial intelligence model (800).

[0180] According to one embodiment, the electronic device (101; 200) may further include a display (900).

[0181] In one embodiment, the instructions may cause the electronic device (101; 200) to provide, through the display (900), a user interface (UI) 910 including information about an audio source corresponding to the at least one type and information about an audio source corresponding to the at least one noise type.

[0182] In one embodiment, the instructions may cause the electronic device (101; 200) to determine a first weight corresponding to a first type corresponding to the first audio source, based on an output value of a first audio source among the at least one audio source and output values ​​of the remaining first audio signals excluding the first audio source.

[0183] In one embodiment, the instructions may cause the electronic device (101; 200) to signal process a first audio source corresponding to the first type based on the identified weight.

[0184] According to one embodiment, the instructions may cause the electronic device (101; 200) to identify context information corresponding to video content corresponding to the first audio signal.

[0185] In one embodiment, the instructions may cause the electronic device (101; 200) to determine a weight corresponding to the at least one type based on the determined context information.

[0186] In one embodiment, the instructions may cause the electronic device (101; 200) to identify a third type associated with the video content based on the identified context information.

[0187] In one embodiment, the instructions may cause the electronic device (101; 200) to update the setting value such that a weight corresponding to the identified third type increases and a weight corresponding to a type other than the third type decreases.

[0188] In one embodiment, the instructions may cause the electronic device (101; 200) to update a weight corresponding to each of the at least one type based on receiving a user input for changing at least one of the setting values ​​corresponding to each of the at least one type.

[0189] In one embodiment, the instructions may cause the electronic device (101; 200) to signal process an audio source corresponding to the identified at least one type based on the updated weights.

[0190] A method of operating an electronic device (101; 200) according to one embodiment of the present disclosure may include an operation of identifying an audio source corresponding to at least one type from a first audio signal including a plurality of audio sources.

[0191] According to one embodiment, the operating method may include an operation of signal processing an audio source corresponding to the at least one type identified based on a setting value corresponding to the at least one type.

[0192] According to one embodiment, the operating method may include providing a second audio signal based on at least a portion of a signal remaining from the first audio signal excluding the identified audio source and the signal-processed audio source.

[0193] In one embodiment, the operation of identifying an audio source corresponding to the at least one type may include an operation of identifying at least one audio source included in the first audio signal based on inputting the first audio signal to the first artificial intelligence model (1310).

[0194] In one embodiment, the operation of identifying an audio source corresponding to at least one type may include the operation of identifying a type corresponding to each of the identified audio sources.

[0195] According to one embodiment, the signal processing operation may include an operation of signal processing an audio source corresponding to at least one identified type based on a filter and a weight corresponding to the identified type, respectively.

[0196] In one embodiment, the act of providing the second audio signal may include providing an audio signal in which at least a portion of the remaining signal and the signal-processed audio source are mixed as the second audio signal.

[0197] According to one embodiment, the signal processing operation may include an operation of identifying a first audio source filtered through a first filter corresponding to the first type based on identifying a first type corresponding to the first audio source among the at least one audio source.

[0198] According to one embodiment, the signal processing operation may include an operation of signal processing the first audio source based on the first weight corresponding to the first type and the filtered first audio source.

[0199] In one embodiment, the operation of checking each of the types may include an operation of checking a second type corresponding to the second audio source based on checking that an output value of the second audio source among the at least one audio source is greater than or equal to the first value.

[0200] In one embodiment, the operation of verifying each of the types may include verifying the second audio source with the remaining signal based on verifying that the output value of the second audio source is less than the first value.

[0201] According to one embodiment, the method of operation may include an operation of identifying an audio source corresponding to at least one noise type from the first audio signal based on inputting the remaining signals of the first audio signal to a second artificial intelligence model (800).

[0202] According to one embodiment, the method of operation may include providing a user interface (UI) 910 that includes information about an audio source corresponding to the at least one type and information about an audio source corresponding to the at least one noise type.

[0203] According to one embodiment, the operation of signal processing the audio source may include an operation of determining a first weight corresponding to a first type corresponding to the first audio source based on an output value of a first audio source among the at least one audio source and an output value of the remaining first audio signals excluding the first audio source.

[0204] According to one embodiment, the operation of signal processing the audio source may include the operation of signal processing a first audio source corresponding to the first type based on the identified weight.

[0205] In a storage medium storing computer-readable instructions according to one embodiment of the present disclosure, the instructions, when executed by at least one processor (120; 220) of an electronic device (101; 200), can cause the electronic device (101; 200) to identify an audio source corresponding to at least one type from a first audio signal including a plurality of audio sources.

[0206] According to one embodiment, the instructions may cause the electronic device (101; 200) to signal process an audio source corresponding to the at least one identified type based on a setting value corresponding to the at least one type.

[0207] In one embodiment, the instructions may cause the electronic device (101; 200) to provide a second audio signal based on at least a portion of the remaining signal of the first audio signal excluding the identified audio source and the signal-processed audio source.

[0208] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains.

[0209] As used herein, the term “if” will be understood to mean “when, upon,” “in response to deciding,” or “in response to detecting,” depending on the context. Similarly, “if it is decided to do,” or “if [the stated condition or event] is detected,” will optionally be understood to mean “upon deciding,” or “in response to deciding,” “upon detecting [the stated condition or event],” or “in response to detecting [the stated condition or event].”

[0210] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. A processing device (or processing circuit) may execute an operating system (OS) and one or more software applications running on the operating system. In addition, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, a single processing device is sometimes described, but one of ordinary skill in the art will recognize that a processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0211] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.

[0212] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include an app store that distributes applications, a site that supplies or distributes various software, and a recording or storage medium managed by a server.

[0213] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components such as the described systems, structures, devices, and circuits are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0214] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

[0215] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0216] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0217] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0218] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0219] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0220] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In electronic devices (101; 200), Memory for storing instructions (130; 210); and Contains at least one processor (120; 220), The above instructions, when executed by the at least one processor (120; 220), cause the electronic device (101; 200) to: Identifying a first audio signal including multiple audio sources, Identifying a first audio source corresponding to at least one type from the first audio signal, Processing the first audio source based on a setting value corresponding to each of the at least one type, Generating a second audio signal based on at least a portion of the remaining signals excluding the first audio source among the first audio signals and the processed first audio source, An electronic device (101; 200) causing the second audio signal to be output.

2. In paragraph 1, To verify the first audio source, the instructions, when executed by the at least one processor (120; 220), cause the electronic device (101; 200) to: Based on inputting the first audio signal to the first artificial intelligence model (1310), at least one second audio source included in the first audio signal is identified, Each of the types corresponding to at least one of the second audio sources is checked, Causing the first audio source to be identified as corresponding to at least one type, The above first artificial intelligence model (1310) is An electronic device (101; 200) learned such that the sum of the first output value of the remaining signal and the second output value of the at least one second audio source is equal to the third output value of the first audio signal.

3. In paragraph 1 or 2, To signal process the first audio source, the instructions, when executed by the at least one processor (120; 220), cause the electronic device (101; 200) to: Obtaining a signal-processed audio source by signal-processing the first audio source based on a filter and a first weight corresponding to at least one type, An electronic device (101; 200) causing the second audio signal to be generated by mixing at least a portion of the remaining signal and the signal-processed audio source.

4. In any one of paragraphs 1 to 3, To signal process the first audio source, the instructions, when executed by the at least one processor (120; 220), cause the electronic device (101; 200) to: Obtaining a first audio source by filtering through a first filter corresponding to the first type based on confirming a first type corresponding to the first audio source among at least one second audio source, An electronic device (101; 200) that causes signal processing of the filtered first audio source based on the first weight corresponding to the first type and the filtered first audio source.

5. In any one of paragraphs 1 to 4, The above instructions, when executed by the at least one processor (120; 220), cause the electronic device (101; 200) to: Based on confirming that the fourth output value of the third audio source among the at least one second audio source is greater than or equal to the first value, a fourth type corresponding to the third audio source is confirmed, An electronic device (101; 200) that causes the third audio source to be identified as the remaining signal based on determining that the fourth output value is less than the first value.

6. In any one of paragraphs 1 to 5, The above instructions, when executed by the at least one processor (120; 220), cause the electronic device (101; 200) to: An electronic device (101; 200) that causes a third audio source corresponding to at least one noise type to be identified from the remaining signal based on inputting the remaining signal to a second artificial intelligence model (800).

7. In any one of paragraphs 1 to 6, The electronic device further comprises a display (900); The above instructions, when executed by the at least one processor (120; 220), cause the electronic device (101; 200) to: An electronic device (101; 200) that causes a UI (user interface, 910) including information about the first audio source and information about the third audio source to be output through the display (900).

8. In any one of paragraphs 1 to 7, The above instructions, when executed by the at least one processor (120; 220), cause the electronic device (101; 200) to: Based on the output value of the first audio source among the plurality of audio sources and the output value of the remaining signals, a first weight corresponding to the first type corresponding to the first audio source is confirmed, An electronic device (101; 200) that causes the first audio source to be signal processed based on the first weight.

9. In any one of paragraphs 1 to 8, The above instructions, when executed by the at least one processor (120; 220), cause the electronic device (101; 200) to: Check context information corresponding to the video content, and the video content corresponds to the first audio signal, An electronic device (101; 200) that causes a first weight corresponding to at least one type to be identified based on the context information.

10. In paragraph 9, The above instructions, when executed by the at least one processor (120; 220), cause the electronic device (101; 200) to: Based on the above context information, identify the second type related to the video content, An electronic device (101; 200) that causes a second weight corresponding to the second type to increase and a third weight corresponding to a third type different from the second type to decrease.

11. In any one of paragraphs 1 to 10, The above instructions, when executed by the at least one processor (120; 220), cause the electronic device (101; 200) to: Based on receiving a user input for changing at least one of the at least one setting value corresponding to each of the at least one type, updating a plurality of weights corresponding to each of the at least one type, An electronic device (101; 200) that causes the first audio source to be signal processed based on the updated plurality of weights.

12. In the method of operating an electronic device (101; 200), An operation of identifying a first audio signal including multiple audio sources; An operation of identifying a first audio source corresponding to at least one type from the first audio signal; An operation of signal processing the first audio source based on a setting value corresponding to each of the at least one type; An operation of generating a second audio signal based on at least a portion of the remaining signal excluding the first audio source among the first audio signals and the signal-processed first audio source; and An operating method comprising an operation of outputting the second audio signal.

13. In paragraph 12, The operation of checking the above first audio source is: An operation of identifying at least one second audio source included in the first audio signal based on inputting the first audio signal to the first artificial intelligence model (1310); An operation of checking a type corresponding to each of the at least one second audio source; and An operation of identifying the first audio source corresponding to at least one type, The above first artificial intelligence model (1310) is An operating method, wherein the sum of the first output value of the remaining signal and the second output value of the at least one second audio source is learned to be equal to the third output value of the first audio signal.

14. In paragraph 12 or 13, The above signal processing operation is, Obtaining a signal-processed audio source by signal-processing the first audio source based on a filter and a first weight corresponding to at least one type, The operation of generating the second audio signal is as follows: A method of generating the second audio signal by mixing at least a portion of the remaining signal and the signal-processed audio source.

15. In a storage medium (130) storing computer-readable instructions, the instructions, when executed by at least one processor (120; 220) of an electronic device (101; 200), cause the electronic device (101; 200) to: Identifying a first audio signal including multiple audio sources, Identifying a first audio source corresponding to at least one type from the first audio signal, Processing the first audio source based on a setting value corresponding to each of the at least one type, Generating a second audio signal based on at least a portion of the remaining signal excluding the first audio source among the first audio signals and the signal-processed first audio source, A storage medium that causes the second audio signal to be output.

Citation Information

Patent Citations

  • Apparatus and method for providing sound

    KR1020150084192A

  • Method for processing audio signal and electronic device supporting the same

    KR1020170019738A

  • The guiding system for location of the acupuncture point and the method thereof

    KR1020230142059A

  • Road traffic safety facility with linear light device

    KR1020240176093A

  • Electronic device and method for controlling the same

    KR102446694B1