Audio signal processing method and electronic device supporting same
The collaborative method between an audio output device and a mobile device improves noise component removal and audio signal quality by synthesizing enhanced audio signals, addressing the limitations of existing audio output devices.
Patent Information
- Application Number
- PCT/KR2024/015306
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-18
- Filing Date
- 2024-10-08
- Publication Date
- 2025-06-12
AI Technical Summary
Existing audio output devices face challenges in perfectly distinguishing and removing noise components due to their physical size limitations, which affects the quality of audio signals.
A method and system that collaboratively use an audio output device and a mobile device to process audio signals. The mobile device obtains enhanced audio signals by removing noise from signals collected by the audio output device's microphones and synthesizes these signals to improve noise component removal.
The collaborative approach enhances the quality of audio signals by effectively removing noise components, resulting in clearer and more focused audio output.
Smart Images

Figure KR2024015306_12062025_PF_FP_ABST
Abstract
Description
Audio signal processing method and electronic device supporting the same
[0001] Embodiments disclosed in this document relate to an audio signal processing method and an electronic device supporting the same.
[0002] Various types of audio output devices (e.g., earphones, headsets) are available for use with mobile devices such as smartphones and tablet PCs. Audio output devices can be paired wirelessly with a mobile device via wireless communication or connected via a wired connection (e.g., an earphone jack).
[0003] Recently, audio output devices that fit comfortably in the user's ear canal with eartips inserted into the ear canal have been released. These audio output devices output audio signals from a mobile device through speakers, and during voice calls, they can collect audio signals through a microphone and transmit them to the mobile device.
[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0005] The aforementioned audio output device can acquire audio signals using a plurality of microphones. The plurality of microphones can include a first microphone configured to collect audio signals in a relatively narrow frequency band (e.g., a voice signal band) and a second microphone configured to collect audio signals in a relatively wide frequency band.
[0006] In this regard, the audio output device can collect a high-quality audio signal by controlling the operation of multiple microphones based on detected noise. The high-quality audio signal may include an audio signal with relatively little noise signal or an audio signal that emphasizes at least a portion of a relatively specific frequency band (e.g., a voice signal band).
[0007] For example, an audio output device can collect a high-quality audio signal by extracting noise components and voice components using an audio signal collected through a first microphone and an audio signal collected through a second microphone, and generating a signal having an opposite phase to the extracted noise components.
[0008] However, due to the performance limitations of audio output devices with physical size limitations, it is somewhat difficult to perfectly distinguish and remove noise components.
[0009] Accordingly, at least one example among various embodiments is intended to provide an audio signal processing method for improving noise component removal through collaboration between an audio output device and a mobile device, and an electronic device supporting the same.
[0010] According to various embodiments, a mobile device includes a communication circuit configured to establish communication with an audio output device, at least one processor, and a memory, wherein the memory may store instructions that, when executed by the at least one processor, cause the mobile device to: obtain a first enhanced audio signal from which noise is removed for a first audio signal obtained through a first audio collection device of the audio output device, obtain a second audio signal obtained through a second audio collection device of the audio output device, generate a second enhanced audio signal based on the second audio signal and first characteristic information about a speaker in a low-noise environment stored in the memory, and generate a composite audio signal by synthesizing the first enhanced audio signal and the second enhanced audio signal.
[0011] An audio signal processing system according to various embodiments includes a first electronic device and a second electronic device forming a communication with the first electronic device, wherein the first electronic device provides a first enhanced audio signal obtained by removing noise from a first audio signal obtained through a first audio collection device and a second audio signal obtained through a second audio collection device to the second electronic device, and the second electronic device (220) generates a second enhanced audio signal based on the second audio signal and first characteristic information about a speaker in a low-noise environment stored in the second electronic device, and generates a synthesized audio signal by synthesizing the first enhanced audio signal and the second enhanced audio signal.
[0012] A method of operating a mobile device according to various embodiments may include an operation of obtaining a first enhanced audio signal from which noise is removed for a first audio signal obtained through a first audio collection device of an audio output device, an operation of obtaining a second audio signal obtained through a second audio collection device of the audio output device, an operation of generating a second enhanced audio signal based on the second audio signal and first characteristic information about a speaker in a low-noise environment stored in the mobile device, and an operation of generating a synthesized audio signal by synthesizing the first enhanced audio signal and the second enhanced audio signal.
[0013] According to various embodiments, a computer-readable recording medium stores instructions for generating a synthetic audio signal by selecting a first method or a second method based on a type of an application running on a mobile device when activation of a microphone function is detected, wherein the first method includes an operation of obtaining a first enhanced audio signal from which noise is removed for a first audio signal obtained through a first audio collection device of an audio output device, obtaining a second audio signal obtained through a second audio collection device of the audio output device, generating a second enhanced audio signal based on the second audio signal and first feature information about a speaker in a low-noise environment stored in the mobile device, and synthesizing the first enhanced audio signal and the second enhanced audio signal to generate a synthetic audio signal, and wherein the second method includes an operation of converting the first audio signal obtained through the first audio collection device of the audio output device into text data, obtaining a second audio signal obtained through the second audio collection device of the audio output device, generating a third enhanced audio signal based on the text data and the first feature information, and synthesizing the first enhanced audio signal and the third enhanced audio signal to generate a synthetic audio signal. there is.
[0014] Electronic devices according to various embodiments disclosed in this document can collect high-quality audio signals with further noise components removed by using a plurality of audio collection devices.
[0015] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs from the description below.
[0016] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.
[0017] FIG. 2 is a diagram illustrating an audio signal processing system according to various embodiments.
[0018] FIG. 3A is a schematic diagram illustrating a configuration of a first electronic device according to various embodiments.
[0019] FIG. 3b is a schematic diagram illustrating a configuration of a second electronic device according to various embodiments.
[0020] FIG. 4 is a drawing for explaining the operation of a first electronic device according to various embodiments.
[0021] FIGS. 5 to 8 are drawings for explaining the operation of a second electronic device according to various embodiments.
[0022] FIGS. 9 to 11 are drawings for explaining the operation of a second electronic device according to various embodiments.
[0023] FIG. 12 is a flowchart illustrating the operation of a second electronic device according to various embodiments.
[0024] FIG. 13 is a flowchart illustrating the operation of a second electronic device according to various embodiments.
[0025] FIG. 14 is a flowchart illustrating the operation of a second electronic device according to various embodiments.
[0026] FIG. 15 is a flowchart illustrating the operation of a second electronic device according to various embodiments.
[0027] FIG. 16 is a flowchart illustrating the operation of an audio signal processing system according to various embodiments.
[0028] FIG. 17 is a flowchart illustrating the operation of an audio signal processing system according to various embodiments.
[0029] Hereinafter, various embodiments of this document are described with reference to the attached drawings. However, this is not intended to limit the technology described in this document to specific embodiments, and it should be understood that various modifications, equivalents, and / or alternatives of the embodiments of this document are included. In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0030]
[0031] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments.
[0032] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0033] The processor (120) may control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing, for example, software (e.g., a program (140)), and may perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculation, the processor (120) may store a command or data received from another component (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the command or data stored in the volatile memory (132), and store the resulting data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that may operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0034] The auxiliary processor (123) may control at least a part of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0035] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0036] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0037] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0038] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0039] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0040] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., an electronic device (102), a speaker, or headphones) directly or wirelessly connected to the electronic device (101).
[0041] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0042] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0043] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0044] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0045] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0046] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0047] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0048] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0049] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0050] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. According to some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0051] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0052] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0053] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service by itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0054]
[0055] According to one embodiment, the processor (120) (e.g., processing circuit) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing. The processor (120) may include at least one electrical circuit and may individually or collectively perform distributed processing of instructions (or programs, data) stored in the memory (130). The processor (120) may include a processor assembly including one or more processing circuits. The processor (120) may include any processing circuit operative to control the performance and operations of one or more components of the electronic device (101) (e.g., the memory (130), the display module (160), the sensor module (176) (e.g., a sensor), the camera module (180) (e.g., an image sensor), and / or the communication module (190) (e.g., a communication circuit)). For example, the processor (120) (e.g., an application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or chipset). For example, the processor (120) may be implemented as multiple cores (or at least one core circuit), multiple chips, or multiple chipsets. For example, the processor (120) may include one or more processing circuits. For example, the processor (120) may include one or more processing circuits configured to individually and / or collectively perform various functions of the present disclosure. As a non-limiting example, at least a portion of the processor (120) may be included in a first chip of the electronic device (101), and at least another portion of the processor (120) may be included in a second chip of the electronic device (101) that is different from the first chip of the electronic device (101).
[0056]
[0057] FIG. 2 is a diagram illustrating an audio signal processing system (20) according to various embodiments.
[0058] Referring to FIG. 2, an audio signal processing system (20) according to various embodiments may be configured with a first electronic device (210) (e.g., at least one first electronic device (210)) and a second electronic device (220) (e.g., at least one second electronic device (220)). According to one embodiment, each of the first electronic device (210) and the second electronic device (220) may be the same type of device as or different from the electronic device (101) illustrated in FIG. 1.
[0059] According to various embodiments, the first electronic device (210) may form a communication channel (e.g., a wired or wireless communication channel) with the second electronic device (220) and transmit an audio signal to the second electronic device (220) or receive an audio signal from the second electronic device (220).
[0060] For example, the first electronic device (210) may be a wireless earphone-type device (e.g., an audio output device) capable of establishing a communication channel with the second electronic device (220). However, this is merely an example, and various embodiments are not limited thereto. For example, the first electronic device (210) may also be a wearable device worn on a part of the body (e.g., a wrist or head).
[0061] According to various embodiments, the second electronic device (220) may form a communication channel with the first electronic device (210) and transmit an audio signal to the first electronic device (210) or receive an audio signal from the first electronic device (210).
[0062] For example, the second electronic device (220) may be a variety of devices (e.g., mobile devices) such as a portable terminal, terminal device, smartphone, tablet PC, or wearable electronic device that can form a communication channel with the first electronic device (210).
[0063] An audio signal processing system (20) according to various embodiments can generate a better quality audio signal by using audio signals collected through a plurality of audio collection devices. As mentioned above, a good quality audio signal may include an audio signal with relatively less noise signals or an audio signal that emphasizes at least a portion of a relatively specific frequency band (e.g., a voice signal band). For example, at least one of the plurality of audio collection devices may be configured as a microphone.
[0064] According to one embodiment, the audio signal processing system (20) can generate a good quality audio signal by synthesizing a first enhanced audio signal based on a first audio signal collected through at least one audio collection device and a second enhanced audio signal based on a second audio signal collected through at least one other audio collection device.
[0065] For example, the first enhanced audio signal is an audio signal from which noise has been removed from the first audio signal, and can provide naturalness similar to the speaker's original voice. Furthermore, the second enhanced audio signal is an audio signal that reflects the speaker's characteristics in a low-noise environment, and can provide the speaker's speech intention more clearly (e.g., clearly). The synthesis of the first enhanced audio signal and the second enhanced audio signal can increase the clarity of the audio signal.
[0066] In this regard, the first electronic device (210) according to various embodiments may include a plurality of audio collection devices. At least one of the plurality of audio collection devices may be used to generate a first enhanced audio signal, and at least another of the plurality of audio collection devices may be used to generate a second enhanced audio signal. For example, at least one of the plurality of audio collection devices may include a first audio collection device configured to collect an audio signal in a relatively wide frequency band. For example, at least another of the plurality of audio collection devices may include a second audio collection device configured to collect an audio signal in a relatively narrow frequency band (e.g., a voice signal band).
[0067] According to various embodiments, the first electronic device (210) may generate a first enhanced audio signal by removing noise from an audio signal collected through at least one audio collection device (e.g., the first audio collection device). According to one embodiment, the first electronic device (210) may obtain an audio signal of a first frequency band collected through at least one audio collection device as the first enhanced audio signal.
[0068] A second electronic device (220) according to various embodiments may generate a second enhanced audio signal that reflects speaker characteristics in a low-noise environment from an audio signal collected through at least one other audio collection device (e.g., a second audio collection device) provided in the first electronic device (210). The speaker characteristics in a low-noise environment may include speaker characteristics extracted from a first audio signal containing noise below a certain level. The generation of the second enhanced audio signal will be described in detail with reference to FIG. 5 below.
[0069] According to one embodiment, the second electronic device (220) can generate a good quality audio signal by synthesizing a first enhanced audio signal with noise removed and a second enhanced audio signal reflecting speaker characteristics in a low-noise environment.
[0070] As described above, the functions of the audio signal processing system (20) according to various embodiments can be performed through collaboration between the first electronic device (210) and the second electronic device (220).
[0071] However, this is merely an example, and various embodiments are not limited thereto. For example, the functions of the audio signal processing system (20) according to various embodiments may be independently applied to the first electronic device (210). For example, the generation of the first enhanced audio signal, the generation of the second enhanced audio signal, and the synthesis thereof may be performed by the first electronic device (210).
[0072] In addition, similarly, the function of the aforementioned audio signal processing system (20) can be independently applied to the second electronic device (220). For example, the generation of the first enhanced audio signal, the generation of the second enhanced audio signal, and the synthesis thereof can be performed by the second electronic device (220). In this regard, the second electronic device (220) can have a plurality of audio collection devices or can generate the first enhanced audio signal and the second enhanced audio signal based on the first audio signal and the second audio signal acquired by the first electronic device (210).
[0073] Depending on the embodiment, some of the functions of the audio signal processing system (20) according to various embodiments may be performed by a third electronic device (e.g., server (108) of FIG. 1) other than the first electronic device (210) and the second electronic device (220).
[0074] The specific operation of the audio signal processing system (20) according to the various embodiments described above will be described in detail with reference to FIGS. 3A to 17 below. In addition, at least one embodiment (e.g., at least a part of the embodiment) among the various embodiments described with reference to FIGS. 3A to 17 below may be combined with another embodiment (e.g., at least a part of the other embodiment).
[0075]
[0076] FIG. 3A is a diagram schematically illustrating the configuration of a first electronic device (210) according to various embodiments. FIG. 4 is a diagram for explaining the operation of the first electronic device (210) according to various embodiments.
[0077] Referring to FIGS. 2, 3a, and 4, a first electronic device (210) of an earphone type according to various embodiments may be configured with a first audio collection device (311), a second audio collection device (312), a first speaker (313), a first audio processing module (314), a first memory (315), a first communication circuit (316), and a first processor (317).
[0078] The components of the first electronic device (210) described above are one embodiment, and various embodiments are not limited thereto. For example, the first electronic device (210) may be implemented to have more or fewer components than the components illustrated in FIG. 3A. For example, at least some of the components of the electronic device (101) illustrated in FIG. 1 (e.g., the input module (150), the display module (160), or the sensor module (176)) may be included as a component of the first electronic device (210). In addition, at least one component of the components of the first electronic device (210) illustrated in FIG. 3A (e.g., the first audio processing module (314)) may be integrated with another component (e.g., the processor (317)).
[0079] According to various embodiments, the first audio collection device (311) may be configured to collect audio signals in a relatively wide first frequency band (e.g., at least a portion of a range of approximately 1 Hz to 20 kHz). In one embodiment, the first audio collection device (311) may include a microphone configured to collect audio signals in the first frequency band.
[0080] For example, the first audio collection device (311) may be designed to collect signals across the entire frequency band that can be collected in relation to voice. For example, the first audio collection device (311) may be equipped to collect external audio signals while the first electronic device (210) is worn on the user's ear.
[0081] According to various embodiments, the second audio collection device (312) may have different characteristics than the first audio collection device (311). In one embodiment, the second audio collection device (312) may be configured to collect audio signals in a second frequency band narrower than the first frequency band (e.g., at least a portion of a range of approximately 0.1 kHz to 3 kHz). In one embodiment, the second audio collection device (312) may be configured to collect signals transmitted inside the outer ear when the first electronic device (210) is worn on the user's ear.
[0082] According to an embodiment, when a user wears the first electronic device (210) and speaks, at least a portion of the vibration caused by the speech is transmitted through the user's skin, muscles, bones, etc., and the transmitted vibration can be collected as an audio signal by the second audio collection device (312) inside the ear. For example, the second audio collection device (312) may include a sensor (e.g., a bone conduction sensor) having a performance capable of collecting an audio signal that is relatively better (or higher than a specified quality value) than the first audio collection device (311). However, this is merely an example, and various embodiments are not limited thereto. For example, an in-ear microphone or a bone conduction microphone may be used as the second audio collection device (312), and according to an embodiment, a microphone implemented with micro-electro-mechanical system technology may be used as the second audio collection device (312).
[0083] According to various embodiments, the first speaker (313) may output audio signals to the outside of the first electronic device (210). According to one embodiment, the first speaker (313) may be used for general purposes, such as multimedia playback or recording playback, and may also be used for receiving incoming calls (e.g., as a receiver). For example, the first speaker (313) may be the audio output module (155) illustrated in FIG. 1.
[0084] According to various embodiments, the first audio processing module (314) may support the audio signal processing function of the first electronic device (210).
[0085] According to one embodiment, the first audio processing module (314) may perform functions related to generating a first enhanced audio signal (415), as illustrated in FIG. 4. As described above, the first enhanced audio signal (415) may provide naturalness similar to the speaker's original voice.
[0086] For example, the first audio processing module (314) may perform noise suppression as part of generating the first enhanced audio signal (415). The first audio processing module (314) may generate the first enhanced audio signal (415) by performing noise processing (413) on the first audio signal (411) collected through the first audio collection device (311).
[0087] According to an embodiment, the first audio processing module (314) may optionally perform noise processing (413) on the first audio signal (411) collected from the first audio collection device (311). For example, if the magnitude of noise included in the first audio signal (411) is greater than or equal to a specified value, the first audio processing module (314) may perform noise processing (413) on the first audio signal (411). In addition, if the magnitude of noise included in the first audio signal (411) is less than the specified value, the first audio processing module (314) may omit noise processing (413) on the first audio signal (411). According to an embodiment, the first enhanced audio signal (415) generated by the first audio processing module (314) may be provided to the second electronic device (220).
[0088] According to one embodiment, the first audio processing module (314) may assist in functions related to generating a second enhanced audio signal (e.g., the second enhanced audio signal (509) of FIG. 5). As described above, the second enhanced audio signal (415) may provide a clearer representation of the speaker's speech intent.
[0089] For example, the first audio processing module (314) may provide a second audio signal (421) collected from the second audio collection device (312), as illustrated in FIG. 4, to the second electronic device (220). This second audio signal (421) may be used to generate a second enhanced audio signal. For example, the second electronic device (220) may generate the second enhanced audio signal by reflecting speaker characteristics in a low-noise environment to the second audio signal (421). The generation of the second enhanced audio signal will be described in detail with reference to FIGS. 3b, 5, 7, and 8 below.
[0090] According to one embodiment, the first audio processing module (314) may extract speaker features (437) (e.g., feature vectors) from the first audio signal (411) as part of an operation that assists in the function related to the generation of the second enhanced audio signal. The speaker features (437) may be unique components for the speaker. For example, the first audio processing module (314) may extract the speaker features (437) by converting the time domain-based first audio signal (411) into a signal on the frequency domain and differently transforming the frequency energy of the converted signal. For example, the speaker features (437) may be extracted based on Mel-Frequency Cepstral Coefficients or Filter Bank Energy, but are not limited thereto, and the speaker features (437) may be extracted from the first audio signal (411) in various ways. According to an embodiment, the first audio processing module (314) may extract pitch components, timbre, voice length, loudness, formant, or a combination thereof as speaker features (437).
[0091] According to an embodiment, the first audio processing module (314) may perform a noise evaluation (431) on the first audio signal (411), as illustrated in FIG. 4, and selectively extract speaker features (437) from the first audio signal (411) based on the noise evaluation result (433). For example, if the first audio signal (411) contains a noise component above a certain level, the extracted speaker features may also contain a certain level of noise components, making them somewhat unsuitable for use as speaker features in a low-noise environment.
[0092] In this regard, when the size of the noise included in the first audio signal (411) is less than a specified value (e.g., a low-noise environment in which noise is generated at a level that is practically imperceptible to a human), the first audio processing module (314) may extract a speaker feature (437) based on the first audio signal (411). In addition, when the size of the noise included in the first audio signal (411) is greater than a specified value (e.g., a noisy environment in which noise is generated at a level that can be perceived by a human), the first audio processing module (314) may omit the speaker feature extraction (437). According to an embodiment, the speaker feature (437) extracted in the low-noise environment may be provided to the second electronic device (220), which may be used to generate a second enhanced audio signal.
[0093] In this regard, the first audio processing module (314) may provide the extracted speaker feature (437) to the second electronic device (220) together with the first enhanced audio signal (415) and the second audio signal (421). However, this is merely an example, and various embodiments are not limited thereto. For example, the extracted speaker feature (437) may be provided to the second electronic device (220) together with the first enhanced audio signal (415) or may be provided to the second electronic device (220) together with the second audio signal (421). In addition, the extracted speaker feature (437) may be provided separately from the first enhanced audio signal (415) and the second audio signal (421).
[0094] As described above, the first audio processing module (314) may perform a noise processing operation (413), a noise evaluation operation (431), and a speaker feature extraction operation (435) when processing an audio signal. In this regard, the first audio processing module (314) according to various embodiments may be configured with at least one module (e.g., a noise processing module, a noise evaluation module, a speaker feature extraction module) related to the noise processing operation (413), the noise evaluation operation (431), and the speaker feature extraction operation (435).
[0095] According to various embodiments, the first memory (315) may store various data used by components of the first electronic device (210). According to one embodiment, the first memory (315) may store instructions that cause the first electronic device (210) to perform functions (e.g., operations). For example, the first memory (315) may be the memory (130) illustrated in FIG. 1.
[0096] According to various embodiments, the first communication circuit (316) may support the communication function of the first electronic device (210). In one embodiment, the first communication circuit (316) may be a device including hardware and software for transmitting and receiving signals (e.g., commands or data) between the first electronic device (210) and the second electronic device (220). For example, the first communication circuit (316) may be the communication module (190) illustrated in FIG. 1.
[0097] According to various embodiments, the first processor (317) may be operatively connected to the first audio collection device (311), the second audio collection device (312), the speaker (313), the first audio processing module (314), the first memory (315), and the first communication circuit (316), and may control various components (e.g., hardware or software components) of the first electronic device (210). According to one embodiment, the first processor (317) may control functions related to audio signal processing of the first electronic device (210). For example, the first processor (317) may include circuitry such as a central processing unit (CPU), a microprocessor unit (MPU), an application processor (AP), a communication processor (CP), a System On Chip (SoC), and an Integrated Circuit (IC). For example, the first processor (317) may be the processor (120) illustrated in FIG. 1.
[0098] As described above, the first electronic device (210) according to various embodiments may include one first processor (317). In this case, the first processor (317) may execute instructions stored in the first memory (315) to perform functions (e.g., operations) of the first electronic device (210).
[0099] However, this is merely an example, and various embodiments are not limited thereto. For example, a first electronic device (210) according to various embodiments may include a plurality of first processors (317). In this case, some of the plurality of first processors (317) may execute instructions stored in the first memory (315) to perform some functions of the first electronic device (210), and other of the plurality of first processors (317) may execute instructions stored in the first memory (315) to perform other functions of the first electronic device (210).
[0100] According to an embodiment, a first electronic device (210) according to various embodiments may include a plurality of first memories (315). In this case, some of the plurality of first memories (315) may store instructions for performing some functions of the first electronic device (210), and other of the plurality of first memories (315) may store instructions for performing other functions of the first electronic device (210).
[0101]
[0102] FIG. 3b is a diagram schematically illustrating the configuration of a second electronic device (220) according to various embodiments. FIGS. 5 to 8 are diagrams for explaining the operation of the second electronic device (220) according to various embodiments.
[0103] Referring to FIGS. 2, 3b, and 5 to 8, a second electronic device (220) according to various embodiments may be configured with a second communication circuit (321), a second audio processing module (322), a second memory (323), a second speaker (324), and a second processor (325).
[0104] The components of the second electronic device (220) described above are only one embodiment, and various embodiments are not limited thereto. For example, the second electronic device (220) may be implemented to have more or fewer components than the components illustrated in FIG. 3B. For example, at least some of the components of the electronic device (101) illustrated in FIG. 1 (e.g., the input module (150), the display module (160), or the sensor module (176)) may be included as a component of the second electronic device (220). In addition, at least one component of the components of the second electronic device (220) illustrated in FIG. 3B (e.g., the second audio processing module (322)) may be integrated with another component (e.g., the second processor (325)).
[0105] According to various embodiments, the second communication circuit (321) may support the communication function of the second electronic device (220). According to one embodiment, the second communication circuit (321) may be a device including hardware and software for transmitting and receiving signals (e.g., commands or data) between the first electronic device (210) and the second electronic device (220). For example, the second communication circuit (321) may be the communication module (190) illustrated in FIG. 1.
[0106] According to various embodiments, the second audio processing module (322) may support an audio signal processing function of the second electronic device (220). For example, the second audio processing module (322) may generate an audio signal of better quality by using audio signals collected through a plurality of audio collection devices provided in the first electronic device (210).
[0107] According to one embodiment, the second audio processing module (322) can generate a good quality audio signal by synthesizing the first enhanced audio signal (415) and the second enhanced audio signal (509), as illustrated in FIG. 5. As described above, the first enhanced audio signal (415) is obtained (e.g., generated) from the first electronic device (210) and can provide naturalness like the speaker's original voice. In addition, the second enhanced audio signal (415) is obtained by the second electronic device (220) and can provide the speaker's speech intention more clearly.
[0108] In this regard, the second audio processing module (322) may perform a function related to generating a second enhanced audio signal (509). As illustrated in FIG. 5, the second audio processing module (322) may generate the second enhanced audio signal (509) by reflecting speaker characteristics (437) into the second audio signal (421).
[0109] For example, the second audio processing module (322) may store (501) speaker features (437) provided from the first electronic device (210) in the second memory (323). Speaker features acquired over a certain period of time (e.g., from 30 days ago to the present) may be stored in the second memory (323) as a single feature vector. For example, the second audio processing module (322) may reflect previously acquired and stored speaker features (437) in the currently collected second audio signal (421). According to an embodiment, the second audio processing module (322) may use an artificial intelligence model that inputs the second audio signal (421) and the speaker features (437) and outputs a second enhanced audio signal (509).
[0110] According to an embodiment, when an event related to an audio signal collection request (e.g., execution of a recording function, execution of a call function, or execution of a video storage function) is detected, the second audio processing module (322) may reflect a first speaker feature (503) corresponding to the speaker among the stored speaker features into the second audio signal (421). For example, the second audio processing module (322) may generate a second enhanced audio signal (509) by reflecting a speaker feature (407) extracted in a low-noise environment into the second audio signal (421).
[0111] According to an embodiment, the second audio processing module (322) may output a synthesized audio signal (513) obtained by synthesizing (511) the first enhanced audio signal (415) and the second enhanced audio signal (509). This synthesized audio signal (513) may be output through the second speaker (323), stored in the second memory (324), or provided to another electronic device (e.g., the first electronic device (210) or the server (108)).
[0112] As described above, as part of the operation of generating a good quality audio signal, the second audio processing module (322) may synthesize the second enhanced audio signal (509) with the first enhanced audio signal (415). However, although the second enhanced audio signal (509) may provide a clearer representation of the speaker's speech intention, it may sound somewhat less like a human voice and more like a machine-like sound.
[0113] In this regard, the second audio processing module (322) according to various embodiments may adjust the synthesis ratio of the first enhanced audio signal (415) and the second enhanced audio signal (509) as part of an operation of generating a good quality audio signal.
[0114] For example, the second audio processing module (322) may perform a quality evaluation (601) on the first enhanced audio signal (415) provided from the first electronic device (210), as illustrated in FIG. 6, and determine (605) a first ratio for the first enhanced audio signal (415) and a second ratio for the second enhanced audio signal (509) based on the evaluation result (603).
[0115] For example, if the quality of the first enhanced audio signal (415) is higher than a specified value, the second audio processing module (322) can synthesize a first enhanced audio signal (607) with a relatively high first ratio and a second enhanced audio signal (609) with a relatively low second ratio. In this case, the synthesized audio signal (513) can eliminate the problem of not sounding like a human voice or a robotic voice, and can provide a natural sound like the speaker's original voice and a clearer intention of speech.
[0116] Additionally, if the quality of the first enhanced audio signal is below a specified value, the second audio processing module (322) may synthesize a first audio signal (607) having a relatively low first ratio and a second enhanced audio signal (609) having a relatively high second ratio. In this case, the synthesized audio signal (513) may not sound like a human voice or may sound like a robot, but may provide a clearer sense of speech intent.
[0117] According to an embodiment, a Mean Opinion Score (MOS) algorithm may be used to evaluate the quality of an audio signal (415). The MOS algorithm may express the recognition quality of human speech as a specified number (e.g., a number ranging from 1 indicating the lowest recognition quality to 5 indicating the highest recognition quality). However, this is merely an example, and various embodiments are not limited thereto. For example, other algorithms may be used in addition to the MOS algorithm to evaluate the quality (601) of the audio signal (415).
[0118] In this regard, when the quality evaluation (601) for the audio signal (415) corresponds to a relatively good value 4, the second audio processing module (322) may calculate a percentage (e.g., 80%) for the result (603) of the quality evaluation (601) (e.g., value 4) and, based on this, determine (605) a first ratio for the first enhanced audio signal (415) and a second ratio for the second enhanced audio signal (509).
[0119] For example, if the quality evaluation (601) for the first enhanced audio signal (415) corresponds to a relatively good value of 4, the second audio processing module (322) may output a synthetic audio signal (513) by using 80% of the first enhanced audio signal (415) and 20% of the second enhanced audio signal (509). In this case, the synthetic audio signal (513) may emphasize the naturalness of the speaker's original voice more than the speaker's intention to speak.
[0120] In addition, the second audio processing module (322) may output a synthetic audio signal (513) by using 20% of the first enhanced audio signal (415) and 80% of the second enhanced audio signal (509) when the quality evaluation (601) for the first enhanced audio signal (415) corresponds to a relatively poor value of 2. In this case, the synthetic audio signal (513) may emphasize the speaker's speech intention more than the speaker's natural voice.
[0121] Additionally or optionally, the second audio processing module (322) according to various embodiments may, as part of the operation of generating the second enhanced audio signal (509), extract a speaking style associated with the speaker and reflect it in the second audio signal (421).
[0122] In this regard, the second audio processing module (312) according to various embodiments may obtain (701) a speaking style from the first enhanced audio signal (415) and reflect (507) the same in the second audio signal (421), as illustrated in FIG. 7. The speaking style may be associated with the speaker's current speaking state. For example, the speaking style may include at least one of the speaker's emotion (e.g., happiness, sadness, anger, disgust, fear, frustration, excitement, or depression), speaking rate, accent, or speaking volume. For example, the second audio processing module (322) may obtain the second enhanced audio signal (509) by using the speaking style (703), the first speaker feature (503), and the second audio signal (421). According to an embodiment, the second audio processing module (322) may use an artificial intelligence model that inputs the collected second audio signal (421), speaker characteristics (437) and speech style (703) and outputs a second enhanced audio signal (509).
[0123] Additionally or optionally, the second audio processing module (322) according to various embodiments may, as part of the operation of generating the second enhanced audio signal (509), adjust the reflection ratio for the speaking style (703) associated with the speaker and the characteristics (503) of the first speaker.
[0124] For example, the second audio processing module (322) may perform a quality evaluation (801) on the first enhanced audio signal (415) provided from the first electronic device (210), and determine (805) a first ratio for the speaking style (703) and a second ratio for the first speaker feature (503) based on the quality evaluation result (803).
[0125] For example, if the quality of the first enhanced audio signal (415) is greater than a specified value, the second audio processing module (322) may utilize a relatively high first rate of speech style (807) and a relatively low second rate of first speaker features (809). In this case, the second enhanced audio signal (509) may emphasize the speaker's current speech state more than the speaker's speech intention.
[0126] Additionally, if the quality of the first enhanced audio signal (415) is below a specified value, the second audio processing module (322) may utilize a first speech style (807) with a relatively low first ratio and a first speaker feature (809) with a relatively high second ratio. In this case, the second enhanced audio signal (509) may emphasize the speaker's speech intention more than the speaker's current speech state.
[0127] According to an embodiment, the second audio processing module (322) may use a Mean Opinion Score (MOS) algorithm to determine (805) the first ratio for the speech style (703) and the second ratio for the first speaker feature (503). For example, the second audio processing module (322) may increase the reflection ratio of the speech style (703) and decrease the reflection ratio of the first speaker feature (503) as the quality assessment (801) for the first enhanced audio signal (415) becomes better.
[0128] As described above, the second audio processing module (322) may perform a speaker feature storage operation (501), a speaker feature reflection operation (507), an audio signal synthesis operation (511), a quality evaluation operation (601), a synthesis ratio determination operation (605), a speech style acquisition operation (701), a quality evaluation operation (801), and a reflection ratio determination operation (805). In this regard, the second audio processing module (322) may be configured with at least one module related to the speaker feature storage operation (501), the speaker feature reflection operation (507), the audio signal synthesis operation (511), a quality evaluation operation (601), a synthesis ratio determination operation (605), a speech style acquisition operation (701), a quality evaluation operation (801), and a reflection ratio determination operation (805).
[0129] According to various embodiments, the second memory (323) may store various data used by components of the second electronic device (220). According to one embodiment, the second memory (323) may store instructions that cause the second electronic device (220) to perform functions (e.g., operations). For example, the second memory (323) may be the memory (130) illustrated in FIG. 1.
[0130] According to various embodiments, the second speaker (324) can output audio signals to the outside of the second electronic device (220). According to one embodiment, the second speaker (324) can be used for general purposes, such as multimedia playback or recording playback, and can also be used for receiving incoming calls (e.g., as a receiver). For example, the second speaker (324) can be the audio output module (155) illustrated in FIG. 1.
[0131] According to various embodiments, the second processor (325) may be operatively connected to the second communication circuit (321), the second audio processing module (322), the second memory (323), and the second speaker (324), and may control various components (e.g., hardware or software components) of the second electronic device (220). For example, the second processor (325) may include circuitry such as a CPU, an MPU, an AP, a CP, an SoC, and an IC. For example, the second processor (325) may be the processor (120) illustrated in FIG. 1.
[0132] As described above, the second electronic device (220) according to various embodiments may include one second processor (325). In this case, the second processor (325) may execute instructions stored in the second memory (323) to perform functions (e.g., operations) of the second electronic device (220).
[0133] However, this is merely an example, and various embodiments are not limited thereto. For example, a second electronic device (220) according to various embodiments may include a plurality of second processors (325). In this case, some of the plurality of second processors (325) may execute instructions stored in the second memory (323) to perform some functions of the second electronic device (220), and other some of the plurality of second processors (325) may execute instructions stored in the second memory (323) to perform other functions of the second electronic device (220).
[0134] According to an embodiment, a second electronic device (220) according to various embodiments may include a plurality of second memories (323). In this case, instructions for performing some functions of the second electronic device (220) may be stored in some of the plurality of second memories (323), and instructions for performing other functions of the second electronic device (220) may be stored in other of the plurality of second memories (323).
[0135] As described above, the audio signal processing system (20) according to various embodiments can generate a better quality audio signal by using a first enhanced audio signal (415) obtained by removing noise from an audio signal (411) collected through a first audio collection device (311). This first enhanced audio signal (415) can enable generation of a better quality audio signal without delay.
[0136] An audio signal processing system (20) according to various embodiments can convert a first enhanced audio signal (415) into text and use it to generate a high-quality audio signal. This text conversion can further improve the quality of the audio signal. This will be described in detail with reference to FIGS. 9 to 11 below.
[0137]
[0138] FIGS. 9 to 11 are drawings for explaining the operation of a second electronic device (220) according to various embodiments.
[0139] Referring to FIGS. 2, 3b and 9, a second audio processing module (322) according to various embodiments can generate text data (901) based on a first enhanced audio signal (415) and generate a synthetic audio signal (915) based thereon.
[0140] According to one embodiment, the second audio processing module (322) can reflect (905) a first speaker feature (903), which is a unique component for a speaker, into text data (902), and generate (909) a third enhanced audio signal based on the text data (907) in which the first speaker feature (903) is reflected.
[0141] For example, the second audio processing module (322) can generate a third enhanced audio signal (911) that significantly removes noise and provides the speaker's speech intention more clearly by reflecting the first speaker feature (903) into the converted text data (902). According to an embodiment, the second audio processing module (322) can use an artificial intelligence model that inputs the first speaker feature (903) and the text data (902) and outputs the third enhanced audio signal (911).
[0142] According to one embodiment, the second audio processing module (322) may synthesize (913) the first enhanced audio signal (415) and the third enhanced audio signal (911) and output a synthesized audio signal (915). This synthesized audio signal (915) may be output through the second speaker (323), stored in the second memory (324), or provided to another electronic device (e.g., the first electronic device (210) or the server (108)).
[0143] Additionally or optionally, the second audio processing module (322) according to various embodiments may, as part of the operation of generating the third enhanced audio signal (911), extract a speaking style associated with the speaker and reflect it in the text data (907).
[0144] In this regard, the second audio processing module (312) according to various embodiments may obtain (1001) a speech style from the first enhanced audio signal (415), as illustrated in FIG. 10. The speech style (1003) may be associated with a current speech state of the speaker. For example, at least one of the speaker's emotion (e.g., happiness, sadness, anger, disgust, fear, frustration, excitement, or depression), speech rate, accent, or speech volume may be obtained as the speech style (1003). For example, the second audio processing module (322) may obtain a third enhanced audio signal (911) using text data (902), the first speaker features (903), and the speech style (1003). According to an embodiment, the second audio processing module (322) may utilize an artificial intelligence model that inputs text data (902), first speaker features (903) and speech style (1003) and outputs a third enhanced audio signal (911).
[0145] Additionally or optionally, the second audio processing module (322) according to various embodiments may adjust the reflection ratio for the first speaker's characteristics (903) and speech style (1003) as part of the operation of generating the third enhanced audio signal (911).
[0146] For example, the second audio processing module (322) may perform a quality evaluation (1101) on the first enhanced audio signal (415) provided from the first electronic device (210), as illustrated in FIG. 11, and determine (1105) a first ratio for the speaking style (1003) and a second ratio for the first speaker feature (903) based on the evaluation result (1103).
[0147] For example, if the quality of the first enhanced audio signal (415) is greater than a specified value, the second audio processing module (322) may utilize a relatively high first rate of speech style (1107) and a relatively low second rate of first speaker features (1109). In this case, the third enhanced audio signal (911) may emphasize the speaker's current speech state more than the speaker's speech intention.
[0148] Additionally, if the quality (415) of the first enhanced audio signal is below a specified value, the second audio processing module (322) may utilize a first speech style (1007) with a relatively low first ratio and a first speaker feature (1009) with a relatively high second ratio. In this case, the third enhanced audio signal (911) may emphasize the speaker's speech intention more than the speaker's current speech state.
[0149] As described above, the second audio processing module (322) can perform a text data generation operation (901), a speaker feature reflection operation (905), a third enhanced audio signal generation operation (909), a speech style acquisition operation (1001), a quality evaluation operation (1101), and a reflection ratio determination operation (1105). In this regard, the second audio processing module (322) may be configured with at least one module related to the text data generation operation (901), the speaker feature reflection operation (905), the third enhanced audio signal generation operation (909), the speech style acquisition operation (1001), the quality evaluation operation (1101), and the reflection ratio determination operation (1105).
[0150]
[0151] Figure 12 is a flowchart illustrating the operation of a second electronic device (220) according to various embodiments. While the operations in the following embodiments may be performed sequentially, they are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.
[0152] Referring to FIG. 12, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) according to various embodiments may, in operation 1210, obtain a first enhanced audio signal (415) from which noise is removed from an audio signal obtained through an audio collection device. The first enhanced audio signal (415) may be generated using a first audio signal (411) collected through at least one (e.g., the first audio collection device (311)) of a plurality of audio collection devices (e.g., the first audio collection device (311) and the second audio collection device (312)). According to one embodiment, the first enhanced audio signal (415) is obtained (e.g., generated) from the first electronic device (210) and may provide naturalness like the original voice of a speaker.
[0153] According to various embodiments, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) may, in operation 1220, obtain a second enhanced audio signal (509) in which characteristics of a low-noise environment are reflected in an audio signal obtained through an audio collection device. The second enhanced audio signal (415) may be generated using a second audio signal (421) collected through at least another one (e.g., the second audio collection device (312)) of a plurality of audio collection devices (e.g., the first audio collection device (311) and the second audio collection device (312)). According to one embodiment, the second enhanced audio signal (415) is obtained by the second electronic device (220) and may provide naturalness like the speaker's original voice. For example, the second electronic device (220) can generate a second enhanced audio signal (509) by reflecting speaker features (407) extracted in a low-noise environment into a second audio signal (421).
[0154] According to various embodiments, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) may synthesize a first enhanced audio signal (415) and a second enhanced audio signal (509) in operation 1230. This synthesis of the first enhanced audio signal (415) and the second enhanced audio signal (509) may increase clarity of the audio signal. Depending on the embodiment, the synthesized audio signal (513) may be output through a second speaker (323), stored in a second memory (324), or provided to another electronic device (e.g., the first electronic device (210) or a server (108)).
[0155]
[0156] Figure 13 is a flowchart illustrating the operation of a second electronic device (220) according to various embodiments. While the operations in the following embodiments may be performed sequentially, they are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.
[0157] Referring to FIGS. 12 and 13, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) according to various embodiments may obtain a first enhanced audio signal (415) from which noise is removed from an audio signal obtained through an audio collection device in operation 1210.
[0158] According to various embodiments, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) may obtain speaker characteristics (503) and speech styles (703) in operation 1310. Speaker characteristics (503) may be associated with unique components of the speaker, such as pitch, timbre, length of voice, and loudness. In addition, speech styles (703) may be associated with current speech states of the speaker, such as emotions (e.g., happiness, sadness, anger, disgust, fear, frustration, excitement, or depression), speech rate, accent, or speech volume. According to one embodiment, speaker characteristics (503) may be obtained from a first audio signal (411) containing noise less than a specified value, and speech styles (703) may be obtained from a first enhanced audio signal (415).
[0159] According to various embodiments, the second electronic device (220) (e.g., the second audio processing module (322) or the second processor (325)) may determine, in operation 1320, a reflection ratio for the speaker feature (503) and the speech style (703) based on the quality of the first enhanced audio signal (415). According to one embodiment, the second electronic device (220) may increase the reflection ratio for the speech style (703) and decrease the reflection ratio for the speaker feature (503) as the quality of the first enhanced audio signal (415) improves.
[0160] According to various embodiments, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) may generate a second enhanced audio signal (509) reflecting the speaker characteristics (503) and the speech style (703) based on the reflection ratio in operation 1330. According to one embodiment, the second electronic device (220) may generate the second enhanced audio signal (509) using the speech style (807) of the first ratio, the speaker characteristics (509) of the second ratio, and the second audio signal (421).
[0161] A second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) according to various embodiments may synthesize a first enhanced audio signal (415) and a second enhanced audio signal (509) in operation 1230.
[0162]
[0163] Figure 14 is a flowchart illustrating the operation of a second electronic device (220) according to various embodiments. While the operations in the following embodiments may be performed sequentially, they are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.
[0164] Referring to FIGS. 12 and 14, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) according to various embodiments may, in operation 1210, obtain a first enhanced audio signal (415) from which noise is removed from an audio signal obtained through an audio collection device.
[0165] A second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) according to various embodiments may, in operation 1220, obtain a second enhanced audio signal (509) in which the characteristics of a low-noise environment are reflected in an audio signal obtained through an audio collection device.
[0166] According to various embodiments, the second electronic device (220) (e.g., the second audio processing module (322) or the second processor (325)) may determine a synthesis ratio of the first enhanced audio signal (415) and the second enhanced audio signal (509) based on the quality of the first enhanced audio signal (415) in operation 1410. According to one embodiment, the second electronic device (220) may increase the synthesis ratio of the first enhanced audio signal (415) and decrease the synthesis ratio of the second enhanced audio signal (509) as the quality of the first enhanced audio signal (415) is better.
[0167] According to various embodiments, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) may synthesize a first enhanced audio signal (415) and a second enhanced audio signal (509) based on a synthesis ratio in operation 1420. According to one embodiment, the second electronic device (220) may output a synthesized audio signal (513) using a first enhanced audio signal (607) of a first ratio and a second enhanced audio signal (609) of a second ratio.
[0168]
[0169] Figure 15 is a flowchart illustrating the operation of a second electronic device (220) according to various embodiments. While the operations in the following embodiments may be performed sequentially, they are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.
[0170] Referring to FIGS. 12 to 15, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) according to various embodiments may, in operation 1210, obtain a first enhanced audio signal (415) from which noise is removed from an audio signal obtained through an audio collection device.
[0171] According to various embodiments, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) may acquire speaker characteristics (503) and speech styles (703) in operation 1310. Speaker characteristics (503) may be associated with unique components of the speaker, such as pitch, timbre, length of voice, and loudness. In addition, speech styles (703) may be associated with current speech states of the speaker, such as emotions (e.g., happiness, sadness, anger, disgust, fear, frustration, excitement, or depression), speech rate, accent, or speech volume.
[0172] According to various embodiments, the second electronic device (220) (e.g., the second audio processing module (322) or the second processor (325)) may determine, in operation 1320, a reflection ratio for the speaker feature (503) and the speech style (703) based on the quality of the first enhanced audio signal (415). According to one embodiment, the second electronic device (220) may increase the reflection ratio for the speech style (703) and decrease the reflection ratio for the speaker feature (503) as the quality of the first enhanced audio signal (415) improves.
[0173] A second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) according to various embodiments may generate a second enhanced audio signal (509) reflecting speaker characteristics (503) and speech style (703) based on a reflection ratio at operation 1330.
[0174] According to various embodiments, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) may determine a synthesis ratio of the first enhanced audio signal (415) and the second enhanced audio signal (509) based on the quality of the first enhanced audio signal (415) in operation 1410.
[0175] According to various embodiments, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) may synthesize a first enhanced audio signal (415) and a second enhanced audio signal (509) based on a synthesis ratio in operation 1420. According to one embodiment, the second electronic device (220) may output a synthesized audio signal (513) using a first enhanced audio signal (607) of a first ratio and a second enhanced audio signal (609) of a second ratio.
[0176]
[0177] Figure 16 is a flowchart illustrating the operation of an audio signal processing system (20) according to various embodiments. While the operations in the following embodiments may be performed sequentially, they are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.
[0178] Referring to FIG. 16, a first electronic device (210) (e.g., a first audio processing module (314) or a first processor (317)) according to various embodiments may obtain a first audio signal (411) in operation 1601. According to one embodiment, the first electronic device (210) may obtain the first audio signal (411) through a first audio collection device (311).
[0179] According to various embodiments, the first electronic device (210) (e.g., the first audio processing module (314) or the first processor (317)) may, in operation 1603, evaluate the quality of the first audio signal (411). According to one embodiment, the first electronic device (210) may determine whether the first audio signal (411) contains a noise component below a certain level.
[0180] According to various embodiments, when the quality of the first audio signal (411) does not satisfy a specified condition (e.g., when it contains noise components exceeding a certain level), the first electronic device (210) (e.g., the first audio processing module (314) or the first processor (317)) may generate a first enhanced audio signal (415) in operation 1609.
[0181] According to various embodiments, when the quality of the first audio signal (411) satisfies a specified condition (e.g., when it contains noise components below a certain level), the first electronic device (210) (e.g., the first audio processing module (314) or the first processor (317)) may transmit first information including speaker characteristics (437) to the second electronic device (220) in operation 1607.
[0182] According to various embodiments, the second electronic device (220) (e.g., the second audio processing module (322) or the second processor (325)) may, in operation 1611, compare the first information with the stored second information. For example, the second electronic device (220) may store speaker characteristics previously provided from the first electronic device (210) as the second information. According to one embodiment, the second electronic device (220) may determine whether the first information and the second information include the same speaker characteristics.
[0183] According to various embodiments, if the first information and the second information include features of the same speaker, the second electronic device (220) (e.g., the second audio processing module (322) or the second processor (325)) may store the first information and the second information as a single feature vector in operation 1613. According to an embodiment, the second electronic device (220) may store feature vectors for multiple speakers. In this case, the second electronic device (220) may assign a speaker identifier to each feature vector.
[0184] According to various embodiments, if the first information and the second information do not include characteristics of the same speaker, the second electronic device (220) (e.g., the second audio processing module (322) or the second processor (325)) may, in operation 1615, store the first information and the second information separately. In this case, the second electronic device (220) may assign a first identifier for the first speaker to the first information, and a second identifier for the second speaker to the second information.
[0185]
[0186] Figure 17 is a flowchart illustrating the operation of an audio signal processing system (20) according to various embodiments. While the operations in the following embodiments may be performed sequentially, they are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.
[0187] Referring to FIG. 16, a second electronic device (220) (e.g., a second audio processing module (322) or a second processor (325)) according to various embodiments may detect microphone function activation in operation 1710. Microphone function activation may be related to execution of a recording function, execution of a call function, or execution of a video storage function.
[0188] According to various embodiments, in response to activation of the microphone function, the second electronic device (220) (e.g., the second audio processing module (322) or the second processor (325)) may determine, in operation 1720, whether a first type of application or a second type of application is being executed. The first type of application may be an application requiring real-time processing of an audio signal, such as a call application. In addition, the second type of application may be an application allowing delayed processing of an audio signal, such as a recording application and a video storage application.
[0189] According to various embodiments, when the first type of application is executed, the second electronic device (220) (e.g., the second audio processing module (322) or the second processor (325)) may generate a synthetic audio signal based on the first method in operation 1730. According to one embodiment, the first method may be a method of generating an audio signal of better quality by using a first enhanced audio signal (415) obtained by removing noise from an audio signal (411) collected through the first audio collection device (311) as described with reference to FIGS. 3A to 8.
[0190] According to various embodiments, when a second type of application is executed, the second electronic device (220) (e.g., the second audio processing module (322) or the second processor (325)) may, in operation 1740, generate a synthetic audio signal based on the second method. According to one embodiment, the second method may be a method of converting the first enhanced audio signal (415) into text and using it to generate a better quality audio signal, as described with reference to FIGS. 3A, 3B, and 9 to 11.
[0191]
[0192] A mobile device (220) (e.g., a second electronic device (220)) according to various embodiments may include a communication circuit (321) configured to establish communication with an audio output device (210) (e.g., a first electronic device (210)), at least one processor (325), and a memory (323). For example, the memory (323) may store instructions that, when executed by the at least one processor (325), cause the mobile device (220) to: obtain a first enhanced audio signal (415) from which noise has been removed from a first audio signal (411) obtained through a first audio collection device (311) of the audio output device (210), obtain a second audio signal (421) obtained through a second audio collection device (312) of the audio output device (210), generate a second enhanced audio signal (509) based on the second audio signal (421) and first characteristic information (503) about a speaker in a low-noise environment stored in the memory (323), and generate a composite audio signal (513) by synthesizing the first enhanced audio signal (415) and the second enhanced audio signal (509).
[0193] According to various embodiments, the instructions may cause the mobile device (220) to: determine a synthesis ratio of the first enhanced audio signal (415) and the second enhanced audio signal (509) based on a quality of the first enhanced audio signal (415).
[0194] According to various embodiments, the instructions may cause the mobile device (220) to: increase a synthesis ratio of the first enhanced audio signal (415) relative to a synthesis ratio of the second enhanced audio signal (509) as the quality of the first enhanced audio signal (415) improves.
[0195] According to various embodiments, the instructions may cause the mobile device (220) to: increase a synthesis ratio of the second enhanced audio signal (509) relative to a synthesis ratio of the first enhanced audio signal (415) as the quality of the first enhanced audio signal (415) deteriorates.
[0196] According to various embodiments, the instructions may cause the mobile device (220) to: extract second feature information (701) related to a current speaker's speaking style based on the first enhanced audio signal (415), and generate the second enhanced audio signal (509) based on the second audio signal (421), the first feature information (503), and the second feature information (701).
[0197] According to various embodiments, the instructions may cause the mobile device (220) to: determine a reflection ratio of the first characteristic information (503) and the second characteristic information (701) based on a quality of the first enhanced audio signal (415).
[0198] According to various embodiments, the instructions may cause the mobile device (220) to: increase the reflection ratio of the second characteristic information (701) more than the reflection ratio of the first characteristic information (503) as the quality of the first enhanced audio signal (415) improves.
[0199] According to various embodiments, the instructions may cause the mobile device (220) to: increase the reflection ratio of the first characteristic information (503) more than the reflection ratio of the second characteristic information (701) as the quality of the first enhanced audio signal (415) deteriorates.
[0200] According to various embodiments, the mobile device (220) may be included. According to one embodiment, the instructions may cause the mobile device (220) to: output the synthesized audio signal (513) through the speaker (324) or store it in the memory (323).
[0201] According to various embodiments, the first characteristic information (503) may be related to at least one of a pitch component, a timbre, a length of a voice, or a loudness of a sound.
[0202] According to various embodiments, the second characteristic information (701) may be related to at least one of the speaker's emotion, speech rate, accent, or speech volume.
[0203] An audio signal processing system (20) according to various embodiments may include a first electronic device (210) and a second electronic device (220) that establishes communication with the first electronic device (210).
[0204] According to one embodiment, the first electronic device (210) may provide a first enhanced audio signal (415) obtained by removing noise from a first audio signal (411) obtained through a first audio collection device (311) and a second audio signal (421) obtained through a second audio collection device (312) to the second electronic device (220).
[0205] According to one embodiment, the second electronic device (220) may generate a second enhanced audio signal (509) based on the second audio signal (421) and the first characteristic information (503) about a speaker in a low-noise environment stored in the second electronic device (220), and may synthesize the first enhanced audio signal (415) and the second enhanced audio signal (509) to generate a synthesized audio signal (513).
[0206] According to various embodiments, the second electronic device (220) may determine a synthesis ratio of the first enhanced audio signal (415) and the second enhanced audio signal (509) based on the quality of the first enhanced audio signal (415).
[0207] According to various embodiments, the second electronic device (220) may extract second characteristic information (701) related to a current speaker's speaking style based on the first enhanced audio signal (415), determine a reflection ratio of the first characteristic information (503) and the second characteristic information (701) based on a quality of the first enhanced audio signal (415), and generate the second enhanced audio signal (509) based on the first characteristic information (503) and the second characteristic information (701) based on the second audio signal (421) and the reflection ratio.
[0208] According to various embodiments, the first electronic device (210) may include the first audio collection device (311) configured to collect an audio signal of a first frequency band and the second audio collection device (312) configured to collect an audio signal of a second frequency band narrower than the first frequency band.
[0209] A method of operating a mobile device (220) according to various embodiments may include an operation of obtaining a first enhanced audio signal (415) from which noise is removed from a first audio signal (411) obtained through a first audio collection device (311) of an audio output device (210), an operation of obtaining a second audio signal (421) obtained through a second audio collection device (312) of the audio output device (210), an operation of generating a second enhanced audio signal (509) based on the second audio signal (421) and first characteristic information (503) about a speaker in a low-noise environment stored in the mobile device (220), and an operation of generating a synthesized audio signal (513) by synthesizing the first enhanced audio signal (415) and the second enhanced audio signal (509).
[0210] According to various embodiments, the method of operating the mobile device (220) may include adjusting a synthesis ratio of the first enhanced audio signal (415) and the second enhanced audio signal (509) based on a quality of the first enhanced audio signal (415).
[0211] According to various embodiments, the operating method of the mobile device (220) may include an operation of increasing a synthesis ratio of the first enhanced audio signal (415) relative to a synthesis ratio of the second enhanced audio signal (509) as the quality of the first enhanced audio signal (415) improves, and an operation of increasing a synthesis ratio of the second enhanced audio signal (509) relative to a synthesis ratio of the first enhanced audio signal (415) as the quality of the first enhanced audio signal (415) deteriorates.
[0212] According to various embodiments, the method of operating the mobile device (220) may include an operation of extracting second feature information (701) related to a current speaker's speaking style based on the first enhanced audio signal (415) and an operation of generating the second enhanced audio signal (509) based on the second audio signal (421), the first feature information (503) and the second feature information (701).
[0213] According to various embodiments, the method of operating the mobile device (220) may include an operation of determining a reflection ratio of the first characteristic information (503) and the second characteristic information (701) based on the quality of the first enhanced audio signal (415).
[0214] According to various embodiments, the operating method of the mobile device (220) may include an operation of increasing a reflection ratio of the second characteristic information (701) more than a reflection ratio of the first characteristic information (503) as the quality of the first enhanced audio signal (415) improves, and an operation of increasing a reflection ratio of the first characteristic information (503) more than a reflection ratio of the second characteristic information (701) as the quality of the first enhanced audio signal (415) deteriorates.
[0215] A computer-readable recording medium according to various embodiments may store instructions for generating a synthetic audio signal by selecting a first method or a second method based on a type of application running on a mobile device (220) when activation of a microphone function is detected.
[0216] According to one embodiment, the first method may include an operation of obtaining a first enhanced audio signal (415) from which noise is removed from a first audio signal (411) obtained through a first audio collection device (311) of an audio output device (210), obtaining a second audio signal (421) obtained through a second audio collection device (312) of the audio output device (210), generating a second enhanced audio signal (509) based on the second audio signal (421) and first characteristic information (503) about a speaker in a low-noise environment stored in the mobile device (220), and synthesizing the first enhanced audio signal (415) and the second enhanced audio signal (509) to generate a synthesized audio signal (513).
[0217] According to one embodiment, the second method may include converting a first audio signal (411) acquired through a first audio collection device (311) of the audio output device (210) into text data, acquiring a second audio signal (421) acquired through a second audio collection device (312) of the audio output device (210), generating a third enhanced audio signal (911) based on the text data (902) and the first feature information (503), and generating a synthesized audio signal (915) by synthesizing the first enhanced audio signal (415) and the third enhanced audio signal (911).
[0218]
[0219] The electronic device (101) according to various embodiments disclosed in this document may be a device of various forms. The electronic device (101) may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a home appliance device. The electronic device (101) according to the embodiments of this document is not limited to the aforementioned devices.
[0220] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0221] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0222] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more commands stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one command among the one or more commands stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one command called. The one or more commands may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0223] According to one embodiment, the method according to the various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0224] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In a mobile device (220), A communication circuit (321) configured to establish communication with an audio output device (210); at least one processor (325); and Contains memory (323), The above memory (323) includes at least one instruction, and when the at least one instruction is executed by the at least one processor (325), the mobile device (220), Obtaining a first enhanced audio signal (415) from which noise has been removed from a first audio signal (411) obtained through a first audio collection device (311) of the above-mentioned audio output device (210), Obtaining a second audio signal (421) obtained through the second audio collection device (312) of the above-mentioned audio output device (210), Generating a second enhanced audio signal (509) based on the second audio signal (421) and the first characteristic information (503) about the speaker in a low-noise environment stored in the memory (323), A mobile device configured to generate a composite audio signal (513) by synthesizing the first enhanced audio signal (415) and the second enhanced audio signal (509).
2. In paragraph 1, The above instructions cause the mobile device (220) to: A mobile device that determines a synthesis ratio of the first enhanced audio signal (415) and the second enhanced audio signal (509) based on the quality of the first enhanced audio signal (415).
3. In paragraph 2, The above instructions cause the mobile device (220) to: As the quality of the first enhanced audio signal (415) is better, the synthesis ratio of the first enhanced audio signal (415) is increased compared to the synthesis ratio of the second enhanced audio signal (509). A mobile device that increases the synthesis ratio of the second enhanced audio signal (509) compared to the synthesis ratio of the first enhanced audio signal (415) as the quality of the first enhanced audio signal (415) deteriorates.
4. In paragraph 1, The above instructions cause the mobile device (220) to: Based on the first enhanced audio signal (415) above, second feature information (701) related to the current speaker's speaking style is extracted, A mobile device that generates the second enhanced audio signal (509) based on the second audio signal (421), the first feature information (503) and the second feature information (701).
5. In paragraph 4, The above instructions cause the mobile device (220) to: A mobile device that determines a reflection ratio of the first feature information (503) and the second feature information (701) based on the quality of the first enhanced audio signal (415).
6. In paragraph 5, The above instructions cause the mobile device (220) to: As the quality of the first enhanced audio signal (415) improves, the reflection ratio of the second feature information (701) is increased more than the reflection ratio of the first feature information (503). A mobile device that increases the reflection ratio of the first feature information (503) more than the reflection ratio of the second feature information (701) as the quality of the first enhanced audio signal (415) deteriorates.
7. In at least one of paragraphs 1 to 6, Including additional speakers (324), The above instructions cause the mobile device (220) to: A mobile device that outputs the above-mentioned synthetic audio signal (513) through the speaker (324) or stores it in the memory (323).
8. In at least one of paragraphs 1 to 6, The above first feature information (503) is a mobile device related to at least one of pitch component, tone, length or intensity of voice.
9. In at least one of paragraphs 1 to 6, The second characteristic information (701) is a mobile device related to at least one of the speaker's emotion, speech rate, accent, or speech volume.
10. In the audio signal processing system (20), A first electronic device (210); and a second electronic device (220) forming communication with the first electronic device (210), The above first electronic device (210) is, A first enhanced audio signal (415) obtained by removing noise from a first audio signal (411) obtained through a first audio collection device (311) and a second audio signal (421) obtained through a second audio collection device (312) are provided to the second electronic device (220). The above second electronic device (220) is, Generating a second enhanced audio signal (509) based on the second audio signal (421) and the first characteristic information (503) about the speaker in a low-noise environment stored in the second electronic device (220), An audio signal processing system that generates a composite audio signal (513) by synthesizing the first enhanced audio signal (415) and the second enhanced audio signal (509).
11. In paragraph 10, The above second electronic device (220) is, An audio signal processing system that determines a synthesis ratio of the first enhanced audio signal (415) and the second enhanced audio signal (509) based on the quality of the first enhanced audio signal (415).
12. In paragraph 10, The above second electronic device (220) is, Based on the first enhanced audio signal (415) above, second feature information (701) related to the current speaker's speaking style is extracted, Based on the quality of the first enhanced audio signal (415), the reflection ratio of the first feature information (503) and the second feature information (701) is determined, An audio signal processing system that generates the second enhanced audio signal (509) based on the first feature information (503) and the second feature information (701) based on the second audio signal (421) and the reflection ratio.
13. In paragraph 10, The above first electronic device (210) is, An audio signal processing system comprising a first audio collection device configured to collect an audio signal of a first frequency band and a second audio collection device configured to collect an audio signal of a second frequency band narrower than the first frequency band.
14. In the operating method of a mobile device (220), An operation of obtaining a first enhanced audio signal (415) from which noise has been removed from a first audio signal (411) obtained through a first audio collection device (311) of an audio output device (210); An operation of obtaining a second audio signal (421) obtained through a second audio collection device (312) of the above-mentioned audio output device (210); An operation of generating a second enhanced audio signal (509) based on the second audio signal (421) and the first characteristic information (503) about the speaker in a low-noise environment stored in the mobile device (220); and, A method comprising an operation of generating a composite audio signal (513) by synthesizing the first enhanced audio signal (415) and the second enhanced audio signal (509).
15. In paragraph 14, A method comprising an operation for adjusting a synthesis ratio of the first enhanced audio signal (415) and the second enhanced audio signal (509) based on the quality of the first enhanced audio signal (415).
Citation Information
Patent Citations
Method for processing audio signal for improving call quality and apparatus for the same
KR1020160026535A
Hearing aid with voice synthesis function considering speaker characteristics and method thereof
KR1020180087038A
the Sound Outputting Device including a plurality of microphones and the Method for processing sound signal using the plurality of microphones
KR102565882B1
Wireless Ear-Phone and Portable Terminal Using the Same
US20080153556A1
KR20200085030A