Electronic apparatus for binaural speech enhancement and operating method thereof
The electronic device uses a neural network-based speech enhancement model to process binaural audio signals, addressing the challenge of noise reduction while maintaining spatial cues, resulting in enhanced audio quality and improved user experience.
Patent Information
- Application Number
- PCT/KR2024/020127
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-07
- Filing Date
- 2024-12-10
- Publication Date
- 2025-08-14
AI Technical Summary
Existing audio devices struggle to effectively enhance binaural speech while preserving spatial characteristics and reducing noise, particularly interference and diffuse noise, leading to distorted audio experiences.
An electronic device employs a neural network-based speech enhancement model that processes binaural input signals to reduce noise components while maintaining spatial cues, using a learning database to select and process speech, spatial impulse response, and noise data to generate improved audio output.
The solution enhances audio quality by reducing noise and preserving spatial characteristics, providing clearer and more accurate sound localization, thereby improving the user experience.
Smart Images

Figure KR2024020127_14082025_PF_FP_ABST
Abstract
Description
Electronic device for binaural speech enhancement and method of operation thereof
[0001] The present disclosure relates to an electronic device for binaural speech enhancement and a method of operating the same.
[0002] Electronic devices may provide functions related to audio signal processing. For example, the electronic devices may provide functions such as a call function that collects and transmits audio signals, a recording function that records audio signals, and an audio output function that outputs audio signals. The electronic devices may output audio through an external audio output device, such as earphones or headphones, or through an audio output module built into the electronic device. Earphones and headphones are examples of binaural devices that include a left channel that outputs audio to the user's left ear and a right channel that outputs audio to the user's right ear.
[0003] An electronic device according to one embodiment includes at least one processor including a processing circuit and at least one memory storing instructions executable by the at least one processor, wherein the at least one processor individually and / or collectively executes the instructions to cause the electronic device to select speech data, spatial impulse response data and noise data from a learning database, determine a binaural input signal based on the selected speech data, the selected spatial impulse response data and the selected noise data, obtain a binaural output signal from a neural network-based speech enhancement model having the binaural input signal as an input, determine a target binaural signal based on the selected speech data, the selected spatial impulse response data and the selected noise data, and update parameters of the speech enhancement model based on the binaural output signal obtained from the speech enhancement model and the determined target binaural signal.
[0004] A method of operating an electronic device according to one embodiment may include selecting speech data, spatial impulse response data, and noise data from a learning database, determining a binaural input signal based on the selected speech data, the selected spatial impulse response data, and the selected noise data, obtaining a binaural output signal from a neural network-based speech enhancement model that has the binaural input signal as an input, determining a target binaural signal based on the selected speech data, the selected spatial impulse response data, and the selected noise data, and updating parameters of the speech enhancement model based on the binaural output signal obtained from the speech enhancement model and the determined target binaural signal.
[0005] The above and other aspects, features and advantages of specific embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings.
[0006] FIG. 1 is a block diagram illustrating an exemplary electronic device according to various embodiments;
[0007] FIG. 2 is a block diagram illustrating an exemplary configuration of an audio module according to various embodiments;
[0008] FIG. 3 is a diagram illustrating an exemplary audio signal processing system including a binaural device and an electronic device according to various embodiments;
[0009] FIG. 4 is a block diagram illustrating an exemplary configuration of an electronic device capable of performing learning operations of a voice enhancement model according to various embodiments;
[0010] FIG. 5 is a diagram illustrating an exemplary operation for selecting learning data according to various embodiments;
[0011] FIG. 6 is a diagram illustrating an exemplary structure and exemplary learning operation of a voice enhancement model according to various embodiments;
[0012] FIG. 7 is a flowchart illustrating an exemplary method of operating an electronic device for training a voice enhancement model according to various embodiments;
[0013] FIG. 8 is a flowchart illustrating exemplary operations for determining a binaural input signal and a target binaural signal for learning a voice enhancement model according to various embodiments;
[0014] FIG. 9 is a flowchart illustrating exemplary operations for obtaining diffusion noise data from a recorded audio signal according to various embodiments;
[0015] FIG. 10 is a flowchart illustrating an exemplary method for providing an audio signal using a voice enhancement model according to various embodiments; and
[0016] FIG. 11 is a diagram for explaining quality improvement of an audio signal by binaural voice enhancement processing according to various embodiments.
[0017] Hereinafter, various embodiments will be described in more detail with reference to the attached drawings. In describing various embodiments with reference to the attached drawings, identical components will be assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof may not be provided.
[0018] FIG. 1 is a block diagram illustrating an exemplary electronic device capable of performing the operations described within the present disclosure according to various embodiments.
[0019] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In various embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178), the subscriber identification module (196)), or may have one or more other components added. In various embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0020] The processor (120) may include various processing circuits and / or multiple processors. For example, as used in this disclosure, including in the claims, the term “processor” may include various processing circuits including at least one processor, wherein one or more of the at least one processor may be configured to perform various functions described herein in an individually and / or collectively distributed manner. When “processor,” “at least one processor,” and “one or more processors” are described herein as being configured to perform multiple functions, these terms include, but are not limited to, situations where one processor performs some of the recited functions and another processor performs other of the recited functions, and situations where a single processor may perform all of the recited functions. Furthermore, the at least one processor may include a combination of processors that perform various recited / disclosed functions, for example, in a distributed manner. The at least one processor may execute program instructions to achieve or perform various functions. The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134).The processor (120) may be implemented as a system on chip (SoC) or an integrated circuit (IC) that performs processing. The processor (120) may include one or more processors. The operations of the electronic device (101) described in the present disclosure may be performed by a single processor or by a combination of multiple processors. When the operations of the electronic device (101) are performed by a combination of multiple processors, any one processor included in the combination of processors may perform some of the operations of the electronic device (101).
[0021] According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0022] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0023] The memory (130) can store various data used by at least one component (e.g., the processor (120) or the sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., the program (140)), input data or display data for commands related thereto, and visual contents that can be displayed through the display module (160). The memory (130) can include a volatile memory (132) or a non-volatile memory (134). The memory (130) can store at least one instruction that can be executed by the processor (120). The memory (130) can include one or more storage devices (storage spaces). Instructions for controlling the processor (120) to perform operations of the electronic device (101) described in the present disclosure can be stored in one memory or can be divided and stored in multiple memories.
[0024] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0025] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0026] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0027] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display (e.g., a flexible touch display), a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch. The display module (160) may be implemented as, for example, a bendable structure, a foldable structure, and / or a rollable structure.
[0028] The audio module (170) can convert sound into an electrical signal, or vice versa. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., a speaker, earphone, hearing aid, or headphone) directly or wirelessly connected to the electronic device (101).
[0029] The sensor module (176) may include one or more sensors. The sensor module (176) may detect an operating state (e.g., power or temperature) of the electronic device (101) or an external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a bending sensor, an inertia measurement unit (IMU), a biosignal sensor, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an infrared (IR) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0030] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0031] The connection terminal (178) may include a connector that allows the electronic device (101) to be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., an earphone connector or a headphone connector).
[0032] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0033] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0034] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0035] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0036] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a Wi-Fi communication module, a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0037] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)).
[0038] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0039] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0040] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0041] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0042] FIG. 2 is a block diagram illustrating an exemplary configuration of an audio module according to various embodiments.
[0043] Referring to FIG. 2, the audio module (170) may include an audio input interface (210) (e.g., including an audio input circuit), an audio input mixer (220), an analog to digital converter (ADC) (230), an audio signal processor (240) (e.g., including an audio signal processing circuit), a digital to analog converter (DAC) (250), an audio output mixer (260), and an audio output interface (270) (e.g., including an audio output circuit). Some of the components of the audio module (170) may be omitted, and other components may be further included in the audio module (170).
[0044] The audio input interface (210) may include various circuits and may receive an audio signal corresponding to a sound acquired from the outside of the electronic device (101) as part of the input module (150) or through a microphone (e.g., a dynamic microphone, a condenser microphone, or a piezo microphone) configured separately from the electronic device (101). For example, when the audio signal is acquired from an external electronic device (102) (e.g., an earphone, a hearing aid, a headset, or a microphone), the audio input interface (210) may be directly connected to the external electronic device (102) through a connection terminal (178) or wirelessly (e.g., Bluetooth communication) through a wireless communication module (192) to receive the audio signal. According to one embodiment, the audio input interface (210) may receive a control signal (e.g., a volume control signal received through an input button) related to the audio signal acquired from the external electronic device (102). The audio input interface (210) includes a plurality of audio input channels and can receive different audio signals for each corresponding audio input channel among the plurality of audio input channels. According to one embodiment, additionally or alternatively, the audio input interface (210) can receive audio signals from other components of the electronic device (101), such as the processor (120) or the memory (130).
[0045] The audio input mixer (220) can synthesize a plurality of input audio signals into at least one audio signal. For example, the audio input mixer (220) can synthesize a plurality of analog audio signals input through the audio input interface (210) into at least one analog audio signal.
[0046] The ADC (230) can convert an analog audio signal into a digital audio signal. For example, the ADC (230) can convert an analog audio signal received through an audio input interface (210) or, additionally or alternatively, an analog audio signal synthesized through an audio input mixer (220) into a digital audio signal.
[0047] The audio signal processor (240) may include various audio signal processing circuits and may perform various processing on a digital audio signal input through the ADC (230) or a digital audio signal received from another component of the electronic device (101). For example, the audio signal processor (240) may change a sampling rate, apply one or more filters, perform interpolation processing, amplify or attenuate all or part of a frequency band, process noise (e.g., noise or echo reduction), change a channel (e.g., switching between mono and stereo), mix, or extract a specified signal on one or more digital audio signals. According to one embodiment, one or more functions of the audio signal processor (240) may be implemented in the form of an equalizer.
[0048] The DAC (250) can convert a digital audio signal into an analog audio signal. For example, the DAC (250) can convert a digital audio signal processed by an audio signal processor (240) or a digital audio signal obtained from another component of the electronic device (101) (e.g., a processor (120) or a memory (130)) into an analog audio signal.
[0049] The audio output mixer (260) can synthesize a plurality of audio signals to be output into at least one audio signal. For example, the audio output mixer (260) can synthesize an audio signal converted into analog through the DAC (250) and another analog audio signal (e.g., an analog audio signal received through the audio input interface (210)) into at least one analog audio signal.
[0050] The audio output interface (270) may include various circuits and may output an analog audio signal converted through the DAC (250) or, additionally or alternatively, an analog audio signal synthesized by the audio output mixer (260) to the outside of the electronic device (101) through the audio output module (155). The audio output module (155) may include, for example, a speaker such as a dynamic driver or a balanced armature driver, or a receiver. According to one embodiment, the audio output module (155) may include a plurality of speakers. In this case, the audio output interface (270) may output an audio signal having a plurality of different channels (e.g., stereo or 5.1 channels) through at least some of the plurality of speakers. According to one embodiment, the audio output interface (270) can output an audio signal by being connected directly to an external electronic device (102) (e.g., earphones, hearing aids, headsets, or external speakers) through a connection terminal (178) or wirelessly through a wireless communication module (192).
[0051] According to one embodiment, the audio module (170) can generate at least one digital audio signal by synthesizing a plurality of digital audio signals using at least one function of the audio signal processor (240) without separately having an audio input mixer (220) or an audio output mixer (260).
[0052] According to one embodiment, the audio module (170) may include an audio amplifier (not shown) (e.g., a speaker amplifier circuit) capable of amplifying an analog audio signal input through the audio input interface (210) or an audio signal to be output through the audio output interface (270). According to one embodiment, the audio amplifier may be configured as a separate module from the audio module (170).
[0053] FIG. 3 is a diagram illustrating an exemplary audio signal processing system including a binaural device and an electronic device according to various embodiments.
[0054] Referring to FIG. 3, the audio signal processing system (300) may include an electronic device (101) and a binaural device (310, 320). According to one embodiment, the binaural device (310, 320) may be connected to the electronic device (101) by wire or wirelessly and may output an audio signal transmitted by the electronic device (101). The binaural device (310, 320) may include a device capable of outputting two-channel binaural audio signals to both ears of a user. The binaural device (310, 320) may include, but is not limited to, earphones (wireless earphones or wired earphones), a headset, a hearing aid, or smart glasses, for example.
[0055] According to one embodiment, the binaural device (310, 320) may be a wireless earphone capable of forming a short-range communication channel (e.g., a communication channel based on a Bluetooth module) with the electronic device (101), as illustrated. For example, the binaural device (310, 320) may include a true-wireless stereo (TWS) earphone, a wireless headphone, and / or a wireless headset. Although the binaural device (310, 320) is illustrated as a kernel-type wireless earphone in FIG. 3 , it is not limited thereto. For example, the binaural device (310, 320) may be a stem-type wireless earphone in which at least a portion of the housing protrudes in a specific direction to collect a good user voice signal. According to another embodiment, the binaural device (310, 320) may be a wired earphone connected to the electronic device (101) via a wire.
[0056] The binaural device (310, 320) may include a binaural device (310) that outputs an audio signal of a left channel corresponding to the user's left ear and a binaural device (320) that outputs an audio signal of a right channel corresponding to the user's right ear. The binaural device (310, 320) may obtain (or collect) an external audio signal (or audio signal) using a plurality of microphones and transmit the obtained audio signal to the electronic device (101).
[0057] According to one embodiment, a binaural device (310, 320) exemplified as an earphone type may include a housing (301, 301-1) (or case) having an insertion part (303, 303-1) that can be inserted into a user's ear, and a receiving part (305, 305-1) that is connected to the insertion part (303, 303-1) and that can be at least partially received on the user's auricle. The binaural device (310, 320) may include a plurality of microphones (350, 350-1, 355, 355-1). For example, the binaural device (310) may include microphones (350, 355), and the binaural device (320) may include microphones (350-1, 355-1).
[0058] According to one embodiment, the binaural device (310, 320) may include an input interface (377, 377-1) including various interface circuits for receiving user input. The input interface (377, 377-1) may include, for example, a physical interface (e.g., a physical button or a touch button) and a virtual interface (e.g., a gesture, object recognition, or voice recognition). According to one embodiment, the binaural device (310, 320) may include a touch sensor capable of detecting contact with the user's skin. For example, the binaural device (310, 320) may include an area where the touch sensor is arranged (e.g., an area corresponding to the input interface (377, 377-1)). According to one embodiment, a user input may be applied by the user touching the area with a body part. The touch input may include, for example, a single touch, multiple touches, a swipe, and / or a flick.
[0059] According to one embodiment, the microphones (350, 350-1, 355, 355-1) mounted on the binaural device (310, 320) may perform the function of the input module (150) described above with reference to FIG. 1. The description of the input module (150) described with reference to FIG. 1 may not be repeated here. For example, among the microphones (350, 350-1, 355, 355-1), the first microphone (350, 350-1) may be placed on the holder (305, 305-1) so that at least a portion of the sound hole is exposed to the outside, based on the inner side of the ear, so as to collect external ambient audio while the binaural device (310, 320) is worn on the user's ear. Among the microphones (350, 350-1, 355, 355-1), the second microphone (355, 355-1) may be placed near the insertion portion (303, 303-1). The second microphone (355, 355-1) may be placed so that at least a portion of the phonatory cavity is exposed toward the inside of the external auditory canal or at least a portion thereof is in contact with the inner wall of the external auditory canal, based on the opening toward the auricle of the external auditory canal, so as to collect a signal transmitted into the external auditory canal (or, external auditory meatus) while the binaural device (310, 320) is worn on the user's ear. For example, when a user wears the binaural device (310, 320) and speaks, at least a portion of the vibration due to the speech is transmitted through the user's skin, muscles, bones, etc., and the transmitted vibration can be collected as ambient audio by the second microphone (355, 355-1) inside the ear.
[0060] According to one embodiment, the second microphone (355, 355-1) may include various types of microphones (e.g., an in-ear microphone, an inner microphone, or a bone conduction microphone) that can collect sound from the inner cavity of the user's ear. For example, the second microphone (355, 355-1) may include at least one air conduction microphone and / or a bone conduction microphone for detecting voice. The air conduction microphone may detect voice transmitted through the air (e.g., the user's speech) and output a voice signal corresponding to the detected voice. The bone conduction microphone may measure vibration of a bone (e.g., the skull) caused by the user's speech and output a voice signal corresponding to the measured vibration. The bone conduction microphone may be referred to as a bone conduction sensor or various other names. The voice detected by the air conduction microphone may be a voice mixed with external noise while the user's speech is transmitted through the air. Since the voice detected by the bone conduction microphone is a voice detected according to the vibration of the bone, it may be a voice with less external noise (or, the influence of noise).
[0061] In FIG. 3, the first microphone (350, 350-1) and the second microphone (355, 355-1) are illustrated as being mounted one each in the binaural devices (310, 320), but this is not limited thereto, and the first microphone (350, 350-1), which is an external microphone, and the second microphone (355, 355-1), which is an internal microphone, may be mounted in multiple numbers in each of the binaural devices (310, 320). Although omitted in FIG. 3, the binaural devices (310, 320) may further include an accelerator and a vibration sensor (e.g., a VPU (voice pickup unit) sensor) for voice activity detection (VAD).
[0062] According to one embodiment, the binaural device (310, 320) may include a sensor capable of detecting a state in which the device is worn on the user's ears. For example, the binaural device (310, 320) may include a sensor capable of detecting a distance from an object (e.g., an infrared sensor or a laser sensor), or a sensor capable of detecting contact with an object (e.g., a touch sensor). When the binaural device (310, 320) is worn on the user's ears, the binaural device (310, 320) may detect a distance from or contact with the skin through the sensor to generate a signal, and may recognize whether the binaural device (310, 320) is currently being worn. The binaural device (310, 320) may correspond to the binaural device (400) of FIG. 4.
[0063] According to one embodiment, the electronic device (101) may include an audio module (170) as described above with reference to FIGS. 1 and 2. The description provided with reference to FIGS. 1 and 2 may not be repeated herein. The electronic device (101) may perform audio signal processing, such as noise processing (e.g., noise suppression processing), frequency band adjustment, or gain adjustment, through the audio module (170) (e.g., through the audio signal processor (240) of FIG. 2).
[0064] According to one embodiment, the electronic device (101) can form a communication channel with the binaural device (310, 320), transmit a designated audio signal to the binaural device (310, 320), or receive an audio signal from the binaural device (310, 320). For example, the electronic device (101) can be various electronic devices such as a portable terminal, a terminal device, a smartphone, a personal computer (PC), a server, a laptop, a tablet PC, a VR (virtual reality) / AR (augmented reality) device, or a pad-type electronic device that can form a communication channel (e.g., a wired or wireless communication channel) with the binaural device (310, 320).
[0065] According to one embodiment, the electronic device (101) can perform binaural speech enhancement processing on an audio signal to be output through the binaural device (310, 320). Through the binaural speech enhancement processing, an audio signal with improved audio quality can be provided to a user wearing the binaural device (310, 320). In one embodiment, the binaural speech enhancement processing can reduce noise components in the audio signal, thereby improving the signal-to-noise ratio (SNR) of the speech signal, and preserve spatial characteristics or spatial cues of the noise components. In the binaural speech enhancement processing, information on the sound image location of noise components that may be included in the audio signal, for example, interference noise and / or diffuse noise, can be preserved. Interference noise may be noise that is transmitted directionally from the sound source that generates the noise signal, while diffuse noise may be noise that is transmitted with the same audio intensity from all directions without directionality.
[0066] In the case of interference noise, the direction from which the noise signal is transmitted can be valuable information for the user. For example, if an object such as a car approaches the user or a gunshot is fired nearby, the sound of the car or gunshot constitutes noise, but the direction from which the sound or gunshot is heard can be significant to the user. For user safety, it may be important to provide the user with information about the direction from which the car sound or gunshot is heard. If the directional information of the noise signal is not preserved during the noise reduction process, the user may have difficulty determining the origin of the car sound or gunshot. Similarly, in the case of diffusion noise, if spatial characteristics or spatial cues are not preserved during the noise reduction process, the sound image may be centered, creating a distorted experience for the user, as if the noise signal originated from a point source. Therefore, preserving the omnidirectional nature of diffusion noise during the noise reduction process may be necessary.
[0067] According to various embodiments described in the present disclosure, a binaural voice enhancement process can be provided that can reduce the noise component of an audio signal to be output through a binaural device (310, 320) while preserving the spatial characteristics (e.g., directionality) of the noise component. In one embodiment, audio signals acquired from microphones (350, 350-1, 355, 355-1) included in the binaural device (310, 320) are transmitted to an electronic device (101), and the electronic device (101) can generate an audio signal in which the noise component is reduced and the spatial characteristics of the noise component are preserved using a neural network-based voice enhancement model. The voice enhancement model can adjust the ratio between the voice component and the noise component included in the audio signal, and can preserve the binaural cues of the noise component as well as the voice component (e.g., preserving the directionality and sound image of the voice component and the noise component). According to one embodiment, the voice enhancement model may be learned by the training operations of the voice enhancement model described in FIGS. 4, 5, 6, 7, and 8. The electronic device (101) may provide the generated audio signal to the user through the binaural device (310, 320). The audio signal for the left channel may be provided through the binaural device (310), and the audio signal for the right channel may be provided through the binaural device (320).
[0068] FIG. 4 is a block diagram illustrating an exemplary configuration of an electronic device capable of performing learning operations of a voice enhancement model according to various embodiments.
[0069] Referring to FIG. 4, the electronic device (101) may include an input module (150) for receiving user input, an audio output module (155) for outputting sound to the outside, an audio module (170) for adjusting the output volume of audio output from the electronic device (101), a communication module (190) (e.g., including a communication circuit) for communicating with a binaural device (400) (e.g., the binaural devices (310, 320) of FIG. 3), one or more processors (120) (e.g., including a processing circuit), and / or one or more memories (130) for storing computer-executable instructions executable by the one or more processors (120). The electronic device (101) may include fewer or more configurations than the various configurations described above with reference to FIG. 1.
[0070] The electronic device (101), the processor (120), the memory (130), the input module (150), the audio output module (155), the audio module (170), and the communication module (190) may correspond to the electronic device (101), the processor (120), the memory (130), the input module (150), the audio output module (155), the audio module (170), and the communication module (190) described above with reference to FIGS. 1, 2, and 3 (also referred to as FIGS. 1 to 3), respectively, and the description provided with reference to FIGS. 1 to 3 may not be repeated here. The binaural device (400) may correspond to the binaural device (310, 320) described in FIG. 3, and may be a two-channel audio output device such as wired / wireless earphones, hearing aids, headsets, or smart glasses.
[0071] In one embodiment, the electronic device (101) can communicate with the binaural device (400) and perform various operations. For example, the electronic device (101) can establish a short-range communication channel (e.g., a communication channel based on a Bluetooth module) with the binaural device (400) through the communication module (190). The electronic device (101) can output an audio signal through the audio output module (155) and can receive an audio signal acquired through a microphone of the binaural device (400) (e.g., the external microphone (350, 350-1) and / or the second microphone (355, 355-1) which is an internal microphone as described above with reference to FIG. 3). The audio module (170) of the electronic device (101) can adjust the audio volume output to the binaural device (400).
[0072] An electronic device (101) according to one embodiment may perform learning operations for training a neural network-based voice enhancement model (420) (e.g., voice enhancement model (600)). The electronic device (101) for performing the learning operations may include one or more processors (120) and one or more memories (130) that store instructions executable by the one or more processors (120). When the instructions stored in the one or more memories (130) are executed by the one or more processors (120), the executed instructions may cause the electronic device (101) to perform the following operations.
[0073] In one embodiment, the processor (120) may select speech data, room impulse response data, and noise data from a learning database (410). The learning database (410) may be a database storing learning data for training a speech enhancement model (420). For example, the speech data, room impulse response data, and noise data may be pre-prepared learning data.
[0074] In one embodiment, the noise data may include interference noise data and diffuse noise data. The interference noise data may be training data for a noise signal having directionality, and the diffuse noise data may be training data for a noise signal having no directionality. The spatial impulse response data may be data representing impulse response characteristics of an audio signal in space. The spatial impulse response data may include spatial impulse response data from the sound image location of the speech data to the binaural device (400) and spatial impulse response data from the sound image location of the interference noise data to the binaural device (400).
[0075] In one embodiment, the processor (120) can determine a binaural input signal based on the selected speech data, the selected spatial impulse response data, and the selected noise data. In one embodiment, the processor (120) can determine a binaural input signal by applying the selected speech data, the selected spatial impulse response data, and the selected noise data to a defined mathematical model (e.g., Equation 2 below). The binaural input signal is an audio signal input to the speech enhancement model (420), and can include an audio signal for a left channel and an audio signal for a right channel.
[0076] In one embodiment, the processor (120) may obtain a binaural output signal from a neural network-based voice enhancement model (420) that takes a binaural input signal as input. The voice enhancement model (420) may be a model that outputs a binaural output signal in which the intensity of noise components included in the binaural input signal is reduced while preserving the directionality of the noise components. The binaural output signal may be an audio signal output from the voice enhancement model (420) after binaural voice enhancement processing is performed by the voice enhancement model (420).
[0077] In one embodiment, a neural network model, such as a voice enhancement model (420), may be created through learning. "Created through learning" may mean that a basic neural network model is trained using a learning algorithm using a plurality of learning data, thereby creating a neural network model or operating rules defined to perform a desired characteristic (or purpose). For example, such learning may be performed on the device itself on which the learning operations according to the present disclosure are performed, or may be performed through a separate server and / or system. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0078] In one embodiment, a neural network model may include a plurality of neural network layers. Each of the plurality of neural network layers has a plurality of weight values, and can perform neural network operations through operations between the operation results of the previous layer and the plurality of weights. The plurality of weights of the plurality of neural network layers may be optimized based on the learning results of the neural network model. For example, the plurality of weights may be updated so that a loss value or cost value obtained from the neural network model is reduced or minimized during the learning process. The neural network model may include a deep neural network (DNN). For example, a deep neural network may include a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or deep Q-networks, but is not limited to the examples described above.
[0079] In one embodiment, the speech enhancement model (420) may include an encoder (e.g., encoders (612, 614) of FIG. 6) that converts the domain-above binaural input signal into a frequency-domain binaural input signal, a mask estimator (e.g., mask estimator (630) of FIG. 6) that determines a mask value to be applied to the frequency-domain binaural input signal based on a noise suppression parameter related to a degree of noise suppression, and a decoder (e.g., decoders (652, 654) of FIG. 6) that converts a result signal obtained by applying the determined mask value to the frequency-domain binaural input signal into a time-domain binaural output signal. In one embodiment, the noise suppression parameter may be randomly selected within a range of specified values. For example, the noise suppression parameter may be selected as any value within a range from 0 to 1. Meanwhile, the voice enhancement model (420) used in various embodiments may be implemented in various embodiments depending on the manufacturer of the electronic device (101) or the user of the electronic device (101), and is not limited by the above example.
[0080] In one embodiment, the processor (120) may determine a target binaural signal based on the selected speech data, the selected spatial impulse response data, and the selected noise data. The target binaural signal may include a target binaural signal for a left channel and a target binaural signal for a right channel. The target binaural signal for the left channel may be determined based on the spatial impulse response data for the left channel and the noise data for the left channel. The target binaural signal for the right channel may be determined based on the spatial impulse response data for the right channel and the noise data for the right channel.
[0081] In one embodiment, the processor (120) can determine a target binaural signal based on the selected speech data, spatial impulse response data for the selected speech data, the selected noise data, spatial impulse response data for the selected noise data, and a noise suppression parameter related to a degree of noise suppression. In one embodiment, the noise suppression parameter can be randomly selected within a range of specified values. The noise suppression parameter used to determine the target binaural signal can be the same as the noise suppression parameter applied to the speech enhancement model (420) to obtain the binaural output signal. In one embodiment, the processor (120) can determine the target binaural signal by applying the selected speech data, spatial impulse response data for the selected speech data, the selected noise data, spatial impulse response data for the selected noise data, and the noise suppression parameter to a defined mathematical model (e.g., Equations 3 and 4).
[0082] In one embodiment, the processor (120) may control to update parameters of the speech enhancement model (420) based on the binaural output signal obtained from the speech enhancement model (420) and the determined target binaural signal. For example, the speech enhancement model (420) may include weight coefficients and / or biases of a neural network.
[0083] In one embodiment, the processor (120) may determine a loss (or loss data) based on the binaural output signal obtained from the speech enhancement model (420) and the determined target binaural signal, and control the parameters of the speech enhancement model (420) to be updated based on the determined loss. The loss may include a loss based on the difference between the result obtained by the neural network model performing forward propagation using the input data and the actual result. For example, the loss may include a loss based on the difference between the binaural output signal output by the speech enhancement model (420) based on the input binaural input signal and the target binaural signal corresponding to the actual result. In addition, the loss may include a loss based on the difference between the two channels of the binaural output signal and / or the difference between the two channels of the target binaural signal.
[0084] In one embodiment, the processor (120) can update parameters of the speech enhancement model (420) based on at least one of a speech-to-distortion ratio loss, a loss based on inter-channel level difference (ICLD), and a loss based on inter-channel time difference (ICTD) based on a binaural output signal obtained from the speech enhancement model (420) and the determined target binaural signal. The inter-channel level difference is information that characterizes a signal level difference of an audio signal between a left channel and a right channel, and the inter-channel time difference is information that characterizes a timing difference between the left channel and the right channel.
[0085] In one embodiment, the processor (120) may gradually update the parameters of the speech enhancement model (420) to desirable values by performing an error backpropagation-based machine learning algorithm that utilizes the gradient of the determined loss. In the error backpropagation-based machine learning algorithm, the parameters of the speech enhancement model (420) may be updated in a direction in which the loss calculated through the loss function is reduced.
[0086] FIG. 5 is a diagram illustrating an exemplary operation for selecting learning data according to various embodiments.
[0087] Referring to FIG. 5, a learning database according to one embodiment (e.g., the learning database (410) of FIG. 4) may store a sound source database (510), a spatial impulse response database (520), and a background noise database (530). Data stored in the sound source database (510), the spatial impulse response database (520), and the background noise database (530) may correspond to raw data.
[0088] In one embodiment, the audio source database (510) may store a speech dataset and a noise dataset for interference noise that serve as targets for binaural voice enhancement. For example, the speech dataset and noise dataset may be publicly available datasets or datasets acquired through direct recording in a recording studio.
[0089] In one embodiment, the spatial impulse response database (520) can store spatial impulse response datasets in various environments. The spatial impulse response data to be used for learning can be extracted, for example, through a process such as inverse filtering, from audio signals acquired by microphones of a binaural device when test signals such as a sine sweep or a maximum length sequence (MLS) are played from various distances and various angles (e.g., azimuth or elevation) while the binaural device is attached to a simulator device (e.g., a head and torso simulator device). In addition to the method of directly acquiring spatial impulse response data, spatial impulse response data can also be acquired through an acoustic simulator device that can simulate multi-channel microphone signals for both ears under various reverberant environments.
[0090] In one embodiment, the background noise database (530) may store a noise dataset for diffuse noise. As the noise dataset for diffuse noise, for example, a dataset obtained directly from a binaural device (e.g., the binaural devices (310, 320) of FIG. 3 or the binaural device (400) of FIG. 4) in various environments or a dataset generated through an acoustic simulator may be used.
[0091] The processor (120) of the electronic device (101) can select data to be used for learning from each of a sound source database (510), a spatial impulse response database (520), and a background noise database (530), and select (540) learning data to be used for learning a voice enhancement model (e.g., the voice enhancement model (420) of FIG. 4) based on the selected data. For example, the processor (120) can determine a binaural input signal to be input to the voice enhancement model and a target binaural signal corresponding to a desired signal to be output by the voice enhancement model based on the selected data.
[0092] In one embodiment, assume a binaural device equipped with M multi-channel microphones on each of the left and right sides. In this case, the input signal y(t) incident on the microphones of the binaural device can be expressed as follows: Equation 1.
[0093] [Mathematical Formula 1]
[0094]
[0095] Here, L represents left, R represents right. 1, 2, ..., M represent the number of microphone channels, and T represents transpose.
[0096] The above input signal y(t) can be modeled as shown in the following mathematical expression 2.
[0097] [Equation 2]
[0098]
[0099] Here, represents the voice component of the voice data (or target voice) selected for learning, represents the interference noise component of the interference noise data selected for learning. can represent multi-channel noise components including diffusion noise components of diffusion noise data selected for learning and self-noise components generated by the microphone of a binaural device. can represent the spatial impulse response from the sound location of the speech data to the binaural device (or the microphone of the binaural device). can represent the spatial impulse response from the sound location of the interference noise data to the binaural device (or the microphone of the binaural device).
[0100] In a binaural device, based on a reference microphone channel (assumed to be microphone channel 1), the target binaural signal targeted in binaural voice enhancement processing can be expressed by the following mathematical expressions 3 and 4. Mathematical expression 3 represents the target binaural signal for the left channel, and mathematical expression 4 represents the target binaural signal for the right channel.
[0101] [Equation 3]
[0102]
[0103] [Equation 4]
[0104]
[0105] Here, represents the noise suppression parameter that determines the degree of suppression against interference noise, represents a noise suppression parameter that determines the degree of suppression of diffusion noise. In one embodiment, and may be an adjustable constant. For example, When it is set to 0.1, the interference noise component in the target binaural signal can be reduced by 20 dB.
[0106] FIG. 6 is a diagram illustrating an exemplary structure and exemplary learning operation of a voice enhancement model according to various embodiments.
[0107] Referring to FIG. 6, a speech enhancement model (600) according to one embodiment (e.g., the speech enhancement model (420) of FIG. 4) may include encoders (612, 614), a mask estimator (630), and decoders (652, 654).
[0108] The voice enhancement model (600) can receive a multi-channel binaural input signal obtained from a binaural device (e.g., the binaural device (310, 320) of FIG. 3 or the binaural device (400) of FIG. 4). The multi-channel binaural input signal , where T represents the length of the audio signal and M represents the number of microphone channels included in the binaural device) can be input to the encoders (612, 614). The encoders (612, 614) output the binaural input signal of the left channel (e.g., y L1 ) and an encoder (612) receiving a binaural input signal of the right channel (e.g., y R1 ) may include an encoder (614) that receives input. In one embodiment, the encoders (612, 614) may exist as many as the number of microphone channels corresponding to the left channel and the right channel, respectively.
[0109] Encoders (612, 614) can convert a binaural input signal in the time domain into a binaural input signal in the frequency domain (or latency domain). In one embodiment, encoders (612, 614) can be implemented as modules or convolutional layers that perform a short-time Fourier transform (STFT) on the binaural input signal. The binaural input signal y in the time domain is converted by encoders (612, 614). L1 , y R1 A frequency spectrum signal can be generated for .
[0110] After binaural input signals in the frequency domain are output from the encoders (612, 614), a concatenation operation (620) may be performed on the binaural input signals in the frequency domain. In one embodiment, the binaural input signals in the frequency domain may be connected or combined with each other to form a single signal through the concatenation operation (620).
[0111] The mask estimator (630) can determine a mask value to be applied to a binaural input signal in the frequency domain based on a given noise suppression parameter. Here, the noise suppression parameter is a noise suppression parameter used when determining the target binaural signal in Equations 3 and 4. and The mask estimator (630) may multiply the mask value to be multiplied to the left reference channel based on the noise suppression parameter. and the mask value to be multiplied by the right reference channel can output the mask value class These are values that determine how much to reduce the influence of noise components in the binaural input signal, and can be a complex valued mask or a real valued mask. Mask values class For example, it can be composed of values in the range of 0 to 1. In one embodiment, temporal convolutional networks (TCN) or dual path recurrent neural networks (DPRNN) can be used as the mask estimator (630), but is not limited to the examples described above.
[0112] A mask value determined by a mask estimator (630) on a binaural input signal in the frequency domain output from encoders (612, 614) class Each of these can be applied. For example, in operation (642), a mask value is applied to the binaural input signal in the frequency domain output from the encoder (612). This is multiplied and the mask value is applied to the binaural input signal in the frequency domain output from the encoder (614) in operation (644). This can be multiplied.
[0113] The decoders (652, 654) can convert the result signal with the mask value applied to the binaural input signal in the frequency domain back into a binaural output signal in the time domain and output it. The decoders (652, 654) can output the binaural output signal of the left channel. A decoder (652) outputting a binaural output signal of the right channel It may include a decoder (654) that outputs. In one embodiment, the decoders (652, 654) may be implemented as a module or a transposed convolutional layer that performs an inverse short-time Fourier transform (ISTFT) on the result signal. The binaural output signal in the time domain on which binaural speech enhancement processing is performed by the decoders (652, 654) , and can be obtained. Binaural output signal , and is the binaural input signal y L1 , y R1 It can be a signal with improved SNR and preserved spatial characteristics of noise components compared to the original signal.
[0114] In the learning operation of the voice enhancement model (600) according to one embodiment, the binaural output signal , and and the target binaural signal defined in Equations 3 and 4. , and The loss can be determined based on the speech to distortion ratio loss, the loss based on the inter-channel level difference, and the loss based on the inter-channel time difference. At least one of the speech to distortion ratio loss can be used for training the speech enhancement model (600). The speech to distortion ratio loss can be determined based on, for example, the binaural output signal. and target binaural signals Differences between and / or binaural output signals and target binaural signals can be determined based on the difference between the channels. Loss based on the level difference between channels is, for example, a binaural output signal. and signal level difference between the two and / or the target binaural signal and can be determined based on the signal level difference between the channels. Loss based on the time difference between channels is, for example, a binaural output signal. and Timing differences between the interoceptive and / or target binaural signals and It can be determined based on the timing difference between the two.
[0115] In one embodiment, a loss function may be determined based on the sum of a speech-to-distortion ratio loss, a loss based on a level difference between channels, and a loss based on a time difference between channels, and an operation of updating the parameters of a speech enhancement model (600) based on the loss function may be performed. For example, the parameters of the speech enhancement model (600) may be updated in a direction in which the value of the loss function decreases through a machine learning algorithm based on error backpropagation.
[0116] Figure 7 is a flowchart illustrating an exemplary method of operating an electronic device for training a voice enhancement model according to various embodiments. In one embodiment, at least one of the operations in Figure 7 may be performed concurrently or in parallel with other operations, and the order of the operations may be changed. Furthermore, at least one of the operations may be omitted, and other operations may be additionally performed.
[0117] Referring to FIG. 7, in operation (710), the electronic device (101) may select learning data to be used for learning from a learning database (e.g., the learning database (410) of FIG. 4). For example, the electronic device (101) may select speech data, spatial impulse response data, and noise data from the learning database. In one embodiment, the noise data may include interference noise data and diffuse noise data. The spatial impulse response data may include spatial impulse response data from a sound image location of the speech data to a binaural device (e.g., 300) and spatial impulse response data from a sound image location of the interference noise data to a binaural device (e.g., the binaural devices (310, 320) of FIG. 3 or the binaural device (400) of FIG. 4).
[0118] In operation (720), the electronic device (101) may determine a binaural input signal based on selected learning data, for example, the selected speech data, the selected spatial impulse response data, and the selected noise data. In one embodiment, the electronic device (101) may determine the binaural input signal by applying the selected speech data, the selected spatial impulse response data, and the selected noise data to a defined mathematical model (e.g., Equation 2).
[0119] In operation (730), the electronic device (101) can obtain a binaural output signal from a neural network-based voice enhancement model (e.g., the voice enhancement model (420) of FIG. 4 or the voice enhancement model (600) of FIG. 6) that inputs a binaural input signal. In one embodiment, the electronic device (101) can obtain a binaural output signal on which binaural voice enhancement processing has been performed from the voice enhancement model (600) by inputting a binaural input signal into the voice enhancement model (600) described in FIG. 6.
[0120] In operation (740), the electronic device (101) may determine a target binaural signal based on the selected learning data, for example, the speech data selected in operation (710), the selected spatial impulse response data, and the selected noise data. The target binaural signal may include a target binaural signal for a left channel and a target binaural signal for a right channel. The electronic device (101) may determine the target binaural signal for the left channel based on the spatial impulse response data for the left channel and the noise data for the left channel. The electronic device (101) may determine the target binaural signal for the right channel based on the spatial impulse response data for the right channel and the noise data for the right channel. In one embodiment, the electronic device (101) may determine the target binaural signal by applying the selected speech data, the spatial impulse response data for the selected speech data, the selected noise data, the spatial impulse response data for the selected noise data, and a noise suppression parameter to a defined mathematical model (e.g., Equations 3 and 4).
[0121] In operation (750), the electronic device (101) may update parameters of the voice enhancement model based on the binaural output signal obtained from the voice enhancement model and the determined target binaural signal. In one embodiment, the electronic device (101) may determine a loss based on the binaural output signal obtained from the voice enhancement model and the determined target binaural signal, and control to update parameters of the voice enhancement model based on the determined loss.
[0122] In one embodiment, the electronic device (101) may update parameters of the speech enhancement model based on at least one of a speech-to-distortion ratio loss, a loss based on a level difference between channels, and a loss based on a time difference between channels based on a binaural output signal obtained from the speech enhancement model and the determined target binaural signal. In one embodiment, the electronic device (101) may perform an update on parameters of the speech enhancement model using an error backpropagation-based machine learning algorithm. The parameters of the speech enhancement model may be adjusted so that the binaural output signal output from the speech enhancement model becomes similar to the target binaural signal.
[0123] FIG. 8 is a flowchart illustrating exemplary operations for determining a binaural input signal and a target binaural signal for training a voice enhancement model according to various embodiments. In one embodiment, at least one of the operations in FIG. 8 may be performed concurrently or in parallel with other operations, and the order of the operations may be changed. Furthermore, at least one of the operations may be omitted, and other operations may be additionally performed.
[0124] Referring to FIG. 8, in operation (810), the electronic device (101) may select (or sample) voice data and interference noise data from a sound source database (e.g., the sound source database (510) of FIG. 5). The sound source database may store various voice data and various interference noise data that are targets of binaural voice enhancement. In one embodiment, the voice data and interference noise data may be randomly selected from the sound source database (510).
[0125] In operation (820), the electronic device (101) may select (or sample) spatial impulse response data from a spatial impulse response database (e.g., the spatial impulse response database (520) of FIG. 5). The spatial impulse response database may store spatial impulse response data representing spatial impulse responses in various environments. In one embodiment, the electronic device (101) may select spatial impulse response data to be applied to speech data and spatial impulse response data to be applied to interference noise data from the spatial impulse response database.
[0126] In operation (830), the electronic device (101) may select (or sample) background noise data from a background noise database (e.g., the background noise database (530) of FIG. 5). The background noise database may store various diffuse noise data.
[0127] Actions (810), (820) and (830) may be included in action (710) of FIG. 7.
[0128] In operation (840), the electronic device (101) may select a noise suppression parameter to determine the degree of noise suppression for the binaural input signal. For example, the electronic device (101) may select any value between 0 and 1 as the noise suppression parameter. The noise suppression parameter may include a noise suppression parameter for determining the degree of suppression of interference noise components and a noise suppression parameter for determining suppression information of diffusion noise components.
[0129] In operation (850), the electronic device (101) can determine a binaural input signal and a target binaural signal based on the selected speech data, the selected interference noise data, the selected spatial impulse response data, the selected background noise data, and the selected noise suppression parameter. In one embodiment, the electronic device (101) can determine a binaural input signal input to the speech enhancement model through the above mathematical expression 2, and can determine a target binaural signal through, for example, mathematical expressions 3 and 4.
[0130] Action (840) and action (850) may include action (720) and action (740) of FIG. 7.
[0131] FIG. 9 is a flowchart illustrating exemplary operations for obtaining diffusion noise data from a recorded audio signal according to various embodiments. In one embodiment, at least one of the operations in FIG. 9 may be performed concurrently or in parallel with other operations, and the order of the operations may be changed. Furthermore, at least one of the operations may be omitted, and other operations may be additionally performed.
[0132] Referring to FIG. 9, in operation (910), a binaural device (e.g., a binaural device (310, 320) of FIG. 3 or a binaural device (400) of FIG. 4) may record an audio signal including a background sound. The background sound may include a noise signal, and the noise signal may include a directional interference noise signal and / or a non-directional diffuse noise signal. The recorded signal may include an audio signal for a right channel and an audio signal for a left channel. The audio signal for the left channel may include an audio signal recorded by microphones (350, 355) of the binaural device (310) of FIG. 3, for example, and the audio signal for the right channel may include an audio signal recorded by microphones (350-1, 355-1) of the binaural device (320) of FIG. 3.
[0133] In operation (920), the electronic device (101) can extract a random section from an audio signal recorded by a binaural device. For example, a partial audio signal at a specific time interval can be randomly extracted from the entire time interval of the recorded audio signal. For example, a partial audio signal at the same time interval can be extracted between an audio signal for a left channel and an audio signal for a right channel.
[0134] In operation (930), the electronic device (101) may determine inter-channel coherence (or inter-channel coherence value) between the extracted partial audio signal for the left channel and the partial audio signal for the right channel. The inter-channel coherence may indicate a correlation between the left channel and the right channel. The more similar the partial audio signal for the left channel and the partial audio signal for the right channel are to each other, the greater the inter-channel coherence may be.
[0135] In operation (940), the electronic device (101) can determine whether the determined inter-channel consistency is greater than or equal to a threshold value. If the determined inter-channel consistency is greater than or equal to the threshold value ('yes' in operation (940)), in operation (950), the electronic device (101) can store an audio signal of a random section extracted in operation (920) as a noise signal. The noise signal can be stored in a background noise database (e.g., a background noise database (530) of FIG. 5) and can be used as diffuse noise data used for training a voice enhancement model (e.g., a voice enhancement model (420) of FIG. 4 or a voice enhancement model (600) of FIG. 6).
[0136] If the determined inter-channel consistency is below the threshold ('No' in operation (940)), the electronic device (101) may return to operation (920) to extract a partial audio signal of another arbitrary section from the audio signal recorded by the binaural device, and perform the operations from operation (930) again.
[0137] Figure 10 is a flowchart illustrating an exemplary method for providing an audio signal using a voice enhancement model according to various embodiments. In one embodiment, at least one of the operations in Figure 10 may be performed concurrently or in parallel with other operations, and the order of the operations may be changed. Furthermore, at least one of the operations may be omitted, and other operations may be additionally performed.
[0138] Referring to FIG. 10, in operation (1010), an electronic device (101) can obtain an audio signal from a binaural device (e.g., a binaural device (310, 320) of FIG. 3 or a binaural device (400) of FIG. 4). The binaural device can obtain multi-channel audio signals through equipped multi-channel microphones and transmit the obtained multi-channel audio signals to the electronic device (101).
[0139] In operation (1020), the electronic device (101) may input a multi-channel audio signal received from a binaural device into a learned speech enhancement model. In one embodiment, the speech enhancement model may be learned using the learning method described with reference to FIGS. 4 to 8.
[0140] In operation (1030), the electronic device (101) may obtain an audio signal processed by a speech enhancement model. The speech enhancement model may have a structure, for example, such as the speech enhancement model (600) of FIG. 6. The speech enhancement model may perform binaural speech enhancement processing on an input multi-channel audio signal (e.g., a binaural input signal) to provide a binaural output signal in which noise is reduced while preserving spatial characteristics of the noise. The binaural output signal may include a binaural output signal for a left channel and a binaural output signal for a right channel.
[0141] In operation (1040), the electronic device (101) can output an audio signal processed by a voice enhancement model through a binaural device. The electronic device (101) can provide a user with an audio signal with an improved SNR while maintaining the acoustic impression of a voice signal, interference noise, and diffusion noise through binaural voice enhancement processing using the voice enhancement model as described above.
[0142] FIG. 11 is a diagram for explaining quality improvement of an audio signal by binaural voice enhancement processing according to various embodiments.
[0143] Referring to FIG. 11, a case (1110) in which binaural voice enhancement processing is not performed and a case (1150) in which binaural voice enhancement processing as exemplified in the present disclosure is performed are illustrated.
[0144] According to case (1110), an audio signal including a target voice component (1120), an interference noise component (1130), and a diffusion noise component (1140) is provided to a user through a binaural device (310, 320). Since the audio signal provided to the user does not undergo reduction processing for the interference noise component (1130) and the diffusion noise component (1140), it may be difficult for the user to recognize the target voice component (1120) from the audio signal.
[0145] In a case (1150) according to one embodiment, an audio signal including a target speech component (1120), an interference noise component (1135), and a diffuse noise component (1145) is provided to a user through a binaural device (310, 320). However, by binaural speech enhancement processing using a speech enhancement model, the interference noise component (1135) and the diffuse noise component (1145) can be reduced compared to the interference noise component (1130) and the diffuse noise component (1140) in the case (1110). Therefore, an audio signal (or speech signal) with an improved signal-to-noise ratio compared to the case (1110) can be provided to the user. Even when binaural speech enhancement processing is performed, the spatial cues of the interference noise component (1135) and the diffuse noise component (1145) can be preserved (e.g., sound image position can be preserved) and provided. Even after binaural speech enhancement processing is performed, the user can still perceive the directionality of the interference noise component (1135) and the non-directivity of the diffusion noise component (1145).
[0146] An electronic device according to one embodiment includes at least one processor including a processing circuit and at least one memory storing instructions executable by the at least one processor, wherein the at least one processor individually and / or collectively executes the instructions to cause the electronic device to select speech data, spatial impulse response data and noise data from a learning database (410), determine a binaural input signal based on the selected speech data, the selected spatial impulse response data and the selected noise data, obtain a binaural output signal from a neural network-based speech enhancement model (420; 600) having the binaural input signal as an input, determine a target binaural signal based on the selected speech data, the selected spatial impulse response data and the selected noise data, and update a parameter of the speech enhancement model (420; 600) based on the binaural output signal obtained from the speech enhancement model (420; 600) and the determined target binaural signal.
[0147] In one embodiment, at least one processor may individually and / or collectively cause the electronic device to determine the target binaural signal based on the selected speech data, spatial impulse response data for the selected speech data, the selected noise data, spatial impulse response data for the selected noise data, and a noise suppression parameter related to a degree of noise suppression.
[0148] In one embodiment, the noise suppression parameter may be randomly selected within a range of specified values.
[0149] In one embodiment, the target binaural signal may include a target binaural signal for a left channel and a target binaural signal for a right channel. The target binaural signal for the left channel may be determined based on spatial impulse response data for the left channel and noise data for the left channel. The target binaural signal for the right channel may be determined based on spatial impulse response data for the right channel and noise data for the right channel.
[0150] In one embodiment, the speech enhancement model may include an encoder that converts the binaural input signal in the time domain into a binaural input signal in the frequency domain, a mask estimator that determines a mask value to be applied to the binaural input signal in the frequency domain based on a noise suppression parameter, and a decoder that converts a result signal in which the determined mask value is applied to the binaural input signal in the frequency domain into a binaural output signal in the time domain.
[0151] In one embodiment, at least one processor may individually and / or collectively cause the electronic device to determine a loss based on a binaural output signal obtained from the speech enhancement model and the determined target binaural signal, and to update parameters of the speech enhancement model based on the loss.
[0152] In one embodiment, at least one processor may individually and / or collectively cause the electronic device to update parameters of the speech enhancement model based on at least one of a speech-to-distortion ratio loss, a loss based on a level difference between channels, and a loss based on a time difference between channels based on a binaural output signal obtained from the speech enhancement model and the determined target binaural signal.
[0153] In one embodiment, the voice enhancement model may include a model that outputs the binaural output signal by reducing the intensity of the noise component included in the binaural input signal while preserving the directionality of the noise component.
[0154] In one embodiment, the noise data may include interference noise data and diffusion noise data.
[0155] In one embodiment, the spatial impulse response data may include spatial impulse response data from a sound location of the speech data to a binaural device and spatial impulse response data from a sound location of the interference noise data to the binaural device.
[0156] A method of operating an electronic device according to one embodiment may include selecting speech data, spatial impulse response data and noise data from a learning database, determining a binaural input signal based on the selected speech data, the selected spatial impulse response data and the selected noise data, obtaining a binaural output signal from a neural network-based speech enhancement model that has the binaural input signal as an input, determining a target binaural signal based on the selected speech data, the selected spatial impulse response data and the selected noise data, and updating parameters of the speech enhancement model based on the binaural output signal obtained from the speech enhancement model and the determined target binaural signal.
[0157] In one embodiment, the operation of determining the target binaural signal may include an operation of determining the target binaural signal based on the selected speech data, spatial impulse response data for the selected speech data, the selected noise data, spatial impulse response data for the selected noise data, and a noise suppression parameter related to a degree of noise suppression.
[0158] In one embodiment, the operation of updating the parameters of the voice enhancement model may include the operation of determining a loss based on a binaural output signal obtained from the voice enhancement model and the determined target binaural signal, and the operation of updating the parameters of the voice enhancement model based on the loss.
[0159] In one embodiment, the operation of updating the parameters of the speech enhancement model may include the operation of updating the parameters of the speech enhancement model based on at least one of a speech-to-distortion ratio loss, a loss based on a level difference between channels, and a loss based on a time difference between channels based on a binaural output signal obtained from the speech enhancement model and the determined target binaural signal.
[0160] The various embodiments of the present disclosure and the terminology used therein are not intended to limit the technical features described in the present disclosure to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In the present disclosure, each of the phrases "A or B," "at least one of A and B," "at least one of A or B," "A, B, or C," "at least one of A, B, and C," and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among the phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0161] The term "module" used in various embodiments of the present disclosure may include a unit implemented in hardware, software, or firmware, or any combination thereof, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0162] Various embodiments of the present disclosure may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101) of FIG. 1). For example, a processor (e.g., a processor (120)) of a machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, a 'non-transitory' storage medium is a tangible device and may not contain signals (e.g., electromagnetic waves), but the term does not distinguish between cases where data is stored semi-permanently on the storage medium and cases where it is stored temporarily.
[0163] According to one embodiment, the method according to various embodiments disclosed in the present disclosure may be provided as included in a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0164] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0165] While the present disclosure has been illustrated and described with reference to various embodiments, it is to be understood that the various embodiments are intended to be illustrative rather than limiting. It will be further understood by those skilled in the art that various changes in form and detail may be made without departing from the true spirit and scope of the present disclosure, including the appended claims and their equivalents. It will also be understood that any embodiment(s) described herein may be used in conjunction with any other embodiment(s) described herein.
Claims
1. In an electronic device (101), At least one processor (120) comprising a processing circuit; and comprising one or more memories (130) storing instructions executable by at least one processor (120); At least one processor (120) individually and / or collectively executes the instructions to cause the electronic device (101) to: Select voice data, spatial impulse response data and noise data from the learning database (410), Determine a binaural input signal based on the selected voice data, the selected spatial impulse response data, and the selected noise data, Obtain a binaural output signal from a neural network-based voice enhancement model (420; 600) that inputs the above binaural input signal, Determine a target binaural signal based on the selected voice data, the selected spatial impulse response data, and the selected noise data, To update the parameters of the voice enhancement model (420; 600) based on the binaural output signal obtained from the voice enhancement model (420; 600) and the determined target binaural signal. Electronic device (101).
2. In paragraph 1, The at least one processor (120) individually and / or collectively causes the electronic device (101) to: Determine the target binaural signal based on the selected voice data, spatial impulse response data for the selected voice data, the selected noise data, spatial impulse response data for the selected noise data, and a noise suppression parameter related to the degree of noise suppression. Electronic device (101).
3. In paragraph 2, The above noise suppression parameters are, which is randomly selected within a range of specified values, Electronic device (101).
4. In any one of paragraphs 1 to 3, The above target binaural signal is, Contains a target binaural signal for the left channel and a target binaural signal for the right channel, The target binaural signal for the above left channel is: It is determined based on the spatial impulse response data for the left channel and the noise data for the left channel, The target binaural signal for the above right channel is: Determined based on the spatial impulse response data for the right channel and the noise data for the right channel, Electronic device (101).
5. In any one of paragraphs 1 to 4, The above voice enhancement model (420; 600) is An encoder (612; 614) that converts the binaural input signal in the time domain into a binaural input signal in the frequency domain; A mask estimator (630) for determining a mask value to be applied to a binaural input signal in the frequency domain based on a noise suppression parameter; and A decoder (652; 654) that converts a result signal in which the determined mask value is applied to a binaural input signal in the frequency domain into a binaural output signal in the time domain. An electronic device (101) comprising:
6. In any one of paragraphs 1 to 5, The instructions executed above cause the electronic device (101) to: The loss is determined based on the binaural output signal obtained from the above voice enhancement model (420; 600) and the determined target binaural signal, Controlling to update the parameters of the voice enhancement model (420; 600) based on the above loss, Electronic device (101).
7. In paragraph 6, The at least one processor (120) individually and / or collectively causes the electronic device (101) to: Update the parameters of the speech enhancement model (420; 600) based on at least one of a speech-to-distortion ratio loss, a loss based on a level difference between channels, and a loss based on a time difference between channels based on the determined target binaural signal and a binaural output signal obtained from the speech enhancement model (420; 600). Electronic device (101).
8. In any one of paragraphs 1 to 7, The above voice enhancement model (420; 600) is A model that outputs the binaural output signal while preserving the directionality of the noise component included in the binaural input signal and reducing the intensity of the noise component, Electronic device (101).
9. In any one of paragraphs 1 to 8, The above noise data is, Containing interference noise data and diffuse noise data, Electronic device (101).
10. In paragraph 9, The above spatial impulse response data is, Including spatial impulse response data from the sound position of the above voice data to the binaural device (400) and spatial impulse response data from the sound position of the above interference noise data to the binaural device (400). Electronic device (101).
11. In a method of operating an electronic device (101), An operation of selecting voice data, spatial impulse response data and noise data from a learning database (410); An operation of determining a binaural input signal based on the selected voice data, the selected spatial impulse response data and the selected noise data; An operation of obtaining a binaural output signal from a neural network-based voice enhancement model (420; 600) that inputs the above binaural input signal; An operation of determining a target binaural signal based on the selected voice data, the selected spatial impulse response data, and the selected noise data; and An operation of updating the parameters of the voice enhancement model (420; 600) based on the binaural output signal obtained from the voice enhancement model (420; 600) and the determined target binaural signal. How to include.
12. In paragraph 11, The operation of determining the above target binaural signal is as follows: An operation of determining the target binaural signal based on the selected voice data, spatial impulse response data for the selected voice data, the selected noise data, spatial impulse response data for the selected noise data, and a noise suppression parameter related to the degree of noise suppression. How to include.
13. In either of paragraphs 11 or 12, The above voice enhancement model (420; 600) is An encoder (612; 614) that converts the binaural input signal in the time domain into a binaural input signal in the frequency domain; A mask estimator (630) for determining a mask value to be applied to a binaural input signal in the frequency domain based on a noise suppression parameter; and A decoder (652; 654) that converts a result signal in which the determined mask value is applied to a binaural input signal in the frequency domain into a binaural output signal in the time domain. How to include.
14. In any one of paragraphs 11 to 13, The operation of updating the parameters of the above voice enhancement model (420; 600) is as follows: An operation of determining a loss based on the binaural output signal obtained from the above-determined voice enhancement model (420; 600) and the determined target binaural signal; and An operation of updating the parameters of the voice enhancement model (420; 600) based on the above loss. How to include.
15. In paragraph 14, The operation of updating the parameters of the above voice enhancement model (420; 600) is as follows: An operation of updating parameters of the speech enhancement model (420; 600) based on at least one of a speech-to-distortion ratio loss, a loss based on a level difference between channels, and a loss based on a time difference between channels based on a binaural output signal obtained from the speech enhancement model (420; 600) and the determined target binaural signal. How to include.
Citation Information
Patent Citations
Method and device for removing noise using neural network model
KR1020180111271A
Method and terminal for reconstructing speech signal, and computer storage medium
US20200251124A1
Hearing device comprising a speech presence probability estimator
US20210352415A1
Fully customizable ear worn devices and associated development platform
US20230300532A1
Speech enhancement
WO2023287773A1