Electronic device and method for improving voice recognition rate of audio signal

The method enhances voice recognition rates in mobile devices by processing signals from internal and external microphones using beamforming and machine learning, addressing the challenge of noise and distortion in external audio devices during voice calls.

WO2026095467A1PCT designated stage Publication Date: 2026-05-07SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing voice recognition systems in mobile devices struggle to maintain high voice recognition rates during voice calls, especially in noisy environments, due to the challenge of balancing noise reduction and voice distortion when using external audio devices.

Method used

A method involving a processor that acquires and processes microphone signals from both internal and external microphones, applying techniques like beamforming, mixing parameters, and machine learning models to enhance voice signals, optimizing for both voice recognition rate and call quality.

Benefits of technology

Improves voice recognition rates during real-time translation functions by effectively reducing noise and maintaining stable call quality, even in challenging acoustic conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025016737_07052026_PF_FP_ABST
    Figure KR2025016737_07052026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to various embodiments of the present document may comprise at least one processor and a memory. The memory may store instructions which may be executed by the at least one processor and, when executed, instruct the electronic device to: acquire a first microphone signal and a second microphone signal, both of which include a user voice; generate a first mixed signal by mixing, on the basis of a first mixed parameter, a first input signal which is generated on the basis of the first microphone signal and a second input signal which is generated on the basis of the second microphone signal; generate a noise-canceled signal by canceling noise from the first mixed signal; generate a second mixed signal by mixing, on the basis of a second mixed parameter, the noise-canceled signal and the first mixed signal; and output the second mixed signal. Various other embodiments are possible.
Need to check novelty before this filing date? Find Prior Art

Description

Method for improving the speech recognition rate of electronic devices and audio signals

[0001] This document relates to an electronic device and, for example, to a method for improving the voice recognition rate of an audio signal including a user voice acquired through a microphone.

[0002] Mobile devices, such as smartphones, can perform audio functions, such as voice calls, by utilizing external audio devices capable of providing audio input and output capabilities. During voice calls, background noise can degrade voice quality and clarity. Therefore, voice enhancement technology is required that can preserve the voice as much as possible while eliminating noise acquired alongside the user's voice signal. When a mobile device is conducting a voice call through an external audio device, consideration of the environmental and structural conditions of the external audio device may be necessary.

[0003] Mobile devices can provide voice recognition capabilities during voice calls. For example, a function may be required to extract the user's voice from an incoming audio signal and convert it into text information, such as when performing real-time translation during a call. As such, when a mobile device requires real-time voice recognition during a voice call, it is necessary to process the voice signal in a way that improves the voice recognition rate while providing stable call quality.

[0004] Call solutions using external audio devices are designed to secure high-quality voice by minimizing voice distortion and residual noise in the input audio signal. Since such call solutions do not consider the voice recognition rate, it may be difficult to provide a high voice recognition rate in high-noise environments. Furthermore, if the call solution is designed to strongly remove noise from the audio signal, the voice recognition rate may be lower than when the audio signal is passed through the voice recognition device as is due to the accompanying voice distortion.

[0005] An electronic device according to various embodiments of this disclosure (or specification, invention) may include at least one processor and memory.

[0006] According to one embodiment, the memory may be executed by at least one processor, and the electronic device may store instructions for acquiring a first microphone signal and a second microphone signal including a user's voice, mixing a first input signal generated based on the first microphone signal and a second input signal generated based on the second microphone signal based on a first mixing parameter to generate a first mixing signal, removing noise from the first mixing signal to generate a noise removal signal, mixing the noise removal signal and the first mixing signal based on a second mixing parameter to generate a second mixing signal, and outputting the second mixing signal.

[0007] A method performed by an electronic device according to various embodiments of the present document may include: an operation of acquiring a first microphone signal and a second microphone signal including a user's voice; an operation of mixing a first input signal generated based on the first microphone signal and a second input signal generated based on the second microphone signal based on a first mixing parameter to generate a first mixing signal; an operation of removing noise from the first mixing signal to generate a noise removal signal; an operation of mixing the noise removal signal and the first mixing signal based on a second mixing parameter to generate a second mixing signal; and instructions for outputting the second mixing signal.

[0008] According to various embodiments of the present document, when a real-time translation function is executed during a voice call using a microphone of an external audio device, an electronic device and a method for improving the voice recognition rate of an audio signal can be provided, which can improve the voice recognition rate while maintaining call quality.

[0009] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments.

[0010] FIG. 2 illustrates a mobile device and an audio device according to one embodiment.

[0011] FIG. 3 is a block diagram of a mobile device according to one embodiment.

[0012] FIG. 4 is a block diagram of an audio device according to one embodiment.

[0013] FIG. 5 is a block diagram showing the process of an electronic device processing an audio signal according to one embodiment.

[0014] FIG. 6 is a block diagram showing the process of an electronic device processing an audio signal according to one embodiment.

[0015] FIG. 7 is a flowchart of a method for enhancing a voice signal of an electronic device according to one embodiment.

[0016] FIG. 8 illustrates models for determining mixing parameters according to one embodiment.

[0017] FIGS. 9A, FIGS. 9B, and FIGS. 9C illustrate a foldable device according to one embodiment.

[0018] FIGS. 10a, FIGS. 10b, FIGS. 10c and FIGS. 10d illustrate a multi-foldable device according to one embodiment.

[0019] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.

[0020] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments.

[0021] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).

[0022] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.

[0023] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0024] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).

[0025] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0026] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0027] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0028] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0029] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).

[0030] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0031] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0032] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0033] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0034] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0035] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0036] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0037] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).

[0038] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.

[0039] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).

[0040] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0041] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0042] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0043] FIG. 2 illustrates a mobile device and an audio device according to one embodiment.

[0044] According to one embodiment, a mobile device (300) (e.g., the electronic device (101) of FIG. 1) can provide various audio experiences such as voice calls, real-time translation, media playback, voice commands, and / or listening to ambient sounds by using an external audio device (400).

[0045] In FIG. 2, the audio device (400) is depicted as an earbud, but is not limited thereto, and various types of devices that provide input and / or output of audio signals, such as earphones, headsets, or wireless microphones, may be implemented as the audio device (400) of this document. In FIG. 2, one audio device (400) is depicted, but two audio devices worn on the user's left and right ears, respectively, may be configured as a set.

[0046] According to one embodiment, the audio device (400) may include an audio output unit and at least one microphone (480, 492, 494). The audio output unit may be in a protruding shape so as to be inserted into a user's ear, and sound generated and amplified by the audio device (400) may be audibly perceived by the user through the audio output unit. The microphones (480, 492, 494) may recognize ambient sounds including the user's voice.

[0047] According to one embodiment, the audio device (400) may include an inner microphone (480) positioned near the user's ear when worn by the user, and outer microphones (492, 494) positioned opposite the user's ear. In FIG. 2, the audio device (400) is illustrated as including one inner microphone (480) and two outer microphones (494, 494), but the number and / or location of the inner microphone (480) and outer microphones (492, 494) included in the audio device (400) are not limited thereto. For example, the audio device (400) may include three or more outer microphones.

[0048] According to one embodiment, the internal microphone signal obtained from the internal microphone (480) and the external microphone signal obtained from the external microphone (492, 494) may have different audio characteristics. For example, when the internal microphone signal is worn by a user, it may be physically isolated from the external space and noise may be shielded by contact with the user's ear. Accordingly, the internal microphone signal may have a relatively high signal-to-noise ratio (SNR) compared to the external microphone signal, and the signal may pass through the low-frequency range, allowing voice to be distributed in the low-band range. Conversely, the external microphone signal may be placed in a position open to the outside, allowing a large amount of noise to enter, resulting in a relatively low signal-to-noise ratio, and voice may be distributed across the entire band.

[0049] According to one embodiment, a mobile device (300) and an audio device (400) can transmit and receive audio signals via near-field wireless communication (e.g., Bluetooth, Wi-Fi Direct, NFC). For example, in a voice call scenario, the audio device (400) can transmit the user's voice obtained through a microphone to the mobile device (300) via near-field wireless communication, and the mobile device (300) can transmit the voice of the call partner received from a network to the audio device (400) so that it is output through an audio output unit. In addition, the mobile device (300) and the audio device (400) can transmit and receive various data via near-field wireless communication, such as battery status, pairing or connection status, touch or gesture input on the audio device (400), location information, and / or firmware updates, in addition to audio signals.

[0050] According to one embodiment, a mobile device (300) or an audio device (400) may perform an operation (or voice enhancement operation) to enhance a voice signal from an audio signal (or microphone signal) obtained from a microphone of the audio device (400) in various usage scenarios. For example, the operation to enhance the audio signal may include, but is not limited to, pre-processing, beamforming, noise removal processing using a machine learning model, and post-processing. Various filtering and / or amplification techniques may be used in the operation to enhance the audio signal. The operations to enhance the audio signal during a voice call may be defined as a call solution.

[0051] In this document, a series of operations for enhancing an audio signal (e.g., preprocessing, beamforming, noise removal processing using a machine learning model, postprocessing) is described as being performed by an electronic device, and the electronic device may mean a mobile device (300) or an audio device (400). In some embodiments, at least some of the series for enhancing the audio signal may be performed by the audio device (400) and the remaining parts may be performed by the mobile device (300).

[0052] According to one embodiment, when making a voice call through the microphone of the audio device (400) (e.g., internal microphone (480), external microphone (492, 494)), various environmental and structural conditions may need to be considered. For example, when making a voice call using the mobile device (300) without the external audio device (400), the user's voice can be transmitted directly to the microphone of the mobile device (300) at a short distance, but when making a voice call using the audio device (400), the distance between the user's mouth and the microphone is long, resulting in a poor signal-to-noise ratio environment, and the distance between the audio output and the microphone is short, resulting in exposure to strong energy echoes and / or wind noise may be introduced into the microphone.

[0053] According to one embodiment, the mobile device (300) may provide a real-time translation function during a voice call. The real-time translation function may include recognizing the voice of the user of the mobile device (300) and / or the call partner, converting it into text, and converting the converted text into another language. In order to provide the real-time translation function, it may be necessary to accurately recognize the user's voice included in the audio signal obtained through the microphone.

[0054] According to one embodiment, when applying voice enhancement technology to an audio signal, there is a risk of voice loss if noise reduction is set too high, and a concern that residual noise may remain if noise reduction is set too low; therefore, it is necessary to find appropriate values ​​when designing parameters used for voice enhancement. In a general call scenario, the algorithm can be designed to increase voice preservation and reduce the amount of residual noise; however, in a real-time translation scenario during a voice call, it is necessary to design the algorithm in a way that maximizes the voice recognition rate while providing stable call quality.

[0055] Hereinafter, various embodiments for improving translation recognition rates during voice calls using an external audio device (400) will be described through FIGS. 3 to 8.

[0056] FIG. 3 is a block diagram of a mobile device according to one embodiment.

[0057] Referring to FIG. 3, the mobile device (300) may include a communication circuit (340), a display (350), a microphone (330), a processor (310), and a memory (320). Various embodiments of this document may be implemented even if at least some of the configurations shown in FIG. 3 are omitted or replaced with other configurations. In addition to the configurations shown, the mobile device (300) may further include at least some of the configurations and / or functions of the electronic device (101) of FIG. 1. At least some of each configuration of the mobile device (300) shown (or not shown) (e.g., processor (310), memory (320)) may be placed within the housing of the mobile device (300), and at least some of the other configurations (e.g., display (350), microphone (330)) may be visually exposed to the outside of the housing. At least some of each configuration of the mobile device (300) may be operatively, functionally, and / or electrically connected to one another.

[0058] According to one embodiment, the display (350) can display image information provided by the processor (310). The display (350) may be implemented as any one of a liquid crystal display (LCD), a light-emitting diode (LED) display, or an organic light-emitting diode (OLED) display, but is not limited thereto. The display (350) may be configured as a touch screen that detects touch and / or proximity touch (or hovering) input using a part of the user's body (e.g., finger) or an input device (e.g., stylus pen). The display (350) may include at least some of the configuration and / or functions of the display module (160) of FIG. 1.

[0059] According to one embodiment, the display (350) may be a flexible display in which at least a portion is flexible. The mobile device (300) may be implemented in various form factors, such as a foldable device or a rollable device, in which the size of the display area can be changed by utilizing the characteristics of the flexible display. An example in which the mobile device (300) is implemented as a foldable device or a multi-foldable device will be described in more detail through FIGS. 9 and FIGS. 10.

[0060] According to one embodiment, the communication circuit (340) may include various configurations to support wireless communication with various external devices. For example, a mobile device (300) may perform cellular wireless communication (e.g., 4G LTE (long term evolution), 5G NR (new radio)) and / or short-range wireless communication (e.g., Wi-Fi, Bluetooth) through the communication circuit (340), and there is no fixed type of wireless communication supported by the mobile device (300). The communication circuit (340) may include at least some of the configurations and / or functions of the communication module (190) of FIG. 1.

[0061] According to one embodiment, a mobile device (300) can transmit and receive audio signals with an audio device via near-field wireless communication (e.g., Bluetooth, WFD (Wi-Fi Direct), NFC (near field communication)) using a communication circuit (340). In addition, the mobile device (300) and the audio device can transmit and receive various data, such as battery status, pairing or connection status, touch or gesture input on the audio device, location information, and / or firmware updates, in addition to audio signals, via near-field wireless communication.

[0062] According to one embodiment, the mobile device (300) may include at least one microphone (330). The mobile device (300) may acquire an audio signal through the microphone (330) of the mobile device (300) or the microphone (330) of the audio device.

[0063] According to one embodiment, the memory (320) may include volatile memory and non-volatile memory, and may store various data temporarily or permanently. The memory (320) may include at least some of the configuration and / or functions of the memory (130) of FIG. 1 and may store the program (140) of FIG. 1. The memory (320) may store various instructions that can be executed by the processor (310). Such instructions may include control commands such as arithmetic and logical operations, data movement, and input / output that can be recognized by the processor (310).

[0064] According to one embodiment, the processor (310) may be configured to perform operations or data processing regarding the control and / or communication of each component of the mobile device (300), and may be composed of one or more processors. The processor (310) may include at least some of the configuration and / or functions of the processor (120) of FIG. 1. Although there are no limitations on the operations and data processing functions that the processor (310) can implement on the mobile device (300), this document describes various embodiments for improving the voice recognition rate and the quality of the voice signal by using a real-time translation call solution (or voice recognition mode) when executing a real-time translation function during a voice call. The operations of the processor (310) described below may be performed by loading instructions stored in the memory (320).

[0065] In this document, the description that the processor (310) can perform a certain operation (or function, task, or operation) may be interpreted substantially as meaning that an instruction (or command, computer program) causing the mobile device (300) (or processor (310)) to perform said operation is stored in memory (320) (e.g., non-volatile memory, storage). Additionally, the description that the processor (310) can perform a certain operation may be interpreted substantially as meaning that at least one processor, without a fixed number, can perform said operation.

[0066] According to one embodiment, the processor (310) can execute a real-time translation function during a voice call. The real-time translation function may be a function that translates the user's voice signal and / or the voice signal of the call partner into a language selected by the user in real time during a voice call, provides it as text information through a display (350), and / or transmits the translated text information to the other party's device.

[0067] According to one embodiment, the processor (310) may execute a call solution for processing audio signals obtained from a microphone (330) (e.g., internal microphone (480), external microphone (492, 494) of FIG. 2) of a mobile device (300) or an external audio device (e.g., audio device (400) of FIG. 2) for smooth voice exchange with the other party during a voice call. The call solution may include processes and / or algorithms for providing voice enhancement actions such as voice quality improvement, noise removal, and / or echo suppression. In a general call scenario, the call solution should be designed to increase voice preservation and reduce residual noise, whereas in a real-time translation scenario, the call solution needs to be designed to maximize voice recognition rates while providing stable call quality. According to one embodiment, the processor (310) may process the audio signal with a real-time translation call solution (e.g., FIG. 6) when the real-time translation function is running during a voice call and the user's voice is input through an audio device, and may process the audio signal with a general call solution (e.g., FIG. 5) when the real-time translation function is not running and / or the user's voice is input through a microphone (330) of a mobile device (300). The real-time translation call solution and the general call solution may be referred to as a voice recognition mode and a general mode, respectively.

[0068] Hereinafter, an embodiment will be described regarding the processing of a user's voice signal according to a real-time translation call solution (or voice recognition mode) when a real-time translation function is executed during a voice call (or video call) with another party using an audio device by a mobile device (300).

[0069] According to one embodiment, the processor (310) can acquire a microphone signal from an audio device. For example, the audio device may include an inner microphone (e.g., inner microphone (490) of FIG. 2) that contacts the user's ear when worn by the user, and at least one outer microphone (e.g., outer microphone (492, 494) of FIG. 2) located opposite the user's ear, and may transmit the audio signal acquired from each microphone to a separate channel. The processor (310) can receive the inner microphone signal and at least one outer microphone signal acquired from the audio device in real time through a communication circuit (340). Hereinafter, the processing operation of the processor (310) is described when the mobile device (300) receives a first outer microphone signal and a second outer microphone signal from the audio device, but one channel of outer microphone signals or three or more channels of outer microphone signals may be used.

[0070] According to one embodiment, the processor (310) may generate a signal in the short-time Fourier transform domain by applying a short-time Fourier transform (STFT) to each of the internal microphone signal and at least one external microphone signal. The short-time Fourier transform divides the audio signal into short segments and represents the frequency components of each segment, and can be used for time-frequency analysis during the processing of the audio signal. According to one embodiment, the processor (310) may generate a signal in the frequency domain by performing a Fourier transform on each microphone signal instead of the short-time Fourier transform, and may perform the processing process described below using the signal in the frequency domain.

[0071] According to one embodiment, the processor (310) may perform a predetermined pre-processing on each of the short-time Fourier transform microphone signals. Here, the pre-processing is an operation to increase the accuracy of signal recognition or analysis in a subsequent processing step, and may include operations to remove unnecessary noise or / or emphasize important features in the signal. For example, the pre-processing process may include, but is not limited to, equalization, echo removal, gain adjustment, and / or high-pass filtering for each microphone signal.

[0072] According to one embodiment, the processor (310) can generate a beamforming signal by applying beamforming to a first external microphone signal and a second external microphone signal. Beamforming of audio signals is a multi-channel filtering technique that uses signals from multiple microphones to strengthen signals coming from a specific direction and suppress signals coming from other directions to reduce background noise and to process audio at a specific location clearly. The processor (310) can input preprocessed external microphone signals into a beamforming module to perform beamforming and generate a beamforming signal. If only one channel of external microphone signals is received from the audio device, the beamforming operation may be omitted. According to one embodiment, if the audio device includes three or more external microphones and transmits three or more channels of external microphone signals to the electronic device (300), the processor (310) can perform beamforming using three or more external microphone signals.

[0073] According to one embodiment, the processor (310) can generate a third mixing signal by mixing the beamforming signal with the signal prior to beamforming. For example, the processor (310) can mix the beamforming signal generated by applying beamforming to the first external microphone signal and the second external microphone signal with the signal having a higher SNR among the first external microphone signal and the second external microphone signal.

[0074] When performing beamforming in a beamforming module, data such as the steering vector, relative transfer function (RTF), noise covariance matrix, speech covariance matrix, or power spectral density (PSD) may be incorrectly estimated. In this case, distortion may occur in the beamforming signal, and such distortion can lead to a decrease in speech recognition rate. The signal prior to beamforming (e.g., the first external microphone signal) has a relatively high SNR but may experience speech distortion. By mixing the signal prior to beamforming with the previous signal, signal distortion is reduced, which can improve the speech recognition rate.

[0075] According to one embodiment, the processor (310) can mix the beamforming signal and the first external microphone signal (or the second external microphone signal) based on a third mixing parameter. Here, the third mixing parameter includes a weight applied to the beamforming signal and the first external microphone signal during mixing and may have a value of 0 to 1. The third mixing parameter may be determined by considering the SNR and voice distortion of the microphone signals.

[0076] According to one embodiment, the processor (310) can generate a first mixing signal by mixing a first input signal generated based on an internal microphone signal and a second input signal generated based on a first external microphone signal based on a first mixing parameter. Here, the first input signal is a signal generated by performing a predetermined preprocessing process on a signal obtained by performing a short-time Fourier transform (STFT) on the internal microphone signal, and may be a signal obtained by mixing the signals before and after beamforming for the preprocessed first external microphone signal and the second external microphone signal.

[0077] The first input signal, based on the internal microphone signal acquired through the internal microphone, has a relatively high SNR but may have weak high-band energy. Additionally, the second input signal, generated based on the external microphone signal, has a relatively low SNR but may have energy distributed across the entire band. Accordingly, when the first and second input signals are appropriately mixed according to external noise and wind intensity, a voice signal of better quality can be obtained than when using the individual signals alone.

[0078] According to one embodiment, the processor (310) can generate a first mixing parameter including weights to be applied to the first input signal and the second input signal by frequency band based on the noise intensity and wind intensity identified from the microphone signals.

[0079] According to one embodiment, the first mixing parameter may include a first weight applied to a first input signal and a second weight applied to a second input signal. The second weight may be a value obtained by subtracting the first weight from 1. The higher the intensity of noise and / or wind detected in the internal microphone signal and / or external microphone signal, the lower the value of the first weight and the higher the value of the second weight. That is, when generating the first mixing signal, the processor (310) may give a higher weight to the first input signal based on the internal microphone signal in a noisy environment and give a higher weight to the second input signal based on the external microphone signal in a low-noise environment.

[0080] According to one embodiment, the processor (310) can generate a noise-removing signal by removing noise from a first mixing signal output from a first mixing module. The electronic device can input the first mixing signal into a machine learning model and generate a noise-removing signal as the output value of the machine learning model. The machine learning model can be trained to extract a clear speech signal from an audio signal containing various types of noise. The machine learning model may be operated by the processor (310) or by a neural processing unit (NPU) independent of the processor (310). The machine learning model may be a deep neural network (DNN) model, but is not limited thereto, and may include other forms of machine learning models such as a convolutional neural network (CNN) model, a recurrent neural network (RNN) model, a graph neural network (GNN) model, a restricted Boltzmann machine (RBM) model, a deep belief network (DBN) model, or a bidirectional recurrent deep neural network (BRDNN) model, or may include a hybrid model in which two or more models are combined.

[0081] According to one embodiment, the processor (310) can generate a second mixing signal by mixing a noise removal signal and a first mixing signal based on a second mixing parameter. According to one embodiment, the second mixing parameter may be tuned according to noise and wind intensity and may have a value between 0 and 1. The second mixing parameter includes a first weight applied to the first mixing signal and a second weight applied to the noise removal signal, wherein the first weight has the same value as the second mixing parameter and the second weight has a value of 1 minus the first weight.

[0082] According to one embodiment, the processor (310) may output the second mixing signal as an audio signal generated according to the real-time translation call solution. For example, the processor (310) may transmit the audio signal corresponding to the second mixing signal to the other party of a voice call via a network using a communication circuit (340), and / or extract text information through a speech recognizer.

[0083] According to one embodiment, at least one of the first mixing parameter, the second mixing parameter, or the third mixing parameter may be determined through a parameter tuning model. To determine at least one of the first mixing parameter, the second mixing parameter, or the third mixing parameter, the parameter tuning model may measure speech recognition rate and speech quality by applying a plurality of parameters to a test vector having a predetermined noise intensity. Based on the measured speech recognition rate and speech quality, the parameter tuning model may select any one of the plurality of parameters. For example, the parameter tuning model may weight the speech recognition rate and speech quality based on weights determined according to the user's selection, and select the parameter applied when the weighted average value among the plurality of parameters is calculated to be the highest.

[0084] According to one embodiment, the parameter tuning model may be provided by a mobile device (300) (e.g., the mobile device (300) of FIG. 3), an audio device (e.g., the audio device (400) of FIG. 4), or by a server device on a network.

[0085] According to one embodiment, the processor (310) can select a parameter set corresponding to the noise intensity and / or wind intensity of an audio signal input through microphones from among the parameter sets output through a parameter tuning model and apply it to a call solution.

[0086] The method by which the parameter tuning model determines the mixing parameters and model parameters used in the real-time translation call solution will be explained in more detail through Fig. 8.

[0087] According to one embodiment, the processor (310) may perform an operation to generate a first mixing signal and / or a second mixing signal when the real-time translation function is running during a voice call. When the real-time translation function is not running, the processor (310) may use a general call solution to perform operations such as preprocessing, beamforming, and noise reduction of microphone signals without mixing operations.

[0088] Table 1 compares the speech recognition WER (word error rate) when a mobile device (300) applies a general call solution and a real-time translation call solution to the same audio signal.

[0089] Existing Mode Recognition Dedicated Mode Speaker 1 Clean 4.05% 4.58% Speaker 1 Cafe Noise 13.40% 7.36% Speaker 1 Subway 9.19% 7.24% Speaker 2 Clean 6.04% 3.55% Speaker 2 Cafe 21.85% 16.16% Speaker 2 Pub 25.81% 16.13% Speaker 2 Subway 15.81% 8.70% Speaker 3 Clean 8.61% 7.88% Speaker 3 Cafe 9.46% 6.43% Speaker 3 Subway 12.20% 6.74% Average 12.64% 8.48%

[0090] Referring to Table 1 above, it can be seen that when a real-time translation call solution (or voice recognition mode) is applied, the voice recognition rate is improved in most situations compared to when a general call solution is applied.

[0091] FIG. 4 is a block diagram of an audio device according to one embodiment.

[0092] Referring to FIG. 4, the audio device (400) may include a plurality of microphones, a communication circuit (440), an audio output unit (460), a processor (410), and a memory (420). Various embodiments of the present document may be implemented even if at least some of the configurations shown in FIG. 4 are omitted or replaced with other configurations.

[0093] Hereinafter, an embodiment in which a real-time translation call solution is performed by an audio device (400) will be described through FIG. 4. Technical features identical to those described above may be omitted from the description below.

[0094] According to one embodiment, the first microphone (432) may be an inner microphone (e.g., the inner microphone (480) of FIG. 2) placed near the user's ear when worn by the user, and the second microphone (434) may be an outer microphone (e.g., the outer microphone (492) of FIG. 2) placed on the opposite side of the user's ear. According to one embodiment, the audio device (400) may include two or more outer microphones. According to one embodiment, the audio device (400) may acquire a first microphone signal (or internal microphone signal) using a first microphone (432) (or internal microphone), acquire a second microphone signal (or first external microphone signal) using a second microphone (434) (or first external microphone), and / or acquire a third microphone signal (or second external microphone signal) using a third microphone (or second external microphone).

[0095] According to one embodiment, the communication circuit (440) may support near-field wireless communication with a mobile device (e.g., mobile device (300) of FIG. 2, mobile device (400) of FIG. 3). For example, near-field wireless communication between the mobile device and the audio device (400) may be exemplified by Bluetooth, WFD (Wi-Fi Direct), and NFC (near field communication), but is not limited thereto.

[0096] According to one embodiment, the audio output unit (460) can output audio data buffered in memory (420) under the control of the processor (410). For example, during a voice call, the other party's voice may be transmitted to the audio device (400) via a mobile device, processed by the processor (410), and output through the audio output unit (460). The audio output unit (460) may include a speaker and an ear tip.

[0097] According to one embodiment, the memory (420) may include volatile memory and non-volatile memory, and may store various data temporarily or permanently. The memory (420) may include at least some of the configuration and / or functions of the memory (130) of FIG. 1 and may store the program (140) of FIG. 1. The memory (420) may store various instructions that can be executed by the processor (410). Such instructions may include control commands such as arithmetic and logical operations, data movement, and input / output that can be recognized by the processor (410).

[0098] According to one embodiment, the processor (410) may be configured to perform operations or data processing regarding the control and / or communication of each component of the audio device (400), and may be composed of one or more processors. Although there are no limitations on the operations and data processing functions that the processor (410) can implement on the audio device (400), this document describes various embodiments for improving the voice recognition rate and the quality of the voice signal using a real-time translation call solution when executing a real-time translation function during a voice call. The operations of the processor (410) described below may be performed by loading instructions stored in the memory (420).

[0099] In this document, the description that the processor (410) can perform a certain operation (or function, task, or operation) may be interpreted substantially as meaning that an instruction (or command, computer program) causing the audio device (400) (or processor (410)) to perform said operation is stored in memory (420) (e.g., non-volatile memory, storage). Additionally, the description that the processor (410) can perform a certain operation may be interpreted substantially as meaning that at least one processor, without a fixed number, can perform said operation.

[0100] According to one embodiment, when a real-time translation function is activated while performing a voice call with a counterparty via a mobile device, the information can be received by the processor (410) through the communication circuit (440). When the real-time translation function is activated, the processor (410) can process microphone signals according to the real-time translation call solution described below.

[0101] According to one embodiment, the processor (410) may acquire an internal microphone signal from the first microphone (432) and acquire a first external microphone signal from the second microphone (434). If the audio device (400) further includes an external microphone, the processor (410) may further acquire a second external microphone signal.

[0102] According to one embodiment, the processor (410) can generate a signal in the short-time Fourier transform region by applying a short-time Fourier transform (STFT) to each of an internal microphone signal and at least one external microphone signal (e.g., a first external microphone signal, a second external microphone signal).

[0103] According to one embodiment, the processor (410) may perform pre-processing on each of the short-time Fourier transformed microphone signals. For example, the pre-processing process may include, but is not limited to, equalization, echo removal, gain adjustment, and / or high-pass filtering for each microphone signal.

[0104] According to one embodiment, the processor (410) can generate a beamforming signal by applying beamforming to a first external microphone signal and a second external microphone signal. The processor (410) can input the preprocessed external microphone signals into a beamforming module to perform beamforming and generate a beamforming signal.

[0105] According to one embodiment, the processor (410) can generate a third mixing signal by mixing the beamforming signal with the signal prior to beamforming. For example, the processor (410) can mix the beamforming signal generated by applying beamforming to the first external microphone signal and the second external microphone signal with the signal having a higher SNR among the first external microphone signal and the second external microphone signal.

[0106] According to one embodiment, the processor (410) can mix the beamforming signal and the first external microphone signal (or the second external microphone signal) based on a third mixing parameter. Here, the third mixing parameter includes a weight applied to the beamforming signal and the first external microphone signal during mixing and may have a value of 0 to 1. The third mixing parameter may be determined by considering the SNR and voice distortion of the microphone signals.

[0107] According to one embodiment, the processor (410) can generate a first mixing signal by mixing a first input signal generated based on an internal microphone signal and a second input signal generated based on a first external microphone signal based on a first mixing parameter. Here, the first input signal is a signal generated by performing a predetermined preprocessing process on a signal obtained by performing a short-time Fourier transform (STFT) on the internal microphone signal, and may be a signal obtained by mixing the signals before and after beamforming for the preprocessed first external microphone signal and the second external microphone signal.

[0108] According to one embodiment, the processor (410) can generate a first mixing parameter including weights to be applied to the first input signal and the second input signal by frequency band based on the noise intensity and wind intensity identified from the microphone signals.

[0109] According to one embodiment, the first mixing parameter may include a first weight applied to a first input signal and a second weight applied to a second input signal. The second weight may be a value obtained by subtracting the first weight from 1. The higher the intensity of noise and / or wind detected in the internal microphone signal and / or external microphone signal, the lower the value of the first weight and the higher the value of the second weight.

[0110] According to one embodiment, the processor (410) can generate a noise-removing signal by removing noise from a first mixing signal output from a first mixing module. The electronic device can input the first mixing signal into a machine learning model (e.g., a deep neural network (DNN) model) and generate a noise-removing signal as the output value of the machine learning model.

[0111] According to one embodiment, the processor (410) can generate a second mixing signal by mixing a noise removal signal and a first mixing signal based on a second mixing parameter. According to one embodiment, the second mixing parameter may be tuned according to noise and wind intensity and may have a value between 0 and 1. The second mixing parameter includes a first weight applied to the first mixing signal and a second weight applied to the noise removal signal, wherein the first weight has the same value as the second mixing parameter and the second weight has a value of 1 minus the first weight.

[0112] According to one embodiment, the processor (410) can transmit an audio signal corresponding to a second mixing signal output from a second mixing module to a mobile device via short-range wireless communication using a communication circuit (440). The mobile device can transmit the received audio signal to the other party of a voice call.

[0113] According to one embodiment, at least one of the first mixing parameter, the second mixing parameter, or the third mixing parameter may be determined through a parameter tuning model. To determine at least one of the first mixing parameter, the second mixing parameter, or the third mixing parameter, the parameter tuning model may measure speech recognition rate and speech quality by applying a plurality of parameters to a test vector having a predetermined noise intensity. Based on the measured speech recognition rate and speech quality, the parameter tuning model may select any one of the plurality of parameters. For example, the parameter tuning model may weight the speech recognition rate and speech quality based on weights determined according to the user's selection, and select the parameter applied when the weighted average value among the plurality of parameters is calculated to be the highest.

[0114] According to one embodiment, the processor (410) can select a parameter set corresponding to the noise intensity and / or wind intensity of an audio signal input through microphones from among the parameter sets output through a parameter tuning model and apply it to a call solution.

[0115] FIG. 5 is a block diagram illustrating the process of an electronic device according to one embodiment processing an audio signal according to a general call solution.

[0116] According to one embodiment, an electronic device (e.g., the mobile device (300) of FIG. 3 or the audio device (400) of FIG. 4) can perform a call solution including a voice enhancement operation for an audio signal (or microphone signal) input through at least one microphone of the audio device during a voice call.

[0117] Referring to FIG. 5, the microphone signal y(n) (510) input through the microphone may be mixed with the user's voice signal s(n) (512) and noise signal d(n) (514). According to one embodiment, the electronic device may acquire a plurality of microphone signals obtained through a plurality of microphones of the audio device (e.g., the internal microphone (480) and external microphones (492, 494) of FIG. 2).

[0118] According to one embodiment, the electronic device may apply a short-time Fourier transform (STFT) (560) to each microphone signal y(n) (510) to generate a signal Y(l, k) (520) in the short-time Fourier transform domain. In the signal Y(l, k) (520) in the short-time Fourier transform domain, l may represent a frame index and k may represent a frequency bin index. The electronic device may input the signal Y(l, k) (520) in the short-time Fourier transform domain to a voice enhancement module (570).

[0119] According to one embodiment, a voice enhancement module (570) can generate a gain function G(l, k) by performing various operations to perform voice enhancement on a signal Y(l, k) in a signal short-time Fourier transform domain. For example, the voice enhancement module (570) can perform pre-processing such as equalization, echo removal, gain control, and / or high-pass filtering, beamforming, and / or noise removal processing using a machine learning model. The electronic device can calculate an enhanced signal S(l, k) (525) in a short-time Fourier transform domain by multiplying the gain function G(l, k) obtained by the voice enhancement module (570) through the above operations by the signal Y(l, k) (520) in a short-time Fourier transform domain.

[0120] According to one embodiment, the electronic device can obtain an enhanced voice signal s(n) (530) by performing an inverse short-time Fourier transform (iSTFT) (580) on the signal S(l, k) (525) of the enhanced short-time Fourier transform region.

[0121] In the embodiment of FIG. 5, the voice enhancement module (570) can perform various processing to minimize voice distortion and residual noise on the input signal Y(l, k) (520) in the short-time Fourier transform domain. In a real-time translation scenario during a voice call, if a conventional call solution such as FIG. 5 is used as a preprocessor for a voice recognizer, it may be difficult to guarantee a high recognition rate in a noisy environment because the voice recognition rate is not considered separately.

[0122] The embodiment of FIG. 5 may be a comparative example to the embodiment of FIG. 6 and may be referred to as a general currency solution below. However, the embodiment of FIG. 5 is not considered to be prior art.

[0123] FIG. 6 is a block diagram illustrating the process of an electronic device according to one embodiment processing an audio signal according to a real-time translation call solution.

[0124] According to one embodiment, an electronic device (e.g., the mobile device (300) of FIG. 3 or the audio device (400) of FIG. 4) can acquire an internal microphone signal (616) containing the user's voice, a first external microphone signal (612), and a second external microphone signal (614). For example, if the electronic device is implemented as an audio device (400), the microphone signals can be acquired using each microphone of the audio device (400) (e.g., the microphones (480, 492, 494) of FIG. 2), and if the electronic device is implemented as a mobile device (300), the microphone signals acquired through each microphone of the audio device (400) can be received via short-range wireless communication.

[0125] According to one embodiment, the audio device may include an external microphone (or a second microphone) located opposite the user's ear when worn by the user, and an internal microphone (or a first microphone) that contacts the user's ear when worn by the user. For example, the audio device may include two external microphones and one internal microphone, but is not limited thereto. FIG. 6 illustrates that a first external microphone signal (612) and a second external microphone signal (614) obtained through the two external microphones of the audio device are input, but only one external microphone signal may be input, or three or more external microphone signals may be input.

[0126] According to one embodiment, the internal microphone signal (616), the first external microphone signal (612), and the second external microphone signal (612) may be digital signals.

[0127] According to one embodiment, the electronic device may apply a short-time Fourier transform (STFT) to an internal microphone signal (616) (or a first microphone signal) to generate a signal in the short-time Fourier transform region and then input it to a preprocessing module (660). The electronic device may apply a short-time Fourier transform to a first external microphone signal (612) (or a second microphone signal) and a second external microphone signal (614) and then input them to the preprocessing module (660).

[0128] The process of obtaining a signal in the time-frequency domain by applying a short-time Fourier transform to microphone signals can be represented as shown in the following mathematical equation 1.

[0129] [Mathematical Formula 1]

[0130] ,

[0131] In the above mathematical formula 1, i represents the microphone index, STFT(·) represents the short-time Fourier transform, l represents the frame index, and k represents the frequency bin index.

[0132] According to one embodiment, an electronic device (e.g., a mobile device or an audio device) may input microphone signals into a pre-processing module (660) to perform a predetermined pre-processing. Here, the pre-processing is an operation to increase the accuracy of signal recognition or analysis in a subsequent processing step, and may include operations to remove unnecessary noise or / or emphasize important features in the signal. For example, the pre-processing process may include, but is not limited to, equalization, echo removal, gain adjustment, and / or high-pass filtering for each microphone signal.

[0133] According to one embodiment, the electronic device may input a pre-processed first external microphone signal and a pre-processed second external microphone signal to a beamforming module (670). Beamforming of the audio signal is a multi-channel filtering technique that utilizes signals from multiple microphones to strengthen signals coming from a specific direction and suppress signals coming from other directions, thereby reducing background noise and processing audio at a specific location clearly. When there are two or more external microphones that acquire audio signals, such as the first external microphone signal (612) and the second external microphone signal (614), the electronic device may apply beamforming to improve the signal-to-noise ratio (SNR) of the voice signal. The beamforming module (670) may be designed using filtering algorithms such as GSC (generalized sidelobe canceller), MVDR (minimum variance distortion response), and GEV (generalized eigenvalue), but is not limited thereto.

[0134] According to one embodiment, the beamforming module (670) can generate and output a beamforming signal (620) by applying beamforming to a preprocessed first external microphone signal and a preprocessed second external microphone signal.

[0135] The beamforming signal (620) can be represented as in Equation 2.

[0136] [Mathematical Formula 2]

[0137]

[0138] Y in the above mathematical formula 2 i (l, k) is the i-th input signal, W i (l, k) is the beamforming filter applied to the i-th input signal, M is the number of external microphones, Y bf (l, k) can be a beamforming signal with beamforming applied.

[0139] According to one embodiment, the third mixing module (692) can mix the signal prior to beamforming with the beamforming signal. The electronic device can generate the third mixing signal (625) by mixing the beamforming signal (620) to which beamforming has been applied and the first external microphone signal (or second external microphone signal) based on the third mixing parameters. The electronic device can input the signal with a higher SNR among the first external microphone signal and the second external microphone signal to the third mixing module (692). For example, the electronic device can input the beamforming signal (620) output from the beamforming module (670) and the preprocessed first external microphone signal (or preprocessed second external microphone signal) output from the preprocessing module (660) to the third mixing module (692). The third mixing module (692) can weight-combine the input signals using a third mixing parameter that includes a weight to be applied to each input signal.

[0140] When performing beamforming in the beamforming module (670), data such as the steering vector, relative transfer function (RTF), noise covariance matrix, speech covariance matrix, or speech power spectral density (PSD) may be incorrectly estimated. In this case, distortion may occur in the beamforming signal, and distortion of the beamforming signal may lead to a decrease in speech recognition rate. The signal prior to beamforming (e.g., the first external microphone signal) has a relatively high SNR, but speech distortion may occur.

[0141] According to one embodiment, the electronic device may determine a third mixing parameter including a weight to be applied to a first external microphone signal and a beamforming signal, respectively, taking into account SNR and voice distortion. For example, the third mixing parameter may have a value between 0 and 1, the weight applied to the beamforming signal may have the same value as the third mixing parameter, and the weight applied to the first external microphone signal may have a value of 1 minus the third mixing parameter.

[0142] According to one embodiment, the electronic device can determine a third mixing parameter corresponding to noise and wind intensity using a parameter tuning module. The operation of the parameter tuning module will be explained in more detail with reference to FIG. 8.

[0143] The process of the third mixing module (692) mixing signals before and after beamforming can be represented as Equation 3.

[0144] [Mathematical Formula 3]

[0145]

[0146] In the above mathematical equation 3, α1(k) is the third mixing parameter, Y mic It may be the signal with the higher SNR between the first external microphone signal and the second external microphone signal.

[0147] According to one embodiment, the electronic device may generate a first mixing signal (630) by mixing a first input signal generated based on an internal microphone signal and a second input signal generated based on a second microphone signal based on a first mixing parameter. Here, the first input signal is a signal (618) generated by performing a predetermined preprocessing process on a signal obtained by performing a short-time Fourier transform (STFT) on the internal microphone signal (or the first microphone signal), and may be output from a preprocessing module (660) and input to a second mixing module (696). The second input signal may be a signal (625) to which beamforming is applied after performing a predetermined preprocessing process on a signal obtained by performing a short-time Fourier transform on a first external microphone signal and a second external microphone signal. The second input signal may be a signal obtained by mixing the signals before and after beamforming output from a third mixing module (692).

[0148] The first input signal, based on the internal microphone signal acquired through the internal microphone, has a relatively high SNR but may have weak high-band energy. Additionally, the second input signal, generated based on the external microphone signal, has a relatively low SNR but may have energy distributed across the entire band. Accordingly, when the first and second input signals are appropriately mixed according to external noise and wind intensity, a voice signal of better quality can be obtained than when using the individual signals alone.

[0149] According to one embodiment, the electronic device can generate a first mixing parameter including weights to be applied to a first input signal and a second input signal by frequency band, based on noise intensity and wind intensity identified from microphone signals.

[0150] According to one embodiment, the first mixing parameter may include a first weight applied to the first input signal and a second weight applied to the second input signal. The second weight may be a value obtained by subtracting the first weight from 1. The higher the intensity of noise and / or wind detected in the internal microphone signal and / or external microphone signal, the lower the value of the first weight and the higher the value of the second weight. That is, when generating the first mixing signal through the first mixing module (694), a higher weight may be given to the first input signal based on the internal microphone signal in a noisy environment, and a higher weight may be given to the second input signal based on the external microphone signal in a low-noise environment.

[0151] The process of the second mixing module (696) mixing the first input signal and the second input signal to generate input values ​​for a machine learning model (680) (e.g., a DNN model) can be represented as Equation 4.

[0152] [Mathematical Formula 4]

[0153]

[0154] In the above mathematical equation 4, α2(k) is the first mixing parameter, Y inner It could be an internal microphone signal.

[0155] According to one embodiment, the electronic device can determine a first mixing parameter corresponding to noise and wind intensity using a parameter tuning module. The operation of the parameter tuning module will be explained in more detail with reference to FIG. 8.

[0156] According to one embodiment, the electronic device can generate a noise removal signal (635) by removing noise from a first mixing signal (630) output from a first mixing module (694). The electronic device can input the first mixing signal (630) into a machine learning model (680) and generate the noise removal signal (635) as the output value of the machine learning model (680).

[0157] According to one embodiment, a machine learning model (680) receives a short-time Fourier transformed audio signal as input, distinguishes the spectrum of the voice signal and the noise signal in the audio signal to block the noise signal, and outputs a restored signal in a form with the noise removed. The machine learning model (680) can be trained to extract a clear voice signal from an audio signal containing various types of noise.

[0158] The process of the machine learning model (680) generating a noise removal signal can be represented as Equation 5.

[0159] [Mathematical Formula 5]

[0160]

[0161] In the above mathematical formula 5, DNN[·] can be any noise removal DNN model.

[0162] According to one embodiment, the electronic device can generate a second mixing signal (640) by mixing a noise removal signal (635) and a first mixing signal (630) based on a second mixing parameter. According to one embodiment, the noise removal signal (635) output from a machine learning model (680) and the first mixing signal (630) output from a first mixing module (694) can be input to a second mixing module (696).

[0163] According to one embodiment, the second mixing parameter may be tuned according to noise and wind intensity and may have a value between 0 and 1. The second mixing parameter includes a first weight applied to the first mixing signal and a second weight applied to the noise removal signal, wherein the first weight has the same value as the second mixing parameter and the second weight may have a value obtained by subtracting the first weight from 1.

[0164] The process of the second mixing module (696) generating the second mixing signal (640) can be represented as Equation 6.

[0165] [Mathematical Formula 6]

[0166]

[0167] In the above mathematical formula 6, α3(k) may be the second mixing parameter.

[0168] According to one embodiment, the electronic device can determine a second mixing parameter corresponding to noise and wind intensity using a parameter tuning module.

[0169] According to one embodiment, the electronic device may output a second mixing signal (640) output from a second mixing module (696) as an enhanced audio signal. For example, the electronic device may input the output audio signal into a speech recognition device to obtain text information, and / or transmit it to the counterparty device of a voice call. According to one embodiment, the speech recognition device may include various algorithms for receiving the audio signal and converting it into text information.

[0170] According to one embodiment, when a real-time translation function is running during a voice call, the electronic device may run a real-time translation call solution including the embodiment of FIG. 6. For example, when the real-time translation function is running, the electronic device may generate a first mixing signal using a first mixing module (694), generate a second mixing signal through a second mixing module (696), and / or generate a third mixing signal through a third mixing module (692). Alternatively, when the real-time translation function is not running, the electronic device may process microphone signals as in the general call solution of FIG. 5 without performing the operations of generating a first mixing signal using the first mixing module (694), generating a second mixing signal through the second mixing module (696), and / or generating a third mixing signal through the third mixing module (692).

[0171] FIG. 7 is a flowchart of a method for enhancing a voice signal of an electronic device according to one embodiment.

[0172] The method (700) illustrated in FIG. 7 can be performed by an electronic device (e.g., the mobile device (300) of FIG. 3, the audio device (400) of FIG. 4), and the technical features described above may be omitted from the description below.

[0173] According to one embodiment, in operation 710, an electronic device (e.g., the mobile device (300) of FIG. 3, the audio device (400) of FIG. 4) can check whether a real-time translation mode is running during a voice call. The real-time translation function may include recognizing the voice of the user of the mobile device and the call partner, converting it into text, and converting the converted text into another language. The mobile device and the audio device are connected via near-field wireless communication (e.g., Bluetooth, Wi-Fi Direct (WFD), near field communication (NFC)), so that the mobile device transmits the voice of the other party received from the other party device via a network to the audio device and outputs it through the audio device, and transmits an audio signal including the user's voice obtained from the microphone of the audio device (e.g., an internal microphone, at least one external microphone) to the mobile device.

[0174] According to one embodiment, when the real-time translation mode is not running, in operation 750, the electronic device may process an audio signal including a user's voice obtained through a microphone using a general call solution (e.g., FIG. 5). When the general call solution (or general mode) is running, the mixing operations of the signal performed in the real-time translation call solution (or voice recognition mode) (e.g., the mixing operations of the third mixing module (692), the first mixing module (694), and the second mixing module (696) of FIG. 6) may not be performed.

[0175] According to one embodiment, when a real-time translation mode is running, in operation 720, the electronic device can estimate the noise and wind intensity of the audio signal. For example, the noise intensity can be divided into multiple stages, such as high noise, medium noise, and low noise, and the wind intensity can be divided into multiple stages, such as wind presence and wind absence.

[0176] According to one embodiment, in operation 730, the electronic device can determine optimal mixing parameters and model parameters based on the noise and wind intensity of the estimated audio signal. Here, the parameter tuning model can determine first mixing parameters, second mixing parameters, and third mixing parameters, including weights applied to input signals in the first mixing module (694), second mixing module (696), and third mixing module (692) of FIG. 6. The model parameters may include various parameters defined in various modules used in the voice call solution, such as preprocessing, beamforming, noise estimation, speech detection, and pre- and post-processing of machine learning models.

[0177] According to one embodiment, a parameter tuning model may measure speech recognition rate and speech quality by applying a plurality of parameters to a test vector having a predetermined noise intensity to determine at least one of a first mixing parameter, a second mixing parameter, or a third mixing parameter. Based on the measured speech recognition rate and speech quality, the parameter tuning model may select any one of the plurality of parameters. For example, the parameter tuning model may weight the speech recognition rate and speech quality based on weights determined by user selection, and select the parameter applied when the weighted average value among the plurality of parameters is calculated to be the highest.

[0178] According to one embodiment, the parameter tuning model may be provided by a mobile device (e.g., the mobile device (300) of FIG. 3), an audio device (e.g., the audio device (400) of FIG. 4), or by a server device on a network.

[0179] According to one embodiment, the electronic device can determine mixing parameters and model parameters corresponding to the noise and wind intensity of the estimated audio signal.

[0180] According to one embodiment, in operation 740, the electronic device can perform a real-time translation mode call solution. The electronic device can perform a processing process of an audio signal according to the real-time translation mode call solution described through FIG. 6 using mixing parameters (e.g., first mixing parameter, second mixing parameter, third mixing parameter) and model parameters (e.g., preprocessing parameter, beamforming parameter, machine learning model parameter) determined in operation 730.

[0181] According to one embodiment, an electronic device can transmit an audio signal processed according to a real-time translation mode call solution to a counterparty device via a network. The counterparty device can input the audio signal received from the electronic device into a speech recognizer for real-time translation to perform a speech recognition operation.

[0182] Instructions for performing each operation constituting the above method can be stored on a tangible and non-transitory computer-readable recording medium.

[0183] FIG. 8 illustrates models for determining mixing parameters according to one embodiment.

[0184] According to one embodiment, at least one of the first mixing parameter used in the first mixing module (e.g., the first mixing module (694) of FIG. 6), the second mixing parameter used in the second mixing module (e.g., the second mixing module (696) of FIG. 6), or the third mixing parameter used in the third mixing module (e.g., the third mixing module (692) of FIG. 6) may be determined through a parameter tuning model (810). The parameter tuning model (810) may be provided by a mobile device (e.g., the mobile device (300) of FIG. 3), an audio device (e.g., the audio device (400) of FIG. 4), or a server device on a network.

[0185] According to one embodiment, the parameter tuning model (810) can determine mixing parameters and model parameters used in various modules used in voice enhancement operations. Here, the parameter tuning model (810) can determine a first mixing parameter, a second mixing parameter, and a third mixing parameter, including weights applied to input signals in the first mixing module (694), the second mixing module (696), and the third mixing module (692) of FIG. 6. The model parameters may include various parameters defined in various modules used in a voice call solution, such as preprocessing, beamforming, noise estimation, voice detection, and pre- and post-processing of machine learning models. For example, the model parameters may include parameters used in the preprocessing module (660), beamforming module (670), and / or machine learning model (680) of FIG. 6.

[0186] According to one embodiment, a parameter tuning model (810) may measure speech recognition rate and speech quality by applying a plurality of parameters to test vectors having a predetermined noise intensity and / or wind intensity to determine at least one of a first mixing parameter, a second mixing parameter, or a third mixing parameter. Based on the measured speech recognition rate and speech quality, the parameter tuning model (810) may select any one of the plurality of parameters. For example, the parameter tuning model (810) may weight the speech recognition rate and speech quality based on weights determined by user selection, and select the parameter applied when the weighted average value among the plurality of parameters is calculated to be the highest.

[0187] According to one embodiment, the parameter tuning model (810) can explore mixing parameters and model parameters in a direction that maximizes the voice recognition rate while providing stable call quality. The parameter tuning model (810) can calculate a result value using a plurality of parameter sets that combine several mixing parameters and model parameters, and can determine one of the plurality of parameter sets based on the calculated result value. The parameter tuning model (810) can explore an optimized parameter set for each situation using test vectors corresponding to noise intensity and / or wind speed.

[0188] The parameter tuning model (810) can search for the optimal set of parameters among various parameter candidate groups according to an objective function such as Equation 7.

[0189] [Mathematical Formula 7]

[0190]

[0191] In the above mathematical formula 7, F(·) may be a predefined speech recognizer. is a parameter set It could be a voice call solution tuned for T j is the j-th test vector, and N may be the total number of test vectors. Score(·) is the speech recognition rate calculated by the speech recognition model (820), and Quality(·) may be the speech quality evaluation result calculated by the speech quality evaluation model (830). δ is a weight applied to the importance of speech quality relative to the speech recognition rate and may be determined according to the user's choice.

[0192] According to one embodiment, the speech recognition model (820) includes a pre-recognized speech recognizer and can perform a speech recognition operation that converts an audio signal received from the parameter tuning model (810) into text information. As a result of the speech recognition operation, the speech recognition model (820) can convert the speech recognition rate into a score and transmit it to the parameter tuning model (810).

[0193] According to one embodiment, a voice quality evaluation model (830) can evaluate the quality of an audio signal corrected according to a call solution (or voice enhancement operation) using indicators such as DNSMOS (deep noise suppression mean opinion score), PESQ (perceptual evaluation of speech quality), and PQLQA (perceptual objective listening quality analysis). The indicators used for voice quality evaluation are not limited thereto.

[0194] According to one embodiment, the parameter tuning model (810) can obtain at least one test vector from input audio signals (e.g., internal microphone signal, at least one external microphone signal). Since the at least one test vector obtained corresponds to an audio signal measured in a short time in the same environment, the noise intensity and / or wind intensity may have substantially the same value.

[0195] According to one embodiment, the parameter tuning model (810) may obtain test vectors corresponding to each noise intensity and / or wind intensity in advance, and use each test vector to obtain an optimal parameter set for each noise intensity and / or wind intensity environment. For example, noise intensity may be divided into multiple stages such as high noise, medium noise, and low noise, and wind intensity may be divided into multiple stages such as wind presence and wind absence. The parameter tuning model (810) may input at least one N test vectors having a specific stage of noise intensity and wind intensity into an objective function such as Equation 6 and determine an optimal parameter set.

[0196] According to one embodiment, the parameter tuning model (810) selects N first test vectors and includes mixing parameters and model parameters to be applied to the first test vectors. Parameter set 1 can be selected from among the parameter sets. The parameter tuning model (810) can transmit the output signal to the speech recognition model (820) and the speech quality evaluation model (830) after performing a call solution (or speech enhancement operation) using parameter set 1 for test vector 1.

[0197] According to one embodiment, the speech recognition model (820) can perform a speech recognition operation on an audio signal processed using the first parameter set of the received first test vector, calculate the speech recognition rate as a score, and transmit it to the parameter tuning model (810). Additionally, the speech quality evaluation model (830) can evaluate the speech quality of the corresponding audio signal using indicators such as DNSMOS, PESQ, and PQLQA and transmit it to the parameter tuning model (810).

[0198] Next, the parameter tuning model (810) performs a call solution (or voice enhancement operation) using the first parameter set for the second test vector and transmits the output signal to the voice recognition model (820) and the voice quality evaluation model (830), and can receive the voice recognition rate score and the voice quality evaluation result. The parameter tuning model (810) can repeat this operation up to the Nth test vector and finally obtain the final result value when using the first parameter set.

[0199] According to one embodiment, the parameter tuning model (810) may apply the above process to several parameter sets and then select the parameter set with the highest final result value. According to one embodiment, the parameter tuning model (810) may utilize a greedy algorithm or a genetic algorithm to find the optimal parameter set among the candidate sets of parameters.

[0200] According to one embodiment, the electronic device can apply mixing parameters and model parameters included in the optimal parameter set obtained by the parameter tuning model (810) to apply a call solution for an input audio signal.

[0201] FIGS. 9A, FIGS. 9B, and FIGS. 9C illustrate a foldable device according to one embodiment.

[0202] According to one embodiment, the electronic device (900) may be configured as a foldable device that includes a flexible display (930) and can be folded along at least one folding axis. For example, the electronic device (900) may include a first housing (910) and a second housing (920) so as to be rotatable relative to each other through a hinge structure, and the display (930) may be positioned across the first housing (910) and the second housing (920). In the unfolded state, the display (930) is exposed to the outside, and when folded, a first area positioned in the first housing (910) and a second area positioned in the second housing (920) may be folded in a direction facing each other.

[0203] According to one embodiment, the electronic device (900) may include at least one microphone (961, 962, 963). Referring to FIG. 9a, the electronic device (900) may include a microphone (961) at the top of the housing, a microphone (962) at the bottom, and a microphone (963) at the back. The number and placement locations of the microphones are not limited thereto.

[0204] According to one embodiment, the electronic device (900) can determine a call solution to be used for processing an audio signal based on a folding state and / or a call mode during a voice call.

[0205] According to one embodiment, when an electronic device (900) receives a user voice using a microphone of an external audio device (e.g., the audio device (400) of FIG. 2, the audio device (400) of FIG. 4) during a voice call, it can process the audio signal obtained from the audio device using a corresponding call solution. For example, the electronic device (900) can process the audio signal using a general call solution described through FIG. 5 or a real-time translation call solution described through FIG. 6.

[0206] According to one embodiment, when the electronic device (900) receives the user's voice using the microphone of the electronic device (900) during a voice call, it can determine a call solution to process the audio signal based on the current folding state. For example, as shown in FIG. 9a, the electronic device (900) may be in a fully folded state where the first housing (910) and the second housing (920) are completely folded, and in this case, the electronic device (900) can process the audio signal obtained through the microphones using a call solution corresponding to the fully folded state.

[0207] According to one embodiment, as shown in FIG. 9b, when the first housing (910) and the second housing (920) are in a fully unfolded state, the electronic device (900) can process audio signals obtained through microphones using a call solution corresponding to the unfolded state.

[0208] According to one embodiment, as shown in FIG. 9c, when the first housing (910) and the second housing (920) are in an intermediate state forming a predetermined angle between 0 and 180 degrees, the electronic device (900) can process an audio signal obtained through microphones using a call solution corresponding to the angle between the first housing (910) and the second housing (920).

[0209] According to one embodiment, the electronic device (900) can estimate the noise and wind intensity of an audio signal obtained through microphones and determine optimal model parameters corresponding to the noise and wind intensity in each folding state. The model parameters may include various parameters defined in various modules used in a voice call solution, such as preprocessing, beamforming, noise estimation, voice detection, and pre- and post-processing of a machine learning model.

[0210] FIGS. 10a, FIGS. 10b, FIGS. 10c and FIGS. 10d illustrate a multi-foldable device according to one embodiment.

[0211] According to one embodiment, the electronic device (1000) may be configured as a multi-foldable device that includes a flexible display (1050) and can be folded along two or more folding axes. For example, the electronic device (1000) may include a first housing (1010), a second housing (1020), and a third housing (1030) so as to be rotatable relative to each other through a hinge structure, and the display (1050) may be positioned across the first housing (1010), the second housing (1020), and the third housing (1030).

[0212] According to one embodiment, when an electronic device (1000) receives a user's voice using the microphone of the electronic device (1000) during a voice call, it can determine a call solution to process the audio signal based on the current folding state.

[0213] According to one embodiment, when the first housing (1010), the second housing (1020), and the third housing (1030) are in a completely folded state as in FIG. 10a, the electronic device (1000) can process audio signals obtained through microphones using a call solution corresponding to the folded state.

[0214] According to one embodiment, when the first housing (1010), the second housing (1020), and the third housing (1030) are in a fully unfolded state as in FIG. 10b, the electronic device (1000) can process audio signals obtained through microphones using a call solution corresponding to the unfolded state.

[0215] According to one embodiment, as shown in FIG. 10c, when the first housing (1010) and the second housing (1020) are folded to form a predetermined angle and the second housing (1020) and the third housing (1030) are folded to form a predetermined angle, the electronic device (1000) can process an audio signal obtained through microphones using a call solution corresponding to the angle between the first housing (1010) and the second housing (1020) and the angle between the second housing (1020) and the third housing (1030).

[0216] According to one embodiment, when the first housing (1010)-second housing (1020) and the second housing (1020)-third housing (1030) are folded to form an angle smaller than that of FIG. 10b, the electronic device (1000) can process an audio signal obtained through microphones using a call solution corresponding to the angle between the first housing (1010) and the second housing (1020) and the angle between the second housing (1020) and the third housing (1030).

[0217] According to one embodiment, the electronic device (1000) can estimate the noise and wind intensity of an audio signal obtained through microphones and determine optimal model parameters corresponding to the noise and wind intensity in each folding state. The model parameters may include various parameters defined in various modules used in a voice call solution, such as preprocessing, beamforming, noise estimation, speech detection, and pre- and post-processing of a machine learning model.

[0218] An electronic device according to various embodiments of the present document may include at least one processor and memory.

[0219] According to one embodiment, the memory may be executed by at least one processor, and the electronic device may store instructions for acquiring a first microphone signal and a second microphone signal including a user's voice, mixing a first input signal generated based on the first microphone signal and a second input signal generated based on the second microphone signal based on a first mixing parameter to generate a first mixing signal, removing noise from the first mixing signal to generate a noise removal signal, mixing the noise removal signal and the first mixing signal based on a second mixing parameter to generate a second mixing signal, and outputting the second mixing signal.

[0220] According to one embodiment, the memory may store instructions that cause the electronic device to perform the operation of generating the first mixing signal and / or the operation of generating the second mixing signal when the real-time translation function is running during a voice call.

[0221] According to one embodiment, the memory may store instructions for the electronic device to further acquire a third microphone signal including a user's voice, to apply beamforming to the second microphone signal and the third microphone signal to generate a beamforming signal, to mix the beamforming signal and the second microphone signal based on a third mixing parameter to generate a third mixing signal, and to mix the first input signal and the third mixing signal based on the first mixing parameter to generate the first mixing signal.

[0222] According to one embodiment, the memory may store instructions for the electronic device to apply beamforming after performing a predetermined preprocessing process on a signal obtained by short-time Fourier transforming the second microphone signal and the third microphone signal.

[0223] According to one embodiment, the memory may store instructions that cause the electronic device to generate the first input signal by performing a predetermined preprocessing process on a signal obtained by performing a short-time Fourier transform on the first microphone signal.

[0224] According to one embodiment, the memory may store instructions that cause the electronic device to input the first mixing signal into a machine learning model and generate the noise removal signal as the output value of the machine learning model.

[0225] According to one embodiment, the first mixing parameter includes a first weight applied to the first input signal and a second weight applied to the second input signal, and the higher the intensity of the noise detected in the first microphone signal and / or the second microphone signal, the lower the value of the first weight and the higher the value of the second weight.

[0226] According to one embodiment, at least one of the first mixing parameter, the second mixing parameter, or the third mixing parameter can be determined through a parameter tuning model.

[0227] According to one embodiment, the parameter tuning model measures a voice recognition rate and voice quality by applying a plurality of parameters to a test vector having a predetermined noise intensity in order to determine at least one of the first mixing parameter, the second mixing parameter, or the third mixing parameter, and can select any one of the plurality of parameters based on the measured voice recognition rate and voice quality.

[0228] According to one embodiment, the parameter tuning model may weight the measured voice recognition rate and voice quality based on weights determined by user selection, and select the applied parameter when the weighted average value among the plurality of parameters is calculated to be the highest.

[0229] According to one embodiment, the first microphone signal may include a signal acquired by an internal microphone, and the second microphone signal may include a signal acquired by an external microphone.

[0230] According to one embodiment, the electronic device is an audio device wearable by a user and may include a first microphone that contacts the user's ear when worn by the user and acquires the first microphone signal, and a second microphone that is located opposite the user's ear when worn by the user and acquires the second microphone signal.

[0231] According to one embodiment, the electronic device is a mobile device that provides voice calls and may include a communication circuit for receiving the first microphone signal and the second microphone signal from an external audio device.

[0232] A method performed by an electronic device according to various embodiments of the present document may include: an operation of acquiring a first microphone signal and a second microphone signal including a user's voice; an operation of mixing a first input signal generated based on the first microphone signal and a second input signal generated based on the second microphone signal based on a first mixing parameter to generate a first mixing signal; an operation of removing noise from the first mixing signal to generate a noise removal signal; an operation of mixing the noise removal signal and the first mixing signal based on a second mixing parameter to generate a second mixing signal; and instructions for outputting the second mixing signal.

[0233] According to one embodiment, the method further includes an operation to check whether a real-time translation function is running during a voice call, and the operation to generate the first mixing signal and / or the operation to generate the second mixing signal may be performed when the real-time translation function is running.

[0234] According to one embodiment, the method comprises: acquiring a third microphone signal including a user's voice; applying beamforming to the second microphone signal and the third microphone signal to generate a beamforming signal; mixing the beamforming signal and the second microphone signal based on a third mixing parameter to generate a third mixing signal; and mixing the first input signal and the third mixing signal based on the first mixing parameter to generate the first mixing signal.

[0235] According to one embodiment, the operation of generating the noise removal signal may include inputting the first mixing signal into a machine learning model and generating the noise removal signal as the output value of the machine learning model.

[0236] According to one embodiment, the first mixing parameter includes a first weight applied to the first input signal and a second weight applied to the second input signal, and the higher the intensity of the noise detected in the first microphone signal and / or the second microphone signal, the lower the value of the first weight and the higher the value of the second weight.

[0237] According to one embodiment, the method may further include the operation of measuring a voice recognition rate and voice quality by applying a plurality of parameters to a test vector having a predetermined noise intensity to determine at least one of the first mixing parameter, the second mixing parameter, or the third mixing parameter using the parameter tuning model, and the operation of selecting any one of the plurality of parameters based on the measured voice recognition rate and voice quality.

[0238] A computer-readable non-transient recording medium according to various embodiments of the present document may store instructions for: acquiring a first microphone signal and a second microphone signal including a user's voice; mixing a first input signal generated based on the first microphone signal and a second input signal generated based on the second microphone signal based on a first mixing parameter to generate a first mixing signal; removing noise from the first mixing signal to generate a noise removal signal; mixing the noise removal signal and the first mixing signal based on a second mixing parameter to generate a second mixing signal; and outputting the second mixing signal.

[0239] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0240] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0241] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0242] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0243] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0244] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device, At least one processor; and It includes memory for storing multiple instructions, When the above plurality of instructions are executed individually or collectively by the at least one processor, the electronic device, Acquire a first microphone signal and a second microphone signal including the user's voice, A first input signal generated based on the first microphone signal and a second input signal generated based on the second microphone signal are mixed based on a first mixing parameter to generate a first mixing signal, and A noise removal signal is generated by removing noise from the first mixing signal above, and The noise removal signal and the first mixing signal are mixed based on the second mixing parameter to generate a second mixing signal, and An electronic device that outputs the above second mixing signal.

2. In Paragraph 1, When the above plurality of instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that performs the operation of generating the first mixing signal and / or the operation of generating the second mixing signal when a real-time translation function is running during a voice call.

3. In Paragraph 1, When the above plurality of instructions are executed individually or collectively by the at least one processor, the electronic device, Further acquire a third microphone signal including the user's voice, and A beamforming signal is generated by applying beamforming to the second microphone signal and the third microphone signal, and The beamforming signal and the second microphone signal are mixed based on the third mixing parameter to generate a third mixing signal, and An electronic device that generates the first mixing signal by mixing the first input signal and the third mixing signal based on the first mixing parameter.

4. In Paragraph 3, When the above plurality of instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that applies beamforming after performing a predetermined preprocessing process on a signal obtained by short-time Fourier transforming the second microphone signal and the third microphone signal.

5. In Paragraph 1, When the above plurality of instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that generates the first input signal by performing a predetermined preprocessing process on a signal obtained by performing a short-time Fourier transform on the first microphone signal.

6. In Paragraph 1, When the above plurality of instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that inputs the first mixing signal into a machine learning model and generates the noise removal signal as the output value of the machine learning model.

7. In Paragraph 1, The first mixing parameter above includes a first weight applied to the first input signal and a second weight applied to the second input signal, and An electronic device in which the higher the intensity of noise detected in the first microphone signal and / or the second microphone signal, the lower the first weight and the higher the second weight.

8. In Paragraph 1, An electronic device in which at least one of the first mixing parameter, the second mixing parameter, or the third mixing parameter is determined through a parameter tuning model.

9. In Paragraph 8, The above parameter tuning model is, To determine at least one of the first mixing parameter, the second mixing parameter, or the third mixing parameter, a plurality of parameters are applied to a test vector having a predetermined noise intensity to measure the speech recognition rate and speech quality, and An electronic device that selects one of the plurality of parameters based on the above-mentioned measured voice recognition rate and voice quality.

10. In Paragraph 9, The above parameter tuning model is, An electronic device that calculates a weighted average of the measured voice recognition rate and voice quality based on weights determined by user selection, and selects the applied parameter when the weighted average value among the plurality of parameters is calculated to be the highest.

11. In Paragraph 1, An electronic device in which the first microphone signal comprises a signal acquired by an internal microphone, and the second microphone signal comprises a signal acquired by an external microphone.

12. In Paragraph 1, The above electronic device is, It is an audio device that can be worn by the user, and An electronic device comprising a first microphone that contacts the user's ear when worn by the user and acquires the first microphone signal, and a second microphone that is positioned opposite the user's ear when worn by the user and acquires the second microphone signal.

13. In Paragraph 1, The above electronic device is, It is a mobile device that provides voice calls, and An electronic device comprising a communication circuit for receiving the first microphone signal and the second microphone signal from an external audio device.

14. In a method performed by an electronic device, An operation of acquiring a first microphone signal and a second microphone signal including the user's voice; An operation to generate a first mixing signal by mixing a first input signal generated based on the first microphone signal and a second input signal generated based on the second microphone signal based on a first mixing parameter; An operation to generate a noise-removed signal by removing noise from the first mixing signal above; An operation to generate a second mixing signal by mixing the noise removal signal and the first mixing signal based on a second mixing parameter; and A method including the operation of outputting the second mixing signal.

15. In a computer-readable non-transient recording medium, An operation of acquiring a first microphone signal and a second microphone signal including the user's voice; An operation to generate a first mixing signal by mixing a first input signal generated based on the first microphone signal and a second input signal generated based on the second microphone signal based on a first mixing parameter; An operation to generate a noise-removed signal by removing noise from the first mixing signal above; An operation to generate a second mixing signal by mixing the noise removal signal and the first mixing signal based on a second mixing parameter; and A recording medium storing instructions for outputting the second mixing signal.

Citation Information

Patent Citations

  • Speaker recognition based on an inside microphone of a headphone

    US10896682B1

  • Microphone mixing for wind noise reduction

    US20170251299A1

  • Two channel headset-based own voice enhancement

    US20180268798A1

  • Smartphone-based telephone translation system

    US20190347331A1

  • Speech Signal Processing Method and Apparatus

    US20230029267A1