Electronic device, operating method thereof, and storage medium
The electronic device uses an ASR module to convert voice input to a second language and adjusts speaker settings to reduce loopback sound, addressing usability issues and enhancing the recognition rate of translated sound.
Patent Information
- Application Number
- PCT/KR2024/096753
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-29
- Filing Date
- 2024-12-12
- Publication Date
- 2025-08-21
AI Technical Summary
Existing electronic devices face challenges in effectively translating voice input due to issues with loopback sound, which reduces the recognition rate of the translated sound, particularly when the translated sound is input back into the microphone, affecting user usability.
The electronic device employs an automatic speech recognition (ASR) module to convert voice input from a first language to a second language and outputs it through a plurality of speakers, with settings adjusted to minimize loopback sound interference by turning off or reducing the volume of speakers adjacent to the microphone.
This approach enhances the recognition rate of translated sound by minimizing loopback sound interference, thereby improving user usability and effectiveness of voice translation functions.
Smart Images

Figure KR2024096753_21082025_PF_FP_ABST
Abstract
Description
Electronic device and its operating method, and storage medium
[0001] The present disclosure relates to an electronic device for outputting sound according to one embodiment, a method of operating the same, and a storage medium.
[0002] Thanks to remarkable advancements in information and communication technology and semiconductor technology, the proliferation and use of various electronic devices is rapidly increasing. Electronic devices are being developed to enable users to carry and communicate with one another. An electronic device can refer to any device that performs a specific function based on its embedded software, such as a mobile communication terminal, tablet PC, audio / video device, desktop / laptop computer, or in-car navigation system.
[0003] Meanwhile, the variety of services and additional features offered through electronic devices, such as smartphones, is steadily increasing. To enhance the utility of these devices and satisfy the diverse needs of users, telecommunications service providers and electronic device manufacturers are competitively developing electronic devices that offer diverse features and differentiate themselves from competitors. Consequently, the various functions offered through electronic devices are also becoming increasingly sophisticated. For example, electronic devices now offer the ability to recognize voice input, convert the recognized voice into text-based data, or perform translations of the recognized voice.
[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0005] An electronic device according to one embodiment of the present disclosure may include a plurality of speakers, at least one microphone, a memory storing instructions, and at least one processor.
[0006] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to execute an application related to translation.
[0007] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine a setting for each of the plurality of speakers, the setting corresponding to the application.
[0008] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to output, through at least one speaker of the plurality of speakers, a second sound corresponding to a second language, the second sound being converted through an automatic speech recognition (ASR) module, based on the identified settings, in response to a first sound corresponding to a first language being input through the at least one microphone.
[0009] A method of operating an electronic device according to one embodiment of the present disclosure may include an operation of executing an application related to translation.
[0010] According to one embodiment, the method of operation may include an operation of checking a setting for each of a plurality of speakers set in response to the application.
[0011] According to one embodiment, the operating method may include an operation of, when a first sound corresponding to a first language is input through at least one microphone, converting the input first sound through an automatic speech recognition (ASR) module and outputting a second sound corresponding to a second language through at least one speaker among the plurality of speakers based on the identified settings.
[0012] In a storage medium storing computer-readable instructions according to one embodiment of the present disclosure, the instructions, when individually or collectively executed by at least one processor of an electronic device, can cause the electronic device to execute an application related to translation.
[0013] In one embodiment, the instructions, when executed individually or collectively by at least one processor, may cause the electronic device to determine a setting for each of a plurality of speakers configured in response to the application.
[0014] In one embodiment, the instructions, when executed individually or collectively by at least one processor, may cause the electronic device to, when a first sound corresponding to a first language is input through at least one microphone, output a second sound corresponding to a second language, converted from the input first sound through an automatic speech recognition (ASR) module, through at least one speaker among the plurality of speakers based on the identified settings.
[0015] In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0016] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.
[0017] FIG. 2 is a block diagram of electronic device configurations according to one embodiment.
[0018] FIG. 3 is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0019] FIG. 4A is a drawing for explaining the settings for each of a plurality of speakers according to one embodiment.
[0020] FIG. 4b is a drawing for explaining a method of outputting sound according to one embodiment.
[0021] FIG. 5 is a flowchart illustrating a method for outputting a second sound according to one embodiment.
[0022] Figure 6 is a flowchart illustrating a method for outputting a second sound according to one embodiment.
[0023] FIG. 7 is a flowchart illustrating a method for verifying weights associated with a transformation according to one embodiment.
[0024] FIG. 8 is a flowchart illustrating a method for confirming a second sound according to one embodiment.
[0025] FIG. 9 is a flowchart illustrating a method for identifying a second sound based on noise according to one embodiment.
[0026] Fig. 10 is a flowchart for explaining an operating method of an electronic device according to one embodiment.
[0027] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.
[0028] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100), according to one embodiment.
[0029] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0030] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0031] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0032] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0033] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0034] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0035] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0036] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0037] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0038] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0039] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0040] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0041] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0042] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0043] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0044] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0045] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0046] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0047] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0048] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0049] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0050] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0051] In the detailed description below, reference numerals in the drawings may be used interchangeably or omitted for components that can be easily understood through the preceding embodiments, and their detailed descriptions may also be omitted. An electronic device according to an embodiment disclosed in this document may be implemented by selectively combining components of different embodiments, and components of one embodiment may be replaced by components of another embodiment. For example, it should be noted that the present disclosure is not limited to any specific drawing or embodiment.
[0052] FIG. 2 is a block diagram of electronic device configurations according to one embodiment.
[0053] According to FIG. 2, according to one embodiment, an electronic device (200, e.g., electronic device (101) of FIG. 1) may include a speaker (210), a microphone (220), a memory (230, e.g., memory (130) of FIG. 1) for storing instructions, a communication module (240, e.g., communication module (190) of FIG. 1)) and at least one processor (250).
[0054] In one embodiment, the speaker (210) may have at least a portion of the same or similar configuration as the audio output module (155) of FIG. 1. In one embodiment, there may be a plurality of speakers (210). In one embodiment, each of the speakers (210, or the plurality of speakers) may be positioned at different locations within the electronic device (200). In one embodiment, the plurality of speakers (210) may output audio signals (or sounds) to the outside of the electronic device (200).
[0055] In one embodiment, the microphone (220) may have at least a portion of the same or similar configuration as the input module (150) of FIG. 1. In one embodiment, the microphone (220) may receive an acoustic signal. In one embodiment, the electronic device (200) may include one or more microphones.
[0056] In one embodiment, the memory (230) may have at least a portion of the same or similar configuration as the memory (130) of FIG. 1. For example, the memory (230) may be configured to temporarily or permanently store digital data and may include at least a portion of the configuration and / or functions of the memory (130) of FIG. 1.
[0057] A memory (230) according to one embodiment can store various instructions that can be executed by at least one processor (250). In addition, the memory (230) can store at least a portion of the program (140) of FIG. 1. Such instructions can include control commands such as logical operations and data input / output that can be recognized and executed by the processor (250). There is no limitation on the type and / or amount of data that the memory (230) can store, but in this document, the configuration and function of the memory related to the method of outputting sound according to various embodiments and the operation of the processor (250) that performs the method will be described. The memory (230) can store various information, and the various information stored by the memory (230) will be described in detail below.
[0058] The communication module (240) according to one embodiment may have at least a portion identical or similar to the communication module (190) of FIG. 1. In one embodiment, the electronic device (200) may receive sound (or an electrical signal corresponding to sound) corresponding to the other party of a call through the communication module (240). In one embodiment, the electronic device (200) may transmit and receive information with an external device (e.g., the server (108) of FIG. 1) through the communication module.
[0059] In one embodiment, at least one processor (250, hereinafter, processor) may have at least a portion of the same or similar configuration as the processor (120) of FIG. 1. In one embodiment, the processor (250) may include one or more processors.
[0060] In one embodiment, the processor (250) may perform various operations by executing instructions stored in the memory (230).
[0061] In one embodiment, the processor (250) may execute an application related to translation. In one embodiment, the application related to translation may be an application for performing translation on sound input during a call. In one embodiment, the processor (250) may perform translation on sound (e.g., sound of a caller) input through a microphone (220, or at least one microphone) included in the electronic device (200) using the application related to translation. Alternatively, in one embodiment, the application related to translation may perform translation on sound (e.g., sound of a caller) received through a communication module (250, e.g., communication module (190) of FIG. 1). However, the present invention is not limited thereto, and the processor (250) may receive sound through a different type of input module (e.g., input module (150) of FIG. 1) other than the communication module (240) and perform translation on the received sound. In one embodiment, the processor (250) may translate sound input during a call into a specified language through a translation-related application and output the translated sound through at least one of a plurality of speakers (210).
[0062] In one embodiment, the processor (250) may check the settings for each of the plurality of speakers (210) when executing an application. In one embodiment, the settings for each of the plurality of speakers (210) may be settings for turning on / off at least one of the plurality of speakers (210). Alternatively, in one embodiment, the settings for each of the plurality of speakers (210) may be settings for controlling the size of an output signal corresponding to at least one of the plurality of speakers (210). In one embodiment, when an application related to translation is executed, the processor (250) may check the settings for each of the plurality of speakers (210) set corresponding to the executed application. For example, the settings for each of the plurality of speakers (210) corresponding to the application related to translation may be stored in advance in a memory (e.g., the memory (130) of FIG. 1). However, it is not limited thereto, and in one embodiment, the processor (250) may check or adjust the settings for each of the plurality of speakers (210) in real time based on the acquired sound.
[0063] In one embodiment, the settings for each of the plurality of speakers (210) when executing the application may be settings for improving the recognition rate for the input sound. In one embodiment, it may be assumed that a translation-related application is executing while the user (or the caller) is on a call. When the caller's first sound is input through the microphone (220, or at least one microphone), the processor (250) may output a second sound (or loopback sound) converted (or translated) into a specified language from the caller's first sound to at least one of the plurality of speakers (210) through an ASR (automatic speech recognition) module described below.
[0064] In one embodiment, as the converted second sound of the speaker is output to at least one of the plurality of speakers (210), the user can check whether the first sound has been correctly translated or whether the user's speech corresponding to the second sound has ended. In one embodiment, the processor (250) can check whether the user's speech corresponding to the second sound has ended. For example, if the time without speech is longer than a specified time (e.g., 1.5 seconds), the processor (250) can check that the speech corresponding to the second sound has ended.
[0065] In one embodiment, when a second sound converted into at least one of a plurality of speakers (210) is output, the converted second sound may be input into at least one microphone (220), and as the converted second sound and the sound of the speaker are input together, the recognition rate of the sound of the speaker may decrease. In one embodiment, in order to improve the recognition rate of the input sound, the setting for each of the plurality of speakers (210) according to executing the application may be a setting that reduces the size of the output of at least one speaker adjacent to at least one microphone (220). Alternatively, the setting for each of the plurality of speakers (210) may be a setting that turns off at least one speaker adjacent to at least one microphone (220). The setting for each of the plurality of speakers (210) will be described in detail with reference to FIGS. 4A and 4B.
[0066] In one embodiment, the processor (250) may use an automatic speech recognition (ASR) module to convert a first sound corresponding to a first language into a second sound corresponding to a second language. In one embodiment, the ASR module may be a module that analyzes an input sound, converts it into text (or text-type data), and converts the converted text-type data into a sound corresponding to a specified language (e.g., a text-to-speech (TTS) function) and outputs it.
[0067] In one embodiment, the ASR module may be stored in the electronic device (200), but is not limited thereto, and in one embodiment, the ASR module may be stored in an external device (e.g., the server (108) of FIG. 1). In one embodiment, when the ASR module is stored in the server, the processor (250) may transmit the first sound to the server through the communication module (240), or receive the second sound converted by the ASR module from the server. For example, when the ASR module is stored in the electronic device (200), the processor (250) may identify the second sound through the ASR module stored in the electronic device (200). Alternatively, when the ASR module is stored in the server, the processor (250) may obtain information about the second sound from the server through the communication module (240).
[0068] In one embodiment, when a first sound corresponding to a first language is input (or received) through at least one microphone (220), the processor (250) can obtain a second sound converted into a second language from the first sound through the ASR module. For example, it can be assumed that a sound including Korean speech is received through at least one microphone (220). The ASR module can convert the received Korean speech into text, and convert the converted text into a sound corresponding to a specified language (e.g., English). According to one embodiment, the process of converting into text may be omitted.
[0069] In one embodiment, the processor (250) may output the converted sound through at least one of a plurality of speakers based on the identified settings corresponding to the translation-related application.
[0070] FIG. 3 is a flowchart for explaining an operation method of an electronic device (e.g., electronic device (200) of FIG. 2) according to one embodiment.
[0071] Hereinafter, a method of operating an electronic device according to various embodiments will be described in detail. According to various embodiments, the operations performed by the electronic device described below may be executed by a processor (e.g., processor 250 of FIG. 2 ) including at least one processing circuitry of the electronic device. According to one embodiment, the operations performed by the electronic device may be stored in a memory (e.g., memory 230 of FIG. 2 ) and, when executed, may be executed by instructions that cause the processor 250 to operate. In the embodiments below, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Depending on the implementation, certain operations may be omitted.
[0072] Referring to FIG. 3, according to one embodiment, in operation 301, the electronic device may execute an application related to translation. In one embodiment, the electronic device may perform a translation of sound (e.g., a speaker's sound) input through a microphone included in the electronic device (e.g., at least one microphone (220) of FIG. 2) using the application related to translation.
[0073] According to one embodiment, in operation 303, the electronic device may check the settings for each of a plurality of speakers (e.g., the plurality of speakers (210) of FIG. 2) set in response to the application. In one embodiment, when an application related to translation is executed, the electronic device may check the settings for each of the plurality of speakers corresponding to the application. In one embodiment, the settings for each of the plurality of speakers may be stored in a memory (e.g., the memory (130) of FIG. 1). However, the present invention is not limited thereto, and in one embodiment, the settings for each of the plurality of speakers may be changed in real time.
[0074] In one embodiment, the settings for each of the plurality of speakers when the application is executed may be settings that reduce the output volume of at least one speaker adjacent to at least one microphone. Alternatively, the settings for each of the plurality of speakers may be settings that turn off at least one speaker adjacent to at least one microphone.
[0075] In one embodiment, in operation 305, the electronic device may output a second sound (or loopback sound) corresponding to a second language through at least one speaker among a plurality of speakers based on the identified settings when a first sound corresponding to a first language is input through at least one microphone. In this case, the second sound corresponding to the second language may be a sound converted from the input first sound through an automatic speech recognition (ASR) module.
[0076] In one embodiment, it can be assumed that a sound including Korean speech is input through at least one microphone. The electronic device can convert the sound including Korean speech into text-type data using an ASR module. The electronic device can convert the text-type data into a second sound corresponding to English using the ASR module. The electronic device can output the second sound corresponding to English based on the settings for each of the plurality of speakers described above. For example, the electronic device can control the output of the plurality of speakers so that the second sound is not output from a first speaker adjacent to a microphone included in the electronic device, but is output through a second speaker other than the first speaker.
[0077] Accordingly, it is possible to address issues that may arise due to loopback sound (e.g., reduced recognition rate associated with translated sound) in order to improve user usability in relation to the voice translation function.
[0078] FIG. 4A is a diagram for explaining the settings for each of a plurality of speakers (e.g., the plurality of speakers (210) of FIG. 2) according to one embodiment. FIG. 4B is a diagram for explaining a method for outputting sound according to one embodiment.
[0079] Referring to FIG. 4A, according to one embodiment, an electronic device (400, e.g., electronic device (200) of FIG. 2) may include a plurality of speakers (45 and 46) including a first speaker (45) and a second speaker (46) and a plurality of microphones (401, 402, 403, and 404, e.g., at least one microphone (220) of FIG. 2).
[0080] In one embodiment, the electronic device may determine the settings for each of the plurality of speakers (45 and 46) based on a recognition rate (e.g., the recognition rate of FIG. 2) corresponding to the second sound (or loopback sound). In one embodiment, the recognition rate corresponding to the second sound may refer to the accuracy of conversion corresponding to the second sound when the first sound corresponding to the first language (e.g., the first sound of FIG. 2) input through at least one microphone (at least one of 401, 402, 403, and 404) included in the electronic device (400) is converted into the second sound corresponding to the second language. However, the present invention is not limited thereto.
[0081] In one embodiment, when a second sound converted into at least one of a plurality of speakers (45 and 46) is output, the converted sound may be input into at least one microphone (at least one of 401, 402, 403, and 404), and as the converted sound and the speaker's sound are input together, the recognition rate for the speaker's sound may decrease. For example, it may be assumed that a translation-related application (e.g., the translation-related application of FIG. 2) is executed. When the second sound is output into a plurality of speakers (45 and 46), the output second sound may be input again through the microphones (401 and 402) adjacent to the plurality of speakers (45 and 46). In this case, a situation (or, Double talk condition) may occur in which the second sound output through the speaker is input into the same microphone together with the user's sound, and the recognition rate corresponding to the second sound may decrease.
[0082] According to one embodiment, the electronic device (400) may check the settings for each of the plurality of speakers (45 and 46) to ensure that the recognition rate corresponding to the second sound is not reduced. For example, when a translation-related application is executed, the electronic device (400) may check the settings so that the output for each of the plurality of speakers (45 and 46) is less than a specified size. Alternatively, for example, the electronic device (400) may check the first microphone (402) having the highest recognition rate corresponding to the second sound, and check the settings to turn off the second speaker (46) adjacent to the checked first microphone (402). This will be described in detail with reference to FIG. 5.
[0083] Referring to FIG. 4b, according to one embodiment, the electronic device (400) may include a plurality of modules for outputting sound.
[0084] In one embodiment, the electronic device (400) may include a first module (49) that acquires the sound of a caller and a second module (47) that acquires and outputs the sound of a caller (or a callee). In one embodiment, the first module (49) may include a third module (42) that processes the acquired sound of the caller when the sound of the caller is acquired. In one embodiment, the second module (47) may include an ASR module (41, for example, the ASR module of FIG. 2) that performs conversion on the input sound. In one embodiment, the second module (47) may include a fourth module (44) that processes the acquired sound of the caller (40) when the sound of the caller is acquired. In one embodiment, the second module (47) may include a mixing module (43) that mixes the sound (or audio signal corresponding to the sound) converted through the ASR module (41) and the sound of the caller processed through the fourth module (44). In one embodiment, the second module (47) may include a plurality of speakers (45, 46, and 48, e.g., the plurality of speakers (210) of FIG. 2) that output the sound mixed through the mixing module (43).
[0085] In one embodiment, when the electronic device (400) receives the sound (40) of the signer corresponding to the first language, the electronic device (400) can use the ASR module (41) to confirm that the received sound (40) is converted into a sound of the second language, and output the confirmed sound through at least one speaker among the plurality of speakers (45, 46, 47, and 48). In one embodiment, when the sound of the caller corresponding to the second language is acquired through the third module (42), the electronic device (400) can use the ASR module (41) to confirm that the acquired sound of the caller is converted into a sound of the first language, and output the confirmed sound through at least one speaker among the plurality of speakers (45, 46, 47, and 48).
[0086] In one embodiment, when a first sound of a speaker is input into at least one microphone, the electronic device (400) can identify a second sound corresponding to a second language using the ASR module (41). In this case, an echo that may affect a recognition rate corresponding to the second sound may be input into at least one microphone together with the first sound. In one embodiment, the echo that may affect a recognition rate corresponding to the second sound may include at least one of the speaker's sound (40), a sound converted from the speaker's sound (40) through the ASR module (41), and a second sound converted from the speaker's first sound. In one embodiment, the electronic device (400) can identify a microphone with the highest recognition rate corresponding to the second sound by checking the ratio of echoes input into each of the at least one microphones. This will be described in detail with reference to FIG. 5 described below.
[0087] FIG. 5 is a flowchart illustrating a method for outputting a second sound (e.g., the second sound of FIG. 2) according to one embodiment.
[0088] According to FIG. 5, in one embodiment, in operation 501, an electronic device (e.g., electronic device (200) of FIG. 2) can identify, among at least one microphone (e.g., at least one microphone (220) of FIG. 2), a first microphone having the highest recognition rate corresponding to a second sound (e.g., recognition rate corresponding to a second sound of FIG. 4) is identified.
[0089] In one embodiment, the electronic device can identify a first microphone having the highest recognition rate corresponding to a second sound based on a signal-to-echo ratio (SER) value corresponding to each of at least one microphone. In one embodiment, the electronic device can obtain the SER value using an echo (e.g., an echo of FIG. 4) corresponding to each of at least one microphone. In one embodiment, the echo may include at least one of a sound of a listener (e.g., a sound of a listener (40) of FIG. 4), a sound converted from a sound of a listener through an ASR module (e.g., an ASR module of FIG. 2), and a second sound converted from a first sound of a speaker (e.g., a first sound of FIG. 2). In one embodiment, the electronic device can calculate a ratio of echoes among sounds input through each of at least one microphone to identify a SER value, and identify a microphone having the highest identified SER value as the first microphone. However, the present invention is not limited thereto.
[0090] In one embodiment, the electronic device may identify the microphone with the relatively greatest distance from the plurality of speakers as the first microphone. Alternatively, in one embodiment, the electronic device may identify the microphone with the largest difference between the size of the echo and the size of the sound excluding the echo as the first microphone.
[0091] In one embodiment, in operation 503, the electronic device may identify a setting for turning off a first speaker adjacent to the identified first microphone or reducing the size of an output signal of the first speaker among a plurality of speakers (e.g., the plurality of speakers (210) of FIG. 2). In one embodiment, the electronic device may identify a first speaker adjacent to the first microphone. In one embodiment, a memory (e.g., the memory (130) of FIG. 1) may include information about each of at least one microphone included in the electronic device and an adjacent speaker. The electronic device may identify a first speaker adjacent to the first microphone and identify a setting for turning off the identified first speaker or reducing the size of an output signal of the first speaker.
[0092] In one embodiment, in operation 505, the electronic device may output a second sound based on a setting for the identified first speaker. In one embodiment, if a setting for turning off the first speaker is identified, the electronic device may output the second sound through the remaining speakers excluding the first speaker among the plurality of speakers. In one embodiment, if a setting for reducing the size of the output signal of the first speaker to a specified size is identified, the electronic device may control the first speaker so that the second sound is output with an output signal of the specified size.
[0093] According to the above example, when a second sound (Loopback sound) is output, the electronic device can identify a designated microphone and, based on the designated microphone, determine the settings for each of the multiple speakers to increase the recognition rate corresponding to the second sound. Accordingly, even under double talk conditions, where the translated sound of the speaker is input along with the speaker's sound, the sound can be translated with a high recognition rate.
[0094] As shown in Table 1 below, according to the above-described embodiments, it can be confirmed that the recognition rate corresponding to the second sound is relatively higher when the first speaker is turned off or the size of the output signal of the first speaker is reduced compared to the case where the setting value is not applied. Here, the 'Pub noise environment' means an environment where the electronic device is located in a cafe, and the 'xRoad noise environment' means an environment where the electronic device is located on the side of the road. In one embodiment, the case where the setting value is not applied may mean a case where the output signal of the first speaker is not changed (for example, a case where the size of the output signal is reduced or the first speaker is not turned off).
[0095] When the recognition rate setting value is not applied, the size of the output signal of the first speaker is reduced, and the first speaker is turned off. Quiet environment 88.9% 91.0% 95.6% Pub noise environment 80.7% 82.2% 92.6% x Road noise environment 71.8% 75.4% 87.4%
[0096] FIG. 6 is a flowchart illustrating a method for outputting a second sound (e.g., the second sound of FIG. 2) according to one embodiment.
[0097] According to FIG. 6, in one embodiment, in operation 601, an electronic device (e.g., electronic device (200) of FIG. 2) can identify a second speaker among a plurality of speakers (e.g., a plurality of speakers (210) of FIG. 2) that lowers a recognition rate corresponding to a second sound (e.g., a recognition rate corresponding to a second sound of FIG. 2).
[0098] In one embodiment, the electronic device may identify at least one speaker adjacent to at least one microphone included in the electronic device (e.g., at least one microphone (220) of FIG. 2) among a plurality of speakers as a second speaker. In one embodiment, a memory (e.g., a memory (130) of FIG. 1) may contain information about each of at least one microphone included in the electronic device and a speaker adjacent thereto, and the electronic device may identify the second speaker based on the information stored in the memory. Alternatively, in one embodiment, the electronic device may identify a speaker, among a plurality of speakers, whose recognition rate corresponding to the second sound decreases as the magnitude of the output signal increases, as the second speaker.
[0099] In one embodiment, at operation 603, the electronic device may determine a setting that turns off the identified second speaker or reduces the magnitude of an output signal of the second speaker.
[0100] In one embodiment, at operation 605, the electronic device may output a second sound based on the settings for the identified second speaker.
[0101] For example, referring to FIG. 4A, the electronic device can identify a speaker (45) adjacent to a microphone (401) as a second speaker, and can identify a speaker (46) adjacent to a microphone (402) as a second speaker. The electronic device can turn off the identified second speakers (45 and 46) or reduce the size of an output signal of the second speakers to a specified size. However, the present invention is not limited thereto, and the electronic device can also control the output of any one of the identified second speakers (45 and 46). For example, the electronic device can identify a speaker (46) among the identified second speakers (45 and 46) that has a relatively greater influence on the recognition rate, and turn off the identified speaker (46) or reduce the size of an output signal of the second speaker to a specified size.
[0102] In one embodiment, the electronic device may output a third sound, different from the second sound, through the second speaker based on the detection of the second speaker. In one embodiment, the third sound may be a beep sound, but is not limited thereto. In one embodiment, the third sound may be various types of sound to inform the user that the second sound is currently being provided to the receiver.
[0103] FIG. 7 is a flowchart illustrating a method for verifying weights associated with a transformation according to one embodiment.
[0104] According to FIG. 7, in one embodiment, in operation 701, an electronic device (e.g., the electronic device (200) of FIG. 2) may identify a second speaker (e.g., the second speaker of FIG. 6) that reduces a recognition rate corresponding to a second sound (e.g., the recognition rate of FIG. 2) among a plurality of speakers (e.g., the plurality of speakers (210) of FIG. 2). In one embodiment, the electronic device may identify at least one speaker adjacent to at least one microphone (e.g., at least one microphone (220) of FIG. 2) included in the electronic device as the second speaker among the plurality of speakers. Alternatively, in one embodiment, the electronic device may identify a speaker, of which a recognition rate corresponding to a second sound (e.g., the second sound of FIG. 2) reduces as a magnitude of an output signal of the speaker increases among the plurality of speakers, as the second speaker.
[0105] In one embodiment, in operation 703, the electronic device may identify a second microphone coupled to a second speaker among at least one microphone. In one embodiment, the second microphone coupled to the second speaker may be a microphone in which the volume of sound input to the second microphone decreases as the volume of sound output through the second speaker decreases, or the volume of sound input to the second microphone increases as the volume of sound output through the second speaker increases. In one embodiment, there may be a plurality of second microphones coupled to the second speaker.
[0106] According to one embodiment, in operation 705, the electronic device may determine a weight associated with the conversion corresponding to the identified second microphone. In one embodiment, the weight associated with the conversion may be a ratio at which sounds input from each of at least one microphone are used for conversion when a first sound (e.g., the first sound of FIG. 2) is converted into a second sound through an ASR module (e.g., the ASR module of FIG. 2). For example, it may be assumed that there are three microphones included in the electronic device, and the weights associated with the conversion corresponding to each of the three microphones are 0.1, 0.3, and 0.6, respectively. In this case, the ASR module may apply a weight corresponding to each microphone to sounds input through each microphone, and may obtain a second sound using each sound to which the weights have been applied.
[0107] In one embodiment, the weight corresponding to each of at least one microphone may be a preset value. In one embodiment, the weight corresponding to each of at least one microphone may be stored in a memory (e.g., memory (130) of FIG. 1). When a second microphone is identified, the electronic device may change the weight corresponding to the second microphone. For example, the electronic device may lower the weight corresponding to the second microphone to a designated value. However, the present invention is not limited thereto, and even when the weight corresponding to each of at least one microphone is not stored in the memory, the electronic device may confirm the weight corresponding to the identified second microphone to a designated value when the second microphone is identified.
[0108] In one embodiment, the electronic device can identify a second sound corresponding to a second language through the ASR module based on the identified weights. For example, suppose that there are three microphones included in the electronic device, and the weights associated with the conversion corresponding to each of the three microphones are 0.1, 0.3, and 0.6, respectively. If the weight corresponding to the second microphone among the three microphones is 0.3, the electronic device can lower the weight corresponding to the second microphone to a specified value (e.g., 0.1) and increase the weights corresponding to microphones other than the second microphone. The electronic device can identify a second sound corresponding to the second language through the ASR module based on the changed weights.
[0109] FIG. 8 is a flowchart illustrating a method for identifying a second sound (e.g., the second sound of FIG. 2) according to one embodiment.
[0110] According to FIG. 8, in one embodiment, in operation 801, an electronic device (e.g., the electronic device (200) of FIG. 2) may boost the size of sound input to a third microphone (e.g., the second microphone of FIG. 7) that is different from a second microphone (e.g., the second microphone of FIG. 2) among at least one microphone (e.g., the at least one microphone (220) of FIG. 2).
[0111] In one embodiment, the electronic device may identify a second speaker (e.g., the second speaker of FIG. 6) that reduces a recognition rate (e.g., the recognition rate of FIG. 2) corresponding to a second sound among a plurality of speakers (e.g., the plurality of speakers (210) of FIG. 2). In one embodiment, the electronic device may identify a second microphone coupled with the identified second speaker among at least one microphone. In one embodiment, the electronic device may identify at least one microphone among the remaining microphones excluding the second microphone among at least one microphone included in the electronic device as a third microphone, and may boost the volume of a sound input through the identified third microphone. For example, the electronic device may identify the remaining microphones excluding the second microphone as the third microphone.
[0112] According to one embodiment, in operation 803, the electronic device can identify the second sound based on the boosted sound. Accordingly, the relative loudness of the sound input through the second microphone coupled with the second speaker, which lowers the recognition rate, is reduced, thereby improving the phenomenon of the recognition rate being lowered due to the loopback sound output through the second speaker.
[0113] FIG. 9 is a flowchart illustrating a method for identifying a second sound (e.g., the second sound of FIG. 2) based on noise according to one embodiment.
[0114] According to FIG. 9, according to one embodiment, in operation 901, an electronic device (e.g., electronic device (200) of FIG. 2) may determine a weight (e.g., weight of FIG. 7) associated with a transformation corresponding to each of at least one microphone based on a ratio of noise included in sound input to each of at least one microphone (e.g., at least one microphone (220) of FIG. 2).
[0115] In one embodiment, the noise included in the sound may be noise generated in the surrounding environment. In one embodiment, the noise ratio may be the ratio between the first sound input through the microphone and the noise input through the microphone. In one embodiment, the electronic device may determine the noise ratio corresponding to each microphone based on the sound input through each of at least one microphones.
[0116] However, the present invention is not limited thereto, and noise included in the sound may be an echo (e.g., an echo in FIG. 4). In one embodiment, the electronic device may determine a SER value (e.g., SER in FIG. 5) corresponding to each of at least one microphone as a ratio of noise based on sound input through each of at least one microphone.
[0117] In one embodiment, the electronic device may lower the weight corresponding to a microphone with a relatively high noise ratio among at least one microphone to a designated value, and relatively increase the weights corresponding to the remaining microphones. Alternatively, in one embodiment, the electronic device may turn off the microphone with a relatively high noise ratio.
[0118] According to one embodiment, at operation 903, the electronic device may identify a second sound corresponding to the second language through an ASR module (e.g., the ASR module of FIG. 2) based on the identified weights.
[0119] For example, let's assume that there are three microphones included in an electronic device, and the weights associated with the transformation corresponding to each of the three microphones are 0.1, 0.3, and 0.6, respectively. In one embodiment, the electronic device can check the ratio of noise included in the sound input to each of the three microphones. If the ratio of noise corresponding to the third microphone among the three microphones is relatively the highest, the electronic device can lower the weight (0.6) corresponding to the third microphone to a designated value (e.g., 0.1) and increase the weights corresponding to microphones other than the third microphone. The electronic device can check the second sound corresponding to the second language through the ASR module based on the changed weights.
[0120] FIG. 10 is a flowchart for explaining an operation method of an electronic device (e.g., the electronic device (200) of FIG. 2) according to one embodiment.
[0121] Referring to FIG. 10, according to one embodiment, in operation 1001, the electronic device may check the settings (e.g., the settings for each of the plurality of speakers of FIG. 3) for each of the plurality of speakers (e.g., the plurality of speakers (210) of FIG. 2) based on satisfying a first condition for performing translation. In one embodiment, the first condition for performing translation may be a condition in which an input corresponding to the execution of a call application is confirmed. For example, the electronic device may determine that the first condition is satisfied when an input corresponding to the execution of the call application is received through an input module (e.g., the input module (150) of FIG. 1). When the electronic device determines that the first condition is satisfied, the electronic device may check the settings for each of the plurality of speakers.
[0122] In one embodiment, in operation 1003, when a first sound corresponding to a first language (e.g., the first sound of FIG. 3) is input to at least one microphone (e.g., at least one microphone (220) of FIG. 2), the electronic device may output a second sound corresponding to a second language (e.g., the first sound of FIG. 3) through at least one speaker among a plurality of speakers based on the identified settings. In this case, the second sound corresponding to the second language may be a sound converted from the input first sound through an automatic speech recognition (ASR) module.
[0123] An electronic device (101; 200) according to one embodiment of the present disclosure may include a plurality of speakers (210), at least one microphone (220), a memory (230) for storing instructions, and at least one processor (250).
[0124] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to execute an application related to translation.
[0125] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to determine a setting for each of the plurality of speakers (210) set in response to the application.
[0126] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to output, through at least one speaker of the plurality of speakers (210), a second sound corresponding to a second language, converted from the input first sound through an automatic speech recognition (ASR) module, based on the identified setting, in response to a first sound corresponding to a first language being input to the at least one microphone (220).
[0127] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to determine a setting for each of the plurality of speakers (210) based on a recognition rate corresponding to the second sound.
[0128] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to identify, among the at least one microphone (220), a first microphone having a highest recognition rate corresponding to the second sound.
[0129] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to identify a setting that turns off a first speaker adjacent to the identified first microphone among the plurality of speakers (210) or reduces the size of an output signal of the first speaker.
[0130] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to output the second sound based on the settings for the identified first speaker.
[0131] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to identify a second speaker among the plurality of speakers (210) that lowers the recognition rate corresponding to the second sound.
[0132] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to determine a setting that turns off the identified second speaker or reduces the magnitude of an output signal of the second speaker.
[0133] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to output the second sound based on the settings for the identified second speaker.
[0134] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to output a third sound, different from the second sound, through the second speaker based on identifying the second speaker.
[0135] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to identify a second speaker among the plurality of speakers (210) that lowers the recognition rate corresponding to the second sound.
[0136] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to identify a second microphone, among the at least one microphone (220), that is coupled to the second speaker.
[0137] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to identify a weight associated with the transformation corresponding to the identified second microphone.
[0138] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to identify a second sound corresponding to the second language through the ASR module based on the identified weight.
[0139] According to one embodiment, the second microphone coupled to the second speaker may be a microphone in which the volume of sound input to the second microphone decreases as the volume of sound output through the second speaker decreases, or in which the volume of sound input to the second microphone increases as the volume of sound output through the second speaker increases.
[0140] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to boost the volume of sound input to a third microphone, different from the second microphone, among the at least one microphone (220).
[0141] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to identify the second sound based on the boosted sound.
[0142] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to determine a weight associated with the transformation corresponding to each of the at least one microphone (220) based on a ratio of noise included in sound input to each of the at least one microphone (220).
[0143] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device (101; 200) to identify a second sound corresponding to the second language through the ASR module based on the identified weight.
[0144] A method of operating an electronic device (101; 200) according to one embodiment of the present disclosure may include an operation of executing an application related to translation.
[0145] According to one embodiment, the method of operation may include an operation of checking a setting for each of a plurality of speakers (210) set in response to the application.
[0146] According to one embodiment, the operating method may include an operation of, when a first sound corresponding to a first language is input to at least one microphone (220), outputting a second sound corresponding to a second language, which is converted from the input first sound through an automatic speech recognition (ASR) module, through at least one speaker among the plurality of speakers (210) based on the confirmed setting.
[0147] According to one embodiment, the operation of checking the setting may check the setting for each of the plurality of speakers (210) based on the recognition rate corresponding to the second sound.
[0148] According to one embodiment, the operation of checking the setting may include an operation of checking a first microphone (220) having the highest recognition rate corresponding to the second sound among the at least one microphone.
[0149] According to one embodiment, the operation of checking the setting may include an operation of checking a setting of turning off a first speaker adjacent to the checked first microphone among the plurality of speakers (210) or reducing the size of an output signal of the first speaker.
[0150] In one embodiment, the outputting operation may output the second sound based on the settings for the identified first speaker.
[0151] According to one embodiment, the operation of checking the setting may include an operation of checking a second speaker among the plurality of speakers (210) that lowers the recognition rate corresponding to the second sound.
[0152] In one embodiment, the action of checking the setting may include an action of checking a setting that turns off the checked second speaker or reduces the size of an output signal of the second speaker.
[0153] In one embodiment, the outputting operation may output the second sound based on the settings for the identified second speaker.
[0154] In one embodiment, the outputting operation may output a third sound different from the second sound through the second speaker based on the identification of the second speaker.
[0155] According to one embodiment, the operating method may include an operation of identifying a second speaker among the plurality of speakers (210) that lowers the recognition rate corresponding to the second sound.
[0156] According to one embodiment, the method of operation may include an operation of identifying a second microphone, among the at least one microphone (220), that is coupled to the second speaker.
[0157] According to one embodiment, the method of operation may include an operation of identifying a weight associated with the transformation corresponding to the identified second microphone.
[0158] According to one embodiment, the operating method may include an operation of identifying a second sound corresponding to the second language through the ASR module based on the identified weight.
[0159] According to one embodiment, the second microphone coupled to the second speaker may be a microphone in which the volume of sound input to the second microphone decreases as the volume of sound output through the second speaker decreases, or in which the volume of sound input to the second microphone increases as the volume of sound output through the second speaker increases.
[0160] According to one embodiment, the operating method may include an operation of boosting the size of sound input to a third microphone, which is different from the second microphone, among the at least one microphone (220).
[0161] According to one embodiment, the operating method may further include an operation of identifying the second sound based on the boosted sound.
[0162] In a storage medium (130) storing computer-readable instructions according to one embodiment of the present disclosure, the instructions, when individually or collectively executed by at least one processor (250) of an electronic device (101; 200), can cause the electronic device (101; 200) to execute an application related to translation.
[0163] According to one embodiment, the instructions, when executed individually or collectively by at least one processor (250) of the electronic device (101; 200), may cause the electronic device (101; 200) to determine a setting for each of a plurality of speakers (210) set in response to the application.
[0164] According to one embodiment, the instructions, when individually or collectively executed by at least one processor (250) of the electronic device (101; 200), may cause the electronic device (101; 200) to, when a first sound corresponding to a first language is input to at least one microphone (220), output a second sound corresponding to a second language, converted from the input first sound through an automatic speech recognition (ASR) module, through at least one speaker of the plurality of speakers (210) based on the identified settings.
[0165] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains.
[0166] As used herein, the term “if” will be understood to mean “when, upon,” “in response to deciding,” or “in response to detecting,” depending on the context. Similarly, “if it is decided to do,” or “if [the stated condition or event] is detected,” will optionally be understood to mean “upon deciding,” or “in response to deciding,” “upon detecting [the stated condition or event],” or “in response to detecting [the stated condition or event].”
[0167] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. A processing device (or processing circuit) may execute an operating system (OS) and one or more software applications running on the operating system. In addition, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0168] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.
[0169] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include an app store that distributes applications, a site that supplies or distributes various software, or a recording or storage medium managed by a server.
[0170] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components such as the described systems, structures, devices, and circuits are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0171] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.
[0172] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0173] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0174] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0175] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0176] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0177] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device (101; 200), Multiple speakers (210); At least one microphone (220); Memory (230) for storing instructions; and comprising at least one processor (250), The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device (101; 200) to: Run the application related to translation, Check the settings for each of the plurality of speakers (210) set in response to the above application, An electronic device (101; 200) that causes a second sound corresponding to a second language, converted from the input first sound through an automatic speech recognition (ASR) module based on the confirmed setting, to be output through at least one speaker among the plurality of speakers (210), in response to the input of a first sound corresponding to a first language through at least one microphone (220).
2. In paragraph 1, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device (101; 200) to: An electronic device (101; 200) that causes the setting for each of the plurality of speakers (210) to be checked based on the recognition rate corresponding to the second sound.
3. In paragraph 1 or 2, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device (101; 200) to: Among the above at least one microphone (220), the first microphone having the highest recognition rate corresponding to the second sound is identified, Among the plurality of speakers (210), check the setting for turning off the first speaker adjacent to the confirmed first microphone or reducing the size of the output signal of the first speaker, An electronic device (101; 200) that causes the second sound to be output based on the settings for the first speaker that have been confirmed above.
4. In any one of paragraphs 1 to 3, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device (101; 200) to: Among the above plurality of speakers (210), a second speaker that lowers the recognition rate corresponding to the second sound is identified, Check the setting to turn off the second speaker or reduce the size of the output signal of the second speaker, An electronic device (101; 200) that causes the second sound to be output based on the settings for the second speaker confirmed above.
5. In any one of paragraphs 1 to 4, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device (101; 200) to: An electronic device (101; 200) that causes a third sound different from the second sound to be output through the second speaker based on the confirmation of the second speaker.
6. In any one of paragraphs 1 to 5, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device (101; 200) to: Among the above plurality of speakers (210), a second speaker that lowers the recognition rate corresponding to the second sound is identified, Among the above at least one microphone (220), a second microphone coupled with the second speaker is identified, An electronic device (101; 200) causing the weight associated with the transformation corresponding to the second microphone identified above to be verified.
7. In any one of paragraphs 1 to 6, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device (101; 200) to: An electronic device (101; 200) that causes the ASR module to identify a second sound corresponding to the second language based on the identified weight.
8. In any one of paragraphs 1 to 7, The second microphone coupled with the second speaker, An electronic device (101; 200) which is a microphone in which the size of sound input to the second microphone decreases as the size of sound output through the second speaker decreases, or in which the size of sound input to the second microphone increases as the size of sound output through the second speaker increases.
9. In any one of paragraphs 1 to 8, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device (101; 200) to: Boosting the size of the sound input to a third microphone different from the second microphone among the at least one microphone (220), An electronic device (101; 200) that causes the second sound to be identified based on the boosted sound.
10. In any one of paragraphs 1 to 9, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device (101; 200) to: Based on the ratio of noise included in the sound input to each of the at least one microphone (220), the weight associated with the transformation corresponding to each of the at least one microphone (220) is determined, An electronic device (101; 200) that causes the ASR module to identify a second sound corresponding to the second language based on the identified weight.
11. In the operating method of an electronic device (101; 200), The action of running an application related to translation; An operation of checking the settings for each of the plurality of speakers (210) set in response to the above application; and An operating method comprising: when a first sound corresponding to a first language is input to at least one microphone (220), an operation of converting the input first sound through an automatic speech recognition (ASR) module and outputting a second sound corresponding to a second language through at least one speaker among the plurality of speakers (210) based on the confirmed settings.
12. In paragraph 11, The action to check the above settings is: An operating method for checking the settings for each of the plurality of speakers (210) based on the recognition rate corresponding to the second sound.
13. In paragraph 11 or 12, The action to check the above settings is: An operation of identifying a first microphone (220) among the above at least one microphone (220) having the highest recognition rate corresponding to the second sound; and An operation of checking a setting for turning off the first speaker adjacent to the confirmed first microphone among the plurality of speakers (210) or reducing the size of the output signal of the first speaker; The above output action is, An operating method for outputting the second sound based on the settings for the first speaker confirmed above.
14. In any one of paragraphs 11 to 13, The action to check the above settings is: An operation of identifying a second speaker among the plurality of speakers (210) that lowers the recognition rate corresponding to the second sound; and An operation of checking a setting for turning off the second speaker or reducing the size of the output signal of the second speaker; The above output action is, An operating method for outputting the second sound based on the settings for the second speaker confirmed above.
15. In any one of paragraphs 11 to 14, The above output action is, An operating method for outputting a third sound different from the second sound through the second speaker based on the confirmation of the second speaker.
Citation Information
Patent Citations
Remote control device
JP2021052250A
Method of providing mother tongue service for multicultural family using communication terminal
KR101679825B1
Method and apparatus for controlling audio signal in portable terminal
KR1020140081445A
Composition for enamel, method for preparation thereof and cooking appliance
KR102172460B1
Blackbox device for vehicle supporting smart report and control method using the same
KR102215629B1