Electronic device for providing interpretation function, operation method thereof, and storage medium
The electronic device and wearable devices use bone conduction sensors and intelligent microphone management to address noise interference and exposure challenges, ensuring accurate real-time interpretation by distinguishing user and other party's speech and optimizing content delivery.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-09-15
- Publication Date
- 2026-05-15
AI Technical Summary
Existing real-time simultaneous interpretation systems face challenges in distinguishing between a user's speech and the other party's speech, managing ambient noise interference, and providing interpretation content effectively when the electronic device is not exposed, such as being in a bag or pocket.
An electronic device and wearable devices, like earphones and watch phones, use bone conduction sensors and intelligent microphone management to identify user speech, minimize noise interference, and adapt output based on device state, ensuring accurate real-time interpretation.
The system effectively distinguishes between user and other party's speech, reduces noise interference, and optimizes interpretation content delivery based on device exposure, enhancing the reliability and usability of real-time simultaneous interpretation.
Smart Images

Figure KR2025014343_15052026_PF_FP_ABST
Abstract
Description
Electronic device for providing interpretation function, method of operation thereof, and storage medium
[0001] One embodiment disclosed in this document relates to an electronic device for providing an interpretation function, a method of operation thereof, and a storage medium.
[0002] Various services and additional functions provided through user terminals, such as electronic devices like smartphones, are gradually increasing. In order to enhance the utility value of these electronic devices and satisfy the needs of various users, telecommunications service providers or electronic device manufacturers are competitively developing electronic devices that provide various functions.
[0003] As artificial intelligence technology advances, users input voice into electronic devices, and the devices perform actions based on that voice input. A voice recognition service is a service that provides various content services to the user in response to the user's voice received by the electronic device. To provide voice recognition services, technologies that recognize and analyze human language (e.g., machine translation, automatic speech recognition, question answering, and / or speech synthesis) are implemented in electronic devices; for example, the electronic device can provide a real-time interpretation function.
[0004] According to one embodiment, the electronic device may include at least one first microphone, a communication circuit configured to support a short-range communication method, at least one processor, and a memory for storing instructions.
[0005] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to receive a first voice signal corresponding to a self-utterance from a first wearable electronic device connected to the electronic device through the communication circuit while an interpretation application is executed on the electronic device, and the self-utterance is obtained through at least one second microphone of the first wearable electronic device.
[0006] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to output first interpretation information for the first voice signal using the interpretation application in response to the reception of the first voice signal.
[0007] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to receive a second voice signal corresponding to a counterpart's speech through the at least one first microphone while the interpretation application is executed on the electronic device.
[0008] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to acquire second interpretation information for the second voice signal using the interpretation application in response to the reception of the second voice signal.
[0009] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to transmit second interpretation information for the second voice signal to the first wearable electronic device through the communication circuit.
[0010] According to one embodiment, a method for providing an interpretation function in an electronic device includes receiving a first voice signal corresponding to a self-utterance from a first wearable electronic device connected to the electronic device while an interpretation application is running in the electronic device, and the self-utterance may be obtained through at least one second microphone of the first wearable electronic device.
[0011] According to one embodiment, the method may include an operation of outputting first interpretation information for the first voice signal using the interpretation application in response to the reception of the first voice signal.
[0012] According to one embodiment, the method may include the operation of receiving a second voice signal corresponding to a counterpart's speech through at least one first microphone of the electronic device while the interpretation application is running on the electronic device.
[0013] According to one embodiment, the method may include an operation of obtaining second interpretation information for the second voice signal using the interpretation application in response to the reception of the second voice signal.
[0014] According to one embodiment, the method may include the operation of transmitting second interpretation information for the second voice signal to the first wearable electronic device.
[0015] According to one embodiment, in a non-transient storage medium storing computer-readable instructions, the instructions are configured to cause the electronic device to perform at least one operation when executed by at least one processor of the electronic device, wherein the at least one operation includes receiving a first voice signal corresponding to a self-utterance from a first wearable electronic device connected to the electronic device while an interpretation application is executed on the electronic device, and the self-utterance may be obtained through at least one second microphone of the first wearable electronic device.
[0016] According to one embodiment, the at least one operation may include an operation of outputting first interpretation information for the first voice signal using the interpretation application in response to the reception of the first voice signal.
[0017] According to one embodiment, the at least one operation may include receiving a second voice signal corresponding to a counterpart's speech through at least one first microphone of the electronic device while the interpretation application is running on the electronic device.
[0018] According to one embodiment, the at least one operation may include an operation of obtaining second interpretation information for the second voice signal using the interpretation application in response to the reception of the second voice signal.
[0019] According to one embodiment, the at least one operation may include transmitting second interpretation information for the second voice signal to the first wearable electronic device.
[0020] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment.
[0021] FIG. 2 is a configuration diagram of a system for providing an interpretation function according to one embodiment.
[0022] FIG. 3 is an internal block diagram of each of an electronic device, a first wearable electronic device, and a second wearable electronic device for providing an interpretation function according to one embodiment.
[0023] FIG. 4 is an operation flowchart of an electronic device for providing an interpretation function according to one embodiment.
[0024] FIG. 5 is a diagram showing signals transmitted and received to provide an interpretation function between an electronic device according to one embodiment, a first wearable electronic device, and a second wearable electronic device.
[0025] FIG. 6 is an operation flowchart of a first wearable electronic device according to one embodiment.
[0026] FIG. 7 is a diagram illustrating microphone optimization based on the other party's speech according to one embodiment.
[0027] FIG. 8 is a diagram illustrating a method for outputting interpretation information according to one embodiment.
[0028] FIG. 9 is an example diagram illustrating a method for outputting interpretation information using a first wearable electronic device in a state where the electronic device according to one embodiment is not exposed to the outside.
[0029] FIG. 10 is an example diagram illustrating a method of outputting interpretation information using a second wearable electronic device in a state where the electronic device according to one embodiment is not exposed to the outside.
[0030] FIG. 11 is an example diagram illustrating a method of outputting interpretation information using an electronic device while the electronic device is exposed to the outside according to one embodiment.
[0031] FIG. 12 is an example diagram illustrating a method of outputting interpretation information using a first wearable electronic device while the electronic device according to one embodiment is exposed to the outside.
[0032] FIG. 13 is an example diagram illustrating a method for outputting information that induces a user to speak interpretation information using a first wearable electronic device according to one embodiment.
[0033] FIG. 14 is an example diagram illustrating the operation of an electronic device when a person speaks in the other person's language according to one embodiment.
[0034] FIG. 15a is an example diagram illustrating a method of sharing interpretation information with a counterparty device according to one embodiment.
[0035] Figure 15b is a drawing that follows Figure 15a.
[0036] FIG. 16 is an example diagram illustrating the output direction of interpretation information corresponding to the position of the speaker according to one embodiment.
[0037] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0038] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to one embodiment. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through the server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0039] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0040] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0041] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, software (e.g., program (140)) and input data or output data for related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0042] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0043] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0044] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0045] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0046] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0047] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0048] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0049] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0050] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0051] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0052] The power management module (188) can manage the power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0053] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0054] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0055] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.
[0056] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0057] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0058] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0059] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used.
[0060] The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199). The electronic device (101) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0061] In the following detailed description, reference numbers in the drawings may be assigned identically or omitted for configurations that can be easily understood through prior embodiments, and detailed descriptions thereof may also be omitted. An electronic device (101) according to one embodiment disclosed in this document may be implemented by selectively combining configurations of different embodiments, and a configuration of one embodiment may be replaced by a configuration of another embodiment. For example, it should be noted that the present invention is not limited to specific drawings or embodiments.
[0062] In systems for simultaneous interpretation that support real-time simultaneous interpretation technology, the performance of speech recognition and automatic translation technologies can be critical. Since the environment where simultaneous interpretation takes place requires converting a foreign language into the native language for the user and the user's language into the other person's language for the other person, it is important to distinguish between the user's speech and the other person's speech, and to translate them into the corresponding languages. For example, when using an interpretation application, users must manually control the microphone and change the settings to select the speaker every time they converse, so it can be cumbersome to manually adjust the timing settings in a real-time conversation environment.
[0063] In order to provide real-time interpretation functions while a user is wearing a wearable electronic device (e.g., earphones) connected to an electronic device, it must be possible to recognize whether the voice signal received through the wearable electronic device is the user's own speech or the other party's speech. Additionally, when a user speaks while wearing the wearable electronic device (e.g., earphones), the user's voice may be input not only through the microphone of the wearable electronic device but also through the microphone of the electronic device; therefore, it is necessary to minimize interference from ambient noise other than the user's voice. Furthermore, when the electronic device is placed in a bag or pocket, there may be limitations in receiving voice signals through the electronic device's microphone or outputting interpretation content through the electronic device's speaker.
[0064] According to one embodiment, when a user runs an interpretation application while wearing a wearable electronic device (e.g., earphones) connected to an electronic device, an electronic device for providing an interpretation function, a method of operation thereof, and a storage medium may be provided so that interpretation content can be provided through the electronic device or the wearable electronic device according to the state of the electronic device while minimizing noise interference according to the user's speech or the other party's speech.
[0065] According to one embodiment, in a real-time simultaneous interpretation environment, while a user is wearing a wearable electronic device (e.g., earphones), it is possible to check whether the user's speech or the other party's speech is being spoken, and when the user speaks, only the microphone of the wearable electronic device is turned on, thereby minimizing interference and mistranslation caused by ambient noise other than the voice signal from the user's speech.
[0066] FIG. 2 is a configuration diagram of a system for providing an interpretation function according to one embodiment.
[0067] Referring to FIG. 2, a system (200) for providing interpretation functions may include an electronic device (201) and at least one external device (210) connected to the electronic device (201).
[0068] According to one embodiment, the electronic device (201) may be connected to at least one external device (210) (e.g., a first wearable electronic device (e.g., earphones) (211), a second wearable electronic device (e.g., a watch phone) (212), or a third wearable electronic device (213)) based on a short-range wireless communication method (e.g., Bluetooth) to provide a real-time simultaneous interpretation function.
[0069] According to one embodiment, at least one external device (210) may have a configuration similar to or identical to at least a part of the electronic device (201).
[0070] According to one embodiment, the first wearable electronic device (211) may be in a state of being connected to (or paired with) the electronic device (201) in communication while being worn on the user's ear. The first wearable electronic device (211) may check whether it is being worn before the interpretation application is executed on the electronic device (201) or in response to the execution of the interpretation application, and the timing of checking whether it is being worn is not limited thereto. While the interpretation application is being executed on the electronic device (201), the first wearable electronic device (211) may receive a voice signal through at least one microphone and check whether the user is speaking from the voice signal.
[0071] According to one embodiment, a bone conduction sensor may be used to check whether user speech is being produced in the first wearable electronic device (211). The first wearable electronic device (211) can check whether there is a vibration signal detected through the bone conduction sensor while receiving a voice signal. If a vibration signal is detected through the bone conduction sensor while receiving a voice signal, the first wearable electronic device (211) can check whether the detected vibration signal is due to user speech. For example, since a vibration signal can be generated by chewing motions rather than user speech, the first wearable electronic device (211) can check whether the pattern of the vibration signal based on deep learning corresponds to a vibration pattern representing user speech. If the pattern of the vibration signal corresponds to a vibration pattern representing user speech, the first wearable electronic device (211) can determine that the voice signal received through the microphone is a voice signal corresponding to user speech.
[0072] The electronic device (201) receives a voice signal corresponding to the user's own speech from the first wearable electronic device (211) and can provide interpretation information that translates the received voice signal using an interpretation application.
[0073] According to one embodiment, the electronic device (201) can convert a voice signal into text data using a speech-to-text (STT) model by using an interpretation application, and can generate (or obtain) interpretation information by translating the converted text data into a first language (e.g., the other person's language). The interpretation information corresponding to one's own speech can be output in the form of text or voice.
[0074] According to one embodiment, when a first wearable electronic device (211) connected to an electronic device (201) acquires a voice signal corresponding to a user's speech, it may perform an operation to check whether the acquired voice signal corresponds to the user's speech. In response to the first wearable electronic device (211) checking whether the speech is the user's speech, the electronic device (201) may generate interpretation information for the spoken sentence. When the interpretation of the spoken sentence is completed by generating interpretation information, the electronic device (201) may determine a device to output the interpretation information. For example, the electronic device (201) may output interpretation information corresponding to the user's speech through the speaker or display of the electronic device (201), but in a situation where the electronic device (201) is located in a bag or pocket, the interpretation information may be output through the speaker of another device connected to the electronic device (201), such as the second wearable electronic device (212).
[0075] According to one embodiment, when a user runs an interpretation application while wearing a first wearable electronic device (211) connected to an electronic device (201), the first wearable electronic device (211) located close to the user's mouth confirms whether the user is speaking, and in response to confirming that it is a user's speech, interpretation content corresponding to the user's speech can be provided through the electronic device (201) or the second wearable electronic device (212) depending on the state of the electronic device (201).
[0076] FIG. 3 is an internal block diagram of each of an electronic device, a first wearable electronic device, and a second wearable electronic device for providing an interpretation function according to one embodiment.
[0077] According to one embodiment, the electronic device (201) can control the operation of a first wearable electronic device (211) connected via communication to provide a real-time simultaneous interpretation function. Additionally, when a second wearable electronic device (212) is connected via communication, the electronic device (201) can control the operation of the first wearable electronic device (211) as well as the operation of the second wearable electronic device (212).
[0078] According to one embodiment, the first wearable electronic device (211) (e.g., the first wearable electronic device (211) of FIG. 2) may include at least one processor (321), memory (331), communication circuit (391), at least one second speaker (356), at least one second microphone (351), and / or bone conduction sensor (377). For example, if the first wearable electronic device (211) can be worn on a user's body (e.g., ear), the first wearable electronic device (211) may also be referred to as an earphone, earpiece, earbud, or auditory device. If the first wearable electronic device (211) is a wireless earphone that can be worn on a user's body (e.g., ear), the first wearable electronic device (211) may be composed of a pair of devices (e.g., an earphone that can be worn on the right ear, an earphone that can be worn on the left ear), and the pair of devices may include substantially identical configurations. Without being limited to what is described, the first wearable electronic device (211) may be a wireless earphone of various forms.
[0079] According to one embodiment, the memory (331) can store instructions that control the processor (321) to perform various operations during execution.
[0080] According to one embodiment, when the first wearable electronic device (211) is placed on the user's ear and the interpretation application is executed, the processor (321) can turn on (or activate) at least one second microphone (351). When the processor (321) receives a voice signal through at least one second microphone (351) during the execution of the interpretation application, it can determine whether the received voice signal is a voice signal corresponding to the user's speech.
[0081] According to one embodiment, the processor (321) can verify whether the user is speaking by using a bone conduction sensor (377). For example, the processor (321) can verify whether there is a vibration signal detected through the bone conduction sensor (377) while receiving a voice signal through at least one second microphone (351). If a vibration signal is detected through the bone conduction sensor (377) while receiving a voice signal, the processor (321) can verify whether the detected vibration signal is due to the user's speech. In response to confirming that the detected vibration signal is due to the user's speech, the processor (321) can transmit a voice signal corresponding to the user's speech to the electronic device (201) through the communication circuit (391). According to one embodiment, if the processor (321) confirms that the detected vibration signal is not due to the user's speech, it may not transmit the voice signal received through at least one second microphone (351) to the electronic device (201). For example, if the voice signal is not a voice signal produced by one's own speech, there is no need to translate the voice signal, so the voice signal may not be transmitted to the electronic device (201).
[0082] In the case of self-speech, when a user speaks while wearing the first wearable electronic device (211), a situation may occur where the voice becomes excessively loud because the user cannot hear their own voice well due to the wearing of the first wearable electronic device (211). The first wearable electronic device (211) may provide an active noise canceling (ANC) function to block external noise and an ambient sound listening function. When self-speech is detected, the electronic device (201) may control the ambient sound listening function to be turned on (or activated) so that the user can recognize their own voice. Additionally, there may be a situation where the other party continues to speak while the interpretation content corresponding to the other party's speech is being played through the first wearable electronic device (211). In such cases, since the voice corresponding to the other party's speech and the interpretation content corresponding to the other party's speech may enter the user's ears simultaneously, the inflow of voice heard from outside the first wearable electronic device (211) can be blocked by turning on (or activating) the ANC function.
[0083] According to one embodiment, in response to confirming that it is a self-speech, the processor (321) may control the at least one second microphone (351) closest to the speaker's mouth to remain on (or activated). According to one embodiment, the operation of other microphones other than the at least one second microphone (351) of the first wearable electronic device (211) may be restricted so as to minimize interference and mistranslation caused by ambient noise while the at least one second microphone (351) remains on.
[0084] For example, in response to the electronic device (201) receiving a voice signal corresponding to its own speech received through at least one second microphone (351), the receiving sensitivity (or receiving level) of at least one first microphone (350) may be lowered or turned off to reduce interference caused by a signal received through other microphones, such as at least one first microphone (350) of the electronic device (201). For example, the operation of at least one first microphone (350) may be restricted while the voice signal corresponding to its own speech is being received.
[0085] According to one embodiment, when the electronic device (201) and the second wearable electronic device (212) are in a state of communication connection, the electronic device (201) can control the operation of at least one third microphone (352) of the second wearable electronic device (212) to lower or turn off the reception sensitivity of at least one third microphone (352) while receiving a voice signal corresponding to self-speech from the first wearable electronic device (211).
[0086] According to one embodiment, the second wearable electronic device (212) (e.g., the second wearable electronic device (212) of FIG. 2) may include at least one processor (322), a memory (332), a communication circuit (392), at least one third speaker (357), and / or at least one third microphone (352). For example, if the second wearable electronic device (212) can be worn on a user's body (e.g., wrist), the second wearable electronic device (212) may be referred to as a watch phone.
[0087] According to one embodiment, the processor (322) can communicate with the electronic device (201) based on a short-range wireless communication method through the communication circuit (392). The processor (322) can receive a voice signal corresponding to the other party's speech through at least one third microphone (352) and can transmit the received voice signal to the electronic device (201) through the communication circuit (392). According to one embodiment, the processor (322) can control the operation of at least one third microphone (352) of the second wearable electronic device (212) to lower or turn off the reception sensitivity of at least one third microphone (352) while receiving a voice signal corresponding to one's own speech from the electronic device (201).
[0088] According to one embodiment, the electronic device (201) (e.g., the electronic device (201) of FIG. 2) may include at least one first microphone (350), at least one processor (320), memory (330), and / or a communication circuit (390). According to one embodiment, the electronic device (201) may further include at least one first speaker (355), a display (360), and / or at least one sensor (376). The electronic device (201) in FIG. 3 may be the electronic device (101) of FIG. 1. Additionally, the electronic device (201) in FIG. 3 may have the same configuration as the electronic device (101) of FIG. 1. Here, not all components shown in FIG. 3 are essential components of the electronic device (201), and the electronic device (201) may be implemented by more or fewer components than those shown in FIG. 3. In describing the electronic device (201) of FIG. 3, detailed descriptions of configurations similar to the embodiment of FIG. 1 or easily understood through the embodiment of FIG. 1 may be omitted.
[0089] According to one embodiment, the memory (330) (e.g., the memory (130) of FIG. 1) may store a control program for controlling the electronic device (201), a UI related to an application provided by the manufacturer or downloaded from an external source, images for providing the UI, user information, documents, databases, or related data.
[0090] According to one embodiment, the memory (330) can store instructions that control the processor (320) to perform various operations during execution.
[0091] According to one embodiment, the communication circuit (390) can communicate with at least one external device (e.g., a first wearable electronic device (e.g., earphones) (211), a second wearable electronic device (e.g., a watch phone) (212)) based on a short-range wireless communication method such as Bluetooth or NFC (near field communication) under the control of the processor (320).
[0092] According to one embodiment, when an interpretation application for real-time simultaneous interpretation is executed under the control of the processor (320), the display (360) may display a UI related to the application being executed. The display (360) may output (or display) the interpretation result (or interpretation content) for the speaker's utterance as text so that the user or the other party can check it. For example, the interpretation content based on one's own utterance output through the display (360) may be provided primarily for the other party. On the other hand, the interpretation content based on the other party's utterance output through the display (360) may be provided for the user.
[0093] According to one embodiment, at least one first speaker (355) can receive an electrical signal from a processor (320) to generate sound and output it externally. At least one first speaker (355) can output interpretation information translated from the speaker's voice signal to the outside of the electronic device (201). For example, interpretation information in text form translated from the speaker's voice signal can be converted into voice, and the converted voice can be output through at least one first speaker (355).
[0094] According to one embodiment, the processor (320) can recognize an utterance (or voice input) received through at least one first microphone (350) and convert it into text based on the recognized voice input. The voice recognition processing for the utterance according to one embodiment may partially include automatic speech recognition (ASR) and / or natural language understanding (NLU) processing. For example, the utterance voice can be converted into text using an ASR module (or ASR engine), and an NLU module (or NLU engine) can extract the meaning of the utterance from the recognition result of the ASR module.
[0095] According to one embodiment, at least one sensor (376) may detect whether the electronic device (201) is in a first state and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, at least one sensor (376) may include an illuminance sensor, a proximity sensor, and / or a grip sensor. The first state of the electronic device (201) indicates a state in which the electronic device (201) is not exposed to the outside, for example, a state in which the electronic device (201) is placed in a place such as a bag or pocket without the user holding the electronic device (201) in their hand. Additionally, the processor (320) may use at least one sensor (376) to recognize a state in which the user is not holding the electronic device (201) even when the display (360) is not activated (or turned off) (e.g., an always-on-display (AOD) state, a screen lock state). For example, the first state of the electronic device (201) can be determined based on sensor information detected (or collected) through an illuminance sensor included in at least one sensor (376), or sensor information detected through a proximity sensor or a grip sensor. For example, if the sensing value by the illuminance sensor is less than a specified value, the processor (320) may recognize that the surroundings of the electronic device (201) are very dark, and if the sensing value by the grip sensor is less than a specific value, the processor may recognize that the user is not holding the electronic device (201) in their hand. Additionally, if the sensing value by the illuminance sensor is less than a specified value and the sensing value by the grip sensor is less than a specific value, the processor (320) may determine that the first state of the electronic device (201) is that the surroundings are dark and the user is not holding it in their hand.
[0096] According to one embodiment, the processor (320) may turn on (or activate) at least one second microphone (351) of the first wearable electronic device (211) to receive a voice signal when the interpretation application is executed. According to one embodiment, the processor (320) may receive a request for the execution of the interpretation application by a user through the electronic device (201). For example, the processor (320) may execute the interpretation application in response to a user selection of an object representing the interpretation application among objects representing applications displayed on the display (360).
[0097] According to one embodiment, the processor (320) may execute an interpretation application in response to receiving a request to execute an interpretation application from the first wearable electronic device (211) via a communication circuit (390). For example, in response to a designated touch input from the first wearable electronic device (211) (e.g., when a touch input is maintained for N seconds or more while one of the pair of first wearable electronic devices (211) is removed), at least one second microphone (351) may be activated in the first wearable electronic device (211), and a request to execute an interpretation application may be transmitted to the electronic device (201) via the communication circuit (391) to execute an interpretation application in the electronic device (201).
[0098] According to one embodiment, the processor (320) may receive a first voice signal corresponding to one's own speech obtained through at least one second microphone (351) of the first wearable electronic device (211) while the interpretation application is running. For example, even if a voice signal is received through at least one second microphone (351) of the first wearable electronic device (211), if it is determined that the received voice signal is not a voice signal corresponding to one's own speech, the first wearable electronic device (211) may not transmit the received voice signal to the electronic device (201). Therefore, if the processor (320) receives a second voice signal through at least one first microphone (350) while the first voice signal is not received from the first wearable electronic device (211), it may recognize that the received second voice signal is a voice signal corresponding to the other party's speech.
[0099] According to one embodiment, the processor (320) can convert the first voice signal into text using an interpretation application in response to the reception of a first voice signal corresponding to a self-utterance from the first wearable electronic device (211). For example, the processor (320) can use a voice recognition engine (or voice recognition algorithm) to recognize the voice corresponding to the utterance and generate a recognition result. For example, the processor (320) can detect the actual voice segment included in the input voice by detecting the start and end points from the voice signal using an STT function. Here, the actual voice segment may correspond to the utterance sentence.
[0100] The processor (320) can convert the first voice signal into text and then check whether the language of the converted text is the user's language. In response to confirming that the language of the first voice signal is the user's language, the processor (320) can translate the text converted in relation to the first voice signal into the other party's language (or a pre-set language) using an interpretation application, generate translated interpretation information, and output the generated interpretation information.
[0101] According to one embodiment, the processor (320) can convert the second voice signal into text using the STT function of an interpretation application in response to receiving a second voice signal corresponding to a speech from at least one first microphone (350). After converting the second voice signal into text, the processor (320) can check whether the language of the converted text is the language of the other party. In response to confirming that the language of the second voice signal is the language of the other party, the processor (320) can translate the text converted in relation to the second voice signal into a user language (or a preset language) using an interpretation application, generate (or obtain) translated interpretation information, and output the generated interpretation information.
[0102] According to one embodiment, since the generated interpretation information has a text form, the processor (320) can output (or display) the interpretation information in text form through the display (360). Additionally, the processor (320) can convert the interpretation information corresponding to the utterance into speech using a text-to-speech (TTS) function and output it.
[0103] According to one embodiment, the processor (320) can display the interpretation information (or interpretation content, interpretation result) generated in response to one's own speech through a display (360) in text form or output (or play) it through at least one first speaker (355) in voice form so that the information can be provided to the other party.
[0104] According to one embodiment, the processor (320) can transmit the interpretation information generated in response to the other party's speech to the first wearable electronic device (201) through the communication circuit (390) so that the interpretation information generated in response to the other party's speech can be provided to the user. Accordingly, the interpretation information generated in response to the other party's speech can be output in the form of voice through at least one second speaker (356) of the first wearable electronic device (201). According to one embodiment, the interpretation information generated in response to the other party's speech may also be displayed in the form of text through the display (360).
[0105] According to one embodiment, the processor (320) can determine whether the electronic device (201) is in a first state in which it is not exposed to the outside based on sensor information detected through at least one sensor (376). The first state of the electronic device (201) indicates a state in which the electronic device (201) is not exposed to the outside, for example, a state in which a user places the electronic device (201) in a place such as a bag or pocket without holding it in their hand.
[0106] According to one embodiment, the processor (320) can perform operations to minimize interference and mistranslation caused by ambient noise other than voice signals in response to confirming that the state of the electronic device (201) is a first state. When the state of the electronic device (201) is a first state, the sensitivity of the first microphone (350) of the electronic device (201) may be lower than the sensitivity of the second microphone (351) of the first wearable electronic device (211), so the first microphone (350) may be turned off or processing of the voice signal received through the first microphone (350) may not be performed while receiving the first voice signal from the first wearable electronic device (211). For example, if the processor (320) determines that the electronic device (201) is stored in a place such as a bag or pocket, it may turn off (or disable) the first microphone (350) so that during self-speaking, voice signals are received only through the second microphone (351) of the first wearable electronic device (211) instead of the first microphone (350) of the electronic device (201).
[0107] On the other hand, if the processor (320) determines that the state of the electronic device (201) is not the first state, that is, if the electronic device (201) is exposed to the outside, it may output interpretation information for the first voice signal corresponding to the user's own speech through the first speaker (355) or display (360). Additionally, if the processor (320) determines that the state of the electronic device (201) is not the first state, it may transmit interpretation information for the second voice signal corresponding to the other party's speech to the first wearable electronic device (211), but may also output interpretation information related to the other party's speech through the first speaker (355) or display (360).
[0108] According to one embodiment, the processor (320) can receive a voice signal corresponding to the other party's speech through the first microphone (350), but when connected to a second wearable electronic device (212) for communication, it can also receive a voice signal corresponding to the other party's speech through at least one third microphone (352) of the second wearable electronic device (212) when the other party speaks. For example, when the second wearable electronic device (212) is worn on the user's wrist, the processor (320) can receive a voice signal corresponding to the other party's speech received through at least one third microphone (352) of the second wearable electronic device (212) via the communication circuit (390).
[0109] According to one embodiment, the processor (320) may recognize (or determine) that the voice signal from the second wearable electronic device (212) is a voice signal corresponding to the other party's speech when no voice signal is received from the first wearable electronic device (211) while receiving a voice signal from the second wearable electronic device (212). While receiving a voice signal from the second wearable electronic device (212), the processor (320) may increase the reception sensitivity (or reception level) while maintaining the ON state of the third microphone (352) of the second wearable electronic device (212), and may turn off another microphone, such as the first microphone (350), or lower the reception sensitivity. By doing so, the processor (320) can obtain a high-quality voice signal corresponding to the other party's speech, thereby minimizing mistranslation and providing an improved interpretation result. The processor (320) performs translation of a voice signal corresponding to a counterpart's speech received from the second wearable electronic device (212) and can output interpretation information related to the counterpart's speech through the first speaker (355) or display (360).
[0110] According to one embodiment, the processor (320) can confirm that the state of the electronic device (201) is not exposed to the outside while the second wearable electronic device (212) is worn on the user's wrist in a first state. For example, the processor (320) can transmit the interpretation information to the second wearable electronic device (212) so that the interpretation information, which translates the voice signal corresponding to the other party's speech received through the second wearable electronic device (212), is output through at least one third speaker (357) of the second wearable electronic device (212), because the second wearable electronic device (212) is worn on the user's wrist but the electronic device (201) is in a bag or pocket.
[0111] According to one embodiment, the processor (320) can confirm that the state of the electronic device (201) is not exposed to the outside when the second wearable electronic device (212) is not being worn on the user's wrist, in a first state. For example, when the electronic device (201) is in a bag or pocket when the second wearable electronic device (212) is not being worn on the user's wrist, the processor (320) may be restricted from outputting interpretation information translated from a voice signal according to one's own speech through the electronic device (201) or the second wearable electronic device (212). Therefore, the processor (320) can control the output of information that induces the user to directly speak interpretation information related to one's own speech to the other party in the form of voice through the second speaker (356) of the first wearable electronic device (211). For example, a voice that induces the user to say "Hello," which is a sentence translated from "Hello" into the other person's language, can be output through the second speaker (356).
[0112] According to one embodiment, when a user runs an interpretation application while wearing a first wearable electronic device (211) (e.g., earphones) connected to an electronic device (201), the interpretation content can be provided through the electronic device (201) or the first wearable electronic device (211) according to the state of the electronic device (201) while minimizing noise interference according to the user's speech or the other party's speech, thereby enabling real-time simultaneous interpretation.
[0113] According to one embodiment, the electronic device (101, 201) may include at least one first microphone (350), a communication circuit (390) configured to support a short-range communication method, at least one processor (320), and a memory (330) for storing instructions.
[0114] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to receive a first voice signal corresponding to a self-utterance from a first wearable electronic device (211) connected to the electronic device through the communication circuit while an interpretation application is executed on the electronic device, and the self-utterance is obtained through at least one second microphone (351) of the first wearable electronic device.
[0115] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to output first interpretation information for the first voice signal using the interpretation application in response to the reception of the first voice signal.
[0116] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to receive a second voice signal corresponding to a counterpart's speech through the at least one first microphone while the interpretation application is executed on the electronic device.
[0117] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to acquire second interpretation information for the second voice signal using the interpretation application in response to the reception of the second voice signal.
[0118] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to transmit second interpretation information for the second voice signal to the first wearable electronic device through the communication circuit.
[0119] According to one embodiment, the instructions may be configured to acquire second interpretation information for the second voice signal in response to receiving the second voice signal through the at least one first microphone while the first voice signal is not received from the first wearable electronic device, when the electronic device receives the second voice signal individually or collectively by the at least one processor.
[0120] According to one embodiment, at least one sensor (376) is further included, and when the instructions are executed individually or collectively by the at least one processor, the electronic device determines whether the state of the electronic device is a first state in which the electronic device is not exposed to the outside based on sensor information detected through the at least one sensor, and in response to determining that the state of the electronic device is the first state, the at least one first microphone may be set to turn off while receiving the first voice signal from the first wearable electronic device.
[0121] According to one embodiment, the apparatus further comprises at least one first speaker (335) and a display (360), and the instructions may be configured to output first interpretation information for the first voice signal through the at least one first speaker or the display in response to the electronic device confirming that the state of the electronic device is not the first state when executed individually or collectively by the at least one processor.
[0122] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to output second interpretation information for the second voice signal through the at least one first speaker or the display in response to confirming that the state of the electronic device is not the first state.
[0123] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to check whether a second wearable electronic device (212) connected to the electronic device through the communication circuit is being worn by a user, and in response to checking that the second wearable electronic device is being worn by the user, receive a third voice signal corresponding to the other party's speech obtained through at least one third microphone (352) of the second wearable electronic device through the communication circuit.
[0124] According to one embodiment, the instructions may be configured such that, when executed individually or collectively by the at least one processor, the electronic device turns off the at least one first microphone in response to confirming that the second wearable electronic device is connected to the electronic device and the state of the electronic device is the first state while the first voice signal is not received from the first wearable electronic device, receives a third voice signal corresponding to the other party's speech obtained through the at least one third microphone (352) of the second wearable electronic device via the communication circuit, obtains third interpretation information for the third voice signal using the interpretation application, and outputs the third interpretation information for the third voice signal through the at least one first speaker or the display.
[0125] According to one embodiment, the instructions may be configured to transmit to the second wearable electronic device, so as to output third interpretation information for the third voice signal through at least one third speaker (357) of the second wearable electronic device, in response to the electronic device confirming that the second wearable electronic device is being worn by the user and that the state of the electronic device is the first state when executed individually or collectively by the at least one processor.
[0126] According to one embodiment, the instructions may be configured to output information that induces the user to speak the first interpretation information through at least one second speaker (356) of the first wearable electronic device in response to the electronic device confirming that the second wearable electronic device is not being worn by the user and that the state of the electronic device is the first state when executed individually or collectively by the at least one processor.
[0127] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to analyze the second voice signal received through the at least one first microphone using an artificial intelligence (AI) model, confirm based on the analysis result that the second voice signal is a voice signal corresponding to the other party's utterance, and acquire second interpretation information for the second voice signal using the interpretation application in response to confirming that the second voice signal is a voice signal corresponding to the other party's utterance.
[0128] FIG. 4 is a flowchart of the operation of an electronic device for providing an interpretation function according to one embodiment. Referring to FIG. 4, the operation method may include operations 405 through 425. Each operation of the operation method of FIG. 4 may be performed by at least one processor (e.g., processor (120) of FIG. 1 and processor (320) of FIG. 3) of an electronic device (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2 and FIG. 3). In one embodiment, at least one of operations 405 through 425 may be omitted, the order of some operations may be changed, or other operations may be added.
[0129] In operation 405, the electronic device (201) can receive a first voice signal corresponding to a self-utterance from a first wearable electronic device (211) connected to the electronic device (201) while an interpretation application is running on the electronic device (201). Here, the self-utterance may be obtained through at least one second microphone (351) of the first wearable electronic device (211).
[0130] In operation 410, the electronic device (201) can output first interpretation information for the first voice signal using the interpretation application in response to the reception of the first voice signal.
[0131] In operation 415, the electronic device (201) can receive a second voice signal corresponding to the other party's speech through at least one first microphone (350) of the electronic device (201) while the interpretation application is running on the electronic device (201).
[0132] According to one embodiment, the operation of receiving a first voice signal (405), the operation of outputting first interpretation information for the first voice signal (410), and the operation of receiving a second voice signal (415) are shown as being sequential, but it should be understood that the order of the 405, 410, and 415 operations may be changed or performed simultaneously. For example, since the other party may speak before the user speaks, the second voice signal corresponding to the other party's speech may be received through the first microphone (350) in the 415 operation, and the first voice signal corresponding to the user's own speech may be received through the second microphone (351) in the 405 operation, so the order of operations is not limited thereto.
[0133] In operation 420, the electronic device (201) can obtain second interpretation information for the second voice signal using the interpretation application in response to the reception of the second voice signal. According to one embodiment, the electronic device (201) can obtain second interpretation information for the second voice signal in response to receiving the second voice signal through at least one first microphone (350) while the first voice signal is not received from the first wearable electronic device (211). For example, the electronic device (201) may recognize that the voice signal received through the first microphone (350) is a voice signal corresponding to the other party's speech when the voice signal is not received from the first wearable electronic device (211) while the voice signal is received through at least one first microphone (350) of the electronic device (201).
[0134] In operation 425, the electronic device (201) can transmit second interpretation information for the second voice signal to the first wearable electronic device (211). According to one embodiment, the electronic device (201) can control the output of the second interpretation information related to the other party's speech through at least one second speaker (356) of the first wearable electronic device (211) so that the user can hear it.
[0135] According to one embodiment, the electronic device (201) can determine whether the state of the electronic device (201) is a first state in which the electronic device (201) is not exposed to the outside, based on sensor information detected through at least one sensor (376). In response to confirming that the state of the electronic device (201) is the first state, the electronic device (201) can control the at least one first microphone (350) to turn off while receiving the first voice signal from the first wearable electronic device (211). According to one embodiment, the electronic device (201) may lower the reception sensitivity of the at least one first microphone (350) while receiving the first voice signal.
[0136] According to one embodiment, the electronic device (201) may output first interpretation information for the first voice signal through at least one first speaker (355) of the electronic device (201) or a display (360) of the electronic device (201) in response to confirming that the state of the electronic device (201) is not the first state. For example, the electronic device (201) may display the interpretation information, which translates a sentence spoken by a user, in text form through the display (360) of the electronic device (201) or output it in voice form through the first speaker (355) so that the other party can see or hear it.
[0137] According to one embodiment, the electronic device (201) may output second interpretation information for the second voice signal through at least one first speaker (355) of the electronic device (201) or a display (360) of the electronic device (201) in response to confirming that the state of the electronic device (201) is not the first state. For example, the electronic device (201) may output interpretation information related to the other party's speech through the first wearable electronic device (211) so that the user can also hear the sentence spoken by the other party through the first wearable electronic device (211), but may also output it through at least one first speaker (355) or a display (360) of the electronic device (201).
[0138] According to one embodiment, the electronic device (201) can determine whether the second wearable electronic device (212) connected to the electronic device (201) is being worn by the user. In response to determining that the second wearable electronic device (212) is being worn by the user, the electronic device (201) can receive a third voice signal corresponding to the other party's speech obtained through at least one third microphone (352) of the second wearable electronic device (212).
[0139] According to one embodiment, the electronic device (201) may turn off the at least one first microphone (350) in response to confirming that the second wearable electronic device (212) is connected to the electronic device (201) and that the state of the electronic device (201) is the first state while the first voice signal is not received from the first wearable electronic device (211). For example, the electronic device (201) may turn off the first microphone (350) when it is difficult to obtain a voice signal through the first microphone (350) in the first state where the electronic device (201) is not being held by a user or is not exposed to the outside, while confirming that the second wearable electronic device (212) is communication-connected to the electronic device (201). Additionally, the electronic device (201) can turn off the first microphone (350) of the electronic device (201) based on the fact that the interpretation application is also running on the second wearable electronic device (212) while confirming that the second wearable electronic device (212) is connected to the electronic device (201).
[0140] The electronic device (201) can receive a third voice signal corresponding to the other party's speech obtained through at least one third microphone of the second wearable electronic device (212) while the first microphone (350) is turned off. The electronic device (201) can obtain third interpretation information for the third voice signal using the interpretation application and output the third interpretation information for the third voice signal through at least one first speaker (355) of the electronic device (201) or the display (360) of the electronic device (201).
[0141] According to one embodiment, the electronic device (201) can transmit to the second wearable electronic device (212) third interpretation information for the third voice signal to be output through at least one third speaker (357) of the second wearable electronic device (212) in response to confirming that the second wearable electronic device (212) is being worn by the user and that the state of the electronic device (201) is the first state.
[0142] According to one embodiment, the electronic device (201) can output information that induces the user to speak the first interpretation information through at least one second speaker (356) of the first wearable electronic device (211) in response to confirming that the second wearable electronic device (212) is not being worn by the user and that the state of the electronic device (201) is the first state.
[0143] FIG. 5 is a diagram showing signals transmitted and received to provide an interpretation function between an electronic device according to one embodiment, a first wearable electronic device, and a second wearable electronic device.
[0144] Referring to FIG. 5, the electronic device (201) may be in a state of being connected to a communication network (505) with a first wearable electronic device (211) and a second wearable electronic device (212). The electronic device (201) may execute an interpretation application in response to a user request in operation 510. While the interpretation application is executed, the first wearable electronic device (211) may check in operation 515 whether a first voice signal is received through the second microphone (351). When the first voice signal is received through the second microphone (351), the electronic device (201) may check whether the received first voice signal corresponds to the user's speech by using a bone conduction sensor (377) in operation 520. The electronic device (201) may check whether the user's speech is being made by checking whether the pattern of the vibration signal detected through the bone conduction sensor (377) corresponds to the pattern of the vibration signal corresponding to the user's speech.
[0145] For example, since a vibration signal can be generated by chewing motions rather than user speech, the first wearable electronic device (211) can check whether the pattern of the vibration signal based on deep learning corresponds to a vibration pattern representing the user's speech. If the pattern of the vibration signal corresponds to a vibration pattern representing the user's speech, the first wearable electronic device (211) can check that the voice signal received through the second microphone (351) in operation 525 is a voice signal corresponding to the user's speech. According to one embodiment, the electronic device (201) may also check whether the received voice signal is a voice signal corresponding to the user's speech by using an AI model when the vibration pattern is not used. In operation 530, the first wearable electronic device (211) can transmit a first voice signal corresponding to the user's speech to the electronic device (201).
[0146] According to one embodiment, the electronic device (201) can check whether it is in a first state that is not exposed to the outside during operation 535. In FIG. 5, the electronic device (201) is shown to perform an operation to check the state of the electronic device (201) in response to the reception of a first voice signal corresponding to its own speech, but is not limited thereto. For example, the electronic device (201) may perform an operation to check the state of the electronic device (201) in response to the execution of an interpretation application.
[0147] In operation 535, the electronic device (201) may output first interpretation information for the first voice signal in operation 540 when the state of the electronic device (201) is not in a first state where the state is not exposed to the outside. For example, the electronic device (201) may operate by setting (or changing) the interpretation language from the user's language to the other party's language for voice signals classified as the user's own speech. For example, the electronic device (201) may output (or display) the first interpretation information, which translates the first voice signal corresponding to the user's own speech into the other party's language, through a display (360) or through a first speaker (355) so that the other party can see or hear it.
[0148] According to one embodiment, the electronic device (201) can check whether a second voice signal corresponding to the other party's speech is received through the first microphone (350) in operation 545. In response to confirming that a second voice signal corresponding to the other party's speech is received through the first microphone (350), the electronic device (201) can obtain (or generate) second interpretation information for the second voice signal using an interpretation application in operation 550. The second interpretation information may be text obtained by converting the second voice signal into text and then translating the converted text into a user language (or a pre-set language). For example, the electronic device (201) may operate by setting (or changing) the interpretation language from the other party's language to the user language for voice signals classified as the other party's speech.
[0149] The electronic device (201) can transmit the second interpretation information obtained from the 560 operation to the first wearable electronic device (211) so that the user can listen to it in the user's language. The first wearable electronic device (211) can output the second interpretation information received from the 565 operation in the form of voice through the second speaker (356). For example, when the interpretation of a spoken sentence according to the other party's speech is completed, the content interpreted in the user's language can be output through the second speaker (356) of the first wearable electronic device (211). While the first wearable electronic device (211) is outputting the content interpreted in the user's language, it controls the light output in various ways, such as blinking or light intensity control, using a light-emitting element (e.g., an LED element) included in the first wearable electronic device (211), so that the other party can recognize that the interpreted content is being output.
[0150] According to one embodiment, when the electronic device (201) is in a state of communication connection with the second wearable electronic device (212), it can be controlled to output interpretation information according to user speech or receive voice signals according to the other party's speech through the second wearable electronic device (212).
[0151] According to one embodiment, in operation 570, the second wearable electronic device (212) can be checked whether it is being worn on the user's body (e.g., wrist). According to one embodiment, in operation 535, in response to confirming that the electronic device (201) is in a first state where it is not exposed to the outside, the electronic device (201) can check the wearing state of the second wearable electronic device (212).
[0152] When the second wearable electronic device (212) is being worn, the second wearable electronic device (212) can output first interpretation information for a first voice signal corresponding to a user's speech in operation 575. For example, when the electronic device (201) is being stored in a bag or pocket, control can be made so that the first interpretation information related to the user's speech is output in voice form through the third speaker (37) of the second wearable electronic device (212) instead of the electronic device (201).
[0153] According to one embodiment, in operation 580, the second wearable electronic device (212) can check whether a third voice signal corresponding to the other party's speech is received through the third microphone (352). If a third voice signal corresponding to the other party's speech is received through the third microphone (352), the third voice signal corresponding to the other party's speech can be transmitted to the electronic device (201) in operation 585. The electronic device (201) can recognize that the third voice signal is a voice signal corresponding to the other party's speech if there is no voice signal received from the first wearable electronic device (211) while receiving the third voice signal. Accordingly, the electronic device (201) can translate the third voice signal into the user's language using an interpretation application in response to the reception of the third voice signal, and output third interpretation information for the third voice signal in operation 590.
[0154] FIG. 6 is a flowchart of the operation of a first wearable electronic device according to one embodiment. Referring to FIG. 6, the operation method may include operations 605 through 625. Each operation of the operation method of FIG. 6 may be performed by at least one processor (e.g., processor (321) of FIG. 3) of the first wearable electronic device (e.g., the first wearable electronic device (2211) of FIG. 2 and FIG. 3). In one embodiment, at least one of operations 605 through 625 may be omitted, the order of some operations may be changed, or other operations may be added.
[0155] In operation 605, the first wearable electronic device (211) can turn on (or activate) the second microphone (351) of the first wearable electronic device (211) in response to the execution of an interpretation application. In operation 610, the first wearable electronic device (211) can check whether a first voice signal is received through the second microphone (351). In response to the reception of the first voice signal, the first wearable electronic device (211) can check in operation 615 whether a vibration signal is detected through the bone conduction sensor. If a vibration signal is detected through the bone conduction sensor (377) in operation 615, the first wearable electronic device (211) can check in operation 620 whether the pattern of the vibration signal based on deep learning (or a deep learning model) corresponds to a vibration pattern representing self-speech. In operation 625, the first wearable electronic device (211) can transmit the first voice signal to the electronic device (201) to output first interpretation information for the first voice signal corresponding to the speaker's speech.
[0156] According to one embodiment, if a vibration signal through the bone conduction sensor (377) is not detected, the first voice signal is determined not to be related to the user's speech, and the operation of determining whether the user is speaking can be terminated. According to one embodiment, there may be cases where the first voice signal is received through the second microphone (351), but the vibration signal detected through the bone conduction sensor (377) is not a vibration signal corresponding to the user's speech.
[0157] According to one embodiment, when a vibration signal is detected through the bone conduction sensor (377) of the first wearable electronic device (211), the first wearable electronic device (211) can use a deep learning model to determine whether the detected vibration signal is a vibration signal due to actual speech or a vibration signal due to jaw joint movement. For example, even if a first voice signal is received through the second microphone (351), if it is not a vibration pattern due to the user's own speech, the first wearable electronic device (211) may consider that a voice signal due to the other party's speech has been input. If the first wearable electronic device (211) determines that the first voice signal does not correspond to the user's own speech, it may not transmit the first voice signal to the electronic device (201). For example, while the user is wearing the first wearable electronic device (211) on their ear, they can run an interpretation application through a designated touch input to the first wearable electronic device (211) without taking out the electronic device (201) from their pocket. When the interpretation application is executed, the second microphone (351) of the first wearable electronic device (211) is turned on first to detect the presence or absence of surrounding voice signals. When surrounding voice signals are received through the second microphone (351), the bone conduction sensor (377) of the first wearable electronic device (211) can be used to determine whether the voice signal is from one's own speech.
[0158] Although the above description explains a method using a bone conduction sensor to determine whether the speaker is speaking, various methods other than the method using a bone conduction sensor may be used. For example, there may be a method of determining whether the speaker is speaking based on the analysis results by analyzing the correlation between the signal indicating jaw movement detected by an accelerometer and the waveform change of the signal input through the second microphone (351) of the first wearable electronic device (211). Additionally, there may be a method of determining whether the speaker is speaking by analyzing heart rate characteristics during speaking using a heart rate sensor included in the first wearable electronic device (211) and detecting a specific heart rate pattern. Furthermore, there may be a method of determining whether the speaker is speaking based on the similarity between the two voice signals by comparing the user's own voice signal learned through machine learning with the voice signal received through the second microphone (351). Additionally, an AI (artificial intelligence) model may be used to determine whether the speaker is speaking. According to one embodiment, as a method for determining whether self-speech is being made, various methods described above, along with a method using a bone conduction sensor, may be used individually or in combination to suit the actual speech situation.
[0159] According to one embodiment, when the first wearable electronic device (211) determines that it is a self-speech, it can notify the electronic device (201) of the determination result regarding the self-speech by transmitting a voice signal corresponding to the self-speech to the electronic device (201). According to one embodiment, the first wearable electronic device (211) can transmit a signal indicating that it is a self-speech to the electronic device (201) before transmitting the voice signal. According to one embodiment, the first wearable electronic device (211) can transmit a signal indicating that it is a self-speech to the electronic device (201) at the same time as transmitting the voice signal corresponding to the self-speech. According to one embodiment, when the first wearable electronic device (211) receives a voice signal through the second microphone (351) but determines through the bone conduction sensor (377) that it is not a self-speech, it can transmit a signal indicating that it is a counterpart's speech to the electronic device (201). The electronic device (201) can control other devices, such as the first wearable electronic device (211), to turn off or lower the reception sensitivity of the second microphone (351) of the electronic device (201) located close to the other party in order to receive a voice signal corresponding to the other party's speech through the first microphone (350) of the electronic device (201) in response to a signal indicating that the other party is speaking, so as to obtain a high-quality voice signal corresponding to the other party's speech.
[0160] According to one embodiment, when it is determined that the speech is self-speech, only the second microphone (351) of the first wearable electronic device (211), which is closest to the user's mouth and can receive a high-quality voice signal, is turned on, and the remaining other devices, such as the first microphone (350) of the electronic device (201) and / or the third microphone (352) of the second wearable electronic device (212), may be turned off or the reception sensitivity (or level) may be lowered. According to one embodiment, the electronic device (201) processes only the voice signal corresponding to the self-speech received through the second microphone (351) of the first wearable electronic device (211), and may not process or remove voice signals received through other microphones, such as the first microphone (350) and the third microphone (352).
[0161] FIG. 7 is a diagram illustrating microphone optimization based on a counterpart's speech according to one embodiment. Referring to FIG. 7, the operation method may include operations 705 through 730. Each operation of the operation method of FIG. 7 may be performed by at least one processor (e.g., processor (120) of FIG. 1 and processor (320) of FIG. 3) of an electronic device (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2 and FIG. 3). In one embodiment, at least one of operations 705 through 730 may be omitted, the order of some operations may be changed, or other operations may be added.
[0162] FIG. 7 illustrates a method for controlling each microphone to more clearly detect the content of the other party's speech and minimize interference from ambient sound depending on the state of the electronic device (201) and other devices, such as the first wearable electronic device (211) and the second wearable electronic device (212).
[0163] Referring to FIG. 7, the electronic device (201) may turn on (or activate) the first microphone (350) of the electronic device (201) in response to the execution of an interpretation application in operation 705. For example, the electronic device (201) may turn on the first microphone (350) to check the state of the electronic device (201) and ambient sounds when selecting a microphone for the other party's speech. Here, the time at which the first microphone (350) is turned on may correspond to the start of execution of the interpretation application or the time at which it is confirmed that the voice signal received through the second microphone (351) from the first wearable electronic device (211) does not correspond to the user's own speech.
[0164] In operation 710, the electronic device (201) can determine whether the electronic device (201) is in a first state where it is not exposed to the outside. For example, the electronic device (201) can compare the reception sensitivity of a voice signal received from the first microphone (350) with the reception sensitivity of a voice signal received from the second microphone (351) of the first wearable electronic device (211). If the reception sensitivity of the first microphone (350) is lower than the reception sensitivity of the second microphone (351), it can be determined that the electronic device (201) is in a first state where it is not exposed to the outside. Additionally, for example, the electronic device (201) can determine whether the electronic device (201) is in a state where it is placed in a place such as a bag or a pocket based on sensor information detected using at least one sensor. As described above, the electronic device (201) can check (or determine, estimate) the state of the electronic device (201) based on the comparison result by comparing the reception sensitivity of each microphone (350, 351), but the state of the electronic device (201) can also be checked using sensor information.
[0165] In response to confirming that the electronic device (201) is not in a first state where it is not exposed to the outside, in operation 715, the electronic device (201) can be controlled to keep the first microphone (350) on and to turn off the second microphone (351) and the third microphone (352). For example, when the state of the electronic device (201) is exposed to the outside, the electronic device (201) can be controlled to turn on only the first microphone (350) so that it can receive a high-quality voice signal according to the other party's speech. Additionally, the electronic device (201) can process only the voice signal received through the first microphone (350) and ignore the voice signals received through the remaining microphones, such as the second microphone (351) and the third microphone (352), without processing them.
[0166] On the other hand, in response to the first state in which the electronic device (201) is not exposed to the outside, in operation 720, the electronic device (201) can check whether the user is wearing the second wearable electronic device (212). For example, when the state of the electronic device (201) is the first state, the electronic device (201) can check whether there is another microphone to receive a voice signal according to the other party's speech instead of the first microphone (350) of the electronic device (201). To do this, the electronic device (201) can check whether the user is wearing the second wearable electronic device (212) which is connected to the electronic device (201) through communication.
[0167] According to one embodiment, as described above, an operation to check whether the second wearable electronic device (212) is worn by a user is shown in a first state in which the electronic device (201) is not exposed to the outside, but the order of operations may not be limited thereto. For example, regardless of the state of the electronic device (201), by performing an operation to check whether the electronic device (201) is worn by a user is shown that the target microphone for outputting interpretation information is determined, and control of the determined microphone can be performed.
[0168] In response to confirming that the user is wearing the second wearable electronic device (212), in operation 725, the electronic device (201) can be controlled to turn off the first microphone (350) and the second microphone (351), and turn on the third microphone (352) of the second wearable electronic device (212). For example, when the user is wearing the second wearable electronic device (212) on their wrist, it is possible to position the second wearable electronic device (212) closer to the other person's mouth, so the third microphone (352) of the second wearable electronic device (212) can be turned on to receive a voice signal in response to the other person's speech. On the other hand, the first microphone (350) and the second microphone (351) can be temporarily turned off until the other party's spoken sentence is completed so that the reception of voice signals through other microphones, such as the first microphone (350) and the second microphone (351), is restricted in addition to the voice signal received through the third microphone (352) of the second wearable electronic device (212).
[0169] According to one embodiment, in response to confirming that the user is not wearing the second wearable electronic device (212), in operation 730, the electronic device (201) can be controlled to turn off the first microphone (350) and turn on the second microphone (351) and the third microphone (352). For example, if the user is not wearing the second wearable electronic device (212) on the wrist, it may be difficult to receive a voice signal through the third microphone (352) of the second wearable electronic device (212), so the third microphone (352) can be turned off. On the other hand, if the user is not wearing the second wearable electronic device (212) but the state of the electronic device (201) is the first state, it may be difficult to receive a voice signal through the first microphone (350) of the electronic device (201), so the first microphone (350) can also be turned off. Therefore, when the first microphone (350) and the third microphone (352) are off, the second microphone (351) can be turned on so that a voice signal can be received through the second microphone (351) of the first wearable electronic device (211).
[0170] According to one embodiment, when turning on the target microphone to receive the voice signal corresponding to the utterance and turning off the other microphones that are not the target microphone, a delay of a specified time may be applied between the operations of controlling the on / off of each microphone. For example, if a voice signal corresponding to the other party's utterance is received through the first microphone (350) while confirming that the voice signal received through the second microphone (351) corresponds to the user's utterance, a part of the user's utterance sentence and a part of the other party's utterance sentence may overlap. Therefore, in the overlay section, both the voice signals received through the microphone to be kept on (or turned on) and the microphone to be turned off can be processed to minimize the loss of the voice signal in the overlay section.
[0171] FIG. 8 is a diagram illustrating a method for outputting interpretation information according to one embodiment. To aid in understanding the explanation of FIG. 8, the explanation will be described with reference to FIG. 9. FIG. 9 is an example diagram illustrating a method for outputting interpretation information using a first wearable electronic device in a state where the electronic device according to one embodiment is not exposed to the outside.
[0172] Referring to FIG. 8, the electronic device (201) can check whether a second voice signal is received through the first microphone (350) while a first voice signal corresponding to self-utterance from the first wearable electronic device (211) is not received during operation 805.
[0173] In response to confirming that a second voice signal is received through the first microphone (350), in operation 810, the electronic device (201) can check whether the display (360) of the electronic device (201) is on. If the display (360) is on, the electronic device (201) can output second interpretation information for the second voice signal through the first speaker (355) or the display (360) in operation 815. For example, since the display (360) being on indicates that the screen is turned on because the user is using the electronic device (201), interpretation information can be output through the first speaker (355) or the display (360) of the electronic device (201) in use.
[0174] On the other hand, if the display (360) is not on, the electronic device (201) can determine that the user is not using the electronic device (201) and search for other available devices. Accordingly, in response to confirming in operation 810 that the display (360) of the electronic device (201) is not on, the electronic device (201) can check in operation 820 whether the display of the second wearable electronic device (212) is on.
[0175] In response to confirming that the display of the second wearable electronic device (212) is on, in operation 825, the electronic device (201) can be controlled to output second interpretation information for the second voice signal through the third speaker (352). For example, the electronic device (201) can transmit the second interpretation information to the second wearable electronic device (212) so that the second interpretation information is output in voice form through the third speaker (352) of the second wearable electronic device (212).
[0176] On the other hand, in response to confirming that the display of the second wearable electronic device (212) is not on, in operation 830, the electronic device (201) may transmit information inducing speech to the first wearable electronic device (211) so as to output information in the form of speech through the second speaker (251) that induces the user to speak second interpretation information for the second voice signal. For example, a sentence inducing the user to speak directly in the other person's language, such as a tutorial, can be provided for the user to speak the sentence 'Hello' that the user has spoken. Thus, a voice inducing the user to say 'Hello', which is the sentence interpreted from 'Hello' into the other person's language, can be output through the second speaker (356).
[0177] FIG. 9 illustrates a case in which a user (A) wears a first wearable electronic device (211) on their ear and a second wearable electronic device (212) on their wrist, and the electronic device (201) is placed in a bag or pocket and is not exposed to the outside, while simultaneously interpreting with another person (B) in real time.
[0178] Referring to FIG. 9, when an interpretation application (hereinafter, interpretation app) is executed (905), the first microphone (350) of the electronic device (201), the second microphone (351) of the first wearable electronic device (211), and the third microphone (352) of the second wearable electronic device (212) can be turned on.
[0179] For example, when the user is not using the electronic device (201), the screen resulting from the execution of the interpretation app may be displayed on the second wearable electronic device (212). Here, since the second wearable electronic device (212) is located close to the other party, the third microphone (352) may be used to receive a voice signal corresponding to the other party's speech. Additionally, since the first wearable electronic device (211) is positioned close to the user's mouth while worn on the user's ear, the second microphone (351) may be used to receive a voice signal corresponding to the user's speech.
[0180] According to one embodiment, the electronic device (201) can confirm that the electronic device (201) is in a first state in which it is located in a bag or pocket and is not exposed to the outside, based on the result of comparing the reception sensitivity of a voice signal received through a first microphone (350) with the reception sensitivity of a voice signal received through a third microphone (352) or based on at least one of sensor information detected through at least one sensor (376).
[0181] In response to detecting that the state of the electronic device (201) is the first state, the electronic device (201) can maintain the ON state of the third microphone (352). For example, the ON state of the third microphone (352) can be indicated by displaying a microphone-shaped object on the screen (915) of the second wearable electronic device (212).
[0182] On the other hand, when a voice signal corresponding to one's own speech is not received from the first wearable electronic device (211), if a voice signal is received through the third microphone (352), the voice signal received through the third microphone (352) can be recognized as corresponding to the other party's speech. Therefore, the electronic device (201) can be controlled so that the second microphone (351) of the first wearable electronic device (211) is turned off while receiving a voice signal through the third microphone (352), so that the electronic device (201) can receive a voice signal corresponding to the other party's speech only through the third microphone (352).
[0183] According to one embodiment, the electronic device (201) can perform the operation of converting a voice signal corresponding to a counterpart's speech received from the second wearable electronic device (212) into text using an STT function and then translating the counterpart's language into the user's language. For example, a screen (920) indicating that translation is in progress can be displayed on the second wearable electronic device (212). When the translation of the counterpart's speech sentence is completed, the electronic device (201) can generate interpretation information related to the counterpart's speech sentence. For example, the electronic device (201) can transmit the interpretation information to the second wearable electronic device (212) so that the generated interpretation information is displayed in text form on the screen (925) of the second wearable electronic device (212). Additionally, the electronic device (201) can play interpretation information (930), and the interpretation information can also be transmitted to the first wearable electronic device (211) so that the interpretation information is played in voice form through the second speaker (356) of the first wearable electronic device (211). When interpretation information translated into the user's language in response to the other party's speech is being played through the second speaker (356) of the first wearable electronic device (356), the first wearable electronic device (211) can output a visual effect (935) so that the other party can recognize that interpretation information is being played.
[0184] FIG. 10 is an example diagram illustrating a method of outputting interpretation information using a second wearable electronic device in a state where the electronic device according to one embodiment is not exposed to the outside.
[0185] Referring to FIG. 10, when the first wearable electronic device (211) detects the user's own speech (1005), the second microphone (351) of the first wearable electronic device (211) is kept on so that only the voice signal corresponding to the user's speech is received when the user's speech is detected, but the third microphone (352) of the second wearable electronic device (212) is controlled to be off. When the electronic device (201) receives the voice signal corresponding to the user's speech received from the first wearable electronic device (211), it may display a screen (1010) on the second wearable electronic device (212) indicating that translation is in progress. Subsequently, when the translation is completed, the interpretation information translated in response to the user's own speech may be displayed in text form on the screen (1015) of the second wearable electronic device (212). For example, since the state of the electronic device (201) can be detected as a first state, the electronic device (201) can be controlled so that interpretation information corresponding to the user's own speech is output through the third speaker (357) of the second wearable electronic device (212) rather than the electronic device (201). Accordingly, interpretation information (1020) in the form of voice translated into the other party's language in response to the user's speech can be output through the third speaker (357) of the second wearable electronic device (212).
[0186] FIG. 11 is an example diagram illustrating a method of outputting interpretation information using an electronic device while the electronic device is exposed to the outside according to one embodiment, and FIG. 12 is an example diagram illustrating a method of outputting interpretation information using a first wearable electronic device while the electronic device is exposed to the outside according to one embodiment.
[0187] FIGS. 11 and 12 illustrate a case in which a user (A) wears a first wearable electronic device (211) on their ear and a second wearable electronic device (212) on their wrist, and simultaneously interprets with another person (B) in real time while the electronic device (201) is exposed to the outside.
[0188] Referring to FIG. 11, when the screen of the electronic device (201) is turned on (1105) while the interpretation app is running, if a voice signal corresponding to the voice is received by the first wearable electronic device (211) and the second microphone (351) of the first wearable electronic device (211) is turned on, the electronic device (201) can control the microphones of other devices, such as the first microphone (350) and the third microphone (352), to be turned off.
[0189] In response to the reception of a voice signal corresponding to one's own speech from the first wearable electronic device (211), the electronic device (201) may display a screen (1115) indicating that translation is in progress on the display (360) and generate interpretation information that translates the spoken sentence in the user's language into the other person's language.
[0190] When the electronic device (201) detects that the screen of the electronic device (201) is turned on, it can control the output of interpretation information through the electronic device (201). Accordingly, interpretation information in text form may be displayed through the display (360) of the electronic device (201), or interpretation information in voice form may be output (or played) through the first speaker (355). For example, by displaying a speaker-shaped object on the screen (1120) of the electronic device (201), it may be indicated that interpretation information is being output through the first speaker (355).
[0191] Referring to FIG. 12, when the electronic device (201) receives a voice signal through the first microphone (350) while no voice signal is received from the first wearable electronic device (211), it can recognize (or detect) that the received voice signal is due to the other party's speech. When the other party's speech is detected (1205), the electronic device (201) can detect that the user is using the electronic device (201) using at least one sensor (376), and can control the first microphone (350) to be turned on so that only the voice signal corresponding to the other party's speech can be received, and the microphones of other devices, such as the second microphone (351) and the third microphone (352), can be turned off. For example, the first microphone (350) can be indicated as being turned on by displaying a microphone-shaped object on the screen (1210) of the electronic device (201). The electronic device (201) can display a screen (1215) indicating that translation is in progress on the display (360) and can generate interpretation information that translates a spoken sentence of the other party's language into the user's language. The electronic device (201) can display the interpretation information (or interpretation content) translated into the user's language in text form through the screen (1220) of the display (360), and the interpretation content in voice form can be output (or played) (1225) through the second speaker (356) so that the user can hear the interpretation content by transmitting the interpretation information to the first wearable electronic device (211).
[0192] FIG. 13 is an example diagram illustrating a method for outputting information that induces a user to speak interpretation information using a first wearable electronic device according to one embodiment, and FIG. 14 is an example diagram illustrating the operation of an electronic device when the user speaks in the other person's language according to one embodiment.
[0193] FIGS. 13 and 14 illustrate a case in which a user (A) wears a first wearable electronic device (211) on their ear and a second wearable electronic device (212) on their wrist, and the electronic device (201) is exposed to the outside, but the screen of the second wearable electronic device (212) and the screen of the electronic device (201) are both turned off, and simultaneous interpretation is performed in real time with another person (B).
[0194] Referring to FIG. 13, when the screen of the second wearable electronic device (212) and the screen of the electronic device (201) are both turned off (1305) and a self-speaking signal is detected (1310) by the first wearable electronic device (211), the electronic device (201) can control the second microphone (351) of the first wearable electronic device (211) to remain on, and since the second wearable electronic device (212) and the electronic device (201) are not in use, the microphones of other devices, such as the first microphone (350) and the third microphone (352), can be turned off.
[0195] When the electronic device (201) receives a voice signal corresponding to its own speech from the first wearable electronic device (211), it can convert it into text using an STT function and then perform translation on the converted text. The electronic device (201) can select a device to output the interpretation information generated through translation. For example, since the second wearable electronic device (212) and the electronic device (201) are not in use, the electronic device (201) can transmit the generated interpretation information to the first wearable electronic device (211). The first wearable electronic device (211) plays back (1315) the interpretation information translated into the other party's language in response to the user's speech, and the interpretation information must be shown or heard by the other party. For example, the electronic device (201) can control the user to directly speak the sentence 'Hello' that the user said 'Hello' into the other person's language by using TTS to convert the sentence 'Hello' into a voice form and then output it through the second speaker (356).
[0196] According to one embodiment, with reference to FIG. 14, when a user (A) speaks, the first wearable electronic device (211) can detect the user's own speech (1405). However, there may be cases where the user speaks in the other person's language rather than the user's language. Therefore, even if the electronic device (201) receives a voice signal corresponding to the user's speech from the first wearable electronic device (211), if the language of the converted text is determined to be the other person's language when the voice signal is converted into text, the electronic device (201) may not interpret the user's speech.
[0197] FIG. 15a is an example diagram illustrating a method of sharing interpretation information with a counterparty device according to one embodiment, and FIG. 15b is a diagram following FIG. 15a.
[0198] FIGS. 15a and 15b illustrate a case in which a first wearable electronic device (211a) (e.g., the first wearable electronic device (211) of FIGS. 2 and 3) is connected to a user's electronic device (201a) (e.g., the electronic device (201) of FIGS. 2 and 3) via a short-range communication method (1505), and the first wearable electronic device (211b) is connected to the other party's electronic device (201b) via communication.
[0199] Referring to FIG. 15a and FIG. 15b, according to one embodiment, when the other party possesses an electronic device (201b) to which the first wearable electronic device (211b) is connected, the user can use the electronic device (201a) to share the input screen of an interpretation app running on the user's electronic device (201a) by using the other party's electronic device (201b) and the first wearable electronic device (211b) without pairing the other party's first wearable electronic device (211b).
[0200] Referring to FIG. 15a, in response to a user selecting the ‘listen together to interpretation’ item on the electronic device (201a), the electronic device (201a) can locate the other electronic device (201b). When the other electronic device (201b) is detected within a short-range wireless communication range, information indicating that the interpretation content can be listened to together through the other party’s first wearable electronic device (211b) and an object representing the detected device, such as the first wearable electronic device (211b), can be displayed. In response to the selection (1515) of the object, the electronic device (201a) can transmit an interpretation sharing request to the other electronic device (201b).
[0201] If the other party chooses to allow the interpretation sharing request (1520), the microphones of the first wearable electronic device (211a) and the first wearable electronic device (211b) can each be turned on (or activated) to receive voice signals according to user speech or other party speech. Accordingly, a screen (1530, 1535) indicating that translation is in progress can be displayed on the electronic device (201a) and the electronic device (201b).
[0202] According to one embodiment, when a voice signal from the other party is detected (1540) in the first wearable electronic device (211b), only the microphone of the first wearable electronic device (211b) may be kept on. On the other hand, while a voice signal corresponding to the voice signal from the other party is received, the electronic device (201a) may control the microphone of the user's first wearable electronic device (211a) to be temporarily turned off.
[0203] The electronic device (201a) can receive a voice signal corresponding to the other party's speech from the electronic device (201b) and can generate an interpretation content by translating the voice signal into the user's language. The electronic device (201a) can output a screen (1550) that displays the interpretation content through a display. Additionally, the electronic device (201a) can transmit the interpretation content to the first wearable electronic device (211a), thereby allowing the interpretation content to be played back in voice form (1555) through the speaker of the first wearable electronic device (211a).
[0204] According to one embodiment, when a user's speech is detected (1560) in the first wearable electronic device (211a), the microphone of the first wearable electronic device (211a) may be kept on. On the other hand, while a voice signal corresponding to the user's speech is being received, the electronic device (201a) may control the microphone of the electronic device (201a) to be temporarily turned off. The electronic device (201a) may generate an interpretation content that translates the voice signal corresponding to the user's speech into the other party's language. By transmitting the interpretation content to the electronic device (201b), the interpretation content translated into the other party's language may be displayed on the screen of the electronic device (201b), and may also be output in the form of voice (1570) through the speaker of the first wearable electronic device (211b).
[0205] For example, whether the other party speaks and the reception of the voice signal corresponding to the other party's speech can be performed through the other party's first wearable electronic device (211b). Additionally, the user's electronic device (201a) can generate interpretation content corresponding to the other party's speech and transmit the generated output content to the first wearable electronic device (211b) so that the generated interpretation content is output through the other party's first wearable electronic device (211b). According to one embodiment, during real-time simultaneous interpretation, the interpretation content translated into the user's language according to the other party's speech is displayed on the user's electronic device (201a), and the other party can minimize screen complexity by displaying the interpretation content translated into the other party's language according to the user's speech on the other party's electronic device (201b).
[0206] FIG. 16 is an example diagram illustrating the output direction of interpretation information corresponding to the position of the speaker according to one embodiment.
[0207] According to one embodiment, the electronic device (201) may provide the following functions while providing multi-party real-time interpretation. For example, when multiple counterparts (B, C, D) are located in different directions relative to user (A), the electronic device (201) may check the user's head movement based on sensor information of a gyroscope sensor using head tracking technology included in the first wearable electronic device (211). The electronic device (201) may adjust the direction of sound output when outputting interpretation content through the first wearable electronic device (211) according to the head movement of the user wearing the first wearable electronic device (211) on their ear. According to one embodiment, by automatically adjusting the direction of sound output according to head movement, the user may feel a sense of direction as if the sound were actually occurring in that space.
[0208] According to one embodiment, the electronic device (201) can detect the direction of a speaker within a space using a stereotype microphone. When the electronic device (201) outputs the interpretation content for the speaker through the first wearable electronic device (211), it can use head tracking technology to output the interpretation content as if it were spoken from the actual location (or direction) of the speaker. Accordingly, the user can easily distinguish the speaker during multi-party simultaneous interpretation, thereby enabling natural simultaneous interpretation.
[0209] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.
[0210] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, each of the phrases such as “A or B”, “at least one of A and B”, “at least one of A or B”, “A, B or C”, “at least one of A, B and C”, and “at least one of A, B, or C” may include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as “first,” “second,” or “first” or “second” may be used simply to distinguish a component from another component and do not limit the components in any other aspect (e.g., importance or order). Where any (e.g., first) component is referred to as “coupled” or “connected” to another (e.g., second) component, with or without the terms “functionally” or “communicationally,” it means that said component may be connected to said other component directly (e.g., wired), wirelessly, or through a third component.
[0211] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0212] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0213] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0214] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0215] According to one embodiment, in a non-transient storage medium storing computer-readable instructions, the instructions are configured to cause the electronic device to perform at least one operation when executed by at least one processor (320) of the electronic device (101, 201), wherein the at least one operation includes receiving a first voice signal corresponding to a self-utterance from a first wearable electronic device (211) connected to the electronic device while an interpretation application is executed on the electronic device, and the self-utterance may be obtained through at least one second microphone (351) of the first wearable electronic device.
[0216] According to one embodiment, the at least one operation may include an operation of outputting first interpretation information for the first voice signal using the interpretation application in response to the reception of the first voice signal.
[0217] According to one embodiment, the at least one operation may include receiving a second voice signal corresponding to the other party's speech through at least one first microphone (350) of the electronic device while the interpretation application is running on the electronic device.
[0218] According to one embodiment, the at least one operation may include an operation of obtaining second interpretation information for the second voice signal using the interpretation application in response to the reception of the second voice signal.
[0219] According to one embodiment, the at least one operation may include transmitting second interpretation information for the second voice signal to the first wearable electronic device.
Claims
1. In an electronic device (101, 201), At least one first microphone (350); A communication circuit (390) configured to support a short-range communication method; At least one processor (320); and It includes a memory (330) for storing instructions, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, While an interpretation application is running on the electronic device, a first voice signal corresponding to a self-utterance is received from a first wearable electronic device (211) connected to the electronic device through the communication circuit, and the self-utterance is obtained through at least one second microphone (351) of the first wearable electronic device. In response to the reception of the first voice signal, first interpretation information for the first voice signal using the interpretation application is output, and While the interpretation application is running on the electronic device, a second voice signal corresponding to the other party's speech is received through the at least one first microphone, and In response to the reception of the second voice signal, second interpretation information for the second voice signal is obtained using the interpretation application, and An electronic device configured to transmit second interpretation information for the second voice signal to the first wearable electronic device through the communication circuit.
2. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device configured to acquire second interpretation information for the second voice signal in response to receiving the second voice signal through the at least one first microphone while the first voice signal is not received from the first wearable electronic device.
3. In paragraph 1 or 2, further comprising at least one sensor (376), When the above instructions are executed individually or collectively by the at least one processor, the electronic device, Based on sensor information detected through the above at least one sensor, determining whether the state of the electronic device is a first state in which the electronic device is not exposed to the outside, and An electronic device configured to turn off at least one first microphone while receiving the first voice signal from the first wearable electronic device in response to confirming that the state of the electronic device is the first state.
4. In any one of paragraphs 1 to 3, further comprising at least one first speaker (335) and a display (360), When the above instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device configured to output first interpretation information for the first voice signal through the at least one first speaker or the display in response to confirming that the state of the electronic device is not the first state.
5. In any one of claims 1 to 4, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device configured to output second interpretation information for the second voice signal through the at least one first speaker or the display in response to confirming that the state of the electronic device is not the first state.
6. In any one of claims 1 to 5, when the instructions are executed individually or collectively by the at least one processor, the electronic device, Checking whether the second wearable electronic device (212) connected to the electronic device through the communication circuit is being worn by the user, An electronic device configured to receive, through the communication circuit, a third voice signal corresponding to the other party's speech obtained through at least one third microphone (352) of the second wearable electronic device in response to confirming that the second wearable electronic device is being worn by the user.
7. In any one of claims 1 to 6, when the instructions are executed individually or collectively by the at least one processor, the electronic device, In response to confirming that the second wearable electronic device is connected to the electronic device and the state of the electronic device is the first state while the first voice signal is not received from the first wearable electronic device, the at least one first microphone is turned off, A third voice signal corresponding to the other party's speech obtained through at least one third microphone (352) of the second wearable electronic device is received through the communication circuit, and Obtaining third interpretation information for the third voice signal using the above interpretation application, and An electronic device configured to output third interpretation information for the above third voice signal through at least one first speaker or the display.
8. In any one of claims 1 to 7, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device configured to transmit to the second wearable electronic device, in response to confirming that the second wearable electronic device is being worn by the user and that the state of the electronic device is the first state, third interpretation information for the third voice signal is output through at least one third speaker (357) of the second wearable electronic device.
9. In any one of claims 1 through 8, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device configured to output information that induces the user to speak the first interpretation information through at least one second speaker (356) of the first wearable electronic device in response to confirming that the second wearable electronic device is not being worn by the user and that the state of the electronic device is the first state.
10. In any one of claims 1 to 9, when the instructions are executed individually or collectively by the at least one processor, the electronic device, The second voice signal received through the above at least one first microphone is analyzed using an AI (artificial intelligence) model, and Based on the above analysis results, it is confirmed that the second voice signal is a voice signal corresponding to the counterpart's speech, and An electronic device configured to acquire second interpretation information for the second voice signal using the interpretation application in response to confirming that the second voice signal is a voice signal corresponding to the other party's speech.
11. A method for providing an interpretation function in an electronic device, While an interpretation application is running on the electronic device, the operation of receiving a first voice signal corresponding to a self-utterance from a first wearable electronic device connected to the electronic device; the self-utterance is acquired through at least one second microphone of the first wearable electronic device, and An operation of outputting first interpretation information for the first voice signal using the interpretation application in response to the reception of the first voice signal; While the interpretation application is running on the electronic device, the operation of receiving a second voice signal corresponding to the other party's speech through at least one first microphone of the electronic device; An operation of obtaining second interpretation information for the second voice signal using the interpretation application in response to the reception of the second voice signal; and A method for providing an interpretation function, comprising the operation of transmitting second interpretation information for the second voice signal to the first wearable electronic device.
12. In paragraph 11, the operation of obtaining second interpretation information for the second voice signal is, A method for providing an interpretation function, comprising the operation of obtaining second interpretation information for the second voice signal in response to receiving the second voice signal through the at least one first microphone while the first voice signal is not received from the first wearable electronic device.
13. In claim 11 or 12, an operation of determining whether the state of the electronic device is a first state in which the electronic device is not exposed to the outside, based on sensor information detected through at least one sensor of the electronic device; and A method for providing an interpretation function, further comprising the operation of turning off the at least one first microphone while receiving the first voice signal from the first wearable electronic device in response to confirming that the state of the electronic device is the first state.
14. In any one of paragraphs 11 to 13, the operation of outputting first interpretation information for the first voice signal is, A method for providing an interpretation function, comprising the operation of outputting first interpretation information for the first voice signal through at least one first speaker of the electronic device or a display of the electronic device in response to confirming that the state of the electronic device is not the first state.
15. In a non-transient storage medium storing computer-readable instructions, said instructions are configured to cause said electronic device (101, 201) to perform at least one operation when executed by at least one processor (120, 320), said at least one operation, said operation being, While an interpretation application is running on the electronic device, the operation of receiving a first voice signal corresponding to a self-utterance from a first wearable electronic device connected to the electronic device; the self-utterance is acquired through at least one second microphone of the first wearable electronic device, and An operation of outputting first interpretation information for the first voice signal using the interpretation application in response to the reception of the first voice signal; While the interpretation application is running on the electronic device, the operation of receiving a second voice signal corresponding to the other party's speech through at least one first microphone of the electronic device; An operation of obtaining second interpretation information for the second voice signal using the interpretation application in response to the reception of the second voice signal; and A storage medium comprising the operation of transmitting second interpretation information for the second voice signal to the first wearable electronic device.