Electronic device and wearable electronic device for simultaneous interpretation
The wearable device addresses the resource-intensive training and cost issues of LLMs by using on-device machine learning for efficient, real-time translation with user-specific adaptation.
Patent Information
- Application Number
- PCT/KR2025/006572
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-26
- Filing Date
- 2025-05-15
- Publication Date
- 2026-01-02
AI Technical Summary
Large language models (LLMs) for natural language processing require significant time and hardware resources for training, and using them on edge devices or accessing servers for translation incur high costs and resource consumption.
A wearable electronic device equipped with a speaker, microphone, transceiver, and processor that uses machine learning models to translate user voice inputs and outputs translated information, with features for noise removal and user voice identification, and can update models based on external data.
Enables efficient, localized translation without high hardware demands or communication costs, providing real-time translation and user-specific language adaptation.
Smart Images

Figure KR2025006572_02012026_PF_FP_ABST
Abstract
Description
Electronic devices and wearable electronic devices for simultaneous interpretation
[0001] The present disclosure relates to an electronic device and a wearable electronic device for simultaneous interpretation.
[0002] Today, mobile electronic devices offer a variety of features for user convenience. In particular, with the recent emergence of large language models (LLMs) for natural language processing in the field of artificial intelligence, mobile electronic devices have been able to offer a variety of features utilizing LLMs. For example, mobile electronic devices recently introduced a translation function that uses LLMs to translate a user's voice during a voice call and transmit it to the caller.
[0003] However, while LLM demonstrates remarkable performance in natural language processing, training it requires time and hardware resources to process large amounts of training data. For example, a single training session can take at least several months. Furthermore, directly using LLM on edge devices requires high-spec hardware resources, or indirectly accessing a server containing LLM at the expense of communication costs.
[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0005] According to one embodiment of the present disclosure, a wearable electronic device comprises: a speaker; a microphone; a transceiver; a memory; and at least one processor including a processing circuit, wherein the memory may store instructions that, when individually or collectively executed by the at least one processor, cause the wearable electronic device to: receive a first user voice input through the microphone; generate first translation information in which the received first user voice input is translated using a machine learning model stored in the memory based on the received first user voice input, wherein the machine learning model includes a first machine learning model trained to output the translated translation information based on the voice input; transmit the first translation information to an external electronic device connected to the wearable electronic device through the transceiver; receive second translation information in which a second user voice input is translated from the external electronic device through the transceiver; and output the second translation information through the speaker.
[0006] In one embodiment, the machine learning model may include a second machine learning model trained to identify the user's voice based on other voice inputs received through the microphone.
[0007] According to one embodiment, the memory may store instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to: obtain user voice information based on a portion of the first user voice input using the second learning model, based at least in part on a determination that a portion of the first user voice input corresponds to a user voice; and translate the obtained user voice information using the first machine learning model to generate the first translation information.
[0008] According to one embodiment, the memory may store instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to: transmit user identification information corresponding to the identified user's voice and the first translation information to the external electronic device.
[0009] According to one embodiment, the memory may store instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to: convert the first user voice input into text using the machine learning model, and perform a translation on the converted text to generate the first translation information, at least as a part of generating the first translation information.
[0010] According to one embodiment, the memory may store instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to: receive update information from the external electronic device, and update the machine learning model based on the received update information.
[0011] According to one embodiment, the first translation information includes at least one of translated data, an index of the data, or a flag for the end of utterance, wherein the index of the data indicates the order in which the translated data is positioned in a sentence structure, and the flag may indicate whether the utterance has ended as identified by voice activity detection (VAD).
[0012] According to one embodiment, the memory may store instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to: detect noise input in a portion of the first user voice input where the user's voice does not exist, remove an input corresponding to the detected noise in the portion of the first user voice input where the user's voice does not exist to obtain user voice information, and translate the obtained user voice information using the first machine learning model to generate the first translation information.
[0013] According to one embodiment, a method of operating a wearable electronic device may be provided. The method may receive a first user voice input through a microphone of the wearable electronic device. The method may generate first translation information by translating the received first user voice input using a machine learning model stored in the memory based on the received first user voice input, wherein the machine learning model includes a first machine learning model trained to output translated translation information based on the voice input. The method may transmit the first translation information to an external electronic device connected to the wearable electronic device through the transceiver. The method may receive second translation information by translating a second user voice input from the external electronic device through the transceiver. The method may output the second translation information through the speaker.
[0014] According to one embodiment, a storage medium storing at least one computer-readable instruction may be provided. The at least one instruction, when executed by at least a portion of at least one processor of a wearable electronic device, may cause the wearable electronic device to perform at least one operation. The at least one operation may receive a first user voice input through a microphone of the wearable electronic device. The at least one operation may generate first translation information by translating the received first user voice input using a machine learning model stored in the memory based on the received first user voice input, wherein the machine learning model includes a first machine learning model trained to output translated translation information based on a voice input. The at least one operation may transmit the first translation information to an external electronic device connected to the wearable electronic device through the transceiver. The at least one operation may receive second translation information by translating a second user voice input from the external electronic device through the transceiver. At least one of the above actions may output the second translation information through the speaker.
[0015] According to another embodiment of the present disclosure, an electronic device includes a microphone; a display; a transceiver; a memory; and at least one processor including a processing circuit, wherein the memory may store instructions that, when individually or collectively executed by the at least one processor, cause the wearable electronic device to: receive a first user voice input through the microphone; generate first translation information in which the received first user voice input is translated using a machine learning model based on the received first user voice input, wherein the machine learning model includes a first machine learning model trained to output the translated translation information based on the voice input; transmit the first translation information to an external electronic device connected to the electronic device through the transceiver; receive second translation information in which a second user voice input is translated from the external electronic device through the transceiver; and display the second translation information in a first area of the display.
[0016] According to one embodiment, the memory may store instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to: display the first translation information in a second area of the display, and cause the first translation information displayed in the second area to be displayed in an opposite direction to the second translation information displayed in the first area.
[0017] According to one embodiment, the memory may store instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to: display the second translation information including user identification information, and a graphical object corresponding to the user identification information, together with the first translation information, in a second area of the display.
[0018] In one embodiment, the machine learning model may include a second machine learning model trained to identify a user's voice based on other voice inputs received via the microphone.
[0019] According to one embodiment, the memory may store instructions that, when individually or wholly executed by the at least one processor, cause the electronic device to: obtain user voice information by filtering out a portion corresponding to the other user's voice based on a portion of the first user's voice input using the second machine learning model, based on a determination that at least part of the first user's voice input corresponds to another user's voice, and translate the obtained user voice information using the first machine learning model to generate the first translation information.
[0020] According to one embodiment, the memory may store instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to: receive the second translation information, convert the received second translation information into speech information (Text to Speech) using the first machine learning model, and output the converted text through the speaker.
[0021] According to one embodiment, the memory may store instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to: convert the second translation information into voice information using at least one user voice information previously stored in the electronic device and then output the converted second translation information through the speaker.
[0022] According to one embodiment, the electronic device further comprises at least one sensor including a global positioning system (GPS), and the memory may store instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to: generate update information including a portion of the first machine learning model based on information input from the at least one sensor; and control the transceiver to transmit the update information to the external electronic device.
[0023] In one embodiment, the update information may include update information associated with a translation support language.
[0024] According to one embodiment, the electronic device further comprises at least one camera, and the memory stores instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to: acquire a mouth shape based on a lip image captured through the at least one camera, and identify a user's voice using the machine learning model based on the mouth shape, wherein the machine learning model may include a third learning model learned to identify the content of speech based on the mouth shape.
[0025] According to one embodiment, the second translation information includes at least one of translated data, an index of the data, or a flag for the end of utterance, and commands for causing the translated data included in the second translation information to be rearranged and displayed in the first area of the display based on the index of the data may be stored.
[0026] According to one embodiment, a method of operating an electronic device may be provided. The method may receive a first user voice input through a microphone of the electronic device. The method may generate first translation information by translating the received first user voice input using a machine learning model based on the received first user voice input, wherein the machine learning model includes a first machine learning model trained to output translated translation information based on the voice input. The method may transmit the first translation information to an external electronic device connected to the electronic device through the transceiver. The method may receive second translation information by translating a second user voice input from the external electronic device through the transceiver. The method may display the second translation information in a first area of the display.
[0027] According to one embodiment, a storage medium storing at least one computer-readable instruction may be provided. The at least one instruction, when executed by at least a part of at least one processor of an electronic device, may cause the electronic device to perform at least one operation. The at least one operation may receive a first user voice input through a microphone of the electronic device. The at least one operation may generate first translation information translated from the received first user voice input using a machine learning model based on the received first user voice input, wherein the machine learning model includes a first machine learning model trained to output translated translation information based on the voice input. The at least one operation may transmit the first translation information to an external electronic device connected to the electronic device through the transceiver. The at least one operation may receive second translation information translated from a second user voice input from the external electronic device through the transceiver. The at least one operation may display the second translation information in a first area of the display.
[0028] According to another embodiment of the present disclosure, a translation system may include a wearable electronic device; and an electronic device connected to the wearable electronic device. The wearable electronic device may: receive a first user voice input through a microphone of the wearable electronic device, and generate first translation information in which the received first user voice input is translated using a machine learning model of the wearable electronic device based on the received first user voice input, wherein the machine learning model includes a model trained to output translation information based on a voice input, transmit the first translation information to the electronic device, receive second translation information in which a second user voice input is translated from the electronic device, and output the second translation information through a speaker of the wearable electronic device. The electronic device may: receive the second user voice input through a microphone of the electronic device, and generate the second translation information using a machine learning model of the electronic device based on the received second user voice input, wherein the machine learning model of the electronic device includes a model trained to output translated information based on a voice input; transmit the second translation information to the wearable electronic device, receive the first translation information from the wearable electronic device, and display the first translation information on a first area of a display of the electronic device.
[0029] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.
[0030] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments.
[0031] FIG. 2 is a block diagram of a wearable electronic device (201) according to one embodiment of the present disclosure.
[0032] FIG. 3 is an example of simultaneous interpretation operations of an electronic device (101) and a wearable electronic device (201) according to one embodiment of the present disclosure.
[0033] FIG. 4 illustrates a data flow according to simultaneous interpretation operations of an electronic device (101) and a wearable electronic device (201) according to one embodiment of the present disclosure.
[0034] FIG. 5 is a flowchart for explaining a simultaneous interpretation method of an electronic device (101) according to one embodiment of the present disclosure.
[0035] FIG. 6 is a flowchart for explaining a simultaneous interpretation method of a wearable electronic device (201) according to one embodiment of the present disclosure.
[0036] FIG. 7 is a block diagram of an AI (artificial intelligence) interpretation module according to one embodiment of the present disclosure.
[0037] FIG. 8 is a flowchart illustrating an operation of an electronic device (101) and a wearable electronic device (201) performing simultaneous interpretation according to one embodiment of the present disclosure.
[0038] FIG. 9 is an example of translation information data according to one embodiment of the present disclosure.
[0039] FIG. 10 is an example of user voice input according to one embodiment of the present disclosure.
[0040] FIG. 11 illustrates a data flow according to a correction translation operation of an electronic device (101) and a wearable electronic device (201) according to an embodiment of the present disclosure.
[0041] FIG. 12 is a flowchart illustrating an operation of an electronic device (101) and a wearable electronic device (201) to correct translation according to an embodiment of the present disclosure.
[0042] FIG. 13 is a flowchart for explaining an operation of an electronic device (101) and a wearable electronic device (201) to output an interpretation result according to an embodiment of the present disclosure.
[0043] FIGS. 14a and 14b are examples of a real-time interpretation result display screen of an electronic device (101) according to one embodiment of the present disclosure.
[0044] FIG. 15 is an example of a multi-party simultaneous interpretation operation of an electronic device (101) and a plurality of wearable electronic devices (201) according to one embodiment of the present disclosure.
[0045] FIGS. 16a and 16b are examples of a multi-party simultaneous interpretation result display screen of an electronic device (101) according to one embodiment of the present disclosure.
[0046] FIG. 17 is a flowchart illustrating a method for an electronic device (101) to perform simultaneous interpretation using a wearable electronic device (201) according to one embodiment of the present disclosure.
[0047] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.
[0048] An embodiment of the present disclosure will be described below with reference to the attached drawings.
[0049] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments.
[0050] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0051] The processor (120) may control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing, for example, software (e.g., a program (140)), and may perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculation, the processor (120) may store a command or data received from another component (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the command or data stored in the volatile memory (132), and store the resulting data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0052] The auxiliary processor (123) may control at least a part of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0053] The processor (120) can control the operations of the electronic device (101) by executing instructions stored in the memory (130). For example, the processor (120) can correspond to a plurality of processors that collectively perform a plurality of operations by dividing them among the processors.
[0054] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0055] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0056] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0057] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0058] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0059] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0060] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0061] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) to an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multi-media interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0062] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0063] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0064] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0065] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0066] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0067] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0068] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0069] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the selected at least one antenna. According to some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0070] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0071] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0072] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service by itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0073] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0074] FIG. 2 is a block diagram of a wearable electronic device (201) according to one embodiment of the present disclosure.
[0075] According to one embodiment, a wearable electronic device (201) may include an audio module (210), a processor (220), a memory (230), or a transceiver (240). In some embodiments, the wearable electronic device (201) may omit at least one of these components, or may have one or more other components added. In some embodiments, some of these components may be integrated into a single component.
[0076] The audio module (210) may include at least one microphone (211), at least one speaker (212), or an audio processor (213). According to one embodiment, the wearable electronic device (201) may be manufactured in a form that can be worn on the ears, and the microphone (211) and the speaker (212) may be arranged in two physically separate electronic devices, respectively. For example, the wearable electronic device (201) may have a first structure (a form that can be worn on the left ear) and a second structure (a form that can be worn on the right ear) corresponding to the first structure so that it can be worn on both ears of the user. The first structure of the wearable electronic device (201) may include a first microphone and a second speaker, and the second structure may include a second microphone and a second speaker. The wearable electronic device (201) may receive an audio signal corresponding to a sound acquired from the outside through a plurality of microphones. A wearable electronic device (201) can output audio signals through multiple speakers.
[0077] The audio signal processor (213) receives an analog audio signal input through a microphone (211) and converts it into a digital audio signal through an analog to digital converter (ADC), and can perform various processing on the received audio signal. For example, according to one embodiment, the audio signal processor (213) can change a sampling rate, apply one or more filters, perform interpolation processing, amplify or attenuate all or part of a frequency band, process noise (e.g., noise or echo reduction), change a channel (e.g., switching between mono and stereo), mix, or extract a specified signal on one or more digital audio signals. According to one embodiment, one or more functions of the audio signal processor (213) can be implemented in the form of an equalizer.
[0078] The processor (220) may, for example, execute software (e.g., a program) to control at least one other component (e.g., a hardware or software component) of the wearable electronic device (201) connected to the processor (220) and perform various data processing or calculations. According to one embodiment, as part of the data processing or calculation, the processor (220) may store commands or data received from other components (e.g., an audio module (210) or a transceiver (240)) in a volatile memory, process the commands or data stored in the volatile memory, and store result data in a non-volatile memory. According to one embodiment, the processor (220) may include a main processor (e.g., a central processing unit or an application processor) or an auxiliary processor (e.g., a neural processing unit (NPU)) that can operate independently or together therewith. For example, the auxiliary processor may perform an operation of the machine learning model (231) included in the memory (230) on an audio signal input from the audio module (210) and transmit the result of the operation to the processor (220). The processor (200) may control the operations of the wearable electronic device (201) by executing instructions stored in the memory (230). For example, the processor (220) may correspond to a plurality of processors that collectively perform a plurality of operations by dividing them among the processors.
[0079] The memory (230) can store various data used by at least one component (e.g., a processor (220) or an audio module (210)) of the wearable electronic device (201). The data can include, for example, input data or output data for software (e.g., a program) and commands related thereto.
[0080] In one embodiment, the memory (230) may store a machine learning model (231) that performs at least one operation. According to one embodiment, the machine learning model (231) may include various components that perform detailed operations to achieve an interpretation function. For example, the machine learning model (231) may include at least one of: ASR (automation speech recognition), STT (speech-to-text), LLM (large language model), S2ST (speech-to-speech translation), a first machine learning model trained to output translated translation information based on a voice input (hereinafter, referred to as a translation learning model), a second machine learning model trained to identify a user's voice based on a user's voice input (hereinafter, referred to as a voice identification learning model), or a third machine learning model trained to recognize a mouth shape based on a continuous image or video input of a user's mouth and identify an utterance through the mouth shape (hereinafter, referred to as a mouth shape identification learning model). According to an embodiment, the first machine learning model may include at least one of ASR, STT, LLM, or S2ST. For example, the first machine learning model may include S2ST and output a voice signal that has interpreted an input voice signal. Alternatively, the first machine learning model may include ASR, LLM, and STT and output text data that has interpreted an input voice signal (or voice data that has converted text data into voice). According to an embodiment, the voice recognition learning model may be trained to identify a specific user voice (e.g., a user of a wearable electronic device (201)) acquired through a microphone, such as during a call or a voice input. The voice recognition learning model may perform supervised learning, semi-supervised learning, or re-training by continuously using the user voice input through the microphone as learning data.In one embodiment, the wearable electronic device (201) can identify a user voice from among voice signals acquired through a microphone (211) and transmit information about the identified user voice to the electronic device (101).
[0081] According to one embodiment, the electronic device (101) may include a machine learning model. The machine learning model included in the electronic device (101) and the machine learning model (231) included in the wearable electronic device (201) may differ in some aspects. The machine learning model (231) included in the wearable electronic device (201) may have more limited functions than the machine learning model of the electronic device (101). For example, the translation learning model included in the electronic device (101) may be capable of translating multiple languages (e.g., Korean, English, Spanish), while the translation learning model of the wearable electronic device (201) may only have the function of translating a smaller number of languages (e.g., English). Alternatively, the translation learning model included in the electronic device (101) may be trained with more training data than the training data of the translation learning model included in the wearable electronic device (201). The translation learning model included in the electronic device (101) may provide more accurate or sophisticated expressions, expressions that reflect the user's language habits, etc. However, the translation learning model of the wearable electronic device (201) is not limited to the translation learning model of the electronic device (101), and the translation learning model of the wearable electronic device (201) may be optimized for the user or may provide functions that are substantially the same as or higher than the translation learning model included in the electronic device (101).
[0082] The translation learning model of the wearable electronic device (201) may be periodically updated by being connected to the translation learning model of the functionally connected electronic device (101). The machine learning model (231) of the wearable electronic device (201) may include a personalized translation learning model trained to perform personalized translation based on the user's speech patterns or frequently used words.
[0083] In one embodiment, the machine learning model (231) may include neural networks, transformers, sequence-to-sequence models, large language models, and / or bidirectional embeddings.
[0084] The transceiver (240) can support the establishment of a long-distance or short-distance wireless communication channel between the wearable electronic device (201) and an external electronic device (e.g., the electronic device (101) or a server), and the performance of communication through the established communication channel. In one embodiment, the long-distance wireless communication may be, for example, a cellular communication module or a GNSS (global navigation satellite system) communication module. In one embodiment, the short-distance wireless communication may be, for example, Bluetooth, WiFi (wireless fidelity) direct, UWB (ultra-wideband), or IrDA (infrared data association).
[0085] According to one embodiment, the transceiver (240) can perform pairing between the wearable electronic device (201) and another electronic device (e.g., the electronic device (101)). For example, the transceiver (240) can perform pairing of the wearable electronic device (201) with another electronic device of the user (e.g., a mobile electronic device (101) or a wearable electronic device (e.g., a smart ring, a smart watch)). In one embodiment, the wearable electronic device (201) can automatically connect to another electronic device (101) paired through the transceiver (240) while being worn by the user.
[0086] The configuration of the wearable electronic device (201) may be partially or entirely identical to the configuration of the electronic device (101) of FIG. 1.
[0087] FIG. 3 is an example for explaining a simultaneous interpretation situation by an electronic device (101) and a wearable electronic device (201) according to one embodiment of the present disclosure.
[0088] According to one embodiment, the electronic device (101) and the wearable electronic device (201) may output the result of simultaneous interpretation of a conversation between a user (A) and a counterpart (B) to the speaker of the wearable electronic device (201), the speaker of the electronic device (101), or the display. In the dictionary, “interpretation” means translating words so that meaning can be communicated between people who do not speak the same language, and “translation” means translating a text in one language into another language. In the embodiments of the present disclosure, interpretation may mean converting the speech of the user (A) into another language appropriate for the situation and outputting it. The electronic device (101) or the wearable electronic device (201) may interpret and process an audio signal and output it as an audio signal. Specifically, the interpretation process may include a process of converting the audio signal into text, translating the text, and then converting it back into an audio signal. In various embodiments of the present disclosure, the terms interpretation and translation are used interchangeably to refer to processing a user's speech (audio signal) and outputting it in another language (audio signal or text signal). Alternatively, the term "interpretation" may be interpreted as "translation," and even if it is described as "translation," it may be interpreted as "interpretation" in terms of meaning.
[0089] In various embodiments of the present disclosure, an electronic device (101) and a wearable electronic device (201) can simultaneously interpret a conversation between two speakers. In various embodiments of the present disclosure, an electronic device (101) and a plurality of wearable electronic devices (201) (e.g., wearable electronic devices in the form of clothes, glasses, watches, rings, earphones, etc.) can interpret a conversation between a plurality of speakers in real time.
[0090] According to one embodiment, a wearable electronic device (201) may receive a user's (A) speech, perform an interpretation, transmit the interpreted result to the electronic device (101), receive information translating the speech of the other party (B), and output it through a speaker. According to one embodiment, an electronic device (101) may receive a user's (B) speech, perform an interpretation, transmit the interpreted result to the wearable electronic device (201), and receive information translating the speech of the user (A), and output it through a speaker or a display.
[0091] Referring to FIG. 3, for example, a user (A) may speak in Korean, and a counterpart (B) conversing with the user (A) may speak in Spanish. The electronic device (101) and the wearable electronic device (201) may receive audio signals for the user's (A) speech and the counterpart's (B) speech, process the interpretation, and output the result as an audio signal or text from the electronic device (101) or the wearable electronic device (201) as appropriate for the situation. The wearable electronic device (201) may be worn on the body (e.g., the ear) of the user (A) so that the user (A) may use the interpretation function. The electronic device (101) may be positioned within the visual range of the counterpart (B) and / or within the reception range of a microphone (e.g., the input module (150) of FIG. 1) so that the counterpart (B) may use the interpretation function.
[0092] According to one embodiment, an electronic device (101) and a wearable electronic device (201) may be connected to each other based on short-range wireless communication (e.g., Bluetooth). For example, the electronic device (101) and the wearable electronic device (201) may perform pairing including user authentication.
[0093] According to one embodiment, the electronic device (101) and the wearable electronic device (201) can receive audio signals regarding speech of the user (A) or the other party (B) through their respective microphones (e.g., the microphone (211) of the wearable electronic device (201) or the microphone (150) of the electronic device (101)) while the user (A) is carrying or wearing the electronic device.
[0094] In one embodiment, the wearable electronic device (201) may be positioned closer to the mouth of the user (A) than to the mouth of the counterpart (B) while being worn on the ear of the user (A). The wearable electronic device (201) worn by the user (A) may interpret and process utterances made by the user (A). The electronic device (101) of the user (A) (e.g., a mobile electronic device (101) placed on the hand of the user (A)) may interpret and process utterances made by the counterpart (B) with whom the user (A) is conversing in close proximity.
[0095] According to one embodiment, the electronic device (101) or the wearable electronic device (201) may determine whether a voice input is the voice of a user (e.g., user A) by using a machine learning model (hereinafter, referred to as a voice identification learning model) trained to identify the voice of a user (e.g., user A) based on a voice input previously received through each microphone (e.g., microphone (150) of the electronic device (101) or microphone (212) of the wearable electronic device (201). Through a conversation between the user (A) and the other party (B), the wearable electronic device (201) and the electronic device (101) may each receive voice signals for speech of the two people. In one embodiment, the wearable electronic device (201) may filter the voice signal for the speech of the other party (B) in order to process interpretation for the user (A). In one embodiment, the electronic device (101) may filter a voice signal of a user's (A) speech in order to process an interpretation for the other party (B). For example, the wearable electronic device (201) may identify the user's voice from a voice input acquired through a microphone (212) and perform an interpretation for the voice information of the user (A) filtered out except for the identified user's voice. The electronic device (101) may identify the user's voice from a voice input acquired through a microphone (150) and perform an interpretation for the voice information of the other party (B) filtered out except for the portion including the identified user's voice.
[0096] According to one embodiment, the electronic device (101) may receive a user input for a speech of the other party (B), process the interpretation, display it on the display (160) of the electronic device (101), and transmit it to the wearable electronic device (201) so that it can be output to the speaker (212) of the wearable electronic device (201). The electronic device (101) may receive translation information that translates the speech of the user (A) from the wearable electronic device (201), display it as text on the display (160), or output it as an audio signal through the speaker (155).
[0097] According to one embodiment, a wearable electronic device (201) may receive a user input for a user's (A) speech, interpret it, and transmit it to the electronic device (101) so that it can be displayed on the display (160) of the electronic device (101). The wearable electronic device (201) may receive translation information for a counterpart's (B) speech from the electronic device (101) and output it as an audio signal through a speaker (212).
[0098] Referring to FIG. 3, the electronic device (101) may obtain a first user voice input for a Spanish utterance (310) of a counterpart (B) such as “Hay algun lugar cerca donde pueda comer pasta deliciosa?” through a microphone (150), generate first translation information (e.g., “Is there a place nearby where I can eat delicious pasta?”) translated into Korean of the first user voice input using a machine learning model included in the memory (130) of the electronic device (101), display the first translation information on a first part (311) of a display (160), and transmit the first translation information to a wearable electronic device (201) through a transceiver (190). The wearable electronic device (201) may output the first translation information translated into Korean for the Spanish utterance (310) of the counterpart (B) received from the electronic device (101) through a speaker (212).
[0099] Referring to FIG. 3, the wearable electronic device (201) may obtain a second user voice input for a Korean utterance (320) of a user (A) saying, “I know a delicious pasta restaurant nearby” through a microphone (211), generate second translation information (e.g., “Conozco un delicioso restautant de pasta cerca.”) in which the second user voice input is translated into Spanish using a machine learning model (231) included in a memory (230) of the wearable electronic device (201), and transmit the second translation information to the electronic device (101) through a transceiver (240). The electronic device (101) may display the second translation information translated into Spanish for the Korean utterance (320) of the user (A) received from the wearable electronic device (201) on a second part (321) of the display (160).
[0100] In various embodiments, when the other party (B) wears his / her wearable electronic device, the electronic device (101) can transmit a text or audio signal to the other party's wearable electronic device so that the translation result of the user (A) can be output through the speaker or display of the other party's wearable electronic device. For example, when the other party (B) wears a wearable electronic device on his / her ear, the wearable electronic device (201) of the user (A) can transmit an audio signal, which is the result of translating the user's speech, to the other party's wearable electronic device so that the audio signal is output through the speaker of the wearable electronic device of the other party (B). The translation of the other party's speech can be performed by the electronic device (101) of the user (A). Simply, the display or speaker of the electronic device carried or worn by the other party (B) can be used as an output device for the simultaneous interpretation function of the electronic device (101) and the wearable electronic device (201).
[0101] In various embodiments, the wearable electronic device (201) may be manufactured as a first structure and a second structure that are physically distinct so as to be wearable on both ears of a user. The user (A) and the other party (B) may each wear the first structure and the second structure of the wearable electronic device (201). In order to simultaneously interpret a conversation between the user (A) and the other party (B), in a situation where the wearable electronic device (201) interprets and processes the user's (A) speech and the electronic device (101) interprets and processes the other party's (B) speech, some speakers of the wearable electronic device (201) (e.g., the second structure wearable on the right ear) may be worn by the other party (B) to use the speaker function. For example, when the first structure of the wearable electronic device (201) is worn by the user (A), the interpretation result for the other party's (B) speech may be output through the first speaker included in the first structure. Conversely, when the second structure of the wearable electronic device (201) is worn by the other party (B), the interpretation result for the user's (A) speech can be output through the second speaker included in the second structure. For a method of outputting different sound signals to each of the two speakers (212) included in one wearable electronic device (201), reference can be made to the multi-device communication method in a Bluetooth communication environment disclosed in Korean Patent Publication No. KR10-2021-01509019.
[0102] FIG. 4 illustrates a data flow according to simultaneous interpretation operations of an electronic device (101) and a wearable electronic device (201) according to one embodiment of the present disclosure.
[0103] According to an embodiment, an electronic device (101) and a wearable electronic device (201) can interpret and process voice signals generated by speech of two people in a conversation. Referring to FIG. 4, as shown in the data flow indicated by a solid line, the wearable electronic device (201) according to an embodiment can receive and process speech (410) of a first user (hereinafter, referred to as a user) who wears the wearable electronic device (201) and speaks in a first language through a microphone (211). Referring to FIG. 4, as shown in the data flow indicated by a dotted line, the electronic device (101) connected to the wearable electronic device (201) can receive and process speech (420) of a second user (hereinafter, referred to as a counterpart) who does not wear the wearable electronic device (201) and speaks in a second language through a microphone (150).
[0104] According to one embodiment, a wearable electronic device (201) may obtain a voice input (410) of a first user speaking in a first language through a microphone (211). The wearable electronic device (201) (or the processor (220) of the wearable electronic device (201)) may identify the user's voice from the received voice input (410) of the first user using a voice recognition learning model of a machine learning model (231), obtain user voice information by filtering out the remainder excluding the user's voice, and perform a translation of the user voice information. The machine learning model (231) may include speech-to-text (STT) or text-to-speech (TTS). For example, the wearable electronic device (201) may convert an audio signal into text using STT, translate the converted text into a target language (e.g., a second language spoken by the other party) to generate translation information, and output the translation information in the form of text as an audio signal through TTS. The wearable electronic device (201) can transmit first translation information, which is a translation of a first user's voice input (410), to the electronic device (101) via the transceiver (240). In one embodiment, the translation information can include voice data (audio) and text data (text) translated into a target language of the user's voice input, original text data corresponding to the user's voice input, or information about the user's voice input.
[0105] According to an embodiment, an electronic device (101) may transmit first translation information received through a transceiver (190) to a processor (120) or a speaker (155). The speaker (155) of the electronic device (101) according to an embodiment may output voice data included in the received first translation information as an audio signal. The processor (120) of the electronic device (101) according to an embodiment may output text data included in the first translation information on a screen through a display (160). A counterparty conversing with a user may confirm the user's translated speech through the voice output through the speaker (155) of the electronic device (101) or the interpretation function screen displayed on the display (160).
[0106] According to an embodiment, an electronic device (101) may obtain a voice input from a second user (e.g., a counterpart) through a microphone (150). According to an embodiment, the electronic device (101) (or the processor (120) of the electronic device (101)) may identify the user's voice based on the voice input (420) of the second user received using a voice recognition learning model of a machine learning model, obtain the counterpart's voice information by filtering the first user's voice, and perform a translation of the counterpart's voice information. The machine learning model may include STT or TTS. For example, the electronic device (101) may convert an audio signal into text using STT, translate the converted text into a target language (e.g., a first language spoken by the user) to generate translation information, and output the translation information in the form of text as an audio signal through TTS. According to one embodiment, the electronic device (101) may transmit second translation information, in which a second user voice input (420) is translated, to the wearable electronic device (201) through the transceiver (190), and simultaneously output text data included in the second translation information on the screen through the display (160). According to one embodiment, the electronic device (101) may use a machine learning model stored in the memory (130) to perform a translation function, or request a translation from a server (not shown) providing an AI interpretation function, and receive the translation result.
[0107] According to one embodiment, the machine learning model (231) of the wearable electronic device (201) may be updated based on information detected by the electronic device (101) (e.g., GPS information). For example, the machine learning model of the electronic device (101) may provide translation functions for Spanish, English, Chinese, Japanese, and French. The machine learning model (231) of the wearable electronic device (201) may include a learning model that is lighter than the machine learning model of the electronic device (101), that is, provides translation functions for a smaller number of languages, depending on the size of the memory (230) of the wearable electronic device (201) or the performance of the processor (220). In a situation where translation is required for a language not supported by the machine learning model (231) of the wearable electronic device (201), the electronic device (101) may support an update to the machine learning model (231) of the wearable electronic device (201). The wearable electronic device (201) can add necessary information (e.g., languages to be additionally supported for translation) from the electronic device (101) or delete some information (e.g., languages to be translated that are currently unnecessary) considering the hardware status of the wearable electronic device (201). For example, if a user is traveling to Spain, the electronic device (101) can support Korean-to-Spanish translation. At this time, the necessary language can be determined using pre-recognized travel information (e.g., travel destinations identified using messages or airplane ticket images) using GPS information, calendar information, or a machine learning model of the electronic device (101). If the wearable electronic device (201) does not support Korean-to-Spanish translation, the electronic device (101) can transmit a learning model for Spanish-to-Korean translation of the electronic device (101) to the machine learning model of the wearable electronic device (201), thereby updating the machine learning model of the wearable electronic device (201).
[0108] According to one embodiment, the electronic device (101) may include a separate machine learning model for each language that it supports for translation. For example, the electronic device (101) may include a machine learning model for translating Spanish to Korean, a machine learning model for translating Chinese to English, and a machine learning model for translating Spanish to French to English. According to one embodiment, the electronic device (101) may identify a language that requires translation based on information detected by the electronic device (e.g., GPS information). According to one embodiment, the electronic device (101) may select a machine learning model corresponding to the language that requires translation and activate the translation function. According to one embodiment, if the wearable electronic device (201) does not include a machine learning model corresponding to the language that requires translation, the wearable electronic device (201) may receive the corresponding machine learning model from the electronic device (101). For example, if a user is traveling to Spain, the electronic device (101) can determine that the language requiring translation is Korean-Spanish by using pre-recognized travel information (e.g., a travel destination identified using a message or an airplane ticket image) using GPS information, calendar information, or a machine learning model of the electronic device (101). If the machine learning model included in the wearable electronic device (201) does not support Korean-Spanish translation, the electronic device (101) can transmit a machine learning model that supports Spanish-Korean translation of the electronic device (101) to the wearable electronic device (201).
[0109] According to various embodiments of the present disclosure, when a user carries or wears a plurality of wearable electronic devices (201), at least one wearable electronic device (201) that supports a translation function may provide a translation function to the user. For example, when a first wearable electronic device (e.g., Buzz) does not support a translation function and a second wearable electronic device (e.g., Watch) supports a translation function, a message and / or notification (e.g., vibration) indicating that the second wearable electronic device provides a translation function may be output through the first wearable electronic device or the second wearable electronic device. For example, the second wearable electronic device may output a guidance message such as “Please speak closer” on the display of the second wearable electronic device to enable the user to use its translation function. In one embodiment, the first wearable electronic device may receive a voice signal corresponding to a user’s speech and transmit the received voice signal to the second wearable electronic device. The second wearable electronic device can receive a voice signal transmitted from the first wearable electronic device, process a real-time translation of the received voice signal, and then transmit the translated information generated by the translation to the electronic device and the first wearable electronic device, respectively.
[0110] According to various embodiments of the present disclosure, the electronic device (101) may include one or more cameras, and may be positioned on the same surface as the display (160) of the electronic device (101). When two or more speakers speak simultaneously, the electronic device (101) may use an image captured of the mouth shape of speaker B (e.g., speaker B of FIG. 3)) through a camera positioned on the same surface as the display (160) to determine which of the multiple voices included in the voice signal acquired through the microphone is the speech of speaker B. According to an embodiment, the machine learning model of the electronic device (101) may include a lip shape identification learning model trained to identify the general content of speech through the lip shapes of users during self-camera shooting or video calls. The lip shape identification learning model may be trained and stored in the memory (130), or may be retrained using data acquired while using the translation function as learning data.
[0111] An electronic device (101) according to an embodiment may capture an image of speaker B through a camera, and generate a character of speaker B using a generative AI model based on the captured image of speaker B. The electronic device (101) according to an embodiment may use the character of speaker B for a translation function. For example, in order to display an area for outputting translation information for speaker B's utterance on a translation result screen output by the electronic device (101), the character of speaker B may be displayed together. The generated character may be an image or a moving image. When the character is generated using a moving image, the character may be displayed in a form of speaking in accordance with the utterance of speaker B, which is the target of the character. The generative AI model may be an AI model trained to output another image of a similar form to an input image. The generative AI model may generate a plurality of characters, and the user's character may be determined from among the plurality of characters generated by the user's selection. The character of speaker A (e.g., speaker A in FIG. 3) may be generated through generative AI using an image pre-stored in the electronic device (101). An electronic device (101) according to one embodiment may receive identification information identifying a voice for a user input from a wearable electronic device (201) and display a character (graphical object) corresponding to the identification information along with translation information.
[0112] According to various embodiments of the present disclosure, when a wearable electronic device (201) interprets a speaker A's speech and an electronic device (101) interprets a speaker B's speech, the interpretation may be performed using active noise cancelling (ANC). For example, the wearable electronic device (201) that has received translation information for a speaker B's speech from the electronic device (101) may output audio for the translation information through the speaker (212) of the wearable electronic device (201) while automatically activating the ANC function. According to one embodiment, the wearable electronic device (201) may detect ambient noise through the ANC function and remove the detected noise while receiving a voice input for a speaker A's speech through the microphone (211), thereby obtaining a clearer voice input signal. In one embodiment, the wearable electronic device (201) can detect ambient noise using a signal acquired through a microphone (211) while speaker A is not speaking, or a signal acquired through a microphone (150) of the electronic device (101) while speakers A and B are not conversing.
[0113] The wearable electronic device (201) can transmit the second translation information received through the transceiver (240) to the speaker (212). The speaker (212) of the wearable electronic device (201) can output voice data included in the received second translation information as an audio signal. The user (410) can confirm the translated speech of the other party (420) through the voice output through the speaker (212) of the wearable electronic device (201) or the interpretation function screen displayed on the display (160).
[0114] FIG. 5 is a flowchart for explaining a simultaneous interpretation method of an electronic device (101) according to one embodiment of the present disclosure.
[0115] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0116] According to one embodiment, steps S510 to S541 may be understood to be performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0117] According to one embodiment, an electronic device (101) can perform simultaneous interpretation of a conversation between two speakers. The operation of the electronic device (101) of FIG. 5 may correspond to the embodiment of FIG. 4 described above. The electronic device (101) according to one embodiment may include at least a portion of a microphone (150), a display (160), a memory (130), a processor (120), or a transceiver (190).
[0118] According to one embodiment, at step S510, the electronic device (101) may execute an interpretation function. For example, the interpretation function may be executed based on a user input for executing the interpretation function. According to one embodiment, the electronic device (101) may complete the settings for the interpretation function based on initial settings including input language settings, target language settings, or microphone settings.
[0119] According to one embodiment, at step S520, the electronic device (101) can activate the microphone and speaker. It can check whether the wearable electronic device (201) is being worn, and if not, a guidance message such as "Please wear earbuds for interpretation" can be output through the display or speaker of the electronic device (101).
[0120] According to one embodiment, the electronic device (101) can detect whether a user's voice is input from a microphone (150) as in step S530 while executing an interpretation function, and can confirm whether translation information is received from a wearable electronic device (201) as in step S540.
[0121] According to one embodiment, in step S530, the electronic device (101) may detect that a user voice is received from the microphone (150). The electronic device (101) according to one embodiment may receive a first user voice input through the microphone (150). The electronic device (101) according to one embodiment may generate translation information based on a machine learning model in response to the user voice being input (S531). The electronic device (101) may generate first translation information translated from the received first user voice input using the machine learning model based on the received first user voice input. The machine learning model may include a translation learning model trained to output translated translation information based on a voice input. The machine learning model may include a voice identification learning model trained to identify a user voice based on a voice input previously received by the microphone (150). An electronic device (101) according to an embodiment may identify a user's voice from among a first user's voice input using a voice recognition learning model, and filter a portion of the first user's voice input containing the identified user's voice to obtain user voice information (e.g., a speaker not wearing a wearable electronic device (201)) to be translated. An electronic device (101) according to an embodiment may translate the user's voice information using a translation learning model to generate first translation information.
[0122] According to one embodiment, at step S532, the electronic device (101) can output the first translation information to the display (160) of the electronic device (101) and simultaneously transmit it to an external electronic device (e.g., a wearable electronic device (201)) connected to the electronic device (101).
[0123] According to one embodiment, in step S540, the electronic device (101) may receive second translation information from an external electronic device (e.g., a wearable electronic device (201)) connected to the electronic device (101). The second translation information may include user identification information. The electronic device (101) according to one embodiment may receive the second translation information, convert the second translation information into voice information (text to speech) using a translation learning model, and output the converted text through the speaker (155). The electronic device (101) according to one embodiment may convert the second translation information into voice information using at least one piece of user voice information pre-stored in the electronic device (101), and then output the converted text through the speaker (155).
[0124] In response to receiving second translation information, the electronic device (101) according to an embodiment may output the received second translation information to the display (160) and output the second translation information as an audio signal through the speaker (155) (S541). The electronic device (101) according to an embodiment may display the second translation information in a first area of the display (160), and the display (160) may display the first translation information in a second area. The first translation information displayed in the second area may be displayed in an opposite direction to the second translation information displayed in the first area. The electronic device (101) according to an embodiment may display a graphical object corresponding to user identification information together with the second translation information. The second translation information may include at least one of translated data, an index of the data, or a flag for the end of utterance. The electronic device (101) according to an embodiment may rearrange translated data included in the second translation information based on the index of the data and display the rearranged translated data in the first area of the display (160).
[0125] An electronic device (101) according to one embodiment may further include at least one sensor including a GPS, and may generate update information including a portion of the translation learning model based on information input from the at least one sensor, and transmit the update information to the external electronic device (e.g., a wearable electronic device (201)). The update information may include update information associated with a translation support language.
[0126] An electronic device (101) according to one embodiment may further include at least one camera, acquire a lip shape based on a lip image captured through the at least one camera, and store instructions for identifying a user's voice using a machine learning model based on the lip shape. The machine learning model may include a third machine learning model trained to identify the content of speech based on the lip shape.
[0127] FIG. 6 is a flowchart for explaining a simultaneous interpretation method of a wearable electronic device (201) according to one embodiment of the present disclosure.
[0128] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0129] According to one embodiment, steps S610 to S641 may be understood to be performed in a processor (e.g., processor (220) of FIG. 2) of a wearable electronic device (e.g., wearable electronic device (201) of FIG. 2).
[0130] According to one embodiment, a wearable electronic device (201) can perform simultaneous interpretation of a conversation between two speakers. The operation of the wearable electronic device (201) of FIG. 6 may correspond to the embodiment of FIG. 4 described above.
[0131] According to one embodiment, in step S610, the wearable electronic device (201) can execute an interpretation function. For example, the wearable electronic device (201) can execute an interpretation function when a specific gesture preset for the interpretation function is input.
[0132] According to one embodiment, in step S620, the wearable electronic device (201) may activate the microphone (211) and the speaker (212). In one embodiment, the wearable electronic device (201) may be paired with the electronic device (101). According to one embodiment, when the wearing of the wearable electronic device (201) is detected, the wearable electronic device (201) may be set to automatically establish a communication connection with the electronic device (101).
[0133] According to one embodiment, the wearable electronic device (201) can detect whether a user's voice is input from a microphone (211) as in step S630 while executing an interpretation function, and can check whether translation information is received from a paired electronic device (101) as in step S640.
[0134] According to one embodiment, in step S630, the wearable electronic device (201) may detect that a user voice is received from a microphone (211). The wearable electronic device (201) may receive a second user voice input through the microphone (211). In one embodiment, the wearable electronic device (201) may detect noise input in a portion where the user voice does not exist among the second user voice input, remove an input corresponding to the noise detected in the portion where the user voice does not exist among the second user voice input, obtain user voice information, and translate the obtained user voice information to generate second translation information.
[0135] According to an embodiment, a wearable electronic device (201) may generate translation information based on a machine learning model in response to a user voice input (S631). According to an embodiment, the wearable electronic device (201) may generate second translation information translated from a second user voice input using a machine learning model stored in a memory (230) of the wearable electronic device (201). The machine learning model may include a translation learning model trained to output translated translation information based on a voice input. The machine learning model may include a voice identification learning model trained to identify a user's voice based on a voice input previously received through a microphone (211). The wearable electronic device (201) may convert the second user voice input into text using the machine learning model and translate the converted text to generate the second translation information. A wearable electronic device (201) according to one embodiment may receive update information from an external electronic device (e.g., an electronic device (101) or a server connected to the wearable electronic device (201)) and update a machine learning model based on the received update information.
[0136] According to an embodiment, a wearable electronic device (201) may identify a user's voice from among a second user's voice input based on a voice recognition learning model, and filter out the remaining voices except for the identified user's voice from among the second user's voice inputs to obtain user voice information (e.g., a user wearing the wearable electronic device (201). According to an embodiment, a wearable electronic device (201) may translate the obtained user voice information based on a translation learning model to generate first translation information. The first translation information may include at least one of translated data, an index of the data, or a flag for the end of utterance. Here, the index of the data may indicate the order in which the translated data is positioned in a sentence structure, and the flag may indicate whether the utterance has ended as identified by voice activity detection (VAD).
[0137] According to one embodiment, in step S632, the wearable electronic device (201) may transmit second translation information to an external electronic device (e.g., electronic device (101)) connected to the wearable electronic device (201). The second translation information may include an audio signal. The wearable electronic device (201) may transmit user identification information corresponding to the identified user's voice to the electronic device (101) together with the second translation information.
[0138] According to one embodiment, at step S640, the wearable electronic device (201) may receive first translation information translated from a first user voice input from an external electronic device (e.g., electronic device (101)) connected to the wearable electronic device (201). In response to receiving the first translation information, the wearable electronic device (201) may output the first translation information as an audio signal through the speaker (212) (S641).
[0139] FIG. 7 is a block diagram of a machine learning model according to an embodiment of the present disclosure.
[0140] According to one embodiment, the electronic device (101) or the wearable electronic device (201) may store a machine learning model in a memory (e.g., the memory (130) of the electronic device (101) or the memory (230) of the wearable electronic device (201), respectively). The machine learning model may be stored in the form of a learning model or a program. The machine learning model may be compiled in whole or in part (e.g., a user voice identification learning model) and executed by a processor (e.g., the processor (120) of the electronic device (101) or the processor (220) of the wearable electronic device (201)). Hereinafter, for convenience of explanation, the operation by the processor (120) of the electronic device (101) will be described with a focus on the operation. The embodiments described below may be operated by the processor (220) of the wearable electronic device (201).
[0141] According to one embodiment, the electronic device (101) or the processor (120) of the electronic device (101) (hereinafter, processor (120)) may process audio input using a machine learning model stored in the memory (130) and output a translated audio signal.
[0142] A machine learning model may include various components that perform detailed operations to achieve the interpretation function. For example, the machine learning model may include automated speech recognition (ASR) (701), speech-to-text (STT) (702), a large language model (LLM) (703), translation (704), proofreading (705), or text-to-speech (TTS) (706). Translation (704) or proofreading (705) may be included in LLM (703).
[0143] The processor (120) can recognize the user's voice from a sound signal acquired through a microphone (150) using ASR (automation speech recognition) (701) or STT (speech-to-text) (702) and convert it into a form (e.g., text) that can be processed by the electronic device (101).
[0144] The LLM (large language model) (703) is a natural language processing learning model trained based on a large amount of text data, and can understand and generate sentences. The LLM (703) can process converted text, and for example, can perform functions according to a user's request, answer a user's question, or interpret (or translate) the user's speech into another language. The processor (120) can use the LLM (703) to identify a user's voice from an input signal, translate the text converted from the user's voice, and, if necessary, correct the translated text. For example, the LLM (703) may include a voice recognition learning model trained with the user's voice. The voice recognition learning model is trained with user voice data (user voice) acquired by the microphone (150) of the electronic device (101) according to the use of the electronic device (101) (e.g., a call), and can identify whether a voice input is a user's voice. The electronic device (101) can continuously collect user voice input through the microphone (150) as learning data. The electronic device (101) or a server (not shown) can retrain a voice recognition learning model based on the user voice collected through the microphone (150).
[0145] The processor (120) can generate translation information by translating text converted from a voice input into a target language using the LLM (703). In one embodiment, the LLM (703) can support translation for a limited number of input languages or target languages. The LLM (703) can include a separate learning model for translation for each language. The electronic device (101) can train the learning model for translation using user voice data to support personalized translation by taking into account the user's speech habits and frequently used words.
[0146] The processor (120) can translate data converted into text into a target language using LLM (703) or translation (704).
[0147] The processor (120) can correct translated information using LLM (703) or correction (705). The processor (120) can correct the translation information using a personalized correction learning model (705) that takes into account the user's speech patterns. For example, the processor (120) can correct the translation information using the user's speech patterns.
[0148] The processor (120) can convert the translated information of the voice input into an audio signal using TTS (706). The electronic device (101) can transmit the audio signal of the translated voice input to the wearable electronic device (201), so that the audio signal can be directly output through the speaker of the wearable electronic device (201). The electronic device (101) and the wearable electronic device (201) can interpret and process the user's speech and the other party's speech in real time during a conversation between the user and the other party.
[0149] The processor (120) can convert the interpretation result into an audio signal based on the voice of the user or the other party. The processor (120) can utilize a machine learning model (hereinafter referred to as a voice conversion learning model) trained to input a specific person's voice and text and output an audio signal that utters the text in the specific person's voice.
[0150] In one embodiment, the machine learning model may include speech-to-speech translation (S2ST) (not shown). S2ST includes an ASR (701) and an LLM (703) that convert an input speech signal into text to quickly process interpretation, and can translate the input speech signal and output a result in the form of a speech signal. According to one embodiment, a wearable electronic device (201) may use S2ST to output a translated speech signal from a user's voice input. In one embodiment, S2ST may include a lightweight LLM (703).
[0151] An electronic device (101) according to one embodiment can convert an input voice signal into text using ASR (701), analyze a text sentence using LLM (703), and provide a translation appropriate to the situation.
[0152] FIG. 8 is a flowchart illustrating an operation of an electronic device (101) and a wearable electronic device (201) performing simultaneous interpretation according to one embodiment of the present disclosure.
[0153] According to one embodiment, the electronic device (101) may receive the interpretation result interpreted by the wearable electronic device (201) and output it to the display (160). In order to process interpretation in real time in response to a user's speech, the wearable electronic device (201) may interpret an incomplete user speech at regular intervals.
[0154] Referring to FIG. 8, the entire sentence spoken by speaker A is "I go to school after work." According to an embodiment, a wearable electronic device (201) may input (801) an utterance (e.g., "I go") of speaker A among the entire sentences, first interpret the utterance for "I go to" (802), and transmit the interpretation result ("I go") to an electronic device (101) connected to the wearable electronic device (201) (803). The electronic device (101) may output the interpretation result received from the wearable electronic device (201) to a translation display screen (8041) through a display (160) (804).
[0155] According to an embodiment, a wearable electronic device (201) may receive "school after work" uttered by speaker A following "I go to" (805). The wearable electronic device (201) may later interpret the remaining utterance of "school after work" among the entire sentence (806), and transmit the interpretation result ("I go to school after work") to the electronic device (101) by combining it with the previously interpreted result (e.g., "I go") (807). The electronic device (101) may output the interpretation result received from the wearable electronic device (201) to a translation display screen (8081) via the display (160) (808).
[0156] Referring to FIG. 8, an electronic device (101) according to one embodiment can output the translation information received from a wearable electronic device (201) as is through a display (160) without any separate operation. An electronic device (101) according to one embodiment can only function as an output device that outputs the translation information translated from speaker A's speech as is.
[0157] FIG. 9 is an example of translation information data according to one embodiment of the present disclosure.
[0158] According to an embodiment, the electronic device (101) or the wearable electronic device (201) may use index information to transmit translation information obtained by translating the user's utterance in real time to the wearable electronic device (201) or the electronic device (101). In various embodiments, in order to output the user's utterance by translating it in real time, the electronic device (101) or the wearable electronic device (201) may first translate and transmit the utterance for some sentences before the sentence corresponding to the user's utterance is completed, and then transmit the interpretation for the remaining sentences later. Depending on the input language or the target language, the word order of the input language at the time of the user's utterance may be different from the word order of the target language in the translated text. If a part of the entire sentence is translated first, and the translation result for the part of the sentence translated first is different from the translation result for the sentence in which the utterance is completed, the translation result that has already been output may need to be modified. If the word order of the input language being translated differs from the target language, the structure or order of the sentence may continue to change until the translation of a single sentence is completed. For example, on the translation screen displayed on the display (160) of the electronic device (101), the translated text may display parts of the sentence in chronological order, delete parts of the displayed sentence, and then display the entire sentence again.
[0159] For example, in the previously described Fig. 8, the sentence "I go to" uttered in English has a different subject / predicate position in the sentence structure when translated into Korean than in the English sentence structure. Therefore, when the translation of the entire sentence is completed, the results of translating "I go to" and "school after work" separately are different from the results of translating the entire sentence "I go to school after work." When the electronic device (101) sequentially outputs the translation information received from the wearable electronic device (201) to the display (160), the translation may not be smooth. For example, when the electronic device (101) receives "I go" and outputs it to the display (160), and then receives "I go to school after work" and outputs it to the display (160), "I go. To school after work" may be displayed on the display (160) screen. To address this issue, according to one embodiment, the wearable electronic device (201) may first transmit translation information for a portion of a sentence ("I am going") (901), and then, upon completion of the sentence, retransmit translation information for the entire sentence ("I am going to school after work") back to the electronic device (101) including the portion of the previously translated sentence (902). However, some of the translation information may be transmitted redundantly.
[0160] According to one embodiment, a wearable electronic device (201) may transmit an interpretation result (e.g., translation information) for a voice input received in real time to the electronic device (101) using index information. The interpretation result may be generated for an entire sentence or a portion of a sentence. In one embodiment, the translation information may be stored in the form of an index and data for each word (or phrase), and may include a flag indicating whether the sentence is a completed utterance (903, 904). The index may indicate the order in which the word (or phrase) is located in the sentence structure. For example, data with an index of 100 may be located before data with an index of 300 in the sentence structure.
[0161] According to one embodiment, a wearable electronic device (201) may first transmit translation information (903) for a portion of a sentence, and then, when the sentence is completed, transmit only translation information for the remaining sentence to the electronic device (101) (904).
[0162] According to an embodiment, the electronic device (101) may output translation information according to the word order structure of the target language by referring to the flags and indices included in the translation information received from the wearable electronic device (201). For example, the electronic device (101) may determine that the first received translation information (903) corresponds to an incomplete sentence and the later received translation information (904) corresponds to a completed sentence, and may combine the two pieces of translation information (903, 904) to output them as a single sentence. In various embodiments, the utterance for the entire sentence may be divided into multiple pieces and interpreted sequentially according to real-time translation. According to an embodiment, the electronic device (101) may combine the two pieces of translation information (903, 904) and rearrange the data in the order of 100, 200, 300, and 1000 according to the indices, and then output them to the display (160) or the speaker (155).
[0163] FIG. 10 is an example of user voice input according to one embodiment of the present disclosure.
[0164] According to one embodiment, an electronic device (101) or a wearable electronic device (201) may receive a continuous analog voice signal for a user's speech and use voice activity detection (VAD) to identify the end of the speech or the completion of a sentence.
[0165] According to one embodiment, a wearable electronic device (201) may process (e.g., interpret) a voice signal input in units of windows (e.g., 500 ms) for continuous analog voice input. According to one embodiment, a wearable electronic device (201) may determine that speech has ended by storing a VAD flag as 0 when there is no sound for a specific period of time (e.g., 300 ms) for an input voice signal.
[0166] For example, referring to FIG. 10, a wearable electronic device (201) according to an embodiment may process interpretation of a voice signal for each window, and determine the end of a user's speech by checking whether a silent section (e.g., 1001 or 1002) lasts for a specific period of time. The wearable electronic device (201) according to an embodiment may determine that a first silent section (1001) lasts within a specific, predetermined period of time (e.g., a threshold value preset according to the user's language habits), and thus determine that the user is speaking. The wearable electronic device (201) according to an embodiment may determine that a second silent section (1002) lasts beyond a specific, predetermined period of time, and thus determine that the speech has ended. The wearable electronic device (201) according to an embodiment may determine that the speech for the first sentence has ended based on a silent section (1002) with a flag of 0 using VAD. The wearable electronic device (201) can then process the input voice signal as a new sentence, a second sentence.
[0167] According to an embodiment, a wearable electronic device (201) may transmit a text translated using a translation learning model (e.g., S2ST) to the electronic device (101). The wearable electronic device (201) according to an embodiment may store an original text for which ASR (701) is performed on a user voice input while transmitting the text translated using S2ST to the electronic device (101). The wearable electronic device (201) may transmit the original text accumulated in the wearable electronic device (201) to the electronic device (101) at the time when the utterance is terminated using VAD. The electronic device (101) may receive a flag for the termination of the user utterance and the original text from the wearable electronic device (201), perform a translation on the original text using a machine learning model (e.g., a correction learning model) of the electronic device (101), and display the translated text by replacing the sentences displayed in real time.
[0168] The wearable electronic device (201) can interpret continuously by cutting it into a certain unit (e.g., 500 ms) while receiving continuous utterances of speaker A. For example, the wearable electronic device (201) can first interpret the input "I go to," store the translation information and the original text, and then transmit the interpretation result "I am going" to the electronic device (101) and display it on the screen. The wearable electronic device (201) can interpret the continuously uttered "school after work," and then transmit the interpretation result "after work, to school" to the electronic device (101). The wearable electronic device (201) can use the VAD to confirm the end of the utterance after the utterance of "school after work," and transmit the translation result and the original text together to the electronic device (101). The electronic device (101) can rearrange the translation results received from the wearable electronic device (201) into complete sentences using index information and then display them again. Alternatively, the electronic device (101) can correct the translation results using the received original text and then display them on the screen.
[0169] FIG. 11 illustrates a data flow according to a correction translation operation of an electronic device (101) and a wearable electronic device (201) according to an embodiment of the present disclosure.
[0170] An electronic device (101) and a wearable electronic device (201) according to an embodiment may interpret a conversation between two people in real time and output the interpretation result through a display or speaker. To ensure smooth conversation, the interpretation result output in real time may differ from a complete sentence translated from the original text. An electronic device (101) or a wearable electronic device (201) according to an embodiment may perform an interpretation of an input voice signal, display the interpretation result (e.g., translation information) in real time, perform a translation of the original text, and then display a complete sentence by replacing the sentence displayed in real time.
[0171] For example, referring to FIG. 11, a wearable electronic device (201) can acquire a user voice input (1110) spoken in a first language through a microphone (211), and transmit translated information (text) and original data interpreted using a machine learning model to the electronic device (101) through a transceiver (240). The electronic device (101) can output the translated information text received from the wearable electronic device (201) in real time through a display (160).
[0172] In response to determining that a user has finished speaking, the wearable electronic device (201) may transmit translation information and the original text to the electronic device (101) along with an end-of-speech flag. The electronic device (101) may verify the end-of-speech flag, translate the original text using a machine learning model, and correct the translation information. The electronic device (101) may display the corrected sentence as a replacement for the real-time displayed interpretation result translation information.
[0173] According to one embodiment, the processor (120) of the electronic device (101) can perform translation or proofreading of the original text using a machine learning model. If a translated sentence is incomplete or short, the machine learning model can infer the user's intent and correct it to a complete sentence. For example, if the user utters a short sentence and the translation is mistranslated contrary to the user's intent, the machine learning model can correct the translated sentence to a complete sentence by considering the previous sentence. Alternatively, if the user utters a sentence with a different word order, the translated sentence can be corrected by correcting it to the correct word order. Furthermore, the machine learning model can smooth the translation by considering the relationship between the user and the other party and the current conversational context. For example, real-time interpretation translates the original text as is, so it can be a general translation depending on the initial settings (e.g., polite expressions). In one embodiment, the electronic device (101) can perform a personalized translation using the machine learning model. For example, if the relationship between the user and the other party is close, the electronic device (101) can perform the translation by utilizing friendly expressions. If the user is conversing with a child, the electronic device (101) can perform translation using expressions related to children. In various embodiments, a machine learning model can be trained to perform personalized translation based on the user's speech patterns and frequently used words.
[0174] FIG. 12 is a flowchart illustrating an operation of an electronic device (101) and a wearable electronic device (201) to correct translation according to an embodiment of the present disclosure.
[0175] According to one embodiment, a wearable electronic device (201) may receive an utterance ("I go to") from speaker A (1201). The wearable electronic device (201) may interpret the input user speech into a target language and generate translation information ("I go") (1202). The wearable electronic device (201) may transmit the translation information ("I go") and the original text ("I go to") together to the electronic device (101) (1203).
[0176] The electronic device (101) checks the received translation information and the flag for the original text, and if it is not the end of speech (flag 0), it can output the interpretation result of speaker A as it is on the display (160) screen (12041) (1204).
[0177] The wearable electronic device (201) can continuously receive speaker A's speech ("school after work") (1205). The wearable electronic device (201) can interpret the input user speech into a target language and generate translation information ("school after work") (1206). The wearable electronic device (201) can transmit the translation information ("school after work") and the original text ("school after work") together to the electronic device (101) (1207).
[0178] The electronic device (101) checks the received translation information and the flag for the original text, and if it corresponds to the end of the speech (flag 1), it can output the interpretation result of speaker A (“I’m going to school after work”) as it is on the display (160) screen (12081) (1208). In response to the end of the speech, the electronic device (101) can translate the original text to correct the translation information. Referring to FIG. 12, the electronic device (101) can generate the corrected sentence “I’m going to school after work.” If the translation information already outputted is different from the corrected sentence, the electronic device (101) can delete the translation information from the display (160) screen (12091). After deleting the translation information already outputted, the electronic device (101) can newly output the corrected sentence on the display (160) screen (12092) (1209).
[0179] FIG. 13 is a flowchart for explaining an operation of an electronic device (101) and a wearable electronic device (201) to output an interpretation result according to an embodiment of the present disclosure.
[0180] According to an embodiment, an electronic device (101) and a wearable electronic device (201) can output translation information translated from a real-time conversation to a display or a speaker. The electronic device (101) or the wearable electronic device (201) can process interpretation for continuous analog voice inputs in regular units, and can output translation information for a portion of a sentence interpreted before the user's speech ends through the display. When the user's speech ends, a translation of the entire sentence can be performed in consideration of the original text, thereby modifying (or correcting) at least a portion of the translation information displayed in real time. The translation information displayed on the display can be re-displayed to indicate the revised translation result, but if the audio signal is output as is during the translation process, it may be difficult for the user to understand the entire sentence.
[0181] According to one embodiment, the electronic device (101) and the wearable electronic device (201) can display and modify in real time each unit that processes voice input through a display. In response to the termination of continuous voice input utterance, the electronic device (101) and the wearable electronic device (201) can output information translating a complete sentence through a speaker.
[0182] Referring to FIG. 13, the electronic device (101) can receive and interpret the first voice input for speaker B's speech in chronological order (1301).
[0183] The electronic device (101) can output the interpretation result for the first voice input of speaker B on the display screen (1302). The electronic device (101) can check whether speaker B has finished speaking. The electronic device (101) can determine that the speech has not ended after the first voice input. The electronic device (101) can output the interpretation result on the display screen before the speech ends.
[0184] After confirming speaker B's speech through the display (160) of the electronic device (101), speaker A can begin speaking. The wearable electronic device (201) can receive a second voice input for speaker A's speech and interpret it (1303).
[0185] The electronic device (101) can receive and interpret a third voice input for the subsequent utterances of speaker B (1304). The electronic device (101) can confirm the end of utterance for the input voice signal. The electronic device (101) can determine the end of utterance using the VAD after the third voice utterance of speaker B.
[0186] The wearable electronic device (201) can transmit the interpretation result for the second voice input to the electronic device (101) (1305).
[0187] The electronic device (101) can display the interpretation results for the second voice input of speaker A and the interpretation results for the third voice input of speaker B together on the display screen (1306). Regardless of the interpretation and display output operations for the third voice input of speaker B, the electronic device (101) can display the interpretation results for the second voice input of speaker A received from the wearable electronic device (201) as is on the display screen and speaker.
[0188] The electronic device (101) can output the interpretation result for the second voice input of speaker A through the speaker (155) (1307).
[0189] In response to the end of speaker B's speech, the electronic device (101) can transmit the interpretation results for the first voice input and the third voice input of speaker B as audio signals to the wearable electronic device (201) (1308).
[0190] The wearable electronic device (201) can output the interpretation results for the first voice input and the third voice input of the received speaker B together through the speaker (212).
[0191] FIGS. 14a and 14b are examples of a real-time interpretation result display screen of an electronic device (101) according to one embodiment of the present disclosure.
[0192] According to one embodiment, an electronic device (101) and a wearable electronic device (201) can each interpret a real-time conversation between two people and output it together through the display (160) of the electronic device (101). Even before the user's speech ends, the electronic device (101) and the wearable electronic device (201) can process user voice inputs received continuously and display them on the display (160). The electronic device (101) and the wearable electronic device (201) can re-edit or delete the translation information displayed on the display (160) and display it again.
[0193] According to an embodiment, an electronic device (101) may display translation information (1411, 1412, 1413, 1414) translated from a first user's utterance on a first part (1410) of a display (160), and may display translation information (1421, 1422, 1423, 1424) translated from a second user's utterance on a second part (1420) of the display (160). The electronic device (101) may split the screen of the display (160) and display the translation information in opposite directions on the split screen so that the first user and the second user, who are conversing facing each other, can conveniently check the screens displaying the translation information. For example, the electronic device (101) may display translation information translated from a first user's utterance in the +y direction on the first part (1410), and may display translation information translated from a second user's utterance in the -y direction on the second part (1420).
[0194] Referring to Figures 14a and 14b, this is an example of displaying translation information for a conversation spoken by a first user, starting with an utterance by a second user.
[0195] According to an embodiment, an electronic device (101) may receive voice inputs by consecutive utterances of a second user, perform interpretation in regular units, and output translation information. Before the second user's utterances end, the electronic device (101) may output translation information for the received voice inputs on the second part (1420) of the display (160) at each of time points t1, t3, t5, and t7. The electronic device (101) may translate continuously received utterances of the second user while adding utterance content according to the translation process, and output the translation result. For example, the electronic device (101) may output the beginning of the sentence "와인도" at time point t1 on the second part (1420) of the display (160), and may additionally output the underlined part of "I wish I could have wine too" at time point t3, following the part output at time point t1. Furthermore, the electronic device (101) can additionally output the underlined part of “I wish I could have some wine with you, and the view was nice” at time t5 in response to real-time interpretation processing, and can additionally output the underlined part of “I wish I could have some wine with you, and the view was nice” at time T7.
[0196] The first user and the second user can check the real-time translation results of each other's speech through the display (160) of the electronic device (101). For example, before the second user finishes speaking, the electronic device (101) can display the interpretation results for the second user's speech at time t1. The first user can check some of the translation results before the second user finishes speaking and then begin speaking.
[0197] According to an embodiment, a wearable electronic device (201) may receive voice inputs by continuous utterances of a first user, perform interpretation in regular units, and output translation information. Before the first user's utterances end, the wearable electronic device (201) may output translation information for the received voices on the first part (1410) of the display (160) at time points t2, t4, and t6. The wearable electronic device (201) may translate continuously received utterances of the first user while changing the utterance content according to the translation process and output the translation result. For example, the electronic device (101) may output the beginning of the sentence "Estory cerca" on the first part (1410) of the display (160) at time point t2, delete the previous sentence ("Estory cerca") at time point t4, and then newly output the sentence "Fui a un delicioso restaurant de pasta cercano". The electronic device (101) can output, at time t6, a portion of the previous sentence (“Fui a un” and “cercano”), and an added portion of “Conozco undelicioso restaurant de pastacerca.” together with the underlined portion of the previous sentence (“delicioso restaurant de pasta”).
[0198] Since the electronic device (101) and the wearable electronic device (201) translate and output in real time at regular intervals even before the first user or the second user finishes speaking, the translated information for the speaking can be modified or added until the speaking ends. Since the electronic device (101) and the wearable electronic device (201) simultaneously process the speaking of the first user and the second user during a real-time conversation, the translation results can be simultaneously output even when the users' speaking overlap. The screen (e.g., the first part (1410) or the second part (1420)) that displays the translation results of the display (160) of the electronic device (101) can have text information corresponding to the translation results added or modified at short time intervals.
[0199] FIG. 15 is an example of a multi-party simultaneous interpretation operation of an electronic device (101) and a plurality of wearable electronic devices (201) according to one embodiment of the present disclosure.
[0200] According to an embodiment, an electronic device (101) (e.g., electronic device (1510)) may perform multi-party interpretation using a plurality of wearable electronic devices (201). The electronic device (101) may be paired with a wearable electronic device (201) of a user of the electronic device (101) (hereinafter, a first wearable electronic device (1501)).
[0201] A wearable electronic device (1501) according to one embodiment can perform interpretation of a user's speech while the wearable electronic device (1501) is worn by the user.
[0202] The electronic device (101) can be connected to wearable electronic devices (201) of other users (e.g., 1502, 1503, 1504, 1505, 1506, 1507). The electronic device (101) can receive a voice input signal by a user's speech from a wearable electronic device (201) of other users in which the user is speaking (e.g., 1503, 1506). Each of the wearable electronic devices (201) of other users can be connected to the electronic device (101) based on short-range wireless communication, and when a voice signal by a user's speech is input through each microphone while the wearable electronic device (201) is worn by each other user, the voice signal acquired through the microphone can be transmitted to the electronic device (101).
[0203] The electronic device (101) can interpret the received voice input signal and output the translation information along with the identification information (e.g., speaker name) of the wearable electronic device (201) that transmitted the voice input signal to the display (1520).
[0204] The display (1520) may include a large screen so that multiple other users can view it. The display (1520) may translate and output the speech of each user on a portion of the screen corresponding to the respective user's location. For example, the electronic device (101) may use beamforming to determine the location of each user at the time each speaker speaks using multiple microphones included in the electronic device (101), and may translate and display the content of the user's speech on a portion of the display screen corresponding to the location of each user.
[0205] FIGS. 16a and 16b are examples of a multi-party simultaneous interpretation result display screen of an electronic device (101) according to one embodiment of the present disclosure.
[0206] As in the embodiment of FIG. 15, an electronic device (101) (e.g., electronic device (1510)) can perform multi-party simultaneous interpretation using a plurality of wearable electronic devices (201) and output the interpretation result on one screen.
[0207] Referring to FIGS. 16A and 16B , the electronic device (101) may output a simultaneous interpretation result screen through the display (160) of the electronic device (101) or a separate display. The simultaneous interpretation result screen may include a first portion (1610, 1611) that displays multiple speakers and a second portion (1620, 1621) that displays the utterance content and translation information according to the speaker. The second portion (1620, 1621) of the simultaneous interpretation result screen may display the translation information of the speaker in a downward direction according to time. In the embodiment of FIG. 16 , the left simultaneous interpretation result screen (1610, 1620) is a screen displayed at a first time, and the right simultaneous interpretation result screen (1611, 1621) is a screen displayed at a second time point that is a time point after the first time point as the conversation progresses.
[0208] The electronic device (101) may consider a wearable electronic device (201) connected to the electronic device (101) via wireless communication as participating in a multi-party conversation and display the result on the first part (1610) of the simultaneous interpretation result screen. For example, at the first time, two speakers (speaker 1, speaker 2) may participate in the multi-party conversation. In response to specifying the user of the wearable electronic device (201) participating in the multi-party conversation, the electronic device (101) may display the specified user on the first part (1611) of the simultaneous interpretation result screen.
[0209] An electronic device (101) may receive a voice input for a speech of a first speaker from the first speaker's wearable electronic device (201), and display translation information obtained by interpreting the received voice input together with the original text and the translation information on a second portion (1620) of a display (1601). The original text displayed on the second portion (1620) may be a text converted from an input voice signal corresponding to the speaker's native language, and the translation information may be a text translated into a target language. The original text may not be an interpretation or translation, but may be a text signal converted from a voice signal.
[0210] The electronic device (101) can process multiple input languages (e.g., Korean, English, Spanish), but can interpret or translate into a single target language (e.g., English) without considering the individual languages of the multiple participants. The electronic device (101) can output translation information by interpreting an input voice signal on a display, and simultaneously generate an audio signal for the translation information in the target language and output it through a speaker. In various embodiments, the electronic device (101) can interpret or translate into a single target language, and then convert it back into the native language of each participant and output it.
[0211] At the second point in time, the second part (1620) can designate a divided area for each speaker participating in the multi-party conversation, and translation information for a specific speaker's utterance can be displayed in the specific speaker's conversation display area. In a multi-party conversation, if the utterances of multiple speakers overlap temporally, the timelines of the translation information displayed in the second part (1620) can overlap. For example, in the simultaneous interpretation result screen (1621) at the second point in time, Kim Cheol-su's utterance (1603) begins before Danyell Mercer's utterance (1602) ends, and thus the timelines of the translation information (1602 and 1603) displayed for the utterances can be displayed overlappingly. In addition, in the simultaneous interpretation result screen (1621) at the second point in time, a new speaker (speaker3)'s utterance can be input before Kim Cheol-su's utterance (1603) ends, and translation information (1604) for this can be displayed.
[0212] FIG. 17 is a flowchart illustrating a method for an electronic device (101) to perform simultaneous interpretation using a wearable electronic device (201) according to one embodiment of the present disclosure.
[0213] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0214] According to one embodiment, steps S1701 to S1707 may be understood to be performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (201) of FIG. 1) or a processor (e.g., processor (220) of FIG. 2) of a wearable electronic device (e.g., wearable electronic device (201) of FIG. 1).
[0215] According to one embodiment, the electronic device (101) can simultaneously interpret a conversation between a user wearing the wearable electronic device (201) and the other party using the wearable electronic device (201), and output the interpretation result to a speaker or display. The electronic device (101) and the wearable electronic device (201) may be paired with each other.
[0216] According to one embodiment, in step S1701, the electronic device (101) can receive a voice signal through a microphone (150) of the electronic device (101) or a microphone (211) of the wearable electronic device (201).
[0217] According to one embodiment, in step S1702, the electronic device (101) may determine whether the input voice signal is a single voice. According to one embodiment, the electronic device (101) may determine whether the voice signal is composed of a single voice using a machine learning model. For example, the electronic device (101) may determine whether the voice of one speaker has ended and the voice of the next speaker has been input. Determining the end of the voice in the input signal may be the same as the embodiment of FIG. 10 described above. If the voice of another speaker begins before the voice of one speaker ends, the electronic device (101) may determine that the input voice signal includes multiple voices. Alternatively, the electronic device (101) may determine that the input voice signal includes multiple voices even if another voice signal is input before the interpretation processing for the input voice signal is completed. In this case as well, the electronic device (101) may consider it as multiple voices because it must perform interpretation on the multiple voice signals.
[0218] According to one embodiment, the electronic device (101) can process interpretation of the voice signal in response to determining that the input voice signal is a single voice (step S1703).
[0219] According to one embodiment, in step S1704, the electronic device (101) may output translation information that interprets the input voice signal to a display (160) or a speaker (e.g., a speaker (155) of the electronic device (101) or a speaker (212) of the wearable electronic device (201)).
[0220] According to one embodiment, in step S1705, in response to determining that the input voice signal is a plurality of voices, the electronic device (101) may extract a first voice from among the plurality of voices included in the voice signal. The electronic device (101) may identify the user voice from among the plurality of voices based on the user's voice stored in the memory (130) and extract the user voice as the first voice.
[0221] According to one embodiment, in step S1706, the electronic device (101) may transmit a request for interpretation processing for a first voice extracted from among a plurality of voices to an external electronic device (e.g., a wearable electronic device (201)). In response to the request of the electronic device (101), the wearable electronic device (201) may transmit translation information obtained by processing the interpretation for the first voice to the electronic device (101). In various embodiments, the electronic device (101) may be equipped with a separate processor (e.g., an NPU) for simultaneous interpretation, and when conversations between speakers overlap, the separate processor may process simultaneous interpretation for the corresponding voice signals. In this case, simultaneous interpretation for a plurality of speakers may be processed in parallel within the electronic device (101).
[0222] According to one embodiment, in step S1707, the electronic device (101) may receive translation information processed for interpretation of the first voice from the wearable electronic device (201) and output it to a display (160) or a speaker (e.g., a speaker (155) of the electronic device (101) or a speaker (212) of the wearable electronic device (201)).
[0223] According to one embodiment of the present disclosure, a wearable electronic device (201) comprises: a speaker (212); a microphone (211); a transceiver (240); a memory (230); And at least one processor (220) including a processing circuit, wherein the memory (230) is configured to cause the wearable electronic device (201) to: receive a first user voice input through the microphone (211), and generate first translation information in which the received first user voice input is translated using a machine learning model (231) stored in the memory (230) based on the received first user voice input, wherein the machine learning model (231) includes a first machine learning model trained to output translated translation information based on a voice input, and transmit the first translation information to an external electronic device (e.g., electronic device (101)) connected to the wearable electronic device (201) through the transceiver (240), and output second translation information in which a second user voice input is translated from the external electronic device (101). It is possible to store instructions that cause the second translation information to be received through the transceiver (240) and output through the speaker (212).
[0224] According to one embodiment, the machine learning model (231) may include a second machine learning model trained to identify the user's voice based on other voice inputs received through the microphone (211).
[0225] According to one embodiment, the memory (230) may store instructions that, when individually or collectively executed by the at least one processor (220), cause the wearable electronic device to: obtain user voice information based on a portion of the first user voice input using the second machine learning model, based at least in part on a determination that a portion of the first user voice input corresponds to a user voice; and translate the obtained user voice information using the first machine learning model to generate the first translation information.
[0226] According to one embodiment, the memory (230) may store instructions that, when individually or collectively executed by the at least one processor (220), cause the wearable electronic device (201) to: transmit user identification information corresponding to the identified user's voice and the first translation information to the external electronic device (101).
[0227] According to one embodiment, the memory (230) may store instructions that, when individually or collectively executed by the at least one processor (220), cause the wearable electronic device (201) to: convert the first user voice input into text using the machine learning model (231), and perform a translation on the converted text to generate the first translation information, as at least part of generating the first translation information.
[0228] According to one embodiment, the memory (230) may store instructions that, when individually or collectively executed by the at least one processor (220), cause the wearable electronic device (201) to: receive update information from the external electronic device (101) and update the machine learning model (231) based on the received update information.
[0229] According to one embodiment, the first translation information includes at least one of translated data, an index of the data, or a flag for the end of utterance, wherein the index of the data indicates the order in which the translated data is positioned in a sentence structure, and the flag may indicate whether the utterance has ended as identified by voice activity detection (VAD).
[0230] According to one embodiment, the memory (230) may store instructions that, when individually or collectively executed by the at least one processor (220), cause the wearable electronic device (201) to: detect noise input in a portion of the first user voice input where the user's voice does not exist, remove an input corresponding to the detected noise in the portion of the first user voice input where the user's voice does not exist to obtain user voice information, and translate the obtained user voice information using the first machine learning model to generate the first translation information.
[0231] According to another embodiment of the present disclosure, an electronic device (101) includes a microphone (150); a display (160); a transceiver (190); a memory (130); and at least one processor (120) including a processing circuit, wherein the memory (130) is configured to cause the electronic device (101) to:
[0232] The device may store instructions for receiving a first user voice input through the microphone (150), generating first translation information in which the first user voice input is translated using a machine learning model based on the received first user voice input - the machine learning model including a first machine learning model trained to output translated translation information based on a voice input - transmitting the first translation information to an external electronic device (e.g., a wearable electronic device (201)) connected to the electronic device (101) through the transceiver (190), receiving second translation information in which a second user voice input is translated from the external electronic device (201) through the transceiver (190), and causing the second translation information to be displayed in a first area of the display (160).
[0233] According to one embodiment, the memory (130) may store instructions that, when individually or collectively executed by the at least one processor (220), cause the electronic device (101) to: display the first translation information in a second area of the display (160), and cause the first translation information displayed in the second area to be displayed in an opposite direction to the second translation information displayed in the first area.
[0234] According to one embodiment, the memory (130) may store instructions that, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: display the second translation information including user identification information, and a graphical object corresponding to the user identification information, together with the first translation information, in a second area of the display.
[0235] In one embodiment, the machine learning model may include a second machine learning model trained to identify a user's voice based on other voice inputs received via the microphone (150).
[0236] According to one embodiment, the memory (130) may store instructions that, when individually or wholly executed by the at least one processor (120), cause the electronic device (101) to: obtain user voice information by filtering out a portion corresponding to the other user's voice based on a portion of the first user's voice input using the second machine learning model, based on a determination that at least a portion of the first user's voice input corresponds to another user's voice, and translate the obtained user voice information using the first machine learning model to generate the first translation information.
[0237] According to one embodiment, the memory (130) may store instructions that, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: receive the second translation information, convert the received second translation information into voice information (Text to Speech) using the first machine learning model, and output the converted information through the speaker.
[0238] According to one embodiment, the memory (130) may store instructions that, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: convert the second translation information into voice information using at least one user voice information pre-stored in the electronic device (101) and then output the converted second translation information through the speaker.
[0239] According to one embodiment, the electronic device (101) may further include at least one sensor including a global positioning system (GPS), and the memory (130) may store instructions that, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: generate update information including a portion of the first learning model based on information input from the at least one sensor, and transmit the update information to the external electronic device (201) via the transceiver (190).
[0240] In one embodiment, the update information may include update information associated with a translation support language.
[0241] According to one embodiment, the electronic device (101) further includes at least one camera, and the memory stores instructions that, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: acquire a mouth shape based on a lip image captured through the at least one camera, and identify a user's voice using the machine learning model based on the mouth shape, wherein the machine learning model may include a third machine learning model trained to identify the content of speech based on the mouth shape.
[0242] According to one embodiment, the second translation information includes at least one of translated data, an index of the data, or a flag for the end of utterance, and commands for causing the translated data included in the second translation information to be rearranged and displayed in the first area of the display based on the index of the data may be stored.
[0243] According to another embodiment of the present disclosure, a translation system may include a wearable electronic device (201); and an electronic device (101) connected to the wearable electronic device (201). The wearable electronic device (201) may: receive a first user voice input through a microphone (211) of the wearable electronic device (201), and generate first translation information in which the received first user voice input is translated using a machine learning model (231) of the wearable electronic device (201) based on the received first user voice input, wherein the machine learning model (231) includes a model trained to output translation information based on a voice input, transmit the first translation information to the electronic device (101), receive second translation information in which a second user voice input is translated from the electronic device (101), and output the second translation information through a speaker (212) of the wearable electronic device (201). The electronic device (101) may: receive the second user voice input through the microphone (150) of the electronic device (101), and generate the second translation information using a machine learning model of the electronic device (101) based on the received second user voice input, wherein the machine learning model of the electronic device (101) includes a model trained to output translated information based on a voice input; transmit the second translation information to the wearable electronic device (201), receive the first translation information from the wearable electronic device (201), and display the first translation information on a first area of a display (160) of the electronic device (101).
[0244] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0245] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0246] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In wearable electronic devices, speaker; mike; Transmitter and receiver; memory; and comprising at least one processor comprising a processing circuit; The memory, when executed individually or collectively by the at least one processor, causes the wearable electronic device to: Receive a first user voice input through the above microphone, Generating first translation information translated from the received first user voice input using a machine learning model stored in the memory, wherein the machine learning model includes a first machine learning model trained to output translated translation information based on the voice input; Transmitting the first translation information to an external electronic device connected to the wearable electronic device through the transceiver; Receive second translation information translated from the second user voice input from the external electronic device through the transceiver, and A wearable electronic device storing instructions that cause the second translation information to be output through the speaker.
2. In paragraph 1, A wearable electronic device, wherein the machine learning model comprises a second machine learning model trained to identify a user's voice based on other voice inputs received through the microphone.
3. In paragraph 2, The memory, when individually or collectively executed by the at least one processor, causes the wearable electronic device to: Obtaining user voice information based on a portion of the first user voice input using the second machine learning model, based at least in part on a determination that a portion of the first user voice input corresponds to the user voice; Translating the acquired user voice information using the first machine learning model to generate the first translation information, and A wearable electronic device storing user identification information corresponding to the identified user's voice and commands causing the first translation information to be transmitted to the external electronic device.
4. In paragraph 1, The memory, when individually or collectively executed by the at least one processor, causes the wearable electronic device to: As at least a part of generating the first translation information, converting the first user voice input into text using the machine learning model, and A wearable electronic device storing commands that cause the above-described converted text to be translated to generate the first translation information.
5. In paragraph 1, The memory, when individually or collectively executed by the at least one processor, causes the wearable electronic device to: A wearable electronic device that receives update information from the external electronic device and stores commands that cause the machine learning model to be updated based on the received update information.
6. In paragraph 1, The first translation information includes at least one of translated data, an index of the data, or a flag for the end of utterance, The index of the above data indicates the order in which the translated data is positioned in the sentence structure, A wearable electronic device, wherein the above flag indicates whether speech has ended as identified by voice activity detection (VAD).
7. In paragraph 1, The memory, when individually or collectively executed by the at least one processor, causes the wearable electronic device to: Detecting noise input from a part of the first user voice input where the user's voice does not exist, Obtain user voice information by removing the input corresponding to the detected noise from the part where the user's voice does not exist among the first user voice input, A wearable electronic device storing commands that cause the acquired user voice information to be translated using the first machine learning model to generate the first translation information.
8. In electronic devices, mike; display; Transmitter and receiver; memory; and comprising at least one processor comprising a processing circuit; The memory, when executed individually or collectively by the at least one processor, causes the electronic device to: Receive a first user voice input through the above microphone, Generating first translation information translated from the received first user voice input using a machine learning model, wherein the machine learning model includes a first machine learning model trained to output translated translation information based on the voice input; Transmitting the first translation information to an external electronic device connected to the electronic device through the transceiver; Receive second translation information translated from the second user voice input from the external electronic device through the transceiver, and An electronic device storing instructions that cause the second translation information to be displayed in the first area of the display.
9. In paragraph 8, The memory, when individually or collectively executed by the at least one processor, causes the electronic device to: Display the above first translation information in the second area of the display, An electronic device that stores commands that cause the first translation information displayed in the second area to be displayed in a direction opposite to the second translation information displayed in the first area.
10. In paragraph 9, The memory, when individually or collectively executed by the at least one processor, causes the electronic device to: The second translation information includes user identification information, An electronic device storing commands that cause a graphical object corresponding to the user identification information to be displayed in a second area of the display together with the first translation information.
11. In paragraph 8, An electronic device, wherein the machine learning model comprises a second machine learning model trained to identify a user's voice based on other voice inputs received through the microphone.
12. In paragraph 11, The memory, when individually or collectively executed by the at least one processor, causes the electronic device to: Obtaining user voice information by filtering out the portion corresponding to the other user's voice based on a portion of the first user's voice input using the second machine learning model, based on a judgment that at least a portion of the first user's voice input corresponds to another user's voice; An electronic device storing commands that cause the acquired user voice information to be translated using the first machine learning model to generate the first translation information.
13. In paragraph 8, further comprising at least one sensor including GPS, The memory, when individually or collectively executed by the at least one processor, causes the electronic device to: An electronic device storing commands that cause the transceiver to control the transmitter and receiver to generate update information including a portion of the first machine learning model based on information input from at least one sensor and transmit the update information to the external electronic device.
14. In paragraph 8, Including at least one more camera, The memory, when individually or collectively executed by the at least one processor, causes the electronic device to: Obtaining a mouth shape based on a lip image captured through at least one camera, Store commands that cause the machine learning model to identify the user's voice based on the above lip shape, An electronic device, wherein the machine learning model includes a third learning model trained to identify the content of speech based on the lip shape.
15. In the translation system, wearable electronic devices; and An electronic device connected to the wearable electronic device, The above wearable electronic device: Receive a first user voice input through a microphone of the wearable electronic device, Using a machine learning model of the wearable electronic device based on the received first user voice input, first translation information is generated by translating the received first user voice input, and the machine learning model includes a model trained to output translation information based on the voice input. Transmitting the above first translation information to the electronic device, Receive second translation information translated from the second user voice input from the electronic device, The above second translation information is output through the speaker of the wearable electronic device, The above electronic device: Receive the second user voice input through the microphone of the electronic device, Generating the second translation information using a machine learning model of the electronic device based on the second user voice input received above, wherein the machine learning model of the electronic device includes a model trained to output translated information based on the voice input; Transmitting the second translation information to the wearable electronic device, Receiving the first translation information from the wearable electronic device, A translation system that displays the first translation information in a first area of a display of the electronic device.
Citation Information
Patent Citations
Apparatus and method for language expression using context and intent awareness
KR1020100126004A
Device and method for voice translation
KR1020170112713A
Organic light emitting display device
KR1020240013888A
Wearable device and translation system
US20160267075A1
Headphones for a real time natural language machine interpretation
US20200125646A1