Electronic device and method for interpretation during call, and storage medium
The electronic device addresses the challenge of language barriers in calls by converting speech to text, translating, and outputting translated voice signals, improving communication clarity through simultaneous display and playback of original and translated content.
Patent Information
- Application Number
- PCT/KR2024/021030
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-06
- Filing Date
- 2024-12-24
- Publication Date
- 2025-07-10
AI Technical Summary
There is a need for effective communication between individuals who speak different languages during calls, as existing technologies do not adequately facilitate real-time language interpretation and translation, especially in scenarios where both parties may not share a common language.
An electronic device equipped with a communication interface, display, speaker, and processor that can convert speech to text, translate text between languages, and output translated text or voice signals, allowing for simultaneous display and playback of both original and translated content, with options for muting original speech.
Enables real-time language interpretation during calls by converting speech to text, translating between languages, and outputting translated text or voice signals, enhancing communication clarity and understanding for individuals with different linguistic backgrounds.
Smart Images

Figure KR2024021030_10072025_PF_FP_ABST
Abstract
Description
Electronic device, method and storage medium for interpreting during a call
[0001] Embodiments of this document relate to an electronic device, method and storage medium for interpreting during a call.
[0002] With the advancement of transportation and communication, interaction with foreigners is increasing. To do so, communication using a mutually understandable language is necessary. For example, if people of different nationalities speak a common language (e.g., English or French), they can communicate using that common language. Alternatively, if one person (or a group of people) speaks the other person's language, they can communicate using that language. However, if people of different nationalities are unfamiliar with the common language or the other person's language, they may need the assistance of an interpreter to communicate.
[0003] The above information may be provided solely as background information to aid in understanding the present disclosure. None of the above-described matters are claimed as prior art related to the present disclosure or can be used in determining prior art.
[0004] The electronic device, method and storage medium for interpreting during a call of this document are intended to provide an interpretation function during a call.
[0005] An electronic device according to various embodiments of the present document may include a communication interface, a display, a speaker, at least one processor, and a memory storing instructions executed by the at least one processor. The instructions stored in the memory may be configured to cause the electronic device to, during a call, convert a first language speech of a caller into a first text and display the first text on a script screen of the display. The instructions stored in the memory may be configured to cause the electronic device to translate the first text into a second language to generate a second text, and display the second text together with the first text on the script screen. The instructions stored in the memory may be configured to cause the electronic device to convert the second text into a first text-to-speech (TTS) voice signal. The instructions stored in the memory may be configured to cause the electronic device to, when a first mute option for the caller's speech is turned on, identify the caller's speech and output the first TTS voice signal through the speaker without outputting the caller's speech. The command stored in the memory may be configured to cause the electronic device to output the other party's speech and the first TTS voice signal through the speaker when the first mute option for the other party's speech is off.
[0006] A method performed by an electronic device including a screen according to various embodiments of the present document may include, during a call, converting a speech of a counterpart in a first language into a first text and displaying the first text on a script screen. The method may include translating the first text into a second language to generate a second text, and displaying the second text together with the first text on the script screen. The method may include converting the second text into a first text-to-speech (TTS) voice signal. The method may include, when a first mute option for the counterpart's speech is in an on state, identifying the counterpart's speech and outputting the first TTS voice signal without outputting the counterpart's speech. The method may include, when a first mute option for the counterpart's speech is in an off state, outputting the counterpart's speech and the first TTS voice signal.
[0007] A non-transitory computer-readable storage medium having recorded thereon a program for performing a method performed by an electronic device including a screen according to various embodiments of the present document may include instructions for performing an operation of converting, during a call, a speech of a counterpart in a first language into a first text, and displaying the first text on a script screen. The storage medium may include instructions for performing an operation of translating the first text into a second language to generate a second text, and displaying the second text together with the first text on the script screen. The storage medium may include instructions for performing an operation of converting the second text into a first text-to-speech (TTS) voice signal. The storage medium may include instructions for performing an operation of identifying the speech of the counterpart and outputting the first TTS voice signal without outputting the speech of the counterpart when a first mute option for the speech of the counterpart is turned on. The storage medium may include instructions for performing an operation of outputting the other party's speech and the first TTS voice signal when the first mute option for the other party's speech is off.
[0008] Various embodiments of this document can optionally provide an interpretation function during a call, depending on the user's preference. The effects of this disclosure are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.
[0009] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.
[0010] FIG. 2 is a block diagram illustrating the configuration of an electronic device according to various embodiments.
[0011] FIGS. 3a, 3b, 3c, 3d, 3e, 3f, 3g and 3h are drawings illustrating operations for installing language data according to various embodiments.
[0012] FIGS. 4a, 4b, 4c, 4d, 4e, 4f, 4g, 4h, 4i, and 4j are drawings illustrating an operation of setting the language of the other party according to various embodiments.
[0013] Figure 5a is a flowchart illustrating an interpreting operation according to various embodiments.
[0014] FIG. 5b is a timing diagram illustrating an interpreting operation according to various embodiments.
[0015] FIGS. 6A, 6B, 6C, and 6D are diagrams illustrating operations for executing an interpretation function in a transmission according to various embodiments.
[0016] FIGS. 7a, 7b, 7c, 7d, 7e and 7f are diagrams illustrating operations for executing an interpretation function in a receiver according to various embodiments.
[0017] Figure 8 is a diagram illustrating ignition options according to various embodiments.
[0018] FIGS. 9A, 9B, 9C, and 9D are diagrams illustrating interpretation actions displayed in a dialogue script according to various embodiments.
[0019] FIGS. 10A and 10B are diagrams illustrating a dialogue script displaying utterances and translations according to various embodiments.
[0020] Figure 11 is a flowchart explaining an operation of executing an interpretation function based on information of the other party confirmed through a server according to various embodiments.
[0021] FIG. 12a and FIG. 12b are diagrams illustrating an operation of setting the language of the other party based on the number information of the other party according to various embodiments.
[0022] FIGS. 13a, 13b, and 13c are diagrams illustrating an operation of setting the language of the other party based on additional information according to various embodiments.
[0023] FIGS. 14a, 14b, 14c, 14d, and 14e are diagrams illustrating screens displayed on a user's electronic device when an interpretation function is executed according to various embodiments.
[0024] FIG. 15a and FIG. 15b are drawings explaining a screen displayed on the other party's electronic device when an interpretation function is executed according to various embodiments.
[0025] FIG. 16A is a diagram illustrating a dialogue script screen including setting options according to various embodiments.
[0026] FIG. 16b is a diagram illustrating the volume of speech and / or TTS voice signals according to various embodiments.
[0027] FIG. 17a and FIG. 17b are flowcharts illustrating operations in which a user's electronic device and a counterpart's electronic device perform an interpretation function together according to various embodiments.
[0028] FIG. 18 is a timing diagram illustrating an operation in which a user's electronic device and a counterpart's electronic device perform an interpretation function together according to various embodiments.
[0029] FIGS. 19a, 19b, and 19c are diagrams illustrating operations for modifying text displayed in a dialogue script according to various embodiments.
[0030] FIGS. 20A, 20B, and 20C are diagrams illustrating operations for modifying text displayed in a call history according to various embodiments.
[0031] FIG. 21a and FIG. 21b are diagrams illustrating operations for executing an interpretation function in a conference call according to various embodiments.
[0032] Figure 22 is a flowchart illustrating a method of interpretation during a call according to various embodiments.
[0033] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.
[0034] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0035] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0036] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0037] The memory (130) can store various data used by at least one component (e.g., the processor (120) or the sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., the program (140)) and input data or output data for commands related thereto. The memory (130) can include a volatile memory (132) or a non-volatile memory (134). The non-volatile memory (134) can include at least one internal memory (136) and an external memory (138).
[0038] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0039] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0040] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0041] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0042] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0043] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0044] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0045] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0046] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0047] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0048] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0049] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0050] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0051] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0052] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0053] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0054] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0055] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0056] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0057] FIG. 2 is a block diagram illustrating the configuration of an electronic device according to various embodiments.
[0058] Referring to FIG. 2, an electronic device (200) (e.g., the electronic device (101) of FIG. 1) may include a memory (210), a speaker (220), a display (230), a communication interface (240), and a processor (250).
[0059] The memory (210) (e.g., the memory (130) of FIG. 1) may store data, algorithms, programs, and / or instructions that perform functions of the electronic device (200) (e.g., the electronic device (101) of FIG. 1). Instructions stored in the memory (210) may be loaded into the processor (250) and executed by the processor (250). For example, the memory (210) may store data of a language set to the user's language and / or the other party's language. The language data may include a language pack, and the language pack may be received from a server (e.g., the server (108) of FIG. 1).
[0060] The speaker (220) (e.g., the audio output module (155) of FIG. 1) may output speech and / or text-to-speech (TTS) voice signals. For example, the speech may include a user's speech and / or a counterpart's speech. The TTS voice signal may include a first TTS voice signal that translates the user's speech and is generated based on the translated user's speech, and a second TTS voice signal that translates the counterpart's speech and is generated based on the translated counterpart's speech. Terms such as "first" or "second" in this document may be used simply to distinguish the corresponding component from other corresponding components, and are not intended to limit the corresponding components in any other aspect (e.g., importance or order).
[0061] The display (230) (e.g., the display module (160) of FIG. 1) may display a user interface (UI) and / or a dialogue script related to the interpretation function. For example, the dialogue script may include a first text converted from the user's speech, a second text translated from the user's speech (or the first text), a third text converted from the other party's speech, and / or a fourth text translated from the other party's speech (or the third text).
[0062] The communication interface (240) (e.g., the communication module (190) of FIG. 1) can transmit the user's speech and / or the first TTS voice signal (e.g., the voice signal translated from the user's speech) to the other party's electronic device, and receive the other party's speech and / or the second TTS voice signal (e.g., the voice signal translated from the other party's speech) from the other party's electronic device. As an example, if only the user's electronic device performs the interpretation function, the communication interface (240) can receive only the other party's speech from the other party's electronic device. As an example, if the user's electronic device and the other party's electronic device (e.g., the electronic devices (102, 104) of FIG. 1) perform the interpretation function together, the communication interface (240) can transmit and receive TTS voice signals (e.g., the first TTS voice signal and the second TTS voice signal) or transmit and receive speech signals (e.g., the user's speech and the other party's speech).
[0063] A processor (250) (e.g., processor (120) of FIG. 1) can control each component of the electronic device (200). The electronic device (200) can include one or more processors (250). For example, the processor (250) can correspond to multiple processors that collectively perform multiple functions by dividing them among the processors.
[0064] As an example, the electronic device (200) can execute an interpretation function between electronic devices without a server. When the electronic device (200) executes an interpretation function without a server, the electronic device (200) can translate the user's speech and the other party's speech based on information verifiable in the electronic device (200) and / or information set in the electronic device (200). In addition, the electronic device (200) can transmit the user's speech and / or a first TTS voice signal generated based on the user's speech to the other party's electronic device, and output the other party's speech and / or a second TTS voice signal generated based on the other party's speech.
[0065] As an example, the electronic device (200) can execute an interpretation function through a server. When the electronic device (200) executes the interpretation function through the server, the electronic device (200) can execute the interpretation function based on information confirmed through the server, information confirmed in the electronic device (200), and / or information set in the electronic device (200). If the other party's electronic device supports the interpretation function, the user's electronic device (200) and the other party's electronic device can execute the interpretation function together.
[0066] For example, the user's electronic device (200) can translate the user's speech and generate a first TTS voice signal based on the translated user's speech. The user's electronic device (200) can transmit the user's speech and / or the first TTS voice signal to the other party's electronic device. The other party's electronic device can translate the other party's speech and generate a second TTS voice signal based on the translated other party's speech. The other party's electronic device can transmit the other party's speech and / or the second TTS voice signal to the user's electronic device (200).
[0067] As another example, the user's electronic device (200) can transmit the user's speech to the other party's electronic device. The other party's electronic device can translate the user's speech and generate a first TTS voice signal based on the translated user's speech. The other party's electronic device can output the user's speech and / or the first TTS voice signal. The other party's electronic device can transmit the other party's speech to the user's electronic device (200). The user's electronic device (200) can translate the other party's speech and generate a second TTS voice signal based on the translated other party's speech. The user's electronic device (200) can output the other party's speech and / or the second TTS voice signal. Even if both the user's electronic device (200) and the other party's electronic device support an interpretation function, an interpretation function can be performed by translating all speech and transmitting and outputting the translation result on one electronic device (e.g., the user's electronic device (200) or the other party's electronic device).
[0068] A user's electronic device (200) can interpret the user's speech and the other party's speech. The user's electronic device (200) can generate a text that translates the user's speech and convert the generated text into a voice signal using a TTS method. In addition, the user's electronic device (200) can generate a text that translates the received speech of the other party and convert the generated text into a voice signal using a TTS method. The user's electronic device (200) can output the other party's speech and / or a TTS voice signal that translates the other party's speech, and transmit the user's speech and / or a TTS voice signal that translates the user's speech to the other party's electronic device.
[0069] When the interpretation function is executed, the processor (250) can convert the user's speech into a first text and display the converted first text on the display (230). For example, the processor (250) can convert the user's speech into the first text using STT (speech-to-text) method. The processor (250) can generate a dialogue script and sequentially display the generated text in the dialogue script. For example, the processor (250) can display the converted first text in the dialogue script.
[0070] The processor (250) can generate a second text by translating the user's utterance into a second language. For example, the second language may be a language set as the other party's language. Furthermore, the processor (250) can display the second text generated together with the displayed first text on the display (230). For example, the second text may be displayed in an adjacent area of the first text. The first language may include a language set as the other party's language or a default language. If the other party's language is set, the processor (250) can translate the user's utterance (or the first text) into the other party's language. If the other party's language is not set, the processor (250) can translate the user's utterance into a language set as the default language. For example, the first text may be a text that displays the user's utterance in the user's language, and the second text may be a text that displays the user's utterance by translating it into the other party's language or a default language.
[0071] For example, the other party's language can be set based on information about the other party. For example, the processor (250) can identify the other party's country code, area code, system language information of the other party's electronic device, and / or language information of the calling application from the other party's number. The processor (250) can determine the language based on the identified other party's information and set the determined language as the other party's language. For example, the processor (250) can check a message received by the other party's number and determine the language of the message. The processor (250) can set the determined language of the message as the other party's language. For example, the processor (250) can determine the other party's language from the other party's speech.
[0072] The processor (250) can change the language of the other party during a call from a second language to a third language based on the user's input. If the other party's language changes, the processor (250) can translate (or interpret) the user's speech into the third language. The processor (250) may require the other party's language data (e.g., a language pack) to translate the user's speech into the other party's language. If the other party's language changes to a third language and there is no third language data, the processor (250) can control the communication interface (240) to download third language data from a server, and store and install the downloaded third language data in the memory (210). Once the third language data is installed, the processor (250) can translate the user's speech into the third language to generate a second text. The processor (250) can display the second text in the third language together with the first text displayed in the conversation script.
[0073] The processor (250) can convert the second text into a first TTS voice signal. Then, the processor (250) can transmit the user's speech and / or the first TTS voice signal to the other party's electronic device using the communication interface (240). For example, the electronic device (200) can include a mute option for the user's speech. If the mute option for the user's speech is turned on, the processor (250) can transmit only the converted first TTS voice signal to the other party's electronic device without transmitting the user's speech to the other party's electronic device. For example, if the mute option for the user's speech is turned on, the processor (250) can transmit only the converted first TTS voice signal to the communication interface (240) without transmitting the user's speech to the communication interface (240). The communication interface (240) can transmit only the transmitted first TTS voice signal to the other party's electronic device. If the mute option for the user's speech is turned off, the processor (250) can transmit the user's speech and / or the converted first TTS voice signal to the other party's electronic device. For example, if the mute option for the user's speech is turned off, the processor (250) can transmit the user's speech and / or the converted first TTS voice signal to the communication interface (240). The communication interface (240) can transmit the transmitted user's speech and / or the first TTS voice signal to the other party's electronic device.
[0074] As an example, the processor (250) may transmit the user's speech to the communication interface (240) when the user speaks, and transmit the user's speech to the other party's electronic device through the communication interface (240). In addition, the processor (250) may transmit the first TTS speech signal to the communication interface (240) when the first TTS speech signal is generated, and transmit the first TTS speech signal to the other party's electronic device through the communication interface (240). As an example, the processor (250) may record the user's speech. As an example, the processor (250) may activate a recording function when the user speaks. When the recording function is activated, the processor (250) may store the speech signal of the user's speech in the memory (210) or a buffer. In addition, the processor (250) transmits the recorded user's speech and the first TTS speech signal to the communication interface (240) when the first TTS speech signal is generated, and can simultaneously transmit the user's speech and the first TTS speech signal to the other party's electronic device through the communication interface (240).
[0075] The processor (250) can receive the other party's speech through the communication interface (240). The processor (250) can convert the received other party's speech into a third text. Then, the processor (250) can display the converted third text on the display (230). The processor (250) can translate the received other party's speech into a set user's language to generate a fourth text. Then, the processor (250) can display the fourth text generated in an area adjacent to the displayed third text on the display (230). For example, the third text may be a text that displays the other party's speech in the other party's language, and the fourth text may be a text that displays the other party's speech by translating it into the user's language.
[0076] For example, the electronic device (200) may include a translation option for speech. When the translation option is turned off, the processor (250) may not translate the user's speech and / or the other party's speech.
[0077] For example, a user may disable (or deactivate) the translation option for their speech in order to speak in the other party's language. If the translation option for the user's speech is disabled, the processor (250) may transmit the user's speech to the other party's electronic device without translating it. Furthermore, the processor (250) may display the first text in the dialogue script and not display the second text.
[0078] For example, if a user can understand the other party's language, the translation option for the other party's speech can be disabled. If the translation option for the other party's speech is disabled, the processor (250) can output the other party's speech without translating it. Furthermore, the processor (250) can display a third text in the dialogue script and not display a fourth text.
[0079] The processor (250) can convert the fourth text into a second TTS voice signal. Furthermore, the processor (250) can output the other party's speech and / or the converted second TTS voice signal through the speaker (220). For example, if the mute option for the other party's speech is turned on, the processor (250) can identify the other party's speech and not output the identified other party's speech. For example, if the mute option for the other party's speech is turned off, the processor (250) can output the other party's speech and the converted second TTS voice signal.
[0080] For example, when the processor (250) receives the other party's speech through the communication interface (240), it can output the other party's speech, and when the second TTS voice signal is generated, it can output the second TTS voice signal. As an example, when the processor (250) receives the other party's speech through the communication interface (240), it can record the other party's speech. Then, when the second TTS voice signal is generated, the processor (250) can sequentially output the recorded other party's speech and the second TTS voice signal.
[0081] When the processor (250) outputs the other party's speech and the second TTS voice signal, the processor (250) may sequentially output the other party's speech and the second TTS voice signal. The processor (250) may output the other party's speech at a second level of volume, and output the second TTS voice signal at a third level of volume that is greater than the second level. As an example, the processor (250) may sequentially output the first TTS voice signal, the other party's speech, and the second TTS voice signal. The processor (250) may output the first TTS voice signal at a first level of volume, output the other party's speech at a second level of volume that is greater than the first level, and output the second TTS voice signal at a third level of volume that is greater than the second level.
[0082] When the processor (250) outputs the second TTS voice signal, the processor (250) can apply an image effect to the fourth text in synchronization with the time at which the second TTS voice signal is output. For example, if the fourth text is “How are you?” and the processor (250) outputs “How are you?” as the second TTS voice signal, when outputting “How,” an indicator indicating “How” of the fourth text can be displayed, when outputting “are,” an indicator indicating “are” of the fourth text can be displayed, and when outputting “you,” an indicator indicating “you” of the fourth text can be displayed. For example, the indicator can include a change in color, a change in font size, a change in font weight, a change in font type, an inversion effect, and / or an underbar.
[0083] The other party's electronic device can output a first TTS voice signal. For example, when the other party's electronic device outputs the first TTS voice signal, the user's electronic device (200) can also output the same first TTS voice signal to inform the user and the other party of the time of speech. The processor (250) can output the first TTS voice signal in synchronization with the time at which the other party's electronic device outputs the first TTS voice signal. For example, while the other party's electronic device outputs the first TTS voice signal, the user's electronic device (200) can maintain a silent state. For example, while the other party's electronic device outputs the first TTS voice signal, the user's electronic device (200) can output a beep sound. For example, when the other party's electronic device outputs the first TTS voice signal, the user's electronic device (200) can output a beep sound at the time of outputting and ending the first TTS voice signal of the other party's electronic device.
[0084] FIGS. 3A, 3B, 3C, 3D, 3E, 3F, 3G, and 3H are diagrams illustrating an operation of installing language data according to various embodiments. The electronic device (200) can download, store, and install user language data and / or the other party's language data before making a call to the other party. The electronic device (200) can set one language as the default language according to a user's command. When the interpretation function is selected during a call to the other party whose language has not been set, the electronic device (200) can perform the interpretation function in the default language.
[0085] Referring to FIG. 3A, as an example, the electronic device (200) may include an option to activate an interpretation function. When the interpretation function is turned on (11) (or activated), the electronic device (200) may convert a user's speech into text, translate the user's speech into a set language of the other party to generate text, and convert the translated text of the user's speech into a TTS voice signal. The electronic device (200) may transmit the user's speech and / or the converted TTS voice signal to the other party's electronic device. In addition, the electronic device (200) may convert the received speech of the other party into text, translate the other party's speech into a set language of the user to generate text, and convert the translated text of the other party's speech into a TTS voice signal. The electronic device (200) may output the other party's speech and / or the converted TTS voice signal.
[0086] The electronic device (200) may set various information or change set information according to a user's command before performing the interpretation function. For example, when the interpretation function is turned on (11), the electronic device (200) may activate interpretation-related setting items (11-1, 11-2). The interpretation-related setting items (11-1, 11-2) may include an item for the user (11-2) and an item for the other party (11-1). For example, the item for the user (11-2) may include an item for setting the user's language and / or the voice of the TTS voice signal output to the user. The item for the other party (11-1) may include an item for setting the other party's language and / or the voice of the TTS voice signal output to the other party. As an example, the electronic device (200) may select the item for the other party (11-1) according to the user's command.
[0087] Referring to FIG. 3B, as an example, a screen showing a list of languages that can be selected as the other party's language is illustrated. For example, English may be set as the user's language and German may be set as the other party's language. The electronic device (200) may set (or change) the user's language and / or the other party's language according to the user's command. For example, in FIG. 3A, when the other party's language setting item is selected among the items (11-1) for the other party, the electronic device (200) may display a list of languages that can be selected as the other party's language, as shown in FIG. 3B. The user may select one language from the displayed language list.
[0088] As illustrated in FIG. 3c, when a selected language is selected (selected), the electronic device (200) can change the other party's language from the currently set language to the selected language. For example, if French (12) is selected from the displayed language list and French is selected (13), the electronic device (200) can determine whether language data for the selected language (e.g., French (12)) is installed. Referring to FIG. 3d, as an example, the electronic device (200) can display a message related to the download of language data. For example, if the language data is installed in the electronic device (200), the electronic device (200) can set (or change) the selected language to the other party's language. If the language data is not installed in the electronic device (200), the electronic device (200) can display a message related to the download of language data, as illustrated in FIG. 3d.
[0089] If Cancel is selected, the electronic device (200) may display the previous screen (e.g., FIG. 3c). If Download is selected, the electronic device (200) may download language data for the selected counterpart's language. For example, if French (12) is selected and the electronic device (200) does not have language data for French (12) installed, the electronic device (200) may download language data for French (12) from the server.
[0090] As illustrated in FIG. 3e, for example, when the electronic device (200) completes downloading (or installing) language data, it may display an interpretation function setting screen. As an example, the electronic device (200) may change the other party's language setting to the newly selected French (12).
[0091] The electronic device (200) can select a setting (change) item (15) of the user language according to a user's command. As an example, when the user's language setting item (15) is selected among the items (11-2) for the user according to the user's command, the electronic device (200) can set (or change) the user's language similarly to what was described in FIGS. 3A to 3E.
[0092] For example, when the user's language setting item (15) is selected, the electronic device (200) can display a list of languages that can be selected as the user's language. When a language is selected by the user from the displayed language list, the electronic device (200) can change the user's language from the currently set language to the selected language. As an example, when English is selected from the displayed language list, the electronic device (200) can determine whether language data for the selected language (e.g., English) is installed. If the language data is installed in the electronic device (200), the electronic device (200) can set (or change) the selected language to the user's language. If the language data is not installed in the electronic device (200), the electronic device (200) can display a message related to downloading the language data. When download is selected, the electronic device (200) can download the language data for the selected user's language. For example, if English is selected and language data for English is not installed in the electronic device (200), the electronic device (200) can download English language data from the server. Once the download (or installation) of the language data is complete, the electronic device (200) can change the user's language setting to the newly selected English.
[0093] Referring to FIG. 3F, as an example, a UI related to changing a user language is illustrated. The electronic device (200) may display a user language setting area (17) including a set user language and / or selectable user languages. In addition, the electronic device (200) may display a list of languages from which language data can be downloaded. For example, the languages that can be downloaded may be languages that can be added as user languages. For example, the user language setting area (17) may include one or more languages for which language data is installed. As an example, as illustrated in FIG. 3F, if language data for English and French is installed in the electronic device (200), the user language setting area (17) may include English and French. The electronic device (200) may set (or change) a language selected by the user from among the languages included in the user language setting area (17) as the user's language. The list of languages from which downloadable data can be downloaded may include one or more languages that can be added as the user's language. For example, a language included in the language list may be a language for which language data can be downloaded (or installed) from a server.
[0094] As an example, the electronic device (200) may select Korean (16) from the list of languages that can be added according to a user's command. Since Korean (16) is included in the list of languages that can be added, language data for Korean (16) may not be installed in the electronic device (200). The electronic device (200) may download language data for Korean (16) from a server.
[0095] Referring to FIG. 3g, the electronic device (200) can include the language for which the language data has been downloaded in the user language setting area (17). As an example, since the electronic device (200) has downloaded (or installed) the language data for Korean (16), the electronic device (200) can include Korean (16) in the user language setting area (17). The electronic device (200) can set English, French, or Korean included in the user language setting area (17) as the user's language according to the user's command.
[0096] Referring to FIG. 3H, as an example, the electronic device (200) may display a call partner language setting pop-up window (18) during a call according to a user's command. The electronic device (200) may change the user's language and / or the call partner's language during a call. As an example, when the electronic device (200) receives a command from the user to change the call partner's language setting, the electronic device (200) may display a call partner language setting pop-up window (18). The call partner language setting pop-up window (18) may include the language of the installed language data. The electronic device (200) may set the language selected in the call partner language setting pop-up window (18) to the call partner's language according to the user's command. As an example, when the call partner's language is set to English and a command to select French is received, the electronic device (200) may change the call partner's language from English to French. In addition, the electronic device (200) may translate the user's speech from English to French and then translate it.
[0097] As an example, setting (or changing) the other party's language can also be performed with similar operations as described in FIGS. 3e to 3h.
[0098] For example, the electronic device (200) can select a setting (change) item of the counterpart language on the interpretation function setting screen according to a user's command. The electronic device (200) can display a counterpart language setting area including the set counterpart language and / or selectable counterpart languages. In addition, the electronic device (200) can display a list of languages that can be added as the counterpart language. For example, the counterpart language setting area can include one or more languages for which language data is installed. The electronic device (200) can set (or change) a language selected by the user among the languages included in the counterpart language setting area as the counterpart language. The list of languages that can be added as the counterpart language can include one or more languages that can be added. For example, the languages included in the language list can be languages for which language data can be downloaded (or installed) from a server.
[0099] If language data for the selected language is not installed in the electronic device (200), the electronic device (200) can download the language data from the server. The electronic device (200) can include the downloaded language data in the other party's language setting area. The electronic device (200) can set the selected language from among at least one language included in the other party's language setting area as the other party's language according to a user's command.
[0100] FIGS. 4a, 4b, 4c, 4d, 4e, 4f, 4g, 4h, 4i, and 4j are drawings illustrating an operation of setting the language of the other party according to various embodiments.
[0101] Referring to FIG. 4A, as an example, an interpretation function setting screen is illustrated. The interpretation function setting screen may include items (21) for setting the interpretation function on / off, the user's language, the voice of the TTS voice signal output to the user, user speech mute, the other party's language, the voice of the TTS voice signal output from the other party's electronic device, the other party's speech mute, and / or the language of each other party. For example, when the user's speech mute item is turned on, the electronic device (200) may not transmit the user's speech to the other party's electronic device, but may only transmit the converted first TTS voice signal to the other party's electronic device. When the user's speech mute item is turned off, the electronic device (200) may transmit the user's speech and / or the converted first TTS voice signal to the other party's electronic device. When the other party's speech mute item is turned on, the electronic device (200) may not output the other party's speech, but may only output the converted second TTS voice signal. When the other party's speech mute item is turned off, the electronic device (200) can output the other party's speech and / or the converted second TTS voice signal.
[0102] The language setting item for the other party is an item that sets the default language of the other party, and the item (21) for setting the language of an individual other party may be an item that sets a specific language used by an individual other party. When the item (21) for setting the language of an individual other party is selected, the electronic device (200) can set a specific language for an individual user as the language of a specific other party.
[0103] As illustrated in FIG. 4B, for example, when the language setting option (21) for the other party is selected, the electronic device (200) may display a list of other parties with set languages (or an interpretation list) and / or an option to add a party (22). For example, a party included in the list may be a party for which a specific language has been selected. When a user makes a call to a party included in the list, the electronic device (200) may interpret the call in the set language of the party included in the list. When the option to add a party (22) is selected, the electronic device (200) may perform an operation to add the party to the list of parties and set the language of the party.
[0104] Referring to FIGS. 4C and 4D , as an example, when the add counterpart option (22) is selected, the electronic device (200) can select a counterpart for which a language has been set according to a user's command. The electronic device (200) can display a number input window (23) and / or a list of stored contacts. As an example, as illustrated in FIG. 4C , the electronic device (200) can select one contact from the displayed list of contacts, or as illustrated in FIG. 4D , the electronic device (200) can directly input a number (25) in the number input window (23). The electronic device (200) can determine whether the selected (or input) counterpart is a counterpart for which a language has already been set. If the selected (or input) counterpart is a counterpart for which a language has been set, the electronic device (200) can display a message notifying the user.
[0105] Referring to FIG. 4e, when a counterparty whose language is to be set is selected, the electronic device (200) may display a list of configurable languages. The list of configurable languages may include one or more languages. For example, the languages included in the list of configurable languages may be languages for which language data has been installed. Furthermore, the electronic device (200) may receive an input from the user to select a language (26) or an input to add a new language (26-1).
[0106] Referring to FIG. 4f, as an example, the electronic device (200) can display a list of languages (27) that can be set and a list of languages that can be added. The list of languages that can be added may include languages for which language data is not installed in the electronic device (200). If German (28) is selected from the list of languages that can be added by the user, the electronic device (200) can download language data for German (28) from the server. Furthermore, referring to FIG. 4g, the electronic device (200) can add German to the list of languages (27) that can be set.
[0107] Referring to FIG. 4H , as an example, the electronic device (200) may include a setting menu for a language and a voice of a TTS voice signal. As an example, the electronic device (200) may display a language set as a user language and a voice to output a first TTS voice signal in a user item (e.g., Me), and may display a language set as a counterpart language and a voice to transmit a second TTS voice signal in a counterpart item (e.g., Andrew Lee). When a language and / or a voice of a TTS voice signal included in a user item is selected, the electronic device (200) may change the user's language and / or the voice of the first TTS voice signal according to a user's command. When a language and / or a voice of a TTS voice signal included in a counterpart item is selected, the electronic device (200) may change the counterpart's language and / or the voice of the second TTS voice signal according to a user's command.
[0108] Referring to FIG. 4i, as an example, multiple numbers are displayed. If there are multiple numbers corresponding to the selected counterpart, the electronic device (200) can display multiple numbers corresponding to the selected counterpart. The electronic device (200) can select one number (29) from the multiple numbers according to a user command. The electronic device (200) can set the language and / or voice corresponding to the selected number (29) through the above-described operation.
[0109] Referring to FIG. 4J, the electronic device (200) can display a list of counterparts for which a language has been set. For example, the electronic device (200) can display the counterpart's name or number, and the set language. For example, if there are multiple counterparts, the electronic device (200) can display the counterpart's name, the number for which the language has been set, and the set language (30).
[0110] Figure 5a is a flowchart illustrating an interpreting operation according to various embodiments.
[0111] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0112] According to one embodiment, steps 510 to 580 may be understood to be performed in a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (200) of FIG. 2).
[0113] Referring to FIG. 5a, the electronic device (200) can perform a call function with the other party (510) and determine whether an interpretation function is being executed (520). If the interpretation function is not executed (520-NO), the electronic device (200) can perform a general call function (510). If the interpretation function is executed (520-YES), the electronic device (200) can determine whether the other party being called is included in the interpretation list (530). The other party included in the interpretation list may be a party for which a corresponding language has been set and for which the set language data has been downloaded.
[0114] If the other party is not included in the interpretation list (530-NO), the electronic device (200) can perform the interpretation function in the set default language (540). In addition, the electronic device (200) can add the other party on the call to the interpretation list (550). If the other party is included in the interpretation list (530-YES), the electronic device (200) can perform the interpretation function in the set language (560).
[0115] The electronic device (200) can determine whether the other party's language setting has changed during a call (570). If the other party's language setting has not changed during the call (570-NO), the electronic device (200) can perform an interpretation function using the previously set language (560). If the other party's language setting has changed during the call (570-YES), the electronic device (200) can perform an interpretation function using the changed language. In addition, the electronic device (200) can update information regarding the changed language setting (580).
[0116] FIG. 5b is a timing diagram illustrating an interpreting operation according to various embodiments.
[0117] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0118] According to one embodiment, steps 605 to 675 may be understood to be performed in a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (200) of FIG. 2 and / or electronic device (102, 104) of FIG. 1).
[0119] Referring to FIG. 5b, a first electronic device (200-1) (e.g., a user's electronic device) and a second electronic device (200-2) (e.g., a counterpart's electronic device) can perform a call function (605). During a call, the first electronic device (200-1) can execute an interpretation function according to the user's command (610).
[0120] The first electronic device (200-1) can convert the user's utterance into a first text and display the converted first text (615). For example, the first electronic device (200-1) can convert the user's utterance into the first text using STT. The first electronic device (200-1) can generate a dialogue script and display the text sequentially generated in the dialogue script.
[0121] The first electronic device (200-1) can translate the user's speech into a set second language and generate and display a second text (620). For example, the second language may be the language set as the other party's language. The first electronic device (200-1) can translate the user's speech (or first text) and generate a second text. The first electronic device (200-1) can display the generated second text in an area adjacent to the displayed first text.
[0122] For example, the user's language (e.g., first language) may be set to Korean, and the other party's language (e.g., second language) may be set to English. When the user utters "Hello" in Korean, the first electronic device (200) may convert the user's utterance into a first text and display it in the dialogue script. In addition, the first electronic device (200-1) may translate the user's utterance (or first text) into English. In example, the first electronic device (200-1) may generate a second text that translates "Hello." into "Hello." In addition, the first electronic device (200-1) may display "Hello." in an area adjacent to "Hello." displayed in the dialogue script.
[0123] The first electronic device (200-1) can convert the second text into a first TTS voice signal (625) and transmit the user's speech and / or the first TTS voice signal (630). For example, the first electronic device (200-1) can include a mute option for the user's speech. When the mute option for the user's speech is turned on, the first electronic device (200-1) can transmit only the converted first TTS voice signal to the second electronic device (200-2) without transmitting the user's speech. When the mute option for the user's speech is turned off, the first electronic device (200-1) can transmit the user's speech and the converted first TTS voice signal to the second electronic device (200-2).
[0124] The second electronic device (200-2) can output the received user's speech and / or the first TTS voice signal (635). For example, when the user speaks, the second electronic device (200-2) can receive and output the user's speech from the first electronic device (200-1). When the first TTS voice signal is generated in the first electronic device (200-1), the second electronic device (200-2) can receive and output the first TTS voice signal from the first electronic device (200-1). For example, the second electronic device (200-2) can simultaneously receive the user's speech and the first TTS voice signal from the first electronic device (200-1), and sequentially output the received user's speech and the first TTS voice signal. For example, the second electronic device (200-2) can output the user's speech at a volume level of 5 and output the first TTS voice signal at a volume level of 6, which is greater than the 5th level. The second electronic device (200-2) can receive the other party's speech. As an example, the second electronic device (200-2) can output the second TTS voice signal in the process 675 described below. The second electronic device (200-2) can output the second TTS voice signal at a volume level of 4, which is less than the 5th level.
[0125] As an example, when the second electronic device (200-2) outputs the first TTS voice signal, the first electronic device (200-1) may also output the same first TTS voice signal to inform the user and the other party of the time of speech (640). The first electronic device (200-1) may output the first TTS voice signal in synchronization with the time at which the first TTS voice signal is output by the second electronic device (200-2). For example, the first TTS voice signal transmitted to the second electronic device (200-2) may include time information. For example, the time information may include current time and delay time information or information about the output time. The first electronic device (200-1) and the second electronic device (200-2) may output the first TTS voice signal in synchronization based on the time information included in the first TTS voice signal. As an example, the first electronic device (200-1) and the second electronic device (200-2) may not perform a separate synchronization process. The first electronic device (200-1) may transmit the first TTS voice signal to the second electronic device (200-2) and then output it, and the second electronic device (200-2) may output the first TTS voice signal immediately after receiving it. As an example, while the second electronic device (200-2) outputs the first TTS voice signal, the first electronic device (200-1) may maintain a silent state. As an example, while the second electronic device (200-2) outputs the first TTS voice signal, the first electronic device (200-1) may output a beep sound. As an example, when the second electronic device (200-2) outputs the first TTS voice signal, the first electronic device (200-1) can output a beep sound at the time of outputting and the time of ending the first TTS voice signal of the second electronic device (200-2).
[0126] The first electronic device (200-1) can receive the other party's speech from the second electronic device (102) (645) and convert the received speech into third text and display it (650). As an example, the first electronic device (200-1) can receive the other party's speech, "I go to the fish market." from the second electronic device (200-2). The first electronic device (200-1) can convert the other party's speech into text and display "I go to the fish market." in the dialogue script.
[0127] The first electronic device (200-1) can translate the other party's utterance (or third text) into the user's language and generate and display a fourth text (655). The first electronic device (200-1) can translate the other party's utterance (or third text) to generate the fourth text. Then, the first electronic device (200-1) can display the generated fourth text in an area adjacent to the displayed third text. As an example, the first electronic device (200-1) can translate the other party's utterance (or third text) into Korean. The first electronic device (200-1) can generate the fourth text that translates "I go to the fish market." into "I go to the fish market." Then, the first electronic device (200-1) can display "I go to the fish market." in an area adjacent to "I go to the fish market." displayed in the dialogue script.
[0128] The first electronic device (200-1) can convert the fourth text into a second TTS voice signal (660). The first electronic device (200-1) transmits the second TTS voice signal, which is a translation of the other party's speech (or third text), to the second electronic device (200-2) (665), and the second electronic device (200-2) can output the received second TTS voice signal (675). For example, the first electronic device (200-1) can output the received other party's speech and / or the second TTS voice signal (670).
[0129] For example, when the other party speaks, the first electronic device (200-1) can receive and output the other party's speech from the second electronic device (200-2). When a second TTS voice signal is generated, the first electronic device (200-1) can output the second TTS voice signal in synchronization with the second TTS voice signal output from the second electronic device (102). In other words, when the second electronic device (200-2) outputs the second TTS voice signal, the first electronic device (200-1) can also output the second TTS voice signal. For example, the second TTS voice signal transmitted to the second electronic device (200-2) can include time information. For example, the time information can include current time and delay time information or information about the output time point. The first electronic device (200-1) and the second electronic device (200-2) can output the second TTS voice signal in synchronization based on the time information included in the second TTS voice signal. As an example, the first electronic device (200-1) and the second electronic device (200-2) may not perform a separate synchronization process. The first electronic device (200-1) can transmit the second TTS voice signal to the second electronic device (200-2) and then output it, and the second electronic device (200-2) can output the first TTS voice signal immediately after receiving it.
[0130] As an example, the first electronic device (200-1) may record the other party's speech, and when outputting a second TTS voice signal from the second electronic device (200-2), the other party's speech and the second TTS voice signal may be sequentially output. For example, the first electronic device (200-1) may output the other party's speech at a second level of volume, and output the second TTS voice signal at a third level of volume that is greater than the second level. As an example, the first electronic device (200-1) may output the first TTS voice signal at step 640. The first electronic device (200-1) may output the first TTS voice signal at a first level of volume that is less than the second level of volume.
[0131] For example, the first electronic device (200-1) may include a mute option for the other party's speech. If the mute option for the other party's speech is turned on, the first electronic device (200-1) may not output the other party's speech, but may only output the converted second TTS voice signal. If the mute option for the other party's speech is turned off, the first electronic device (200-1) may output the other party's speech and / or the converted second TTS voice signal.
[0132] As an example, the first electronic device (200-1) may output the received other party's speech and / or the second TTS voice signal without transmitting the second TTS voice signal to the second electronic device (200-2) (670).
[0133] FIGS. 6A, 6B, 6C, and 6D are diagrams illustrating operations for executing an interpretation function in a transmission according to various embodiments.
[0134] Referring to Fig. 6a, as an example, a call UI is illustrated. The call UI may include a call assist item (33). For example, the call assist item (33) may be an item for selecting additional functions during a call, and may include an on / off function for the interpretation function.
[0135] Referring to FIG. 6B, as an example, the electronic device (200) may display a text call item and a live translate item (34) depending on the selection of the call assist item (33). As an example, the live translate item (34) may include an on / off function for the interpretation function. When the live translate item (34) is selected, the electronic device (200) may perform the interpretation function.
[0136] Referring to FIGS. 6C and 6D , the UI for the interpretation function is illustrated. The electronic device (200) can generate a separate conversation script screen from the outgoing UI screen. For example, FIG. 6C illustrates the outgoing UI screen with the interpretation function turned on. The electronic device (200) can display a stop translating item (34-3). When the stop translating item (34-3) is selected, the electronic device (200) can stop the interpretation. When the electronic device (200) receives a gesture (34-1) for switching screens or an input (34-2) for selecting another screen, the electronic device (200) can switch to the conversation script screen.
[0137] For example, FIG. 6D illustrates a conversation script screen. The conversation script screen may include information about the other party, a set language of the other party, and / or information about the set user language. In addition, the electronic device (200) may display, on the conversation script screen, a first text converted from the user's speech into text, a second text generated by translating the user's speech, a third text converted from the other party's speech into text, and / or a fourth text generated by translating the other party's speech. The electronic device (200) may switch to the outgoing UI screen when it receives a gesture (35-1) for switching screens or an input (35-2) for selecting another screen.
[0138] FIGS. 7a, 7b, 7c, 7d, 7e and 7f are diagrams illustrating operations for executing an interpretation function in a receiver according to various embodiments.
[0139] Referring to Fig. 7a, as an example, a receiving UI is illustrated. The receiving UI may include a call assist item (36). For example, the call assist item (33) may be an item for selecting additional functions during a call and may include an on / off function for the interpretation function.
[0140] Referring to FIG. 7b, as an example, the electronic device (200) may display a text call item and a live translate item (37) depending on the selection of the call assist item (36). As an example, the live translate item (37) may include an on / off function for the interpretation function. When the live translate item (37) is selected, the electronic device (200) may turn on the interpretation function and receive a call.
[0141] Referring to FIG. 7C, as an example, a dialogue script screen is illustrated. For example, the dialogue script screen may include information about the other party, a set language of the other party, and / or a set user language. Furthermore, the electronic device (200) may display, on the dialogue script screen, a first text converted from the user's speech into text, a second text generated by translating the user's speech, a third text converted from the other party's speech into text, and / or a fourth text generated by translating the other party's speech.
[0142] For example, when the user's language (38-2) is selected, the electronic device (200) may display a menu for changing the user's language. The electronic device (200) may display a menu (39-2) for selecting one of the languages for which language data is installed. When a different language is selected according to the user's command, the electronic device (200) may change the user's language from the currently set language (e.g., Korean) to the selected language (e.g., English (Indian)).
[0143] For example, when the other party's language (38-1) is selected, the electronic device (200) can display a menu for changing the other party's language. The electronic device (200) can display a menu (39-1) for selecting one of the languages in which language data is installed. When another language is selected according to a user's command, the electronic device (200) can change the other party's language from the currently set language (e.g., English) to the selected language (e.g., Korean). As an example, when the Add Language (39-3) item is selected, the electronic device (200) can add a new language during a call.
[0144] Referring to FIG. 7D , as an example, a settings screen is illustrated when language addition (39-3) is selected. The settings screen may include a selectable language area (40) and a list of languages that can be added. As an example, the selectable language area (40) may include English (United States), English (India), and Korean. The electronic device (200) may select French (41) as the language to be added according to a user's command. When French is selected, the electronic device (200) may receive French (41) language data from the server. The electronic device (200) may store and install the received French (41) language data.
[0145] Referring to Figure 7e, as an example, a settings screen is shown when French is installed. As an example, since French (41) is installed, the selectable language areas (40) may include English (United States), English (India), Korean, and French.
[0146] Referring to FIG. 7F, as an example, when the other party's language (38-1) is selected, the electronic device (200) can display the other party's language setting menu in a pop-up window (42). The pop-up window (42) can include the languages of the installed language data. For example, since French (41) is installed, the electronic device (200) can display English (US), English (India), Korean, and French in the pop-up window (42). The electronic device (200) can set the language selected in the pop-up window (42) to the other party's language according to the user's command. As an example, when the other party's language is set to English and French is selected, the electronic device (200) can change the other party's language from English to French. In addition, the electronic device (200) can translate the user's speech from English to French and then translate it.
[0147] Figure 8 is a diagram illustrating ignition options according to various embodiments.
[0148] Referring to FIG. 8, the interpretation function setting menu may include a mute option (46) for the user's speech and / or a mute option (47) for the other party's speech. For example, when the mute option (46) for the user's speech is turned on, the electronic device (200) may not transmit the user's speech to the other party's electronic device, but may transmit only the first TTS voice signal converted from the user's speech to the other party's electronic device. When the mute option for the user's speech is turned off, the electronic device (200) may transmit the user's speech and / or the first TTS voice signal converted from the user's speech to the other party's electronic device. For example, when the mute option (47) for the other party's speech is turned on, the electronic device (200) may not output the other party's speech, but may output only the second TTS voice signal converted from the other party's speech. When the mute option (47) for the other party's speech is turned off, the electronic device (200) can output the other party's speech and / or a second TTS voice signal converted from the other party's speech.
[0149] FIGS. 9A, 9B, 9C, and 9D are diagrams illustrating interpretation actions displayed in a dialogue script according to various embodiments.
[0150] Referring to FIG. 9A, for example, the electronic device (200) may display visual information (51) indicating that it is recognizing the utterance in the dialogue script when the other party speaks. The electronic device (200) may recognize the utterance, convert the utterance into text, and display it in the dialogue script.
[0151] Referring to FIG. 9B, the electronic device (200) can partially recognize speech and convert the recognized speech into text. The electronic device (200) can display the converted text (52) in the dialogue script. The electronic device (200) can check whether the other party has stopped speaking at regular intervals. For example, the electronic device (200) can check whether the other party has stopped speaking at approximately every 1,500 ms. If the other party has stopped speaking, the electronic device (200) can convert the speech into text at regular intervals. For example, the electronic device (200) can convert the speech into text at the word level and display it in the dialogue script.
[0152] Referring to FIG. 9c, the electronic device (200) can display text generated by translating the other party's speech. The electronic device (200) can display text (56) translated into the user's language based on the other party's speech or text (55) converted from the other party's speech. For example, the electronic device (200) can identify the first sentence in the other party's speech. Then, the electronic device (200) can translate the identified first sentence into the user's language. The electronic device (200) can display the text (56) of the translated first sentence.
[0153] For example, the electronic device (200) can transmit a second TTS voice signal to the other party's electronic device. The second TTS voice signal can be a voice signal with the same content as the text (56) translated into the user's language. The other party's electronic device can output the received second TTS voice signal. As an example, the second electronic device (102) can output the second TTS voice signal at a lower volume than the volume at which the user's speech is output.
[0154] Referring to FIG. 9D , the electronic device (200) can display text generated by translating the second sentence of the other party's speech. For example, the electronic device (200) can identify the second sentence following the identified first sentence. Then, the electronic device (200) can translate the identified second sentence into the user's language. The electronic device (200) can display text (58) including the second sentence following the translated first sentence. In other words, the electronic device (200) can identify and translate the user's speech on a sentence-by-sentence basis. The electronic device (200) can display text translated on a sentence-by-sentence basis.
[0155] Additionally, the electronic device (200) can convert text (58) translated into the user's language into a TTS voice signal and output it. The TTS voice signal can also be converted and output in sentence units. When outputting the TTS voice signal, the electronic device (200) can apply an image effect to the text (58) translated into the user's language in synchronization with the time at which the TTS voice signal is output.
[0156] FIGS. 10A and 10B are diagrams illustrating a dialogue script displaying utterances and translations according to various embodiments.
[0157] Referring to FIG. 10A, a dialogue script (59) from which interpretation begins is illustrated. As an example, the electronic device (200) may display text from the bottom of the dialogue script (59). When the next text is input (or generated), the electronic device (200) may move the currently displayed text upward and display the next text input below.
[0158] Referring to FIG. 10b, the electronic device (200) may delete text associated with a previous utterance. For example, the electronic device (200) may delete text (60) associated with a previous utterance after a preset period of time has elapsed. For example, if the electronic device (200) displays text exceeding the text that can be displayed in the dialogue script, the electronic device (200) may delete the text (60) associated with the previous utterance that was displayed first among the texts displayed in the dialogue script and display new text. For example, the text (60) associated with the previous utterance may be deleted from the screen and / or memory (210). For example, the electronic device (200) may delete the text by gradually disappearing the text over a predetermined period of time (e.g., fading out).
[0159] Figure 11 is a flowchart explaining an operation of executing an interpretation function based on information of the other party confirmed through a server according to various embodiments.
[0160] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0161] According to one embodiment, steps 1110 to 1190 may be understood to be performed in a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (101) of FIG. 2).
[0162] Referring to FIG. 11, the electronic device (200) can perform a call function with the other party (1110) and determine whether the interpretation function is turned on (or activated) (1120). If the interpretation function is not turned on (1120-NO), the electronic device (200) can terminate operations related to the interpretation function (or perform no special operations) (1130).
[0163] When the interpretation function is turned on (1120-YES), the electronic device (200) can check the information of the other party's electronic device via the server (1140). For example, the information of the other party's electronic device may include the other party's number's country code, region code, system language information of the other party's electronic device, the other party's electronic device's manufacturer, the other party's electronic device's model name, and / or interpretation function support information of the other party's electronic device.
[0164] The electronic device (200) can determine whether the other party's electronic device supports the interpretation function based on the information of the identified other party's electronic device (1150). If the other party's electronic device does not support the interpretation function (1150-NO), the electronic device (200) can perform the interpretation function alone (1190).
[0165] If the other party's electronic device supports the interpretation function (1150-YES), the electronic device (200) can determine whether the other party's language and the user's language are different (1160). If the other party's language and the user's language are the same (1160-NO), the electronic device (200) can terminate the interpretation function-related operation (or perform no special operation) (1130). For example, if the user and the other party's languages are the same, the electronic device (200) may not perform the interpretation function.
[0166] If the other party's language and the user's language are different (1160-YES), the electronic device (200) can determine whether the other party's language data exists (or is installed) (1170). If the other party's language data exists (or is installed) in the electronic device (200) (1170-YES), the electronic device (200) can execute an interpretation function (1190). The electronic device (200) can generate a text translated from the user's speech into the other party's language based on the determined other party's language, and transmit a TTS voice signal generated from the user's speech and / or the translated text to the other party. In addition, the electronic device (200) can generate a text translated from the other party's speech into the user's language based on the set user's language, and output a TTS voice signal generated from the other party's speech and / or the translated text. The electronic device (200) can add the other party to the interpretation list.
[0167] If the other party's language data does not exist (1170-NO), the electronic device (200) can download the other party's language data (1180). Once the language data is downloaded, the electronic device (200) can store and install the language data and execute the interpretation function (1190).
[0168] FIG. 12a and FIG. 12b are diagrams illustrating an operation of setting the language of the other party based on the number information of the other party according to various embodiments.
[0169] Referring to FIG. 12A, the received phone number of the other party may include a country code (61) and / or an area code (62). The electronic device (200) may set the other party's language based on the country code (61) and / or the area code (62). For example, if the country code is 1 and the area code is 418, the electronic device (200) may determine that the other party's location is the Quebec region of Canada. The Quebec region of Canada may be a French-speaking region. The electronic device (200) may set the other party's language to French based on the language used according to the determined location. If the French language is not installed on the electronic device (200), the electronic device (200) may download and install French language data through the server. Then, the electronic device (200) may set the other party's language to French.
[0170] Referring to FIG. 12b, the electronic device (200) can execute an interpretation function based on the user's language and the other party's language. For example, the electronic device (200) can display a dialogue script. The dialogue script can include information related to the other party's number, the other party's language, and / or the user's language. As an example, since the electronic device (200) determines that the other party's language is French, it can display French (63) as information related to the other party's language.
[0171] The electronic device (200) can convert a user's utterance into a first text, and translate the user's utterance (or the first text) into the other party's language to generate a second text. For example, the first text may be Korean text, and the second text may be French text. The electronic device (200) can display the first text and / or the second text in a dialogue script. In addition, the electronic device (200) can receive the other party's utterance. The electronic device (200) can convert the other party's utterance into a third text, and translate the other party's utterance (or the third text) into the user's language to generate a fourth text. For example, the third text may be French text, and the fourth text may be Korean text. The electronic device (200) can display the third text and / or the fourth text in a dialogue script.
[0172] FIGS. 13a, 13b, and 13c are diagrams illustrating an operation of setting the language of the other party based on additional information according to various embodiments.
[0173] Referring to FIG. 13A, the electronic device (200) can receive a call from a specific number (66). The specific number (66) may be a number not stored in the electronic device (200). If the electronic device (200) does not receive information about the other party from the server, the other party's language is not set, or the other party's location cannot be determined from the phone number, the electronic device (200) can use additional information to determine the other party's language and set the other party's language.
[0174] Referring to FIG. 13B, as an example, the electronic device (200) may receive a message (67) from a specific number (66) prior to a call. The electronic device (200) may determine the language of the other party of the specific number (66) based on the language in which the message is written. As an example, the message received from the specific number (66) may be a message (67) written in French. The electronic device (200) may determine that the other party is a French-speaking user and set the language for the specific number (66) to French.
[0175] Referring to FIG. 13c, the electronic device (200) can execute an interpretation function based on the user's language and the other party's language. For example, the electronic device (200) can display a dialogue script. The dialogue script can include information related to the other party's number, the other party's language, and / or the user's language. As an example, since the electronic device (200) determines that the other party's language is French, it can display French (68) as information related to the other party's language.
[0176] FIGS. 14a, 14b, 14c, 14d, and 14e are diagrams illustrating screens displayed on a user's electronic device when an interpretation function is executed according to various embodiments.
[0177] Referring to Fig. 14a, a call UI is illustrated. As an example, the call UI may include an item (71) for turning the interpretation function on / off. When the on / off item (71) of the interpretation function is turned on, the electronic device (200) (e.g., the first electronic device) of a user (e.g., a first user) can check the information of the other party through the server. For example, the user's electronic device (200) can check the country code of the other party's (e.g., a second user's) number, the area code, the system language information of the other party's electronic device, the manufacturer of the other party's electronic device (e.g., the second electronic device), the model name of the other party's electronic device, and / or the interpretation function support information of the other party's electronic device.
[0178] The user's electronic device (200) can determine whether the other party's electronic device supports the interpretation function based on the verified information about the other party. For example, if the other party's electronic device does not support the interpretation function, the electronic device (200) can execute the interpretation function on its own without a separate notification. If the other party's electronic device supports the interpretation function, the user's electronic device (200) can provide the other party with a notification regarding the execution of the interpretation function. For example, if the other party's electronic device supports the interpretation function, the other party's electronic device can also activate the interpretation function. If the other party's electronic device's interpretation function is activated, the other party's electronic device can provide the user's electronic device (200) with a notification regarding the execution of the interpretation function.
[0179] Referring to FIG. 14b, the user's electronic device (200) can display a notification (72) received from the other party's electronic device. If the user's electronic device (200) and the other party's electronic device support an interpretation function and both execute the interpretation function, the user's electronic device (200) and the other party's electronic device can interpret only the utterances they transmit or receive, respectively.
[0180] For example, the user's electronic device (200) and the other party's electronic device can only interpret the utterances they transmit.
[0181] As an example, the user's electronic device (200) can convert a user's (e.g., a first user's) speech into text and display the converted text in a dialogue script. The user's electronic device (200) can generate a text translated into the language of the other party (e.g., a second user) from the user's speech (or the converted text) and convert the generated text into a TTS voice signal. The user's electronic device (200) can display the text translated into the other party's language in the dialogue script. In addition, the user's electronic device (200) can transmit the user's speech, a text for the user's speech, a text translated from the user's speech into the other party's language, and / or a TTS voice signal translated from the user's speech to the other party's electronic device. In addition, the other party's electronic device can convert the other party's speech into text and display the converted text in the dialogue script.
[0182] The electronic device of the other party can generate a text translated into the user's language from the other party's speech (or converted text) and convert the generated text into a TTS voice signal. The electronic device of the other party can display the text translated into the user's language in a conversation script. In addition, the electronic device of the other party can transmit the other party's speech, a text regarding the other party's speech, a text translated from the other party's speech into the user's language, and / or a TTS voice signal translated from the other party's speech to the user's electronic device (200). The electronic device of the user (200) can display the text converted from the other party's speech and / or the text translated from the other party's speech into the user's language in a conversation script.
[0183] For example, the user's electronic device (200) and the other party's electronic device can only interpret the utterances they each receive.
[0184] As an example, the user's electronic device (200) may convert the user's speech into text and display the converted text in a dialogue script. Furthermore, the user's electronic device (200) may transmit the user's speech and / or text related to the user's speech to the other party's electronic device.
[0185] The other party's electronic device can convert the received user's utterance into text and display it in a conversation script, or display text corresponding to the received user's utterance in the conversation script. The other party's electronic device can generate text translated into the other party's language from the user's utterance (or text corresponding to the user's utterance) and convert the generated text into a TTS voice signal. The other party's electronic device can then display the translated text in the conversation script.
[0186] The other party's electronic device can convert the other party's speech into text and display the converted text in a conversation script. Furthermore, the other party's electronic device can transmit the other party's speech and / or text related to the other party's speech to the user's electronic device (200).
[0187] The user's electronic device (200) can convert the received speech of the other party into text and display it in a dialogue script, or display text regarding the received speech of the other party in the dialogue script. The user's electronic device (200) can generate text translated into the user's language from the other party's speech (or text regarding the other party's speech) and convert the generated text into a TTS voice signal. The user's electronic device (200) can display the text translated into the user's language in the dialogue script.
[0188] Referring to FIG. 14c, the electronic device (200) can automatically perform an interpretation function. For example, when receiving a call from the other party's electronic device, the electronic device (200) can check whether the interpretation function of the other party's electronic device is turned on or off. The electronic device (200) can check information about the other party's electronic device from a server. For example, the electronic device (200) can check the other party's electronic device's interpretation function support information and whether the interpretation function is turned on or off through the server. For example, if the other party's electronic device supports the interpretation function, the electronic device (200) can directly check whether the interpretation function is turned on or off through the other party's electronic device.
[0189] If the interpretation function of the other party's electronic device is turned on, the electronic device (200) can display a notification (72-2) indicating that the interpretation function of the other party's electronic device is turned on. For example, if the interpretation function of the other party's electronic device is turned on, the electronic device (200) can change the receive button (73) to an image (or icon) indicating that the interpretation function is automatically performed when the other party's call is received. In addition, when the receive button (call button) (73) is selected, the electronic device (200) can automatically perform the interpretation function when the call is received. Referring to FIG. 14d, as an example, a dialogue script displayed on the user's electronic device (200) is illustrated when the user's electronic device (200) and the other party's electronic device each interpret only the utterances received. Since the user's electronic device (200) only interprets the utterances it receives (e.g., the other party's utterances), it can display a text (74-1) converted from the user's utterances, rather than displaying a text translated into the other party's language in the dialogue script.
[0190] Referring to FIG. 14e, as an example, a dialogue script displayed on the user's electronic device (200) is illustrated when only the utterances transmitted by the user's electronic device (200) and the other party's electronic device are interpreted. Since the user's electronic device (200) interprets only the utterances transmitted (e.g., the user's utterance), the dialogue script can display a text converted from the user's utterance and a text translated into the other party's language (74-2).
[0191] FIG. 15a and FIG. 15b are drawings explaining a screen displayed on the other party's electronic device when an interpretation function is executed according to various embodiments.
[0192] Referring to FIG. 15A, the other party's electronic device (e.g., the second electronic device) may display a notification (76) received from the user's electronic device (200) (e.g., the first electronic device). For example, the received notification (76) may be a notification inquiring whether the interpretation function is turned on in the other party's electronic device. For example, if the interpretation function is turned off (or the interpretation function execution is not selected) in the other party's electronic device according to a command from the other party (e.g., the second user), the other party's electronic device does not perform the interpretation function, and only the user's electronic device (200) may perform the interpretation function. For example, if the interpretation function is turned on (or the interpretation function execution is selected) in the other party's electronic device according to a command from the other party, the user's electronic device (200) and the other party's electronic device may only interpret the speech they transmit or receive, respectively.
[0193] Similar to the description in FIG. 14c, the other party's electronic device can automatically perform the interpretation function. For example, when the other party's electronic device receives a call from the user's electronic device (200), the other party's electronic device can check whether the interpretation function of the user's electronic device (200) is turned on or off. The other party's electronic device can check information about the user's electronic device (200) from the server. For example, the other party's electronic device can check the interpretation function support information of the user's electronic device (200) and whether the interpretation function is turned on or off through the server. For example, if the user's electronic device (200) supports the interpretation function, the other party's electronic device can directly check whether the interpretation function is turned on or off through the user's electronic device (200).
[0194] If the interpretation function of the user's electronic device (200) is turned on, the other party's electronic device may display a notification indicating that the interpretation function of the user's electronic device (200) is turned on. For example, if the interpretation function of the user's electronic device (200) is turned on, the other party's electronic device may change the receive button to an image (or icon) indicating that the interpretation function is automatically performed upon receiving a call. In addition, when the receive button (call button) is selected, the other party's electronic device may automatically perform the interpretation function upon receiving a call.
[0195] Referring to FIG. 15B, as an example, when only the utterances received by the user's electronic device (200) and the other party's electronic device are interpreted, a dialogue script displayed on the other party's electronic device is illustrated. The other party's electronic device may convert the received user's utterance into text (77-1) and display it in the dialogue script, or display the text (77-1) for the received user's utterance in the dialogue script. The other party's electronic device may generate a text (77-2) translated into the language of the other party (e.g., the second user) from the user's (e.g., the first user's) utterance (or the text for the user's utterance) and convert the generated text (77-2) into a TTS voice signal. The other party's electronic device may display the text (77-2) translated into the other party's language in the dialogue script.
[0196] FIG. 16A is a diagram illustrating a dialogue script screen including setting options according to various embodiments.
[0197] Referring to FIG. 16A, the electronic device (200) may display options related to interpretation on the dialogue script screen. For example, the options related to interpretation may include a first option for selecting whether to output a TTS voice signal interpreting the other party's speech, a second option for selecting whether to interpret the user's speech, and / or a third option for selecting whether to mute the user's speech.
[0198] As an example, the other party's information (78) displayed on the dialogue script screen may be selected. When the other party's information (78) is selected, the electronic device (200) may display a first option (78-1) for selecting whether to output a TTS voice signal that interprets the other party's speech. If the first option is turned on (or turned on) (78-1), the electronic device (200) may translate the other party's speech and output a TTS voice signal generated based on the translation result. The electronic device (200) may display an image effect on the translated text displayed on the dialogue script screen in synchronization with the output of the TTS voice signal. If the first option is turned off (or turned off) (78-2), the electronic device (200) may not output the TTS voice signal. The electronic device (200) may not display an image effect on the translated text displayed on the dialogue script screen.
[0199] As an example, the user's information (79) displayed on the dialogue script screen may be selected. When the user's information (79) is selected, the electronic device (200) may display a second option for selecting whether to interpret the user's speech and / or a third option for selecting whether to mute the user's speech.
[0200] For example, if the second option is turned on (or turned on) (79-1), the electronic device (200) can translate the user's speech and transmit a TTS voice signal generated based on the translation result to the other party's electronic device. The other party's electronic device can output the received TTS voice signal. If the second option is turned on (79-1), the third option can be selected. For example, if the third option is turned off (or turned off), the electronic device (200) can transmit the TTS voice signal generated based on the user's speech and the translation result of the user's speech to the other party's electronic device. The other party's electronic device can sequentially output the received user's speech and the TTS voice signal. If the third option (79-2) is turned on (or turned on), the electronic device (200) can transmit the TTS voice signal generated based on the translation result of the user's speech to the other party's electronic device without transmitting the user's speech to the other party's electronic device. The other party's electronic device can output the received TTS voice signal.
[0201] For example, if the second option is in the off state (or turned off) (79-2), the electronic device (200) may not interpret the user's speech. And, since the electronic device (200) does not interpret the user's speech, the third option for selecting whether to mute the user's speech may be dimmed (or disabled).
[0202] As an example, the interpretation option related to the other party's information (78) may include a fourth option that selects whether to transmit a second TTS voice signal. The second TTS voice signal may be a voice signal generated by translating the other party's speech into the user's language. When the other party's information (78) is selected, the electronic device (200) may display the fourth option along with the first option (78-1).
[0203] When the fourth option is turned on, the electronic device (200) can transmit the second TTS voice signal to the other party's electronic device. The other party's electronic device can output the received second TTS voice signal. For example, the other party's electronic device can output the second TTS voice signal at a lower volume than the volume at which the user's speech is output. When the fourth option is turned off, the electronic device (200) may not transmit the second TTS voice signal to the other party's electronic device.
[0204] FIG. 16b is a diagram illustrating the volume of speech and / or TTS voice signals according to various embodiments.
[0205] Referring to FIG. 16b, as an example, the volume levels of the other party's speech and the second TTS voice signal output through a loudspeaker and the volume levels of the speech and the TTS voice signal output through another output device (e.g., earphones) are illustrated.
[0206] For example, the electronic device (200) can translate the other party's speech and generate a second TTS voice signal corresponding to the translated other party's speech. The electronic device (200) can output the other party's speech and / or the second TTS voice signal. As an example, the electronic device (200) can output the other party's speech at a second level of volume and output the second TTS voice signal at a third level of volume that is greater than the second level. As illustrated in FIG. 16B, as an example, the electronic device (200) can output the other party's speech at a volume of 3 and output the second TTS voice signal at a volume of 10. When the volume level is above a certain level, the electronic device (200) can limit the output volume of the loudspeaker to prevent howling and / or echo phenomena that occur when the output of the loudspeaker is input to the microphone. As illustrated in FIG. 16b, if the output volume of the loudspeaker is 11, the electronic device (200) may not increase the volume of the second TTS voice signal even if the volume of the other party's speech increases. If the other party's speech and / or the second TTS voice signal is output to another output device, the electronic device (200) may output the other party's speech and / or the second TTS voice signal at a larger volume because howling and / or echo phenomena do not occur.
[0207] The electronic device (200) can translate a user's speech and generate a first TTS voice signal corresponding to the translated user's speech. As an example, the electronic device (200) can output the first TTS voice signal. The electronic device (200) can output the first TTS voice signal at a first level of volume that is lower than a second level of volume. As an example, when the electronic device (200) outputs the other party's speech at a volume of 3, it can output the first TTS voice signal at a volume of 1. As an example, the electronic device (200) can map (or group) the volume of the other party's speech, the volume of the first TTS voice signal, and / or the volume of the second TTS voice signal to each other. As the volume increases or decreases, the volume of the other party's speech, the volume of the first TTS voice signal, and / or the volume of the second TTS voice signal can increase or decrease together.
[0208] For example, the electronic device (200) can adjust the volume based on input to a volume key located on the side. As an example, the electronic device (200) can adjust the volume based on a touch gesture input on the screen.
[0209] For example, the other party's electronic device may output a first TTS voice signal corresponding to the user's speech, a translated user's speech, and / or a second TTS voice signal corresponding to the translated other party's speech. As an example, the other party's electronic device may output the user's speech at a volume level of 5, output the first TTS voice signal at a volume level of 6, which is greater than the 5th level, and output the second TTS voice signal at a volume level of 4, which is less than the 5th level.
[0210] FIG. 17a and FIG. 17b are flowcharts illustrating operations in which a user's electronic device and a counterpart's electronic device perform an interpretation function together according to various embodiments.
[0211] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0212] According to one embodiment, steps 1705 to 1760 may be understood to be performed in a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (200) of FIG. 2).
[0213] Referring to FIG. 17a, the electronic device (200) can activate the interpretation function (1705) and check the information of the other party's electronic device via the server (1710). For example, the information of the other party's electronic device may include the other party's number's country code, region code, system language information of the other party's electronic device, the manufacturer of the other party's electronic device, the other party's electronic device's model name, and / or interpretation function support information of the other party's electronic device.
[0214] The electronic device (200) can determine whether the other party's electronic device supports the interpretation function based on the information of the identified other party's electronic device (1715). If the other party's electronic device does not support the interpretation function (1715-NO), the electronic device (200) can execute the interpretation function without notification (1720).
[0215] If the other party's electronic device supports the interpretation function (1715-YES), the electronic device (200) can provide the other party with a notification regarding the execution of the interpretation function (1725). The other party's electronic device can display a notification received from the user's electronic device (200). For example, the received notification may be a notification asking whether the interpretation function is enabled on the other party's electronic device. The electronic device (200) can determine whether the other party consents (1730). For example, if the interpretation function is enabled on the other party's electronic device, the electronic device (200) can determine that the other party consents. If the interpretation function is not enabled on the other party's electronic device, the electronic device (200) can determine that the other party does not consent.
[0216] If the other party does not agree with the interpretation function (1730-NO), the electronic device (200) executes the interpretation function, and the other party's electronic device can display a dialogue script screen (1735). For example, the electronic device (200) can convert the user's speech into a first text and display the converted first text. In addition, the electronic device (200) can generate and display a second text that translates the user's speech (or the first text) into the other party's language. The electronic device (200) can convert the second text into a first TTS voice signal. The electronic device (200) can transmit the user's speech, the first TTS voice signal, the first text, and / or the second text to the other party's electronic device. The other party's electronic device can output the user's speech and / or the first TTS voice signal based on the data received from the electronic device (200), and display the first text and / or the second text on the dialogue script screen.
[0217] The server can manage conversation scripts. For example, the server can generate a single conversation script that includes text data received from the user's electronic device (200) and text data received from the other party's electronic device, and manage the generated conversation script. As an example, the server can generate a first conversation script displayed on the user's electronic device (200) and a second conversation script displayed on the other party's electronic device. The server can synchronize and manage the first conversation script and the second conversation script.
[0218] If the other party agrees to the interpretation function (1730-YES), the electronic device (200) can turn off the interpretation of the user's speech or the other party's speech (1740). When the interpretation of the user's speech is turned off, the electronic device (200) can interpret the other party's speech. Furthermore, the other party's electronic device can interpret the user's speech. When the interpretation of the other party's speech is turned off, the electronic device (200) can interpret the user's speech. Furthermore, the other party's electronic device can interpret the other party's speech.
[0219] Referring to FIG. 17b, the electronic device (200) can determine whether the text displayed in the dialogue script has been modified (1745). If the electronic device (200) determines that the text displayed in the dialogue script has not been modified (1745-NO), the electronic device (200) may not perform any special action (1750). If the electronic device (200) determines that the text displayed in the dialogue script has been modified (1745-YES), the electronic device (200) may update the modified text and translation result (1755). In addition, the dialogue script updated in the electronic device (200) may be synchronized with the dialogue script displayed on the other party's electronic device by the server (1760).
[0220] For example, if the text of a dialogue script displayed on an electronic device (200) is modified, the electronic device (200) may transmit information about the modified text to the server. For example, if the server manages a single dialogue script, the server may update the text of the single dialogue script based on the information about the modified text. The server may transmit the updated dialogue script or an update notification to the electronic device of the other party. The electronic device of the other party may receive information about the updated dialogue script or the modified text and update the dialogue script displayed on the electronic device of the other party. For example, if the server manages a first dialogue script and a second dialogue script, the server may update the first dialogue script based on information about the modified text. Then, the server may synchronize the first dialogue script and the second dialogue script. The server may transmit the updated second dialogue script or an update notification to the electronic device of the other party. The electronic device of the other party may receive information about the updated second dialogue script or the modified text and update the dialogue script displayed on the electronic device of the other party.
[0221] FIG. 18 is a timing diagram illustrating an operation in which a user's electronic device and a counterpart's electronic device perform an interpretation function together according to various embodiments.
[0222] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0223] According to one embodiment, steps 1805 to 1885 may be understood to be performed in an electronic device (e.g., electronic device (200) of FIG. 2 and / or electronic devices (102, 104) of FIG. 1) and a server (e.g., server (108) of FIG. 1).
[0224] Referring to FIG. 18, a first electronic device (200-1) (e.g., a user electronic device) can turn on an interpretation function (1805) and check whether a second electronic device (200-2) (e.g., a counterparty electronic device) supports the interpretation function (1810). For example, the first electronic device (200-1) can check information on the second electronic device (200-2) from the server (300). The information on the second electronic device (200-2) can include the country code of the counterparty's number, the area code, the system language information of the second electronic device (200-2), the manufacturer of the second electronic device (200-2), the model name of the second electronic device (200-2), and / or interpretation function support information of the second electronic device (200-2).
[0225] If the second electronic device (200-2) supports the interpretation function, the first electronic device (200-1) can transmit a notification to the second electronic device (200-2) regarding the execution of the interpretation function (1815) and receive consent from the second electronic device (200-2) regarding the execution of the interpretation function (1820).
[0226] For example, if the second electronic device (200-2) agrees to execute the interpretation function, the first electronic device (200-1) can turn off the speech translation function of the second user (e.g., the user of the second electronic device (200-2)) (1825), and the second electronic device (200-2) can turn on the speech translation function of the second user (1830). In other words, the first electronic device (200-1) and the second electronic device (200-2) can perform interpretation of the transmitted speech. As another example, the first electronic device (200-1) can turn off the speech translation function of the first user (e.g., the user of the first electronic device (200-1)), and the second electronic device (200-2) can turn on the speech translation function of the first user. In other words, the first electronic device (200-1) and the second electronic device (200-2) can perform interpretation of the received speech.
[0227] When the first electronic device (200-1) and the second electronic device (200-2) perform interpretation of the transmitted speech, the first electronic device (200-1) can translate the first user's speech (1835) and transmit the translation result to the server (300) and the second electronic device (200-2) (1840). The server (300) can manage a conversation script between the first user and the second user (1845).
[0228] The second electronic device (200-2) can display the translation result for the received first user's speech (1850). The second electronic device (200-2) can translate the second user's speech (1855) and transmit the translation result to the server (300) and the first electronic device (200-1) (1860).
[0229] The first electronic device (200-1) can display the translation results for the received second user's speech (1865). The first electronic device (200-1) can modify the text displayed in the dialogue script based on the user's input and translate the modified text (1870). The first electronic device (200-1) can transmit the modified text and the translation results to the server (300) (1875).
[0230] The server (300) can update and / or synchronize the managed dialogue script based on the received revised text and translation results (1880). For example, if the server (300) manages one dialogue script, it can update one dialogue script. As an example, the server (300) can manage a first dialogue script displayed on a first electronic device (200-1) and a second dialogue script displayed on a second electronic device (200-2). The server (300) can update the first dialogue script and synchronize the second dialogue script based on the received revised text and translation results. The server (300) can transmit the synchronized revised text and translation results to the second electronic device (200-2) (1885). The second electronic device (200-2) can update the displayed dialogue script.
[0231] FIGS. 19a, 19b, and 19c are diagrams illustrating operations for modifying text displayed in a dialogue script according to various embodiments.
[0232] FIG. 19A illustrates a dialogue script screen (3010) displayed on a first electronic device (200-1) (e.g., the user's electronic device (200)) and a dialogue script screen (3020) displayed on a second electronic device (200-2) (e.g., the other party's electronic device). The first electronic device (200-1) and the second electronic device (200-2) can each translate a transmitted or received utterance and display the translation result. As an example, the first electronic device (200-1) and the second electronic device (200-2) can display the text "How are you? Good to hear that." uttered by the first user and the text "Hello? Nice to hear from you." translated from the first user's utterance. The first electronic device (200-1) can receive an input (81) for modifying the first user's utterance.
[0233] Referring to FIG. 19b, when the first electronic device (200-1) receives user input on the displayed text, it may display a modification menu (82). When the modification menu is selected, the first electronic device (200-1) may display a UI (83) for modifying the text of the first user's utterance (or a translated text of the first user's utterance). As an example, the first electronic device (200-1) may receive an input from the first user to modify "How are you? Good to hear that." to "How are you going to get here?"
[0234] Referring to FIG. 19c, a dialogue script screen (3030) of an electronic device (200-1) and a dialogue script screen (3040) of a second electronic device (200-2) are illustrated, in which the modified text is displayed in synchronization. Once the text modification is complete, the first electronic device (200-1) can translate the modified text. As an example, the first electronic device (200-1) can display the translated text “What are you going to ride?” from the modified text “How are you going to get here?” The first electronic device (200-1) can transmit the modified text and the translated text to the server.
[0235] For example, the server can manage a dialogue script. As an example, the server (300) can generate a dialogue script including text data received from the first electronic device (200-1) and text data received from the second electronic device (200-2), and manage the generated dialogue script. As another example, the server (300) can generate a first dialogue script displayed on the first electronic device (200-1) and a second dialogue script displayed on the second electronic device (200-2). The server (300) can synchronize and manage the first dialogue script and the second dialogue script. When the server (300) receives modified text and translated text from the first electronic device (200-1), the server (300) can update a dialogue script. As an example, the server (300) can update the first dialogue script and synchronize the second dialogue script with the first dialogue script.
[0236] The server (300) may transmit information about the updated second dialogue script to the second electronic device (200-2). For example, the server (300) may transmit information about a notification of the dialogue script, an updated dialogue script, and / or a modified / translated text to the second electronic device (200-2). The second electronic device (200-2) may update the displayed dialogue script based on the received information. For example, if the spoken text and translated text of the dialogue script of the first electronic device (200-1) are modified to the text (82-1) of “How are you going to get here?” and “What are you going to ride?”, the second electronic device (200-2) may synchronize with the dialogue script of the first electronic device (200-1) and update it with the modified text (82-2).
[0237] FIGS. 20A, 20B, and 20C are diagrams illustrating operations for modifying text displayed in a call history according to various embodiments.
[0238] Referring to FIG. 20A, a conversation history (84) stored in an electronic device (200) is illustrated. The electronic device (200) can modify text included in the stored conversation history (84). The electronic device (200) can receive an input from a user to select (85) a single conversation history.
[0239] Referring to FIG. 20B, a dialogue script screen (86) of a selected dialogue history is illustrated. When a dialogue history is selected, the electronic device (200) may display the dialogue script (86) of the selected dialogue history. When the electronic device (200) receives an input (87) for modifying text on the displayed dialogue script, the electronic device (200) may display a menu (88) for modifying the selected text. When the modification menu (88) is selected, the electronic device (200) may display a UI (89) for modifying the text of the first user's utterance (or the translated text of the first user's utterance). As an example, the electronic device (200) may receive an input from the first user for modifying "How are you? Good to hear that." to "How are you going to get here?"
[0240] Referring to FIG. 20c, a dialogue script with modified text is illustrated. Once the text modification is complete, the electronic device (200) can translate the modified text. For example, the electronic device (200) can display the text (90-2) of "What are you going to ride?", which is a translation of the modified text (90-1) of "How are you going to get here?" The electronic device (200) can transmit the modified text and the translated text to the server (300). The server (300) can update the dialogue script displayed on the electronic device (200) and synchronize the dialogue script displayed on the other party's electronic device. When the server transmits information about the updated second dialogue script to the other party's electronic device, the other party's electronic device can update the corresponding dialogue script.
[0241] FIG. 21a and FIG. 21b are diagrams illustrating operations for executing an interpretation function in a conference call according to various embodiments.
[0242] Referring to FIG. 21a, the call screen of the electronic device (200) may include an item (91) for executing an interpretation function. The electronic device (200) may receive a user command for executing an interpretation function during a conference call (e.g., a multi-party call).
[0243] Referring to Figure 21b, a dialogue script for executing the interpretation function is illustrated. The dialogue script may include an area (92) displaying participant information and an area displaying text. As an example, a user (Me) may speak Korean, a first participant (e.g., Christina Adams) may speak English, and a second participant (e.g., Andrew Lee) may speak French.
[0244] When the interpretation function is executed, the electronic device (200) can display a notification according to the execution of the interpretation function. The electronic device (200) can receive Korean speech from the user. The electronic device (200) can convert the received Korean speech into text. The electronic device (200) can generate a translation text that translates the received Korean speech (or the converted text) into English, the language of the first participant, and a translation text that translates the received Korean speech into French, the language of the second participant. In addition, the electronic device (200) can display the text converted from the Korean speech, the translation text translated into English, and the translation text translated into French (93) in the dialogue script. The electronic device (200) can receive English speech from the first participant and convert the received English speech into text. The electronic device (200) can generate a translation text that translates the received English speech (or the converted text) into Korean, the language of the user. In addition, the electronic device (200) can display the text converted from the English utterance and the translated text (94) translated into Korean in the dialogue script. The electronic device (200) can receive a French utterance from a second participant and convert the received French utterance into text. The electronic device (200) can generate a translated text that translates the received French utterance (or the converted text) into Korean, the user's language. In addition, the electronic device (200) can display the text converted from the French utterance and the translated text (95) translated into Korean in the dialogue script.
[0245] When a user speaks, the electronic device (200) can convert the user's speech into text. The electronic device (200) can translate the user's speech (or the converted text) into English, the language of the first participant, and generate an English TTS voice signal. The electronic device (200) can transmit the generated English TTS voice signal to the first participant. Simultaneously or sequentially, the electronic device (200) can translate the user's speech (or the converted text) into French, the language of the second participant, and generate a French TTS voice signal. The electronic device (200) can transmit the generated French TTS voice signal to the second participant.
[0246] Figure 22 is a flowchart illustrating a method of interpretation during a call according to various embodiments.
[0247] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0248] According to one embodiment, steps 2210 to 2280 may be understood to be performed in a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (200) of FIG. 2).
[0249] Referring to FIG. 22, the electronic device (200) can convert a user's speech in a first language into first text and display it on the screen (2210). For example, the electronic device (200) can generate a dialogue script and display the first text in the dialogue script.
[0250] The electronic device (200) can display a second text generated by translating a first text into a second language on the screen (2220). As an example, the electronic device (200) can translate a user's speech into the second language. The second language may be a language set to the other party's language. As an example, the user's language may be Korean, and the other party's language may be English. The user may speak in Korean, and the first text may be in Korean. The electronic device (200) can translate the first text (or the user's speech) into English to generate a second text. Then, the electronic device (200) can display the second text generated together with the first text. As an example, the second text may be displayed in an area adjacent to the first text. The electronic device (200) can convert the second text into a first TTS voice signal (2230). For example, the converted first TTS voice signal may be in English.
[0251] The electronic device (200) can transmit voice data according to the mute option status of the user's speech to the other party's electronic device (2240). For example, if the mute option for the user's speech is turned on, the electronic device (200) can transmit only the converted first TTS voice signal, excluding the user's speech, to the communication interface. The electronic device (200) can transmit the converted first TTS voice signal to the other party's electronic device through the communication interface. If the mute option for the user's speech is turned off, the electronic device (200) can transmit the user's speech and the converted first TTS voice signal to the communication interface. The electronic device (200) can transmit the user's speech and the converted first TTS voice signal to the other party's electronic device through the communication interface. As an example, when the user speaks, the electronic device (200) can transmit the user's speech to the other party's electronic device, and when the first TTS voice signal is generated, the electronic device (200) can transmit the first TTS voice signal. As an example, the electronic device (200) can record the user's speech when the user speaks. In addition, when the first TTS voice signal is generated, the electronic device (200) can simultaneously transmit the user's speech and the first TTS voice signal.
[0252] The electronic device (200) can convert the other party's speech in a second language into a third text and display it on the screen (2250). The electronic device (200) can receive the other party's speech in the second language from the other party's electronic device. The electronic device (200) can convert the received speech into a third text. For example, since the other party's language is English, the third text may be in English. The electronic device (200) can display the converted third text in the dialogue script.
[0253] The electronic device (200) can display a fourth text generated by translating a third text (or the other party's speech) into a first language on the screen (2260). For example, the first language may be a language set as the user's language. As an example, the electronic device (200) can translate the other party's speech in English (the other party's language) into Korean (the user's language) to generate the fourth text. The electronic device (200) can display the fourth text generated together with the third text. As an example, the fourth text may be displayed in an area adjacent to the third text.
[0254] The electronic device (200) can convert the fourth text into a second TTS voice signal (2270). For example, the converted second TTS voice signal may be in Korean.
[0255] The electronic device (200) can output voice data according to the status of the mute option of the other party's speech (2280). For example, if the mute option of the other party's speech is on, the electronic device (200) can output only the converted second TTS voice signal excluding the other party's speech. If the mute option of the other party's speech is off, the electronic device (200) can output the other party's speech and the converted second TTS voice signal. For example, when the other party speaks, the electronic device (200) can output the other party's speech, and when the second TTS voice signal is generated, the electronic device (200) can output the second TTS voice signal. For example, when the other party speaks, the electronic device (200) can record the other party's speech. In addition, when the second TTS voice signal is generated, the electronic device (200) can sequentially output the other party's speech and the second TTS voice signal. For example, the electronic device (200) may output the other party's speech at a second level of volume and output the second TTS voice signal at a third level of volume that is greater than the second level. In addition, the electronic device (200) may apply an image effect to the fourth text in synchronization with the time at which the second TTS voice signal is output.
[0256] For example, the electronic device (200) may include an option to select whether to transmit a second TTS voice signal to the other party. The second TTS voice signal may be a voice signal generated by translating the other party's speech into a first language (the user's language). As an example, if the transmission option for the second TTS voice signal is turned off, the electronic device (200) may not transmit the second TTS voice signal to the other party's electronic device. As an example, if the transmission option for the second TTS voice signal is turned on, the electronic device (200) may transmit the second TTS voice signal to the other party's electronic device.
[0257] The other party's electronic device can output the received second TTS voice signal. In addition, the other party's electronic device can output the user's speech and the first TTS voice signal translated from the electronic device (200). As an example, the other party's electronic device can output the user's speech at a volume level of 5, output the first TTS voice signal at a volume level of 6, which is greater than the 5th level, and output the second TTS voice signal at a volume level of 4, which is less than the 5th level.
[0258] For example, the electronic device (200) can output a voice signal in synchronization with the output time of the voice signal from the other party's electronic device. For example, when the other party's electronic device outputs the first TTS voice signal, the electronic device (200) can also output the first TTS voice signal. The electronic device (200) can output the second TTS voice signal in synchronization with the time at which the other party's electronic device outputs the second TTS voice signal. As an example, the electronic device (200) can output the other party's speech at a second level of volume, output the first TTS voice signal at a first level of volume lower than the second level, and output the second TTS voice signal at a third level of volume higher than the second level.
[0259] As an example, the electronic device (200) may include a communication interface (240), a display (230), a speaker (220), at least one processor (250), and a memory (210) that stores instructions executed by the at least one processor (250). The instructions stored in the memory (210) may be configured to cause the electronic device (200) to convert, during a call, an utterance of the other party in a first language into a first text and to display the first text on a script screen of the display. The instructions stored in the memory (210) may be configured to cause the electronic device (200) to translate the first text into a second language to generate a second text, and to display the second text together with the first text on the script screen. The instructions stored in the memory (210) may be configured to cause the electronic device (200) to convert the second text into a first text-to-speech (TTS) voice signal. The command stored in the memory (210) may be set to cause the electronic device (200) to identify the other party's speech and output the first TTS voice signal through the speaker (220) without outputting the other party's speech when the first mute option for the other party's speech is on. The command stored in the memory (210) may be set to cause the electronic device (200) to output the other party's speech and the first TTS voice signal through the speaker (220) when the first mute option for the other party's speech is off.
[0260] As an example, the command stored in the memory (210) may be configured to cause the electronic device (200) to convert a user's speech in a second language into a third text during a call and to display the third text on the script screen. The command stored in the memory (210) may be configured to cause the electronic device (200) to translate the third text into the first language to generate a fourth text, and to display the fourth text together with the third text on the script screen. The command stored in the memory (210) may be configured to cause the electronic device (200) to convert the fourth text into a second TTS voice signal. The command stored in the memory (210) may be configured to cause the electronic device (200) to transmit the second TTS voice signal to the other party's device through the communication interface (240) without transmitting the user's speech if a second mute option for the user's speech is turned on. The command stored in the memory (210) may be set to cause the electronic device (200) to transmit the user's speech and the second TTS voice signal to the other party's device through the communication interface (240) when the second mute option for the user's speech is in the off state.
[0261] As an example, the command stored in the memory (210) may be configured to cause the electronic device (200) to transmit the user's speech to the other party's device when the user speaks, if the second mute option is off, and to transmit the second TTS speech signal to the other party's device when the second TTS speech signal is converted.
[0262] As an example, the command stored in the memory (210) may be configured to cause the electronic device (200) to store the user's speech in the memory (210) when the user speaks, and to transmit the user's speech and the second TTS speech signal to the other party's device when the second TTS speech signal is converted, if the second mute option is in an off state.
[0263] As an example, the command stored in the memory (210) may be set to cause the electronic device (200) to exclude translation of the first text if the translation option of the other party's speech is off, and to exclude translation of the third text if the translation option of the user's speech is off.
[0264] As an example, the command stored in the memory (210) may be configured to cause the electronic device (200) to change the language of the other party from the first language to a third language according to a user input during a call with the other party. The command stored in the memory (210) may be configured to cause the electronic device (200) to download data of the third language if the data of the third language is not stored in the memory. The command stored in the memory (210) may be configured to cause the electronic device (200) to translate the user's speech into the third language to generate a fifth text, and to display the fifth text together with the third text on the script screen.
[0265] As an example, the command stored in the memory (210) may be set to cause the electronic device (200) to output the second TTS voice signal in synchronization with the time at which the first TTS voice signal is output from the other party's device.
[0266] As an example, the command stored in the memory (210) may be set to cause the electronic device (200) to output the other party's speech at a first level of volume and to output the first TTS voice signal at a second level of volume greater than the first level.
[0267] As an example, the command stored in the memory (210) may be set to cause the electronic device (200) to apply an image effect to the second text in synchronization with the time at which the first TTS voice signal is output.
[0268] As an example, the command stored in the memory (210) may be set to cause the electronic device (200) to delete the first text displayed on the script screen and display new text when the number of texts displayed exceeds the number of texts that can be displayed without scrolling in the entire area of the script screen.
[0269] As an example, a method performed by an electronic device including a screen may include, during a call, converting a speech of a counterpart in a first language into a first text and displaying the first text on a script screen. The method may include translating the first text into a second language to generate a second text, and displaying the second text together with the first text on the script screen. The method may include converting the second text into a first text-to-speech (TTS) voice signal. The method may include, when a first mute option for the counterpart's speech is turned on, identifying the counterpart's speech and outputting the first TTS voice signal without outputting the counterpart's speech. The method may include, when a first mute option for the counterpart's speech is turned off, outputting the counterpart's speech and the first TTS voice signal.
[0270] As an example, the method may include, during a call, converting a user's utterance in a second language into a third text and displaying the third text on the script screen. The method may include translating the third text into the first language to generate a fourth text, and displaying the fourth text together with the third text on the script screen. The method may include converting the fourth text into a second TTS voice signal. The method may include, if a second mute option for the user's utterance is turned on, transmitting the second TTS voice signal to a device of the other party without transmitting the user's utterance. The method may include, if the second mute option for the user's utterance is turned off, transmitting the user's utterance and the second TTS voice signal to the device of the other party.
[0271] As an example, the operation of transmitting to the other party's device may be such that when the second mute option is off, the user's speech may be transmitted to the other party's device when the user speaks, and the second TTS speech signal may be transmitted to the other party's device when the second TTS speech signal is converted.
[0272] As an example, the operation of transmitting to the other party's device may be such that when the second mute option is off, the user's speech may be stored in memory when the user speaks, and when the second TTS voice signal is converted, the user's speech and the second TTS voice signal may be transmitted to the other party's device.
[0273] As an example, the method may include an operation for excluding translation of the first text if the translation option for the other party's speech is turned off. The method may include an operation for excluding translation of the third text if the translation option for the user's speech is turned off.
[0274] As an example, the method may include an operation of changing the language of the other party from the first language to a third language based on the user's input during a call with the other party. The method may include an operation of downloading data for the third language if data for the third language is not stored. The method may include an operation of translating the user's speech into the third language to generate a fifth text, and displaying the fifth text together with the third text on the script screen.
[0275] As an example, the method may include an operation of outputting the second TTS voice signal in synchronization with a time at which the first TTS voice signal is output from the other party's device.
[0276] As an example, the operation of outputting the other party's speech and the first TTS voice signal may output the other party's speech at a first level of volume and output the first TTS voice signal at a second level of volume that is greater than the first level.
[0277] As an example, the method may include an operation of applying an image effect to the second text in synchronization with the time at which the first TTS voice signal is output.
[0278] As an example, a non-transitory computer-readable storage medium having recorded thereon a program for performing a method performed by an electronic device including a screen may include instructions for performing an operation of converting, during a call, a speech of a counterpart in a first language into a first text and displaying the first text on a script screen. The storage medium may include instructions for performing an operation of translating the first text into a second language to generate a second text, and displaying the second text together with the first text on the script screen. The storage medium may include instructions for performing an operation of converting the second text into a first text-to-speech (TTS) voice signal. The storage medium may include instructions for performing an operation of identifying the speech of the counterpart and outputting the first TTS voice signal without outputting the speech of the counterpart when a first mute option for the speech of the counterpart is turned on. The storage medium may include instructions for performing an operation of outputting the other party's speech and the first TTS voice signal when the first mute option for the other party's speech is off.
[0279] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0280] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0281] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0282] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0283] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0284] The effects of this document are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the above description.
Claims
1. In electronic devices, communication interface; display; speaker; at least one processor; and A memory storing instructions executed by at least one processor; The instructions stored in the above memory cause the electronic device to: During a call, convert the other party's speech in the first language into a first text and display said first text on the script screen of said display; Translating the first text into a second language to generate a second text, and displaying the second text together with the first text on the script screen; Converting the above second text into a first text-to-speech (TTS) voice signal, If the first mute option for the other party's speech is turned on, the other party's speech is identified and the first TTS voice signal is output through the speaker without outputting the other party's speech. An electronic device set to output the other party's speech and the first TTS voice signal through the speaker when the first mute option for the other party's speech is off.
2. In paragraph 1, The instructions stored in the above memory cause the electronic device to: During a call, convert the user's speech in the second language into a third text and display said third text on the script screen; Translating the third text into the first language to generate a fourth text, and displaying the fourth text together with the third text on the script screen; Convert the above fourth text into a second TTS voice signal, If the second mute option for the user's speech is turned on, the second TTS voice signal is transmitted to the other party's device through the communication interface without transmitting the user's speech, An electronic device configured to transmit the user's speech and the second TTS voice signal to the other party's device via the communication interface when the second mute option for the user's speech is off.
3. In paragraph 2, The instructions stored in the above memory cause the electronic device to: An electronic device configured to transmit the user's speech to the other party's device when the user speaks, and to transmit the second TTS speech signal to the other party's device when the second TTS speech signal is converted, when the second mute option is off.
4. In paragraph 2, The instructions stored in the above memory cause the electronic device to: An electronic device configured to store the user's speech in the memory when the user speaks, and transmit the user's speech and the second TTS speech signal to the other party's device when the second TTS speech signal is converted, when the second mute option is off.
5. In paragraph 2, The instructions stored in the above memory cause the electronic device to: An electronic device set to exclude translation of the first text if the translation option of the other party's speech is turned off, and to exclude translation of the third text if the translation option of the user's speech is turned off.
6. In paragraph 2, The instructions stored in the above memory cause the electronic device to: During a call with the other party, change the language of the other party from the first language to the third language according to the user's input, If the data of the third language is not stored in the memory, the data of the third language is downloaded, An electronic device configured to translate the user's utterance into the third language to generate a fifth text, and to display the fifth text together with the third text on the script screen.
7. In paragraph 2, The instructions stored in the above memory cause the electronic device to: An electronic device set to output the second TTS voice signal in synchronization with the time at which the first TTS voice signal is output from the other party's device.
8. In paragraph 1, The instructions stored in the above memory cause the electronic device to: An electronic device set to output the other party's speech at a first level of volume and output the first TTS voice signal at a second level of volume greater than the first level.
9. In paragraph 1, The instructions stored in the above memory cause the electronic device to: An electronic device set to apply an image effect to the second text in synchronization with the time at which the first TTS voice signal is output.
10. In paragraph 1, The instructions stored in the above memory cause the electronic device to: An electronic device set to delete previously displayed text and display new text when the number of texts displayed exceeds the number of texts that can be displayed without scrolling in the entire area of the above script screen.
11. A method performed by an electronic device including a screen, During a call, an action of converting the other party's speech in a first language into a first text and displaying the first text on a script screen; An action of translating the first text into a second language to generate a second text, and displaying the second text together with the first text on the script screen; An operation of converting the second text into a first text-to-speech (TTS) voice signal; If the first mute option for the other party's speech is turned on, an operation of identifying the other party's speech and outputting the first TTS voice signal without outputting the other party's speech; and A method comprising: an operation of outputting the other party's speech and the first TTS voice signal when the first mute option for the other party's speech is off; 12. In paragraph 11, During a call, an action of converting the utterance of a user of a second language into a third text and displaying said third text on the script screen; An action of translating the third text into the first language to generate a fourth text, and displaying the fourth text together with the third text on the script screen; An operation of converting the above fourth text into a second TTS voice signal; If the second mute option for the user's speech is turned on, the operation of transmitting the second TTS voice signal to the other party's device without transmitting the user's speech; and A method further comprising: transmitting the user's speech and the second TTS voice signal to the other party's device if the second mute option for the user's speech is off; 13. In paragraph 12, The action of transmitting to the other party's device is: A method of transmitting the user's speech to the other party's device when the user speaks, and transmitting the second TTS speech signal to the other party's device when the second TTS speech signal is converted, when the second mute option is off.
14. In paragraph 12, The action of transmitting to the other party's device is: A method of storing the user's speech in memory when the user speaks when the second mute option is off, and transmitting the user's speech and the second TTS speech signal to the other party's device when the second TTS speech signal is converted.
15. A non-transitory computer-readable storage medium having recorded thereon a program for performing a method performed by an electronic device including a screen, Instructions for performing, during a call, an action of converting the other party's utterance in a first language into a first text and displaying the first text on a script screen; Instructions for performing an action of translating the first text into a second language to generate a second text, and displaying the second text together with the first text on the script screen; Instructions for performing an operation of converting the second text into a first text-to-speech (TTS) voice signal; An instruction for identifying the other party's speech and outputting the first TTS voice signal without outputting the other party's speech, if the first mute option for the other party's speech is turned on; and A non-transitory computer-readable storage medium having recorded thereon a program for performing a method, the method comprising: instructions for performing an operation of outputting the counterpart's speech and the first TTS voice signal when the first mute option for the counterpart's speech is off;
Citation Information
Patent Citations
Translation method and electronic device
JP2023029846A
System and method for operating language training electronic device and real-time translation training apparatus operated thereof
KR1020110062738A
Electronic device shortening chat window scroll and method thereof
KR1020150008642A
User terminal apparatus and two-way translation method thereof
KR1020150025750A
Message translation system with smartphone and operating method thereof
KR1020160049947A