Electronic device and method for determining text-to-speech conversion output during translation
The electronic device uses ASR to identify sentence endpoints and perform real-time translation with text-to-speech conversion, addressing inefficiencies in existing translation technologies by ensuring seamless and synchronized output during calls.
Patent Information
- Application Number
- PCT/KR2024/016266
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-14
- Filing Date
- 2024-10-24
- Publication Date
- 2025-07-10
AI Technical Summary
Existing translation technologies in electronic devices struggle to seamlessly integrate text-to-speech conversion during real-time translation, particularly in calls, leading to inefficiencies and disruptions in communication.
An electronic device and method that utilize automatic speech recognition (ASR) to identify sentence endpoints based on pause intervals, translate text between languages, and perform text-to-speech conversion to generate synthesized speech before the utterance ends, ensuring smooth and synchronized output during translation.
Enables synchronized and uninterrupted text-to-speech conversion during translation, enhancing communication clarity and reducing interruptions in real-time translation processes.
Smart Images

Figure KR2024016266_10072025_PF_FP_ABST
Abstract
Description
How to Determine Text-to-Speech Output in Electronic Devices and Translation
[0001] Embodiments of the present invention relate to an electronic device and a method for determining text-to-speech output during translation.
[0002] Applications running on electronic devices (e.g., translation apps, voice assistants, messenger apps, and call apps) can provide translation (or interpretation) technology. These translation technologies can translate sentences entered by a user through a translation function UI and display the translated results on the electronic device's screen.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0004] Aspects of the present disclosure address at least the problems and / or disadvantages described above, and provide at least the advantages described below. Accordingly, aspects of the present invention provide an electronic device and method for determining text-to-speech output in translation.
[0005] Additional aspects will be partly set forth in the description which follows, and partly will be apparent from the description or may be learned by practice of the embodiments presented.
[0006] A method performed by an electronic device during a call is provided according to one embodiment. The method may include an operation in which the electronic device receives an utterance from a user of the electronic device through a microphone. The method may include an operation in which the electronic device performs ASR based on a voice signal corresponding to a portion of the utterance to generate a first text in a first language. The method may include an operation in which the electronic device identifies an endpoint of a sentence included in the first text based on one or more pause intervals associated with the first text. The method may include an operation in which the electronic device translates a portion of the first text corresponding to the sentence into a second text in a second language based on the identified endpoint of the sentence included in the first text. The method may include an operation in which the electronic device performs text-to-speech conversion on the second text. The method may include an operation in which the electronic device generates a synthesized sound corresponding to the portion of the utterance before the end of the utterance received from the user based on the text-to-speech conversion.
[0007] According to one embodiment, one or more non-transitory computer-readable storage media are provided storing one or more computer programs including computer-executable instructions that are executed by one or more processors of an electronic device to cause the electronic device to perform operations. The operations may include: the electronic device receiving an utterance from a user of the electronic device through a microphone; the operations may include: the electronic device performing ASR based on a voice signal corresponding to a portion of the utterance to generate a first text in a first language; the operations may include: the electronic device identifying an end point of a sentence included in the first text based on one or more pause sections associated with the first text; the operations may include: the electronic device translating a portion of the first text corresponding to the sentence into a second text in a second language based on the identified end point of the sentence included in the first text; the operations may include: the electronic device performing text-to-speech conversion on the second text. The above operations may include an operation in which the electronic device generates a synthesized sound corresponding to a portion of an utterance received from the user before the end of the utterance based on the text-to-speech conversion. An electronic device according to one embodiment is provided. The electronic device may include a microphone. The electronic device may include a memory storing one or more computer programs. The electronic device may include one or more processors communicatively coupled to the microphone and the memory. The one or more computer programs may include computer execution instructions.The computer-executable instructions, when collectively or individually executed by the one or more processors, may cause the electronic device to receive an utterance from a user of the electronic device through the microphone. The computer-executable instructions, when collectively or individually executed by the one or more processors, may cause the electronic device to perform ASR based on a speech signal corresponding to a portion of the utterance to generate a first text in a first language. The computer-executable instructions, when collectively or individually executed by the one or more processors, may cause the electronic device to identify an endpoint of a sentence included in the first text based on one or more pause intervals associated with the first text. The computer-executable instructions, when collectively or individually executed by the one or more processors, may cause the electronic device to translate a portion of the first text corresponding to the sentence into a second text in a second language based on the identified endpoint of the sentence included in the first text. The computer-executable instructions, when collectively or individually executed by the one or more processors, may cause the electronic device to perform text-to-speech conversion on the second text. The computer-executable instructions, when collectively or individually executed by the one or more processors, may cause the electronic device to generate, based on the text-to-speech conversion, a synthesized sound corresponding to a portion of an utterance received from the user before the end of the utterance.
[0008] Other aspects, advantages and salient features of the present disclosure will become apparent to those skilled in the art from the following detailed description of various embodiments of the present disclosure taken in conjunction with the accompanying drawings.
[0009] FIG. 1 is a block diagram of an electronic device within a network environment, according to one embodiment.
[0010] FIG. 2 is a block diagram illustrating an integrated intelligence system according to one embodiment.
[0011] FIG. 3 is a diagram showing a form in which relationship information between concepts and actions is stored in a database according to one embodiment.
[0012] FIG. 4 is a diagram illustrating a screen for processing voice input received through an intelligent app by an electronic device according to one embodiment.
[0013] FIG. 5 is a diagram for explaining a concept of outputting a synthesized sound according to text-to-speech conversion during translation during a call according to one embodiment.
[0014] FIG. 6 is a drawing for explaining an electronic device that executes a method for outputting synthesized sound according to text-to-speech conversion during translation during a call according to one embodiment.
[0015] FIG. 7 is a diagram illustrating a method for displaying results generated by executing a translation service according to one embodiment.
[0016] Figure 8 is a schematic block diagram of a translation service according to one embodiment.
[0017] FIG. 9 is a diagram illustrating an example of a method for determining the output of text-to-speech conversion during a call according to one embodiment.
[0018] FIG. 10 is a diagram illustrating an example of a method for determining the output of text-to-speech conversion during a call according to one embodiment.
[0019] Fig. 11 is a flowchart for explaining an operating method of an electronic device according to one embodiment.
[0020] Fig. 12 is a flowchart for explaining an operating method of an electronic device according to one embodiment.
[0021] FIG. 13 is a drawing for explaining an example of what an electronic device displays to a user during real-time translation according to one embodiment.
[0022] Throughout the drawings, identical or similar reference numerals are used to describe identical or similar components, features, and structures.
[0023] The following description, with reference to the accompanying drawings, is provided to facilitate a comprehensive understanding of various embodiments of the present invention as defined by the claims and their corresponding claims. While various specific details are included herein to facilitate this understanding, they are to be considered merely exemplary. Accordingly, those skilled in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. Furthermore, descriptions of well-known functions and structures may be omitted for clarity and brevity.
[0024] The terms and words used in the following description and claims are not limited to their bibliographical meanings, but are merely used by the inventors to facilitate a clear and consistent understanding of the present invention. Therefore, it will be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustrative purposes only and is not intended to limit the present disclosure as defined by the appended claims and their equivalents.
[0025] Unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" should be understood to include the plural. Thus, for example, reference to "a component surface" includes reference to one or more of such surfaces.
[0026] It should be understood that the blocks and combinations of flowcharts in each flowchart can be performed by one or more computer programs containing instructions. One or more computer programs may be stored entirely in a single memory device, or one or more computer programs may be divided into different parts stored in multiple different memory devices.
[0027] Any function or operation described herein may be processed by a single processor or a combination of processors. A single processor or a combination of processors is a circuit that performs processing, and is an application processor (AP) (e.g., a central processing unit (CPU)), a communication processor (CP) (e.g., a modem), a graphics processing unit (GPU)), a neural processing unit (NPU) (e.g., an artificial intelligence (AI) chip), a Wi-Fi chip, a Bluetooth® chip, a global positioning system (GPS) chip, a near field communication (NFC) chip, a connectivity chip, a sensor controller, a touch controller, a fingerprint sensor controller, a display driver integrated circuit (IC), an audio codec chip, a universal serial bus (USB) controller, a camera controller, an image processing IC, a microprocessor unit (MPU), a system on a chip (SoC), an IC, or a similar chip.
[0028]
[0029] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100), according to one embodiment.
[0030] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an external electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an external electronic device (104) or a server (108) via a network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the external electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0031] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0032] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0033] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0034] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0035] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0036] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0037] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0038] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., an external electronic device (102)) (e.g., a speaker or headphones) directly or wirelessly connected to the electronic device (101).
[0039] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0040] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) to an external electronic device (e.g., the external electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0041] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., an external electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0042] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0043] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0044] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0045] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0046] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., external electronic device (102), external electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as a plurality of separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0047] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., an external electronic device (104)), or a network system (e.g., a second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0048] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0049] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0050] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0051] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0052]
[0053] Referring to FIG. 2, an integrated intelligent system (20) of one embodiment may include an electronic device (201) (e.g., electronic device (101) of FIG. 1), an intelligent server (200) (e.g., server (108) of FIG. 1), and a service server (300) (e.g., server (108) of FIG. 1).
[0054] An electronic device (201) of one embodiment may be a terminal device (or electronic device) that can connect to the Internet, and may be, for example, a mobile phone, a smart phone, a personal digital assistant (PDA), a laptop computer, a TV, white goods, a wearable device, an HMD, or a smart speaker.
[0055] According to the illustrated embodiment, the electronic device (201) may include a communication interface (202) (e.g., interface (177) of FIG. 1), a microphone (206) (e.g., input module (150) of FIG. 1), a speaker (205) (e.g., audio output module (155) of FIG. 1), a display module (204) (e.g., display module (160) of FIG. 1), a memory (207) (e.g., memory (130) of FIG. 1), or a processor (203) (e.g., processor (120) of FIG. 1). The above-listed components may be operatively or electrically connected to each other.
[0056] The communication interface (202) of one embodiment may be configured to connect to an external device and transmit and receive data. The microphone (206) of one embodiment may receive sound (e.g., user speech) and convert it into an electrical signal. The speaker (205) of one embodiment may output the electrical signal as sound (e.g., voice).
[0057] The display module (204) of one embodiment may be configured to display an image or video. The display module (204) of one embodiment may also display a graphical user interface (GUI) of a running app (or application program). The display module (204) of one embodiment may receive touch input via a touch sensor. For example, the display module (204) may receive text input via a touch sensor in an on-screen keyboard area displayed within the display module (204).
[0058] In one embodiment, the memory (207) may store a client module (209), a software development kit (SDK) (208), and a plurality of apps (211). The client module (209) and the SDK (208) may constitute a framework (or solution program) for performing general-purpose functions. In addition, the client module (209) or the SDK (208) may constitute a framework for processing user input (e.g., voice input, text input, touch input).
[0059] The plurality of apps (211) stored in the memory (207) of one embodiment may be programs for performing a specified function. According to one embodiment, the plurality of apps (211) may include a first app (211_1) and a second app (211_2). According to one embodiment, each of the plurality of apps (211) may include a plurality of operations for performing a specified function. For example, the apps may include an alarm app, a message app, and / or a schedule app. According to one embodiment, the plurality of apps (211) may be executed by the processor (203) to sequentially execute at least some of the plurality of operations.
[0060] In one embodiment, the processor (203) can control the overall operation of the electronic device (201). For example, the processor (203) can be electrically connected to a communication interface (202), a microphone (206), a speaker (205), and a display module (204) to perform a specified operation.
[0061] The processor (203) of one embodiment may also execute a program stored in the memory (207) to perform a designated function. For example, the processor (203) may execute at least one of the client module (209) or the SDK (208) to perform the following operations for processing user input. The processor (203) may control the operations of multiple apps (211), for example, through the SDK (208). The following operations described as operations of the client module (209) or the SDK (208) may be operations executed by the processor (203).
[0062] The client module (209) of one embodiment can receive user input. For example, the client module (209) can receive a voice signal corresponding to a user utterance detected through the microphone (206). Alternatively, the client module (209) can receive a touch input detected through the display module (204). Alternatively, the client module (209) can receive a text input detected through a keyboard or a visual keyboard. In addition, the client module (209) can receive various forms of user input detected through an input module included in the electronic device (201) or an input module connected to the electronic device (201). The client module (209) can transmit the received user input to the intelligent server (200). The client module (209) can transmit status information of the electronic device (201) to the intelligent server (200) together with the received user input. The status information can be, for example, execution status information of an app.
[0063] The client module (209) of one embodiment may receive a result corresponding to the received user input. For example, the client module (209) may receive a result corresponding to the received user input if the intelligent server (200) can produce a result corresponding to the received user input. The client module (209) may display the received result on the display module (204). Additionally, the client module (209) may output the received result as audio through the speaker (205).
[0064] In one embodiment, the client module (209) can receive a plan corresponding to the received user input. The client module (209) can display the results of executing multiple operations of the app according to the plan on the display module (204). For example, the client module (209) can sequentially display the results of executing multiple operations on the display module (204) and output audio through the speaker (205). The electronic device (201) can, for example, display only some results of executing multiple operations (e.g., the result of the last operation) on the display module (204) and output audio through the speaker (205).
[0065] In one embodiment, the client module (209) may receive a request from the intelligent server (200) to obtain information necessary to produce a result corresponding to a user input. In one embodiment, the client module (209) may transmit the necessary information to the intelligent server (200) in response to the request.
[0066] In one embodiment, the client module (209) can transmit result information of executing multiple operations according to a plan to the intelligent server (200). The intelligent server (200) can use the result information to confirm that the received user input has been processed correctly.
[0067] In one embodiment, the client module (209) may include a voice recognition module. In one embodiment, the client module (209) may recognize voice inputs that perform limited functions through the voice recognition module. For example, the client module (209) may execute an intelligent app that processes voice inputs to perform organic actions based on a specified input (e.g., "Wake up!").
[0068] An intelligent server (200) of one embodiment can receive information related to a user voice input from an electronic device (201) via a communication network. According to one embodiment, the intelligent server (200) can convert data related to the received voice input into text (e.g., text data). According to one embodiment, the intelligent server (200) can generate a plan for performing a task corresponding to the user voice input based on the text.
[0069] In one embodiment, the plan may be generated by an artificial intelligence (AI) system. The AI system may be a rule-based system, a neural network-based system (e.g., a feedforward neural network (FNN) or a recurrent neural network (RNN)), or a combination of the aforementioned or another AI system. In one embodiment, the plan may be selected from a set of predefined plans or may be generated in real time in response to a user request. For example, the AI system may select at least one plan from a plurality of predefined plans.
[0070] In one embodiment, the intelligent server (200) can transmit the results according to the generated plan to the electronic device (201), or transmit the generated plan to the electronic device (201). In one embodiment, the electronic device (201) can display the results according to the plan on the display module (204). In one embodiment, the electronic device (201) can display the results of executing an operation according to the plan on the display module (204).
[0071] An intelligent server (200) of one embodiment may include a front end (215), a natural language platform (220), a capsule database (230), an execution engine (240), an end user interface (250), a management platform (260), a big data platform (270), or an analytic platform (280).
[0072] A front end (215) of one embodiment can receive user input from an electronic device (201). The front end (215) can transmit a response corresponding to the user input.
[0073] According to one embodiment, the natural language platform (220) may include an automatic speech recognition module (ASR module) (221), a natural language understanding module (NLU module) (223), a planner module (225), a natural language generator module (NLG module) (227), or a text to speech module (TTS module) (229).
[0074] An automatic speech recognition module (221) of one embodiment can convert data related to a voice input received from an electronic device (201) into text (e.g., text data). A natural language understanding module (223) of one embodiment can use the text of the voice input to determine a user's intent. For example, the natural language understanding module (223) can perform syntactic analysis or semantic analysis on user input in the form of text data to determine a user's intent. The natural language understanding module (223) of one embodiment can use linguistic features (e.g., grammatical elements) of morphemes or phrases to determine the meaning of words extracted from user input, and can match the meaning of the determined words to the intent to determine the user's intent. That is, the natural language understanding module (223) can obtain intent information corresponding to a user's utterance. The intent information can be information indicating the user's intent determined by interpreting text. The intent information can include information indicating an action (or function) that the user intends to execute using the device. Intention information can also be referred to as goal information. Slots can be detailed information related to intent information. Slots can be parameters required to perform actions based on user intent. Slots can also be variable information required to perform actions.
[0075] In one embodiment, the planner module (225) can generate a plan using the intent and parameters (e.g., slots) determined by the natural language understanding module (223). According to one embodiment, the planner module (225) can determine a plurality of domains necessary to perform a task based on the determined intent. The planner module (225) can determine a plurality of operations included in each of the plurality of domains determined based on the intent. According to one embodiment, the planner module (225) can determine parameters necessary to execute the determined plurality of operations or result values output by the execution of the plurality of operations. The parameters and the result values can be defined as concepts of a specified format (or class). Accordingly, the plan can include a plurality of operations and a plurality of concepts determined by the user's intent. The planner module (225) can determine the relationships between the plurality of operations and the plurality of concepts in a stepwise (or hierarchical) manner. For example, the planner module (225) can determine the execution order of a plurality of actions based on the user's intention, based on a plurality of concepts. In other words, the planner module (225) can determine the execution order of a plurality of actions based on parameters required for the execution of the plurality of actions and results output by the execution of the plurality of actions. Accordingly, the planner module (225) can generate a plan including association information (e.g., ontology) between a plurality of actions and a plurality of concepts. The planner module (225) can generate the plan using information stored in a capsule database (230) in which a set of relationships between concepts and actions is stored.
[0076] The natural language generation module (227) of one embodiment can convert specified information into text format. The information converted into text format may be in the form of natural language speech. The text-to-speech conversion module (229) of one embodiment can convert information in text format into information in speech format.
[0077] According to one embodiment, some or all of the functions of the natural language platform (220) may also be implemented in the electronic device (201).
[0078] The capsule database (230) may store information on the relationships between a plurality of concepts and actions corresponding to a plurality of domains. According to one embodiment, a capsule may include a plurality of action objects (or action information) and concept objects (or concept information) included in a plan. According to one embodiment, the capsule database (230) may store a plurality of capsules in the form of a concept action network (CAN). According to one embodiment, the plurality of capsules may be stored in a function registry included in the capsule database (230).
[0079] The capsule database (230) may include a strategy registry that stores strategy information necessary for determining a plan corresponding to a voice input. The strategy information may include reference information for determining a single plan when there are multiple plans corresponding to a user input. According to one embodiment, the capsule database (230) may include a follow-up registry that stores information on follow-up actions for suggesting follow-up actions to a user in a specified situation. The follow-up actions may include, for example, follow-up utterances. According to one embodiment, the capsule database (230) may include a layout registry that stores layout information of information output through the electronic device (201). According to one embodiment, the capsule database (230) may include a vocabulary registry that stores vocabulary information included in capsule information. According to one embodiment, the capsule database (230) may include a dialog registry that stores information on a dialogue (or interaction) with a user. The capsule database (230) can update stored objects through a developer tool. The developer tool may include, for example, a function editor for updating action objects or concept objects. The developer tool may include a vocabulary editor for updating vocabulary. The developer tool may include a strategy editor for creating and registering strategies that determine plans. The developer tool may include a dialog editor for creating conversations with users.The developer tool may include a follow-up editor that activates follow-up goals and allows editing of follow-up utterances that provide hints. The follow-up goals may be determined based on currently set goals, user preferences, or environmental conditions. In one embodiment, the capsule database (230) may also be implemented within the electronic device (201).
[0080] The execution engine (240) of one embodiment can produce a result using the generated plan. The end user interface (250) can transmit the produced result to the electronic device (201). Accordingly, the electronic device (201) can receive the result and provide the received result to the user. The management platform (260) of one embodiment can manage information used in the intelligent server (200). The big data platform (270) of one embodiment can collect user data. The analysis platform (280) of one embodiment can manage the quality of service (QoS) of the intelligent server (200). For example, the analysis platform (280) can manage the components and processing speed (or efficiency) of the intelligent server (200).
[0081] In one embodiment, a service server (300) may provide a service (e.g., food ordering or hotel reservation) specified to an electronic device (201). In one embodiment, the service server (300) may be a server operated by a library. Services of the service server (300), such as CP Service A (301) and CP Service B (302), may interact with the front end (210) of the intelligent server (200). The service server (300) in one embodiment may provide information to the intelligent server (200) for generating a plan corresponding to a received user input. The provided information may be stored in a capsule database (230). In addition, the service server (300) may provide result information according to the plan to the intelligent server (200).
[0082] In the integrated intelligence system (20) described above, the electronic device (201) can provide various intelligent services to the user in response to user input. The user input may include, for example, input via a physical button, touch input, or voice input.
[0083] In one embodiment, the electronic device (201) may provide a voice recognition service through an intelligent app (or voice recognition app) stored within the device. In this case, for example, the electronic device (201) may recognize a user utterance or voice input received through the microphone and provide the user with a service corresponding to the recognized voice input.
[0084] In one embodiment, the electronic device (201) may perform a designated action based on the received voice input, either alone or in conjunction with the intelligent server and / or service server. For example, the electronic device (201) may execute an app corresponding to the received voice input and perform a designated action through the executed app.
[0085] In one embodiment, when an electronic device (201) provides a service together with an intelligent server (200) and / or a service server (300), the electronic device (201) can detect a user's speech using the microphone (206) and generate a signal (or voice data) corresponding to the detected user's speech. The electronic device (201) can transmit the voice data to the intelligent server (200) using the communication interface (202).
[0086] An intelligent server (200) according to one embodiment may generate a plan for performing a task corresponding to a voice input received from an electronic device (201), or a result of performing an operation according to the plan, in response to a voice input. The plan may include, for example, a plurality of operations for performing a task corresponding to a user's voice input, and a plurality of concepts related to the plurality of operations. The concept may define parameters input to the execution of the plurality of operations, or result values output by the execution of the plurality of operations. The plan may include association information between the plurality of operations and the plurality of concepts.
[0087] An electronic device (201) of one embodiment can receive the response using a communication interface (202). The electronic device (201) can output a voice signal generated within the electronic device (201) to the outside using the speaker (205), or can output an image generated within the electronic device (201) to the outside using the display module (204).
[0088]
[0089] FIG. 3 is a diagram showing a form in which relationship information between concepts and actions is stored in a database according to various embodiments.
[0090] The capsule database (e.g., the capsule database (230) of FIG. 2) of the intelligent server (e.g., the intelligent server (200) of FIG. 2) may store capsules in the form of a CAN (concept action network) (400). The capsule database may store operations for processing tasks corresponding to a user's voice input and parameters necessary for the operations in the form of a CAN (concept action network).
[0091] The capsule database may store a plurality of capsules (capsule (A) (401), capsule (B) (404)) corresponding to each of a plurality of domains. In one embodiment, one capsule (e.g., capsule (A) (401)) may correspond to one domain (e.g., location (geo)). In addition, one capsule may correspond to at least one service provider (e.g., CP 1 (402) or CP 2 (403)) for performing a function for a domain related to the capsule. In one embodiment, one capsule may include at least one operation (410) and at least one concept (420) for performing a designated function. CAN (400) may store other information such as CP 3 (406). In addition, capsule B (404) may correspond to a service provider (e.g., CP 4 (405)).
[0092] The above natural language platform (e.g., the natural language platform (220) of FIG. 2) can generate a plan for performing a task corresponding to a received voice input using capsules stored in a capsule database. For example, the planner module of the natural language platform (e.g., the planner module (225) of FIG. 2) can generate a plan using capsules stored in a capsule database. For example, a plan (407) can be generated using actions (4011, 4013) and concepts (4012, 4014) of capsule A (401) and actions (4041) and concepts (4042) of capsule B (404).
[0093]
[0094] FIG. 4 is a diagram illustrating a screen for processing voice input received through an intelligent app by an electronic device according to various embodiments.
[0095] The electronic device (201) can execute an intelligent app to process user input through an intelligent server (e.g., the intelligent server (200) of FIG. 2).
[0096] According to one embodiment, on screen 310, when the electronic device (201) recognizes a designated voice input (e.g., wake up!) or receives an input via a hardware key (e.g., a dedicated hardware key), the electronic device (201) may execute an intelligent app for processing the voice input. For example, the electronic device (201) may execute the intelligent app while the schedule app is running. According to one embodiment, the electronic device (201) may display an object (e.g., an icon) (311) corresponding to the intelligent app on the display module (204) (e.g., the display module (160) of FIG. 1, the display module (204) of FIG. 2). According to one embodiment, the electronic device (201) may receive a voice input by a user's speech. For example, the electronic device (201) may receive a voice input such as "Tell me my schedule for this week!" According to one embodiment, the electronic device (201) may display a user interface (UI) (313) (e.g., an input window) of an intelligent app displaying text (e.g., text data) of a received voice input on a display module (204).
[0097] According to one embodiment, on the 320 screen, the electronic device (201) may display a result corresponding to the received voice input on the display module (204). For example, the electronic device (201) may receive a plan corresponding to the received user input and display 'this week's schedule' on the display module (204) according to the plan.
[0098]
[0099] FIG. 5 is a diagram for explaining a concept of outputting a synthesized sound according to text-to-speech conversion during a call translation according to one embodiment.
[0100] Referring to FIG. 5, according to one embodiment, an electronic device (501) (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2) may provide a translation function (or an interpretation function) during a call with an external electronic device (601). During a call, the electronic device (501) may translate in real time a voice of a first user of the electronic device (501) that is input and / or a voice of a second user of the external electronic device (601) that is received. The electronic device (501) may display a text translated from the first user's voice on a display module (595) (e.g., a screen) of the electronic device (501), convert the text translated from the first user's voice into text-to-speech, generate a synthesized sound, and provide the synthesized sound to the external electronic device (601). Additionally, the electronic device (501) can display a text translated from the second user's voice on the display module (595) of the electronic device (501), convert the text translated from the second user's voice into text-to-speech to generate a synthesized sound, and provide the synthesized sound to the first user.
[0101] According to one embodiment, during a call, when translating a voice signal according to a user's speech in real time, the electronic device (501) may pause the speech at a certain point, convert the translated text up to the pause point into text-to-speech, generate a synthesized sound, and output the synthesized sound. For convenience of explanation, it is assumed that the user of the electronic device (501) utters "I'm going to invite Jane to a birthday party on Friday evening. Can you give me Jane's contact information?" in a first language. The electronic device (501) may receive a voice signal according to the speech of the user of the electronic device (501) in real time and display (610, 620, 630) a real-time translated sentence on the display module (595) together with the ASR result. The electronic device (501) may convert speech from "I'm going to invite Jane to a birthday party on Friday evening. Can you give me Jane's contact information?" to a text translated in real time into a second language (e.g., "I'm going to invite to Jane at a birthday party on Friday evening") and output the synthesized sound. The electronic device (501) may display an indicator (640) (e.g., UI) for controlling the output of a synthesized voice at a time when the text is to be converted into speech and output. The electronic device (501) may mix the speech signal of “I’m going to invite Jane to a birthday party on Friday evening” with the synthesized voice of “I’m going to invite to Jane at a birthday party on Friday evening” and transmit the result to an external electronic device (601). The electronic device (501) may mix the speech signal of “I’m going to invite Jane to a birthday party on Friday evening” with the synthesized voice of “I’m going to invite to Jane at a birthday party on Friday evening” at a time when the text is to be converted into speech and output."The synthesized sound of "I'm going to invite to Jane at a birthday party on Friday evening" can be mixed and automatically transmitted to the external electronic device (601), and can also be transmitted to the external electronic device (601) in response to a user input through an indicator (640) (e.g., UI). When the synthesized sound of "I'm going to invite to Jane at a birthday party on Friday evening" is created, the voice signal of "I'm going to invite Jane to a birthday party on Friday evening" and the synthesized sound of "I'm going to invite to Jane at a birthday party on Friday evening" can be automatically mixed and transmitted to the external electronic device (601), or manually transmitted to the external electronic device (601) according to a user input input through the indicator (640) (e.g., UI). After the electronic device (501) generates and outputs the synthesized sound for "I'm going to invite to Jane at a birthday party on Friday evening," it can also process and transmit "Can you give me Jane's contact information?" to the external electronic device (601) in the above-described manner.
[0102]
[0103] FIG. 6 is a drawing for explaining an electronic device for executing a method of outputting a synthesized sound according to text-to-speech conversion during a call translation according to one embodiment, and FIG. 7 is a drawing for explaining a method for displaying a result generated by executing a translation service according to one embodiment.
[0104] Referring to FIGS. 6 and 7, according to one embodiment, an electronic device (501) may convert a translation text into a voice signal and output it when translating (e.g., real-time translation) during a call with an external electronic device (e.g., the external electronic device (601) of FIG. 5). At this time, the electronic device (501) may determine the output of text-to-speech conversion for the translation text (e.g., the output timing of text-to-speech conversion).
[0105] According to one embodiment, the electronic device (501) may be implemented as at least one of a smartphone, a tablet personal computer, a mobile phone, a speaker (e.g., an AI speaker), a video phone, an e-book reader, a desktop personal computer, a laptop personal computer, a netbook computer, a workstation, a server, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, a wearable device, a virtual reality (VR) device, or an augmented reality (AR) device.
[0106] According to one embodiment, the electronic device (501) may include a processor (520) (e.g., processor (120) of FIG. 1, processor (203) of FIG. 2), a memory (530) (e.g., memory (130) of FIG. 1, memory (207) of FIG. 2), an input module (591) (e.g., input module (150) of FIG. 1, microphone (206) of FIG. 2), an audio output module (593) (e.g., audio output module (155) of FIG. 1, speaker (205) of FIG. 2), a display module (595) (e.g., display module (160) of FIG. 1, display module (204) of FIG. 2), and an antenna module (597) (e.g., antenna module (197) of FIG. 1). The first signal processing module (541), the second signal processing module (545), the Tx mixer (543), the Rx mixer (547), and the translation service (550) may be executable by the processor (520) and may be configured as one or more of a program code including instructions that can be stored in the memory (530), an application, an algorithm, a routine, a set of instructions, or an artificial intelligence learning model. Tx may represent transmit or transmission, and Rx may represent receive or reception. One or more of the first signal processing module (541), the second signal processing module (545), the Tx mixer (543), the Rx mixer (547), and the translation service (550) may be implemented by hardware and / or a combination of hardware and software.
[0107] According to one embodiment, the memory (530) may include one or more memories. The instructions stored in the memory (530) may be stored in a single memory. The instructions stored in the memory (530) may be divided and stored in multiple memories. The instructions stored in the memory (530) may be individually or collectively executed by the processor (520) to cause the electronic device (501) to perform and / or control the method of outputting synthesized sound according to text-to-speech conversion during translation during a call, as described with reference to FIGS. 5 to 13 .
[0108] According to one embodiment, the processor (520) may be implemented as a circuit (e.g., a processing circuit) such as a system on chip (SoC) or an integrated circuit (IC). The processor (520) may include one or more processors. For example, the processor (520) may include a combination of one or more processors such as a CPU, a GPU, an MPU, an AP, and a CP. The instructions stored in the memory (520) may be individually or collectively executed by one processor to cause the electronic device (501) to perform and / or control the method for outputting a synthesized sound according to text-to-speech conversion during a call translation described with reference to FIGS. 5 to 13. The instructions stored in the memory (530) may be individually or collectively executed by a plurality of processors to cause the electronic device (501) to perform and / or control the method for outputting a synthesized sound according to text-to-speech conversion during a call translation described with reference to FIGS.
[0109] According to one embodiment, the electronic device (501) can perform a transmission / reception process (or a transmission / reception process) for a call with an external electronic device (e.g., the external electronic device (601) of FIG. 5). The transmission / reception process can include a Tx process (or a transmission process) that processes a voice signal according to a speech input by an input module (591) (e.g., a microphone) and an Rx process (or a reception process) that receives and processes a voice signal according to a speech from the external electronic device.
[0110] According to one embodiment, the Tx process may process a voice signal through a first signal processing module (541), a translation service (550), and a Tx mixer (543). The input module (591) may receive a voice signal according to speech in a first language. The first signal processing module (541) may perform signal processing on the voice signal received from the input module (591). The first signal processing module (541) may perform signal processing on the voice signal using at least one of microphone array processing (MAP), adaptive echo canceller (AEC), noise suppression (NS), and automatic gain control or adaptive gain control (AGC). The translation service (550) may receive a voice signal (e.g., an intermediate voice signal and / or an entire voice signal) processed by the first signal processing module (541) in real time, convert it into text in a first language, and translate the text in the first language into a second language. The translation service (550) can convert the result translated into a second language into a voice signal of the second language. The Tx mixer (543) can mix the voice signal processed by the first signal processing module (541) (e.g., the intermediate voice signal of the first language and / or the entire voice signal of the first language) and the voice signal of the second language at a set ratio to generate one output signal (e.g., an output audio signal). The output signal can be transmitted to an external electronic device (e.g., the external electronic device (601) of FIG. 5) that is in a call with the electronic device (501) through the antenna module (597). At this time, the mixing ratio can be x:y, where x+y satisfies 1, and x and y can be set. For example, the mixing ratio can include 1:0, 0:1 (e.g., when only one signal is output).
[0111] According to one embodiment, in the Tx process, the result processed by the translation service (550) may be displayed via the display module (595). For example, the result of a speech signal (e.g., an intermediate speech signal and / or a full speech signal of a first language) processed by the first signal processing module (541) being converted into text of a first language in real time and the result of a text of the first language being translated into a second language may be displayed via the display module (595).
[0112] According to one embodiment, the Rx process may process a voice signal through a second signal processing module (545), a translation service (550), and an Rx mixer (547). The antenna module (597) may receive a voice signal according to a second language utterance from an external electronic device (e.g., the external electronic device (601) of FIG. 5) that is in a call with the electronic device (501). The second signal processing module (545) may perform signal processing on the voice signal received from the antenna module (597). The second signal processing module (545) may perform signal processing on the voice signal using at least one of noise suppression (NS) and automatic gain control or adaptive gain control (AGC). The translation service (550) may receive a voice signal (e.g., an intermediate voice signal and / or a full voice signal) processed by the second signal processing module (545) in real time, convert it into text of a second language, and translate the text of the second language into a first language. The translation service (550) can convert the result translated into the first language into a voice signal in the first language. The Rx mixer (547) can mix the voice signal processed by the second signal processing module (545) (e.g., the intermediate voice signal of the second language and / or the entire voice signal of the second language) and the voice signal of the first language at a set ratio to generate one output signal (e.g., an output audio signal). The output signal can be output to the user of the electronic device (501) through the audio output module (593) (e.g., a speaker). At this time, the mixing ratio can be x:y, where x+y satisfies 1, and x and y can be set. For example, the mixing ratio can include 1:0, 0:1 (e.g., when only one signal is output).
[0113] According to one embodiment, in the Rx process, the result processed by the translation service (550) may be displayed via the display module (595). For example, the result of a speech signal (e.g., an intermediate speech signal and / or a full speech signal of a second language) processed by the second signal processing module (545) being converted into text of a second language in real time and the result of a text of the second language being translated into a first language may be displayed via the display module (595).
[0114] According to one embodiment, the electronic device (501) can process the Tx process and the Rx process in various ways. For example, the electronic device (501) can sequentially process the Tx process and then process the Rx process according to the input order of the voice signals. In addition, if the electronic device (501) receives a voice signal from an external electronic device (e.g., the external electronic device (601) of FIG. 5) while processing the Tx process, the electronic device (501) can process the Rx process after completing the processing of the Tx process or can process the Rx processes simultaneously (or in parallel). In addition, if the electronic device (501) receives a voice signal from a user of the electronic device (501) while processing the Rx process, the electronic device (501) can process the Tx process after completing the processing of the Rx process or can process the Tx processes simultaneously (or in parallel). When the electronic device (501) processes the Tx process and the Rx process simultaneously, the processing result of the Tx process (e.g., ASR result, translation result, indicator (640)) and the processing result of the Rx process (e.g., ASR result, translation result, indicator (640)) can be displayed simultaneously on the display module (595), and the two processing results can be displayed separately.
[0115]
[0116] Figure 8 is a schematic block diagram of a translation service according to one embodiment.
[0117] Referring to FIG. 8, according to one embodiment, the translation service (550) may be used for an application of the translation service (550) itself (e.g., a translation app). In addition, the translation service (550) may be used for an application (750) that requires the translation service (550). The application (750) may include one or more applications running on the electronic device (501) (e.g., applications that can use the translation service (550), such as a Call App, a Message App, a Note App, a video conferencing App, a recording App, and a chat App). The application (750) may use an API to transmit and receive information to use the translation service (550). For example, the application (750) may call an API to use a real-time translation service.
[0118] According to one embodiment, the first voice signal (710) and / or the second voice signal (730) may be processed through the translation service (550). The first voice signal (710) may be a signal that is processed by receiving an utterance in a first language spoken by a first user using the electronic device (501) through the input module (591) of the electronic device (501). For example, the first voice signal (710) may include a voice signal (e.g., an intermediate voice signal and / or a full voice signal) that is processed by the first signal processing module (541) in the Tx process. The second voice signal (730) may be a signal that is processed by receiving an utterance in a second language spoken by a second user who is on a call with the first user of the electronic device (501) through an external electronic device (e.g., the external electronic device (601)) through the antenna module (597) of the electronic device (501). The second voice signal (730) may include a voice signal (e.g., an intermediate voice signal and / or a full voice signal) processed by the second signal processing module (545) in the Rx process.
[0119] According to one embodiment, the translation service (550) may include a language pack (560). The translation service (550) may include a speech information extractor (571), an ASR module (572), a translator (577), a TTS output determiner (579), and a TTS module (580).
[0120] According to one embodiment, the language pack (560) can support languages for a real-time translation service provided by the translation service (550). The user can select the languages used by the first user (e.g., the call transmitter) and the second user (e.g., the call receiver) to use the real-time translation service during a call. In addition, the user can select the languages used based on people stored in the contacts and / or address book of the electronic device (501). The user can set the languages used in the settings screen. For example, when Jane (e.g., an English-speaking user) receives a call, if the user (e.g., a Korean-speaking user) wants to use the "real-time translation during a call" service, Jane's language can be set to English→Korea, and the user's language can be set to Korea→English. In this case, the language pack (560) can pre-store the languages it supports, and if there are no supported languages, it can download new ones. The language pack (560) may include a voice information extractor (571), an ASR module (572) (e.g., a first ASR module (573) and a second ASR module (575)), a translator (577), a TTS output determiner (579), and a TTS module (580) to be used according to the language set for the first user (e.g., a call transmitter) and the second user (e.g., a call receiver) performing the call.
[0121] According to one embodiment, the voice information extractor (571) can extract voice information from the voice signal (710 or 730). The voice information extractor (571) can receive the voice signal (710 or 730) in real time, extract voice information of the voice signal (710 or 730) from the voice signal (710 or 730), and output the voice signal (710 or 730) together with the extracted voice information to the ASR module (572) in real time. In addition, the voice information extractor (571) can output the extracted voice information to the TTS output determiner (579).
[0122] According to one embodiment, the voice information extractor (571) can extract voice information from the voice signal (710 or 730) through various methods such as voice activity detection (VAD) and / or endpoint detection (EPD). For example, the voice information can be determined through signal processing or statistical pattern recognition (classification) methods using acoustic information of the voice signal or feature information (e.g., information such as zero-crossing rate, energy, MFCC, and pitch). The voice information can include one or more combinations of information on a voice section, information on a pause section, start time of speech, information on the time of speech (e.g., start time of speech, end time of speech), intonation information (e.g., pitch and / or low pitch information), and ASR end time.
[0123] According to one embodiment, the speech information extractor (571) can determine a speech section (e.g., a speech signal section) on which ASR decoding is to be performed through VAD and / or EPD, and can determine a pause section (e.g., pause information, pause indicator) in which silence exists. The pause (rest) section can include a silence section. The pause section can include a long silence (or long silence section), a short silence (or short silence section), a short pause (or short pause section), and / or a long pause (or long pause section) determined according to the length of time in which silence exists.
[0124] In one embodiment, the pause interval may be determined prior to the ASR decoding operation. For example, the voice information extractor (571) may measure the energy of frames in units of approximately 10 msec, and determine that a pause interval (e.g., a short pause) exists when frames with energy lower than a reference value occur continuously for approximately 500 msec.
[0125] According to one embodiment, the pause interval may be determined by the ASR module (572) during the ASR decoding operation. The pause interval may be determined as one of the units (e.g., phonemes, syllables, or N-grams using the same) determined during ASR decoding. In this case, information about the pause interval may be additionally included in the speech information.
[0126] According to one embodiment, the voice information extractor (571) may be implemented in the ASR module (572).
[0127] According to one embodiment, the ASR module (572) may perform ASR (e.g., ASR decoding) on a voice signal (710 or 730) received in real time using voice information extracted in real time. For example, the ASR module (572) may perform ASR decoding (e.g., first-pass, second-pass decoding) on a voice signal from the start of speech to the occurrence of a pause period. Similarly, the ASR module (572) may perform ASR decoding on a voice signal from a pause period to the occurrence of another pause period, or from a pause period to the detected end of speech.
[0128] According to one embodiment, the ASR module (572) can be implemented with various algorithms (e.g., hidden markov model (HMM), weighted finite-state transducers (WFST), artificial neural networks (ANN), and support vector machines (SVM)). For example, the ASR module (572) can be implemented with an artificial neural network. The artificial neural network can be an RNN, a long short-term memory (LSTM), or a Transformer, and there is no limitation on the model method. It can be implemented in the form of an end-to-end model (E2E) that implements acoustic, pronunciation, and language models as a single network. For example, artificial neural network models such as RNN-T (RNN transduction), LAS (listen, attend, and spell), and ConformerT (transformer-based conformer) can be used as an E2E ASR. The ASR module (572) can be implemented as a streaming ASR capable of outputting intermediate recognition results during the user's speech input through the artificial neural network model-based ASR mentioned above.
[0129] According to one embodiment, the ASR module (572) may utilize multi-pass decoding to reduce the recognition result output delay time and improve speech recognition performance. For example, the ConformerT method including two-pass decoding may generate a plurality of primary candidate results (hypotheses) based on the ConformerT neural-net model in the first pass, and in the second pass, rescore (e.g., lattice rescoring, N-best reranking) the plurality of candidate results to determine and output the final result. This is an example of multi-pass decoding, and the present invention is not limited thereto. For example, during the first-pass operation, the optimal primary result may be first output based on the plurality of candidate results, and then modified and output as the final result after the second-pass decoding operation.
[0130] According to one embodiment, the ASR module (572) may include one or more ASR modules. For example, there may be one or more ASR modules (572) depending on the number and / or language of users performing the call (e.g., call transmitter, call receiver). The ASR module (572) may include a first ASR module (573) and a second ASR module (575). The first ASR module (573) may perform ASR on a voice signal (710), and the second ASR module (575) may perform ASR on a voice signal (730). The first ASR module (573) may support a first language of a first user (e.g., call transmitter), and the second ASR module (575) may support a second language of a second user (call receiver). The first ASR module (573) can perform ASR on a speech signal (710) of a first language and convert it into text of the first language. The second ASR module (575) can perform ASR on a speech signal (730) of a second language and convert it into text of the second language.
[0131] According to one embodiment, the translator (577) may receive the result of performing ASR from the ASR module (572) and translate the result of performing ASR into a target language based on a language supported by the ASR module (572). The result of performing ASR may be a text output by completing ASR decoding in real time by the ASR module (572). For example, in the case of two-pass decoding ASR, the result of performing ASR may be a text output by completing both first-pass decoding and second-pass decoding in real time, or may include a first-pass decoded result and a second-pass decoded final result.
[0132] According to one embodiment, in the case of two-pass decoding ASR, the translator (577) performs translation on the output text by completing both first-pass decoding and second-pass decoding in real time, and the translation result can be displayed through the display module (595). The translator (577) performs translation based on the first-pass decoded result, and the initial translation result can be displayed through the display module (595). When the final second-pass decoded ASR result is generated, the translator (577) performs translation thereon, and the final translation result can be displayed through the display module (595) instead of the initial translation result. All translation results can be displayed through the display module (595) in real time. In addition, the initial translation result or the temporary translation result may not be displayed through the display module (595), and only the final translation result may be displayed through the display module (595). The display of the translation result through the display module (595) can be set by a user (e.g., a user of the electronic device (501)).
[0133] According to one embodiment, the translator (577) may receive a result of performing ASR (e.g., text in a first language) from the first ASR module (573) and translate the result of performing ASR into a second language (e.g., text in the second language). In addition, the translator (577) may receive a result of performing ASR (e.g., text in a second language) from the second ASR module (575) and translate the result of performing ASR into a first language (e.g., text in the first language).
[0134] According to one embodiment, the TTS output determiner (579) can receive the result of performing ASR (e.g., text converted while performing ASR) from the ASR module (572). The TTS output determiner (579) can receive voice information (e.g., information about a short pause section, the start time of speech, the end time of speech, the end time of ASR) in real time. The voice information can be received directly from the voice information extractor (571), extracted from the voice information extractor (571) and received through the ASR module (572), or extracted and received from the ASR module (572) having a built-in voice information extraction function (e.g., the voice information extractor (571)). The voice information can include information about a portion of the voice signal (710 or 730) corresponding to the result of performing ASR, rather than the entire voice signal (710 or 730).
[0135] According to one embodiment, the TTS output determiner (579) may determine to output the translation result (e.g., the final translation result) up to the section where the ASR no longer changes by using the text (e.g., partial text) and voice information converted while performing ASR by converting the text to speech (TTS). For example, the TTS output determiner (579) may control the TTS module (580) to output the final translation result translated by converting the ASR output to speech (TTS) because the ASR output does not change. When two-pass decoding is performed in pause units, the TTS output determiner (579) may determine whether to convert the final translation result to speech based on at least one of the pause time, the speech start / end time, and the ASR end time.
[0136] According to one embodiment, the TTS output determiner (579) can identify the end point of a sentence (e.g., a complete sentence) in a text based on a pause section (e.g., a short pause section and / or an EPD) in the text being converted while performing ASR. The TTS output determiner (579) can identify the end point of a sentence in order to separate the sentence included in the text into complete sentences. The TTS output determiner (579) can separate the complete sentences in the text. The TTS output determiner (579) can determine to convert text to speech and output the sentence based on the point in time at which the end point of the sentence in the text is identified based on the pause section in the text being converted while performing ASR. The point in time at which the end point of the sentence is identified can include the point in time at which the complete sentence is separated from the text. The TTS output determiner (579) can control the TTS module (580) to convert text to speech and output the sentence. The TTS output determiner (579) can determine to output text-to-speech conversion for each sentence at each point in time when the end point of a sentence in the text is separated. The TTS output determiner (579) can control the TTS module (580) to output text-to-speech conversion for each sentence at each point when the end point of a sentence in the text is identified.
[0137] According to one embodiment, the TTS output determiner (579) can identify the end point of a sentence (e.g., a complete sentence) in a text being converted while performing ASR based on one or more combinations of information about a pause interval, token information coming after the pause interval, and punctuation information. The TTS output determiner (579) can separate a complete sentence from a text being converted while performing ASR based on one or more combinations of information about a pause interval, token information coming after the pause interval, and punctuation information. The TTS output determiner (579) can analyze the text being converted while performing ASR to determine which section of the text is "a sentence or not." Alternatively, the TTS output determiner (579) can determine how far to break up continuously input sentences. The TTS output determiner (579) can separate a sentence (e.g., a complete sentence) from a text. The TTS output determiner (579) determines whether a text is a sentence or not based on a pause section in the text or token information received after the pause section, or whether a text is a complete sentence or a sentence connected to a subsequent sentence based on subsequent token information, and can predict the symbols of "punctuation marks," "exclamation marks," and "question marks," and separate complete sentences from the text. The TTS output determiner (579) can also analyze question marks based on intonation information such as pitch and low pitch. For example, the TTS output determiner (579) can predict whether a punctuation mark or a question mark will be included in the last sentence based on the intonation information from “Have you eaten?” and “Have you eaten?” The TTS output determiner (579) can transmit the predicted punctuation mark information and the complete sentence together as text information to the translator (577). In addition to the complete sentence including the predicted punctuation mark information, the TTS output determiner (579) can also output voice information (e.g., including intonation information such as low and high pitch) to the translator (577).The translator (577) can translate a portion of the text corresponding to a sentence based on the identified end point of the sentence included in the text. The translator (577) can translate a result (e.g., a complete sentence including punctuation marks) received from the TTS output determiner (579) and transmit the translation result to the TTS module (580). The translator (577) can translate a result of ASR performance received from the ASR module (572) and / or a result (e.g., a complete sentence including punctuation marks) received from the TTS output determiner (579). The translation for the result of ASR performance received from the ASR module (572) may be a temporary translation result, and the translation for the result (e.g., a complete sentence including punctuation marks) received from the TTS output determiner (579) may be a final translation result. The temporary translation result and the final translation result may be displayed on the display module (595) in real time. The temporary translation result may be first displayed on the display module (595), and then the final translation result may be displayed on the display module (595) to replace the temporary translation result. In addition, the temporary translation result may not be displayed on the display module (595), and only the final translation result may be displayed on the display module (595). When a sentence is separated from the text, only the final translation result translated as a sentence unit may be displayed on the display module (595). The display of the translation result through the display module (595) may be set by the user (e.g., the user of the electronic device (501)).
[0138] According to one embodiment, the TTS output determiner (579) can transmit text information (e.g., predicted punctuation information and complete sentences) and / or control information to the TTS module (580). The control information may be for controlling the TTS module (580) to convert text to speech for complete sentences and output them. The TTS module (580) can generate and output text corresponding to the complete sentences translated from the translator (577) as synthesized speech under the control of the TTS output determiner (579). The text information output from the TTS output determiner (579) to the TTS module (580) may also serve as control information.
[0139] According to one embodiment, the TTS module (580) may receive a translated text corresponding to a complete sentence from the translator (577) (e.g., a text translated into a second language and / or a text translated into a first language), and may generate and output the received text as a synthetic voice through text analysis (581), prosody prediction (583), and vocoder (585). The TTS module (580) may generate and output the translated text corresponding to the complete sentence received from the translator (577) as a synthetic voice under the control of the TTS output determiner (579). The synthetic voice generated by the TTS module (580) corresponds to a portion of the voice signal (710) or the voice signal (730), and the TTS module (580) may output the synthetic voice while the utterance corresponding to the voice signal (710) or the voice signal (730) is not finished. For example, in a situation where the entire utterance corresponding to the voice signal (710) or the voice signal (730) is not finished, if a punctuation mark (e.g., a punctuation mark such as a period, a comma, an exclamation point, or a question mark) is detected in a part of the entire utterance (e.g., a partial utterance) and it is determined to be a complete sentence, reproduction of a synthesized sound corresponding to the complete sentence may be started before the end of utterance (EPD) of the entire utterance. In the case of the Tx process, the synthesized sound may be transmitted to an external electronic device (e.g., the external electronic device (601)) that is in a call. In the case of the Rx process, the synthesized sound may be output to the user of the electronic device (501) through the audio output module (593) (e.g., a speaker).
[0140]
[0141] FIG. 9 is a diagram illustrating an example of a method for determining the output of text-to-speech conversion during a call according to one embodiment.
[0142] In FIG. 9, it is assumed that a Tx process is processed when a user of an electronic device (501) utters "I'm going to invite Jane to my birthday party on Friday evening. Can you give me Jane's contact information?" during a call. The voice signal according to the user's utterance "I'm going to invite Jane to my birthday party on Friday evening. Can you give me Jane's contact information?" may be processed by a signal processing module (e.g., the first signal processing module (541) of FIG. 6) and input to an ASR module (e.g., the first ASR module (573) of FIG. 8).
[0143] The first ASR module (573) can perform ASR (e.g., ASR decoding) on a voice signal received in real time, and as soon as the ASR is completed, the partial texts “Friday evening” (911), “To the birthday party” (912), “Jane” (913), and “I will invite you” (914) as ASR results can be sequentially output to the translator (577) and the TTS output determiner (579). The first ASR module (573) can sequentially output “Jane’s contact information” (915) and “Can you tell me?” (916) to the translator (577) and the TTS output determiner (579). Additionally, the first ASR module (573) may sequentially output the partial texts “Friday evening”, “Friday evening birthday party”, “Jane to the Friday evening birthday party”, and “I will invite Jane to the Friday evening birthday party” to the translator (577) and the TTS output determiner (579) as soon as the ASR is completed.
[0144] The first ASR module (573) can output voice information along with the ASR result to the TTS output determiner (579). The voice information can be extracted by the voice information extractor (571) or extracted by the first ASR module (573) during ASR decoding. The first ASR module (573) can transmit information on pause sections (e.g., [SP], [EOS]) along with text information to the TTS output determiner (579). For example, the first ASR module (573) transmits to the TTS output determiner (579) together with text information information about the pause interval between “Friday evening” (911) and “Birthday party” (912) (e.g., [SP] (921)), information about the pause interval between “Birthday party” (912) and “Jane” (913) (e.g., [SP] (922)), information about the pause interval between “Jane” (913) and “I will invite” (914) (e.g., [SP] (923)), information about the pause interval between “I will invite” (914) and “Jane’s contact” (915) (e.g., [SP] (924)), information about the pause interval between “Jane’s contact” (915) and “Can you tell me?” (916) (e.g., [SP] (925)), and information about the pause interval after “I can tell you” (e.g., [EOS] (926)). It can be given. Preferably, the EPD should detect [EOS] after "I will invite you" (914), but if there is not enough pause in the utterance, the EPD may not detect [EOS].
[0145] The TTS output determiner (579) can determine the point at which the sentence no longer changes based on the text (e.g., partial text) and voice information (e.g., information on a pause section, intonation information) that are ASR results.
[0146] For example, if an EOS result is included in the ASR result, the TTS output determiner (579) may determine that the user's speech has ended and request the TTS module (580) to convert the translation of the ASR result into text-to-speech and output it.
[0147] For another example, if short pause information (e.g., [SP]) comes as information of a pause section in the ASR result, the TTS output determiner (579) can analyze the previous text at the time when SP is printed to determine that it is a complete sentence. The TTS output determiner (579) can perform sentence segmentation and punctuation prediction (or punctuation insertion) operations. A complete sentence can be defined as a sentence that has a grammatical and semantic structure sufficient to perform translation. The punctuation prediction operation (or punctuation insertion operation) determines at which point in the sentence punctuation marks such as a period, comma, question mark, or exclamation mark can be inserted, and it can be assumed that a complete sentence is determined based on the predicted punctuation point. The TTS output determiner (579) can analyze the previous text from the time when the pause section (e.g., [SP] (921)) is printed to the time when the pause section (e.g., [EOS]) is printed to determine that it is a complete sentence. The TTS output determiner (579) can determine that the previous text (e.g., "Friday evening" (911), "Friday evening birthday party" (911, 912), "Jane at the Friday evening birthday party" (911-913)) at the time when the pause sections (e.g., [SP](921), [SP](922), [SP](923)) are printed is not a complete sentence. In the previous text (e.g., "I will invite Jane to the Friday evening birthday party" (911-914)) at the time when the pause section (e.g., [SP](924)) is printed, the TTS output determiner can determine that the punctuation mark of the period can be included at the point "I will invite" (914), and thus "I will invite Jane to the Friday evening birthday party" (911-914)) is a complete sentence. If the TTS output decision unit (579) determines that the sentence entered so far is a complete sentence based on SP, it can request the TTS module (580) to convert the translation of the determined complete sentence into text-to-speech and output it.
[0148] Before the TTS output determiner (579) requests the TTS module (580) to output text-to-speech conversion, the display module (595) can display in real time the text converted by the first ASR module (573) (e.g., "Friday Evening" (911), "At a Birthday Party on Friday Evening" (911, 912), "Jane at a Birthday Party on Friday Evening" (911-913), "I'm going to invite Jane to a Birthday Party on Friday Evening" (911-914)) and the text translated by the translator (577) (e.g., "Friday Evening" (931), "At a Birthday Party on Friday Evening" (932), "Jane at a Birthday Party on Friday Evening" (933), "I'm going to invite to Jane at a Birthday Party on Friday Evening" (934)). The results displayed on the display module (595) may include temporary results (e.g., "Friday evening" (911), "Friday evening birthday party" (911, 912), "Jane at the Friday evening birthday party" (911-913)) and / or final results (e.g., "I will invite Jane to the Friday evening birthday party" (911-914)) that continuously change as the first ASR module (573) continuously streams out the ASR results. The point in time when the TTS output determiner (579) requests the output of the text-to-speech conversion to the TTS module (580) may be the point in time when the ASR results and / or the translation results (e.g., the translation of the ASR results) no longer change. At the point in time when the TTS module (580) requests the output of the text-to-speech conversion, an indicator (e.g., the indicator (640) of FIG. 5) (e.g., a UI) regarding the output of the text-to-speech conversion may be generated and displayed on the display module (595). The indicator is intended to control the output of the synthesized sound generated by text-to-speech conversion, and the user can control the output of the synthesized sound through the indicator.The indicator may include functions for controlling the speed of the sound of the synthesized sound (e.g., the speed of the sound output), volume, play, and stop. In addition, the indicator may further include a function for controlling the output of a signal that mixes the synthesized sound and a voice signal. When the TTS output determiner (579) separates a complete sentence, the TTS module (580) may automatically output the synthetic sound that is the result of converting the translation result of the complete sentence into text-to-speech (e.g., output to an external electronic device (601) and / or an audio output module (593)), or may output it manually according to a user input input through the indicator.
[0149] If the first ASR module (573) supports two-pass decoding, the first ASR module (573) may stream out temporary ASR results (e.g., "Friday evening" (911), "At a birthday party on Friday evening" (911, 912), "Jane at a birthday party on Friday evening" (911-913)) even while performing this first-pass decoding. At this time, the translator (577) may perform translation based on the temporary ASR results (e.g., "Friday evening" (911), "At a birthday party on Friday evening" (911, 912), "Jane at a birthday party on Friday evening" (911-913)) to perform temporary translation results (e.g., "Friday Evening" (931), "At a birthday party on Friday evening" (932), "Jane at a birthday party on Friday evening" (933)). The provisional translation result may be continuously updated and output to the display module (595). Even if the first ASR module (573) performs second-pass decoding up to the SP point to determine the final ASR result, the translation result by the translator (577) may not be final. A complete sentence determined through a sentence separation operation may include an ASR recognition unit divided into multiple SPs. Alternatively, an ASR recognition unit composed of SPs may include multiple complete sentences.
[0150]
[0151] FIG. 10 is a diagram illustrating an example of a method for determining the output of text-to-speech conversion during a call according to one embodiment.
[0152] In FIG. 10, it is assumed that the user of the external electronic device (601) utters "He is my friend who is a professor at CMU." (1010) during a call, and the Rx process is received and processed by the electronic device (501). The voice signal according to the utterance "He is my friend who is a professor at CMU." (1010) of the user of the external electronic device (601) may be signal processed by a signal processing module (e.g., the second signal processing module (545) of FIG. 6) and input to an ASR module (e.g., the second ASR module (575) of FIG. 8).
[0153] The second ASR module (575) can analyze the voice signal coming in in real time and convert it into text. The second ASR module (575) can sequentially output the sentence "He does my friends" (1020) as a result of the 1st pass decoding, and can output the finally recognized "He is my friend" (1031) as a result of the 2nd pass decoding instead of the 1st pass decoding result. The second ASR module (575) can correct "He does my friends" (1020) in the 2nd pass decoding to output "He is my friend" (1031).
[0154] The TTS output decision unit (579) can determine the point in time to convert text into speech and output it by additionally considering the context with the next sentence based on the token information coming in after the pause section.
[0155] For example, the TTS output determiner (579) may determine a short pause point (e.g., or) together with the 2nd pass decoding result (e.g., "He is my friend" (1031)) transmitted from the second ASR module (575). (1040)) After that, sentences can be separated by checking up to N word or token information. For example, N can be set to a natural number greater than or equal to 1.
[0156] The TTS output decider (579) can determine a complete sentence from "He is my friend who is a professor at CMU." (1060) to "he is my friend." (1030) based on the sentence segmentation operation, but can further consider the two (N=2) token information "who is" (1053) that come in later to decide whether to segment the complete sentence. Here, "who is" (1053) may be the first decoding result in the second ASR decoding section (e.g., the section for ASR decoding "who is a professor at CMU" (1023)). For example, since the two token information "Who is" (1053) can be connected to "He does my friends" to form one sentence, the TTS output decider (579) does not determine the output of the text-to-speech conversion and instead considers the short pause (e.g., or (1040)) or wait for EPD to determine the completed sentence section. The TTS output decider (579) can determine whether the sentence is completed or whether the sentence should be cut off even if it is not completed and request the TTS module (580) to convert the translation of the ASR result into text-to-speech and output it.
[0157] Additionally, the TTS output determiner (579) may be configured to determine the pause interval (e.g., or (1040)) You can use the cache when considering the context with the next sentence.
[0158] For example, the TTS output determiner (579) may determine a pause interval (e.g., a short pause or (1040)) The previous sentence "he is my friend" (1031) can be predicted to be "he is my friend." (1030) through punctuation prediction, and the punctuation mark is the punctuation mark "." (1033). This data is passed to the cache and stored until the next ASR section (e.g., the section for ASR decoding "who is a professor at CMU" (1023)), and then the sentence can be separated by additionally considering the context of the next sentence that comes in. Since the next sentence can be composed of "Who is a professor at CMU" (1051) when judged by the sentence itself, a "?" question mark (1057) can be added. However, the TTS output decider (579) may consider the context with the sentence (e.g., "He is my friend" (1031)) stored in the previous cache to separate the complete sentence "He is my friend who is a professor at CMU." (1060), decide to output the translation of this sentence by converting it into text-to-speech, and request the TTS module (580). At the point where it is decided to output it by converting it into text-to-speech, the synthesized voice of "He is my friend who is a professor at CMU." (1060), which is the translated version of "He is my friend who is a professor at CMU," (1070), may be output from the TTS module (580).
[0159] In addition, the TTS output determiner (579) can separate sentences by considering grammar rules in order to consider the context with the sentences that come in later. In the following, it is assumed that the user of the external electronic device (601) utters "He is my friend who is a professor at CMU. He is a smart guy." during a call and the utterance is received by the electronic device (501). If the user of the external electronic device (601) utters "He is my friend who is a professor at CMU. He is a smart guy." <sp>who is a professor at CMU. He is a smart guy. <eos>"If uttered as "He is my friend", the second ASR module (575) decode the first ASR decoding section. <sp>" and in the second ASR decoding section, "who is a professor at CMU. He is a smart guy. <eos>" can be decoded. The TTS output determiner (579) is a pause section <sp>(short pause) The previous sentence "he is my friend" may be predicted to be "he is my friend." with the punctuation mark "." through punctuation prediction and this data may be stored in the cache. The TTS output determiner (579) may segment "who is a professor at CMU. He is a smart guy." through a segmentation operation and, considering the context with the sentence stored in the cache (e.g., "He is my friend"), merge "who is a professor at CMU." with the sentence stored in the cache to determine it as a single complete sentence, "He is my friend who is a professor at CMU." The TTS output determiner (579) may determine to output the translation of "He is my friend who is a professor at CMU." by converting it into text-to-speech. In addition, the TTS output determiner (579) may determine to output the translation of "He is a smart guy." by converting it into text-to-speech by determining "He is a smart guy." as another complete sentence. The TTS module (580) can generate and output a synthetic sound for the translation of "He is my friend who is a professor at CMU." and then generate and output a synthetic sound for the translation of "He is a smart guy."
[0160] In addition, the TTS output determiner (579) can decide to wait for an EPD to be input and determine the sentence to be input afterward to separate the sentence and output the translation of the separated sentence by converting it into text-to-speech if it is determined that the sentence is to be connected using the token information that comes after it. At this time, the symbol “…..” can be inserted into the ASR result while waiting and displayed on the display module (595). When the user utters “I had rice…spaghetti yesterday,” the TTS output determiner (579) can decide to separate the sentence and output the translation of the separated sentence by converting it into text-to-speech by looking at the “spaghetti~” that comes after it even if an EPD after “I had rice yesterday” is input.
[0161] For example, the first pause interval (e.g., short pause interval) may be determined in a range of about 200 ms to 500 ms, and the second pause interval (e.g., EPD Time) may be determined in a range of about 500 ms to 2 sec, but may not be limited thereto.
[0162]
[0163] Fig. 11 is a flowchart for explaining an operating method of an electronic device according to one embodiment.
[0164] In Fig. 11, the electronic device (501) may be related to an Rx process for processing a voice signal received from an external electronic device (601) during a call.
[0165] Actions 1110 to 1190 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of each action (1110 to 1190) may be changed, and at least two actions may be performed in parallel.
[0166] Instructions stored in the memory (530) of the electronic device (501) can cause the electronic device (501) to perform operations 1110 to 1190 when executed by the processor (520).
[0167] According to one embodiment, in operation 1110, the electronic device (501) may receive a call from an external electronic device (601).
[0168] According to one embodiment, at operation 1120, the electronic device (501) may check whether a translation service (e.g., translation service (550)) is on when a call comes in or a call is made to a person designated for translation during a call, or a call is initiated to a new number.
[0169] According to one embodiment, in operation 1125, the electronic device (501) may display a UI regarding turning on the translation service (e.g., translation service (550)) so that the user can turn on the translation service during a call on the screen if the translation service (e.g., translation service (550)) is not turned on.
[0170] According to one embodiment, in operation 1130, the electronic device (501) may perform ASR on a speech signal (e.g., a speech signal of a second language) received from an external electronic device (601). The electronic device (501) may extract speech information from the speech signal before and / or during the ASR. The speech information may include one or more combinations of information about a speech segment, information about a pause segment, a start time of speech, information about a start time of speech (e.g., information about a start time of speech, information about a end time of speech), intonation information (e.g., information about a pitch and / or a low pitch), and an ASR end time.
[0171] According to one embodiment, in operation 1140, the electronic device (501) may perform sentence segmentation (or sentence classification) and / or sign prediction (e.g., determining the location of periods, question marks, exclamation marks, etc.) on the text based on the result of performing ASR (e.g., text in a second language converted while performing ASR) and voice information. Based on this, the electronic device (501) may determine whether or not it is time to output text-to-speech (TTS).
[0172] According to one embodiment, in operation 1150, the electronic device (501) may translate (e.g., translate into a first language) the result of performing ASR (e.g., text in a second language converted during performing ASR). The electronic device (501) may translate the results of performing ASR in parallel or sequentially.
[0173] According to one embodiment, in operation 1160, the electronic device (501) may determine whether a condition is applicable for outputting TTS synthesized sound. The condition for outputting TTS synthesized sound is a condition in which the translation result no longer changes (e.g., a complete sentence), and the electronic device (501) may determine whether the texts preceding the point in time when EOS comes or SP is printed are sentences or not, or determine the point in time to output text-to-speech (TTS) based on at least one piece of token information that comes in after EOS or SP is printed.
[0174] According to one embodiment, in operation 1170, the electronic device (501) may convert (or synthesize) the translated text into a synthesized voice through text-to-speech conversion if the condition is such that TTS synthesized voice output is possible.
[0175] According to one embodiment, in operation 1180, when a synthetic sound is output, the electronic device (501) may change the text display to indicate that the already displayed text is output and display it on the display module (595). At this time, the UI may change the color of the text according to the time when the synthetic sound is output or may display an indicator (e.g., indicator (640) of FIG. 5) on the display module (595) for controlling the display or output of the synthetic sound so as to indicate the speed at which the synthetic sound is output, but is not limited thereto.
[0176] According to one embodiment, in operation 1190, if the sentence is one for which TTS synthesis sound output is impossible, the electronic device (501) may display the ASR result and / or the translated text sentence on the display module (595). The displayed sentence or word may change variably depending on the time flow of the user speaking.
[0177]
[0178] Fig. 12 is a flowchart for explaining an operating method of an electronic device according to one embodiment.
[0179] FIG. 12 may relate to a Tx process for processing a voice signal according to a speech of a user of an electronic device (501) when the user of an electronic device (501) and an external electronic device (601) are talking, and an Rx process for processing a voice signal received from an external electronic device (601).
[0180] Actions 1210 to 1290 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of each action (1210 to 1290) may be changed, and at least two actions may be performed in parallel.
[0181] Instructions stored in the memory (530) of the electronic device (501) can cause the electronic device (501) to perform operations 1210 to 1290 when executed by the processor (520).
[0182] According to one embodiment, in operation 1210, the electronic device (501) may receive a call from an external electronic device (601).
[0183] According to one embodiment, in operation 1215, the electronic device (501) may check whether a translation service (e.g., translation service (550)) is on when a call comes in or a call is made to a person designated for translation during a call, or a call is initiated to a new number.
[0184] According to one embodiment, in operation 1217, the electronic device (501) may display a UI regarding turning on the translation service (e.g., translation service (550)) so that the user can turn on the translation service during a call on the screen if the translation service (e.g., translation service (550)) is not turned on.
[0185] According to one embodiment, in operation 1220, the electronic device (501) may perform ASR on the Tx voice. For example, the electronic device (501) may perform ASR on a voice signal according to a first language utterance of a user (e.g., a user of the electronic device (501)). The electronic device (501) may extract voice information from the voice signal before and / or during the ASR. The voice information may include one or more combinations of information on a voice section, information on a pause section, a start time of utterance, information on a time of utterance (e.g., information on a start time of utterance, information on a end time of utterance), intonation information (e.g., information on pitch and / or low pitch), and an ASR end time.
[0186] According to one embodiment, at operation 1230, the electronic device (501) may translate a result of performing ASR (e.g., text in a first language converted while performing ASR) into a second language.
[0187] According to one embodiment, in operation 1240, the electronic device (501) displays a translation result translated into a second language in real time on the display module (595), determines a time point at which a complete sentence is separated from the text being converted while performing ASR as a time point for outputting the text-to-speech conversion, and converts the translated text corresponding to the complete sentence (e.g., text translated into the second language) into text-to-speech. The text being converted while performing ASR may also be displayed on the display module (595) together with the translation result translated into the second language. The electronic device (501) may display an indicator (e.g., UI) for controlling the output of the synthesized sound generated according to the text-to-speech conversion on the display module (595).
[0188] According to one embodiment, in operation 1250, the electronic device (501) may mix a synthesized sound (e.g., a synthesized sound of a second language) and a voice signal according to the user's speech (e.g., a voice signal corresponding to the synthesized sound) at a certain ratio and transmit the mixed signal to the external electronic device (601). The mixing operation may be implemented as a weighted sum of the two voice signals. The mixing ratio may be changed by user input. The mixing ratio may include a ratio of 1:0 or 0:1. For example, the electronic device (501) may transmit only a voice signal according to the user's speech or only a synthesized sound.
[0189] According to one embodiment, in operation 1260, the electronic device (501) may perform ASR on the Rx voice. For example, the electronic device (501) may perform ASR on a voice signal according to a second language utterance received from an external electronic device (601). The electronic device (501) may extract voice information from the voice signal before and / or during the ASR. The voice information may include one or more combinations of information on a voice section, information on a pause section, a start time of utterance, information on a time of utterance (e.g., information on a start time of utterance, information on a end time of utterance), intonation information (e.g., information on pitch and / or low pitch), and an ASR end time.
[0190] According to one embodiment, at operation 1270, the electronic device (501) may translate a result of performing ASR (e.g., text in a second language converted while performing ASR) into a first language.
[0191] According to one embodiment, in operation 1280, the electronic device (501) displays the translation result translated into the first language in real time on the display module (595), determines the time point at which a complete sentence is separated from the text being converted while performing ASR as the time point at which to convert the text into speech and output it, and converts the translated text corresponding to the complete sentence (e.g., the text translated into the first language) into text. The text being converted while performing ASR may also be displayed on the display module (595) together with the translation result translated into the first language. The electronic device (501) may display an indicator (e.g., a UI) for controlling the output of the synthesized sound generated according to the text-to-speech conversion on the display module (595).
[0192] According to one embodiment, in operation 1290, the electronic device (501) can mix a synthesized sound (e.g., a synthesized sound of a first language) and a voice signal according to a user's speech of the external electronic device (601) (e.g., a voice signal corresponding to the synthesized sound) at a predetermined ratio and output the result to the user of the electronic device (501) through an audio output module (593) (e.g., a speaker). The mixing operation can be implemented as a weighted sum of the two voice signals. The mixing ratio can be changed by a user input. The mixing ratio can include a ratio of 1:0 or 0:1. For example, the electronic device (501) can transmit only a voice signal according to a user's speech of the external electronic device (601) or only a synthesized sound.
[0193]
[0194] FIG. 13 is a drawing for explaining an example of what an electronic device displays to a user during real-time translation according to one embodiment.
[0195] In Fig. 13, in a call between an electronic device (501) and a user and a user of an external electronic device (601), it is assumed that after the electronic device (501) and the user make a first utterance, the user of the external electronic device (601) makes a second utterance (e.g., a one-way call), or in the middle of the first utterance by the user of the electronic device (501), the user of the external electronic device (601) also makes a second utterance (e.g., a two-way call).
[0196] The electronic device (501) can receive a sentence uttered in a first language by a user of the electronic device (501), such as “I’m going to invite Jane to a birthday party on Friday evening. Can you give me Jane’s contact information?” The electronic device (501) can receive a voice signal in real time and display a real-time translated sentence on the screen together with the ASR result. The electronic device (501) can determine “I’m going to invite Jane to a birthday party on Friday evening” in “I’m going to invite Jane to a birthday party on Friday evening. Can you give me Jane’s contact information?” as a complete sentence and convert the text translated into a second language of “I’m going to invite Jane to a birthday party on Friday evening” (e.g., “I’m going to invite to Jane at a birthday party on Friday evening”) into speech to output a synthesized speech. The electronic device (501) can display text on the screen according to the timing at which the synthesized speech is output, and can, for example, display information such as the color, font thickness, inclination, size, and font of some of the text differently. The electronic device (501) may display a portion of the text displayed on the screen differently according to the timing of the synthetic sound being generated to indicate the current progress based on the generation and / or output of the synthetic sound. The electronic device (501) may display an indicator (e.g., UI) for controlling the output of the synthetic sound on the display module (595) at the timing of converting the text to speech and outputting it. The electronic device (501) may mix the voice signal of “I’m going to invite Jane to a birthday party on Friday evening” with the synthesized voice of “I’m going to invite to Jane at a birthday party on Friday evening” and transmit the result to the external electronic device (601). The electronic device (501) may output “I’m going to invite to Jane at a birthday party on Friday evening.”After generating and outputting a synthesized sound for "Can you give me Jane's contact information?", a synthesized sound can be generated and output for a text translated into a second language (e.g., Can you give me Jane's contact information?). The electronic device (501) can mix the voice signal of "Can you give me Jane's contact information?" with the synthesized sound of "Can you give me Jane's contact information?" and transmit the mixed sound to the external electronic device (601). The electronic device (501) can transmit the voice signal of "Can you give me Jane's contact information?" and the synthesized sound of "Can you give me Jane's contact information?" to the external electronic device (601) at different volume levels. The electronic device (501) can transmit the voice signal of "Can you give me Jane's contact information?" at a first volume level and the sound obtained by mixing the synthesized sound of "Can you give me Jane's contact information?" at a second volume level to the external electronic device (601). The first volume level may be lower than the second volume level, but is not limited thereto.
[0197] An electronic device (501) can receive "OK. Wait a minute. Instagram ID is happyJane." spoken in a second language by a user of an external electronic device (601). The electronic device (501) can receive a voice signal in real time and display a real-time translated sentence on the screen together with the ASR result. The electronic device (501) can determine "OK. Wait a minute." in "OK. Wait a minute. Instagram ID is happyJane." as a complete sentence and convert the text translated into a first language (e.g., "Okay. Just a moment") of "OK. Wait a minute." into speech to output a synthesized sound. The electronic device (501) can display text on the screen according to the timing of the synthesized sound, and can display information such as the color, font, incline, size, and font of some of the text differently, for example. The electronic device (501) can change and display some of the text displayed on the screen differently according to the timing of the synthesized sound to indicate the current progress based on the generation and / or output of the synthesized sound. The electronic device (501) may display an indicator (e.g., UI) on the display module (595) to control the output of the synthesized sound at the time of converting the text into speech and outputting it. The electronic device (501) may mix the voice signal of “OK. Wait a minute.” with the synthesized sound of “Okay. Wait a minute” and output it to the voice output module (593). After generating and outputting the synthesized sound for “Okay. Wait a minute,” the electronic device (501) may generate and output the synthesized sound for the text translated into the first language of “Instagram ID is happyJane.” (e.g., “Instagram ID is happyJane”). The electronic device (501) may mix the voice signal of “Instagram ID is happyJane.” with the synthesized sound of “Instagram ID is happyJane” and output it to the voice output module (593).
[0198] In FIG. 13, after the electronic device (501) and the user process the "I'm going to invite Jane to my birthday party on Friday night. Can you tell me Jane's contact information?" spoken in a first language, the external electronic device (601) processes the "OK. Wait a minute. Instagram ID is happyJane." spoken in a second language, but this is not limited thereto. The electronic device (501) can simultaneously display on the screen of the electronic device (501) the result of processing the "I'm going to invite Jane to my birthday party on Friday night. Can you tell me Jane's contact information?" spoken in a first language by the electronic device (501) and the user and the result of processing the "OK. Wait a minute. Instagram ID is happyJane." spoken in a second language by the user of the external electronic device (601), and can simultaneously display the results by dividing the screen of the electronic device (501) into an area of the electronic device (501) and the user and an area of the user of the external electronic device (601). When an electronic device (501) processes "OK. Wait a minute. Instagram ID is happyJane." spoken by a user of an external electronic device (601) in a second language while processing "I'm going to invite Jane to my birthday party on Friday night. Can you give me Jane's contact information?" spoken by a user of the electronic device (501) in a first language, the electronic device (501) can display text on the screen at the time when the synthesized sound is produced, and can display information such as the color, font, slant, size, and font of some of the text differently, for example. In this way, the current progress of processing "I'm going to invite Jane to my birthday party on Friday night. Can you give me Jane's contact information?" spoken by a user of the electronic device (501) in a first language and "OK. Wait a minute. Instagram ID is happyJane." spoken by a user of the external electronic device (601) in a second language are compared."The current status of the processing may be displayed.
[0199] The electronic device (501) can accumulate and display on the screen the processing results (e.g., translated sentences, ASR results, indicators) for the user's utterance of the electronic device (501) and / or the processing results (e.g., translated sentences, ASR results, indicators) for the user's utterance of the external electronic device (601). The accumulated results can be displayed on the screen based on the criteria for generating synthesized sounds. In addition, the electronic device (501) can accumulate and display the processing results for the user's utterance of the electronic device (501) in one bubble (e.g., speech bubble, memo), and can accumulate and display the processing results for the user's utterance of the external electronic device (601) in one bubble (e.g., speech bubble, memo). In addition, the electronic device (501) can display the processing results for the user's utterance of the electronic device (501) and / or the processing results for the user's utterance of the external electronic device (601) in separate bubbles based on the criteria for generating synthesized sounds, and then make them disappear from the screen.
[0200]
[0201] According to one embodiment, a method performed by an electronic device (e.g., electronic device 101 of FIG. 1, electronic device 201 of FIG. 2, electronic device 501 of FIG. 5) during a call may include an operation in which the electronic device receives an utterance from a user of the electronic device through a microphone. The method may include an operation in which the electronic device performs ASR based on a voice signal corresponding to a portion of the utterance to generate a first text in a first language. The method may include an operation in which the electronic device identifies an end point of a sentence included in the first text based on one or more pause sections associated with the first text. The method may include an operation in which the electronic device translates a portion of the first text corresponding to the sentence into a second text in a second language based on the identified end point of the sentence included in the first text. The method may include an operation in which the electronic device performs text-to-speech conversion on the second text. The method may include an operation in which the electronic device generates a synthetic sound corresponding to a portion of an utterance received from the user before the utterance ends, based on the text-to-speech conversion.
[0202] According to one embodiment, the method may further comprise transmitting the synthesized sound toward a counterpart device before the remaining portion of the utterance ends.
[0203] According to one embodiment, the method may further include an operation of determining a TTS conversion at each point in time when an end point of a sentence in the first text is identified.
[0204] According to one embodiment, the identifying action may include an action of identifying an end point of the sentence included in the first text based on a combination of one or more of information about the pause section, token information coming after the pause section, and punctuation information.
[0205] According to one embodiment, the method may further include an operation of mixing and outputting a voice signal corresponding to the sentence in the voice signal and the synthesized sound.
[0206] According to one embodiment, the method may further include an operation of displaying an indicator for controlling output of the synthesized sound.
[0207] According to one embodiment, the indicator may include a user interface (UI) for controlling one or more combinations of speed, volume, play, and stop of the synthesized sound.
[0208] According to one embodiment, the method may further include an operation of automatically outputting the synthesized sound when the synthesized sound is generated or an operation of outputting the synthesized sound according to a user input input through the indicator.
[0209] According to one embodiment, the method may further include an action of displaying a portion of text displayed on the display differently based on a point in time associated with the synthesized sound.
[0210] According to one embodiment, the displaying action may include an action of changing one or more combinations of the color, font, slant, size, and font of some text displayed on the display.
[0211] An electronic device according to one embodiment (e.g., an electronic device (101) of FIG. 1, an electronic device (201) of FIG. 2, an electronic device (501) of FIG. 5) may include a microphone (e.g., an input module (150) of FIG. 1, a microphone (206) of FIG. 2, an input module (591) of FIG. 6). The electronic device (101, 201, 501) may include a processor (e.g., a processor (120) of FIG. 1, a processor (203) of FIG. 2, a processor (520) of FIG. 6). The electronic device (101, 201, 501) may include a memory (e.g., a memory (130) of FIG. 1, a memory (207) of FIG. 2, a memory (530) of FIG. 6)) that stores one or more computer programs. The electronic device (101, 201, 501) may include one or more processors (e.g., a processor (150) of FIG. 1, a microphone (206) of FIG. 2, a microphone (591) of FIG. 6)) that are communicatively coupled with the microphone and the memory. The electronic device (101, 201, 501) may include a processor (120) of FIG. 1, a processor (203) of FIG. 2, and a processor (520) of FIG. 6. The one or more computer programs may include computer executable instructions. The computer executable instructions, when collectively or individually executed by the one or more processors (120, 203, 520), may cause the electronic device (101, 201, 501) to receive an utterance from a user of the electronic device through the microphone. The computer executable instructions, when collectively or individually executed by the one or more processors (120, 203, 520), may cause the electronic device (101, 201, 501) to perform ASR based on a voice signal corresponding to a portion of an utterance to generate a first text in a first language.The computer executable instructions, when collectively or individually executed by the one or more processors (120, 203, 520), may cause the electronic device (101, 201, 501) to identify an end point of a sentence included in the first text based on one or more pause sections associated with the first text. The computer executable instructions, when collectively or individually executed by the one or more processors (120, 203, 520), may cause the electronic device (101, 201, 501) to translate a portion of the first text corresponding to a sentence included in the first text into a second text in a second language based on the identified end point of the sentence included in the first text. The computer executable instructions, when executed by the one or more processors (120, 203, 520), may cause the electronic device (101, 201, 501) to perform text-to-speech conversion on the second text. The computer executable instructions, when executed collectively or individually by the one or more processors (120, 203, 520), may cause the electronic device (101, 201, 501) to generate a synthesized sound corresponding to a portion of an utterance received from the user before the end of the utterance, based on the text-to-speech conversion.
[0212] In one embodiment, the one or more computer programs may further include computer executable instructions. The computer executable instructions, when collectively or individually executed by the one or more processors (120, 203, 520), may cause the electronic device (101, 201, 501) to transmit the synthesized sound to a counterpart device before the remaining portion of the speech is finished.
[0213] According to one embodiment, the one or more computer programs may further include computer executable instructions. The computer executable instructions, when collectively or individually executed by the one or more processors (120, 203, 520), may cause the electronic device (101, 201, 501) to determine TTS conversion at each point in time when the end point of a sentence in the first text is identified.
[0214] According to one embodiment, the one or more computer programs may further include computer executable instructions. The computer executable instructions, when collectively or individually executed by the one or more processors (120, 203, 520), may cause the electronic device (101, 201, 501) to identify an end point of the sentence included in the first text based on a combination of one or more of information about the pause interval, token information coming after the pause interval, and punctuation information.
[0215] According to one embodiment, the one or more computer programs may further include computer executable instructions. The computer executable instructions, when collectively or individually executed by the one or more processors (120, 203, 520), may cause the electronic device (101, 201, 501) to mix and output a voice signal corresponding to the sentence in the voice signal and the synthesized sound.
[0216] According to one embodiment, the one or more computer programs may further include computer executable instructions. The computer executable instructions, when collectively or individually executed by the processor (120, 203, 520), may cause the electronic device (101, 201, 501) to display an indicator for controlling the output of the synthesized sound.
[0217] According to one embodiment, the indicator may include a user interface (UI) for controlling one or more combinations of speed, volume, play, and stop of the synthesized sound.
[0218] According to one embodiment, the one or more computer programs may further include computer executable instructions. The computer executable instructions, when collectively or individually executed by the one or more processors (120, 203, 520), may cause the electronic device (101, 201, 501) to automatically output the synthesized sound when generating the synthesized sound, or to output the synthesized sound according to a user input input through the indicator.
[0219] According to one embodiment, the one or more computer programs may further include computer executable instructions. The computer executable instructions, when collectively or individually executed by the processor (120, 203, 520), may cause the electronic device (101, 201, 501) to display one or more combinations of color, font thickness, slant, size, and font of some text displayed on the display differently based on a point in time associated with the synthesized sound.
[0220]
[0221] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0222] The various embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B, or C," "at least one of A, B, and C," and "at least one of A, B, or C" can each include any one of the items listed together in the corresponding phrase among the phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish the corresponding element from other corresponding elements and do not limit the corresponding elements in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as being “coupled” or “connected” to another component (e.g., a second component), with or without the terms “functionally” or “communicatively,” it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0223] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0224] Various embodiments of the present document may be implemented as software (e.g., a program) including one or more instructions stored in a storage medium (e.g., built-in memory or external memory) readable by a machine (e.g., an electronic device). For example, a processor (e.g., a processor) of the machine (e.g., an electronic device) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one instruction called. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' only means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.
[0225] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0226] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0227] While the present disclosure has been illustrated and described with reference to various embodiments, it will be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents.< / sp> < / eos> < / sp> < / eos> < / sp>
Claims
1. A method performed by an electronic device (101; 201; 501) during a call, An operation in which the electronic device receives a speech from a user of the electronic device through a microphone (150; 206; 591); An operation in which the electronic device performs ASR based on a speech signal corresponding to a portion of an utterance to generate a first text in a first language; An operation of the electronic device identifying an end point of a sentence included in the first text based on one or more pause sections associated with the first text; An action of the electronic device to translate a portion of the first text corresponding to a sentence in the first text into a second text in a second language based on the identified endpoint of the sentence included in the first text; The operation of the electronic device performing text-to-speech conversion on the second text; and The operation of the electronic device generating a synthetic sound corresponding to a portion of the utterance received from the user before the end of the utterance based on the text-to-speech conversion. A method comprising:
2. In paragraph 1, An action of transmitting the synthesized sound to a counterpart device before the remaining portion of the speech ends. A method further comprising:
3. In either of paragraphs 1 and 2, Action to determine TTS conversion at each point where the end point of a sentence in the above first text is identified A method further comprising:
4. In any one of paragraphs 1 to 3, A method wherein the operation of identifying the end point of the sentence is based on a combination of one or more of information about the pause section, token information coming after the pause section, and punctuation information.
5. In any one of paragraphs 1 to 4, An operation of mixing and outputting the voice signal corresponding to the sentence in the above voice signal and the synthetic sound A method further comprising:
6. In any one of paragraphs 1 to 5, An action to display an indicator for controlling the output of the above synthetic sound. A method further comprising:
7. In any one of paragraphs 1 to 6, The above indicators are, A method comprising a user interface (UI) for controlling one or more combinations of speed, volume, play, and stop of the synthesized sound.
8. In any one of paragraphs 1 to 7, An action for automatically outputting the synthesized sound when the synthesized sound is generated; or An action of outputting the synthesized sound according to the user input entered through the above indicator. A method further comprising:
9. In any one of paragraphs 1 to 8, An action to display a portion of the text displayed on the display differently based on the point in time associated with the above synthesized sound. A method further comprising:
10. In any one of paragraphs 1 to 9, The actions indicated above are: An action to change one or more combinations of the color, font, slant, size, and font of some of the text displayed on the above display. A method comprising:
11. In electronic devices (101; 201; 501), Mike (150; 206; 591); A memory (130; 207; 530) storing one or more computer programs; and One or more processors (120; 203; 520) communicatively coupled to said microphone and said memory; Including, The one or more computer programs include computer execution instructions, Computer execution instructions, when collectively or individually executed by said one or more processors (120; 203; 520), cause said electronic device (101; 201; 501) to: Receiving a speech from a user of said electronic device through said microphone, To generate the first text in the first language, ASR is performed based on the speech signal corresponding to the portion of the utterance, Identifying an end point of a sentence included in the first text based on one or more pause sections associated with the first text; Translate a portion of the first text corresponding to the sentence based on the identified end point of the sentence included in the first text into a second text in a second language, Perform text-to-speech conversion on the above second text, An electronic device (101; 201; 501) for generating a synthetic sound corresponding to a portion of an utterance received from the user before the end of the utterance based on the text-to-speech conversion.
12. In paragraph 11, The one or more computer programs further comprise computer execution instructions, Computer execution instructions, when collectively or individually executed by said one or more processors (120; 203; 520), cause said electronic device (101; 201; 501) to: An electronic device (101; 201; 501) for transmitting the synthesized sound toward a counterpart device before the remaining portion of the speech is finished.
13. In either of paragraphs 11 and 12, The one or more computer programs further comprise computer execution instructions, Computer execution instructions, when collectively or individually executed by said one or more processors (120; 203; 520), cause said electronic device (101; 201; 501) to: An electronic device (101; 201; 501) that determines TTS conversion at each point in time when the end point of a sentence in the first text is identified.
14. In any one of paragraphs 11 to 13, The one or more computer programs further comprise computer execution instructions, Computer execution instructions, when collectively or individually executed by said one or more processors (120; 203; 520), cause said electronic device (101; 201; 501) to: An electronic device (101; 201; 501) that identifies the end point of the sentence included in the first text based on a combination of one or more of information about the pause section, token information coming after the pause section, and punctuation information.
15. In any one of paragraphs 11 to 14, The one or more computer programs further comprise computer execution instructions, Computer execution instructions, when collectively or individually executed by said one or more processors (120; 203; 520), cause said electronic device (101; 201; 501) to: An electronic device (101; 201; 501) that mixes and outputs a voice signal corresponding to the sentence in the voice signal and the synthesized sound.
Citation Information
Patent Citations
Device and method of translating a language
KR101827773B1
Telephone service system and method supporting interpreting and translation
KR1020160097406A
Device and method of translating a language into another language
KR1020180020368A
Artificial Intelligence-Based Security Event Analysis System and Its Method Using Semi-Supervised Machine Learning
KR102089688B1
System and method for simultaneous multilingual dubbing of video-audio programs
US20200211565A1