Electronic device, method, and non-transitory computer-readable storage medium for outputting speech data
By adjusting speech output rate based on the amount of text data waiting for conversion, the system optimizes voice translation efficiency and timeliness, addressing inefficiencies in existing voice translation systems.
Patent Information
- Application Number
- PCT/KR2025/009649
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-06
- Filing Date
- 2025-07-04
- Publication Date
- 2026-01-08
AI Technical Summary
Existing voice translation systems do not effectively manage speech output rates based on the amount of text data waiting for text-to-speech conversion, leading to inefficiencies and potential delays in speech output.
The system determines a speech output rate for translated text data based on the amount of text data waiting for text-to-speech conversion, adjusting the speech output speed accordingly to optimize the conversion process.
This approach ensures efficient and timely speech output by adapting the speech rate to the amount of text data, enhancing user experience in voice translation applications.
Smart Images

Figure KR2025009649_08012026_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transitory computer-readable storage medium for outputting voice data
[0001] The following descriptions relate to electronic devices, methods, and non-transitory computer-readable storage media for outputting voice data.
[0002] An electronic device can convert a voice signal into text. The electronic device can convert text into a voice signal. The electronic device can translate text in a first language into text in a second language and then provide the translated text in the second language to the user.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.
[0004] In embodiments, an electronic device is provided. The electronic device includes at least one processor including a processing circuit; and a memory storing instructions, wherein the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to generate first text data for first speech data, generate second text data corresponding to the first text data through translation of the first text data from a first language to a second language, determine a speech output rate for the second text data based on an amount of text data waiting for text-to-speech conversion, perform the text-to-speech conversion based on the speech output rate, thereby generating second speech data for the second text data, and output the second speech data.
[0005] In embodiments, a method performed by an electronic device is provided. The method may include: generating first text data for first speech data; generating second text data corresponding to the first text data by translating the first text data from a first language to a second language; determining a speech output speed for the second text data based on an amount of text data waiting for text-to-speech conversion; generating second speech data for the second text data by performing the text-to-speech conversion based on the speech output speed; and outputting the second speech data.
[0006] In embodiments, an electronic device is provided. The electronic device may include at least one processor including a processing circuit; and a memory storing instructions. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate first text data for first speech data, generate second text data corresponding to the first text data through translation of the first text data from a first language to a second language, determine a speech output rate for the second text data based on a speech rate corresponding to a speech time of the first speech data and a text length of the first text data, perform text-to-speech conversion based on the speech output rate, thereby generating second speech data for the second text data, and output the second speech data.
[0007] In embodiments, a method performed by an electronic device is provided. The electronic device may include an operation of generating first text data for first voice data, an operation of generating second text data corresponding to the first text data through translation of the first text data from a first language to a second language, an operation of determining a voice output speed for the second text data based on a voice rate corresponding to a speech time of the first voice data and a text length of the first text data, an operation of generating second voice data for the second text data by performing the text-to-speech conversion based on the voice output speed, and an operation of outputting the second voice data.
[0008] In embodiments, an electronic device is provided. The electronic device includes at least one processor including a processing circuit; and a memory storing instructions, wherein the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to generate first text data based on first speech data, generate second text data corresponding to the first text data through translation of the first text data from a first language to a second language, determine a speech output rate for the second text data based on an amount of text data waiting for text-to-speech conversion, generate second speech data for the second text data by performing the text-to-speech conversion, and cause output of the second speech data, and at least one of generating second speech data for the second text data by performing the text-to-speech conversion and outputting the second speech data is performed based on the determined speech output rate.
[0009] In embodiments, a method performed by an electronic device is provided. The electronic device includes an operation of generating first text data based on first speech data, an operation of generating second text data corresponding to the first text data through translation of the first text data from a first language to a second language, an operation of determining a speech output speed for the second text data based on a speech time of the first speech data and a speech speed corresponding to a text length of the first text data, an operation of generating second speech data for the second text data by performing the text-to-speech conversion, and an operation of outputting the second speech data, wherein at least one of generating the second speech data for the second text data by performing the text-to-speech conversion and outputting the second speech data is performed based on the determined speech output speed.
[0010] In embodiments, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium can cause an electronic device to generate first text data for first speech data, generate second text data corresponding to the first text data through translation of the first text data from a first language to a second language, determine a speech output rate for the second text data based on an amount of text data waiting for text-to-speech conversion, perform the text-to-speech conversion based on the speech output rate, thereby generating second speech data for the second text data, and output the second speech data.
[0011] In embodiments, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium can cause an electronic device to generate first text data for first speech data, generate second text data corresponding to the first text data through translation of the first text data from a first language to a second language, determine a speech output rate for the second text data based on a speech rate corresponding to an utterance time of the first speech data and a text length of the first text data, and perform text-to-speech conversion based on the speech output rate to generate second speech data for the second text data, and output the second speech data.
[0012] In embodiments, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium generates first text data based on first speech data, generates second text data corresponding to the first text data through translation of the first text data from a first language to a second language, determines a speech output rate for the second text data based on a speech rate corresponding to an utterance time of the first speech data and a text length of the first text data, generates second speech data for the second text data by performing the text-to-speech conversion, and causes an electronic device to output the second speech data, wherein at least one of generating the second speech data for the second text data by performing the text-to-speech conversion and outputting the second speech data is performed based on the determined speech output rate.
[0013] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.
[0014] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.
[0015] FIG. 2A illustrates an example of speech language conversion for converting speech in a first language into speech in a second language, according to various embodiments.
[0016] FIG. 2b illustrates components of an electronic device for speech language conversion according to various embodiments.
[0017] FIG. 3 illustrates an example of a buffer for text-to-speech conversion according to various embodiments.
[0018] FIG. 4 illustrates an example of speech language conversion performed based on adaptive speech output rate according to various embodiments.
[0019] FIG. 5 illustrates components of a text-to-speech conversion unit for controlling speech output speed according to various embodiments.
[0020] FIG. 6 illustrates an operational flow of an electronic device for performing speech language conversion based on adaptive speech output rate according to various embodiments.
[0021] FIG. 7 illustrates an operation flow of an electronic device for adjusting a voice output speed according to a speech speed according to various embodiments.
[0022] Figures 8a to 8c illustrate examples of speech language conversion according to various embodiments.
[0023] FIG. 9 illustrates functional components of an electronic device utilizing a voice language conversion function according to various embodiments.
[0024] FIG. 10 illustrates an example of a screen in a conversation mode in an application for speech language conversion according to various embodiments.
[0025] FIGS. 11A and 11B illustrate examples of a user interface for entering a listening mode in an application for speech language conversion, according to various embodiments.
[0026] FIG. 12 illustrates an example screen of a listening mode of an application for voice language conversion according to various embodiments.
[0027] FIG. 13A illustrates an example of a screen of a conversation mode of an application for voice language conversion in a foldable type electronic device according to various embodiments.
[0028] FIG. 13b illustrates an example screen of a listening mode of an application for voice language conversion in a foldable type electronic device according to various embodiments.
[0029] FIG. 14 illustrates an example screen of a listening mode of an application for voice language conversion in a foldable type electronic device according to various embodiments.
[0030] FIG. 15 illustrates an example screen of a function for converting voice language during a call, according to various embodiments.
[0031] The terms used in this disclosure are used only to describe specific embodiments and may not be intended to limit the scope of other embodiments. The singular expression may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by those of ordinary skill in the art described in this disclosure. Terms defined in general dictionaries among the terms used in this disclosure may be interpreted as having the same or similar meaning as the meaning they have in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined in this disclosure. In some cases, even if a term is defined in this disclosure, it cannot be interpreted to exclude embodiments of the present disclosure.
[0032] The various embodiments of the present disclosure described below illustrate a hardware-based approach as an example. However, since the various embodiments of the present disclosure include techniques utilizing both hardware and software, the various embodiments of the present disclosure do not exclude a software-based approach.
[0033] In the following description, terms referring to signals (e.g., signal, information, message, signaling), terms referring to components that perform functions (e.g., code, function, instruction, module, unit, circuit, program, command), terms referring to data types (e.g., data, data segment, data part, data set), terms for operational states (e.g., step, operation, procedure), terms referring to data (e.g., packet, user stream, information, bit, symbol, codeword), etc. are examples for convenience of explanation. Therefore, the present disclosure is not limited to the terms described below, and other terms having equivalent technical meanings may be used. In addition, terms such as '... part', '... device', '... thing', '... body' used below may mean at least one shape structure or a unit that processes a function.
[0034] In addition, in the present disclosure, expressions such as "more than" or "less than" may be used to determine whether a specific condition is satisfied or fulfilled, but this is merely a description for expressing an example and does not exclude descriptions such as "more than" or "less than." A condition described as "more than" may be replaced with "more than," a condition described as "less than" may be replaced with "less than," and a condition described as "more than and less than" may be replaced with "more than and less than." In addition, hereinafter, "A" to "B" mean at least one of elements from A (including A) to B (including B). hereinafter, "C" and / or "D" mean at least one of "C" or "D," that is, including {"C", "D", "C" and "D"}.
[0035] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.
[0036] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0037] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculation, the processor (120) may store a command or data received from another component (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the command or data stored in the volatile memory (132), and store the resulting data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0038] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0039] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0040] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0041] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0042] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0043] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0044] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0045] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0046] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0047] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0048] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0049] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0050] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0051] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0052] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0053] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0054] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0055] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0056] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0057] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0058] FIG. 2A illustrates an example of speech language conversion for converting speech in a first language into speech in a second language, according to various embodiments. An electronic device (101) may provide a service for converting speech in a first language into speech in a second language (hereinafter, speech language conversion) (e.g., interpretation). For example, the electronic device (101) may execute an application for speech language conversion. For example, the electronic device (101) may execute a function for speech language conversion while executing another application.
[0059] Referring to FIG. 2A, speech-to-language conversion may refer to converting a speech signal (201) of a first language (e.g., English) into a speech signal (207) of a second language (e.g., Korean). The speech-to-language conversion may include speech-to-text conversion (211), translation (213), and text-to-speech conversion (215). The electronic device (101) may perform speech-to-text conversion (211). The electronic device (101) may obtain a speech signal (201) of a first language (e.g., English). The electronic device (101) may perform speech-to-text conversion (211) on the speech signal (201) of the first language. The speech-to-text conversion (211) may refer to converting at least a portion of the speech signal (201) into text. Both the speech signal and the text may correspond to the first language (e.g., English). A unit in which speech-to-text conversion (211) is performed may be referred to as a speech unit. For example, a speech unit may be a syllable or a word. The speech data of a speech signal may include a plurality of speech units. For example, a speech unit may include a speech signal section divided based on a silent section (e.g., a pause, a short pause, a long pause, a section where no speech exists, a section where only background noise exists). In this case, the speech signal section may be determined through EPD (end-point detection). The electronic device (101) may generate a text (203) in a first language through speech-to-text conversion (211) for the speech signal (201). The text (203) in the first language may include a text unit corresponding to each speech unit. For example, the text (203) in the first language may include a plurality of text units. As a non-limiting example, each speech unit may correspond to a plurality of text units. The electronic device (101) can generate text units of a first language through speech-to-text conversion (211) for a speech signal (201).For example, the electronic device (101) can generate text (203) in a first language through a speech-to-text conversion module (e.g., an automatic speech recognition (ASR) module). The output unit of the speech-to-text conversion module may be the text unit. As a non-limiting example, the unit of a speech signal input to the ASR module may be different from the text unit output from the ASR module. The number of text units output from the ASR module may be different from the number of speech units processed by the ASR module. For example, in streaming ASR, different processing units and output units may be identified through continuous updating of the recognized result.
[0060] The electronic device (101) can perform translation (213). The electronic device (101) can perform translation (213) on text (203) in a first language. The translation (213) may represent (or include) converting text in the first language (e.g., English) into text in a second language (e.g., Korean). The text (203) in the first language may include a plurality of text units. For example, each text unit may correspond to a speech unit, which is an output unit of speech-to-text conversion (211). The electronic device (101) can sequentially perform translation (213) on each of the text units. The electronic device (101) can generate text (205) in a second language through translation (213) on text (203) in the first language. For example, instead of performing translation (213) for each distinct character or word, the electronic device (101) may perform translation (213) by grouping the characters and / or words into sentence units or paragraph units. The output unit of the translation (213) may be larger than the output unit of the voice-to-text conversion (211). For example, the output unit of the translation (213) may be a sentence unit, and the output unit of the voice-to-text conversion (211) (or the input unit of the translation (213)) may be a word unit. The output unit of the translation (213) may be understood as a single text data. The electronic device (101) may generate text data of a second language through the translation (213). The text data may correspond to text (205) of the second language. For example, the electronic device (101) can obtain a text (205) of a second language corresponding to the text (203) of the first language through a translation module (e.g., a machine translator (MT) module). The output unit (e.g., a sentence, a paragraph) of the MT module may be different from the input unit (e.g., a character, a word, a phrase) of the MT module.
[0061] The electronic device (101) can perform text-to-speech conversion (215). The electronic device (101) can perform text-to-speech conversion (215) on text (205) in a second language. Text-to-speech conversion (215) can refer to converting text data into voice data. The text data can correspond to an output unit of translation (213). The text data can be a unit of a request for text-to-speech conversion (215). The electronic device (101) can perform text-to-speech conversion for each text data. In other words, the unit on which the text-to-speech conversion (215) is performed can be text data. Both the language of the voice data and the language of the text data can correspond to a second language (e.g., Korean). The electronic device (101) can generate voice data in the second language (e.g., Korean). For each request for text-to-speech conversion (215), voice data may be generated through text-to-speech conversion (215). The electronic device (101) may obtain voice data of a second language through text-to-speech conversion (215). The generated voice data may be output as a voice signal (207) of the second language. For example, the electronic device (101) may obtain the voice signal (207) of the second language through a text-to-speech conversion module (e.g., a text-to-speech (TTS) module). The electronic device (101) may output the voice signal (207) of the second language. As a non-limiting example, at least one of generating the voice data of the second language or outputting the voice signal (207) of the second language may be performed based on a voice output speed, and the voice output speed is determined based on an amount of text data of the second language waiting for text-to-speech conversion 215.
[0062] FIG. 2B illustrates components of an electronic device (e.g., electronic device (101)) for speech language conversion. While FIG. 2B illustrates components according to an application for speech language conversion, these components are merely exemplary and are not intended to limit the embodiments of the present disclosure. The electronic device illustrated in FIG. 2B can perform the speech language conversion (or aspects thereof) described above with respect to FIG. 2A.
[0063] Referring to FIG. 2B, the electronic device (101) (e.g., processor (120)) can execute an application for voice language conversion or a function for voice language conversion. For example, the electronic device (101) can execute a translation application. The translation application may include a function for voice language conversion. For example, the electronic device (101) can execute a function for voice language conversion while another application (e.g., a call application) is running. The electronic device (101) may include a control unit (230) (e.g., processor (120)) for controlling the application for voice language conversion or the function for voice language conversion. The control unit (230) may be a component (e.g., a module, a program, a code, an instruction) for controlling the voice-to-text conversion unit (251), the translation unit (253), and / or the text-to-speech conversion unit (255). For example, the control unit (230) can be understood as the operation of a processor (e.g., processor (120)) for executing codes of a program for the above-described voice language conversion or codes of a function of the above-described program.
[0064] The electronic device (101) may include a voice-to-text conversion unit (251). The voice-to-text conversion unit (251) may be a unit (module, function code, separate device, circuit, or set of instructions) configured to perform voice-to-text conversion (e.g., voice-to-text conversion (211)) in the electronic device (101). For example, the voice-to-text conversion unit (251) may be configured to generate first text data based on the first voice data. For example, the voice-to-text conversion unit (251) may include an ASR module. The voice-to-text conversion unit (251) may obtain a voice signal (220). The voice-to-text conversion unit (251) may include a plurality of voice units corresponding to the voice data of the voice signal (220) (or may perform voice-to-text conversion on the plurality of voice units). The voice signal (220) may be divided into a plurality of voice units. The speech-to-text conversion unit (251) can divide the input speech signal (220) into designated units (e.g., syllables, phrases, words, sentences). For example, the speech unit may include a speech signal section divided based on a silent section (e.g., a pause, a short pause, a long pause, a section without speech, a section with only background noise). In this case, the speech signal section may be determined through EPD (end-point detection). The unit may be referred to as a speech unit. The speech-to-text conversion unit (251) may generate a text unit corresponding to the divided speech unit. The text unit may correspond to a first language (e.g., English). The speech-to-text conversion unit (251) may provide text units (261) to the control unit (230). For example, a speech signal (220) of "Hello, I am going to visit your office this afternoon" may be input.If the speech-to-text conversion unit (251) divides the speech signal (220) into speech units distinguished by EPD, the speech-to-text conversion unit (251) can convert the divided speech units into text units. For example, the speech-to-text conversion unit (251) can generate text units corresponding to 'Hello', 'I am', 'going to', 'visit', 'your', 'office', 'this', and 'afternoon', respectively. The speech-to-text conversion unit (251) can output the generated text units (261) to the control unit (230).
[0065] The electronic device (101) may include a translation unit (253). The translation unit (253) may be a unit (module, function code, separate device, circuit, or set of instructions) configured to perform translation (e.g., translation (213)) in the electronic device (101). For example, the translation unit (253) may be configured to generate second text data corresponding to the first text data by translating the first text data from a first language to a second language. For example, the translation unit (253) may include an MT module. The electronic device (101) may generate second text data (263b) by translating the first text data (263a). The translation unit (253) may obtain the first text data (263a). The first text data (263a) may include a plurality of text units (e.g., text units (261)). The above multiple text units may be composed of a first language. The translation unit (253) may generate second text data (263b) corresponding to the multiple text units of the first text data (263a). The second text data (263b) may be composed of a second language. The translation unit (253) may provide the second text data (263b) to the control unit (230). Instead of performing the translation (213) for each distinct character or word, the translation unit (253) may perform the translation (213) by grouping the characters and / or words into sentence units or paragraph units. The output unit (e.g., sentence, paragraph) of the translation unit (253) may be the same as or different from the input unit (e.g., character, word) of the translation unit (253). For example, text units corresponding to 'Hello', 'I am', 'going to', 'visit', 'your', 'office', 'this', and 'afternoon' can be input into the translation unit (253) as first text data (263a). The translation unit (253) can generate "Hello. I am planning to visit your office this afternoon" as second text data (263b).The translation unit (253) can provide text data (263b) to the control unit (230).
[0066] The input unit of the translation unit (253) (e.g., each text unit of the first text data (263a)) may be the same as the output unit of the speech-to-text conversion unit (251) (e.g., each text unit of the text units (261)). For example, the control unit (230) may directly provide the text unit provided from the speech-to-text conversion unit (251) to the translation unit (253). As a non-limiting example, the input unit of the translation unit (253) may be different from the output unit of the speech-to-text conversion unit (251). The control unit (230) may also provide two or more text units provided from the speech-to-text conversion unit (251) together to the translation unit (253). The output unit (or translation unit) (e.g., the second text data (263b)) of the translation unit (253) may be different from the output unit of the speech-to-text conversion unit (251). For example, the translation unit (253) may be configured to collect input text units and perform translation (213). The translation unit (253) may output text data (263b) having a second language through the translation (213).
[0067] The control unit (230) may obtain text data (263b) from the translation unit (253), and then store the text data (263b) in a buffer (240) for text-to-speech conversion. In general, since the text speed input to the text-to-speech conversion unit (255) is faster than the output speed from the text-to-speech conversion unit (255) (or the resulting output speed of the text-to-speech conversion unit, for example, the audio output speed of the second voice data, is faster), buffering of the text may be required. The buffer (240) may be a space where text data waiting for text-to-speech conversion is stored. The buffer (240) may be referred to as a text queue, a text data queue, a TTS queue, a conversion waiting queue, a voice conversion waiting queue, a text buffer, a text data buffer, a TTS buffer, a conversion waiting buffer, a voice conversion waiting buffer, a text queue, a text data queue, a TTS queue, a conversion queue, a voice conversion queue, and / or equivalent technical terms therefor. For example, the buffer (240) may correspond to a FIFO (first input first output) type.
[0068] The electronic device (101) may include a text-to-speech conversion unit (255). The text-to-speech conversion unit (255) may be a unit (module, function code, separate device, circuit, or set of instructions) configured to perform text-to-speech conversion (e.g., text-to-speech conversion (215)) in the electronic device (101). For example, the text-to-speech conversion unit (255) may be configured to perform text-to-speech conversion to generate second voice data for second text data. For example, the text-to-speech conversion unit (255) may include a TTS module. The text-to-speech conversion unit (255) may obtain input text data. The text-to-speech conversion unit (255) may obtain text data from a buffer (240) for text-to-speech conversion. The text-to-speech conversion unit (255) may generate voice data corresponding to the text data. The text-to-speech conversion unit (255) can output voice data. The text-to-speech conversion unit (255) can output a voice signal (270) including voice data. Specifically, the text-to-speech conversion unit (255) can generate a sequence of synthesis units by analyzing text data. The sequence can be a set of synthesis units. A synthesis unit can be a unit that forms a sound. The text-to-speech conversion unit (255) can generate prosody information for each synthesis unit (e.g., a frame (e.g., a certain time interval, a set unit of data samples in signal processing), a phoneme, a syllable, a character, a word) of the sequence. The prosody information can represent information for reflecting voice characteristics such as rhythm, intonation, stress, and speed to a specified synthesis unit. The text-to-speech conversion unit (255) can generate a synthesis sound corresponding to a plurality of synthesis units based on the prosody information. The text-to-speech conversion unit (255) can be configured to output the synthesis sound. For example, the text-to-speech conversion unit (255) can generate a waveform corresponding to the synthesized sound.The above synthetic sound can correspond to a voice signal (270). The text-to-speech conversion unit (255) is described in detail through Fig. 5.
[0069] The process of speech-to-text conversion involves speech-to-text conversion, translation, and text-to-speech conversion. The larger the data size requiring conversion, the greater the potential for delay. This delay can reduce the real-time nature of the speech-to-text conversion service. For example, in a live translation service, assume a situation where there are multiple sentences requiring translation, such as a speech or sermon. The electronic device (101) may output speech in a second language (e.g., Korean) that corresponds to content significantly earlier than the content of the first language (e.g., English) currently heard by the user through the speaker. For example, there may be a significant delay between the time a portion of the first language speech is received by the electronic device (101) and the time the corresponding translated speech in the second language is output to the user of the electronic device (or output to the user of an external device connected to the electronic device). This may result in inconvenience to the user of the electronic device (101) (and / or the external electronic device). Additionally, if the speech of the first language is output at a faster rate than the speech of the second language, the buffer (e.g., buffer (240)) that stores the text data for text-to-speech conversion may fill up or start to consume a large amount of memory, which may affect the translation performance and / or the overall operation of the electronic device (101). To alleviate this problem, embodiments of the present disclosure propose a technique for adjusting the speech output rate to prevent or alleviate the inconvenience (and / or buffer issues). For example, the electronic device (101) may adjust the speech output rate depending on the length of the content that the user of the electronic device (101) is listening to. For example, the speech output rate may be adjusted depending on the difference between the content that the user of the electronic device (101) is listening to and the content that the electronic device (101) is outputting through the text-to-speech conversion unit (255).In order to determine the length of the content that the user is listening to or the difference above, text data waiting in a buffer for text-to-speech conversion (e.g., buffer (240)) can be used. When the speaker's speech ends, the voice output speed can be controlled so that the voice output through the electronic device (101) also ends at a similar time (e.g., the end time is within a threshold range), thereby improving the user's real-time experience. The voice output speed, which is a control element according to embodiments of the present disclosure, may be referred to as a sound source playback speed, a voice-to-text conversion output speed, a translation sound source playback speed, a translation playback speed, a converted voice speed, a converted speech speed, a converted playback speed, a converted sound source speed, a synthesized voice speed, a TTS output speed, a TTS synthesis speed, a TTS synthesis sound speed, a TTS playback speed, a TTS voice speed, and / or equivalent technical terms therefor. A faster TTS synthesis speed may mean a shorter overall length of the TTS synthesis sound. As the TTS synthesis speed increases, the overall length of the TTS synthesis can also become shorter.
[0070] The control unit (230) according to embodiments of the present disclosure may determine a voice output speed for second text data based on the amount of text data awaiting text-to-speech conversion. The control unit (230) according to embodiments of the present disclosure may identify text data pending in a buffer (240) for text-to-speech conversion. The buffer (240) for text-to-speech conversion may store text data awaiting conversion from text to speech (hereinafter, text-to-speech conversion). The control unit (230) may determine the amount of text data pending in the buffer (240) for text-to-speech conversion. According to one embodiment, the control unit (230) may determine the number of requests for text-to-speech conversion. Text data input into the buffer (240) for text-to-speech conversion may be understood as a request for text-to-speech conversion. For example, if the number of text data input into the buffer (240) is three, the number of requests for the text-to-speech conversion may be three. According to one embodiment, the control unit (230) may determine the length of text waiting for text-to-speech conversion. For example, the text length may indicate the number of characters included in the text. The control unit (230) may determine the number of characters included in the text data stored in the buffer (240) for text-to-speech conversion. The text length may be expressed in various ways other than the number of characters. For example, the text length may correspond to the number of characters. For example, the text length may correspond to the number of syllables. For example, the text length may correspond to the number of phrases. For example, the text length may correspond to the number of words. The control unit (230) may determine the voice output speed based on the amount of text data pending in the buffer (240) for text-to-speech conversion. The above voice output speed may indicate the speed at which voice data is output from the text-to-speech conversion unit (255).For example, the text-to-speech conversion unit (255) may be configured to adjust the duration of each synthesis unit based on the speech output speed. For example, the text-to-speech conversion unit (255) may be configured to adjust the total time length of speech data corresponding to the sum of the synthesis units based on the speech output speed. The electronic device (101) may control the speech output speed in the text-to-speech conversion unit (255) according to the amount of text data stored in the buffer (240) (e.g., the number of input texts, the number of requests, the number of characters in the input text).
[0071] The speed of the TTS synthesized sound may be related to the overall duration of the TTS synthesized sound. As described above, a faster TTS synthesized sound speed may mean a shorter overall duration of the TTS synthesized sound. However, the above relationship does not necessarily indicate that the speed of the TTS synthesized sound and the overall duration of the TTS synthesized sound are reciprocals of each other. In the present disclosure, changes in the speed or duration of the TTS synthesized sound may vary for each phoneme that constitutes the synthesized sound. In other words, even if the speed of the TTS synthesized sound increases by N times (N is an integer) (e.g., 2 times), it cannot be determined that the length of the TTS synthesized sound necessarily decreases by 1 / N times (e.g., 1 / 2 times).
[0072] FIG. 3 illustrates an example of a buffer (e.g., buffer (240)) for text-to-speech conversion according to various embodiments. Text data waiting for text-to-speech conversion may be stored in the buffer. The buffer may correspond to a FIFO type. The electronic device (101) may input a plurality of output requests to the buffer (240) based on an application or function for voice language conversion. Each output request of the buffer (240) may correspond to text data. While a TTS synthesis operation is performed or while a TTS synthesized sound is output, text data and / or requests thereof may be accumulated in the buffer for text-to-speech conversion.
[0073] Referring to FIG. 3, the buffer (240) can store first text data (311) (S1), second text data (312) (S2), third text data (313) (S3), and fourth text data (314) (S4). The first text data (311) (S1) contains five words (e.g., W 11 , W 12 , W 13 , W 14 , W 15 ) may be included. For example, five words may be formed with a total of 25 characters. For example, the first text data (311) may be 'Device could improve your life'. The second text data (312) (S2) may be composed of seven words (e.g., W 21 , W 22 , W 23 , W 24 , W 25 , W 26 , W 27 ) may be included. For example, seven words may be formed with a total of 45 characters. For example, the second text data (312) may be 'Eligible users can select which devices participate'. The third text data (313) (S3) may be four words (e.g., W 31 , W 32 , W 33 , W 34 ) may be included. For example, four words may be formed with a total of 21 characters. For example, the third text data (313) may be 'New product was released'. The fourth text data (314) (S4) may be composed of six words (e.g., W 41 , W 42 , W 43 , W 44 , W 45 , W 46) may be included. For example, six words may be formed with a total of 37 characters. As an example, the fourth text data (314) may be 'Customers can conserve energy with program'. In the above example, words for English sentences are exemplified. Meanwhile, in some languages (e.g., Korean, Japanese), words and particles may be used together. In such cases, phrases may be used instead of words. The phrases are units separated by spaces and may include only words or words and particles. As an example, the text data of "Customers like products with good performance" may include five phrases. The descriptions for the 'word' unit in FIG. 3 may also be applied to text data in the 'phrase' unit.
[0074] According to one embodiment, the electronic device (101) may determine a voice output speed based on the number of text data stored in the buffer (240) for text-to-speech conversion. The electronic device (101) may determine a request level corresponding to the number of text data, and determine a voice output speed according to the request level. For example, if the number of requests (i.e., the number of text data) stored in the buffer (240) for text-to-speech conversion is 1, the electronic device (101) may determine the voice output speed as a value that is 1.2 times the currently set value (v). If the number of requests (i.e., the number of text data) stored in the buffer (240) for text-to-speech conversion is 2, the electronic device (101) may determine the voice output speed as a value that is 1.5 times the currently set value (v). If the number of requests (i.e., the number of text data) stored in the buffer (240) for text-to-speech conversion is 3, the electronic device (101) may determine the voice output speed to be twice the currently set value (v). If the number of requests (i.e., the number of text data) stored in the buffer (240) for text-to-speech conversion is 4 or more, the electronic device (101) may determine the voice output speed to be 2.5 times the currently set value (v). For example, the electronic device (101) may determine that a total of 4 text-to-speech conversion requests stored in the buffer (240) of FIG. 3 are present. The electronic device (101) may determine the voice output speed corresponding to '4'. The electronic device (101) may change the voice output speed to be 2.5 times the currently set value (v). As a non-limiting example, the electronic device (101) may determine the voice output speed to be a value proportional to the number of requests for text data.
[0075] According to one embodiment, the electronic device (101) may determine the voice output speed based on the text length of the text data stored in the buffer (240) for text-to-speech conversion. The electronic device (101) may determine the length level of the text length and determine the voice output speed according to the length level. For example, if the total length of the text data stored in the buffer (240) for text-to-speech conversion is less than 50 characters, the electronic device (101) may determine the voice output speed as a value that is 1.2 times the currently set value (v). If the total length of the text data stored in the buffer (240) for text-to-speech conversion is 50 characters or more and less than 100 characters, the electronic device (101) may determine the voice output speed as a value that is 1.5 times the currently set value (v). If the total length of the text data stored in the buffer (240) for text-to-speech conversion is less than 150 characters, the electronic device (101) may determine the voice output speed to be twice the currently set value (v). If the total length of the text data stored in the buffer (240) for text-to-speech conversion is 150 characters or more, the electronic device (101) may determine the voice output speed to be 2.5 times the currently set value (v). For example, the electronic device (101) may determine that the text length of the text data stored in the buffer (240) of FIG. 3 is 125 characters in total. The electronic device (101) may determine the voice output speed corresponding to '125 characters'. The electronic device (101) may change the voice output speed to be twice the currently set value (v). As a non-limiting example, the electronic device (101) may determine the voice output speed to be a value proportional to the amount of text data. In the above example, the speech output rate may be set based on a currently set value (v), but it is understood that the present disclosure is not limited thereto. For example, the speech output rate may alternatively be set based on a preset value or a default value.For example, if the total length of text data stored in the buffer (240) for text-to-speech conversion is less than 50 characters, the electronic device (101) can determine the voice output speed to be 1.2 times the default value, and if the total length of text data stored in the buffer (240) for text-to-speech conversion is more than 50 characters but less than 100 characters, the electronic device 101 can determine the voice output speed to be 1.5 times the default value, and the voice output speed can be determined in this manner.
[0076] Figure 4 illustrates an example of speech language conversion performed based on adaptive speech output speed according to various embodiments. Figure 4 illustrates a situation in which an electronic device (101) listens to a speaker's speech provided in a first language and outputs speech in a second language.
[0077] Referring to FIG. 4, the horizontal axis represents the flow of time. A speaker (or a source of voice data) may start speaking at a first time point (401) and end speaking at a second time point (402). The electronic device (101) may receive a voice signal (410) between the first time point (401) and the second time point (402). The voice data of the voice signal (410) may include a plurality of voice units. The electronic device (101) may generate a text unit corresponding to each voice unit through voice-to-text conversion (211). For example, the text unit may include a sentence. The voice-to-text conversion unit (251) may output text in sentence units. Hereinafter, an example in which the text unit is a sentence is described, but the text unit may correspond to text in character units, word units, or phrase units. The electronic device (101) may generate a text unit through the voice-to-text conversion unit (251) (e.g., ASR module). Once a text unit corresponding to a portion of a voice signal (410) is generated, the generated text unit can be input to a translation unit (253) (e.g., MT module). As a non-limiting example, the text unit can be input to the translation unit (253) (e.g., MT module) in real time.
[0078] The electronic device (101) can translate text units output through the voice-to-text conversion unit (251). The electronic device (101) can provide the text units to the translation unit (253). The input of the translation unit (253) may be in units of text units. For example, the translation unit (253) may input text in units of characters. The translation unit (253) may collect the text units and perform translation (213) together. The translation unit (253) may generate a single translated text based on one or more text units. The unit in which the translation (213) is performed may be referred to as text data. For example, the translation unit (253) may perform translation for a single sentence by combining a plurality of characters and / or words. The unit output from the translation unit (253) may be text data. The translation unit (253) may output text data. For example, the electronic device (101) can generate text data (463a) of a second language corresponding to a set of text units (461a) of a first language through the translation unit (253). The electronic device (101) can generate text data (463b) of a second language corresponding to a set of text units (461b) of the first language through the translation unit (253). The electronic device (101) can generate text data (463c) of a second language corresponding to a set of text units (461c) of the first language through the translation unit (253). The electronic device (101) can generate text data (463d) of a second language corresponding to a set of text units (461d) of the first language through the translation unit (253). The electronic device (101) can generate text data (463e) of a second language corresponding to a set of text units (461e) of a first language through the translation unit (253). The electronic device (101) can generate text data (463f) of a second language corresponding to a set of text units (461f) of a first language through the translation unit (253).The electronic device (101) can generate text data (463g) of a second language corresponding to a set of text units (461g) of a first language through the translation unit (253). The electronic device (101) can generate text data (463h) of a second language corresponding to a set of text units (461h) of a first language through the translation unit (253).
[0079] The electronic device (101) can store text data output from the translation unit (253) in a buffer (240). One or more text data can be stored in the buffer (240). After text-to-speech conversion is performed, the text data of the buffer (240) that was input first can be output from the buffer (240). The output text data can be converted into voice data for the next voice output.
[0080] The electronic device (101) can convert text data into voice data in the order stored in the buffer (e.g., the order input). The electronic device (101) can provide the text data to the text-to-speech conversion unit (255) in the order stored in the buffer (e.g., the order input). The electronic device (101) can generate voice data (465a) corresponding to the text data (463a) through the text-to-speech conversion unit (255) (e.g., the TTS module). The electronic device (101) can generate voice data (465b) corresponding to the text data (463b) through the text-to-speech conversion unit (255) (e.g., the TTS module). The electronic device (101) can generate voice data (465c) corresponding to the text data (463c) through the text-to-speech conversion unit (255) (e.g., the TTS module). The electronic device (101) can generate voice data (465d) corresponding to text data (463d) through a text-to-speech conversion unit (255) (e.g., a TTS module). The electronic device (101) can generate voice data (465e) corresponding to text data (463e) through a text-to-speech conversion unit (255) (e.g., a TTS module). The electronic device (101) can generate voice data (465f) corresponding to text data (463f) through a text-to-speech conversion unit (255) (e.g., a TTS module). The electronic device (101) can generate voice data (465g) corresponding to text data (463g) through a text-to-speech conversion unit (255) (e.g., a TTS module). The electronic device (101) can generate voice data (465h) corresponding to text data (463h) through a text-to-speech conversion unit (255) (e.g., TTS module).
[0081] The electronic device (101) can determine a voice output speed to generate each voice data. When generating voice data, the electronic device (101) according to embodiments of the present disclosure can check the amount of text data currently pending in the buffer (240) for text-to-speech conversion. The electronic device (101) can generate voice data based on the amount of text data. For example, the electronic device (101) can generate voice data (465a) for text data (463a). After outputting the voice data (465a), the electronic device (101) can check the buffer (240). The electronic device (101) can check the amount of text data stored in the buffer (240). For example, the electronic device (101) can check and determine that text data (463b) is waiting in the buffer (240). The electronic device (101) can determine that the number of requests stored in the buffer (240) is '1'. The electronic device (101) can adjust the voice output speed to a value (v2) that is 1.2 times higher than the previous voice output speed (v1). The electronic device (101) can generate voice data (465b) corresponding to the text data (463b) based on the adjusted voice output speed. The electronic device (101) can output the voice data (465b). After outputting the voice data (465b), the electronic device (101) can check the buffer (240) again. The electronic device (101) can check and determine that text data (463c) is waiting in the buffer (240). The electronic device (101) can determine that the number of requests stored in the buffer (240) is '1'. The electronic device (101) can adjust the voice output speed to a value (v3) that is 1.2 times higher than the previous voice output speed (v2).In this manner, the electronic device (101) can adjust the speech output rate each time it outputs speech data (or prepares speech data for output) to reflect the amount of text data waiting for text-to-speech conversion (e.g., the number of stored text data requests or units, or the text length of the stored text data). In some examples, the electronic device (101) can adjust the speech output rate to reflect the status of the current buffer (240) (e.g., the remaining capacity, the number of stored text data, and / or the text length of the stored text data).
[0082] Whenever a request comes to the text-to-speech conversion unit (255) (e.g., TTS module), the electronic device (101) can store text data corresponding to the request (e.g., text data (463a), text data (463b), text data (463c), text data (463d), text data (463e), text data (463f), text data (463g), text data (463h)) in a buffer (240) for text-to-speech conversion. The electronic device (101) according to embodiments of the present disclosure can adaptively control the voice output speed according to the number of requests accumulated in the buffer (240) (in other words, the number of text data transmitted from the translation unit (253) and stored in the buffer (240). In the present disclosure, text data, which is a unit of an output request, can be configured in terms of a word unit, a phrase unit, a sentence unit, or a sentence unit in actual voice.
[0083] Although FIG. 4 describes an example of controlling a voice output speed based on the amount of text data stored in the buffer (240) (e.g., the number of text data, the text length of the text data), embodiments of the present disclosure are not limited thereto. According to one embodiment, the electronic device (101) may determine a voice output speed based on whether a voice in a first language is currently input to the electronic device (101) (e.g., whether a voice signal is being input to the ASR module), whether a translation for the voice in the first language is requested (e.g., whether a text unit is being input to the MT module), and / or the number of text inputs in the first language waiting for translation (e.g., the number of text units output from the voice-to-text conversion unit (251) but not input to the translation unit (253). For example, as the number of text units in the first language waiting for translation increases, the voice output speed may become faster. The above elements may be used independently of the amount of text data in the buffer (240) described through FIGS. 2a, 2b, 3, and 4, or may be used together with the amount of text data in the buffer (240) to determine the speech output rate.
[0084] Although the input and output units of the speech-to-text conversion unit (251), the translation unit (253), and the text-to-speech conversion unit (255) are shown to be the same in FIG. 4, the embodiments of the present disclosure are not limited thereto. For example, the output unit of the speech-to-text conversion unit (251) may be different from the output unit of the text-to-speech conversion unit (255). For example, the output unit of the translation unit (253) may be different from the input unit of the text-to-speech conversion unit (255).
[0085] Although an example of outputting voice data according to a determined voice output speed is described in FIG. 4, embodiments of the present disclosure are not limited thereto. The electronic device (101) may determine the voice output speed of current voice data based on the voice output speed of previous voice data. In order to allow the user of the electronic device (101) to relatively feel less of a sudden change in the voice output speed, the electronic device (101) may output the current voice data at a voice output speed corresponding to a value between the output speed of the previous voice data and the determined voice speed. For example, even if the determined voice output speed is 1.5 times ('x1.5'), if the voice output speed of the previous voice data is 1 times ('x1'), the output speed for the current voice data may be determined to be 1.2 times ('x1.2'). As a non-limiting example, when text data includes multiple sentences, the electronic device (101) may gradually increase the voice output speed each time it outputs one sentence. For example, the electronic device (101) can output the voice data of the first sentence at 1.2 times the speed and the voice data of the second sentence at 1.4 times the speed. The electronic device (101) can generate the voice data to be output at the above voice output speed. For another example, even if the determined voice output speed is 1.0 times ('x1.0'), if the voice output speed of the previous voice data is 1.5 times ('x1.5'), the output speed for the current voice data can be determined to be 1.3 times ('x1.3'). For example, the electronic device (101) can output the voice data of the first sentence at 1.3 times the speed and the voice data of the second sentence at 1.0 times the speed. The electronic device (101) can generate the voice data to be output at the above voice output speed.
[0086] As a non-limiting example, the voice output speed may vary even within a single sentence. For example, the voice output speed for voice data corresponding to the previous sentence may be 1x, and the determined voice output speed may be 1.3x. The voice output speed for a portion (e.g., a first voice segment) of a sentence of the voice data to be output may be 1.2x, and the voice output speed for a portion (e.g., a second voice segment) following the portion of the sentence may be 1.3x. The electronic device (101) may generate voice data to be output at the voice output speed.
[0087] FIG. 5 illustrates components of a text-to-speech conversion unit (e.g., text-to-speech conversion unit (255)) for controlling voice output speed according to various embodiments.
[0088] Referring to FIG. 5, the text-to-speech conversion unit (255) may include a text analyzer (511), a prosody predictor (513), and a vocoder (515). The text analyzer (511) may be configured to analyze input text data. The text analyzer (511) may generate a sequence of synthesis units by analyzing the text data. The sequence may be a set of synthesis units. For example, the synthesis unit may be a phoneme. For example, the synthesis unit may be a syllable. For example, the synthesis unit may be an N-gram (N consecutive unit entities (e.g., word, morpheme, syllable, character)). The text analyzer (511) may transmit the sequence or input text data to the prosody predictor (513). As a non-limiting example, each synthesis unit of the sequence may have a duration set as a default. The duration may indicate the time for which a sound corresponding to the synthesis unit is played. According to one embodiment, the electronic device (101) may set the basic duration of the synthesis unit based on at least one of linguistic characteristics including the text length of the input text data (e.g., the number of characters in the text data), the number of sentences in the input text data, the number of words in the input text data, the number of vowels in the input text data, the number of consonants in the input text data, the nationality of the speaker, the nationality of the listener, and / or other languages (e.g., English, Chinese, Japanese, Korean).
[0089] A prosody predictor (513) can generate prosody information. The prosody information can represent information for reflecting voice characteristics, such as rhythm, intonation, stress, and speed, to a designated synthesis unit. For example, the prosody information can include information about duration (e.g., the number of frames corresponding to each synthesis unit), mel-spectrum, power, and / or energy. Prosody information for speech synthesis can include other voice feature values that can be replaced (or modified) in addition to (or instead of) the information listed above. The duration can represent the duration of a synthesis unit. The mel-spectrum can represent a change in frequency over time of a voice to be output. The power can represent energy in a frequency component. The energy can represent power in the entire frequency domain. According to one embodiment, the electronic device (101) may generate prosody information based on a speech output speed determined based on an amount of text data for text-to-speech conversion (e.g., text-to-speech text data in a buffer (e.g., buffer (240))). The electronic device (101) may generate the prosody information so that speech data may be output according to the determined speech output speed. For example, the prosody information may include a duration of a sound corresponding to a synthesis unit. The electronic device (101) may adjust the speed of the synthesis sound by changing the duration of at least some of the synthesis units (e.g., changing from 3 frames to 5 frames).
[0090] The vocoder (515) can generate a synthesized sound based on prosody information corresponding to a plurality of synthesized units. The vocoder (515) can generate voice data (e.g., second voice data of a second language) based on the generated synthesized sound. The vocoder (515) can generate a waveform corresponding to the generated synthesized sound. According to one embodiment, the electronic device (101) can change the length of the synthesized sound based on a voice output speed determined according to the amount of text data in a buffer for text-to-speech conversion (e.g., buffer (240)). The prosody predictor (513) can calculate the length (e.g., frame information) of the synthesized sound to be generated for each synthesized unit (e.g., phoneme) and then transmit the calculation result to the vocoder (515). The vocoder (515) can generate the synthesized sound based on the length and prosody information of the synthesized sound. For example, the vocoder (515) can calculate the length of the final synthesized sound by adding the number of frames (e.g., duration) of the synthesis units transmitted from the prosody predictor. The vocoder (515) can generate speech data based on the length of the final synthesized sound. For example, if the length of the final synthesized sound is in a first range, the vocoder (515) can generate speech data by multiplying the length of the final synthesized sound by a weight (e.g., 0.8) corresponding to the first range (e.g., 10 seconds or more). The vocoder can generate speech data by multiplying the length of the final synthesized sound by a weight (e.g., 0.5) corresponding to the second range (e.g., 20 seconds or more). As a non-limiting example, the weights may differ depending on the characteristics of the synthesis units. For example, the weights may be applied only to phonemes corresponding to voiced sounds among the synthesis units. The vocoder (515) can change the length of the final synthesized sound by applying the weights to each frame length. By changing the length of the final synthesized sound, voice data can be output according to a predetermined voice output speed (or a voice output speed adjusted in consideration of the previous sentence).As a non-limiting example, the speech output rate according to the output target may be set before or at the time the text data is input to the text analyzer (511).
[0091] FIG. 6 illustrates an operational flow of an electronic device (e.g., electronic device (101)) for performing speech language conversion based on adaptive speech output speed according to various embodiments.
[0092] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0093] Referring to FIG. 6, in operation (601), the electronic device (101) (e.g., processor (120)) may generate first text data for first voice data. For example, the first text data may be generated based on the first voice data. The electronic device (101) may receive voice data of a voice signal in a first language (e.g., English). The electronic device (101) may perform voice-to-text conversion for each voice unit of the voice data. For example, the electronic device (101) may perform voice-to-text conversion through an ASR module. The electronic device (101) may generate text units through the voice-to-text conversion. A set of the text units may correspond to the first text data. The language type of the text units of the first text data may be the first language.
[0094] In operation (603), the electronic device (101) (e.g., the processor (120)) may obtain second text data corresponding to the first text data through translation from a first language (e.g., English) to a second language (e.g., Korean). The second language may be different from the first language. The electronic device (101) may translate text units of the first text data. The first text data may indicate a unit in which translation is performed. The electronic device (101) may translate the text units through an MT module. The electronic device (101) may generate second text data as a result of the translation. The language type of the second text data may be the second language.
[0095] In operation (605), the electronic device (101) (e.g., processor (120)) may determine a speech output speed for second text data based on the amount of text data waiting for text-to-speech conversion (e.g., waiting in a buffer (240)). The text data, which is a unit of speech output, may correspond to one sentence or two or more sentences (e.g., paragraph).
[0096] The electronic device (101) can determine the amount of text data waiting for text-to-speech conversion. There may be a difference between the speed of text data input to a text-to-speech conversion module (e.g., a TTS module) and the speed of voice data output. Therefore, the text data may be temporarily stored in a buffer (e.g., buffer (240)). The electronic device (101) can determine the amount of text data stored in the buffer. In one embodiment, the electronic device (101) can determine the number of text data stored in the buffer. After translation is performed, text data in a second language may be input into the buffer. The input to the buffer may be understood as a request for text-to-speech conversion. In other words, the electronic device (101) can determine the number of text-to-speech conversion requests as the amount of text data stored in the buffer. In addition, in one embodiment, the electronic device (101) can determine the text length of the text data stored in the buffer. For example, the text length may indicate the number of characters in the text data. As a non-limiting example, spacing (' ') may be treated as a single character. As another example, the text length may represent the number of words in the text data. Furthermore, according to one embodiment, the electronic device (101) may determine the amount of text data waiting for text-to-speech conversion based on both the number of text data and the length of the text data.
[0097] The electronic device (101) can determine the voice output speed. The electronic device (101) can determine the voice output speed for the second text data according to the amount of text data waiting for text-to-speech conversion. The electronic device (101) can set the voice output speed to be faster as the amount of text data pending in the buffer increases. The electronic device (101) can set the voice output speed to be slower as the amount of text data pending in the buffer decreases. For example, if the number of requests stored in the buffer (240) for text-to-speech conversion (i.e., the number of text data waiting in the buffer (240)) is 1, the electronic device (101) can determine the voice output speed to be 1.5 times the currently set value (v). If the number of requests stored in the buffer (240) for text-to-speech conversion (i.e., the number of text data) is 2, the electronic device (101) can determine the voice output speed to be 2 times the currently set value (v). If the number of requests (i.e., the number of text data) stored in the buffer (240) for text-to-speech conversion is 3, the electronic device (101) may determine the voice output speed as a value that is 3 times the currently set value (v). If the number of requests (i.e., the number of text data) stored in the buffer (240) for text-to-speech conversion is 4 or more, the electronic device (101) may determine the voice output speed as a value that is 4 times the currently set value (v). For another example, if the total length of the text data stored in the buffer (240) for text-to-speech conversion is less than 50 characters, the electronic device (101) may determine the voice output speed as a value that is 1.2 times the currently set value (v). If the total length of the text data stored in the buffer (240) for text-to-speech conversion is 50 or more characters and less than 100 characters, the electronic device (101) may determine the voice output speed as a value that is 1.5 times the currently set value (v).If the total length of the text data stored in the buffer (240) for text-to-speech conversion is less than 150 characters, the electronic device (101) may determine the voice output speed as a value that is twice the currently set value (v). If the total length of the text data stored in the buffer (240) for text-to-speech conversion is 150 characters or more, the electronic device (101) may determine the voice output speed as a value that is 2.5 times the currently set value (v). As another example, the electronic device (101) may determine the voice output speed as a value that is proportional to the amount of text data. For example, if the amount of text data exceeds a specified value (e.g., 200 characters), the electronic device (101) may determine (or maintain) the voice output speed as a fixed output speed (e.g., 3 times). In the above example, the voice output speed may be set based on the currently set value (v), but it will be understood that the present disclosure is not limited thereto. For example, the voice output speed may alternatively be set based on a predetermined value or a default value. For example, if the total length of the text data stored in the buffer (240) for text-to-speech conversion is less than 50 characters, the electronic device (101) may determine the voice output speed to be 1.2 times the default value, and if the total length of the text data stored in the buffer (240) for text-to-speech conversion is 50 characters or more but less than 100 characters, the electronic device (101) may determine the voice output speed to be 1.5 times the default value, and the voice output speed may be determined in this manner.
[0098] In operation (607), the electronic device (101) (e.g., the processor (120)) may generate second voice data for the second text data by performing text-to-speech conversion based on a voice output speed. The electronic device (101) may perform text-to-speech conversion. For example, the electronic device (101) (e.g., the processor (120)) may generate second voice data for the second text data by performing text-to-speech conversion based on the determined voice output speed. The text-to-speech conversion may include text analysis (e.g., operation of a text analyzer (511)), generation of prosody information (e.g., operation of a prosody predictor (513)), and generation of synthesized sound (e.g., operation of a vocoder (515)). The electronic device (101) may perform text analysis for the second text data. The electronic device (101) may determine a sequence of synthesized units corresponding to the second text data through the text analysis. A synthesis unit may be a unit that forms a sound. For example, a synthesis unit may be a phoneme. For example, when the second text data is 'Hello', the synthesis units may be 'A', 'N', 'N', 'Y', 'Y', 'H', 'A', 'S', 'E', 'Yo'. For example, a synthesis unit may be a syllable. For example, when the second text data is 'Hello', the synthesis units may be 'An', 'Nyeong', 'Ha', 'Se', 'Yo'. For example, a synthesis unit may be an N-gram (N consecutive unit entities (e.g., word, morpheme, syllable, character)). As a non-limiting example, a method of determining synthesis units may vary depending on which language the second language is (e.g., Italian, Spanish, Chinese, Japanese, or Korean).
[0099] According to one embodiment, the electronic device (101) may apply a voice output speed in the step of generating prosody information. The electronic device (101) may generate prosody information of a synthesis unit. The prosody information may include information about a duration (e.g., number of frames), a mel-spectrum, power, and / or energy. For example, the electronic device (101) may determine a duration for each synthesis unit. The electronic device (101) may determine the duration of each of 'ㅏ', 'ㄴ', 'ㄴ', 'ㅕ', 'ㅇ', 'ㅎ', 'ㅏ', 'ㅅ', 'ㅔ', and 'ㅛ'. The electronic device (101) may generate the prosody information so that voice data can be output according to the determined voice output speed. For example, the electronic device (101) may reduce or increase the duration of at least some of the synthesis units. The electronic device (101) can generate voice data according to a determined voice output rate by changing the duration of individual synthesis units. As a non-limiting example, the method of changing the parameters (e.g., duration, power, energy) of the prosodic information of each synthesis unit may vary depending on the language of the second language (e.g., Italian, Spanish, Chinese, Japanese, or Korean).
[0100] According to one embodiment, the electronic device (101) may apply a voice output speed in the synthetic sound generation step. The electronic device (101) may generate the synthetic sound based on the synthesis units and prosody information. The electronic device (101) may determine a length corresponding to the synthetic sound. If the length is greater than or equal to a threshold value, the electronic device (101) may adjust the length corresponding to the synthetic sound to be output by applying a weight. For example, if the length is greater than or equal to 10 seconds, the electronic device (101) may reduce the overall length by multiplying each frame (e.g., duration) forming the synthetic sound by a weight of 0.8. For example, if the length is greater than or equal to 20 seconds, the electronic device (101) may reduce the overall length by multiplying each frame (e.g., duration of each synthesis unit) forming the synthetic sound by a weight of 0.5. In this manner, the electronic device (101) may generate voice data reflecting the determined voice output speed.
[0101] As a non-limiting example, the electronic device (101) can generate voice data according to the voice output speed by using both a method of controlling the voice output speed in the step of generating prosody information and a method of controlling the voice output speed in the step of generating synthesized sound.
[0102] In operation (609), the electronic device (101) (e.g., the processor (120)) may output second voice data. In some examples, the electronic device (101) (e.g., the processor (120)) may output the second voice data according to the determined voice output speed. The language type of the second voice data may be a second language. According to one embodiment, the electronic device (101) may output the second voice data through a speaker of the electronic device (101). A voice according to the second voice data may be output through the speaker. In addition, according to one embodiment, the electronic device (101) may output the second voice data through an external electronic device (e.g., a wearable device, wireless earphones, an external speaker) connected to the electronic device (101). The electronic device (101) may transmit information about the second voice data to the external electronic device. The external electronic device may output the second voice data. When a user of the electronic device (101) wears the external electronic device, the user of the electronic device (101) can hear a voice according to the second voice data through the external electronic device (101). In addition, according to one embodiment, the electronic device (101) can output the second voice data to an external electronic device (e.g., a server, a cloud) so that another user's electronic device (101) can directly output the second voice data. The electronic device (101) can transmit information about the second voice data to the external electronic device. The other user's electronic device can access the external electronic device and stream the second voice data.
[0103] Although FIG. 6 describes an example of generating voice data based on the determined voice output speed after determining the voice output speed, the embodiments of the present disclosure are not limited thereto. In some embodiments, at least one of the step of generating second voice data by performing text-to-speech conversion or the step of outputting the second voice data may be based on the determined voice output speed. For example, the electronic device (101) may convert text data of a second language (i.e., second text data) into second voice data through a TTS module, and then change the playback speed of the second voice data to match the determined voice output speed. As a non-limiting example, the electronic device (101) may convert text data of a second language (i.e., second text data) into second voice data through a TTS module, and then modify the second voice data so that the playback speed of the second voice data matches the determined voice output speed, and store the modified second voice data. For example, the voice output speed can be adjusted through a predetermined algorithm (e.g., PSOLA (pitch synchronous overlap and add), WSOLA (waveform similarity overlap and add)) for the waveform that is the output of the vocoder of the TTS module.
[0104] FIG. 7 illustrates an operational flow of an electronic device (e.g., electronic device (101)) for adjusting a speech output speed according to a speech rate according to various embodiments. Unlike FIG. 6 , FIG. 7 describes operations of the electronic device (101) for adaptively changing a speech output speed based on the speech rate of a speaker (or data source) in addition to the amount of text data waiting for text-to-speech conversion.
[0105] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0106] Referring to FIG. 7, in operation (701), the electronic device (101) (e.g., the processor (120)) may generate first text data for first voice data. For example, the electronic device (101) (e.g., the processor (120)) may generate the first text data based on the first voice data. The electronic device (101) may receive voice data of a voice signal in a first language (e.g., English). The electronic device (101) may perform voice-to-text conversion for each voice unit of the voice data. For example, the electronic device (101) may perform voice-to-text conversion through an ASR module. The electronic device (101) may generate text units through the voice-to-text conversion. A set of the text units may correspond to the first text data. The language type of the text units of the first text data may be the first language.
[0107] In operation (703), the electronic device (101) (e.g., the processor (120)) may obtain second text data corresponding to the first text data through translation from a first language (e.g., English) to a second language (e.g., Korean). The second language may be different from the first language. The electronic device (101) may translate text units of the first text data. The first text data may indicate a unit in which translation is performed. The electronic device (101) may translate the text units through an MT module. The electronic device (101) may generate second text data as a result of the translation. The language type of the second text data may be the second language.
[0108] In operation (705), the electronic device (101) (e.g., processor (120)) may determine a voice output speed based on a speech speed corresponding to the speech time of the first voice data and / or the text length of the first text data.
[0109] The electronic device (101) can determine the speech time of the first speech data. For example, the electronic device (101) can obtain information about the speech time of the speech units of the first speech data from the ASR module. The electronic device (101) can determine the speech length provided by the text units of the first language corresponding to the first text data to be translated. In other words, the electronic device (101) can determine the speech time of the first language associated with the second text data. The electronic device (101) can determine the text length of the first text data. For example, the electronic device (101) can determine the text length of the first text data from the MT module. The first text data can correspond to a unit in which translation is performed. For example, when the second text data is generated in the MT module, the electronic device (101) can identify the text units of the first language corresponding to the second text data. The text units of the first language can be referred to as the first text data. The electronic device (101) can determine the length of characters of the text units.
[0110] The electronic device (101) can determine a speech rate. The electronic device (101) can determine the text length for the speech time based on the speech rate. The electronic device (101) can determine a voice output rate. The electronic device (101) can determine a voice output rate corresponding to the speech rate. For example, the electronic device (101) can determine a voice output rate as a value proportional to the speech rate. For example, the electronic device (101) can determine a speed level to which the speech rate belongs among a plurality of speed levels. The electronic device (101) can determine a voice output rate corresponding to the determined speed level. The electronic device (101) can determine a voice output rate for text data of a second language (e.g., second text data) for which text-to-speech conversion is to be performed. As a non-limiting example, the electronic device (101) can determine the synthesized voice rate depending on whether first voice data is input. For example, if the electronic device (101) determines that first voice data is being input at the time of determining the voice synthesis speed (or at a time close to the time of determining the speed), it can adjust the speed of the synthesized voice to become faster.
[0111] In operation (707), the electronic device (101) (e.g., the processor (120)) may generate second voice data for the second text data by performing text-to-speech conversion based on the voice output speed. For example, the electronic device (101) (e.g., the processor (120)) may generate second voice data for the second text data by performing text-to-speech conversion. For operation (707), reference may be made to the description for operation (607). According to one embodiment, the speech rate used to determine the speech output speed may be used as a basic speed in the speech data generation step. For example, the basic playback speed of the second text data input for text-to-speech conversion may be set according to the speech rate. Thereafter, by adjusting the duration in the prosody information and / or applying a weight to the output of the vocoder to the basic playback speed, voice data according to the intended voice output speed may be generated.
[0112] In operation (709), the electronic device (101) (e.g., processor (120)) may output second voice data. In some examples, the electronic device (101) (e.g., processor (120)) may output the second voice data according to the determined voice output speed. For operation (709), reference may be made to the description of operation (609).
[0113] Although FIG. 7 illustrates an example of measuring the speech time associated with the second text data to be translated and determining a voice output rate corresponding to the speech time, embodiments of the present disclosure are not limited thereto. For example, the electronic device (101) may also determine the speech rate based on how many characters are requested for translation during a specified period. The electronic device (101) may determine the number of characters of text units for which translation is requested during the specified period as the speech rate. Thereafter, the electronic device (101) may determine the voice output rate corresponding to the speech rate (or corresponding to the level of the speech rate).
[0114] In FIG. 6, an example of determining a speech output rate for second text data based on the amount of text data waiting in a buffer for text-to-speech conversion is described, and in FIG. 7, an example of determining a speech output rate for second text data based on the utterance time and the text length of the first text data is described. However, embodiments of the present disclosure are not limited thereto. According to various embodiments of the present disclosure, the elements used to determine the speech output rate for the second text data in FIGS. 6 and 7 may be used together. For example, the electronic device (101) may determine the speech output rate for the second text data based on the amount of text data waiting in a buffer for text-to-speech conversion and the text length of the first text data. For example, the electronic device (101) may determine the speech output rate for the second text data based on the amount of text data waiting in a buffer for text-to-speech conversion and the utterance time of the first voice data. For example, the electronic device (101) can determine a voice output speed for second text data based on the amount of text data waiting in a buffer for text-to-speech conversion, the speech time of the first voice data, and the text length of the first text data.
[0115] Figures 8a to 8c illustrate examples of speech language conversion according to various embodiments.
[0116] Referring to FIG. 8A, a user (803) of an electronic device (101) can utilize a voice language conversion service (810). For example, the electronic device (101) can execute a translation application. The electronic device (101) can convert a voice signal of a first language (e.g., Korean) acquired through the electronic device (101) into a voice signal of a second language (e.g., Spanish). The electronic device (101) can output voice data corresponding to the voice signal of the second language. For example, the electronic device (101) can execute a video application. The electronic device (101) can convert a voice signal of a first language (e.g., English) provided by the video application into a voice signal of a second language (e.g., Korean). The electronic device (101) can output voice data corresponding to the voice signal of the second language. The user (803) can wear a wearable device (811) (e.g., wireless earphones, headphones). An electronic device (101) may transmit a signal including voice data generated through a voice language conversion function to a wearable device (811). The signal may cause the wearable device (811) to output the voice data. In response to receiving the signal, the wearable device (811) may provide a voice signal in a second language to the user (803). The user (803) may hear the voice in the second language through the wearable device (811) instead of the voice in the first language received by the electronic device (101). In some examples, the electronic device (101) may not transmit the voice signal in the first language to the wearable device (811) (i.e., may transmit only the voice signal in the second language). That is, when audio is output through the wearable device (811), the original voice signal of the video application (i.e., English voice) is replaced with the voice signal in the second language (i.e., voice translated into Korean). Although illustrated as a wearable device in FIG. 8a, embodiments of the present disclosure are not limited thereto.If it is a device (e.g., earphone, wired speaker) that is connected to an electronic device (101) that receives a first language voice and outputs a second language voice signal from the electronic device (101), it can be understood as one embodiment of the present disclosure.
[0117] Referring to FIG. 8B, a user (803) of an electronic device (101) can utilize a voice language conversion service (820). For example, the electronic device (101) can execute a phone application. The electronic device (101) can be performing a phone service through a call connection with an external electronic device. While the phone application is executing, the electronic device (101) can convert a voice signal of a first language (e.g., Korean) transmitted from the external electronic device into a voice signal of a second language (e.g., Spanish). The electronic device (101) can output voice data corresponding to the voice signal of the second language. As a non-limiting example, the electronic device (101) can be configured to output a separate signal to cancel out the voice of the first language so that the user (803) does not feel uncomfortable due to the synthesis of the voice of the first language and the voice of the second language.
[0118] Referring to FIG. 8C, a user (803) of an electronic device (101) can utilize a voice language conversion service (830). Unlike the user (803) translating the voice of a received call, the user's (803) voice can also be transmitted to the caller in a translated state. For example, the user (803) of the electronic device (101) can utter a voice in a first language. The voice in the first language is input to the electronic device (101), and the voice input to the electronic device (101) can be generated as voice data (841) in a second language. The electronic device (101) can transmit the voice data (841) in the second language to a call server (850). The server (850) can transmit the voice data (841) in the second language to an external electronic device (860) (e.g., another user's mobile phone). For another example, an electronic device (101) may execute a conference application. The electronic device (101) may be performing an online conference service by connecting to a server (850) that provides a conference. While the conference application is executing, the user (803) of the electronic device (101) may speak a voice in a first language. The voice in the first language may be input into the electronic device (101), and the voice input into the electronic device (101) may be generated as voice data (841) in a second language. The electronic device (101) may transmit the voice data (841) in the second language to the server (850). The server (850) may transmit the voice data (841) in the second language to an external electronic device (860) (e.g., a conference room speaker, another electronic device connected to the server (850).
[0119] FIG. 9 illustrates functional components of an electronic device (e.g., electronic device (101)) that utilizes a voice language conversion function according to various embodiments.
[0120] Referring to FIG. 9, the electronic device (101) may include components for an input voice signal (901). For example, the electronic device (101) may include a microphone (910) (e.g., an input module (150)), a processing circuit (920), a filter circuit (930), a transmitting mixer (940), and an encoder / decoder circuit (950). The electronic device (101) may include components for an output voice signal (991). For example, the electronic device (101) may include a gain controller (960), a receiving mixer (970), and a speaker (980). At least some of the components for the input voice signal (901) and the components for the output voice signal (991) may be understood as components of the audio module (170) of FIG. 1.
[0121] According to one embodiment, the electronic device (101) may include a voice language conversion function (990) as an application, module, and / or a function (or program and / or code) associated with the application. Operations for the voice language conversion function (990) may be understood to be performed by a control unit (e.g., a control unit (230)) configured to perform the function (which may be an application and / or module equipped with the function).
[0122] The electronic device (101) can obtain an input voice signal (901) of a first language through a microphone (however, it should be understood that the embodiment is not limited thereto, and the voice signal can be obtained from an application or function of the electronic device or an external device). The electronic device (101) can generate the voice signal through a processing circuit (920) and a filter circuit (930). For example, the filter circuit (930) can include an acoustic echo canceller (AEC) for echo cancellation and / or a noise suppressor (NS) for ambient noise cancellation. The filtered voice signal can be provided to a voice language conversion function (990). As a non-limiting example, a voice signal (994) translated according to the voice language conversion function (990) can be provided to a transmission mixer (940). The transmission mixer (940) can mix the filtered speech signal of the first language and the speech signal (994) translated into the second language according to a specified standard and provide the result to the encoder / decoder circuit (950). As a non-limiting example, the operation in the transmission mixer (940) can be omitted, and the speech signal (994) of the translated language (second language) can be transmitted to the encoder / decoder circuit (950).
[0123] The encoder / decoder circuit (950) can provide information related to a speech signal to be output to the gain controller (960). The encoder / decoder circuit (950) can provide information related to a speech signal to be output to the speech language conversion function (990). The gain controller (960) can provide gain control information to the receiving mixer. The receiving mixer (970) can obtain speech data of a first language, which is an output of the speech language conversion function (990). The receiving mixer (970) can transmit an output speech signal (991) to a speaker (980) (e.g., an audio output module (155)). The receiving mixer (970) can mix a gain-controlled speech signal (e.g., a signal of a second language) and a translated speech signal (e.g., a signal of a first language) according to a specified criterion and output the result to the speaker (980). As a non-limiting example, the operation of the receiving mixer (970) may be omitted, and the translated voice signal may be output directly to the speaker (980).
[0124] Although FIG. 9 illustrates an example in which voice data is output through a speaker of the electronic device (101), embodiments of the present disclosure are not limited thereto. If voice data of a second language is output through an external electronic device connected to the electronic device (101), as in FIG. 8A or FIG. 8C , the electronic device (101) may be configured to transmit a signal including the voice data to an external electronic device (e.g., an external audio output device such as a speaker or another electronic device via a server) through a wireless communication circuit instead of the speaker (980).
[0125] FIG. 10 illustrates an example of a screen in a conversation mode in an application for speech language conversion according to various embodiments.
[0126] Referring to FIG. 10, an electronic device (e.g., electronic device (101)) can execute an application for voice language conversion. For example, the electronic device (101) can execute an interpretation application. The electronic device (101) can execute a conversation mode of the interpretation application. The electronic device (101) can execute (or activate) an ASR module in response to the execution of the call application. A user of the electronic device (101) can speak in a first language (e.g., Korean). The electronic device (101) can obtain voice data of the first language (e.g., Korean) through a microphone and the ASR module. The electronic device (101) can convert voice data of the first language (e.g., Korean) through the ASR module. The electronic device (101) can convert the voice data of the first language into first text data. The electronic device (101) can convert the first text data into second text data by translating it from the first language to a second language (e.g., English). The electronic device (101) can generate second voice data corresponding to the second text data through a TTS module. The electronic device (101) can output the second voice data.
[0127] The electronic device (101) may display a screen (1000) while executing the above conversation mode. The screen (1000) may include a first area (1011) and a second area (1012). The first area (1011) may be used to record the speech of a user of the electronic device (101). For example, if the user of the electronic device (101) says, “Hello. Nice to meet you. How are you?”, text data (e.g., “Hello. Nice to meet you. How are you?”) in a first language (e.g., Korean) may be displayed in the first area (1011). For example, the first language may be set to Korean. The second area (1012) may be used to record text to be output through a TTS module of the electronic device (101). For example, the second area (1012) may display text data (e.g., Hello nice to meet you. How are you...) in a second language (e.g., English). For example, the second language may be set to English.
[0128] The electronic device (101) may output second voice data corresponding to the second text data while displaying text data before translation (i.e., first text data) and text data after translation (i.e., second text data). While the electronic device (101) outputs the second voice data in the second language, the electronic device (101) may display an emphasis effect (e.g., bold, color) to indicate the second text data corresponding to the second voice data among the entire text. For example, the electronic device (101) may display the second text data in blue. According to one embodiment, the electronic device (101) may use the emphasis effect to inform the user of the voice output speed. The electronic device (101) may display a screen (1050) while executing the translation application. The screen (1050) may include a first area (1061) and a second area (1062). The first area (1061) of the screen (1050) can be used to record the user's speech of the electronic device (101). For the first area (1061), reference may be made to the description of the first area (1011). The second area (1062) of the screen (1050) can be used to record text to be output through the TTS module of the electronic device (101). For the second area (1062), reference may be made to the description of the second area (1012). For example, the voice output speed of the second text data on the screen (1000) may be 1.3 times the speed. The voice output speed of the second text data on the screen (1050) may be 1.5 times the speed. For example, the second text data on the screen (1000) may be underlined with a solid line, and the second text data on the screen (1050) may be underlined with a dotted line. For example, the thickness of the second text data on the screen (1000) may be '1', and the thickness of the second text data on the screen (1050) may be '3'.For example, the color of the second text data on the screen (1000) may be 'blue', and the color of the second text data on the screen (1050) may be 'red'.
[0129] Although FIG. 10 illustrates an example in which text data in a second language is displayed in the upper area of the user interface and text data in a first language is displayed in the lower area, embodiments of the present disclosure are not limited thereto. For example, the electronic device (101) may display text data in a first language in the upper area and text data in a second language in the lower area in the user interface related to the translation application. As a non-limiting example, the user interface illustrated in FIG. 10 may be displayed in a manner that is easy to view for the other party of the user of the electronic device (101) and rotated approximately 180 degrees. For example, text data in a first language may be displayed in the lower area and text data in a second language may be displayed in the lower area. For example, text data in the first language may be displayed in the lower area, and text data in the second language may be rotated 180 degrees and displayed in the upper area.
[0130] Although FIG. 10 illustrates an example of applying an emphasis effect to text in a language after translation, the embodiments of the present disclosure are not limited thereto. In addition to the language after translation, the emphasis effect may also be applied to the language before translation. For example, the electronic device (101) may apply an emphasis effect to text in the first region (1061) in addition to the second region (1062). Furthermore, for example, the electronic device (101) may apply an emphasis effect to texts in both the first region (1061) and the second region (1062).
[0131] FIGS. 11A and 11B illustrate examples of a user interface for entering a listening mode in an application for speech language conversion, according to various embodiments.
[0132] Referring to FIGS. 11A and 11B , an electronic device (101) may display a screen (1110) for controlling an application for voice language conversion. The electronic device (101) may display a visual object (1111) on the screen (1110). The visual object (1111) may be used to enter a listening mode in the application.
[0133] The electronic device (101) can display a screen (1120). The electronic device (101) can display the screen (1120) in response to a user input (e.g., touch input, voice input) to the visual object (1111). The screen (1120) can include a pop-up screen (1121) and a visual object (1122) for displaying a description of the listening mode. In some embodiments, the screen (1120) may not be displayed, and the electronic device can directly navigate from the screen (1110) to the screen (1130). In some embodiments, the screen (1120) may only be displayed when the listening mode is first used, and / or the user can configure the electronic device to stop displaying the screen (1120).
[0134] The electronic device (101) can display a screen (1130). The electronic device (101) can display the screen (1130) in response to a user input (e.g., touch input, voice input) to a visual object (1122). The screen (1130) can include a control interface (1131) for specifying a translation language, a text area (1132), and a visual object (1133) (e.g., a microphone icon) for executing a listening function. The control interface (1131) can be used to set a language requiring translation (i.e., a first language (e.g., Spanish)) and a language in which translation is performed (i.e., a second language (e.g., English)). Since the translation is not performed, a guide (e.g., 'Tap to mic button to start translating') for guiding the execution of voice translation can be displayed in the text area (1132). The visual object (1133) can be used by the device (101) to execute a listening function to acquire external voice in listening mode on the translation application.
[0135] The electronic device (101) may display a screen (1140). The electronic device (101) may display the screen (1140) in response to a user input (e.g., touch input, voice input) to a visual object (1133). On the screen (1140), the electronic device (101) may be waiting for a voice input. The electronic device (101) may be configured to monitor whether a voice signal is input via an activated ASR module. As a non-limiting example, the electronic device (101) may display a visual object (1141) indicating that it is monitoring.
[0136] The electronic device (101) can display a screen (1150). When a voice signal is input, the electronic device (101) can generate text data (e.g., “Hoy, estoy muy Feliz de...”) in a first language corresponding to the voice data of the voice signal. The electronic device (101) can display the text data in the first language on a text area (1132). The electronic device (101) can generate text data in a second language corresponding to the text data in the first language through translation. The electronic device (101) can display the text data in the second language (e.g., “Today. I am very happy to introduce...”) on the text area (1132). While displaying the text data in the second language, the electronic device (101) can output a voice signal corresponding to the text data in the second language. As a non-limiting example, the electronic device (101) can display a plurality of text data (e.g., a plurality of sentences) on the text area (1132).
[0137] FIG. 12 illustrates an example screen of a listening mode of an application for voice language conversion according to various embodiments.
[0138] Referring to FIG. 12, an electronic device (e.g., electronic device (101)) may display a screen (1200). The screen (1200) may include a text area (1210) and a visual object (1233) (e.g., a microphone icon) for executing a listening function. The text area (1210) may display first text data (e.g., text in Spanish) before translation and second text data (e.g., text in English) after translation. The visual object (1233) may be used for the device (101) to execute a listening function for acquiring external voice in a listening mode on the translation application. Since the listening mode corresponds to one-way translation rather than a conversation, the electronic device (101) may continuously receive a voice signal in a first language through an input module (e.g., a microphone (910)). The speaker may provide a voice signal in the first language. The electronic device (101) can provide a voice signal in a second language by performing interpretation on a section-by-section basis.
[0139] In one embodiment, the electronic device (101) may use highlighting (e.g., bolding) to indicate a translated portion on the display area. For example, the text area (1210) may include a first area (1211), a second area (1212), a third area (1213), and a fourth area (1214). A speaker may provide a voice signal in a first language. The first area (1211) may display text data in the first language (e.g., Spanish, the language before translation) corresponding to voice data of the voice signal. The second area (1212) may display text data in the second language (e.g., English, the language after translation) corresponding to the text data displayed in the first area (1211). The third area (1213) may display text data in the first language (e.g., Spanish, the language before translation) corresponding to subsequent voice data of the voice signal. The fourth area (1214) may display text data in a second language (e.g., English, which is a translated language) corresponding to the text data displayed in the third area (1213). For example, the electronic device (101) may display texts in the first area (1211) and the third area (1213) without underlining. The electronic device (101) may display texts in the second area (1212) and the fourth area (1214) with underlining. As another example, the electronic device (101) may display texts in the first area (1211) and the third area (1213) in a default bold color. The electronic device (101) may display texts in the second area (1212) and the fourth area (1214) in a bolder color than the default bold color.
[0140] The electronic device (101) may output a voice corresponding to some of the texts in the second language while displaying texts in the display area. According to one embodiment, the electronic device (101) may apply a specific visual effect (e.g., bold, color) to the text portion corresponding to the voice currently output to the user or display an additional visual effect (e.g., underline, shadow) around the text portion to distinguish the text portion corresponding to the voice currently output to the user from other text portions. To display the visual effects, a screen (1250) may be exemplified. The text area (1260) of the screen (1250) may include a first area (1261), a second area (1262), a third area (1263), and a fourth area (1264). For the first region (1261), the second region (1262), the third region (1263), and the fourth region (1264), reference may be made to the descriptions for the first region (1211), the second region (1212), the third region (1213), and the fourth region (1214), respectively. The regions on the screen (1200) and the corresponding regions on the screen (1250) may contain the same text data, but may be distinguished through different effects (e.g., underlining, color, boldness, markup). For example, while the electronic device (101) displays the screen (1200), the electronic device (101) may output a voice signal at 1.5 times the default speed. While the electronic device (101) displays the screen (1250), the electronic device (101) may output a voice signal at the default speed. For example, on the screen (1200), the electronic device (101) may highlight (markup) a portion of text corresponding to the currently output voice with a first pattern. On the screen (1250), the electronic device (101) may highlight (markup) a portion of text corresponding to the currently output voice with a second pattern. As another example, on the screen (1200), the electronic device (101) may highlight (markup) a portion of text corresponding to the currently output voice in red.On the screen (1250), the electronic device (101) may display a portion of text corresponding to the currently output voice in green. As another example, on the screen (1200), the electronic device (101) may display a portion of text corresponding to the currently output voice with an underline in a thickness of '10'. On the screen (1250), the electronic device (101) may display a portion of text corresponding to the currently output voice with an underline in a thickness of '3'.
[0141] FIG. 13A illustrates an example of a screen of a conversation mode of an application for voice language conversion in a foldable type electronic device (e.g., electronic device (101)) according to various embodiments.
[0142] Referring to FIG. 13A, an electronic device (101) may include a first housing portion (1301), a second housing portion (1302), and a hinge structure (1303) for rotatably connecting the first housing portion (1301) and the second housing portion (1302). The electronic device (101) may include a display (1304) (e.g., a display module (160)) arranged across the first housing portion (1301) and the second housing portion (1302). A display area of the display may include a first display area (e.g., a screen (1310)) corresponding to the first housing portion and a second display area (e.g., a screen (1320)) corresponding to the second housing portion.
[0143] The electronic device (101) can display a screen (1300). The electronic device (101) can display the screen (1310) through a first display area corresponding to the first housing portion. The screen (1310) can include a control interface for determining an operation mode in an application for voice language conversion (e.g., a translation or interpretation application). For example, the screen (1310) can include a visual object for entering a conversation mode and / or a visual object for entering a listening mode. The electronic device (101) can provide a user input to the visual object for entering the conversation mode.
[0144] The screen (1320) may include a user interface according to a currently set operation mode. For example, in response to a user input for a visual object for entering the conversation mode, the electronic device (101) may display a user interface for the conversation mode in a second display area. The screen (1320) may include a user interface in the conversation mode. The screen (1320) may include a first area (1321) for a first language and a second area (1322) for a second language. Text data in the first language (e.g., English displayed as text before translation) may be displayed in the first area (1321). The electronic device (101) may perform a translation from the first language to the second language. Through the translation, the electronic device (101) may generate text data in the second language (e.g., English displayed as text after translation). The second area (1322) may display text data in the second language (e.g., Spanish displayed as text after translation). The user interface of FIG. 10 may be referenced for the user interface for the above conversation mode. For example, the first area (1321) may correspond to the first area (1011) of FIG. 10. The second area (1322) may correspond to the second area (1012) of FIG. 10.
[0145] FIG. 13b shows an example of a screen in a listening mode of an application for voice language conversion in a foldable type electronic device (e.g., electronic device (101)).
[0146] Referring to FIG. 13B, the electronic device (101) may include a first housing portion (1301), a second housing portion, and a hinge structure (1303) for rotatably connecting the first housing portion (1301) and the second housing portion (1302). The electronic device (101) may include a display (1304) arranged across the first housing portion (1301) and the second housing portion (1302). A display area of the display may include a first display area (e.g., screen (1310)) corresponding to the first housing portion (1301) and a second display area (e.g., screen (1320)) corresponding to the second housing portion (1302). The electronic device (101) may display a visual object (1311) on the screen (1310). The visual object (1311) can be used to enter listening mode in the above application.
[0147] The electronic device (101) can display a screen (1350). The electronic device (101) can display the screen (1310) through a first display area corresponding to the first housing portion. The screen (1310) can include a control interface for determining an operation mode in an application for voice language conversion (e.g., a translation or interpretation application). For example, the screen (1350) can include a visual object for entering a conversation mode and / or a visual object for entering a listening mode. The electronic device (101) can provide a user input to the visual object for entering the conversation mode. The screen (1350) can include a user interface according to a currently set operation mode. For example, in response to a user input for the visual object for entering the listening mode, the electronic device (101) can display a user interface for the listening mode in a second display area. The screen (1350) can include a user interface for the listening mode. The screen (1350) can include a display area. The electronic device (101) may sequentially display text data of the first language and text data of the second language when translating a speech signal of a first language into a speech signal of a second language. Specifically, the electronic device (101) may generate text data of the first language corresponding to the speech data of the first language in a listening mode (e.g., displaying text in Spanish as text before translation). The electronic device (101) may display the text data of the first language. The electronic device (101) may perform a translation from the first language to the second language. Through the translation, the electronic device (101) may generate text data of the second language (e.g., displaying text in English as text after translation).
[0148] FIG. 14 illustrates an example of a screen of a listening mode of an application for voice language conversion in a foldable type electronic device (e.g., electronic device 101) according to various embodiments. The electronic device 101 may include a first housing portion, a second housing portion, and a hinge structure for rotatably connecting the first housing portion and the second housing portion. The electronic device 101 may include a display arranged across the first housing portion and the second housing portion. A display area of the display may include a first display area corresponding to the first housing portion and a second display area corresponding to the second housing portion. A state of the electronic device 101 may be defined according to an angle formed by the first housing portion and the second housing portion. The electronic device 101 may operate in a mode corresponding to the state of the electronic device 101. For example, a state in which the first housing portion and the second housing form an angle of about 180 degrees so that the first display area and the second display area face the same direction may be referred to as an unfolded state. A state in which the first housing portion and the second housing form an angle of less than 180 degrees so that the first display area and the second display area face different directions may be referred to as a flexed state. A state in which the first housing portion and the second housing form an angle of less than about 15 degrees so that the first display area and the second display area face opposite directions may be referred to as a folded state. The electronic device (101) may display different screens in each of the unfolded state, the flexed state, and the folded state.
[0149] Referring to FIG. 14, the electronic device (101) may display a screen (1410) in an unfolded state. The screen (1410) may include a control interface (1411) for setting a translation language, a text area (1413), and a visual object (1433) (e.g., a microphone icon) for executing a listening function. The state of the electronic device (101) may change from the unfolded state to a flex state. In response to the change (1415) to the flex state, the electronic device (101) may display a screen (1420). In the screen (1420), the text area (1413) may be displayed in a first display area of a first housing portion of the electronic device (101). The control interface (1411) and the visual object (1433) for executing the listening function may be displayed in a second display area of a second housing portion of the electronic device (101).
[0150] The electronic device (101) can receive a user input for a visual object (1433). In response to the user input, the electronic device (101) can execute a voice language conversion function. As the electronic device (101) executes the voice language conversion function, the electronic device (101) can monitor a voice signal. For example, the electronic device (101) can check a voice signal received from the outside through an ASR module. As a non-limiting example, the electronic device (101) can display a screen (1430) including a visual object (1439) indicating that the electronic device (101) is currently monitoring a voice signal.
[0151] When the electronic device (101) detects a voice signal, it can display a screen (1440). For example, when the electronic device (101) detects a voice signal, it can generate first text data corresponding to the voice signal. The language type of the voice signal can be a first language (e.g., English). The language type of the first text data can be a first language (e.g., English). The electronic device (101) can perform a translation on the first text data. The electronic device (101) can perform a translation from the first language to a second language (e.g., Spanish). Through the translation, the electronic device (101) can generate second text data corresponding to the first text data. The electronic device (101) can display the first text data and the second text data. The electronic device (101) can display the first text data and the second text data in a text area (1413).
[0152] Figure 15 illustrates an example screen for a function for voice language conversion during a call, according to various embodiments. While a separate application for voice language conversion is running, as well as a separate application (e.g., a call application, a video application), the voice language conversion function can provide improved services to other applications. For example, the voice language conversion function can run simultaneously with another application (e.g., a call application or a media application) to provide the voice language conversion function described below to the other applications.
[0153] Referring to FIG. 15, an electronic device (101) can execute a call application. For a call between a user of the electronic device (101) and a user of another electronic device (hereinafter, an external electronic device) (e.g., electronic device (102), electronic device (104), server (108)), the electronic device (101) can execute the call application. The electronic device (101) can display a screen (1500). The screen (1500) can include a user interface (1510) for the call application. The electronic device (101) can execute a voice language conversion function while the call application is executing. For example, the electronic device (101) can execute the voice language conversion function in response to a specified input (e.g., touch input, voice input). For another example, the electronic device (101) may execute the voice language conversion function in response to a designated event (e.g., a call connection with a counterpart designated separately for voice language conversion, a call connection with a counterpart designated with a name different from the device settings in the contact list). As a non-limiting example, the electronic device (101) may display a query message asking the user whether to execute the voice language conversion function. For example, the query message may be displayed before a session is established with the counterpart through the call application or within a certain period of time after the session is established. As another example, the query message may be displayed in response to a separate input on the electronic device (101).
[0154] The electronic device (101) may, in response to executing the voice language conversion function while the call application is running, display a user interface (1520) for the voice language conversion function. The user interface (1520) may include a control interface (1521) and a conversation area (1522). The control interface (1521) may be used to set a language that requires translation (i.e., a first language (e.g., English)) and a language in which translation is performed (or will be performed) (i.e., a second language (e.g., Korean)). The conversation area (1522) may be used to display a call between a user of the electronic device (101) and the other party of the user (i.e., a user of an external electronic device (101)) as text data. For example, the user of the electronic device (101) may provide a voice signal in Korean. The user of the external electronic device may provide a voice signal in English. Assuming a situation where a user of an electronic device (101) speaks, the electronic device (101) can convert the Korean speech signal into an English speech signal through a speech language conversion function according to embodiments of the present disclosure. The electronic device (101) can adjust the speed of the English speech signal converted through the disclosed methods and transmit it to an external electronic device. The transmitted English speech signal can be output through the external electronic device. The electronic device (101) can output the English speech signal to the external electronic device through a server. According to one embodiment, the electronic device (101) can mix the input Korean speech signal and the speed-adjusted English speech signal in the audio signal domain and transmit the result to the external electronic device. A weighted sum can be used for the mixing. For example, the electronic device (101) can perform mixing by applying a greater weight to the English speech signal.The electronic device (101) can display text data corresponding to the Korean voice signal and text data corresponding to the English voice signal on the conversation area (1522) in response to what the user said. Let us assume a situation in which a user of an external electronic device (i.e., a counterpart of the user of the electronic device (101)) speaks. The electronic device (101) can obtain an English voice signal from the external electronic device. The electronic device (101) can convert the English voice signal into a Korean voice signal through a voice language conversion function according to embodiments of the present disclosure. The electronic device (101) can display text data corresponding to the English voice signal and text data corresponding to the Korean voice signal on the conversation area (1522) in response to what the user of the external electronic device said.
[0155] If the voice output speed of the TTS module varies depending on the situation while the electronic device (101) performs interpretation, the implementation of the present disclosure can be confirmed. For example, if the voice output speed of the voice signal of the TTS module of the second language in a situation where the user rapidly inputs a plurality of sentences in a first language into the electronic device (101) is different from the voice output speed of the voice signal of the TTS module of the second language in a situation where the user relatively slowly inputs a plurality of sentences in the first language into the electronic device (101), the implementation of the present disclosure can be confirmed. The faster the sentences in the first language are input, the faster the voice output speed of the voice signal of the second language can be.
[0156] An electronic device (101) according to embodiments of the present disclosure can adjust the voice output speed based on the amount of text data requested for text-to-speech conversion (e.g., the number of text data, the length of text within the text data). This can reduce the problem of a gap between what the speaker provides and what the listener receives when the speaker speaks quickly or has many sentences when the interpretation listening mode is played back through text-to-speech conversion. As this gap is reduced, the user can experience improved real-time performance in voice language conversion services such as the interpretation listening mode. Furthermore, by outputting pending text data more quickly, the memory performance of the electronic device can be improved.
[0157] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs from the description below.
[0158] In embodiments, an electronic device (101) is provided. The electronic device (101) includes at least one processor (120) including a processing circuit; and a memory storing instructions, which, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to generate (601) first text data for first speech data, generate (603) second text data corresponding to the first text data through translation of the first text data from a first language to a second language, determine (605) a speech output rate for the second text data based on an amount of text data waiting for text-to-speech conversion, generate (607) second speech data for the second text data by performing the text-to-speech conversion based on the speech output rate, and output (609) the second speech data.
[0159] For example, the instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to control a text-to-speech (TTS) module (or a set of function codes, devices, circuits, or instructions) to determine, in response to output of previous speech data, an amount of text data waiting for text-to-speech conversion in a buffer for text-to-speech conversion, determine a speech output rate corresponding to the amount of text data, and generate the second speech data based on the speech output rate.
[0160] For example, the instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate a signal including the second voice data, and to transmit the generated signal to the external electronic device (101), such that the external electronic device (101) connected to the electronic device (101) outputs the second voice data.
[0161] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a sequence of synthesis units for sound formation based on analysis of the second text data, generate prosody information for each synthesis unit of the sequence based on the speech output rate, and output, as the second speech data, a synthesis sound corresponding to the synthesis units based on the prosody information.
[0162] For example, each of the above synthesis units may correspond to a phoneme, a syllable, or a word. The prosodic information may indicate the duration of the corresponding synthesis unit.
[0163] For example, the electronic device (101) may include a speaker. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to output the second voice data through the speaker.
[0164] For example, the amount of text data waiting for the text-to-speech conversion may correspond to the number of text data stored for the text-to-speech conversion in a buffer configured for the text-to-speech conversion in the specified application.
[0165] For example, the amount of text data waiting for the text-to-speech conversion may correspond to the number of characters of the text data waiting for the text-to-speech conversion.
[0166] For example, the instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to, in response to execution of a designated application, obtain a voice signal, and, while the designated application is being executed, obtain the first text data for the first voice data from the voice signal through an automatic speech recognition (ASR) module (or a set of function codes, devices, circuits, or instructions). The second voice data may be output while the designated application is being executed.
[0167] For example, the instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to determine whether a difference between a first value, which is a voice output rate of previously output voice data, and a second value, which is the determined voice output rate, is greater than or equal to a threshold value, and, upon a determination that the difference between the first value and the second value is greater than or equal to the threshold value, generate the second voice data for the second text data based on a third value between the first value and the second value, and, upon a determination that the difference between the first value and the second value is less than the threshold value, generate the second voice data for the second text data based on the second value.
[0168] For example, the instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate a first voice segment corresponding to a first portion of the second text data according to a third value between a first value, which is a voice output rate of previously output voice data, and a second value, which is the determined voice output rate, and to generate a second voice segment corresponding to a second portion of the second text data according to a second value, which is the determined voice output rate. The second voice data may include the first voice segment and the second voice segment.
[0169] For example, the speech output speed for the second text data may be determined based on at least one of whether speech data according to the first language is being input for speech-to-text conversion, whether text data according to the first language is being input for translation, or the number of inputs of text data according to the first language waiting for translation.
[0170] For example, speech-to-text conversion from the first speech data to the first text data may be performed through an automatic speech recognition (ASR) module (or a set of function codes, devices, circuits, or instructions) of the electronic device (101). Translation of the first text data from a first language to a second language may be performed through a translation module. Text-to-speech conversion from the second text data to the second speech data may be performed through a text-to-speech (TTS) module of the electronic device (101). The buffer for the text-to-speech conversion may correspond to a first input first output (FIFO) type. The buffer for the text-to-speech conversion may use text data output from the translation module as input, and the text data input into the buffer may be provided to the TTS module after the TTS module outputs speech data.
[0171] For example, the instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate a first voice segment corresponding to first sentence data from among a plurality of sentence data of the second text data according to a first value, and to generate a second voice segment corresponding to second sentence data from among a plurality of sentence data of the second text data according to a second value different from the first value. The second voice data may include the first voice segment and the second voice segment.
[0172] For example, the electronic device (101) may include a display. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to display the first text data and the second text data through the display while outputting the second voice data, to display an emphasis effect on a first text portion of the first text data corresponding to a portion of the second voice data while a portion of the second voice data is output, and to display an emphasis effect on a second text portion of the second text data corresponding to the portion through the display.
[0173] For example, the electronic device (101) may include a display. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to establish a call connection with an external electronic device (101) through a call application, and while the call connection is in progress, to display a user interface for translation through the display, and to display, through the display, the first text data for the first voice data from the external electronic device (101) on the user interface, and to display, through the display, the second text data for the second voice data from the external electronic device (101) on the user interface.
[0174] For example, the instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain a voice signal according to the first language and, through voice-to-text conversion, generate text data corresponding to each voice data of a designated unit from the voice signal. The text data may include the first text data. The designated unit may be a sentence unit.
[0175] In embodiments, a method performed by an electronic device (101) is provided. The method may include an operation (601) of generating first text data for first voice data, an operation (603) of generating second text data corresponding to the first text data by translating the first text data from a first language to a second language, an operation (605) of determining a voice output speed for the second text data based on an amount of text data waiting for text-to-speech conversion, an operation (607) of generating second voice data for the second text data by performing the text-to-speech conversion based on the voice output speed, and an operation (609) of outputting the second voice data.
[0176] For example, the operation of generating the second voice data may include, in response to the output of the previous voice data, determining an amount of text data waiting for text-to-speech conversion in a buffer for the text-to-speech conversion, determining a voice output speed corresponding to the amount of text data, and controlling a text-to-speech (TTS) module (or a set of function codes, devices, circuits, or instructions) to generate the second voice data based on the voice output speed.
[0177] For example, the operation of outputting the second voice data may include an operation of generating a signal including the second voice data so that an external electronic device (101) connected to the electronic device (101) outputs the second voice data, and an operation of transmitting the generated signal to the external electronic device (101).
[0178] For example, the method may include an operation of acquiring a voice signal in response to execution of a specified application, and an operation of acquiring first text data for the first voice data from the voice signal through an automatic speech recognition (ASR) module (or a set of function codes, devices, circuits, or instructions) while the specified application is executed. The second voice data may be output while the specified application is executed.
[0179] In embodiments, an electronic device (101) is provided. The electronic device (101) may include at least one processor (120) including a processing circuit; and a memory storing instructions. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate first text data for first voice data (601), generate second text data corresponding to the first text data through translation of the first text data from a first language to a second language (603), determine a voice output rate for the second text data based on an utterance time of the first voice data and a utterance rate corresponding to a text length of the first text data (605), generate second voice data for the second text data by performing the text-to-speech conversion based on the voice output rate (607), and output the second voice data (609).
[0180] In embodiments, a method performed by an electronic device (120) is provided. The electronic device (120) may include an operation of generating first text data for first voice data, an operation of generating second text data corresponding to the first text data through translation of the first text data from a first language to a second language, an operation of determining a voice output speed for the second text data based on a voice speed corresponding to a voice time of the first voice data and a text length of the first text data, an operation of generating second voice data for the second text data by performing the text-to-speech conversion based on the voice output speed, and an operation of outputting the second voice data.
[0181] In embodiments, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium can cause an electronic device (120) to generate first text data for first speech data, generate second text data corresponding to the first text data through translation of the first text data from a first language to a second language, determine a speech output rate for the second text data based on an amount of text data waiting for text-to-speech conversion, perform the text-to-speech conversion based on the speech output rate, thereby generating second speech data for the second text data, and output the second speech data.
[0182] In embodiments, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium can cause an electronic device (120) to generate first text data for first speech data, generate second text data corresponding to the first text data through translation of the first text data from a first language to a second language, determine a speech output rate for the second text data based on a speech rate corresponding to an utterance time of the first speech data and a text length of the first text data, generate second speech data for the second text data by performing the text-to-speech conversion based on the speech output rate, and output the second speech data.
[0183] In embodiments, an electronic device is provided. The electronic device includes at least one processor including a processing circuit; and a memory storing instructions, wherein the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to generate first text data based on first speech data, generate second text data corresponding to the first text data through translation of the first text data from a first language to a second language, determine a speech output rate for the second text data based on an amount of text data waiting for text-to-speech conversion, generate second speech data for the second text data by performing the text-to-speech conversion, and cause output of the second speech data, and at least one of generating second speech data for the second text data by performing the text-to-speech conversion and outputting the second speech data is performed based on the determined speech output rate.
[0184] In embodiments, a method performed by an electronic device is provided. The electronic device includes an operation of generating first text data based on first speech data, an operation of generating second text data corresponding to the first text data through translation of the first text data from a first language to a second language, an operation of determining a speech output speed for the second text data based on a speech time of the first speech data and a speech speed corresponding to a text length of the first text data, an operation of generating second speech data for the second text data by performing the text-to-speech conversion, and an operation of outputting the second speech data, wherein at least one of generating the second speech data for the second text data by performing the text-to-speech conversion and outputting the second speech data is performed based on the determined speech output speed.
[0185] In embodiments, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium generates first text data based on first speech data, generates second text data corresponding to the first text data through translation of the first text data from a first language to a second language, determines a speech output rate for the second text data based on a speech rate corresponding to an utterance time of the first speech data and a text length of the first text data, generates second speech data for the second text data by performing the text-to-speech conversion, and causes an electronic device to output the second speech data, wherein at least one of generating the second speech data for the second text data by performing the text-to-speech conversion and outputting the second speech data is performed based on the determined speech output rate.
[0186] For one or more embodiments, at least one of the components described in one or more of the preceding drawings may be configured to perform one or more operations, techniques, processes, and / or methods as described herein. For example, a processor (e.g., a baseband processor) described herein with respect to one or more of the preceding drawings may be configured to operate according to one or more examples described herein. For another example, circuitry associated with a user equipment (UE), a base station, or a network element as described above with respect to one or more of the preceding drawings may be configured to operate according to one or more examples described herein.
[0187] Any of the embodiments described above may be combined with any other embodiment (or combination of embodiments) unless explicitly stated otherwise. The foregoing description of one or more implementations provides examples and descriptions, but is not intended to be exhaustive or limit the scope of the embodiments to the precise forms disclosed. Modifications and variations are possible in light of the above teachings or may be learned from practicing various embodiments.
[0188] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, electronic devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0189] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0190] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0191] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0192] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0193] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In electronic devices; At least one processor comprising a processing circuit; and An electronic device comprising a memory storing instructions, wherein the instructions, when individually or collectively executed by the at least one processor: Generate first text data based on first voice data, By translating the first text data from a first language to a second language, second text data corresponding to the first text data is generated, Determine a speech output speed for the second text data based on the amount of text data waiting for text-to-speech conversion, Generating second voice data for the second text data by performing the text-to-speech conversion based on the voice output speed, causing the second voice data to be output, Electronic devices.
2. In claim 1, When the above instructions are individually or collectively executed by the at least one processor, the electronic device: In response to the output of the previous speech data, determine the amount of text data waiting for the text-to-speech conversion within the buffer for the text-to-speech conversion, Determine the speech output rate corresponding to the amount of text data waiting in the above buffer, Causing a text-to-speech (TTS) function to be controlled to generate the second voice data based on the voice output speed. Electronic devices.
3. In claim 1 or 2, When the above instructions are individually or collectively executed by the at least one processor, the electronic device: Generating a signal including the second voice data so that an external electronic device connected to the electronic device outputs the second voice data, causing the generated signal to be transmitted to the external electronic device; Electronic devices.
4. In any one of the above claims, When the above instructions are individually or collectively executed by the at least one processor, the electronic device: Based on the analysis of the above second text data, a sequence of synthetic units for sound formation is obtained, Generate prosody information for each synthesis unit of the sequence based on the above speech output rate, As the second voice data, causing the synthesized sound corresponding to the synthesized units to be output based on the prosody information, Electronic devices.
5. In claim 4, Each of the above synthesis units corresponds to a phoneme, syllable, or word, The above rhyme information indicates the duration of the corresponding synthetic unit. Electronic devices.
6. In any one of the above claims, The amount of text data waiting for the text-to-speech conversion is determined based on the number of characters of the text data waiting for the text-to-speech conversion. Electronic devices.
7. In any one of the above claims, When the above instructions are individually or collectively executed by the at least one processor, the electronic device: In response to the execution of a specified application, acquire a voice signal, While the above-mentioned application is running, the ASR (automatic speech recognition) function causes the first text data for the first voice data to be acquired from the voice signal, The above second voice data is output while the above specified application is running. Electronic devices.
8. In any one of the above claims, When the above instructions are individually or collectively executed by the at least one processor, the electronic device: Determine whether the difference between the first value, which is the voice output rate of previously output voice data, and the second value, which is the determined voice output rate, is greater than or equal to a threshold value; Based on a determination that the difference between the first value and the second value is greater than or equal to the threshold value, the second voice data for the second text data is generated based on a third value between the first value and the second value, Causing to generate the second voice data for the second text data based on the second value, based on a determination that the difference between the first value and the second value is less than the threshold value. Electronic devices.
9. In any one of the above claims, When the above instructions are individually or collectively executed by the at least one processor, the electronic device: Generate a first voice segment corresponding to a first part of the second text data according to a third value between a first value, which is a voice output rate of previously output voice data, and a second value, which is a determined voice output rate; According to the second value, which is the determined voice output rate, a second voice segment corresponding to the second part of the second text data is caused to be generated, The second voice data includes the first voice segment and the second voice segment, Electronic devices.
10. In any one of the above claims, The speech output speed for the second text data is determined based on at least one of whether speech data according to the first language is currently being input for speech-to-text conversion, whether text data according to the first language is currently being input for translation, the number of inputs of text data according to the first language waiting for translation, the speech time of the first speech data, or the speech speed corresponding to the text length of the first text data. Electronic devices.
11. In any one of the above claims, The speech-to-text conversion from the first voice data to the first text data is performed through the ASR (automatic speech recognition) function of the electronic device, The translation of the first text data from the first language to the second language is performed through a translation function, The text-to-speech conversion from the second text data to the second voice data is performed through the TTS (text-to-speech) function of the electronic device, The buffer for the above text-to-speech conversion corresponds to the FIFO (first input first output) type, The buffer for the above text-to-speech conversion uses text data output from the translation function as input, and the text data input into the buffer is provided to the TTS function after the output of voice data from the TTS function. Electronic devices.
12. In any one of the above claims, When the above instructions are individually or collectively executed by the at least one processor, the electronic device: Generate a first voice segment corresponding to the first sentence data among a plurality of sentence data of the second text data according to the first value, Causes to generate a second speech segment corresponding to the second sentence data among a plurality of sentence data of the second text data according to a second value different from the first value, The second voice data includes the first voice segment and the second voice segment, Electronic devices.
13. In any one of the above claims, Including more displays, When the above instructions are individually or collectively executed by the at least one processor, the electronic device: While outputting the second voice data, the first text data and the second text data are displayed through the display, While a portion of the above second voice data is being output: Displaying an emphasis effect on the first text portion of the first text data corresponding to the above part through the display, Causing the second text portion of the second text data corresponding to the above part to be displayed with an emphasis effect through the display, Electronic devices.
14. In any one of the above claims, Including more displays, When the above instructions are individually or collectively executed by the at least one processor, the electronic device: Establish a call connection with an external electronic device through a calling application; While the above call connection is in progress, the user interface for translation is displayed through the above display, Through the display, displaying the first text data for the first voice data from the external electronic device on the user interface, Causing the second text data for the second voice data from the external electronic device to be displayed on the user interface through the display; Electronic devices.
15. In a method performed by an electronic device, An operation for generating first text data based on first voice data, An operation of generating second text data corresponding to the first text data by translating the first text data from a first language to a second language; An operation of determining a speech output speed for the second text data based on the amount of text data waiting for text-to-speech conversion; An operation of generating second voice data for the second text data by performing the text-to-speech conversion based on the voice output speed; Including an operation of outputting the second voice data, method.
Citation Information
Patent Citations
Synthetic speech text inputting device and program
JP2011059412A
Interpretation apparatus controlling method, interpretation server controlling method, interpretation system controlling method and user terminal
KR1020140120560A
Simultaneous interpretation system for generating a synthesized voice similar to the native talker''s voice and method thereof
KR1020170103209A
Hybrid electric vehicle and method of uphill drive control for the same
KR1020220144425A
Organic light emitting device
KR1020230039010A