Electronic device, method, and non-transitory computer-readable storage medium for converting audio data into text data
The electronic device uses STT and language summarization to convert multi-lingual audio data into coherent text, addressing language alignment issues and ensuring accurate translation.
Patent Information
- Application Number
- PCT/KR2025/006800
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-19
- Filing Date
- 2025-05-19
- Publication Date
- 2026-01-08
AI Technical Summary
Existing technologies face challenges in accurately converting audio data containing multiple languages into text data, particularly when the language of the utterance differs from the designated language, leading to inconveniences for users due to discrepancies in content and context.
The electronic device employs a speech-to-text (STT) function using an artificial intelligence model to identify and convert utterances in multiple languages into text, followed by a language summarization model to generate coherent text data, ensuring alignment with the intended language and context.
This approach ensures that the generated text accurately reflects the speaker's intention by aligning language and context, enhancing user understanding and reducing discrepancies between the original utterance and the translated text.
Smart Images

Figure KR2025006800_08012026_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transitory computer-readable storage medium for converting audio data into text data
[0001] The following descriptions relate to electronic devices, methods, and non-transitory computer-readable storage media for converting audio data into text data.
[0002] Electronic devices can generate text data by providing audio data to an AI model. The audio data may include recording data of the user's speech. The electronic device can record the user's speech through a microphone.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0004] An electronic device is provided. The electronic device may include a memory storing instructions and including one or more storage media. The electronic device may include at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input indicating that audio data is to be converted into text data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a first utterance in a first language and a second utterance in a second language subsequent to the first utterance from the audio data based on the input, by identifying the language of each of the utterances included in the audio data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a first text in the first language by applying a speech-to-text (STT) operation to the first utterance in the first language. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a second text in the second language by applying STT to the second utterance in the second language based on the first text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate text data comprising the first text and the second text.
[0005] A method is provided. The method can be executed within an electronic device. The method can include receiving an input indicating that audio data is to be converted into text data. The method can include identifying a first utterance in a first language and a second utterance in a second language subsequent to the first utterance from the audio data based on the input, by identifying the language of each of the utterances included in the audio data. The method can include generating a first text in the first language by applying a speech-to-text (STT) operation to the first utterance in the first language. The method can include generating a second text in the second language by applying the STT operation to the second utterance in the second language based on the first text. The method can include generating text data including the first text and the second text.
[0006] A non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device, cause the electronic device to receive an input indicating that audio data is to be converted into text data. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, from the audio data, a first utterance in a first language and a second utterance in a second language subsequent to the first utterance, based on the input identifying the language of each of the utterances included in the audio data. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a first text in the first language by applying a speech-to-text (STT) operation to the first utterance in the first language. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a second text in the second language by applying STT to the second utterance in the second language based on the first text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate text data including the first text and the second text.
[0007] An electronic device is provided. The electronic device may include a memory storing instructions and including one or more storage media. The electronic device may include at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input indicating that first text data is to be summarized. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, from the first text data, a first text in a first language and a second text in a second language subsequent to the first text, based on identifying a language of each of the texts included in the first text data based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, based on the language of each of the texts included in the first text data, a third language corresponding to a maximum number of languages among the texts. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a third text in the third language based on performing a summary of text generated by performing a translation of the first text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a fourth text in the third language based on performing a summary of text generated by performing a translation of the second text based on the first text.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate second text data including the third text and the fourth text.
[0008] A method is provided. The method can be executed in an electronic device. The method may include receiving an input indicating that first text data is to be summarized. The method may include identifying a language of each of the texts included in the first text data based on the input, and identifying, from the first text data, a first text in a first language and a second text in a second language following the first text. The method may include identifying, based on the language of each of the texts included in the first text data, a third language corresponding to a maximum number of languages among the texts. The method may include generating a third text in the third language based on summarizing a text generated by performing a translation of the first text. The method may include generating a fourth text in the third language based on summarizing a text generated by performing a translation of the second text based on the first text. The method may include generating second text data including the third text and the fourth text.
[0009] A non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive an input indicating that first text data is to be summarized. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, from the first text data, a first text in a first language and a second text in a second language subsequent to the first text, based on identifying a language of each of the texts included in the first text data based on the input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, based on the language of each of the texts included in the first text data, a third language corresponding to a maximum number of languages among the texts. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a third text in the third language based on performing a summary of text generated by performing a translation of the first text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a fourth text in the third language based on performing a summary of text generated by performing a translation of the second text based on the first text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate second text data including the third text and the fourth text.
[0010] Figure 1 illustrates an example of an electronic device having a speech to text (STT) function.
[0011] Figure 2 illustrates an example of text data generated by applying STT (speech to text) to audio data.
[0012] Figure 3 is a simplified block diagram of an exemplary electronic device.
[0013] Figure 4 illustrates an example of operations for converting audio data of an electronic device into text data.
[0014] Figure 5 illustrates an example of an operation of applying STT (speech to text) to a second utterance based on a first text of an electronic device.
[0015] Figures 6a and 6b illustrate examples of electronic devices that convert audio data into text data.
[0016] Figures 7a and 7b illustrate examples of operations for summarizing text data of an electronic device.
[0017] Figure 8 illustrates an example of an operation of performing a translation of a fifth text based on a fourth text of an electronic device.
[0018] Figure 9 illustrates an example of an electronic device that summarizes text data.
[0019] Figures 10a and 10b illustrate examples of electronic devices for translating text data.
[0020] FIG. 11 is a block diagram of an electronic device within a network environment according to various embodiments.
[0021] FIG. 1 illustrates an example of an electronic device (100) having a STT (speech to text) function.
[0022] Referring to FIG. 1, the electronic device (100) can convert audio data (110) (e.g., recording data, voice data) into text data (120). For example, the electronic device (100) can generate text data (120) based on the audio data (110). For example, the electronic device (100) can generate text data (120) by applying STT (speech to text) to the audio data (110). For example, applying STT to the audio data (110) can be referred to as converting utterances (or voices) of the audio data (110) into text. For example, applying STT to the audio data (110) can include an operation of identifying utterances in the audio data and an operation of converting the identified utterances into text. For example, the operation of converting the identified utterances into text can be represented as converting the utterances into characters corresponding to pronunciation.
[0023] For example, audio data (110) may include recording data and / or voice data. For example, audio data (110) may include user speech. For example, audio data (110) may represent data in which user speech is recorded. For example, audio data (110) may include speeches of multiple users. For example, the speeches of multiple users may be composed of multiple languages. For example, audio data (110) may include speeches in multiple languages. For example, audio data (110) may include speeches in Korean, English, Japanese, Chinese, and / or French. However, the present invention is not limited thereto.
[0024] For example, the electronic device (100) can generate text data (120) by applying STT (speech to text) to audio data (110). For example, the electronic device (100) can provide the audio data (110) to an artificial intelligence model that performs the STT function. For example, the electronic device (100) can apply STT to the audio data (110) by providing the audio data (110) to an artificial intelligence model that performs the STT function. For example, the artificial intelligence model that performs the STT function can be generated through machine learning. For example, the learning can be performed in the electronic device (100) itself or through a separate server. For example, the artificial intelligence model that performs the STT function can be represented as a deep-learning model. For example, the artificial intelligence model that performs the STT function can be included in the electronic device (100).
[0025] As a non-limiting example, an artificial intelligence model performing the STT function may be included in an external electronic device (e.g., a server). For example, the electronic device (100) may transmit audio data (110) to the external electronic device (e.g., a server) via a communication circuit (not shown) to input audio data (110) to the artificial intelligence model performing the STT function within the external electronic device (e.g., a server). For example, the electronic device (100) providing audio data (110) to the artificial intelligence model performing the STT function may include the electronic device (100) transmitting the audio data (110) to the external electronic device (e.g., a server) via a communication circuit to input the audio data (110) to the artificial intelligence model performing the STT function within the external electronic device (e.g., a server). For example, the electronic device (100) may receive output data of the artificial intelligence model performing the STT function within the external electronic device (e.g., a server) from the external electronic device (e.g., a server) via the communication circuit.
[0026] For example, text data (120) may be represented as being generated based on audio data (110). For example, text data (120) may be represented as a character representation of the voice of audio data (110). For example, the electronic device (100) may perform translation, summary, and keyword extraction of text data (120) by inputting the text data (120) into a generative artificial intelligence (AI) model.
[0027] FIG. 2 illustrates an example of text data (e.g., text data (120)) generated by applying STT (speech to text) to audio data (e.g., audio data (110)).
[0028] Referring to FIG. 2, the electronic device (100) can display a screen (210) through a display. The screen (210) can be represented as a screen corresponding to text data (120). For example, the screen (210) can be described as a screen that displays text (e.g., first text (211), second text (212)) within the text data (120). For example, the text data (120) can be generated by applying STT (speech to text) to audio data (110) (e.g., recording data, voice data).
[0029] In the present disclosure, STT may be described as a technology for converting utterances (or voices) within audio (e.g., audio data (110)) into text (e.g., text data (120)). STT may include an operation in which an electronic device (100) converts utterances identified through a voice recognition function into text form. In an embodiment of the present disclosure, applying STT to audio data (e.g., audio data (110)) may be referred to as converting utterances (or voices) of the audio data into text. For example, the electronic device (100) may generate text data (120) by converting the audio data (110) into text. Applying STT to the audio data (110) may include an operation for identifying utterances within the audio data (110) and an operation for converting the identified utterances into text. For example, the operation for converting the identified utterances into text may be represented as converting the utterances into characters corresponding to their pronunciation.
[0030] According to one embodiment, the electronic device (100) can set a language representing text data (120) to be generated using the STT function. For example, the screen (210) can indicate that the electronic device (100) has set the language of the text data (120) to a specified language (e.g., English).
[0031] For example, audio data (110) may include utterances in multiple languages. For example, audio data (110) may include utterances in a first language (e.g., English) and utterances in a second language (e.g., Japanese). For example, text data (120) may include first text (211) and second text (212). For example, first text (211) may correspond to the utterance in the first language within the audio data (110). For example, first text (211) may be data in which the utterance in the first language within the audio data (110) is converted into text. For example, second text (212) may correspond to the utterance in the second language within the audio data (110). For example, second text (212) may be data in which the utterance in the second language within the audio data (110) is converted into text.
[0032] For example, the electronic device (100) can generate the first text (211) by applying STT to the utterance of the first language in the audio data (110). For example, the electronic device (100) can set the language of the text generated through the STT function to the specified language (e.g., English). For example, the electronic device (100) can set the language of the text generated through the STT function in response to a user setting. For example, the electronic device (100) can automatically set the language of the text generated through the STT function by analyzing the audio data (110). For example, the electronic device (100) can set the language of the text generated through the STT function to the maximum number of languages among the utterances in the audio data (110) by analyzing the audio data (110).
[0033] For example, the electronic device (100) can generate a first text (211) converted into the specified language (e.g., English) by applying STT to the utterance in the first language (e.g., English). For example, since the first text (211) converted into the specified language (e.g., English) is generated based on the utterance in the first language (e.g., English), the content of the utterance in the first language and the content of the first text (211) can correspond. For example, the context derived from the first text (211) can be identical to the context derived from the utterance in the first language.
[0034] For example, the electronic device (100) can generate a second text (212) by applying STT to the utterance of the second language (e.g., Japanese) in the audio data (110). For example, the electronic device (100) can generate a second text (212) converted into the designated language (e.g., English) by applying STT to the utterance of the second language (e.g., Japanese). For example, since the second text (212) converted into the designated language (e.g., English) is generated based on the utterance of the second language (e.g., Japanese), the content of the utterance of the second language may be different from the content of the first text (211). For example, since the second text (212) is a transcription of the utterance in the second language (e.g., Japanese) using characters (e.g., alphabets) of the designated language according to pronunciation, the content of the second text (212) may be different from the content of the utterance in the second language. For example, the context derived from the second text (212) may be different from the context derived from the utterance in the second language. For example, since the second language and the designated language are different, it may be relatively difficult for a user to understand the content of the second text (212).
[0035] For example, the utterance in the second language may include a user utterance such as "Hello, Kim. Reservations can be made after 1 p.m. tomorrow." For example, the content of the utterance in the second language may include content such as "Hello, Kim. Reservations can be made after 1 p.m. tomorrow." For example, the second text (212) may include text such as "konnichi yakim yooakua asunoichikoni kkanotest." For example, the second text (212) may represent the utterance in the second language as being written in characters (e.g., alphabets) of the designated language. For example, the second text (212) may be written in characters (e.g., alphabet) of the specified language according to the pronunciation of the utterance in the second language.
[0036] According to one embodiment, the electronic device (100) may receive user input based on the difference between the second language and the designated language. Based on receiving the user input, the electronic device (100) may apply STT to the utterance in the second language (e.g., Japanese) within the audio data (110), thereby generating text converted into a language other than the designated language (e.g., Japanese).
[0037] According to one embodiment, the electronic device (100) can perform a summary of text data (120). For example, the electronic device (100) can perform a summary of text data (120) by inputting the text data (120) into a language summary model. For example, the language summary model can be included in the electronic device (100). For example, the screen (220) can display data summarizing the text data (120). For example, the screen (220) can include text (221) and text (222). For example, the electronic device (100) can display the screen (220) through a display.
[0038] As a non-limiting example, the language summarization model may be included in an external electronic device (e.g., a server). For example, the electronic device (100) may transmit text data (120) to the external electronic device (e.g., the server) via a communication circuit of the electronic device (100) in order to input the text data (120) into the language summarization model within the external electronic device (e.g., the server). The external electronic device (e.g., the server) may generate data summarizing the text data (120) by inputting the received text data (120) into the language summarization model. The external electronic device (e.g., the server) may transmit the data summarizing the text data (120) to the electronic device (100). The electronic device (100) may display a screen (220) through a display using the data summarizing the text data (120). For example, the text (221) may be represented as text generated based on the electronic device (100) performing a summary of the first text (211). For example, text (221) may be represented as data generated by the electronic device (100) inputting the first text (211) into a language summarization model. For example, the content of text (221) may be identical to the content of the utterance in the first language (e.g., English) because the context derived from the first text (211) is identical to the context derived from the utterance in the first language. For example, the context derived from text (221) may be identical to the context derived from the utterance in the first language.
[0039] For example, the text (222) may be represented as text generated based on the electronic device (100) performing a summary of the second text (212). For example, the text (222) may be represented as data generated when the electronic device (100) inputs the second text (212) into a language summarization model. For example, the content of the text (222) may be different from the content of the utterance in the second language (e.g., Japanese) because the context derived from the second text (212) is different from the context derived from the utterance in the second language. For example, the content of the text (222) may be distinguished from the content of the utterance in the second language. For example, the context derived from the text (222) may be different from the context derived from the utterance in the second language.
[0040] For example, the second text (212) may include text such as "konnichi yakim yooakua asunoichikoni kkanotest." For example, the text (222) may include text such as "kim is going to take a test." For example, the text (222) may be generated by inputting the second text (212) into a language summarization model. For example, the text (222) may differ from the speaker's intention of the utterance in the second language.
[0041] According to one embodiment, the electronic device (100) can generate a first text (211) in a designated language (e.g., English) by applying STT to an utterance in a first language (e.g., English) within audio data (110). The electronic device (100) can generate a second text (212) in a designated language (e.g., English) by applying STT to an utterance in a second language (e.g., Japanese) within audio data (110). The electronic device (100) can identify whether a language (e.g., a designated language) of text (e.g., the first text (211), the second text (212)) is the same as a language of an utterance corresponding to a text within the audio data (110). For example, the electronic device (100) can identify whether a language (e.g., a designated language) of the first text (211) is the same as the first language. For example, the electronic device (100) can identify whether the language of the second text (212) (e.g., a specified language) and the second language are the same.
[0042] The electronic device (100) may perform a summary of a text if the language (e.g., a designated language) of the text (e.g., a first text (211), a second text (212)) and the language of the utterance corresponding to the text are the same. For example, the electronic device (100) may input the first text (211) to a language summary model based on the fact that the language (e.g., English) of the first text (211) and the first language (e.g., English) are the same. The electronic device (100) may generate a text (221) by inputting the first text (211) to a language summary model within the electronic device (1001). The electronic device (100) may refrain from summarizing the text if the language (e.g., a designated language) of the text (e.g., a first text (211), a second text (212)) and the language of the utterance corresponding to the text are different. For example, the electronic device (100) may refrain from or bypass inputting the second text (212) to the language summarization model depending on whether the language of the second text (212) (e.g., English) is different from the second language (e.g., Japanese).
[0043] According to one embodiment, when converting audio data (110) into text data (120) of a specified language (e.g., English), the electronic device (100) may generate text with content different from the speaker's intention. For example, the electronic device (100) converting audio data (110) into text data (120) of a specified language may cause inconvenience to a user who uses the text data (120). For example, when converting audio data (110) into text data (120), a method may be required to generate text in the corresponding language by applying STT to the corresponding utterance by identifying the language of each utterance included in the text data (120). For example, the electronic device described below may include components (or hardware components) for providing such a method. The components are described and exemplified in more detail with reference to FIG. 3.
[0044] FIG. 3 is a simplified block diagram of an exemplary electronic device (301) (e.g., electronic device (101), electronic device (1101)).
[0045] Referring to FIG. 3, the electronic device (301) may include at least one processor (300), a display (310), and a memory (320).
[0046] At least one processor (300) may include a hardware component for processing data based on executing instructions. The hardware component for processing data may include, for example, a central processing unit (CPU) (e.g., including processing circuitry). For example, the hardware component for processing data may include a graphic processing unit (GPU) (e.g., including processing circuitry). For example, the hardware component for processing data may include a display processing unit (DPU) (e.g., including processing circuitry). For example, the hardware component for processing data may include a neural processing unit (NPU) (e.g., including processing circuitry).
[0047] At least one processor (300) may include one or more cores. For example, at least one processor (300) may have a multi-core processor structure such as a dual core, quad core, or hexa core.
[0048] The display (310) may include hardware components of the electronic device (301) used to display a screen. For example, the display (310) may include light-emitting elements and circuits (e.g., transistors) that control the light-emitting elements to emit light. For example, each of the light-emitting elements may include an organic light emitting diode (OLED) or a micro LED. However, the present invention is not limited thereto. For example, the display (310) may include a liquid crystal display (LCD).
[0049] The memory (320) may include a hardware component for storing data and / or instructions input to and / or output from at least one processor (300). The memory (320) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). The volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disc, and embedded multimedia card (EMMC).
[0050] For example, at least one processor (300) may execute instructions stored in the memory (320) to convert audio data (110) into text data (120). For example, the instructions, when individually or collectively executed by the at least one processor (300), may cause the electronic device (301) to identify the language of each of the utterances contained in the audio data (110) and generate text in the language of the utterance by applying STT to the utterance. These operations will be described and illustrated in more detail with reference to FIGS. 4 to 6B.
[0051] For example, at least one processor (300) may execute instructions stored in the memory (320) to summarize text data (120). For example, the instructions, when individually or collectively executed by the at least one processor (300), may cause the electronic device (301) to identify the language of each of the texts contained in the text data (120) and perform a summary of the text.
[0052] FIG. 4 illustrates examples of operations for converting audio data (e.g., audio data (110)) of an electronic device (e.g., electronic device (100), electronic device (301), electronic device (1101)) into text data (e.g., text data (120)).
[0053] Referring to FIG. 4, at operation 410, at least one processor (300) may receive an input related to processing of audio data (110) (e.g., recording data, voice data). For example, the input may receive an input indicating that audio data (110) is to be converted into text data (120). For example, converting audio data (110) into text data (120) may include generating text data (120) having the same content as the audio data (110). For example, the input may include a touch input to an executable object (e.g., an executable object (613)) within a screen (e.g., screen (610) of FIG. 6A) displayed on a display (310) of an electronic device (301). For example, the input will be described and exemplified in more detail with reference to FIG. 6A.
[0054] In operation 420, at least one processor (300) may identify the language of each utterance included in the audio data (110) (e.g., recording data, voice data) based on the input. For example, at least one processor (300) may identify the language of each utterance included in the audio data (110) by inputting the audio data (110) to a language detection model. For example, the language detection model may be included in the electronic device (301). As a non-limiting example, the language detection model may be included in an external electronic device (e.g., a server). For example, at least one processor (300) may transmit the audio data (110) to the external electronic device (e.g., a server) in order to utilize the language detection model of the external electronic device (e.g., a server). For example, at least one processor (300) may receive output data of the language detection model for the audio data (110) from the external electronic device (e.g., a server). For example, the language detection model can be generated through machine learning. For example, the learning can be performed on the electronic device (301) itself or through a separate server. For example, the language detection model can be an example of a deep learning model. For example, at least one processor (300) can identify a first utterance in a first language and a second utterance in a second language following the first utterance from the audio data (110) by identifying the language of each of the utterances.
[0055] At operation 430, at least one processor (300) may generate a first text in a first language by applying an STT to a first utterance in a first language. For example, applying an STT to the first utterance may be referred to as converting the first utterance into a first text. Applying an STT to the first utterance may include an operation of identifying the first utterance and an operation of converting the identified first utterance into text. For example, the operation of converting the identified first utterance into text may be represented as converting the utterance into characters in the first language corresponding to the pronunciation.
[0056] For example, at least one processor (300) may generate a first text in a first language by inputting a first utterance within audio data (110) into an artificial intelligence model performing an STT function. For example, the content of the first text may be identical to the content of the first utterance.
[0057] For example, an artificial intelligence model performing the STT function can be generated through machine learning. For example, the learning can be performed within the electronic device (301) itself or through an external electronic device (e.g., a server). For example, an artificial intelligence model performing the STT function can be represented as a deep learning model.
[0058] For example, at least one processor (300) may derive context from the first utterance by providing the first utterance to a large language model (LLM). For example, at least one processor (300) may apply STT to the first utterance based on the context derived from the first utterance. For example, the large language model may be included in the electronic device (301). For example, the large language model may be included in an external electronic device (e.g., a server). For example, the electronic device (301) may transmit the first utterance to the external electronic device via a communication circuit (not shown) to input the first utterance to the large language model within the external electronic device. For example, obtaining (or deriving) other data from the data by providing data (e.g., a first utterance) to a large-scale language model by having at least one processor (300) transmit the data to an external electronic device via a communication circuit for inputting the data into the large-scale language model within the external electronic device. For example, the at least one processor (300) may receive, via the communication circuit, from the external electronic device the other data selected by the large-scale language model within the external electronic device.
[0059] A large-scale language model can be referred to as a language model comprised of an artificial neural network pre-trained on a massive amount of text data. A large-scale language model can contain more than ten times as many parameters (for example, more than 100 billion parameters) as a conventional general language model. A large-scale language model can utilize a transformer artificial neural network structure based on an attention mechanism. The attention mechanism is a technique that helps an artificial intelligence model focus on important parts of input data. The attention mechanism can be used to predict output data by predicting the degree to which at least a portion of time-series input data (e.g., input data such as voice or video, or input data of a layer of a neural network) contributes to the intermediate or final output of the neural network. The recurrent neural network (RNN) structure, which sequentially processes each element of a sequence, has poor prediction performance when there is information dependency between long time series distances, but the attention mechanism can consider information dependency between long time series distances by controlling the degree of weight concentration within the overall (or partial) context of the input data.
[0060] For example, a large-scale language model may include a transformer with an encoder-decoder structure. The encoder processes input data and outputs compressed information (e.g., an attention mechanism), and the decoder processes the compressed information and outputs token-based output data. Each encoder and decoder may include an independent attention network, or a cross-attention network connecting the encoder and decoder.
[0061] For example, large-scale language models can be trained in two stages: pre-training and fine-tuning. Pre-training involves training a large-scale language model to process massive amounts of text data and acquire general linguistic knowledge. For example, this could involve self-supervised learning, such as predicting the next word using a sequence of previous words in a text sequence. Fine-tuning involves training a large-scale language model to fit a specific domain (e.g., chatbot, translation, summarization, Q&A) or task. This can be further supervised (or adaptive) based on a pre-trained model using a dataset tailored to the domain's purpose. Large-scale language models can perform tasks with text inputs containing natural language, called prompts. For example, large-scale language models can include BERT (bidirectional encoder representations from transformer) and GPT (generative pre-trained transformer). The term "LLM" can refer to the neural network model itself, but can also refer to the model of an LLM-based application (e.g., chatbot, translation, summarization, text classification, sentence generation). For example, an LLM-based chatbot such as chatGPT can also be referred to as an LLM. "LLM" can also include an inference engine that utilizes the LLM neural network model. For example, "entering an input prompt into an LLM" can refer to "entering an input prompt into an LLM-based inference engine."
[0062] At operation 440, at least one processor (300) may generate a second text in a second language by applying STT to a second utterance in a second language. For example, applying STT to the second utterance may be referred to as converting the second utterance into a second text. Applying STT to the second utterance may include an operation of identifying the second utterance and an operation of converting the identified second utterance into text. For example, the operation of converting the identified second utterance into text may be represented as converting the utterance into characters in the second language corresponding to the pronunciation. For example, the content of the second text may be identical to the content of the second utterance. For example, at least one processor (300) may apply STT to the second utterance in the second language based on the first text in the first language. For example, the operation of applying STT to a second utterance in a second language based on a first text in a first language will be described and illustrated in more detail with reference to FIG. 5.
[0063] At operation 450, at least one processor (300) may generate text data (120) by applying STT to utterances included in audio data (110). For example, applying STT to utterances included in audio data (110) may be referred to as converting utterances (or voices) included in audio data (110) into texts (or text data). Applying STT to utterances included in audio data (110) may include an operation of identifying utterances included in audio data (110) and an operation of converting the identified utterances into texts. For example, the operation of converting the identified utterances into texts may be represented as converting the utterances into characters corresponding to pronunciations.
[0064] For example, the text data (120) may include texts converted into the language of the utterances by applying STT to the utterances by at least one processor (300). For example, the text data (120) may include a first text in a first language and a second text in a second language.
[0065] For example, at least one processor (300) may receive an input indicating that the text data (120) should be translated into a language after generating the text data (120). For example, at least one processor (300) may perform a translation of texts within the text data (120) based on the input. For example, at least one processor (300) may generate other text data by performing a translation of each of the texts.
[0066] FIG. 5 illustrates an example of an operation of applying STT (speech to text) to a second utterance based on a first text of an electronic device (e.g., electronic device (100), electronic device (301), electronic device (1101)).
[0067] Referring to FIG. 5, the operations of FIG. 5 may be related to operation 440 of FIG. 4.
[0068] In operation 510, at least one processor (300) may identify whether the first language is the same as the second language. For example, at least one processor (300) may identify whether the first language of a first utterance and the second language of a second utterance following the first utterance are the same.
[0069] At operation 520, at least one processor (300) may execute operations 530 and 540 based on identifying that the first language is different from the second language. For example, at least one processor (300) may execute operation 550 based on identifying that the first language is the same as the second language.
[0070] In operation 530, at least one processor (300) may perform a translation of a first text in a first language based on identifying that the first language is different from the second language.
[0071] According to one embodiment, at least one processor (300) may translate a first text in a first language into a second language of a second utterance following the first utterance. For example, by translating the first text, the at least one processor (300) may obtain a third text in which the first text in the first language is translated into the second language. For example, the at least one processor (300) may generate a third text in the second language by translating the first text in the first language. For example, the first language of the first text may be different from the second language of the third text, but the content of the first text may be identical to the content of the third text.
[0072] According to one embodiment, at least one processor (300) may translate a first text in a first language into a language (e.g., a second language) designated by the user. For example, at least one processor (300) may generate a third text in a designated language (e.g., a second language) by translating the first text in the first language.
[0073] In one embodiment, at least one processor (300) may determine a maximum number of languages among utterances in audio data (e.g., audio data (110)) as translation languages. For example, at least one processor (300) may generate a third text in the translation language by performing a translation of a first text in a first language.
[0074] In operation 540, at least one processor (300) may apply STT to a second utterance in a second language based on a third text in the second language. For example, at least one processor (300) may generate a second text in the second language by applying STT to a second utterance in the second language based on the third text in the second language. For example, at least one processor (300) may generate a second text in the second language from a second utterance in the second language based on the third text in the second language. For example, at least one processor (300) may convert or change a second utterance in the second language into a second text in the second language based on the third text in the second language.
[0075] In one embodiment, applying STT to the second utterance may be referred to as converting the second utterance into a second text. Applying STT to the second utterance may include identifying the second utterance and converting the identified second utterance into text. For example, converting the identified second utterance into text may be referred to as converting the utterance into characters in a second language corresponding to its pronunciation.
[0076] In one embodiment, at least one processor (300) may obtain the frequency of use of words within a third text in a second language by inputting the third text into a large-scale language model. For example, at least one processor (300) may apply STT to a second utterance in the second language using the frequency of use obtained from the third text.
[0077] In one embodiment, at least one processor (300) may input a third text in a second language into a large-scale language model to derive context from the third text. For example, at least one processor (300) may apply STT to a second utterance in the second language using the context derived from the third text.
[0078] In one embodiment, at least one processor (300) may convert a second utterance in a second language into a second text in the second language. For example, at least one processor (300) may correct the second text in the second language using a third text in the second language (or a context derived from the third text). For example, at least one processor (300) may correct the second text in the second language by inputting the second text in the second language and the third text in the second language into a large-scale language model.
[0079] For example, in operations 530 and 540, at least one processor (300) is shown generating a second text by applying STT to a second utterance based on a first text, but this is merely an example. For example, at least one processor (300) may apply STT to the second utterance based on texts generated by applying STT to utterances prior to the second utterance (e.g., texts including the first text).
[0080] For example, in operation 550, at least one processor (300) may refrain from (or bypass) performing a translation of the first text based on identifying that the first language is identical to the second language. For example, the at least one processor (300) may derive context from the first text by inputting the first text in the first language into a large-scale language model. For example, the at least one processor (300) may apply STT to a second utterance in the second language using the context derived from the first text. For example, applying STT to the second utterance may be referred to as converting the second utterance into a second text. For example, applying STT to the second utterance may be referred to as converting the second utterance into a second text. Applying STT to the second utterance may include identifying the second utterance and converting the identified second utterance into text. For example, the act of converting an identified second utterance into text can be represented as converting the utterance into characters of a second language corresponding to the pronunciation.
[0081] FIGS. 6A and 6B illustrate examples of electronic devices that convert audio data (e.g., audio data (110)) into text data (e.g., text data (120)).
[0082] Referring to FIG. 6A, at least one processor (300) may display a screen (610) via the display (310). For example, the screen (610) may include a list of audio data (110) (e.g., recording data, voice data). For example, the screen (610) may represent the audio data (110) as text (612). For example, the text (612) may include keywords of the audio data (110). For example, the text (612) may represent a title of the audio data (110). For example, the text (612) may represent a summary of the audio data (110).
[0083] For example, the screen (610) may include an executable object (611) and an executable object (613). For example, at least one processor (300) may select audio data (110) corresponding to the executable object (611) based on receiving an input for the executable object (611). For example, the input may include a touch input on the executable object (611).
[0084] For example, at least one processor (300) may receive input for an executable object (613) after receiving input for an executable object (611). For example, the input for the executable object (613) may include a touch input on the executable object (613).
[0085] For example, at least one processor (300) may display a screen (620) through the display (310) based on receiving an input for an executable object (613). For example, the screen (620) may include an executable object (621), an executable object (622), an executable object (623), and an executable object (624). For example, at least one processor (300) may receive an input for an executable object (621), an executable object (622), or an executable object (623).
[0086] For example, at least one processor (300) may set the language of text data (120) to a designated language (e.g., English) based on receiving an input for an executable object (621). For example, at least one processor (300) may set the language of text data (120) to another designated language (e.g., Korean) based on receiving an input for an executable object (622).
[0087] For example, at least one processor (300) may identify the language of each of the utterances included in the audio data (110) based on receiving an input for the executable object (623). For example, at least one processor (300) may identify the language of each of the utterances included in the audio data (110) upon receiving an input for the executable object (623) and then receiving an input for the executable object (624). For example, at least one processor (300) may generate text representing the language of each of the utterances included in the audio data (110) based on identifying the language of each of the utterances included in the audio data (110). For example, the input for the executable object (624) may be an example of an input indicating converting the audio data (110) of operation 410 of FIG. 4 into text data (120).
[0088] Referring to FIG. 6B, at least one processor (300) may display a screen (630) through the display (310) based on receiving an input for an executable object (621), an executable object (622), or an executable object (623) and then receiving an input for an executable object (624). For example, the screen (630) may indicate that the at least one processor (300) is applying STT to audio data (110). For example, the screen (630) may indicate that the at least one processor (300) is generating text data (120). For example, the at least one processor (300) may display the screen (630) through the display (310) while applying STT to the audio data (110).
[0089] According to one embodiment, at least one processor (300) may display, through the display (310), a screen (630) indicating the type of at least one language in which STT is applied to audio data (110). For example, when at least one processor (300) receives an input for an executable object (621) and then receives an input for an executable object (624), the screen (630) including text such as “Converting text to English... 29%” may be displayed through the display (310). For example, when at least one processor (300) receives an input for an executable object (624) and then receives an input for an executable object (623), the screen (630) including text such as “Converting text to English and Japanese... 29%” may be displayed through the display (310).
[0090] According to one embodiment, at least one processor (300) may display a screen (640) through the display (310) based on generating text data (120). For example, the screen (640) may include text (641), text (642), and text (643).
[0091] For example, an utterance corresponding to text (641) in audio data (110) may be represented as an utterance recorded in a first language (e.g., English). For example, since an utterance corresponding to text (641) in audio data (110) is recorded in a first language (e.g., English), the text (641) may be represented in the first language (e.g., English). For example, the text (641) may include text in the first language, such as "Hello, I'm calling to make an appointment to remove my wisdom tooth. I'm Kim and I'm a regular customer of Mr. Park."
[0092] For example, an utterance corresponding to text (642) in audio data (110) may be represented as an utterance recorded in a second language (e.g., Japanese). For example, since an utterance corresponding to text (642) in audio data (110) is recorded in a second language (e.g., Japanese), the text (642) may be represented in the second language (e.g., Japanese). For example, the text (642) may include text in the second language, such as "こんにちは, キム. The reservation can be made by 1 o'clock tomorrow." For example, the content of the text (642) may be the same as the content of the utterance corresponding to the text (642). For example, a user of the electronic device (301) can easily understand the text (642) than the second text (212) of FIG. 2. For example, a user of the electronic device (301) may prefer the text (642) over the second text (212) because the text (642) better reflects the content of the recorded speech in the second language than the second text (212).
[0093] For example, an utterance corresponding to text (643) in audio data (110) may be represented as an utterance recorded in a first language (e.g., English). For example, because an utterance corresponding to text (643) in audio data (110) is recorded in the first language (e.g., English), the text (643) may be represented in the first language (e.g., English). For example, the text (643) may include text in the first language, such as, "Then, please make a reservation for 2 o'clock tomorrow. I have some questions. Is there a visitor's parking lot? How much pain does the wisdom tooth removal hurt?"
[0094] FIGS. 7A and 7B illustrate examples of operations for summarizing text data of an electronic device (301) (e.g., electronic device (100), electronic device (1101)).
[0095] Referring to FIG. 7A, in operation 701, at least one processor (300) may receive an input indicating that first text data is to be summarized. For example, the first text data may be composed of multiple languages. For example, the first text data may include text data generated by performing operation 450 of FIG. 4. For example, operation 701 may be represented as an operation subsequent to operation 450 of FIG. 4, but is not limited thereto.
[0096] For example, an input indicating that the first text data is summarized may include a touch input to an executable object (e.g., an executable object (901) of FIG. 9) within a screen (e.g., a screen (900) of FIG. 9) displayed on a display (310) of the electronic device (301).
[0097] In operation 703, at least one processor (300) may identify the language of each of the texts included in the first text data based on an input indicating that the first text data is summarized. For example, the at least one processor (300) may identify a first text in a first language (e.g., English) and a second text in a second language (e.g., Japanese) from the first text data based on the identification. For example, the second text may be represented as text following the first text. For example, the first text may include text (641) of FIG. 6B . For example, the second text may include text (642) of FIG. 6B .
[0098] In operation 705, at least one processor (300) may generate a third text in the first language by performing a translation of the second text. For example, at least one processor (300) may generate the third text by performing a translation of the second text into the first language.
[0099] In operation 707, at least one processor (300) may generate second text data by providing the first text and the third text to a summary model. The second text data may include a summary text of the first text and a summary text of the third text. For example, the second text data may be represented as data that summarizes the first text data. The summary model may include a generative AI model. For example, a generative AI model may generally refer to an artificial intelligence neural network that creates new types of data based on user input information. For example, a generative AI model may include a language generating model. For example, a language generating model is a model trained to output the most statistically appropriate output value based on input values, and representative examples thereof include models such as CHAT-GPT 3 and CHAT-GPT 4. For example, a language-generating model can refer to a language model that can perform inference without fine-tuning using methods such as few-shot learning, and can have more than 10 times more parameters (e.g., about 100 billion or more parameters) than existing general language models. For example, large-scale language models such as GPT-3 (generative pre-trained transformer 3) and GPT-4 (generative pre-trained transformer 4) are excellent few-shot learners that can be controlled through natural text prompts. They can solve NLP (natural language processing) problems by understanding patterns with only a small amount of data through prompts, which is possible with in-context learning.
[0100] In one embodiment, the summary model may be included within the electronic device (301). As a non-limiting example, the summary model may be included within an external electronic device (e.g., a server). Providing the first text and the third text to the summary model by the at least one processor (300) may include transmitting the first text and the third text to the external electronic device (e.g., a server) via a communication circuit (not shown) of the electronic device (301). The at least one processor (300) may receive output data of the summary model of the external electronic device from the external electronic device (e.g., the server).
[0101] Referring to FIG. 7B, at operation 710, at least one processor (300) may receive an input indicating that first text data is to be summarized. For example, the first text data may be composed of multiple languages. For example, the first text data may include text data generated by performing operation 450 of FIG. 4. For example, operation 710 may be represented as an operation subsequent to operation 450 of FIG. 4, but is not limited thereto.
[0102] For example, an input indicating that the first text data is summarized may include a touch input to an executable object (e.g., an executable object (901) of FIG. 9) within a screen (e.g., a screen (900) of FIG. 9) displayed on a display (310) of the electronic device (301).
[0103] In operation 720, at least one processor (300) may identify the language of each of the texts included in the first text data based on an input indicating that the first text data is summarized. For example, the at least one processor (300) may identify a fourth text in the first language and a fifth text in the second language from the first text data based on the identification. For example, the fifth text may be represented as text following the fourth text. For example, the fourth text may include text (641) of FIG. 6B. For example, the fifth text may include text (642) of FIG. 6B.
[0104] In operation 730, at least one processor (300) may identify a third language in which to present a summary of the first text data. For example, the at least one processor (300) may identify the third language based on the language of each text included in the first text data. For example, the at least one processor (300) may identify the language having the largest number of languages among the texts in the first text data as the third language. For example, the at least one processor (300) may identify the language of the longest text among the texts in the first text data as the third language. As a non-limiting example, the third language may be a language preset by the user. For example, the at least one processor (300) may present a summary of the first text data in the third language (e.g., Korean) set by the user.
[0105] In operation 740, at least one processor (300) may perform a summary of the fourth text. For example, the at least one processor (300) may generate a sixth text in a third language based on the summary of the fourth text. For example, the at least one processor (300) may generate a sixth text in a third language representing a summary of the fourth text by providing the fourth text to a generative AI model.
[0106] For example, at least one processor (300) may perform a summary of a text generated by performing a translation of a fourth text. For example, at least one processor (300) may perform a summary of a text translated from a fourth text in a first language into a third language, thereby generating a sixth text in a third language representing a summary of the fourth text.
[0107] For example, at least one processor (300) may bypass (or refrain from) performing a translation of the fourth text based on identifying that the language of the fourth text is the same as the third language. For example, at least one processor (300) may generate a sixth text in the third language by performing a summary of the fourth text based on identifying that the language of the fourth text is the same as the third language.
[0108] In operation 750, at least one processor (300) may perform a summary of the fifth text. For example, the at least one processor (300) may generate a seventh text in a third language based on the summary of the fifth text. For example, the at least one processor (300) may generate a seventh text in a third language representing a summary of the fifth text by providing the fifth text to a generative AI model.
[0109] For example, at least one processor (300) may perform a summary of a fifth text based on a fourth text. For example, at least one processor (300) may perform a summary of a text generated by performing a translation of a fifth text based on the fourth text. For example, at least one processor (300) may perform a summary of a text translated from a fifth text in a second language into a third language based on the fourth text, thereby generating a seventh text in a third language representing a summary of the fifth text. For example, performing a translation of the fifth text based on the fourth text will be described and exemplified in more detail with reference to FIG. 8.
[0110] In operation 760, at least one processor (300) may generate second text data including the sixth text and the seventh text. For example, the second text data may be represented as data that summarizes the first text data. For example, at least one processor (300) may generate summary texts of the texts in the first text data by inputting each text in the first text data to a generative AI model. For example, at least one processor (300) may generate the second text data using the summary texts.
[0111] FIG. 8 illustrates an example of an operation of performing a translation of a fifth text based on a fourth text of an electronic device (e.g., electronic device (100), electronic device (301), electronic device (1101)). The operations of FIG. 8 may be related to operation 750 of FIG. 7b.
[0112] Referring to FIG. 8, in operation 810, at least one processor (300) may generate an eighth text in a second language by performing a translation of a fourth text in a first language. For example, the content of the fourth text and the content of the eighth text may be identical.
[0113] In operation 820, at least one processor (300) may perform a translation of a fifth text based on an eighth text. For example, at least one processor (300) may derive a context from the eighth text by inputting the eighth text into a large-scale language model. For example, at least one processor (300) may perform a translation of the fifth text using the context derived from the eighth text. For example, since the language of the eighth text and the language of the fifth text are the same as the second language, performing a translation of the fifth text using the context derived from the eighth text may result in a higher degree of translation completion than performing a translation of the fifth text without using the context. For example, at least one processor (300) may generate a ninth text in a third language by performing a translation of the fifth text.
[0114] In operation 830, at least one processor (300) may generate a seventh text based on summarizing a ninth text in a third language. For example, at least one processor (300) may perform a summary of the ninth text by inputting the ninth text into a generative AI model.
[0115] FIG. 9 illustrates an example of an electronic device (301) (e.g., electronic device (100), electronic device (1101)) that summarizes text data.
[0116] Referring to FIG. 9, at least one processor (300) may display a screen (900) via a display (310). For example, the screen (900) may include text (902), text (903), and text (904). For example, the text (902) and text (904) may be expressed in a first language (e.g., English). For example, the text (903) may be expressed in a second language (e.g., Japanese).
[0117] For example, the screen (900) may include an executable object (901). For example, at least one processor (300) may receive an input for the executable object (901). For example, the input may include a touch input on the executable object (901), but is not limited thereto.
[0118] For example, at least one processor (300) may perform a summary of text data based on an input to an executable object (901). For example, the text data may include text (902), text (903), and text (904). For example, performing the summary of the text data may include performing a summary of text (902), text (903), and text (904). For example, at least one processor (300) may display a screen (910) via the display (310) based on performing the summary of the text data. For example, at least one processor (300) may identify a third language in which the summary of the text data (e.g., text (920)) is to be represented.
[0119] For example, since the first language corresponds to the largest number among the first language (e.g., English) of the text (902), the second language (e.g., Japanese) of the text (902), and the first language (e.g., English) of the text (903), at least one processor (300) can determine the third language as the first language (e.g., English).
[0120] For example, at least one processor (300) may display a screen (910) through the display (310) based on an input to an executable object (901). For example, the screen (910) may include text (920). For example, the text (920) may be expressed in a third language (e.g., English).
[0121] For example, text (920) may represent other text data generated as at least one processor (300) performs summarization of the text data. For example, text (920) may include keywords of the other text data generated. For example, text (920) may include a title of the other text data generated.
[0122] For example, text (920) may include text (921), text (922), and text (923). For example, text (921) may be represented as text generated by at least one processor (300) performing a summary of text (902). For example, text (902) may include text in a first language, such as "Hello, I'm calling to make an appointment to remove my wisdom tooth. I'm Kim and I'm a regular customer of Mr. Park." For example, text (921) may include text in a third language, such as "Kim, a regular patient, called dental clinic for dental appointment."
[0123] For example, text (922) may be represented as text generated by at least one processor (300) performing a summary of text (903). For example, text (903) may include text in a second language, such as "This is Kim. Reservations can be made by 1 o'clock tomorrow." For example, text (922) may include text in a third language, such as "The reservation is available after 1 o'clock tomorrow." For example, text (903) may be represented as text in a second language generated by applying STT to a second utterance in the second language. For example, since text (922) is text generated by performing a summary of text (903) in the second language, it may include the content of the second utterance more accurately than text (222) of FIG. 2. For example, text (922) may provide a user with an enhanced user experience than text (222) of FIG. 2.
[0124] For example, text (923) may be represented as text generated by at least one processor (300) performing a summary of text (904). For example, text (904) may include text in a first language, such as "Then, please make a reservation for 2 o'clock tomorrow. I have some questions. Is there a visitor's parking lot? How much pain does the wisdom tooth removal hurt?" For example, text (923) may include text in a third language, such as "Kim had two questions: parking availability for visitors, and the pain level of the surgery."
[0125] For example, at least one processor (300) may receive an input for translation of text (920) while displaying a screen (910) through a display (310). For example, at least one processor (300) may translate text (920) into another language (e.g., Korean) based on the input for translation of text (920). For example, at least one processor (300) may display a screen (930) through the display (310) based on the input for translation of text (920). For example, the screen (930) may include text (940). For example, the text (940) may be represented as text translated from text (920) into another language (e.g., Korean). For example, the text (940) may include text (941), text (942), and text (943).
[0126] For example, text (941) may be represented as text translated from text (921) into another language (e.g., Korean). For example, text (941) may include text in another language, such as "Kim, a regular patient at the dental clinic, called to make a dental appointment."
[0127] For example, text (942) may be represented as text translated from text (922) into another language (e.g., Korean). For example, text (941) may include text in another language, such as "Reservations are available after 1:00 PM tomorrow."
[0128] For example, text (943) may be translated from text (923) into another language (e.g., Korean). For example, the other language may be specified by the user. For example, text (941) may include text in another language, such as, "Kim had two questions: whether visitor parking was available and how painful the surgery was."
[0129] FIGS. 10A and 10B illustrate examples of electronic devices (301) (e.g., electronic devices (100), electronic devices (1101)) that translate text data.
[0130] Referring to FIG. 10A, at least one processor (300) can display a screen (1000) through a display (310). For example, the screen (1000) can include text (1001), text (1002), and text (1003).
[0131] For example, at least one processor (300) may receive input for translation of texts (e.g., text (1001), text (1002), and text (1003)) while displaying the screen (1000) through the display (310). For example, the input may include a touch input for each of the texts. For example, the input may include a long press input for each of the texts. For example, the input may include a double tap input for each of the texts. For example, the input may include a touch input for an executable object (not shown) for translation. However, the present invention is not limited thereto.
[0132] For example, at least one processor (300) may perform a translation of text (1001) into a designated language (e.g., Korean) based on the input for text (1001). For example, at least one processor (300) may display a screen (1010) through the display (310) based on the input for text (1001). For example, the screen (1010) may include text (1001), text (1002), and text (1011). For example, the text (1011) may be represented as text generated as the at least one processor (300) performs a translation of text (1001). For example, the text (1011) may be located between text (1001) and text (1002) within the screen (1010). However, the present invention is not limited thereto.
[0133] For example, at least one processor (300) may perform a translation of text (1001) from a language represented by an executable object (1020) to a language represented by an executable object (1030). For example, at least one processor (300) may identify the language of the text (1001). For example, the executable object (1020) may represent the language of the text (1001).
[0134] For example, the executable object (1020) may represent an estimated language by inputting text (1001) to a language detection model. For example, at least one processor (300) may receive an input for the executable object (1020). For example, at least one processor (300) may change the language estimated by the language detection model to a language specified by the user based on the input for the executable object (1020).
[0135] For example, the executable object (1030) may represent a language of text (1011) generated as at least one processor (300) performs a translation of text (1001). For example, the at least one processor (300) may receive an input to the executable object (1030). For example, the at least one processor (300) may change the language represented by the executable object (1030) based on the input to the executable object (1030). For example, the at least one processor (300) may perform a translation of text (1001) into the language represented by the executable object (1030).
[0136] Referring to FIG. 10B, at least one processor (300) can display a screen (1000) through a display (310). For example, at least one processor (300) can perform translation on each of the texts (e.g., text (1001), text (1002), and text (1003)) in the screen (1000). For example, at least one processor (300) can receive an input (1041) for an executable object (1040) for each of the texts. For example, the input (1041) for the executable object (1040) can include a touch input for the executable object (1040). For example, the input (1041) for the executable object (1040) can include a long press input for the executable object (1040). For example, input (1041) for an executable object (1040) may include, but is not limited to, a double-tap input for the executable object (1040).
[0137] An executable object (1040) may represent a speaker corresponding to a text (e.g., text (1001), text (1002), text (1003)). For example, if text (1001) is converted from a first utterance in audio data, an executable object (1040-1) may represent a speaker of the first utterance (e.g., speaker 1). For example, if text (1002) is converted from a second utterance in audio data, an executable object (1040-2) may represent a speaker of the second utterance (e.g., speaker 2). For example, if text (1003) is converted from a third utterance in audio data, an executable object (1040-3) may represent a speaker of the third utterance (e.g., speaker 1).
[0138] At least one processor (300) may display a pop-up window (1050) through the display (310) based on receiving an input (1041) for an executable object (1040-1) for each of the texts. For example, at least one processor (300) may display a pop-up window (1050) as an overlay on the screen (1000) based on receiving an input (1041) for an executable object (1040-1) for each of the texts. The pop-up window (1050) may be described as content for specifying a language in which to perform translation of the text (1001). For example, the pop-up window (1050) may include an executable object (1051), an executable object (1052), an executable object (1053), and an executable object (1054). For example, at least one processor (300) may receive an input for an executable object (1051), an executable object (1052), or an executable object (1053). For example, at least one processor (300) may set a language in which to perform translation of text (1001) to a designated language (e.g., English) based on receiving an input for the executable object (1051). For example, at least one processor (300) may set a language in which to perform translation of text (1001) to a designated language (e.g., Korean) based on receiving an input for the executable object (1052).
[0139] At least one processor (300) may determine a language in which to perform translation of text (1001) based on receiving an input for the executable object (1052). For example, at least one processor (300) may determine a language in which to perform translation of text (1001) based on a translation history. For example, if there is a history of texts in a first language (e.g., the language of the text (1001)) being translated into a second language (e.g., Korean), the at least one processor (300) may determine a language in which to perform translation of text (1001) as the second language. For example, if texts in the first language (e.g., the language of the text (1001)) have been most frequently translated into the second language (e.g., Korean), the at least one processor (300) may determine a language in which to perform translation of text (1001) as the second language.
[0140] At least one processor (300) can perform translation of text (1001) based on receiving input for an executable object (1051), an executable object (1052), or an executable object (1053) and then receiving input for an executable object (1054).
[0141] FIG. 11 is a block diagram of an electronic device (1101) (e.g., electronic device (100), electronic device (301)) within a network environment (1100) according to various embodiments.
[0142] For example, in a network environment (1100), an electronic device (1101) may communicate with an electronic device (1102) via a first network (1198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (1104) or a server (1108) via a second network (1199) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (1101) may communicate with the electronic device (1104) via the server (1108). According to one embodiment, the electronic device (1101) may include a processor (1120), a memory (1130), an input module (1150), an audio output module (1155), a display module (1160), an audio module (1170), a sensor module (1176), an interface (1177), a connection terminal (1178), a haptic module (1179), a camera module (1180), a power management module (1188), a battery (1189), a communication module (1190), a subscriber identification module (1196), or an antenna module (1197). In some embodiments, the electronic device (1101) may omit at least one of these components (e.g., the connection terminal (1178)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1176), camera module (1180), or antenna module (1197)) may be integrated into a single component (e.g., display module (1160)).
[0143] The processor (1120) may, for example, execute software (e.g., a program (1140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (1101) connected to the processor (1120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1120) may store commands or data received from other components (e.g., a sensor module (1176) or a communication module (1190)) in a volatile memory (1132), process the commands or data stored in the volatile memory (1132), and store result data in a non-volatile memory (1134). According to one embodiment, the processor (1120) may include a main processor (1121) (e.g., a central processing unit or an application processor) or an auxiliary processor (1123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1121). For example, when the electronic device (1101) includes the main processor (1121) and the auxiliary processor (1123), the auxiliary processor (1123) may be configured to use less power than the main processor (1121) or to be specialized for a given function. The auxiliary processor (1123) may be implemented separately from the main processor (1121) or as a part thereof.
[0144] The auxiliary processor (1123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (1160), a sensor module (1176), or a communication module (1190)) of the electronic device (1101), for example, on behalf of the main processor (1121) while the main processor (1121) is in an inactive (e.g., sleep) state, or together with the main processor (1121) while the main processor (1121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1180) or a communication module (1190)). In one embodiment, the auxiliary processor (1123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0145] The memory (1130) can store various data used by at least one component (e.g., the processor (1120) or the sensor module (1176)) of the electronic device (1101). The data can include, for example, software (e.g., the program (1140)) and input data or output data for commands related thereto. The memory (1130) can include a volatile memory (1132) or a non-volatile memory (1134).
[0146] The program (1140) may be stored as software in memory (1130) and may include, for example, an operating system (1142), middleware (1144), or an application (1146).
[0147] The input module (1150) can receive commands or data to be used in a component of the electronic device (1101) (e.g., a processor (1120)) from an external source (e.g., a user) of the electronic device (1101). The input module (1150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0148] The audio output module (1155) can output audio signals to the outside of the electronic device (1101). The audio output module (1155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0149] The display module (1160) can visually provide information to an external party (e.g., a user) of the electronic device (1101). The display module (1160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (1160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0150] The audio module (1170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (1170) can acquire sound through the input module (1150), output sound through the sound output module (1155), or an external electronic device (e.g., electronic device (1102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1101).
[0151] The sensor module (1176) can detect the operating status (e.g., power or temperature) of the electronic device (1101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0152] The interface (1177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1101) with an external electronic device (e.g., the electronic device (1102)). In one embodiment, the interface (1177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0153] The connection terminal (1178) may include a connector through which the electronic device (1101) may be physically connected to an external electronic device (e.g., the electronic device (1102)). In one embodiment, the connection terminal (1178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0154] The haptic module (1179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1179) may include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0155] The camera module (1180) can capture still images and videos. In one embodiment, the camera module (1180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0156] The power management module (1188) can manage the power supplied to the electronic device (1101). According to one embodiment, the power management module (1188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0157] A battery (1189) may power at least one component of the electronic device (1101). In one embodiment, the battery (1189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0158] The communication module (1190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1101) and an external electronic device (e.g., electronic device (1102), electronic device (1104), or server (1108)), and the performance of communication through the established communication channel. The communication module (1190) may operate independently from the processor (1120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1190) may include a wireless communication module (1192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1194) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1104) via a first network (1198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1199) (e.g., a long-range communication network such as a legacy cellular network, a fifth generation (5G) network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1192) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1196) to verify or authenticate the electronic device (1101) within a communication network such as the first network (1198) or the second network (1199).
[0159] The wireless communication module (1192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1192) may support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1192) may support various requirements specified in the electronic device (1101), an external electronic device (e.g., the electronic device (1104)), or a network system (e.g., the second network (1199)). According to one embodiment, the wireless communication module (1192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB implementation, a loss coverage (e.g., 164 dB or less) for mMTC implementation, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC implementation.
[0160] The antenna module (1197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1198) or the second network (1199), may be selected from the plurality of antennas by, for example, the communication module (1190). A signal or power may be transmitted or received between the communication module (1190) and an external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1197).
[0161] According to various embodiments, the antenna module (1197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0162] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0163] According to one embodiment, commands or data may be transmitted or received between the electronic device (1101) and an external electronic device (1104) via a server (1108) connected to a second network (1199). Each of the external electronic devices (1102 or 1104) may be the same or a different type of device as the electronic device (1101). According to one embodiment, all or part of the operations executed in the electronic device (1101) may be executed in one or more of the external electronic devices (1102, 1104, or 1108). For example, when the electronic device (1101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1101). The electronic device (1101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (1104) may include an Internet of Things (IoT) device. The server (1108) may be an intelligent server utilizing machine learning and / or a neural network.According to one embodiment, an external electronic device (1104) or server (1108) may be included within the second network (1199). The electronic device (1101) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.
[0164] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0165] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0166] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0167] Various embodiments of the present document may be implemented as software (e.g., a program (1140)) including one or more instructions stored in a storage medium (e.g., an internal memory (1136) or an external memory (1138)) readable by a machine (e.g., an electronic device (1101)). For example, a processor (e.g., a processor (1120)) of the machine (e.g., an electronic device (1101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0168] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0169] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0170] An electronic device as described above may include a memory storing instructions and including one or more storage media. The electronic device may include at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input indicating that audio data is to be converted into text data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, from the audio data, a first utterance in a first language and a second utterance in a second language subsequent to the first utterance, based on the input identifying the language of each of the utterances included in the audio data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a first text in the first language by applying a speech-to-text (STT) operation to the first utterance in the first language. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a second text in the second language by applying STT to the second utterance in the second language based on the first text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate text data comprising the first text and the second text.
[0171] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a third text in the second language by performing a translation of the first text based on identifying that the first language is different from the second language. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second text by applying STT to the second utterance in the second language based on the third text.
[0172] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second text by applying STT to the second utterance in the second language using a context derived from the third text.
[0173] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to refrain from performing a translation of the first text based on identifying that the first language is identical to the second language. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second text by applying STT to the second utterance in the second language using context derived from the first text.
[0174] According to one embodiment, the electronic device may include a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display a screen including the first text and the second text through the display based on the text data.
[0175] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the first text by applying STT to the first utterance in the first language based on a context derived from the first utterance.
[0176] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input indicating that the text data includes the first text and the second text, and then translate the language of the text data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to, based on the input, perform a translation of the first text to generate a third text in a third language, and to perform a translation of the second text to generate a fourth text in a fourth language. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate other text data including the third text and the fourth text.
[0177] According to one embodiment, the text data including the first text and the second text may be first text data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input indicating that the first text data is to be summarized after generating the first text data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a third language in which to present a summary of the first text data based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a third text in the third language by performing a translation of the first text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a fourth text in the third language by performing a translation of the second text based on the first text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate second text data including the third text and the fourth text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate third text data by performing a summary of the second text data.
[0178] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a fifth text in the second language by performing a translation of the first text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the fourth text by performing a translation of the second text based on the fifth text.
[0179] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the fourth text by performing a translation of the second text using context derived from the fifth text.
[0180] A method performed by an electronic device as described above may include receiving an input indicating that audio data is to be converted into text data. The method may include identifying, from the audio data, a first utterance in a first language and a second utterance in a second language subsequent to the first utterance by identifying the language of each of the utterances included in the audio data based on the input. The method may include generating a first text in the first language by applying a speech-to-text (STT) operation to the first utterance in the first language. The method may include generating a second text in the second language by applying a STT operation to the second utterance in the second language based on the first text. The method may include generating text data including the first text and the second text.
[0181] In one embodiment, the method may include generating a third text in the second language by performing a translation of the first text based on identifying that the first language is different from the second language. The method may further include generating the second text by applying STT to the second utterance in the second language based on the third text.
[0182] In one embodiment, the method may include generating the second text by applying STT to the second utterance in the second language using a context derived from the third text.
[0183] In one embodiment, the method may include an action of refraining from performing a translation of the first text based on identifying that the first language is identical to the second language. The method may further include an action of generating the second text by applying STT to the second utterance in the second language using context derived from the first text.
[0184] According to one embodiment, the electronic device may further include a display. The method may include an operation of displaying a screen including the first text and the second text through the display based on the text data.
[0185] In one embodiment, the method may include generating the first text by applying STT to the first utterance in the first language based on a context derived from the first utterance.
[0186] In one embodiment, the method may include an operation of receiving an input indicating that a language of the text data is to be translated after generating the text data including the first text and the second text. The method may include an operation of generating a third text in a third language by performing a translation of the first text based on the input, and a fourth text in a fourth language by performing a translation of the second text. The method may include an operation of generating other text data including the third text and the fourth text.
[0187] According to one embodiment, the text data including the first text and the second text may be first text data. The method may include an operation of receiving, after generating the first text data, an input indicating that the first text data is to be summarized. The method may include an operation of identifying, based on the input, a third language in which a summary of the first text data is to be presented. The method may include an operation of generating a third text in the third language by performing a translation of the first text. The method may include an operation of generating a fourth text in the third language by performing a translation of the second text based on the first text. The method may include an operation of generating second text data including the third text and the fourth text. The method may include an operation of generating the third text data by performing a summary of the second text data.
[0188] In one embodiment, the method may include generating a fifth text in the second language by performing a translation of the first text. The method may include generating a fourth text by performing a translation of the second text based on the fifth text.
[0189] In one embodiment, the method may include generating the fourth text by performing a translation of the second text using a context derived from the fifth text.
[0190] In a computer-readable storage medium having one or more programs stored thereon, as described above, the one or more programs may include instructions that, when executed by an electronic device, cause the electronic device to receive an input indicating that audio data is to be converted into text data. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, from the audio data, a first utterance in a first language and a second utterance in a second language subsequent to the first utterance, based on the input identifying the language of each of the utterances included in the audio data. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a first text in the first language by applying a speech-to-text (STT) operation to the first utterance in the first language. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a second text in the second language by applying STT to the second utterance in the second language based on the first text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate text data including the first text and the second text.
[0191] In one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a third text in the second language by performing a translation of the first text based on identifying that the first language is different from the second language. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate the second text by applying STT to the second utterance in the second language based on the third text.
[0192] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate the second text by applying STT to the second utterance in the second language using a context derived from the third text.
[0193] In one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to refrain from performing a translation of the first text based on identifying that the first language is identical to the second language. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate the second text by applying STT to the second utterance in the second language using context derived from the first text.
[0194] According to one embodiment, the electronic device may include a display. The one or more programs, when executed by the electronic device, may include instructions that cause the electronic device to display a screen including the first text and the second text through the display based on the text data.
[0195] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate the first text by applying STT to the first utterance in the first language based on a context derived from the first utterance.
[0196] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive an input indicating that a language of the text data is to be translated after generating the text data including the first text and the second text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to, based on the input, perform a translation of the first text to generate a third text in a third language, and to perform a translation of the second text to generate a fourth text in a fourth language. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate other text data including the third text and the fourth text.
[0197] According to one embodiment, the text data including the first text and the second text may be first text data. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive an input indicating that the first text data is to be summarized after generating the first text data. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify a third language in which to present a summary of the first text data based on the input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a third text in the third language by performing a translation of the first text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a fourth text in the third language by performing a translation of the second text based on the first text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate second text data including the third text and the fourth text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate third text data by performing a summary of the second text data.
[0198] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a fifth text in the second language by performing a translation of the first text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate the fourth text by performing a translation of the second text based on the fifth text.
[0199] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate the fourth text by performing a translation of the second text using a context derived from the fifth text.
[0200] An electronic device as described above may include a memory storing instructions and including one or more storage media. The electronic device may include at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input indicating that first text data is to be summarized. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, from the first text data, a first text in a first language and a second text in a second language subsequent to the first text, based on identifying a language of each of the texts included in the first text data based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, based on the language of each of the texts included in the first text data, a third language corresponding to a maximum number of languages among the texts. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a third text in the third language based on performing a summary of text generated by performing a translation of the first text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a fourth text in the third language based on performing a summary of text generated by performing a translation of the second text based on the first text.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate second text data including the third text and the fourth text.
[0201] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to bypass performing a translation of the first text based on identifying that the first language is identical to the third language, and to generate the third text based on performing a summary of the first text.
[0202] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a fifth text in the second language by performing a translation of the first text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a sixth text in the third language by performing a translation of the second text based on the fifth text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the fourth text based on performing a summary of the sixth text.
[0203] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the sixth text by performing a translation of the second text using a context derived from the fifth text.
[0204] According to one embodiment, the electronic device may include a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display a screen including the third text and the fourth text through the display, based on the second text data.
[0205] A method performed by an electronic device as described above may include receiving an input indicating that first text data is to be summarized. The method may include identifying a language of each of the texts included in the first text data based on the input, and identifying, from the first text data, a first text in a first language and a second text in a second language following the first text. The method may include identifying, based on the language of each of the texts included in the first text data, a third language corresponding to a maximum number of languages among the texts. The method may include generating a third text in the third language based on performing a summary of a text generated by performing a translation of the first text. The method may include generating a fourth text in the third language based on performing a summary of a text generated by performing a translation of the second text based on the first text. The method may include generating second text data including the third text and the fourth text.
[0206] In one embodiment, the method may include bypassing performing a translation of the first text based on identifying that the first language is identical to the third language, and generating the third text based on performing a summary of the first text.
[0207] In one embodiment, the method may include generating a fifth text in the second language by performing a translation of the first text. The method may include generating a sixth text in the third language by performing a translation of the second text based on the fifth text. The method may include generating the fourth text based on performing a summary of the sixth text.
[0208] According to one embodiment, the method may include generating the sixth text by performing a translation of the second text using a context derived from the fifth text.
[0209] According to one embodiment, the electronic device may include a display. The method may include an operation of displaying a screen including the third text and the fourth text through the display based on the second text data.
[0210] In a computer-readable storage medium having one or more programs stored thereon, as described above, the one or more programs may include instructions that, when executed by an electronic device, cause the electronic device to receive an input indicating that first text data is to be summarized. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, from the first text data, a first text in a first language and a second text in a second language subsequent to the first text, based on identifying a language of each of the texts included in the first text data based on the input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, based on the language of each of the texts included in the first text data, a third language corresponding to a maximum number of languages among the texts. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a third text in the third language based on performing a summary of text generated by performing a translation of the first text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a fourth text in the third language based on performing a summary of text generated by performing a translation of the second text based on the first text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate second text data including the third text and the fourth text.
[0211] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to bypass performing a translation of the first text based on identifying that the first language is identical to the third language, and to generate the third text based on performing a summary of the first text.
[0212] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a fifth text in the second language by performing a translation of the first text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate a sixth text in the third language by performing a translation of the second text based on the fifth text. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate the fourth text based on performing a summary of the sixth text.
[0213] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate the sixth text by performing a translation of the second text using a context derived from the fifth text.
[0214] According to one embodiment, the electronic device may include a display. The one or more programs, when executed by the electronic device, may include instructions that cause the electronic device to display, through the display, a screen including the third text and the fourth text based on the second text data.
Claims
1. In electronic devices, A memory storing instructions and including one or more storage media; and At least one processor comprising a processing circuit, The above instructions, when individually or collectively executed by the at least one processor, Receives an input indicating that audio data is to be converted into text data, Based on the above input, identifying the language of each of the utterances included in the audio data, and identifying a first utterance of a first language and a second utterance of a second language following the first utterance from the audio data, By applying STT (speech to text) to the first utterance of the first language, a first text of the first language is generated, By applying STT to the second utterance of the second language based on the first text, a second text of the second language is generated, and To generate text data including the first text and the second text, causing the above electronic device, Electronic devices.
2. In claim 1, the instructions, when individually or collectively executed by the at least one processor, generating a third text in the second language by performing a translation of the first text based on identifying that the first language is different from the second language, and By applying STT to the second utterance of the second language based on the third text, the second text is generated. causing the above electronic device, Electronic devices.
3. In claim 2, the instructions, when individually or collectively executed by the at least one processor, By applying STT to the second utterance of the second language using the context derived from the third text, the second text is generated. causing the above electronic device, Electronic devices.
4. In claim 1, the instructions, when individually or collectively executed by the at least one processor, Refrain from performing a translation of the first text based on identifying that the first language is identical to the second language, and To generate the second text by applying STT to the second utterance of the second language using the context derived from the first text, causing the above electronic device, Electronic devices.
5. In claim 1, Including more displays, The above instructions, when individually or collectively executed by the at least one processor, Based on the above text data, a screen including the first text and the second text is displayed through the display. causing the above electronic device, Electronic devices.
6. In claim 1, the instructions, when individually or collectively executed by the at least one processor, By applying STT to the first utterance of the first language based on the context derived from the first utterance, the first text is generated. causing the above electronic device, Electronic devices.
7. In claim 1, the instructions, when individually or collectively executed by the at least one processor, After generating the text data including the first text and the second text, receiving an input indicating that the language of the text data is to be translated, Based on the input, a third text in a third language is generated by performing a translation of the first text, and a fourth text in a fourth language is generated by performing a translation of the second text, and To generate other text data including the third text and the fourth text, causing the above electronic device, Electronic devices.
8. In claim 1, The text data including the first text and the second text is first text data, The above instructions, when individually or collectively executed by the at least one processor, After generating the first text data, receiving an input indicating that the first text data is summarized, Based on the above input, identify a third language in which to present a summary of the first text data, By performing a translation of the first text, a third text in the third language is generated, By performing a translation of the second text based on the first text, a fourth text in the third language is generated, Generating second text data including the third text and the fourth text, and To generate third text data by performing a summary of the second text data, causing the above electronic device, Electronic devices.
9. In claim 8, the instructions, when individually or collectively executed by the at least one processor, Generating a fifth text in the second language by performing a translation of the first text, and By performing a translation of the second text based on the fifth text, the fourth text is generated. causing the above electronic device, Electronic devices.
10. In claim 9, the instructions, when individually or collectively executed by the at least one processor, By performing a translation of the second text using the context derived from the fifth text, the fourth text is generated. causing the above electronic device, Electronic devices.
11. In electronic devices, A memory storing instructions and including one or more storage media; and At least one processor comprising a processing circuit, The above instructions, when individually or collectively executed by the at least one processor, Receives an input indicating that the first text data is to be summarized, Based on the input, identifying the language of each of the texts included in the first text data, and identifying a first text in a first language and a second text in a second language following the first text from the first text data, Based on the language of each of the texts included in the first text data, a third language corresponding to the largest number of languages among the texts is identified, Generating a third text in the third language based on performing a summary of the text generated by performing the translation of the first text, Generating a fourth text in the third language based on summarizing the text generated by performing a translation of the second text based on the first text, and To generate second text data including the third text and the fourth text, causing the above electronic device, Electronic devices.
12. In claim 11, the instructions, when individually or collectively executed by the at least one processor, Based on identifying that the first language is identical to the third language: Bypassing the translation of the above first text, and To generate the third text based on performing the summary of the first text, causing the above electronic device, Electronic devices.
13. In claim 11, the instructions, when individually or collectively executed by the at least one processor, Generating a fifth text in the second language by performing a translation of the first text, By performing a translation of the second text based on the fifth text, a sixth text in the third language is generated, and Based on performing the summary of the above 6th text, to generate the above 4th text, causing the above electronic device, Electronic devices.
14. In claim 13, the instructions, when individually or collectively executed by the at least one processor, By performing a translation of the second text using the context derived from the fifth text, the sixth text is generated. causing the above electronic device, Electronic devices.
15. In claim 11, Including more displays, The above instructions, when individually or collectively executed by the at least one processor, Based on the second text data, a screen including the third text and the fourth text is displayed through the display. causing the above electronic device, Electronic devices.
Citation Information
Patent Citations
Information processor and information processing program
JP2023120030A
Speech recognition method, apparatus, and program
JP4000095B2
Device for extracting information from a dialog
KR1020140142280A
Transparent coating removal through laser ablation
KR1020230023589A
device for recognizing voice and method for recognizing voice
KR102084646B1