Text-providing method and electronic device for performing method

The text providing method and electronic device address the challenges of language barriers by using a trained model for real-time voice-to-text translation with high accuracy, enhancing communication efficiency and reducing latency.

WO2026155564A1PCT designated stage Publication Date: 2026-07-23SCONAI INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SCONAI INC
Filing Date
2026-01-15
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing text translation methods, particularly those using artificial intelligence models, face challenges with accuracy and latency due to error propagation and language barriers, making real-time communication difficult across different languages.

Method used

A text providing method and electronic device that utilizes a model trained to convert voice input of a first language into text in a second language, providing real-time translation with high accuracy by determining the meaning of voice input and outputting text simultaneously.

Benefits of technology

Enables real-time, high-accuracy translation of voice input into text, reducing latency and user discomfort by eliminating the need for separate input and verification steps, and accurately identifying speakers across multiple languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2026000910_23072026_PF_FP_ABST
    Figure KR2026000910_23072026_PF_FP_ABST
Patent Text Reader

Abstract

A text-providing method and an electronic device for performing the method according to various embodiments are disclosed. An electronic device according to an embodiment may comprise: a processor; and a memory electrically connected to the processor and storing at least one instruction executed by the processor, wherein, when the at least one instruction is executed, the processor causes the electronic device to receive a voice input of a first language from a first terminal, determine text of a second language corresponding to the voice input of the first language, and provide the text of the second language to a second terminal, wherein at least a part of the text of the second language is provided to the second terminal while the voice input of the first language is being received.
Need to check novelty before this filing date? Find Prior Art

Description

Text provision method and electronic device performing the above method

[0001] The present invention relates to a text provision method and an electronic device for performing said method.

[0002] Due to the development of internet technology, mobile technology, and information and communication technology, the frequency of communication, such as conversations and chatting with people who speak other languages, is increasing compared to the past.

[0003] Furthermore, due to the development of the tourism industry, the number of people experiencing overseas travel has increased compared to the past; however, there are many cases where travelers experience inconvenience due to language barriers, and these barriers also limit the experiences and opportunities available from overseas travel.

[0004] Furthermore, in order to provide goods and services in the global market, collaboration and transactions with companies in countries using different languages ​​are increasing. In the preparatory stage for entering the global market, products and services can be promoted through various events such as exhibitions and trade fairs. However, when promoting products and services through exhibitions and trade fairs, it can be difficult to properly promote or respond to people who speak a different language due to language barriers.

[0005] The background technology described above is possessed or acquired by the inventor in the process of deriving the content of the disclosure of the present application, and cannot necessarily be considered as prior art disclosed to the general public prior to the filing of this application.

[0006] In the case of services that simply provide text translation, there is the inconvenience of each user having to input the text into a device and then verify the translation for other users. Furthermore, unlike real-time conversations, text translation services require procedures such as text input, translation output, and translation verification, making it difficult to ensure smooth communication.

[0007] When using an artificial intelligence model to convert voice input in a first language into text in the first language and convert the converted text in the first language into text in a second language, the accuracy of the provided text in the second language may decrease and the latency may be long due to error propagation and distribution mismatch.

[0008] A text providing method and electronic device according to one embodiment of the present invention can provide text in a second language with high translation accuracy by using a model trained to input voice input of a first language and output text in a second language.

[0009] A text providing method and electronic device according to one embodiment of the present invention can provide voice input of a first language as text of a second language.

[0010] A text providing method and electronic device according to one embodiment of the present invention can provide text in a second language corresponding to a voice input with a determined meaning while voice is being input.

[0011] A text providing method and electronic device according to one embodiment of the present invention can provide text in a second language by using a model trained to input voice input of a first language and output text in a second language.

[0012] However, technical challenges are not limited to the technical challenges described above, and other technical challenges may exist.

[0013] An electronic device according to various embodiments includes a processor and a memory electrically connected to the processor and storing at least one instruction executed by the processor, wherein when the at least one instruction is executed, the processor causes the electronic device to receive voice input of a first language from a first terminal, determine text of a second language corresponding to the voice input of the first language, and provide the text of the second language to a second terminal, wherein at least a portion of the text of the second language may be provided to the second terminal while the voice input of the first language is received.

[0014] The processor may provide text of the second language through at least one chat window among the first terminal and the second terminal.

[0015] When a plurality of voice inputs are received from the first terminal, the processor provides text corresponding to the plurality of voice inputs to the first terminal, and can determine the speaker of the voice input of the first language based on the input received from the first terminal.

[0016] When a speaker corresponding to the second terminal is registered, the processor determines a voice input excluding the voice input of the speaker corresponding to the second terminal from among the voice inputs input from the first terminal, and can determine the speaker of the voice input of the first language using the voice input excluding the voice input of the speaker corresponding to the second terminal.

[0017] The above processor provides an interface for determining the speaker of a voice input of the first language through the first terminal, and can determine the speaker by using the voice input received in response to the interface.

[0018] The processor can determine the text of the second language based on at least one of the text corresponding to the previously entered voice input of the first language, the previously provided text of the second language, and a set keyword.

[0019] The processor can determine the characteristics of the voice input of the first language, generate a voice output corresponding to the text of the second language based on the characteristics, and provide the voice output corresponding to the text of the second language to the second terminal.

[0020] The processor may provide text of a third language corresponding to voice input of the first language to a third terminal, provided that at least a portion of the text of the third language is provided while voice input of the first language is being received.

[0021] The processor receives voice input of a second language through a second terminal and provides text of a first language corresponding to the voice input of the second language to the first terminal, wherein at least a portion of the text of the first language may be provided while the voice input of the second language is being received.

[0022] The processor may provide at least one of a conversation record and a conversation summary to the first terminal or the second terminal based on the text of the first language and the text of the second language.

[0023] The processor can generate text in the second language by using a model trained to input voice input in the first language and output text in the second language.

[0024] The processor may, using the model, determine at least one of a part of the text of the second language corresponding to the input voice input and a part of the text of the second language with a determined meaning and a part of the text of the second language with an undetermined meaning while the voice input of the first language is received, and provide at least one of the text of the second language with a determined meaning and the text of the second language with an undetermined meaning to the second terminal while the voice input of the first language is received.

[0025] By using a model trained to input and output text in a second language, it is possible to provide text in a second language with high translation accuracy.

[0026] According to one embodiment of the present invention, a text providing method and an electronic device can reduce the latency of providing text in a second language by providing at least a portion of text in a second language while voice input in a first language is being input.

[0027] According to one embodiment of the present invention, a text providing method and an electronic device can identify an accurate user by identifying a speaker who inputs voice to a terminal.

[0028] According to one embodiment of the present invention, a text providing method and an electronic device can provide text in a plurality of languages ​​(e.g., a second language, a third language, ...) corresponding to the voice input of a first language when voice input of a first language is input.

[0029] According to one embodiment of the present invention, a text providing method and an electronic device can reduce discomfort and confusion caused by a speaker's voice mismatch by converting text of a second language into voice output and providing it according to the characteristics of voice input of a first language.

[0030] According to one embodiment of the present invention, the text providing method and electronic device can improve the convenience of use because, when users of different languages ​​converse, separate button control, input, etc. are not required, and only voice is input to each terminal.

[0031] FIGS. 1 and FIGS. 2 are schematic block diagrams of electronic devices according to various embodiments.

[0032] FIG. 3 is a schematic block diagram of a model according to various embodiments.

[0033] FIG. 4 is a flowchart of the operation of a text provision method according to various embodiments.

[0034] FIG. 5 is a flowchart of the operation of a text provision method according to various embodiments.

[0035] FIGS. 6 and 7 are drawings showing a chat window provided by an electronic device according to various embodiments.

[0036] FIG. 8 is a diagram showing the time when an electronic device according to various embodiments outputs text in response to voice input.

[0037] FIG. 9 is a diagram showing voice input of a first language and text of a second language according to various embodiments.

[0038] FIGS. 10, FIGS. 11, FIGS. 12, FIGS. 13, FIGS. 14, FIGS. 15 and FIGS. 16 are drawings illustrating text of a first language and text of a second language provided by an electronic device according to various embodiments.

[0039] FIGS. 17 and 18 are drawings illustrating voice input of a first language and text of a second language according to various embodiments.

[0040] FIGS. 19, 20, and 21 are drawings showing text of a first language and text of a second language provided by an electronic device according to various embodiments.

[0041] FIG. 22 is a diagram showing an interface provided by an electronic device according to various embodiments.

[0042] FIGS. 23 and 24 are drawings illustrating voice input of a first language and text of a second language according to various embodiments.

[0043] FIGS. 25 and 26 are drawings illustrating a conversation record and a conversation summary provided by an electronic device according to various embodiments.

[0044] FIGS. 27, 28, and 29 are diagrams illustrating the operation flowcharts of an electronic device identifying a user of a terminal according to various embodiments.

[0045] Hereinafter, embodiments are described in detail with reference to the attached drawings. However, various modifications may be made to the embodiments, and thus the scope of the patent application is not limited or restricted by these embodiments. It should be understood that all modifications, equivalents, and substitutions to the embodiments are included within the scope of the rights.

[0046] The terms used in the embodiments are for illustrative purposes only and should not be interpreted as intended to be limiting. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "comprising" or "having" are intended to indicate the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0047]

[0048] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the embodiments pertain. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.

[0049] In addition, when describing with reference to the attached drawings, identical components are assigned the same reference numeral regardless of drawing symbols, and redundant descriptions thereof are omitted. In describing the embodiments, if it is determined that a detailed description of related prior art could unnecessarily obscure the essence of the embodiments, such detailed description is omitted.

[0050] FIGS. 1 and FIGS. 2 are schematic block diagrams of an electronic device (100) according to various embodiments.

[0051] FIG. 1 is a schematic block diagram of an electronic device (100) according to various embodiments.

[0052] Referring to FIG. 1, an electronic device (100) according to one embodiment may include a processor (110), memory (120), model (130) and / or a communication circuit (140).

[0053] The processor (110) can, for example, execute software (e.g., a program) to control at least one other component (e.g., a hardware or software component) of the electronic device (100) connected to the processor (110) and perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (110) can store commands or data received from other components (e.g., a sensor module or a communication module) in volatile memory, process the commands or data stored in volatile memory, and store the resulting data in non-volatile memory. According to one embodiment, the processor (110) may include a main processor (e.g., a central processing unit or an application processor) or an auxiliary processor (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with it. For example, if the electronic device (100) includes a main processor and an auxiliary processor, the auxiliary processor may be configured to use less power than the main processor or to be specialized for a designated function. The auxiliary processor can be implemented separately from the main processor or as part of it.

[0054] An auxiliary processor can control at least some of the functions or states associated with at least one component (e.g., a display module, a sensor module, or a communication module) of the electronic device (100), for example, on behalf of the main processor while the main processor is in an inactive (e.g., sleep) state, or together with the main processor while the main processor is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (e.g., an image signal processor or a communication processor) may be implemented as part of another functionally related component (e.g., a camera module or a communication module). According to one embodiment, the auxiliary processor (e.g., a neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (100) itself where the artificial intelligence model is executed, or through a separate server.

[0055] The memory (120) can store various data used by at least one component of the electronic device (100) (e.g., processor (110) or sensor module). The data may include, for example, software (e.g., program) and input data or output data for related commands. The memory may include volatile memory or non-volatile memory.

[0056] For example, the model (130) can be trained to input a voice input of a first language and output text of a second language. For example, the training data of the model (130) may include a voice input of a first language, text of a first language, and text of a second language. Among the training data, the text of the second language may be ground truth data.

[0057] In the following description, the description and operation regarding the model (130) may be applied substantially identically to the description and operation regarding the electronic device (100) or the processor (110). For the convenience of explanation, the model (130) is described as outputting text in a second language using voice input of a first language, but this can be understood as substantially identical to the operation of the electronic device (100) or the processor (110). Furthermore, the description regarding the operation of the model (130) determining whether the sentence of the voice input of the first language (or the text of the second language) has ended or whether its meaning has been extended may also be understood as an explanation regarding the operation of the electronic device (100) or the processor (110).

[0058] According to one embodiment, the model (130) may be trained to receive voice input of a first language and output text of the first language and / or text of the second language. The model (130) may output text of the first language and / or text of the second language using the input voice input of the first language. A method (first method) in which the model (130) outputs text of the first language and / or text of the second language may be distinguished from a method (second method) of converting voice input of the first language into text of the first language and converting (or translating) the converted text of the first language into text of the second language.

[0059] The second method may be a method of cascading two models. The two models may be a model that converts speech input of the first language into text of the first language and a model that converts the converted text of the first language into text of the second language.

[0060] The first method uses a learned model (130), the input of the model (130) may be a speech input of the first language, and the output may be text of the first language and / or text of the second language.

[0061] Since the first method uses one model (130), it can prevent or mitigate the decrease in accuracy of the output due to error propagation and / or distribution mismatch.

[0062] For example, while voice input of the first language is being received, the model (130) can determine a part of text of the second language corresponding to the input voice input that has a determined meaning and / or a part that has not a determined meaning. Alternatively, while voice input of the first language is being received, the model (130) can determine a part of text of the input voice input that has a determined meaning and / or a part that has not a determined meaning.

[0063] For example, the voice input of the first language may include a first sentence, a second sentence, and a third sentence. The first sentence, the second sentence, and the third sentence may each be input sequentially and may each have a fixed meaning. The text of the second language corresponding to the voice input of the first language may include a fourth sentence, a fifth sentence, and a sixth sentence. The fourth sentence, the fifth sentence, and the sixth sentence may each correspond to the first sentence, the second sentence, and the third sentence.

[0064] The model (130) may determine that when voice input is input, if voice input of the first sentence is input, the meaning of the text of the second language (or the fourth sentence) corresponding to the first sentence in the voice input is determined. Alternatively, the model (130) may determine that when voice input of the first sentence is input, the meaning of the first sentence in the voice input is determined. The model (130) may output the fourth sentence using the first sentence in which the meaning is determined in the voice input.

[0065] For example, the model (130) can determine whether a sentence included within a voice input is completed and / or whether its meaning is determined as a voice input is received. The model (130) can calculate a probability regarding whether a sentence is completed and / or whether its meaning is determined as a voice input is received. By comparing the probability with a set value, the model (130) can determine whether a sentence is completed and / or whether its meaning is determined within the voice input received so far.

[0066] After the first sentence of the voice input is input, if the voice input of the second sentence is input, the model (130) may determine that the meaning of the text of the second language (or the fifth sentence) corresponding to the second sentence of the voice input is determined. Alternatively, the model (130) may determine that the meaning of the second sentence of the voice input is determined when the voice input of the second sentence is input. The model (130) may output the fifth sentence using the second sentence of the voice input whose meaning has been determined.

[0067] When a third voice input is input after a second voice input has been input, the above description may be applied substantially the same way to the operation of the model (130).

[0068] For example, the communication circuit (140) can establish a direct communication channel or a wireless communication channel between the electronic device (100) and a terminal (e.g., a first terminal (201-1), a second terminal (201-2), ..., an nth terminal (200-n)), and support wired / wireless communication through the established communication channel.

[0069] For example, the electronic device (100) can receive voice input of a first language from a first terminal (200-1). For example, the first terminal (200-1) may include a voice input device.

[0070] For example, the electronic device (100) can determine text in a second language corresponding to voice input in a first language. For example, the electronic device (100) can determine text in a second language using a model (130). The text in the second language may be text that translates voice input in the first language into the second language.

[0071] For example, the electronic device (100) may provide text of a second language to a second terminal (200-2). The electronic device (100) may provide at least a portion of the text of the second language to the second terminal (200-2) while voice input of the first language is being received. The electronic device (100) may simultaneously perform the operation of receiving voice input of the first language and / or the operation of determining text of the second language and the operation of providing at least a portion of the text of the second language to the second terminal (200-2).

[0072] For example, at least a portion of the text in the second language provided while receiving voice input in the first language (or while determining text in the second language) may correspond to a portion of the voice input in the first language whose meaning has been determined.

[0073] Alternatively, at least a portion of the text in the second language provided while receiving voice input in the first language (or while determining text in the second language) may represent a portion of the text in the second language whose meaning has been determined. The portion of the text in the second language whose meaning has been determined may correspond to the voice input in the first language entered up to that point in time.

[0074] The electronic device (100) can provide interpretation / translation for voice input in real time by receiving voice input of a first language and simultaneously providing at least a portion of text of a second language that translates the voice input of the first language.

[0075] As another example, at least a portion of the text in the second language provided while receiving voice input in the first language (or while determining text in the second language) may correspond to parts of the voice input in the first language where the meaning is not determined and / or parts where the meaning is determined. For example, as voice input in the first language is received, the meaning of the voice input in the first language may change from an undetermined state to a determined state.

[0076] For example, the electronic device (100) may provide text in a second language corresponding to a voice input of a first language whose meaning has not been determined. As the voice input of the first language is input, when the meaning of the voice input of the first language is determined, the electronic device (100) may provide text in a second language corresponding to the voice input of the first language whose meaning has been determined. The electronic device (100) may convert the text in the second language corresponding to the voice input of the first language whose meaning has been determined into a voice output. The electronic device (100) may generate a voice output using the text in the second language corresponding to the voice input of the first language whose meaning has been determined. The electronic device (100) may provide the voice output to the voice input of the first language whose meaning has been determined to the first terminal (200-1) and / or the second terminal (200-2).

[0077] For example, the electronic device (100) can determine text of the first language corresponding to a voice input of the first language. For example, the electronic device (100) can determine text of the first language and / or text of the second language using a model (130). The model (130) can be trained to output text of the first language and / or text of the second language when a voice input of the first language is input.

[0078] For example, the electronic device (100) may provide text of a first language to a first terminal (200-1) and / or a second terminal (200-2). For example, the electronic device (100) may provide text of a second language to a first terminal (200-1) and / or a second terminal (200-2).

[0079] For example, the electronic device (100) can provide text of a first language and / or text of a second language to a first terminal (200-1) and / or a second terminal (200-2) through an interface such as a chat room.

[0080] The first terminal (200-1) and / or the second terminal (200-2) may include a display. The first terminal (200-1) and / or the second terminal (200-2) may display text in the first language and / or text in the second language on the display through an interface such as a chat room.

[0081] The text of the first language and / or the text of the second language displayed on the display of the first terminal (200-1) may be text displaying voice input spoken by the speaker (or user of the first terminal (200-1)).

[0082] The text of the first language and / or the text of the second language displayed on the display of the second terminal (200-2) may be text displaying voice input spoken by the other party (or the user of the second terminal (200-2)).

[0083] For example, the electronic device (100) can determine the characteristics of the voice input of the first language. For example, the electronic device (100) can extract the characteristics of the voice input from the voice input of the first language. Regarding the operation of the electronic device (100) for determining the characteristics of the voice input, known methods, algorithms, etc. for determining and extracting the characteristics of the voice input may be applied.

[0084] For example, the electronic device (100) can generate a voice output corresponding to text in a second language based on features. The electronic device (100) can provide the voice output corresponding to text in a second language to a second terminal (200-2).

[0085] The electronic device (100) can reduce user inconvenience caused by differences in voice, gender, tone, etc. between the voice output and the voice input by providing voice output according to the characteristics of the input voice input.

[0086] For example, the electronic device (100) can determine text in a third language corresponding to voice input in a first language. The electronic device (100) can determine text in a third language corresponding to voice input in a first language by using a model. The electronic device (100) can provide text in a third language to another terminal (e.g., a third terminal).

[0087] For example, the model (130) can be trained to output text in a second language, text in a third language, ..., text in the nth language when voice input in a first language is input. The model (130) can be trained to output text in multiple languages ​​when voice input in one language is input.

[0088] For example, the type of the first language may not be fixed. For example, when Korean voice input is input, the model (130) may output text in multiple languages ​​(e.g., Japanese, Chinese, English, French, etc.). For example, when Japanese voice input is input, the model (130) may output text in multiple languages ​​(e.g., Korean, Chinese, English, French, etc.).

[0089] As described above, voice inputs to the model (130) may be voice inputs in various languages. When voice input in one of the various languages ​​is input, the model (130) may output text in a language different from the input language.

[0090] In other words, the model (130) can be trained to output text in a language different from the input language when a voice input in one of the various languages ​​is input. The language of the output text can be determined according to the training data set. For example, the language of the text output by the model (130) can be determined according to the voice input of the first language, the text of the first language, and the text of the second language included in the training data set. For example, if the number of voice inputs of the first language, the text of the first language, and the text of the second language is 10, when a voice input of one of the 10 languages ​​is input, the model (130) can output text in the remaining 9 languages.

[0091] In the above example, the number and types of the first language and the second language are exemplary and are not limited to the example described above. For example, the number and types of the first language and the second language may be determined differently depending on the training data of the model (130).

[0092] FIG. 2 is a schematic block diagram of an electronic device (100) according to various embodiments.

[0093] In the description of the electronic device (100) illustrated in FIG. 2, the same content as the description of the electronic device (100) illustrated in FIG. 1 may be omitted. Therefore, even if the description of the electronic device (100) in FIG. 2 is omitted, the content described for the electronic device (100) illustrated in FIG. 1 may be applied substantially identically to the electronic device (100) in FIG. 2.

[0094] Referring to FIG. 2, an electronic device (100) according to one embodiment can receive voice input of a first language from a terminal (e.g., a first terminal (200-1).

[0095] The electronic device (100) can determine text in a second language corresponding to voice input in a first language. The electronic device (100) can provide text in the second language to a terminal. The electronic device (100) can provide at least a portion of the text in the second language while voice input in the first language is being received.

[0096] In the embodiment illustrated in FIG. 1, the electronic device (100) receives voice input of a first language from a first terminal (200-1) and can provide text of a second language corresponding to the voice input of the first language to a second terminal (200-2). In the embodiment illustrated in FIG. 2, the electronic device (100) receives voice input of a first language from a first terminal (200-1) and can provide text of a second language corresponding to the voice input of the first language to the first terminal (200-1).

[0097] In the embodiment illustrated in FIG. 1, the operation of the electronic device (100) is shown when a user of the first language uses the first terminal (200-1) and a user of the second language uses the second terminal (200-2).

[0098] In the embodiment illustrated in FIG. 2, the operation of an electronic device (100) is shown when a user using a first language and a user using a second language use the first terminal (200-1) together.

[0099] FIG. 3 is a schematic block diagram of a model (130) according to various embodiments.

[0100] The description of the illustrated model (130) in FIG. 3 may be substantially applied to the description of the electronic device (100) or the processor (110). For example, the electronic device (100) or the processor (110) may perform the operation of the model (130) substantially the same.

[0101] Referring to FIG. 3, a model (130) according to one embodiment may include a speech encoder (131) and / or a text decoder (133).

[0102] For example, a speech encoder (131) can extract features (or feature vectors) of a first language voice input (135) using the input voice input (135) of the first language. A text encoder (165) can output text (137) of a second language using the features (or feature vectors) of the voice input (135).

[0103] For example, the text decoder (133) can output text (137) of a second language using at least one of the keyword (139) and the previously output text (137), or a combination thereof.

[0104] For example, the text decoder (133) can generate text (137) in a second language according to a set keyword (139). For example, the keyword (139) may represent a field, subject, topic, etc. related to a conversation, speech, or voice input (135).

[0105] For example, if the keyword (139) is mart, supermarket and voice input (133) "Where is the 'A' display?" is input, the text decoder (133) can output text (137) in a second language "Where is the display stand for the 'A' product?".

[0106] For example, if the keyword (139) is exhibition and the voice input (133) "Where is the 'A' display?" is input, the text decoder (133) can output text (137) in a second language "Where is the display stand of company 'A'?".

[0107] For example, the text decoder (133) can output text (137) in a second language based on previously output text (137). For example, after outputting text (137) "He was a great tennis player.", if voice input (133) in the first language "And, still love tennis." is input, the text decoder (133) can output text (137) "And, he still loves tennis." instead of text "And, I still love tennis."

[0108] For example, the model (130) can be trained to output text in a second language by inputting voice input in a first language. Additionally, the model (130) can be trained to output text in a first language and text in a second language by inputting voice input in a first language.

[0109] A training data set for training the model (130) may include voice input of a first language, text of a first language, and text of a second language. For example, the voice input of the first language may be data synthesized into speech from text of the first language.

[0110] For example, there may be multiple languages ​​included in the training data set. For example, there may be multiple types of languages ​​for the speech input of the first language, the text of the first language, and the text of the second language.

[0111] For example, training data may include combinations of voice input and text in various languages, such as (Korean voice input, Korean text, English text), (Korean voice input, Korean text, Japanese text), (English voice input, English text, Korean text), (English voice input, English text, Japanese text), (English voice input, English text, Chinese text), (English voice input, English text, French text), (English voice input, English text, Russian text), etc.

[0112] For example, when voice input of a first language is input, the model (130) can output text in multiple languages ​​(e.g., text of the first language, text of the second language, text of the third language, ..., text of the nth language). For example, when Korean voice input is input, the model (130) can output Korean text, Japanese text, Chinese text, English text, etc. Also, when text of another language (e.g., Japanese, Chinese, English, etc.) is input, the model (130) can output text in multiple languages ​​in substantially the same way as when Korean voice input is input.

[0113] For example, while receiving voice input of the first language, the model (130) can determine the parts of the text of the second language corresponding to the input voice input that have a determined meaning and / or parts of the text that have not a determined meaning. For example, the model (130) can generate text of the second language while receiving voice input of the first language. The model (130) can determine the parts of the text of the second language corresponding to the voice input of the first language received up to a specific point in time that have a determined meaning.

[0114] For example, among the text of the second language, the remaining part of the portion whose meaning has been determined may be a portion whose meaning has not been determined. Alternatively, among the text of the second language, the text of the second language corresponding to the voice input of the first language entered after the point in time when the meaning was determined may be a portion whose meaning has not been determined.

[0115] For example, while a voice input of the first language is being received, the model (130) may output text of the second language corresponding to a voice input of the first language whose meaning is not determined and / or text of the second language corresponding to a voice input of the first language whose meaning is determined.

[0116] For example, the model (130) can output text in a second language corresponding to the voice input of the first language input received up to a specific point in time, while the meaning of the voice input of the first language received up to a specific point in time has not been determined. For example, the text in the second language output by the model (130) may represent an output estimated from the voice input of the first language received up to a specific point in time.

[0117] For example, while voice input of the first language is being received, the model (130) can provide text of the second language corresponding to the voice input of the first language, whose meaning has been determined, to the second terminal. For example, as voice input of the first language is input, the meaning of the voice input of the first language input entered up to a specific point in time can be determined. The model (130) can output text of the second language corresponding to the voice input of the first language input entered up to a specific point in time.

[0118] In addition, voice input of the first language may continue to be input even after the point in time when the meaning of the voice input of the first language is determined. If voice input of the first language continues to be input after the point in time when the meaning is determined, as the voice input of the first language continues, the meaning of additional voice input of the first language may change from an undetermined state to a determined state.

[0119] The model (130) can output text in a second language corresponding to the additionally input voice input of the first language at a time when the meaning of the additionally input voice input of the first language is not determined. The text in the second language output at a time when the meaning of the additionally input voice input of the first language is not determined can correspond to the estimated meaning of the additionally input voice input of the first language. For example, the model (130) can estimate the meaning of the voice input of the first language whose meaning is not determined and determine text in the second language corresponding to the estimated meaning.

[0120] The model (130) can output text in a second language corresponding to the additionally input voice input of the first language at the time when the meaning of the additionally input voice input of the first language is determined.

[0121] As described above, in response to voice input of the first language input in real time, the model (130) can output text of the second language corresponding to the voice input of the first language in real time. For example, the model (130) can provide text of the second language corresponding to a part of the first language voice input in real time where the meaning is determined and text of the second language corresponding to a part where the meaning is not determined. For example, the text of the second language corresponding to a part where the meaning is not determined may represent a result estimated by the model (130) for the voice input of the first language where the meaning is not determined.

[0122] The electronic device (100) can reduce the delay time required to provide text by providing text of a second language with a determined meaning while voice input is being received.

[0123] For example, while voice input of the first language is being received, the model (130) can determine the parts of the voice input that have a determined meaning and / or the parts of the voice input that have not been determined. When the meaning of the voice input received up to a specific point in time is determined, the model (130) can determine text of the second language corresponding to the voice input with a determined meaning. The electronic device (100) can provide the text of the second language corresponding to the voice input with a determined meaning to the second terminal while voice input of the first language is being received.

[0124] For example, the model (130) can convert text in a second language corresponding to a voice input with a determined meaning into voice output. The electronic device (100) can provide text and / or voice output in a second language corresponding to a voice input with a determined meaning to a second terminal while a voice input in a first language is being received.

[0125] The model (130) can determine whether the meaning of at least some of the text in the second language is determined, or whether the meaning of at least some of the voice input in the first language is determined. For example, the model (130) can calculate the probability that the meaning of at least some of the text in the second language corresponding to the voice input entered up to a specific point in time is determined. For example, the model (130) can calculate the probability that the meaning of at least some of the voice input entered up to a specific point in time is determined. The model (130) can determine whether the meaning of at least some of the text in the second language is determined, or whether the meaning of at least some of the voice input in the first language is determined, by comparing the probability with a set threshold value.

[0126] When the meaning of at least part of the text of the second language is determined in response to voice input entered up to a specific point in time, the model (130) can determine the text of the second language corresponding to the voice input up to the specific point in time where the meaning is determined. The model (130) can continue to determine the text of the second language using voice input entered after the specific point in time.

[0127] FIG. 4 is a flowchart of the operation of a text provision method according to various embodiments.

[0128] The operations (310) to (330) illustrated in FIG. 4 can be performed substantially identically by the processor (110) of the electronic device (100).

[0129] The operations (310) to (330) illustrated in FIG. 4 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0130] Referring to FIG. 4, an electronic device (100) according to various embodiments can receive voice input of a first language from a first terminal (200-1) in operation (310).

[0131] For example, the electronic device (100) can determine text of a second language corresponding to voice input of a first language in operation (320). For example, the first terminal (200-1) and / or the second terminal (200-2) may include a voice input device (e.g., a microphone).

[0132] For example, an electronic device (100) can determine text of a second language corresponding to a voice input of a first language using a model (130). The model (130) can be trained to output at least one of text of the first language, text of the second language, or text of the nth language, or a combination thereof, when a voice input of the first language is input.

[0133] For example, the electronic device (100) may provide text of a second language to a second terminal (200-2) in operation (330). The electronic device (100) may provide at least a portion of the text of the second language to the second terminal (200-2) while voice input of the first language is being received.

[0134] Accordingly, in FIG. 3, the operation (310), operation (320) and operation (330) are shown as being performed sequentially, but at least some of the operations (310), operation (320) and operation (330) can be performed simultaneously by the electronic device (100).

[0135] For example, operations (310), (320), and (330) may be performed simultaneously by the electronic device (100). Additionally, after operation (310) is completed, operations (320) and (330) may be performed simultaneously by the electronic device (100).

[0136] FIG. 5 is a flowchart of the operation of a text provision method according to various embodiments.

[0137] The operations (410) to (480) illustrated in FIG. 5 can be performed substantially identically by the processor (110) of the electronic device (100).

[0138] The operations (410) to (480) illustrated in FIG. 2 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0139] Referring to FIG. 5, an electronic device (100) according to various embodiments can, in operation (410), allow a first terminal (200-1) and a second terminal (200-2) to enter a set chat room. The electronic device (100) can provide text in a first language and / or text in a second language to at least one of the first terminal (200-1) and the second terminal (200-2) through the chat room.

[0140] For example, the first terminal (200-1) and / or the second terminal (200-2) can enter the chat room via a link, a URL (Uniform Resource Locator), a QR code, etc. Either of the first terminal (200-1) and / or the second terminal (200-2) can send a message to the other terminal containing an access method (e.g., QR code, link, URL, etc.) for entering the chat room. As another example, the first terminal (200-1) and / or the second terminal (200-2) may also enter the chat room through a separate application.

[0141] For example, keywords may be set for a chat room. For example, the administrator of the chat room may be a user of the first terminal (200-1) or a user of the second terminal (200-2). The administrator of the chat room may set keywords for the chat room.

[0142] For each terminal, the language used by each user can be set. For example, each user can set the language used for each terminal.

[0143] For example, the electronic device (100) can determine the language of use of the first terminal (200-1) and / or the second terminal (200-2) based on voice input received through the first terminal (200-1) and / or the second terminal (200-2). The electronic device (100) can determine what language the voice input received through the first terminal (200-1) and / or the second terminal (200-2) is (using the model (130).

[0144] For example, for each of the first terminal (200-1) and / or the second terminal (200-2), one or more languages ​​may be set. For example, for the first terminal (200-1), multiple languages ​​may be set, such as a first language, a second language, and a third language.

[0145] Hereinafter, the operation of the electronic device (100) is described when the language of use for the first terminal (200-1) is set to the first language and the language of use for the second terminal (200-2) is set to the second language.

[0146] For example, the electronic device (100) can determine the terminal receiving the voice input in operation (420). When the first terminal (200-1) and / or the second terminal (200-2) receives the voice input, they can transmit the voice input to the electronic device (100).

[0147] For example, in operation (420), if it is determined that voice input is received from the first terminal (200-1), the electronic device (100) can receive the first voice input of the first language through the first terminal (200-1) in operation (430).

[0148] For example, the electronic device (100) can determine text in a second language corresponding to a first voice input in operation (440). For example, the electronic device (100) can determine text in a second language corresponding to a first voice input by using a model (130).

[0149] For example, the electronic device (100) may provide text in a second language through a chat room in operation (450). For example, the electronic device (100) may provide text in a second language to the first terminal (200-1) and / or the second terminal (200-2) through a chat room.

[0150] For example, in operation (420), if it is determined that voice input is received from the second terminal (200-2), the electronic device (100) can receive a second voice input of the second language through the second terminal (200-2) in operation (460).

[0151] For example, the electronic device (100) can determine text of the first language corresponding to the second voice input in operation (470). For example, the electronic device (100) can determine text of the first language corresponding to the second voice input by using a model (130).

[0152] For example, the electronic device (100) may provide text in the first language through a chat room in operation (480). For example, the electronic device (100) may provide text in the first language to the first terminal (200-1) and / or the second terminal (200-2) through a chat room.

[0153] For example, the electronic device (100) can provide at least a portion of the text of the first language to the first terminal (200-1) while the second voice input is being received.

[0154] For example, the electronic device may provide at least a portion of the text of the second language to the second terminal (200-2) while the first voice input is being received.

[0155] The operations (410) to (480) illustrated in FIG. 5 are shown as being performed sequentially, but are not limited thereto. For example, some of the operations (410) to (480) may be performed simultaneously. For example, the first terminal (200-1) and the second terminal (200-2) may receive voice input simultaneously, and the electronic device (100) may perform operations (430) to (480) simultaneously. For example, the electronic device (100) may perform at least some of the operations (430) to (450) simultaneously. Alternatively, the electronic device (100) may perform at least some of the operations (460) to (480) simultaneously.

[0156] FIGS. 6 and FIGS. 7 are drawings showing a chat window provided by an electronic device (100) according to various embodiments.

[0157] Referring to FIGS. 6 and FIGS. 7, an electronic device (100) according to one embodiment can provide text in a second language through a chat window of at least one of a first terminal (200-1) and a second terminal (200-2).

[0158] For example, the first terminal (200-1) and / or the second terminal (200-2) can access the chat window through an application or a web browser.

[0159] In the following description, the electronic device (100) is described as providing text of a first language, text of a second language and / or text of a third language, etc., through a chat window, but is not limited thereto.

[0160] For example, the electronic device (100) can provide text of a first language, text of a second language and / or text of a third language, etc. to a first terminal (200-1) and / or a second terminal (200-2) through a chat window and other interfaces.

[0161] FIG. 6 is a drawing showing a screen (510) in which a chat window is displayed on the display of a terminal according to one embodiment.

[0162] As shown in FIG. 6, the electronic device (100) can provide chat (511, 513, 515) to the terminal through a chat window.

[0163] Chat (511) and chat (513) may represent chat provided by the electronic device (100) when voice input of a first language (e.g., Korean) is input. The electronic device (100) may determine text of the first language and text of the second language (e.g., English) corresponding to the voice input of the first language. The electronic device (100) may provide a chat, such as chat (511) and chat (513), to the first terminal (200-1) and / or the second terminal (200-2), in which text of the first language is displayed above and text of the second language is displayed below.

[0164] The chat (515) may represent a chat provided by the electronic device (100) when voice input of a second language (e.g., English) is input. The electronic device (100) may determine text of a first language (e.g., Korean) and text of a second language corresponding to voice input of the second language. The electronic device (100) may provide a chat, such as the chat (515), to a first terminal (200-1) and / or a second terminal (200-2), in which text of the second language is displayed above and text of the first language is displayed below.

[0165] The electronic device (100) can provide a list of chat rooms (517) through a chat room interface. The user can select a chat room they wish to display from the list of chat rooms (517). The electronic device (100) and / or the terminal can make the selected chat room appear on the terminal's display.

[0166] FIG. 7 is a drawing showing a screen (520) in which a chat window is displayed on the display of a terminal according to one embodiment. As shown in FIG. 7, the electronic device (100) can provide a chat room including chats (521, 523, 525) to a terminal (e.g., a first terminal (200-1), a second terminal (200-2)).

[0167] For chat (521), chat (523) and chat (525), the descriptions for chat (511), chat (513), and chat (515) can be applied substantially identically, respectively.

[0168] Chat (521) and chat (523) may represent chat provided by the electronic device (100) when voice input of a first language (e.g., Korean) is input. The electronic device (100) may determine text of the first language and text of the second language (e.g., English) corresponding to the voice input of the first language. The electronic device (100) may provide a chat, such as chat (521) and chat (523), to the first terminal (200-1) and / or the second terminal (200-2), in which text of the first language is displayed above and text of the second language is displayed below.

[0169] The chat (525) may represent a chat provided by the electronic device (100) when voice input of a second language (e.g., English) is input. The electronic device (100) may determine text of a first language (e.g., Korean) and text of a second language corresponding to voice input of the second language. The electronic device (100) may provide a chat, such as the chat (525), to a first terminal (200-1) and / or a second terminal (200-2), in which text of the second language is displayed above and text of the first language is displayed below.

[0170] FIG. 8 is a diagram showing the time when an electronic device (100) according to various embodiments outputs text in response to voice input.

[0171] In the example of FIG. 8, the arrows indicated on the row corresponding to the voice input may represent sentences included in the voice input. For example, in FIG. 8, the voice input may include four sentences (e.g., sentences 1 through 4). The voice input may be input over time. For example, each sentence may represent a unit in which the meaning is determined within the voice input.

[0172] In the example of FIG. 8, the arrows indicated on the lines corresponding to the text may represent sentences included in the texts of the second language. For example, in FIG. 8, the text of the second language may include four sentences (e.g., sentences 1 through 4).

[0173] In the following description regarding FIG. 8, the operation of the electronic device (100) is described when the voice input is a first language and the text is a second language. In the following description regarding FIG. 8, it is assumed that the electronic device (100) receives voice input from a first terminal (200-1) and provides text to a second terminal (200-2).

[0174] For example, the electronic device (100) can receive voice input. At time T1, the input of the first sentence may be completed. The electronic device (100) can perform operations from the time the voice input is received until time T1. For example, the electronic device (100) can use the model (130) to determine text of the second language corresponding to the voice input from the time the voice input is received until time T1.

[0175] At time T1, the electronic device (100) may provide at least a portion of the text to the second terminal (200-1). At time T1, the electronic device (100) may determine whether there is a part of the text corresponding to the input voice input that has a determined meaning. At time T1, the electronic device (100) may provide a sentence with a determined meaning (e.g., a first sentence) to the second terminal (200-2).

[0176] After voice input is received, similar to the operation of the electronic device (100) up to time point T1, the electronic device (100) can determine text corresponding to the voice input received from time point T1 to time point T2. The electronic device (100) can provide a second sentence of text with a determined meaning at time point T2 to the terminal (200-2).

[0177] In the intervals from time point T2 to time point T3 and from time point T3 to time point T4, the electronic device (100) can provide the third and fourth sentences of text to the terminal (200-2) substantially in the same manner as described above. For example, the electronic device (100) can provide the third sentence of text, whose meaning is determined at time point T3, to the terminal (200-2). For example, the electronic device (100) can provide the fourth sentence of text, whose meaning is determined at time point T4, to the terminal (200-2).

[0178] In the example of FIG. 8, the first to fourth sentences of text may correspond to the first to fourth sentences included in the voice input. For example, the first sentence of text may represent a translation of the first sentence in the first language included in the voice input into the second language.

[0179] FIG. 9 is a diagram illustrating voice input (610) of a first language and text (620) of a second language according to various embodiments. FIG. 9 illustrates an example where the first language is Korean and the second language is English. FIG. 9

[0180] For example, in FIG. 9, voice input (611), voice input (613), and voice input (615) can be input sequentially to the electronic device (100).

[0181] The electronic device (100) can determine text (621) corresponding to voice input (611). The electronic device (100) can provide the text (621) to a second terminal (200-1).

[0182] The electronic device (100) can determine text (623) corresponding to voice input (613). The electronic device (100) can provide the text (623) to a second terminal (200-1).

[0183] The electronic device (100) can determine text (621) corresponding to voice input (615). The electronic device (100) can provide the text (625) to a second terminal (200-1).

[0184] FIGS. 10, FIGS. 11, FIGS. 12, FIGS. 13, FIGS. 14, FIGS. 15 and FIGS. 16 are drawings illustrating text of a first language and text of a second language provided by an electronic device according to various embodiments.

[0185] In the following description regarding FIGS. 10 to 16, when the electronic device (100) provides text of a first language, text of a second language and / or text of a third language to a second terminal (200-2) or a third terminal, a screen displayed on the display of the second terminal (200-2) or the third terminal is shown, but is not limited thereto. For example, the electronic device (100) may provide text of a first language, text of a second language and / or text of a third language to at least one terminal.

[0186] FIGS. 10 to 13 show text in a first language and text in a second language (620) provided by the electronic device (100) through a chat room when the voice input (610) of FIG. 9 is input.

[0187] Referring to FIGS. 10 and 11, an electronic device (100) according to one embodiment may provide a portion of text (620) of a second language to a second terminal (200-2) while voice input (610) of a first language is received.

[0188] FIG. 10 is an example showing text (711) of a first language and text (713) of a second language provided by an electronic device (100) through a chat room at the time when voice input (611) is input. The second terminal (200-1) can display a screen (710) showing text (711) of the first language and text (713) of the second language through the display of the second terminal (200-1).

[0189] In FIG. 10, the text of the second language (713) shown may represent a part of the text of the second language (620) where the meaning is determined (e.g., the text of FIG. 9 (621)).

[0190] As shown in FIG. 10, the electronic device (100) can determine whether there is a part with a determined meaning among the voice input (610) of the first language. Alternatively, the electronic device (100) can determine whether there is a part with a determined meaning among the text (620) of the second language.

[0191] For example, the electronic device (100) can determine a part of the voice input (610) of the first language in which the meaning is determined (e.g., voice input (611) of FIG. 9). The electronic device (100) can determine a text (621) of the second language corresponding to the voice input (611) in which the meaning is determined. The electronic device (100) can provide the text (621) of the second language to the second terminal (200-2).

[0192] For example, the electronic device (100) can determine a part of the text (620) of the second language in which the meaning is determined (e.g., the text (621) of FIG. 9). The electronic device (100) can provide the text (621) in which the meaning is determined to the second terminal (200-2).

[0193] FIG. 11 is a diagram showing text in a second language provided by an electronic device (100) when a portion of the voice input (611) of FIG. 9 is input. The second terminal (200-2) can display a chat room containing text in a first language (721) and text in a second language (723) through a display, such as a screen (720).

[0194] As shown in FIG. 11, when some of the voice inputs (611) are input into the electronic device (100), the electronic device (100) can provide text (721) of the first language and text (723) of the second language corresponding to the voice input of the first language (e.g., "Hello. No. 2345") input up to that point to the second terminal (200-2).

[0195] As shown in FIG. 11, at the point when part of a sentence (or voice input of a unit whose meaning can be determined) is input, the electronic device (100) can provide text (721) of a first language and / or text (723) of a second language corresponding to the voice input input up to that point to a second terminal (200-2).

[0196] For example, as the voice input (611) of FIG. 9 is input, the electronic device (100) can provide text of the first language and text of the second language to the second terminal (200-2) as in FIG. 10 and FIG. 11. As the voice input (611) is input, the second terminal (200-2) can display a screen (720) such as FIG. 11 on the display, and then display a screen (710) such as FIG. 10 on the display.

[0197] Referring to FIGS. 10 and FIGS. 11 above, an electronic device (100) according to one embodiment may provide text of an estimated second language for an incomplete voice input and may provide a voice utterance of a second language corresponding to a voice input with a determined meaning.

[0198] For example, the text of the first language (721) shown in FIG. 11 may represent an incomplete voice input, and the text of the second language (723) may represent an estimated text of the second language for the incomplete voice input.

[0199] For example, the text of the first language (711) shown in FIG. 10 may represent a voice input with a determined meaning, and the text of the second language (713) may represent text of the second language for the voice input with a determined meaning. The electronic device (100) may convert the text of the second language (713) for the voice input with a determined meaning into a voice output to provide voice utterance of the second language.

[0200] FIG. 12 is an example showing text (731) of a first language and text (733) of a second language provided by an electronic device (100) through a chat room at the time when voice input (611) and voice input (613) are input. The second terminal (200-1) can display a screen (730) showing text (731) of the first language and text (733) of the second language through the display of the second terminal (200-1).

[0201] In FIG. 12, the text of the second language (733) shown may represent a part of the text of the second language (620) whose meaning has been determined (e.g., text (621) and text (623) of FIG. 9).

[0202] In the example of FIG. 12, the electronic device (100) can determine whether there is a part with a determined meaning in the voice input (610) of the first language. Or, the electronic device (100) can determine whether there is a part with a determined meaning in the text (620) of the second language.

[0203] For example, the electronic device (100) can determine a part of the voice input (610) of the first language in which the meaning is determined (e.g., voice input (611) and voice input (613) of FIG. 9). The electronic device (100) can determine text (621, 623) of the second language corresponding to the voice input (611, 613) in which the meaning is determined. The electronic device (100) can provide the text (621, 623) of the second language to the second terminal (200-2).

[0204] For example, the electronic device (100) can determine a part of the text (620) of the second language in which the meaning is determined (e.g., text (621), text (623) of FIG. 9). The electronic device (100) can provide the text in which the meaning is determined (621, 623) to the second terminal (200-2).

[0205] FIG. 13 is an example showing text (741) of a first language and text (743) of a second language provided by an electronic device (100) through a chat room at the time when voice input (611), voice input (613) and voice input (615) are input. The second terminal (200-1) can display a screen (740) showing text (741) of the first language and text (743) of the second language through the display of the second terminal (200-1).

[0206] As shown in FIG. 13, the electronic device (100) can provide text (741) of a first language and text (743) of a second language corresponding to the entire voice input (610) to a second terminal (200-2).

[0207] In the above example, the example illustrated in FIGS. 10, FIGS. 12, and FIGS. 13 can be understood as representing the sequential operation of the electronic device (100) or the screen displayed on the terminal (200-2) in response to the voice input (610) of FIG. 9. For example, as the voice input (610) is input, the electronic device (100) may provide the text illustrated in FIG. 10, the text illustrated in FIG. 12, and the text illustrated in FIG. 13 to the second terminal (200-2).

[0208] As illustrated in the examples in FIGS. 10, 12 and 13, the electronic device (100) may provide text of a first language (711, 731, 741) and / or text of a second language (713, 733, 743) to a second terminal (200-2) while receiving voice input (610).

[0209] In the above example, the example illustrated in FIGS. 10 and FIGS. 11 can be understood as representing the sequential operation of an electronic device (100) or a screen displayed on a terminal (200-2) as a voice input of one sentence is input. For example, as a voice input of one sentence (611) is input, the electronic device (100) may provide the text illustrated in FIG. 11 and the text illustrated in FIG. 10 to the second terminal (200-2).

[0210] The text of the second language shown in FIGS. 10, 12, and 13 above is illustrated as including text of the second language corresponding to a voice input of the first language with a determined meaning, but is not limited thereto.

[0211] For example, as shown in FIG. 11, the electronic device (100) can provide text (723) in a second language corresponding to a voice input of a first language whose meaning is not determined. The electronic device (100) illustrated in FIG. 11 can estimate text (723) in a second language using the input voice input of the first language.

[0212] For example, the electronic device (100) may provide text in a second language corresponding to a voice input of a first language with a determined meaning and / or text in a second language corresponding to a voice input of a first language with an undetermined meaning to a terminal (e.g., first terminal (200-1), second terminal (200-2)).

[0213] For example, as shown in FIG. 12, after voice input (611) and voice input (613) have been input, and while voice input (615) is being input, the electronic device (100) may provide text in a second language (621, 622) corresponding to voice input (611, 613) of a first language with a determined meaning, and text in a second language corresponding to voice input (615) of a first language with an undetermined meaning (e.g., the part "breakfast in the room" in voice input (615)) (e.g., the part "breakfast in the room" in text (625) of a second language) to a terminal (e.g., a first terminal (200-1), a second terminal (200-2)).

[0214] FIG. 14 is a diagram showing an example in which an electronic device (100) according to one embodiment provides text (753) of a third language (e.g., Japanese) to a third terminal.

[0215] For example, FIG. 14 may be an example showing a screen (750) displayed on the display of a third terminal. As in FIG. 14, when the voice input (610) of FIG. 9 is input, the electronic device (100) may provide text (751) of the first language and / or text (753) of the third language to the third terminal.

[0216] The third terminal may display text (751) of the first language and / or text (753) of the third language on the display of the third terminal. For example, in the examples of FIGS. 10 to 13, the language set for the second terminal (200-2) may be English, and in the example of FIG. 14, the language set for the third terminal may be Japanese.

[0217] According to one embodiment, the electronic device (100) can provide at least a portion of the text of a third language to a third terminal while voice input is being received, substantially the same as the example shown in FIGS. 10 to 13.

[0218] FIG. 15 is a drawing showing a screen (760) displayed on the display of the second terminal (200-2) when the first language is Japanese and the second language is Korean.

[0219] FIG. 16 is a drawing showing a screen (770) displayed on the display of the second terminal (200-2) when the first language is Japanese and the second language is English.

[0220] As illustrated in the examples in FIGS. 10 to 16, the first language and the second language can each be set in various ways. For example, the types of the language of the voice input and the language of the output text input into the model (130) can each be multiple.

[0221] For example, if the types of languages ​​supported by the model (130) are Korean, Japanese, English, French, and Chinese, when voice input is input, the model (130) can determine at least one of the text in Korean, Japanese, English, French, and Chinese. The electronic device (100) can provide at least one of the text in Korean, Japanese, English, French, and Chinese to the terminal according to the language of use set for the terminal.

[0222] For example, the electronic device (100) (or model (130)) can determine the language of the text to be output according to the language set for the terminal connected to the chat room. For example, when three terminals, each with the language set to Korean, Japanese, and English, are connected to the chat room, the electronic device (100) (or model (130)) can determine Korean text, Japanese text, and English text in response to the input voice input. The electronic device (100) can provide at least one of the determined Korean text, Japanese text, and English text, or a combination thereof, to each terminal according to the language set for each terminal and the language of the input voice input.

[0223] According to one embodiment, in the example of FIGS. 14 and FIGS. 15, the language of use (first language) set for the first terminal may be Korean, and the language of use (second language) set for the second terminal may be Japanese.

[0224] According to one embodiment, FIG. 14 may show an example in which an electronic device (100) provides text in a second language corresponding to voice input of a first language to a second terminal (200-2). According to one embodiment, FIG. 15 may show an example in which an electronic device (100) provides text in a first language corresponding to voice input of a second language to a first terminal (200-1).

[0225] FIGS. 17 and 18 are drawings illustrating voice input of a first language and text of a second language according to various embodiments.

[0226] In FIG. 17, voice input (810), voice input (820), voice input (830), and voice input (840) can be input to an electronic device (100). As shown in FIG. 17, the voice inputs (810, 820, 830, 840) are identical up to time T1, and different from time T1 to time T2.

[0227] FIG. 18 shows text (900, 910, 920, 930, 940) of a second language provided by an electronic device (100) to at least one of a plurality of terminals (200-1, 200-2, ..., 200-n).

[0228] Text (900) represents text provided by the electronic device (100) at time T1 when voice input (810, 820, 830, 840) is input. Text (900) is an example of text provided by the electronic device at time T1, and the electronic device may provide text different from the text shown in FIG. 18 (e.g., "I home", etc.).

[0229] Text (910) represents the text provided by the electronic device (100) at time T2 when voice input (810) is input.

[0230] Text (920) represents the text provided by the electronic device (100) at time T2 when voice input (820) is input.

[0231] Text (930) represents the text provided by the electronic device (100) at time T2 when voice input (830) is input.

[0232] Text (940) represents the text provided by the electronic device (100) at time T2 when voice input (840) is input.

[0233] When voice input (810, 820, 830, 840) is input, the text provided by the electronic device (100) at time T1 is the same, but depending on the voice input from time T1 to time T2, the electronic device (100) may maintain, modify, or change a part of the text (900) provided at time T1 to provide text (910, 920, 930, 940).

[0234] The electronic device (100) can retain ('am') or modify and / or change ('go', 'want to go', 'can go') parts of the text (900) according to the content ("go", "am", "want to go", "can go") of voice input (810, 820, 830, 840) input from time T1 to time T2.

[0235] As shown in FIGS. 17 and 18, the electronic device (100) may provide at least a portion of text in a second language while voice input is being input. The electronic device (100) may determine at least a portion of text in a second language while voice input is being input.

[0236] FIGS. 19, 20 and 21 are drawings showing text of a first language and text of a second language provided by an electronic device (100) according to various embodiments.

[0237] FIG. 19 is a diagram showing text (1010) of a first language and text (1020) of a second language provided by an electronic device (100) at time T1 when voice input (810, 820, 830, 840) of FIG. 17 is input. FIG. 19 shows a screen (1000) in which text (1010) of a first language and text (1020) of a second language are displayed on the display of a terminal.

[0238] FIG. 20 is a diagram showing the text of a first language (1110) and the text of a second language (1120) provided by the electronic device (100) at time T2 when the voice input (810) of FIG. 17 is input. FIG. 20 shows a screen (1100) in which the text of the first language (1110) and the text of the second language (1120) are displayed on the display of the terminal.

[0239] FIG. 21 is a diagram showing the text of a first language (1210) and the text of a second language (1220) provided by the electronic device (100) at time T2 when the voice input (840) of FIG. 17 is input. FIG. 21 shows a screen (1000) in which the text of the first language (1210) and the text of the second language (1220) are displayed on the display of the terminal.

[0240] As shown in FIGS. 19 to 21, the electronic device (100) can determine text of a first language and / or text of a second language while voice input is being input. The electronic device (100) can maintain, change, and / or modify the text of the first language and / or text of the second language determined at a specific point in time according to voice input input after a specific point in time.

[0241] FIG. 22 is a drawing showing an interface provided by an electronic device (100) according to various embodiments.

[0242] The interface illustrated in FIG. 22 represents an interface for registering a user's voice corresponding to each terminal. The electronic device (100) can enable each terminal to display an interface such as a screen (1300) on a display.

[0243] As shown in FIG. 22, the interface may include instructions (1310) and example sentences (1320). A user may speak according to the example sentences (1320). A terminal may receive the user's speech. The electronic device (100) and / or the terminal may process the speech of the example sentences (1320) to determine and / or extract the user's voice features. The electronic device (100) and / or the terminal may identify the user's speech using the user's voice features.

[0244] For example, when a terminal receives multiple utterances (or voice inputs) from multiple users, the electronic device (100) and / or the terminal can identify the user's utterance (or voice input) among the multiple utterances (or voice inputs) based on the user's voice characteristics.

[0245] Regarding the method by which the electronic device (100) and / or terminal determines and / or extracts voice features from a user's utterance, various known voice feature determination / extraction methods or algorithms may be applied. Additionally, regarding the method by which the electronic device (100) and / or terminal identifies a user's utterance based on voice features, various known identification methods or algorithms may be applied.

[0246] FIGS. 23 and 24 are drawings illustrating voice input (1410, 1510) of a first language and text (1420, 1520) of a second language according to various embodiments.

[0247] Referring to FIG. 23 and FIG. 24, an electronic device (100) according to various embodiments can determine text of a second language based on at least one of text corresponding to a previously entered voice input of a first language, text of a second language previously provided, and a set keyword.

[0248] In FIG. 23, text (1421) corresponds to voice input (1411), and text (1423) may correspond to voice input (1413).

[0249] The electronic device (100) can determine the text (1423) of the second language based on the text ("He was a great tennis player") corresponding to the previously entered voice input (1411) of the first language and / or the previously provided text (1421) of the second language ("He was a great tennis player.").

[0250] For example, the electronic device (100) can determine the text of the second language corresponding to the voice input (1413) as text (1423) ("And he still loves tennis.") instead of "And I still love tennis."

[0251] Referring to FIG. 24, the electronic device (100) can determine text of a second language by considering set keywords.

[0252] For example, when the set keyword (in relation to the chat room) is market or mart, when voice input (1511) is entered, the electronic device (100) can determine text (1521) ("Where is the display stand?") as text corresponding to the voice input (1511).

[0253] For example, when a keyword set (regarding a chat room) is exhibition, when voice input (1511) is entered, the electronic device (100) can determine text (1523) ("Where is the exhibition hall?") ​​as text corresponding to voice input (1511).

[0254] As shown in the example illustrated in FIG. 24, the electronic device (100) can determine different text based on keywords even when the same voice input is input.

[0255] For example, keywords can be set by the chat room administrator (e.g., the user of the device). Keywords may include one or more words.

[0256] FIGS. 25 and 26 are drawings showing a conversation record and a conversation summary provided by an electronic device (100) according to various embodiments.

[0257] FIG. 25 is an example showing a conversation record provided by an electronic device (100). The terminal can display the conversation record received from the electronic device (100) through a display, such as a screen (1600).

[0258] For example, the conversation record may include text in a first language and / or text in a second language corresponding to voice input in a first language provided by the electronic device (100) to the terminal. Additionally, the conversation record may include text in a first language and / or text in a second language corresponding to voice input in a second language provided by the electronic device (100) to the terminal.

[0259] FIG. 26 is an example showing a conversation summary provided by an electronic device (100). A terminal can display the conversation summary received from the electronic device (100) through a display, such as a screen (1700).

[0260] A conversation summary may represent a summary of the conversation content illustrated in FIG. 25. Various known methods and algorithms may be applied to the method of generating a conversation summary by summarizing the conversation record.

[0261] FIGS. 27, 28, and 29 are drawings showing the operation flowcharts of an electronic device (100) identifying a user of a terminal according to various embodiments.

[0262] The operations (1810) to (1820) illustrated in FIG. 27 can be performed substantially identically by the processor (110) of the electronic device (100).

[0263] The operations (1810) to (1820) illustrated in FIG. 27 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0264] Referring to FIG. 27, when a plurality of voice inputs are received in operation (1810), the electronic device (100) can provide text corresponding to the plurality of voice inputs to the first terminal (200-1). For example, the electronic device (100) can provide a plurality of texts of a first language corresponding to the plurality of voice inputs to the first terminal (200-1).

[0265] The first terminal (200-1) can display an interface on the display for selecting text corresponding to the content spoken by the user among a plurality of texts.

[0266] For example, the electronic device (100) can determine the speaker of the voice input of the first language based on the input received from the first terminal (200-1) in operation (1820).

[0267] When the first terminal (200-1) receives input from a user selecting text corresponding to the content spoken by the user, the electronic device (100) and / or the first terminal (200-1) can determine and / or extract voice features using the voice input of the selected text.

[0268] The electronic device (100) and / or the first terminal (200-1) can determine the speaker of the voice input of the first language according to voice characteristics. For example, the first terminal (200-1) can transmit only the voice input corresponding to the voice characteristics of the speaker among a plurality of voice inputs to the electronic device (100). For example, the electronic device (100) can determine only the voice input corresponding to the voice characteristics of the speaker among a plurality of voice inputs as the voice input of the first language.

[0269] For example, the electronic device (100) can perform the operation (310) of FIG. 3 or the operation (410) of FIG. 4 after the operation (1820).

[0270] The operations (1910) to (1920) illustrated in FIG. 28 can be performed substantially identically by the processor (110) of the electronic device (100).

[0271] The operations (1910) to (1920) illustrated in FIG. 28 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0272] Referring to FIG. 28, an electronic device (100) according to one embodiment may provide an interface for determining the speaker of a voice input of a first language through a first terminal (200-1) in operation (1910).

[0273] For example, the electronic device (100) can provide an interface such as the screen (1300) shown in FIG. 22 through the first terminal (200-1).

[0274] For example, the electronic device (100) can determine the speaker in operation (1920) by using voice input corresponding to the interface.

[0275] For example, when the first terminal (200-1) receives a voice input corresponding to an interface, the electronic device (100) and / or the first terminal (200-1) can determine and / or extract voice features using the voice input.

[0276] The electronic device (100) and / or the first terminal (200-1) can determine the speaker using voice features. The electronic device (100) and / or the first terminal (200-1) can determine a voice input corresponding to (or matching) the voice features as a voice input of the first language.

[0277] Regarding the operation for determining the speaker of the operation (1920), the description of the operation for determining the speaker in the operation (1820) of FIG. 27 can be applied substantially the same way.

[0278] For example, the electronic device (100) can perform the operation (310) of FIG. 3 or the operation (410) of FIG. 4 after the operation (1920).

[0279] The operations (2010) to (2040) illustrated in FIG. 29 can be performed substantially identically by the processor (110) of the electronic device (100).

[0280] The operations (2010) to (2040) illustrated in FIG. 29 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0281] Referring to FIG. 29, an electronic device (100) according to one embodiment can determine in operation (2010) whether a speaker corresponding to the second terminal (200-2) is registered. For example, if voice features for identifying voice input are registered (or stored) in the electronic device (100) and / or the second terminal (200-2), the electronic device (100) can determine that a speaker corresponding to the second terminal (200-2) is registered.

[0282] For example, if a speaker corresponding to the second terminal (200-2) is registered in operation (2010), the electronic device (100) (or the first terminal (200-1)) can determine, in operation (2020), voice inputs excluding the voice input of the speaker corresponding to the second terminal (200-2) from among the voice inputs input from the first terminal (200-1).

[0283] For example, if there is a voice input to the first terminal (200-1) and a voice input of a speaker registered in the second terminal (200-2), and another voice input, the electronic device (100) (or the first terminal (200-1)) can identify the other voice input.

[0284] For example, the electronic device (100) (or the first terminal (200-1)) can determine the speaker of the first language voice input in operation (2030) by using voice inputs other than the voice input of the speaker corresponding to the second terminal (200-2).

[0285] For example, the electronic device (100) (or the first terminal (200-1)) can determine / extract voice features from another voice input. The electronic device (100) (or the first terminal (200-1)) can determine the speaker of the voice input of the first language using the voice features. The electronic device (100) (or the first terminal (200-1)) can determine the voice input of the first language using the voice features.

[0286] For example, if a speaker corresponding to the second terminal (200-2) is not registered in operation (2010), the electronic device (100) (or the first terminal (200-1)) can determine whether multiple voice inputs are received in operation (2040). The electronic device (100) (or the first terminal (200-1)) can determine whether there are multiple voice inputs received from the first terminal (200-1).

[0287] For example, if it is determined that there are multiple voice inputs received from the first terminal (200-1) in operation (2040), the electronic device (100) can perform operation (1810) of FIG. 27 after operation (2040).

[0288] For example, if it is determined that there is only one voice input received from the first terminal (200-1) in operation (2040), the electronic device (100) can perform operation (1910) of FIG. 28 after operation (2040).

[0289] In FIGS. 27 and 28 above, the operation of determining a speaker of the first terminal (200-1), the operation of determining voice characteristics and determining voice input of the first language according to the determined voice characteristics, and the operation of determining voice input of the first language among a plurality of voice inputs input to the first terminal (200-1) are described, but are not limited thereto.

[0290] For example, regarding the operation of an electronic device (100) (or terminal) that determines the speaker of a terminal different from the first terminal (200-1), such as the second terminal (200-1), the third terminal, ..., the nth terminal (200-n), the contents described in FIGS. 27 to 29 may be applied substantially in the same way.

[0291] Meanwhile, the method according to the present invention is written as a program executable on a computer and can be implemented on various recording media such as magnetic storage media, optical reading media, and digital storage media.

[0292] Implementations of the various technologies described herein may be implemented as digital electronic circuits, or as computer hardware, firmware, software, or combinations thereof. Implementations may be implemented as computer program products, i.e., computer programs tangibly embodied in information carriers, such as machine-readable storage devices (computer-readable media) or radio signals, for processing by the operation of data processing devices, e.g., programmable processors, computers, or multiple computers, or for controlling such operation. Computer programs such as the computer program(s) described above may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. Computer programs may be deployed to be processed on one computer or multiple computers at one site, or distributed across multiple sites and interconnected by a communication network.

[0293] Processors suitable for processing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Generally, the processor will receive instructions and data from read-only memory or random access memory, or both. The elements of the computer may include at least one processor that executes instructions and one or more memory devices that store instructions and data. Generally, the computer may include one or more mass storage devices that store data, for example, magnetic, magneto-optical disks, or optical disks, or may be combined to receive data from these, transmit data to these, or both. Information carriers suitable for embodying computer program instructions and data include, for example, semiconductor memory devices, magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROMs (Compact Disk Read Only Memory) and DVDs (Digital Video Disks); magneto-optical media such as floptical disks; ROMs (Read Only Memory); RAMs (Random Access Memory); flash memory; EPROMs (Erasable Programmable ROM); EEPROMs (Electrically Erasable Programmable ROM); etc. Processors and memory may be supplemented by or included in special-purpose logic circuit organizations.

[0294] Additionally, a computer-readable medium may be any available medium accessible by a computer and may include both computer storage media and transmission media.

[0295] Although this specification contains details of a number of specific embodiments, they should not be understood as limiting the scope of any invention or claimables, but rather as descriptions of features that may be characteristic of a specific embodiment of a specific invention. Specific features described in this specification in the context of individual embodiments may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any appropriate sub-combination. Furthermore, while features may operate in a specific combination and be described as initially claimed, one or more features from the claimed combination may be excluded from the combination in some cases, and the claimed combination may be changed to a sub-combination or a variation of the sub-combination.

[0296] Likewise, although operations are depicted in the drawings in a specific order, this should not be understood as requiring that such operations be performed in that specific or sequential order depicted to obtain a desirable result, or that all depicted operations must be performed. In certain cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various device components of the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and devices can generally be integrated together into a single software product or packaged into multiple software products.

[0297] Meanwhile, the embodiments of the present invention disclosed in this specification and drawings are merely specific examples provided to aid understanding and are not intended to limit the scope of the present invention. It is obvious to those skilled in the art that other variations based on the technical concept of the present invention are possible in addition to the embodiments disclosed herein.

[0298] 100: Electronic device

[0299] 110: Processor

[0300] 120: Memory

[0301] 130: Model

[0302] 131: Speech Encoder

[0303] 133: Text Decoder

[0304] 140: Communication circuit

[0305] 200-1, 200-2, ..., 200-n: terminal

Claims

In electronic devices, processor; and A memory electrically connected to the processor and storing at least one instruction executed by the processor. Includes, The above processor is, When the above at least one command is executed, the electronic device receives voice input of a first language from a first terminal; Determining text of a second language corresponding to voice input of the first language; The text of the second language is provided to a second terminal, wherein at least a portion of the text of the second language is provided to the second terminal while voice input of the first language is received. Electronic device. In paragraph 1, The above processor is, Providing text of the second language through at least one chat window of the first terminal and the second terminal, Electronic device. In paragraph 1, The above processor is, When a plurality of voice inputs are received from the first terminal, text corresponding to the plurality of voice inputs is provided to the first terminal; An electronic device that determines the speaker of the voice input of the first language based on the input received from the first terminal. In paragraph 1, The above processor is, When a speaker corresponding to the second terminal is registered, among the voice inputs input from the first terminal, a voice input is determined excluding the voice input of the speaker corresponding to the second terminal; Determining the speaker of the voice input of the first language by using voice input excluding the voice input of the speaker corresponding to the second terminal. Electronic device. In paragraph 1, The above processor is, An interface for determining the speaker of the voice input of the first language through the first terminal is provided, and Determining the speaker using voice input corresponding to the above interface, Electronic device. In paragraph 1, The above processor is, Determining the text of the second language based on at least one of the text corresponding to the previously entered voice input of the first language, the previously provided text of the second language, and a set keyword. Electronic device. In paragraph 1, The above processor is, Determining the characteristics of the voice input of the first language above; Based on the above features, generate a voice output corresponding to the text of the second language; Providing voice output corresponding to the text of the second language to the second terminal Electronic device. In paragraph 1, The above processor is, Provides text of a third language corresponding to voice input of the first language to a third terminal, wherein at least a portion of the text of the third language is provided while voice input of the first language is received. Electronic device. In paragraph 1, The above processor is, Receiving voice input of the second language through the second terminal; Provides text of a first language corresponding to voice input of the second language to a first terminal, wherein at least a portion of the text of the first language is provided while voice input of the second language is received. Electronic device. In Paragraph 9, The above processor is, Based on the text of the first language and the text of the second language, providing at least one of a conversation record and a conversation summary to the first terminal or the second terminal. Electronic device. In paragraph 1, The above processor is, Generating text in the second language by using a model trained to input voice input in the first language and output text in the second language. Electronic device. In Paragraph 11, The above processor is, Using the above model, while voice input of the first language is received, at least one of a part of the text of the second language corresponding to the input voice input that has a determined meaning and a part of the text that has not a determined meaning is determined; While voice input of the first language is being received, at least one of the text of the second language with a determined meaning and the text of the second language with an undetermined meaning is provided to the second terminal. Electronic device.