Electronic device and control method thereof
The dual neural network model in the electronic device addresses the challenge of real-time translation by individually processing sentences and extracting keywords, improving accuracy and speed in translating multiple sentences.
Patent Information
- Application Number
- PCT/KR2024/021158
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-02
- Filing Date
- 2024-12-26
- Publication Date
- 2025-08-07
AI Technical Summary
Existing electronic devices struggle with maintaining connectivity and accuracy in real-time translation of multiple sentences, often leading to slower processing speeds when considering entire sentences and losing connectivity between preceding and following sentences.
An electronic device employs a dual neural network model architecture that processes user input by first translating individual sentences and then extracting keywords, using a first neural network for translation and a second for keyword extraction, to improve connectivity and speed in real-time translation.
This approach enhances translation accuracy and speed by maintaining connectivity between sentences while processing multiple sentences concurrently, ensuring timely and coherent translation results.
Smart Images

Figure KR2024021158_07082025_PF_FP_ABST
Abstract
Description
Electronic device and method of controlling the same
[0001] The present disclosure relates to an electronic device and a control method thereof, and more particularly, to an electronic device providing a translation function for user input and a control method thereof.
[0002] When providing a translation function for user input, the user may speak multiple sentences consecutively. If the user speaks multiple sentences before performing the translation function, the translation of these sentences must be processed in real time.
[0003] When translating multiple sentences simultaneously, each sentence may be translated individually to speed up processing. Individual translations can lead to a loss of connectivity between preceding and following sentences.
[0004] Translation operations can be divided into preset time units. Even when user input (user voice) is divided into preset time units, the problem of poor connectivity between preceding and following sentences can arise.
[0005] Translating the next sentence by considering the entire translated sentence can improve accuracy. However, considering the entire translated sentence can lead to longer processing times and slower processing speeds.
[0006] The present disclosure is designed to improve the above-described problems, and an object of the present disclosure is to provide an electronic device and a control method thereof that translates current data by considering keywords of previous data.
[0007] According to one embodiment, an electronic device includes a memory, a processor connected to the memory, wherein the processor, when a first user voice in a first language is received, inputs the first user voice into a first neural network model to obtain a first translation text in a second language corresponding to the first user voice, inputs the first user voice into a second neural network model to obtain a first keyword text in the second language corresponding to the first user voice, and when a second user voice in the first language is received, inputs the second user voice and the first keyword text into the first neural network model to obtain a second translation text in the second language corresponding to the second user voice, and provides the first translation text and the second translation text.
[0008] The processor can identify the first language based on the first user voice, obtain the first translation text in the second language, which is a preset target language, based on the first user voice in the first language through the first neural network model, and obtain the second translation text in the second language based on the second user voice in the first language and the first keyword text through the first neural network model.
[0009] The processor may obtain an audio signal including the first user voice and the second user voice, divide the audio signal into first audio data including the first user voice and second audio data including the second user voice based on a preset unit time, input the first audio data into the first neural network model to obtain the first translation text, input the first audio data into the second neural network model to obtain the first keyword text, and input the second audio data and the first keyword text into the first neural network model to obtain the second translation text.
[0010] The second user voice may be a voice spoken following the first user voice.
[0011] The processor may obtain a second keyword text corresponding to the second user voice based on the second audio data and the first keyword text, obtain third audio data including a third user voice in the first language spoken following the second user voice, input the third audio data and the second keyword text into the first neural network model to obtain a third translation text corresponding to the third user voice, and provide the first translation text, the second translation text, and the third translation text.
[0012] The memory stores an artificial intelligence model including the first neural network model and the second neural network model, and the processor can obtain the first translated text and the second translated text by inputting the first user voice and the second user voice into the artificial intelligence model.
[0013] The artificial intelligence model may be a model trained based on first learning audio data including a first learning user voice in the first language, second learning audio data including a second learning user voice in the first language, first learning keyword text in the second language corresponding to the first learning user voice, learning translation text in the second language corresponding to the second learning user voice, and second learning keyword text in the second language corresponding to the second learning user voice.
[0014] The artificial intelligence model may be a model trained to receive the second learning audio data and the first learning keyword text and output the learning translation text by performing a function of translating into the second language.
[0015] The artificial intelligence model may be a model trained to receive the second learning audio data, the first learning keyword text, and the noise keyword text and output the learning translation text by performing a function of translating into the second language.
[0016] The artificial intelligence model may be a model trained to receive the second learning audio data and the first learning keyword text and output the second learning keyword text by performing a function of extracting keywords in the second language.
[0017] According to one embodiment, a method for controlling an electronic device includes the steps of: when a first user voice in a first language is received, inputting the first user voice into a first neural network model to obtain a first translation text in a second language corresponding to the first user voice; inputting the first user voice into a second neural network model to obtain a first keyword text in the second language corresponding to the first user voice; when a second user voice in the first language is received, inputting the second user voice and the first keyword text into the first neural network model to obtain a second translation text in the second language corresponding to the second user voice; and providing the first translation text and the second translation text.
[0018] The control method further includes a step of identifying the first language based on the first user voice, and the step of obtaining the first translation text may obtain the first translation text in the second language, which is a preset target language, based on the first user voice in the first language through the first neural network model, and the step of obtaining the second translation text may obtain the second translation text in the second language based on the second user voice in the first language and the first keyword text through the first neural network model.
[0019] The control method further includes a step of obtaining an audio signal including the first user voice and the second user voice, and a step of dividing the audio signal into first audio data including the first user voice and second audio data including the second user voice based on a preset unit time, wherein the step of obtaining the first translation text may input the first audio data into the first neural network model to obtain the first translation text, the step of obtaining the first keyword text may input the first audio data into the second neural network model to obtain the first keyword text, and the step of obtaining the second translation text may input the second audio data and the first keyword text into the first neural network model to obtain the second translation text.
[0020] The second user voice may be a voice spoken following the first user voice.
[0021] The control method may further include a step of obtaining a second keyword text corresponding to the second user voice based on the second audio data and the first keyword text, a step of obtaining third audio data including a third user voice in the first language spoken following the second user voice, a step of inputting the third audio data and the second keyword text into the first neural network model to obtain a third translation text corresponding to the third user voice, and a step of providing the first translation text, the second translation text, and the third translation text.
[0022] The electronic device stores an artificial intelligence model including the first neural network model and the second neural network model, and the step of obtaining the first translated text and the second translated text may obtain the first translated text and the second translated text by inputting the first user voice and the second user voice into the artificial intelligence model.
[0023] The artificial intelligence model may be a model trained based on first learning audio data including a first learning user voice in the first language, second learning audio data including a second learning user voice in the first language, first learning keyword text in the second language corresponding to the first learning user voice, learning translation text in the second language corresponding to the second learning user voice, and second learning keyword text in the second language corresponding to the second learning user voice.
[0024] The artificial intelligence model may be a model trained to receive the second learning audio data and the first learning keyword text and output the learning translation text by performing a function of translating into the second language.
[0025] The artificial intelligence model may be a model trained to receive the second learning audio data, the first learning keyword text, and the noise keyword text and output the learning translation text by performing a function of translating into the second language.
[0026] The artificial intelligence model may be a model trained to receive the second learning audio data and the first learning keyword text and output the second learning keyword text by performing a function of extracting keywords in the second language.
[0027] FIG. 1 is a diagram for explaining an operation of providing a translation function according to one embodiment.
[0028] FIG. 2 is a block diagram illustrating an electronic device according to one embodiment.
[0029] FIG. 3 is a drawing for explaining a device performing a translation function according to one embodiment.
[0030] FIG. 4 is a diagram for explaining an operation of providing a translated text according to one embodiment.
[0031] FIG. 5 is a diagram for explaining an operation of providing translation text for a first user voice and a second user voice, according to one embodiment.
[0032] FIG. 6 is a drawing for explaining an operation of analyzing a first user voice in the embodiment of FIG. 5, according to one embodiment.
[0033] FIG. 7 is a drawing for explaining an operation of analyzing a second user voice in the embodiment of FIG. 5, according to one embodiment.
[0034] FIG. 8 is a diagram for explaining an operation of providing a translation text corresponding to an audio signal, according to one embodiment.
[0035] FIG. 9 is a diagram for explaining an operation of providing a translation text using keyword text according to one embodiment.
[0036] FIG. 10 is a diagram for explaining an operation of providing a translated text using multiple neural network models according to one embodiment.
[0037] FIG. 11 is a diagram illustrating an operation for obtaining additional audio data according to one embodiment.
[0038] FIG. 12 is a diagram for explaining learning data according to one embodiment.
[0039] FIG. 13 is a diagram for explaining an operation of learning a first neural network model and a second neural network model using learning data of the embodiment of FIG. 12, according to one embodiment.
[0040] FIG. 14 is a diagram for explaining learning data according to one embodiment.
[0041] FIG. 15 is a diagram for explaining an operation of learning a first neural network model and a second neural network model using learning data of the embodiment of FIG. 14, according to one embodiment.
[0042] FIG. 16 is a diagram for explaining learning data according to one embodiment.
[0043] FIG. 17 is a diagram for explaining an operation of learning a first neural network model with learning data of the embodiment of FIG. 16, according to one embodiment.
[0044] FIG. 18 is a diagram for explaining learning data according to one embodiment.
[0045] FIG. 19 is a diagram for explaining an operation of learning a first neural network model with learning data of the embodiment of FIG. 18, according to one embodiment.
[0046] FIG. 20 is a drawing for explaining a method for controlling an electronic device according to one embodiment.
[0047] Hereinafter, the present disclosure will be described in detail with reference to the attached drawings.
[0048] The terms used in the embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of this disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant disclosure. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on the meanings of the terms and the overall content of this disclosure.
[0049] In this specification, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a feature (e.g., a number, function, operation, or component such as a part), and do not exclude the presence of additional features.
[0050] The expression "at least one of A and / or B" should be understood to mean either "A" or "B" or "A and B".
[0051] As used herein, the expressions “first,” “second,” “first,” or “second,” etc., may describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0052] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that the component may be directly coupled to the other component, or may be connected through another component (e.g., a third component).
[0053] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this application, terms such as "comprise" or "consist of" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0054] In the present disclosure, a "module" or "part" performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Furthermore, multiple "modules" or multiple "parts" may be integrated into at least one module and implemented as at least one processor, excluding any "modules" or "parts" that need to be implemented as specific hardware.
[0055] In this specification, the term user may refer to a person using an electronic device or a device using an electronic device (e.g., an artificial intelligence electronic device).
[0056] An embodiment of the present disclosure will be described in more detail with reference to the attached drawings below.
[0057] FIG. 1 is a diagram for explaining an operation of providing a translation function according to one embodiment.
[0058] Referring to embodiment (10) of FIG. 1, the electronic device (100) can perform a function of providing a translation result corresponding to a user input. The electronic device (100) can translate the user input using a preset language. The electronic device (100) can provide the user with output data in the preset language.
[0059] For example, the electronic device (100) can receive a user input (e.g., [EN] hello, world) in a first language (e.g., English). The electronic device (100) can perform a translation function on the user input (e.g., [EN] hello, world) to provide a translation result (e.g., [KR] hello, world).
[0060] The electronic device (100) can receive user input. The user input may include at least one of an audio signal including the user's voice or text information entered by the user. For example, the electronic device (100) can receive the user input in the form of an audio signal. For example, the electronic device (100) can receive the user input in the form of text information.
[0061] FIG. 2 is a block diagram illustrating an electronic device according to one embodiment.
[0062] An electronic device (100) may include a memory (110) and a processor (120) connected to the memory (110).
[0063] The memory (110) may be described as at least one memory. The processor (120) may be described as at least one processor.
[0064] The memory (110) may be implemented as an internal memory such as a ROM (e.g., an electrically erasable programmable read-only memory (EEPROM)) or RAM included in the processor (120), or may be implemented as a separate memory from the processor (120). The memory (110) may be implemented as a memory embedded in the electronic device (100) or as a memory that can be detachably attached to the electronic device (100) depending on the purpose of data storage. For example, data for driving the electronic device (100) may be stored in a memory embedded in the electronic device (100), and data for expanding the functions of the electronic device (100) may be stored in a memory that can be detachably attached to the electronic device (100).
[0065] In the case of memory embedded in the electronic device (100), it may be implemented as at least one of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), etc.), hard drive, or solid state drive (SSD), and in the case of memory that can be attached or detached to the electronic device (100), it may be implemented in the form of a memory card (e.g., compact flash (CF), secure digital (SD), micro secure digital (Micro-SD), mini secure digital (Mini-SD), extreme digital (xD), multi-media card (MMC), etc.), external memory that can be connected to a USB port (e.g., USB memory), etc.
[0066] The memory (110) can store at least one instruction. Based on the instruction stored in the memory (110), the processor (120) can perform various operations.
[0067] The processor (120) may be implemented as a digital signal processor (DSP), a microprocessor, or a time controller (TCON) that processes digital signals. However, the present invention is not limited thereto, and may include one or more of a central processing unit (CPU), a micro controller unit (MCU), a micro processing unit (MPU), a controller, an application processor (AP), a graphics-processing unit (GPU), a communication processor (CP), or an advanced reduced instruction set computer (RISC) machines (ARM) processor, or may be defined by the relevant terminology. The processor (120) may be implemented as a system on chip (SoC) or large scale integration (LSI) having a processing algorithm built in, or may be implemented in the form of a field programmable gate array (FPGA). The processor (120) may perform various functions by executing computer executable instructions stored in a memory.
[0068] When a first user voice in a first language is received, the processor (120) inputs the first user voice into a first neural network model (45-1) to obtain a first translation text in a second language corresponding to the first user voice, inputs the first user voice into a second neural network model (45-2) to obtain a first keyword text in a second language corresponding to the first user voice, and when a second user voice in the first language is received, the processor (120) inputs the second user voice and the first keyword text into the first neural network model (45-1) to obtain a second translation text in the second language corresponding to the second user voice, and can provide the first translation text and the second translation text.
[0069] The first language may be the language of the user's spoken voice. The first language may be the voice spoken by the user. The processor (120) may identify the first language based on the first user's voice. The processor (120) may identify the first language corresponding to the first user's voice.
[0070] The user's spoken speech may include at least one language. The first language may represent the primary language of the user's spoken speech. The processor (120) may identify the language with the largest proportion of the at least one language included in the first user's speech as the first language. The first language may be described as the identification language.
[0071] For example, assume a user's speech consists of 90% English and 10% Korean. The user's primary language may be English.
[0072] The memory (110) can store an artificial intelligence model including a first neural network model (45-1) and a second neural network model (45-2).
[0073] The processor (120) can perform a function of translating a user's voice using an artificial intelligence model. The processor (120) can input the user's voice in a first language into the artificial intelligence model to obtain a translated text (or translation result) in a second language corresponding to the user's voice.
[0074] A second language may be the target language of translation when performing a translation function. The second language may be referred to as the translation language, the target language, or the target language.
[0075] The artificial intelligence model may include at least one detailed model. The artificial intelligence model may include at least one of the first neural network model (45-1) or the second neural network model (45-2).
[0076] The first neural network model (45-1) may be a model that performs a translation function. The second neural network model (45-2) may be a model that performs a keyword extraction function. A description of this is provided in Fig. 4.
[0077] The processor (120) can obtain a first translation text in a second language, which is a preset target language, based on a first user voice in a first language through a first neural network model (45-1).
[0078] The first translation text may include text information translated from the first user's speech input (or spoken) in the first language into the second language.
[0079] For example, when receiving input data containing words, the first neural network model (45-1) can obtain words translated into a second language as output data. The input data containing words may be data containing user speech without a sentence structure. For example, the input data may include a user speech uttering "apple, egg, pear."
[0080] For example, when receiving input data containing a sentence, the first neural network model (45-1) can obtain a sentence translated into a second language as output data. For example, the input data may include a user's voice uttering "I am looking for an apple, an egg, and a pear."
[0081] The processor (120) can input the first user voice into the second neural network model (45-2) to obtain a first keyword text in a second language corresponding to the first user voice.
[0082] The first keyword text may represent the result of extracting keywords in a second language from the first user's speech in the first language. The keyword text may represent a word identified as having relatively high importance among multiple words included in the user's speech.
[0083] The processor (120) can identify the importance of each of the multiple words included in the user's voice, and identify words with an importance exceeding a threshold as keywords. Keywords can be identified or stored in text form. Accordingly, keywords can be described as keyword text.
[0084] In one embodiment, the keyword text may include words in the user's speech whose importance exceeds a threshold. For example, in a user's speech that utters "I'm looking for apples, eggs, and pears," the keyword text may be "apples, eggs, pears."
[0085] In one embodiment, the keyword text may include alternative words that represent what the user's speech means. For example, in a user's speech utterance, "I'm looking for apples, eggs, and pears," the keyword text may be "buy."
[0086] The processor (120) can obtain (or extract) a first keyword text in a second language from a first user voice in a first language through a second neural network model (45-2).
[0087] The first keyword text may include keywords in a second language extracted from a first user speech input (or spoken) in a first language.
[0088] For example, when receiving input data containing words, the second neural network model (45-2) can obtain words extracted in the second language as output data. The input data containing words may be data containing user speech without a sentence structure. For example, the input data may include a user speech uttering "apple, egg, pear." The extracted keyword text may be "apple, egg, pear."
[0089] For example, when receiving input data containing sentences, the second neural network model (45-2) can obtain words extracted in the second language as output data. For example, the input data may include a user's voice uttering "I'm looking for an apple, an egg, and a pear." The extracted keyword text may be "apple, egg, pear."
[0090] In the first neural network model (45-1), the form of the output data (word or sentence) can be determined according to the form of the input data (word or sentence).
[0091] In the second neural network model (45-2), the form of the output data (word) can be fixed regardless of the form of the input data (word or sentence).
[0092] When a second user voice in a first language is received, the processor (120) can input the second user voice and the first keyword text into the first neural network model (45-1) to obtain a second translation text in a second language corresponding to the second user voice.
[0093] The second user voice may be a voice spoken following the first user voice. The user may speak the first user voice and then speak the second user voice in succession. The second user voice may be a follow-up voice to the first user voice.
[0094] The processor (120) can obtain a second translation text in a second language based on the second user voice and the first keyword text in the first language through the first neural network model (45-1).
[0095] The second translation text may include text information translated from a second user's speech input (or spoken) in a first language into a second language.
[0096] The processor (120) can provide the first translated text and the second translated text to the user.
[0097] The processor (120) can generate a translation screen including a first translation text and a second translation text. The processor (120) can provide the generated translation screen to a user.
[0098] According to various embodiments, the processor (120) may provide text representing the user's speech (or user input) along with the translation result. For example, the processor (120) may provide the user's speech displayed in a first language along with the translation result displayed in a second language.
[0099] The processor (120) can obtain a first input text representing a first user's speech in a first language. The processor (120) can obtain a second input text representing a second user's speech in the first language.
[0100] The processor (120) can provide a first input text in a first language, a first translation text in a second language corresponding to the first input text, a second input text in the second language, and a second translation text in the second language corresponding to the second input text.
[0101] The processor (120) can generate a translation screen including a first input text, a first translated text, a second input text, and a second translated text.
[0102] The translation screen can be described as a translation result screen or a translation GUI (Graphical User Interface).
[0103] The processor (120) obtains an audio signal including a first user voice and a second user voice, and divides the audio signal into first audio data including the first user voice and second audio data including the second user voice based on a preset unit time,
[0104] The first audio data is input into the first neural network model (45-1) to obtain the first translated text,
[0105] The first audio data is input into the second neural network model (45-2) to obtain the first keyword text,
[0106] The second audio data and the first keyword text can be input into the first neural network model (45-1) to obtain the second translated text.
[0107] The processor (120) can obtain an audio signal including at least one user voice. The audio signal can include audio data. The audio signal can include information representing a waveform of a received sound (e.g., a user voice).
[0108] The audio data may be data that includes at least a portion of an entire audio signal. The audio data may include information for representing at least a portion of the audio signal. The audio data may include at least a portion of the audio signal and identification information (e.g., audio data number 1) representing at least a portion of the audio signal. The audio data may include a sound waveform and additional information (e.g., identification information) related to the sound waveform.
[0109] The processor (120) can segment an audio signal based on preset unit information. The preset unit information may include at least one of time, number of words, data size, sentences, and ratio. A description thereof is provided in FIG. 4.
[0110] For example, assume that an audio signal includes 4 seconds of user speech. The processor (120) may divide the 4 seconds of audio signal into first audio data and second audio data. The first audio data may include the user speech between 0 and 2 seconds and first identification information indicating 0 and 2 seconds. The second audio data may include the user speech between 2 and 4 seconds and second identification information indicating 2 and 4 seconds.
[0111] The processor (120) may divide an audio signal into first audio data and second audio data based on a preset time. The first audio data may include a first user's voice. The second audio data may include a second user's voice.
[0112] The first audio data may include an audio signal representing a first user's voice. The second audio data may include an audio signal representing a second user's voice.
[0113] The processor (120) may receive an additional voice after receiving the second user voice. The additional voice may be described as a subsequent voice. The processor (120) may obtain a translation result for the additional voice using the additional voice and keyword text extracted from the previous voice.
[0114] If it is identified that an additional voice exists, the processor (120) can obtain a second keyword text corresponding to the second user voice based on the second audio data and the first keyword text.
[0115] The processor (120) can receive a third user voice that is received following the second user voice. The processor (120) can obtain third audio data including the third user voice in the first language spoken following the second user voice.
[0116] The processor (120) can input third audio data and second keyword text into the first neural network model (45-1) to obtain third translation text corresponding to the third user voice.
[0117] The processor (120) can provide a first translated text, a second translated text, and a third translated text. The processor (120) can generate a translation screen including the first translated text, the second translated text, and the third translated text.
[0118] An example of obtaining a third translation text is additionally described in FIG. 11.
[0119] The processor (120) can provide translation results corresponding to the user's voice through an artificial intelligence model. The processor (120) can obtain a first translation text and a second translation text by inputting a first user's voice and a second user's voice into the artificial intelligence model.
[0120] An AI model can be a trained model. Training data can be used to train the AI model.
[0121] The artificial intelligence model may be a model trained based on first training audio data including a first learning user voice in a first language, second training audio data including a second learning user voice in the first language, first learning keyword text in the second language corresponding to the first learning user voice, learning translation text in the second language corresponding to the second learning user voice, and second learning keyword text in the second language corresponding to the second learning user voice.
[0122] The artificial intelligence model can learn a first neural network model (45-1) that directly performs a translation function and a second neural network model (45-2) that performs a keyword extraction function. The artificial intelligence model can learn the first neural network model (45-1) and the second neural network model (45-2) using a training data set. The training data set can include first training audio data, second training audio data, first training keyword text, second training keyword text, and training translation text.
[0123] The artificial intelligence model can be trained based on first learning audio data, second learning audio data, first learning keyword text, second learning keyword text, and training translation text.
[0124] The first learning audio data may include a first learning user's speech in a first language.
[0125] The second learning audio data may include the second learning user's speech in the first language.
[0126] The second learning user voice may be a voice spoken following the first learning user voice.
[0127] The first learning keyword text may include keywords in a second language extracted from the first learning user's speech. The first learning keyword text may include keyword text in a second language corresponding to the first learning user's speech in the first language.
[0128] The second learning keyword text may include keywords in a second language extracted from the second learning user's speech. The second learning keyword text may include keyword text in a second language corresponding to the second learning user's speech in the first language.
[0129] The learning translation text may be a learning translation text in a second language that corresponds to the second learning user's speech in the first language.
[0130] Additional descriptions regarding the learning data set are provided in Figures 12 and 14.
[0131] The artificial intelligence model may be a model trained based on a training data set. The artificial intelligence model may be a model trained to receive second training audio data and first training keyword text and output a training translation text by performing a function of translating into a second language through the first neural network model (45-1). The training operation of the artificial intelligence model related to the translation function may represent the training operation of the first neural network model (45-1). Additional descriptions related to this are described in the embodiment (1310) of FIG. 13 and the embodiment (1510) of FIG. 15.
[0132] The artificial intelligence model can learn the first neural network model (45-1) that performs the translation function using noise keywords.
[0133] Noise keywords can represent learning keywords that are input to provide consistent quality translation results even when the translation results include keywords with low relevance. Noise keyword text can represent text representing noise keywords.
[0134] In one embodiment, the noise keyword may be a preset keyword. The noise keyword may be at least a keyword preset according to the user's settings.
[0135] In one embodiment, the noise keyword may be a randomly generated keyword. The noise keyword may be a keyword identified based on a random number during the learning process. The noise keyword may be randomly generated each time a learning operation is performed.
[0136] The artificial intelligence model may be a model trained to receive second learning audio data, first learning keyword text, and noise keyword text and output a learning translation text by performing a function of translating into a second language through the first neural network model (45-1). Descriptions related to the noise keyword are described in FIGS. 18 and 19.
[0137] The artificial intelligence model may be a model trained to receive second learning audio data and first learning keyword text and output second learning keyword text by performing a function of extracting keywords in a second language using a second neural network model (45-2). The learning operation of the artificial intelligence model related to the keyword extraction function may represent the learning operation of the second neural network model (45-2). Additional descriptions related thereto are described in the embodiment (1320) of FIG. 13 and the embodiment (1520) of FIG. 15.
[0138] The electronic device (100) described an operation of analyzing user voice or audio data and providing a translation result.
[0139] According to various embodiments, the electronic device (100) may analyze input data in text format and provide analysis results. The user voice described in this disclosure may be replaced with user input. The audio data described in this disclosure may be described as text data.
[0140] The electronic device (100) may additionally include at least one of a communication interface (130) or a display (140).
[0141] The communication interface (130) is a configuration that performs communication with various types of external devices according to various types of communication methods. The communication interface (130) may include a wireless communication module or a wired communication module. Each communication module may be implemented in the form of at least one hardware chip.
[0142] A wireless communication module may be a module that communicates wirelessly with an external device. For example, the wireless communication module may include at least one of a Wi-Fi module, a Bluetooth module, an infrared communication module, or other communication modules.
[0143] Wi-Fi and Bluetooth modules can communicate via Wi-Fi and Bluetooth, respectively. When using a Wi-Fi or Bluetooth module, various connection information, such as the service set identifier (SSID) and session key, is first transmitted and received. This information is then used to establish a communication connection before various other information can be transmitted and received.
[0144] Infrared communication modules perform communication based on infrared communication (IrDA, infrared Data Association) technology, which transmits data wirelessly over short distances using infrared light, which is between visible light and millimeter waves.
[0145] In addition to the above-described communication method, other communication modules may include at least one communication chip that performs communication according to various wireless communication standards such as zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), 5G (5th Generation), etc.
[0146] A wired communication module may be a module that communicates with an external device via a wire. For example, the wired communication module may include at least one of a Local Area Network (LAN) module, an Ethernet module, a paired cable, a coaxial cable, a fiber optic cable, or an Ultra Wide-Band (UWB) module.
[0147] According to various embodiments, the communication interface (130) may utilize the same communication module (e.g., a Wi-Fi module) to communicate with an external device such as a remote control device and an external server.
[0148] According to various embodiments, the communication interface (130) may utilize different communication modules to communicate with external devices, such as remote control devices, and external servers. For example, the communication interface (130) may utilize at least one of an Ethernet module or a Wi-Fi module to communicate with an external server, and may also utilize a Bluetooth module to communicate with an external device, such as a remote control device. However, this is merely an embodiment, and the communication interface (130) may utilize at least one of various communication modules when communicating with multiple external devices or external servers.
[0149] The display (140) may be implemented as a variety of displays such as a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display panel (PDP), etc. The display (140) may also include a driving circuit, a backlight unit, etc., which may be implemented as a form such as an a-si TFT (amorphous silicon thin film transistor), an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc. The display (140) may be implemented as a touch screen combined with a touch sensor, a flexible display, a three-dimensional display (3D display, three-dimensional dispaly), etc. According to an embodiment of the present disclosure, the display (140) may include not only a display panel that outputs an image, but also a bezel that houses the display panel. In particular, according to an embodiment of the present disclosure, the bezel may include a touch sensor for detecting user interaction.
[0150] FIG. 3 is a drawing for explaining a device performing a translation function according to one embodiment.
[0151] Referring to the embodiment (310) of FIG. 3, the electronic device (100) may include at least one of an audio receiving module (40), a translation module (45), or a result providing module (47).
[0152] The audio receiving module (40) may be a module that receives an audio signal. The audio receiving module (40) may be a module that converts an audio signal in the form of an analog signal into an audio signal in the form of a digital signal. For example, the audio receiving module (40) may include a microphone.
[0153] The translation module (45) may be a module that provides translation results corresponding to user input. The translation module (45) may be a module that performs a translation function. The translation module (45) may include at least one neural network model. For example, the translation module (45) may include at least one of the first neural network model (45-1) or the second neural network model (45-2). A description thereof is provided in FIG. 4.
[0154] The result provision module (47) may be a module that provides translation results to the user. The result provision module (47) may be a module for providing the results of a translation function. For example, the result provision module (47) may include a display, a speaker, etc.
[0155] The electronic device (100) can provide a translation function in an on-device form without using an external device.
[0156] Referring to embodiment (320) of FIG. 3, unlike embodiment (310) of FIG. 3, the server (200) can perform a translation function.
[0157] The electronic device (100) may include at least one of an audio receiving module (40) or a result providing module (47). The server (200) may include a translation module (45). Since the description of each module is described in the embodiment (310), a redundant description is omitted.
[0158] The electronic device (100) can transmit an audio signal acquired through the audio receiving module (40) to the server (200). The server (200) can obtain a translation result by performing a translation function on the received audio signal. The server (200) can transmit the translation result to the electronic device (100).
[0159] The electronic device (100) can provide the translation results obtained from the server (200) to the user through the result provision module (47).
[0160] FIG. 4 is a diagram for explaining an operation of providing a translated text according to one embodiment.
[0161] The electronic device (100) may include at least one of an audio receiving module (40), an audio segmentation module (41), a signal management module (42), an audio encoder (43), a text encoder (44), a translation module (45), a result storage module (46), and a result providing module (47). The translation module (45) may include a first neural network model (45-1) and a second neural network model (45-2).
[0162] The audio receiving module (40) may be a module that receives an audio signal. A description thereof is provided in FIG. 3. The audio signal may include at least one user voice. The audio receiving module (40) may transmit the audio signal to the audio segmentation module (41) (S1).
[0163] The audio segmentation module (41) can obtain an audio signal from the audio receiving module (40). The audio segmentation module (41) can segment the audio signal. The audio segmentation module (41) can segment the audio signal into at least one piece of audio data based on preset unit information.
[0164] The audio data may represent a portion of a segmented audio signal. The electronic device (100) may segment the audio signal into first audio data and second audio data.
[0165] The preset unit information may include at least one of time, number of words, data size, sentence, and ratio.
[0166] For example, the audio segmentation module (41) can segment an audio signal based on a preset time.
[0167] For example, the audio segmentation module (41) can segment an audio signal based on a preset number of words.
[0168] For example, the audio segmentation module (41) can segment an audio signal based on a preset data size.
[0169] For example, the audio segmentation module (41) can segment an audio signal into sentence units. Sentence units can be separated by periods.
[0170] For example, the audio segmentation module (41) can segment an audio signal based on a preset ratio from the entire audio signal.
[0171] The audio segmentation module (41) can transmit the first audio data and the second audio data to the signal management module (42) (S2).
[0172] The signal management module (42) may be a module that controls a signal for performing a translation function. The signal management module (42) may be described as a signal control module, a translation management module, a translation control module, etc.
[0173] The signal management module (42) can receive first audio data and second audio data from the audio segmentation module (41). The signal management module (42) can transmit the first audio data to the audio encoder (43) (S3).
[0174] An audio encoder (43) may be a module that converts audio data (or audio signals) into feature data. The feature data may be data containing features necessary for analyzing (translating, extracting keywords, etc.) user voice contained in the audio data. The feature data may include at least one of acoustic features and linguistic features. Acoustic features may include pronunciation, intonation, stress, rhythm, etc. Linguistic features may include words, grammar, sentence structure, etc.
[0175] The value converted (or output) by the audio encoder (43) may not be a text result of audio data conversion. The audio encoder (43) may obtain a feature vector for analyzing the user's voice included in the audio data, and generate feature data including the feature vector. The feature vector may represent a mathematical value representing a feature used to analyze a voice signal.
[0176] An audio encoder (43) can generate first audio feature data including a feature vector of first audio data. The audio encoder (43) can transmit the first audio feature data to a first neural network model (45-1) (S4). The audio encoder (43) can transmit the first audio feature data to a second neural network model (45-2) (S5).
[0177] The first neural network model (45-1) may be a model that performs a translation function. The first neural network model (45-1) may be a model that outputs a translated text corresponding to input data. The first neural network model (45-1) may be a model that receives input data and provides a translated text corresponding to the input data as output data. The first neural network model (45-1) may be described as a translation decoder.
[0178] The first neural network model (45-1) can receive first audio feature data from the audio encoder (43). The first neural network model (45-1) can obtain a first translation text based on the first audio feature data. The first neural network model (45-1) can transmit the first translation text to the result storage module (46) (S6).
[0179] The first neural network model (45-1) can utilize keyword feature data transmitted from the text encoder (44) to perform the translation function. When audio data is first received, previously analyzed keyword text may not exist. When audio data is first received, the first neural network model (45-1) can output the translated text using only audio feature data, without utilizing keyword feature data.
[0180] The second neural network model (45-2) may be a model that performs a keyword extraction function. The second neural network model (45-2) may be a model that outputs keyword text corresponding to input data. The second neural network model (45-2) may be a model that receives input data and provides keyword text corresponding to the input data as output data. The second neural network model (45-2) may be described as a context text extraction decoder.
[0181] The second neural network model (45-2) can receive first audio feature data from the audio encoder (43). The second neural network model (45-2) can obtain first keyword text based on the first audio feature data. The second neural network model (45-2) can transmit the first keyword text to the result storage module (46) (S7).
[0182] The second neural network model (45-2) can utilize keyword feature data transmitted from the text encoder (44) to perform the keyword extraction function. When audio data is first received, previously analyzed keyword text may not exist. When audio data is first received, the second neural network model (45-2) can output keyword text using only audio feature data without utilizing the keyword feature data.
[0183] The result storage module (46) may be a module that stores various received information. The result storage module (46) may store the received information until a preset event occurs. The result storage module (46) may be described as a buffer.
[0184] The result storage module (46) can receive the first translation text from the first neural network model (45-1). The result storage module (46) can receive the first keyword text from the second neural network model (45-2). The result storage module (46) can store the first translation text and the first keyword text.
[0185] The result storage module (46) can transmit the first keyword text to the signal management module (42) (S8).
[0186] The signal management module (42) can receive the first keyword text from the result storage module (46).
[0187] When the first keyword text is received, the signal management module (42) can transmit the second audio data to the audio encoder (43) (S9).
[0188] The audio encoder (43) can receive second audio data from the signal management module (42).
[0189] The audio encoder (43) can generate second audio feature data including a feature vector of second audio data. The audio encoder (43) can transmit the second audio feature data to a first neural network model (45-1) (S10). The audio encoder (43) can transmit the second audio feature data to a second neural network model (45-2) (S11).
[0190] When the first keyword text is received, the signal management module (42) can transmit the first keyword text to the text encoder (44) (S12).
[0191] A text encoder (44) may be a module that converts text data into feature data. The feature data may be data containing features necessary for analyzing (translating, extracting keywords, etc.) user speech contained in the text data. The feature data may include linguistic features. Linguistic features may include word embeddings, sentence embeddings, contextual information, grammar, sentence structure, etc. The text encoder (44) may be described as a contextual text encoder.
[0192] Word embeddings can involve converting words into vectors in a space of a specific dimension. Word embeddings can numerically represent the meaning of a specific word. Words with similar meanings can be positioned close to each other in a space of a specific dimension.
[0193] Sentence embedding can involve converting a sentence into a vector in a specific dimensional space. A sentence embedding can numerically represent the meaning of a specific sentence. Sentences with similar meanings can be positioned close to each other in a specific dimensional space.
[0194] Contextual information can encompass meanings recognized not through individual words but through their relationships with surrounding words. The same word can have different meanings depending on the preceding and following sentences. Contextual information can indicate the meaning of a specific word, considering the preceding and following sentences.
[0195] Grammar can represent information related to grammar contained in text data.
[0196] Sentence structure can represent information related to the structure of sentences contained in text data.
[0197] A text encoder (44) can obtain a feature vector for representing text data and generate feature data including the feature vector. The feature vector can represent a mathematical value representing a feature used to analyze a voice signal.
[0198] The text encoder (44) can receive the first keyword text from the signal management module (42). The text encoder (44) can generate first keyword feature data including a feature vector of the first keyword text. The text encoder (44) can transmit the first keyword feature data to the first neural network model (45-1) (S13). The text encoder (44) can transmit the first keyword feature data to the second neural network model (45-2) (S14).
[0199] The first neural network model (45-1) can receive second audio feature data from the audio encoder (43). The first neural network model (45-1) can receive first keyword feature data from the text encoder (44).
[0200] The first neural network model (45-1) can obtain a second translation text based on the second audio feature data and the first keyword feature data. The first neural network model (45-1) can transmit the second translation text to the result storage module (46) (S15).
[0201] The second neural network model (45-2) can receive second audio feature data from the audio encoder (43). The second neural network model (45-2) can receive first keyword feature data from the text encoder (44).
[0202] The second neural network model (45-2) can obtain second keyword text based on the second audio feature data and the first keyword feature data. The second neural network model (45-2) can transmit the second keyword text to the result storage module (46) (S16).
[0203] The result storage module (46) can receive a second translation text from the first neural network model (45-1). The result storage module (46) can receive a second keyword text from the second neural network model (45-2).
[0204] When the second translation text is received or the second keyword text is received, the result storage module (46) can transmit the second keyword text to the signal management module (42) (S17).
[0205] The signal management module (42) can receive the second keyword text from the result storage module (46).
[0206] When the second translation text is received, the signal management module (42) can determine whether subsequent audio data exists. If there is no subsequent audio data, the signal management module (42) can request the translation text from the result storage module (46) (S18).
[0207] The result storage module (46) can transmit the first translation text and the second translation text to the signal management module (42) in response to a translation text request from the signal management module (42) (S19).
[0208] The signal management module (42) can receive the first translation text and the second translation text from the result storage module (46). When the first translation text and the second translation text are received, the signal management module (42) can transmit the first translation text and the second translation text to the result provision module (47) (S20).
[0209] The result provision module (47) may be a module that provides translation results to the user. A description related to this is provided in Fig. 3.
[0210] The result provision module (47) can generate a translation screen including a first translation text and a second translation text. The result provision module (47) can provide a translation screen.
[0211] In various embodiments, step S17 may be replaced with an operation of transmitting a notification signal indicating that the second translated text has been obtained.
[0212] Depending on various embodiments, the operation of obtaining the second keyword text may be omitted. If it is determined that subsequent audio data to be translated does not exist, the signal management module (42) may not perform the operation of obtaining the keyword text through the second neural network model (45-2).
[0213] For example, if the second audio data is the final translation target, the signal management module (42) may not perform the operation (S11) of transmitting the second audio feature data to the second neural network model (45-2), the operation (S14) of transmitting the first keyword feature data to the second neural network model (45-2), etc. The second neural network model (45-2) may not generate the second keyword text.
[0214] If the second neural network model (45-2) does not generate the second keyword text, step S17 may be replaced with an operation of transmitting the second translation text.
[0215] FIG. 5 is a diagram for explaining an operation of providing translation text for a first user voice and a second user voice, according to one embodiment.
[0216] Referring to the embodiment (510) of FIG. 5, the electronic device (100) can receive an audio signal including a user's voice.
[0217] Referring to the embodiment (520) of FIG. 5, the electronic device (100) can divide an audio signal according to a preset method. The electronic device (100) can divide an audio signal into a plurality of audio data.
[0218] The preset method may utilize preset unit information. The preset unit information may include at least one of time, word count, data size, sentences, and ratio. A description of this is provided in Figure 4.
[0219] The electronic device (100) can divide the audio signal into first audio data (521) and second audio data (522) based on preset unit information.
[0220] For example, assume that an audio signal includes a user voice saying “Hello, my name is Chulsoo. I develop next generation technology to make a better world and society.” Assume that the user voice is in a first language (EN). The electronic device (100) can divide the audio signal into first audio data (521) including the first user voice (Hello, my name is Chulsoo.) and second audio data (522) including the second user voice (I develop next generation technology to make a better world and society.). The first language (EN) can represent English.
[0221] The first audio data (521) may be described as head audio data.
[0222] The second audio data (522) may be described as tail audio data.
[0223] FIG. 6 is a drawing for explaining an operation of analyzing a first user voice in the embodiment of FIG. 5, according to one embodiment.
[0224] Referring to the embodiment (610) of FIG. 6, the electronic device (100) can perform a translation function using the first neural network model (45-1).
[0225] The electronic device (100) can transmit first audio data (521) to the first neural network model (45-1) through the audio encoder (43).
[0226] The first audio data (521) can be converted into feature data through an audio encoder (43). The audio encoder (43) can generate first audio feature data corresponding to the first audio data (521). The audio encoder (43) can transmit the first audio feature data to a first neural network model (45-1).
[0227] The electronic device (100) can transmit temporary keyword text (611) to the first neural network model (45-1) via the text encoder (44). The temporary keyword text (611) can indicate that there is no text information corresponding to the keyword (none). Since the first audio data (521) corresponds to the first sentence of a sentence, the keyword text acquired in the previous step may not exist.
[0228] Temporary keyword text (611) can be converted into feature data through a text encoder (44). The text encoder (44) can generate temporary feature data corresponding to the temporary keyword text (611). The text encoder (44) can transmit the temporary keyword feature data to the first neural network model (45-1).
[0229] The first neural network model (45-1) can obtain the first translation text (612) based on the first audio data (521) and the temporary keyword text (611). Since the temporary keyword text (611) does not contain a keyword, the first neural network model (45-1) can substantially perform the translation function using only the first audio data (521).
[0230] For example, the first neural network model (45-1) may be a model that performs a translation function into a second language (KR). The second language (KR) may represent Korean.
[0231] For example, the first neural network model (45-1) may be a model that performs a translation function from a first language (EN) to a second language (KR).
[0232] Through the first neural network model (45-1), the electronic device (100) can obtain a first translation text (Hello, my name is Chulsoo. I develop.) in a second language (KR) based on a first user speech (Hello, my name is Chulsoo. I develop.) in a first language (EN).
[0233] Referring to the embodiment (620) of FIG. 6, the electronic device (100) can perform a keyword extraction function using the second neural network model (45-2).
[0234] The electronic device (100) can transmit first audio data (521) to a second neural network model (45-2) via an audio encoder (43). The audio encoder (43) can transmit first audio feature data to the second neural network model (45-2).
[0235] The electronic device (100) can transmit temporary keyword text (611) to the second neural network model (45-2) via the text encoder (44). The text encoder (44) can transmit temporary keyword feature data to the second neural network model (45-2).
[0236] The second neural network model (45-2) can obtain the first keyword text (622) based on the first audio data (521) and the temporary keyword text (611). Since the temporary keyword text (611) does not contain a keyword, the second neural network model (45-2) can substantially perform the keyword extraction function using only the first audio data (521).
[0237] For example, the second neural network model (45-2) may be a model that performs a keyword extraction function in a second language (KR).
[0238] For example, the second neural network model (45-2) may be a model that performs a keyword extraction function from a first language (EN) to a second language (KR).
[0239] Through the second neural network model (45-2), the electronic device (100) can obtain the first keyword text (hello, name, Chulsoo, development) of the second language (KR) based on the first user voice (Hello, my name is Chulsoo. I develop) of the first language (EN).
[0240] The electronic device (100) can perform a translation function for the second audio data (522) using the first keyword text (622).
[0241] FIG. 7 is a drawing for explaining an operation of analyzing a second user voice in the embodiment of FIG. 5, according to one embodiment.
[0242] Referring to the embodiment (710) of FIG. 7, the electronic device (100) can perform a translation function using the first neural network model (45-1).
[0243] The electronic device (100) can transmit second audio data (522) to the first neural network model (45-1) via the audio encoder (43).
[0244] Second audio data (522) can be converted into feature data via an audio encoder (43). The audio encoder (43) can generate second audio feature data corresponding to the second audio data (522). The audio encoder (43) can transmit the second audio feature data to a first neural network model (45-1).
[0245] The electronic device (100) can transmit the first keyword text (622) to the first neural network model (45-1) via the text encoder (44). The first keyword text (622) can include text representing at least one keyword (hello, name, withdrawal, development).
[0246] The first keyword text (622) can be converted into feature data through a text encoder (44). The text encoder (44) can generate feature data corresponding to the first keyword text (622). The text encoder (44) can transmit the first keyword feature data to the first neural network model (45-1).
[0247] The first neural network model (45-1) can obtain a second translation text (712) based on the second audio data (522) and the first keyword text (622).
[0248] For example, the first neural network model (45-1) may be a model that performs a translation function into a second language (KR).
[0249] For example, the first neural network model (45-1) may be a model that performs a translation function from a first language (EN) to a second language (KR).
[0250] Through the first neural network model (45-1), the electronic device (100) can obtain a second translation text (the developed prior research will further develop the world and society.) in a second language (KR) based on the first user's speech (next generation technology to make a better world and society.) in a first language (EN).
[0251] Referring to the embodiment (720) of FIG. 7, the electronic device (100) can perform a keyword extraction function using the second neural network model (45-2).
[0252] The electronic device (100) can transmit second audio data (522) to a second neural network model (45-2) via an audio encoder (43). The audio encoder (43) can transmit second audio feature data to the second neural network model (45-2).
[0253] The electronic device (100) can transmit the first keyword text (622) to the second neural network model (45-2) via the text encoder (44). The text encoder (44) can transmit the first keyword feature data to the second neural network model (45-2).
[0254] The second neural network model (45-2) can obtain second keyword text (722) based on second audio data (522) and first keyword text (622).
[0255] For example, the second neural network model (45-2) may be a model that performs a keyword extraction function in a second language (KR).
[0256] For example, the second neural network model (45-2) may be a model that performs a keyword extraction function from a first language (EN) to a second language (KR).
[0257] Through the second neural network model (45-2), the electronic device (100) can obtain a second keyword text (previous research, world, social development) in a second language (KR) based on a first user voice (next generation technology to make a better world and society.) in a first language (EN).
[0258] The electronic device (100) can obtain a first translation text (612) through the embodiment (610) of FIG. 6, and can obtain a second translation text (712) through the embodiment (710) of FIG. 7.
[0259] FIG. 8 is a diagram for explaining an operation of providing a translation text corresponding to an audio signal, according to one embodiment.
[0260] Referring to FIG. 8, the electronic device (100) can obtain an audio signal (S810). The audio signal can include a user's voice.
[0261] The electronic device (100) can segment an audio signal (S820). The electronic device (100) can segment an audio signal into at least one audio data. The segmentation criterion can be determined based on preset unit information. A description related to this is provided in FIG. 4.
[0262] The electronic device (100) can obtain a first translation text based on the segmented audio signal (S830). The first translation text can include a translation result corresponding to a specific audio signal among the segmented audio signals.
[0263] The electronic device (100) can obtain a first keyword text based on the segmented audio signal (S840).
[0264] The electronic device (100) can obtain a second translation text corresponding to a subsequent audio signal based on the first keyword text (S850). The subsequent audio signal may be an audio signal obtained after a specific audio signal among the segmented audio signals.
[0265] The electronic device (100) can provide a translation text (S860). The electronic device (100) can provide the translation text based on the first translation text and the second translation text (S860). The electronic device (100) can generate a translation result screen including the first translation text and the second translation text. The electronic device (100) can provide the translation result screen to the user.
[0266] FIG. 9 is a diagram for explaining an operation of providing a translation text using keyword text according to one embodiment.
[0267] Steps S910, S920, and S960 of FIG. 9 may correspond to steps S810, S820, and S860 of FIG. 8. Duplicate descriptions are omitted.
[0268] After segmenting the audio signal, the electronic device (100) can obtain a first translation text corresponding to the first user's speech (S930). The first translation text may represent the result of translating the first user's speech in a first language into a second language.
[0269] The electronic device (100) can obtain a first keyword text corresponding to the first user's voice (S940). The first keyword text can represent the result of extracting keywords in a second language from the first user's voice in a first language.
[0270] The electronic device (100) can obtain a second translation text corresponding to the second user's voice based on the first keyword text (S950). The electronic device (100) can obtain a second translation text corresponding to the second user's voice using the first keyword text obtained based on the first user's voice.
[0271] The electronic device (100) can provide translated text (S960).
[0272] FIG. 10 is a diagram for explaining an operation of providing a translated text using multiple neural network models according to one embodiment.
[0273] Steps S1010 and S1060 of Fig. 10 may correspond to steps S810 and S860 of Fig. 8. Duplicate explanation is omitted.
[0274] After acquiring an audio signal, the electronic device (100) can divide the audio signal into first audio data including a first user voice and second audio data including a second user voice (S1020). Each of the divided audio data may include a different user voice. The electronic device (100) can acquire a plurality of audio data including each user voice.
[0275] The electronic device (100) can input first audio data into a first neural network model (45-1) to obtain a first translation text (S1030).
[0276] The electronic device (100) can input first audio data into a second neural network model (45-2) to obtain a first keyword text (S1040).
[0277] The electronic device (100) can input second audio data and first keyword text into the first neural network model (45-1) to obtain a second translation text (S1050).
[0278] The electronic device (100) can provide translated text (S1060).
[0279] FIG. 11 is a diagram illustrating an operation for obtaining additional audio data according to one embodiment.
[0280] Steps S1110, S1120, S1130, S1140, S1150, and S1160 of FIG. 11 may correspond to steps S1010, S1020, S1030, S1040, S1050, and S1060 of FIG. 10. Duplicate explanations are omitted.
[0281] After obtaining the second translation text, the electronic device (100) can identify whether additional audio data is received (S1151).
[0282] If no additional audio data is received (S1151-N), the electronic device (100) can provide the first translated text and the second translated text (S1160).
[0283] When additional audio data is received (S1151-Y), the electronic device (100) can input the second audio data and the first keyword text into the second neural network model (45-2) to obtain the second keyword text (S1152).
[0284] The electronic device (100) may acquire third audio data (S1153). The third audio data may include a third user voice received subsequent to the second user voice included in the second audio data. Based on time order, the first user voice, the second user voice, and the third user voice may be received sequentially.
[0285] The electronic device (100) can input third audio data and second keyword text into the first neural network model (45-1) to obtain a third translation text (S1154). The third translation text may include the result of translating the third user's speech into a second language.
[0286] The electronic device (100) can provide a first translation text, a second translation text, and a third translation text (S1160).
[0287] In various embodiments, operation S1152 may be performed before operation S1151.
[0288] The first neural network model (45-1) and the second neural network model (45-2) may each be models generated according to a learning operation. Below, the learning process for each of the first neural network model (45-1) and the second neural network model (45-2) is described. For convenience of explanation, the artificial intelligence model including the first neural network model (45-1) and the second neural network model (45-2) is described as the subject of the operation.
[0289] FIG. 12 is a diagram for explaining learning data according to one embodiment.
[0290] Referring to embodiment 1210 of FIG. 12, an artificial intelligence model can be trained based on an audio signal including user voices with relatively low correlation.
[0291] Referring to the embodiment (1220) of FIG. 12, the audio signal can be divided into first training audio data (1221) including a first user voice ("Who is going to be Mr. President of the United States") and second training audio data (1222) including a second user voice ("Java is a programming language that is widely used"). The first user voice and the second user voice may be voices with relatively low correlation. The artificial intelligence model can be trained to output accurate translation text despite the low correlation. The second user voice may be a voice spoken (or received) after the first user voice.
[0292] Referring to embodiment 1230 of FIG. 12, an artificial intelligence model may be trained based on a learning data set. The learning data set may include multiple learning data.
[0293] The training data set may include second training audio data (1222), training translation text (1231), first training keyword text (1232), and second training keyword text (1233).
[0294] The second learning audio data (1222) may include a second user's speech in the first language (EN) (Java is a programming language that is widely used.).
[0295] The learning translation text (1231) may be a text translated from the second user's speech in the first language (EN) (Java is a programming language that is widely used.) to the second language (KR) (Java is a widely used programming language.).
[0296] The first learning keyword text (1232) may include a keyword (President of the United States) in a second language (KR) extracted from a first user speech (Who is going to be Mr. President of the United States) in a first language (EN).
[0297] The second learning keyword text (1233) may include a second user speech in the first language (EN) and a keyword (Java, a programming language) in the second language (KR).
[0298] FIG. 13 is a diagram for explaining an operation of learning a first neural network model and a second neural network model using learning data of the embodiment of FIG. 12, according to one embodiment.
[0299] Referring to the embodiment (1310) of FIG. 13, the first neural network model (45-1) can be trained based on second learning audio data (1222), first learning keyword text (1232), and learning translation text (1231).
[0300] Referring to the embodiment (1310) of FIG. 13, the artificial intelligence model can perform a translation function using the first neural network model (45-1).
[0301] The artificial intelligence model can transmit the second learning audio data (1222) to the first neural network model (45-1) via the audio encoder (43).
[0302] The second learning audio data (1222) can be converted into feature data via an audio encoder (43). The audio encoder (43) can generate second audio feature data corresponding to the second learning audio data (1222). The audio encoder (43) can transmit the second audio feature data to the first neural network model (45-1).
[0303] The artificial intelligence model can transmit the first learning keyword text (1232) to the first neural network model (45-1) via the text encoder (44). The first learning keyword text (1232) can include text representing at least one keyword (US President).
[0304] The first learning keyword text (1232) can be converted into feature data through a text encoder (44). The text encoder (44) can generate feature data corresponding to the first learning keyword text (1232). The text encoder (44) can transmit the first learning keyword feature data to the first neural network model (45-1).
[0305] The first neural network model (45-1) can obtain a translation text (1312, Java is a programming language used by the President of the United States) based on the second learning audio data (1222) and the first learning keyword text (1232).
[0306] For example, the first neural network model (45-1) may be a model that performs a translation function into a second language (KR).
[0307] For example, the first neural network model (45-1) may be a model that performs a translation function from a first language (EN) to a second language (KR).
[0308] The first neural network model (45-1) can be trained so that the output data becomes the training translation text (1231, Java is a widely used programming language).
[0309] When the second learning audio data (1222) and the first learning keyword text (1232) are input to the first neural network model (45-1), the artificial intelligence model can be trained so that the first neural network model (45-1) outputs the learning translation text (1231).
[0310] The artificial intelligence model can change (or update) at least one parameter used in the first neural network model (45-1) based on the difference between the translated text (1312) and the learning translated text (1231). The artificial intelligence model can identify the optimal value of the parameter used in the first neural network model (45-1) through the learning process.
[0311] The artificial intelligence model can change (or update) at least one parameter used in at least one of the audio encoder (43) or the text encoder (44) based on the difference between the translated text (1312) and the learning translated text (1231). The artificial intelligence model can identify an optimal value of at least one parameter used in at least one of the audio encoder (43) or the text encoder (44) through a learning process.
[0312] Referring to the embodiment (1320) of Fig. 13, the artificial intelligence model can perform a keyword extraction function using the second neural network model (45-2).
[0313] The artificial intelligence model can transmit second learning audio data (1222) to the second neural network model (45-2) via the audio encoder (43). The audio encoder (43) can transmit second audio feature data to the second neural network model (45-2).
[0314] The artificial intelligence model can transmit the first learning keyword text (1232) to the second neural network model (45-2) via the text encoder (44). The text encoder (44) can transmit the first learning keyword feature data to the second neural network model (45-2).
[0315] The second neural network model (45-2) can obtain second keyword text (1322, Java, popularity) based on second learning audio data (1222) and first learning keyword text (1232).
[0316] For example, the second neural network model (45-2) may be a model that performs a keyword extraction function in a second language (KR).
[0317] For example, the second neural network model (45-2) may be a model that performs a keyword extraction function from a first language (EN) to a second language (KR).
[0318] The second neural network model (45-2) can be trained so that the output data becomes the second learning keyword text (1233).
[0319] When the second learning audio data (1222) and the first learning keyword text (1232) are input to the second neural network model (45-2), the artificial intelligence model can be trained so that the second neural network model (45-2) outputs the second learning keyword text (1233).
[0320] The artificial intelligence model can change (or update) at least one parameter used in the second neural network model (45-2) based on the difference between the keyword text (1322) and the second learning keyword text (1233). The artificial intelligence model can identify the optimal value of the parameter used in the second neural network model (45-2) through the learning process.
[0321] The artificial intelligence model can change (or update) at least one parameter used in at least one of the audio encoder (43) or the text encoder (44) based on the difference value between the keyword text (1322) and the second learning keyword text (1233). The artificial intelligence model can identify an optimal value of at least one parameter used in at least one of the audio encoder (43) or the text encoder (44) through a learning process.
[0322] FIG. 14 is a diagram for explaining learning data according to one embodiment.
[0323] Referring to the embodiment (1410) of FIG. 14, an artificial intelligence model can be trained based on an audio signal including user voices with relatively high correlation.
[0324] Referring to an embodiment (1420) of FIG. 14, the audio signal may be divided into first training audio data (1421) including a first user voice ("Who is going to be Mr. President of the United States") and second training audio data (1422) including a second user voice ("We are not sure yet. President Joe Biden probably will go for the next presidential election."). The first user voice and the second user voice may be voices with relatively high correlation. The second user voice may be a voice spoken (or received) after the first user voice.
[0325] Referring to embodiment 1430 of FIG. 14, an artificial intelligence model may be trained based on a learning data set. The learning data set may include multiple learning data.
[0326] The training data set may include second training audio data (1422), training translation text (1431), first training keyword text (1432), and second training keyword text (1433).
[0327] The second training audio data (1422) may include a second user speech in the first language (EN) (We are not sure yet. President Joe Biden probably will go for the next presidential election.).
[0328] The training translation text (1431) could be a translation of the second language (KR) text (We are not sure yet. President Joe Biden probably will go for the next presidential election.) from the first language (EN) to the second language (KR).
[0329] The first learning keyword text (1432) may include a keyword (President of the United States) in a second language (KR) extracted from a first user speech (Who is going to be Mr. President of the United States) in a first language (EN).
[0330] The second learning keyword text (1433) may include a second user's speech in the first language (EN) with a keyword in the second language (KR) (President Joe Biden, possibility of running for president).
[0331] FIG. 15 is a diagram for explaining an operation of learning a first neural network model and a second neural network model using learning data of the embodiment of FIG. 14, according to one embodiment.
[0332] Referring to the embodiment (1510) of FIG. 15, the first neural network model (45-1) can be trained based on second learning audio data (1422), first learning keyword text (1432), and learning translation text (1431).
[0333] Referring to the embodiment (1510) of FIG. 15, the artificial intelligence model can perform a translation function using the first neural network model (45-1).
[0334] The artificial intelligence model can transmit the second learning audio data (1422) to the first neural network model (45-1) via the audio encoder (43).
[0335] The second learning audio data (1422) can be converted into feature data via an audio encoder (43). The audio encoder (43) can generate second audio feature data corresponding to the second learning audio data (1422). The audio encoder (43) can transmit the second audio feature data to the first neural network model (45-1).
[0336] The artificial intelligence model can transmit the first learning keyword text (1432) to the first neural network model (45-1) via the text encoder (44). The first learning keyword text (1432) can include text representing at least one keyword (US President).
[0337] The first learning keyword text (1432) can be converted into feature data through a text encoder (44). The text encoder (44) can generate feature data corresponding to the first learning keyword text (1432). The text encoder (44) can transmit the first learning keyword feature data to the first neural network model (45-1).
[0338] The first neural network model (45-1) can obtain a translated text (1512, not yet certain. Joe Biden is likely to run for president) based on the second learning audio data (1422) and the first learning keyword text (1432).
[0339] For example, the first neural network model (45-1) may be a model that performs a translation function into a second language (KR).
[0340] For example, the first neural network model (45-1) may be a model that performs a translation function from a first language (EN) to a second language (KR).
[0341] The first neural network model (45-1) can be trained so that the output data is the training translation text (1431, not yet certain. It is likely that President Joe Biden will run for the next presidential election).
[0342] When the second learning audio data (1422) and the first learning keyword text (1432) are input to the first neural network model (45-1), the artificial intelligence model can be trained so that the first neural network model (45-1) outputs the learning translation text (1431).
[0343] The artificial intelligence model can change (or update) at least one parameter used in the first neural network model (45-1) based on the difference between the translated text (1512) and the learning translated text (1431). The artificial intelligence model can identify the optimal value of the parameter used in the first neural network model (45-1) through the learning process.
[0344] The artificial intelligence model can change (or update) at least one parameter used in at least one of the audio encoder (43) or the text encoder (44) based on the difference between the translated text (1512) and the learning translated text (1431). The artificial intelligence model can identify an optimal value of at least one parameter used in at least one of the audio encoder (43) or the text encoder (44) through a learning process.
[0345] Referring to the embodiment (1520) of FIG. 15, the artificial intelligence model can perform a keyword extraction function using the second neural network model (45-2).
[0346] The artificial intelligence model can transmit second learning audio data (1422) to the second neural network model (45-2) via the audio encoder (43). The audio encoder (43) can transmit second audio feature data to the second neural network model (45-2).
[0347] The artificial intelligence model can transmit the first learning keyword text (1432) to the second neural network model (45-2) via the text encoder (44). The text encoder (44) can transmit the first learning keyword feature data to the second neural network model (45-2).
[0348] The second neural network model (45-2) can obtain second keyword text (1522, next presidential candidate, President Joe Biden, uncertain) based on second learning audio data (1422) and first learning keyword text (1432).
[0349] For example, the second neural network model (45-2) may be a model that performs a keyword extraction function in a second language (KR).
[0350] For example, the second neural network model (45-2) may be a model that performs a keyword extraction function from a first language (EN) to a second language (KR).
[0351] The second neural network model (45-2) can be trained so that the output data becomes the second learning keyword text (1433).
[0352] When the second learning audio data (1422) and the first learning keyword text (1432) are input to the second neural network model (45-2), the artificial intelligence model can be trained so that the second neural network model (45-2) outputs the second learning keyword text (1433).
[0353] The artificial intelligence model can change (or update) at least one parameter used in the second neural network model (45-2) based on the difference between the keyword text (1522) and the second learning keyword text (1433). The artificial intelligence model can identify the optimal value of the parameter used in the second neural network model (45-2) through the learning process.
[0354] The artificial intelligence model can change (or update) at least one parameter used in at least one of the audio encoder (43) or the text encoder (44) based on the difference between the keyword text (1522) and the second learning keyword text (1433). The artificial intelligence model can identify an optimal value of at least one parameter used in at least one of the audio encoder (43) or the text encoder (44) through a learning process.
[0355] In FIGS. 12 to 15, an example is described in which the first neural network model (45-1) and the second neural network model (45-2) are trained using one training data set.
[0356] According to various embodiments, the first neural network model (45-1) can be trained independently. This is described in FIGS. 16 to 19.
[0357] FIG. 16 is a diagram for explaining learning data according to one embodiment.
[0358] Referring to embodiment 1610 of FIG. 16, an artificial intelligence model can be trained based on an audio signal including a user's voice.
[0359] Referring to the embodiment (1620) of FIG. 16, the audio signal can be divided into first learning audio data (1621) including a first user voice (I develop the software) and second learning audio data (1622) including a second user voice (to make a better world). The second user voice can be a voice spoken (or received) after the first user voice.
[0360] Referring to embodiment 1630 of FIG. 16, an artificial intelligence model may be trained based on a training data set. The training data set may include multiple training data.
[0361] The training data set may include second training audio data (1622), training translation text (1631), and first training keyword text (1632).
[0362] The second learning audio data (1622) may include a second user voice (to make a better world) in the first language (EN).
[0363] The learning translation text (1631) may be a text translated from the first user voice (I develop the software) and the second user voice (to make a better world) in the first language (EN) to the second language (KR) (I develop the software to make a better world).
[0364] The first learning keyword text (1632) may include keywords (development, software) in a second language (KR) extracted from the first user speech (I develop the software) in the first language (EN).
[0365] FIG. 17 is a diagram for explaining an operation of learning a first neural network model with learning data of the embodiment of FIG. 16, according to one embodiment.
[0366] Referring to the embodiment (1710) of FIG. 17, the first neural network model (45-1) can be trained based on second learning audio data (1622), first learning keyword text (1632), and learning translation text (1631).
[0367] Referring to the embodiment (1710) of FIG. 17, the artificial intelligence model can perform a translation function using the first neural network model (45-1).
[0368] The artificial intelligence model can transmit the second learning audio data (1622) to the first neural network model (45-1) via the audio encoder (43).
[0369] The second learning audio data (1622) can be converted into feature data via an audio encoder (43). The audio encoder (43) can generate second audio feature data corresponding to the second learning audio data (1622). The audio encoder (43) can transmit the second audio feature data to the first neural network model (45-1).
[0370] The artificial intelligence model can transmit the first learning keyword text (1632) to the first neural network model (45-1) via the text encoder (44). The first learning keyword text (1632) can include text representing at least one keyword (development, software).
[0371] The first learning keyword text (1632) can be converted into feature data through a text encoder (44). The text encoder (44) can generate feature data corresponding to the first learning keyword text (1632). The text encoder (44) can transmit the first learning keyword feature data to the first neural network model (45-1).
[0372] The first neural network model (45-1) can obtain a translation text (1712, Developing software for a better world) based on the second learning audio data (1622) and the first learning keyword text (1632).
[0373] For example, the first neural network model (45-1) may be a model that performs a translation function into a second language (KR).
[0374] For example, the first neural network model (45-1) may be a model that performs a translation function from a first language (EN) to a second language (KR).
[0375] The first neural network model (45-1) can be trained so that the output data becomes the learning translation text (1631, I develop software to make a better world.).
[0376] When the second learning audio data (1622) and the first learning keyword text (1632) are input to the first neural network model (45-1), the artificial intelligence model can be trained so that the first neural network model (45-1) outputs the learning translation text (1631).
[0377] The artificial intelligence model can change (or update) at least one parameter used in the first neural network model (45-1) based on the difference between the translated text (1712) and the learning translated text (1631). The artificial intelligence model can identify the optimal value of the parameter used in the first neural network model (45-1) through the learning process.
[0378] The artificial intelligence model can change (or update) at least one parameter used in at least one of the audio encoder (43) or the text encoder (44) based on the difference between the translated text (1712) and the learning translated text (1631). The artificial intelligence model can identify an optimal value of at least one parameter used in at least one of the audio encoder (43) or the text encoder (44) through a learning process.
[0379] FIG. 18 is a diagram for explaining learning data according to one embodiment.
[0380] Referring to embodiment 1810 of FIG. 18, an artificial intelligence model can be trained based on an audio signal including a user's voice.
[0381] Referring to the embodiment (1820) of FIG. 18, the audio signal can be divided into first learning audio data (1821) including a first user voice (I develop the software) and second learning audio data (1822) including a second user voice (to make a better world). The second user voice can be a voice spoken (or received) after the first user voice.
[0382] Referring to embodiment 1830 of FIG. 18, an artificial intelligence model may be trained based on a training data set. The training data set may include multiple training data.
[0383] The training data set may include second training audio data (1822), training translation text (1831), and first training keyword text (1832).
[0384] The second learning audio data (1822) may include a second user voice (to make a better world) in the first language (EN).
[0385] The learning translation text (1831) may be a text translated from the first user voice (I develop the software) and the second user voice (to make a better world) in the first language (EN) to the second language (KR) (I develop the software to make a better world).
[0386] The first learning keyword text (1832) may include keywords (development, software) and noise keywords (cat, shrimp) in the second language (KR) extracted from the first user speech ("I develop the software") in the first language (EN). The noise keywords may represent learning keywords that are input to provide a consistent quality of translation results even when keywords with low relevance to the translation result are included.
[0387] FIG. 19 is a diagram for explaining an operation of learning a first neural network model with learning data of the embodiment of FIG. 18, according to one embodiment.
[0388] Referring to the embodiment (1910) of FIG. 19, the first neural network model (45-1) can be trained based on second learning audio data (1822), first learning keyword text (1832), and learning translation text (1831).
[0389] Referring to the embodiment (1910) of FIG. 19, the artificial intelligence model can perform a translation function using the first neural network model (45-1).
[0390] The artificial intelligence model can transmit the second learning audio data (1822) to the first neural network model (45-1) via the audio encoder (43).
[0391] The second learning audio data (1822) can be converted into feature data via an audio encoder (43). The audio encoder (43) can generate second audio feature data corresponding to the second learning audio data (1822). The audio encoder (43) can transmit the second audio feature data to the first neural network model (45-1).
[0392] The artificial intelligence model can transmit the first learning keyword text (1832) to the first neural network model (45-1) via the text encoder (44). The first learning keyword text (1832) can include text representing at least one keyword (development, software, cat, shrimp).
[0393] The first learning keyword text (1832) can be converted into feature data through a text encoder (44). The text encoder (44) can generate feature data corresponding to the first learning keyword text (1832). The text encoder (44) can transmit the first learning keyword feature data to the first neural network model (45-1).
[0394] The first neural network model (45-1) can obtain a translation text (1912, Developing software for a better world for cats) based on the second learning audio data (1822) and the first learning keyword text (1832).
[0395] For example, the first neural network model (45-1) may be a model that performs a translation function into a second language (KR).
[0396] For example, the first neural network model (45-1) may be a model that performs a translation function from a first language (EN) to a second language (KR).
[0397] The first neural network model (45-1) can be trained so that the output data becomes the learning translation text (1831, I develop software to make the world a better place.).
[0398] When the second learning audio data (1822) and the first learning keyword text (1832) are input to the first neural network model (45-1), the artificial intelligence model can be trained so that the first neural network model (45-1) outputs the learning translation text (1831).
[0399] The artificial intelligence model can change (or update) at least one parameter used in the first neural network model (45-1) based on the difference between the translated text (1912) and the training translated text (1831). The artificial intelligence model can identify the optimal value of the parameter used in the first neural network model (45-1) through the learning process.
[0400] The artificial intelligence model can change (or update) at least one parameter used in at least one of the audio encoder (43) or the text encoder (44) based on the difference between the translated text (1912) and the learning translated text (1831). The artificial intelligence model can identify an optimal value of at least one parameter used in at least one of the audio encoder (43) or the text encoder (44) through a learning process.
[0401] Even if the first learning keyword text (1832) includes noise keywords (cat, shrimp), the first neural network model (45-1) can be trained to output the learning translation text (1831) regardless of the noise keywords.
[0402] FIG. 20 is a drawing for explaining a method for controlling an electronic device according to one embodiment.
[0403] Referring to FIG. 20, a method for controlling an electronic device includes a step (S2005) of, when a first user voice in a first language is received, inputting the first user voice into a first neural network model to obtain a first translation text in a second language corresponding to the first user voice, a step (S2010) of inputting the first user voice into a second neural network model to obtain a first keyword text in a second language corresponding to the first user voice, a step (S2015) of, when a second user voice in the first language is received, inputting the second user voice and the first keyword text into the first neural network model to obtain a second translation text in the second language corresponding to the second user voice, and a step (S2020) of providing the first translation text and the second translation text.
[0404] The control method further includes a step of identifying a first language based on a first user voice, and the step of obtaining a first translation text (S2005) obtains a first translation text in a second language, which is a preset target language, based on the first user voice in the first language through a first neural network model, and the step of obtaining a second translation text (S2015) obtains a second translation text in the second language based on the second user voice in the first language and the first keyword text through the first neural network model.
[0405] The control method further includes a step of obtaining an audio signal including a first user voice and a second user voice, and a step of dividing the audio signal into first audio data including the first user voice and second audio data including the second user voice based on a preset unit time, wherein the step of obtaining a first translation text (S2005) inputs the first audio data into a first neural network model to obtain a first translation text, the step of obtaining a first keyword text (S2010) inputs the first audio data into a second neural network model to obtain a first keyword text, and the step of obtaining a second translation text (S2015) inputs the second audio data and the first keyword text into the first neural network model to obtain a second translation text.
[0406] The second user voice may be a voice spoken following the first user voice.
[0407] The control method may further include a step of obtaining a second keyword text corresponding to a second user voice based on the second audio data and the first keyword text, a step of obtaining third audio data including a third user voice in a first language spoken following the second user voice, a step of inputting the third audio data and the second keyword text into a first neural network model to obtain a third translated text corresponding to the third user voice, and a step of providing the first translated text, the second translated text, and the third translated text.
[0408] The electronic device stores an artificial intelligence model including a first neural network model and a second neural network model, and the step of obtaining a first translated text and a second translated text (S2005, 2015) can obtain the first translated text and the second translated text by inputting a first user voice and a second user voice into the artificial intelligence model.
[0409] The artificial intelligence model may be a model trained based on first training audio data including a first learning user voice in a first language, second training audio data including a second learning user voice in the first language, first learning keyword text in the second language corresponding to the first learning user voice, learning translation text in the second language corresponding to the second learning user voice, and second learning keyword text in the second language corresponding to the second learning user voice.
[0410] The artificial intelligence model may be a model trained to receive second learning audio data and first learning keyword text and output a learning translation text by performing a function of translating into a second language.
[0411] The artificial intelligence model may be a model trained to receive second learning audio data, first learning keyword text, and noise keyword text and output a learning translation text by performing a function of translating into a second language.
[0412] The artificial intelligence model may be a model trained to receive second learning audio data and first learning keyword text and output second learning keyword text by performing a function of extracting keywords in a second language.
[0413] The methods according to the various embodiments of the present disclosure described above can be implemented in the form of an application that can be installed on an existing electronic device.
[0414] The methods according to the various embodiments of the present disclosure described above can be implemented only with a software upgrade or a hardware upgrade for an existing electronic device.
[0415] The various embodiments of the present disclosure described above may also be performed through an embedded server provided in an electronic device, or an external server of at least one of the electronic device and the display device.
[0416] According to an example embodiment of the present disclosure, the various embodiments described above may be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device may include an electronic device according to the disclosed embodiments, which is a device that can call instructions stored in the storage medium and operate according to the called instructions. When the instructions are executed by a processor, the processor may directly or under the control of the processor use other components to perform a function corresponding to the instructions. The instructions may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' means that the storage medium does not contain signals and is tangible, but does not distinguish between data being stored semi-permanently or temporarily in the storage medium.
[0417] According to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as included in a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0418] Each of the components (e.g., modules or programs) according to the various embodiments described above may be composed of a single or multiple entities, and some of the sub-components described above may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the respective components prior to integration. Operations performed by modules, programs or other components according to various embodiments may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
[0419] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea of the present disclosure.
Claims
1. In electronic devices, memory; a processor connected to the above memory; The above processor, When a first user speech in a first language is received, the first user speech is input into a first neural network model to obtain a first translation text in a second language corresponding to the first user speech, By inputting the first user voice into a second neural network model, a first keyword text in a second language corresponding to the first user voice is obtained, When a second user voice in the first language is received, the second user voice and the first keyword text are input into the first neural network model to obtain a second translation text in the second language corresponding to the second user voice, An electronic device providing the first translated text and the second translated text.
2. In paragraph 1, The above processor, Identifying the first language based on the first user voice, Through the first neural network model, the first translation text in the second language, which is a preset target language, is obtained based on the first user voice in the first language, An electronic device that obtains the second translated text in the second language based on the second user voice and the first keyword text in the first language through the first neural network model.
3. In paragraph 1, The above processor, Obtaining an audio signal including the first user voice and the second user voice, Splitting the audio signal into first audio data including the first user voice and second audio data including the second user voice based on a preset unit time, Inputting the first audio data into the first neural network model to obtain the first translated text, Inputting the first audio data into the second neural network model to obtain the first keyword text, An electronic device that inputs the second audio data and the first keyword text into the first neural network model to obtain the second translated text.
4. In paragraph 1, The second user voice above is, An electronic device, wherein the voice is spoken following the first user voice.
5. In paragraph 1, The above processor, Based on the second audio data and the first keyword text, a second keyword text corresponding to the second user voice is obtained, Obtain third audio data including a third user voice of the first language spoken following the second user voice, By inputting the third audio data and the second keyword text into the first neural network model, a third translation text corresponding to the third user voice is obtained, An electronic device providing the first translated text, the second translated text, and the third translated text.
6. In paragraph 1, The above memory is, Store an artificial intelligence model including the first neural network model and the second neural network model, The above processor, An electronic device that obtains the first translated text and the second translated text by inputting the first user voice and the second user voice into the artificial intelligence model.
7. In paragraph 6, The above artificial intelligence model, An electronic device, wherein the model is learned based on first learning audio data including a first learning user voice in the first language, second learning audio data including a second learning user voice in the first language, first learning keyword text in the second language corresponding to the first learning user voice, learning translation text in the second language corresponding to the second learning user voice, and second learning keyword text in the second language corresponding to the second learning user voice.
8. In paragraph 7, The above artificial intelligence model, An electronic device, which is a model trained to receive the second learning audio data and the first learning keyword text and output the learning translation text by performing a function of translating into the second language.
9. In paragraph 7, The above artificial intelligence model, An electronic device, which is a model trained to receive the second learning audio data, the first learning keyword text, and the noise keyword text and output the learning translation text by performing a function of translating into the second language.
10. In paragraph 7, The above artificial intelligence model, An electronic device, wherein the model is trained to receive the second learning audio data and the first learning keyword text and output the second learning keyword text by performing a function of extracting keywords in the second language.
11. In a method for controlling an electronic device, When a first user voice in a first language is received, a step of inputting the first user voice into a first neural network model to obtain a first translation text in a second language corresponding to the first user voice; A step of inputting the first user voice into a second neural network model to obtain a first keyword text in a second language corresponding to the first user voice; When a second user voice in the first language is received, a step of inputting the second user voice and the first keyword text into the first neural network model to obtain a second translation text in the second language corresponding to the second user voice; and A control method comprising: providing the first translated text and the second translated text.
12. In paragraph 11, The above control method is, further comprising a step of identifying the first language based on the first user voice; The step of obtaining the above first translation text is: Through the first neural network model, the first translation text in the second language, which is a preset target language, is obtained based on the first user voice in the first language, The step of obtaining the second translated text is as follows: A control method for obtaining the second translation text in the second language based on the second user voice and the first keyword text in the first language through the first neural network model.
13. In paragraph 11, The above control method is, A step of obtaining an audio signal including the first user voice and the second user voice; and Further comprising a step of dividing the audio signal into first audio data including the first user voice and second audio data including the second user voice based on a preset unit time; The step of obtaining the above first translation text is: Inputting the first audio data into the first neural network model to obtain the first translated text, The step of obtaining the first keyword text is: Inputting the first audio data into the second neural network model to obtain the first keyword text, The step of obtaining the second translated text is as follows: A control method for obtaining the second translated text by inputting the second audio data and the first keyword text into the first neural network model.
14. In paragraph 11, The second user voice above is, A control method, wherein the voice is spoken following the first user voice.
15. In paragraph 11, The above control method is, A step of obtaining a second keyword text corresponding to the second user voice based on the second audio data and the first keyword text; A step of obtaining third audio data including a third user voice of the first language spoken following the second user voice; A step of inputting the third audio data and the second keyword text into the first neural network model to obtain a third translation text corresponding to the third user voice; and A control method further comprising the step of providing the first translation text, the second translation text, and the third translation text.
Citation Information
Patent Citations
Method for unifying keyword translation
CN103678287A
Apparatus and method of interpreting an international conference based speech recognition
KR1020110066622A
Reducer for Steering Column of Vehicle
KR1020200118961A
Foreign matter shredder in a drainage ditch operated by a water level sensor
KR1020220161530A
Method For Manufacturing Gelatin Hydrogel And Gelatin Hydrogel Comprising Drug Which Is Adjustable Release Behavior
KR1020240033829A