Electronic device and control method
The electronic device and method improve translation accuracy by using an AI model to identify and confirm sentence-ending tokens, addressing error propagation and latency in conventional translation models.
Patent Information
- Application Number
- PCT/KR2025/095281
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-14
- Filing Date
- 2025-04-22
- Publication Date
- 2025-10-30
AI Technical Summary
Conventional translation models, both cascaded and end-to-end, suffer from error propagation and latency, with the end point of an utterance often not being accurately extracted, limiting translation performance.
An electronic device and method that utilize an artificial intelligence model to input voice signals, determine tokens within the recognized and translated texts, identify a token terminating a sentence, and re-input the signal if necessary to improve translation accuracy by confirming the token multiple times or within a preset period, using quality assessment models to ensure the final text meets a threshold score.
Enhances translation accuracy by precisely extracting the end point of sentences, reducing errors and latency in multilingual communication systems.
Smart Images

Figure KR2025095281_30102025_PF_FP_ABST
Abstract
Description
Electronic devices and control methods
[0001] The present disclosure relates to an electronic device and a control method, and more particularly, to an electronic device and a control method thereof that receive a user's voice signal, extract the end point of a sentence based on ASR (Automatic Speech Recognition), and translate the same.
[0002] To facilitate communication between users who speak multiple languages, the use of translation models that translate between different languages is increasing. In particular, interest in multilingual models that translate between two or more languages is also growing.
[0003] Conventional translation models utilize a cascaded approach. Because the cascaded approach involves two translation models, it suffers from error propagation and latency. Therefore, the need for an end-to-end model emerged to address translation errors and accuracy.
[0004] However, even when using an end-to-end model, there was a problem in that the translation performance was limited because the end point of the utterance could not be accurately extracted.
[0005] Accordingly, a method is requested to improve the accuracy of translation by receiving the user's voice signal and extracting the end point of the sentence based on Automatic Speech Recognition (ASR).
[0006] The aspects according to the present disclosure are intended to provide an electronic device and a control method thereof that can improve the accuracy of translation by solving at least the above-described problems.
[0007] Additional aspects will be disclosed in part in the following description, and in part will be obvious from the description or may be learned by practice of the embodiments presented.
[0008] According to one aspect of the present disclosure, an electronic device includes a communication interface, a memory storing instructions, and at least one processor, wherein the instructions, when collectively or individually executed by the at least one processor, cause the electronic device to input a first voice signal into an artificial intelligence model to obtain a first text in a source language obtained by voice-recognizing the first voice signal and a first text in the target language obtained by translating the first voice signal into a target language, input a second voice signal including the first voice signal into the artificial intelligence model to obtain a second text in the source language obtained by voice-recognizing the second voice signal and a second text in the target language obtained by translating the second voice signal into the target language, determine at least one token among a plurality of tokens included in the second text in the source language based on the first text in the source language and the second text in the source language, identify whether the determined at least one token includes a token terminating a sentence, and if the token terminating a sentence is included, re-input the second voice signal into the artificial intelligence model to obtain a final text in the target language.
[0009] The instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to determine an identified token when the same token is identified a preset number of times or more among a plurality of tokens included in a first text of the source language and a plurality of tokens included in a second text of the source language.
[0010] The instructions, when collectively or individually executed by the at least one processor, cause the electronic device to obtain a third speech signal including the second speech signal if the second text of the source language does not include a token terminating the sentence, input the third speech signal to the artificial intelligence model to obtain a third text of the source language obtained by speech-recognizing the third speech signal and a third text of the target language obtained by translating the third speech signal into a target language, determine at least one token among a plurality of tokens included in the third text of the source language based on the second text of the source language and the third text of the source language, and identify whether the determined at least one token includes a token terminating the sentence, wherein the third speech signal may be a speech signal including a speech signal received for a preset period of time after receiving the second speech signal.
[0011] The above instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to output the determined token among a plurality of tokens included in the second text of the source language and not output the remaining tokens.
[0012] The instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to output an identified token when the same token is identified a preset number of times or more among a plurality of tokens included in a first text of the target language and a plurality of tokens included in a second text of the target language.
[0013] The second voice signal may be a voice signal including the first voice signal and a voice signal received for a preset period of time after receiving the first voice signal.
[0014] The above artificial intelligence model may be characterized in that, when the first voice signal is input, it is learned to obtain in parallel a first text in a source language that recognizes the first voice signal and a first text in a target language that translates the first voice signal into the target language.
[0015] The instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to obtain a score representing the quality of the second text of the target language if the token terminating the sentence is included, and if the score is above a preset threshold value, obtain the second text of the target language as the final text of the target language, and if the score is below the threshold value, re-input the second speech signal into the artificial intelligence model to obtain the final text of the target language.
[0016] The instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to obtain the score based on at least one of a first score obtained by inputting a second text of the target language into a quality assessment model and a second score obtained by applying a predefined metric to the second text of the target language.
[0017] The electronic device further comprises a display, and the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to control the display to display a user interface including a second text in the target language, receive a user input through the user interface to evaluate a quality of the second text in the target language, and obtain the score based on the user input.
[0018] According to one aspect of the present disclosure, a control method of an electronic device includes the steps of: inputting a first voice signal into an artificial intelligence model to obtain a first text in a source language obtained by voice-recognizing the first voice signal and a first text in the target language obtained by translating the first voice signal into a target language; inputting a second voice signal including the first voice signal into the artificial intelligence model to obtain a second text in the source language obtained by voice-recognizing the second voice signal and a second text in the target language obtained by translating the second voice signal into the target language; determining at least one token among a plurality of tokens included in the second text of the source language based on the first text of the source language and the second text of the source language; identifying whether a token terminating a sentence is included in the determined at least one token; and if the token terminating a sentence is included, re-inputting the second voice signal into the artificial intelligence model to obtain a final text of the target language.
[0019] The method may include a step of confirming the identified token when the same token is identified a preset number of times or more among a plurality of tokens included in a first text of the source language and a plurality of tokens included in a second text of the source language.
[0020] If the second text of the source language does not include a token terminating the sentence, the method comprises: obtaining a third speech signal including the second speech signal; inputting the third speech signal into the artificial intelligence model to obtain a third text of the source language obtained by speech-recognizing the third speech signal and a third text of the target language obtained by translating the third speech signal into the target language; determining one or more tokens among a plurality of tokens included in the third text of the source language based on the second text of the source language and the third text of the source language; and identifying whether the determined one or more tokens include a token terminating the sentence, wherein the third speech signal may include a speech signal received for a preset period of time after receiving the second speech signal.
[0021] It may include a step of outputting the confirmed token among the plurality of tokens included in the second text and not outputting the remaining tokens.
[0022] The method may include a step of outputting the identified token when the same token is identified a preset number of times or more among a plurality of tokens included in a first text of the target language and a plurality of tokens included in a second text of the target language.
[0023] The second voice signal may be a voice signal including the first voice signal and a voice signal received for a preset period of time after receiving the first voice signal.
[0024] The above artificial intelligence model may be characterized in that, when the first voice signal is input, it is learned to obtain in parallel a first text in a source language that recognizes the first voice signal and a first text in a target language that translates the first voice signal into the target language.
[0025] If a token terminating the sentence is included, the method may further include a step of obtaining a score representing the quality of the second text of the target language, if the score is greater than or equal to a preset threshold value, a step of obtaining the second text of the target language as the final text of the target language, and if the score is less than the threshold value, a step of re-inputting the second speech signal into the artificial intelligence model to obtain the final text of the target language.
[0026] The step of obtaining the score may further include a step of obtaining the score based on at least one of a first score obtained by inputting the second text of the target language into a quality evaluation model and a second score obtained by applying a predefined metric to the second text of the target language.
[0027] According to one aspect of the present disclosure, a non-transitory computer-readable recording medium including a program for executing a method for controlling an electronic device may include a step of inputting a first voice signal into an artificial intelligence model to obtain a first text in a source language obtained by voice-recognizing the first voice signal and a first text in the target language obtained by translating the first voice signal into a target language; a step of inputting a second voice signal including the first voice signal into the artificial intelligence model to obtain a second text in the source language obtained by voice-recognizing the second voice signal and a second text in the target language obtained by translating the second voice signal into the target language; a step of determining at least one token among a plurality of tokens included in the second text of the source language based on the first text of the source language and the second text of the source language; a step of identifying whether a token terminating a sentence is included in the determined at least one token; and a step of re-inputting the second voice signal into the artificial intelligence model to obtain a final text of the target language if the token terminating a sentence is included.
[0028] FIG. 1 is a drawing for explaining an electronic device according to one embodiment of the present disclosure;
[0029] FIG. 2 is a block diagram illustrating a configuration of an electronic device according to at least one embodiment of the present disclosure;
[0030] FIGS. 3 to 7 are drawings for explaining a process of determining a token according to one embodiment of the present disclosure.
[0031] FIG. 8 is a drawing for explaining an artificial intelligence model according to one embodiment of the present disclosure.
[0032] FIG. 9 is a flowchart for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0033] FIG. 10 is a flowchart for explaining a control method of an electronic device according to one embodiment of the present disclosure;
[0034] FIG. 11 is a flowchart illustrating one or more embodiments related to quality assessment.
[0035] The present embodiments may be modified and have various embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope to specific embodiments, but should be understood to encompass various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0036] In describing the present disclosure, if it is determined that a specific description of a related known function or configuration may unnecessarily obscure the gist of the present disclosure, a detailed description thereof will be omitted.
[0037] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concepts of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to further faithfully and completely convey the technical concepts of the present disclosure to those skilled in the art.
[0038] The terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the scope of the rights. Singular expressions include plural expressions unless the context clearly dictates otherwise.
[0039] In this disclosure, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a corresponding feature (e.g., a component such as a number, function, operation, or part), and do not exclude the presence of additional features.
[0040] In this disclosure, expressions such as “A or B,” “at least one of A and / or B,” or “one or more of A or / and B” can include all possible combinations of the listed items. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” can all refer to (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
[0041] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0042] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that said component may be directly coupled to said other component, or may be coupled via another component (e.g., a third component).
[0043] On the other hand, when it is said that a component (e.g., a first component) is "directly connected" or "directly connected" to another component (e.g., a second component), it can be understood that no other component (e.g., a third component) exists between said component and said other component.
[0044] The expression "configured to" as used in the present disclosure may be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" may not necessarily mean only "specifically designed to" in terms of hardware.
[0045] Instead, in some contexts, the phrase "a device configured to" may mean that the device, in conjunction with other devices or components, is "capable of" performing A, B, and C. For example, the phrase "a processor configured (or set) to perform A, B, and C" may refer to a dedicated processor (e.g., an embedded processor) for performing those operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform those operations by executing one or more software programs stored in a memory device.
[0046] In the embodiments, a 'module' or 'part' performs at least one function or operation, and may be implemented as hardware or software, or as a combination of hardware and software. Furthermore, a plurality of 'modules' or 'parts' may be integrated into at least one module and implemented as at least one processor, except for a 'module' or 'part' that needs to be implemented as a specific hardware.
[0047] Meanwhile, the various elements and areas in the drawings are schematically drawn. Therefore, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0048] Hereinafter, various embodiments of the present invention will be described in detail using the attached drawings.
[0049] FIG. 1 is a drawing illustrating an electronic device (100) according to an embodiment of the present disclosure. In FIG. 1, the electronic device (100) is depicted as a smartphone, but this is merely an example, and the electronic device (100) may be implemented as a user terminal such as a tablet, a notebook PC, or the like. Furthermore, the electronic device (100) may of course be implemented as a server device.
[0050] For example, a user (30) may utter (30-1) the sentence "I went to the grocery store." while holding an electronic device (100) nearby. The electronic device (100) may input the sentence (30-1) uttered by the user (30) as a voice signal, and may obtain a text (100-1) in a source language obtained by recognizing the user's voice signal through ASR (Automatic Speech Recognition) technology and a text (100-2) obtained by translating the user's voice signal into a target language through ST (Speech Translation) technology.
[0051] Here, the source language refers to the language in which the voice is expressed before being translated, and the target language refers to the language into which the source language is translated.
[0052] Meanwhile, ASR (Automatic Speech Recognition) is an automatic speech recognition technology that converts a user's (30) voice signal into text in a source language.
[0053] For example, if a user (30) utters the voice “I went to the grocery store.”, the ASR system can obtain the text “I went to the grocery store.” (100-1) by recognizing the voice signal of the user (30).
[0054] ST (Speech Translation) refers to a technology that converts a user's (30) voice signal into text translated into a target language.
[0055] For example, a speech signal spoken in English by a user (30) can be converted into text translated into Korean. If the user (30) speaks "I went to the grocery store" (30-1) as shown in Fig. 1, the ST (Speech Translation) system can obtain the text "I went to the grocery store yesterday" (100-2) translated into Korean.
[0056] At this time, the electronic device (100) may include an artificial intelligence model that outputs voice-recognized text and text translated into a target language when a voice signal of a user (30) is input.
[0057] When a user's (30) voice signal is input into an artificial intelligence model at preset intervals, at least one token included in the text recognized by the voice signal and the text translated into the target language can be identified a preset number of times or more. At this time, the electronic device (100) can output the identified token through a display (130) or a speaker (not shown).
[0058] The electronic device (100) can output the text "I went to the grocery store." (100-1) obtained through ASR technology and the text "I went to the grocery store yesterday." (100-2) obtained through ST technology.
[0059] Meanwhile, although the above description assumes that the source language is English and the target language is Korean, this is only an example, and it is of course possible for the source language or target language to be Japanese, Chinese, or Spanish.
[0060] FIG. 2 is a block diagram illustrating a configuration of an electronic device (100) according to at least one embodiment of the present disclosure. The configuration illustrated in FIG. 2 is merely an example of various embodiments, and some configurations may be omitted and new configurations may be added.
[0061] As illustrated in FIG. 2, the electronic device (100) may include a communication interface (110), a memory (120), a display (130), and a processor (140). The configuration illustrated in FIG. 2 is merely an example, and it is to be understood that some components may be deleted or added depending on the configuration of the electronic device (100).
[0062] The communication interface (110) is a configuration that performs communication with various types of external devices according to various types of communication methods. The communication interface (110) can obtain data for translation from an external device. In addition, the communication interface (110) can receive information about a model that recognizes a voice signal and translates it into text in a source language or text in a target language from an external server. In addition, the communication interface (110) can also receive information about a method for determining tokens in text in a source language for which a voice signal has been recognized and information about a method for determining tokens in text in a target language for which a voice signal has been recognized.
[0063] A wireless communication module may be a module that communicates wirelessly with an external device. For example, the wireless communication module may include at least one of a Wi-Fi module, a Bluetooth module, an infrared communication module, or other communication modules.
[0064] Wi-Fi modules and Bluetooth modules can communicate via Wi-Fi and Bluetooth, respectively. When using a Wi-Fi or Bluetooth module, various connection information, such as the service set identifier (SSID) and session key, is first transmitted and received. This information is then used to establish a communication connection before various other information can be transmitted and received.
[0065] Infrared communication modules perform communication based on infrared communication (IrDA, infrared Data Association) technology, which transmits data wirelessly over short distances using infrared light, which is between visible light and millimeter waves.
[0066] In addition to the above-described communication method, other communication modules may include at least one communication chip that performs communication according to various wireless communication standards such as zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), 5G (5th Generation), etc.
[0067] A wired communication module may be a module that communicates with an external device via a wire. For example, the wired communication module may include at least one of a Local Area Network (LAN) module, an Ethernet module, a paired cable, a coaxial cable, a fiber optic cable, or an Ultra Wide-Band (UWB) module.
[0068] The memory (120) may store an operating system (OS) for controlling the overall operation of the components of the electronic device (100) and instructions or data related to the components of the electronic device (100). In particular, the memory (120) may store a voice signal spoken by a user (30), a source language text obtained by recognizing the voice signal spoken by the user (30), a text translated into a target language by the voice signal spoken by the user (30), or a plurality of tokens included in the source language text. In addition, the memory (120) may include an artificial intelligence model. In this case, the artificial intelligence model has been described above on the premise that it is stored in the memory (120), but it may be mounted on an external device and may be received from an external device.
[0069] The memory (120) may be implemented in various forms, such as volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD)).
[0070] The display (130) can display various information. Specifically, the display (130) can display text in the source language or text translated into the target language based on the voice signal spoken by the user (30).
[0071] The display (130) may be implemented as a display in various forms, such as a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display panel (PDP), etc. The display may also include a driving circuit, a backlight unit, etc., which may be implemented in forms such as an amorphous silicon thin film transistor (a-si TFT), a low temperature poly silicon (LTPS) TFT, and an organic TFT (OTFT). The display may be implemented as a touch screen combined with a touch sensor, a flexible display, a three-dimensional display (3D display), etc. According to various embodiments of the present disclosure, the display (130) may include not only a display panel that outputs an image, but also a bezel that houses the display panel.
[0072] The processor (140) may include one or more processors. Specifically, the one or more processors may include one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), a Many Integrated Core (MIC), a Digital Signal Processor (DSP), a Neural Processing Unit (NPU), a hardware accelerator, or a machine learning accelerator. The processor (140) may control one or any combination of other components of the electronic device, and may perform operations related to communication or data processing. The one or more processors may execute one or more programs or instructions stored in a memory. For example, the one or more processors may perform a method according to an embodiment of the present disclosure by executing one or more instructions stored in a memory.
[0073] One or more processors may be implemented as a single core processor including one core, or may be implemented as one or more multicore processors including multiple cores (e.g., homogeneous multicores or heterogeneous multicores). When one or more processors are implemented as a multicore processor, each of the multiple cores included in the multicore processor may include internal processor memory, such as cache memory or on-chip memory, and a common cache shared by the multiple cores may be included in the multicore processor. In addition, each of the multiple cores (or some of the multiple cores) included in the multicore processor may independently read and execute a program instruction for implementing a method according to an embodiment of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to an embodiment of the present disclosure.
[0074] In particular, the processor (140) inputs a first voice signal into an artificial intelligence model (143) to obtain a first text in a source language obtained by voice-recognizing the first voice signal and a first text in a target language obtained by translating the first voice signal into a target language, inputs a second voice signal including the first voice signal into an artificial intelligence model to obtain a second text in the source language obtained by voice-recognizing the second voice signal and a second text in the target language obtained by translating the second voice signal into the target language, determines at least one token among a plurality of tokens included in the second text of the source language based on the first text of the source language and the second text of the source language, identifies whether a token terminating a sentence is included in the determined at least one token, and if a token terminating a sentence is included, the second voice signal can be re-input into the artificial intelligence model to obtain a final text of the target language.
[0075] More specifically, the processor (140) may include a voice signal acquisition module (141), a text acquisition module (142), a token confirmation module (143), or a text output module (144) to acquire a token terminating a sentence included in a text of a source language of the user (30), a text of the source language, or a text of the target language.
[0076] The voice signal acquisition module (141) can acquire a voice signal from an external device or memory (120). The text acquisition module (142) can input the voice signal into an artificial intelligence model to acquire text in a source language and text in a target language. The token confirmation module (143) can confirm tokens identified a preset number of times or more among a plurality of tokens included in the voice signal. At this time, a method for confirming a plurality of tokens among tokens included in text in a source language and a method for confirming a plurality of tokens among tokens included in text in a target language will be described in detail below. Meanwhile, the text output module (144) can output confirmed tokens through a speaker or display (130).
[0077] The functions of each module will be described in detail later with reference to Figures 3 to 7.
[0078] The electronic device (100) can acquire a voice signal through a voice signal acquisition module (141). The voice signal acquisition module (141) can acquire a voice signal in chunk units at preset intervals. A chunk refers to a unit in which a portion of a sentence is expressed in the form of a phrase.
[0079] The electronic device (100) can acquire text in the source language and text in the target language from which the voice signal is voice-recognized through the text acquisition module (142). At this time, the text acquisition module (142) can input the voice signal into an artificial intelligence model to acquire text in the source language and text in the target language from which the voice signal is voice-recognized.
[0080] First, the electronic device (100) can acquire a first text of a source language and a first text of a target language by recognizing a first voice signal (Chunk #1) through a text acquisition module (142).
[0081] For example, for the convenience of explanation, the speech signal "I went" is defined as the first speech signal, the text obtained by recognizing the first speech signal is defined as the first text in the source language, and the text obtained by translating the first speech signal into the target language is defined as the first text in the target language, which will be described later.
[0082] The electronic device (100) can acquire a first text of a source language, “I went,” and a first text of a target language, “I went,” through a text acquisition module (142).
[0083] At this time, the electronic device (100) can identify that the first text of the source language and the first text of the target language are composed of a plurality of tokens.
[0084] For example, the electronic device (100) may identify that a first text of a source language, "I went," is composed of a plurality of tokens. Specifically, the electronic device (100) may identify that the first text of the source language, "I went," is composed of an "I" token and a "went" token.
[0085] Additionally, the electronic device (100) can identify that the first text of the target language, "I went," is composed of multiple tokens. Specifically, the electronic device (100) can identify that the first text of the target language is composed of the "I" token and the "went" token.
[0086] The electronic device (100) can acquire a second text in the source language that has been recognized as a second voice signal through the text acquisition module (142). In this case, the second voice signal may be a voice signal including the first voice signal and a voice signal (Chunk #2) received for a preset period of time after receiving the first voice signal.
[0087] For example, the case where the second speech signal (Chunk #1 and Chunk #2) is the speech "I went to the grocery" will be described below.
[0088] The electronic device (100) can input "I went to the grocery" into an artificial intelligence model through the text acquisition module (142) to obtain the text "I went to the grocery." In addition, the electronic device (100) can input "I went to the grocery" into an artificial intelligence model through the text acquisition module (142) to obtain the text "I went to the grocery" translated into a target language.
[0089] The electronic device (100) can determine at least one token from among a plurality of tokens included in a second text of the source language based on a first text of the source language and a second text of the source language through a token determination module (143). The token determination module (143) can determine an identified token if the same token is identified a preset number of times or more among the plurality of tokens included in the first text of the source language and the plurality of tokens included in the second text of the source language. The token determination module (143) can fix a token that has been identified a preset number of times or more so that it cannot be changed.
[0090] For example, the electronic device (100) can identify the tokens included in the text "I went" as the "I" token and the "went" token through the token confirmation module (143). In addition, the electronic device (100) can identify the tokens included in the text "I went to the grocery" as the "I" token, the "went" token, the "to" token, the "the" token, and the "grocery" token through the token confirmation module (143). At this time, the electronic device (100) can determine that the "I" token and the "went" token have been identified twice. The electronic device (100) can confirm the "I" token and the "went" token through the token confirmation module (143).
[0091] Meanwhile, the electronic device (100) can confirm tokens included in the text translated into the target language. For example, the electronic device (100) can identify the tokens included in "I went" as the "I" token and the "went" token through the token confirmation module (143). Furthermore, the electronic device (100) can identify the tokens included in "I went to the grocery store" as the "I" token, the "grocery" token, and the "went" token through the token confirmation module (143). At this time, the electronic device (100) can determine that the "I" token has been identified twice. The electronic device (100) can confirm the "I" token through the token confirmation module (143).
[0092] At this time, the electronic device (100) can output a confirmed token among a plurality of tokens included in the second text of the source language through the text output module (144) and not output the remaining tokens. The text output module (144) can output the confirmed token through a display or speaker.
[0093] That is, the electronic device (100) can output the confirmed "I" token and "went" token through the text output module (144) as illustrated in FIG. 3. At this time, the electronic device (100) may not output the "to" token, the "the" token, and the "grocery" token, excluding the confirmed "I" token and the "went" token.
[0094] In addition, the electronic device (100) can output an identified token when the same token is identified a preset number of times or more among a plurality of tokens included in a first text of the target language and a plurality of tokens included in a second text of the target language through the text output module (144).
[0095] Meanwhile, the electronic device (100) may output an "I" token through the text output module (144). At this time, the electronic device (100) may not output the "to groceries" token and the "went" token, excluding the confirmed "I" token.
[0096] When the electronic device (100) is identified as a case where the second text of the source language does not include a token that terminates a sentence, the electronic device (100) may obtain a third speech signal (Chunk #1, Chunk #2, and Chunk 33) including the second speech signal, input the third speech signal into an artificial intelligence model to obtain a third text of the source language obtained by voice-recognizing the third speech signal and a third text of the target language obtained by translating the third speech signal into the target language, determine at least one token among a plurality of tokens included in the third text of the source language based on the second text of the source language and the third text of the source language, and identify whether the determined at least one token includes a token that terminates a sentence. In addition, the third speech signal may be a speech signal including a speech signal received for a preset period of time after receiving the second speech signal.
[0097] This will be described later with reference to Fig. 4.
[0098] For example, if the electronic device (100) is identified as not including a token terminating the sentence “I went to the grocery,” the electronic device (100) may acquire a voice signal (Chunk #3) called “store” through the voice signal acquisition module (141).
[0099] At this time, for the convenience of explanation, "I went to the grocery store" is defined as the third speech signal, and the text in the source language that recognizes the third speech signal is defined as the third text in the source language, and this will be described later. In addition, the text in the target language that translates the third speech signal into the target language is defined as the third text in the target language, and this will be described later.
[0100] The electronic device (100) can acquire a third text in the source language, “I went to the grocery store,” and a third text in the target language, “I went to the grocery store,” by recognizing the third speech signal, “I went to the grocery store,” through the text acquisition module (142).
[0101] At this time, the electronic device (100) can determine, through the token confirmation module (145), a token identified twice among the tokens included in the second text of the source language and the tokens included in the third text of the source language.
[0102] For example, the electronic device (100) may identify the tokens included in "I went to the grocery" as an "I" token, a "went" token, a "to" token, a "the" token, and a "grocery" token. Additionally, the electronic device (100) may identify the tokens included in "I went to the grocery store" as an "I" token, a "went" token, a "to" token, a "the" token, a "grocery" token, and a "store" token.
[0103] The electronic device (100) can confirm the “to” token, the “the” token, and the “grocery” token, excluding the “I” token and the “went” token, which are already confirmed tokens, through the token confirmation module (145).
[0104] Meanwhile, the electronic device (100) can determine a token identified twice among the tokens included in the second text of the target language and the tokens included in the third text of the target language through the token confirmation module (145).
[0105] For example, the electronic device (100) may identify the tokens included in "I went to the grocery store" as the "I" token, the "to the grocery store" token, and the "went" token. Furthermore, the electronic device (100) may identify the tokens included in "I went to the grocery store" as the "I" token, the "to the grocery store" token, and the "went" token.
[0106] The electronic device (100) can determine that the "Food" token, excluding the "I" token already confirmed through the token confirmation module (145), has been identified twice. At this time, the electronic device (100) can confirm the "Food" token through the token confirmation module (145).
[0107] The electronic device (100) can output the confirmed "to" token, "the" token, and "grocery" token through the text output module (144). At this time, the "I" token and the "went" token can be prevented from being output.
[0108] Meanwhile, the electronic device (100) may output the "I" token, which is a confirmed token, through the text output module (144). At this time, the electronic device (100) may not output the "to groceries" token and the "went" token, excluding the confirmed "I" token.
[0109] If the electronic device (100) identifies that the sentence "I went to the grocery store" does not include a token that terminates the sentence, the electronic device (100) can acquire the speech signal "yesterday" (Chunk #4) through the speech signal acquisition module (141). At this time, for the convenience of explanation, "I went to the grocery store yesterday" is defined as the fourth speech signal (Chunk #1, Chunk #2, Chunk #3, and Chunk #4) and will be described later.
[0110] As shown in FIG. 5, the electronic device (100) inputs the fourth voice signal “I went to the grocery store yesterday” into an artificial intelligence model, which is the fourth text of the source language that has been recognized through the text acquisition module (142), to acquire the fourth text of the source language “I went to the grocery store yesterday” and the fourth text of the target language “I went to the grocery store yesterday.”
[0111] At this time, the electronic device (100) can determine, through the token confirmation module (145), a token identified twice among the tokens included in the third text of the source language and the tokens included in the fourth text of the source language.
[0112] For example, the electronic device (100) may identify tokens included in a third text of the source language as an "I" token, a "went" token, a "to" token, a "the" token, a "grocery" token, and a "store" token. Additionally, the electronic device (100) may identify tokens included in a fourth text of the source language as an "I" token, a "went" token, a "to" token, a "the" token, a "grocery" token, a "store" token, and a "yesterday" token.
[0113] The electronic device (100) can confirm the “store” token excluding the “I” token, “went” token, “to” token, “the” token, and “grocery” token, which are already confirmed tokens, through the token confirmation module (145).
[0114] The electronic device (100) can identify the tokens included in the third text of the target language as the "I" token, the "to the grocery store" token, and the "went" token. Furthermore, the electronic device (100) can identify the tokens included in the fourth text of the target language as the "I" token, the "to the grocery store" token, the "went" token, and the "yesterday" token.
[0115] The electronic device (100) can determine that the "point" token has been identified twice, excluding the "I" token and the "groceries" token, which have already been confirmed through the token confirmation module (145). At this time, the electronic device (100) can confirm the "point" token through the token confirmation module (145).
[0116] The electronic device (100) can output the confirmed "store" token through the text output module (144). At this time, the "I" token, the "went" token, the "to" token, the "the" token, and the "grocery" token can be prevented from being output.
[0117] Additionally, the electronic device (100) can output the "at the point" token and the "went" token, which are confirmed tokens, through the text output module (144). At this time, the "I" token and the "groceries" token, excluding the confirmed "at the point" token and the "went" token, can be prevented from being output.
[0118] The electronic device (100) can obtain tokens that end sentences such as "I went to the grocery store yesterday" and "". For the convenience of explanation, "I went to the grocery store yesterday" will be defined as the fifth voice signal, and the text of the source language obtained by voice recognition of the fifth voice signal will be defined as the fifth text of the source language and will be described later. Also, the text of the target language obtained by translating the fifth voice signal into the target language will be defined as the fifth text of the target language and will be described later.
[0119] On the other hand, in order to identify "", which is a token that ends a sentence, there are methods of silence detection, phoneme and prosody pattern analysis, language model-based prediction, or methods of recognizing special expressions.
[0120] In one embodiment, the electronic device (100) can identify a token that ends a sentence by using the method of silence detection. Specifically, when the electronic device (100) detects only noise and no voice for a certain period of time in the voice signal, it can be recognized as a sentence end signal. For example, when the electronic device (100) detects only noise and no voice of the user for a certain period of time, it can be considered that the sentence has ended and a token that ends the sentence can be generated.
[0121] In one or more embodiments, the electronic device (100) can identify a token that ends a sentence by using the method of phoneme and prosody pattern analysis. Specifically, the electronic device (100) can identify a token that ends a sentence by recognizing a specific pattern or phoneme that appears at the end of the sentence in the voice signal. For example, when the user utters the voice "입니다." in Korean, since this corresponds to an ending suffix that ends the sentence, the electronic device (100) can identify it as a token that ends the sentence.
[0122] In one or more embodiments, the electronic device (100) can identify a token that terminates a sentence using a language model-based prediction method. Specifically, the electronic device (100) can identify the end of a sentence in a speech signal based on context and grammar. For example, if a user utters "I'm home today" in Korean, the sentence can be identified as not ending, as there is a high probability that a verb form such as "went."
[0123] In one or more embodiments, the electronic device (100) may identify a token that terminates a sentence using a special expression. Specifically, the electronic device (100) may recognize a specific word as a sentence end signal when detected. For example, the electronic device (100) may recognize the phrase "That's all" as the end of a sentence when detected.
[0124] As illustrated in FIG. 6, the electronic device (100) inputs the fifth speech signal “I went to the grocery store yesterday” into the artificial intelligence model through the text acquisition module (142), thereby acquiring the fifth text “I went to the grocery store yesterday” in the source language and the fifth text “I went to the grocery store yesterday” in the target language.
[0125] At this time, the electronic device (100) can determine, through the token confirmation module (145), a token identified twice among the tokens included in the fourth text of the source language and the tokens included in the fifth text of the source language.
[0126] For example, the electronic device (100) can identify that the tokens included in the fourth text of the source language are an "I" token, a "went" token, a "to" token, a "the" token, a "grocery" token, a "store" token, and a "" token, and that the tokens included in the fifth text of the source language are also an "I" token, a "went" token, a "to" token, a "the" token, a "grocery" token, a "store" token, a "yesterday" token, and a "" token.
[0127] The electronic device (100) can determine that the "yesterday" token and the "" token have been identified twice through the token confirmation module (145). The electronic device (100) can confirm the "yesterday" token and the "" token through the token confirmation module (145).
[0128] Meanwhile, the electronic device (100) can identify that the tokens included in the fourth text of the target language are the "I" token, the "to the grocery store" token, the "went" token, and the "" token, and that the tokens included in the fifth text of the target language are the "I" token, the "to the grocery store" token, the "went" token, the "yesterday" token, and the "" token.
[0129] The electronic device (100) can determine that the "Yesterday" token and the "" token, excluding the "I" token, the "To the grocery store" token, and the "Went" token, which have already been confirmed through the token confirmation module (145), have been identified twice. At this time, the electronic device (100) can confirm the "Yesterday" token and the "" token through the token confirmation module (145).
[0130] The electronic device (100) can output the confirmed "yesterday" token through the text output module (144) and identify that the sentence has ended through the " " token. At this time, the "I" token, the "went" token, the "to" token, the "the" token, the "grocery" token, and the "store" token can be prevented from being output.
[0131] Meanwhile, when outputting the target language, the electronic device (100) can output the "yesterday" token, which is a confirmed token, through the text output module (144), and can identify that the sentence has ended through the " " token. At this time, the "I" token, the "to the grocery store" token, and the "went" token, excluding the confirmed "yesterday" token, can be prevented from being output.
[0132] Meanwhile, Fig. 8 is a diagram for explaining an artificial intelligence model (143) according to one embodiment of the present disclosure. The artificial intelligence model is characterized in that it is trained to acquire a first text in a source language and a first text in a target language corresponding to a first speech signal in parallel based on a first speech signal input to the artificial intelligence model.
[0133] As illustrated in FIG. 8, when a user (30) utters the voice "I went to the grocery store yesterday" and the uttered voice is input into the artificial intelligence model (143), the first text in the source language (English) that recognizes the voice signal, i.e., the text "I went to the grocery store yesterday.", can be obtained. In addition, the first text translated into the target language (Korean) of the voice signal, i.e., the text "I went to the grocery store yesterday.", can be obtained.
[0134] At this time, the artificial intelligence model that outputs text in the source language that recognizes the voice signal and text in the target language that translates the voice signal into the target language was described above assuming a model based on the transformer structure. However, this is only one example, and various types of neural network models such as CNN (Convolution Neural Network), 1DCNN (1-Dimension Convolution Neural Network), R-CNN (Region with Convolution Neural Network), RPN (Region Proposal Network), RNN (Recurrent Neural Network), S-DNN (Stacking-based deep Neural Network), S-SDNN (State-Space Dynamic Neural Network), Deconvolution Network, DBN (Deep Belief Network), RBM (Restricted Boltzman Machine), Fully Convolutional Network, LSTM (Long Short-Term Memory) Network, Bi-LSTM (Bidirectional-Long Short-Term Memory) Network Classification Network, Plain Residual Network, Dense Network, Hierarchical Pyramid Network, Fully Convolutional Network, SENet (Squeeze and Excitation Network), Transformer Network, Encoder, Decoder, Auto Encoder, or a combination thereof. It may include at least one, and the artificial neural network in the present disclosure is not limited to the examples described above.
[0135] FIG. 9 is a flowchart illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0136] First, the electronic device (100) can acquire a first voice signal and obtain a first text in a source language and a first text in a target language (S900). By inputting the first voice signal into an artificial intelligence model (143), the first text in the source language and the first text in the target language can be acquired through voice recognition.
[0137] Thereafter, the electronic device (100) can acquire a second voice signal to acquire a second text in the source language and a second text in the target language (S905). At this time, the second voice signal can be input into an artificial intelligence model (143) to acquire a second text in the source language and a second text in the target language through voice recognition.
[0138] At this time, the artificial intelligence model (143) has been described above, so a detailed description will be omitted.
[0139] The electronic device (100) can determine whether the same token among the multiple tokens included in the two received voice signals has been identified a preset number of times or more (S910).
[0140] That is, the electronic device (100) can determine whether the same token is identified at least twice among a plurality of tokens included in a first text that recognizes a first voice signal and a plurality of tokens included in a second text that recognizes a second voice signal.
[0141] At this time, the electronic device (100) can also determine the tokens included in the text of the target language independently of the source language text. That is, the electronic device (100) can determine whether the same token is identified at least twice among the plurality of tokens included in the first text translated from the first speech signal into the target language and the plurality of tokens included in the second text translated from the second speech signal into the target language. Since the method of identifying and confirming the tokens included in the text of the source language and the text of the target language has been described above, a detailed description thereof will be omitted.
[0142] The electronic device (100) can output text of the source language and text of the target language corresponding to tokens identified more than a preset number of times (S915).
[0143] The electronic device (100) may output tokens identified a preset number of times or more among tokens included in the text of the source language, and may output tokens identified a preset number of times or more among tokens included in the text of the target language. However, this is only an example, and the electronic device (100) may not output tokens included in the text of the source language, but may output only tokens identified a preset number of times or more among tokens included in the text of the target language.
[0144] The electronic device (100) can determine whether at least one confirmed token includes a token that ends a sentence (S920).
[0145] If at least one token that terminates a sentence is included in the confirmed tokens (S920-Y), the electronic device (100) can determine whether the token that terminates a sentence has been identified a preset number of times or more.
[0146] At this time, the electronic device (100) can identify that the sentence has ended if the token that ends the sentence is identified more than a preset number of times (S925).
[0147] The electronic device (100) can obtain the final text in the target language (S930). The electronic device (100) can also obtain the final text in the source language.
[0148] However, the electronic device (100) can receive the third voice signal (S935) if at least one token that terminates the sentence is not included in the confirmed token (S920-N).
[0149] That is, if the electronic device (100) does not include a token that terminates the sentence, it can additionally receive a voice signal as it is determined that the sentence is not terminated.
[0150] At this time, the electronic device (100) can obtain a third text in the source language that has recognized the third voice signal and a third text in the target language that has translated the third voice signal into the target language (S940). At this time, the third voice signal refers to a voice signal that includes a voice signal additionally received after a preset time has elapsed since receiving the second voice signal and the second voice signal.
[0151] The electronic device (100) can determine whether the same token has been identified a preset number of times or more based on the third text of the source language and the third text of the target language (S910).
[0152] FIG. 10 is a flowchart for explaining a control method of an electronic device according to one embodiment of the present disclosure.
[0153] First, the electronic device (100) inputs a first voice signal into an artificial intelligence model to obtain a first text in a source language that recognizes the first voice signal and a first text in a target language that translates the first voice signal into the target language (S1010).
[0154] The electronic device (100) can input a second voice signal including a first voice signal into an artificial intelligence model to obtain a second text in a source language in which the second voice signal is voice-recognized and a second text in a target language in which the second voice signal is translated into the target language (S1020).
[0155] The electronic device (100) can determine at least one token among a plurality of tokens included in the second text of the source language based on the first text of the source language and the second text of the source language (S1030).
[0156] The electronic device (100) can identify whether at least one confirmed token includes a token that terminates a sentence (S1040). At this time, the token that terminates a sentence (End of Sentence token or EOS token) is described as ''.
[0157] If the electronic device (100) includes a token that terminates a sentence, the second voice signal can be re-input into the artificial intelligence model to obtain the final text in the target language (S1050).
[0158] Meanwhile, the control method of the electronic device (100) according to the above-described embodiment may be implemented as a program and provided to the electronic device (100). In particular, the program including the control method of the electronic device (100) may be stored and provided in a non-transitory computer readable medium.
[0159] Specifically, in a non-transitory computer-readable recording medium including a program for executing a control method of an electronic device (100), the control method of the electronic device (100) may include a step of inputting a first voice signal into an artificial intelligence model to obtain a first text in a source language obtained by voice-recognizing the first voice signal and a first text in the target language obtained by translating the first voice signal into a target language, a step of inputting a second voice signal including the first voice signal into an artificial intelligence model to obtain a second text in the source language obtained by voice-recognizing the second voice signal and a second text in the target language obtained by translating the second voice signal into the target language, a step of determining at least one of a plurality of tokens included in the second text of the source language based on the first text of the source language and the second text of the source language, a step of identifying whether a token terminating a sentence is included in the determined at least one token, and a step of re-inputting the second voice signal into the artificial intelligence model to obtain a final text of the target language if a token terminating a sentence is included.
[0160]
[0161] FIG. 11 is a flowchart illustrating one or more embodiments related to quality assessment.
[0162] The above described embodiment re-inputs the second speech signal into the AI model without considering the quality of the second text in the target language, translated from the second speech signal, if at least one of the confirmed tokens includes a sentence-terminating token. However, according to another embodiment, the electronic device (100) can obtain the final text while considering the quality of the second text.
[0163] As illustrated in FIG. 11, the electronic device (100) can identify whether at least one confirmed token includes a token that terminates a sentence (S1110).
[0164] If at least one of the confirmed tokens does not include a token that terminates a sentence (S1110-N), as described above, the electronic device (100) obtains a third voice signal including a second voice signal (e.g., a voice signal received for a preset time after receiving the second voice signal), inputs the third voice signal into an artificial intelligence model to obtain a third text in a source language obtained by voice-recognizing the third voice signal and a third text in a target language obtained by translating the third voice signal into a target language, and determines one or more tokens among a plurality of tokens included in the third text in the source language based on the second text in the source language and the third text in the source language, and identifies whether the confirmed one or more tokens include a token that terminates a sentence.
[0165] If at least one of the confirmed tokens includes a token that terminates a sentence (S1110-Y), the electronic device (100) can obtain a score indicating the quality of the second text in the target language (S1120). The electronic device (100) can obtain a score indicating the quality of the second text in the target language based on at least one of a learned neural network model, a predefined rule, and a user input.
[0166] In one embodiment, the electronic device (100) may input a second text in a target language into a quality assessment model to obtain a first score representing the quality of the second text in the target language. The quality assessment model may refer to a neural network model trained based on training data including an original text in a source language and a translated text in a target language. For example, the quality assessment model may include, but is not limited to, a neural network architecture such as BERT (Bidirectional Encoder Representations from Transformers) for classifying or regressively predicting the similarity or relationship between the original text and the translated text.
[0167] The quality assessment model can be implemented as a separate model from the artificial intelligence model described above, or it can be implemented as a single model integrated into the artificial intelligence model.
[0168] In one embodiment, the electronic device (100) can obtain a second score indicating the quality of the second text in the target language by applying a predefined metric to the second text in the target language. The electronic device (100) can store information about the predefined metric so as to evaluate the naturalness of a sentence, and can obtain a second score indicating the quality of the second text in the target language using the information about the stored metric. For example, the metric can be defined based on rules such as whether there is a lot of overlapping content in a sentence, or how evenly words used in a sentence are distributed.
[0169] In one embodiment, the electronic device (100) may obtain a third score indicating the quality of a second text in a target language based on a user input. Specifically, the electronic device (100) may display a user interface including the second text in the target language. The electronic device (100) may receive a user input evaluating the quality of the second text in the target language through the user interface. The electronic device (100) may obtain a third score based on the user input.
[0170] The user input or third score can be considered as the user's feedback to the artificial intelligence model, and the electronic device (100) can also retrain the quality assessment model based on the user input or the third score.
[0171] In one embodiment, the electronic device (100) may obtain a final score using two or more of the first, second, and third scores obtained as described above. For example, the electronic device (100) may obtain a final score based on a weighted sum of two or more of the first, second, and third scores.
[0172] The electronic device (100) can determine whether the acquired score is greater than or equal to a preset threshold value (S1130). If the score is greater than or equal to the preset threshold value (S1130-Y), the electronic device (100) can acquire the second text in the target language as the final text in the target language (S1140). On the other hand, if the score is less than the threshold value, the electronic device (100) can re-input the second speech signal into the artificial intelligence model to acquire the final text in the target language (S1150).
[0173] In other words, the electronic device (100) performs a quality evaluation on the second text in the target language, identifies whether re-inference using an artificial intelligence model is necessary based on the quality evaluation results, and improves the quality using the artificial intelligence model only when the quality of the second text in the target language is low, and provides the second text in the target language as the final text when the quality of the second text in the target language is high.
[0174] Meanwhile, whether to obtain the final text by considering the quality of the second text, i.e., whether to perform a quality evaluation process, can be turned ON or OFF depending on the user settings.
[0175] In the above, a method for controlling an electronic device (100) and a computer-readable recording medium including a program for executing the method for controlling an electronic device (100) have been briefly described, but this is only to omit redundant descriptions, and it goes without saying that various embodiments of the electronic device (100) can also be applied to a method for controlling an electronic device (100) and a computer-readable recording medium including a program for executing the method for controlling an electronic device (100).
[0176] According to the embodiments described above, the electronic device (100) performs speech recognition and interpretation / translation in parallel using an artificial intelligence model, and identifies the end point of a sentence based on the speech recognition result that shows higher performance than the interpretation, and inputs a speech signal corresponding to a specific sentence based on the identified end point back into the artificial intelligence model to obtain the final text (translation). Accordingly, the quality of the final text resulting from the interpretation / translation can be significantly improved.
[0177] In other words, the electronic device (100) can find the end point of a sentence based on the result of voice recognition and use the corresponding voice for voice translation at the same time, thereby significantly improving the relatively difficult voice translation performance without any decrease in speed.
[0178] Furthermore, the electronic device (100) can display the interpretation result for a portion of a sentence in real time without delay, and can replace the interpretation result for a portion of the sentence with a new result by going through a re-inference process through an artificial intelligence model once again using the separated voice signal according to the voice recognition result, thereby providing a high-quality interpretation result.
[0179]
[0180] The artificial intelligence-related function according to the present disclosure is operated through the processor (140) and memory (120) of the electronic device (100).
[0181] The processor (140) may be composed of one or more processors. In this case, the one or more processors may include at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an NPU (Neural Processing Unit), but is not limited to the examples of the processors described above.
[0182] CPUs are general-purpose processors capable of performing not only general calculations but also artificial intelligence calculations. Their multi-layered cache structure allows for the efficient execution of complex programs. CPUs are advantageous for serial processing, enabling organic linking of previous and subsequent calculation results through sequential calculations. General-purpose processors are not limited to the examples described above, except where specifically identified as CPUs.
[0183] A GPU is a processor designed for large-scale computations, such as floating-point operations used in graphics processing. It integrates a large number of cores to perform large-scale computations in parallel. In particular, GPUs may be advantageous over CPUs in parallel processing methods, such as convolution operations. Furthermore, GPUs can be used as coprocessors to supplement the functions of CPUs. Processors for large-scale computations are not limited to the examples described above, except in cases where they are specifically referred to as GPUs.
[0184] An NPU is a processor specialized in artificial intelligence computation using artificial neural networks, and each layer of the artificial neural network can be implemented in hardware (e.g., silicon). Since NPUs are designed specifically according to the company's specifications, they have less freedom than CPUs or GPUs, but can efficiently process the AI computations requested by the company. Meanwhile, as a processor specialized in AI computation, an NPU can be implemented in various forms, such as a Tensor Processing Unit (TPU), an Intelligence Processing Unit (IPU), or a Vision Processing Unit (VPU). Except as specifically designated as an NPU, an AI processor is not limited to the examples described above.
[0185] Additionally, one or more processors may be implemented as a System on Chip (SoC). In this case, in addition to one or more processors, the SoC may further include memory and a network interface, such as a bus, for data communication between the processor and the memory.
[0186] When a plurality of processors are included in a SoC (System on Chip) included in an electronic device (100), the electronic device (100) may perform operations related to artificial intelligence (e.g., operations related to learning or inference of an artificial intelligence model) by using some of the plurality of processors. For example, the electronic device (100) may perform operations related to artificial intelligence by using at least one of a GPU, an NPU, a VPU, a TPU, and a hardware accelerator specialized in artificial intelligence operations such as convolution operations and matrix multiplication operations among the plurality of processors. However, this is merely an example, and it is of course possible to process operations related to artificial intelligence by using a CPU or a general-purpose processor.
[0187] Additionally, the electronic device (100) can perform operations related to functions related to artificial intelligence by utilizing multiple cores (e.g., dual cores, quad cores, etc.) included in a single processor. In particular, the electronic device (100) can perform artificial intelligence operations, such as convolution operations and matrix multiplication operations, in parallel by utilizing multiple cores included in the processor.
[0188] One or more processors are controlled to process input data according to predefined operating rules or artificial intelligence models stored in memory. The predefined operating rules or artificial intelligence models are characterized by being created through learning.
[0189] Here, "created through learning" means that a predefined set of behavioral rules or an AI model with desired characteristics is created by applying a learning algorithm to a large number of learning data. This learning may be performed on the device itself, where the AI according to the present disclosure is implemented, or through a separate server / system.
[0190] An artificial intelligence model may be composed of multiple neural network layers. At least one layer has at least one weight value and performs its operation through the operation result of the previous layer and at least one defined operation. Examples of neural networks include a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, and a transformer. The neural networks in the present disclosure are not limited to the above-described examples unless otherwise specified.
[0191] A learning algorithm is a method for training a target device (e.g., a robot) using a large amount of learning data, enabling the target device to make decisions or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Unless otherwise specified, the learning algorithms in this disclosure are not limited to the aforementioned examples.
[0192] A device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0193] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM) or directly between two user devices (e.g., smartphones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be at least temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0194] Each of the components (e.g., modules or programs) according to the various embodiments of the present disclosure as described above may be composed of a single or multiple entities, and some of the sub-components described above may be omitted, or other sub-components may be further included in the various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the respective components prior to integration.
[0195] According to various embodiments, operations performed by a module, program or other component may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
[0196] Meanwhile, the terms "part" or "module" used in the present disclosure include units composed of hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A "part" or "module" may be an integrally composed component, a minimum unit performing one or more functions, or a portion thereof. For example, a module may be composed of an application-specific integrated circuit (ASIC).
[0197] Various embodiments of the present disclosure may be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device may include an electronic device (e.g., an electronic device (100)) according to the disclosed embodiments, which is a device capable of calling instructions stored in the storage medium and operating according to the called instructions.
[0198] When the above instruction is executed by the processor, the processor may perform the function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or interpreter.
[0199] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.
Claims
1. In electronic devices, communication interface; Memory that stores instructions; and comprising at least one processor; The above instructions, when collectively or individually executed by the at least one processor, cause the electronic device to: By inputting a first voice signal into an artificial intelligence model, a first text in a source language obtained by recognizing the first voice signal and a first text in a target language obtained by translating the first voice signal into the target language are obtained, By inputting a second voice signal including the first voice signal into the artificial intelligence model, a second text in a source language obtained by recognizing the second voice signal and a second text in a target language obtained by translating the second voice signal into the target language are obtained, Determine at least one token among a plurality of tokens included in the second text of the source language based on the first text of the source language and the second text of the source language, Identify whether at least one of the above-determined tokens contains a token that terminates the sentence, An electronic device that re-inputs the second speech signal into the artificial intelligence model to obtain a final text in the target language, if a token terminating the sentence is included.
2. In paragraph 1, The above instructions, when collectively or individually executed by the at least one processor, cause the electronic device to: An electronic device that determines the identified token when the same token is identified a preset number of times or more among a plurality of tokens included in a first text of the source language and a plurality of tokens included in a second text of the source language.
3. In paragraph 1, The above instructions, when collectively or individually executed by the at least one processor, cause the electronic device to: If the second text of the source language does not contain a token that terminates the sentence, a third speech signal including the second speech signal is obtained, By inputting the third voice signal into the artificial intelligence model, a third text in the source language in which the third voice signal is recognized and a third text in the target language in which the third voice signal is translated into the target language are obtained, Determine one or more tokens among a plurality of tokens included in the third text of the source language based on the second text of the source language and the third text of the source language, Identify whether one or more of the above-determined tokens contain a token that terminates a sentence, An electronic device in which the third voice signal is a voice signal including a voice signal received for a preset period of time after receiving the second voice signal.
4. In paragraph 2, The above instructions, when collectively or individually executed by the at least one processor, cause the electronic device to: An electronic device that outputs the confirmed token among a plurality of tokens included in a second text of the source language and does not output the remaining tokens.
5. In paragraph 4, The above instructions, when collectively or individually executed by the at least one processor, cause the electronic device to: An electronic device configured to output the identified token when the same token is identified a preset number of times or more among a plurality of tokens included in a first text of the target language and a plurality of tokens included in a second text of the target language.
6. In paragraph 1, An electronic device in which the second voice signal is a voice signal including the first voice signal and a voice signal received for a preset time after receiving the first voice signal.
7. In paragraph 1, The above artificial intelligence model is, An electronic device characterized in that, when the first voice signal is input, the device is trained to acquire in parallel a first text in a source language in which the first voice signal is voice-recognized and a first text in a target language in which the first voice signal is translated into the target language.
8. In paragraph 1, The above instructions, when collectively or individually executed by the at least one processor, cause the electronic device to: If a token that terminates the above sentence is included, a score representing the quality of the second text in the target language is obtained, If the above score is greater than or equal to a preset threshold, the second text in the target language is obtained as the final text in the target language, An electronic device that re-inputs the second speech signal into the artificial intelligence model to obtain a final text in the target language if the score is less than the threshold value.
9. In paragraph 8, The above instructions, when collectively or individually executed by the at least one processor, cause the electronic device to: An electronic device that obtains the score based on at least one of a first score obtained by inputting a second text of the target language into a quality assessment model and a second score obtained by applying a predefined metric to the second text of the target language.
10. In paragraph 8, display; including more, The above instructions, when collectively or individually executed by the at least one processor, cause the electronic device to: Controlling the display to display a user interface including a second text in the target language; Receiving user input for evaluating the quality of a second text in the target language through the user interface; An electronic device that obtains the score based on the user input.
11. In a method for controlling an electronic device, A step of inputting a first voice signal into an artificial intelligence model to obtain a first text in a source language obtained by voice-recognizing the first voice signal and a first text in a target language obtained by translating the first voice signal into the target language; A step of inputting a second voice signal including the first voice signal into the artificial intelligence model to obtain a second text in a source language in which the second voice signal is voice-recognized and a second text in a target language in which the second voice signal is translated into the target language; A step of determining at least one token among a plurality of tokens included in a second text of the source language based on a first text of the source language and a second text of the source language; A step of identifying whether at least one of the above-determined tokens includes a token that terminates a sentence; and A control method comprising: a step of re-inputting the second speech signal into the artificial intelligence model to obtain a final text in the target language, if a token terminating the sentence is included; 12. In paragraph 11, A control method comprising: a step of confirming an identified token when the same token is identified a preset number of times or more among a plurality of tokens included in a first text of the source language and a plurality of tokens included in a second text of the source language.
13. In paragraph 11, If the second text of the source language does not contain a token terminating the sentence, obtaining a third speech signal including the second speech signal; A step of inputting the third voice signal into the artificial intelligence model to obtain a third text in the source language in which the third voice signal is voice-recognized and a third text in the target language in which the third voice signal is translated into the target language; A step of determining at least one token among a plurality of tokens included in a third text of the source language based on a second text of the source language and a third text of the source language; and A step of identifying whether one or more of the above-determined tokens includes a token that terminates a sentence; A control method in which the third voice signal is a voice signal including a voice signal received for a preset period of time after receiving the second voice signal.
14. In paragraph 12, A control method comprising: a step of outputting the confirmed token among the plurality of tokens included in the second text and not outputting the remaining tokens; 15. In paragraph 14, A control method comprising: a step of outputting the identified token when the same token is identified a preset number of times or more among a plurality of tokens included in a first text of the target language and a plurality of tokens included in a second text of the target language.
Citation Information
Patent Citations
Mobile terminal and method for translating voice using thereof
KR101949731B1
Device and method of translating a language into another language
KR1020180020368A
Foreign matter shredder in a drainage ditch operated by a water level sensor
KR1020220161530A
Pharmaceutical composition for enhancing hyperthermia for improving gut microbiota
KR1020240046153A
A method of operating platform for advertising an shopping malls
KR1020250137899A