Character string prediction device, and character string prediction method
The character string prediction device uses a predictive translation unit and selective execution to ensure real-time coherent translation by discarding or continuing processing based on match thresholds, addressing delays in existing systems.
Patent Information
- Application Number
- JP2024048848
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-26
- Publication Date
- 2025-10-08
AI Technical Summary
Existing character string processing systems struggle with real-time translation and understanding due to delays in processing units smaller than sentences, leading to difficulties in maintaining coherent communication, especially in situations requiring immediate interpretation.
A character string prediction device that includes a predictive translation unit using a pre-trained language model to generate and output related character strings in a second language, with a selective execution unit that discards or continues processing based on the match between predicted and subsequent input strings, ensuring real-time coherence.
Enables real-time translation by predicting and displaying coherent character strings only when the match exceeds a threshold, reducing calculation time and improving understanding of the input text.
Smart Images

Figure 2025148634000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to natural language processing technology, and more particularly to a character string prediction device and a character string prediction method for predicting a character string that has a predetermined relationship with an input character string. [Background technology]
[0002] People often communicate through spoken language. However, communication can go wrong between people who speak different languages. This problem is not limited to one-to-one situations, but can also occur in one-to-many and many-to-many situations. For example, in international conferences, simultaneous or consecutive interpreters are used to address this issue. However, it is difficult to provide such means for person-to-person communication.
[0003] To solve these problems, it would be desirable to have a device that uses a speech recognition device, an automatic translation device, and a speech synthesis device to assist communication between different people. Even in the international conferences mentioned above, a device that could replace simultaneous interpreters would be advantageous in many ways. However, it is extremely difficult to perform automatic translation to replace simultaneous interpreters. This is because automatic translation is often based on the premise that it processes each sentence. When translating sentence by sentence, if a speaker is speaking continuously, the automatic translation may not be able to keep up with what the speaker is saying. This situation can cause problems in communication between interlocutors and in understanding the opinions of speakers in conferences, etc. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-071761 Summary of the Invention [Problem to be solved by the invention]
[0005] A proposal to solve the above-mentioned problem is disclosed in the above-mentioned Patent Document 1. The machine translation device disclosed in Patent Document 1 analyzes the grammar of an input utterance and divides the utterance into processing units smaller than sentences. This machine translation device then translates each processing unit into a target language and controls the display order, etc., according to the analysis results.
[0006] However, in the device disclosed in Patent Document 1, controlling the display order causes delays in translation, making it difficult to translate the content of the utterance in real time. If the display order were to match the order of the utterance, each processing unit would be translated and displayed separately. This creates the problem of making the content of the utterance difficult to understand.
[0007] These problems can arise not only when attempting to perform simultaneous interpretation using machine translation, but also when, for example, trying to rephrase difficult phrases into simpler expressions within the same language.
[0008] Therefore, an object of the present invention is to provide a character string processing device and a character string processing method that can convert an input character string into another character string that has a predetermined relationship with the input character string while maintaining real-time processing of the input character string. [Means for solving the problem]
[0009] A character string prediction device according to a first aspect of the present invention includes: a prediction unit that, in response to input of a first character string in a first language, predicts and outputs at least a second character string in the first language and a related character string, which is a character string in the second language that has a predetermined relationship with a concatenated character string, which is a character string obtained by concatenating the first character string and the second character string; and an execution unit that, in response to the prediction unit outputting at least the second character string, executes a process of outputting the related character string output by the prediction unit or a process of inputting the subsequent input character string following the first character string to the prediction unit, depending on whether a subsequent input character string input following the first character string and the second character string satisfy a predetermined condition.
[0010] Preferably, the prediction unit includes a predictive translation unit that, in response to an input of the first character string, predicts and outputs at least a second character string and a related character string that is a translation in the second language for the concatenated character string.
[0011] More preferably, the second language is a language different from the first language, and the predictive translation unit includes a pre-trained language model to predict and output at least the second string and related strings in response to input of the first string.
[0012] More preferably, the execution unit includes a selective execution unit that, in response to the prediction unit outputting at least the second string, executes a process of outputting the related string output by the prediction unit, or a process of inputting the subsequent input string to the prediction unit following the first string, depending on whether or not the condition that the degree of match between the subsequent input string and the second string is greater than a threshold value is satisfied.
[0013] Preferably, the character string prediction device further includes an input discarding unit that discards the second character string and the character string in the first language that is input to the character string prediction device following the first character string in response to the condition that the degree of match is greater than a threshold being satisfied.
[0014] A string prediction method according to a second aspect of the present invention includes the steps of: a computer, in response to input of a first string in a first language, predicting and outputting at least a second string in the first language and a related string, which is a string in the second language that has a predetermined relationship with a concatenated string, which is a string obtained by concatenating the first string and the second string; and a computer, in response to output of at least the second string in the predicting step, executing a process of outputting the related string output in the predicting step, or a process of inputting a subsequent input string following the first string into the predicting step, according to whether or not a subsequent input string input following the first string and the second string satisfy a predetermined condition.
[0015] The above and other objects, features, aspects and advantages of the present invention will become apparent from the following detailed description of the invention taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0016] [Figure 1] FIG. 1 is a block diagram showing the functional configuration of a pre-reading translation device according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a schematic diagram showing an example of learning data for training the predictive translation unit shown in FIG. [Figure 3] FIG. 3 is a flowchart representing a control structure of a program for implementing the pre-reading translation apparatus shown in FIG. 1 in cooperation with computer hardware. [Figure 4] FIG. 4 is a schematic diagram showing an example of a translation process according to the prior art. [Figure 5] FIG. 5 is a schematic diagram showing an example of a translation process performed by the pre-reading translation device shown in FIG. [Figure 6] FIG. 6 is an external view of a computer system for realizing a pre-reading translation device according to the first embodiment of the present invention. [Figure 7] FIG. 7 is a hardware block diagram showing an example of the configuration of the computer system shown in FIG. DETAILED DESCRIPTION OF THE INVENTION
[0017] In the following description and drawings, the same parts are designated by the same reference numerals, and therefore detailed descriptions thereof will not be repeated.
[0018] 1. First embodiment 1.1 Overview of the Pre-reading Translation System 50 Referring to FIG. 1 , the pre-reading translation system 50 according to the first embodiment of the present invention has the following general functions. Specifically, the pre-reading translation system 50 receives a speech signal 60 in a first language and outputs a translation 66 in a second language that corresponds to the semantic content of the speech signal 60. The pre-reading translation system 50 receives the initial portion (referred to as a first string) of the speech recognition result for the speech signal 60 (a character string in the first language) and predicts the subsequent character string in the first language (referred to as a second string). The pre-reading translation system 50 further generates and outputs a translation in the second language of a character string formed by concatenating the first and second character strings (referred to as a concatenated character string). This translation is referred to as a related character string, meaning a character string related to the first or concatenated character string. Based on the initial first character string of the speech signal 60, the pre-reading translation system 50 has the function of predicting the subsequent second character string in the first language and the related character string for the concatenated character string.
[0019] The look-ahead translation system 50 further compares the character string input following the first character string (referred to as the "subsequent input character string") with the second character string as a result of speech recognition of the input speech signal. The look-ahead translation system 50 calculates the degree of match between the second character string and the subsequent input character string and determines whether the condition that this degree of match is greater than a predetermined threshold is met. If this condition is not met, the look-ahead translation system 50 repeats the above-described prediction of the second character string and related character string for the character string that follows the first character string and the subsequent input character string. If this condition is met, the look-ahead translation system 50 outputs the related character string generated by the look-ahead translation system 50, discards the remaining part of the speech signal that begins with the first character string, and moves on to processing the next utterance.
[0020] The degree of agreement between the second character string and the subsequent input character string can be determined, for example, by using the degree of agreement between the words contained in both characters, or by using a machine learning model for determining whether the meanings of the two sentences match. This machine learning model can be trained using training data consisting of the two sentences and labels indicating whether they have the same meaning.
[0021] In the example shown in this embodiment, the first language is Japanese and the second language is English. However, this invention is not limited to such an embodiment and can be applied to any combination of languages. As will be described later, it is also possible to specify the language to be translated.
[0022] 1.2 Configuration of the pre-reading translation system 50 1.2.1 Overall structure 1, a pre-translation system 50 according to a first embodiment of the present invention includes an automatic speech recognition device 62 that, in response to an input speech signal 60 in a first language, performs speech recognition on the speech signal 60 and outputs an input string 100 in the first language, and a pre-translation device 64 that receives the input string 100 from the automatic speech recognition device 62, performs pre-translation from the first language to a second language as described below, and outputs a translation 66 consisting of a string obtained by translating the input string 100. In this embodiment, the language into which the input is to be translated is specified by adding a token to the beginning of the input sentence.
[0023] 1.2.2 Automatic Speech Recognition Device 62 The automatic speech recognizer 62 may be any type that is in practical use for the first language. However, it is preferable that the automatic speech recognizer 62 has the function of chunking the character string resulting from speech recognition and outputting each chunk. A chunk is a group of characters that expresses a certain meaning. It is preferable that the granularity of the chunks can be adjusted. The end of a sentence that constitutes an utterance is always the end of a chunk.
[0024] The automatic speech recognizer 62 may have a function to divide the speech recognition result into chunks by itself and output the chunks, or the automatic speech recognizer 62 may not have such a function, and a function to chunk the speech recognition result may be provided downstream of the output of the automatic speech recognizer 62. This embodiment utilizes such a function, as will be described later.
[0025] To determine the end of a chunk (chunk separation point), a neural network can be used that has been trained in advance using training data in which chunk end labels are attached to the ends of chunks and sentence end labels are attached to the ends of sentences. The granularity of chunks can be controlled by how chunk end labels are attached in the training data. The end of a sentence is always the end of a chunk.
[0026] In this embodiment, the string segmentation unit 80 also has the function of adding a token specifying the language to be translated to the beginning of the string of characters in the output sentence. In the following explanation, a token specifying translation into English will be called 2en, a token specifying translation into Japanese will be called 2jp, etc.
[0027] 1.2.3 Pre-reading translation device 64 1.2.3.1 Overall configuration of the pre-reading translation device 64 The predictive translation device 64 includes a string segmentation unit 80 that receives an input string 100 that is a speech recognition result from the automatic speech recognition device 62, chunks the input string 100, and outputs it as a string 102; and a predictive translation unit 82 that receives the chunked string 102 from the string segmentation unit 80, generates a predicted string 106 that follows the string 102, and a string (concatenated string) that is a concatenation of the string 102 and the predicted string 106, and outputs the translation string 108.
[0028] The pre-reading translation device 64 further includes a selective execution unit 84 that receives the character string 102 from the character string segmentation unit 80 and the predicted character string 106 and the translated character string 108 from the predictive translation unit 82, and outputs the translated character string 108 as the translation 66 if the translated character string 108 is suitable as a translation of the utterance beginning with the character string 102. If the translated character string 108 is not suitable as a translation of the utterance, the selective execution unit 84 discards the translated character string 108. In this case, the selective execution unit 84 further inputs to the predictive translation unit 82 the subsequent character string following the character string input to the predictive translation unit 82 in the immediately preceding process, from the character string 102 from the character string segmentation unit 80, to regenerate the predicted character string 106 and the translated character string 108.
[0029] 1.2.3.2 String splitter 80 As described above, the string segmentation unit 80 has the function of segmenting the input string 100 into chunks and inputting them as string 102 to the comparison unit 90 and predictive translation unit 82. At this time, the string segmentation unit 80 adds a token indicating the end of a sentence to a chunk that corresponds to the end of a sentence.
[0030] 1.2.3.3 Predictive Translation Unit 82 The predictive translation unit 82 is composed of a large-scale language model that is equivalent to a so-called generative AI (Artificial Intelligence) that is made up of a decoder-type neural network. When the predictive translation unit 82 receives input of the initial character string of a Japanese sentence, it at least predicts the subsequent character string of the sentence, and further has the function of generating a translation of the character string obtained by concatenating the input character string and the predicted character string into the language specified by the initial token of the character string 102.
[0031] 2 shows an example of training data 200 used to train the predictive translation unit 82. Referring to FIG. 2, the training data 200 in this example includes many pairs of Japanese sentences and sentences in various languages, such as pairs of bidirectional translations created from Japanese-English translations.
[0032] In Figure 2, the first token of each pair (2en, 2ja, 2zh, 2ko, etc.) is the token that specifies the language to which it is to be translated.
[0033] Taking a Japanese-English pair as an example, for a given translation, training data called "2en Japanese English" and training data called "2ja English Japanese" are generated. is a token that indicates the separation part of a sentence. The first training data is training data for generating English sentences from Japanese, and the second training data is training data for generating Japanese sentences from English.
[0034] It is known that training a large-scale language model using such multilingual training data makes it possible to generate a single translation engine capable of translating multiple languages. Furthermore, the predictive translation unit 82 has the function of generating a character string corresponding to the latter half of the training data from a portion of the beginning of the input. This is because, by training the predictive translation unit 82 using training data such as that shown in FIG. 2, even if the input to the predictive translation unit 82 is only partway through the original text, any number of predictions can be generated from the internal state of the predictive translation unit 82 by beam search, such as "input → original text prediction n + translated text prediction n." In this embodiment, the number of predictions is N. These N predictions change moment by moment as the input continues.
[0035] In this embodiment, the leading portion of the predicted character string 106 output by the predictive translation unit 82 does not include the leading portion of the character string 102 received from the character string segmentation unit 80. The leading portion of the character string 102 is added to the beginning of the predicted character string 106 by the selective execution unit 84. However, the present invention is not limited to such an embodiment. The leading portion of the character string 102 may also be included in the predicted character string 106.
[0036] 1.2.3.4 Selective Execution Unit 84 The selective execution unit 84 includes a comparison unit 90 that compares the predicted character string 106 output from the predictive translation unit 82 with the portion of the character string 102 output by the character string segmentation unit 80 that follows the initial portion, calculates the degree of match between the two, and outputs a permission / prohibition signal 94 indicating whether to permit or prohibit the output of the translated character string 108 output by the predictive translation unit 82 according to whether the degree of match is greater than a threshold value; and an output control unit 92 that is connected to the predictive translation unit 82 to receive the translated character string 108 from the predictive translation unit 82, and that selectively executes a process of outputting the translated character string 108 as a translated sentence 66 or a process of discarding the output of the translated character string 108 in response to receiving the permission / prohibition signal 94 from the comparison unit 90.
[0037] The permission / prohibition signal 94 is also provided to the predictive translation unit 82. When the value of the permission / prohibition signal 94 prohibits the output of the translation string 108, the predictive translation unit 82 further reads the character string of the chunk that follows the character string that was processed immediately before from the character string 102, and outputs the predicted character string 106 and the translation string 108. When the value of the permission / prohibition signal 94 permits the character string of the translation string 108, the predictive translation unit 82 discards the character string that has already been read and the character string of the sentence that follows that character string, and waits for the character string of the next sentence from the character string segmentation unit 80.
[0038] 1.2.3.5 Program Structure 3 is a flowchart showing the control structure of a program that implements the read-ahead translation device 64 in cooperation with computer hardware. Referring to FIG. 3, this program includes step 250, which acquires the output of the automatic speech recognition device 62, and step 252, which breaks down the input character string 100 acquired in step 250 into chunks. In step 252, the program reads and accumulates the input character string 100 sequentially from the beginning, and outputs the character string read up to that point as a single chunk each time it detects a character that marks the end of a chunk. Similarly, each time it detects a character that marks the end of a sentence, it outputs the character string read up to that point as a single chunk, adding a tag to the end.
[0039] Following step 252, the program further includes step 254 for saving the obtained chunk in buffer CH, and step 256 for branching the flow of control depending on whether the chunk at the position of interest (being processed) in buffer CH is the end of a sentence or not.
[0040] This program includes step 258, in which, when the determination in step 256 is positive, i.e., when the chunk at the position of interest (being processed) in buffer CH is the end of a sentence, each chunk stored in buffer CH is read out, shaped into a sentence in the input language (Japanese in this example), and input to predictive translation unit 82, thereby translating it into a sentence in the target language; and, following step 258, step 260, in which buffer CH is cleared and control is returned to step 250.
[0041] This program further includes step 262, when the determination in step 256 is negative, of predicting (reading ahead) the input sentence using predictive translation unit 82 based on the contents of buffer CH and, based on the result, generating a translation sentence in English, which is specified as the translation target language, and step 263, of branching the control flow depending on whether the chunk being processed is the end of the chunk sequence. If the chunk sequence does not contain the end of the sentence, the next chunk sequence is required for processing. Therefore, this program includes step 264, when the determination in step 263 is positive, of obtaining the next output of automatic speech recognizer 62, and step 266, of dividing the character string obtained in step 264 into chunk sequences.
[0042] This program further includes step 268, which compares the input sentence predicted in step 262 with the string of chunks up to the end of the sentence following the chunk being processed and calculates the degree of match when the determination in step 263 is negative, and when the determination in step 263 is positive and the processing in steps 264 and 266 is completed, and step 270, which branches the flow of control in accordance with the degree of match calculated in step 268. In step 270, if the degree of match calculated in step 268 is greater than the threshold value, control returns to step 254. As a result, the processing of step 254 is executed for the first chunk in the string of chunks divided in step 266.
[0043] The program further includes step 274, which, when the determination in step 270 is affirmative, clears the buffer CH, skips chunks up to the end of the chunk string being processed, thereby discarding the character string being processed, and returns control to step 254.
[0044] 1.3 Operation of the pre-reading translation system 50 1.3.1 Operation of the Prior Art Before explaining the operation of the above-mentioned pre-reading translation system 50, for comparison, a display example when translation is performed for each chunk in a device according to the prior art will be explained.
[0045] Referring to Figure 4, let us assume that the chunk sequence processed in the prior art includes chunks 300, 304, 308, 312, and 316. First, chunk 300 is translated, and translation result 302 is displayed. Next, chunk 304 is translated, and translation result 306 is displayed. Similarly, translation result 310 is displayed as the translation of chunk 308, translation result 314 is displayed as the translation of chunk 312, and translation result 318 is displayed as the translation of chunk 316, in that order. After the translation of all chunks is complete, translation sentence 320 is displayed as the result of translating a sentence combining all of these chunks.
[0046] As is clear from this result, when translation results 302 to 318 for each chunk are displayed in order, it is difficult to understand the overall translation. Only when translation 320 is displayed after processing for all chunks has been completed does the meaning of the original text finally become comprehensible.
[0047] 1.3.2 Operation of the pre-reading translation system 50 In contrast, the pre-reading translation system 50 of the present application operates as follows, making it possible to express the meaning of the source text in a target language through translation before all processing of the source text is completed. Note that the target language is assumed to be specified in advance.
[0048] 1, automatic speech recognition device 62 performs speech recognition on speech signal 60, and sequentially provides character strings resulting from the speech recognition to character string segmentation unit 80 as input character strings 100. Character string segmentation unit 80 acquires this input character string 100 (step 250 in FIG. 3; hereinafter, each step number is as shown in FIG. 3), breaks it down into a sequence of chunks (step 252), and provides this as character string 102 consisting of the sequence of chunks to predictive translation unit 82 and comparison unit 90 of selective execution unit 84. Of this sequence of chunks, character string segmentation unit 80 adds a token indicating the translation target language to the beginning of a chunk that corresponds to the beginning of a sentence, and adds a token indicating the end of the sentence to a chunk that corresponds to the end of the sentence.
[0049] The predictive translation unit 82 reads the first chunk of the string of chunks input as the character string 102 and saves it in the buffer CH (254). The predictive translation unit 82 further determines whether this chunk is the end of a sentence or not, and branches control according to the result (step 256). In the current example, it is assumed that the first chunk is the beginning of a sentence but not the end of the sentence. Since the determination in step 256 is negative, control proceeds to step 262. In step 262, the predictive translation unit 82 predicts the sentence that will follow in the source text based on the contents of the buffer CH and the words stored in the buffer CH, and further translates the source text, which is a combination of the chunk being processed and the prediction result, into the target words and outputs the translated text.
[0050] A specific example is shown in Figure 5. Assume that the original text 400 is the target of translation. In this example, the original text 400 is divided into a sequence of chunks: "Tanaka-san," "At the cafe," "While reading a book," "Tasting the coffee," and "It's fun."
[0051] First, input string 410, with 2en added to the beginning of the first chunk, is the target of processing in step 262. An example of a prediction result for the succeeding sentence for input string 410 and its translation result is shown as output string 412. Output string 412 includes a prediction result and a translation result separated from each other by a token. Referring again to FIG. 3, in step 262, of output string 412, the predicted string is output to comparison unit 90 as predicted string 106, and the translation result is output to output control unit 92 as translation string 108.
[0052] Returning to FIG. 3 , following step 262, in step 263, it is determined whether the chunk being processed is the end of the chunk sequence. In this example, since the first chunk is being processed, the determination is negative, and control proceeds to step 268. In step 268, the prediction result in step 262 is compared with the subsequent chunk sequence, and the degree of match between the two is calculated. Because the input string 410 is limited, the output string 412 differs from the input string 410 except for its beginning. As a result, the determination in step 270 is negative. That is, the comparison unit 90 sets the value of the enable / disable signal to a value indicating disable. Control returns to step 254. As a result, the output control unit 92 does not output the translation string 108.
[0053] In step 254, the first chunk in the chunk string being processed (in the current example, the second chunk in the input chunk string) (i.e., the character string "at the cafe" shown in FIG. 5) is saved in buffer CH. The determination in step 256 is negative, and in step 262, based on the contents of buffer CH (input character string 414 shown in FIG. 5), the subsequent part of the input sentence is pre-read by prediction, and a translation is generated using the result. The pre-read input sentence and the translation generated based on the result are output character string 416 shown in FIG. 5. Control proceeds to step 263. Here too, the determination in step 263 is negative, and step 268 is executed. Because output character string 416 differs from the character string following input character string 414 in original text 400, the determination in step 270 is negative, and control returns to step 254. As a result, the next chunk (the character string "while reading a book" shown in FIG. 5) becomes the next target for processing.
[0054] This process is repeated until the determination in step 270 becomes positive. In the example shown in FIG. 5 , for example, when output string 420 is obtained for input string 418, the determination in step 270 becomes negative. When output string 424 is obtained for input string 422, the determination in step 270 becomes positive. As a result, when processing up to input string 422 is completed, the process in step 272 in FIG. 3 is executed. That is, the translation obtained for the predictive input in step 262 from output string 424 is displayed on a display device or synthesized as speech, for example. At this time, the permission / prohibition signal 94 shown in FIG. 3 becomes a value indicating permission, and the translation string 108 output by the predictive translation unit 82 is output as translation 66. Furthermore, in response to the permission / prohibition signal 94 becoming a value indicating permission, in step 274, the predictive translation unit 82 skips up to the end-of-sentence chunk in the chunk sequence being processed and resumes the above-described process from the first chunk of the next sentence. As a result, the above-described processing can be omitted up to the end-of-sentence chunk.
[0055] If the determination in step 270 is not positive until the end of the sentence, the determination in step 256 becomes positive, and in step 258, translation is performed based on the contents of buffer CH, and the result is output. That is, the permission / prohibition signal 94 shown in FIG. 3 becomes a value indicating permission, and the translation string 108 output by predictive translation unit 82 is output as translated sentence 66. Furthermore, in step 260, buffer CH is cleared. Thereafter, control returns to step 250, where a new output of automatic speech recognition device 62 is obtained, and step 252 and the subsequent processes are executed for the next chunk string.
[0056] 1.4 Computer implementation Fig. 6 is an external view of a computer system 600 that realizes the pre-reading translation system 50 (Fig. 1) according to the first embodiment. Fig. 7 is a hardware block diagram of the computer system 600. The hardware configuration of the computer system 600 will be described below.
[0057] 6, this computer system 600 includes a computer 650 having a DVD (Digital Versatile Disc) drive 662, and a keyboard 654, a mouse 656, a monitor 652, a microphone 660, and a pair of speakers 658 for interacting with a user, all of which are connected to the computer 650. These are examples of devices for interaction, and any general hardware and software (e.g., a touch panel, voice input, or a general pointing device) that can be used for interacting with a user can also be used.
[0058] 6 and 7 , the computer 650 includes a central processing unit (CPU) 710, a graphics processing unit (GPU) 712, and a bus 720 connected to the CPU 710, the GPU 712, and the DVD drive 662, in addition to a DVD drive 662. The computer 650 further includes a read-only memory (ROM) 714 connected to the bus 720 and storing a boot-up program of the computer 650, a random access memory (RAM) 716 connected to the bus 720 and storing instructions constituting a program, a system program, working data, and the like, and a solid state drive (SSD) 718, which is nonvolatile memory, connected to the bus 720. The SSD 718 is used to store programs executed by the CPU 710 and the GPU 712, as well as data used by the programs executed by the CPU 710 and the GPU 712. The computer 650 further includes a network I / F (Interface) 726 that provides connection to a network enabling communication with other terminals, and a USB port 664 to which a USB (Universal Serial Bus) memory 702 can be attached / detached and that provides communication between the USB memory 702 and each part within the computer 650.
[0059] The computer 650 further includes an audio I / F 722 that is connected to the microphone 660, the speaker 658, and the bus 720, and has the function of reading out audio signals, image signals, and text data generated by the CPU 710 and stored in the RAM 716 or the SSD 718 according to instructions from the CPU 710, converting them to analog, amplifying them, and driving the speaker 658, and digitizing the analog audio signal from the microphone 660 and storing it at any address in the RAM 716 or the SSD 718 specified by the CPU 710.
[0060] In the above embodiment, the programs and the like that realize the various functions of the pre-reading translation system 50 according to the first embodiment are all stored in, for example, the ROM 714, SSD 718, DVD 700, or USB memory 702 shown in Fig. 7, or in a storage medium of an external device (not shown) connected via the network I / F 726 and the network 704. Typically, these data and parameters are written to the SSD 718 from the outside, for example, and loaded into the RAM 716 when the computer 650 is executed.
[0061] Computer programs for operating this computer system 600 to realize the functions of each component of the pre-reading translation system 50 are stored on a DVD 700 inserted into a DVD drive 662, and are transferred from the DVD drive 662 to the SSD 718. Alternatively, these programs may be stored on a USB memory 702, which may be inserted into a USB port 664 and the programs transferred to the SSD 718. Alternatively, the programs may be transmitted to the computer 650 via the network 704 and stored on the SSD 718.
[0062] The program is loaded into RAM 716 when executed. A machine learning model such as a deep neural network is used in the automatic speech recognition device 62, the string segmentation unit 80, and the predictive translation unit 82 shown in Fig. 1. In the computer system 600, a machine learning model that has already been trained in another device may be used, or the computer system 600 may be used as a training device to train a machine learning model.
[0063] The CPU 710 reads a program from the RAM 716 according to an address indicated by an internal register called a program counter (not shown) and interprets the instructions. The CPU 710 reads data required to execute the instructions from the RAM 716, the SSD 718, or another device according to the address specified by the instruction, and executes the processing specified by the instruction. The CPU 710 stores the execution result data at an address specified by the program, such as in the RAM 716, the SSD 718, or a register within the CPU 710. Depending on the address, the execution result data is output from the computer to an external device via, for example, the network I / F 726. At this time, the program counter value is also updated by the program. The computer program may be loaded directly into the RAM 716 from the DVD 700, the USB memory 702, or via the network 704. Note that some tasks (mainly numerical calculations) of the program executed by the CPU 710 are issued to the GPU 712 according to instructions included in the program or according to the analysis results obtained when the CPU 710 executes the instructions.
[0064] The program that enables the computer 650 to implement the functions of each of the above-described parts of the pre-reading translation system 50 includes a plurality of instructions written and arranged to cause the computer 650 to operate to implement those functions. Some of the basic functions required to execute these instructions may be provided by the operating system (OS) or third-party programs running on the computer 650, by various toolkit modules installed on the computer 650, or by the program execution environment. Therefore, the program does not necessarily include all of the functions required to implement the system and method of this embodiment. The program may include only instructions that execute the operations of the above-described devices and their components by statically linking appropriate functions or modules at compile time or by dynamically calling them at runtime in a controlled manner to achieve the desired results. The method for operating the computer 650 to achieve this is well known. Therefore, a description of the method for operating the computer 650 will not be repeated here.
[0065] 1.5 Effects of the First Embodiment In the example shown in FIG. 5, the actual translation is not output until the latter half of the original sentence 400, so it takes a relatively long time for the translation to be displayed. This is because the original sentence 400 is a relatively long sentence. If the original sentence 400 is a shorter sentence or if the internal state of the predictive translation unit 82 is suitable for prediction based on the content of the input sentence up to the input sentence, the determination in step 270 may become positive at a relatively early stage, and a translation based on the predicted input sentence may be output. As a result, the possibility of a translation that does not adequately reflect the content of the input sentence being displayed, as in the prior art, is reduced. Furthermore, by predicting and displaying an appropriate translation at an early stage, processing for the input sentence can be terminated early and processing for the next input sentence can be started. As a result, the amount of calculation is reduced, and a translation that expresses the meaning of the input text can be obtained almost in real time.
[0066] In the first embodiment, the translation result is not displayed unless the degree of match between the predicted input and the actual input is equal to or greater than a threshold value. However, the present invention is not limited to such an embodiment. The translation result may be always displayed regardless of the degree of match between the predicted input and the actual input.
[0067] 2. Variations 2.1 First Modification In the above embodiment, the translation of the input sentence is displayed only when the determination in step 270 in Figure 5 is positive or when processing of the input sentence has been completed to the end of the sentence. However, the present invention is not limited to such an embodiment. For example, one or more processes similar to step 270 may be provided before step 270, and the thresholds in those processes may be set to be smaller than the threshold in step 270.
[0068] With this configuration, when the degree of match between the predicted sentence and the input sentence is smaller than the threshold value set in step 270 but still has a certain degree of reliability, a translation based on the predicted sentence can be displayed. In this case, processing of the chunk sequence being processed continues unless the determination in step 270 is affirmative. Also, in this example, the translation displayed at each stage becomes a single sentence that is somewhat consistent with the input content, and the content is updated so that it gradually approaches the input sentence. As a result, compared to the prior art where individual chunks are displayed separately, the user can more easily and naturally understand the content of the original text.
[0069] 2.2 Second variant In the above embodiment, each chunk and the sentence end are used for processing. However, the present invention is not limited to such an embodiment. For example, when sentences are connected by conjunctions, it is also possible to detect the boundaries between the sentences and process them separately. To do this, all that is required is a neural network with a configuration similar to the neural network for detecting sentence ends described above, trained using learning data for detecting not only sentence end but also boundaries between sentences connected by conjunctions, etc.
[0070] 2.3 Third variant In the above embodiment, it is assumed that an input sentence is translated into only one language. However, the present invention is not limited to such an embodiment. By providing multiple predictive translation units 82 and preparing the same input with tokens indicating translation into different languages and inputting them into each of these multiple predictive translation units 82, it becomes possible to simultaneously translate an input sentence into multiple languages. In this case, a simultaneous automatic translation system for multilingual conferences can be realized by providing a device that automatically identifies the language of the speech signal 60 and, based on the results, attaching a token to each input specifying a language different from the language of the input speech as the translation target language.
[0071] 2.4 Fourth Variant In the above embodiment, a translation of a predicted input sentence is output only when the degree of match between the predicted input sentence and the actual subsequently input sentence exceeds a threshold value. However, the present invention is not limited to such an embodiment. A reliability level is assigned to the predicted input sentence. A translation of the predicted input sentence may be output when the reliability level exceeds a certain threshold value. Alternatively, these two may be combined, and a translation of the predicted input sentence may be output when both the reliability level and the degree of match exceed their respective threshold values.
[0072] 2.5 Fifth Variant In the above embodiment, for each chunk that has been speech-recognized, the input (chunk sequence) that follows that chunk is predicted. However, this embodiment is not limited to this. The latter half of a speech that has been speech-recognized may be predicted when the input of the first half of that speech is completed. In short, this modification divides the speech into two chunks, a first half and a second half.
[0073] To divide an utterance into two chunks, the first half and the second half, the time from the beginning of the utterance when each phoneme of the speech signal is input to the speech recognizer is assigned to each chunk of the speech recognition result. The first half and the second half of the utterance can be divided by comparing the duration to the end of the utterance with the elapsed time from the beginning of the utterance to the end of each chunk output by the speech recognizer. In this case, it is desirable to select the end of the chunk whose end point is closest to the boundary between the first and second halves of the utterance rather than simply determining the division point based on time.
[0074] 2.6 Sixth Variant The sixth modification is an application of the above-described embodiment to language learning support such as English composition. The system disclosed in the first embodiment can be used. Learning is performed as follows.
[0075] (1) The learner enters the original text partway through (2) The pre-reading translation system 50 completes the source text up to the end of the sentence and performs pre-reading translation. (3) The pre-translation system 50 reverse-translates the results of the pre-translation. (4) Compare the original text completed by the pre-reading translation system 50 in (2) above with the result of the back-translation obtained in (3). (5) If the two match and are correct, the translation obtained in (2) is successfully created; otherwise, the source text is further input. (6) Repeat steps (1) to (4) above until the translation is successfully completed.
[0076] When using conventional machine translation, the entire original text is input, and then translation and back-translation are performed to determine whether the translation is correct. However, with this modified version, the original text is completed when it has been input partially, and the translation result and its back-translation are obtained, and the two can be compared. If the two match, the translation is obtained, and there is no need to input the original text in its entirety.
[0077] This modification improves input efficiency by complementing the original text. Furthermore, learners can confirm what kind of original text they should input to increase the probability of successful translation, allowing them to efficiently learn how to create sentences that are easy to translate.
[0078] 2.7 Seventh Variant The seventh modified example relates to a device that paraphrases a long input sentence into easier-to-understand sentences in almost real time. The system in this modified example can be the same as that in the first embodiment. However, it differs from the first embodiment in that, instead of the predictive translation unit 82, a large-scale language model is used to predict the following part based on the beginning of the input sentence, and then to "translate" the resulting sentence into easier-to-understand Japanese. Except for the different language model used, this seventh modified example can be realized using almost the same mechanism as the first embodiment. The language model in this seventh embodiment can be trained in the same way as the predictive translation unit 82 in the first embodiment.
[0079] This seventh variant can be used, for example, in schools or government offices, when a native Japanese speaker speaks with someone who is not yet fluent in Japanese. Phrases that seem common sense to native Japanese speakers are often completely incomprehensible to non-native Japanese speakers. When a speaker does not know which part is difficult to understand, they are unable to paraphrase appropriately, making communication difficult. In such cases, the system according to this seventh variant can be used to facilitate mutual communication. Compared to simply translating speech into the language of the interlocutor, the use of the system according to this seventh variant may facilitate the acquisition of Japanese by non-native Japanese speakers.
[0080] 3. Summary As described above, according to each embodiment of the present invention, a subsequent portion of an original text is predicted based on the input of a portion of the original text, and the predicted result is then converted to generate a sentence having a predetermined relationship with the original text, such as a translation or a rephrased sentence using simpler expressions, in almost real time. If the predicted portion of the original text closely matches the actually input character string, the converted sentence obtained from the prediction can be used. For example, in simultaneous translation of spoken audio, compared to simply translating only a portion of the original text, such as speech, the entire content of the speech can be output before the input of the original text is completed. As a result, a character string processing device and a character string processing method can be provided that can convert an input character string into another character string having a predetermined relationship to the input character string while maintaining real-time performance.
[0081] The embodiments disclosed herein are merely examples, and the present invention is not limited to the above-described embodiments. The scope of the present invention is defined by the claims in the appended claims, taking into consideration the detailed description of the invention, and includes all modifications within the meaning and scope equivalent to the wordings described therein. [Explanation of symbols]
[0082] 50 Pre-reading translation system 60 Audio Signal 62 Automatic speech recognition device 64 Pre-reading translation device 66, 320 translations 80 String splitter 82 Predictive Translation Department 84 Selective Execution Department 90 Comparison Section 92 Output control section 94 Permission / Prohibition Signal 100, 410, 414, 418, 422 Input string 102 Strings 106 Predicted Strings 108 Translation Strings 200 training data 300, 304, 308, 312, 316 chunks 302, 306, 310, 314, 318 Translation results 400 original text 412, 416, 420, 424 Output string
Claims
1. a prediction unit that, in response to an input of a first character string in a first language, predicts and outputs at least a second character string in the first language and a related character string that is a character string in the second language and has a predetermined relationship with a concatenated character string that is a character string obtained by concatenating the first character string and the second character string; an execution unit that, in response to the prediction unit outputting at least the second character string, executes a process of outputting the related character string output by the prediction unit, or a process of inputting the subsequent input character string after the first character string to the prediction unit, according to whether a subsequent input character string input following the first character string and the second character string satisfy a predetermined condition.
2. 2. The character string prediction device according to claim 1, wherein the prediction unit includes a predictive translation unit that, in response to an input of the first character string, predicts and outputs at least the second character string and the related character string that is a translation of the concatenated character string in the second language.
3. the second language is a language different from the first language; The character string prediction device according to claim 2 , wherein the predictive translation unit includes a pre-trained language model that predicts and outputs at least the second character string and the related character string in response to an input of the first character string.
4. 2. The character string prediction device according to claim 1, wherein the execution unit includes a selective execution unit that, in response to the prediction unit outputting at least the second character string, executes a process of outputting the related character string output by the prediction unit or a process of inputting the subsequent input character string to the prediction unit following the first character string, according to whether a condition that a degree of match between the subsequently input character string and the second character string is greater than a threshold is satisfied.
5. 5. The character string prediction device according to claim 4, further comprising: an input discarding unit configured to discard the second character string and the character string in the first language input to the character string prediction device following the first character string in response to a condition that the degree of match is greater than the threshold value being satisfied.
6. a step in which, in response to an input of a first character string in a first language, the computer predicts and outputs at least a second character string in the first language and a related character string, which is a character string in the second language and has a predetermined relationship with a concatenated character string, which is a character string obtained by concatenating the first character string and the second character string; a step of executing, by a computer, in response to output of at least the second character string in the predicting step, a process of outputting the related character string output in the predicting step, or a process of inputting the subsequent input character string following the first character string into the predicting step, according to whether or not a subsequent input character string input following the first character string and the second character string satisfy a predetermined condition.
Citation Information
Patent Citations
Machine translation device, method, and program
JP2016071761A