Translation device
The translation device addresses the challenge of translating noisy sentences by generating a single normalization/translation model from combined learning data, improving processing speed and accuracy while reducing computational costs.
Patent Information
- Application Number
- JP2021565556
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-17
- Filing Date
- 2020-12-11
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-12-11
AI Technical Summary
Existing translation devices face challenges in accurately translating sentences with noise, such as fillers or paraphrases, especially between languages with insufficient corpus data, leading to high computational costs and low translation accuracy.
A translation device that stores learning data associating source texts with normalized and translated texts, using a normalization learning unit and a translation learning unit to generate a single normalization/translation model that outputs normalized and translated texts, with translation learning performed after normalization learning to reduce noise influence.
This approach improves processing speed and accuracy by generating a single model from combined learning results, reducing computational costs, and enhancing translation accuracy by suppressing noise during the translation process.
Smart Images

Figure 0007696296000001 
Figure 0007696296000002 
Figure 0007696296000003
Abstract
Description
Technical Field
[0001] One aspect of the present invention relates to a translation device.
Background Art
[0002] Conventionally, a technique for improving the translation accuracy of a translation device by learning a translation sentence for an input sentence (for example, natural speech) is known (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Here, for translation between languages without a sufficient amount of corpus (for example, Japanese ⇔ Chinese, etc.), there is a problem that a sentence (natural speech) including noise such as a filler or a paraphrase cannot be accurately translated. For such a problem, for example, it is conceivable to perform translation using a translation model after removing noise from natural speech using a normalization model (a model that grammatically correctly converts natural speech).
[0005] However, when using a plurality of independent models as described above, the computational cost at both the time of model generation (learning time) and the time of model utilization (translation time) becomes high, and the processing takes time. Further, since they are still separate models, the synergistic effect of each model is small, and the translation accuracy has not been sufficiently improved.
[0006] One aspect of the present invention has been made in view of the above circumstances, and an object thereof is to improve the processing speed and accuracy related to translation.
Means for Solving the Problems
[0007] A translation device according to one aspect of the present invention stores a plurality of pieces of learning data in which a source text for learning in a first language, a normalized text for learning obtained by grammatically correctly converting the source text for learning, and a translated text for learning obtained by translating the source text for learning into a second language different from the first language are associated with each other. A normalization learning unit that learns by combining the source text for learning and the corresponding normalized text for learning for a plurality of pieces of learning data, a translation learning unit that learns by combining the source text for learning and the corresponding translated text for learning for a plurality of pieces of learning data, and based on the learning results of the normalization learning unit and the translation learning unit, a model generation unit that generates one normalization / translation model configured to be able to output a normalized text for an input text in the first language and a translated text into the second language. For at least some of the learning data, learning by the translation learning unit is performed after learning by the normalization learning unit.
[0008] In the translation device according to one aspect of the present invention, for a plurality of pieces of learning data, a combination of the source text for learning and the corresponding normalized text for learning is learned, and a combination of the source text for learning and the corresponding translated text for learning is learned. Then, based on these learning results, one normalization / translation model that outputs a normalized text and a translated text into the second language from the input text in the first language is generated. In this way, by generating a common one output model (normalization / translation model) from the learning results of normalization and translation, the period required for model generation (the total period required for learning and model generation) can be shortened compared to the case where output models are generated individually, and the output speed of the normalized text and the translated text can be improved. Further, in the translation device according to one aspect of the present invention, for at least some of the learning data, learning by the translation learning unit is performed after learning by the normalization learning unit is performed first. Thereby, for example, in the case where learning is performed using an encoder-decoder model, for at least some of the learning data, translation learning can be performed while suppressing the influence of noise in the source text for learning by using the parameters learned in the normalization learning (that is, the parameters suitable for normalization). This can improve the translation accuracy in the normalization / translation model.
Advantages of the Invention
[0009] According to one aspect of the present invention, the processing speed and accuracy related to translation can be improved.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the description of the drawings, the same or equivalent elements are denoted by the same reference numerals, and redundant descriptions are omitted.
[0012] First, referring to FIGS. 1 and 2, the outline of the translation device according to this embodiment will be described. FIG. 1 is a diagram for explaining the outline of the normalization / translation model of the translation device according to this embodiment. As shown in FIG. 1, in the translation device according to this embodiment, the original text in the first language is input into the normalization / translation model, and from the normalization / translation model, the normalized text (normalization result) in the first language and the translated text (translation result) in the second language corresponding to the normalized text are output. That is, in the example shown in FIG. 1, the original text "The improvement of the company's business efficiency is really great. It is achieved by the utilization of IT." in the first language is input into the normalization / translation model, and the normalized text "The utilization of IT can achieve the improvement of the company's business efficiency." and the translated text in the second language are output from the normalization / translation model. The normalized text is a sentence in the first language obtained by grammatically correctly converting the input sentence (the original text in the first language). The first language and the second language are different languages from each other. In the example shown in FIG. 1, the first language is Japanese and the second language is Chinese.
[0013] As shown in FIG. 1, in the translation device according to this embodiment, for example, learning data in which the original text for learning in the first language, which is natural speech, the normalized text for learning obtained by grammatically correctly converting the original text for learning, and the translated text for learning obtained by translating the original text for learning into a second language different from the first language are associated with each other is being learned. More specifically, in the translation device, for a plurality of pieces of learning data, normalization text learning is performed in which the original text for learning as the input and the corresponding normalized text for learning as the output are combined for learning, and for a plurality of pieces of learning data, translation text learning is performed in which the original text for learning as the input and the corresponding translated text for learning as the output are combined for learning. Then, based on these learning results, a single normalization / translation model is generated in the translation device, which is configured to be able to output the normalized text for the input sentence in the first language and the translated text into the second language. Using the normalization / translation model generated in this way, the derivation of the above-described translated text is performed.
[0014] FIG. 2 is a diagram for explaining an overview of the effects of the translation device according to the present embodiment. FIG. 2(a) shows an example of a translated sentence when the technique described in the present embodiment (the technique of the translation device according to the present embodiment) is not used, and FIG. 2(b) shows an example of a translated sentence when the technique described in the present embodiment (the technique of the translation device according to the present embodiment) is used. In the examples shown in FIGS. 2(a) and 2(b), the original sentence shown in FIG. 1 ("Enterprise business efficiency is great, you know. It succeeds by leveraging IT.") is input as the input sentence. In the example shown in FIG. 2(a), since the normalization / translation model described in the present embodiment is not used, the original sentence (input sentence) ("Enterprise business efficiency is great, you know. It succeeds by leveraging IT.") is directly translated. The original sentence is natural speech containing many noises such as fillers and paraphrases. Therefore, when the original sentence is directly translated, as shown in the back translation of the translated sentence in FIG. 2(a), an accurate translation cannot be performed. On the other hand, in the example shown in FIG. 2(b), the normalization / translation model described in the present embodiment is used. After the original sentence (input sentence) is normalized and noises such as fillers and paraphrases are removed, the translation into the second language is performed based on the normalized sentence. As a result, as shown in the back translation of the translated sentence in FIG. 2(b), an accurate translation can be performed for the sentence that is originally intended to be translated.
[0015] Next, with reference to FIG. 3, the configuration of the translation device 10 according to the present embodiment will be described. FIG. 3 is a functional block diagram of the translation device 10 according to the present embodiment. The translation device 10 shown in FIG. 3 is a device that generates a translated sentence in a second language from an input sentence in a first language that is the translation target. As described above, the first language is, for example, Japanese, and the second language is, for example, Chinese. The first language and the second language may be different from each other, and are not limited to natural languages, and may be human languages and formal languages (computer program languages), etc. A sentence is a single unit of a language expression that is unified by one statement and is complete in form. A sentence may be rewritten as something consisting of one or more sentences (for example, a paragraph, an article, etc.).
[0016] As shown in FIG. 3, the translation device 10 includes a storage unit 11, a normalization learning unit 12, a translation learning unit 13, a model generation unit 14, and an evaluation unit 15 as functions related to the learning and generation of the normalization / translation model 70.
[0017] Regarding the functions related to the learning of the translation device 10, reference will be made to FIGS. 4 to 7 for explanation. The translation device 10 generates one normalization / translation model 70 configured to be able to output a normalized sentence for an input sentence in the first language and a translated sentence into the second language by learning a plurality of learning data. Such a normalization / translation model 70 is, for example, a machine translation model (e.g., NMT), and is generated by performing learning using, for example, an encoder-decoder model. The encoder-decoder model is composed of two recurrent neural networks called an encoder and a decoder. The encoder converts the input sequence into an intermediate representation, and the decoder generates a sequence that becomes the output from the intermediate representation.
[0018] FIG. 4 is a diagram for explaining the learning related to normalization and translation of the translation device 10 according to the present embodiment, and is a diagram for explaining the learning using the encoder-decoder model according to the present embodiment. As shown in FIG. 4, in the learning using the encoder-decoder model in the present embodiment, one common encoder related to normalization and translation, a decoder related to normalization (Decoder1 in FIG. 4), and a decoder related to translation (Decoder2 in FIG. 4) are used. The encoder converts an input sentence, which is natural speech, into a fixed-length vector representation. Such a vector representation is an intermediate representation that is passed on to the decoder. In the encoder-decoder model, an attention function is adopted, and the decoder can decode while referring to the history of the hidden state of the encoder. Note that the attention supports the hidden state of the encoder and also has functions such as storing the order of words (word position information).
[0019] FIG. 5 is a diagram for explaining the learning related to the normalization and translation of the translation device according to the comparative example, and is a diagram for explaining the learning using the encoder-decoder model according to the comparative example. As shown in FIG. 5, usually, since the model related to normalization and the model related to translation are generated separately (individually), even in the learning using the encoder-decoder model, the encoder and the decoder are provided individually. Compared with such a case, in the encoder-decoder model according to the present embodiment shown in FIG. 4, since one common encoder related to normalization and translation is used, the calculation cost in learning can be reduced and the processing can be speeded up. In addition, when the second language is a plurality of languages in the normalization / translation model 70, the number of decoders may be increased according to the number of languages. In this way, by learning using the encoder-decoder model, it is possible to easily cope with the case where the second language is a plurality of languages.
[0020] FIG. 6 is a diagram for explaining the learning related to the normalization and translation of the translation device 10 according to the present embodiment, and is a diagram for explaining the learning using the encoder-decoder model according to the present embodiment. As shown in FIG. 6, in the learning using the encoder-decoder model in the present embodiment, for the learning data, after the conversion from the learning source text to the learning normalization text is learned, the conversion from the learning source text to the learning translation text is learned. By learning the conversion from the learning source text to the learning normalization text, it is learned which words in the learning source text are not important (are noise), and the hidden state of the encoder becomes robust to noise. Then, after the conversion to the learning normalization text is learned, by learning the conversion from the learning source text to the learning translation text, it is possible to learn the conversion to the learning translation text while inheriting (utilizing) the hidden state of the encoder learned at the time of the conversion to the learning normalization text. In this way, it is possible to learn the conversion to the learning translation text while suppressing the influence of noise, and improve the translation accuracy.
[0021] FIG. 7 is a diagram for explaining the learning related to the normalization and translation of the translation device according to the comparative example, and is a diagram for explaining the learning using the encoder-decoder according to the comparative example. In the example shown in FIG. 7, unlike the embodiment described with reference to FIG. 6, after the conversion from the learning source text to the learning normalized text is learned, the conversion from the learning source text to the learning translation text is not learned (for example, the conversion to the learning translation text has been learned in advance). In such an embodiment, when learning the conversion to the learning translation text, since the hidden state of the encoder that is robust to noise as described in FIG. 6 cannot be used, noise such as fillers is likely to remain in the translation result. Compared with such an embodiment, as described above, in the embodiment shown in FIG. 6, it is possible to learn the conversion to the learning translation text with the noise influence suppressed, and the translation accuracy can be improved.
[0022] Returning to FIG. 3, the storage unit 11 stores a plurality of pieces of learning data in which the learning source text in the first language, the learning normalized text obtained by grammatically correctly converting the learning source text, and the learning translation text obtained by translating the learning source text into a second language different from the first language are associated with each other. Such learning data is a corpus (database of sentences) in which sentences are associated with each other, constructed for machine learning.
[0023] The normalization learning unit 12 learns by combining the learning source text and the corresponding learning normalized text for a plurality of pieces of learning data. That is, the normalization learning unit 12 learns the conversion from the learning source text to the learning normalized text for each piece of learning data stored in the storage unit 11. The normalization learning unit 12 learns, for example, which words in the learning source text are not important (which words are noise such as fillers). The normalization learning unit 12 and the translation learning unit 13 perform learning alternately with each other. That is, for each piece of learning data, for example, after the learning by the normalization learning unit 12 is performed, the learning by the translation learning unit 13 is continuously performed. In this way, for at least some of the learning data, after the learning by the normalization learning unit 12 is performed, the learning by the translation learning unit 13 is performed.
[0024] The normalization learning unit 12 uses the same encoder as the translation learning unit 13 and uses a separately provided decoder (separate from the decoder used by the translation learning unit 13), and performs learning using an encoder-decoder model. The normalization learning unit 12 may perform learning multiple times repeatedly for each learning data. The normalization learning unit 12 basically performs learning alternately with the translation learning unit 13 as described above, but when the evaluation unit 15 evaluates that the value of the loss function regarding normalization is greater than the first threshold (details will be described later), separately from the learning performed alternately with the translation learning unit 13, it may perform repeated learning for each learning data independently. The normalization learning unit 12 outputs the learning result to the model generation unit 14.
[0025] The translation learning unit 13 learns by combining the learning source text and the corresponding learning translation text for a plurality of learning data. That is, the translation learning unit 13 learns the conversion from the learning source text to the learning translation text for each learning data stored in the storage unit 11. The translation learning unit 13 performs learning alternately with the normalization learning unit 12. That is, for each learning data, for example, after the learning by the normalization learning unit 12 is performed, the learning by the translation learning unit 13 is continuously performed. In this way, for at least some of the learning data, after the learning by the normalization learning unit 12 is performed, the learning by the translation learning unit 13 is performed.
[0026] The translation learning unit 13 uses the same encoder as the normalization learning unit 12 and uses a separately provided decoder (separate from the decoder used by the normalization learning unit 12), and performs learning using an encoder-decoder model. The translation learning unit 13 may perform learning for each learning data using the hidden state of the encoder learned by the normalization learning unit 12. The translation learning unit 13 may perform learning a plurality of times repeatedly for each learning data. The translation learning unit 13 basically performs learning alternately with the normalization learning unit 12 as described above, but when the evaluation unit 15 evaluates that the value of the loss function regarding translation is larger than the second threshold (details will be described later), separately from the learning performed alternately with the normalization learning unit 12, it may perform repeated learning for each learning data independently. The translation learning unit 13 outputs the learning result to the model generation unit 14.
[0027] The model generation unit 14 generates one normalization / translation model 70 configured to be able to output a normalized sentence for the input sentence in the first language and a translated sentence into the second language based on the learning results of the normalization learning unit 12 and the translation learning unit 13. The model generation unit 14 outputs the generated normalization / translation model 70 to the evaluation unit 15 and the translation unit 17.
[0028] The evaluation unit 15 derives a loss function regarding normalization and a loss function regarding translation for the normalization / translation model 70 generated by the model generation unit 14, and evaluates the normalization / translation model 70 based on the values of the respective loss functions. Specifically, the evaluation unit 15 derives the loss function by comparing the softmax output value of each word output on the decoder side with the embedding of the correct word. Although it is common to use softmax cross entropy as the loss function, other loss functions may be used. The loss function is a function that represents the magnitude of the deviation between the prediction and the actual value, and is a function used when evaluating the prediction accuracy of the model. It can be said that the smaller the value of the loss function, the more accurate the model. That is, for the normalization / translation model 70, the smaller the value of the loss function regarding normalization, the higher the accuracy of normalization, and the smaller the value of the loss function regarding translation, the higher the accuracy of translation.
[0029] When, with respect to a plurality of learning data, in the case where learning is repeatedly performed a plurality of times by the normalization learning unit 12 and the translation learning unit 13, at least one of the following conditions is satisfied: the value of the loss function regarding normalization is greater than a predetermined first threshold, and the value of the loss function regarding translation is greater than a predetermined second threshold, the evaluation unit 15 evaluates that the normalization-translation model 70 is in a first state with low prediction accuracy. And when it is evaluated by the evaluation unit 15 that it is in the first state with low prediction accuracy and the value of the loss function regarding normalization is greater than the first threshold, the normalization learning unit 12 performs repeated learning with respect to the learning data alone, separately from the learning performed alternately with the translation learning unit 13. Further, when it is evaluated by the evaluation unit 15 that it is in the first state with low prediction accuracy and the value of the loss function regarding translation is greater than the second threshold, the translation learning unit 13 performs repeated learning with respect to the learning data alone, separately from the learning performed alternately with the normalization learning unit 12.
[0030] As shown in FIG. 3, the translation device 10 includes an acquisition unit 16, a translation unit 17, and an output unit 18 as functions related to translation using the normalization-translation model 70. The functions related to translation are realized on the premise that the normalization-translation model 70 is generated by the functions related to the learning and generation of the normalization-translation model 70 described above.
[0031] The acquisition unit 16 acquires an input sentence in a first language that is the object to be translated. The input sentence may be, for example, a sentence obtained by converting the result of voice recognition of the voice uttered by the user into text. When a voice recognition result or the like is used as the input sentence, the input sentence may include noises such as fillers, paraphrases, and hesitations. The input sentence may be, for example, a sentence input by the user using an input device such as a keyboard. Also in such a case, the input sentence may include noises such as input errors. The acquisition unit 16 outputs the input sentence to the translation unit 17.
[0032] The translation unit 17 has the normalization and translation model 70 generated by the model generation unit 14. The translation unit 17 inputs the input sentence acquired by the acquisition unit 16 into the normalization and translation model 70 to generate a normalized sentence in the first language. Further, the translation unit 17 inputs the normalized sentence into the normalization and translation model 70 to generate a translated sentence in the second language corresponding to the normalized sentence. The translation unit 17 outputs the generated normalized sentence and translated sentence to the output unit 18.
[0033] The output unit 18 outputs the translated sentence. The output unit 18 may output the normalized sentence together with the translated sentence. For example, when the output unit 18 receives the translated sentence from the translation unit 17, it outputs the translated sentence (and the normalized sentence) to the outside of the translation device 10. The output unit 18 may output the translated sentence (and the normalized sentence) to an output device such as a display and a speaker.
[0034] Next, with reference to FIG. 8, the learning process of the translation device 10 will be described. FIG. 8 is a flowchart showing the learning process of the translation device 10.
[0035] As shown in FIG. 8, in the translation device 10, first, each sentence of a plurality of learning data is word-segmented, and one learning data is selected (step S1). The learning data is data in which a learning original text in the first language, a learning normalized text obtained by grammatically correctly converting the learning original text, and a learning translated text obtained by translating the learning original text into a second language different from the first language are associated with each other. Hereinafter, it will be described on the assumption that the learning original text is a natural utterance sentence.
[0036] Subsequently, the translation device 10 learns by combining the learning original text, which is a natural utterance sentence, and the learning normalized text for the selected one learning data, and learns the conversion from the natural utterance sentence to the normalized sentence (step S2). Subsequently, the translation device 10 learns by combining the learning original text, which is a natural utterance sentence, and the learning translated text for the same learning data, and learns the conversion from the natural utterance sentence to the translated sentence (step S3).
[0037] Subsequently, the translation device 10 determines whether or not all the learning data has been learned a predetermined number of times (learning regarding normalization and translation) respectively (step S4). If it is determined in step S4 that there is learning data that has not been learned the predetermined number of times, the process returns to step S1 and is executed again.
[0038] On the other hand, if it is determined in step S4 that all the learning data has been learned the predetermined number of times, the translation device 10 generates one normalization / translation model 70 based on the learning result, and derives a loss function regarding normalization and a loss function regarding translation for the normalization / translation model 70 (step S5).
[0039] Subsequently, the translation device 10 determines whether or not the values of the two derived loss functions are equal to or less than a predetermined threshold value (step S6). That is, the translation device 10 determines whether the value of the loss function regarding normalization is equal to or less than a predetermined first threshold value and the value of the loss function regarding translation is equal to or less than a predetermined second threshold value. If it is determined in step S6 that the value of any of the loss functions is equal to or less than the predetermined threshold value, the learning process ends assuming that the loss function has converged.
[0040] On the other hand, if it is determined in step S6 that the value of at least one of the loss functions is greater than the threshold value, the translation device 10 determines whether or not the number of learning loop times of individual learning (the number of learning loop times including step S8 described later) is equal to or less than a predetermined threshold value (step S7).
[0041] In step S7, when it is determined that the number of learning loop times of individual learning is equal to or less than a predetermined threshold, the translation device 10 performs individual learning on the learning items for which it is determined that the value of the loss function is equal to or less than the predetermined threshold (step S8). Specifically, when the value of the loss function related to normalization is evaluated to be greater than the first threshold (for example, when the loss function is gradually increasing), the translation device 10 separately and repeatedly performs learning related to normalization for each learning data, separately from the learning performed alternately with the learning related to translation. Similarly, when the value of the loss function related to translation is evaluated to be greater than the second threshold (for example, when the loss function is gradually increasing), the translation device 10 separately and repeatedly performs learning related to translation for each learning data, separately from the learning performed alternately with the learning related to normalization. After the individual learning in step S8 is performed, the process returns to step S5 and is executed again.
[0042] On the other hand, in step S7, when it is determined that the number of learning loop times (the number of executions of step S8) of individual learning is greater than a predetermined threshold, the translation device 10 determines that it is not possible to converge both of the two loss functions by individual learning, and performs exception handling (step S9). In the exception handling, the translation device 10 performs learning processing so that the sum of the value of the loss function related to normalization and the value of the loss function related to translation becomes equal to or less than a predetermined threshold (third threshold). When the processing in step S9 is completed, the learning processing ends. In this way, the learning processing ends when the values of the two loss functions become equal to or less than the predetermined threshold (the loss functions converge), or when the sum of the values of the two loss functions becomes equal to or less than the predetermined threshold by exception handling. The above is the learning processing.
[0043] Next, with reference to FIG. 9, the translation processing of the translation device 10 will be described. FIG. 9 is a flowchart showing the translation processing of the translation device 10.
[0044] As shown in FIG. 9, in the translation device 10, first, an input sentence in the first language to be translated is acquired (step S101). Subsequently, the translation device 10 generates a normalized sentence in the first language corresponding to the input sentence by inputting the acquired input sentence into the normalization / translation model 70 (step S102).
[0045] Subsequently, the translation device 10 generates a translated sentence in the second language corresponding to the normalized sentence by inputting the normalized sentence into the normalization / translation model 70 (step S103). Finally, the translation device 10 outputs the generated translated sentence to the outside (step S104). The translation device 10 may output the normalized sentence together with the translated sentence. The above is the translation process.
[0046] Next, the operation and effect of the translation device 10 according to the present embodiment will be described.
[0047] The translation device 10 according to the present embodiment includes a storage unit 11 that stores a plurality of pieces of learning data in which a learning original text in the first language, a learning normalized text obtained by grammatically correctly converting the learning original text, and a learning translated text obtained by translating the learning original text into a second language different from the first language are associated with each other, a normalization learning unit 12 that learns by combining the learning original text and the corresponding learning normalized text for a plurality of pieces of learning data, a translation learning unit 13 that learns by combining the learning original text and the corresponding learning translated text for a plurality of pieces of learning data, and a model generation unit 14 that generates one normalization / translation model 70 configured to be able to output a normalized sentence for an input sentence in the first language and a translated sentence into the second language based on the learning results of the normalization learning unit 12 and the translation learning unit 13. For at least some of the learning data, learning by the translation learning unit 13 is performed after learning by the normalization learning unit 12.
[0048] In the translation device 10 according to this embodiment, for a plurality of learning data, a combination of a learning source text and a corresponding normalized learning text is learned, and a combination of the learning source text and a corresponding translated learning text is learned. Then, based on these learning results, one normalization / translation model 70 that outputs a normalized text and a translated text into a second language from an input text in a first language is generated. In this way, by generating a common single output model (normalization / translation model 70) from the learning results of normalization and translation, the period required for model generation (the total period required for learning and model generation) can be shortened compared to the case where output models are generated individually, and the output speeds of the normalized text and the translated text can be improved. Further, in the translation device 10 according to this embodiment, for at least some of the learning data, translation learning is performed after normalization learning has been performed first. Thereby, for example, in the case where learning is performed using an encoder-decoder model, for at least some of the learning data, translation learning can be performed in a state where the influence of noise in the learning source text is suppressed by using the parameters learned in the normalization learning (that is, the parameters suitable for normalization). This can improve the translation accuracy in the normalization / translation model 70.
[0049] The normalization learning unit 12 and the translation learning unit 13 perform learning alternately with each other. For each learning data, after the learning by the normalization learning unit 12 is performed, the learning by the translation learning unit 13 may be continuously performed. In this way, the learning by the normalization learning unit 12 and the learning by the translation learning unit 13 are alternately performed, and for each learning data, the learning by the normalization learning unit 12 is always performed first, and then the learning by the translation learning unit 13 is continuously performed. Thus, in a case where learning is performed using, for example, an encoder-decoder model, for each learning data, parameters suitable for both normalization and translation can be learned. For example, in a case where learning for normalization is performed for all learning data and then learning for translation is performed for all learning data, for each learning data, parameters suitable for both normalization and translation cannot be learned (when performing learning for translation, the parameters will be learned in a state where the influence of the previously learned normalization has weakened). In this regard, as described above, by performing the learning by the normalization learning unit 12 first for each learning data and then continuously performing the learning by the translation learning unit 13, parameters suitable for both normalization and translation can be appropriately learned. By this, the translation accuracy can be further improved.
[0050] The normalization learning unit 12 and the translation learning unit 13 perform learning using an encoder-decoder model that uses a common encoder and decoders provided individually. The translation learning unit 13 may perform learning using the hidden state of the encoder learned by the normalization learning unit 12 for each learning data. When the learning by the normalization learning unit 12 and the learning by the translation learning unit 13 are continuously performed for each learning data, since the encoder is shared and the hidden state learned in the normalization learning is used in the translation learning, translation learning with reduced influence of noise (grammatically correctly converted) can be performed, and the translation accuracy can be further improved.
[0051] The normalization learning unit 12 and the translation learning unit 13 may repeatedly learn a plurality of learning data a plurality of times. By repeatedly learning, parameters suitable for both normalization and translation can be learned more effectively, and the translation accuracy can be further improved.
[0052] The translation device 10 further includes an evaluation unit 15 that derives a loss function related to normalization and a loss function related to translation for the normalization / translation model 70 generated by the model generation unit 14, and evaluates the normalization / translation model 70 based on the values of the respective loss functions. When at least one of the following conditions is satisfied for a plurality of learning data when the normalization learning unit 12 and the translation learning unit 13 have repeatedly learned a plurality of times: the value of the loss function related to normalization is greater than a predetermined first threshold, and the value of the loss function related to translation is greater than a predetermined second threshold, the evaluation unit 15 evaluates that the normalization / translation model 70 is in a first state with low prediction accuracy. When it is evaluated that it is in the first state and the value of the loss function related to normalization is greater than the first threshold, the normalization learning unit 12 separately performs repeated learning for each learning data alone, separately from the learning performed alternately with the translation learning unit 13. When it is evaluated that it is in the first state and the value of the loss function related to translation is greater than the second threshold, the translation learning unit 13 may separately perform repeated learning for each learning data alone, separately from the learning performed alternately with the normalization learning unit 12. In this way, separate from the normal learning (normalization learning and translation learning performed alternately with each other), by separately performing intensive learning for processes where the value of the loss function is large and the prediction accuracy is assumed to be low, the loss function can be effectively converged and the accuracy of the model can be improved. As a result, the translation accuracy can be further improved.
[0053] The translation device 10 further includes an acquisition unit 16 that acquires an input sentence in a first language, and a translation unit 17 that has a normalization / translation model 70. The translation unit 17 may generate a normalized sentence by inputting the input sentence acquired by the acquisition unit 16 into the normalization / translation model 70, and generate a translated sentence in a second language corresponding to the normalized sentence by inputting the normalized sentence into the normalization / translation model 70. Thereby, using the generated one normalization / translation model 70, the normalization and translation of natural speech (input sentence) can be smoothly performed, and translation can be performed at high speed and with high accuracy.
[0054] Finally, the hardware configuration of the translation device 10 will be described with reference to FIG. 10. Physically, the above-described translation device 10 may be configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like.
[0055] In the following description, the term "device" can be read as a circuit, a device, a unit, etc. The hardware configuration of the translation device 10 may be configured to include one or more of each device shown in the figure, or may be configured without including some devices.
[0056] Each function in the translation device 10 is realized by causing a processor 1001 to load a predetermined software (program) onto hardware such as the processor 1001 and the memory 1002, so that the processor 1001 performs calculations and controls communication by the communication device 1004 and reading and / or writing of data in the memory 1002 and the storage 1003.
[0057] The processor 1001 controls the entire computer by operating an operating system, for example. The processor 1001 may be composed of a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic device, a register, and the like. For example, the control function of the normalization learning unit 12 of the translation device 10 may be realized by the processor 1001.
[0058] Also, the processor 1001 reads a program (program code), software module, and data from the storage 1003 and / or the communication device 1004 into the memory 1002, and executes various processes according to them. As the program, a program that causes a computer to execute at least a part of the operations described in the above-described embodiments is used. For example, control functions such as the normalization language learning unit 12 of the translation device 10 may be stored in the memory 1002 and realized by a control program operating on the processor 1001, and other functional blocks may be realized in the same manner. Although it has been described that the above-described various processes are executed by one processor 1001, they may be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented with one or more chips. Note that the program may be transmitted from a network via a telecommunication line.
[0059] The memory 1002 is a computer-readable recording medium and may be composed of at least one of, for example, ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), and the like. The memory 1002 may be referred to as a register, cache, main memory (main storage device), and the like. The memory 1002 can store a program (program code), software module, etc. executable for implementing the wireless communication method according to an embodiment of the present invention.
[0060] Storage 1003 is a computer-readable recording medium, which may be composed of at least one of, for example, optical discs such as CD-ROM (Compact Disc ROM), hard disk drives, flexible disks, magneto-optical disks (e.g., compact discs, digital versatile discs, Blu-ray (registered trademark) discs), smart cards, flash memories (e.g., cards, sticks, key drives), floppy (registered trademark) disks, magnetic strips, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate media including memory 1002 and / or storage 1003.
[0061] Communication device 1004 is hardware (a transceiver device) for performing communication between computers via a wired and / or wireless network, and is also referred to as, for example, a network device, a network controller, a network card, a communication module, etc.
[0062] Input device 1005 is an input device for receiving external input (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.). Output device 1006 is an output device for performing external output (e.g., a display, a speaker, an LED lamp, etc.). Note that input device 1005 and output device 1006 may have an integrated configuration (e.g., a touch panel).
[0063] Also, each device such as processor 1001 and memory 1002 is connected by a bus 1007 for communicating information. Bus 1007 may be composed of a single bus or may be composed of different buses between devices.
[0064] Further, the translation device 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be implemented by the hardware. For example, the processor 1001 may be implemented by at least one of these hardware components.
[0065] As described above in detail, it is obvious to those skilled in the art that the present embodiment is not limited to the embodiments described in this specification. The present embodiment can be implemented as a modified and changed form without departing from the spirit and scope of the present invention defined by the claims. Therefore, the description in this specification is for illustrative purposes and has no restrictive meaning for the present embodiment.
[0066] Each aspect / embodiment described in this specification may be applicable to systems using LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G, 5G, FRA (Future Radio Access), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broad-band), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), IEEE 802.20, UWB (Ultra-Wide Band), Bluetooth (registered trademark), and other appropriate systems and / or next-generation systems extended based on these.
[0067] The processing procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this specification may be reordered as long as there is no contradiction. For example, for the methods described in this specification, the elements of various steps are presented in an exemplary order and are not limited to the specific order presented.
[0068] The input / output information, etc. may be stored in a specific location (e.g., memory) or may be managed by a management table. The input / output information, etc. may be overwritten, updated, or appended. The output information, etc. may be deleted. The input information, etc. may be transmitted to other devices.
[0069] The determination may be made by a value represented by 1 bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (e.g., comparison with a predetermined value).
[0070] Each aspect / embodiment described in this specification may be used alone, in combination, or switched and used during execution. Also, the notification of predetermined information (e.g., notification of "being X") is not limited to being explicitly performed and may be performed implicitly (e.g., by not performing the notification of the predetermined information).
[0071] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., regardless of whether it is called software, firmware, middleware, microcode, a hardware description language, or by another name.
[0072] Also, software, instructions, etc. may be transmitted and received via a transmission medium. For example, when software is transmitted from a website, server, or other remote source using wired technologies such as coaxial cables, fiber optic cables, twisted pairs, and digital subscriber lines (DSL) and / or wireless technologies such as infrared, wireless, and microwave, these wired and / or wireless technologies are included within the definition of the transmission medium.
[0073] The information, signals, etc. described in this specification may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0074] Note that for the terms described in this specification and / or terms necessary for understanding this specification, they may be replaced with terms having the same or similar meanings.
[0075] Also, the information, parameters, etc. described in this specification may be represented by absolute values, relative values from a predetermined value, or in terms of corresponding other information.
[0076] The user terminal may be referred to by those skilled in the art as a mobile communication terminal, subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other suitable term.
[0077] As used herein, the terms "determining" and "determination" may encompass a wide variety of actions. "Determining" and "determination" can include, for example, calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, database, or other data structure), and ascertaining something as having been "determined". Also, "determining" and "determination" can include receiving (e.g., receiving information), transmitting (e.g., transmitting information), inputting, outputting, accessing (e.g., accessing data in a memory), and ascertaining something as having been "determined". Further, "determining" and "determination" can include resolving, selecting, choosing, establishing, comparing, etc., and ascertaining something as having been "determined". That is, "determining" and "determination" can include ascertaining that some action has been "determined".
[0078] As used herein, the description "based on" does not mean "based only on" unless otherwise specified. In other words, the description "based on" means both "based only on" and "based at least on".
[0079] When the terms "first", "second", etc. are used herein, any reference to those elements does not generally limit the quantity or order of those elements. These terms can be used herein as a convenient way to distinguish between two or more elements. Thus, a reference to the first and second elements does not mean that only two elements can be employed there, or that the first element must precede the second element in any way.
[0080] As long as the terms "include", "including", and their variants are used in this specification or the claims, these terms are intended to be inclusive, just like the term "comprising". Further, the term "or" used in this specification or the claims is not intended to be an exclusive disjunction.
[0081] In this specification, unless the context or the technology clearly indicates that there is only one device, multiple devices are also included.
[0082] Throughout this disclosure, unless the context clearly indicates a singular form, multiple entities are included.
Description of Reference Numerals
[0083] 10... Translation device, 11... Memory unit, 12... Normalization language learning unit, 13... Translation language learning unit, 14... Model generation unit, 15... Evaluation unit, 16... Acquisition unit, 17... Translation unit, 18... Output unit, 70... Normalization-translation model.
Claims
1. A storage unit that stores a plurality of pieces of learning data in which an original text for learning in a first language, a normalized learning text obtained by grammatically correctly converting the learning original text, and a translated learning text obtained by translating the learning original text into a second language different from the first language are associated with each other; A normalized text learning unit that learns by combining the learning original text and the corresponding normalized learning text for a plurality of the learning data; A translation text learning unit that learns by combining the learning original text and the corresponding translated learning text for a plurality of the learning data; A model generation unit that generates one normalization / translation model configured to be able to output a normalized text and a translated text into a second language for an input text in the first language based on the learning results of the normalized text learning unit and the translation text learning unit; For at least some of the learning data, learning by the translation text learning unit is performed after learning by the normalized text learning unit. The normalized text learning unit and the translation text learning unit perform learning alternately with each other. For each of the learning data, learning by the translation text learning unit is continuously performed after learning by the normalized text learning unit. The normalized text learning unit and the translation text learning unit repeatedly learn a plurality of times with respect to the plurality of learning data. The evaluation unit further includes: for the normalization / translation model generated by the model generation unit, a loss function related to normalization and a loss function related to translation are derived, and the normalization / translation model is evaluated based on the values of the respective loss functions. When at least one of the following conditions is satisfied for the plurality of learning data when the normalized text learning unit and the translation text learning unit repeatedly learn a plurality of times: the value of the loss function related to normalization is greater than a predetermined first threshold, and the value of the loss function related to translation is greater than a predetermined second threshold, the evaluation unit evaluates that the normalization / translation model is in a first state with low prediction accuracy. When it is evaluated to be in the first state and the value of the loss function regarding the normalization is greater than the first threshold, the normalization learning unit performs iterative learning separately for each of the learning data, separately from the learning performed alternately with the translation learning unit. The translation learning unit, when it is evaluated to be in the first state and the value of the loss function regarding the translation is greater than the second threshold, performs iterative learning separately for each of the learning data, separately from the learning performed alternately with the normalization learning unit. A translation device.
2. A storage unit that stores a plurality of learning data in which a learning original text in a first language, a learning normalized text obtained by grammatically correctly converting the learning original text, and a learning translation text obtained by translating the learning original text into a second language different from the first language are associated with each other; A normalization learning unit that learns by combining the learning original text and the corresponding learning normalized text for a plurality of the learning data; A translation learning unit that learns by combining the learning original text and the corresponding learning translation text for a plurality of the learning data; A model generation unit that generates one normalization / translation model configured to be able to output a normalized text and a translation text into a second language for an input text in the first language based on the learning results of the normalization learning unit and the translation learning unit. For at least some of the learning data, after learning by the normalization learning unit, learning by the translation learning unit is performed. The normalization learning unit and the translation learning unit perform learning using an encoder-decoder model that uses a common encoder and individually provided decoders. The translation learning unit performs learning using the hidden state of the encoder learned by the normalization learning unit for each of the learning data. The normalization learning unit and the translation learning unit perform learning a plurality of times repeatedly for a plurality of the learning data. For the normalization / translation model generated by the model generation unit, a loss function related to normalization and a loss function related to translation are derived, and an evaluation unit that evaluates the normalization / translation model based on the values of the respective loss functions is further provided. When the evaluation unit determines that, for a plurality of the learning data, when the normalization / translation model is repeatedly learned a plurality of times by the normalization learning unit and the translation learning unit, at least one of the following conditions is satisfied: the value of the loss function related to normalization is greater than a predetermined first threshold, and the value of the loss function related to translation is greater than a predetermined second threshold, the evaluation unit evaluates that the normalization / translation model is in a first state with low prediction accuracy. When the normalization learning unit is evaluated to be in the first state and the value of the loss function related to normalization is greater than the first threshold, the normalization learning unit separately performs repeated learning for each of the learning data independently, separately from the learning performed alternately with the translation learning unit. When the translation learning unit is evaluated to be in the first state and the value of the loss function related to translation is greater than the second threshold, the translation learning unit separately performs repeated learning for each of the learning data independently, separately from the learning performed alternately with the normalization learning unit. A translation device.
3. An acquisition unit that acquires the input sentence in the first language; A translation unit having the normalization / translation model, and further comprising: The translation unit: By inputting the input sentence acquired by the acquisition unit into the normalization / translation model, a normalized sentence is generated. By inputting the normalized sentence into the normalization / translation model, a translated sentence in the second language corresponding to the normalized sentence is generated. The translation device according to claim 1 or 2.
Citation Information
Patent Citations
Pseudo parallel translation data generation device, machine translation processing device, and pseudo parallel translation data generation method
JP2019153023A
Apparatus and method for generating translation model, apparatus and method for automatic translation
US20170139905A1