Translator, translation learning device, translation method, translation learning method, and program
The translation device addresses the challenge of generating syntactically diverse translations by using a fine-tuned multilingual model to generate syntax trees and employing prefix constrained decoding and two-level beam search, resulting in controlled and accurate syntactic variations in machine translation outputs.
Patent Information
- Application Number
- JP2023184451
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-27
- Publication Date
- 2025-05-13
AI Technical Summary
Conventional machine translation methods struggle to generate syntactically diverse translations, as they either lack control over the output syntax or result in limited lexical and syntactic diversity.
A translation device that generates a syntax tree for input statements using a fine-tuned multilingual model, allowing for explicit control over the output syntax through prefix constrained decoding and two-level beam search.
The approach enables the generation of syntactically diverse translations while maintaining translation accuracy, allowing for controlled manipulation of the syntactic structure of output sentences.
Smart Images

Figure 2025073547000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a translation device, a translation learning device, a translation method, a translation learning method, and a program. [Background technology]
[0002] Recent advances in deep learning have dramatically improved the accuracy of neural machine translation (NMT). However, conventional NMT methods lack the ability to generate diverse translations.
[0003] The commonly used beam search can generate different translation candidates from an NMT model, but in many cases it only brings small lexical changes to the output sentence and cannot significantly change the syntactic structure of the output sentence.
[0004] One example of an approach to generate diverse translations is diverse beam search (DBS) (Non-Patent Document 1). DBS extends the beam search by dividing the beams into groups and performing a beam search for each group. Nodes that have already been visited in the previous group are assigned a penalty, encouraging diverse output. DBS increases the lexical diversity of the output sentences, but has limited effect on syntactic diversity.
[0005] In order to generate syntactically diverse translations, a method has been proposed that uses a discrete syntactic code that encodes the structure of a sentence (Non-Patent Document 2). In this method, an auto-encoder based on TreeLSTM is used to create an embedding vector from the syntactic tree of the sentence, which is then discretized to generate a syntactic code.
[0006] By adding a syntactic code as a prefix to the beginning of a sentence in the target language of the bilingual data, an NMT model can be trained to generate sentences based on the given syntactic code. By sampling the prefixed syntactic code, it is possible to generate syntactically diverse translations using randomized syntactic codes.
[0007] The drawback of this approach is that it does not allow explicit control over the output syntax: since there is no one-to-one correspondence between syntax codes and sentence structures, it is not guaranteed to output sentences with the desired syntactic structure. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R. Selvaraju, Qing Sun, Stefan Lee, avid Crandall, and Dhruv Batra, "Diverse beam search for improved description of complex scenes", In AAAI, 2018 [Non-Patent Document 2] Raphael Shu, Hideki Nakayama, and Kyunghyun Cho, "Generating diverse translations with sentence codes", In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 1823-1827, Florence, Italy, July 2019. Association for Computational Linguistics Summary of the Invention [Problem to be solved by the invention]
[0009] As mentioned above, conventional methods for enabling machine translation to output diverse sentences are diverse beam search (DBS) and syntactic code. DBS has the effect of diversifying the vocabulary of the output sentences because there is a penalty for outputting the same word or word string. However, there is a problem that it is not very effective in diversifying the syntax, i.e., word order, of the output sentences.
[0010] Since the syntax code is based on the syntax tree, it has the effect of diversifying the syntax of the output sentence. However, since the syntax code is created by creating an embedding vector from the syntax tree and then discretizing it, the correspondence between the syntax code and the syntactic structure is unclear. Therefore, there is a problem that there is no way to externally control the output sentence to have a desired syntactic structure.
[0011] The present invention has been made in view of the above, and aims to provide a new method for generating syntactically diverse translations. [Means for solving the problem]
[0012] In order to solve the above problem, the translation device has a translation unit that is configured to generate a syntax tree for the input sentence by inputting the input sentence into a translation model that has been generated by fine-tuning a trained multilingual model using a pair of a syntax tree of a source language sentence and a target language sentence as training data. Effect of the Invention
[0013] It can provide new ways of generating syntactically diverse translations. [Brief description of the drawings]
[0014] [Figure 1] 1 is a diagram illustrating an example of a hardware configuration of a translation device 10 according to an embodiment of the present invention. [Diagram 2] 1 is a diagram illustrating an example of a functional configuration of a translation device 10 during training according to an embodiment of the present invention. [Diagram 3] 11 is a flowchart illustrating an example of a processing procedure of a translation model training phase. [Figure 4] FIG. 2 is a diagram showing an example of a phrase structure tree output by a phrase structure parser. [Diagram 5] FIG. 13 is a diagram showing an example in which a front terminal node has been removed from a phrase structure tree. [Figure 6] FIG. 13 is a diagram showing an example of flattening a phrase structure tree so that the maximum tree height is 4. [Figure 7] 11 is a flowchart illustrating an example of a processing procedure of an inference phase of a translation model. [Figure 8] Fig. 11 shows BLEU scores and word alignment based diversity scores (AD) for different methods. [Figure 9] FIG. 13 illustrates an example of controlling the syntax of an output sentence by prefix-constrained decoding. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Fig. 1 is a diagram showing an example of a hardware configuration of a translation device 10 in an embodiment of the present invention. The translation device 10 in Fig. 1 has a drive device 100, an auxiliary storage device 102, a memory device 103, a processor 104, and an interface device 105, which are all connected to each other via a bus B.
[0016] The program for implementing the processing in translation device 10 is provided by recording medium 101 such as a CD-ROM. When recording medium 101 storing the program is set in drive device 100, the program is installed from recording medium 101 via drive device 100 into auxiliary storage device 102. However, the program does not necessarily have to be installed from recording medium 101, but may be downloaded from another computer via a network. Auxiliary storage device 102 stores the installed program as well as necessary files, data, etc.
[0017] When an instruction to start a program is received, memory device 103 reads out and stores the program from auxiliary storage device 102. Processor 104 is a CPU or a GPU (Graphics Processing Unit), or a CPU and a GPU, and executes functions related to translation device 10 in accordance with the program stored in memory device 103. Interface device 105 is used as an interface for connecting to a network.
[0018] Fig. 2 is a diagram showing an example of a functional configuration of a translation device 10 according to an embodiment of the present invention. In Fig. 2, the translation device 10 has a syntax analysis unit 11, a multilingual model fine-tuning unit 12, and a translation decoder 13. Each of these units is realized by a process executed by a processor 104 of one or more programs installed in the translation device 10.
[0019] In the training phase, the syntactic analysis unit 11 generates training data from a bilingual corpus for generating a translation model that receives a source language sentence as input and outputs a phrase structure tree (syntax tree) of a target language sentence as a translation result.
[0020] In a training phase, the multilingual model fine-tuning unit 12 fine-tunes the trained multilingual model using training data to generate a translation model that outputs a syntax tree of a target language sentence corresponding to an input sentence.
[0021] In the inference phase, the translation decoder 13 translates the input using a translation model and outputs an output sentence that is the translation result.
[0022] The training and inference phases of the translation model are described below.
[0023] [Training Phase] 3 is a flowchart for explaining an example of a processing procedure in the training phase of a translation model. In the training phase, the translation device 10 functions as a translation learning device. Different computers may be used in the training phase and the inference phase.
[0024] In step S101, the syntax analysis unit 11 generates a constituency tree for each output sentence by parsing each target language sentence in a bilingual corpus, which is a collection of pairs of input sentences (source language sentences) and output sentences (target language sentences) that are translation results of the source language sentences, using a phrase structure syntax parser.
[0025] Fig. 4 is a diagram showing an example of a phrase structure tree output by the phrase structure parser. Fig. 4 shows an example of a phrase structure tree generated when parsing "Meetings for announcing research results were held in Japan and abroad."
[0026] In this embodiment, it is not the ultimate purpose to obtain a linguistic sentence structure. Therefore, the syntax analysis unit 11 simplifies the phrase structure tree of each output sentence to such an extent that the syntactic structures output from the translation model become diverse (S102). Specifically, the syntax analysis unit 11 removes the front terminal node from each phrase structure tree, and flattens the tree height to a maximum of 4. However, the maximum value of the tree height may be other than 4.
[0027] Figure 5 shows an example of removing front terminal nodes from a phrase structure tree. Figure 6 shows an example of flattening a phrase structure tree so that the height of the tree is a maximum of 4.
[0028] Next, the syntax analysis unit 11 linearizes the simplified phrase structure tree (S103), resulting in a token sequence.
[0029] When linearizing a phrase structure tree, component tags are treated as special tokens and are not split during tokenization. An open parenthesis and a component symbol are treated as one token representing the start tag of a component, and a close parenthesis for the start tag of a different component is treated as a different token. Note that a component symbol is a symbol that represents the type of sentence component based on linguistic theory, such as a noun phrase NP or a verb phrase VP. A component tag is one that is treated as one token in this embodiment, such as "(NP". For example, for "(NP", "<" and ">" are treated as "<" and ">". It is written as "<(NP>" surrounded by "".
[0030] In particular, the linearized phrase structure tree "(S (NP We) (VP asked (NP two specialists)) (PP for (NP their opinion)) .)" is changed to "<(S> <(NP> We <(NP> <(VP> asked <(NP> two specialists <)NP> <)VP> <(PP> for <(NP> their opinion <)NP> <)PP> . <)S>".
[0031] Next, the syntactic analysis unit 11 generates, as training data, a set of pairs of input sentences (source language sentences) included in the bilingual corpus and sequence data in which the phrase structure trees of output sentences have been simplified (S104).
[0032] Next, the multilingual model fine-tuning unit 12 uses the training data to fine-tune a trained multilingual model (e.g., an encoder-decoder) including the source language and the target language so as to output sequence data in which the phrase structure tree of an output sentence is simplified for an input sentence, thereby generating a translation model that outputs sequence data in which the phrase structure tree of an output sentence is simplified for an input sentence (S105).
[0033] It is also possible to directly train an encoder-decoder model using pairs of source and target language sentences as training data, but preliminary experiments by the inventors of the present application have shown that in this case, the translation accuracy is significantly lower than when an encoder-decoder model is directly trained using pairs of source and target language sentences as training data. Therefore, it is important to fine-tune a trained multilingual model.
[0034] [Inference Phase] FIG. 7 is a flowchart illustrating an example of a processing procedure of the inference phase of the translation model.
[0035] In step S201, the translation decoder 13 accepts input of an input sentence (sentence to be translated) from the user, and also accepts input of a prefix of the phrase structure tree of the output sentence as a translation condition (constraint). A prefix is a token sequence including a component tag that constitutes a part from the beginning of the output sentence. For example, "Postoperative leg pain has improved, and there has been no recurrence of symptoms" is input as an input sentence, and "<(S> <(PP>))" is input as a prefix. By inputting a prefix, the user can, for example, distinguish between a there construction and a construction with a we as the subject (a construction that expresses the existence of an object in the form of a person's possession). It is assumed that the user has knowledge of component tags.
[0036] Next, the translation decoder 13 uses the translation model to execute a translation according to the input from the user (S202). Specifically, the translation decoder 13 inputs the input sentence and a prefix to the translation model. The target language sequence output by the translation model always includes multiple tokens (constituent tags) that are not words as prefixes. The syntactic structure of the output sentence can be controlled by specifying a prefix including a constituent tag from the outside to the translation model and performing prefix-constrained decoding. At this time, the translation model executes any one of various types of beam searches using the output probability (probability distribution) of each token generated in the decoding process to generate various translation sentences. The beam search method may be a normal beam search or a DBS (Diverse Beam Search). Alternatively, multiple types of beam searches may be executed. The translation model outputs a phrase structure tree (linearized and flattened phrase structure tree) of the output sentence as a result of decoding. In the case of the input sentence and prefix exemplified in step S201, for example, "<(S> <(PP> After <(NP> the operation<)NP><)PP>, <(S> <(NP> the inferior limb pain<)NP> <(VP> was improved<)VP><)S>, and <(S> <(NP> there<)NP> <(VP> is no recurrence of the symptom<)VP><)S>.<)S>" is output from the translation model.
[0037] Next, the translation decoder 13 outputs the phrase structure tree output from the translation model (S203).
[0038] Note that the user does not have to input a prefix, in which case the translation model performs decoding without prefix constraints.
[0039] [Two-Level Beam Search (TLBS)] Here, we propose a two-level beam search (TLBS) as one of the beam searches performed in step S202.
[0040] In two-stage beam search (TLBS), decoding is divided into two stages. In the first stage, a set of multiple types (b token sequences, each of which consists of the first n tokens of the output sentence, is generated by beam search. The n tokens may include not only constituent tags but also non-constituent tokens. In the second stage, each of the b token sequences is used as a prefix for further decoding by beam search. In TLBS, the first n token sequences of all output sentence candidates are forced to be different. Thus, b object sequences with different phrase structure trees are generated.
[0041] For example, when no prefix is input by the user (when no prefix is provided by the translation decoder 13), the translation model performs decoding using TLBS in step S202. At this time, the translation decoder 13 may present to the user a set of n token sequences generated by the translation model in the first stage of TLBS. The translation decoder 13 may provide one or more token sequences selected by the user from the set of n token sequences to the translation model, and cause the translation model to execute the second stage of TLBS.
[0042] In addition, Reference 11 proposes a method of implementing a string-to-linearized-tree model using an encoder-decoder model and outputting the syntactic structure of an input sentence (reference information for each reference will be provided later).
[0043] In Reference 1, the method of Reference 11 is applied to machine translation. That is, in Reference 11, machine translation is regarded as a conversion from a source language sentence to a syntax tree of a target language sentence, and machine translation is realized using a syntax parser of the target language and an encoder-decoder model.
[0044] Reference 8 proposes Translation between Augmented Natural Languages (TANL), a general framework for predicting natural language structures using a trained encoder-decoder model. In this paper, named entity recognition, relation extraction, semantic role assignment, and coreference analysis are presented as examples of applications of TANL, but it is not applied to machine translation.
[0045] This embodiment is similar to the method of Reference 1 in that machine translation is regarded as conversion from a source language sentence to a syntax tree of a target language sentence. Furthermore, by combining this with the method of Reference 8, machine translation is regarded as a problem of predicting a target language sentence structure across languages from a source language sentence. Then, conversion from a source language sentence to a syntax tree of a target language sentence is realized using a trained multilingual model.
[0046] Furthermore, Reference 12 proposes prefix-constrained decoding, which controls the sentence generated by an encoder-decoder model by externally specifying a sequence of consecutive words (i.e., a prefix) from the beginning of a sentence and having the model generate the sequence of words that follows it.
[0047] In this embodiment, the target output is a linearized syntax tree, so the first few tokens on the decoder side always contain constituent tags with syntactic information, and the prefix-constrained decoding framework allows the user to control the syntactic structure of the target sentence by externally specifying prefixes that contain constituent tags.
[0048] [Diversity score based on word alignment] Pairwise-BLEU (Reference 10) and diversity score (DP) (Non-Patent Document 2), which have been used so far as measures to evaluate the diversity of output sentences, are measures based on the matching of word ngrams. Although these methods can certainly evaluate diversity at the word level, they cannot necessarily evaluate syntactic diversity.
[0049] For example, the use of synonyms increases the ngram-based diversity score but does not increase syntactic diversity. Furthermore, outputs that contain both correct and incorrect translations are scored higher in the ngram-based diversity score because they have a higher degree of ngram mismatch.
[0050] To solve these problems, we propose a measure of word alignment-based diversity (AD). Similar to DP, the proposal score (AD) is calculated as follows:
[0051]
number
[0052] Δ k In order to calculate (y,y'), in this embodiment, a list of semantically corresponding word pairs between sentence candidate y and sentence candidate y' is calculated using {(w1,w1'),(w2,w2'),...,(w n ,w nThen, Kendall's rank correlation coefficient τ is calculated from the list of word pair positions in the two sentence candidates {(pos(w1),pos(w1')),...,(pos(wn),pos(wn'))}. For example, assume that the following two sentence candidates are as follows. Note that the number in parentheses indicates the position of the immediately preceding word, and, together with the parentheses, is not a component of the sentence. This (1) is (2) a book (3) (4). This is book (1) (2) and this (3).
[0053] In these two sentence candidates, the semantically corresponding word pairs are {(kore(1), kore(3)), (hon(3), book(1)), (desu(4), desu(2))}. Therefore, the position list of the semantically corresponding word pairs is {(1,3),(3,1),(4,2)}.
[0054] Δ k (y,y') is calculated using τ as follows:
[0055]
number
[0056] That is, the degree of agreement in word order is basically measured by the rank correlation coefficient. However, when syntax is forced, it is possible that a target sentence that is semantically deviated from the source sentence will be generated, resulting in fewer semantically corresponding words, and limiting the number of corresponding words to the same word order. Equation (2) is a measure that takes into account both the degree of agreement in word order and the degree of agreement in vocabulary.
[0057] [Experiment to verify the effectiveness of this embodiment] [1 Data] In this experiment, we used the ASPEC Japanese-English dataset (Reference 7) and the WMT14 German-English dataset (https: / / www.statmt.org / wmt14 / translation-task.html). The ASPEC Japanese-English dataset consists of approximately 3M training sentences, 1.8k development sentences, and 1.8k test sentences, and is made up of abstracts of Japanese and English scientific papers. The WMT14 German-English dataset is made up of the Europarl v7, Common Crawl, and News Commentary corpora, and consists of approximately 4.5M training sentences, 3K development sentences, and 3K test sentences. Following the instructions included with the datasets, we removed date expressions at the end of sentences in the Japanese part of the ASPEC dataset.
[0058] For both datasets, we used the Berkley Neural Parser (Reference 3) to parse the English sentences into phrase structures, removed pre-terminal nodes from the phrase structure trees, and flattened the trees to a maximum height of 4. The flattened trees were paired with the corresponding Japanese and German sentences and used as training data for the translation model.
[0059] [2 Experimental setup] For the translation model, we used the publicly available mBART-large-cc25 (Reference 6) as a pre-trained multilingual model. We learned a new sentencepiece model (Reference 4) from the training data of the dataset, and reduced the vocabulary size of mBART from 250k to about 64k by removing tokens and their corresponding embeddings that were not included in the vocabulary of this new sentencepiece model from the model.
[0060] For ASPEC Japanese-English, the sentencepiece model was trained with a vocabulary size of 64k and character coverage of 0.9995. For WMT14 German-English, the character coverage was changed to 1.0. Due to the constraints of mBART, source and target sentence pairs with sequences longer than 1028 tokens were removed from the training data. The mBART model was fine-tuned using cross-entropy loss with teacher forcing and label smoothing factor of 0.2.
[0061] The hyperparameters were as close as possible to those used in reference 6. After 2.5k warm-up steps, training was performed with a learning rate of 3.0 × 10 -5 The optimizer was Adam, with the parameter ε = 1.0 × 10 -6 ,(β1,β2)=(0.9,0.98) with no weight decay. The batch size was set to approximately 12k tokens per batch. Training took 2 days on a single NVIDIA® A100 GPU. Checkpoints were saved every 2k steps, and the checkpoint with the least validation data loss was used for final evaluation.
[0062] For comparison, we fine-tuned mBART on a normal sentence-sentence pair dataset using the model (vocabulary) size reduction techniques and hyperparameters described above. For decoding, we used normal beam search (BS), diverse beam search (DBS), and two-stage beam search (TLBS). We used the automatic evaluation metric BLEU to evaluate the accuracy of the machine translation, specifically the sacreBLEU toolkit (Reference 9). To quantitatively evaluate the diversity of the output, we used a neural semi-Markov CRF-based word alignment tool (Reference 5) to find word alignments for English sentence pairs and calculated the diversity score (AD) with gamma = 0.2.
[0063] [3 Experimental Results] [3.1 Comparison of diversity scores] The experimental results are shown in Figure 8. Figure 8 shows the BLEU score and word alignment based diversity score (AD) for different methods. BLEU is the BLEU score of the candidate sentence with the highest log probability. The score with - above the BLEU is the average BLEU score of the three candidate sentences, and SD is its standard deviation. The beam size was set to 3 for all methods.
[0064] All models are based on mBART (Reference 6). For each source sentence, we generate three translation candidates and calculate a word alignment based diversity score (AD) for each translation candidate set. The reported diversity score is the average across all translation candidate sets.
[0065] The translation model of this embodiment ("Ours" in FIG. 8) generally has a slightly lower translation accuracy than the baseline mBART model. This is thought to be because the sequence representing the syntax tree is longer than a normal sentence and is difficult for the translation model to generate.
[0066] Regarding the decoding method, BS and DBS showed similar trends for the baseline translation model (sequence-to-sequence) and the translation model of the present embodiment (sequence-to-linearized-tree). Compared to BS, DBS has a lower BLEU score (translation accuracy) but a higher AD score (syntactic diversity).
[0067] However, TLBS results in a greater increase in AD score (syntactic diversity) for the translation model of this embodiment compared to the baseline translation model. The greater increase in syntactic diversity for the translation model of this embodiment is likely due to the presence of component tags. Because TLBS forces the use of different starting component tags, each translation candidate is likely to have a different syntactic structure. Increasing n in TLBS improves the average BLEU score, reduces the standard deviation, and stabilizes the output, but reduces the diversity score AD.
[0068] Below is a translation example of "When the Hungarians fought the Hungarians, they were forced to go to the Rally Obedience Turnaround. (German)". In the following, (A-1) to (A-3) are examples of sentences obtained by the TLBS of mBART, which was fine-tuned using normal bilingual sentence pair data. (B-1) to (B-3) are examples of sentences obtained by the TLBS of the translation model of this embodiment. (A-1)Once again the athletes of the dog-friends Bitz were successful at a Rally-Obedience- Tournament. (A-2)Once again, the athletes of the dog-friends Bitz were successful at a Rally-Obedience- Tournament. (A-3)Once again the athletes from Dog Friends Bitz were successful at a Rally-Obedience- Tournament. (B-1)For the second time, the athletes of dog friends Bitz were successful at a Rally - Obedience Tournament. (B-2)Once again the athletes of dog friends Bitz were successful at a Rally - Obedience Tournament. (B-3)The athletes of dog friends Bitz were successful at a Rally - Obedience - Tournament again. It can be seen that when decoding is performed using a two-stage beam search TLBS based on the translation model of this embodiment, it is possible to generate translations that are more syntactically diverse than the baseline translation model.
[0069] [3.2 Controlling syntactic diversity through prefix constraints] Figure 9 shows an example of controlling the syntax of an output sentence by prefix-constrained decoding. In (A)-(D), the second and subsequent lines show the linearized phrase structure tree output from the model, and the first line shows the sentence obtained from the phrase structure tree.
[0070] (A) is decoded without the prefix. (B), (C), and (D) are decoded with the underlined prefix. Among (B) to (D), (B) and (C) are correctly translated from the source language sentence, but when an impossible prefix is used as in (D), the source language sentence is not translated correctly. When it is difficult to satisfy the given syntactic constraints, the translation model of this embodiment may omit facts or generate hallucinations.
[0071] [effect] As described above, according to this embodiment, a translation model trained to convert from sentences to linearized syntax trees, rather than a model that converts from sentences to sentences, generates tokens of both words and syntax elements, making it possible to generate diverse translations by sampling output probabilities. In other words, a new method for generating syntactically diverse translations can be provided.
[0072] The first important point to achieve this effect is to convert the target language sentences of the bilingual data into phrase structure trees using a phrase structure parser, and to generate a translation model by fine-tuning a trained multilingual model using pairs of source language sentences and target language sentence syntax trees as training data. This method can obtain translation accuracy almost equivalent to that of creating a translation model by fine-tuning a trained multilingual encoder-decoder model using bilingual data.
[0073] The next point is to generate translation candidates with different prefixes by using a two-stage beam search or by specifying a prefix from outside. Since the prefix always contains a component tag, it is more likely to obtain translation candidates with different syntactic structures than with a normal translation model.
[0074] In this embodiment, the translation decoder 13 is an example of a translation unit.
[0075] [References] [Reference 1] Roee Aharoni and Yoav Goldberg, "Towards string-to-tree neural machine translation", In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 132-140, Vancouver, Canada, July 2017. Association for Computational Linguistics. [Reference 2] Hideki Isozaki, Tsutomu Hirao, Kevin Duh, Katsuhito Sudoh, and Hajime Tsukada., "Automatic evaluation of translation quality for distant language pairs", In Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing, pp. 944-952, Cambridge, MA, October 2010. Association for Computational Linguistics. [Reference 3] Nikita Kitaev and Dan Klein., "Constituency parsing with a self-attentive encoder", In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2676-2686, Melbourne, Australia, July 2018. Association for Computational Linguistics. [Reference 4] Taku Kudo and John Richardson, "SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing", In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 66-71, Brussels, Belgium, November 2018. Association for Computational Linguistics. [Reference 5] Wuwei Lan, Chao Jiang, and Wei Xu, "Neural semi-Markov CRF for monolingual word alignment", In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 6815-6828, Online, August 2021. Association for Computational Linguistics. [Reference 6] Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer, "Multilingual denoising pre-training for neural machine translation", Transactions of the Association for Computational Linguistics, Vol. 8, pp. 726-742, 2020. [Figure 7]Toshiaki Nakazawa, Manabu Yaguchi, Kiyotaka Uchimoto, Masao Utiyama, Eiichiro Sumita, Sadao Kurohashi, and Hitoshi Isahara (LREC'16), pp. 101-1 2204-2208, Portoroz' (Z' and Z'ălăkălıka, Slovenia, May 2016. European Language Resources Association(ELRA). [Figure 8]Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Jie Ma, Alessandro Achille, Rishita Anubhai, Cicero Nogueira dos Santos, Bing Xiang, and Stefano Soatto. [Photograph 9]Matt Post、"A Call for Clarity in Reporting BLEU Scores"、In Proceedings of the Third Conference on Machine Translation: Research Papers, pp. 107-111. 186-191, Brussels, Belgium, October 2018. Association for Computational Linguistics. [Reference 10] Tianxiao Shen, Myle Ott, Michael Auli, and Marc'Aurelio Ranzato, "Mixture models for diverse machine translation: Tricks of the trade", In ICML, 2019. [Reference 11] Oriol Vinyals, L'ukasz Kaiser (L' is a stroke to L), Terry Koo, Slav Petrov, Ilya Sutskever, and Geoffrey Hinton, "Grammar as a foreign language", In NeurIPS, 2015. [Reference 12] Joern Wuebker, Spence Green, John DeNero, Sas'a Hasan, and Minh-Thang Luong, "Models and inference for prefix-constrained machine translation", In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 66-75, Berlin, Germany, August 2016. Association for Computational Linguistics. Although the embodiment of the present invention has been described in detail above, the present invention is not limited to such specific embodiment, and various modifications and variations are possible within the scope of the gist of the present invention described in the claims. [Explanation of symbols]
[0076] 10 Translation Device 11 Parser 12 Multilingual Model Fine Tuning Department 13 Translation Decoder 100 Drive device 101 Recording media 102 Auxiliary storage 103 Memory device 104 processors 105 Interface device B Bus
Claims
1. A translation unit configured to generate a syntax tree for an input sentence by inputting the input sentence into a translation model generated by fine-tuning a trained multilingual model using a pair of a syntax tree of a source language sentence and a target language sentence as training data. A translation device comprising:
2. The translation unit is configured to input an input sentence and a prefix including a component tag to the translation model, and to cause the translation model to output a syntax tree with the prefix as a constraint.
2. The translation device according to claim 1,
3. The translation model is configured to generate a plurality of types of token sequences constituting a portion from the beginning of an output sentence by a beam search, and to generate the syntax tree by performing a beam search using each of the token sequences as a prefix.
2. The translation device according to claim 1,
4. a multilingual model fine-tuning unit configured to generate a translation model that outputs a syntax tree of a target language sentence corresponding to an input sentence by fine-tuning a trained multilingual model using a pair of a syntax tree of a source language sentence and a syntax tree of a target language sentence as training data; A translation learning device comprising:
5. A translation procedure in which an input sentence is input to a translation model generated by fine-tuning a trained multilingual model using a pair of a syntax tree of a source language sentence and a target language sentence as training data, and a syntax tree for the input sentence is generated. A translation method characterized in that the above is executed by a computer.
6. A multilingual model fine-tuning procedure for generating a translation model that outputs a syntax tree of a target language sentence corresponding to an input sentence by fine-tuning a trained multilingual model using a pair of a syntax tree of a source language sentence and a syntax tree of a target language sentence as training data; A translation learning method characterized in that the above steps are executed by a computer.
7. A translation procedure in which an input sentence is input to a translation model generated by fine-tuning a trained multilingual model using a pair of a syntax tree of a source language sentence and a target language sentence as training data, and a syntax tree for the input sentence is generated. A program characterized by causing a computer to execute the above.
8. A multilingual model fine-tuning procedure for generating a translation model that outputs a syntax tree of a target language sentence corresponding to an input sentence by fine-tuning a trained multilingual model using a pair of a syntax tree of a source language sentence and a syntax tree of a target language sentence as training data; A program characterized by causing a computer to execute the above.