Machine translation strengthening method based on bilingual dictionary injection

Through the machine translation enhancement method based on bilingual dictionary injection, bilingual alignment and data enhancement of unsupervised monolingual corpus is solved, and the problem of traditional neural machine translation systems performing poorly in long text processing and proprietary field translation is achieved, achieving higher quality and accurate translation effects.

CN120068893AInactive Publication Date: 2025-05-30HARBIN INST OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510107862.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional neural machine translation systems based on self-attention mechanisms have high computational costs when processing long texts, limited long sequence processing capabilities, and lack structural knowledge, especially in proprietary field translation.

Method used

Using a machine translation enhancement method based on bilingual dictionary injection, a bilingual dictionary is generated by bilingual alignment of large-scale unsupervised monolingual corpus, a parallel corpus statistical word hit rate was introduced, a Memory Bank was established and data enhancement was performed, and finally a deep adversarial network model was used for model training.

Benefits of technology

It significantly improves the performance of machine translation systems in proprietary field translation, improves translation quality and accuracy, and solves the problem of lack of long text processing and structural knowledge in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068893A_ABST
    Figure CN120068893A_ABST
Patent Text Reader

Abstract

The invention discloses a machine translation strengthening method based on bilingual dictionary injection, and belongs to the technical field of machine translation strengthening. The problem that in the prior art, a traditional machine translation strengthening method is poor in model performance for translation in the special field is solved. The method comprises the following steps: performing bilingual alignment on large-scale unsupervised monolingual corpora to generate a bilingual dictionary; parallel corpora are introduced into the bilingual dictionary, the hit rate of each word pair in the bilingual dictionary in the parallel corpora is counted, a Memory Bank is established and the hit rate is recorded, the importance of the word pairs is sorted according to the hit rate, and the sorted bilingual dictionary is obtained; and performing data enhancement on the sorted source end data in the bilingual dictionary through Memory Bank, and inputting the data into a deep adversarial network model for model training to obtain a trained deep adversarial network model. According to the method, the parallel corpora are effectively subjected to data enhancement, the generation quality of a machine translation system is improved, and the method can be applied to machine translation modeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for enhancing machine translation, and particularly to a method for enhancing machine translation based on bilingual dictionary injection, belonging to the technical field of machine translation enhancement. Background Art

[0002] In the prior art, Transformer is a neural machine translation system based on the self-attention mechanism, which includes an encoder and a decoder. The encoder and decoder structures of Transformer are both composed of a multi-head self-attention module and a feed-forward neural network. Among them, the introduction of the attention module and position embedding enables the Transformer architecture to process sequence information. Transformer adopts a multi-head attention mechanism, and each head is calculated from three matrices (query Q, key K, value V) obtained by linear transformations (query weight W Q , key weight W K , value weight W V );

[0003] The result of the multi-head attention mechanism Attention i (Q i , K i , V i ) is expressed as:

[0004]

[0005] where i ∈ [1, a], a is the number of heads of the self-attention module, d k is the dimension of the key K, softmax is the attention weight function, T represents transpose, Q i is the query matrix of any self-attention module, K i is the key matrix of any self-attention module, and V i is the value matrix of any self-attention module;

[0006] After that, the attention results of each head are concatenated to obtain the concatenation result MultiAttention(Q, K, V);

[0007] The concatenation result MultiAttention(Q, K, V) is expressed as:

[0008] MultiAttention(Q, K, V) = Concat(Attention i )W O

[0009] where W Ois the output weight matrix of the multi-head attention mechanism, and Concat is the concatenation function;

[0010] The feed-forward neural network of Transformer is a two-layer fully connected layer. Among them, the first fully connected layer uses ReLU as the activation function;

[0011] The feed-forward neural network FFN(x) of Transformer is expressed as:

[0012] FFN(x) = ReLU(0, xW 1 + b 1 )W 2 + b 2

[0013] Among them, W 1 is the first trainable parameter of the first fully connected layer, b 1 is the second trainable parameter of the first fully connected layer, W 2 is the first trainable parameter of the second fully connected layer, b 2 is the first trainable parameter of the second fully connected layer;

[0014] Since self-attention and the fully connected network cannot capture the relative position information of the input sequence, Transformer uses the method of positional encoding PE to add position information at the encoding end and the decoding end. The position is an absolute position information and can also represent the relative position between tokens. Among them, PE (pos,2i) and PE (pos,2i+1) represent the position vectors of even positions and odd positions respectively;

[0015] The position vectors of even positions and odd positions are respectively expressed as:

[0016]

[0017] Among them, pos represents the position of the token in the sequence, and h represents the hidden layer size of the model.

[0018] However, traditional neural machine translation systems based on self-attention mechanisms still have the following problems: (1) High computational cost: When processing long texts, the computational and memory requirements increase sharply, which may lead to slower model training and inference speeds; (2) Limited long-sequence processing ability: Although the self-attention mechanism can capture long-distance dependencies, in practical applications, for very long input sequences, the model may still face difficulties and important information may be lost; (3) Lack of structural knowledge: The self-attention mechanism mainly relies on data for learning. For translations in certain specific fields, such as medical and other professional fields, it may lack necessary domain knowledge and structural information, thus affecting the accuracy and professionalism of translations.

[0019] In summary, a method for enhancing machine translation based on bilingual dictionary injection is required. Summary of the Invention

[0020] A brief overview of the present invention is given below in order to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description that follows.

[0021] In view of this, in order to solve the problem that the model performance of traditional machine translation enhancement methods for domain-specific translation in the prior art is poor, the present invention provides a method for enhancing machine translation based on bilingual dictionary injection.

[0022] The technical solution is as follows: A method for enhancing machine translation based on bilingual dictionary injection includes the following steps:

[0023] S1. Bilingual alignment is performed on a large-scale unsupervised monolingual corpus to generate a bilingual dictionary;

[0024] S2. Parallel corpus is introduced into the bilingual dictionary, the hit rate of each word pair in the parallel corpus is counted, a Memory Bank is established and the hit rate is recorded, and the word pairs are sorted according to the hit rate to obtain a sorted bilingual dictionary;

[0025] S3. The source-side data in the sorted bilingual dictionary is data-augmented through the Memory Bank and input into a deep adversarial network model for model training to obtain a trained deep adversarial network model.

[0026] Further, in S1, for the large-scale unsupervised monolingual corpus collected, namely the source-side monolingual corpus and the target-side monolingual corpus, the 20% words with the highest word frequencies are extracted, input into the deep adversarial network model for training, the linear mapping W between the two languages is initialized, and based on the linear mapping W obtained from the deep adversarial network model, the source language vocabulary corresponding to the source-side monolingual corpus and the target language vocabulary corresponding to the target-side monolingual corpus are traversed. If a certain word pair between the two languages is the nearest neighbor to each other, the word pair is added to the bilingual dictionary.

[0027] Further, in S2, an external information sequence Dict_S is introduced into the source-side sentence Sent_S of the bilingual dictionary, word segmentation is performed, and the external information sequence Dict_S is restricted to obtain the relationship between the word-segmented source-side sentence and the external information sequence;

[0028] The relationship between the word-segmented source-side sentence and the external information sequence is expressed as:

[0029]

[0030] Among them, k is a predefined constant, is the external information sequence after word segmentation, is the external information sequence after word segmentation;

[0031] The hit rate of the bilingual dictionary is defined as: the word pair in the bilingual dictionary after word segmentation ( word S ,word T) matches in a pair of segmented parallel sentences in the training set of the deep adversarial network model, and is regarded as the word pair ( word S ,word T) hits once in the training set;

[0032] Sort the importance of word pairs according to the hit rate to obtain the sorted bilingual dictionary.

[0033] Furthermore, in S3, for the source sequence in each parallel corpus, retrieve the words that appear in the sorted bilingual dictionary, directly splice the word translations involved in the source sentence of the source sequence at the beginning of the input side front end of the deep adversarial network model, train the deep adversarial network model, extract the word corresponding to the current word translation, splice it in front of the word translation, and obtain the trained deep adversarial network model, and output the target side;

[0034] The modeling process of the deep adversarial network model is expressed as:

[0035] Output l =argmax P(Output l |D 1 ,Input,Output [1,l-1] )

[0036] Among them, Input is the source sequence, Output l is the output of the deep adversarial network model at the current moment, Output [1,l-1] is the output sequence before the l-th moment, D 1 is the set of word translations spliced at the source end, and argmax is the optimization function;

[0037] The training process of the deep adversarial network model is expressed as:

[0038] Output l '=argmax P(Output l '|D 2 ,Input,Output [1,l-1] )

[0039] Among them, Output l ' is the output of the trained deep adversarial network model at the current moment, D 2 is a subset of the word translation set D spliced at the source end 1 .

[0040] The beneficial effects of the present invention are as follows: The present invention provides a method for data augmentation of parallel corpora based on bilingual dictionaries and improving the generation quality of machine translation systems. By using the MUSE method to align unsupervised monolingual corpora, a bilingual dictionary is generated; then, based on the bilingual dictionary, the original parallel corpus is augmented with data, and then the augmented parallel corpus is used to train the translation model; the present invention enables the translation model to utilize a large amount of unsupervised monolingual corpora during the training process with supervised parallel corpora, and improves the performance of the translation system by aligning monolingual corpora to generate bilingual dictionaries and injecting bilingual dictionaries into parallel corpora; for translation tasks in specific domains, by using the bilingual dictionary data augmentation method proposed in the present invention, compared with traditional parallel corpus training data, the performance of the translation model in specific domains can be significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0042] Figure 1 is a schematic flow chart of a machine translation enhancement method based on bilingual dictionary injection;

[0043] Figure 2 is a schematic flow chart of an embodiment of a machine translation enhancement method based on bilingual dictionary injection;

[0044] Figure 3 is a schematic flow chart of an embodiment of data augmentation. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] In order to make the technical solutions and advantages in the embodiments of the present invention clearer and more understandable, the following further details the exemplary embodiments of the present invention with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0046] Refer to Figures 1-3 for a detailed description of this embodiment. A machine translation enhancement method based on bilingual dictionary injection specifically includes the following steps:

[0047] S1. Bilingual alignment is performed on large-scale unsupervised monolingual corpora to generate a bilingual dictionary;

[0048] S2. Parallel corpora are introduced into the bilingual dictionary, the hit rates of each word pair in the parallel corpora in the bilingual dictionary are counted, a Memory Bank is established and the hit rates are recorded, and the word pairs are sorted according to the hit rates to obtain a sorted bilingual dictionary;

[0049] S3. Data augmentation is performed on the source-side data in the sorted bilingual dictionary through the Memory Bank and input into a deep adversarial network model for model training to obtain a trained deep adversarial network model. The source-side corpus is input into the trained deep adversarial network model, and inference is performed in combination with the bilingual dictionary to output the target-side corpus.

[0050] Furthermore, in S1, for the large-scale unsupervised monolingual corpora collected, namely the source-side monolingual corpus and the target-side monolingual corpus, the 20% words with the highest word frequencies are extracted and input into the deep adversarial network model for training. The linear mapping W between the two languages is initialized. Based on the linear mapping W obtained from the deep adversarial network model, the source language vocabulary corresponding to the source-side monolingual corpus and the target language vocabulary corresponding to the target-side monolingual corpus are traversed. If a certain word pair between the two languages is the nearest neighbor to each other, the word pair is added to the bilingual dictionary;

[0051] Specifically, the main purpose of step S1 is to construct a bilingual dictionary based on unsupervised monolingual corpora, thereby supporting the data augmentation method proposed in the present invention. The bilingual dictionary involved is derived from the MUSE method open-sourced by facebook research. The present invention uses the MUSE method to perform bilingual word vector alignment. The MUSE method is a method for bilingual word vector alignment from unsupervised monolingual corpora. Step S1 of generating a bilingual dictionary based on unsupervised monolingual corpora accumulates parallel word pairs for subsequent step S2;

[0052] The MUSE method is a method for bilingual word vector alignment from unsupervised monolingual corpora. For a training corpus containing 200k words, 50k words with the highest word frequencies are extracted, and the linear mapping W between the source language and the target language is initialized by training with a deep adversarial network model. Assume that the word embeddings of the source language with a vocabulary size of n and the target language with a vocabulary size of m are X = {x 1 …x n} and Y = {y 1 …y m} respectively. First, a discriminator is defined so that the discriminator can distinguish the linear mapping WX = {Wx 1 …Wx n} randomly sampled from the source language with a vocabulary size of n and the word embeddings Y = {y 1 …ym Elements of}, and then define a generator to train the linear mapping W to make the linear mapping WX of the source language with vocabulary size n as close as possible to the word embedding Y of the target language with vocabulary size m, so that the discriminator is difficult to judge the source of the word embedding. The MUSE method uses a deep adversarial network model for training, and defines the discriminator L D and the generator L W The loss functions of are defined as L D (θ D |W) and L W (W|θ D );

[0053] The loss function L D of the discriminator is defined as L D (θ D |W) is expressed as:

[0054] And are respectively defined as L D (θ D |W) and L W 9W|θ D)

[0055]

[0056] where source represents the data source, represents the profile distribution predicted by the model;

[0057] The loss function L W of the generator is expressed as: W9 W|θ D) is expressed as:

[0058]

[0059] To ensure orthogonality, the linear mapping W is updated to obtain the updated linear mapping W t+1 ;

[0060] The updated linear mapping W t+1 is expressed as:

[0061] W t+1 = ( 1 + β ) W - β ( WW T ) W

[0062] where the value of β is 0.01;

[0063] The updated linear mapping W obtained based on the deep adversarial network t+1Traverse the vocabulary lists of the source language and the target language. If a word pair between the two languages is a nearest neighbor to each other, add the word pair to the bilingual dictionary. In a high-dimensional space, the nearest neighbor points between word pairs are often asymmetric, which will lead to the phenomenon of the hub problem. Some word embeddings are the nearest neighbors of many points, while some word embeddings are not the nearest neighbors of any point;

[0064] To address this issue, in the MUSE method, it is defined that is the set of the K nearest words in the word embedding Y of the target language to the sth source language word x s to obtain the average distance between the sth source language word x s and its K nearest neighbors

[0065] The sth source language word x s and the average distance to its K nearest neighbors is expressed as:

[0066]

[0067] where is the vector of the sth source language word, and y t is the word embedding of the tth target language word;

[0068] Adopt the distance metric method CSLS in cross-lingual word embedding alignment to measure the distance between y t to obtain the distance between y t which is

[0069] y t The distance between them is expressed as:

[0070]

[0071] where r S( y t) represents the average distance between the word embedding y t of the tth target language word and its K nearest neighbors;

[0072] The MUSE method only selects the mutually nearest word pairs based on the distance which significantly reduces the size of the dictionary generation, but improves the accuracy of the dictionary and the overall performance.

[0073] Furthermore, in S2, an external information sequence Dict_S is introduced to the source-side sentence Sent_S of the bilingual dictionary, tokenized, and the external information sequence Dict_S is restricted to obtain the relationship between the tokenized source-side sentence and the external information sequence;

[0074] The relationship between the segmented source - side sentence and the external information sequence is expressed as:

[0075]

[0076] where k is a predefined constant, is the segmented external information sequence, is the segmented external information sequence;

[0077] The hit rate of the bilingual dictionary is defined as: the word pair ( word S ,word T) in the segmented bilingual dictionary matches in a pair of segmented parallel sentences in the training set of the deep adversarial network model, and is regarded as the word pair ( word S ,word T) hitting once in the training set;

[0078] Sort the word pairs according to the hit rate to obtain the sorted bilingual dictionary;

[0079] Specifically, since some word pairs in the bilingual dictionary appear frequently in the parallel corpus of the training set, without adding additional bilingual dictionary information, the model can also model the target word pairs with the help of the original parallel corpus;

[0080] Therefore, before data augmentation, pre - calculate the hit rate of each word pair in the bilingual dictionary in the parallel corpus and store it in the Memory Bank. For the source - side sentence that needs data augmentation, considering the training cost and the length of the model input, and that introducing too much external information may lead to noise introduction, limit the added external information sequence;

[0081] The purpose of step S2 is to increase the information density in the data augmentation step and prepare for step S3.

[0082] Furthermore, in S3, for each source - side sequence in the parallel corpus, retrieve the words that appear in the sorted bilingual dictionary, directly splice the translations of the words involved in the source - side sentence of the source - side sequence at the beginning of the input side of the deep adversarial network model, train the deep adversarial network model, extract the word corresponding to the current word translation, splice it in front of the word translation, and obtain the trained deep adversarial network model, and output the target side;

[0083] Output l =argmax P(Output l |D 1 ,Input,Output [1,l-1] )

[0084] Among them, Input is the source sequence, and Output l is the output of the deep adversarial network model at the current moment. Output [1,l-1] is the output sequence before the l-th moment, and D 1 is the set of word translations concatenated at the source end. argmax is the optimization function, corresponding to source-side data augmentation - Method 1;

[0085] The training process of the deep adversarial network model is expressed as:

[0086] y t ' = argmax P(y t '|D 2 , X, Y [1,t-1] )

[0087] Output l ' = argmax P(Output l '|D 2 , Input, Output [1,l-1] )

[0088] Among them, Output l ' is the output of the trained deep adversarial network model at the current moment, and D 2 is a subset of the set of word translations D concatenated at the source end, corresponding to source-side data augmentation - Method 2; 1 Specifically, referring to

[0089] and adopting the data augmentation method in it to modify the source-side statement. Taking the English-Chinese translation direction as an example, search for word pairs that appear in the bilingual dictionary at the source end of the parallel corpus and concatenate the word pairs in the training corpus. In order to avoid introducing noise into the sequence information of the original corpus as much as possible, do not directly replace the vocabulary in the original sequence, and add word translations at different positions in the source sequence. Among them, " Figure 3 ” is a special token without specific semantic meaning, similar to the [SEP] token in the BERT model, used to separate word translations and the original sequence;

[0090] Based on the bilingual dictionary, the present invention concatenates all possible translations that the source side may involve at the front end of the model input side, enabling the model to assist in the translation task by selecting the added word translations during training. As shown in Source-side Data Augmentation - Method 1, this processing method is expressed as the deep adversarial network model performing "fill in the blanks with words" for the word translations involved in the source side based on the sequence information in the sentence, that is, the deep adversarial network model selects the part that should be output from the possible target language word translations based on the source language input sequence, i.e., the source-side sequence. In contrast, the traditional neural machine translation model only models the relationship between the two languages and outputs the translation results word by word only based on the information of the source language, and the translation effect is far inferior to that of the present invention;

[0091] The modeling process of the traditional neural machine translation model is expressed as:

[0092] y t ”=argmaxP(y t ”|X,Y [1,t-1] )

[0093] where y t ” is the output of the traditional neural machine translation model at the current moment;

[0094] To better represent the corresponding relationship of word translations on the source side, the deep adversarial network model is trained to concatenate the word translations in the form of source language - target language word pairs in the bilingual dictionary on the source side, that is, to further increase the corresponding relationship between the word translations and the vocabulary in the source-side sequence on the basis of Source-side Data Augmentation - Method 1, as shown in Source-side Data Augmentation - Method 2;

[0095] Refer to Figure 2 , in the inference stage, based on the bilingual dictionary generated by the above steps or the bilingual dictionary in the target specific domain, data augmentation is performed on the source-side corpus to be translated, and the translated target-side corpus is obtained based on the trained translation model;

[0096] In the English-Chinese translation task in a proprietary domain, if trained according to traditional English-Chinese parallel corpora, for the source sentence to be translated "The patient was diagnosed with hypertrophic cardiomyopathy.", the traditional neural machine translation model will output: "The patient was diagnosed with cardiac hypertrophy."; among them, the proper noun "hypertrophic cardiomyopathy" is translated as "cardiac hypertrophy.", rather than the translation "hypertrophic cardiomyopathy" that should be in the context of this task. Using the method proposed in the present invention for training the deep adversarial network model and splicing the word pair "hypertrophic cardiomyopathy→hypertrophic cardiomyopathy" in the bilingual dictionary of this proprietary domain during the inference stage, the trained deep adversarial network model will output: "The patient was diagnosed with hypertrophic cardiomyopathy.", presenting a more accurate translation effect.

[0097] Although the present invention has been described based on a limited number of embodiments, those skilled in the art in this technical field will understand, based on the above description, that other embodiments can be envisioned within the scope of the present invention thus described. In addition, it should be noted that the language used in this specification is mainly selected for the purpose of readability and teaching, rather than for the purpose of explaining or limiting the subject matter of the present invention. Therefore, many modifications and variations will be obvious to those of ordinary skill in the art in this technical field without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure of the present invention is illustrative rather than restrictive, and the scope of the present invention is defined by the appended claims. ​

Claims

1. A machine translation enhancement method based on bilingual dictionary injection, characterized in that: The following steps are involved: S1. Perform bilingual alignment on large-scale unsupervised monolingual corpora and generate bilingual dictionaries; S2. Introduce parallel corpora into the bilingual dictionary, count the hit rates of each word pair in the bilingual dictionary in the parallel corpus, establish a MemoryBank and record the hit rates, sort the word pairs by importance according to the hit rates, and obtain a sorted bilingual dictionary; S3. Perform data enhancement on the source data in the sorted bilingual dictionary through MemoryBank and input it into the deep adversarial network model for model training to obtain a trained deep adversarial network model.

2. According to claim 1, a method for enhancing machine translation based on bilingual dictionary injection is characterized in that: In S1, for the collected large-scale unsupervised monolingual corpus, i.e., the source monolingual corpus and the target monolingual corpus, 20% of the words with the highest word frequency are extracted and input into the deep adversarial network model for training, and the linear mapping W between the two languages ​​is initialized. Based on the linear mapping W obtained by the deep adversarial network model, the source language vocabulary corresponding to the source monolingual corpus and the target language vocabulary corresponding to the target monolingual corpus are traversed. If a word pair between the two languages ​​is a nearest neighbor to each other, the word pair is added to the bilingual dictionary.

3. The method for enhancing machine translation based on bilingual dictionary injection according to claim 2, characterized in that: In S2, the source sentence Sent_S of the bilingual dictionary is introduced into the external information sequence Dict_S, segmented, and the external information sequence Dict_S is restricted to obtain the relationship between the source sentence after segmentation and the external information sequence; The relationship between the source sentence after word segmentation and the external information sequence is expressed as: Where k is a predefined constant, is the external information sequence after word segmentation, is the external information sequence after word segmentation; The hit rate of a bilingual dictionary is defined as: the word pair word in the bilingual dictionary after word segmentation S ,word T If a pair of parallel sentences in a pair of segmented words in the training set of the deep adversarial network model matches, it is considered a word pair. S ,word T Hit once in the training set; The word pairs are ranked according to their importance according to the hit rates to obtain a ranked bilingual dictionary.

4. The method for enhancing machine translation based on bilingual dictionary injection according to claim 3, characterized in that: In S3, for each source sequence in the parallel corpus, the words appearing therein are searched in the sorted bilingual dictionary, and the word translation involved in the source sentence of the source sequence is directly spliced ​​at the beginning of the sentence, that is, the front end of the deep adversarial network model input side, and the deep adversarial network model is trained, and the word corresponding to the current word translation is extracted and spliced ​​before the word translation to obtain the trained deep adversarial network model, and output to the target end; The modeling process of the deep adversarial network model is expressed as: Output l =argmaxPOutput l D 1 ,Input,Output 1,l-1 Among them, Input is the source sequence, Output l Output is the output of the deep adversarial network model at the current moment. [1,l-1] is the output sequence before time l, D 1 is the word translation set concatenated at the source end, and argmax is the optimization function; The training process of the deep adversarial network model is expressed as: Output l '=argmaxPOutput l 'D 2 ,Input,Output 1,l-1 Among them, Output l ' is the output of the trained deep adversarial network model at the current moment, D 2 The word translation set D concatenated at the source end 1 A subset of .

Citation Information

Patent Citations

  • Chinese- Vietnamese unsupervised neural machine translation method fusing EMD minimized bilingual dictionary

    CN111753557A

  • Chinese-Vietnamese unsupervised neural machine translation method based on shared encoder

    CN112287694A

  • Language code conversion vocabulary overlapping enhancement method for unsupervised neural machine translation

    CN114492476A

  • Method and system for carrying out data enhancement on neural machine translation model

    CN117540755A

  • Apparatus and method for unsupervised learning translation relationships among words and phrases in the statistical machine translation system

    KR1020080052282A