A language translation method based on a prefix-tuning machine translation model
By building a prefix tuning machine translation model, the poor performance of the machine translation model in low-resource and high-resource scenarios is solved, the learning and expression ability and controllability of the model are enhanced, the amount of training parameters is reduced, and the translation quality is improved.
Patent Information
- Application Number
- CN202310149564.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-02-22
AI Technical Summary
Existing machine translation models perform poorly in low-resource and high-resource scenarios, with problems of text grammar errors, duplication and inconsistencies, and cross-language pre-trained models may destroy the knowledge of the original language model in low-resource scenarios.
The prefix tuning machine translation model is constructed, and the prefix tuning module is constructed, and the control signal matrix is reparameterized by MLP neural network is used to transform the attention layer of the mBART translation model. Only the parameters of the prefix tuning module are trained to keep the original mBART model unchanged, and the model's learning and expression ability and controllability are enhanced.
Improve translation performance in low-resource scenarios, reduce the amount of training parameters, reduce training memory and time, effectively alleviate syntax errors and repetition problems of output text, and adapt to various language scenarios.
Smart Images

Figure CN116151278B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of translation, and particularly relates to a language translation method based on a prefix-tuning machine translation model. Background Art
[0002] A machine translation model is a tool for language translation. To train a high-quality machine translation model, a large amount of high-quality aligned parallel corpus is required. However, the vast majority of languages in the world lack large-scale, high-quality, and high-coverage parallel corpora. Building a high-quality aligned parallel corpus requires expensive human and material costs. Therefore, neural machine translation under the condition of scarce corpus has always been a problem to be solved.
[0003] The prior art uses the large cross-lingual pre-trained model mBART for training. mBART contains rich multilingual knowledge and can improve the translation performance of the translation model in the case of scarce corpus.
[0004] Although the cross-lingual pre-trained model can improve the performance of the translation model to a certain extent, using the cross-lingual pre-trained language model in low-resource scenarios may still lead to the inability to train the translation model well due to too little data, and may also damage the knowledge of the original language model, resulting in an unsatisfactory performance improvement effect.
[0005] In high-resource scenarios, since the number of trainable parameters of the current prefix-tuning strategy is only 10% of the original training parameters, the expressive ability that can be learned is weaker than that of traditional full-parameter fine-tuning. Facing numerous different language scenarios, it is difficult for traditional translation models to effectively control the output to adapt to the current language scenario, and using a neural network to build a machine translation model will cause problems such as text grammar errors, repetitions, and contradictions.
[0006] In summary, the prior art has problems in that the machine translation model has poor performance in low-resource and high-resource scenarios using cross-lingual pre-trained models, and there are text grammar errors, repetitions, and contradictions. Summary of the Invention
[0007] In order to overcome the above deficiencies of the prior art, the present invention provides a language translation method based on a prefix-tuning machine translation model.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] A language translation method based on a prefix-tuning machine translation model, comprising:
[0010] Constructing a prefix-tuning machine translation model, which includes:
[0011] Constructing a prefix-tuning module, which includes:
[0012] Initialize the control attribute tags that tune the word order, length, and language style of the target language in the machine translation model with the control prefix to a vector S with consistent dimensions i , for the vector S i Perform a linear weighted combination to obtain the control signal matrix S;
[0013] Preset three groups of control signal matrices, namely S1, S2, and S3;
[0014] Construct the first MLP neural network, and use the first MLP neural network to reparameterize S1 to output the first prefix sequence key-value pair P 1key -P 1value ;
[0015] Construct the second MLP neural network, and use the second MLP neural network to reparameterize S2 in it to output the second prefix sequence key-value pair P 2key -P 2value ;
[0016] Construct the third MLP neural network, and use the third MLP neural network to convert S3 into the Q value, K value, and V value of the cross-attention layer of the mBART translation model;
[0017] Modify the mBART translation model, which includes:
[0018] Input the Q value, K value, and V value output by the third MLP neural network into the cross-attention layer of the mBART translation model as trainable parameters in the cross-attention layer;
[0019] For P 1key -P 1value Concatenate with the key-value pair in the encoder self-attention layer of the source language text to form a new key-value pair K' l -V' l , and input K' l -V' l Into the self-attention layer of the mBART encoder to achieve the improvement of the self-attention layer of the mBART encoder;
[0020] For the second prefix sequence key-value pair P 2key -P 2value Concatenate with the key-value pair in the decoder self-attention layer of the target language text to form a new key-value pair K' l 1-V' l 1, and input K' l 1-V' l 1 into the self-attention layer of the mBART decoder to achieve the improvement of the self-attention layer of the mBART decoder;
[0021] For K' l -V'l In the cross-attention layer of the input mBART decoder, the improvement of the cross-attention layer of the mBART decoder is realized;
[0022] Train a prefixed-tuning machine translation model to realize the construction of the prefixed-tuning machine translation model;
[0023] Input the source language into the machine translation model, and the machine translation model outputs the target language to realize language translation.
[0024] Furthermore, in the training of the prefixed-tuning machine translation model, only the prefixed-tuning module is trained, and the parameters of the original mBART model remain unchanged.
[0025] Furthermore, the objective function for training the prefixed-tuning module is:
[0026]
[0027] where Ψ are the parameters of mBART and do not change during training; x is the source language; y is the target language output by the model, and α, θ, are the model parameters of the first MLP neural network, the second MLP neural network, and the third MLP neural network.
[0028] Furthermore, set the control attribute label vector S of the three groups of control signal matrices S1, S2, and S3 i to an initial value of 0.
[0029] Furthermore, the linear weighted combination of the vector Si to obtain the control signal matrix S is:
[0030]
[0031] where S i is the control attribute vector; W i is the trainable weight used to adjust the intervention intensity of each control attribute; S is the control signal matrix; N is the number of vectors Si.
[0032] Furthermore, the control attribute for controlling the word order of the target language output by the prefixed-tuning machine translation model is the deviation intensity δ(s) of the non-diagonal alignment {(i, j)} between the source language sentence src = {x1, x2,..., x l} and the target language sentence tgt = {y1, y2,..., y m}, where i ∈ {1, 2,..., l}, j ∈ {1, 2,..., m};
[0033] where the deviation intensity δ(s) is:
[0034]
[0035] Among them, #{(i,j)} represents the alignment base; when δ(s) = 0, the sentence pair is in a strictly monotonic situation. At this time, l = m and {(i,j)} is a strictly increasing bijective mapping; the smaller δ(s) is, the higher the monotonicity between the source language and the target translation is.
[0036] Furthermore, the trainable weight W of the word order i is:
[0037] W i = δ(s) + k,
[0038] where δ(s) is the intensity of the non - diagonal alignment deviation of each vocabulary in the same sentence of the source language and the target language corresponding to a matrix, and k is the offset value.
[0039] Furthermore, the improvement of the self - attention layer of the mBART encoder includes:
[0040] In the l - th Transformer block layer of the encoder, the query Q is obtained by linearly transforming the hidden state of the source - language text sequence l , the key K l and the value V l ;
[0041] The key - value pair P key -P value , is concatenated with the key K l and the value V l to obtain a new key - value pair K' l -V' l :
[0042]
[0043] The K' l -V' l is passed into the encoder self - attention layer to achieve the improvement of the encoder self - attention layer:
[0044]
[0045] Among them, Q l is the query in the encoder self - attention mechanism, generated based on the source language, and K' l -V' l is the key - value pair.
[0046] Furthermore, the improved cross - attention layer of the mBART encoder is:
[0047]
[0048] Among them, It is the query in the decoder self-attention mechanism, generated based on the source language.
[0049] The language translation method based on the prefix-tuning machine translation model provided by the present invention has the following beneficial effects:
[0050] The present invention integrates the control attributes for controlling the output language format into the prefix-tuning module, enhancing the learning and expression ability of the model, enabling it to effectively learn the relationship between the two semantic spaces; due to the control of the output by the control mechanism, the controllability of the model is strengthened, it can effectively adapt to various language scenarios, and effectively alleviate the problems of grammar errors, repetitions, and contradictions in the output text. And in the training process of the model of the present invention, only the parameters in the prefix-tuning module are trained, without changing the parameters in the mBART model, freezing the weights of the cross-lingual pre-trained model, effectively reducing 90% of the trainable parameter quantity, reducing the training memory and time, and improving the translation performance in low-resource scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present invention and their design schemes, the accompanying drawings required for the present embodiments will be briefly introduced below. The accompanying drawings in the following description are only partial embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a schematic structural diagram of a machine translation model based on prefix tuning according to an embodiment of the present invention;
[0053] Figure 2 It is a block diagram of the prefix-tuning part of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] In order to enable those skilled in the art to better understand the technical solutions of the present invention and implement them, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and cannot be used to limit the protection scope of the present invention.
[0055] In addition, terms such as "first", "second", etc. are for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of the present invention, it should be noted that unless otherwise clearly specified or limited, the terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. In the description of the present invention, unless otherwise stated, the meaning of "plural" is two or more, which will not be elaborated here.
[0056] Embodiment:
[0057] The present invention provides a language translation method based on a prefix-tuning machine translation model, specifically as Figure 1 shown, including: constructing a prefix-tuning machine translation model, which includes:
[0058] Constructing a prefix-tuning module, which includes:
[0059] Initializing the control attribute tags that control the word order, length, and language style of the target language output by the prefix-tuning machine translation model into a vector S with consistent dimensions i , performing a linear weighted combination on the vector S i to obtain a control signal matrix S; presetting three groups of control signal matrices, namely S1, S2, and S3; constructing a first MLP neural network, and using the first MLP neural network to reparameterize S1 to output the first prefix sequence key-value pair P 1key -P 1value ; constructing a second MLP neural network, and using the second MLP neural network to reparameterize S2 in it to output the second prefix sequence key-value pair P 2key -P 2value ; constructing a third MLP neural network, and using the third MLP neural network to convert S3 into the Q value, K value, and V value of the cross-attention layer of the mBART translation model.
[0060] Modifying the mBART translation model, which includes:
[0061] Inputting the Q value, K value, and V value output by the third MLP neural network into the cross-attention layer of the mBART translation model as the trainable parameters in the cross-attention layer; splicing P 1key -P 1value with the key-value pairs in the encoder self-attention layer of the source language text to form new key-value pairs K' l -V' l , and using K' l -V' lInto the self-attention layer of the mBART encoder to improve the self-attention layer of the mBART encoder; the second prefix sequence key-value pair P 2key -P 2value Is concatenated with the key-value pairs of the target language text passing through the key-value pair in the decoder self-attention layer to form a new key-value pair K' l 1-V' l 1, and K' l 1-V' l 1 is input into the self-attention layer of the mBART decoder to improve the self-attention layer of the mBART decoder; K' l -V' l Is input into the cross-attention layer of the mBART decoder to improve the cross-attention layer of the mBART decoder;
[0062] Train the prefix-tuning machine translation model to build the prefix-tuning machine translation model.
[0063] Input the source language into the machine translation model, and the machine translation model outputs the target language to achieve language translation.
[0064] The following are the detailed embodiments of the present invention:
[0065] 1. Model Overview
[0066] This model will first apply the control mechanism integrated into the prefix-tuning method to the field of text translation. In the machine translation task, by solidifying the language model and optimizing the prefix encoding for fine-tuning, different trained prefix encodings are used to extract features in the language model, where the prefix encoding is generated by the corresponding prefix network through linear weighting of various control features. Here, the control features are descriptions of the attributes that affect the translation result, such as attributes like word order and length ratio.
[0067] The basic architecture uses the mBART model. In the mBART model structure, the corresponding Prefix for different attention calculation methods is also different, which is more suitable for extracting semantic features and control features stored in the language model through different prefix encodings, thereby training a high-quality machine translation model. That is, on the basis of keeping the mBART weights fixed, some soft-tokens are added to train a Seq2Seq machine translation model.
[0068] 2. Detailed Design
[0069] 2.1 Encoder Module
[0070] 2.1.1 Control Mechanism
[0071] The control mechanism in machine translation is to add special tags before the parallel sentence pairs are input into the model to control the model's output. First, each control attribute tag (such as word order, length, language style, etc.) is initialized to a vector value S with the same dimension. i , If there is no attribute mark during input, all vector values are initialized to 0. This can effectively maintain the training results of incomplete annotation data caused by expensive manual annotation and difficult-to-define attributes, and avoid the impact of removing the control of certain attributes on the model performance in subsequent experiments. And different calculation methods are designed for the characteristics of different control attributes. For example, the word order here refers to the degree of closeness of the word order in the source language sentence and the target language sentence. Therefore, it can be defined as the intensity δ(s) of the non-diagonal alignment deviation. For the language pair s = {src, tgt} and the source language sentence src = {x1, x2, …, x l} and the target language sentence tgt = {y1, y2, …, y m} between the alignment {(i, j)}, where i ∈ {1, 2, …, l} and j ∈ {1, 2, …, m}, the deviation intensity δ(s) is specifically defined as follows:
[0072]
[0073] Among them, #{(i, j)} represents the cardinality of the alignment. When δ(s) = 0, the sentence pair shows a strictly monotonic situation. At this time, l = m and {(i, j)} is a strictly increasing bijective mapping. The smaller δ(s) is, the higher the monotonicity between the source language and the target translation. Generally, in order to avoid collision with the neutral mode, that is, when S i = 0, a small offset value k will be added, that is, W i = δ(s) + k.
[0074] Then the initialized vector values are linearly weighted and combined together to obtain the control signal S. As shown in formula (1).
[0075]
[0076] Among them, W i is a trainable weight used to adjust the intervention intensity of each attribute. Using the method of linear addition can reduce the order dependence of attribute tags and increase the interpretability of multi-attribute control, etc. Finally, for different control attributes, there are corresponding processing strategies to obtain the attribute vector value representation.
[0077] 2.1.2 Prefix Module
[0078] Since the prefix module in the model is equivalent to the prompt part in prompt learning, and prompt learning has proven that conditioning on appropriate contexts can control the output of the language model without changing the language model parameters.
[0079] Therefore, three prefix ids (corresponding to token ids) are predicted in the experiment. Let the prefix length be 10. The three prefix ids are: p1,..., p 10 ,p 11 ,..., p 20 ,p 21 ,..., p 30 ,corresponding to three embedding matrices S1, S2, and S3. These prefix sequences are control signals obtained by randomly initializing the embeddings of the attribute annotations of the parallel language pairs.
[0080] To make the training stable, S1, S2, and S3 are reparameterized through MLP so that they can stably learn knowledge during the prefix training process.
[0081] First, S1 and S2 are respectively input into the first MLP neural network and the second MLP neural network, and then the outputs of the MLP are used as P key 、P value , and then concatenated with the s_key and s_value of the text tokenid to be used as the Key and Value when calculating the attention in mBART.
[0082] Secondly, S3 is input into the third MLP neural network, and the third MLP neural network is used to transform S3 into the Q value, K value, and V value of the cross-attention layer of the mBART translation model.
[0083] During the training process, only the MLP parameters are iteratively updated, keeping the parameters of the original mBART unchanged. Finally, the initial parameters of the MLP are obtained, and this MLP is used to map the initial embeddings of the prefix representations of each Transformer layer, which are used in both the encoder and the decoder.
[0084] The following is the MLP reparameterization process of Prefix-tuning:
[0085]
[0086] where i ∈ P idx .
[0087] 2.1.3 Encoder module
[0088] First, the source text sequence X is used as part of the input sequence of the encoder and fed into the encoder based on the mBART encoder, which contains multiple Transformer block layers. In this paper, by adding a prefix sequence key-value pair for Chinese-English machine translation and concatenating it with the key-value pair obtained from the source language text representation, the multi-head self-attention mechanism is jointly modified, and this behavior is used to introduce the influence of the prefix weight on the hidden layer features of the model. No changes are made to the queries during the attention calculation for the text sequence. The prefix sequence learns knowledge from the pre-trained model through the interaction with the source language text to execute the entire task and achieve overall optimization.
[0089] For example, in the l-th encoder Transformer block layer, the query (Q l ), key (K l ), and value (V l ) are obtained through the linear transformation of the hidden state of the source language text sequence. The key-value pair P θ -P key is obtained through the linear transformation of P' value , and then they are concatenated to obtain the new key-value pair K' l -V' l .
[0090]
[0091] Among them Finally, they are fed into the self-attention layer for further calculation:
[0092]
[0093] 2.2 Decoder module
[0094] For the decoder module, we also added a prefix-tuned MLP reparameterization module, and its multi-head self-attention mechanism and cross-attention mechanism are enhanced in a similar way as in the encoder module. The implementation of the self-attention layer directly uses the same method as the encoder, and then it is fed into the cross-attention mechanism. In the l-th layer decoder cross-attention, the K' l -V' l input by the decoder and the Q l of the encoder itself are used as shown in formula (4).
[0095]
[0096] Among them It is calculated by linearly transforming the translated text Y and the decoder hidden state.
[0097] 2.3 Training strategy
[0098] In the prefix module, the parameter set of all linear transformations is denoted as α. For the training strategy of this module, this paper performs gradient updates on the following log-likelihood objective:
[0099]
[0100] where the mBART parameters Ψ are fixed. The prefix parameters α, θ, and are the only trainable parameters. After training is completed, only all the parameters of the prefix module are retained, and the reparameterized parameters are deleted.
[0101] Advantages of the present invention:
[0102] Since large cross-lingual pre-trained models are trained using a large amount of monolingual corpora and have a huge scale of training parameters, strong hardware conditions are required for generalization when migrating to the field of machine translation; although cross-lingual pre-trained models can improve the performance of translation models to a certain extent, using cross-lingual pre-trained language models in low-resource scenarios may still lead to the inability to train the translation model well due to too little data and may also damage the knowledge of the original language model, resulting in an unsatisfactory performance improvement effect.
[0103] Based on the mBART model, the present invention adopts the idea of prefix tuning. During the training process, only the parameters in the prefix module are trained, without changing the parameters in the mBART model, and the weights of the cross-lingual pre-trained model are frozen, which can effectively reduce 90% of the number of trainable parameters and reduce the memory and time for training; moreover, since all the parameters of the pre-trained model are fixed in the present invention, the semantic information of the trained pre-trained language model is completely retained. In addition, the additional prefix module can learn the relationship between language pairs. It can learn translation knowledge additionally on the basis of completely protecting the knowledge of the language model.
[0104] In high-resource scenarios, since the number of trainable parameters of the current prefix tuning strategy is only 10% of the original training parameters, the expressive ability that can be learned is weaker than that of traditional full-parameter fine-tuning, and using a neural network to build a machine translation model will cause problems such as text grammar errors, repetitions, and contradictions.
[0105] The present invention integrates a controllable mechanism for controlling output attributes into the prefix tuning module, enhancing the learning and expressive ability of the model, enabling it to effectively learn the relationship between two semantic spaces, and effectively alleviating the grammar problems of the output text.
[0106] In addition, the present invention adopts a control mechanism of continuous vector-valued linear weighted superposition to manage the output attributes of the model, and adopts the method of initializing the control vector to zero, which can effectively avoid the problem of translation performance degradation caused by unclear attribute decision boundaries and incomplete attribute annotations.
[0107] The above-described embodiments are only preferred specific embodiments of the present invention, and the protection scope of the present invention is not limited thereto. Any simple changes or equivalent replacements of technical solutions that can be obviously obtained by those skilled in the art within the technical scope disclosed by the present invention all belong to the protection scope of the present invention.
Claims
1. A language translation method based on a prefix-tuning machine translation model, characterized in that Including: Constructing a prefix-tuning machine translation model, which includes: Initialize the control attribute labels of word order, length, and language style of the translated target language output by the prefix tuning machine translation model into a vector S with consistent dimensions i , for the vector S i Perform linear weighted combination to obtain the control signal matrix S, and set three groups of control signal matrices, namely S1, S2, and S3; Reparameterize S1 using the first MLP neural network and output the first prefix sequence key-value pair P 1key -P 1value ; Reparameterize S2 using the second MLP neural network and output the second prefix sequence key-value pair P 2key -P 2value ; Convert S3 into the Q value, K value, and V value of the cross-attention layer of the mBART translation model using the third MLP neural network; Feeding the Q value, K value, and V value output by the third MLP neural network into the cross-attention layer of the mBART translation model as trainable parameters in the cross-attention layer; Concatenate P 1key -P 1value with the key-value pairs in the encoder self-attention layer of the source language text to form new key-value pairs K' l -V' l , and pass K' l -V' l into the self-attention layer of the mBART encoder to achieve the improvement of the self-attention layer of the mBART encoder; Concatenate the second prefix sequence key-value pair P 2key -P 2value with the key-value pair in the key-value pair of the decoder self-attention layer of the target language text to form a new key-value pair K' l 1-V' l 1. Pass K' l 1-V' l 1 into the self-attention layer of the mBART decoder to implement the improvement of the self-attention layer of the mBART decoder; Input K' l -V' l into the cross-attention layer of the mBART decoder to achieve an improvement in the cross-attention layer of the mBART decoder; Inputting the source language into the trained prefix-tuning machine translation model to output the translated target language.
2. A language translation method based on a prefix-tuned machine translation model according to claim 1, characterized in that During the training of the prefix-tuning machine translation model, only the prefix-tuning module is trained, and the parameters of the original mBART model remain unchanged.
3. A language translation method based on a prefix-tuned machine translation model according to claim 2, characterized in that The objective function for training the prefix-tuning module is: Among them, Ψ is the parameter of mBART and does not change during training; x is the source language; y is the target language output by the model, and α, θ, are the model parameters of the first MLP neural network, the second MLP neural network, and the third MLP neural network.
4. A language translation method based on a prefix-tuned machine translation model according to claim 1, characterized in that, Set the control attribute label vector S of the three groups of control signal matrices S1, S2, and S3 i to the initial value of 0.
5. A language translation method based on a prefix-tuned machine translation model according to claim 1, characterized in that, The vector S i is linearly weighted and combined to obtain the control signal matrix S as follows: Among them, S i is the control attribute vector; W i is the trainable weight, used to adjust the intervention intensity of each control attribute; S is the control signal matrix; N is the number of vectors S i .
6. A language translation method based on a prefix-tuned machine translation model according to claim 1, characterized in that, The control attribute that controls the word order of the target language output by the prefix tuning machine translation model is the deviation intensity δ(s) of the off-diagonal alignment {(i, j)} between the source language sentence src = {x1, x2, …, x l} and the target language sentence tgt = {y1, y2, …, y m}, where i ∈ {1, 2, ..., l}, j ∈ {1, 2, ..., m}; where the deviation intensity δ(s) is: where #{(i,j)} represents the cardinality of the alignment; when δ(s) = 0, the sentence pair is in a strictly monotonic situation, where l = m and {(i,j)} is a strictly increasing bijective mapping; the smaller δ(s) is, the higher the monotonicity between the source language and the target translation.
7. A language translation method based on a prefix-tuned machine translation model according to claim 6, characterized in that, Trainable weight W of the word order i is as follows: W i = δ(s) + k where δ(s) is the intensity of the off-diagonal alignment deviation of each word in the same sentence of the source language and the target language corresponding to a matrix, and k is the offset value.
8. A language translation method based on a prefix-tuned machine translation model according to claim 1, characterized in that, The improvement of the self-attention layer of the mBART encoder includes: In the l-th Transformer block layer of the encoder, the query Q is obtained through a linear transformation of the hidden states of the source language text sequence l , the key K l and the value V l ; Combine the key-value pair P key -P value , with the key K l and the value V l to splice and obtain a new key-value pair K' l -V' l : Input K' l -V' l into the encoder self-attention layer to achieve the improvement of the encoder self-attention layer: Among them, Q l is the query in the encoder self-attention mechanism, generated based on the source language, and K' l -V' l is the key-value pair.
9. A language translation method based on a prefix-tuned machine translation model according to claim 1, characterized in that, The improved cross-attention layer of the mBART encoder is: Among them, is the query in the decoder self-attention mechanism, generated based on the source language.
Citation Information
Patent Citations
Machine translation model training method, machine translation method and related equipment
CN113515959A
Simultaneous translation device and computer program
WO2022181040A1