Large language model translation system based on encoder and decoder architecture
Through a large language model translation system based on the encoder decoder architecture, combining pre-trained models and dynamic layer jump strategies, the number of layers and connection methods on the decoder end are optimized, and the problems of high computing resources and slow inference speed are solved, and efficient and robust machine translation is achieved.
Patent Information
- Application Number
- CN202510335480.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
Large language models have high demand for computing resources, slow inference speed, insufficient support for low-resource languages, poor domain adaptability, and lack of controllability.
A large language model translation system based on the encoder decoder architecture is adopted. Through two-stage training, deep encoding-shallow decoding mode and dynamic layer jump strategy, the encoder-decoder structure is built with pre-trained large language models, the number of layers and connection methods at the decoder end are optimized, and a large number of multilingual bilingual corpus and high-quality fine-tuned parallel corpus are used for training.
It improves the translation quality and efficiency of the model, reduces the demand for computing resources, improves the robustness and domain adaptability of the model, and meets the real-time translation needs.
Smart Images

Figure CN120258005A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a large language model translation system, specifically a large language model translation system based on an encoder-decoder architecture. Background Art
[0002] In the field of natural language processing, models based on the encoder-decoder framework are generally applied to conditional generation tasks such as machine translation, text generation, and intelligent dialogue; the pre-training method trains a basic model through a large amount of general data, enabling the model to have good generalization ability in downstream tasks. Especially when the task-specific data is scarce, it can significantly improve the model performance and accelerate the convergence speed.
[0003] In neural networks, the pre-training method refers to training a basic model through a large amount of general data. This general and sufficient data can encourage the model to have good generalization ability in downstream tasks in the same field. Subsequently, for downstream tasks, task-specific data is used to fine-tune the pre-trained model, enabling the model to focus more on task-related features and perform better on this task. In the case of a small amount of task-specific data, the pre-training method can effectively improve the model performance. Moreover, since the pre-trained model already has general feature extraction ability, the fine-tuned model can achieve a faster convergence speed and stronger robustness.
[0004] With the emergence of large-scale pre-trained models (such as GPT, BERT, T5, etc.), the field of natural language processing has entered a new era. Through large-scale unsupervised pre-training and then supervised fine-tuning, large language models can demonstrate excellent performance in various natural language processing tasks, including machine translation. Compared with traditional translation methods, large language models have stronger language understanding and generation abilities, especially performing well in complex context and long text translation. Through pre-training on a large amount of corpora, large language models can capture the implicit associations between different languages, especially significantly improving the translation ability of low-resource languages. Due to their strong context understanding and generation abilities, large language models can provide more natural and fluent translations, reducing the rigid and mechanical translations that occurred in the past.
[0005] Although using large language models for machine translation tasks has many advantages, there are also some obvious disadvantages and challenges. First, large language models have high computational resource requirements. Large language models usually have a huge number of parameters, and training and inference require a large amount of computational resources and memory, making it difficult to deploy in resource-limited environments; second, the inference speed of large language models is relatively slow, especially in generative tasks (such as translation), where output needs to be generated word by word, making it difficult to meet the real-time translation requirements, especially in long text or high-concurrency scenarios.
[0006] However, no technical solutions capable of solving the above problems have been reported yet. Summary of the Invention
[0007] Aiming at the deficiencies in the prior art, such as the high computing resource requirements of large language models, the slow inference speed of large language models, insufficient support for low-resource languages, poor domain adaptability, and lack of controllability, the technical problem to be solved by the present invention is to provide a large language model translation system based on an encoder-decoder architecture, which combines a pre-trained large language model and an encoder-decoder architecture, combines the advantages of both to achieve better results.
[0008] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:
[0009] The present invention provides a large language model translation system based on an encoder-decoder architecture, including the following steps:
[0010] 1) In the data processing stage, a two-stage training method is used. For the first-stage training, a large amount of multilingual bilingual corpora are collected and preprocessed; for the second-stage training, high-quality fine-tuning parallel corpora are constructed.
[0011] 2) Model architecture selection: Use the pre-trained large language model to construct an encoder-decoder structure, adopt a deep encoding - shallow decoding mode, and confirm the number of layers retained at the decoder end and the connection method between the encoder and the decoder.
[0012] 3) Machine translation model training: Use the large amount of multilingual bilingual corpora and high-quality fine-tuning parallel corpora obtained in the data processing stage to train the model to obtain a machine translation model.
[0013] 4) Decoding stage: The encoder of the machine translation model encodes the source language sentence, and then decodes it through the decoder to generate the target language sentence.
[0014] In step 2) of model architecture selection, use the pre-trained large language model to construct an encoder-decoder structure, adopt a deep encoding - shallow decoding mode, and confirm the number of layers retained at the decoder end, specifically as follows:
[0015] Select some layers of the pre-trained large language model at the decoder end of the machine translation model for decoding to improve the decoding efficiency. Before training the machine translation model, determine the number of layers retained at the decoder end and which layers to retain, specifically:
[0016] Conduct pre-experiments to test various fixed layer skip strategies, including retaining one layer every N layers in the pre-trained large language model, retaining all intermediate layers, retaining the bottom layer close to the input end of the pre-trained large language model, the top layer close to the output end of the pre-trained large language model, or retaining some layers of the above bottom layer, top layer, and intermediate layers.
[0017] In step 2), for the model architecture selection, a pre-trained large language model is used to construct an encoder-decoder structure, adopting a deep encoding - shallow decoding mode, and determining the number of layers retained at the decoder end. Specifically:
[0018] A dynamic skip-layer strategy is adopted, including dynamic skip-layer based on input complexity and dynamic skip-layer based on gating mechanism. For the dynamic skip-layer based on input complexity, for simple inputs with a sentence length less than or equal to 20 words, 1 / 4 to 1 / 3 of the layers of the pre-trained large language model are retained; for complex inputs with a sentence length greater than 20 words, 3 / 5 to 2 / 3 of the layers of the pre-trained large language model are retained. For the dynamic skip-layer based on gating mechanism, a learnable gating mechanism is introduced for each layer of the pre-trained large language model, that is, a neural network is used before each layer to predict the skip probability of the current layer based on the input representation of the current layer. When the skip probability is greater than 0.5, the current layer is skipped.
[0019] In step 2), the connection method between the encoder and the decoder is as follows:
[0020] The input content is independently encoded at the encoder end of the machine translation model to obtain the input representation of each layer at the encoder end. Subsequently, during decoding, for each layer at the decoder end, the input representation of the corresponding layer at the encoder end is concatenated with the target output representation of the previous layer at the decoder end, so that self-attention calculation is performed on the input and output sequences of each layer to generate the target output representation of that layer.
[0021] The concatenation of the input representation of the corresponding layer at the encoder end and the target output representation of the previous layer at the decoder end is specifically:
[0022] Given the source sequence X and the target sequence Y, this process is represented by the following formula:
[0023] X l = FFN(SAtt(X l-1 , M c ))
[0024] Y l = FFN(SAtt([X l , Y l-1 , M c ))
[0025] where l represents the layer index, X l is the input representation of the l-th layer, Y l is the target output representation of the l-th layer, X l-1 is the input representation of the (l - 1)-th layer, Y l-1 is the target output representation of the (l - 1)-th layer, and M cIt is a causal masking pattern, SAtt is the self-attention mechanism, and FFN is the feed-forward neural network;
[0026] Each layer of the encoder first uses the causal masking pattern M c to perform self-attention calculation, and then obtains the input representation of this layer through the feed-forward neural network FFN;
[0027] Each layer of the decoder combines the input representation X of the corresponding encoder layer l and the target output representation Y of the previous decoder layer l-1 After merging, the causal masking pattern M is used again c to perform self-attention calculation, and then the target output representation of this layer is obtained through the feed-forward neural network FFN.
[0028] In step 3) of machine translation model training, the machine translation model is trained using the massive multilingual bilingual corpus and high-quality fine-tuning parallel corpus obtained in the data processing stage, specifically:
[0029] In the first-stage training, the parameters of the encoder end of the machine translation model are frozen, and the decoder end of the machine translation model is trained using the massive multilingual bilingual corpus obtained in step 1) of the data processing stage;
[0030] In the second-stage training, the entire machine translation model is trained using the high-quality fine-tuning parallel corpus collected in step 1) of the data processing stage to update all the parameters of the encoder end and the decoder end.
[0031] The present invention has the following beneficial effects and advantages:
[0032] 1. The present invention is a brand-new large language model translation system. It uses a pre-trained large language model to construct an encoder-decoder architecture, which is used for machine translation tasks after two-stage training. By using the pre-trained large language model, constructing an encoder-decoder architecture, and adopting a deep encoding - shallow decoding format, while fully utilizing the powerful context understanding and generation ability of the large language model, it overcomes the disadvantage of its slow inference speed, thereby improving the translation quality and effect of the model.
[0033] 2. The present invention applies the pre-trained model to the neural machine translation model. At the same time, the same pre-trained large language model is used to construct the encoder end and the decoder end of the machine translation model. By using the layer skipping technology to reduce the number of layers at the decoder end, it avoids the problem of inconsistency between the encoder and the decoder during the training process, speeds up the convergence speed of the model, and improves the robustness of the model and the benefits brought by the pre-training method. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic diagram of the internal structure of each layer of the pre-trained encoder large model in the present invention;
[0035] Figure 2 This is a diagram of the internal structure of a layer of the pre-trained decoder large model in the present invention;
[0036] Figure 3 This is a diagram of the present invention based on a pre-trained encoder-decoder architecture. Detailed implementation manners
[0037] The present invention will be further described below in conjunction with the accompanying drawings of the specification.
[0038] The present invention provides a large language model translation system based on an encoder-decoder architecture, including the following steps:
[0039] 1) In the data processing stage, a two-stage training method is used. For the first-stage training, a large amount of multilingual bilingual corpora are collected and preprocessed; for the second-stage training, high-quality fine-tuning parallel corpora are constructed;
[0040] 2) In the model architecture selection, a pre-trained large language model is used to construct an encoder-decoder structure, adopting a deep encoding - shallow decoding mode, and determining the number of layers retained at the decoder end and the connection method between the encoder and the decoder;
[0041] 3) In the machine translation model training, the machine translation model is obtained by training using the large amount of multilingual bilingual corpora and high-quality fine-tuning parallel corpora obtained in the data processing stage;
[0042] 4) In the decoding stage, the encoder of the machine translation model encodes the source language sentence, and then the decoder decodes to generate the target language sentence.
[0043] The present invention proposes an innovative machine translation architecture that cleverly uses a pre-trained large language model to construct an encoder-decoder structure. In this architecture, the large language model is used as the encoder and the decoder respectively, but at the decoder end, not all layers of the entire large language model are used. Instead, several of them are carefully selected to perform the decoding task. This design not only retains the powerful representation ability of the large language model but also enhances the flexibility and efficiency of the model through hierarchical selection.
[0044] Specifically, when the input sequence passes through the encoder side of the pre-trained large language model, each layer generates corresponding output representations. These rich hierarchical information provides multi-level context clues for the decoder. At the decoder side, instead of simply passing the final output of the encoder to the decoder, the present invention adopts a more refined strategy: each layer of the decoder combines the output representation of the corresponding encoder layer and the output representation of the previous layer of the decoder, and obtains the output representation of the current layer through complex calculations. This progressive processing method enables the decoder to capture and integrate information at different abstraction levels, thereby performing decoding operations more accurately.
[0045] In step 1), for the first-stage training, it is necessary to collect and process a large amount of multilingual bilingual parallel corpora, specifically:
[0046] First, collect a large amount of multilingual bilingual parallel corpora covering multiple language pairs; these corpora can come from publicly available multilingual datasets, bilingual texts crawled from the web, or other multilingual resources;
[0047] Add a special identifier in front of each sentence to indicate the language type of the sentence; this identifier helps the model to identify and process inputs in different languages during training, so as to perform specific encoding and generation;
[0048] Clean and filter the collected bilingual parallel corpora, remove noise data, and then use the sentence length ratio filtering method to remove sentence pairs in the bilingual parallel corpora without noise data where the length ratio of the source language sentence to the target language sentence is greater than 1.5;
[0049] Process the cleaned same-language sentence pair data to ensure the alignment quality between sentence pairs, use tools such as COMET to score and screen the cleaned and filtered parallel sentence pairs to ensure the data quality; then mix the corpora of different language pairs to form multilingual training data; this mixed data helps the model learn cross-language general representations.
[0050] This step is to construct high-quality fine-tuning parallel corpora to further improve the model performance. These corpora should cover the language species in the previous stage of training and have high translation quality and domain coverage. On this basis, collect the test sets and validation sets of WMT over the years, as well as other open-source high-quality datasets, such as the data of datasets like Flores. Randomly select some samples as the validation set and test set to ensure that the training set, validation set, and test set have high quality while covering multiple fields.
[0051] In this embodiment, parallel corpora covering multiple language pairs such as Chinese-English, German-English, Russian-English, etc. are collected from public multilingual data sets such as OPUS, WMT, etc., bilingual texts crawled from the web such as news, Wikipedia, etc., and other multilingual resources. A special identifier is added in front of each sentence, such as English for English and Chinese for Chinese, to indicate the language type of the sentence. For example: for the sentence "The weather is nice today.", after adding the identifier, it becomes "Chinese: The weather is nice today."; for the sentence "The weather is nice today.", after adding the identifier, it becomes "English: The weather is nice today.". Then, the collected bilingual parallel corpus is cleaned and preprocessed to ensure data quality, including removing noise data, including repeated sentences, sentences that are too short (such as less than 3 words) or too long (such as more than 100 words), sentences that do not conform to grammatical rules, etc. Then, the sentence length ratio filtering method is used to remove sentence pairs in the bilingual parallel corpus that does not contain noise data, where the ratio of the source sentence to the target sentence length is greater than 1.5; then, tools (such as COMET) are used to score and screen the parallel sentence pairs to ensure the alignment quality between the sentence pairs. Then, the cleaned data of the same language pairs is processed, and the corpus of different language pairs is mixed to form multilingual training data.
[0052] In step 1), for the second stage of training, it is necessary to build high-quality fine-tuning parallel corpus, specifically:
[0053] In order to further improve the translation performance of the model, it is necessary to build high-quality fine-tuned parallel corpora that cover the languages trained in the previous stage to achieve higher translation quality and domain coverage.
[0054] On this basis, we collected the test sets and validation sets of WMT over the years, as well as other open source high-quality datasets, such as the data of Flores dataset;
[0055] Some samples are randomly selected as validation sets and test sets to ensure that the training sets, validation sets, and test sets are of high quality and cover a variety of fields.
[0056] In step 2), the model architecture selection uses the pre-trained large language model to build the encoder-decoder structure, adopts the deep encoding-shallow decoding mode, and confirms the number of layers retained on the decoder side, as follows:
[0057] On the decoder side of the machine translation model (e.g. Figure 2 As shown in the figure, select some layers of the pre-trained large language model for decoding to improve the decoding efficiency. Before training the machine translation model, determine the number of layers to be retained on the decoder side and which layers to retain, specifically:
[0058] Conduct a pre-experiment to test various fixed layer skipping strategies, including retaining one layer out of every N layers in the pre-trained large language model, retaining all intermediate layers, retaining the bottom layers close to the input end of the pre-trained large language model, the top layers close to the output end of the pre-trained large language model, or retaining some layers of the above bottom, top, and intermediate layers.
[0059] This step takes into account that there are tasks of different difficulties in the translation task. For example, the difficulty of short text translation is relatively low, and more layers can be skipped appropriately; while the difficulty of text translation, domain translation, etc. is relatively high, and more layers need to be retained. Therefore, various aspects need to be considered when determining the layer skipping strategy.
[0060] The invention adopts a deep encoding - shallow decoding mode. At the decoder end, the entire pre-trained large language model is not used for decoding, but some of its layers are selected to improve the decoding efficiency. Therefore, before training, it is necessary to first determine the number of layers to be retained at the decoder end and which layers to retain.
[0061] In the pre-experiment, various fixed layer skipping strategies were tested, including skipping once every N layers, skipping intermediate layers, skipping bottom or top layers; the bottom, intermediate, and top layers each account for one-third of the total number of layers of the entire pre-trained large language model.
[0062] Among them, various fixed layer skipping strategies include skipping once every N layers: for example, in a 32-layer large language model, retain one layer out of every 4 layers (retain layers 4, 8, 12, etc.);
[0063] Skipping intermediate layers: In a deep model, the intermediate layers may contribute less to the final output of the decoder. Therefore, several intermediate layers can be selected to be skipped (such as skipping layers 12 - 24);
[0064] Skipping bottom or top layers: Skip the bottom layers close to the input end of the decoder (such as layers 1 - 2) or the top layers close to the output end of the decoder (such as layers 31 - 32);
[0065] Retaining bottom, top, and intermediate layers: Retain the bottom layers close to the input end of the decoder (such as layers 1 - 2) and the top layers close to the output end of the decoder (such as layers 31 - 32) and some intermediate layers.
[0066] Step 2) can also adopt a dynamic layer skipping strategy, including dynamic layer skipping based on input complexity and dynamic layer skipping based on a gating mechanism. For dynamic layer skipping based on input complexity, for simple inputs with a sentence length less than or equal to 20 words, retain 1 / 4 to 1 / 3 of the number of layers of the pre-trained large language model. For complex inputs with a sentence length greater than 20 words, retain 3 / 5 to 2 / 3 of the number of layers of the pre-trained large language model. For dynamic layer skipping based on a gating mechanism, introduce a learnable gating mechanism for each layer of the pre-trained large language model, that is, use a neural network before each layer to predict the skipping probability of that layer based on the input representation of the current layer. When the skipping probability is greater than 0.5, skip the current layer.
[0067] Therefore, through experiments, a dynamic layer skipping strategy was also tried, including dynamic layer skipping based on input complexity: for simple inputs (sentence length less than or equal to 20 words), skip more layers; for complex inputs (sentence length greater than 20 words), skip fewer layers; and dynamic layer skipping based on a gating mechanism: introduce a learnable gating mechanism for each layer and dynamically decide whether to skip that layer based on the input. For example, use a lightweight neural network to predict the skipping probability of each layer.
[0068] After testing and tuning, considering various factors such as translation effect and efficiency, this embodiment finally selects the method of retaining the bottom layer, top layer, and middle layers, specifically retaining the bottom layer (layers 1 and 2) near the input end of the decoder, the top layer (layers 31 and 32) near the output end of the decoder, and some middle layers (layers 7, 14, 19, 26).
[0069] After determining the layer skipping strategy, the present invention needs to determine how to connect the encoder end and the decoder end, that is, use the output of the encoder end (as Figure 1 shown) for decoding at the decoder end. The connection method between the encoder and the decoder in step 2) is as Figure 3 shown, specifically:
[0070] Independently encode the input content at the encoder end of the machine translation model to obtain the input representation of each layer at the encoder end. Subsequently, during decoding, for each layer at the decoder end, concatenate the input representation of the corresponding layer at the encoder end with the target output representation of the previous layer at the decoder end, so that self-attention calculation is performed on the input and output sequences of each layer to generate the target output representation of that layer.
[0071] The present invention uses a pre-trained large language model to construct an encoder-decoder structure. Therefore, after determining the layer skipping strategy, it is necessary to determine how to connect the encoder end and the decoder end, that is, use the output of the encoder end for decoding at the decoder end.
[0072] Different from the Transformer-based encoder-decoder architecture, in the traditional architecture, the decoder generates the output based on the source representation generated by the top layer of the encoder and the target representation of the previous layer.
[0073] In the present invention, due to the nature of the causal attention mechanism, the input representation is independent of the output. Therefore, the encoder can be used to independently encode the input to obtain the input representation of each layer. Subsequently, during decoding, each layer of the decoder combines the input representation of the corresponding encoder layer and the target output representation of the previous decoder layer, enabling the decoder to perform self-attention calculations on the input and output sequences for generating the target representation output of that layer. Given the source sequence X and the target sequence Y, this process can be expressed by the following formula:
[0074] X l = FFN(SAtt(X l-1 , M c ))
[0075] Y l = FFN(SAtt([X l , Y l-1 , M c ))
[0076] where l represents the layer index, X l is the input representation of the l-th layer, Y l is the target output representation of the l-th layer, X l-1 is the input representation of the (l - 1)-th layer, Y l-1 is the target output representation of the (l - 1)-th layer, M c is the causal mask pattern, SAtt is the self-attention mechanism, and FFN is the feed-forward neural network;
[0077] Each layer of the encoder first performs self-attention calculation using the causal mask pattern M c and then obtains the input representation of that layer through the feed-forward neural network FFN.
[0078] Step 3) Machine translation model training. Use the massive multilingual bilingual corpus and high-quality fine-tuning parallel corpus obtained in the data processing stage to train the machine translation model. Specifically:
[0079] First-stage training. Freeze the parameters of the encoder end of the machine translation model and train the decoder end of the machine translation model using the massive multilingual bilingual corpus obtained in Step 1) data processing stage.
[0080] When performing the first-stage training, the main purpose is to activate and enhance the translation ability of the decoder. In this way, the decoder can learn how to generate sentences in the target language based on the output of the encoder while retaining the multilingual representation ability of the encoder.
[0081] In the second - stage training, use the high - quality fine - tuned parallel corpus collected in step 1) to train the entire machine translation model, which is used to update all the parameters of the encoder side and the decoder side.
[0082] When performing the second - stage training, the main purpose is to improve the performance and effect of the entire translation system. Different from the first - stage training, in this stage, all the parameters of the encoder side and the decoder side will be updated simultaneously to make the model better adapt to a specific translation task; use the high - quality fine - tuned parallel corpus collected in step 1) to train the pre - trained encoder - decoder architecture. During the training process, a smaller learning rate is adopted to avoid overfitting, and techniques such as early stopping are used to prevent the performance of the model from degrading on the validation set.
[0083] First, freeze the parameters of the encoder side and only train the decoder side. Use the Llama - 3.1 - 8B open - source large model, which has a total of 32 layers. Finally, select and retain the bottom layers (layers 1 and 2), top layers (layers 31 and 32), and intermediate layers (layers 7, 14, 19, 26). Then freeze the parameters of the encoder side and use the processed multilingual training data to train the decoder side.
[0084] Then, use the high - quality fine - tuned parallel corpus to train the model, and update all the parameters of the encoder side and the decoder side simultaneously. Use a smaller learning rate (such as 5e - 5) to avoid overfitting. At the same time, use the early - stopping technique to stop training when the performance on the validation set no longer improves. Through training, make the model better adapt to a specific translation task and improve the translation quality and domain adaptability.
[0085] 4) Decoding stage: The encoder of the machine translation model encodes the source - language sentence, and then the decoder decodes to generate the target - language sentence.
[0086] In this step, the encoder encodes the source - language sentence to generate the context representation of the source - language sentence; then, the decoder gradually generates the target - language sentence according to the output of the encoder. Techniques such as beam search can be used during the decoding process to improve the quality of the generated sentence.
[0087] Use the trained model for translation generation. Use the encoder to encode the source - language sentence to generate the context representation of the source - language sentence. Then the decoder decodes and gradually generates the target - language sentence according to the output of the encoder. Use the beam - search technique and set the beam width to 5 to improve the quality of the generated sentence.
[0088] Verification process:
[0089] The invention was verified by multilingual translation tasks in eight directions, including English-German, English-Russian, English-Czech, and English-Chinese. In the first stage of training, all the data of WMT23 was used to train the decoder side of the model after processing. The number of layers retained by the decoder side selected the method of retaining the bottom layer, top layer, and middle layer, specifically retaining the bottom layer close to the input (1st and 2nd layers) and the top layer close to the output (31st and 32nd layers) and some layers in the middle (7th, 14th, 19th, and 26th layers).
[0090] Afterwards, the entire model is trained using the collected and organized high-quality data sets, and then tested on the test set. As shown in Table 1, the experimental results show that compared with the traditional NMT model (NMT-40-8), although the efficiency of translation decoding is somewhat different, the translation quality has been greatly improved, and the BLEU value and COMET value have increased by 2.00 and 1.84 on average, respectively; compared with the fine-tuned large model of the same type (Llama-3.1-8B-SFT), the present invention has greatly improved both translation efficiency and translation quality, with an efficiency increase of about 4.65 times, and the BLEU value and COMET value have also increased by 2.23 and 1.36 on average, respectively.
[0091] Table 1 Translation effect and speed comparison
[0092]
[0093] Since the encoder-decoder architecture has the advantages of high efficiency and domain adaptability in translation tasks, combining the pre-trained large language model with the encoder-decoder architecture can not only fully utilize the powerful capabilities of the large language model, but also overcome its shortcomings such as slow reasoning speed, thereby achieving higher quality machine translation. The present invention innovatively proposes an encoder-decoder architecture based on a pre-trained large language model, which cleverly combines the pre-trained large language model and the hierarchical decoding strategy, which not only improves the translation quality, but also optimizes the performance of the model, bringing new breakthroughs to the field of machine translation.
Claims
1. A large language model translation system based on an encoder-decoder architecture, characterized in that It includes the following steps: 1) Data processing stage: Using a two-stage training method, collect a large amount of multilingual bilingual corpora for the first-stage training and perform preprocessing; construct high-quality fine-tuning parallel corpora for the second-stage training. 2) Model architecture selection: Use a pre-trained large language model to construct an encoder-decoder structure, adopt a deep encoding - shallow decoding mode, and confirm the number of layers retained at the decoder end and the connection method between the encoder and the decoder. 3) Machine translation model training: Use the large amount of multilingual bilingual corpora and high-quality fine-tuning parallel corpora obtained in the data processing stage to train the model to obtain a machine translation model. 4) Decoding stage: The encoder of the machine translation model encodes the source language sentence, and then decodes it through the decoder to generate the target language sentence.
2. The large language model translation system based on the encoder-decoder architecture according to claim 1, characterized in that: In step 2), for the model architecture selection, use a pre-trained large language model to construct an encoder-decoder structure, adopt a deep encoding - shallow decoding mode, and confirm the number of layers retained at the decoder end, specifically as follows: Select some layers of the pre-trained large language model at the decoder end of the machine translation model for decoding to improve the decoding efficiency. Before training the machine translation model, determine the number of layers retained at the decoder end and which layers to retain, specifically: Conduct preliminary experiments to test various fixed layer skipping strategies, including retaining one layer out of every N layers in the pre-trained large language model, retaining all intermediate layers, retaining the bottom layers close to the input end of the pre-trained large language model, the top layers close to the output end of the pre-trained large language model, or retaining some layers of the above bottom layers, top layers, and intermediate layers.
3. The large language model translation system based on the encoder-decoder architecture according to claim 1, wherein: In step 2), for the model architecture selection, use a pre-trained large language model to construct an encoder-decoder structure, adopt a deep encoding - shallow decoding mode, and confirm the number of layers retained at the decoder end, specifically: Adopt dynamic layer skipping strategies, including dynamic layer skipping based on input complexity and dynamic layer skipping based on a gating mechanism. Among them, for dynamic layer skipping based on input complexity, for simple inputs with a sentence length less than or equal to 20 words, retain 1 / 4 - 1 / 3 of the layers of the pre-trained large language model; for complex inputs with a sentence length greater than 20 words, retain 3 / 5 - 2 / 3 of the layers of the pre-trained large language model; for dynamic layer skipping based on a gating mechanism, introduce a learnable gating mechanism for each layer of the pre-trained large language model, that is, use a neural network in front of each layer to predict the skipping probability of the current layer based on the input representation of the current layer. When the skipping probability is greater than 0.5, skip the current layer.
4. The large language model translation system based on the encoder-decoder architecture according to claim 1, characterized in that: In step 2), the connection method between the encoder and the decoder is: Independently encode the input content at the encoder end of the machine translation model to obtain the input representation of each layer at the encoder end; subsequently, during decoding, for each layer at the decoder end, concatenate the input representation of the corresponding layer at the encoder end with the target output representation of the previous layer at the decoder end, so that each layer performs self-attention calculation on the input and output sequences to generate the target output representation of that layer.
5. The large language model translation system based on the encoder-decoder architecture according to claim 4, characterized in that: Concatenate the input representation of the corresponding layer at the encoder end with the target output representation of the previous layer at the decoder end, specifically: Given the source sequence X and the target sequence Y, this process is expressed by the following formula: X l = FFN(SAtt(X l-1 , M c )) Y l = FFN(SAtt([X l , Y l-l , M c )) where l represents the layer index, X l is the input representation of the l-th layer, Y l is the target output representation of the l-th layer, X l-1 is the input representation of the (l-1)-th layer, Y l-1 is the target output representation of the (l-1)-th layer, M c is the causal mask pattern, SAtt is the self-attention mechanism, and FFN is the feed-forward neural network; Each layer of the encoder first uses the causal masking pattern M c to perform self-attention calculations, and then obtains the input representation of the layer through the feed-forward neural network FFN; Each layer of the decoder takes as input the representation X corresponding to the encoder layer l and the target output representation Y of the previous decoder layer l-1 After merging, the causal masking pattern M c is used for self-attention calculation, and then the target output representation of this layer is obtained through the feed-forward neural network FFN 6. The large language model translation system based on the encoder-decoder architecture according to claim 1, characterized in that: In step 3) of machine translation model training, a large amount of multilingual bilingual corpus and high-quality fine-tuning parallel corpus obtained in the data processing stage are used for model training to obtain a machine translation model. Specifically: In the first-stage training, the parameters of the encoder end of the machine translation model are frozen, and the decoder end of the machine translation model is trained using the large amount of multilingual bilingual corpus obtained in step 1) of the data processing stage; In the second-stage training, the overall machine translation model is trained using the high-quality fine-tuning parallel corpus collected in step 1) of the data processing stage to update all the parameters of the encoder end and the decoder end.